{"notes":[{"content":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Video benchmark","Intuitive Physic","Evaluation"]},"supplementary_material":{"value":"/attachment/4932fdafe9cb6b17072f67b2caf246775bce6feb.zip"},"_bibtex":{"value":"@misc{\nbordes2026intphys,\ntitle={IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments},\nauthor={Florian Bordes and Quentin Garrido and Justine T Kao and Adina Williams and Michael Rabbat and Emmanuel Dupoux},\nyear={2026},\nurl={https://openreview.net/forum?id=Xpf5x3mLvn}\n}"},"title":{"value":"IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments"},"paperhash":{"value":"bordes|intphys_2_benchmarking_intuitive_physics_understanding_in_complex_synthetic_environments"},"dataset_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh.zip"},"TLDR":{"value":"We introduced a new intuitive physic benchmark IntPhys 2, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments"},"primary_area":{"value":"datasets_&_benchmarks_for_computer_vision"},"abstract":{"value":"We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to  macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50\\%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies."},"pdf":{"value":"/pdf/c29b03435c4329f0422348326d7ca2b5d606960a.pdf"},"croissant_file":{"value":"/attachment/54b155200f4140df0470921a91b0093300c98b6a.json"},"code_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh_code.zip"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"authorids":{"value":["~Florian_Bordes1","~Quentin_Garrido1","~Justine_T_Kao1","~Adina_Williams1","~Michael_Rabbat1","~Emmanuel_Dupoux1"]},"authors":{"value":["Florian Bordes","Quentin Garrido","Justine T Kao","Adina Williams","Michael Rabbat","Emmanuel Dupoux"]}},"tmdate":1776956109173,"tcdate":1746810448141,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Authors"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Authors"],"forum":"Xpf5x3mLvn","license":"CC BY-NC 4.0","number":1018,"cdate":1746810448141,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Submission","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Post_Submission","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Full_Submission","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Supplementary_Material"],"mdate":1776956109173,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","id":"Xpf5x3mLvn","version":2},{"content":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Intuitive Physics","Computer Vision"]},"supplementary_material":{"value":"/attachment/d183adaec59dcb150929a16c875749851b24a450.pdf"},"_bibtex":{"value":"@inproceedings{\nxue2023dintphys,\ntitle={3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes},\nauthor={Haotian Xue and Antonio Torralba and Joshua B. Tenenbaum and Daniel LK Yamins and Yunzhu Li and Hsiao-Yu Tung},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=Fp5uC6YHwe}\n}"},"title":{"value":"3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes"},"paperhash":{"value":"xue|3dintphys_towards_more_generalized_3dgrounded_visual_intuitive_physics_under_challenging_scenes"},"abstract":{"value":"Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate the scene to achieve desired outcomes without relying on extensive trial and error. In this paper, we present a framework capable of learning 3D-grounded visual intuitive physics models from videos of complex scenes with fluids. Our method is composed of a conditional Neural Radiance Field (NeRF)-style visual frontend and a 3D point-based dynamics prediction backend, using which we can impose strong relational and structural inductive bias to capture the structure of the underlying environment. Unlike existing intuitive point-based dynamics works that rely on the supervision of dense point trajectory from simulators, we relax the requirements and only assume access to multi-view RGB images and (imperfect) instance masks acquired using color prior. This enables the proposed model to handle scenarios where accurate point estimation and tracking are hard or impossible. We generate datasets including three challenging scenarios involving fluid, granular materials, and rigid objects in the simulation. The datasets do not include any dense particle information so most previous 3D-based intuitive physics pipelines can barely deal with that. We show our model can make long-horizon future predictions by learning from raw images and significantly outperforms models that do not employ an explicit 3D representation space. We also show that once trained, our model can achieve strong generalization in complex scenarios under extrapolate settings."},"pdf":{"value":"/pdf/739e599bd87f8fd4fcb09b78dd8827c81af83052.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Haotian_Xue1","~Antonio_Torralba1","~Joshua_B._Tenenbaum1","~Daniel_LK_Yamins1","~Yunzhu_Li1","~Hsiao-Yu_Tung1"]},"authors":{"value":["Haotian Xue","Antonio Torralba","Joshua B. Tenenbaum","Daniel LK Yamins","Yunzhu Li","Hsiao-Yu Tung"]}},"tmdate":1698949739165,"pdate":1695325945585,"tcdate":1683724459033,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission7124/Authors"],"signatures":["NeurIPS.cc/2023/Conference/Submission7124/Authors"],"forum":"Fp5uC6YHwe","number":7124,"cdate":1683724459033,"mdate":1698949739165,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/-/Submission","NeurIPS.cc/2023/Conference/-/Post_Submission","NeurIPS.cc/2023/Conference/Submission7124/-/Revision","NeurIPS.cc/2023/Conference/Submission7124/-/Supplementary_Material_Revision","NeurIPS.cc/2023/Conference/-/Edit","NeurIPS.cc/2023/Conference/Submission7124/-/Camera_Ready_Revision"],"odate":1698949739151,"domain":"NeurIPS.cc/2023/Conference","id":"Fp5uC6YHwe","version":2},{"content":{"summary":{"value":"The paper introduces INPHYRE, a synthetic CLEVR‑style benchmark that gives exemplar videos and asks VQA about outcomes of collisions, including rule violations. The authors report three findings. Limited parametric physics knowledge, weak adaptation to unseen rules, and strong language bias."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- What exact boundary makes this benchmark first? If the novelty is exemplar videos plus VQA under rule changes, state that narrowly and revise the claim.\n\n- Why is the query only a first frame? Do results hold if the model sees the full video at test."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"- Clear motivation about adaptation to out‑of‑distribution physics.\n- Simple, reproducible pipeline built on PyBullet and Blender with a scenario taxonomy.\n- Broad model sweep and a few ablations on number of exemplars, CoT, and quantization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Firstness claim is incorrect or overstated. Prior work already probes counterfactual or violated physics and adaptation in VQA or video settings, for example IntPhys, CLEVRER, CoPhy, ComPhy, Physion and Physion++, ContPhy, PhysBench, Physics Context Builders, and context conditioned physics efforts, and related context driven benchmarks.\n\n- Core novelty is thin. This is another templated synthetic collision suite with attribute rules tied to color. The few‑shot prompting setup is standard. It is unclear why the original CLEVRER dataset cannot be used for this task.\n\n- Evaluation design confounds language bias by construction. Exemplars often include QA pairs, while the test uses only a single image frame. This invites template copying and makes the video‑only drop unsurprising. Claims that models ignore vision are therefore not well supported.\n\n- Many questions require motion information that a single frame cannot provide. Penalizing Not enough data may punish reasonable uncertainty.\n\n- No human baseline, no test of option prior balance, limited statistical analysis, and no side‑by‑side with earlier benchmarks under a shared protocol.\n\nOverall, the paper’s empirical sweep is useful, but the novelty and claims are not strong, and key design choices undermine the main conclusions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924868076,"tcdate":1761846833724,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14468/Reviewer_unjk"],"signatures":["ICLR.cc/2026/Conference/Submission14468/Reviewer_unjk"],"forum":"IIrPoZ28dN","number":3,"license":"CC BY 4.0","cdate":1761846833724,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14468/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924868076,"domain":"ICLR.cc/2026/Conference","replyto":"IIrPoZ28dN","id":"Tqh4Gaovom","forumContent":{"TLDR":{"value":"We evaluate whether large multimodal models can infer and apply unseen physical laws from demonstration samples."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["large multimodal models","physical reasoning","inductive reasoning"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large multimodal models (LMMs) encode universal physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potential collision event from visual input. However, since parametric knowledge includes only the physical laws seen during training, it is insufficient for reasoning when the inference scenario follows physical laws unseen during training. In contrast, humans can adapt their physical reasoning to unseen physical environments with only a few visual examples. This inductive physical reasoning ability is indispensable for LMMs if they are to replace human agents in safety-critical applications. Despite its importance, existing visual benchmarks evaluate only the parametric knowledge in LMMs, and not inductive physical reasoning. To this end, we propose InPhyRe, the first visual question answering benchmark to measure inductive physical reasoning in LMMs. InPhyRe evaluates LMMs on their ability to predict the outcome of collision events in algorithmically generated synthetic videos. By inspecting over 13 open-source and proprietary LMMs, InPhyRe informs us that (1) LMMs struggle to apply their limited parametric knowledge about universal physical laws to reasoning, (2) inductive physical reasoning in LMMs is weak when inference scenarios obey physical laws unseen during training, and (3) inductive physical reasoning in LMMs suffers from language bias and largely ignores the visual inputs, questioning the trustworthiness of LMMs regarding visual inputs."},"_bibtex":{"value":"@misc{\nsreekumar2026inphyre,\ntitle={InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning},\nauthor={Gautam Sreekumar and Vishnu Boddeti},\nyear={2026},\nurl={https://openreview.net/forum?id=IIrPoZ28dN}\n}"},"title":{"value":"InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning"},"pdf":{"value":"/pdf/b87a0e49daf5f5f05caff235d97e823b17b5dbfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sreekumar|inphyre_discovers_large_multimodal_models_struggle_in_inductive_physical_reasoning"},"authorids":{"value":["~Gautam_Sreekumar1","~Vishnu_Boddeti1"]},"authors":{"value":["Gautam Sreekumar","Vishnu Boddeti"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LikePhys, a training-free evaluation framework designed to assess the intuitive physics understanding of video diffusion models (VDMs). Instead of relying on human or vision-language judgments, LikePhys measures how well a model distinguishes physically valid from invalid videos using its denoising loss as a likelihood proxy. The authors construct a controlled benchmark of twelve simulated scenarios across four physics domains—rigid-body, continuum, fluid mechanics, and optical effects—and define the Plausibility Preference Error (PPE) to quantify whether the model assigns higher likelihood to physically plausible sequences.\n\nThrough systematic experiments on twelve state-of-the-art VDMs, the study finds that larger Transformer-based models (e.g., Hunyuan T2V, Wan 2.1–14B) outperform UNet-based ones, showing partial emergence of physics reasoning. PPE correlates strongly with human judgments but remains largely independent of visual quality metrics, confirming that it measures physical plausibility rather than appearance. Nonetheless, current models still struggle with complex or chaotic dynamics—particularly in fluid mechanics and conservation-law scenarios—highlighting the need for future work on physics-aware training and longer temporal modeling."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I do not see evaluation code and data in the supplementary material. Do you have any opensource plan?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper proposes LikePhys, a novel and training-free method that evaluates intuitive physics understanding through a model’s own likelihood estimation. This approach elegantly connects diffusion models’ denoising objective with physical plausibility assessment, avoiding dependence on human annotation or vision–language model judges. It offers a fundamentally objective and interpretable evaluation paradigm for generative models.\n2. LikePhys can be directly applied to any diffusion-based video model without additional training or fine-tuning. By using the denoising loss as an ELBO-based likelihood surrogate, it remains compatible with a wide range of architectures and inference setups, enabling scalable and reproducible benchmarking across different models.\n3. The authors construct a highly systematic simulation dataset of twelve physics scenarios covering four domains—rigid-body, continuum, fluid mechanics, and optical effects. Each valid–invalid pair is carefully designed to isolate specific physics violations (e.g., energy, mass, continuity) while holding appearance constant, ensuring the evaluation reflects genuine physics reasoning rather than visual bias.\n4. The proposed Plausibility Preference Error (PPE) metric quantifies the proportion of cases where a model fails to prefer physically valid samples. It is intuitive, numerically stable, and easily comparable across models. Moreover, the authors demonstrate that PPE aligns strongly with human judgments of physical consistency while being independent from standard visual quality metrics, confirming its validity and specificity."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. A key limitation of the proposed likelihood-preference framework lies in its implicit assumption that higher estimated likelihood (i.e., lower denoising loss) reflects stronger physical understanding. In practice, the likelihood assigned by a diffusion model is influenced by many confounding factors beyond physics correctness. For instance, if a video sample—whether physically valid or invalid—resembles patterns frequently seen in the training data, it may naturally obtain a higher likelihood simply due to data distribution similarity, rather than genuine adherence to physical laws. Conversely, both valid and invalid videos that deviate from the model’s training distribution could be assigned uniformly low likelihoods, making the difference between them statistically insignificant. As a result, the LikePhys metric may sometimes conflate distribution familiarity with physical plausibility, leading to inaccurate or unstable evaluations, especially when the model exhibits strong dataset bias. This issue raises questions about the robustness and interpretability of using likelihood differences alone as a proxy for intuitive physics understanding.\n2. Although the benchmark covers four major physics domains, it remains synthetic and controlled, relying solely on Blender-rendered simulations. While this ensures experimental rigor, it limits the method’s ability to generalize to real-world, noisy, or unstructured videos, where visual complexity, uncertainty, and imperfect physical consistency are common. The results might therefore overestimate models’ real-world physics reasoning ability.\n3. The metric treats intuitive physics understanding as a pairwise likelihood preference problem, which captures surface-level plausibility but may fail to reflect causal reasoning, temporal prediction, or long-horizon dynamics that are essential for deeper physical understanding. As a result, models that memorize motion patterns could perform well on LikePhys without genuinely learning the underlying physical principles.\n4. While the paper reports domain-level PPE scores, it provides limited qualitative analysis or visual diagnosis of why certain models fail under specific laws (e.g., temporal continuity or conservation of mass). More detailed case studies or ablation examples could have strengthened interpretability and clarified whether errors stem from architecture limitations, data bias, or diffusion noise modeling."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918334376,"tcdate":1761464998590,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5897/Reviewer_y5Dy"],"signatures":["ICLR.cc/2026/Conference/Submission5897/Reviewer_y5Dy"],"forum":"6UJf6B8RZ8","number":1,"license":"CC BY 4.0","cdate":1761464998590,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5897/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918334376,"domain":"ICLR.cc/2026/Conference","replyto":"6UJf6B8RZ8","id":"mYaB0rsKA3","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generative Model","Video Diffusion Model","Intuitive Physics Understanding"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human\npreference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale up."},"_bibtex":{"value":"@inproceedings{\nyuan2026likephys,\ntitle={LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference},\nauthor={Jianhao Yuan and Fabio Pizzati and Francesco Pinto and Lars Kunze and Ivan Laptev and Paul Newman and Philip Torr and Daniele De Martini},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6UJf6B8RZ8}\n}"},"title":{"value":"LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference"},"pdf":{"value":"/pdf/be5f5c438a68c2b02423c0092164e93b5f3d245d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|likephys_evaluating_intuitive_physics_understanding_in_video_diffusion_models_via_likelihood_preference"},"authorids":{"value":["~Jianhao_Yuan2","~Fabio_Pizzati1","~Francesco_Pinto1","~Lars_Kunze1","~Ivan_Laptev1","~Paul_Newman1","~Philip_Torr1","~Daniele_De_Martini1"]},"authors":{"value":["Jianhao Yuan","Fabio Pizzati","Francesco Pinto","Lars Kunze","Ivan Laptev","Paul Newman","Philip Torr","Daniele De Martini"]}},"version":2},{"content":{"summary":{"value":"This paper explores the paradigm of Physics informed Machine Learning in the context of tradeoff between efficiency and accuracy. The authors deep dive into how introducing physics biases into pure data-driven methods can reduce carbon footprint of the computational heavy deep learning models. The authors provide in-depth empirical results for the simplest physics problem (oscillation) to more complex physical phenomenon (navier-stokes). The experimental setup shows during training pure data-driven methods like UNet variants perform better in terms of accuracy and strong physics methods like FNO perform better in terms of efficiency with lowest carbon footprint. However, during inference FNO variants and Flow Matching models costs higher than pure data-driven models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. An experimental setup on real world dataset would definitely strengthen the study.\n2. Additional experiments on real world dataset exploring the costs at training and inference times \n3. More insights into why the carbon footprint landscape changes during training and inference time. for example can authors provide an insight into why FM inference cost is so high, does it solely have to do with the increase in roll-out steps?\n4. Figure 8 a shows the predictive performance vs carbon footprint at training time, could authors include the same figure or results for inference time as well for longer rollouts?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper explores an important subject of carbon footprint of pure data-driven methods (deep learning models).\n2. In-depth experiments are performed targeting simplest physics to more complex physics phenomena. \n3. The paper presents both efficacy and efficiency metrics at both training and inference times.\n4. The discussion section of the paper is clearly and well-written highlighting how physics alone is not the answer of reducing carbon footprint."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The authors emphasise on the trade-off for carbon footprint and accuracy, however as the problem gets complex the trade-off seems to disappear as Unet based variants do perform better at inference times as well as have a lower carbon footprint, while FM which is supposed to be a mid point for physics and data-driven has the highest carbon footprint at inference time.\n2. The study is solely done on synthetic dataset, it would be good to see some examples, efficacy and efficiency metrics for some real world datasets.\n3. In my point of view the study is largely experimental and I couldn't pinpoint the innovative aspect of it. Therefore, I think it needs to be backed up by a lot more experiments convincing that there indeed exists a accuracy–carbon trade-off in introducing strong and weak physics inductive bias."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917133904,"tcdate":1760487968827,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4009/Reviewer_Y12R"],"signatures":["ICLR.cc/2026/Conference/Submission4009/Reviewer_Y12R"],"forum":"kfwKNYvjS4","number":1,"license":"CC BY 4.0","cdate":1760487968827,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4009/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917133904,"domain":"ICLR.cc/2026/Conference","replyto":"kfwKNYvjS4","id":"OJGFugluFn","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"We promote evaluation practices that consider both efficacy and efficiency, and characterise the trade-offs through physics-inductive spatio-temporal forecasting models."},"keywords":{"value":["physics informed machine learning","energy consumption","carbon footprint","flow matching"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Development of modern deep learning methods has been driven primarily by the push for improving model efficacy (accuracy metrics). This sole focus on efficacy has steered development of large-scale models that require massive resources, and results in considerable carbon footprint across the model life-cycle. In this work, we explore how physics inductive biases can offer useful trade-offs between model efficacy and model efficiency (compute, energy, and carbon). We study a variety of models for spatio-temporal forecasting, a task governed by physical laws and well-suited for exploring different levels of physics inductive bias. We show that embedding physics inductive biases into the model design can yield substantial efficiency gains while retaining or even improving efficacy for the tasks under consideration. In addition to using standard physics-informed spatio-temporal models, we demonstrate the usefulness of more recent models like flow matching as a general purpose method for spatio-temporal forecasting. Our experiments show that incorporating physics inductive biases offer a principled way to improve the efficiency and reduce the carbon footprint of machine learning models. We argue that model efficiency, along with model efficacy, should become a core consideration driving machine learning model development and deployment."},"_bibtex":{"value":"@misc{\nwilson2026trading,\ntitle={Trading Carbon for Physics:  On the Resource Efficiency of Machine Learning for Spatio-Temporal Forecasting},\nauthor={Sophia N. Wilson and Jens H. Christensen and Raghavendra Selvan},\nyear={2026},\nurl={https://openreview.net/forum?id=kfwKNYvjS4}\n}"},"title":{"value":"Trading Carbon for Physics:  On the Resource Efficiency of Machine Learning for Spatio-Temporal Forecasting"},"pdf":{"value":"/pdf/03fc05a5f58e13e2ed4a153d9b69ccb5a5396409.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wilson|trading_carbon_for_physics_on_the_resource_efficiency_of_machine_learning_for_spatiotemporal_forecasting"},"authorids":{"value":["~Sophia_N._Wilson1","~Jens_H._Christensen1","~Raghavendra_Selvan1"]},"authors":{"value":["Sophia N. Wilson","Jens H. Christensen","Raghavendra Selvan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes T2-PILOT, a method that jointly optimizes k-space trajectories, reconstruction, and T2 estimation by enforcing the exponential decay model. The framework integrates physics constraints into trajectory learning and includes optional test-time refinement. Experiments on CMRxRecon show consistent but small improvements over fixed and unconstrained learned trajectories in PSNR and T2 accuracy. The work highlights the benefit of aligning acquisition design with quantitative imaging objectives."},"review":{"value":"Overall, the paper is well-motivated and technically sound. Incorporating the T2 decay model into trajectory optimization is intuitive and aligns with the goal of task-driven MRI acquisition.\nStrengths:\n- Clear physics-informed formulation \n- Joint optimization of acquisition and estimation \n- Consistent improvements across settings \n- Clinically relevant application \n\nWeaknesses:\n- Gains are relatively small (e.g., ~0.1 dB PSNR) \n- Novelty is somewhat incremental over prior physics-informed / task-driven methods \n- Limited analysis (e.g., ablations, robustness, stronger baselines) \n- Some components (e.g., temporal embedding) lack justification \n- Test-time fine-tuning adds complexity with marginal benefit\n\nOverall, while the idea is solid, the paper would benefit from stronger empirical evidence and deeper analysis of why the method works."},"strengths":{"value":"The paper presents a clean and well-motivated integration of physics constraints into trajectory learning for quantitative MRI. The joint optimization framework is technically sound and aligns well with task-driven imaging goals. The method is evaluated across multiple sampling schemes and consistently improves both reconstruction quality and T2 estimation. The problem is clinically relevant, and the approach is conceptually appealing."},"weaknesses":{"value":"The main limitation is the modest improvement over baselines, which raises questions about practical impact. The novelty is incremental relative to existing physics-informed and task-driven acquisition methods. The paper lacks deeper analysis, including ablations and robustness studies. Some design choices are not well justified, and the benefit of test-time fine-tuning appears limited compared to its added complexity."},"confidence":{"value":4},"rating":{"value":3},"justification_of_rating":{"value":"The method is reasonable and well-executed, but the contribution is incremental and the improvements are modest. Stronger experimental validation and deeper analysis would be needed to clearly justify acceptance."},"title":{"value":"Physics-informed kspace sampling trajectory learning for T2 mapping"}},"parentInvitations":"MIDL.io/2026/Short_Papers/-/Official_Review","nonreaders":[],"tmdate":1778421293010,"tcdate":1777862651401,"writers":["MIDL.io/2026/Short_Papers","MIDL.io/2026/Short_Papers/Submission35/Reviewer_a61h"],"signatures":["MIDL.io/2026/Short_Papers/Submission35/Reviewer_a61h"],"forum":"42yNhuBQQf","number":1,"license":"CC BY 4.0","cdate":1777862651401,"readers":["everyone"],"invitations":["MIDL.io/2026/Short_Papers/Submission35/-/Official_Review","MIDL.io/2026/Short_Papers/-/Edit"],"mdate":1778421293010,"domain":"MIDL.io/2026/Short_Papers","replyto":"42yNhuBQQf","id":"kYXNjSKGAY","forumContent":{"TLDR":{"value":"T2-PILOT reduces per-beat acquisition time for cardiac MRI by integrating the $T_2$ decay model into the joint optimization of trajectories and reconstruction, ensuring superior quantitative accuracy."},"venue":{"value":"MIDL 2026 - Short Papers Poster"},"keywords":{"value":["Cardiac MRI","$T_2$ Mapping","Trajectory Optimization and Reconstruction","Physics-Informed Deep-Learning"]},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Cardiac MRI $T_2$ mapping is essential for diagnosing myocardial pathologies, but prolonged acquisitions often extend beyond the diastolic rest phase, leading to motion artifacts and reduced reliability. While deep learning accelerates imaging via k-space undersampling, existing learned trajectories optimize for image reconstruction and neglect the underlying $T_2$ relaxation physics. We propose T2-PILOT, which jointly optimizes non-Cartesian k-space trajectories and $T_2$ map estimation by enforcing the $T_2$ decay model during training, with additional subject-specific test-time fine-tuning. On the CMRxRecon dataset, T2-PILOT outperforms both fixed and unconstrained reconstruction-guided trajectories. Under high undersampling, it improves image quality and quantitative $T_2$ accuracy while reducing per-beat acquisition time by 54\\% (32 vs. 70 spokes), yielding a PSNR gain of 0.14 dB (35.06 vs. 34.92) and a $T_2$ map accuracy improvement of 0.67 dB (31.96 vs. 31.29). These results demonstrate that incorporating physics-based constraints into trajectory learning enables more accurate, robust, and clinically reliable accelerated $T_2$ mapping."},"_bibtex":{"value":"@inproceedings{\ngavrielov2026tpilot,\ntitle={T2-{PILOT}: Optimized Trajectories for \\$T\\_2\\$ Mapping Acceleration},\nauthor={Naama Gavrielov and Tamir Shor and Alex M. Bronstein and Moti Freiman},\nbooktitle={Medical Imaging with Deep Learning - Short Papers},\nyear={2026},\nurl={https://openreview.net/forum?id=42yNhuBQQf}\n}"},"title":{"value":"T2-PILOT: Optimized Trajectories for $T_2$ Mapping Acceleration"},"pdf":{"value":"/pdf/8f89c53d0df200faa99d6bf86e436e68ec3c03a4.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Short_Papers"},"paperhash":{"value":"gavrielov|t2pilot_optimized_trajectories_for_t_2_mapping_acceleration"},"authorids":{"value":["~Naama_Gavrielov1","~Tamir_Shor1","~Alex_M._Bronstein1","~Moti_Freiman1"]},"registration":{"value":"Yes"},"authors":{"value":["Naama Gavrielov","Tamir Shor","Alex M. Bronstein","Moti Freiman"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"The paper introduces DynSuperCLEVR, a video question answering (VideoQA) dataset that emphasizes understanding dynamic 3D object properties within 4D (3D + time) scenes. \nAdditionally, the authors present NS-4DPhysics, a model that combines neural-symbolic reasoning with physics-based priors to analyze these dynamic properties. The model first constructs an explicit 4D scene representation using a 3D generative model, followed by neural-symbolic reasoning to answer questions. \nExperimental results demonstrate that NS-4DPhysics surpasses existing VideoQA models across various question types (factual, predictive, and counterfactual), underscoring its effectiveness in reasoning about object dynamics in complex, synthetic environments."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Have the authors considered extending the dataset to include articulated or deformable objects? If so, what challenges or limitations do they anticipate with this extension?\n\n2. Could the authors provide an efficiency analysis of the proposed model, including resource usage and runtime under typical conditions?\n\n3. What modifications to the NS-4DPhysics framework would make it more efficient for real-time performance or deployment in resource-limited environments?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"**1. Novel Dataset:** DynSuperCLEVR is a novel dataset that focuses on 4D dynamics, addressing a critical gap in existing VideoQA datasets which typically overlook explicit physics-based scene understanding.\n\n**2. Innovative Model Design:** The NS-4DPhysics model combines 3D generative modeling with physics-informed priors, represents an innovative approach to handling dynamic 4D scene reasoning.\n\n**3. Comprehensive Benchmarking:** Extensive evaluations against baseline models, including video large language models (Video-LLMs) and other symbolic frameworks, highlight the superior performance of NS-4DPhysics in capturing 4D dynamics.\n\n**4. Future and Counterfactual Simulations:** By leveraging physics-based priors, the model excels at simulating both future and hypothetical states, demonstrating practical value and broad application potential."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Synthetic Data Limitations:** While the dataset is suitable for testing dynamic properties, its synthetic nature may limit generalizability to real-world applications. Despite the authors’ efforts to improve aspects like background realism (L201), models trained exclusively on synthetic data often struggle to handle real-world noise and variability.\n\n**2. Computational Complexity:** The NS-4DPhysics model is computationally demanding due to its reliance on 3D generative modeling and physics-based priors, presenting challenges for scalability and use in resource-constrained environments.\n\n**3. Limited Object Diversity:** The dataset is limited to a narrow range of rigid objects, which may not adequately represent the complexity of real-world scenes that often include deformable or articulated objects.\n\n**4. Evaluation of Real-World Applicability:** The paper lacks an analysis of the model’s performance on real-world video data, which is essential for evaluating its practical applicability outside synthetic benchmarks."}},"nonreaders":[],"tmdate":1731427265099,"tcdate":1730707085575,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission479/Reviewer_jmWM"],"signatures":["ICLR.cc/2025/Conference/Submission479/Reviewer_jmWM"],"forum":"6Vx28LSR7f","number":4,"license":"CC BY 4.0","cdate":1730707085575,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission479/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427265099,"domain":"ICLR.cc/2025/Conference","replyto":"6Vx28LSR7f","id":"C11ZWQLsKS","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"We introduce DynSuperCLEVR, a video question answering dataset focused on the dynamic properties of 3D objects. We propose NS-4DPhysics, which use 4D world states with a 3D generative model and uses neural symbolic reasoning to answer questions."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video question answering","Compositional reasoning","Physical scene understanding","3D scene understanding"]},"supplementary_material":{"value":"/attachment/9df5057549ac56979fea8269802d34b5d654e102.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-level temporal and action semantics. Although humans are adept at understanding these properties by constructing 3D and temporal (4D) representations of the world, current video understanding models struggle to extract these dynamic semantics, arguably because these models use cross-frame reasoning without underlying knowledge of the 3D/4D scenes.\nIn this work, we introduce **DynSuperCLEVR**, the first video question answering dataset that focuses on language understanding of the dynamic properties of 3D objects. We concentrate on three physical concepts—*velocity*, *acceleration*, and *collisions*—within 4D scenes. We further generate three types of questions, including factual queries, future predictions, and counterfactual reasoning that involve different aspects of reasoning on these 4D dynamic properties.\nTo further demonstrate the importance of explicit scene representations in answering these 4D dynamics questions, we propose **NS-4DPhysics**, a **N**eural-**S**ymbolic VideoQA model integrating **Physics** prior for **4D** dynamic properties with explicit scene representation of videos. \nInstead of answering the questions directly from the video text input, our method first estimates the 4D world states with a 3D generative model powered by a physical prior, and then uses neural symbolic reasoning to answer the questions based on the 4D world states.\nOur evaluation on all three types of questions in DynSuperCLEVR shows that previous video question answering models and large multimodal models struggle with questions about 4D dynamics, while our NS-4DPhysics significantly outperforms previous state-of-the-art models."},"_bibtex":{"value":"@inproceedings{\nwang2025compositional,\ntitle={Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering},\nauthor={Xingrui Wang and Wufei Ma and Angtian Wang and Shuo Chen and Adam Kortylewski and Alan Yuille},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=6Vx28LSR7f}\n}"},"title":{"value":"Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering"},"pdf":{"value":"/pdf/9ebb092c1ee905ec7df8632cc804269f09324d4a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"wang|compositional_4d_dynamic_scenes_understanding_with_physics_priors_for_video_question_answering"},"authorids":{"value":["~Xingrui_Wang1","~Wufei_Ma1","~Angtian_Wang2","~Shuo_Chen13","~Adam_Kortylewski1","~Alan_Yuille1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xingrui Wang","Wufei Ma","Angtian Wang","Shuo Chen","Adam Kortylewski","Alan Yuille"]}},"version":2},{"content":{"decision":{"value":"Reject"},"comment":{"value":"The work proposes IntPhys 2, a video benchmark for evaluation of the intuitive physics understanding of deep learning models. IntPhys 2 focuses on core principles related to macroscopic objects, namely, permanence, immutability, spatio-temporal continuity, and solidity. These are concepts related to developmental psychology in the context of early childhood. In addition to the benchmark, the paper provides performance evaluations of several state-of-the-art models, which show that these models struggle to capture intuitive physics across the four principles in complex scenes (most models performing at chance levels). As such the benchmark provides a useful resource for future studies. The paper initially received all positive reviews accompanied by a number of questions. A rebuttal phase followed, along with extensive discussion, which clarified various aspects, including comparisons with other existing benchmarks, a question about an extended set of physical principles, architectural issues, etc. The reviewers' final recommendations converged on 4x Borderline Accept. I agree with the recommendations and have indicated accordingly.\n\n===== FINAL UPDATE FROM DB Track PCs ====\n\nThe final decision for this paper has been taken by the program chairs after consultation with the SACs. All Senior Area Chairs have ranked papers according to the feedback from the AC during the review process. We decided to leave the original meta-review to reflect the opinion of the AC in light of the initial discussions with reviewers and SAC."},"title":{"value":"Paper Decision"}},"parentInvitations":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Decision","nonreaders":[],"tmdate":1761796510539,"tcdate":1758198235393,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Program_Chairs"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Program_Chairs"],"forum":"Xpf5x3mLvn","number":1,"license":"CC BY 4.0","cdate":1758198235393,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Decision","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1761796510539,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","replyto":"Xpf5x3mLvn","id":"8XazITjcWa","forumContent":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Video benchmark","Intuitive Physic","Evaluation"]},"supplementary_material":{"value":"/attachment/4932fdafe9cb6b17072f67b2caf246775bce6feb.zip"},"_bibtex":{"value":"@misc{\nbordes2026intphys,\ntitle={IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments},\nauthor={Florian Bordes and Quentin Garrido and Justine T Kao and Adina Williams and Michael Rabbat and Emmanuel Dupoux},\nyear={2026},\nurl={https://openreview.net/forum?id=Xpf5x3mLvn}\n}"},"title":{"value":"IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments"},"paperhash":{"value":"bordes|intphys_2_benchmarking_intuitive_physics_understanding_in_complex_synthetic_environments"},"dataset_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh.zip"},"TLDR":{"value":"We introduced a new intuitive physic benchmark IntPhys 2, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments"},"primary_area":{"value":"datasets_&_benchmarks_for_computer_vision"},"abstract":{"value":"We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to  macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50\\%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies."},"pdf":{"value":"/pdf/c29b03435c4329f0422348326d7ca2b5d606960a.pdf"},"croissant_file":{"value":"/attachment/54b155200f4140df0470921a91b0093300c98b6a.json"},"code_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh_code.zip"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"authorids":{"value":["~Florian_Bordes1","~Quentin_Garrido1","~Justine_T_Kao1","~Adina_Williams1","~Michael_Rabbat1","~Emmanuel_Dupoux1"]},"authors":{"value":["Florian Bordes","Quentin Garrido","Justine T Kao","Adina Williams","Michael Rabbat","Emmanuel Dupoux"]}},"version":2},{"content":{"venue":{"value":"Intuitive interaction"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"blackler|perspectives_on_the_nature_of_intuitive_interaction"},"html":{"value":"https://researchers.mq.edu.au/en/publications/9006062b-91c4-470b-be1f-c0354fd542f5"},"abstract":{"value":"Intuitive interaction is defined as fast, somewhat non-conscious, and generally accurate interaction with an interface that is informed by past experience or technology familiarity (TF). Eighteen years of research into intuitive interaction by various researchers on four different continents using a variety of products, interfaces, and experiment designs has shown that prior experience is the leading contributor to intuitive interaction. In Intuitive Use of User Interfaces (IUUI) continuum, the most basic and broadly possessed knowledge identified is innate knowledge, which has genetic origins and manifests in responses such as reflexes. In the Australian continuum, the most accessible design strategy is to use physical affordances, which take advantage of embodied knowledge of the world established. Tangible and embodied interfaces (TEIs) and natural user interfaces (NUIs) have long been claimed to be intuitive. This intuitiveness is attributed to tactile or haptic interactions in terms of static system properties such as directness, ease of learning and naturalness and speed, simplicity, and effectiveness."},"_bibtex":{"value":"@inbook{9006062b91c4470bbe1fc0354fd542f5,\n  title     = \"Perspectives on the nature of intuitive interaction\",\n  abstract  = \"Intuitive interaction is defined as fast, somewhat non-conscious, and generally accurate interaction with an interface that is informed by past experience or technology familiarity (TF). Eighteen years of research into intuitive interaction by various researchers on four different continents using a variety of products, interfaces, and experiment designs has shown that prior experience is the leading contributor to intuitive interaction. In Intuitive Use of User Interfaces (IUUI) continuum, the most basic and broadly possessed knowledge identified is innate knowledge, which has genetic origins and manifests in responses such as reflexes. In the Australian continuum, the most accessible design strategy is to use physical affordances, which take advantage of embodied knowledge of the world established. Tangible and embodied interfaces (TEIs) and natural user interfaces (NUIs) have long been claimed to be intuitive. This intuitiveness is attributed to tactile or haptic interactions in terms of static system properties such as directness, ease of learning and naturalness and speed, simplicity, and effectiveness.\",\n  author    = \"Alethea Blackler and Shital Desai and Mitchell McEwan and Vesna Popovic and Sarah Diefenbach\",\n  year      = \"2018\",\n  doi       = \"10.1201/b22191\",\n  language  = \"English\",\n  pages     = \"19--39\",\n  editor    = \"Alethea Blackler\",\n  booktitle = \"Intuitive interaction\",\n  publisher = \"CRC Press (Taylor and Francis)\",\n}"},"title":{"value":"Perspectives on the nature of intuitive interaction"},"authors":{"value":[{"fullname":"Alethea Blackler"},{"fullname":"Shital Desai"},{"fullname":"Mitchell McEwan","username":"~Mitchell_McEwan1"},{"fullname":"Vesna Popovic"},{"fullname":"Sarah Diefenbach"}]}},"tmdate":1789090879725,"pdate":1514764800000,"externalIds":["doi:10.1201/b22191"],"tcdate":1762732240840,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Mitchell_McEwan1"],"forum":"f40QFX2fWM","license":"CC BY-SA 4.0","number":12548,"cdate":1578614114253,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789090879725,"domain":"OpenReview.net/Public_Article","id":"f40QFX2fWM","version":2},{"content":{"comment":{"value":"**Q2: Although the benchmark covers four major physics domains, it remains synthetic and controlled, relying solely on Blender-rendered simulations. While this ensures experimental rigor, it limits the method’s ability to generalize to real-world, noisy, or unstructured videos, where visual complexity, uncertainty, and imperfect physical consistency are common. The results might therefore overestimate models’ real-world physics reasoning ability.**\n\n\n\n\n**A2**: Our method follows the classic violation-of-expectation paradigm [1,2], which is mostly built on simulated data and has been successful in evaluating the physics understanding of various vision models. It is essentially impossible to obtain large-scale real-world test data with precisely controlled violations of physical laws, so simulation is a natural choice for isolating physics.\nWhile real-world videos are indeed more visually complex and often contain imperfect physics, this does not exempt models from making errors on much simpler, controlled violations. In our benchmark, each valid–invalid video pair differs only by a specific, well-defined physics violation. This allows us to attribute failures to particular laws or scenarios and provides a lower bound on a model’s physics capacity: if a model struggles on these controlled cases, it is unlikely to handle more noisy, unstructured real-world scenes.\n\n\nIn the evaluation protocol, we deliberately control confounding factors through simulation. Each physics scenario is designed to be relatively simple, with clearly attributable governing dynamics and consistent appearance. Within a variation, valid and invalid videos share the same prompt, camera, lighting, textures, and geometry, and differ only in physics adherence. Our method then compares the relative likelihood within each valid–invalid pair, rather than the absolute likelihood of a single sample. This pairwise design is intended to cancel out model-specific responses to visual appearance or style, so that the remaining likelihood difference arises from the controlled physics violation. Overall, the synthetic setting is used to isolate and probe physics in a controlled way; it does not claim to capture the full complexity of real-world videos, but it provides a rigorous and interpretable test of whether models respect basic physical regularities under matched visual conditions.\n\n\n[1] Riochet, R., Castro, M.Y., Bernard, M., Lerer, A., Fergus, R., Izard, V. and Dupoux, E., 2018. Intphys: A framework and benchmark for visual intuitive physics reasoning. arXiv preprint arXiv:1803.07616.\n\n\n[2] Bordes, F., Garrido, Q., Kao, J.T., Williams, A., Rabbat, M. and Dupoux, E., 2025. IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments. arXiv preprint arXiv:2506.09849.\n\n\n\n\n\n\n\n\n**Q3: The metric treats intuitive physics understanding as a pairwise likelihood preference problem, which captures surface-level plausibility but may fail to reflect causal reasoning, temporal prediction, or long-horizon dynamics that are essential for deeper physical understanding. As a result, models that memorize motion patterns could perform well on LikePhys without genuinely learning the underlying physical principles.**\n\n\n\n\n**A3**: Thanks for raising this interesting point.\nAs discussed in our response to Q1, we take an operational view and define a model’s intuitive physics capacity as its ability to assign higher likelihood to physically valid videos than to invalid ones under controlled violations. LikePhys is therefore designed to test whether the learned distribution aligns with physics-plausible patterns, irrespective of whether the internal mechanism is “pattern memorization” or an explicit causal model.\nRegarding the concern about “surface-level plausibility”, several of our scenarios already require non-trivial temporal reasoning (e.g., tracking momentum exchange over time, delayed collisions, occlusions), so a model cannot succeed purely from a single static frame. As shown in our main results (Sec. 4.1 and Sec. 4.5), current video diffusion models still make systematic errors on these relatively simple controlled valid–invalid pairs, which suggests that their physics capacity is limited even before considering more complex, long-horizon settings.\nWe agree that PPE does not cover all aspects of deep causal reasoning or arbitrarily long-term prediction, and we do not claim it is a complete test of physics understanding. It is a first-step, controlled pairwise probe under matched visual conditions. Extending this framework to longer horizons, multi-step counterfactuals, or explicitly causal tasks is a natural direction for future work. Importantly, our human study shows that PPE already correlates well with human judgments of physical plausibility in downstream text-to-video generation, indicating that it captures a practically relevant component of “physics quality” observed by users."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763657531048,"tcdate":1763657531048,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5897/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission5897/Authors"],"forum":"6UJf6B8RZ8","number":8,"license":"CC BY 4.0","cdate":1763657531048,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5897/-/Official_Comment"],"mdate":1763657531048,"domain":"ICLR.cc/2026/Conference","replyto":"63FhnUfVSy","id":"uZKCzmHEiv","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generative Model","Video Diffusion Model","Intuitive Physics Understanding"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human\npreference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale up."},"_bibtex":{"value":"@inproceedings{\nyuan2026likephys,\ntitle={LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference},\nauthor={Jianhao Yuan and Fabio Pizzati and Francesco Pinto and Lars Kunze and Ivan Laptev and Paul Newman and Philip Torr and Daniele De Martini},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6UJf6B8RZ8}\n}"},"title":{"value":"LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference"},"pdf":{"value":"/pdf/be5f5c438a68c2b02423c0092164e93b5f3d245d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|likephys_evaluating_intuitive_physics_understanding_in_video_diffusion_models_via_likelihood_preference"},"authorids":{"value":["~Jianhao_Yuan2","~Fabio_Pizzati1","~Francesco_Pinto1","~Lars_Kunze1","~Ivan_Laptev1","~Paul_Newman1","~Philip_Torr1","~Daniele_De_Martini1"]},"authors":{"value":["Jianhao Yuan","Fabio Pizzati","Francesco Pinto","Lars Kunze","Ivan Laptev","Paul Newman","Philip Torr","Daniele De Martini"]}},"version":2},{"content":{"summary":{"value":"- The paper proposes IntPhys 2, a video-based benchmark designed to evaluate AI models' understanding of intuitive physics (e.g., object permanence, invariance, spatiotemporal continuity, solidity). Core innovations include:\n  - Constructing complex scenarios: Building high-fidelity environments with Unreal Engine (dynamic lighting, textures, occlusions) and introducing dynamic camera movements (simulating human perspective changes), breaking free from traditional static scenarios (e.g., simple occlusions in the original IntPhys).\n  - Comprehensive testing of MLLMs (e.g., Gemini, GPT4-o) and predictive models (e.g., VideoMAEv2, V-JEPA), establishing new benchmarks and identifying specific challenges in intuitive physical reasoning."},"ethical_considerations":{"value":"No, there are no or only very minor ethics concerns"},"dataset_code_accessibility":{"value":"Yes"},"responsible_reviewing_acknowledgement":{"value":"Yes"},"code_of_conduct_acknowledgement":{"value":"Yes"},"confidence":{"value":3},"rating":{"value":4},"final_justification":{"value":"The response has addressed all my questions. The author is encouraged to include this discussion in the revision."},"limitations_weaknesses":{"value":"- The innovation of this paper seems incremental, as its distinction from the existing IntPhy benchmark seems limited to the complexity of objects, lighting, etc., in scenarios. The authors need to clarify the fundamental differences between their benchmark and other existing methods like IntPhy and GRASP [1].\n- The authors present the benchmark test results, but only stay at performance comparisons, failing to deeply analyze the underlying causes of model failures (e.g., whether continuity errors stem from motion prediction flaws) and also failing to propose targeted improvement directions.\n\nRef: [1] Serwan Jassim, Mario Holubar, Annika Richter, Cornelius Wolff, Xenia Ohmer, and Elia Bruni. 2024. GRASP: a novel benchmark for evaluating language grounding and situated physics understanding in multimodal language models. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI '24)"},"strengths_contributions":{"value":"- Rigorous benchmark design: The paper uses dynamic cameras and complex occlusions (with the longest occlusion duration exceeding existing benchmarks) to simulate real-world scenarios, enhancing the ecological validity of the tasks.\n- The paper is well-written, logically structured, concise, and clear, making it easy for readers to understand."}},"parentInvitations":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Official_Review","nonreaders":[],"tmdate":1761794722998,"tcdate":1751467850027,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_q2be"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_q2be"],"forum":"Xpf5x3mLvn","number":3,"license":"CC BY 4.0","cdate":1751467850027,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Official_Review","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Official_Review3/-/Review_Revision"],"mdate":1761794722998,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","replyto":"Xpf5x3mLvn","id":"jVZdtEOm3Z","forumContent":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Video benchmark","Intuitive Physic","Evaluation"]},"supplementary_material":{"value":"/attachment/4932fdafe9cb6b17072f67b2caf246775bce6feb.zip"},"_bibtex":{"value":"@misc{\nbordes2026intphys,\ntitle={IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments},\nauthor={Florian Bordes and Quentin Garrido and Justine T Kao and Adina Williams and Michael Rabbat and Emmanuel Dupoux},\nyear={2026},\nurl={https://openreview.net/forum?id=Xpf5x3mLvn}\n}"},"title":{"value":"IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments"},"paperhash":{"value":"bordes|intphys_2_benchmarking_intuitive_physics_understanding_in_complex_synthetic_environments"},"dataset_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh.zip"},"TLDR":{"value":"We introduced a new intuitive physic benchmark IntPhys 2, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments"},"primary_area":{"value":"datasets_&_benchmarks_for_computer_vision"},"abstract":{"value":"We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to  macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50\\%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies."},"pdf":{"value":"/pdf/c29b03435c4329f0422348326d7ca2b5d606960a.pdf"},"croissant_file":{"value":"/attachment/54b155200f4140df0470921a91b0093300c98b6a.json"},"code_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh_code.zip"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"authorids":{"value":["~Florian_Bordes1","~Quentin_Garrido1","~Justine_T_Kao1","~Adina_Williams1","~Michael_Rabbat1","~Emmanuel_Dupoux1"]},"authors":{"value":["Florian Bordes","Quentin Garrido","Justine T Kao","Adina Williams","Michael Rabbat","Emmanuel Dupoux"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a geometry-embedded deep cone beam CT reconstruction method for non-circular trajectories, where differentiable backprojection operators derived from the actual source-detector geometry are built directly into the network. The architecture combines a projection-domain preprocessing network, a U-Net-like reconstruction network with multi-scale backprojection skip connections, and a postprocessing network for refinement. On synthetic data with simulated trajectory variation, the method substantially outperforms standard FBP in both MSE and SSIM."},"review":{"value":"The paper’s main strengths are that it tackles a highly relevant CBCT problem, uses a physics-informed architecture that embeds acquisition geometry directly through differentiable backprojection operators, and shows a clear, intuitive design with strong synthetic gains over FBP. The multi-scale geometry-aware reconstruction idea is particularly appealing because it aligns well with how projection data should be mapped into image space. Its main weaknesses are that the evidence is still limited to synthetic experiments, the baselines are relatively weak, and the claim of handling “arbitrary trajectories” is broader than what is actually validated in the paper."},"strengths":{"value":"- CBCT reconstruction for flexible or non-circular trajectories is an important application \n- Geometry is embedded through differentiable backprojection operators rather than treated only implicitly.\n- Strong synthetic results versus FBP, with both quantitative and qualitative improvements.\n- Clearly written paper"},"weaknesses":{"value":"- Validation is limited to synthetic data\n- Baseline methods are not very strong\n- The title/claim of “arbitrary trajectories” is broader than what is experimentally demonstrated.\n- More ablation is needed to isolate the value of the embedded multi-scale backprojection design."},"confidence":{"value":3},"rating":{"value":4},"justification_of_rating":{"value":"The idea is compelling but the experimental validation is still limited relative to the breadth of the arbitrary trajectory claim."},"title":{"value":"strong short paper limited to synthetic results"}},"parentInvitations":"MIDL.io/2026/Short_Papers/-/Official_Review","nonreaders":[],"tmdate":1778421295229,"tcdate":1776970082452,"writers":["MIDL.io/2026/Short_Papers","MIDL.io/2026/Short_Papers/Submission63/Reviewer_EGyn"],"signatures":["MIDL.io/2026/Short_Papers/Submission63/Reviewer_EGyn"],"forum":"60PLqW11ZQ","number":1,"license":"CC BY 4.0","cdate":1776970082452,"readers":["everyone"],"invitations":["MIDL.io/2026/Short_Papers/Submission63/-/Official_Review","MIDL.io/2026/Short_Papers/-/Edit"],"mdate":1778421295229,"domain":"MIDL.io/2026/Short_Papers","replyto":"60PLqW11ZQ","id":"66WnaChQO2","forumContent":{"venue":{"value":"MIDL 2026 - Short Papers Poster"},"keywords":{"value":["CT Reconstruction","Deep Learning","Non-Circular Scanning Trajectories"]},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Flexible scanning trajectories for CBCT systems offer enhanced imaging capabilities but\npose challenges for conventional reconstruction algorithms. We present a novel deep learning-\nbased reconstruction framework that explicitly integrates trajectory geometry information\nthrough differentiable back projection operators embedded within the network architecture.\nOur approach comprises three cascaded networks: a preprocessing network for projection\nfiltering, an encoder-decoder reconstruction network with multi-scale back projection opera-\ntors, enabling transition from projection to image domain, and a postprocessing network for\nfinal refinement. Evaluation on 2,000 test objects demonstrates substantial improvements\nover conventional filtered back projection (FBP), achieving an MSE of 4.8e−3 (vs.2.4e−2\nfor FBP) and SSIM of 0.96 (vs. 0.54 for FBP)."},"_bibtex":{"value":"@inproceedings{\nblum2026geometryembedded,\ntitle={Geometry-Embedded Neural Networks Cone Beam {CT} Reconstruction for Arbitrary Scanning Trajectories},\nauthor={Nele Blum and Max Stickel and Moritz Schaar and Thorsten Buzug and Maik Stille},\nbooktitle={Medical Imaging with Deep Learning - Short Papers},\nyear={2026},\nurl={https://openreview.net/forum?id=60PLqW11ZQ}\n}"},"title":{"value":"Geometry-Embedded Neural Networks Cone Beam CT Reconstruction for Arbitrary Scanning Trajectories"},"pdf":{"value":"/pdf/3fb5f1d30bd010cb43cc79ab4687b1739ec71f32.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Short_Papers"},"paperhash":{"value":"blum|geometryembedded_neural_networks_cone_beam_ct_reconstruction_for_arbitrary_scanning_trajectories"},"authorids":{"value":["~Nele_Blum1","~Max_Stickel1","moritz.schaar@imte.fraunhofer.de","~Thorsten_Buzug1","~Maik_Stille1"]},"registration":{"value":"Yes"},"authors":{"value":["Nele Blum","Max Stickel","Moritz Schaar","Thorsten Buzug","Maik Stille"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"This paper introduces IntPhys 2, a new benchmark designed to evaluate intuitive physics understanding in AI systems using synthetic video data. Compared to prior work, IntPhys 2 features significantly more complex and realistic environments, leveraging Unreal Engine for photorealistic rendering, occlusions, and camera movements. The benchmark tests models on four core intuitive physics principles: Object Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. The dataset includes 1,416 videos, grouped into three splits: Debug - for robustness and noise sensitivity; Main - core evaluation set with three difficulty levels; Held Out - a challenging test set without metadata. \nThe authors evaluate several Multimodal Large Language Models (MLLMs) (e.g., GPT-4o, Gemini 2.5) and predictive models (e.g., V-JEPA, VideoMAEv2), using carefully controlled protocols inspired by the Violation of Expectation paradigm. The results reveal a large performance gap between humans (~96%) and AI models (mostly ~50–56%), emphasizing current models' limitations in physical reasoning."},"ethical_considerations":{"value":"No, there are no or only very minor ethics concerns"},"dataset_code_accessibility":{"value":"Yes"},"responsible_reviewing_acknowledgement":{"value":"Yes"},"code_of_conduct_acknowledgement":{"value":"Yes"},"confidence":{"value":3},"rating":{"value":4},"final_justification":{"value":"The authors have addressed most of my concerns. With the expanded discussion, this work offers additional insights for future research in this direction. My overall evaluation remains positive."},"limitations_weaknesses":{"value":"1. Despite the improved photorealism and complexity of scenes, the benchmark still operates in synthetic environments, limiting direct applicability to real-world physics understanding.\n2. This work only assess four physical principles. Other important concepts like inertia, gravity, and support relations seems omitted, which are tackled by some other benchmarks (e.g., GRASP as in Figure 2).\n3. The evaluations are mainly compared to IntPhys in Table 2. Providing results on other benchmarks of Figure 2 would help with the understanding of the novel challenges and gaps proposed by this work.\n4. The discussion and analysis of why current methods achieve poor performance on the proposed benchmark is a bit limited at this point. Is it due to memory? Or representation and perception? Is it possible to probing in the intermediate reasoning process to evaluate at which step the model fails to achieve intuitive physics understanding. The combination of object actions in the data does seem to require a \"chain of thought\" to understand or predict the content. Is it possible to obtain such reasoning process from human annotators? If yes, is it possible to evaluate the models on alignment with the reasoning process of real humans?\n5. Due to architectural limitations (e.g., short context, no memory), the evaluation of LLMs may depend more on prompt design than true reasoning ability, which also makes the discussion a bit insufficient.\n6. The paper still has some typos and issues in presentation:\n  - Line153, \"wether the video despite a plausible scenario\";\n  - Line 159, requires -> required;\n  - Line 304 , trough -> through;\n  - appendix G and appendix F seem incorrectly labeled;\n  - The caption of Figure 3 seems mismatched to the figures;\n  - Line 220, Table 2 and 3 here should be about ablation studies? \"The results of our evaluation are presented in Table 2 and 3, showcasing the optimal performance 221 outcomes for each model across the various factors we examined.\""},"dataset_code_comments":{"value":"The authors provide a full datasheet for the IntPhys2 benchmark in A.4-Table.5, and include the complete dataset and evaluation code in the supplementary material."},"strengths_contributions":{"value":"1. This paper studies an important and underexplored area in intuitive physics understanding and addressing key limitations of prior work in presenting limited cognitive challenges (e.g., IntPhys 1).\n2. The focus on psychological constructs like object permanence and the use of developmental-inspired methods (e.g., VOE) make the benchmark well-grounded in cognitive science.\n3. This paper adopts human evaluations to provide a strong baseline, demonstrating the validity of the proposed dataset and standards. In contrast, two lines of existing methods (i.e., MLLMs and predictive models) show near chance performance. \n4. The evaluation and ablation studies provide detailed information on the robustness of the existing methods under different circumstances.\n5. The authors provide a full datasheet for the IntPhys2 benchmark in A.4-Table.5, and include the complete dataset and evaluation code in the supplementary material, highlighting their effort in reproducilibity and transparency of their work."}},"parentInvitations":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Official_Review","nonreaders":[],"tmdate":1761794722898,"tcdate":1751532927720,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_9GPH"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_9GPH"],"forum":"Xpf5x3mLvn","number":4,"license":"CC BY 4.0","cdate":1751532927720,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Official_Review","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Official_Review4/-/Review_Revision"],"mdate":1761794722898,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","replyto":"Xpf5x3mLvn","id":"07jQGHXxI9","forumContent":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Video benchmark","Intuitive Physic","Evaluation"]},"supplementary_material":{"value":"/attachment/4932fdafe9cb6b17072f67b2caf246775bce6feb.zip"},"_bibtex":{"value":"@misc{\nbordes2026intphys,\ntitle={IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments},\nauthor={Florian Bordes and Quentin Garrido and Justine T Kao and Adina Williams and Michael Rabbat and Emmanuel Dupoux},\nyear={2026},\nurl={https://openreview.net/forum?id=Xpf5x3mLvn}\n}"},"title":{"value":"IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments"},"paperhash":{"value":"bordes|intphys_2_benchmarking_intuitive_physics_understanding_in_complex_synthetic_environments"},"dataset_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh.zip"},"TLDR":{"value":"We introduced a new intuitive physic benchmark IntPhys 2, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments"},"primary_area":{"value":"datasets_&_benchmarks_for_computer_vision"},"abstract":{"value":"We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to  macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50\\%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies."},"pdf":{"value":"/pdf/c29b03435c4329f0422348326d7ca2b5d606960a.pdf"},"croissant_file":{"value":"/attachment/54b155200f4140df0470921a91b0093300c98b6a.json"},"code_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh_code.zip"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"authorids":{"value":["~Florian_Bordes1","~Quentin_Garrido1","~Justine_T_Kao1","~Adina_Williams1","~Michael_Rabbat1","~Emmanuel_Dupoux1"]},"authors":{"value":["Florian Bordes","Quentin Garrido","Justine T Kao","Adina Williams","Michael Rabbat","Emmanuel Dupoux"]}},"version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_10.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"han|a_novel_measurement_of_structure_properties_in_complex_networks"},"authorids":{"value":["~Yanni_Han1","https://dblp.org/search/pid/api?q=author:Jun_Hu:","https://dblp.org/search/pid/api?q=author:Deyi_Li:","https://dblp.org/search/pid/api?q=author:Shuqing_Zhang:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_10"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/HanHLZ09,\n  author={Yanni Han and Jun Hu and Deyi Li and Shuqing Zhang},\n  title={A Novel Measurement of Structure Properties in Complex Networks},\n  year={2009},\n  cdate={1230768000000},\n  pages={1292-1297},\n  url={https://doi.org/10.1007/978-3-642-02469-6_10},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"Traditional measurements provide an effective tool to study the complex large systems in the real world. These global quantities only analyze the general statistical properties and interconnectivity structure of the entire network. However the complicated interactions among the locals are indeed the origin to emergent complex behavior. So in this paper we present a new measurement to reveal the local structure properties - topology potential, which reflects the differential position of each node in the topology. It is flexible by adjusting the influence factor. We demonstrate our measurement in US politics books network. Experiments confirm that topology potential has inherently implied the traditional measurements to some extent."},"title":{"value":"A Novel Measurement of Structure Properties in Complex Networks"},"authors":{"value":["Yanni Han","Jun Hu","Deyi Li","Shuqing Zhang"]}},"tmdate":1738980683474,"pdate":1230768000000,"tcdate":1738980657810,"writers":["~"],"signatures":["~Han_Yanni1"],"forum":"jpgDu44fLl","license":"CC BY-SA 4.0","number":312263,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1738980683474,"domain":"DBLP.org","id":"jpgDu44fLl","version":2},{"content":{"summary":{"value":"This paper addresses a critical evaluation gap in scientific machine learning, arguing that models for hyperspectral pansharpening that succeed on idealized, synthetic benchmarks often fail in real-world deployment. To solve this, the authors introduce PRISMABENCH, a new physics-aware evaluation ecosystem. This ecosystem features three main contributions: a large, diverse dataset of 10 real-world PRISMA satellite scenes; a suite of principled metrics, including a novel PAN-Conditioned Spatial Score ($D_{\\rho}^{PAN}$) for more robust, no-reference assessment; and insightful visualization tools, like multi-metric radar charts, to expose performance trade-offs that single-score leaderboards hide. Using this new framework, the paper demonstrates a significant disconnect between model rankings on traditional synthetic benchmarks and their actual performance on real-world data, providing a new blueprint to guide the field toward developing more robust and physically plausible models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. You correctly identify the out-of-band problem as a fundamental challenge. Why then does your primary metric contribution, $D_{\\rho}^{PAN}$, explicitly avoid evaluating this? Why dismiss the SWIR analysis to a qualitative visualization instead of proposing a quantitative metric for spatial fidelity in these, non-overlapping bands?\n2. In your related work, you argue that generative models (DDPMs) and self-supervised learning (SSL) are critical and promising paradigms for this task. Why are none of these modern SOTA approaches included in your experimental baselines?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper clearly diagnoses the problems of spectral mismatch, synthetic to real gap, and fragile no‑reference metrics, and motivates why HS‑PAN needs physics‑aware evaluation.\n2. The authors introduce a PRISMABENCH dataset. This new resource addresses key limitations of prior benchmarks by providing a larger and more diverse collection of 10 globally distributed scenes.\n3. The proposal of the PAN-Conditioned Spatial Score ($D_{\\rho}^{PAN}$). Instead of naively comparing all bands, this metric is physics-aware in a practical way, focusing the spatial quality assessment only on the spectral bands where the high-resolution PAN sensor actually provides reliable ground-truth spatial information."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The test set expands prior PRISMA FR corpora from 2–4 to 10 scenes, and tiles are larger (2400×2400). That is welcome but still small for a benchmark intended to “re‑calibrate progress.” The benchmark is single satellite (PRISMA only), so conclusions about robustness and real‑world utility may not transfer to other HSI platforms (e.g., differing PAN SRFs/PSFs, swath, radiometry).\n2. Because $D_{\\rho}^{PAN}$ focuses on PAN-overlapping bands, a method could over-inject PAN structure to improve $D_{\\rho}^{PAN}$ while degrading SWIR (non-overlapping) fidelity. The paper mitigates this with VIS-vs-SWIR visual analysis but lacks a quantitative SWIR-specific metric to complement $D_{\\rho}^{PAN}$.\n3. The paper correctly identifies complex physical challenges like sensor noise, PSF/SRF mismatch, and misregistration as key components of the synthetic-to-real gap. However, the proposed physics-aware solutions do not address these issues. The $D_{\\rho}^{PAN}$ metric is only \"physics-aware\" in the sense that it uses the known spectral range of the PAN sensor. This is a very limited application of physics that ignores the more challenging sensor properties the paper itself brought up."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926315785,"tcdate":1761955691050,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16149/Reviewer_nVdC"],"signatures":["ICLR.cc/2026/Conference/Submission16149/Reviewer_nVdC"],"forum":"rBdGw9PDiD","number":2,"license":"CC BY 4.0","cdate":1761955691050,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16149/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926315785,"domain":"ICLR.cc/2026/Conference","replyto":"rBdGw9PDiD","id":"L6G1oqlYGi","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Hyperspectral Pansharpening"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Progress in scientific machine learning is critically hindered by a pervasive \"evaluation gap\", where models that excel on legacy benchmarks fail in real-world deployment due to a reliance on idealized synthetic data and fragile proxy metrics. We argue that the path forward requires a new paradigm of physics-aware benchmarking, which we instantiate with $\\text{PRISMABench}$ for the challenging inverse problem of hyperspectral pansharpening. Our ecosystem introduces three core contributions: a $\\textbf{physics-enriched dataset}$ that packages real satellite PRISMA hyperspectral (HS) and panchromatic (PAN) pairs by their real physical sensors with 10 challenge scenes; an extended $\\textbf{PAN-centric evaluation metric}$, including a novel physics-consistency score for robust, no-reference assessment; and $\\textbf{insightful visualization tools}$, such as multi-metric radar charts, to move beyond single-score leaderboards and expose performance trade-offs. Using this framework, we reveal a critical disconnect: a model's rank on traditional reduced resolution benchmarks is a limited predictor of its real-world performance. By open-sourcing our ecosystem, we provide a blueprint for creating benchmarks that challenge the community to move beyond optimizing flawed proxies and towards developing models that are demonstrably robust and physically plausible."},"_bibtex":{"value":"@misc{\nluo2025recalibrating,\ntitle={Re-calibrating Progress: A Physics-Aware Benchmark to Expose the Evaluation Gap in Scientific Machine Learning},\nauthor={Chenxi Luo and Xin Gu and Wei Xiang},\nyear={2025},\nurl={https://openreview.net/forum?id=rBdGw9PDiD}\n}"},"title":{"value":"Re-calibrating Progress: A Physics-Aware Benchmark to Expose the Evaluation Gap in Scientific Machine Learning"},"pdf":{"value":"/pdf/4c533b66865818358557a984e1d9ceb1f1fa69fd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"luo|recalibrating_progress_a_physicsaware_benchmark_to_expose_the_evaluation_gap_in_scientific_machine_learning"},"authorids":{"value":["~Chenxi_Luo1","~Xin_Gu1","~Wei_Xiang3"]},"authors":{"value":["Chenxi Luo","Xin Gu","Wei Xiang"]}},"version":2},{"content":{"summary":{"value":"This paper presents NewtonGen, a framework aimed at solving the lack of physical consistency and controllability in text-to-video (T2V) generation. The authors propose a framework that integrates a data-driven T2V model with learnable physical principles. Its core is the Neural Newtonian Dynamics (NND) module, a physics-informed neural ODE that models a 9-dimensional latent physical state (e.g., position, velocity, rotation, size). The NND is first trained on \"physics-clean\" synthetic data to learn these dynamics. During inference, this trained NND predicts a sequence of physical states from user-defined initial conditions, which are then converted to optical flow to guide a motion-controlled T2V generator. Experiments across 12 motion types demonstrate quantitative improvements in physical consistency and precise parameter control. The reviewer regards NewtonGen as a novel and well-motivated two-stage framework for injecting physical realism and controllability into video generation. The core contribution, the NND module, is noted as being cleverly designed, as it combines linear ODEs with a residual MLP to learn complex dynamics from synthetic data efficiently. The reviewer acknowledges that the framework shows significant quantitative improvements on the Physical Invariance Score (PIS) metric and possesses excellent qualitative controllability. However, the reviewer raises concerns about the framework's reliance on synthetic \"physics-clean\" data, which poses significant questions about the sim-to-real gap. Furthermore, the reviewer points out that the method is explicitly limited to continuous dynamics, excluding collisions or discrete interactions, which narrows the scope of its \"Newtonian\" claims"},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The authors should explicitly state in the main paper that the framework's current application is generation from parameters, not prediction from real video. The \"continuous dynamics only\" limitation should also be candidly discussed in the main text, with potential extensions for handling discrete events like collisions.\n\n2. Add a \"Perfect Oracle\" Baseline: It is suggested to add a \"perfect simulator\" baseline to Table 1, where the ground-truth simulator optical flow is fed directly into the generator. This would quantify the performance gap between the learned NND and a \"perfect\" physics oracle.\n\n3. The authors should report standard deviations for the PIS scores in Table 1, rather than only the median, to provide a clearer measure of generation stability."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. The NND module is cleverly designed, combining physics-informed linear neural ODEs with a residual MLP. This allows it to flexibly learn a wide range of dynamical systems, from simple to complex, with high data efficiency.\n\n2. The framework achieves exceptional physical consistency across multiple motion types. In quantitative evaluations using the \"Physical Invariance Score\" (PIS) metric, NewtonGen's results significantly outperform all SOTA baselines and are very close to the simulated ground truth.\n\n3. The model achieves precise control over physical parameters, addressing a key weakness in existing T2V models. Experiments demonstrate it can generate motions that accurately correspond to user-specified initial physical states (e.g., position, velocity, size)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Sim-to-Real Gap: The NND is trained exclusively on synthetic \"physics-clean\" data. The paper does not demonstrate or discuss how the framework would handle real-world, noisy videos, suggesting its current application may be limited to generation from scratch rather than editing or predicting existing real-world videos.\n\n2. Limited Scope of Dynamics: The framework explicitly excludes discrete events such as collisions, rebounds, or multi-object interactions.\n\n3. Approximation of 3D Motion: The paper claims to handle 3D motion, but the method is \"equivalently realized through the combination of position and size control\". This appears to be a 2.5D approximation (i.e., objects scaling larger as they move closer) rather than a true 3D representation capable of handling complex perspective rotations.\n\n4. Misaligned Baseline Comparison: The paper compares NND against general-purpose T2V models like Sora and Veo3. These baselines were not optimized for fine-grained physical control, making the comparison somewhat misaligned."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915566495,"tcdate":1761915647651,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission615/Reviewer_r8si"],"signatures":["ICLR.cc/2026/Conference/Submission615/Reviewer_r8si"],"forum":"rJ6N6sunaU","number":3,"license":"CC BY 4.0","cdate":1761915647651,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission615/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915566495,"domain":"ICLR.cc/2026/Conference","replyto":"rJ6N6sunaU","id":"9vVulWQOT8","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"NewtonGen leverages Neural Newtonian Dynamics to learn general latent dynamics, enabling physically aware and controllable Text-to-Video generation."},"keywords":{"value":["Generative Models","Video Generation"]},"supplementary_material":{"value":"/attachment/9cc815abed32c4037e71149b52140fdc9597336b.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt changes in velocity and direction. Moreover, these models lack precise parameter control, struggling to generate physically consistent dynamics under different initial conditions. We argue that this fundamental limitation stems from current models learning motion distributions solely from appearance, while lacking an understanding of the underlying dynamics. In this work, we propose NewtonGen, a framework that integrates data-driven synthesis with learnable physical principles. At its core lies trainable Neural Newtonian Dynamics (NND), which can model and predict a variety of Newtonian motions, thereby injecting latent dynamical constraints into the video generation process. By jointly leveraging data priors and dynamical guidance, NewtonGen enables physically consistent video synthesis with precise parameter control.  All data and code are available at https://github.com/pandayuanyu/NewtonGen."},"_bibtex":{"value":"@inproceedings{\nyuan2026newtongen,\ntitle={NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics},\nauthor={Yu Yuan and Xijun Wang and Tharindu Wickremasinghe and Zeeshan Nadir and Bole Ma and Stanley H. Chan},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=rJ6N6sunaU}\n}"},"title":{"value":"NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics"},"pdf":{"value":"/pdf/77f56149d8717bd2f3e32f8d954800e0aebadb4a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|newtongen_physicsconsistent_and_controllable_texttovideo_generation_via_neural_newtonian_dynamics"},"authorids":{"value":["~Yu_Yuan4","~Xijun_Wang4","~Tharindu_Wickremasinghe1","~Zeeshan_Nadir1","~Bole_Ma2","~Stanley_H._Chan2"]},"authors":{"value":["Yu Yuan","Xijun Wang","Tharindu Wickremasinghe","Zeeshan Nadir","Bole Ma","Stanley H. Chan"]}},"version":2},{"content":{"summary":{"value":"The CONO is a sophisticated deep learning architecture designed to operate within the complex domain, aiming to capture and represent complex numerical signals effectively. Its architecture is underpinned by several key features:\n\n1. **Complex Domain Operation**: At its core, CONO processes data within the complex plane. By doing so, it taps into the rich information available in both the real and imaginary components of complex data. This enhances its expressive power and capacity for feature extraction.\n\n2. **Point-wise Operators and Transformations**: The model employs several operations and transformations, such as P, Q, and R, which project, convert, and lift data into various domains. Notably, it uses a Complex Convolutional Neural Network (CCNN) for certain transformations and integrates a complex UNET for additional processing.\n\n3. **Fractional Fourier Transforms**: A distinctive feature of the CONO is its utilization of the Discrete Fractional Fourier Transform (FrFT). This allows the model to learn and operate 'in between' the physical and frequency domains, offering a unique perspective and capturing various frequency contents and directional features present in data.\n\n4. **Continuous-Discrete Equivalence**: The model emphasizes maintaining a balance between continuous and discrete operations. This structure-preserving approach ensures that the model remains aligned with foundational principles like the Shannon-Whittaker-Kotel’nikov theorem. This ensures reliable analyses and predictions by the model.\n\n5. **End-to-End Architecture**: From input to output, the CONO model is structured to project, transform, process, and then revert data, ensuring the entire process is smooth and cohesive. The various layers, including complex UNET, CCNN, and point-wise operators, work together in harmony to achieve this.\n\nIn summary, the CONO is a robust and versatile model that capitalizes on the richness of complex domain operations, layered transformations, and a structure-preserving approach to effectively handle and analyze complex numerical signals."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. **Complex Domain Operations**: The ability of the CONO model to operate within the complex domain enables it to effectively capture the intricacies of complex numerical signals. This not only augments its expressive power but also enhances feature extraction, making it adept at representing and analyzing complex data.\n\n2. **Structure-preserving Architecture**: The CONO aims to maintain complex continuous-discrete equivalence, ensuring that the Shannon-Whittaker-Kotel’nikov theorem is obeyed for all continuous operations. This kind of structure preservation ensures that the model remains faithful to the underlying physics or principles, making predictions and analyses more reliable.\n\n3. **Comprehensive Framework**: The CONO encompasses a series of intricate operations, transformations, and layers, such as complex UNET, CCNN with a residual connection, and the use of fractional Fourier transforms. This comprehensive framework makes the model versatile and robust, allowing it to handle a wide variety of tasks and challenges, especially in the context of capturing complex numerical signals."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The articulation and expression of the manuscript require further refinement. In several sections, the clarity of the narrative falls short, making it challenging for the reader to grasp the content.\n\n2. The operations of CONO within the complex domain allow it to effectively capture and represent the nuances of complex signals. This enhances its expressive power and improves feature extraction. However, due to the intricacies of the involved operations and transformations, a significant amount of parameter tweaking and experimentation may be necessary to achieve optimal performance. It would be beneficial if the authors could provide detailed settings from their experiments, including memory usage.\n\n3. I'm particularly interested in the experiments related to the NS equation. To my knowledge, the original FNO paper mentioned three distinct viscosity coefficients. However, the authors seem to have chosen 10e-4 without clearly specifying it. This coefficient may not be the most challenging one. I would suggest the authors consider using the more challenging 10e-5 as the viscosity coefficient.\n\n4. I would encourage the authors to incorporate more visualizations to allow readers to gain a more detailed and intuitive understanding of the predicted outcomes.\n\n5. The selection of baselines for comparison appears to be incomplete. To provide a more comprehensive evaluation, I recommend the authors consider adding models like PINN[1] and LSM[2] to the comparisons.\n\n   [1] Raissi, M., Perdikaris, P. & Karniadakis, G.E. (2019). Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. *Journal of Computational physics*, *378*, pp.686-707.\n\n   [2] Wu, H., Hu, T., Luo, H., Wang, J. & Long, M. (2023). Solving High-Dimensional PDEs with Latent Spectral Models. *arXiv preprint arXiv:2301.12664*."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"see Weaknesses"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636648393,"tcdate":1698405984115,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6030/Reviewer_8MNi"],"signatures":["ICLR.cc/2024/Conference/Submission6030/Reviewer_8MNi"],"forum":"5vJe8XKFv0","number":1,"license":"CC BY 4.0","cdate":1698405984115,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6030/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636648393,"domain":"ICLR.cc/2024/Conference","replyto":"5vJe8XKFv0","id":"1K97PEmCeg","forumContent":{"TLDR":{"value":"We present a complex neural operator for learning the partial differential equations"},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Complex valued neural network","neural operator","partial differential equations","dynamical systems"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Neural operators extend data-driven models to map between infinite-dimensional functional spaces. These models have successfully solved continuous dynamical systems represented by differential equations, viz weather forecasting, fluid flow, or solid mechanics. However, the existing operators still rely on real space, thereby losing rich representations potentially captured in the complex space by functional transforms. In this paper, we introduce a Complex Neural Operator (CoNO), that parameterizes the integral kernel in the complex fractional Fourier domain. Additionally, the model employing a complex-valued neural network along with aliasing-free activation functions preserves the complex values and complex algebraic properties, thereby enabling improved representation, robustness to noise, and generalization. We show that the model effectively captures the underlying partial differential equation with a single complex fractional Fourier transform. We perform an extensive empirical evaluation of CoNO on several datasets and additional tasks such as zero-shot super-resolution, evaluation of out-of-distribution data, data efficiency, and robustness to noise. CoNO exhibits comparable or superior performance to all the state-of-the-art models in these tasks. Altogether, CoNO presents a robust and superior model for modeling continuous dynamical systems, providing a fillip to scientific machine learning. Our code implementation is available at https://anonymous.4open.science/r/anonymous-cono."},"_bibtex":{"value":"@misc{\ntiwari2024cono,\ntitle={Co{NO}: Complex Neural Operator for Continuous Dynamical Systems},\nauthor={Karn Tiwari and N M Anoop Krishnan and Prathosh AP},\nyear={2024},\nurl={https://openreview.net/forum?id=5vJe8XKFv0}\n}"},"title":{"value":"CoNO: Complex Neural Operator for Continuous Dynamical Systems"},"pdf":{"value":"/pdf/4b1e4f1003f798ce804f85a6de6f2db6ef731548.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"tiwari|cono_complex_neural_operator_for_continuous_dynamical_systems"},"authorids":{"value":["~Karn_Tiwari1","~N_M_Anoop_Krishnan1","~Prathosh_AP1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Karn Tiwari","N M Anoop Krishnan","Prathosh AP"]}},"version":2},{"content":{"comment":{"value":"Dear Reviewer Bi5p,\n\n\nThank you very much for your insightful review and constructive comments. We appreciate that you find our method easy to understand and implement, with clear motivation. We also appreciate that you value our analysis on disentanglement of visual appearance, as well as discussions regarding model size, number of frames, CFG, and data scale which provide useful insights. We now address your concerns below.\n\n\n\n\n**Q1: This approach needs to access model params, which makes it impossible for evaluating closed-source video diffusion models. This is also discussed in Section 5**\n\n\n**A1**:  Thank you for raising this point. Our method indeed requires access to the denoising network, which makes it most naturally applicable to open-source or internally accessible models. We argue that this requirement is less restrictive in the open-source community, where open-source video diffusion models play an important role and are widely used as research baselines.\nFor closed-source models, our method is still useful from the provider’s side: it can be integrated into the internal evaluation pipeline to monitor training progress, diagnose physics-related failure modes, and select checkpoints for release to users. \nIn general, we see our approach as a novel contribution.  To our knowledge, the first method that evaluates physics in diffusion models directly through the denoising process, without requiring video generation, and can therefore serve as a valuable tool for both open-source and proprietary model development.\n\n\n\n\n**Q2: All of the scenarios are from simulator, will the models pretrained more on simulated data have advantages over other models? I see some discussion in Section 5, the authors rely on an assumption that the models are mostly trained on real-world recordings rather than animated or synthetic content. But I doubt this may not hold.**\n\n\n**A2**: Our method follows the classic violation-of-expectation paradigm [1,2], which is mostly built on simulated data and has been successful in evaluating the physics understanding of various vision models. It is essentially infeasible to obtain test data from the real world that contains controlled violations of physical laws at scale, so simulation is a natural and widely adopted choice.\nRegarding whether a model trained more on synthetic data would have an advantage, our method is designed to provide an unbiased estimation of physics plausibility preference for the following reason: we do not measure the absolute likelihood of a single video, but rather the likelihood difference within a valid–invalid pair. As long as the two videos in a pair share the same visual appearance and differ only in physics adherence, the pairwise comparison effectively cancels out the influence of visual style and other appearance factors. Thus, the difference in estimated likelihood remains meaningful, can be attributed to physics violations, and yields a metric that is comparable across models, regardless of whether they were trained more on realistic or synthetic data. \n\n\n\n\nRegarding the assumption that models are mostly trained on real-world recordings, our method does not rely on this assumption. Our method aims to assess whether the learned video generative model distribution is close to a physics-plausible distribution by comparing the likelihoods the model assigns to physically valid and invalid video pairs. The training data distribution is one factor that may influence this physics capacity, but we do not make any prior assumptions about it. We have revised the discussion in Section 5 in the updated manuscript to remove potential ambiguity on this point.\n\n\n\n\n\n\n\n\n[1] Riochet, R., Castro, M.Y., Bernard, M., Lerer, A., Fergus, R., Izard, V. and Dupoux, E., 2018. Intphys: A framework and benchmark for visual intuitive physics reasoning. arXiv preprint arXiv:1803.07616.\n\n\n[2] Bordes, F., Garrido, Q., Kao, J.T., Williams, A., Rabbat, M. and Dupoux, E., 2025. IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments. arXiv preprint arXiv:2506.09849."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763657327589,"tcdate":1763657327589,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5897/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission5897/Authors"],"forum":"6UJf6B8RZ8","number":3,"license":"CC BY 4.0","cdate":1763657327589,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5897/-/Official_Comment"],"mdate":1763657327589,"domain":"ICLR.cc/2026/Conference","replyto":"HpsUnum2Bw","id":"xt4A2ziEPN","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generative Model","Video Diffusion Model","Intuitive Physics Understanding"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human\npreference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale up."},"_bibtex":{"value":"@inproceedings{\nyuan2026likephys,\ntitle={LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference},\nauthor={Jianhao Yuan and Fabio Pizzati and Francesco Pinto and Lars Kunze and Ivan Laptev and Paul Newman and Philip Torr and Daniele De Martini},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6UJf6B8RZ8}\n}"},"title":{"value":"LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference"},"pdf":{"value":"/pdf/be5f5c438a68c2b02423c0092164e93b5f3d245d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|likephys_evaluating_intuitive_physics_understanding_in_video_diffusion_models_via_likelihood_preference"},"authorids":{"value":["~Jianhao_Yuan2","~Fabio_Pizzati1","~Francesco_Pinto1","~Lars_Kunze1","~Ivan_Laptev1","~Paul_Newman1","~Philip_Torr1","~Daniele_De_Martini1"]},"authors":{"value":["Jianhao Yuan","Fabio Pizzati","Francesco Pinto","Lars Kunze","Ivan Laptev","Paul Newman","Philip Torr","Daniele De Martini"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LikePhys, a training-free, likelihood-preference-based evaluation method designed to assess the intuitive physics understanding of video diffusion models (VDMs). The core motivation is to address the challenge of disentangling physical plausibility from visual appearance in generated videos—an issue that existing evaluation methods often fail to handle due to biases from visual fidelity or subjective judgments. The key contribution of this work is the development of a new evaluation framework along with a comprehensive benchmark."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- The benchmarks were simulated using Blender. Which physics engine did you use, and what is the duration of each video?\n\n- In Figure 3, the textures appear relatively simple. Have you included more complex textured objects?\n\n- I feel that this approach is somewhat similar to Direct Preference Optimization (DPO). Could the author please elaborate on this point?\n\n- If the model not only trained on realistic videos, how to adapt your work to evaluate this model?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- This work is the first to utilize diffusion model likelihoods for evaluating intuitive physics. The approach of using the denoising objective as a likelihood surrogate for a \"violation-of-expectation\" test is both clever and well-justified, as it effectively examines the model's internal representation of physical concepts.\n\n- Experimental results indicate that the proposed Physics Perceptual Evaluation (PPE) aligns more closely with human preferences compared to existing automatic metrics.\n\n- The benchmark is comprehensive, providing a valuable framework for evaluating these models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The primary concern is that the method's validity depends on a curated set of synthetic simulations. While controlled violations are necessary, this raises questions about how well the findings generalize to the distribution of real-world, natural videos.\n\n- As stated in the paper, this work assumes the learned distribution is physics-plausible, which is actually infeasible for large video diffusion models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918333411,"tcdate":1762189983694,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5897/Reviewer_hTWK"],"signatures":["ICLR.cc/2026/Conference/Submission5897/Reviewer_hTWK"],"forum":"6UJf6B8RZ8","number":4,"license":"CC BY 4.0","cdate":1762189983694,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5897/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918333411,"domain":"ICLR.cc/2026/Conference","replyto":"6UJf6B8RZ8","id":"g95hP3rGHM","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generative Model","Video Diffusion Model","Intuitive Physics Understanding"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human\npreference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale up."},"_bibtex":{"value":"@inproceedings{\nyuan2026likephys,\ntitle={LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference},\nauthor={Jianhao Yuan and Fabio Pizzati and Francesco Pinto and Lars Kunze and Ivan Laptev and Paul Newman and Philip Torr and Daniele De Martini},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6UJf6B8RZ8}\n}"},"title":{"value":"LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference"},"pdf":{"value":"/pdf/be5f5c438a68c2b02423c0092164e93b5f3d245d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|likephys_evaluating_intuitive_physics_understanding_in_video_diffusion_models_via_likelihood_preference"},"authorids":{"value":["~Jianhao_Yuan2","~Fabio_Pizzati1","~Francesco_Pinto1","~Lars_Kunze1","~Ivan_Laptev1","~Paul_Newman1","~Philip_Torr1","~Daniele_De_Martini1"]},"authors":{"value":["Jianhao Yuan","Fabio Pizzati","Francesco Pinto","Lars Kunze","Ivan Laptev","Paul Newman","Philip Torr","Daniele De Martini"]}},"version":2},{"content":{"comment":{"value":"Dear Reviewer hTWK,\n\nThank you very much for your insightful reviews and constructive comments. We appreciate that you recognize our work as the first to leverage diffusion model likelihoods for evaluating intuitive physics, and that you find our approach both clever and well-justified, with a comprehensive benchmark that provides a valuable framework. We now address your concerns below.\n\n**Q1: The primary concern is that the method's validity depends on a curated set of synthetic simulations. While controlled violations are necessary, this raises questions about how well the findings generalize to the distribution of real-world, natural videos.**\n\n**A1**: Thank you for raising this point. Our approach follows the violation-of-expectation paradigm [1,2], which is mostly built on simulated data and has been successful in evaluating the physics understanding of various vision models. As you noted, it is essentially infeasible to obtain real-world videos with controlled physics violations. In our benchmark, we therefore use simulation to isolate the governing physics: within each preference pair, we deliberately match visual appearance (camera, lighting, textures, composition) so that valid and invalid videos differ only in physical plausibility. This design ensures that any likelihood difference within a pair can be attributed to violations of the underlying physics law rather than to confounding visual factors.\nRegarding generalization to real-world, natural videos, our human study directly targets this question. We generate text-to-video samples with various video diffusion models and ask human annotators to rate their physical plausibility. We then compare these human-based rankings to the rankings induced by PPE. We find that PPE serves as a good proxy for human preference on downstream text-to-video generation, indicating that the physics capacity measured on our synthetic benchmark meaningfully transfers to realistic use cases of video generative models in downstream generation.\n\n[1] Riochet, R., Castro, M.Y., Bernard, M., Lerer, A., Fergus, R., Izard, V. and Dupoux, E., 2018. Intphys: A framework and benchmark for visual intuitive physics reasoning. arXiv preprint arXiv:1803.07616.\n\n[2] Bordes, F., Garrido, Q., Kao, J.T., Williams, A., Rabbat, M. and Dupoux, E., 2025. IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments. arXiv preprint arXiv:2506.09849.\n\n**Q2: As stated in the paper, this work assumes the learned distribution is physics-plausible, which is actually infeasible for large video diffusion models.**\n\n**A2**: Our method does not rely on the assumption that the video diffusion models are trained on realistic videos. Measuring if the learned model distribution is aligned with a real-world physics-plausible distribution is the aim of our method, rather than its assumption. Our method measures the video diffusion model’s physics understanding defined from a distribution perspective, where we use the estimated likelihood difference over controlled invalid-valid video pairs as a proxy. The idea is that if a video diffusion model can assign a low likelihood to an invalid video (i.e. low probability on a physically plausible distribution) and vice versa, we consider the learned model distribution to be close to the physically plausible distribution, and thus have a good physics understanding. \n\n\nWe have revised Section 5 of the discussion to address the ambiguity and misunderstanding, better highlighting our motivations and method assumptions.\n\n\n\n**Q3: The benchmarks were simulated using Blender. Which physics engine did you use, and what is the duration of each video?**\n\n**A3**: All twelve scenarios are implemented in Blender 4.4.3 using its native physics engines and standard renderers. For rigid-body dynamics (Ball Drop, Ball Collision, Block Slide, Pyramid Impact), we use Blender’s Bullet rigid-body engine via the scene rigid body world; Ball Drop and Ball Collision are rendered with Cycles, while Block Slide and Pyramid Impact are rendered with Eevee Next. For deformable solids, Cloth Drape and Cloth Waving use Blender’s Cloth solver and are rendered with Cycles, and the gelatine/deformable drop scenario uses the Soft Body solver, also rendered with Cycles. For fluid scenarios (Faucet Flow, Droplet Fall, River Flow), we use Mantaflow in liquid mode as the physics engine and render all three with Cycles. For the Moving/Orbit Shadow scenarios, there is no rigid/cloth/fluid/soft-body solver active; shadows are produced purely by Cycles ray-traced lighting and animated geometry. For video duration, we standardize all clips to consist of 60 rendered frames in 24 fps."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763657248075,"tcdate":1763657248075,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5897/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission5897/Authors"],"forum":"6UJf6B8RZ8","number":1,"license":"CC BY 4.0","cdate":1763657248075,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5897/-/Official_Comment"],"mdate":1763657248075,"domain":"ICLR.cc/2026/Conference","replyto":"g95hP3rGHM","id":"MqtoPUh1jY","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generative Model","Video Diffusion Model","Intuitive Physics Understanding"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human\npreference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale up."},"_bibtex":{"value":"@inproceedings{\nyuan2026likephys,\ntitle={LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference},\nauthor={Jianhao Yuan and Fabio Pizzati and Francesco Pinto and Lars Kunze and Ivan Laptev and Paul Newman and Philip Torr and Daniele De Martini},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6UJf6B8RZ8}\n}"},"title":{"value":"LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference"},"pdf":{"value":"/pdf/be5f5c438a68c2b02423c0092164e93b5f3d245d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|likephys_evaluating_intuitive_physics_understanding_in_video_diffusion_models_via_likelihood_preference"},"authorids":{"value":["~Jianhao_Yuan2","~Fabio_Pizzati1","~Francesco_Pinto1","~Lars_Kunze1","~Ivan_Laptev1","~Paul_Newman1","~Philip_Torr1","~Daniele_De_Martini1"]},"authors":{"value":["Jianhao Yuan","Fabio Pizzati","Francesco Pinto","Lars Kunze","Ivan Laptev","Paul Newman","Philip Torr","Daniele De Martini"]}},"version":2},{"content":{"comment":{"value":"**Q3: Distribution discrepancy among the benchmark videos and tested models: from the demo frames in Fig.1 and appendix page 20~22, it seems the synthesized video from Blender are over simplified, with non-realistic looking and blank background. This distribution would be definitely far from the trainset of video generation models tested in the paper. Though results in Table 2 show that the method proposed is better than Qwen2.5 VL and VideoPhy1/2, I doubt whether this is fair for models with less training data similar to the synthesized video from the paper. From Table 1 in page 6, it seems the calculated PPE scores fluctuate dramatically among the tested 12 physics categories for any of the 12 models listed there, and I doubt the distribution discrepancy might be the hidden problem.**\n\n\n**A3**: First, regarding the simulator-generated data, we design each physics scenario to be relatively simple, with clearly attributable governing physics dynamics and consistent appearance. While the Blender videos are indeed stylized and far from the training distribution of current video generators, our method is constructed to be insensitive to such style differences. We do not use the absolute likelihood of a single video, but the likelihood difference between a valid–invalid pair that shares the same prompt and scene content. By comparing relative likelihood within a controlled pair, model-specific responses to visual style and other confounding factors cancel out, and the remaining difference is driven by the controlled physics violation. In this way, model bias toward particular visual styles should not dominate the PPE metric, and the metric remains comparable across models even when the benchmark distribution is simplified. This follows the classic violation-of-expectation paradigm for evaluating physics understanding in vision models, where there is also a distribution shift between the models’ training data and the simulator-generated test data [1,2,3]. Prior work has shown that this setting can still yield meaningful and robust conclusions, which motivates us to adopt the same practice.\n\n\nSecond, regarding the human study and comparison to other evaluators, we do not use the Blender synthetic data for human annotation. Instead, we use standard text-to-video generation settings to generate videos with various video generative models and then collect human ratings on these generated videos. We follow the standard paradigm [4] for VLM-based video physics plausibility assessment: given text prompts, models generate videos, and human annotators rate their physical plausibility. We then study how well PPE, computed on our synthetic benchmark, correlates with human judgments on these downstream generations, and we find that PPE aligns well with human preference. \nWe have revised Section 4 to better highlight the procedure for the human study, and Sections 1 (Introduction) and 3 (Methodology) to clarify these points and remove potential ambiguity.\n\n\n[1] Riochet, R., Castro, M.Y., Bernard, M., Lerer, A., Fergus, R., Izard, V. and Dupoux, E., 2018. Intphys: A framework and benchmark for visual intuitive physics reasoning. arXiv preprint arXiv:1803.07616.\n\n\n[2] Bordes, F., Garrido, Q., Kao, J.T., Williams, A., Rabbat, M. and Dupoux, E., 2025. IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments. arXiv preprint arXiv:2506.09849.\n\n\n[3] Garrido, Q., Ballas, N., Assran, M., Bardes, A., Najman, L., Rabbat, M., Dupoux, E. and LeCun, Y., 2025. Intuitive physics understanding emerges from self-supervised pretraining on natural videos. arXiv preprint arXiv:2502.11831.\n\n\n[4] Bansal, H., Peng, C., Bitton, Y., Goldenberg, R., Grover, A. and Chang, K.W., 2025. Videophy-2: A challenging action-centric physical commonsense evaluation in video generation. arXiv preprint arXiv:2503.06800.\n\n\n\n\n\n\n**Q4: Following question 2, authors didn't specify which one from the Qwen2.5 VL model family is used.**\n\n\n**A4**: We use Qwen2.5 VL 7B-Instruct, following the same practice as DreamGen [5]. We have revised Section 4 to specify this detail.\n\n\n[5] Jang, J., Ye, S., Lin, Z., Xiang, J., Bjorck, J., Fang, Y., Hu, F., Huang, S., Kundalia, K., Lin, Y.C. and Magne, L., 2025. DreamGen: Unlocking Generalization in Robot Learning through Video World Models. arXiv preprint arXiv:2505.12705."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763657416618,"tcdate":1763657416618,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5897/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission5897/Authors"],"forum":"6UJf6B8RZ8","number":6,"license":"CC BY 4.0","cdate":1763657416618,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5897/-/Official_Comment"],"mdate":1763657416618,"domain":"ICLR.cc/2026/Conference","replyto":"Q63lRXxpDk","id":"dJMBgI6KDu","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generative Model","Video Diffusion Model","Intuitive Physics Understanding"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human\npreference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale up."},"_bibtex":{"value":"@inproceedings{\nyuan2026likephys,\ntitle={LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference},\nauthor={Jianhao Yuan and Fabio Pizzati and Francesco Pinto and Lars Kunze and Ivan Laptev and Paul Newman and Philip Torr and Daniele De Martini},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6UJf6B8RZ8}\n}"},"title":{"value":"LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference"},"pdf":{"value":"/pdf/be5f5c438a68c2b02423c0092164e93b5f3d245d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|likephys_evaluating_intuitive_physics_understanding_in_video_diffusion_models_via_likelihood_preference"},"authorids":{"value":["~Jianhao_Yuan2","~Fabio_Pizzati1","~Francesco_Pinto1","~Lars_Kunze1","~Ivan_Laptev1","~Paul_Newman1","~Philip_Torr1","~Daniele_De_Martini1"]},"authors":{"value":["Jianhao Yuan","Fabio Pizzati","Francesco Pinto","Lars Kunze","Ivan Laptev","Paul Newman","Philip Torr","Daniele De Martini"]}},"version":2},{"content":{"summary":{"value":"This paper try to addresses the gap in MLLMs understanding of intuitive physics, particularly for continuum objects like fluids. The authors first introduce two low-level benchmark tasks, next frame selection and temporal coherence verification, to demonstrate that current MLLMs perform poorly at perceiving physical dynamics. To solve this, they propose SDF, an intermediate representation generated by physics simulators that visually encodes motion (e.g., velocity as color intensity). Through a multi-task fine-tuning strategy, their SDF-enhanced model achieves substantial gains on fluid tasks and shows strong generalization to unseen physical domains like cloth and sand."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See in the weekness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The motivation of the paper correctly focuses on a simpler, core problem, which is just perceiving physical motion, separating it from complex, high-level reasoning.\n\n- Using a visual map (SDF) from a physics simulator to train the model works. This helps the model learn the idea of dynamics, letting it generalize from fluids to unseen materials like cloth or sand.\n\n- The authors built a strong benchmark by mixing simulated data with real-world videos, and they ran thorough experiments to validate their approach.\n\n- The paper is well presented with clear logic."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The SDF representation is very basic, encoding just the projected velocity magnitude into one color channel. It's questionable if this simple map captures enough information for complex interactions, or if it needs to include more data like 3D vector direction or using optical flow (which avoids physics simulation). A discussion comparing the performance, advantages, and disadvantages of these different representations would greatly strengthen the analysis.\n\n- The pipeline is complex. It requires a full-parameter fine-tuning process, data from simulators, and knowledge distilled from expert models, which is computationally expensive and hard to scale."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919791310,"tcdate":1762127704670,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7741/Reviewer_tuv4"],"signatures":["ICLR.cc/2026/Conference/Submission7741/Reviewer_tuv4"],"forum":"Ax02eR2c3d","number":4,"license":"CC BY 4.0","cdate":1762127704670,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7741/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919791310,"domain":"ICLR.cc/2026/Conference","replyto":"Ax02eR2c3d","id":"El4N02cPEa","forumContent":{"TLDR":{"value":"This paper introduces two low-level tasks to test intuitive physics understanding and proposes Scene Dynamic Field, a method to integrate visual representation from physics simulators to MLLMs while showcasing generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multi-modal LLM","Intuitive Physics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. \nIn this work, we investigate the first step of physical reasoning, i.e., **intuitive physics understanding**, revealing substantial limitations in understanding the dynamics of continuum objects. \nTo isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. \nTo address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. \nSDF substantially improves performance, achieving up to $20.7\\%$ gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field."},"_bibtex":{"value":"@inproceedings{\nli2026beyond,\ntitle={Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models},\nauthor={Nanxi Li and Xiang Wang and Yuanjie Chen and Haode Zhang and Hong Li and Yong-Lu Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Ax02eR2c3d}\n}"},"title":{"value":"Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models"},"pdf":{"value":"/pdf/3b199b74bbdbddb8166b5e799fc00ddb71a413ae.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|beyond_static_vision_scene_dynamic_field_unlocks_intuitive_physics_understanding_in_multimodal_large_language_models"},"authorids":{"value":["~Nanxi_Li1","~Xiang_Wang19","~Yuanjie_Chen2","~Haode_Zhang2","~Hong_Li5","~Yong-Lu_Li1"]},"authors":{"value":["Nanxi Li","Xiang Wang","Yuanjie Chen","Haode Zhang","Hong Li","Yong-Lu Li"]}},"version":2},{"content":{"summary":{"value":"This paper explores whether supervised fine tuning works better than reinforcement learning fine tuning for vision language models in the context of intuitive physics tasks. They construct a train and test dataset using the ThreeDWorld simulator where structures are constructed with multiple blocks (i.e cubes), and an agent is required to either make a judgement of whether the structure is stable or how much a given block should be moved to make the structure stable. Vision language models are then trained on this task using supervised finetuning or reinforcement learning and tested to compare relative performance. \n\nThe authors find that there is not much difference in the performance of these two classes of models. They perform at ceiling when tested on the same kind of tasks seen in training and generalize equally poorly to new physical tasks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"My main concerns are about the dataset being too simple. How could this be extended or improved to ensure that the conclusions are reliable?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Training models to understand intuitive physics by interacting with the environment  is a well motivated hypothesis, as it is similar to how babies learn. \n2. The proposed metrics to test models seems reasonable, and the authors conduct rigorous evaluations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The dataset seems a bit simple and contrived. The fact that supervised learning performs as good as reinforcement learning might be because it’s a really easy task, and not because both methods fundamentally work equally well. \n2. I strongly disagree with the statement made in the conclusion that “these results cast doubt on whether posttraining methods are sufficient for developing models that reason about the world in a human-like manner”. The models not generalizing to new tasks, might just be because the training set is nowhere close to the amount of data that babies see, and not because the training algorithm is limited in some way. So we can’t really conclude anything about which model class is better from this result. \n3. It seems like the rewards are really handcrafted for this particular task. This is certainly not how humans would get rewards in the real world, so I’m curious to know what the authors think about how this method would scale to multiple tasks in different environments. Would rewards need to be defined for each task separately?\n4. It’s also not clear whether babies need a particular set of task specifications and goals for being able to learn intuitive physics. Most of intuitive physics might be learnt just by passive object manipulations without any defined goal like stability or placement. So how would we control for this kind of variable in the experiment? It makes me think that there is an inherent limitation in the way the training pipeline is set up here, which again makes me less confident about making any conclusions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943103542,"tcdate":1761940016749,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24501/Reviewer_rGrg"],"signatures":["ICLR.cc/2026/Conference/Submission24501/Reviewer_rGrg"],"forum":"XdLgOm5giq","number":3,"license":"CC BY 4.0","cdate":1761940016749,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24501/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943103542,"domain":"ICLR.cc/2026/Conference","replyto":"XdLgOm5giq","id":"7tuuGAguoe","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Vision language models","Intuitive physics","Interaction","Cognitive Science","Computational Cognitive Science","Human-like machine learning"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"abstract":{"value":"Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple physical tasks. However, fine-tuned models do not appear to learn robust physical rules that can generalize to new contexts. Based on research in cognitive science, we hypothesize that models need to interact with an environment to properly learn its physical dynamics. We train models that learn through interaction with the environment using reinforcement learning, as well as models that learn without interaction using supervised fine-tuning. While both reinforcement learning and supervised fine-tuning appear to improve within-task performance, they fail to produce models with generalizable physical intuitions. Models trained on one task do not reliably generalize to related tasks, even if they share visual statistics and physical principles, and regardless of whether they are trained through interaction."},"_bibtex":{"value":"@misc{\nbuschoff2026can,\ntitle={Can vision language models learn intuitive physics from interaction?},\nauthor={Luca M. Schulze Buschoff and Konstantinos Voudouris and Can Demircan and Eric Schulz},\nyear={2026},\nurl={https://openreview.net/forum?id=XdLgOm5giq}\n}"},"title":{"value":"Can vision language models learn intuitive physics from interaction?"},"pdf":{"value":"/pdf/1d06abe1f86256716c7d6a870bfac945f106010c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"buschoff|can_vision_language_models_learn_intuitive_physics_from_interaction"},"authorids":{"value":["~Luca_M._Schulze_Buschoff1","~Konstantinos_Voudouris1","~Can_Demircan1","~Eric_Schulz1"]},"authors":{"value":["Luca M. Schulze Buschoff","Konstantinos Voudouris","Can Demircan","Eric Schulz"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generative Model","Video Diffusion Model","Intuitive Physics Understanding"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in generation. To the end, we introduce LikePhys, a training-free method that evaluates intuitive physics in video diffusion models by distinguishing physically valid and impossible videos using the denoising objective as an ELBO-based likelihood surrogate on a curated dataset of valid-invalid pairs. By testing on our constructed benchmark of twelve scenarios spanning over four physics domains, we show that our evaluation metric, Plausibility Preference Error (PPE), demonstrates strong alignment with human\npreference, outperforming state-of-the-art evaluator baselines. We then systematically benchmark intuitive physics understanding in current video diffusion models. Our study further analyses how model design and inference settings affect intuitive physics understanding and highlights domain-specific capacity variations across physical laws. Empirical results show that, despite current models struggling with complex and chaotic dynamics, there is a clear trend of improvement in physics understanding as model capacity and inference settings scale up."},"_bibtex":{"value":"@inproceedings{\nyuan2026likephys,\ntitle={LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference},\nauthor={Jianhao Yuan and Fabio Pizzati and Francesco Pinto and Lars Kunze and Ivan Laptev and Paul Newman and Philip Torr and Daniele De Martini},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6UJf6B8RZ8}\n}"},"title":{"value":"LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference"},"pdf":{"value":"/pdf/be5f5c438a68c2b02423c0092164e93b5f3d245d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|likephys_evaluating_intuitive_physics_understanding_in_video_diffusion_models_via_likelihood_preference"},"authorids":{"value":["~Jianhao_Yuan2","~Fabio_Pizzati1","~Francesco_Pinto1","~Lars_Kunze1","~Ivan_Laptev1","~Paul_Newman1","~Philip_Torr1","~Daniele_De_Martini1"]},"authors":{"value":["Jianhao Yuan","Fabio Pizzati","Francesco Pinto","Lars Kunze","Ivan Laptev","Paul Newman","Philip Torr","Daniele De Martini"]}},"tmdate":1775876979649,"pdate":1769435834152,"tcdate":1757944177230,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5897/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission5897/Authors"],"forum":"6UJf6B8RZ8","license":"CC BY 4.0","number":5897,"cdate":1757944177230,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission5897/-/Full_Submission","ICLR.cc/2026/Conference/Submission5897/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission5897/-/Camera_Ready_Revision"],"mdate":1775876979649,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"6UJf6B8RZ8","version":2},{"content":{"summary":{"value":"This manuscript presents a theoretical analysis of physics informed machine learning with data are dependent. The work is an attempt to address the open challenge why incorporating physics prior can benefit data-drive learning."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Minor formatting issue: all equations should be numbered."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. This paper addresses the open challenge: the theoretical soundness of physics informed ML, beyond empirical evidence and intuitive understanding that the prior physics knowledge is conductive to learning. \n2. Theoretical analysis appears to be rigorous (but I didn't check all the proofs)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The organization and exposition of the manuscript can be improved. Without loss of theoretical rigor, it would be great if the intuitive explanations can be provided for deep learning practitioners the meaning of the theoretical results in practice. \n2. Related to the first question, intuitive understanding of the importance of conditions related to key variables such as $T$, as in Theorem 5.2 and $\\lambda_T$ as in theorem 5.1, would greatly benefit the reader.\n3. It's understandable this is a theoretical paper, but the only numerical experiment does add the weight to the paper. A well thought-off experiments to demonstrate the conditions related to $T$ and $\\lambda_T$ would also greatly benefit the readers, see the concern above."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922857823,"tcdate":1761967518241,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11838/Reviewer_7BGW"],"signatures":["ICLR.cc/2026/Conference/Submission11838/Reviewer_7BGW"],"forum":"IvLVPbeoRx","number":1,"license":"CC BY 4.0","cdate":1761967518241,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11838/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922857823,"domain":"ICLR.cc/2026/Conference","replyto":"IvLVPbeoRx","id":"pNqICDZ6FG","forumContent":{"TLDR":{"value":"We prove that adding correct prior domain knowledge to nonparametric learning with dependent data speeds up learning"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["learning with dependent data","physics-informed machine learning","convergence rates","complexity-dependent bounds"]},"supplementary_material":{"value":"/attachment/a00585da11ea98b0c0c19b8e1c5d1c3a2b58963e.pdf"},"primary_area":{"value":"learning theory"},"abstract":{"value":"A major challenge in physics-informed machine learning is to understand how the incorporation of prior domain knowledge affects learning rates when data are dependent. Focusing on empirical risk minimization with physics-informed regularization, we derive complexity-dependent bounds on the excess risk in probability and in expectation. We prove that, when the physical prior information is aligned, the learning rate improves from the (slow) Sobolev minimax rate to the (fast) optimal i.i.d. one without sample-size deflation due to data dependence."},"_bibtex":{"value":"@inproceedings{\nscampicchio2026physicsinformed,\ntitle={Physics-informed learning under mixing: How physical knowledge speeds up learning},\nauthor={Anna Scampicchio and Leonardo Felipe Toso and Rahel Rickenbach and James Anderson and Melanie Zeilinger},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=IvLVPbeoRx}\n}"},"title":{"value":"Physics-informed learning under mixing: How physical knowledge speeds up learning"},"pdf":{"value":"/pdf/9917146046b56820383565485dace6196f5b94c5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"scampicchio|physicsinformed_learning_under_mixing_how_physical_knowledge_speeds_up_learning"},"authorids":{"value":["~Anna_Scampicchio1","~Leonardo_Felipe_Toso1","~Rahel_Rickenbach1","~James_Anderson6","~Melanie_Zeilinger1"]},"authors":{"value":["Anna Scampicchio","Leonardo Felipe Toso","Rahel Rickenbach","James Anderson","Melanie Zeilinger"]}},"version":2},{"content":{"summary":{"value":"The authors propose a physics-informed method to generate synthetic samples of X ray scans of energy materials. The goal is to use this to train deep learning segmentation to be used on such images. The method builds on domain knowledge of the process, starting by generating typically observed shapes to then introduce noise and common artifacts. A U-Net trained on the synthetic data is compared with Otsu-thresholding and shows promising results."},"correctness":{"value":"3: Good"},"soundness":{"value":"3: Good"},"strengths":{"value":"The approach of generating synthetic data through domain knowledge is interesting, and different from common approaches of generating synthetic samples based on the available sample data. The need for such methods is motivated, and it seems to perform well on the test case."},"weaknesses":{"value":"There is a lack of discussion on wether similar methods have been deployed before, e.g. physics informed methods to generate synthetic data, or other methods to generate synthetic data in this problem. Without knowledge of the specific dicipline (X ray for energy materials), it is also hard to evaluate if the used domain knowledge is reasonable, or if there are important aspects missing. Both these issues could be resolved by referring to related works, and are understandable issues in an abstract, but makes my evaluation less certain."},"confidence":{"value":2},"rating":{"value":4}},"parentInvitations":"NLDL.org/2026/Abstracts_Track/-/Official_Review","nonreaders":[],"tmdate":1762345633808,"tcdate":1762174539267,"writers":["NLDL.org/2026/Abstracts_Track","NLDL.org/2026/Abstracts_Track/Submission25/Reviewer_XVwR"],"signatures":["NLDL.org/2026/Abstracts_Track/Submission25/Reviewer_XVwR"],"forum":"79VPSfpWZI","number":3,"license":"CC BY 4.0","cdate":1762174539267,"readers":["everyone"],"invitations":["NLDL.org/2026/Abstracts_Track/Submission25/-/Official_Review","NLDL.org/2026/Abstracts_Track/-/Edit"],"mdate":1762345633808,"domain":"NLDL.org/2026/Abstracts_Track","replyto":"79VPSfpWZI","id":"ip25P9ah5m","forumContent":{},"version":2},{"content":{"summary":{"value":"This paper introduces KineMask, a novel method for integrating physics-guided control and realistic object interaction into Video Diffusion Models (VDMs). It addresses the key limitation of current video generation models, which often fail to produce physically plausible motion and causal object interactions. A crucial innovation is the two-stage training strategy, consisting of Low-Level Control and High-Level Conditioning. KineMask is trained on simple synthetic scenes (boxes and cylinders), and successfully generalizes to complex interactions in real-world images. It achieves strong improvements over comparable state-of-the-art models in terms of synthesizing realistic motion and object interactions."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See Weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. It is interesting that the paper constructs a synthetic dataset of dynamic scenes with simple object (boxes/cylinders) interactions using Blender, and the model trained exclusively on this simple synthetic data can successfully generalize and synthesize complex interactions in real-world scenes.\n2. KineMask successfully integrates two distinct levels of control: precise low-level kinematic control (velocity mask) and semantic high-level textual guidance (predicted outcome description). This is also interesting.\n3. The core strength is enabling the VDM to infer and generate causal physical interactions (like collisions and liquid spilling) without relying on explicit frame-by-frame guidance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Training data is limited to basic rigid body interactions. Generalization to complex non-rigid materials (e.g., fire, smoke, cloth) or articulated bodies (e.g., machines) is unproven and likely limited.\n2. The use of 3-channel (RGB) velocity encoding in a mask may be a non-intuitive control format for end-users compared to simpler input methods (e.g., a single force vector)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918479111,"tcdate":1761890681502,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6120/Reviewer_VAx1"],"signatures":["ICLR.cc/2026/Conference/Submission6120/Reviewer_VAx1"],"forum":"vrY91av397","number":1,"license":"CC BY 4.0","cdate":1761890681502,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6120/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918479111,"domain":"ICLR.cc/2026/Conference","replyto":"vrY91av397","id":"HKdpMNfq3S","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Diffusion","Diffusion Models","Physics","Velocity"]},"supplementary_material":{"value":"/attachment/347ff9b4016f472a95f5a9c18c2b01857fa0c737.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent models for video generation have achieved remarkable progress and are now deployed in film, social media production, and advertising. Beyond their creative potential, such models also hold promise as world simulators for robotics and embodied decision making. Despite strong advances, however, current approaches still struggle to generate physically plausible object interactions and lack physics-grounded control mechanisms. To address this limitation, we introduce KineMask, an approach for physics-guided video generation that enables realistic rigid body control, interactions, and effects. Given a single image and a specified object velocity, our method generates videos with inferred motions and future object interactions. We propose a two-stage training strategy that gradually removes future motion supervision via object masks. Using this strategy we train video diffusion models (VDMs) on synthetic scenes of simple interactions and demonstrate significant improvements of object interactions in real scenes. Furthermore, KineMask integrates low-level motion control with high-level textual conditioning via predictive scene descriptions, leading to effective support for synthesis of complex dynamical phenomena. Extensive experiments show that KineMask achieves strong improvements over recent models of comparable size. Ablation studies further highlight the complementary roles of low- and high-level conditioning in VDMs. Our code, model, and data will be made publicly available"},"_bibtex":{"value":"@misc{\nromero2025learning,\ntitle={Learning to Generate Object Interactions with Physics-Guided Video Diffusion},\nauthor={David Romero and Ariana Bermudez and Hao Li and Fabio Pizzati and Ivan Laptev},\nyear={2025},\nurl={https://openreview.net/forum?id=vrY91av397}\n}"},"title":{"value":"Learning to Generate Object Interactions with Physics-Guided Video Diffusion"},"pdf":{"value":"/pdf/0c6e6df3bc5779ea4365518b7e6397eda28596f1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"romero|learning_to_generate_object_interactions_with_physicsguided_video_diffusion"},"authorids":{"value":["~David_Romero1","~Ariana_Bermudez1","~Hao_Li2","~Fabio_Pizzati1","~Ivan_Laptev1"]},"authors":{"value":["David Romero","Ariana Bermudez","Hao Li","Fabio Pizzati","Ivan Laptev"]}},"version":2},{"content":{"summary":{"value":"This paper presents IntPhys 2, a new video benchmark for evaluating the intuitive physics understanding of AI models. It uses complex, photorealistic synthetic videos to test core principles like object permanence and solidity, based on the \"violation of expectation\" framework. The authors evaluate several state-of-the-art models, including Multimodal Large Language Models (MLLMs) and predictive models. They find that while humans achieve near-perfect accuracy, most models perform poorly, often at chance levels (50%). This highlights a significant gap between current model capabilities and human-like physical reasoning, pointing to the need for better model architectures."},"ethical_considerations":{"value":"No, there are no or only very minor ethics concerns"},"dataset_code_accessibility":{"value":"Yes"},"responsible_reviewing_acknowledgement":{"value":"Yes"},"code_of_conduct_acknowledgement":{"value":"Yes"},"ethical_comments":{"value":"No ethics concern."},"confidence":{"value":3},"rating":{"value":4},"final_justification":{"value":"I am maintaining my score and assigning a borderline accept. The authors' rebuttal addresses mainly part of my concerns, which include:\n\n*   Incorporating comparisons with more recent benchmarks.\n*   Explaining the insufficient context during evaluation.\n*   Potential methods for improving the physics understanding capability.\n*   Additional results regarding the video generation model.\n\nOne suggestion is that this paper primarily focuses on rigid body motion, which I still believe could be improved by incorporating more types of dynamics, such as fluid motion, etc."},"limitations_weaknesses":{"value":"- The primary benchmark used for comparison, IntPhys, is quite dated (published in 2018). The authors should include more recent and relevant benchmarks in the main comparison tables.\n- After reviewing the submitted videos, it seems that rigid body motion accounts for the main part.  The authors may incorporate other motion types for evaluating physics understanding, such as deformable body dynamics, fluid motion, etc for robust evaluation.\n- The performance bottleneck may not lie in the physics understanding capability but in the insufficient context. This raises concerns about the effectiveness of IntPhys 2. For example, some videos require long-term memory to infer, but only 16 frames are provided as context for V-JEPA and VideoMAEv2, which may be why they lag behind on IntPhys 2 compared to their performance on IntPhys. Please justify this.\n- Any ideas to improve the performance of models on this chanllenging IntPhys 2?\n- In addition to predictive models, could the performance of world simulators (video generation models) like MAGI-1[1] or self-forcing[2] on IntPhys 2 be provided?\n\n[1] Teng, Hansi, et al. \"MAGI-1: Autoregressive Video Generation at Scale.\"*arXiv preprint arXiv:2505.13211*(2025).\n\n[2] Huang, Xun, et al. \"Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.\"*arXiv preprint arXiv:2506.08009*(2025)."},"dataset_code_comments":{"value":"The authors properly provide the data and code, providing guideline for using."},"strengths_contributions":{"value":"- Motivation is clear. This paper points out that the existing benchmarks focus on simple environments that lack variations and complexities and propose a more chanllenging benchmark\n- Unlike previous benchmarks that may have included a variety of scenarios, IntPhys 2 exclusively considers occlusions. This makes it more difficult and inquire stronger physics reasoning ability of models to infering.\n- IntPhys 2 advances beyond existing benchmarks by incorporating photo-realistic scenes with sophisticated visual elements, incorporating static and dynamic camera perspective.\n- Experiments are conducted to confirm the difficulty of IntPhys 2 for existing models, highlighting the need for advancements in model architectures and training methodologies."}},"parentInvitations":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Official_Review","nonreaders":[],"tmdate":1761794723195,"tcdate":1751339182038,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_PW3t"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_PW3t"],"forum":"Xpf5x3mLvn","number":1,"license":"CC BY 4.0","cdate":1751339182038,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Official_Review","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Official_Review1/-/Review_Revision"],"mdate":1761794723195,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","replyto":"Xpf5x3mLvn","id":"hiw7bPrz0U","forumContent":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Video benchmark","Intuitive Physic","Evaluation"]},"supplementary_material":{"value":"/attachment/4932fdafe9cb6b17072f67b2caf246775bce6feb.zip"},"_bibtex":{"value":"@misc{\nbordes2026intphys,\ntitle={IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments},\nauthor={Florian Bordes and Quentin Garrido and Justine T Kao and Adina Williams and Michael Rabbat and Emmanuel Dupoux},\nyear={2026},\nurl={https://openreview.net/forum?id=Xpf5x3mLvn}\n}"},"title":{"value":"IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments"},"paperhash":{"value":"bordes|intphys_2_benchmarking_intuitive_physics_understanding_in_complex_synthetic_environments"},"dataset_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh.zip"},"TLDR":{"value":"We introduced a new intuitive physic benchmark IntPhys 2, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments"},"primary_area":{"value":"datasets_&_benchmarks_for_computer_vision"},"abstract":{"value":"We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to  macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50\\%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies."},"pdf":{"value":"/pdf/c29b03435c4329f0422348326d7ca2b5d606960a.pdf"},"croissant_file":{"value":"/attachment/54b155200f4140df0470921a91b0093300c98b6a.json"},"code_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh_code.zip"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"authorids":{"value":["~Florian_Bordes1","~Quentin_Garrido1","~Justine_T_Kao1","~Adina_Williams1","~Michael_Rabbat1","~Emmanuel_Dupoux1"]},"authors":{"value":["Florian Bordes","Quentin Garrido","Justine T Kao","Adina Williams","Michael Rabbat","Emmanuel Dupoux"]}},"version":2},{"content":{"summary":{"value":"The paper presents LoopNav, a benchmark built in Minecraft to study spatial consistency in world models. It collects loop-style navigation videos so that models must reproduce the same scenes when revisiting locations. The benchmark measures how well models maintain spatial coherence over long sequences. The authors evaluate several existing world models using standard video metrics (FVD, LPIPS, SSIM) and show that current approaches still fail to keep consistent scene layouts over time."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- This paper investigates a timely and relevant research question given the popularity of video generation models.\n- The proposed dataset seems to be quite large scale and appears to be a good test-bed for developing video generation models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Whether the A→B exploration context includes a 360° view at B? If not, the excessive amount of new observations during B→A would be highly unpredictable. Including these in the evaluation metrics could lead to biased results.\n\n- Whether the A→B trajectory is guaranteed to be linear? If not, any intermediate point in the trajectory could be regarded as a point C, making A→B and A→B→C→A effectively equivalent?\n\n- Whether solving LoopNav implies a world model that generalizes to other domains (e.g., unseen situations, other simulation environments, or the real world)? Some discussion or qualitative assessment of generalization would strengthen the paper.\n\n- The benchmark cannot accommodate non-generative world models.\n\n- There is no comparison with existing world-model benchmarks. Suggested related works include:\n    - *3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Models*\n    - *World Consistency Score: A Unified Metric for Video Generation Quality*\n    - *VBench: Comprehensive Benchmark Suite for Video Generative Models*\n    - *WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning*\n    - *GRASP: A Novel Benchmark for Evaluating Language Grounding and Situated Physics Understanding in Multimodal Language Models*\n    - *IntPhys 2: Benchmarking Intuitive Physics Understanding in Complex Synthetic Environments*"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927142406,"tcdate":1761495154917,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17158/Reviewer_xpSX"],"signatures":["ICLR.cc/2026/Conference/Submission17158/Reviewer_xpSX"],"forum":"uDmxJ6133n","number":1,"license":"CC BY 4.0","cdate":1761495154917,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17158/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927142406,"domain":"ICLR.cc/2026/Conference","replyto":"uDmxJ6133n","id":"Y6nELL8xzG","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Dataset","Benchmark","WorldModel","Memory"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"The ability to simulate the world in a spatially consistent manner is a crucial requirements for effective world models. Such a model enables high-quality visual generation, and also ensures the reliability of world models for downstream tasks such as simulation and planning. Designing a memory module is a crucial component for addressing spatial consistency: such a model must not only retain long-horizon observational information, but also enables the construction of explicit or implicit internal spatial representations. However, there are no dataset designed to promote the development of memory modules by explicitly enforcing spatial consistency constraints. Furthermore, most existing benchmarks primarily emphasize visual coherence or generation quality, neglecting the requirement of long-range spatial consistency. To bridge this gap, we construct a dataset and corresponding benchmark by sampling 150 distinct locations within the open-world environment of Minecraft, collecting about 250 hours (20 million frames) of loop-based navigation videos with actions. Our dataset follows a curriculum design of sequence lengths, allowing models to learn spatial consistency on increasingly complex navigation trajectories. Furthermore, our data collection pipeline is easily extensible to new Minecraft environments and modules. Four representative world model baselines are evaluated on our benchmark. Dataset, benchmark, and code are open-sourced to support future research."},"_bibtex":{"value":"@misc{\nlian2026toward,\ntitle={{TOWARD} {MEMORY}-{AIDED} {WORLD} {MODELS}: {BENCHMARKING} {VIA} {SPATIAL} {CONSISTENCY}},\nauthor={Kewei Lian and Shaofei Cai and Yilun Du and Yitao Liang},\nyear={2026},\nurl={https://openreview.net/forum?id=uDmxJ6133n}\n}"},"title":{"value":"TOWARD MEMORY-AIDED WORLD MODELS: BENCHMARKING VIA SPATIAL CONSISTENCY"},"pdf":{"value":"/pdf/ca9ef0f2272608b395c9e0d9ace27b1300a3730c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"lian|toward_memoryaided_world_models_benchmarking_via_spatial_consistency"},"authorids":{"value":["~Kewei_Lian1","~Shaofei_Cai2","~Yilun_Du1","~Yitao_Liang1"]},"authors":{"value":["Kewei Lian","Shaofei Cai","Yilun Du","Yitao Liang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a benchmark dataset for evaluating the physics awareness of methods targeting dynamic scene understanding. \nThis dataset involves materials including liquid, gas, rheological substances, and textiles. It also captures multi-object interaction. \nFor the physics awareness assessment, evaluation is made against ground truth 3D trajectories and physics materials."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. In Figure 5, why is the number of trajectories different for different methods?\n\n2. For the calculation of TD, how is a predicted primitive trajectory matched to a GT trajectory?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The overall presentation is clear and easy to follow.\n\n2. The research problem this paper targets is insightful: how to evaluate physics realism beyond photorealism and extend it to more complex scenarios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The scale of the proposed dataset is limited (17 scenes only). With diverse materials and multi-object interaction, the benchmark is expected to be larger.\n\n2. Evaluation of physics awareness is not fully benchmarked in this paper. \n\na. A majority of experiments still focus on photorealism assessment. \n\nb. The proposed metrics only focus on 3D trajectories of primitives. Are matched trajectories equivalent to physics awareness? Will it be useful to consider future trajectory predictions as additional metrics?\n\nc. Experiments for physics parameter prediction are very limited. \n\n3. Only 4D reconstruction methods are tested. Relevant physics learning approaches are not evaluated."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924695386,"tcdate":1761840733042,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14241/Reviewer_HztB"],"signatures":["ICLR.cc/2026/Conference/Submission14241/Reviewer_HztB"],"forum":"kwhk8o3k5O","number":2,"license":"CC BY 4.0","cdate":1761840733042,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14241/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924695386,"domain":"ICLR.cc/2026/Conference","replyto":"kwhk8o3k5O","id":"E2boMg6R5O","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["4D Gaussian Splatting","Physics","Dynamic Novel View Synthesis"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We introduce Phys-Bench, a novel physics-aware benchmark for 3D dynamic scene understanding. This benchmark is designed to evaluate methods for reconstructing 4D scenes and understanding underlying physics from given videos, with a main\nfocus on the Dynamic Novel View Synthesis (DyNVS) task. \nWhile existing algorithms and benchmarks primarily focus on photorealistic reconstruction, they largely overlook physics understanding. This neglect is a critical limitation, as a true understanding of dynamic scenes requires models to reason about physical interactions, not just appearance. \nOur benchmark provides complex dynamic scenarios with rich multi-object interactions, featuring realistic collisions and force exchanges that are faithfully generated to strictly adhere to physical laws. \nFurthermore, it contains a diverse range of physical materials, such as liquid, gas, rheological substances, and textiles, which move beyond the rigid bodies prevalent in existing benchmarks. \nTo enable quantitative evaluation, we provide essential ground-truth information such as 3D particle trajectories and physics parameters and propose two novel metrics tailored to assessing physical realism. \nWe further evaluate existing Dynamic Novel View synthesis and physics parameter estimation method on our benchmark and reveal their overlooked limitations in physics understanding and multi-body dynamics handling. \nWe believe Phys-Bench will serve as a crucial foundation for advancing research in dynamic view synthesis, physics-based scene understanding, and the integration of deep learning with physical simulation, ultimately enabling more faithful reconstruction and interpretation of complex 3D dynamic scenes."},"_bibtex":{"value":"@misc{\nkim2025physbench,\ntitle={Phys-Bench: A Physics-aware Benchmark with Multi-Body Interactions for 3D Dynamic Scene Understanding},\nauthor={Mijeong Kim and Gunhee Kim and Jungyoon Choi and Wonjae Roh and Bohyung Han},\nyear={2025},\nurl={https://openreview.net/forum?id=kwhk8o3k5O}\n}"},"title":{"value":"Phys-Bench: A Physics-aware Benchmark with Multi-Body Interactions for 3D Dynamic Scene Understanding"},"pdf":{"value":"/pdf/7397fc3721e1087578236428bbabf75295e20450.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kim|physbench_a_physicsaware_benchmark_with_multibody_interactions_for_3d_dynamic_scene_understanding"},"authorids":{"value":["~Mijeong_Kim1","~Gunhee_Kim4","~Jungyoon_Choi1","~Wonjae_Roh2","~Bohyung_Han1"]},"authors":{"value":["Mijeong Kim","Gunhee Kim","Jungyoon Choi","Wonjae Roh","Bohyung Han"]}},"version":2},{"content":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Intuitive physics","inverse graphics","particle-based fluid simulation","neural rendering"]},"supplementary_material":{"value":"/attachment/216f05bad7be4ad1bba4435eb0aacfc1d0b8d098.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We introduce latent intuitive physics, a transfer learning framework for physics simulation that can infer hidden properties of fluids from a single 3D video and simulate the observed fluid in novel scenes. Our key insight is to use latent features drawn from a learnable prior distribution conditioned on the underlying particle states to capture the invisible and complex physical properties. To achieve this, we train a parametrized prior learner given visual observations to approximate the visual posterior of inverse graphics, and both the particle states and the visual posterior are obtained from a learned neural renderer. The converged prior learner is embedded in our probabilistic physics engine, allowing us to perform novel simulations on unseen geometries, boundaries, and dynamics without knowledge of the true physical parameters. We validate our model in three ways: (i) novel scene simulation with the learned visual-world physics, (ii) future prediction of the observed fluid dynamics, and (iii) supervised particle simulation. Our model demonstrates strong performance in all three tasks."},"_bibtex":{"value":"@inproceedings{\nzhu2024latent,\ntitle={Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D Video},\nauthor={Xiangming Zhu and Huayu Deng and Haochen Yuan and Yunbo Wang and Xiaokang Yang},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=WZu4gUGN13}\n}"},"title":{"value":"Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D Video"},"pdf":{"value":"/pdf/384fabb8d66fd0f055123d3418529ac196af71ff.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"zhu|latent_intuitive_physics_learning_to_transfer_hidden_physics_from_a_3d_video"},"authorids":{"value":["~Xiangming_Zhu2","~Huayu_Deng1","~Haochen_Yuan1","~Yunbo_Wang2","~Xiaokang_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiangming Zhu","Huayu Deng","Haochen Yuan","Yunbo Wang","Xiaokang Yang"]}},"tmdate":1713672545180,"pdate":1705410902788,"tcdate":1695344345971,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4317/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission4317/Authors"],"forum":"WZu4gUGN13","number":4317,"cdate":1695344345971,"mdate":1713672545180,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/-/Submission","ICLR.cc/2024/Conference/-/Post_Submission","ICLR.cc/2024/Conference/Submission4317/-/Revision","ICLR.cc/2024/Conference/Submission4317/-/Rebuttal_Revision","ICLR.cc/2024/Conference/-/Edit","ICLR.cc/2024/Conference/Submission4317/-/Camera_Ready_Revision"],"odate":1697213872796,"domain":"ICLR.cc/2024/Conference","id":"WZu4gUGN13","version":2},{"content":{"summary":{"value":"The authors tackle a problem on how to make data-driven generative models respect physics when the underlying governing equations are unknown. Most physics-informed ML frameworks (like PINNs) rely on having systems of equations whereas this one doesn’t. Their idea is to first train a Gaussian Process–based Port-Hamiltonian model on scarce observations to learn an energy-based representation of the dynamics. Then, they use that GP model to both generate synthetic, physically consistent training data, and inject physics-awareness into the diffusion model’s training loss. The result is a diffusion model that can generate physically valid spatiotemporal trajectories even in low-data regimes, with calibrated uncertainty via conformal prediction."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"My major questions rise from stage 2 data generation.\n\nIt seems you generate training samples for the diffusion model by drawing Hamiltonian gradients from the GP-dPHS posterior and integrating them to create trajectories. How do you ensure these samples represent physically plausible dynamics rather than artifacts of the GP prior? Since the GP posterior is conditioned on very few observations, random Fourier feature sampling might produce unrealistic dynamics. How sensitive is PHDME to the number of samples or to kernel hyperparameters? Could the diffusion model be learning the GP’s bias instead of the true system’s variability? In addition, once you use GP-based uncertainty in training, and conformal prediction for post-hoc calibration, how do these two uncertainty sources interact? Are they redundant, complementary, or potentially conflicting?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The reasoning linking Hamiltonian structure, energy conservation, and diffusion regularization is coherent. The experiments support the core claim that PHDME can generate physically consistent trajectories without explicit governing equations. The paper is technically sound but logically structured, in particular, the exposition of the Hamiltonian and GP-dPHS background is rigorous, and the two-stage training diagram helps readers grasp the workflow. The writing demonstrates mastery of both physics-based modeling and generative modeling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The experimental diversity is limited: a single synthetic system (the nonlinear string + qualitative check) doesn’t establish generality across different physical domains (e.g., fluid flow, elasticity, robotics, multiphase systems). The reliance on a simulator that’s already physics-based may inflate the gains of PHDME versus standard data-driven baselines. The GP-dPHS prior is treated, more or less, as a black box. I hope to see checks on that energy or momentum are actually conserved (if “physics consistency” is the core contribution, the authors may consider proving it quantitatively, via eg. energy error plots, invariants over time, etc.). While conformal prediction is implemented, the calibration methodology is somewhat surface-level (no ablation or coverage plots). Also, comparing only to plain DDPM and GP-dPHS isn't probably enough, the authors may consider adding more stronger benchmarks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941873182,"tcdate":1762804248800,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21657/Reviewer_bG2L"],"signatures":["ICLR.cc/2026/Conference/Submission21657/Reviewer_bG2L"],"forum":"kTEOG9a2W3","number":4,"license":"CC BY 4.0","cdate":1762804248800,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21657/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941873182,"domain":"ICLR.cc/2026/Conference","replyto":"kTEOG9a2W3","id":"gQsuVNEVu7","forumContent":{"TLDR":{"value":"PHDM: diffusion guided by GP-learned Port-Hamiltonian energy gradients from sparse data. No exact fuction needed; structure is conserved and uncertainty is calibrated for spatiotemporal prediction."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed Learning","Diffusion Model","Port-Hamiltonian system","Uncertainty Quantification","Gaussian process"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Diffusion models are expressive priors for generating and predicting data from high-dimensional dynamical systems. Yet, purely data-driven approaches often lack reliability and trustworthiness, motivating growing interest in physics-informed machine learning (PIML). Most existing PIML methods, however, assume access to exact governing equations during training—an assumption that fails when the dynamics are unknown or too complex to model accurately. To address this gap, we introduce PHDME (Port-Hamiltonian Diffusion Model), a physics-informed diffusion framework that learns system dynamics without requiring exact equations. Our approach first trains a Gaussian process distributed Port-Hamiltonian system (GP-dPHS) on limited observations to capture an energy-based representation of the dynamics. The GP-dPHS is then used to generate a physically consistent and diverse dataset for diffusion training. To enforce physics-consistency, we embed the GP-dPHS structure directly into the diffusion training objective through a loss that penalizes deviations from the learned Hamiltonian dynamics, weighted by the GP’s predictive uncertainty. After training, we employ conformal prediction to provide distribution-free uncertainty quantification of the generated trajectories. In this way, PHDME is designed for regimes with scarce data and unknown equations, enabling data-efficient, physically valid trajectory generation with calibrated uncertainty estimates."},"_bibtex":{"value":"@misc{\ntan2026phdme,\ntitle={{PHDME}: Physics-Informed Diffusion Models without Explicit Governing Equations},\nauthor={Kaiyuan Tan and Kendra Lee Givens and Peilun Li and Thomas Beckers},\nyear={2026},\nurl={https://openreview.net/forum?id=kTEOG9a2W3}\n}"},"title":{"value":"PHDME: Physics-Informed Diffusion Models without Explicit Governing Equations"},"pdf":{"value":"/pdf/caa63d9bc69a07229f5b0d48dcdca801e0cc83ab.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"tan|phdme_physicsinformed_diffusion_models_without_explicit_governing_equations"},"authorids":{"value":["~Kaiyuan_Tan1","~Kendra_Lee_Givens1","~Peilun_Li3","~Thomas_Beckers1"]},"authors":{"value":["Kaiyuan Tan","Kendra Lee Givens","Peilun Li","Thomas Beckers"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PHYCO, a new method for learning how objects behave physically from monocular videos, which are videos taken from a single camera view. The key ideas are to use Edge-Aware Depth Consensus Anchors to get reliable geometric information from limited data, and a Multi-Hypothesis Physics Verifier to incorporate well-known physical laws as guiding hypotheses. The approach effectively handles noisy and sparse supervision and works well even with complex materials and real-world scenes. Experiments on synthetic and real data show that PHYCO outperforms existing methods, producing realistic, physically consistent results while maintaining generalization to new scenarios."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How do the balance factors λm​ and λg​ influence the training stability and convergence?\n\n2. Could you conduct ablation experiments on the components you proposed to demonstrate the effectiveness of each part?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Uses Edge-Aware Depth Consensus Anchors to extract reliable geometric information from limited and noisy data, improving the accuracy of 3D shape and motion understanding.\n\n2. Incorporates classical physical laws as differentiable hypotheses through the Multi-Hypothesis Physics Verifier, ensuring learned models are physically consistent.\n\n3. Performs well even with monocular videos that have limited detail and are affected by noise.\n\n4. Outperforms some existing state-of-the-art methods on both synthetic and real-world datasets, producing better physical simulations and renderings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the method handles single-object dynamics well, its scalability to multi-object interactions or highly complex scenes remains underexplored.\n\n2. Although not explicitly discussed, the integration of multiple components such as the verifiers and anchors potentially increases training complexity and time, which could hinder practical adoption.\n\n3. The method assumes reasonably accurate geometric initializations; cases with severe geometric ambiguities might challenge the approach.\n\n4. The choice and diversity of the classical models used in the physics verifier may limit applicability to certain material classes or behaviors not represented by the hypotheses.\n\n5. The real-world experiments are limited in scale and variety (e.g., only a few objects like dragon, wolf, pudding, etc.). Although comparisons with methods like NeuMA and NCLaw are conducted, the scope of evaluation could be broader by including more recent or diverse approaches, or ablation studies that more directly isolate the contributions of individual components."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923186272,"tcdate":1762204260254,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12246/Reviewer_g9CR"],"signatures":["ICLR.cc/2026/Conference/Submission12246/Reviewer_g9CR"],"forum":"XyHbp7Y2T5","number":4,"license":"CC BY 4.0","cdate":1762204260254,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12246/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923186272,"domain":"ICLR.cc/2026/Conference","replyto":"XyHbp7Y2T5","id":"m74yiNKLJX","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Gaussian splatting","physics-informed learning","implicit constitutive laws"]},"supplementary_material":{"value":"/attachment/831d2a297e1a2ec1a3e64d35f8a5d86acb52a56b.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"We present **PhyCo**, a framework for learning implicit constitutive laws from \\textbf{monocular dynamic observations} of Gaussian splatting. Existing implicit methods often suffer from local minima under noisy supervision and lack physical interpretability, while explicit approaches rely on predefined constitutive equations, limiting generalizability. To address these issues, our framework, **PhyCo**, introduces two key innovations. First, **initializing from a static multi-view scan, we propose *Edge-Aware Depth Consensus Anchors* to establish robust geometric constraints from subsequent monocular dynamic observations**, circumventing unreliable pixel-level supervision. Second, a *Multi-Hypothesis Physics Verifier* integrates classical constitutive models as differentiable hypotheses, providing strong physical priors to regularize the optimization while preserving the flexibility of implicit modeling. This unified approach ensures physical plausibility without sacrificing generality. Extensive experiments on synthetic, real-to-sim, and real-world datasets demonstrate that **PhyCo** significantly outperforms existing methods, achieving state-of-the-art performance in learning accurate and generalizable physical dynamics from monocular videos."},"_bibtex":{"value":"@misc{\nliu2026phyco,\ntitle={PhyCo: Physics-Consistent Learning of Implicit Constitutive Laws via Monocular Observations of 3D Gaussians},\nauthor={Xiaoyang Liu and Kai Han},\nyear={2026},\nurl={https://openreview.net/forum?id=XyHbp7Y2T5}\n}"},"title":{"value":"PhyCo: Physics-Consistent Learning of Implicit Constitutive Laws via Monocular Observations of 3D Gaussians"},"pdf":{"value":"/pdf/230c2a74b4fdefee465eb282bf5252b5dd6e30e3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"liu|phyco_physicsconsistent_learning_of_implicit_constitutive_laws_via_monocular_observations_of_3d_gaussians"},"authorids":{"value":["~Xiaoyang_Liu5","~Kai_Han1"]},"authors":{"value":["Xiaoyang Liu","Kai Han"]}},"version":2},{"content":{"venue":{"value":"Crossref"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-20893-6_44.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"ehrhardt|unsupervised_intuitive_physics_from_visual_observations"},"html":{"value":"https://doi.org/10.1007/978-3-030-20893-6_44"},"_bibtex":{"value":"@incollection{Ehrhardt_2019,\n\tdoi = {10.1007/978-3-030-20893-6_44},\n\turl = {https://doi.org/10.1007%2F978-3-030-20893-6_44},\n\tyear = 2019,\n\tpublisher = {Springer International Publishing},\n\tpages = {700--716},\n\tauthor = {Sebastien Ehrhardt and Aron Monszpart and Niloy Mitra and Andrea Vedaldi},\n\ttitle = {Unsupervised Intuitive Physics from Visual Observations},\n\tbooktitle = {Computer Vision {\\textendash} {ACCV} 2018}\n}"},"abstract":{"value":"While learning models of intuitive physics is an active area of research, current approaches fall short of natural intelligences in one important regard: they require external supervision, such as explicit access to physical states, at training and sometimes even at test time. Some approaches sidestep these requirements by building models on top of handcrafted physical simulators. In both cases, however, methods cannot learn automatically new physical environments and their laws as humans do. In this work, we successfully demonstrate, for the first time, learning unsupervised predictors of physical states, such as the position of objects in an environment, directly from raw visual observations and without relying on simulators. We do so in two steps: (i) we learn to track dynamically-salient objects in videos using causality and equivariance, two non-generative unsupervised learning principles that do not require manual or external supervision. (ii) we demonstrate that the extracted positions are sufficient to successfully train visual motion predictors that can take the underlying environment into account. We validate our predictors on synthetic datasets; then, we introduce a new dataset, Roll4Real, consisting of real objects rolling on complex terrains (pool table, elliptical bowl, and random height-field). We show that it is possible to learn reliable object trajectory extrapolators from raw videos alone, without any external supervision and with no more prior knowledge than the choice of a convolutional neural network architecture."},"title":{"value":"Unsupervised Intuitive Physics from Visual Observations"},"authors":{"value":[{"fullname":"Sebastien Ehrhardt"},{"fullname":"Aron Monszpart"},{"fullname":"Niloy Mitra","username":"~Niloy_Mitra1"},{"fullname":"Andrea Vedaldi"}]}},"tmdate":1789091243377,"pdate":1546300800000,"externalIds":["doi:10.1007/978-3-030-20893-6_44"],"tcdate":1769350425921,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Niloy_Mitra1"],"forum":"YcYhbUXdWh","license":"CC BY-SA 4.0","number":38459,"cdate":1559033326286,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789091243377,"domain":"OpenReview.net/Public_Article","id":"YcYhbUXdWh","version":2},{"content":{"venue":{"value":"CVPR 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10203037/10203050/10204759.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2023"},"paperhash":{"value":"tripathi|3d_human_pose_estimation_via_intuitive_physics"},"authorids":{"value":["~Shashank_Tripathi1","~Lea_Müller1","~Chun-Hao_P._Huang1","https://dblp.org/search/pid/api?q=author:Omid_Taheri:","~Michael_J._Black1","~Dimitrios_Tzionas1"]},"html":{"value":"https://doi.org/10.1109/CVPR52729.2023.00457"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/TripathiMHTBT23,\n  author={Shashank Tripathi and Lea Müller and Chun-Hao P. Huang and Omid Taheri and Michael J. Black and Dimitrios Tzionas},\n  title={3D Human Pose Estimation via Intuitive Physics},\n  year={2023},\n  cdate={1672531200000},\n  pages={4713-4725},\n  url={https://doi.org/10.1109/CVPR52729.2023.00457},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2023}\n}\n"},"abstract":{"value":"Estimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical plausibility, but these are not differentiable, rely on unrealistic proxy bodies, and are difficult to integrate into existing optimization and learning frameworks. In contrast, we exploit novel intuitive-physics (IP) terms that can be inferred from a 3D SMPL body interacting with the scene. Inspired by biomechanics, we infer the pressure heatmap on the body, the Center of Pressure (CoP) from the heatmap, and the SMPL body's Center of Mass (CoM). With these, we develop IPMAN, to estimate a 3D body from a color image in a “stable” configuration by encouraging plausible floor contact and overlapping CoP and CoM. Our IP terms are intuitive, easy to implement, fast to compute, differentiable, and can be integrated into existing optimization and regression methods. We evaluate IPMAN on standard datasets and MoYo, a new dataset with synchronized multi-view images, ground-truth 3D bodies with complex poses, body-floor contact, CoM and pressure. IPMAN produces more plausible results than the state of the art, improving accuracy for static poses, while not hurting dynamic ones. Code and data are available for research at https://ipman.is.tue.mpg.de."},"title":{"value":"3D Human Pose Estimation via Intuitive Physics"},"authors":{"value":["Shashank Tripathi","Lea Müller","Chun-Hao P. Huang","Omid Taheri","Michael J. Black","Dimitrios Tzionas"]}},"tmdate":1772656928042,"pdate":1672531200000,"tcdate":1730698244889,"writers":["~"],"signatures":["~Shashank_Tripathi1"],"forum":"gRxpQ1dHS1","license":"CC BY-SA 4.0","number":169241,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1772656928042,"domain":"DBLP.org","id":"gRxpQ1dHS1","version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2303.18246v3"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"tripathi|3d_human_pose_estimation_via_intuitive_physics"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Shashank_Tripathi:","https://dblp.org/search/pid/api?q=author:Lea_Müller:","https://dblp.org/search/pid/api?q=author:Chun-Hao_P._Huang:","https://dblp.org/search/pid/api?q=author:Omid_Taheri:","~Michael_J._Black1","https://dblp.org/search/pid/api?q=author:Dimitrios_Tzionas:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2303.18246"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2303-18246,\n  publtype={informal},\n  author={Shashank Tripathi and Lea Müller and Chun-Hao P. Huang and Omid Taheri and Michael J. Black and Dimitrios Tzionas},\n  title={3D Human Pose Estimation via Intuitive Physics},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2303.18246},\n  url={https://doi.org/10.48550/arXiv.2303.18246}\n}\n"},"abstract":{"value":"Estimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical plausibility, but these are not differentiable, rely on unrealistic proxy bodies, and are difficult to integrate into existing optimization and learning frameworks. In contrast, we exploit novel intuitive-physics (IP) terms that can be inferred from a 3D SMPL body interacting with the scene. Inspired by biomechanics, we infer the pressure heatmap on the body, the Center of Pressure (CoP) from the heatmap, and the SMPL body's Center of Mass (CoM). With these, we develop IPMAN, to estimate a 3D body from a color image in a \"stable\" configuration by encouraging plausible floor contact and overlapping CoP and CoM. Our IP terms are intuitive, easy to implement, fast to compute, differentiable, and can be integrated into existing optimization and regression methods. We evaluate IPMAN on standard datasets and MoYo, a new dataset with synchronized multi-view images, ground-truth 3D bodies with complex poses, body-floor contact, CoM and pressure. IPMAN produces more plausible results than the state of the art, improving accuracy for static poses, while not hurting dynamic ones. Code and data are available for research at https://ipman.is.tue.mpg.de."},"title":{"value":"3D Human Pose Estimation via Intuitive Physics"},"authors":{"value":["Shashank Tripathi","Lea Müller","Chun-Hao P. Huang","Omid Taheri","Michael J. Black","Dimitrios Tzionas"]}},"tmdate":1731478058690,"pdate":1672531200000,"tcdate":1731476391906,"writers":["~"],"signatures":["~Michael_J_Black1"],"forum":"NclXKGVdEW","license":"CC BY-SA 4.0","number":202332,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731478058690,"domain":"DBLP.org","id":"NclXKGVdEW","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2512.06232v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"su|opinion_learning_intuitive_physics_may_require_more_than_visual_data"},"html":{"value":"https://doi.org/10.48550/arXiv.2512.06232"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2512-06232,\n  publtype={informal},\n  author={Ellen Su and Solim LeGris and Todd M. Gureckis and Mengye Ren},\n  title={Opinion: Learning Intuitive Physics May Require More than Visual Data},\n  year={2025},\n  month={December},\n  cdate={1764547200000},\n  journal={CoRR},\n  volume={abs/2512.06232},\n  url={https://doi.org/10.48550/arXiv.2512.06232}\n}\n"},"abstract":{"value":"Humans expertly navigate the world by building rich internal models founded on an intuitive understanding of physics. Meanwhile, despite training on vast quantities of internet video data, state-of-the-art deep learning models still fall short of human-level performance on intuitive physics benchmarks. This work investigates whether data distribution, rather than volume, is the key to learning these principles. We pretrain a Video Joint Embedding Predictive Architecture (V-JEPA) model on SAYCam, a developmentally realistic, egocentric video dataset partially capturing three children's everyday visual experiences. We find that training on this dataset, which represents 0.01% of the data volume used to train SOTA models, does not lead to significant performance improvements on the IntPhys2 benchmark. Our results suggest that merely training on a developmentally realistic dataset is insufficient for current architectures to learn representations that support intuitive physics. We conclude that varying visual data volume and distribution alone may not be sufficient for building systems with artificial intuitive physics."},"title":{"value":"Opinion: Learning Intuitive Physics May Require More than Visual Data"},"authors":{"value":[{"fullname":"Ellen Su","username":""},{"fullname":"Solim LeGris","username":""},{"fullname":"Todd M. Gureckis","username":""},{"fullname":"Mengye Ren","username":"~Mengye_Ren1"}]}},"tmdate":1786038326263,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2512-06232"],"tcdate":1786038319359,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Mengye_Ren1"],"forum":"fb1sTMCQaE","license":"CC BY-SA 4.0","number":124850,"cdate":1764547200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1786038326263,"domain":"OpenReview.net/Public_Article","id":"fb1sTMCQaE","version":2},{"content":{"TLDR":{"value":"Scaling up the size of visual datasets or using a developmentally realistic dataset is insufficient for current deep learning architectures to learn robust representations that support intuitive physics reasoning."},"venue":{"value":"NeurIPS 2025 Workshop EWM"},"pdf":{"value":"/pdf/daa402dcea0229b404616a40da69d63fc5f7e789.pdf"},"keywords":{"value":["world models","intuitive physics","cognitive science"]},"venueid":{"value":"NeurIPS.cc/2025/Workshop/EWM"},"paperhash":{"value":"su|opinion_learning_intuitive_physics_may_require_more_than_visual_data"},"authorids":{"value":["~Ellen_Su1","~Solim_LeGris1","~Todd_M._Gureckis1","~Mengye_Ren1"]},"abstract":{"value":"Humans expertly navigate the world by building rich internal models founded on an intuitive understanding of physics. Meanwhile, despite training on vast quantities of internet video data, state-of-the-art deep learning models still fall short of human-level performance on intuitive physics benchmarks. This work investigates whether data distribution, rather than volume, is the key to learning these principles. We pretrain a Video Joint Embedding Predictive Architecture (V-JEPA) model on SAYCam, a developmentally realistic, egocentric video dataset partially capturing three children's everyday visual experiences. We find that training on this dataset, which represents 0.01% of the data volume used to train SOTA models, does not lead to significant performance improvements on the IntPhys2 benchmark. Our results suggest that merely training on a developmentally realistic dataset is insufficient for current architectures to learn representations that support intuitive physics. We conclude that varying visual data volume and distribution alone may not be sufficient for building systems with artificial intuitive physics."},"_bibtex":{"value":"@inproceedings{\nsu2025opinion,\ntitle={Opinion: Learning Intuitive Physics Requires More Than Visual Data},\nauthor={Ellen Su and Solim LeGris and Todd M. Gureckis and Mengye Ren},\nbooktitle={NeurIPS 2025 Workshop on Embodied World Models for Decision Making},\nyear={2025},\nurl={https://openreview.net/forum?id=z9WKQF2kJD}\n}"},"title":{"value":"Opinion: Learning Intuitive Physics May Require More Than Visual Data"},"authors":{"value":["Ellen Su","Solim LeGris","Todd M. Gureckis","Mengye Ren"]}},"tmdate":1761430006969,"pdate":1758263921125,"tcdate":1756757437902,"writers":["NeurIPS.cc/2025/Workshop/EWM","NeurIPS.cc/2025/Workshop/EWM/Submission42/Authors"],"signatures":["NeurIPS.cc/2025/Workshop/EWM/Submission42/Authors"],"forum":"z9WKQF2kJD","license":"CC BY 4.0","number":42,"cdate":1756757437902,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Workshop/EWM/-/Submission","NeurIPS.cc/2025/Workshop/EWM/-/Post_Submission","NeurIPS.cc/2025/Workshop/EWM/-/Edit","NeurIPS.cc/2025/Workshop/EWM/Submission42/-/Camera-Ready_Version"],"mdate":1761430006969,"odate":1758263921125,"domain":"NeurIPS.cc/2025/Workshop/EWM","id":"z9WKQF2kJD","version":2},{"content":{"summary":{"value":"This paper investigates whether VLMs can learn generalizable intuitive physics through interaction, as opposed to passive SFT. The authors compare two methods—GRPO (an interactive RL method) and SFT (non-interactive)—on tasks involving block tower stability and construction. They test three hypotheses: whether interactive training improves (1) within-task generalization, (2) cross-task generalization, and (3) sample efficiency on new tasks. The results show that while both methods achieve near-ceiling performance on trained tasks, neither leads to robust generalization to related tasks, suggesting that current fine-tuning methods encourage shortcut learning rather than genuine physical reasoning."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weaknesses.\nI am willing to chat with the authors to improve the paper."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. The paper is one of the first to systematically compare interactive (RL) and non-interactive (SFT) training for intuitive physics in VLMs.\n2. The writing is clear, and the experimental setup is well-explained."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Experimental Scope: The study is confined to a single model family (Qwen2.5-VL) and a narrow set of block-stacking tasks. Broader evaluation across diverse model architectures and more varied physical reasoning tasks would strengthen the conclusions.\n\n2. Limited Innovation: The paper only compares SFT and RL on a small task without further analysis of the underlying reasons or proposing potential solutions—or at least offering improvements specific to the physics subtask. Additionally, it does not explore whether longer training, alternative RL algorithms, or more diverse interaction strategies could enhance generalization.\n\n3. Limited Community Impact: The paper’s point about not overestimating the generalization ability brought by SFT (L459) appears to be a widely recognized observation. Moreover, the study focuses only on the narrow subfield of physical perception and is validated on a small-scale dataset, which limits the significance of its findings.\n\n4. Presentation Issues:\n(a) Appendix A.5 Attention Maps: It is unclear what the authors intend to illustrate with these visualizations.\n(b) Related Work: The related work section is overly verbose and fails to clearly highlight the paper’s significant contributions compared to prior research."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943104098,"tcdate":1760502974562,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24501/Reviewer_aHQU"],"signatures":["ICLR.cc/2026/Conference/Submission24501/Reviewer_aHQU"],"forum":"XdLgOm5giq","number":1,"license":"CC BY 4.0","cdate":1760502974562,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24501/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943104098,"domain":"ICLR.cc/2026/Conference","replyto":"XdLgOm5giq","id":"iBzuCM3Gzh","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Vision language models","Intuitive physics","Interaction","Cognitive Science","Computational Cognitive Science","Human-like machine learning"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"abstract":{"value":"Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple physical tasks. However, fine-tuned models do not appear to learn robust physical rules that can generalize to new contexts. Based on research in cognitive science, we hypothesize that models need to interact with an environment to properly learn its physical dynamics. We train models that learn through interaction with the environment using reinforcement learning, as well as models that learn without interaction using supervised fine-tuning. While both reinforcement learning and supervised fine-tuning appear to improve within-task performance, they fail to produce models with generalizable physical intuitions. Models trained on one task do not reliably generalize to related tasks, even if they share visual statistics and physical principles, and regardless of whether they are trained through interaction."},"_bibtex":{"value":"@misc{\nbuschoff2026can,\ntitle={Can vision language models learn intuitive physics from interaction?},\nauthor={Luca M. Schulze Buschoff and Konstantinos Voudouris and Can Demircan and Eric Schulz},\nyear={2026},\nurl={https://openreview.net/forum?id=XdLgOm5giq}\n}"},"title":{"value":"Can vision language models learn intuitive physics from interaction?"},"pdf":{"value":"/pdf/1d06abe1f86256716c7d6a870bfac945f106010c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"buschoff|can_vision_language_models_learn_intuitive_physics_from_interaction"},"authorids":{"value":["~Luca_M._Schulze_Buschoff1","~Konstantinos_Voudouris1","~Can_Demircan1","~Eric_Schulz1"]},"authors":{"value":["Luca M. Schulze Buschoff","Konstantinos Voudouris","Can Demircan","Eric Schulz"]}},"version":2},{"content":{"summary":{"value":"The paper presents a physics-based simulation framework for generating synthetic X-ray CT datasets of energy materials to enable deep learning-based segmentation without relying on scarce annotated real-world data. The authors simulate realistic polychromatic X-ray projections using the ASTRA Toolbox and SpekPy, modeling the tungsten anode X-ray source and beam-hardening effects. The synthetic data, derived from procedurally generated microstructures, are used to train a U-Net segmentation model. The model achieves high segmentation performance (Dice score: 0.9943, IoU: 0.9886) on test data, outperforming classical methods such as Otsu thresholding."},"correctness":{"value":"3: Good"},"soundness":{"value":"2: Fair"},"strengths":{"value":"- The work addresses an important bottleneck in materials imaging: the lack of annotated training data for deep learning segmentation.\n\n- The combination of physics-based simulation (ASTRA Toolbox + SpekPy) and data-driven segmentation (U-Net) is well-motivated. \n\n- The reported Dice and IoU scores indicate that even with simple synthetic geometries, the U-Net can effectively learn to segment features and outperform classical methods.\n\n- The authors explicitly plan to scale the framework toward more complex, realistic microstructures and real CT data, which suggests meaningful long-term impact and research potential."},"weaknesses":{"value":"- The current simulation is based on simple non-overlapping circular shapes, which poorly represent the complex, irregular, and multi-phase microstructures of actual energy materials. This limits generalization to real data.\n\n- While the paper mentions planned experiments on real µCT data, no results are currently presented. The absence of domain transfer analysis makes it difficult to assess the real-world applicability of the proposed approach.\n\n- The use of a vanilla U-Net with BCE loss, without exploring modern segmentation architectures or domain adaptation strategies, limits the methodological depth. Techniques like attention U-Nets, transformers, or physics-informed regularization could strengthen the contribution.\n\n- The evaluation compares only against Otsu thresholding, which is a weak baseline. Additional baselines (e.g., classical machine learning, morphological segmentation, or learned models trained on limited real data) would better contextualize performance.\n\n- Details on the computational cost of the simulation and training pipeline, as well as robustness to noise, spectrum variation, or parameter changes, are not discussed but are essential for practical adoption."},"confidence":{"value":5},"rating":{"value":4}},"parentInvitations":"NLDL.org/2026/Abstracts_Track/-/Official_Review","nonreaders":[],"tmdate":1762345633939,"tcdate":1762155712688,"writers":["NLDL.org/2026/Abstracts_Track","NLDL.org/2026/Abstracts_Track/Submission25/Reviewer_a8gs"],"signatures":["NLDL.org/2026/Abstracts_Track/Submission25/Reviewer_a8gs"],"forum":"79VPSfpWZI","number":2,"license":"CC BY 4.0","cdate":1762155712688,"readers":["everyone"],"invitations":["NLDL.org/2026/Abstracts_Track/Submission25/-/Official_Review","NLDL.org/2026/Abstracts_Track/-/Edit"],"mdate":1762345633939,"domain":"NLDL.org/2026/Abstracts_Track","replyto":"79VPSfpWZI","id":"EWxQHGGEPu","forumContent":{},"version":2},{"content":{"summary":{"value":"This paper makes two main contributions: (1) introducing DynSuperCLEVR, a novel video question answering dataset that focuses on understanding 4D dynamics (velocity, acceleration, collisions) of objects in 3D scenes, and (2) proposing NS-4DPhysics, a neural-symbolic model that integrates physics priors with 3D scene understanding for dynamics reasoning. Through extensive experiments, their model significantly outperforms existing approaches, including large multimodal models, demonstrating current limitations in physical reasoning capabilities of video-language models."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- It would be great if the author analyze the performance of proposed model on more complex dynamic scenarios such as non-rigid object deformation or fluid dynamics\n- How do the authors plan to address the challenges in real datasets such as motion blur, camera shake, and varying lighting conditions?\n- How do architectural choices  or important hyperparameters such as different CNN backbones or the choice of physics engine parameters impact on the performance of proposed model?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper's main objective of addressing multimodal 4D dynamics understanding is well-motivated.\n- The authors provide comprehensive evaluation results across three types of reasoning tasks (factual, predictive, and counterfactual), demonstrating the model's capabilities in different scenarios.\n- The proposed physics-aware neural-symbolic architecture presents an innovative approach"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The dataset only considers rigid objects with linear velocity and acceleration. Real-world scenarios often involve more complex dynamics like non-rigid deformation, rotation-based motion, and fluid dynamics.\n- The dataset uses synthetic rendering which may not capture real-world challenges like motion blur, camera shake, varying lighting conditions, and partial occlusions.\n- The ablation studies are limited. While the paper shows the importance of physics priors, there could be more detailed analysis of other architectural choices and hyperparameters, like the impact of different CNN backbones or the choice of physics engine parameters."}},"nonreaders":[],"tmdate":1731427265260,"tcdate":1730705688915,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission479/Reviewer_oBU1"],"signatures":["ICLR.cc/2025/Conference/Submission479/Reviewer_oBU1"],"forum":"6Vx28LSR7f","number":3,"license":"CC BY 4.0","cdate":1730705688915,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission479/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427265260,"domain":"ICLR.cc/2025/Conference","replyto":"6Vx28LSR7f","id":"Cn2Us8Nfb8","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"We introduce DynSuperCLEVR, a video question answering dataset focused on the dynamic properties of 3D objects. We propose NS-4DPhysics, which use 4D world states with a 3D generative model and uses neural symbolic reasoning to answer questions."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video question answering","Compositional reasoning","Physical scene understanding","3D scene understanding"]},"supplementary_material":{"value":"/attachment/9df5057549ac56979fea8269802d34b5d654e102.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-level temporal and action semantics. Although humans are adept at understanding these properties by constructing 3D and temporal (4D) representations of the world, current video understanding models struggle to extract these dynamic semantics, arguably because these models use cross-frame reasoning without underlying knowledge of the 3D/4D scenes.\nIn this work, we introduce **DynSuperCLEVR**, the first video question answering dataset that focuses on language understanding of the dynamic properties of 3D objects. We concentrate on three physical concepts—*velocity*, *acceleration*, and *collisions*—within 4D scenes. We further generate three types of questions, including factual queries, future predictions, and counterfactual reasoning that involve different aspects of reasoning on these 4D dynamic properties.\nTo further demonstrate the importance of explicit scene representations in answering these 4D dynamics questions, we propose **NS-4DPhysics**, a **N**eural-**S**ymbolic VideoQA model integrating **Physics** prior for **4D** dynamic properties with explicit scene representation of videos. \nInstead of answering the questions directly from the video text input, our method first estimates the 4D world states with a 3D generative model powered by a physical prior, and then uses neural symbolic reasoning to answer the questions based on the 4D world states.\nOur evaluation on all three types of questions in DynSuperCLEVR shows that previous video question answering models and large multimodal models struggle with questions about 4D dynamics, while our NS-4DPhysics significantly outperforms previous state-of-the-art models."},"_bibtex":{"value":"@inproceedings{\nwang2025compositional,\ntitle={Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering},\nauthor={Xingrui Wang and Wufei Ma and Angtian Wang and Shuo Chen and Adam Kortylewski and Alan Yuille},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=6Vx28LSR7f}\n}"},"title":{"value":"Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering"},"pdf":{"value":"/pdf/9ebb092c1ee905ec7df8632cc804269f09324d4a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"wang|compositional_4d_dynamic_scenes_understanding_with_physics_priors_for_video_question_answering"},"authorids":{"value":["~Xingrui_Wang1","~Wufei_Ma1","~Angtian_Wang2","~Shuo_Chen13","~Adam_Kortylewski1","~Alan_Yuille1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xingrui Wang","Wufei Ma","Angtian Wang","Shuo Chen","Adam Kortylewski","Alan Yuille"]}},"version":2},{"content":{"summary":{"value":"The paper introduces HGTFT, a graph transformer architecture for time series forecasting in multi-physics domains. This approach improves upon drawbacks of existing methods in multi-physics settings, which often struggle to perform under the complexity of disparate spatiotemporal dynamics. The proposed architecture attempts to overcome this challenge by piecing together relevant neural modules (e.g., the temporal, graph, and subtask layers) capable of jointly capturing complex dynamics. This method is compared to many competing baseline models on several common time series benchmarks, as well as on a newly proposed Multiphysics Building System (MBS) dataset."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- In Table 2, there are a few counter-intuitive fluctuations in performance across zero-shot and few-shot settings. For instance, the RCS metric is better for HGTFT zero-shot than it is in the few-shot (in both the \"50 MBS\" and \"Full MBS\" settings). I understand minor fluctuations could very well be noise, but is there a more principled reason behind this?\n- In the main multi-physics evaluation setup, building samples from the MBS dataset are used for pre-training. Is there a significant amount of overlap between MBS and BTS? What dynamics does MBS capture that are expected to be helpful for BTS, or perhaps more importantly, what are the meaningful differences (potentially highlighting ability to generalize to new dynamics seen only in BTS)?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper's stated contributions are clear and address a difficult, high-impact problem in the domain of multi-physics systems. The approach potentially lays the groundwork for tackling broader challenges across connected physics models (not just the building environment).\n- The presentation of the paper is clear and well-organized. I appreciated the comprehensive literature review and logically grouped discussions across Sections 1-4, which made problem setup and methodology easy to compartmentalize and digest.\n- The reported evaluation is very extensive, covering a variety of important dimensions that help position the model's utility. For instance, the model is compared to several baseline methods on common time series datasets (highlighting its comparative advantages), key ablations are reported (justifying architectural decisions), and different model sizes are evaluated, highlighting the impact of parameter scaling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The empirical evaluation of the method on multi-physics settings is somewhat limited, provided only the multi-physics building setting is explored. While results appear strong and there are diverse dynamics present, it is difficult to assess the proposed architecture's general utility as a multi-physics model beyond this domain.\n- The analysis of empirical results would benefit from a discussion that characterizes when a complex architecture like HGTFT is justified, compared to a less complex model, such as LSTMs. It would be very insightful to see practical tradeoff considerations, for instance, highlighting model differences in performance at fixed parameter counts or time spent training.\n- It is claimed that the more common time series benchmark datasets don't capture the multi-domain complexity that HGTFT targets, but there is little to no explanation behind why the architecture underperforms other methods on these benchmarks (e.g., in Table 13). Presumably many of the mechanisms relevant for capturing complex multi-physics interactions would be beneficial in modeling complex multi-variate time series more broadly. Additionally, it is not particularly clear why the other datasets, e.g., traffic, are not considered to exhibit multi-scale dynamics among diverse entities (stated in Section 6.3) when these are common qualities of traffic forecasting settings. Characterizing the performance differences across these settings in more depth would go some way to helping bridge the empirical gap and help characterize model behavior in lieu of an additional multi-physics dataset."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926586193,"tcdate":1762000828764,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16485/Reviewer_uK4E"],"signatures":["ICLR.cc/2026/Conference/Submission16485/Reviewer_uK4E"],"forum":"esM0zdV3NO","number":3,"license":"CC BY 4.0","cdate":1762000828764,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16485/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926586193,"domain":"ICLR.cc/2026/Conference","replyto":"esM0zdV3NO","id":"G5bs1541sw","forumContent":{"TLDR":{"value":"HGTFT, a pre-train and fine-tune framework that integrates heterogeneous spatiotemporal data with physics-informed constraints for accurate, physically consistent time series forecasting in Multi-Domain Physical Systems."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Heterogeneous Graph","Time Series Forecasting","Multiphysics","Physical Systems","Pre-training","Transformer"]},"supplementary_material":{"value":"/attachment/ca3774d7ae4892bac79634cf64764e0b3fc1e6dc.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Existing Transformer-based models effectively capture multivariate dependencies, while pre-trained large models achieve strong generalization but are often confined to single-object or single-physics settings. Spatial-temporal approaches leverage graph structures but fall short in modeling heterogeneous entities with diverse inter-variable interactions, and they often lack mechanisms to enforce physical consistency.\nTo address these challenges, we propose the Heterogeneous Graph Temporal Fusion Transformer (HGTFT), a pre-training and fine-tuning framework tailored for spatially and temporally structured physical environments. HGTFT tokenizes observation points and generates embeddings that capture both temporal patterns and spatial correlations, enabling the integration of heterogeneous static and dynamic information. \nWe further introduce optimized normalization and physics-informed loss functions that enhance predictive accuracy while improving physical plausibility. Applied to temperature, flow, and energy-related datasets in building environments, our approach demonstrates strong zero-shot generalization and achieves substantial accuracy gains through few-shot fine-tuning with domain-specific data."},"_bibtex":{"value":"@misc{\nsun2026heterogeneous,\ntitle={Heterogeneous Graph Temporal Fusion Transformer for Time Series Forecasting in Multi-Domain Physical Systems},\nauthor={Yifu Sun and Xin Li and Qi Shen and Qiang Dou and Yingqiu Qiu and Yuankai Zhao and Qianchuan Zhao},\nyear={2026},\nurl={https://openreview.net/forum?id=esM0zdV3NO}\n}"},"title":{"value":"Heterogeneous Graph Temporal Fusion Transformer for Time Series Forecasting in Multi-Domain Physical Systems"},"pdf":{"value":"/pdf/a883883d2e6fc778a14ab7211a7dda9de53e617e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sun|heterogeneous_graph_temporal_fusion_transformer_for_time_series_forecasting_in_multidomain_physical_systems"},"authorids":{"value":["~Yifu_Sun3","~Xin_Li94","~Qi_Shen4","~Qiang_Dou1","~Yingqiu_Qiu2","~Yuankai_Zhao1","~Qianchuan_Zhao1"]},"authors":{"value":["Yifu Sun","Xin Li","Qi Shen","Qiang Dou","Yingqiu Qiu","Yuankai Zhao","Qianchuan Zhao"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a new framework, the Simulation-Heuristics Model (SHM), which conceptualizes intuitive physics as a dual process: Intuitive Physics Engine (IPE) dominates in short-term simulations, while a heuristic-based approach takes over when the IPE’s simulation extends beyond a certain time boundary."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"refer to weakness."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper introduces a new framework, the Simulation-Heuristics Model (SHM), which conceptualizes intuitive physics as a dual process: Intuitive Physics Engine (IPE) dominates in short-term simulations, while a heuristic-based approach takes over when the IPE’s simulation extends beyond a certain time boundary."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-I don't think this work is suitable for submission to ICLR, as it lacks AI/ML elements, learning representation and primarily consists of human experiments. I would recommend the author consider submitting it to CogSci or another more relevant conference.\n\n-Too simple task scenarios, it would more convincing to see how this SHM can helped with other downstream real-world tasks?\n\n-How well can existing VLM do in the proposed tasks?"}},"nonreaders":[],"tmdate":1731427290866,"tcdate":1731085840738,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission634/Reviewer_u9Qr"],"signatures":["ICLR.cc/2025/Conference/Submission634/Reviewer_u9Qr"],"forum":"BkeJro1xps","number":4,"license":"CC BY 4.0","cdate":1731085840738,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission634/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427290866,"domain":"ICLR.cc/2025/Conference","replyto":"BkeJro1xps","id":"IU3X4i19zn","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Intuitive physics","physical reasoning","mental simulation","heuristic model"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The role of mental simulation in human behavior for various physical tasks is widely acknowledged, attributed to the generality of Intuitive Physics Engine (IPE). However, it remains unclear whether mental simulation is consistently employed across scenarios of different simulation costs and where its boundary is. Moreover, cognitive strategies beyond these boundaries have not been thoroughly investigated. Here, we adopted a pouring-marble task containing various conditions to study IPE's limits and strategies beyond. A human study revealed two distinct error patterns in predicting the pouring angle, differentiated by the simulation time using a boundary. This suggests a possible switching of the underlying reasoning strategies. Our initial experiment on IPE showed that its correlation with human judgments diminished in scenarios requiring extended time of simulation. This observation prompted the exploration of an alternative mechanism based on heuristics for intuitive physics. We uncovered that a linear heuristic model, relying exclusively on empirical data, replicated human prediction more accurately when the simulation time exceeded a certain boundary. Motivated by these observations, we propose a new framework, Simulation-Heuristics Model (SHM), which conceptualizes intuitive physics as a dual process: IPE is predominant only in short-time simulation, whereas a heuristics-based approach is applied as IPE's simulation time extends beyond the simulation boundary. The SHM model aligns more precisely with human behavior across various scenarios and demonstrates superior generalization capabilities under different conditions. Crucially, SHM integrates computational methods previously viewed as separate into a unified model, quantitatively studying their switching mechanism."},"_bibtex":{"value":"@misc{\nli2024a,\ntitle={A simulation-heuristics dual-process model for intuitive physics},\nauthor={Shiqian Li and Yuxi Ma and Bo Dai and Yujia Peng and Chi Zhang and Yixin Zhu},\nyear={2024},\nurl={https://openreview.net/forum?id=BkeJro1xps}\n}"},"title":{"value":"A simulation-heuristics dual-process model for intuitive physics"},"pdf":{"value":"/pdf/93856d41c3aa9841ddf51a3fb5488577d1272d6a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|a_simulationheuristics_dualprocess_model_for_intuitive_physics"},"authorids":{"value":["~Shiqian_Li1","~Yuxi_Ma2","~Bo_Dai5","~Yujia_Peng1","~Chi_Zhang12","~Yixin_Zhu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shiqian Li","Yuxi Ma","Bo Dai","Yujia Peng","Chi Zhang","Yixin Zhu"]}},"version":2},{"content":{"summary":{"value":"This paper revisits Conditional Whitney Forms (CWFs)—a recent framework combining finite-element exterior calculus (FEEC) with neural operator learning—to analyze why prior formulations only trivially satisfy conservation laws without recovering the governing physics. The authors theoretically prove that the original equality-constrained optimization admits infinitely many flux solutions, leading to “formal” but physically meaningless conservation. They then propose a flux-regularized reformulation that introduces a data-driven physics-recovery term and evaluate it on several synthetic advection–diffusion and Poisson systems. Empirically, the regularized model recovers realistic flux fields while preserving structure, confirming the theoretical diagnosis."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Can the proposed formulation be adapted to settings where flux is unobserved—e.g., via PDE residuals, energy-based terms, or self-consistent latent-flux estimation?\n\n2. How sensitive is training to the flux penalty and the choice of optimizer (AdamW vs. Shampoo)?\n\n3. Could the flux regularization be interpreted as enforcing an energy-minimization principle or weak form of the PDE?\n\n4. How does the approach scale to 3D problems or complex geometries, where flux computation becomes quadratic in mesh size?\n\n5. Beyond advection–diffusion, does the method extend to nonlinear or coupled PDE systems (e.g., Navier–Stokes)?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper rigorously shows why existing CWFs reduce to unconstrained regression, explaining the phenomenon of trivial conservation through rank-deficiency analysis (Propositions 1–2). This is not really surprising, but it is good for the authors to point out explicitly.\n\n2. Adds an interpretable flux-regularization term that bridges finite-element constraints and learnable operator spaces, linking structure preservation and physics recovery. But this comes with limitations, listed in the \"Weaknesses\" section.\n\n3. Across 1D and 2D PDEs, the regularized method yields orders-of-magnitude lower flux error while retaining distribution accuracy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Dependence on ground-truth flux (major limitation!)  \nThe proposed flux-reconstruction loss requires known $f_{true}$. In real scientific-ML settings, fluxes are latent. We typically only observe scalar fields. Thus, the approach is only feasible in synthetic or diagnostic experiments, not in realistic PDE inference or discovery tasks.\n\n2. Limited methodological novelty  \nThe regularization essentially adds a supervised flux-matching term, a straightforward and expected extension of existing physics-informed or operator-learning frameworks. The conceptual leap beyond prior PINN/PINO-style constraints is small.\n\n3. Expected improvement  \nSince the loss directly penalizes flux error, the performance gain is tautological rather than emergent. The results confirm that additional supervision helps, but do not reveal new modeling behavior.\n\n4. Narrow experimental scope  \nAll tests involve simple 1D–2D advection–diffusion or Poisson problems with synthetic data and fixed meshes. The method’s stability and benefit for nonlinear, convection-dominated, or real-world PDEs remain untested.\n\n5. Limited practical deployability  \nThe computational cost roughly doubles due to flux evaluation, and no clear path is given for extending the approach to unobserved-flux or high-dimensional cases. Without an unsupervised or self-consistent flux-recovery variant, the method’s real impact is theoretical rather than applied."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941718584,"tcdate":1761418045220,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21357/Reviewer_bHgX"],"signatures":["ICLR.cc/2026/Conference/Submission21357/Reviewer_bHgX"],"forum":"UVgNkQTScp","number":1,"license":"CC BY 4.0","cdate":1761418045220,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21357/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941718584,"domain":"ICLR.cc/2026/Conference","replyto":"UVgNkQTScp","id":"fF4PDaqqPJ","forumContent":{"TLDR":{"value":"We show that existing conditional Whitney form formulations lead to trivial structure preservation, propose incorporating additive structure to achieve meaningful physics recovery and validate our approach experimentally."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["ai4science","physics-informed machine learning","finite elements","operator learning","interpretable scientific discovery"]},"supplementary_material":{"value":"/attachment/f3a488b6854318552fb614324c5cfe32e93e6f30.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Conditional Whitney forms have recently emerged as a promising framework at the intersection of scientific machine learning and finite element analysis. They offer a solid theoretical foundation for enforcing conservation laws in complex machine learning settings. However, their use so far has been restricted to learning tasks where structural constraints can be satisfied with simple, yet inaccurate, physics representations. In this work, we analyze why existing formulations reduce to typical unconstrained reformulations, circumventing physics recovery, and highlight the necessity of incorporating additive structure pertaining to the governing physics of the system. Based on the theoretical insights we first attain, we proceed to the reformulation of the learning problem to enable data-driven physics recovery and employ conditional Whitney forms to turn a Transformer-based architecture into a structure-preserving reduced-order model. We demonstrate the validity of our theoretical insights and the effectiveness of the subsequent proposed reformulation in a range of advection-diffusion systems of increasing difficulty. Our contributions can be viewed as a step towards understanding the capacity of conditional Whitney forms to build reliable structure-preserving models by harnessing the modeling power of state-of-the-art machine learning architectures in physical sciences."},"_bibtex":{"value":"@misc{\nkallinikidis2025revisiting,\ntitle={Revisiting Conditional Whitney Forms: From Structure Preservation to Physics Recovery},\nauthor={Pavlos Kallinikidis and Paris Perdikaris and George J. Pappas},\nyear={2025},\nurl={https://openreview.net/forum?id=UVgNkQTScp}\n}"},"title":{"value":"Revisiting Conditional Whitney Forms: From Structure Preservation to Physics Recovery"},"pdf":{"value":"/pdf/c00d44e75b31a888585ba4b5329f7effe91b703f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kallinikidis|revisiting_conditional_whitney_forms_from_structure_preservation_to_physics_recovery"},"authorids":{"value":["~Pavlos_Kallinikidis1","~Paris_Perdikaris1","~George_J._Pappas1"]},"authors":{"value":["Pavlos Kallinikidis","Paris Perdikaris","George J. Pappas"]}},"version":2},{"content":{"summary":{"value":"This paper presents HGSolver to tackle PDEs with heterogeneous geometries, where the input and output have distinct geometries. Specifically, HGSolver enhances the Physics Attention mechanism proposed by Transolver by introducing “position information”, which is named as position-enhanced physics attention. Besides, an encoder-decoder framework is also presented to tackle the heterogeneous geometries, in which the decoder processes query geometries by a cross-version of position-enhanced physics attention. As a result, HGSolver can handle sparse observations and inverse problems. The position-enhanced physics attention is better than vanilla physics attention, but is over 2x slower than the vanilla version."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"What do you mean by “completers” and “propagators” in Table 2."},"rating":{"value":2},"details_of_ethics_concerns":{"value":"This paper has an ethical issue in the appendix A1.1 and A1.4, which are directly copied from Transolver’s appendix B.1 and B.2 without any modification. I think this is a very serious academic ethical issue. The authors should delete them and directly cite the previous paper."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"(1)\tIn this paper, the authors interpret that the slice weight distribution contains “position” information of physics tokens. And it is reasonable to introduce “position” information into Physics-Attention.\n\n(2)\tThe encoder-decoder framework can empower the model to process heterogeneous geometries."},"flag_for_ethics_review":{"value":["Yes, Research integrity issues (e.g., plagiarism, dual submission)"]},"weaknesses":{"value":"### (1) Vague motivation and missing baselines.\n\nThe motivation for introducing “position” information into physics tokens is reasonable for improving physics learning. However, it is not designed for heterogeneous geometries. Actually, vanilla physics attention with the encoder-decoder architecture can also tackle heterogeneous geometries, which is also an important baseline.\n\n### (2) The efficiency is too bad. \n\nAlthough the authors have mentioned this issue in the limitations section, I do not think this issue should be neglected. It is worth noticing that vanilla Transolver demonstrates favorable scalability. If we align the running time of Transolver and TransolverXP, it may allow us to add layers to Transolver, which can be better than TransolverXP. Also, the GPU memory is not included.\n\n### (3) Missing visualization of position information.\n\nIt is reasonable to introduce the distribution information of slice weight into physics attention. However, there are no intuitive visualizations for the final learned position embedding. I think this is an essential analysis, while the current version does not contain this experiment.\n\n### (4) The current presentation is hard to understand and some statements are without support.\n\n-\tFor all the citations, there should be a blank between two words, such as “LNO \\cite{lno}”.\n\n-\tIn Section 2.2, the introduction of Physics Attention is too long. Why introduce Transolver++? Does HGSolver employ the physics attention with adaptive temperature in Transolver++?\n\n-\tIn Section 2.3, there are confusing and inconsistent usages in Z or \\mathbf{Z}, and v or \\mathbf{v}. Besides, in Section 2.3, no citations for Transolver.\n\n-\tIn Lines 239-240, the authors claim that “Physics Attention cannot encode directional information inherent in the learned bases, which may be crucial for distinguishing certain physical states.” However, this claim is without any evidence in the experiment section.\n\n-\tIn Section 2.4, there are too many notations. I think the authors should include the shape of each symbol.\n\n-\tIn Line 317 of Section 2.5, I would suggest not to reuse the symbol \\mathbf{\\Theta}. In Line 320, the definition of the RoPE operator is wrong, which should be \\phi(Embed, Coor).\n\n-\tIn Eq. (18)-(19), I think you should introduce these two equations separately."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921104620,"tcdate":1761570211208,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9544/Reviewer_7M5A"],"signatures":["ICLR.cc/2026/Conference/Submission9544/Reviewer_7M5A"],"forum":"Msj4TIAkDI","number":1,"license":"CC BY 4.0","cdate":1761570211208,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9544/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921104620,"domain":"ICLR.cc/2026/Conference","replyto":"Msj4TIAkDI","id":"QLKiY76siO","forumContent":{"TLDR":{"value":"Position-Enhansced Physics Attention Informed Heterogeneous Geometry Neural Solver for PDEs"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Neural PDE Solver","Physics Attention","Heterogeneous Geometries","Positional Encoding","Sparse learning"]},"supplementary_material":{"value":"/attachment/46b0ae4c521b4f2be42ee4f55c9c42cd2cca09be.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Partial differential equations (PDEs) provide a fundamental framework for modeling complex physical phenomena. However, modeling PDEs on heterogeneous geometries remains a significant challenge for both traditional numerical solvers and neural operator methods, as sparse observations, multiphysics interactions, and distinct discretizations often produce heterogeneous geometries between the observation and output spaces. In this work, we introduce a unified perspective on physics attention, formulating physical states as projections of observation embeddings onto learnable functional bases in Hilbert space. Building on this formulation, we introduce a position-enhanced physics attention mechanism that incorporates coordinate representations of these bases via rotary position embeddings, thereby enabling more effective modeling of heterogeneous interactions. Leveraging this mechanism, we develop HGsolver, an encoder–decoder framework designed for PDE tasks on heterogeneous domains. Extensive experiments demonstrate that HGsolver achieves state-of-the-art performance across forward, inverse, and reconstruction benchmarks under heterogeneous geometries, while a minimally modified variant, TransolverXP, also delivers competitive results on standard homogeneous benchmarks. These findings highlight the importance of effective interactions among physical states in advancing neural PDE solvers and their potential to address the complexity of the heterogeneous real-world geometries."},"_bibtex":{"value":"@misc{\nli2025hgsolver,\ntitle={{HG}solver: Position-Enhanced Physics Attention Informed Heterogeneous Geometries Neural Solver for {PDE}s},\nauthor={Zhonghao Li and Changhong Zhong and Zhang Qian},\nyear={2025},\nurl={https://openreview.net/forum?id=Msj4TIAkDI}\n}"},"title":{"value":"HGsolver: Position-Enhanced Physics Attention Informed Heterogeneous Geometries Neural Solver for PDEs"},"pdf":{"value":"/pdf/844b075cb7458c85756ca3e14b9a423c8aa63d82.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|hgsolver_positionenhanced_physics_attention_informed_heterogeneous_geometries_neural_solver_for_pdes"},"authorids":{"value":["~Zhonghao_Li2","~Changhong_Zhong1","~Zhang_Qian7"]},"authors":{"value":["Zhonghao Li","Changhong Zhong","Zhang Qian"]}},"version":2},{"content":{"summary":{"value":"The paper proposes **DGNet (Discrete Green Network)** for learning **spatiotemporal PDEs** on irregular meshes, with a core design that **explicitly decouples system evolution from source response** via a **discrete Green’s function** formulation. Concretely, it discretizes ( $\\partial_t u = L_x[u] + f$ ) into a **Crank–Nicolson** update and realizes a per-step Green operator ($G(\\Delta t) = (I-\\tfrac{\\Delta t}{2}L)^{-1}$), enabling superposition of (i) state propagation and (ii) source-term response. The spatial operator is built as a **physics–neural hybrid** ($L=L_{\\text{physics}}+L_{\\text{neural}}$): ($L_{\\text{physics}}$) uses mesh-aware gradient/Laplacian discretizations, while ($L_{\\text{neural}}$) is a GNN correction for discretization errors; a lightweight residual GNN further captures leftover dynamics. An efficient **factorize-once, solve-many** sparse LU scheme makes long rollouts practical. Experiments across **classical PDEs**, **complex geometries**, and **unseen source terms** show consistent SOTA accuracy, with marked robustness when test-time sources differ from training."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. **Precomputation & fairness.**\n   Please report a cost table (params, FLOPs/step, GPU memory, LU pre-factorization time for Eq. (9), wall-clock/epoch, hardware) and confirm matched training budgets across baselines.\n\n2. **Scalability & factorization reuse.**\n   Results on substantially larger meshes (and ideally 3D) would align with the motivation. Also clarify when (L) changes across samples (coefficients/BCs/mesh): must you re-factor each time, and what is the end-to-end impact?\n\n3. **Benchmarks & baseline selection.**\n   Add stronger, recent multiscale/transformer/graph operator baselines for irregular spatiotemporal PDEs, and justify inclusion of elliptic-focused baselines. For **cylinder flow**, reconcile why baseline errors are far worse than commonly reported (detail resolution/Re/splits/targets/configs).\n\n4. **Robustness & hybrid ablations.**\n   Provide stress tests (larger $(\\Delta t)$, stiff sources, mesh noise/density shifts) and ablations separating $(L_{\\text{physics}})$ vs. $(L_{\\text{neural}})$ to show stability and the benefit of the physics–neural hybrid."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"No ethics concerns."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. **Addresses challenging spatiotemporal PDEs on irregular meshes.**\n   The setting combines *both* complex geometry (unstructured meshes) and temporal evolution, which is closer to real CFD/physical simulations than static or grid-regular cases. Tackling this regime is practically meaningful. \n\n2. **Green’s-function–motivated decomposition.**\n   By formulating updates through a discrete Green operator, the method **explicitly separates** (i) the evolution of the initial state and (ii) the accumulated response to time-varying sources. This superposition-friendly design improves interpretability, allows cleaner handling of unseen source patterns, and provides a principled bridge between physics and learning.\n\n3. **Physics-informed design on irregular meshes.**\n   The hybrid operator and loss terms embed **physically consistent priors** (e.g., mesh-aware differential operators, stability-friendly time-stepping), encouraging fidelity to the governing equations while letting the learned components correct discretization errors. This tends to enhance robustness, boundary handling, and generalization across meshing/sampling changes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Heavy reliance on pre-factorization; fairness not quantified.**\n   The method hinges on a **“factorize once, solve many”** sparse LU of ($(I-\\tfrac{\\Delta t}{2}L)$) before rollout, then reuses the factors each step. While efficient per step, this adds a non-negligible **precomputation and memory** burden uncommon in NN-only baselines, yet there’s **no cost table** (params/FLOPs/GPU memory/wall-clock) to establish fairness. The paper mentions hardware (4×RTX4090) and vague training durations (“hours to one day”) but lacks matched-budget reporting across methods. A standardized efficiency comparison is needed.  \n\n2. **Small problem sizes vs. stated motivation.**\n   Core 2D meshes are relatively modest (e.g., **Cylinder ~2.3k nodes**, Sediments ~5.8k, Complex Obstacles ~8.8k; classical PDEs 1.3k–10k). These scales undercut the claim of handling challenging irregular domains; results on **larger meshes/3D** or higher Reynolds/complex boundaries would better substantiate practical robustness. \n\n3. **Baseline scope/positioning leave doubts.**\n   The baseline set (DeepONet, MGN, MP-PDE, PhyMPGN, **BENO**) mixes operator, graph, and hybrid methods, but (i) it’s unclear **how comparability is enforced** beyond a sentence (“tuned to have comparable parameters/training costs”) and (ii) some choices may be **mismatched to spatiotemporal settings** (e.g., BENO targets elliptic PDEs), while **strong recent multiscale/graph/transformer/point-operator** [1-5] baselines on irregular domains are missing. A head-to-head with closer, stronger contemporaries and a **capacity-controlled** comparison would sharpen novelty and fairness. \n\n[1] Feng, Mingquan, et al. \"SINGER: Stochastic Network Graph Evolving Operator for High Dimensional PDEs.\" The Thirteenth International Conference on Learning Representations. 2025.\n\n[2] Zhang, Xuan, et al. \"SineNet: Learning Temporal Dynamics in Time-Dependent Partial Differential Equations.\" The Twelfth International Conference on Learning Representations.\n\n[3] Wang, Qi, et al. \"P $^ 2$ C $^ 2$ Net: PDE-Preserved Coarse Correction Network for efficient prediction of spatiotemporal dynamics.\" Advances in Neural Information Processing Systems 37 (2024): 68897-68925.\n\n[4] Li, Zhihao, et al. \"Harnessing scale and physics: A multi-graph neural operator framework for pdes on arbitrary geometries.\" Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 2025.\n\n[5] Wu, Haixu, et al. \"Transolver: A Fast Transformer Solver for PDEs on General Geometries.\" International Conference on Machine Learning. PMLR, 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931073615,"tcdate":1761550793557,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19031/Reviewer_pbbn"],"signatures":["ICLR.cc/2026/Conference/Submission19031/Reviewer_pbbn"],"forum":"EJ8HnNTEAv","number":1,"license":"CC BY 4.0","cdate":1761550793557,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19031/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931073615,"domain":"ICLR.cc/2026/Conference","replyto":"EJ8HnNTEAv","id":"DS8AxKnwMq","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"A data-efficient method for learning spatiotemporal PDEs that generalizes to unseen source terms."},"keywords":{"value":["Partial Differential Equations","Data-efficient learning","Graph Neural Networks","Physics-Informed Machine Learning","Generalization"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Spatiotemporal partial differential equations (PDEs) underpin a wide range of scientific and engineering applications. Neural PDE solvers offer a promising alternative to classical numerical methods. However, existing approaches typically require large numbers of training trajectories, while high-fidelity PDE data are expensive to generate. Under limited data, their performance degrades substantially, highlighting their low data efficiency. A key reason is that PDE dynamics embody strong structural inductive biases that are not explicitly encoded in neural architectures, forcing models to learn fundamental physical structure from data. A particularly salient manifestation of this inefficiency is poor generalization to unseen source terms. In this work, we revisit Green’s function theory—a cornerstone of PDE theory—as a principled source of structural inductive bias for PDE learning. Based on this insight, we propose DGNet, a discrete Green network for data-efficient learning of spatiotemporal PDEs. The key idea is to transform the Green’s function into a graph-based discrete formulation, and embed the superposition principle into the hybrid physics–neural architecture which reduces the burden of learning physical priors from data, thereby improving sample efficiency. Across diverse spatiotemporal PDE scenarios, DGNet consistently achieves state-of-the-art accuracy using only tens of training trajectories. Moreover, it exhibits robust zero-shot generalization to unseen source terms, serving as a stress test that highlights its data-efficient structural design."},"_bibtex":{"value":"@inproceedings{\ntan2026dgnet,\ntitle={{DGN}et: Discrete Green Networks for Data-Efficient Learning of Spatiotemporal {PDE}s},\nauthor={Yingjie Tan and Quanming Yao and Yaqing Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EJ8HnNTEAv}\n}"},"title":{"value":"DGNet: Discrete Green Networks for Data-Efficient Learning of Spatiotemporal PDEs"},"pdf":{"value":"/pdf/107ff4674b22e4b0b19c7d4c071741bba67df063.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"tan|dgnet_discrete_green_networks_for_dataefficient_learning_of_spatiotemporal_pdes"},"authorids":{"value":["~Yingjie_Tan3","~Quanming_Yao3","~Yaqing_Wang2"]},"authors":{"value":["Yingjie Tan","Quanming Yao","Yaqing Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PHDME, a physics-informed diffusion model designed to predict dynamical systems from sparse data without access to the governing equations. The core idea is a two-stage approach: \n- Train a Gaussian process distributed Port-Hamiltonian system (GP-dPHS) on limited data, and generate synthetic data for diffusion training/validation;\n- Train a diffusion model with dPHS residual and boundary term constrains, weighted by GP uncertainty.\nThe final PHDME model generates entire spatiotemporal fields in a single pass and uses conformal prediction for uncertainty quantification.  \n\nOn a 1D wave/soft-string benchmark with limited observations, PHDME outperforms vanilla DDPM and a weak-physics DDPM and is faster at rollout than GP simulation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"### Regarding the Possibly Missing Core Equation\nThe paper defines the base loss but it never provides the uncertainty-aware physics loss $\\tilde{R}_\\text{phys}$ explicitly. The cited Appendix A.5 only contains 3D plots, not the loss function. Could the authors please provide the precise formula for this uncertainty-weighted loss?\n\n### Confusing Justification for Using Conformal Prediction\nIn section 3.2, the authors state: \"The reverse diffusion in DDPM is stochastic, which supports the exchangeability assumption of conformal prediction\". Could the authors please elaborate more on this reasoning? The validity of CP's grarantee depends on the data exchangeability assumption, not on any property of the model. How would DDPM sampling supports the data exchangeability property?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"### Originality\n- The paper combines GP-dPHS with diffusion models in a new way. The specific approach of using GP posterior samples to generate synthetic training data for diffusion models is novel. The incorporation of GP predictive variance directly into the diffusion training loss as a weighting factor represents a new approach to handle observation uncertainty in physics-informed generation. The application of conformal prediction to diffusion-generated PDE trajectories for uncertainty quantification hasn't been demonstrated before in this context.\n\n### Quality\n- The methodological design is a key strength. The mathematical formulation of GP-dPHS is proper and introduces principled inductive bias. The paper provides complete implementation details including data generation process.The paper reports metrics with variance estimates. \n\n### Clarity\n- The motivation, preliminary section and the two-stage training process are clearly presented. The assumptions are stated explicitly.\n\n### Significane\n- The potential significance of this work is high, as it is promising for data-scarce, equation-unknown regimes. The framework itself is general and will be of high interest to researchers who face similar challenges."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### Misleading Claims and a Mismatch Between Motivation and validation\nThe paper's experimental validation fails to test the very problem it claims to solve, and the presentation of this validation is misleading. \n\nThe method is repeatedly motivated by the significant challenge of modeling \"highly nonlinear and unstructured dynamics\". The authors claim their experiments use a \"physically faithful simulator\" to \"approximate the highly non-linear PDE dynamics of the soft robots\". This language suggest a complex validation. However, the only PDE used for all experiments is a 1D wave equation with fixed Dirichlet boundaries, which is canonical and linear, not a complex, nonlinear one. Also, the paper explicitly states they \"leverage the (1d wave equation) simulator to create 10,000 data samples as the real-world test set\", which is confusing/misleading terminology for synthetic data. The table 1 itself is labeled \"Metrics Test on Real-World Data\" even though the test set is generated by the simulator above. Line 412-413, the authors states \"We present the grid-average metrics for the nonlinear string PDEs in Table (4.2)\", which is again very misleading and confusing.\n\n\nWhile using the wave equation as a testbed is a reasonable for method development, the evaluation is limited to this single 1D system. This is a canonical linear PDE. For this simple system, the Hamiltonian is a known, simple quadratic function. The paper's core idea of learning a complex, unknown Hamiltonian with a GP-dPHS is an overkill. The GP is simply learning a quadratic surface. This provides zero evidence that the method's complex machinery is effective to handle complex, nonlinear dynamics that supposedly motivate this paper. \n\nTo support its central claims, a minimum requirement is that the paper should demonstrate the method's utility beyond a trivial linear case. Furthermore, the authors are suggest to be more precise in the main text. \n\n### UQ claim contradicted by the paper’s own results.\nThe paper repeatedly claims “calibrated uncertainty” and “tighter conformal thresholds,” but the reported Non-Conformity Score (NCS) shows the opposite. In Table 1, PHDME has NCS = 6.41×10⁻³, whereas the baseline DDPM has NCS = 6.82×10⁻⁶ (with lower being better for tighter bounds). That makes PHDME’s non-conformity roughly ~940× larger than DDPM’s, i.e., substantially looser bounds, not tighter. This directly undermines the UQ claim.\n\n### Missing baseline and contradiction in time comparison\nThe paper's method is a two-stage pipeline. The most critical baseline is missing: What is the accuracy of the GP-dPHS model on its own? The paper only compares its speed but not its accuracy. Also, the text in Section 4.2, when summarizing the table, states \"and the generative time of the proposed PHDME is significantly higher than the GP-dPHS\", which is the exact opposite of what the data in Table 1 shows. \n\n### Confusing baseline results: DDPM+Limited Physics\nThe \"DDPM+Limited Physics\" (boundary conditions) baseline performs worse than the pure DDPM baseline. This is highly conter-intuitive: adding correct physical information should improve, not degrade, performance. The authors are suggested to investigate and explain why adding boundary conditions resulted in worse performance. This suggests that the baselines were poorly tuned or implemented, making the results questionable."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941875251,"tcdate":1761855926975,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21657/Reviewer_4tWQ"],"signatures":["ICLR.cc/2026/Conference/Submission21657/Reviewer_4tWQ"],"forum":"kTEOG9a2W3","number":1,"license":"CC BY 4.0","cdate":1761855926975,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21657/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941875251,"domain":"ICLR.cc/2026/Conference","replyto":"kTEOG9a2W3","id":"AOeVyaK6mw","forumContent":{"TLDR":{"value":"PHDM: diffusion guided by GP-learned Port-Hamiltonian energy gradients from sparse data. No exact fuction needed; structure is conserved and uncertainty is calibrated for spatiotemporal prediction."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed Learning","Diffusion Model","Port-Hamiltonian system","Uncertainty Quantification","Gaussian process"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Diffusion models are expressive priors for generating and predicting data from high-dimensional dynamical systems. Yet, purely data-driven approaches often lack reliability and trustworthiness, motivating growing interest in physics-informed machine learning (PIML). Most existing PIML methods, however, assume access to exact governing equations during training—an assumption that fails when the dynamics are unknown or too complex to model accurately. To address this gap, we introduce PHDME (Port-Hamiltonian Diffusion Model), a physics-informed diffusion framework that learns system dynamics without requiring exact equations. Our approach first trains a Gaussian process distributed Port-Hamiltonian system (GP-dPHS) on limited observations to capture an energy-based representation of the dynamics. The GP-dPHS is then used to generate a physically consistent and diverse dataset for diffusion training. To enforce physics-consistency, we embed the GP-dPHS structure directly into the diffusion training objective through a loss that penalizes deviations from the learned Hamiltonian dynamics, weighted by the GP’s predictive uncertainty. After training, we employ conformal prediction to provide distribution-free uncertainty quantification of the generated trajectories. In this way, PHDME is designed for regimes with scarce data and unknown equations, enabling data-efficient, physically valid trajectory generation with calibrated uncertainty estimates."},"_bibtex":{"value":"@misc{\ntan2026phdme,\ntitle={{PHDME}: Physics-Informed Diffusion Models without Explicit Governing Equations},\nauthor={Kaiyuan Tan and Kendra Lee Givens and Peilun Li and Thomas Beckers},\nyear={2026},\nurl={https://openreview.net/forum?id=kTEOG9a2W3}\n}"},"title":{"value":"PHDME: Physics-Informed Diffusion Models without Explicit Governing Equations"},"pdf":{"value":"/pdf/caa63d9bc69a07229f5b0d48dcdca801e0cc83ab.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"tan|phdme_physicsinformed_diffusion_models_without_explicit_governing_equations"},"authorids":{"value":["~Kaiyuan_Tan1","~Kendra_Lee_Givens1","~Peilun_Li3","~Thomas_Beckers1"]},"authors":{"value":["Kaiyuan Tan","Kendra Lee Givens","Peilun Li","Thomas Beckers"]}},"version":2},{"content":{"summary":{"value":"The paper proposes REPST, a spatiotemporal prediction framework that enables PLM (Pre-trained Language Models) to understand complex spatiotemporal patterns through a reprogramming strategy based on physics-aware decomposition. This framework employs a physics-aware decomposer to decompose spatially correlated time series, enhancing the model's ability to comprehend the patterns. Additionally, the paper introduces a selective discrete reprogramming scheme, which projects spatiotemporal series into discrete representations."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Does REPST perform well only for short prediction lengths?\n2. Line 212: If the concatenation of two parts in $X_{dec}$ is necessary, it would be necessary to include this component in the ablation study to assess its impact. Additionally, sensitivity analysis should be conducted to evaluate the effect of the hyperparameter $\\alpha$.\n3. It is recommended to provide the detailed settings of the ablation experiments in the appendix, as the current text only gives a very brief overview of the approach.\n4. In the ablation study, why does the performance of the Solar Energy dataset show relatively little impact when the pretrained model is not used, compared to the other two datasets?\n5. Why is the dimension of E' in Figure 2 inconsistent with its description in the text?\n6. Line 285: What model does \"HI\" represent in line 285?\n7. Since the Koopman theory-based evolutionary matrix can describe the evolution of the system over time $t$, is it possible to use matrix $\\mathcal{A}$ directly for prediction? If so, it is recommended to include it in the comparative experiments."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper is the first to propose a physics-aware spatio-temporal decomposer.\n2. The experimental datasets span the fields of traffic, solar energy, and air quality, offering good diversity.\n3. REPST demonstrates strong performance under the parameters and datasets specified in the paper."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The statement that \"the rich physical semantic information can boost the pretrained physical knowledge of PLMs\" (Line 288-231) lacks experimental or theoretical evidence demonstrating that it was specifically the physical knowledge of PLMs that contributed to the results. It seems more likely that the physical methods applied simply extracted features that were more reasonable and of finer granularity.\n2. The paper, i.e., abstract part, mentions that the physics-aware decomposer allows PLM to understand complex spatiotemporal dynamics using a divide-and-conquer strategy. How exactly does this relate to the divide-and-conquer approach?\n3. The experimental setup is limited to a short prediction length, with no results for other prediction lengths.\n4. The experimental results do not include standard deviations and there is no mention of random seeds used, making it difficult to assess the statistical significance and reproducibility of the results.\n5. The complete numerical results for Figure 3 are not provided in a tabular format, which would have been helpful for detailed comparison.\n6. The paper could benefit from more visualizations to better illustrate how the physics-aware decomposer enhances interpretability.\n7. There are relatively few models compared when evaluating Zero-Shot performance.\n8. There is a typo in the middle of Section A.5, specifically in the last line of page 17 (line 916)."}},"nonreaders":[],"tmdate":1731427308734,"tcdate":1730511907400,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission732/Reviewer_i6Kn"],"signatures":["ICLR.cc/2025/Conference/Submission732/Reviewer_i6Kn"],"forum":"wCNuEA5MSv","number":3,"license":"CC BY 4.0","cdate":1730511907400,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission732/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427308734,"domain":"ICLR.cc/2025/Conference","replyto":"wCNuEA5MSv","id":"Buv7srMkrf","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["spatio-temporal forecasting","time series forecasting"]},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Spatio-temporal forecasting is pivotal in numerous real-world applications, including transportation planning, energy management, and climate monitoring. \nIn this work, we aim to harness the reasoning and generalization abilities of Pre-trained Language Models (PLMs) for more effective spatio-temporal forecasting, particularly in data-scarce scenarios. \nHowever, recent studies uncover that PLMs, which are primarily trained on textual data, often falter when tasked with modeling the intricate correlations inherent in numerical time series, thereby limiting their effectiveness in comprehending spatio-temporal data.\nTo bridge the gap, we propose REPST, a physics-aware PLM reprogramming framework tailored for spatio-temporal forecasting. \nSpecifically, we first propose a physics-aware decomposer that adaptively disentangles spatially correlated time series into interpretable sub-components, which facilitates PLM’s understanding of sophisticated spatio-temporal dynamics via a divide-and-conquer strategy.\nMoreover, we propose a selective discrete reprogramming scheme, which introduces an expanded spatio-temporal vocabulary space to project spatio-temporal series into discrete representations. This scheme minimizes the information loss during reprogramming and enriches the representations derived by PLMs.\nExtensive experiments on real-world datasets show that the proposed REPST outperforms twelve state-of-the-art baseline methods, particularly in data-scarce scenarios, highlighting the effectiveness and superior generalization capabilities of PLMs for spatio-temporal forecasting."},"_bibtex":{"value":"@misc{\nwang2025language,\ntitle={Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming},\nauthor={Hao Wang and Jindong Han and Wei Fan and Hao Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=wCNuEA5MSv}\n}"},"title":{"value":"Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming"},"pdf":{"value":"/pdf/cb7d921d45bfc48169f16c8f000ca68be9a42b3d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|language_model_empowered_spatiotemporal_forecasting_via_physicsaware_reprogramming"},"authorids":{"value":["~Hao_Wang92","~Jindong_Han1","~Wei_Fan6","~Hao_Liu17"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hao Wang","Jindong Han","Wei Fan","Hao Liu"]}},"version":2},{"content":{"summary":{"value":"This paper addresses a critical limitation in MLLMs: their poor understanding of intuitive physics, particularly for continuum objects like fluids. The authors introduce two diagnostic tasks—Next Frame Selection (NFS) and Temporal Coherence Verification (TCV)—to systematically evaluate low-level physical perception. They reveal that even state-of-the-art MLLMs perform poorly on these tasks, often near random baselines. To bridge this gap, the authors propose Scene Dynamic Field (SDF), a method that leverages physics simulators to generate visual prompts representing motion dynamics. Through multi-task fine-tuning, SDF significantly improves model performance and generalizes well to unseen physical domains like cloth and smoke."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses. I am willing to discuss with the authors"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is interesting for addressing the problem of poor low-level physical perception by proposing practical physics simulators for visual prompting to improve performance.\n\n2. The paper is well-organized and accessible, with clear explanations of both the problem and the solution."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Scope of Transfer Experiments: While the transfer experiments across continuum domains (cloth, sand, smoke, plasticine) are promising, the scope remains limited. The paper would be significantly strengthened by evaluating the method's generalization to other fundamental physical phenomena, such as rigid-body dynamics, collisions, or optical effects. \n\n2. Experimental Design and Baselines: To provide a more comprehensive performance comparison, we suggest expanding Table 1 to include additional baseline models. Specifically, it would be informative to compare against larger models with more parameters.\n\n3. Dependence on Synthetic Data and Real-World Applicability: The SDF method's reliance on synthetic data from simulators is a potential limitation, as such data may not fully capture the complexity and noise of real-world physical systems. The paper could be improved by discussing the feasibility and potential challenges of applying this method to real-world physics problems. For instance, how would the SDF approach perform with real sensor data that is often incomplete or noisy?\n\n4. Analysis of MLLM Limitations: The paper identifies that MLLMs struggle with low-level dynamics but does not deeply investigate the root cause. A more thorough analysis is needed to determine whether these failures stem from architectural limitations (e.g., an inductive bias towards high-level semantics) or a bias in the pre-training data (e.g., a lack of low-level physical reasoning examples). Uncovering this would provide valuable insight for future research."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919792420,"tcdate":1760575789632,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7741/Reviewer_HQBD"],"signatures":["ICLR.cc/2026/Conference/Submission7741/Reviewer_HQBD"],"forum":"Ax02eR2c3d","number":1,"license":"CC BY 4.0","cdate":1760575789632,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7741/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919792420,"domain":"ICLR.cc/2026/Conference","replyto":"Ax02eR2c3d","id":"ZLs3U9qGC9","forumContent":{"TLDR":{"value":"This paper introduces two low-level tasks to test intuitive physics understanding and proposes Scene Dynamic Field, a method to integrate visual representation from physics simulators to MLLMs while showcasing generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multi-modal LLM","Intuitive Physics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. \nIn this work, we investigate the first step of physical reasoning, i.e., **intuitive physics understanding**, revealing substantial limitations in understanding the dynamics of continuum objects. \nTo isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. \nTo address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. \nSDF substantially improves performance, achieving up to $20.7\\%$ gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field."},"_bibtex":{"value":"@inproceedings{\nli2026beyond,\ntitle={Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models},\nauthor={Nanxi Li and Xiang Wang and Yuanjie Chen and Haode Zhang and Hong Li and Yong-Lu Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Ax02eR2c3d}\n}"},"title":{"value":"Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models"},"pdf":{"value":"/pdf/3b199b74bbdbddb8166b5e799fc00ddb71a413ae.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|beyond_static_vision_scene_dynamic_field_unlocks_intuitive_physics_understanding_in_multimodal_large_language_models"},"authorids":{"value":["~Nanxi_Li1","~Xiang_Wang19","~Yuanjie_Chen2","~Haode_Zhang2","~Hong_Li5","~Yong-Lu_Li1"]},"authors":{"value":["Nanxi Li","Xiang Wang","Yuanjie Chen","Haode Zhang","Hong Li","Yong-Lu Li"]}},"version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_61.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"zhang|correlation_properties_and_selfsimilarity_of_renormalization_email_networks"},"authorids":{"value":["~Lianming_Zhang1","https://dblp.org/search/pid/api?q=author:Sundong_Liu:","https://dblp.org/search/pid/api?q=author:Yuling_Tang:","https://dblp.org/search/pid/api?q=author:Hualan_Xu:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_61"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/ZhangLTX09,\n  author={Lianming Zhang and Sundong Liu and Yuling Tang and Hualan Xu},\n  title={Correlation Properties and Self-similarity of Renormalization Email Networks},\n  year={2009},\n  cdate={1230768000000},\n  pages={1846-1859},\n  url={https://doi.org/10.1007/978-3-642-02469-6_61},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"A degree-thresholding renormalization method is recently introduced to find topological characteristics of some complex networks. As a matter of fact, the applicability of these characteristics depends on the level or the type of complex networks. Here, a modified version of this original algorithm is presented to unravel ubiquitous characteristics of observed email networks and obtain correct understanding of underlying evolutionary mechanism. Some topology metrics of the email networks under renormalization were analyzed. The results show that renormalization email networks have the power-law distribution with double exponents, are disassortative and become assortative after half of total renormalization steps, have high-clustering coefficients and rich-club phenomena. These characteristics are self-similar both before and after renormalization until half of total renormalization steps, otherwise are self-dissimilar."},"title":{"value":"Correlation Properties and Self-similarity of Renormalization Email Networks"},"authors":{"value":["Lianming Zhang","Sundong Liu","Yuling Tang","Hualan Xu"]}},"tmdate":1770688362656,"pdate":1262217600000,"externalIds":["dblp:conf/complex/ZhangLTX09"],"tcdate":1770688328016,"writers":["~"],"signatures":["~Lianming_Zhang1"],"forum":"mMDw7IWWSl","license":"CC BY-SA 4.0","number":821397,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1770688362656,"domain":"DBLP.org","id":"mMDw7IWWSl","version":2},{"content":{"venue":{"value":"Complex (1) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02466-5_34.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"hui|identifying_social_communities_in_complex_communications_for_network_efficiency"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Pan_Hui_0001:","~Eiko_Yoneki1","https://dblp.org/search/pid/api?q=author:Jon_Crowcroft:","https://dblp.org/search/pid/api?q=author:Shu_Yan_Chan:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02466-5_34"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/HuiYCC09,\n  author={Pan Hui and Eiko Yoneki and Jon Crowcroft and Shu Yan Chan},\n  title={Identifying Social Communities in Complex Communications for Network Efficiency},\n  year={2009},\n  cdate={1230768000000},\n  pages={351-363},\n  url={https://doi.org/10.1007/978-3-642-02466-5_34},\n  booktitle={Complex (1)},\n  crossref={conf/complex/2009-1}\n}\n"},"abstract":{"value":"Complex communication networks, more particular Mobile Ad Hoc Networks (MANET) and Pocket Switched Networks (PSN), rely on short range radio and device mobility to transfer data across the network. These kind of mobile networks contain duality in nature: they are radio networks at the same time also human networks, and hence knowledge from social networks can be also applicable here. In this paper, we demonstrate how identifying social communities can significantly improve the forwarding efficiencies in term of delivery ratio and delivery cost. We verify our hypothesis using data from five human mobility experiments and test on two application scenarios, asynchronous messaging and publish/subscribe service."},"title":{"value":"Identifying Social Communities in Complex Communications for Network Efficiency"},"authors":{"value":["Pan Hui","Eiko Yoneki","Jon Crowcroft","Shu Yan Chan"]}},"tmdate":1730577182951,"pdate":1230768000000,"tcdate":1730577122721,"writers":["~"],"signatures":["~Eiko_Yoneki1"],"forum":"y0HPQQyZbb","license":"CC BY-SA 4.0","number":168438,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1730577182951,"domain":"DBLP.org","id":"y0HPQQyZbb","version":2},{"content":{"summary":{"value":"The paper presents **STANCE**, a controllable image-to-video generation framework designed for physically coherent motion and per-instance editing, built upon the CogVideoX (5B-parameter) backbone. The approach targets the gap between intuitive per-object controls and globally consistent motion in video diffusion models.\n\n**Key Contributions:**\n1. Instance Cues: A mechanism that transforms sparse, user-provided annotations, such as object-level arrows, masks, optional mass, and a scalar ∆z into a dense, pixel-aligned 2.5D motion field used to condition the diffusion model. This formulation grounds the motion control in geometric consistency and user editability.\n2. Dense RoPE: A dense control-token strategy employing rotary positional embeddings to preserve spatially localized motion anchors within a DiT-based video diffusion backbone, ensuring better control fidelity and reduced spatial drift.\n\nAdditionally, the framework supports joint generation of RGB frames and a secondary \"structural witness\" (e.g., depth or segmentation), intended to stabilize temporal dynamics and improve motion realism.\n\nQuantitative evaluations using Physics-IQ and FVD demonstrate consistent improvements over strong SVD-based baselines such as SG-I2V, Drag-Anything, MoFA-Adapter, and MotionPro (all ~1.5B parameters). Qualitative results show noticeably improved motion coherence, temporal stability, and control accuracy in both synthetic and simple real-world scenes."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Computational cost\n    - How much additional training time, GPU memory, or inference latency does joint RGB + structure generation introduce compared to single RGB output?\n    - Is there any measurable efficiency or stability gain that justifies this added cost?\n- Evaluation scope\n    - The paper does not explore how STANCE behaves on longer sequences or more complex real-world scenes. Although the model is trained primarily on simple synthetic data, it would be valuable to understand its ability to generalize or adapt to more complex motion and interactions.\n    - The robustness of the method to imperfect user inputs, such as noisy masks, misaligned arrows, or ambiguous object boundaries, is not analyzed. An evaluation of sensitivity to such input noise would strengthen the empirical section.\n- Missing related work\n    - The paper does not cite Wan-VACE, which is relevant as a recent video generation model emphasizing motion coherence and controllable conditioning.It should be discussed and cited in the related works section to provide a complete comparison landscape.\n- Writing style\n    - Although the paper explicitly mentions the use of language-model assistance under **\"G Writing Assistance (LLM Use Disclosure)\"**, some passages, especially those with repeated em-dash usage and overly polished phrasing, sound distinctly machine-generated. A careful human language revision could improve readability and overall presentation quality."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Clear motivation and problem focus\n    - Addresses the specific gap of achieving physically coherent, instance-controllable motion in image-to-video generation rather than general visual quality.\n- Effective integration of control signals\n    - Combines arrows, masks, ∆z, and optional mass into a unified 2.5D motion field, creating an interpretable and editable conditioning scheme.\n- Strong experimental design\n    - Ablations clearly show the contribution of each proposed component, and the method consistently outperforms all SVD-based baselines (~1.5B) in Physics-IQ and FVD despite being built on a larger 5B backbone."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Model scale and fairness of comparison\n    - Using a 5B-parameter CogVideoX backbone on a relatively simple synthetic dataset seems excessive. The observed gains may partly arise from sheer model capacity rather than the proposed techniques. Since all baselines (SG-I2V, Drag-Anything, MoFA-Adapter, MotionPro) are SVD-based models around 1.5B parameters, a more balanced comparison would involve using an SVD backbone or a smaller-scale CogVideo variant.\n    - Additionally, instead of full fine-tuning, LoRA-based or other efficient adaptation methods could be explored, especially given the strength of the base model, to verify whether the improvements generalize without retraining the entire network.\n- Computational cost of joint generation\n    - The joint RGB + structure (depth/segmentation) generation likely increases training and inference cost, yet the paper doesn’t quantify the overhead or analyze the trade-off between the added cost and the relatively small metric gain.\n- Limited evaluation scope\n    - Experiments are confined to synthetic or very simple real scenes with short sequences and limited dynamics. The method’s effectiveness on more complex, realistic datasets remains untested.\n- Physics-IQ explanation could be clearer\n    - While the metric itself is cited, providing a short subsection or appendix note summarizing its components and interpretation would make the evaluation more self-contained and easier to follow for readers unfamiliar with the metric."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916463981,"tcdate":1761814038529,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2961/Reviewer_itDA"],"signatures":["ICLR.cc/2026/Conference/Submission2961/Reviewer_itDA"],"forum":"FwtKMYHov7","number":2,"license":"CC BY 4.0","cdate":1761814038529,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2961/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916463981,"domain":"ICLR.cc/2026/Conference","replyto":"FwtKMYHov7","id":"SwN91knFeU","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Generation","Generative Model"]},"supplementary_material":{"value":"/attachment/415a700d351d8bd244411f2900432bea49e53253.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation has recently made striking visual progress, but maintaining coherent object motion and interactions remains difficult. We trace two practical bottlenecks: (i) human-provided motion hints (e.g., small 2D maps) often collapse to too few effective tokens after encoding, weakening guidance; and (ii) optimizing for appearance and motion in a single head can favor texture over temporal consistency. We present STANCE, an image-to-video framework that addresses both issues with two simple components.\nFirst, we introduce Instance Cues—a pixel-aligned control signal that turns sparse, user-editable hints into a dense 2.5D (camera-relative) motion field by averaging per-instance flow and augmenting with monocular depth over the instance mask. This reduces depth ambiguity compared to 2D drag/arrow inputs while remaining easy to user. Second, we preserve the salience of these cues in token space with Dense RoPE, which tags a small set of motion tokens (anchored on the first frame) with time-addressable rotary embeddings. Paired with joint RGB + auxiliary-map prediction (segmentation or depth), our model anchors structure while RGB handles appearance, stabilizing optimization and improving temporal coherence without requiring per-frame trajectory scripts."},"_bibtex":{"value":"@misc{\nanonymous2026stance,\ntitle={{STANCE}: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=FwtKMYHov7}\n}"},"title":{"value":"STANCE: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding"},"pdf":{"value":"/pdf/dcb00196385b99164d59c430fb8613633f2432c0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"chen|stance_motion_coherent_video_generation_via_sparsetodense_anchored_encoding"},"authorids":{"value":["~ZhiFei_Chen1","~Tianshuo_Xu1","~Leyi_Wu1","~Luozhou_Wang2","~Dongyu_Yan1","~Zihan_You2","~Wenting_Luo1","~Guo_Zhang2","~Ying-Cong_Chen1"]},"authors":{"value":["ZhiFei Chen","Tianshuo Xu","Leyi Wu","Luozhou Wang","Dongyu Yan","Zihan You","Wenting Luo","Guo Zhang","Ying-Cong Chen"]}},"version":2},{"content":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Intuitive physics","physical reasoning","mental simulation","heuristic model"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The role of mental simulation in human behavior for various physical tasks is widely acknowledged, attributed to the generality of Intuitive Physics Engine (IPE). However, it remains unclear whether mental simulation is consistently employed across scenarios of different simulation costs and where its boundary is. Moreover, cognitive strategies beyond these boundaries have not been thoroughly investigated. Here, we adopted a pouring-marble task containing various conditions to study IPE's limits and strategies beyond. A human study revealed two distinct error patterns in predicting the pouring angle, differentiated by the simulation time using a boundary. This suggests a possible switching of the underlying reasoning strategies. Our initial experiment on IPE showed that its correlation with human judgments diminished in scenarios requiring extended time of simulation. This observation prompted the exploration of an alternative mechanism based on heuristics for intuitive physics. We uncovered that a linear heuristic model, relying exclusively on empirical data, replicated human prediction more accurately when the simulation time exceeded a certain boundary. Motivated by these observations, we propose a new framework, Simulation-Heuristics Model (SHM), which conceptualizes intuitive physics as a dual process: IPE is predominant only in short-time simulation, whereas a heuristics-based approach is applied as IPE's simulation time extends beyond the simulation boundary. The SHM model aligns more precisely with human behavior across various scenarios and demonstrates superior generalization capabilities under different conditions. Crucially, SHM integrates computational methods previously viewed as separate into a unified model, quantitatively studying their switching mechanism."},"_bibtex":{"value":"@misc{\nli2024a,\ntitle={A simulation-heuristics dual-process model for intuitive physics},\nauthor={Shiqian Li and Yuxi Ma and Bo Dai and Yujia Peng and Chi Zhang and Yixin Zhu},\nyear={2024},\nurl={https://openreview.net/forum?id=BkeJro1xps}\n}"},"title":{"value":"A simulation-heuristics dual-process model for intuitive physics"},"pdf":{"value":"/pdf/93856d41c3aa9841ddf51a3fb5488577d1272d6a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|a_simulationheuristics_dualprocess_model_for_intuitive_physics"},"authorids":{"value":["~Shiqian_Li1","~Yuxi_Ma2","~Bo_Dai5","~Yujia_Peng1","~Chi_Zhang12","~Yixin_Zhu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shiqian Li","Yuxi Ma","Bo Dai","Yujia Peng","Chi Zhang","Yixin Zhu"]}},"tmdate":1733031002159,"tcdate":1726285242663,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission634/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission634/Authors"],"forum":"BkeJro1xps","license":"CC BY 4.0","number":634,"cdate":1726285242663,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/-/Withdrawn_Submission"],"mdate":1733031002159,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"BkeJro1xps","version":2},{"content":{"summary":{"value":"This paper proposes a novel deep learning framework for Non-Line-of-Sight (NLOS) gesture recognition. The core of this method is a 'physics-inspired' architecture that attempts to integrate principles from optical physics with deep learning models. The authors build the model using complex mathematical formulations and specific network components to process NLOS data. The evaluation on three different benchmarks demonstrates its effectiveness."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"See Weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":3},"strengths":{"value":"- Proposes a new deep learning architecture specifically designed to solve the challenging, cross-disciplinary task of human gesture recognition from Non-Line-of-Sight imaging data.\n- Designs a 'physics-inspired' core model that embed optical physics phenomena into the neural network structure through specific components.\n- Validates the method's performance on three different benchmarks and demonstrate the performance effectiveness of the proposed components through ablation studies."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  The related work section fails to clarify the connection between existing studies and this research. The authors summarize the developments in NLOS imaging and gesture recognition separately but do not highlight their specific relevance to the proposed work. And, a discussion of related work addressing the specific intersectional need this paper focuses on is missing. Additionally, the phrase \"these methods\" in lines 146-147 lacks corresponding citations.\n2.  The method description lacks significant details.\n    1.  What is the motivation for the outer product expansion in Equs. 1, 4, and 6? Why is it formulated this way? Does this outer product form not lead to a dramatic explosion in the number of tensor elements? Moreover, the shapes on the left and right sides of the equations do not match, as the resulting outer product on the right is not a 2D matrix.\n    2.  The meanings of $\\xi$ and $\\eta$ in Line 233 should be explicitly defined.\n    3.  How is the directional filter designed in Equ. 2 designed? How are the specific angles and the number of angles selected?\n    4.  Line 242 mentions aggregating weighted responses from all directions but does not explain how this aggregation is performed.\n    5.  What does $\\Delta$ represent in Equ. 4?\n    6.  What is the meaning of $S_m$ in Equ. 5? Is it a hyperparameter? How is it set?\n    7.  How is the \"adaptive kernel function\" mentioned in Line 269, which can \"adaptively adjust its shape\", designed? Does it use an existing algorithm?\n    8.  What are $H’$ and $W’$ in Line 296? How do they differ from the output height $\\alpha$ and width $\\beta$ (from Line 208)? What is the original input shape? Why is it necessary to further expand it to $H’ \\times W’$?\n    9.  What does the symbol $\\mathcal{S}^{\\circ 3}$ in Line 383 specifically refer to?\n    10. The paper uses three different datasets but lacks details about the datasets themselves, such as data format, dimensions, and scale.\n3.  The ablation study is too coarse. It merely provides a combinatoric comparison of component performance. The current experiments and corresponding analysis fail to reflect the reasonableness of individual components or link them to their intended design motivations.\n4.  The paper's contribution is obscured by its opaque and dense writing style. The readability is poor, posing a significant barrier to accurately assessing its novelty, technical correctness, and contribution. The authors tend to use lengthy, overly-specialized jargon and forcibly \"stitch\" complex concepts from optical physics and deep learning into single sentences, resulting in convoluted explanations. More critically, the core \"physics-inspired\" design lacks clear, high-level intuition. Instead of first conceptually explaining why specific architectural choices are reasonable analogies for particular physical phenomena, the authors directly present structural descriptions and implementation details, before immediately jumping into complex mathematical derivations. This makes the \"physics-inspired\" claims feel more like post-hoc justifications rather than the driving force behind the design.\n5.  There are formatting issues. The authors have excessively compressed the line spacing in the text (e.g., Lines 203 and 308). I am unsure if this violates the ICLR conference's formatting requirements, but it creates a visually jarring layout."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921709961,"tcdate":1761651032526,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10388/Reviewer_hbsw"],"signatures":["ICLR.cc/2026/Conference/Submission10388/Reviewer_hbsw"],"forum":"AlqrnU93o7","number":1,"license":"CC BY 4.0","cdate":1761651032526,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10388/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921709961,"domain":"ICLR.cc/2026/Conference","replyto":"AlqrnU93o7","id":"kY98p91vWt","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Gesture recognition","Computer Vision","Deep Learning","Pattern recognition"]},"supplementary_material":{"value":"/attachment/271ea332c6e17de92d5de6b15f01e4e1127e02f4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Accurately decoding hidden information in dynamic shadows for Non-Line-of-Sight (NLOS) imaging enables us to overcome visual occlusions and perceive or reconstruct obscured targets. This breakthrough holds significant potential for real-world applications such as disaster rescue, autonomous driving, and security surveillance. Conventional algorithms struggle to model the physical propagation of light in space. Furthermore, the signal distortions introduced by nonlinear transformations incur the loss of geometric information about the source scene, limiting sensitivity to subtle shadow variations. To overcome these challenges, we present Radiation-constraint Network (RacoNet) that marries physical propagation simulation with geometric-information recovery to interpret minute gesture signals embedded in dynamic shadows. In RacoNet, Radiance-Constrained Light-Transportation (RCLT) optical propagation is proposed to capture complete light-space information. Meanwhile, Geometric Information Aliment Operation (GIAO) restores source-scene geometry lost in the modulated shadow through layer-by-layer refined prior attention. Moreover, Kolmogorov-Arnold Enhanced Layerwise Nonlinear Reorganization (KA-ELNR) fuses light-space and geometric cues to produce the final decoded output. Extensive experiments show that RacoNet markedly surpasses existing approaches in both accuracy and robustness for dynamic-shadow decoding, confirming the possibility of gesture-based information interaction via shadows."},"_bibtex":{"value":"@misc{\nzheng2026shadowspeak,\ntitle={ShadowSpeak: Is It Possible to Communicate Cross-Room Solely by Decoding Gesture Shadows?},\nauthor={Zhiwen Zheng and Yubo Chen and Shaowei Jiang and Huiyu Zhou and Zhao Huang and Tao Zhang and Jin Liu and Guangyuan Zhang and Xiaoshuai Zhang and Xingru Huang},\nyear={2026},\nurl={https://openreview.net/forum?id=AlqrnU93o7}\n}"},"title":{"value":"ShadowSpeak: Is It Possible to Communicate Cross-Room Solely by Decoding Gesture Shadows?"},"pdf":{"value":"/pdf/1b1a3d75d4826dc6ce191c5e822701a9f6d2d043.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zheng|shadowspeak_is_it_possible_to_communicate_crossroom_solely_by_decoding_gesture_shadows"},"authorids":{"value":["~Zhiwen_Zheng1","~Yubo_Chen4","~Shaowei_Jiang1","~Huiyu_Zhou3","~Zhao_Huang2","~Tao_Zhang5","~Jin_Liu22","~Guangyuan_Zhang1","~Xiaoshuai_Zhang2","~Xingru_Huang1"]},"authors":{"value":["Zhiwen Zheng","Yubo Chen","Shaowei Jiang","Huiyu Zhou","Zhao Huang","Tao Zhang","Jin Liu","Guangyuan Zhang","Xiaoshuai Zhang","Xingru Huang"]}},"version":2},{"content":{"venue":{"value":"Journal of Experimental Psychology: Human Perception and Performance"},"pdf":{"value":"/pdf/6d2450fc54530bb4f6b6b674956c88352e4e2874.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"yates|temporal_segmentation_and_look_ahead_simulation_physical_events_structure_visual_perception_of_intuitive_physics"},"authorids":{"value":["tristan.yates@columbia.edu","~Shannon_Yasuda1","~Ilker_Yildirim2"]},"abstract":{"value":"How we perceive the physical world is not only organized in terms of objects, but also structured in time as sequences of events. This is especially evident in intuitive physics, with temporally bounded dynamics such as falling, occlusion, and bouncing demarcating the continuous flow of sensory inputs. While the spatial structure and attentional consequences of physical objects have been well-studied, much less is known about the temporal structure and attentional consequences of physical events in visual perception. Previous work has recognized physical events as units in the mind, and used presegmented object interactions to explore physical representations. However, these studies did not address whether and how perception imposes the kind of temporal structure that carves these physical events to begin with, and the attentional consequences of such segmentation during intuitive physics. Here, we use performance-based tasks to address this gap. In Experiment 1, we find that perception not only spontaneously separates visual input in time into physical events, but also, this segmentation occurs in a nonlinear manner within a few hundred milliseconds at the moment of the event boundary. In Experiment 2, we find that event representations, once formed, use coarse “look ahead” simulations to selectively prioritize those objects that are predictively part of the unfolding dynamics. This rich temporal and predictive structure of physical event representations, formed during vision, should inform models of intuitive physics."},"title":{"value":"Temporal segmentation and “look ahead” simulation: Physical events structure visual perception of intuitive physics."},"authors":{"value":["Tristan Yates","Shannon Yasuda","Ilker Yildirim"]}},"tmdate":1776971240545,"pdate":1718856000000,"tcdate":1776971240545,"writers":["tristan.yates@columbia.edu","~Shannon_Yasuda1","~Ilker_Yildirim2"],"signatures":["~Shannon_Yasuda1"],"forum":"ewqClUUmNx","license":"CC BY 4.0","number":47391,"cdate":1776971240545,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1776971240545,"domain":"OpenReview.net/Archive","id":"ewqClUUmNx","version":2},{"content":{"venue":{"value":"AAAI 2021"},"pdf":{"value":"https://ojs.aaai.org/index.php/AAAI/article/download/16159/15966"},"venueid":{"value":"dblp.org/conf/AAAI/2021"},"paperhash":{"value":"alhasoun|probabilistic_programming_bots_in_intuitive_physics_game_play"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Fahad_Alhasoun:","~Sarah_Alnegheimish1"]},"html":{"value":"https://doi.org/10.1609/aaai.v35i1.16159"},"_bibtex":{"value":"@inproceedings{DBLP:conf/aaai/AlhasounA21,\n  author={Fahad Alhasoun and Sarah Alnegheimish},\n  title={Probabilistic Programming Bots in Intuitive Physics Game Play},\n  year={2021},\n  cdate={1609459200000},\n  pages={778-783},\n  url={https://doi.org/10.1609/aaai.v35i1.16159},\n  booktitle={AAAI},\n  crossref={conf/aaai/2021}\n}\n"},"abstract":{"value":"Recent findings suggest that humans deploy cognitive mechanism of physics simulation engines to simulate the physics of objects. We propose a framework for bots to deploy probabilistic programming tools for interacting with intuitive physics environments. The framework employs a physics simulation in a probabilistic way to infer about moves performed by an agent in a setting governed by Newtonian laws of motion. However, methods of probabilistic programs can be slow in such setting due to their need to generate many samples. We complement the model with a model-free approach to aid the sampling procedures in becoming more efficient through learning from experience during game playing. We present an approach where combining model-free approaches (a convolutional neural network in our model) and model-based approaches (probabilistic physics simulation) is able to achieve what neither could alone. This way the model outperforms an all model-free or all model-based approach. We discuss a case study showing empirical results of the performance of the model on the game of Flappy Bird."},"title":{"value":"Probabilistic Programming Bots in Intuitive Physics Game Play"},"authors":{"value":["Fahad Alhasoun","Sarah Alnegheimish"]}},"tmdate":1748006218056,"pdate":1609459200000,"tcdate":1748006209853,"writers":["~"],"signatures":["~Sarah_Alnegheimish1"],"forum":"YLyXFNF7Gi","license":"CC BY-SA 4.0","number":549673,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1748006218056,"domain":"DBLP.org","id":"YLyXFNF7Gi","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_74.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"geng|emergence_of_scalefree_networks_with_seceding_mechanism"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xian-Min_Geng:","~Guanghui_Wen1","https://dblp.org/search/pid/api?q=author:Shu-Chen_Wan:","https://dblp.org/search/pid/api?q=author:Jie-Yu_Xiong:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_74"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/GengWWX09,\n  author={Xian-Min Geng and Guanghui Wen and Shu-Chen Wan and Jie-Yu Xiong},\n  title={Emergence of Scale-Free Networks with Seceding Mechanism},\n  year={2009},\n  cdate={1230768000000},\n  pages={1973-1983},\n  url={https://doi.org/10.1007/978-3-642-02469-6_74},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"In order to explore further the underlying mechanism of the scale-free networks, we study stochastic secession as a mechanism for the creation of complex networks. In this evolution the network growth incorporates the addition of new links between existing nodes, the deleting and rewiring of some existing links, and the stochastic secession of nodes. To random growing networks with preferential attachment, the model yields scale-free behavior for the degree distribution. Furthermore, we get the analytical expression of the power law degree distribution with scaling exponent γ ranges from 1.1 to 9. The analytical expressions are in good agreement with the numerical simulation results."},"title":{"value":"Emergence of Scale-Free Networks with Seceding Mechanism"},"authors":{"value":["Xian-Min Geng","Guanghui Wen","Shu-Chen Wan","Jie-Yu Xiong"]}},"tmdate":1749148047600,"pdate":1230768000000,"tcdate":1749147714000,"writers":["~"],"signatures":["~Guanghui_Wen1"],"forum":"S3y2Nkpcva","license":"CC BY-SA 4.0","number":557214,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1749148047600,"domain":"DBLP.org","id":"S3y2Nkpcva","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_74.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"geng|emergence_of_scalefree_networks_with_seceding_mechanism"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xian-Min_Geng:","~Guanghui_Wen1","https://dblp.org/search/pid/api?q=author:Shu-Chen_Wan:","https://dblp.org/search/pid/api?q=author:Jie-Yu_Xiong:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_74"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/GengWWX09,\n  author={Xian-Min Geng and Guanghui Wen and Shu-Chen Wan and Jie-Yu Xiong},\n  title={Emergence of Scale-Free Networks with Seceding Mechanism},\n  year={2009},\n  cdate={1230768000000},\n  pages={1973-1983},\n  url={https://doi.org/10.1007/978-3-642-02469-6_74},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"In order to explore further the underlying mechanism of the scale-free networks, we study stochastic secession as a mechanism for the creation of complex networks. In this evolution the network growth incorporates the addition of new links between existing nodes, the deleting and rewiring of some existing links, and the stochastic secession of nodes. To random growing networks with preferential attachment, the model yields scale-free behavior for the degree distribution. Furthermore, we get the analytical expression of the power law degree distribution with scaling exponent γ ranges from 1.1 to 9. The analytical expressions are in good agreement with the numerical simulation results."},"title":{"value":"Emergence of Scale-Free Networks with Seceding Mechanism"},"authors":{"value":["Xian-Min Geng","Guanghui Wen","Shu-Chen Wan","Jie-Yu Xiong"]}},"tmdate":1749147959587,"pdate":1230768000000,"tcdate":1749147708879,"writers":["~"],"signatures":["~Guanghui_Wen1"],"forum":"Gni2roZNvc","license":"CC BY-SA 4.0","number":557029,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1749147959587,"domain":"DBLP.org","id":"Gni2roZNvc","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_11.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"shi|a_new_genetic_algorithm_for_community_detection"},"authorids":{"value":["~Chuan_Shi1","https://dblp.org/search/pid/api?q=author:Yi_Wang_0010:","https://dblp.org/search/pid/api?q=author:Bin_Wu_0001:","https://dblp.org/search/pid/api?q=author:Cha_Zhong:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_11"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/ShiWWZ09,\n  author={Chuan Shi and Yi Wang and Bin Wu and Cha Zhong},\n  title={A New Genetic Algorithm for Community Detection},\n  year={2009},\n  cdate={1230768000000},\n  pages={1298-1309},\n  url={https://doi.org/10.1007/978-3-642-02469-6_11},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"With the rapidly grown evidence that various systems in nature and society can be modeled as complex networks, community detection in networks becomes a hot research topic in many research fields. This paper proposes a new genetic algorithm for community detection. The algorithm uses the fundamental measure criterion modularity Q as the fitness function. A special locus-based adjacency encoding scheme is applied to represent the community partition. The encoding scheme is suitable for the community detection based on the reason that it determines the community number automatically and reduces the search space distinctly. In addition, the corresponding crossover and mutation operators are designed. The experiments in three aspects show that the algorithm is effective, efficient and steady."},"title":{"value":"A New Genetic Algorithm for Community Detection"},"authors":{"value":["Chuan Shi","Yi Wang","Bin Wu","Cha Zhong"]}},"tmdate":1734002912726,"pdate":1230768000000,"tcdate":1734002863288,"writers":["~"],"signatures":["~Chuan_Shi1"],"forum":"Z0DwdTocjC","license":"CC BY-SA 4.0","number":256050,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1734002912726,"domain":"DBLP.org","id":"Z0DwdTocjC","version":2},{"content":{"summary":{"value":"The manuscript addresses edge-flow imputation under conservation on graphs. The proposed method comprises three key components: (1) an initial minimum-norm balanced completion that satisfies B f = c while keeping observed edges fixed; (2) a physics-aware group-action subspace for all observed edges from which an orthonormal basis U is constructed (truncated to k columns), and attention over this basis selects divergence-free corrections that cannot alter observed edges; (3) a lightweight Tikhonov refinement solved as a single SPD system with exact implicit differentiation via a Cholesky/CG solve. Experiments on Traffic, Power, and Bike benchmarks demonstrate consistent gains over a broad set of baselines, with ablations for basis size, attention, and the bilevel/implicit components."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.\tA single global λ seems suboptimal if the noise isn't uniform. If some parts of the network are much noisier, how does the model adapt? Dose learning a different λ for each edge provide the flexibility needed to handle different noise levels?\n2.\tMany state of the art variants of GNN are selected as baseline, however, some other physics informed machine learning model (with similar ideas) may also perform well (thus be a good baseline candidate) on such tasks, e.g. [1] (transformer with interpretable basis), [2] (neural operator with learnable basis).\n3.\tThe proposed model is claimed to have a good balance in physics and noisy data, however, the performance on data is verified on better RMSE/MAE/CORR, then what about physics? E.g. is the Kirchhoff’s law respected better in the power case？\n4.\tAre the k basis vectors sorted by singular value? Does this mean you're assuming that large-scale flow patterns are important, while small, local ones can be ignored?\n5.\tYour global attention applies one uniform fix across the entire graph. How does it handle specific local regions that have completely different or more complex physics? Does the model fail to capture these important local dynamics?\n\n[1] Cao, Shuhao. \"Choose a transformer: Fourier or galerkin.\" NeurIPS (2021): 24924-24940.\n[2] Lu, Lu, et al. \"A comprehensive and fair comparison of two neural operators (with practical extensions) based on fair data.\" Computer Methods in Applied Mechanics and Engineering 393 (2022): 114778."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Clear physics prior with strict guarantees: the anchor solution and group-action construction ensure corrections remain on the B f = c manifold and do not alter sensor edges.\n2. Interpretability: attention over ker(B) basis elements (zero on observed edges) admits a physical reading as redistributions consistent with conservation; reported sparsity aids inspection.\n3. Efficient inner solver with exact hypergradients: the Tikhonov layer yields an SPD system amenable to stable Cholesky/CG solves; implicit differentiation avoids unrolling.\n4. Thorough baselines and ablations across physics-free, physics-soft, and bilevel methods; ablations isolate the effect of k, group actions, attention, and bilevel training. Reported gains are consistent across RMSE/MAE/CORR and across domains (traffic, power, bikes)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The core components of the proposed model, such as the single regularization weight and the global attention, assume uniform noise and physics, limiting its flexibility in handling more complex, heterogeneous networks.\n2. While the paper claims physics-awareness, it only evaluates data-fit metrics (like RMSE) and fails to directly measure how well the final predictions actually adhere to the physical conservation laws.\n3. The comparison could be improved by including other relevant physics-informed baselines in addition to GNNs; meanwhile, key details about the basis vector sorting method are unclear."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924959760,"tcdate":1761692504332,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14572/Reviewer_MoiE"],"signatures":["ICLR.cc/2026/Conference/Submission14572/Reviewer_MoiE"],"forum":"vu1IEpdUQh","number":2,"license":"CC BY 4.0","cdate":1761692504332,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14572/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924959760,"domain":"ICLR.cc/2026/Conference","replyto":"vu1IEpdUQh","id":"JOeRraQmQu","forumContent":{"TLDR":{"value":"FlowSymm is an end-to-end graph neural model that completes missing flows by enforcing a divergence-free group-action prior, scoring corrections with attention, and refining with feature-conditioned Tikhonov regularization"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["graphs","networks","flow graphs","graph attention networks","group action","bilevel-optimization","physics-aware graph neural networks"]},"supplementary_material":{"value":"/attachment/64543588602e2348957ff8821e15ddc0f5206c70.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Recovering missing flows on the edges of a network, while exactly respecting local conservation laws, is a fundamental inverse problem that arises in many systems such as transportation, energy, and mobility. We introduce FlowSymm, a novel architecture that combines (i) a group-action on divergence-free flows, (ii) a graph-attention encoder to learn feature-conditioned weights over these symmetry-preserving actions, and (iii) a lightweight Tikhonov refinement solved via implicit bilevel optimization. The method first anchors the given observation on a minimum-norm divergence-free completion. We then compute an orthonormal basis for all admissible group actions that leave the observed flows invariant and parameterize the valid solution subspace, which shows an Abelian group structure under vector addition. A stack of GATv2 layers then encodes the graph and its edge features into per-edge embeddings, which are pooled over the missing edges and produce per-basis attention weights. This attention-guided process selects a set of physics-aware group actions that preserve the observed flows. Finally, a scalar Tikhonov penalty refines the missing entries via a convex least-squares solver, with gradients propagated implicitly through Cholesky factorization. Across three real-world flow benchmarks (traffic, power, bike), FlowSymm substantially outperforms state-of-the-art baselines in RMSE, MAE and correlation metrics."},"_bibtex":{"value":"@inproceedings{\ndemirci2026flowsymm,\ntitle={FlowSymm: Physics{\\textendash}Aware, Symmetry{\\textendash}Preserving Graph Attention for Network Flow Completion},\nauthor={Ege Demirci and Francesco Bullo and Ananthram Swami and Ambuj Singh},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=vu1IEpdUQh}\n}"},"title":{"value":"FlowSymm: Physics–Aware, Symmetry–Preserving Graph Attention for Network Flow Completion"},"pdf":{"value":"/pdf/5f45aa1b88ec06fd6198f68eb0b3063fdd10f8ce.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"demirci|flowsymm_physicsaware_symmetrypreserving_graph_attention_for_network_flow_completion"},"authorids":{"value":["~Ege_Demirci1","~Francesco_Bullo1","~Ananthram_Swami1","~Ambuj_Singh1"]},"authors":{"value":["Ege Demirci","Francesco Bullo","Ananthram Swami","Ambuj Singh"]}},"version":2},{"content":{"summary":{"value":"This paper investigates whether modern large reasoning models, such as DEEPSEEK-R1 and its distilled variants, can effectively solve physics reasoning tasks without extensive prompt engineering or external tools. The authors benchmark these models on three datasets from SciBench (covering classical dynamics, thermodynamics, and fundamental physics) and compare their performance against more general chat-oriented LLMs.\n\nThe authors specifically explore - (1) Whether reasoning models can perform well intrinsically without heavy prompting; (2)  Whether few-shot prompt design still provides measurable benefits; (3) What underlying reasoning mechanisms distinguish reasoning models from standard chat models."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1.  Citations use inconsistent style\n2.  In the accuracy computation, you allow for a 5% tolerance - why is that?\n3.  \"Qwen architecture's superior symbolic processing capabilities\" - this is unclear."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.  Most existing studies on reasoning LLMs concentrate on mathematics, logic puzzles, or code synthesis. This paper’s focus on physics, thereby, broadening the empirical scope of reasoning model evaluation.\n2.  The evaluation pipeline (zero-shot vs. few-shot CoT) is well described, using publicly available datasets (SciBench). The inclusion of multiple Deepseek variants and comparison with baselines like GPT-4-Turbo adds breadth.\n3.  The paper includes appendices with prompt templates, parameter settings, and reproducibility statements, which are helpful for replication.\n4.  The contrast between symbolic derivation and numeric substitution reasoning styles is somewhat insightful and well illustrated."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  This is primarily an evaluation work and lacks novelty. The authors evaluate pre-existing reasoning models (Deepseek-R1 and distill variants) on a known benchmark (SciBench).\n2.  The “symbolic vs. numeric” observation, while intuitive, is anecdotal and not systematically analyzed or quantified.\n3.  The evaluation is restricted to SciBench and unimodal, text-only physics questions, excluding diagrams, visual reasoning, or multimodal tasks that are central to real-world physics understanding.\n4.  It is unclear whether Deepseek-R1’s pretraining might overlap with SciBench content leading to potential contamination.\n\nOverall, this submission adds minimal incremental value and lacks novelty and significant contribution required to extend the current state-of-the-art."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921882819,"tcdate":1761849633015,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10626/Reviewer_GwoV"],"signatures":["ICLR.cc/2026/Conference/Submission10626/Reviewer_GwoV"],"forum":"Gq73XUbtjb","number":4,"license":"CC BY 4.0","cdate":1761849633015,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10626/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921882819,"domain":"ICLR.cc/2026/Conference","replyto":"Gq73XUbtjb","id":"FfknSe97cz","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Large Language Model","Physics Reasoning","Model Evaluation"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Navigating the complexities of physics reasoning has long been a difficult task for Large Language Models (LLMs), requiring a synthesis of profound conceptual understanding and adept problem-solving techniques. In this study, we investigate the application of advanced instruction-tuned reasoning models, such as Deepseek-R1, to address a diverse spectrum of physics problems curated from the challenging SciBench benchmark. Our comprehensive experimental evaluation reveals the remarkable capabilities of reasoning models. Not only do they achieve state-of-the-art accuracy in answering intricate physics questions, but they also generate distinctive reasoning patterns that emphasize on symbolic derivation. Furthermore, our findings indicate that even for these highly sophisticated reasoning models, the strategic incorporation of few-shot prompting can still yield measurable improvements in overall accuracy, highlighting the potential for continued performance gains."},"_bibtex":{"value":"@misc{\ndan2025symbolic,\ntitle={Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning {LLM}s},\nauthor={Nifu Dan and Yujun Cai and Yiwei Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=Gq73XUbtjb}\n}"},"title":{"value":"Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs"},"pdf":{"value":"/pdf/de767fc41a7e66a7cf24f68c360544b2833a341e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"dan|symbolic_or_numerical_understanding_physics_problem_solving_in_reasoning_llms"},"authorids":{"value":["~Nifu_Dan1","~Yujun_Cai1","~Yiwei_Wang2"]},"authors":{"value":["Nifu Dan","Yujun Cai","Yiwei Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a framework for designing novel BCC/B2 superalloys by fine-tuning language models (LMs) using preference learning. The LMs are optimized via Direct Preference Optimization (DPO), uniquely employing a reward signal derived from multi-objective, physics-based feedback from thermodynamic simulations (Thermo-Calc). This work is the first to align LMs toward a practical engineering goal using physics-grounded feedback, moving beyond simple stability to optimize for complex, synthesizeable materials."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please refer to the Weaknesses section above."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is clearly written and well-motivated, convincingly arguing for a shift from optimizing simple stability to complex engineering utility.\n\n2. It demonstrates that preference learning is a promising pathway to achieve this, successfully aligning language models with physics-grounded, multi-objective design goals."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The study's core contribution, preference learning via DPO, yielded only modest gains over the SFT baseline. This method proved inconsistent, as it failed on one of the three test models (OLMo), which showed significant performance degradation. This undermines the claim of a successfully applied and robust preference learning framework.\n\n2. The paper only tested DPO and failed to explore advanced methods such as GRPO, which is widely used in training recent reasoning models. This is a significant omission, especially given the modest results of DPO."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762999983471,"tcdate":1761970783282,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20860/Reviewer_SWw1"],"signatures":["ICLR.cc/2026/Conference/Submission20860/Reviewer_SWw1"],"forum":"nEF9q1UmEZ","number":2,"license":"CC BY 4.0","cdate":1761970783282,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20860/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762999983471,"domain":"ICLR.cc/2026/Conference","replyto":"nEF9q1UmEZ","id":"PvKJcQAdDb","forumContent":{"TLDR":{"value":"We use preference learning to optimize local LMs to generate candidate compositions for BCC/B2 alloys"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Direct preference optimization","preference learning","materials science","alloys","language models"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"We apply preference learning to the task of language model generation of novel structural alloys. Where prior work focuses on generating stable inorganic crystals, our approach optimizes for the synthesizeability of a specific structural class: BCC/B2 superalloys, an underexplored family of materials with applications in extreme environments. Using three open-weight models (LLaMA-3.1, Gemma-2, and OLMo-2), we demonstrate that language models can be optimized for multiple design objectives using a single, unified reward signal through Direct Preference Optimization (DPO). Our reward signal is derived from thermodynamic phase calculations, offering a scientifically-grounded feedback for model tuning. To our knowledge, this is the first demonstration of preference-tuning a language model using physics-grounded feedback for targeted properties (in our case, BCC/B2 alloys). The resulting framework is general and adaptable to any design problem for which the design space is enumerable and simulation-based feedback is available."},"_bibtex":{"value":"@misc{\nghosh2025preference,\ntitle={Preference Learning from Physics-Based Feedback: Tuning Language Models to Design {BCC}/B2 Superalloys},\nauthor={Satanu Ghosh and Collin Holgate and Neal R Brodnik and Doug Downey and Samantha Daly and Tresa Pollock and Samuel Carton},\nyear={2025},\nurl={https://openreview.net/forum?id=nEF9q1UmEZ}\n}"},"title":{"value":"Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys"},"pdf":{"value":"/pdf/e7e19aee2f2570a808b1d3bdd813d60c2e51060f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ghosh|preference_learning_from_physicsbased_feedback_tuning_language_models_to_design_bccb2_superalloys"},"authorids":{"value":["~Satanu_Ghosh1","~Collin_Holgate1","~Neal_R_Brodnik1","~Doug_Downey1","~Samantha_Daly2","~Tresa_Pollock2","~Samuel_Carton1"]},"authors":{"value":["Satanu Ghosh","Collin Holgate","Neal R Brodnik","Doug Downey","Samantha Daly","Tresa Pollock","Samuel Carton"]}},"version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_4.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"feng|cache_allocation_in_cdn_an_evolutionary_game_generalized_particle_model"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Francis_C._M._Lau_0001:","https://dblp.org/search/pid/api?q=author:Daqi_Gao:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_4"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/FengLG09a,\n  author={Xiang Feng and Francis C. M. Lau and Daqi Gao},\n  title={Cache Allocation in CDN: An Evolutionary Game Generalized Particle Model},\n  year={2009},\n  cdate={1230768000000},\n  pages={1226-1237},\n  url={https://doi.org/10.1007/978-3-642-02469-6_4},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"Content distribution networks (CDNs) increasingly have been used to reduce the response times experienced by Internet users through placing surrogates close to the clients. This paper presents an object replacement approach based on an evolutionary game generalized particle model (G-GPM). We first propose a problem model for CDNs. The CDN model is then fit into a gravitational field. The origin servers and surrogates are regarded as two kinds of particles which are located in two force-fields. The cache allocation problem is thus transformed into the kinematics and dynamics of the particles in the annular and the round force-fields. The G-GPM approach is unique in four aspects: 1) direct viewing of individual and overall optimization; 2) parallel computing (lower time complexity); 3) multi-objective solution; and 4) being able to deal with some social interactions behaviors."},"title":{"value":"Cache Allocation in CDN: An Evolutionary Game Generalized Particle Model"},"authors":{"value":["Xiang Feng","Francis C. M. Lau","Daqi Gao"]}},"tmdate":1744105802815,"pdate":1230768000000,"tcdate":1744105521543,"writers":["~"],"signatures":["~Xiang_Feng2"],"forum":"mW58sAuGwh","license":"CC BY-SA 4.0","number":382267,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1744105802815,"domain":"DBLP.org","id":"mW58sAuGwh","version":2},{"content":{"summary":{"value":"The paper investigates causal reasoning over narratives in LLMs by probing two shortcuts: 1) ordering prior (events mentioned earlier are treated as causes) and 2) parametric/world-knowledge prior (typicality from pretraining). Using synthetic chains, synthetic general causal graphs, semi-synthetic (CauseNet chains verbalized by an LLM), and real-world (CauseNet sentences) narratives, the authors show that a) reverse order and atypical relations (vs the model’s prior) significantly degrade performance in causal narrative understanding and b) longer and more complex graphs (with forks/colliders) moderately degrade accuracy. They also report that extracting an estimated causal graph $G'$ and answering from the graph alone helps, but the benefit vanishes when the narrative is reintroduced."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- L148: Synthetic generation (3.1, Setting): How are events linked into $G$, exactly? How did you ensure generated narratives do not contradict pre-existing LLM knowledge, a confounding failure mode that you identified in later sections?\n- L151-154: 50/2500 checks is too small. What were the selection criteria? Are there plans to extend the correctness/validity checks of the labels beyond the synthetic case?\n- Accuracy/Consistency formulas and CIs: Please state the exact formulas, unit of aggregation, seed handling, and CI method.\n- $G'$ quality evaluation: What are $G'$ edge precision/recall/F1, GED($G$, $G'$)?\n- Narr+Graph fusion diagnostics: Did you verify the model actually uses $G'$ efficiently under Narr+Graph? Please provide position swaps (graph before or after narrative), instruction variants (i.e., \"prioritize the graph\", vs neutral), format variants.\n- L245–254 This paragraph is very confusing as to which results it is discussing. Point to the exact figure/panel and report numerical deltas.\n- L261–263 / L345 / L355: Were these datasets examined? What audit criteria were used for semi-synthetic and real narratives (fluency, faithfulness to $G$, label correctness)?\n- Are there automatic sanity checks (and broader human audits) to catch cause/effect inversions when generating or stitching narratives?\n- The semi-synthetic and real world cases seem quite similar, intuitively. One would expect very similar performance. Why, then, are there differences in Figure 4?\n\nMinor:\n- L197 / L310 “randomly”: Specify distributions or use more precise wording.\n- Figure 2 (right) contains complex graphs, including forks and colliders, yet their \"chain size\" is computed. How is this calculated? Graph size usually denotes the number of edges, but I believe you are calculating the graph order here. Please clarify."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper tackles the crucial issue of LLMs' causal reasoning ability, a topic of significant interest.\n- There is clear factorization of failure modes (ordering vs typicality)\n- Breadth across synthetic chains, synthetic general causal graphs, semi-synthetic chains, and real chains.\n- A consistent empirical trend is uncovered: Reversing the ordering hurts, atypical/counter-intuitive chains confuse models.\n- Graph-only prompting consistenly mitigates most of these biases. This could provide a useful prompting insight: explicit $G'$ extraction reduces shortcutting when used alone."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Metrics & aggregation left implicit. Accuracy = agreement with $G$ and Consistency = agreement with $G'$ are briefly stated but the paper lacks explicit formulas, units of aggregation, seed handling, and CI recipe in the main text.\n- Thin human validation; none beyond synthetic. Only 50/2500 synthetic narratives were audited, no inter-annotation agreement is reported, no audits were performed for semi-synthetic and real datasets, even though LLMs were used to modify them. This severely limits confidence in the labels.\n- Narrative vs graph interaction under-specified. Graph-only helps; Narr+Graph removes most of the Graph-only gain under reverse ordering. The fusion protocol isn't ablated (position of $G'$, format, instruction, relative length study). The mitigation claim (extract only $G'$) is fragile.\n- Complex-graph sampling is described in appendix only. There needs to be a brief summary of the fork/collider samplign and the chain-connect step in the main text.\n- Anti-Causal accuracy clustering around $\\approx 0.1-0.2$ for a binary task is alarming and warrants further investigation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942287653,"tcdate":1761003128194,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22578/Reviewer_BDoT"],"signatures":["ICLR.cc/2026/Conference/Submission22578/Reviewer_BDoT"],"forum":"GfVKK5sKit","number":1,"license":"CC BY 4.0","cdate":1761003128194,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22578/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942287653,"domain":"ICLR.cc/2026/Conference","replyto":"GfVKK5sKit","id":"SwsdxH5RMn","forumContent":{"TLDR":{"value":"In this paper, we examine the failure Modes of LLMs for causal reasoning on narratives and the unreliable shortcuts LLMs take to make causal inferences."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Causal Inference","Large Language Models","Reasoning","Narratives"]},"supplementary_material":{"value":"/attachment/bc04136f532b8cb4d21b3fafddf7721600bb8b68.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"The ability to robustly identify causal relationships is essential for autonomous decision-making and adaptation to novel scenarios. However, accurately inferring causal structure requires integrating both world knowledge and abstract logical reasoning. In this work, we investigate the interaction between these two capabilities through the representative task of causal reasoning over narratives. Through controlled synthetic, semi-synthetic and real-world experiments, we find that state-of-the-art large language models (LLMs) often rely on superficial heuristics—for example, inferring causality from event order or recalling memorized world knowledge without attending to context. Furthermore, we show that simple reformulations of the task can elicit more robust reasoning behavior. Our evaluation spans a range of causal structures, from linear chains to complex graphs involving colliders and forks. These findings uncover systematic patterns in how LLMs perform causal reasoning and lay the groundwork for developing methods that better align LLM behavior with principled causal inference."},"_bibtex":{"value":"@inproceedings{\nyamin2026llms,\ntitle={{LLM}s Struggle to Balance Reasoning and World Knowledge in Causal Narrative Understanding},\nauthor={Khurram Yamin and Shantanu Gupta and Gaurav Rohit Ghosal and Zachary Chase Lipton and Bryan Wilder},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=GfVKK5sKit}\n}"},"title":{"value":"LLMs Struggle to Balance Reasoning and World Knowledge in Causal Narrative Understanding"},"pdf":{"value":"/pdf/ecf5f53bf29b6fbdd19eac004f528e64010a7f26.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yamin|llms_struggle_to_balance_reasoning_and_world_knowledge_in_causal_narrative_understanding"},"authorids":{"value":["~Khurram_Yamin1","~Shantanu_Gupta2","~Gaurav_Rohit_Ghosal1","~Zachary_Chase_Lipton1","~Bryan_Wilder2"]},"authors":{"value":["Khurram Yamin","Shantanu Gupta","Gaurav Rohit Ghosal","Zachary Chase Lipton","Bryan Wilder"]}},"version":2},{"content":{"summary":{"value":"IntPhys 2 extends the original IntPhys benchmark with a photorealistic, Unreal Engine–based video corpus designed to assess four fundamental intuitive-physics principles—object permanence, immutability, spatio-temporal continuity, and solidity. The release comprises 1,416 clips organised into easy, medium, and hard subsets, plus a held-out test split; each scene is presented in “possible” and “impossible” variants to minimise reliance on low-level visual cues. Leading multimodal language–vision models (e.g., GPT-4o, Gemini series) and predictive video models (e.g., VideoMAEv2, V-JEPA) achieve roughly random guess accuracy, whereas human annotators reach approximately 96 %, underscoring the current gap between machine and human common-sense physical reasoning."},"dataset_code_accessibility":{"value":"Yes"},"code_of_conduct_acknowledgement":{"value":"Yes"},"confidence":{"value":3},"rating":{"value":4},"final_justification":{"value":"The author addresses most of my concern in the rebuttal stage. Hence, I recommend accepting this paper. However, the evaluation protocol of the VLMs on IntPhys2 is not very fair compared to the jepa-based architectures. The real bottleneck of the VLM on the dataset remain not very clear. Therefore, I choose keep my original score."},"limitations_weaknesses":{"value":"1. Realism vs. Physical Law Trade-off: IntPhys2 tests physical possibility, but it’s unclear if some failures reflect physical law violations or unrealistic scenarios. For example, in Figure 5, a suitcase placed inside a jail may follow physical laws, but as a human observer, I would still judge this event as unlikely to happen since it's abnormal to see a suitcase in the jail. Therefore, I encourage the authors to illustrate how the MLLM models reason when making these decisions, for example by using a common chain-of-thought (CoT) approach.\n\n2. Although IntPhys2 features more photorealistic backgrounds, some object configurations in the scenes remain unrealistic.\n\n3. The paper illustrates the model's poor generalization on IntPhys2. I suggest adding a baseline where a supervised binary classifier (e.g., a shallow CNN/ViT classifier) is trained directly on IntPhys2 for comparison. This would help clarify whether the challenge arises from the task’s complexity or from the models’ poor generalization."},"ethical_considerations":{"value":"No, there are no or only very minor ethics concerns"},"responsible_reviewing_acknowledgement":{"value":"Yes"},"ethical_comments":{"value":"The limitations of the work are discussed in the main paper. There is no societal impact of the work performed."},"additional_feedback":{"value":"1. In Table 2, the overall performance is noticeably lower than the performance on easy and medium videos, suggesting that the dataset may contain a large proportion of hard cases. It would be helpful if the authors could provide additional statistics on the distribution of easy, medium, and hard videos within the IntPhys2 dataset.\n\n2. How about the large video language model on the IntPhys2 dataset e.g., Video-LLaVA?"},"dataset_code_comments":{"value":"The benchmark has been open-sourced on the huggingface. The code to evaluate MLLMs and prediction based models on IntPhys2 are released too."},"strengths_contributions":{"value":"1. As one of the pioneering benchmarks for physical reasoning, IntPhys has become saturated with recent predictive models such as V-JEPA achieving near-ceiling performance. The introduction of IntPhys 2 provides the community with a more challenging benchmark, advancing the evaluation of models' physical understanding capabilities.\n\n2. Compared to IntPhys 1, the videos in IntPhys 2 feature more photorealistic environments and incorporate moving camera settings, making the dataset more challenging and better aligned with real-world scenarios."}},"parentInvitations":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Official_Review","nonreaders":[],"tmdate":1761794723091,"tcdate":1751459514492,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_vk7Z"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_vk7Z"],"forum":"Xpf5x3mLvn","number":2,"license":"CC BY 4.0","cdate":1751459514492,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Official_Review","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Official_Review2/-/Review_Revision"],"mdate":1761794723091,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","replyto":"Xpf5x3mLvn","id":"oMzjsdrfoZ","forumContent":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Video benchmark","Intuitive Physic","Evaluation"]},"supplementary_material":{"value":"/attachment/4932fdafe9cb6b17072f67b2caf246775bce6feb.zip"},"_bibtex":{"value":"@misc{\nbordes2026intphys,\ntitle={IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments},\nauthor={Florian Bordes and Quentin Garrido and Justine T Kao and Adina Williams and Michael Rabbat and Emmanuel Dupoux},\nyear={2026},\nurl={https://openreview.net/forum?id=Xpf5x3mLvn}\n}"},"title":{"value":"IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments"},"paperhash":{"value":"bordes|intphys_2_benchmarking_intuitive_physics_understanding_in_complex_synthetic_environments"},"dataset_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh.zip"},"TLDR":{"value":"We introduced a new intuitive physic benchmark IntPhys 2, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments"},"primary_area":{"value":"datasets_&_benchmarks_for_computer_vision"},"abstract":{"value":"We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to  macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50\\%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies."},"pdf":{"value":"/pdf/c29b03435c4329f0422348326d7ca2b5d606960a.pdf"},"croissant_file":{"value":"/attachment/54b155200f4140df0470921a91b0093300c98b6a.json"},"code_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh_code.zip"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"authorids":{"value":["~Florian_Bordes1","~Quentin_Garrido1","~Justine_T_Kao1","~Adina_Williams1","~Michael_Rabbat1","~Emmanuel_Dupoux1"]},"authors":{"value":["Florian Bordes","Quentin Garrido","Justine T Kao","Adina Williams","Michael Rabbat","Emmanuel Dupoux"]}},"version":2},{"content":{"summary":{"value":"Summary:   \nThis paper addresses the critical limitation of existing text-to-video (T2V) generation models—their inability to adhere to physical laws despite producing photorealistic content—by proposingPhyWorldBench, a comprehensive benchmark for evaluating physical realism in T2V models.  \n\nContributions:  \n（1）Development of a comprehensive physics benchmark: PhyWorldBench fills the gap of lacking holistic tools to evaluate physical realism in text-to-video models. Its hierarchical structure (three levels, 10 main categories, 50 subcategories) covers diverse physical scenarios—from basic motion and energy conservation to anti-physics scenarios—and includes 1,050 well-curated prompts with variations, enabling systematic testing of models’ physical reasoning capabilities.  \n（2）Creation of a zero-shot automatic evaluator: The proposed Context-Aware Prompt method resolves the limitations of traditional evaluation. By guiding large language models to explicitly assess AI-generated videos and use chain-of-thought reasoning, it achieves high accuracy in evaluating physical realism, providing a scalable and objective alternative to costly human evaluation.  \n（3）Extensive evaluation of state-of-the-art models: The paper conducts large-scale tests on 12 leading text-to-video models, generating 12,600 videos. This evaluation identifies key challenges models face—such as struggling with complex interactions and prioritizing cinematic aesthetics over physics—and quantifies performance gaps across different physics categories, offering clear directions for future model improvement."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Questions:   \n（1）Questions About CAP Evaluator’s Aesthetic Bias  \nYou note the Context-Aware Prompt evaluator prefers visually polished videos but provide no quantitative data on this bias. How often does this preference lead to misclassifying physically incorrect yet visually polished videos as physically plausible? Does a visually vivid video with physical violations score higher than a plain but physically correct one? Have you tested if revising CAP prompts to explicitly ignore aesthetics reduces this bias? Clarifying these points will confirm if CAP’s objectivity is compromised and if adjustments can fix it.  \n（2）The comparison with the methods of predecessors is not sufficient.  \nA comparison between the dataset and some previous related datasets, such as VideoREPA[1], WISA[2], NewtonGen[3], etc.?  \n（3）Niche Fundamental Physics Performance\nPhyWorldBench covers 10 main physics categories, but your experiments focus on common phenomena and lack breakdowns for niche subcategories like pressure-dependent phase changes or non-uniform heat transfer. Do you have data on model performance in these niche areas? If not, why were they excluded from detailed analysis? Understanding these gaps will help researchers target specific physical principles models struggle with. If the author's reply is strong and reasonable, I will consider increasing my score.  \n\n[1] VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models  \n[2] Wisa: World simulator assistant for physics-aware text-to-video generation  \n[3] NEWTONGEN: PHYSICS-CONSISTENT AND CONTROL-LABLE TEXT-TO-VIDEO GENERATION VIA NEURAL NEWTONIAN DYNAMICS"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Strengths:   \n（1）Originality  \nIt proposes a three-tier (Fundamental, Composite, Anti-Physics) hierarchical structure with 10 main physics categories, going beyond prior benchmarks by using Anti-Physics scenarios to test true physical understanding. The CAP evaluator creatively combines LLMs with domain constraints, separating aesthetics from physics via two-step reasoning to fix traditional evaluators’ inaccuracies. It also offers three prompt variants to test model adaptability, filling gaps of static-prompt benchmarks.  \n（2）Quality  \nIts benchmark is built via a rigorous three-stage process (literature/expert category definition, LLM-human prompt generation, expert validation) to ensure 1,050 prompts are diverse and accurate. CAP aligns with human evaluation (ROC-AUC: 80.3 for SA, 75.1 for PC) and outperforms baseline LLMs. Experiments cover 12 models (5 proprietary, 7 open-source) with 12,600 videos and systematic analyses to avoid cherry-picking.  \n（3）Clarity  \nIt follows a clear \"problem-solution-validation\" flow, with the introduction outlining existing benchmark limits and methodology linking components to specific issues. Technical details (CAP’s reasoning, \"Yes/No\" criteria) are explained in plain language without jargon. Consistent terminology and condensed key findings make it readable for physics and AI researchers.  \n（4）Significance  \nAs a standardized tool, its open benchmark and CAP provide universal metrics for tracking text-to-video physical realism. Prompt design insights (e.g., Physics-Enhanced Prompts boost PC) guide real-world uses like scientific visualization. It shifts evaluation focus to physical correctness, identifies model weaknesses, and reduces educational misinformation risks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses:   \n（1）CAP Evaluator’s Aesthetic Bias  \nThe Context-Aware Prompt (CAP) evaluator favors visually polished videos (e.g., smooth lighting, dynamic camera movement) over physically correct ones, which conflicts with the benchmark’s goal of prioritizing physical plausibility. However, the paper doesn’t quantify this bias or fix it. For example, a visually vivid but gravity-violating floating apple video might get a higher score than a plain yet physically correct one.  \n（2）Inadequate Niche Fundamental Physics Analysis  \nPhyWorldBench focuses on common physics (e.g., free fall) but ignores niche subcategories of fundamental physics, like pressure-dependent phase changes (water boiling at high altitude) or non-uniform heat transfer (a metal rod heating unevenly). It only reports broad category performance, hiding which specific principles models struggle with.  \n（3）No Long-Duration Video Validation  \nThe paper doesn’t specify video duration or test how performance scales with length. Physical inconsistencies (e.g., trajectory drift) worsen in longer videos, which are needed for real uses (e.g., education). Without this data, the benchmark’s real-world value is limited.  \n（4）Lack of Specialized Benchmark Comparisons  \nThe paper only compares to general physics benchmarks (e.g., VideoPhy) but not specialized ones like Morpheus (real-world experiment focus) or T2VPhysBench (first-principles physics). This hides PhyWorldBench’s unique value."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920345856,"tcdate":1761897841317,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8464/Reviewer_oKee"],"signatures":["ICLR.cc/2026/Conference/Submission8464/Reviewer_oKee"],"forum":"rlZeILv3fm","number":4,"license":"CC BY 4.0","cdate":1761897841317,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8464/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920345856,"domain":"ICLR.cc/2026/Conference","replyto":"rlZeILv3fm","id":"w9QOvqF58L","forumContent":{"venue":{"value":"ICLR 2026 Oral"},"TLDR":{"value":"Large-scale, multidimensional video generation for physics"},"keywords":{"value":["Video Generation","Video Evaluation"]},"supplementary_material":{"value":"/attachment/bd8cdaa60661c40d3e703b3dd09ff6689653ba72.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents $PhyWorldBench$\n, a comprehensive benchmark designed to evaluate video generation models based on their adherence to the laws of physics. The benchmark covers multiple levels of physical phenomena, ranging from fundamental principles like object motion and energy conservation to more complex scenarios involving rigid body interactions and human or animal motion. Additionally, we introduce a novel \"Anti-Physics\" category, where prompts intentionally violate real-world physics, enabling the assessment of whether models can follow such instructions while maintaining logical consistency. Besides large-scale human evaluation, we also design a simple yet effective method that could utilize current MLLM to evaluate the physics realism in a zero-shot fashion. We evaluate 10 state-of-the-art text-to-video generation models, including five open-source and five proprietary models, with a detailed comparison and analysis. we identify pivotal challenges models face in adhering to real-world physics. Through systematic testing of their outputs across 1,050 curated prompts—spanning fundamental, composite, and anti-physics scenarios—we identify pivotal challenges these models face in adhering to real-world physics. We then rigorously examine their performance on diverse physical phenomena with varying prompt types, deriving targeted recommendations for crafting prompts that enhance fidelity to physical principles."},"_bibtex":{"value":"@inproceedings{\ngu2026phyworldbench,\ntitle={\\$PhyWorldBench\\$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models},\nauthor={Jing Gu and Xian Liu and Yu Zeng and Ashwin Nagarajan and Fangrui Zhu and Daniel Hong and Yue Fan and Qianqi Yan and Kaiwen Zhou and Ming-Yu Liu and Xin Eric Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=rlZeILv3fm}\n}"},"title":{"value":"$PhyWorldBench$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models"},"pdf":{"value":"/pdf/6138b5d05836a6d9ee27ae5fc6f3bbbe3667ff02.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gu|phyworldbench_a_comprehensive_evaluation_of_physical_realism_in_texttovideo_models"},"authorids":{"value":["~Jing_Gu2","~Xian_Liu1","~Yu_Zeng1","~Ashwin_Nagarajan1","~Fangrui_Zhu1","~Daniel_Hong1","~Yue_Fan3","~Qianqi_Yan1","~Kaiwen_Zhou3","~Ming-Yu_Liu1","~Xin_Eric_Wang2"]},"authors":{"value":["Jing Gu","Xian Liu","Yu Zeng","Ashwin Nagarajan","Fangrui Zhu","Daniel Hong","Yue Fan","Qianqi Yan","Kaiwen Zhou","Ming-Yu Liu","Xin Eric Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a ODNN framework to disentangle the real physics and disturbances by leverage the orthogonality. The ODNN is evaluated across eight synthetic and real-world datasets and comparability with several baselines."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. In many situations, the physics and disturbances can not be separated by the assumed way mentioned in the paper. Moreover, even if some of them could be separated, there could be multiple contrast pairs. However, the paper only investigate 1 pair at a time. It could be a problem if scaling up to multiple pairs. Could you add one example showing it works for real world complex system, like weather forecast and explain how your method could scale to real world complex system?\n\n2. The assumption is too strong that the physics network only learns the physics while the disturbance network only learn the disturbances. The boundary of it is actually very ambiguous in the real world. For example, some noise/disturbances could be proportional to the actual state variables and make them have the similar trend. In that case, how will you leverage your work to separate the physics and the disturbances. \n\n3. Your comparison of the baseline models are somehow obsoleted. There are more recent algorithm about SINDy, Bayesian regression, symbolic learning. To name a few, [1][2][3]. Could you add comparisons to these SOTA model?\n\n[1] Physics-informed learning of governing equations from scarce data.\n\n[2] Bayesian spline learning for equation discovery of nonlinear dynamics with quantified uncertainty\n\n[3] Symbolic physics learner: Discovering governing equations via monte carlo tree search"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Originality: The novelty lies in proposing an alternative view of separating the disturbances and the physics of the system, while most of the existing work adopts adding regularization terms in the loss functions. \n\nQuality; This paper is rich of details and experiments results. Eight different datasets are investigated and the proposed framework performs the best in these datasets compared with other baseline models.\n\nClarity: In general clear illustrating the methodology and describe the experiments results. \n\nSignificance: It probably could be a valuable framework for physics systems identifications."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Clarity: Too much detail are put in the appendix. It makes the reviewer hard to understand the experiment details without cross-checking. \n\nThe assumption 1 is too strong, which introduced a very big limitation about the current work. In real world environments, the physics and disturbances cannot be just separated by the periodicity vs non-periodicity, monotonic vs non-monotonically and so on. \n\nThe experiment cases are all toy problem without real world complex systems like weather foreseeing, time series prediction in financial supermarket and so on."}},"nonreaders":[],"tmdate":1731428035212,"tcdate":1730685078460,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4095/Reviewer_HTby"],"signatures":["ICLR.cc/2025/Conference/Submission4095/Reviewer_HTby"],"forum":"ZujMVRn7Md","number":3,"license":"CC BY 4.0","cdate":1730685078460,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4095/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428035212,"domain":"ICLR.cc/2025/Conference","replyto":"ZujMVRn7Md","id":"VVpgIjTW9C","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["representation learning","physics identification","orthogonality"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Accurately identifying the underlying physical laws in complex systems is vital for effective control and interpretation. However, many systems are governed by a combination of known physical principles and unobservable or poorly understood components. Traditional model-based methods like Kalman filters and state-space models often rely on oversimplified assumptions, while modern data-driven approaches, such as physics-informed neural networks (PINNs), can suffer from overfitting or lack theoretical guarantees in recovering true physical dynamics. We propose the Orthogonal Deep Neural Network (ODNN) architecture to address these limitations. ODNN disentangles known physical components from unobservable or poorly understood components by imposing orthogonal constraints on the deep neural network. Unlike additive regularization methods, ODNN converts the physical constraints directly into the network structure, ensuring that the DNN focuses on capturing the unknown or complex dynamics without overfitting. This novel approach leverages both explicit orthogonality (e.g., zero inner product) and implicit orthogonality (e.g., contrasting convexity, periodicity, or symmetry) between physical laws and unknown components. Theoretically, we prove that ODNN provides strong guarantees for accurate system identification under mild orthogonality assumptions, building on the universal approximation theorem. Empirically, ODNN is evaluated across eight synthetic and real-world datasets, showcasing its ability to recover governing physical equations with high accuracy and interpretability. Our results demonstrate that ODNN offers significant advantages in terms of generalizability and robustness, making it a valuable framework for physics-based model identification in complex systems."},"_bibtex":{"value":"@misc{\nxiao2025orthogonal,\ntitle={Orthogonal Deep Neural Networks ({ODNN}): Uncovering Hidden Physics in Partially Observable Systems},\nauthor={CHENHAN XIAO and Yang Weng},\nyear={2025},\nurl={https://openreview.net/forum?id=ZujMVRn7Md}\n}"},"title":{"value":"Orthogonal Deep Neural Networks (ODNN): Uncovering Hidden Physics in Partially Observable Systems"},"pdf":{"value":"/pdf/a0db04dd97c2cd93dbed6db975b85ec6e11eff6b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"xiao|orthogonal_deep_neural_networks_odnn_uncovering_hidden_physics_in_partially_observable_systems"},"authorids":{"value":["~CHENHAN_XIAO1","~Yang_Weng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["CHENHAN XIAO","Yang Weng"]}},"version":2},{"content":{"summary":{"value":"This paper proposes APEX, a framework intended to augment LLMs with explicit physics-based foresight for task planning. It introduces a Perception–Graph–Language–Physics–Action pipeline that integrates a graph attention module and a physics simulator to provide physical rollouts as feedback to an LLM during decision making. The authors evaluate APEX on three benchmark domains—synthetic physics QA, Tetris planning, and dynamic obstacle avoidance—claiming substantial improvements over vanilla GPT-4o and VLM baselines."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"See above."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Interesting high-level motivation. Bridging symbolic reasoning in LLMs with physically grounded modeling is an important and timely goal.\n2. Attempt to unify physics reasoning and LLM planning. The modular architecture (graph → simulator → LLM → action) provides a readable system outline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Conceptual novelty is limited. The core idea—using a physics engine to simulate candidate actions and feeding the results back to an LLM—is conceptually straightforward and has appeared in prior “simulation-in-the-loop” or “world-model prompting” works (e.g., Mind’s Eye, PiLoT, PhysVLM). The proposed Perception–Graph–Language–Physics–Action paradigm is mostly a re-labeling of existing perception-simulation-planning loops in robotics; there is no theoretical or algorithmic advance beyond modular composition.\n2. Questionable experimental design and fairness. Benchmarks are non-standard. The “Physics Reasoning Benchmark,” “Tetris,” and “Dynamic Obstacle Avoidance” are all custom setups with unclear data availability or reproducibility.\n3. Paper is not well written. The main text could not fully present the results and the analysis. I suggest putting some result figures from appendix to the main text. Table format should be unified and the figures need further improvement for clarity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924233842,"tcdate":1760685103104,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13665/Reviewer_ToiG"],"signatures":["ICLR.cc/2026/Conference/Submission13665/Reviewer_ToiG"],"forum":"ROB3ALLKIX","number":2,"license":"CC BY 4.0","cdate":1760685103104,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13665/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924233842,"domain":"ICLR.cc/2026/Conference","replyto":"ROB3ALLKIX","id":"nhklx0ckmu","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We build a framework that helps LLMs quantify when a cat will collide with them and assess the physical outcomes of different escape routes (e.g., whether they’ll crash into a table) by feeding them physics-based simulations."},"keywords":{"value":["Physics-Enhanced LLMs","Graph-Based Perception","Task Planning","Predictive Simulation"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Large Language Models (LLMs) demonstrate strong reasoning and task planning capabilities but remain fundamentally limited in physical interaction modeling. Existing approaches integrate perception via Vision-Language Models (VLMs) or adaptive decision-making through Reinforcement Learning (RL), but they fail to capture dynamic object interactions or require task-specific training, limiting their real-world applicability.\nWe introduce APEX (Anticipatory Physics-Enhanced Execution), a framework that equips LLMs with physics-driven foresight for real-time task planning. APEX constructs structured graphs to identify and model the most relevant dynamic interactions in the environment, providing LLMs with explicit physical state updates. Simultaneously, APEX provides low-latency forward simulations of physically feasible actions, allowing LLMs to select optimal strategies based on predictive outcomes rather than static observations.\nWe evaluate APEX on three benchmarks designed to assess perception, prediction, and decision-making: (1) Physics Reasoning Benchmark, testing causal inference and object motion prediction; (2) Tetris, evaluating whether physics-informed prediction enhances decision-making performance in long-horizon planning tasks; (3) Dynamic Obstacle Avoidance, assessing the immediate integration of perception and action feasibility analysis. APEX significantly outperforms standard LLMs and VLM-based models, demonstrating the necessity of explicit physics reasoning for bridging the gap between language-based intelligence and real-world task execution."},"_bibtex":{"value":"@misc{\nhuang2025apex,\ntitle={{APEX}: Empowering {LLM}s with Physics-Based Task Planning for Real-time Insight},\nauthor={Wanjing Huang and Weixiang Yan and Zhen Zhang and Ambuj Singh},\nyear={2025},\nurl={https://openreview.net/forum?id=ROB3ALLKIX}\n}"},"title":{"value":"APEX: Empowering LLMs with Physics-Based Task Planning for Real-time Insight"},"pdf":{"value":"/pdf/0529797af939423f17dcaccea10909babb2a5c4d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|apex_empowering_llms_with_physicsbased_task_planning_for_realtime_insight"},"authorids":{"value":["~Wanjing_Huang1","~Weixiang_Yan1","~Zhen_Zhang16","~Ambuj_Singh1"]},"authors":{"value":["Wanjing Huang","Weixiang Yan","Zhen Zhang","Ambuj Singh"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"venue":{"value":"WM PAI Workshop Poster"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"pdf":{"value":"/pdf/f5caff8ab7e1d48a05e07dfe8ac1e225d6a89993.pdf"},"keywords":{"value":["Intuitive Physics; Video Foundation Models; Frozen-Feature Probing"]},"venueid":{"value":"NeurIPS.cc/2026/Workshop/WM_PAI"},"paperhash":{"value":"punzo|how_do_video_foundation_models_encode_intuitive_physics_probing_across_pretraining_paradigms"},"abstract":{"value":"Do video foundation models encode a usable sense of how the physical world works, and does intuitive physics emerge naturally from large-scale video pretraining? In this work, we study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how its accessibility varies across pretraining paradigms, layers, backbone variants and probe types. Using frozen-feature probing on IntPhys2 and Minimal Video Pairs (MVP), we compare predictive joint-embedding models (V-JEPA), masked reconstruction models (VideoMAE), and diffusion-based video generators (LTX-Video). We find that across linear, MLP, and temporal attentive probes, physics-relevant information is generally weakest in early layers and most accessible at intermediate to late depth. On MVP, temporal attentive probes substantially outperform linear readouts, with V-JEPA achieving 95.0% pair consistency. On IntPhys2, linear probes recover more signal, and VideoMAE attains the strongest temporal attentive result at 73.9% violation-of-expectation accuracy. On both benchmarks, attentive probe performance on LTX consistently underperforms compared to the other two pretraining paradigms. Together, these results suggest that intuitive-physics is broadly decodable from video foundation models, but it is not represented as a single, uniform capability. Its accessibility—and especially its temporal grounding—depends much more on the benchmark and readout than on a simple hierarchy of pretraining objectives or model scale."},"_bibtex":{"value":"@inproceedings{\npunzo2026how,\ntitle={How Do Video Foundation Models Encode Intuitive Physics? Probing Across Pretraining Paradigms},\nauthor={Samuele Punzo and Niccol{\\`o} Caselli and Ippokratis Pantelidis and Francesco Massafra and Salvatore Lo Sardo and Mohammadreza Salehi},\nbooktitle={World Models in Physical AI Workshop},\nyear={2026},\nurl={https://openreview.net/forum?id=MSoPQ5R35q}\n}"},"title":{"value":"How Do Video Foundation Models Encode Intuitive Physics? Probing Across Pretraining Paradigms"},"authors":{"value":[{"username":"~Samuele_Punzo1","fullname":"Samuele Punzo","institutions":[{"name":"University of Amsterdam","domain":"uva.nl","country":"NL"}]},{"username":"~Niccolò_Caselli1","fullname":"Niccolò Caselli","institutions":[{"name":"University of Amsterdam","domain":"uva.nl","country":"NL"}]},{"username":"~Ippokratis_Pantelidis2","fullname":"Ippokratis Pantelidis","institutions":[{"name":"University of Amsterdam","domain":"uva.nl","country":"NL"}]},{"username":"~Francesco_Massafra1","fullname":"Francesco Massafra","institutions":[{"name":"University of Amsterdam","domain":"uva.nl","country":"NL"}]},{"username":"~Salvatore_Lo_Sardo2","fullname":"Salvatore Lo Sardo","institutions":[{"name":"University of Amsterdam","domain":"uva.nl","country":"NL"}]},{"username":"~Mohammadreza_Salehi2","fullname":"Mohammadreza Salehi","institutions":[{"name":"Samsung","domain":"samsung.com","country":"GB"}]}]}},"tmdate":1791063577081,"pdate":1791063575845,"tcdate":1787996661709,"writers":["NeurIPS.cc/2026/Workshop/WM_PAI","NeurIPS.cc/2026/Workshop/WM_PAI/Submission31/Authors"],"signatures":["NeurIPS.cc/2026/Workshop/WM_PAI/Submission31/Authors"],"forum":"MSoPQ5R35q","license":"CC BY 4.0","number":31,"cdate":1787996661709,"readers":["everyone"],"invitations":["NeurIPS.cc/2026/Workshop/WM_PAI/-/Submission","NeurIPS.cc/2026/Workshop/WM_PAI/-/Submission_Change_Before_Bidding","NeurIPS.cc/2026/Workshop/WM_PAI/-/Submission_Change_Before_Reviewing","NeurIPS.cc/2026/Workshop/WM_PAI/-/Submission_Release"],"mdate":1791063577081,"odate":1791063575845,"domain":"NeurIPS.cc/2026/Workshop/WM_PAI","id":"MSoPQ5R35q","version":2},{"content":{"summary":{"value":"The paper does a good job of establishing a theoretical rationale behind using synthetic data for LLM post-training processes like SFT. The modeling of synthetic data, focusing on distributional aspects is elaborate and properly justified as well. It establishes an intuitive information-theoretic based upper bound on the generalization error for an under-aligned LLM fine-tuned on synthetic data, and also delineates how this upper bound compares to the one when the LLM is tuned on just the real world samples (anchor data) - showcasing how the former proves to be a more stringent upper bound, leading to better generalization capabilities. Definitely helps bridge a gaping gap in LLM research!"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Here are suggested experiments that could help validate the paper's hypotheses and theoretical frameworks:\n1. LLM-based Validation Experiments :\n     - ⁠Compare different sizes of anchor data and their corresponding synthetic data generation\n     - ⁠Measure the relationship between model size and synthetic data quality\n     - ⁠Test various prompting strategies and their impact on synthetic data diversity\n     - Track information gain across different LLM architectures \n2. Practical Application Questions:\n    - How can practitioners use your theoretical bounds to improve their synthetic data generation process?\n    - What specific guidance would you give for prompt engineering based on your theoretical findings?\n3. Validation Questions:\n    - How would your bounds change with different LLM architectures or sizes?\n    -  ⁠What metrics would you recommend for measuring the quality of synthetic data in practice?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The theoretical modeling of synthetic data and how it interacts with the output distribution of the LLM is intuitive and complete. Great job there!\n2. The math behind generalization error bounds is solid.\n3. The connection to the classical information bottleneck theory is novel and very intuitive, clearly justifies why synthetic data helps better performance on downstream task generalization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper presents good theoretical foundations of \"why\" it is important to use synthetic data in post-training alignment and \"how\" it aids better generalization capabilities of the model. But it fails to address some really important \"hows\" that it promises in the Introduction section :\n   - How does this study help in developing tailored synthetic data generation - that more effectively address specific gaps in training data, thereby enhancing the overall performance and generalization capabilities of large language models\n   - How does this translate to better post-training practices involving synthetic datasets, from a practical perspective.\n2. The assumption that the transformation function φT is reversible (Section 3.2) is unclear for real LLM prompting scenarios, as task-to-prompt relationships are often many-to-many.A better delineation of this point would be helpful.\n3. ⁠The relationship between \"synthetic factors\" and actual LLM generation processes is not well defined.\n4. The connection between the theoretical \"information gain\" concept and practical improvements in LLM performance is not clearly established. \n5. I think its very crucial to address how the \"diversity\" and \"quality\" in synthetic data samples are a big part of why task generalization actually works when models are post trained on synthetic data. These should be adequately quantified and incorporated in the upper bound equations as well, for a more complete picture."}},"nonreaders":[],"tmdate":1731428539628,"tcdate":1730715963646,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5944/Reviewer_sWKh"],"signatures":["ICLR.cc/2025/Conference/Submission5944/Reviewer_sWKh"],"forum":"UxkznlcnHf","number":3,"license":"CC BY 4.0","cdate":1730715963646,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5944/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428539628,"domain":"ICLR.cc/2025/Conference","replyto":"UxkznlcnHf","id":"oPDgivgRuA","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"This paper explores the critical role of synthetic data in enhancing the post-training performance of large language models (LLMs) from a novel reverse-bottleneck perspective."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["large language models; synthetic data; information bottleneck"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Synthetic data has become a pivotal resource in post-training tasks for large language models (LLMs) due to the scarcity of high-quality, specific data. While various methods have been developed to generate synthetic data, there remains a discernible gap between the practical effects of synthetic data and our theoretical comprehension. To address this challenge, we commence by presenting a detailed modeling of the prevalent synthetic data generation process. Building upon this modeling, we demonstrate that the generalization capability of the post-trained model is critically determined by the information gain derived from the generative model, as analyzed from a novel reverse-bottleneck perspective. Moreover, we introduce the concept of Generalization Gain via Mutual Information (GGMI) and elucidate the relationship between generalization gain and information gain. This analysis serves as a theoretical foundation for synthetic data generation and further highlights its connection with the generalization capability of post-trained models, offering an understanding about the design of synthetic data generation techniques and the optimization of the post-training process. We open-source our code at https://github.com/ZyGan1999/Towards-a-Theoretical-Understanding-of-Synthetic-Data-in-LLM-Post-Training."},"_bibtex":{"value":"@inproceedings{\ngan2025towards,\ntitle={Towards a Theoretical Understanding of Synthetic Data in {LLM} Post-Training: A Reverse-Bottleneck Perspective},\nauthor={Zeyu Gan and Yong Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=UxkznlcnHf}\n}"},"title":{"value":"Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective"},"pdf":{"value":"/pdf/127d76775eb769452b3e1f3cffc5359d9e886a32.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"gan|towards_a_theoretical_understanding_of_synthetic_data_in_llm_posttraining_a_reversebottleneck_perspective"},"authorids":{"value":["~Zeyu_Gan1","~Yong_Liu7"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyu Gan","Yong Liu"]}},"version":2},{"content":{"summary":{"value":"The authors train their own video diffusion models on synthetic video models and run thorough analysis on how the physics based video generation."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"* In what real world cases would the data be significantly outside of the distribution of the training data. Should the solution not be scaling the number of parameters or the number of data, but rather the diversity of the data then?\n\n* What aspects of the behavior you observation are purely artifacts of the latent encoder / decoder?\n\n* Why did you not finetune any existing video models on these tasks? Surely this task is entirely OOD from the training task since it's a synthetic environment. Furthermore, if as a researcher your goal is to train the best video world model possible, why would you not start with a pretrained model. If you are confident that these problems cannot be solved purely by scaling, than pretrained video models on real world images should not generalize to this simple benchmark and should observe similar biases as in this paper. If you are making a claim that this behavior can be observed on all video models, than you should be able to evaluate this lack of generalizations on existing pre-trained models right?\n\n\n* We design datasets which delibrately leave out some latent values, i.e. velocity. After training, we test model’s prediction on both seen and unseen scenarios. We mainly focus on uniform motion and collision processes\nFrom an optimization point of view, what sparsity level is needed for the model to be able to linearly extrapolate then? After all, a diffusion model must also know what NOT to generate. There is surely some level sparsity for which the video model will generalize? What level of sparsity is it?\n\n* \"color > size > velocity > shape\" how much of this is simply affected by this explicit instantiation of the video VAE? What about other ones like Stable Video Diffusions? Or more recently released video model like COGVideo's?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"* The paper train many video diffusion models from scratch on small physical datasets across model parameter sizes. \n* They have thorough evaluation of extrapolation behavior\n* The paper attempts to tackle a very complex problem, by simplifying it to a synthetic data setting.\n* The analysis between in distribution and OOD seems useful, and this distinction could inform the creation of future video models, particularly the experiments in 5.4"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Instead of training a \"small\" video model from scratch, why not try finetuning SOTA models on these video datasets? One issue with this analysis it supposes that there is not a minimum threshold for the number of parameters needed for a useful diffusion video model or for generalization to hold. I would not be surprised if these models generalized when simply having more parameters and train on more data, even out of domain data. Finetuning a model like SVD should be doable on a similar level of compute.\n\n* The rendering of the synthetic examples are overly simplistic and do not have imaging artifacts that video models could exploit in real world use cases to generalize, like motion blur.\n\n* Many of these reasoning weakness of generative models have been brought before in other domains. Such as Arc-AGE Challenge - \"On the Measure of Intelligence\" by François Chollet.\n\n*  \"For example, in Figure 10, it is difficult to determine if a ball can pass through a gap based on vision alone when the size difference\nis at the pixel level, leading to visually plausible but incorrect results. Similarly, visual ambiguity in a ball’s horizontal position relative to a block can result in different outcomes. These findings suggest that relying solely on visual representations, may be inadequate for accurate physics modeling.\" Pixel level differences will be obliterated by the VAE encoder. It is a compressive architecture, information will be lost to train the model more cheaply. Pixel level criteria is not a motivating example as a result. If you want pixel level accuracy, you need pixel level diffusion models, not one trained on latent. I would disregard these experiments or rewrite this entire section 5.5 as this is a fundamental issue with the model architecture, and reveals no new information. Can you at least verify that the VAE is able to reconstruct different images with pixel level differences? I think you will find that it will not, as Stable Diffusion's VAE architecture cannot. The number of channels is way too limiting which is why it's been massively increased in more recent models like BlackForrest's FLUX.\n\n* \" This ranking could explain why current video generation models often struggle with maintaining object consistency\" Please support with evidence? Or at least specify with which video models?\n\n\n* Our in-depth analysis suggests that video model generalization relies more on referencing similar training examples rather than learning universal rules.\nAll generative models exhibit this behavior, and it is well known. Similar observations can be made on small language models, but they usually generalize way better when scaled up in terms of compute. What is the new observation here with respect to physics?\n\n\nRather a better way to frame this paper might be what role do VAE reconstruction issues prevent us from using existing video world models for physics in critical settings? How do the video VAE's and the diffusion model prioritize shape color and velocity? The main issue here is that some of the experiments are clearly demonstrating architectural failures of the VAEs and the authors are attempting to generalize it to all video models, which is an massive overclaim. \n\nThis paper has useful experimental and scientific data, but it needs to be rewritten to support the claims in the paper. Furthermore, it could massively benefit by examining which of these issues are coming from the VAE and which are coming from the latent video diffusion.\n\nClaims are made that scaling cannot solve these physics problem. The evidence this paper shows that scaling parameters and perhaps date in the video diffusion model is ineffective if the VAE removes important information for the video model."}},"nonreaders":[],"tmdate":1731427313191,"tcdate":1730663711328,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission759/Reviewer_ym7u"],"signatures":["ICLR.cc/2025/Conference/Submission759/Reviewer_ym7u"],"forum":"ZyLkNVHBZF","number":3,"license":"CC BY 4.0","cdate":1730663711328,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission759/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427313191,"domain":"ICLR.cc/2025/Conference","replyto":"ZyLkNVHBZF","id":"4DOJ51QmAi","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"we conduct systematic experiments to investigate \"how far is video generation model from world model\" from the physical law perspetive."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video generation","diffusion model","world model"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws. \nHowever, the ability of video generation models to discover such laws purely from visual data without human priors can be questioned.\nA world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios.\nIn this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization.\nWe developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws.\nThis provides unlimited supply of data for large-scale experimentation, and enables quantitative evaluation for the law in generated videos. \nWe trained diffusion-based video generation models to predict object movements based on initial frames.\nOur scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios.\nFurther experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit ``case-based'' generalization behavior, \\textit{i.e.}, mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color $>$ size $>$ velocity $>$ shape.\nOur study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws, despite its role in Sora's broader success."},"_bibtex":{"value":"@misc{\nkang2025how,\ntitle={How Far Is Video Generation from World Model: A Physical Law Perspective},\nauthor={Bingyi Kang and Yang Yue and Rui Lu and Zhijie Lin and Yang Zhao and Kaixin Wang and Gao Huang and Jiashi Feng},\nyear={2025},\nurl={https://openreview.net/forum?id=ZyLkNVHBZF}\n}"},"title":{"value":"How Far Is Video Generation from World Model: A Physical Law Perspective"},"pdf":{"value":"/pdf/61786c392aeb1743704ab9483f7f67d63ce837d2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"kang|how_far_is_video_generation_from_world_model_a_physical_law_perspective"},"authorids":{"value":["~Bingyi_Kang1","~Yang_Yue1","~Rui_Lu2","~Zhijie_Lin1","~Yang_Zhao14","~Kaixin_Wang1","~Gao_Huang1","~Jiashi_Feng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Bingyi Kang","Yang Yue","Rui Lu","Zhijie Lin","Yang Zhao","Kaixin Wang","Gao Huang","Jiashi Feng"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a PINN framework for estimating cerebral blood flow (CBF) and arterial transit time (ATT) from synthetic ASL data. While the topic is relevant, the implementation falls short of expectations for a physics-informed model: the physical constraint is enforced only through derivative matching rather than solving a full PDE, and the architecture separates data fitting and physics modeling without justification. Moreover, the absence of a hyperparameter to balance the hybrid loss raises concerns about the stability and generality of the approach. The study is limited to synthetic 2D data with no validation on real images. Overall, the contribution is too preliminary and methodologically weak to be accepted in its current form."},"strengths":{"value":"The topic is relevant, addressing robustness issues in ASL quantification under noisy conditions.\nThe use of synthetic data from a controlled simulator provides a reproducible benchmark.\nResults indicate improvements over a regularized nonlinear least squares (NLLS) baseline, especially for ATT estimation under low SNR."},"weaknesses":{"value":"The physical constraint is only weakly enforced via derivative matching rather than solving a full PDE or ODE, limiting the \"physics-informed\" nature of the approach. A more rigorous physics enforcement would be expected in a PINN framework.\nThe separation between a data-fitting and a physics-based network complicates the method without clear benefit. No justification is provided for not using a single network with a global hybrid loss.\nThe hybrid loss combines data and physics terms, but no balancing hyperparameter is introduced or discussed. This is a major omission: tuning the relative weight between data and physics is standard in PINNs to ensure convergence and stability.\nThe study is restricted to synthetic 2D slices with no experiments on real (in vivo) data or on full 3D volumes. As a result, the clinical relevance and generalizability of the approach remain highly speculative."},"confidence":{"value":4},"rating":{"value":3}},"parentInvitations":"MIDL.io/2025/Short_Papers/-/Official_Review","nonreaders":[],"tmdate":1746120023972,"tcdate":1745834495596,"writers":["MIDL.io/2025/Short_Papers","MIDL.io/2025/Short_Papers/Submission73/Reviewer_2J7a"],"signatures":["MIDL.io/2025/Short_Papers/Submission73/Reviewer_2J7a"],"forum":"fKxgUmAZx6","number":1,"license":"CC BY 4.0","cdate":1745834495596,"readers":["everyone"],"invitations":["MIDL.io/2025/Short_Papers/Submission73/-/Official_Review","MIDL.io/2025/Short_Papers/-/Edit"],"mdate":1746120023972,"domain":"MIDL.io/2025/Short_Papers","replyto":"fKxgUmAZx6","id":"NQU08q0VIT","forumContent":{"TLDR":{"value":"This study presents a physics-informed neural network approach for estimating cerebral perfusion parameters—CBF and ATT—from simulated time-encoded ASL MRI data, demonstrating improved robustness to noise."},"venue":{"value":"MIDL 2025 - Short Papers"},"pdf":{"value":"/pdf/1045de5afee8d2dbbfd171b7012b83e776b1549d.pdf"},"keywords":{"value":["Arterial Spin Labeling","Physics-Informed Neural Network","Hadamard Encoding","Cerebral perfusion"]},"venueid":{"value":"MIDL.io/2025/Short_Papers"},"paperhash":{"value":"giupponi|physicsinformed_neural_network_for_quantifying_timeencoded_arterial_spin_labeling_a_simulation_study"},"authorids":{"value":["~Alessandro_Giupponi1","chiara.davilla@studenti.unipd.it","mattia.veronese@unipd.it","~Marco_Castellaro1"]},"abstract":{"value":"Arterial Spin Labeling (ASL) MRI enables non-invasive quantification of cerebral perfusion. Hadamard time-encoding improves acquisition efficiency and allows the simultaneous estimation of cerebral blood flow (CBF) and arterial transit time (ATT) via the Buxton model. Physics-informed neural networks (PINNs) integrate physical laws into neural networks, improving parameter estimation under noisy and sparse data conditions. We propose a two-stage PINN framework trained on synthetic ASL data from the Boston ASL Template and Simulator. Leveraging coupled neural networks and differential equation constraints, our method produces smoother and more robust CBF and ATT maps compared to regularized nonlinear least squares (NLLS), demonstrating its potential for clinical ASL quantification. While this work focuses on simulation data, it represents a first step toward extending such models to in vivo applications using a similar architecture."},"_bibtex":{"value":"@inproceedings{\ngiupponi2025physicsinformed,\ntitle={Physics-Informed Neural Network for Quantifying Time-Encoded Arterial Spin Labeling: A Simulation Study},\nauthor={Alessandro Giupponi and Chiara Da Villa and Mattia Veronese and Marco Castellaro},\nbooktitle={Medical Imaging with Deep Learning - Short Papers},\nyear={2025},\nurl={https://openreview.net/forum?id=fKxgUmAZx6}\n}"},"title":{"value":"Physics-Informed Neural Network for Quantifying Time-Encoded Arterial Spin Labeling: A Simulation Study"},"authors":{"value":["Alessandro Giupponi","Chiara Da Villa","Mattia Veronese","Marco Castellaro"]}},"version":2},{"content":{"summary":{"value":"The study introduces AirPhyNet, a physics-guided neural network designed for enhanced air quality prediction. This method incorporates fundamental physics principles into the network architecture, improving predictive performance and interpretability. For this, it draws from existing literature in physics guided ML and neural ODEs. Tests on real-world data showcase its potential to improve over existing methods."},"presentation":{"value":"4 excellent"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"- Putting together multiple complex concepts and methods is indeed a difficult task and requires a thorough understanding of physical dynamics and deep learning.\n\n- The paper addresses a significant and timely problem.\n\n- The narrative is clear and accessible.\n\n- The case study illustrates some of the physics that the model captures."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- My main concern with this work is the lack of contributions to hybrid AI or AI in general. Authors did not identify any technical gaps in our current hybrid AI methods. Instead, authors take what other researchers have developed for physics-guided ML in a variety of domains (e.g., physics) and use them for air quality prediction. Therefore, the method appears to be a combination of multiple well known methods with some developments in how to incorporate the specific physics priors for air quality priors. The air quality priors are just new equations and do not pose a significant technical challenge. Therefore, I do not think this is not a significant contribution for ICLR's research track. Perhaps the paper's contribution is better suited to a domain journal or the applied track of an AI conference.\n\n- Liang et al. 2023 (cited by authors in experimental setup) performed experiments in 342 cities in China and data appears publicly available. However, authors of this paper performed experiments only in 2 cities."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"- What is the reason for selecting so few cities?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700776536399,"tcdate":1699133532337,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2973/Reviewer_pK4K"],"signatures":["ICLR.cc/2024/Conference/Submission2973/Reviewer_pK4K"],"forum":"JW3jTjaaAB","number":4,"license":"CC BY 4.0","cdate":1699133532337,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2973/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700776536399,"domain":"ICLR.cc/2024/Conference","replyto":"JW3jTjaaAB","id":"0BILgPRjmj","forumContent":{"TLDR":{"value":"AirPhyNet is a physics-guided deep learning framework for air quality prediction. It shows superior performance in lead times upto 72-hours especially in sparse data scenarios while generating forecasts with a real physical meaning."},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["air quality prediction","physics-informed","spatiotemporal-learning","interpretability"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Air quality prediction and modelling plays a pivotal role in public health and environment management, for individuals and authorities to make informed decisions. Although traditional data-driven models have shown promise in this domain, their long-term prediction accuracy can be limited, especially in scenarios with sparse or incomplete data and they often rely on black-box deep learning structures that lack solid physical foundation leading to reduced transparency and interpretability in predictions. To address these limitations, this paper presents a novel approach named Physics guided Neural Network for Air Quality Prediction (AirPhyNet). Specifically, we leverage two well-established physics principles of air particle movement (diffusion and advection) by representing them as differential equation networks. Then, we utilize a graph structure to integrate physics knowledge into a neural network architecture and exploit latent representations to capture spatio-temporal relationships within the air quality data. Experiments on two real-world benchmark datasets demonstrate that AirPhyNet outperforms state-of-the-art models for different testing scenarios including different lead time (24h, 48h, 72h), sparse data and sudden change prediction, achieving reduction in prediction errors up to 10\\%. Moreover, a case study further validates that our model captures underlying physical processes of particle movement and generates accurate predictions with real physical meaning. The code is available at: https://github.com/kethmih/AirPhyNet"},"_bibtex":{"value":"@inproceedings{\nhettige2024airphynet,\ntitle={AirPhyNet: Harnessing Physics-Guided Neural Networks for Air Quality Prediction},\nauthor={Kethmi Hirushini Hettige and Jiahao Ji and Shili Xiang and Cheng Long and Gao Cong and Jingyuan Wang},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=JW3jTjaaAB}\n}"},"title":{"value":"AirPhyNet: Harnessing Physics-Guided Neural Networks for Air Quality Prediction"},"pdf":{"value":"/pdf/5d2ef57196b7e029a27fe09b5352d94b4adb75bf.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"hettige|airphynet_harnessing_physicsguided_neural_networks_for_air_quality_prediction"},"authorids":{"value":["~Kethmi_Hirushini_Hettige1","~Jiahao_Ji1","~Shili_Xiang1","~Cheng_Long1","~Gao_Cong1","~Jingyuan_Wang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kethmi Hirushini Hettige","Jiahao Ji","Shili Xiang","Cheng Long","Gao Cong","Jingyuan Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a novel approach called Physics-Aware Spatiotemporal Causal Graph Network (P-STCGN) for integrating physical equations into spatiotemporal models. The idea is to leverage causality to capture the fundamental causal relations present in physics dynamics. The proposed approach uses a causal module to learn causal weights from past observations to current observations and a forecasting module to perform predictions guided by cause-effect relations. Evaluations conducted on synthetic as well as real-world climate datasets demonstrate the superior performance for the proposed method."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The Integration of Physics Knowledge is quite innovative.\n2. The paper provides an extensive evaluation of the proposed method on different datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses/questions\n1.\tHow does the model perform when the prior physics knowledge is ambiguous or not well established? How to verify the accuracy of the physics knowledge being integrated?\n2.\tCan the authors elaborate on why the model is able to handle the noisy data?\n3.\tBesides the climate-related application, how easy it is to extend the model to other domains?"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"See weaknesses above."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636693970,"tcdate":1698715715623,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6312/Reviewer_B5jF"],"signatures":["ICLR.cc/2024/Conference/Submission6312/Reviewer_B5jF"],"forum":"2uHTuvDkLZ","number":3,"license":"CC BY 4.0","cdate":1698715715623,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6312/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636693970,"domain":"ICLR.cc/2024/Conference","replyto":"2uHTuvDkLZ","id":"aAMPG7Jb8b","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["physics-informed deep learning; causal learning; spatiotemporal learning"]},"supplementary_material":{"value":"/attachment/5873a933b20dc180f679318946973106e5a5d0fa.pdf"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Interpretable physics equations are widely recognized as valuable inductive biases for constructing robust spatiotemporal models. To harness these valuable pieces of knowledge, existing approaches often presuppose access to the exact underlying equations. However, such an assumption usually doesn't hold, especially in the context of real-world observations. Conversely, causality systematically captures the fundamental causal relations across space and time that are intrinsically present in physics dynamics. Nevertheless, causality is often ignored as a means of integrating prior physics knowledge. In this work, we propose a novel approach that effectively captures and leverages causality to integrate physics equations into spatiotemporal models, without assuming access to precise physics principles. \nSpecifically, we introduce a physics-aware spatiotemporal causal graph network (P-stCGN). Causal relationships are analytically derived from prior physics knowledge and serve as physics-aware causality labels. A causal module is introduced to learn causal weights from spatially close and temporally past observations to current observations via semi-supervised learning. Given the learned causal structure, a forecasting module is introduced to perform predictions guided by the cause-effect relations. Extensive experiments on time series data show that our semi-supervised causal learning approach is robust with noisy and limited data. Furthermore, our evaluations on real-world graph signals demonstrate superior forecasting performance, achieved by utilizing prior physics knowledge from a causal perspective."},"_bibtex":{"value":"@misc{\nseo2024physicsaware,\ntitle={Physics-aware Causal Graph Network for Spatiotemporal Modeling},\nauthor={Sungyong Seo and Zijun Cui and Sam Griesemer and Joshua Hikida and Yan Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=2uHTuvDkLZ}\n}"},"title":{"value":"Physics-aware Causal Graph Network for Spatiotemporal Modeling"},"pdf":{"value":"/pdf/b680101f5e19ac6f8f83c61e3ae0f38cbe0cd7bc.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"seo|physicsaware_causal_graph_network_for_spatiotemporal_modeling"},"authorids":{"value":["~Sungyong_Seo1","~Zijun_Cui1","~Sam_Griesemer1","joshua.hikida@gmail.com","~Yan_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sungyong Seo","Zijun Cui","Sam Griesemer","Joshua Hikida","Yan Liu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a framework called LLMPhy, which combines large language models (LLMs) with physics engines to tackle complex physical reasoning tasks. The proposed method is tested on a new dataset, TraySim, where the model predicts object stability and interactions on a tray after an impact. LLMPhy operates in two phases: parameter estimation and simulation, with the LLM synthesizing hypotheses and the physics engine verifying them in a feedback loop. \n\nKey Strengths: \n\nTraySim Dataset: A tailored benchmark for multi-object physical interactions.\n\nZero-Shot Reasoning: Achieves notable improvements over Bayesian optimization on complex tasks. \n\nLimitations: \n\nLimited Parameter Scope: Currently models only four physical attributes, reducing generalizability. \n\nNo Real-World Testing: Results are based solely on simulations, lacking validation in physical settings."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"(1) how would LLMPhy perform when scaled to more complex environments with additional physical parameters? Besides, since different physics engines can have unique methods for simulating physical interactions, how might the choice of simulator (e.g., using MuJoCo vs. others) affect the model’s predictions and the general applicability of the method?\n\n(2)  the simulator access setup is interesting, if applicable to real-world tasks, we can use LLM augmented with simulators to perform physical reasoning in real-world. However, the paper is assuming a perfect simulator  that can simulate real-world environment, which makes the results less interesting. It would be nice to add such experiments (studying how sim2real gaps will impact the evaluation)\nSince the paper only use synthetic dataset to evaluate, this raise the question if the method applies to real-world robotics settings\n\n(3) Reproducibility with Open-Source LLMs: Given that GPT-4o is closed-source, how feasible is it to adapt LLMPhy to open-source alternatives, and what changes would be necessary to maintain accuracy?\n\n(4) Can the authors provide insights into how the framework might handle situations where parameter estimation fails due to high variability in initial conditions or other environmental factors?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper proposes a black-box optimization framework that uses In-context learning to enhance the physical reasoning skills of SOTA LLMs.To augment the reasoning, the framework leverages the coding capability of LLMs and give it access to a simulator. This paper is a valuable contribution to physical reasoning using LLMs, presenting an original approach that effectively combines the strengths of language models and physics-based simulations. The results are promising, though some improvements in scalability, dataset credibility, and validation across physical engines would strengthen the framework’s applicability and robustness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) my first concern is the limited physical parameters, the model currently explores only four physical parameters, which restricts its application to broader and more complex real-world scenarios. This constraint is acknowledged by the authors but remains a significant limitation for potential applications in varied physical environments.\n\n(2) Dataset Scale and Generalization: With only 100 sequences, TraySim may be limited in capturing the full complexity of physical interactions. Expanding this dataset or testing LLMPhy on existing physics-based benchmarks could provide a more comprehensive understanding of the model’s generalization ability.\n\n(3) Absence of Real-World Validation: While the simulation results are promising, real-world validations or experiments in a physical setup would enhance the credibility of the proposed method. Without real-world tests, it remains uncertain how well the model's inferred physical parameters would translate to actual physics."}},"nonreaders":[],"tmdate":1731428017628,"tcdate":1730679480767,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12837/Reviewer_sHax"],"signatures":["ICLR.cc/2025/Conference/Submission12837/Reviewer_sHax"],"forum":"qGL6fE1lqd","number":3,"license":"CC BY 4.0","cdate":1730679480767,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12837/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428017628,"domain":"ICLR.cc/2025/Conference","replyto":"qGL6fE1lqd","id":"b0IrVTmedV","forumContent":{"TLDR":{"value":"Combining LLMs and physics engines for solving physical reasoning problems"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["large language models","physics simulators","physical reasoning"]},"supplementary_material":{"value":"/attachment/269b0fa498f32985b40c2d153c2c21b350427de2.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Physical reasoning is an important skill needed for robotic agents when operating in the real world. However, solving such reasoning problems often involves hypothesizing and reflecting over complex multi-body interactions under the effect of a multitude of physical forces and thus learning all such interactions poses a significant hurdle for state-of-the-art machine learning frameworks, including large language models (LLMs). To study this problem, we propose a new physical reasoning task and a dataset, dubbed TraySim. Our task involves predicting the dynamics of several objects on a tray that is given an external impact -- the domino effect of the ensued object interactions and their dynamics thus offering a challenging yet controlled setup, with the goal of reasoning being to infer the stability of the objects after the impact. To solve this complex physical reasoning task, we present LLMPhy, a zero-shot black-box optimization framework that leverages the physics knowledge and program synthesis abilities of LLMs, and synergizes these abilities with the world models built into modern physics engines. Specifically, LLMPhy uses an LLM to  generate code to iteratively estimate the physical hyperparameters of the system (friction, damping, layout, etc.) via an implicit analysis-by-synthesis approach using a (non-differentiable) simulator in the loop and uses the inferred parameters to imagine the dynamics of the scene towards solving the reasoning task.} To show the effectiveness of LLMPhy, we present experiments on our TraySim dataset to predict the steady-state poses of the objects. Our results show that the combination of the LLM and the physics engine leads to state-of-the-art zero-shot physical reasoning performance, while demonstrating superior convergence against standard black-box optimization methods and better estimation of the physical parameters. Further, we show that LLMPhy is capable of solving both continuous and discrete black-box optimization problems."},"_bibtex":{"value":"@misc{\ncherian2025llmphy,\ntitle={{LLMP}hy: Complex Physical Reasoning Using Large Language Models and World Models},\nauthor={Anoop Cherian and Radu Corcodel and Siddarth Jain and Diego Romeres},\nyear={2025},\nurl={https://openreview.net/forum?id=qGL6fE1lqd}\n}"},"title":{"value":"LLMPhy: Complex Physical Reasoning Using Large Language Models and World Models"},"pdf":{"value":"/pdf/1ac1d478c0fe0fe083ebaae5f4be44f529f96183.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"cherian|llmphy_complex_physical_reasoning_using_large_language_models_and_world_models"},"authorids":{"value":["~Anoop_Cherian1","~Radu_Corcodel1","~Siddarth_Jain2","~Diego_Romeres1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Anoop Cherian","Radu Corcodel","Siddarth Jain","Diego Romeres"]}},"version":2},{"content":{"summary":{"value":"The authors propose a novel method to accelerate the inference of myocardial perfusion simulations. Specifically, the method relies on a graph attention UNet, which is trained in an unsupervised manner by optimizing a physics-informed loss function derived from a finite-volume formulation. The method was trained on 2,000 synthetic  perfusion simulations with idealistic ventricular shapes and evaluated on additional simulations over 400 synthetic and 400 patient-specific geometries. The method achieves on both datasets a mean relative absolute error of 0.5% or lower."},"final_rating":{"value":4,"readers":["everyone"]},"justification_of_final_rating":{"value":"I appreciate that the authors have addressed all of my questions and concerns. The changes made to the manuscript have greatly improved the quality of the paper. The authors have now added all necessary information to replicate the results. I am therefore increasing the rating.","readers":["everyone"]},"justification_of_the_preliminary_rating":{"value":"In general, this is a very interesting paper with a novel approach. In the current form of the manuscript, there is, however, too much essential information missing, which raises several question and concerns."},"strengths":{"value":"- The proposed method is very interesting and another great example of physics-informed learning\n- The method is novel in that it is the first fully unsupervised 3D finite-volume-informed graph neural network to be evaluated on both synthetic and patient-specific geometries\n- The derivation of the finite-volume-based loss is well explained and may be understood even by non-experts\n- The generalization to the patient-specific geometries is encouraging despite training the model on synthetic geometries (truncated ellipsoids)"},"weaknesses":{"value":"The proposed method has several benefits, but I found several questions and concerns:\n\n- First of all, the manuscript lacks essential information to replicate the results. Among other things, the paper is missing information on how the parameters for the synthetic geometries were sampled (e.g., range of parameters L1-L5 and the sampling pattern) as well as a clear description of the node encoder and decoder architecture is missing. Similarly, no hyperparameters for training the network were presented.\n\n- Second, it seems to me like an essential limitation/requirement of the proposed method was not discussed at all. In particular, since the method takes as input nodal coordinates, the method would require the geometries to be in a canonical shape or it would require extensive augmentation to cover all possible rotations & translations. \n\n- Similarly, Figure 1 suggests that relatively homogeneous meshes were generated for these experiments. It is therefore unclear what impact mesh resolution has onto the prediction results. Could the authors please also comment on how the receptive field size of the network relate to the overall size of the underlying mesh?\n\n- One particular claim of the proposed method is to have improved generalization capabilities thanks to the physics-informed/finite-volume-informed graph neural network. There is, however, no comparator. While I understand that a fair comparison is not straightforward, I strongly believe that it would be very valuable if the method is at least compared against a regular physics-informed neural network and/or a supervised graph deep learning approach. \n\n- In addition, the presented numerical evaluations only provide a global insight into the prediction accuracy. Neither the table nor the visual examples (since they are only showing a specific front of the geometry) can provide conclusions about any regional biases. Considering that cardiac ventricular geometries are considered, one simple next step may comprise the regional error analysis, e.g., based on the standardized 17-segment subdivision of the left ventricle by the American Heart Association. Moreover, it may be worth showing few more examples (front and back) of other heart geometries\n\n- Finally, one particular advantage of such a method, as mentioned by the authors, may be the inference time once the network is trained. It would strengthen the paper, if a preliminary analysis on run-time differences between the ground truth finite-volume simulations and the network predictions (including the training time) were presented. In addition, I would appreciate if any information on the utilized software packages were presented."},"confidence":{"value":4},"detailed_comments":{"value":"- All Equations: Shouldn't the Darcy velocity in 3D be a vector? If so, please properly reflect it in the equations\n- Equation 4: I think there is a minus missing. Could the authors please double-check the equation?\n- Figure 2: Please add a legend to clarify what f, i, and j are\n- Section 3.3.2: The graph definition is different from the one in the preceding section, please revise the definition.\n- Table 1: Two spaces are missing in the middle column. Please also clarify whether the values after +- denote the standard deviation or something else\n- Figure 5: Please increase the size of the color bars and please use the same value range to color the meshes by the errors\n- Abbreviations: Please be consistent with the abbreviations, use all capital letters for the abbreviations, and also define the abbreviations in the table descriptions so that the tables can be understood without having to read the entire paper \n- Conclusion: non-supervised -> unsupervised"},"recommendation":{"value":"Poster"},"questions_to_address_in_the_rebuttal":{"value":"In addition to the questions and concerns raised above, I would appreciate if the authors could comment on the following topics:\n\n- The manuscript states in the conclusion \"Moreover, the integration of physics-based loss functions with a non-supervised framework holds potential for enhancing the explainability of the models while reducing the reliance on data, thereby supporting their application in medical context\". From my point of view, neither the unsupervised learning nor the physics-based loss helps in any way with explainability since there will be no guarantee that the output of the network is physically plausible. It may, however, help with reducing the errors, improved generalization, and faster convergence. Could the authors please share their opinion on this topic?\n\n- Equation 3 presents the boundary condition of the Darcy model. Were any special considerations in the loss function taken to properly handle/incoporate this constraint?\n\n- In the description of the graph pooling, it was mentioned that a random subset of nodes is chosen. Could the authors please clarify the reasoning behind it and how the connectivity for the coarse graph was computed?"},"preliminary_rating":{"value":2}},"nonreaders":[],"tmdate":1712410985705,"tcdate":1709077921906,"writers":["MIDL.io/2024/Conference","MIDL.io/2024/Conference/Submission325/Reviewer_4GD6"],"signatures":["MIDL.io/2024/Conference/Submission325/Reviewer_4GD6"],"forum":"CTJ5ERXOYE","number":1,"license":"CC BY 4.0","cdate":1709077921906,"readers":["everyone"],"invitations":["MIDL.io/2024/Conference/Submission325/-/Official_Review","MIDL.io/2024/Conference/-/Edit","MIDL.io/2024/Conference/Submission325/Official_Review1/-/Review_Revision"],"mdate":1712410985705,"domain":"MIDL.io/2024/Conference","replyto":"CTJ5ERXOYE","id":"M58QJ9EKP3","forumContent":{"venue":{"value":"MIDL 2024 Oral"},"keywords":{"value":["Graph Neural Network","Partial Differential Equations","Physics-informed Neural Network","Finite Volume method","Computed Tomography","Digital twins","Perfusion simulation"]},"abstract":{"value":"Medical imaging and numerical simulation of partial differential equations (PDEs) representing biophysical processes, have been combined in the past few decades to provide noninvasive diagnostic and treatment prediction tools for various diseases. Most approaches involve solving computationally expensive PDEs, which can hinder their effective deployment in clinical settings. To overcome this limitation, deep learning has emerged as a promising method to accelerate numerical solvers. One challenge persists however in the generalization abilities of these models, given the wide variety of patient morphologies. This study addresses this challenge by introducing a physics-informed graph neural network designed to solve Darcy equations for the simulation of myocardial perfusion. Leveraging a finite volume discretization of the equations as a \"physics-informed\" loss, our model was successfully trained and tested on a 3D synthetic dataset, namely meshes representing simplified myocardium shapes. Subsequent evaluation on a genuine myocardium mesh, extracted from patient Computed Tomography images, demonstrated promising results and generalized capabilities. Such a fast solver, within a differentiable learning framework will enable to tackle inverse problems based on $\\text{H}_2$O-PET perfusion imaging data."},"_bibtex":{"value":"@inproceedings{\nchou2024finite,\ntitle={Finite Volume Informed Graph Neural Network for Myocardial Perfusion Simulation},\nauthor={Raoul Sall{\\'e} de Chou and Matthew Sinclair and Sabrina Lynch and Nan Xiao and Laurent Najman and Irene Vignon-clementel and Hugues Talbot},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2024},\nurl={https://openreview.net/forum?id=CTJ5ERXOYE}\n}"},"title":{"value":"Finite Volume Informed Graph Neural Network for Myocardial Perfusion Simulation"},"latex_code":{"value":"/attachment/8cd8d55a31a88ef808031df0fd7b39f917d856f4.zip"},"pdf":{"value":"/pdf/5d84be2c0a796b1379db541ff70ac21829a45221.pdf"},"copyright_form":{"value":"/attachment/96bb8f2d1ba3fe21abd328030cb86a5692a428b8.pdf"},"venueid":{"value":"MIDL.io/2024/Conference"},"paperhash":{"value":"chou|finite_volume_informed_graph_neural_network_for_myocardial_perfusion_simulation"},"authorids":{"value":["~Raoul_Sallé_de_Chou1","~Matthew_Sinclair1","slynch@heartflow.com","nxiao@heartflow.com","~Laurent_Najman1","~Irene_Vignon-clementel1","~Hugues_Talbot1"]},"authors":{"value":["Raoul Sallé de Chou","Matthew Sinclair","Sabrina Lynch","Nan Xiao","Laurent Najman","Irene Vignon-clementel","Hugues Talbot"]}},"version":2},{"content":{"summary":{"value":"This work proposes a diffusion-based reconstruction framework for dark-field micro-CT that extends the DDIP by integrating physics-based consistency and test-time adaptation via low-rank adaptation (LoRA). Experiments show improved quantitative and qualitative results. Overall, the study demonstrates that diffusion priors combined with physics-based modeling can significantly improve image quality in dose-constrained imaging settings."},"review":{"value":"The paper addresses an important problem in dark-field CT reconstruction under severe undersampling, which is highly relevant for dose-limited pre-clinical imaging. The integration of diffusion priors with physics-based reconstruction and test-time adaptation is technically sound and aligns with recent trends in inverse problems. The methodology is reasonably well-structured, combining DDIP, ADMM-based data consistency, and LoRA adaptation, although the presentation is quite dense and may be challenging for readers unfamiliar with diffusion-based reconstruction frameworks.\n\nIn terms of originality, the work is primarily an adaptation and integration of existing components (DDIP, LoRA, ADMM, TIGRE operators) rather than a fundamentally new methodological contribution. However, its application to dark-field CT and the incorporation of weighted temporal averaging provide incremental novelty. T Overall, the work is technically solid, suitable for acceptance as a short paper, and will generate important discussion during the meeting."},"strengths":{"value":"Addresses a clinically and scientifically important problem: reconstruction under severe undersampling in dark-field CT. Effective integration of diffusion priors with physics-based reconstruction, which is a strong and modern approach to inverse problems. Incorporates test-time adaptation (LoRA), improving robustness to out-of-distribution data—an important practical consideration."},"weaknesses":{"value":"Limited methodological novelty: the framework largely integrates existing techniques (DDIP, LoRA, ADMM) rather than introducing fundamentally new concepts. Evaluation is insufficiently comprehensive—comparisons are limited to FDK, with no benchmarking against recent state-of-the-art reconstruction or diffusion-based methods. Heavy reliance on synthetic phantoms for training raises concerns about generalization to real-world or clinical data. Lack of external validation or reader studies makes it difficult to assess practical or clinical impact. The method is computationally complex, but no analysis of runtime, scalability, or resource requirements is provided."},"confidence":{"value":4},"rating":{"value":4},"justification_of_rating":{"value":"see above"},"title":{"value":"intersting work"}},"parentInvitations":"MIDL.io/2026/Short_Papers/-/Official_Review","nonreaders":[],"tmdate":1778509196717,"tcdate":1777421263679,"writers":["MIDL.io/2026/Short_Papers","MIDL.io/2026/Short_Papers/Submission114/Reviewer_aP8o"],"signatures":["MIDL.io/2026/Short_Papers/Submission114/Reviewer_aP8o"],"forum":"vQMgNm8lFc","number":1,"license":"CC BY 4.0","cdate":1777421263679,"readers":["everyone"],"invitations":["MIDL.io/2026/Short_Papers/Submission114/-/Official_Review","MIDL.io/2026/Short_Papers/-/Edit"],"mdate":1778509196717,"domain":"MIDL.io/2026/Short_Papers","replyto":"vQMgNm8lFc","id":"cexAgdtz6A","forumContent":{"TLDR":{"value":"The method combines a pretrained diffusion prior with physics-based consistency and incorporates low-rank adaptation (LoRA) to adapt the pretrained diffusion model to the measured data for pre-clinical dark-field micro-CT reconstruction."},"venue":{"value":"MIDL 2026 - Short Papers Poster"},"keywords":{"value":["Computed tomography","dark-field imaging","diffusion models"]},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"X-ray dark-field imaging enables visualization of lung microstructure and the detection of pulmonary diseases. However, in pre-clinical micro-CT studies, the reconstruction suffers from severe streak artifacts due to highly undersampled acquisitions constrained by radiation dose. We leverage a diffusion-based reconstruction framework by extending the Deep Diffusion Image Prior (DDIP) to dark-field CT. The method combines a pretrained diffusion prior with physics-based consistency and incorporates low-rank adaptation (LoRA) for improved robustness to out-of-distribution data. Experiments demonstrate increased contrast-to-noise ratio and artifact suppression while preserving edge sharpness, enabling higher-quality dark-field imaging for dose-constrained longitudinal studies."},"_bibtex":{"value":"@inproceedings{\nhiu2026adaptive,\ntitle={Adaptive Diffusion Priors on Pre-clinical {CT} Reconstruction},\nauthor={Theresa Hiu and Daniel Frey and Tina Dorosti and Johannes Thalhammer and Sebastian Peterhansl and Zijin Huang and Simon Zandarco and Franz Pfeiffer and Florian Schaff},\nbooktitle={Medical Imaging with Deep Learning - Short Papers},\nyear={2026},\nurl={https://openreview.net/forum?id=vQMgNm8lFc}\n}"},"title":{"value":"Adaptive Diffusion Priors on Pre-clinical CT Reconstruction"},"pdf":{"value":"/pdf/695db9fec6a64235f4765a588d948ae85c012ff1.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Short_Papers"},"paperhash":{"value":"hiu|adaptive_diffusion_priors_on_preclinical_ct_reconstruction"},"authorids":{"value":["~Theresa_Hiu1","~Daniel_Frey1","~Tina_Dorosti1","~Johannes_Thalhammer1","sebastian.peterhansl@tum.de","~Zijin_Huang2","simon.zandarco@tum.de","~Franz_Pfeiffer1","~Florian_Schaff1"]},"registration":{"value":"Yes"},"authors":{"value":["Theresa Hiu","Daniel Frey","Tina Dorosti","Johannes Thalhammer","Sebastian Peterhansl","Zijin Huang","Simon Zandarco","Franz Pfeiffer","Florian Schaff"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"This paper presents a theoretical analysis of physics-informed learning in the presence of dependent data. The authors analyze empirical risk minimization with a physics-informed regularizer that encodes known physical priors in the form of elliptic PDE constraints. Using the small-ball method and martingale offset complexity, they derive complexity-dependent excess risk bounds both in probability and expectation. The main result shows that when the physical prior regularizer is aligned with the true dynamics, the convergence rate improves from the slow Sobolev minimax rate to the optimal i.i.d. rate, even with dependent samples. The theoretical framework is supported by a unicycle dynamics experiment illustrating the empirical benefit of incorporating physics-informed regularization."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. see in the weakness\n\n2. The phsyics-informed regularizer is limited to linear elliptic PDEs. However, the data is dependent, which is always occured in dynamical systems. Is this two conflicting? The numerical example is an ODE, not elliptic PDEs. How the convergence rate behaves for common PINNs problems such as Poisson equation or Darcy flow problem? How can the proposed theory be connected to the optimization landscape of neural networks trained with PINNs?\n\n3. Does the empirical rate in the numerical example persist when neural architectures differ from MLPs, or with stochastic training?\n\n4. Can the presentation be more friendly to the general ICLR audience without additional intuition or graphical explanation of the key proof mechanisms? The current form is mathematically dense and more suitable to journals like JMLR, not general top conferences."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The paper rigorously extends the small-ball method and offset complexity analysis to dependent data settings with physics-informed regularization, which is technically novel.\n\n2. The paper establishes a clear theoretical connection between physical priors (encoded as elliptic PDE constraints) and improved learning rates, addressing a long-standing gap in understanding the benefits of physics-informed models.\n\n3. The proofs and appendices appear detailed, well-grounded in functional analysis, Sobolev space theory, and dependent process theory.\n\n4. The  experiment, though simple, can demonstrate the theoretical prediction that physical priors accelerate convergence."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The analysis assumes elliptic linear PDE operators; it is unclear whether the results generalize to non-elliptic, nonlinear, or mixed-type operators often seen in physics-informed neural networks (PINNs). The gap or difficulty is not well addressed. \n\n2. The optimization error's influnene is not discussed in the theoretical part and the numerical example. Note that the physics-informed regularizer will increase the stiffness of the Hessian matrix and increase difficulties for the optimization in PINNs, which is not aligned with the main result (convergence rate improves with physics-inforemed regularizer added)\n\n3. Too many assumptions may not hold for complex real-world systems, a more interpretable or verifiable condition would strengthen the practical relevance.\n\n4. Only one low-dimensional example is provided. Additional tests on nonlinear PDE systems or stochastic dynamical systems would reinforce the claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922856344,"tcdate":1762570917225,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11838/Reviewer_tMrX"],"signatures":["ICLR.cc/2026/Conference/Submission11838/Reviewer_tMrX"],"forum":"IvLVPbeoRx","number":4,"license":"CC BY 4.0","cdate":1762570917225,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11838/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922856344,"domain":"ICLR.cc/2026/Conference","replyto":"IvLVPbeoRx","id":"eMfuLA0HVo","forumContent":{"TLDR":{"value":"We prove that adding correct prior domain knowledge to nonparametric learning with dependent data speeds up learning"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["learning with dependent data","physics-informed machine learning","convergence rates","complexity-dependent bounds"]},"supplementary_material":{"value":"/attachment/a00585da11ea98b0c0c19b8e1c5d1c3a2b58963e.pdf"},"primary_area":{"value":"learning theory"},"abstract":{"value":"A major challenge in physics-informed machine learning is to understand how the incorporation of prior domain knowledge affects learning rates when data are dependent. Focusing on empirical risk minimization with physics-informed regularization, we derive complexity-dependent bounds on the excess risk in probability and in expectation. We prove that, when the physical prior information is aligned, the learning rate improves from the (slow) Sobolev minimax rate to the (fast) optimal i.i.d. one without sample-size deflation due to data dependence."},"_bibtex":{"value":"@inproceedings{\nscampicchio2026physicsinformed,\ntitle={Physics-informed learning under mixing: How physical knowledge speeds up learning},\nauthor={Anna Scampicchio and Leonardo Felipe Toso and Rahel Rickenbach and James Anderson and Melanie Zeilinger},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=IvLVPbeoRx}\n}"},"title":{"value":"Physics-informed learning under mixing: How physical knowledge speeds up learning"},"pdf":{"value":"/pdf/9917146046b56820383565485dace6196f5b94c5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"scampicchio|physicsinformed_learning_under_mixing_how_physical_knowledge_speeds_up_learning"},"authorids":{"value":["~Anna_Scampicchio1","~Leonardo_Felipe_Toso1","~Rahel_Rickenbach1","~James_Anderson6","~Melanie_Zeilinger1"]},"authors":{"value":["Anna Scampicchio","Leonardo Felipe Toso","Rahel Rickenbach","James Anderson","Melanie Zeilinger"]}},"version":2},{"content":{"summary":{"value":"PACER is an uncertainty-aware climate emulator that integrates a physics-based advection PDE with a Neural SDE residual corrector. The proposed method enables stable 10-year temperature field emulation on multiple climate models."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Can you provide quantitative evidence that learned velocity fields (after refinement) maintain physical plausibility? For example, the autghors can show visualizations compared to reanalysis winds.\n2. Can you compare performance using Euler and RK4 solvers instead of the dopri5 solver? Can you compare RMSE and wall-clock time for different ODE solvers?\n3. Do PACER and the baselines have the same number of encoder and decoder layers? What are the parameter counts for PACER and each baseline? Could you provide comparisons under a reasonably fair parameter budget?\n4. Why was no direct comparison performed with studies [1,2] that similarly integrated the Advection-Diffusion Equation into a Neural Network framework? Can you discuss further what differentiates your work from these related studies?\n\n\n> [1] \"Climate modeling with neural advection–diffusion equation.\" Knowledge and Information Systems 65.6 (2023): 2403-2427.\n> \n> [2] “Climode: Climate and weather forecasting with physics-informed neural odes.\" ICLR 2024"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. The paper proposes a novel autoregressive framework that combines an advection PDE-based physics-informed backbone with a Neural SDE-based uncertainty-aware residual corrector.\n2. Training the neural SDE under a Gaussian NLL objective with multiple Brownian realizations ($K>1$) enables heteroscedastic aleatoric uncertainty estimation and reduces gradient variance.\n3. The explicit integration of advection dynamics via ODE solver (dopri5) with velocity field inference provides a transparent physics backbone that is computationally lightweight."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Despite claiming to be \"Physics Informed,\" the physical laws integrated into the emulator are limited only to advection, which is far from the complex multi-physics systems of actual GCMs.\n2. PACER emphasizes being \"lightweight\" (2.1M parameters) compared to ACE (200M parameters) and Spherical Dyffusion (200M parameters), but in the results section, it only compares performance with UNet, ConvLSTM, and SFNO instead of these models.\n3. There is only one ablation (Figure 3) that compares \"physics informed vs uninformed\" but doesn't isolate individual components.\n4. The notation lacks consistency, reducing readability. Eq.2 uses $x_t$, while subsequent sections use $u_t$ for the climate state. Forcing is denoted as $F$ in some places and $f$ in others.\n5. The authors claim their proposed method is \"lightweight,\" but there is no wall-clock time comparison, GPU memory usage, or FLOP analysis."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927022188,"tcdate":1761983256563,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17000/Reviewer_3eAm"],"signatures":["ICLR.cc/2026/Conference/Submission17000/Reviewer_3eAm"],"forum":"RhW8yXgIxY","number":4,"license":"CC BY 4.0","cdate":1761983256563,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17000/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927022188,"domain":"ICLR.cc/2026/Conference","replyto":"RhW8yXgIxY","id":"kvayHJczLg","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"A 10-year auto-regressive, advection informed, uncertainty aware climate emulator trained on NLL objective."},"keywords":{"value":["Neural SDE","Uncertainty","Physics Informed Network","Climate Emulator"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics based numerical climate models serve as critical tools for evaluating the effects of climate change and projecting future climate scenarios. However, the reliance on numerical simulations of physical equations renders them computationally intensive and inefficient. While deep learning methodologies have made significant progress in weather forecasting, they are still unstable for longer roll-out climate emulation task. Here, we propose PACER, a relatively lightweight 2.1M parameter Physics Informed Uncertainty Aware Climate EmulatoR. PACER is trained across is trained across varying spatial resolutions and physics based climate models, enabling faithful and stable emulation of temperature fields at multiple surface levels over a 10 year horizon. We propose an auto-regressive ODE–SDE framework for climate emulation that integrates the fundamental physical law of advection, while being trained under a negative log-likelihood objective to enable principled uncertainty quantification of stochastic variability. We show PACER's emulation performance across 20 climate models outperforming relevant baselines and advancing towards explicit physics infusion in ML emulator."},"_bibtex":{"value":"@misc{\nsaleem2025pacer,\ntitle={{PACER}: Physics Informed and Uncertainty Aware Climate Emulator},\nauthor={Hira Saleem and Flora D. Salim and Cormac Purcell},\nyear={2025},\nurl={https://openreview.net/forum?id=RhW8yXgIxY}\n}"},"title":{"value":"PACER: Physics Informed and Uncertainty Aware Climate Emulator"},"pdf":{"value":"/pdf/efe47dca6a1e7f0bb6347335b2bbfb98d2b9ec39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"saleem|pacer_physics_informed_and_uncertainty_aware_climate_emulator"},"authorids":{"value":["~Hira_Saleem1","~Flora_D._Salim1","~Cormac_Purcell1"]},"authors":{"value":["Hira Saleem","Flora D. Salim","Cormac Purcell"]}},"version":2},{"content":{"summary":{"value":"Direct numerical simulations of turbulent flows can be prohibitively expensive to carry out. Fully data-driven or hybrid physics-based machine learning models can be quite promising in reducing the turnaround times for reconstructing fine scale data from coarse grained simulations or long time prediction of flow-field given historical DNS data. Motivated by the potential benefits of fourier neural operators in handling complex spatio-temporal data, the authors propose a physics enhanced neural operator method to model complex flow-field dynamics.\nWhile FNOs work in a purely data driven fashion, the PENO in addition to data, also leverages the physics knowledge in the form of the underlying governing PDEs of turbulent flows. The authors also introduce a self-augmentation technique to enable long-time simulations/roll-outs. They demonstrate the model's capability on  different turbulent flow datasets and test across different resolutions as well."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Could you elaborate on the key differences between PENO and Physics informed neural Operator?\n2. Is there any reason to not compare your model with other operator learning based frameworks and not include them in the survey or in this study?\n3. Could you explain why temporal or spatial spectrum of the flow-fields have not been included in the evaluation? Any justification for why dissipation different has been considered?\n4. Can PENO be used in a fully data-free regime as in the PINO paper? Where given only the initial and boundary conditions, can the model be trained to simulate turbulent flow?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The problem definition is clear with regard to the PENO being trained under a forecasting objective satisfying the physics constraints.\nInstead of having a single network satisfy both data-driven and physics-based constraints, PENO has two branches, one FNO and the other physics based PDE branch. The final prediction / forecast is a weighted combination of the outputs of both branches.  Instead of using continuous derivatives, the authors use temporal and spatial discretisation to estimate the gradient terms in the governing equations. This is beneficial as it reduces the load on the neural network to strongly learn the continuous derivatives in the presence of only sparse data.\nThe authors clearly describe their methodology, datasets used, training and testing protocol.\nThey provide results validating their method across different benchmarks and other data-driven models. \nThe objective function is a simple forecasting MSE based loss function.This makes the learning easy and could prevent the competing objectives problem otherwise encountered in PINNs. Moreover, PENO allows multiple data sources to be combined as well such as DNS and LES through a weighted combination."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the authors claim novelty in the physics enhanced operator, the authors have seemed to ignore previous works on physics informed/enhanced operator learning in this domain. Physics informed neural operator: https://arxiv.org/pdf/2111.03794v3, Physics informed DeepONet https://ar5iv.labs.arxiv.org/html/2207.05748 and its derivatives. Instead of comparing their model with other operator learning frameworks, they infact compare with some of the super-resolution models which in some way seems out of context. It would have been better if the authors could have provided comparison with the other previously proposed Physics based operator learning frameworks which came out much earlier than this work. While the contours look decent in the results, comparison of spatial and temporal spectrum would be worthwhile in showing if the model is capable of overcoming spectral bias that otherwise plagues neural networks in general. It is not clear if PENO can operate in a purely physics based training regime without the data-driven component as in PINO."}},"nonreaders":[],"tmdate":1731429116057,"tcdate":1730374524002,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8511/Reviewer_iKBQ"],"signatures":["ICLR.cc/2025/Conference/Submission8511/Reviewer_iKBQ"],"forum":"5LvTfc4fBz","number":3,"license":"CC BY 4.0","cdate":1730374524002,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8511/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429116057,"domain":"ICLR.cc/2025/Conference","replyto":"5LvTfc4fBz","id":"U0XrG0KhyJ","forumContent":{"TLDR":{"value":"This paper presents a novel physics-enhanced neural operator (PENO) that incorporates physical knowledge of partial differential equations (PDEs) to accurately model flow dynamics."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["turbulent flow","neural operator","knowledge-guided machine learning","sequential simulation"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The precise simulation of turbulent flows is of immense importance in a variety of scientific and engineering fields, including climate science, freshwater science, and the development of energy-efficient manufacturing processes. Within the realm of turbulent flow simulation, direct numerical simulation (DNS) is widely considered to be the most reliable approach, but it is prohibitively expensive for long-term simulation at fine spatial scales. Given the pressing need for efficient simulation, there is an increasing interest in building machine learning models for turbulence, either by reconstructing DNS from alternative low-fidelity simulations or by predicting DNS based on the patterns learned from historical data. However, standard machine learning techniques remain limited in capturing complex spatio-temporal characteristics of turbulent flows, resulting in limited performance and generalizability. This paper presents a novel physics-enhanced neural operator (PENO) that incorporates physical knowledge of partial differential equations (PDEs) to accurately model flow dynamics. The model is further refined by a self-augmentation mechanism to reduce the accumulated error in long-term simulations. The proposed method is evaluated through its performance on two distinct sets of 3D turbulent flow data, showcasing the model's capability to reconstruct high-resolution DNS data, maintain the inherent physical properties of flow transport, and generate flow simulations across various resolutions. Additionally, experimental results on multiple 2D vorticity flow series, generated by different PDEs, highlight the transferability and generalizability of the proposed method. This confirms its applicability to a wide range of real-world scenarios in which extensive simulations are needed under diverse settings."},"_bibtex":{"value":"@misc{\nchen2025physicsenhanced,\ntitle={Physics-enhanced Neural Operator: An Application in Simulating Turbulent Transport},\nauthor={Shengyu Chen and Peyman Givi and Can Zheng and Xiaowei Jia},\nyear={2025},\nurl={https://openreview.net/forum?id=5LvTfc4fBz}\n}"},"title":{"value":"Physics-enhanced Neural Operator: An Application in Simulating Turbulent Transport"},"pdf":{"value":"/pdf/cbb85a4fedc7db15cdd904a4aa99dedf2f8609ce.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|physicsenhanced_neural_operator_an_application_in_simulating_turbulent_transport"},"authorids":{"value":["~Shengyu_Chen1","~Peyman_Givi1","~Can_Zheng1","~Xiaowei_Jia1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shengyu Chen","Peyman Givi","Can Zheng","Xiaowei Jia"]}},"version":2},{"content":{"comment":{"value":"That's a brilliant suggestion. We have crafted a more intuitive explanation here. \n    \nOur framework makes uncertainty quantification for neural-PDE solvers more intuitive and practically useful by leveraging two key insights. First, we use the PDE residual (how well the solution satisfies the underlying physics equations) as a measure of uncertainty, rather than traditional error metrics that require ground truth data. This means we can evaluate the quality of predictions without needing expensive simulation data - if a predicted solution violates conservation laws by a large amount, we know it's likely to be inaccurate. Second, we calibrate these physics-based error estimates using conformal prediction, which provides statistical guarantees about the uncertainty bounds. This is like having a physics-aware \"confidence score\" that tells us how much we can trust each prediction.\n\nThe practical value becomes clear in applications - imagine using a neural-PDE solver for weather forecasting or fusion reactor control. Our method provides calibrated error bounds that tell you when predictions are physically implausible and should be double-checked with more expensive traditional methods. We offer two complementary ways to quantify uncertainty: \"marginal\" coverage that gives error bars for each point in space and time (useful for identifying specific problematic regions), and \"joint\" coverage that evaluates entire predictions (useful for filtering out globally unreliable solutions). This makes the uncertainty estimates both theoretically rigorous and practically actionable for scientists and engineers using these models."},"title":{"value":"Intuitive Explanation"}},"tmdate":1732300129567,"tcdate":1732300129567,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6603/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission6603/Authors"],"forum":"cF6OoaYcRa","number":2,"license":"CC BY 4.0","cdate":1732300129567,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6603/-/Official_Comment"],"mdate":1732300129567,"domain":"ICLR.cc/2025/Conference","replyto":"vMtJXjN2bV","id":"twBbEXVq6M","forumContent":{"TLDR":{"value":"Uncertainty quantification of neural-PDE solvers using physics residual errors with conformal prediction."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Surrogate Models","Uncertainty Quantification","Neural-PDE","Physics-Informed","Conformal Prediction"]},"supplementary_material":{"value":"/attachment/52dd2eca0ed4a7944d1f4297d904bc6087d82e8b.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Neural PDEs have emerged as inexpensive surrogate models for numerical PDE solvers. While they offer efficient approximations, they often lack robust uncertainty quantification (UQ), limiting their practical utility. Existing UQ methods for these models typically have high computational demands and lack guarantees. We introduce a novel framework for calibrated physics-informed uncertainty quantification to address these limitations. Our approach leverages physics residual errors as a nonconformity score within a conformal prediction (CP) framework. This enables data-free, model-agnostic, and statistically guaranteed uncertainty estimates. Our framework utilises convolutional layers as finite difference stencils for gradient estimation, our framework provides inexpensive coverage bounds for the violation of conservation laws within model predictions. In our experiments, we utilise CP to obtain marginal coverage for each cell and joint coverage over the entire prediction domain of various PDEs."},"_bibtex":{"value":"@misc{\ngopakumar2025calibrated,\ntitle={Calibrated Physics-Informed Uncertainty Quantification},\nauthor={Vignesh Gopakumar and Ander Gray and Daniel Giles and Lorenzo Zanisi and Stanislas Pamela and Matt Kusner and Marc Peter Deisenroth},\nyear={2025},\nurl={https://openreview.net/forum?id=cF6OoaYcRa}\n}"},"title":{"value":"Calibrated Physics-Informed Uncertainty Quantification"},"pdf":{"value":"/pdf/6df75a15074729c26bfc97f1b876973872208e8a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"gopakumar|calibrated_physicsinformed_uncertainty_quantification"},"authorids":{"value":["~Vignesh_Gopakumar1","~Ander_Gray1","~Daniel_Giles1","~Lorenzo_Zanisi1","~Stanislas_Pamela1","~Matt_Kusner1","~Marc_Peter_Deisenroth1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Vignesh Gopakumar","Ander Gray","Daniel Giles","Lorenzo Zanisi","Stanislas Pamela","Matt Kusner","Marc Peter Deisenroth"]}},"version":2},{"content":{"summary":{"value":"The authors purport to develop a novel framework for modeling physical systems using deep learning architectures in this work. The authors claim that models that first learn the \"physics\" of a given simple system will generalize better to more complex system variants. The authors utilize metallic alloy strength and brain morphological development as two systems of study, providing simple examples of how learning lower-fidelity models can aid in modeling more complex systems."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"See above:\n1) how does this framework differ from other existing work in the literature regarding transfer learning from low to high fidelity systems?\n2) how exactly does learning a black-box model which can predict the system count as learning the \"physics\" of the system?"},"rating":{"value":1},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"Transfer learning from low to high fidelity systems is an interesting sub-domain within transfer learning at large. I thank the authors for highlighting this problem in their work, even if it is a domain which has already been significantly explored elsewhere.\n\nI do think the core intuition at work in this paper is interesting - the use of lower-fidelity models for developing physics is quite common in the engineering sciences and the sciences at large; however, I think this core insight is marred by a lack of detail or acknowledgement of other related works which approach these problems in similar ways."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I am chiefly concerned with the significance of this contribution to the literature on transfer learning and the deep learning community at large. Many approaches to transfer learning using lower-fidelity systems or simulations exist already in the literature, and it is well-understood that such approaches can provide benefits over training directly on the more complex system. It is not clear to me how the approach in this work differs from these methodologies significantly other than in applications.  See [1], [2], [3], [4] for some examples which utilize a very similar underlying approach to that in this work. I would like to challenge the authors to differentiate their work more from this existing literature and consider a resubmission. If anything, the authors should consider these other works as baseline approaches for the sake of comparison. \n\nAdditionally, this paper suffers from a lack of detail regarding the proposed framework. The most obvious omission is any rigorous definition of what the \"learning the physics\" means within this work. In previous literature, \"learning the physics\" more frequently means learning a set of differential equations which describe the system, rather than learning a black-box model with parameters which can predict system changes, but which doesn't provide any direct physical interpretation. Based on the discussion in section 2, the authors seem to be utilizing the latter kind of approach. It is not at all clear to me how learning a simple convolutional neural network provides any manner of learning of the physics of a system. The authors either need to clarify this connection, or clarify that they are doing something other than learning the physics of the system. I am willing to reconsider this point given a very compelling argument from the authors; however, I think even if an argument is provided here, a more significant revision would be needed to further clarify the details omitted in this work.\n\n[1] De, Subhayan, et al. \"On transfer learning of neural networks using bi-fidelity data for uncertainty propagation.\" International Journal for Uncertainty Quantification 10.6 (2020).\n\n[2] Chakraborty, Souvik. \"Transfer learning based multi-fidelity physics informed deep neural network.\" Journal of Computational Physics 426 (2021): 109942.\n\n[3] Liu, Zeyu, Meng Jiang, and Tengfei Luo. \"Leveraging low-fidelity data to improve machine learning of sparse high-fidelity thermal conductivity data via transfer learning.\" Materials Today Physics 28 (2022): 100868.\n\n[4] Song, Dong H., and Daniel M. Tartakovsky. \"Transfer learning on multifidelity data.\" Journal of Machine Learning for Modeling and Computing 3.1 (2022)."}},"nonreaders":[],"tmdate":1731428175869,"tcdate":1730472484420,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission14270/Reviewer_81CL"],"signatures":["ICLR.cc/2025/Conference/Submission14270/Reviewer_81CL"],"forum":"llW4qRsF0o","number":2,"license":"CC BY 4.0","cdate":1730472484420,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission14270/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428175869,"domain":"ICLR.cc/2025/Conference","replyto":"llW4qRsF0o","id":"AHc8qRoiyY","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-Transfer Learning; Accuracy-Performance Dilemma; Engineering Sciences; Complexity; Materials Strength; Brain Development"]},"supplementary_material":{"value":"/attachment/97b61f4bba847e7d34ccef5405af7080e6f70324.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The development of theoretical sciences traditionally adheres to an observation-assumption-model paradigm, which is effective in simple systems but challenged by the `curse of complexity’ in modern engineering sciences. Advancements in artificial intelligence (AI) and machine learning (ML) offer a data-driven alternative, capable of interpolating and extrapolating scientific inference where direct solutions are intractable. Moreover, feature engineering in ML resembles dimensional analysis in classical physics, suggesting that data-driven ML methods could potentially extract new physics behind complex data. Here we propose a physics-transfer (PT) learning framework to learn physics across digital models of varying fidelities and complexities, which addresses the accuracy-performance dilemma in understanding representative multiscale problems. The capability of our approach is showcased through screening metallic alloys by their strengths and predicting the morphological development of brains. The physics of crystal plasticity is learned from low-fidelity molecular dynamics simulation and the model is then fed by material parameters from high-fidelity, electronic structures level, density functional theory calculations, offering chemically accurate strength predictions with several orders lower computational costs. The physics of bifurcation in the evolution of brain morphologies is learned from simple sphere and ellipsoid models and then applied to predict the morphological development of human brains, showing excellent agreement with longitudinal magnetic resonance imaging (MRI) data. The learned latent variables are shown to be highly relevant to uncovered physical descriptors, explaining the effectiveness of the PT framework, which holds great potential in closing the gaps in understanding complexity problems in engineering sciences."},"_bibtex":{"value":"@misc{\nzhao2025physicstransfer,\ntitle={Physics-Transfer Learning: A Framework to Address the Accuracy-Performance Dilemma in Modeling Complexity Problems in Engineering Sciences},\nauthor={Yingjie Zhao and Zhiping Xu},\nyear={2025},\nurl={https://openreview.net/forum?id=llW4qRsF0o}\n}"},"title":{"value":"Physics-Transfer Learning: A Framework to Address the Accuracy-Performance Dilemma in Modeling Complexity Problems in Engineering Sciences"},"pdf":{"value":"/pdf/b322fb7dd027df5b022615a17baf5f1134a4cc88.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhao|physicstransfer_learning_a_framework_to_address_the_accuracyperformance_dilemma_in_modeling_complexity_problems_in_engineering_sciences"},"authorids":{"value":["~Yingjie_Zhao2","~Zhiping_Xu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yingjie Zhao","Zhiping Xu"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewer for their critical assessment. To address the concerns regarding task validity and experimental rigor, we have conducted significant additional experiments, including a **Human Baseline (App. E.9)**, **Full-Video Evaluation (App. E.10)**, **Attention Map Analysis (App. F.1)**, and **Prompt Perturbation Analysis (App. E.11)**.\n\n**Summary of Rebuttal:**\n\n- **Novelty (W1, W2, Q1):** InPhyRe is the first benchmark where physical laws **change between training and testing** **and** **requires the model to infer the rule from context** rather than relying on memory.\n- **\"Template Copying\" (W3):** Models fail even when QA pairs are provided, proving they cannot simply copy templates. Furthermore, attention analysis reveals they ignore visual tokens, confirming language bias is a *finding*, not a design flaw.\n- **Evaluation Setup (W4, Q2):** Our new experiment shows that providing **all video frames** actually *degrades* performance, validating our single-frame design.\n- **Baselines & Rigor (W5):** Humans achieve **~90% accuracy** with *less* context than the models, proving the task is visually solvable. We also confirm robustness to prompt perturbations.\n\n> W1. Firstness claim is incorrect or overstated. Prior work already probes counterfactual or violated physics and adaptation in VQA or video settings, for example IntPhys, CLEVRER, CoPhy, ComPhy, Physion and Physion++, ContPhy, PhysBench, Physics Context Builders, and context conditioned physics efforts, and related context driven benchmarks.\n\nQ1. What exact boundary makes this benchmark first? If the novelty is exemplar videos plus VQA under rule changes, state that narrowly and revise the claim.\n> \n\nThe reviewer asks for the \"exact boundary\" of our novelty. We define this boundary precisely: **InPhyRe is the first benchmark where the physical laws themselves change between training and testing, requiring the model to infer the rule from context rather than memory**. InPhyRe is compared against existing benchmarks below (also included as Tab. 1 on page 3).\n\n| Benchmark | Task | Example query | Do physical conditions/laws change between training and testing? | Require on-the-fly physical condition/law inference from demos/test sample? |\n| --- | --- | --- | --- | --- |\n| CLEVRER | Factual and counterfactual physical reasoning | “What shape is the object that collides with the cyan cylinder?” | No | No |\n| ComPhy | General physical reasoning requiring latent property reasoning | “Which event would happen if the grey cylinder were not magnetic?” | No | Yes |\n| CoPhy | Counterfactual physical reasoning | “what happens if we push this ball to the right?” | No | No (models are trained to do counterfactual reasoning) |\n| PhysBench | General physical reasoning about property, dynamics, relations, etc. | “Which object will the cart hit first?” “What is the object closest to the teacup in the Figure?” | No | No |\n| IntPhys (v1 and v2) | Physical plausibility prediction | Predict the next frame based on the initial frames of some event | Yes | No |\n| Physion | Object contact prediction | The initial frames of a video showing falling dominoes | No | No |\n| Physion++ | Object contact prediction through physical property understanding | The initial frames of a video showing falling dominoes | No | No |\n| ContPhy | Physical property and dynamics prediction | “Is the density of the orange fluid greater than that of the green fluid?” ”What can we do to guide most of the orange fluid into cyan cylinder?” | No | No |\n| PCB | Train smaller VLMs to support models during inference | Not an evaluation benchmark. PCBs provide scene narration to support the inference model. | No | No |\n| **InPhyRe** | Infer physical laws from demo samples and apply them to the evaluation sample. | “Will the velocity of the green cube increase after collision with the red sphere?” | **Yes** | **Yes** |\n\nPCB is Physics Context Builder (V. Balazadeh et al., CVPR 2025), discussed in the revision.\n\n**IPR vs counterfactual physical reasoning (CPR)**: The physical laws do not change between training and testing in CPR. Thus, CPR checks the memory of the LMM. Compared to CPR, IPR measures a stronger form of generalization.\n\n**IPR vs intuitive physical reasoning**: The goal in intuitive physical reasoning is to evaluate physical reasoning in models through its ability to detect violated physics. **Violation detection does not require understanding the differing rules**. In contrast, IPR requires inferring the underlying rules from demos.\n\nWe have made these points clear in the revised PDF (Sec. 2, page 3, and refer to lines 1167-1179, page 22)."},"title":{"value":"Response to Reviewer unjk (part 1/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763729131437,"tcdate":1763729131437,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14468/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission14468/Authors"],"forum":"IIrPoZ28dN","number":4,"license":"CC BY 4.0","cdate":1763729131437,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14468/-/Official_Comment"],"mdate":1763729131437,"domain":"ICLR.cc/2026/Conference","replyto":"Tqh4Gaovom","id":"UdAyu61y8C","forumContent":{"TLDR":{"value":"We evaluate whether large multimodal models can infer and apply unseen physical laws from demonstration samples."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["large multimodal models","physical reasoning","inductive reasoning"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large multimodal models (LMMs) encode universal physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potential collision event from visual input. However, since parametric knowledge includes only the physical laws seen during training, it is insufficient for reasoning when the inference scenario follows physical laws unseen during training. In contrast, humans can adapt their physical reasoning to unseen physical environments with only a few visual examples. This inductive physical reasoning ability is indispensable for LMMs if they are to replace human agents in safety-critical applications. Despite its importance, existing visual benchmarks evaluate only the parametric knowledge in LMMs, and not inductive physical reasoning. To this end, we propose InPhyRe, the first visual question answering benchmark to measure inductive physical reasoning in LMMs. InPhyRe evaluates LMMs on their ability to predict the outcome of collision events in algorithmically generated synthetic videos. By inspecting over 13 open-source and proprietary LMMs, InPhyRe informs us that (1) LMMs struggle to apply their limited parametric knowledge about universal physical laws to reasoning, (2) inductive physical reasoning in LMMs is weak when inference scenarios obey physical laws unseen during training, and (3) inductive physical reasoning in LMMs suffers from language bias and largely ignores the visual inputs, questioning the trustworthiness of LMMs regarding visual inputs."},"_bibtex":{"value":"@misc{\nsreekumar2026inphyre,\ntitle={InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning},\nauthor={Gautam Sreekumar and Vishnu Boddeti},\nyear={2026},\nurl={https://openreview.net/forum?id=IIrPoZ28dN}\n}"},"title":{"value":"InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning"},"pdf":{"value":"/pdf/b87a0e49daf5f5f05caff235d97e823b17b5dbfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sreekumar|inphyre_discovers_large_multimodal_models_struggle_in_inductive_physical_reasoning"},"authorids":{"value":["~Gautam_Sreekumar1","~Vishnu_Boddeti1"]},"authors":{"value":["Gautam Sreekumar","Vishnu Boddeti"]}},"version":2},{"content":{"summary":{"value":"This paper proposes to use progressively refined differentiable physics, termed as PRDP, to increase the training efficiency while not harnessing the accuracy. The key finding lies in the fact that the full accuracy of the neural network is achievable through insufficiently converged solvers. Several experiments are conducted to validate the effectiveness of PRDP in reducing training time."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"The experiments report the improved efficiency by adopting progressive refinement and incomplete convergence. Does these strategies influence the accuracy?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"The topic this paper wants to tackle seems interesting. It seems intuitive that, considering the noiseness of neural network training and approximative nature of deep models, the physics solver does not need to fully converge for the network to achieve maximum possible accuracy. This paper proposes to use an adaptive strategy to progressively refine the physics solver and thus improve the training efficiency. Several experiments are conducted to verify the efficacy of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- This paper should provide more background information about *differentiable physics* to make readers better understand the core contribution of the proposed method. I am not an expert of this field, and I find this paper a little bit hard to follow, and also unaware of the broader context this paper lies in.\n- The experiment settings in this paper are not clearly presented. Considering that this is paper submitted to ICLR, I want to know what is the role of the neural networks in each experiment."}},"nonreaders":[],"tmdate":1732528234659,"tcdate":1730446922550,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3012/Reviewer_1fLE"],"signatures":["ICLR.cc/2025/Conference/Submission3012/Reviewer_1fLE"],"forum":"9Fh0z1JmPU","number":2,"license":"CC BY 4.0","cdate":1730446922550,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3012/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732528234659,"domain":"ICLR.cc/2025/Conference","replyto":"9Fh0z1JmPU","id":"wt3hpen9nF","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["differentiable physics","iterative PDE solvers","neural surrogate"]},"supplementary_material":{"value":"/attachment/21fd955829533ffcb5edab8123b77632a930e068.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The physics solvers employed for neural network training are primarily iterative, and hence, differentiating through them introduces a severe computational burden as iterations grow large. Inspired by works in bilevel optimization, we show that full accuracy of the network is achievable through physics significantly coarser than fully converged solvers. We propose *progressively refined differentiable physics* (PRDP), an approach that identifies the level of physics refinement sufficient for full training accuracy. By beginning with coarse physics, adaptively refining it during training, and stopping refinement at the level adequate for training, it enables significant compute savings without sacrificing network accuracy. Our focus is on differentiating iterative linear solvers for sparsely discretized differential operators, which are fundamental to scientific computing. PRDP is applicable to both unrolled and implicit differentiation. We validate its performance on a variety of learning scenarios involving differentiable physics solvers such as inverse problems, autoregressive neural emulators, and correction-based neural-hybrid solvers. In the challenging example of emulating the Navier-Stokes equations, we reduce training time by 62%."},"_bibtex":{"value":"@inproceedings{\nbhatia2025prdp,\ntitle={{PRDP}: Progressively Refined Differentiable Physics},\nauthor={Kanishk Bhatia and Felix Koehler and Nils Thuerey},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=9Fh0z1JmPU}\n}"},"title":{"value":"PRDP: Progressively Refined Differentiable Physics"},"pdf":{"value":"/pdf/82abf59987428b8451c36c339fb2fc08f5b0afdf.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"bhatia|prdp_progressively_refined_differentiable_physics"},"authorids":{"value":["~Kanishk_Bhatia1","~Felix_Koehler1","~Nils_Thuerey1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kanishk Bhatia","Felix Koehler","Nils Thuerey"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Decoupled Value Attention (DVA), a new architecture of Prior data Fitted networks (PFN), with the main contribution being an attention model that separates the Key-Query interaction to apply only on the features (x), to provide weights multiplied by the values y.  This decoupling makes the estimation process more intuitive as it measures the similarity between features via the Key-Query matrices to create weights multiplied by the matching labels (y), similarly to a Gaussian process. The authors claim the method they provide leads to the following main contributions: \n\n \n1) The authors show that the DVA reduces the difference between predicted and true posterior distribution in PFN training.  \n2) The paper shows that a CNN based PFN equipped with DVA performs comparably to a Transformer based DVA-PFN  \n3) The authors show that DVA enables PFNs to scale complex, high-dimensional problems. The authors demonstrate this on a 64-dimensional power flow simulation."},"soundness":{"value":1},"confidence":{"value":1},"questions":{"value":"1. A key result of the original PFN paper (Muller et al) was its ability to learn an intractable posterior from a hierarchical GP prior.  Can the authors provide results showing DVA's performance on this task? \n \n\n2. The original PFN paper (Muller et al) performed a high variety of tests and comparisons with a great contribution for tabular data, while this paper focuses mainly on synthetic data sets and the “power flow” problem, is there any reason to choose that specific problem? Could the method be applied to additional cases?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"1. The premise of the paper is reasonable and seems to be well-founded. The summation of the label and feature embeddings in the original PFN appears to be hurting the results in the provided synthetic cases. \n  \n\n2. The \"Attention is More Important than Architecture\" finding is a significant contribution. By showing that a CNN+DVA can match a Transformer+DVA, the authors successfully decouple the PFN concept from a strict reliance on the Transformer, opening the door to other, potentially more efficient options."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The key demonstration of the original PFN's power was its ability to learn a complex hierarchical model, and from my understanding this is the main reason to implement the K data set training they suggest. However, the authors only test the mechanism against a simple, fixed hyperparameter RBF kernel. Thus, significant omission to not test if DVA retains the ability to approximate these more complex, mixed priors, which was a primary advantage of the PFN framework is problematic. \n \n\n2. The original PFN paper validated its method on a wide range of real-world tabular data. This paper ignores those benchmarks and instead introduces a new, highly specific physics problem (power flow). Without a direct comparison on the original paper benchmarks, it is impossible to assess if DVA is an improvement or a specialized architecture that only excels on certain tasks. \n \n\n3. The authors claim DVA succeeds in high dimensional regimes where the VA PFN simply stops learning, they link this threshold of D = 10 in their GP tests, but when looking at the Müller et al works, they consider datasets with a large amount of features (e.g., covertype = 55), of a similar order of the 64 feature space considered in the paper “power flow” problem.  \n\n \n\n4. Section 2.2 sketches the PFN training objective, but a reader unfamiliar with PFNs will still need (Muller et al) to follow training/inference. \n\nThe paper’s central contribution is the DVA method, which adjusts the original PFN method. This idea seems sensible, as the usage of the attention Key-Query to generate weights for the label is highly intuitive and closely matches the established GP methods. Still, as the main claim of the paper is to provide a better alternative to an existing method, it was not clear enough that the advantages appear numerically and that there is a clear advantage for high-dimensional cases."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925817720,"tcdate":1761885172698,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15541/Reviewer_KGsD"],"signatures":["ICLR.cc/2026/Conference/Submission15541/Reviewer_KGsD"],"forum":"noLMXTqgCp","number":1,"license":"CC BY 4.0","cdate":1761885172698,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15541/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925817720,"domain":"ICLR.cc/2026/Conference","replyto":"noLMXTqgCp","id":"4TpkPhpsb0","forumContent":{"TLDR":{"value":"Decoupled-Value Attention (DVA) separates input similarity from label propagation, mirroring Gaussian process updates and enabling scalable, kernel-free PFNs. This achieves architecture-agnostic and scalable PFNs."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Gaussian Process","Meta-Learning","Prior-data Fitted Networks","Learning of Physics"]},"supplementary_material":{"value":"/attachment/0b1eae37f527a9c2e6d3dbc050517429ff3901de.zip"},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"abstract":{"value":"Prior-data fitted networks (PFNs) are a promising alternative to time-consuming Gaussian process (GP) inference for creating fast surrogates of physical systems. PFN reduces the computational burden of GP-training by replacing Bayesian inference in GP with a single forward pass of a learned prediction model. However, with standard Transformer attention, PFNs show limited effectiveness on high-dimensional regression tasks. We introduce Decoupled-Value Attention (DVA)-- motivated by the GP property that the function space is fully characterized by the kernel over inputs and the predictive mean is a weighted sum of training targets. DVA computes similarities from inputs only and propagates labels solely through values. Thus, the proposed DVA mirrors the GP update while remaining kernel-free. We demonstrate that PFNs are backbone architecture invariant and the crucial factor for scaling PFNs is the attention rule rather than the architecture itself. Specifically, our results demonstrate that (a) localized attention consistently reduces out-of-sample validation loss in PFNs across different dimensional settings, with validation loss reduced by more than 50\\% in five- and ten-dimensional cases, and (b) the role of attention is more decisive than the choice of backbone architecture, showing that CNN, RNN and LSTM-based PFNs can perform at par with their Transformer-based counterparts. The proposed PFNs provide 64-dimensional power flow equation approximations with a mean absolute error of the order of $10^{-3}$, while being over $80\\times$ faster than exact GP inference."},"_bibtex":{"value":"@misc{\nsharma2026decoupledvalue,\ntitle={Decoupled-Value Attention for Prior-Data Fitted Networks: {GP}-Inference for Physical Equations},\nauthor={Kaustubh Sharma and Simardeep Singh and Parikshit Pareek},\nyear={2026},\nurl={https://openreview.net/forum?id=noLMXTqgCp}\n}"},"title":{"value":"Decoupled-Value Attention for Prior-Data Fitted Networks: GP-Inference for Physical Equations"},"pdf":{"value":"/pdf/515a69f9475f19e0257ccd4f4ce710e7a16d159f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sharma|decoupledvalue_attention_for_priordata_fitted_networks_gpinference_for_physical_equations"},"authorids":{"value":["~Kaustubh_Sharma1","~Simardeep_Singh1","~Parikshit_Pareek1"]},"authors":{"value":["Kaustubh Sharma","Simardeep Singh","Parikshit Pareek"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2508.01835v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"zhang|diffusionbased_3d_hand_motion_recovery_with_intuitive_physics"},"authorids":{"value":["~Yufei_Zhang1","https://dblp.org/search/pid/api?q=author:Zijun_Cui:","~Jeffrey_O._Kephart1","https://dblp.org/search/pid/api?q=author:Qiang_Ji:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2508.01835"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2508-01835,\n  publtype={informal},\n  author={Yufei Zhang and Zijun Cui and Jeffrey O. Kephart and Qiang Ji},\n  title={Diffusion-based 3D Hand Motion Recovery with Intuitive Physics},\n  year={2025},\n  month={August},\n  cdate={1754006400000},\n  journal={CoRR},\n  volume={abs/2508.01835},\n  url={https://doi.org/10.48550/arXiv.2508.01835}\n}\n"},"abstract":{"value":"While 3D hand reconstruction from monocular images has made significant progress, generating accurate and temporally coherent motion estimates from videos remains challenging, particularly during hand-object interactions. In this paper, we present a novel 3D hand motion recovery framework that enhances image-based reconstructions through a diffusion-based and physics-augmented motion refinement model. Our model captures the distribution of refined motion estimates conditioned on initial ones, generating improved sequences through an iterative denoising process. Instead of relying on scarce annotated video data, we train our model only using motion capture data without images. We identify valuable intuitive physics knowledge during hand-object interactions, including key motion states and their associated motion constraints. We effectively integrate these physical insights into our diffusion model to improve its performance. Extensive experiments demonstrate that our approach significantly improves various frame-wise reconstruction methods, achieving state-of-the-art (SOTA) performance on existing benchmarks."},"title":{"value":"Diffusion-based 3D Hand Motion Recovery with Intuitive Physics"},"authors":{"value":["Yufei Zhang","Zijun Cui","Jeffrey O. Kephart","Qiang Ji"]}},"tmdate":1767986688584,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2508-01835"],"tcdate":1758688321990,"writers":["~"],"signatures":["~Jeffrey_O._Kephart1"],"forum":"5LBRsSktJo","license":"CC BY-SA 4.0","number":627710,"cdate":1754006400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1767986688584,"domain":"DBLP.org","id":"5LBRsSktJo","version":2},{"content":{"venue":{"value":"Humanoids 2017"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/8215882/8239529/08246964.pdf"},"venueid":{"value":"dblp.org/conf/HUMANOIDS/2017"},"paperhash":{"value":"felip|towards_intuitive_rigidbody_physics_through_parameter_search"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Javier_Felip:","https://dblp.org/search/pid/api?q=author:David_Gonzalez-Aguirre:","~Omesh_Tickoo1"]},"html":{"value":"https://doi.org/10.1109/HUMANOIDS.2017.8246964"},"_bibtex":{"value":"@inproceedings{DBLP:conf/humanoids/FelipGT17,\n  author={Javier Felip and David Gonzalez-Aguirre and Omesh Tickoo},\n  title={Towards intuitive rigid-body physics through parameter search},\n  year={2017},\n  cdate={1483228800000},\n  pages={803-810},\n  url={https://doi.org/10.1109/HUMANOIDS.2017.8246964},\n  booktitle={Humanoids},\n  crossref={conf/humanoids/2017}\n}\n"},"abstract":{"value":"The ability to predict the future location of objects is key for robots operating in unstructured and uncertain scenarios. It is even more important for general purpose humanoid robots that are meant to operate and adapt to multiple scenarios. They need to determine possible outcomes of actions, reason about their effect and plan subsequent movements accordingly to act preemptively. The prediction ability of current robotic systems in is far from that of humans. Neuroscience studies point out that humans have a predictive ability, called intuitive physics, to anticipate the behavior of dynamic environments enabling them to predict and take preemptive actions when necessary, for example to catch a flying ball or grab an object that is about to fall off a table. In this paper, we present a system that learns to predict based on previous observations. First, object's physical parameters are learned through observation using parameter search techniques. Second, the learned dynamic model of objects is used to generate probabilistic predictions through physics simulation. The parameter search update rules proposed, are compared to other approaches from the state-of-the-art in physical parameter learning. Finally, the predictive capability is evaluated through simulated and real experiments."},"title":{"value":"Towards intuitive rigid-body physics through parameter search"},"authors":{"value":["Javier Felip","David Gonzalez-Aguirre","Omesh Tickoo"]}},"tmdate":1731507130124,"pdate":1483228800000,"tcdate":1731507127621,"writers":["~"],"signatures":["~Omesh_Tickoo1"],"forum":"g64k02AwLr","license":"CC BY-SA 4.0","number":228582,"cdate":1483228800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731507130124,"domain":"DBLP.org","id":"g64k02AwLr","version":2},{"content":{"summary":{"value":"Summary:   \nThis paper addresses a key limitation of existing diffusion-based video generation models: while these models can produce visually realistic content, they often generate physically implausible results that violate fundamental physical laws. To solve this problem, the paper proposes a training-free framework that enhances physical plausibility during the inference phase.\n\nContributions:  \n（1）Training-Free Physics-Aware Paradigm: The framework eliminates the high costs of retraining or fine-tuning, a major limitation of prior physics-aware video generation methods. It acts as a plug-and-play solution that can be integrated with existing diffusion models, filling the gap in inference-time physical control for video generation.  \n（2）Reasoning-Driven Counterfactual Construction: Unlike generic negative prompts that often lead to irrelevant violations, the Physics-Aware Reasoning pipeline leverages LLMs to generate targeted counterfactuals. These counterfactuals directly violate specific physical laws while preserving the original scene’s entities and context, ensuring that suppression signals are tightly aligned with physical principles.  \n（3）Innovative Synchronized Decoupled Guidance: The Synchronized Decoupled Guidance strategy addresses the inherent flaws of traditional negative prompting through two complementary designs. Synchronized directional normalization enables early-stage suppression by focusing on the direction of suppression signals (rather than magnitude), while trajectory-decoupled denoising eliminates cumulative bias by evolving parallel latent trajectories for user prompts and counterfactual prompts. Ablation studies confirm that both designs are critical for improving physical plausibility.  \n（4）Broad Generalization and Practical Utility: The framework achieves consistent performance improvements across four key physical domains (mechanics, optics, thermodynamics, and material interactions) when applied to two representative backbones. It also remains competitive with physics-aware models that require additional training. By preserving photorealism and inference efficiency, the framework supports practical applications such as scientific visualization and realistic content generation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Questions:   \n（1）Questions About LLM Reliability in PAR  \nYou use LLMs for Physics-Aware Reasoning to generate counterfactual prompts. Do you have data on how often LLMs produce inaccurate counterfactuals related to physical laws? How much do these inaccuracies reduce SDG’s ability to improve physical plausibility? A response will clarify if LLM dependence risks PAR’s effectiveness.  \n（2）Questions About SDG Hyperparameters  \nSDG uses λ and w as key hyperparameters. Do these values perform consistently across all four physical domains tested? If not, what hyperparameter settings work best for each domain? This will resolve confusion about SDG’s plug-and-play adaptability.  \n（3）Questions About Ambiguous Prompt Handling  \nYour experiments use clear prompts. Have you tested how PAR and SDG perform with vague user prompts? Does ambiguity lead to misaligned counterfactuals and worse physical plausibility scores? Answers will show the framework’s real-world utility.  \n（4）Suggestions for Inference-Time Baselines  \nYou compare only to trained physics-aware models. Add comparisons to recent inference-time physics-aware methods that also avoid retraining. This will better position SDG against similar state-of-the-art approaches.  \n（5）Suggestions for Niche Physics Validation  \nYou test PAR on common physics domains. Validate it on niche domains too, and add a lightweight physics verifier to filter inaccurate counterfactuals. This will improve the framework’s robustness across more scenarios."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"Strengths:   \n（1）The paper stands out for its creative reimagining of physics-aware video generation, addressing critical gaps in prior work:  \n1）It introduces atraining-free, inference-only framework—a departure from existing methods (e.g., PhyT2V, WISA) that require costly retraining or fine-tuning. This removes practical barriers to adopting physics enhancement for real-world use.  \n2）Instead of generic negative prompts, it uses aLLM-powered Physics-Aware Reasoning (PAR) pipelineto generate targeted counterfactuals. These prompts violate specific physical laws while preserving scene entities, turning negative prompting into a precise physics-guided tool rather than a vague avoidance mechanism.  \n\n（2）The work maintains high rigor through sound methodology and comprehensive validation:  \n1）Methodological soundness: PAR’s LLM instruction template (with clear rules for entity analysis and counterfactual construction) ensures consistency. SDG’s design is mathematically grounded in diffusion model principles, avoiding arbitrary choices.  \n2）Comprehensive experiments: It validates across two backbones (CogVideoX-5B, Wan2.1-14B) to prove generality, four physical domains (mechanics, optics, thermodynamics, material interactions) to show cross-domain effectiveness, and rigorous ablations to confirm that PAR and SDG’s two components are all necessary.  \n\n（3）The paper is well-organized and accessible, ensuring complex ideas are understandable:  \n1）Logical structure: It follows a tight \"problem→method→result\" flow—Introduction frames limitations of existing models/negative prompting; Methodology links each component (PAR, SDG) to a specific flaw it solves; Experiments organize results by benchmark, domain, and ablation for clarity.  \n2）Technical transparency: Complex concepts (e.g., SDG’s designs) are explained with equations and plain language, and PAR’s workflow is visualized. Key terms (e.g., \"counterfactual prompt\") are defined consistently, avoiding jargon overload."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses:   \n（1）Over-Reliance on LLM Physics Accuracy Without Mitigation  \nThe framework’s Physics-Aware Reasoning (PAR) depends entirely on LLMs to generate valid counterfactual prompts, yet it ignores LLMs’ limitations in niche physical domains (e.g., non-Newtonian fluids, electromagnetism). No experiments test PAR’s ability to handle such scenarios, nor is there a check for LLM-generated counterfactual errors (e.g., misstating viscosity effects).  \n（2）No SDG Hyperparameter Sensitivity Analysis  \nSDG uses hyperparameters λ (directional normalization scale) and w (guidance strength), but their impact across domains/backbones is untested. A fixed λ may work for mechanics but fail for subtle optics (e.g., light scattering), harming “plug-and-play” utility.  \n（3）No Validation on Ambiguous/Low-Quality Prompts  \nExperiments use clear prompts (e.g., “A tennis ball thrown to the ground”) but ignore real-world vague prompts (e.g., “Something falling in water”) or typos. Ambiguity forces LLMs to make arbitrary assumptions, leading to misaligned counterfactuals.  \n（4）Missing Comparisons to Inference-Time Baselines  \nThe paper only compares to trained physics models (e.g., PhyT2V) but omits recent inference-time methods (e.g.,Think Before You Diffuse,VLiPP) that also avoid retraining. This hides how SDG stacks up to peers."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917878793,"tcdate":1761895967781,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5109/Reviewer_7HDW"],"signatures":["ICLR.cc/2026/Conference/Submission5109/Reviewer_7HDW"],"forum":"SYBfaOcbmw","number":3,"license":"CC BY 4.0","cdate":1761895967781,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5109/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917878793,"domain":"ICLR.cc/2026/Conference","replyto":"SYBfaOcbmw","id":"PSqTdpKfn4","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physical plausibility","video generation","reasoning","synchronized decoupled guidance"]},"supplementary_material":{"value":"/attachment/e59c044a6f6c697f122fdc022811822e5f54850b.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Diffusion models can generate realistic videos, but existing methods rely on implicitly learning physical reasoning from large-scale text-video datasets, which is costly, difficult to scale, and still prone to producing implausible motions that violate fundamental physical laws. We introduce a training-free framework that improves physical plausibility at inference time by explicitly reasoning about implausibility and guiding the generation away from it. Specifically, we employ a lightweight physics-aware reasoning pipeline to construct counterfactual prompts that deliberately encode physics-violating behaviors. Then, we propose a novel *Synchronized Decoupled Guidance* (SDG) strategy, which leverages these prompts through synchronized directional normalization to counteract lagged suppression and trajectory-decoupled denoising to mitigate cumulative trajectory bias, ensuring that implausible content is suppressed immediately and consistently throughout denoising. Experiments across different physical domains show that our approach substantially enhances physical fidelity while maintaining photorealism, despite requiring no additional training. Ablation studies confirm the complementary effectiveness of both the physics-aware reasoning component and SDG. In particular, the aforementioned two designs of SDG are also individually validated to contribute critically to the suppression of implausible content and the overall gains in physical plausibility. This establishes a new and plug-and-play physics-aware paradigm for video generation."},"_bibtex":{"value":"@misc{\nhao2026enhancing,\ntitle={Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility},\nauthor={Yutong Hao and Chen Chen and Ajmal Saeed Mian and Chang Xu and Daochang Liu},\nyear={2026},\nurl={https://openreview.net/forum?id=SYBfaOcbmw}\n}"},"title":{"value":"Enhancing Physical Plausibility in Video Generation by Reasoning the Implausibility"},"pdf":{"value":"/pdf/6baf9535937fe830983783d56f2d16b21780bd91.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"hao|enhancing_physical_plausibility_in_video_generation_by_reasoning_the_implausibility"},"authorids":{"value":["~Yutong_Hao1","~Chen_Chen39","~Ajmal_Saeed_Mian1","~Chang_Xu4","~Daochang_Liu1"]},"authors":{"value":["Yutong Hao","Chen Chen","Ajmal Saeed Mian","Chang Xu","Daochang Liu"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"TLDR":{"value":"We extract Concept Activation Vectors from a frozen VideoMAE's Physics Emergence Zone and inject them at inference time to achieve the first effective, bidirectional, training-free causal control of physical plausibility in a video world model."},"venue":{"value":"CVPR 2026 Workshop VideoWorldModel Poster"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"pdf":{"value":"/pdf/ae77e7e2e2225918a7f7688c318a4e1e109b67f1.pdf"},"keywords":{"value":["video world models","interpretability","steering","controllable world models","physics reasoning","concept activation vectors"]},"venueid":{"value":"thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track"},"paperhash":{"value":"alam|causal_physics_steering_in_video_world_models_via_concept_activation_vectors"},"authorids":{"value":["~Nahid_Alam1"]},"abstract":{"value":"Video world models learn rich internal representations of physical dynamics, yet steering what physics governs a predicted scene at inference time remains unsolved. Recent interpretability work identified a Physics Emergence Zone (PEZ), a narrow band of middle transformer layers in VideoMAE where physical plausibility is encoded in a direction-centric population code, nearly orthogonal to other visual features. We present physics steering: a training-free method that extracts a Concept Activation Vector (CAV) from a lightweight linear probe at PEZ layers and injects it — scaled by strength $\\alpha$ — into hidden states at inference time, causally shifting the model's physical expectations without modifying any weights. On the IntPhys benchmark (O1/O2/O3), physics steering achieves a 75% flip rate in physical plausibility predictions at $\\alpha = +5$ and drives $P(\\text{impossible})$ to 1.0 at $\\alpha = +10$, with directional purity 1.00 at the PEZ. A layer-specificity ablation confirms that identical interventions at non-PEZ layers produce zero effect (flip rate 0.00 at layers 6–11), establishing the PEZ as causally necessary. Subspace analysis reveals physics and motion direction are encoded at 90° — perfectly orthogonal in representation space — confirming that physics can be steered without corrupting the model's motion representations."},"title":{"value":"Causal Physics Steering in Video World Models via Concept Activation Vectors"},"authors":{"value":["Nahid Alam"]}},"tmdate":1777300201765,"pdate":1777300201209,"tcdate":1772301247719,"writers":["thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track","thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track/Submission5/Authors"],"signatures":["thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track/Submission5/Authors"],"forum":"tPVyo1rEn3","license":"CC BY 4.0","number":5,"cdate":1772301247719,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track/-/Submission","thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track/-/Submission_Change_Before_Bidding","thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track/-/Submission_Change_Before_Reviewing","thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track/-/Submission_Release"],"mdate":1777300201765,"odate":1777300201209,"domain":"thecvf.com/CVPR/2026/Workshop/VideoWorldModel_Proceedings_Track","id":"tPVyo1rEn3","version":2},{"content":{"summary":{"value":"COPHYBENCH is a benchmark for evaluating Video-LLMs on physics-based reasoning from real-world videos. It includes three tasks: (1) causal future prediction, predicting future events from initial observations, (2) physical calculation, estimating quantitative variables such as time and position, and (3) counterfactual reasoning, answering “what if” questions beyond what is directly seen. The dataset contains over 200 real-world physics videos and 1,300 verified QA pairs across various kinematic and dynamic phenomena. Results show that while models perform moderately well on causal prediction, they struggle with quantitative estimation and counterfactual reasoning, revealing limits in physics-grounded understanding."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. How can readers be confident that the benchmark questions are well-posed and solvable rather than underspecified? Presenting human evaluation results could help demonstrate that the tasks are interpretable and have unambiguous answers.\n2. How should zero or near-zero model scores be interpreted? Could they reflect linguistic misunderstanding rather than failure of physical reasoning?\n3. What principle connects the three tasks? Are they intended as representative axes of physics understanding or illustrative examples? How is the benchmark’s overall scope of “physics understanding” defined, and what are its current limitations?\n4. How are symbolic identification and parameter estimation linked to the benchmark tasks? For example, do counterfactual tasks test the ability to identify symbolic relations or estimate parameters? How does the Hamiltonian framework guide the task design or interpretation?\n\nI would raise my rating if the authors ensure task soundness through human validation, clarify the conceptual framework, and establish the scope and limitations of the benchmark clearly."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The paper tackles an important and timely topic, assessing physics-grounded reasoning in Video-LLMs.\n2. The idea of evaluating models from conditional video observations is intuitive and practical, reflecting how humans reason about physical dynamics from partial information.\n3. The dataset is grounded in real-world videos, covering diverse physics phenomena (kinematics, dynamics, optics).\n4. The results, if valid, are insightful in showing that current Video-LLMs succeed at causal prediction but struggle with quantitative estimation and counterfactual reasoning, exposing clear limits in their physics-based understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The conclusions are not fully convincing. Since no model performs well on counterfactual tasks, it is hard to tell whether the issue lies in the task design or the model limitations. The results could reflect either underspecified questions or genuinely weak model reasoning, and the paper should include more examples or human baselines to clarify this.\n2. The benchmark’s underlying principles are not well defined. Each task makes sense individually, but it is unclear whether they collectively form a coherent framework or just cover arbitrary question types. It is also uncertain whether these questions are meant to be complete or complementary to other benchmarks for evaluating physical understanding.\n3. The writing, particularly in Section 3.3, is confusing and lacks coherence between paragraphs. Some claims are overstated, such as insisting that the model “must” solve symbolic identification."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915850265,"tcdate":1761972167162,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1668/Reviewer_wZLx"],"signatures":["ICLR.cc/2026/Conference/Submission1668/Reviewer_wZLx"],"forum":"rDiKG1xlDV","number":3,"license":"CC BY 4.0","cdate":1761972167162,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1668/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915850265,"domain":"ICLR.cc/2026/Conference","replyto":"rDiKG1xlDV","id":"QZLwGsgTj7","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video-LLMs","Physical Video Reasoning","Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present \\textsc{CoPhyBench}, a \\textsc{Co}nditional reasoning \\textsc{Phy}sics-based \\textsc{Bench}mark. \\textsc{CoPhyBench} evaluates the ability of Video-LLMs to reason about physical events based on conditional observations from real-world videos. It probes physics understanding from \nthree perspectives: 1) Prediction: predicting future events from observable cues, assessing a grasp of causality in real-world scenarios. 2) Physical Calculation: estimating times and positions \nby translating visual conditions into variables of dynamics equations. 3) Counterfactual Reasoning: inferring futures based on hypothetical changes, to distinguish between generalizable physical understanding instead of superficial correlations. \nWe construct a high-quality dataset consisting of 1,300 carefully verified question-answer pairs grounded in 232 diverse, real-world physics videos to support these tasks, spanning various phenomena in \nkinematics and dynamics.\nExtensive benchmarking on leading Video-LLMs reveals that while models perform reasonably on causal prediction, they struggle with precise physical calculations and counterfactual reasoning. \nThese findings highlight the limitations of current models in transitioning from semantic alignment to deeper, physics-grounded reasoning, calling for new training paradigms to incorporate physics reasoning. Our dataset and resources will be released."},"_bibtex":{"value":"@misc{\nwei2025cophybench,\ntitle={CoPhyBench: Benchmarking Physical Reasoning from Conditional Video Observation},\nauthor={Fanyue Wei and Kai Xu and Yizhuo Zhang and Pengzhan Sun and Junbin Xiao and Angela Yao},\nyear={2025},\nurl={https://openreview.net/forum?id=rDiKG1xlDV}\n}"},"title":{"value":"CoPhyBench: Benchmarking Physical Reasoning from Conditional Video Observation"},"pdf":{"value":"/pdf/e580d103f2bc862a37822d083ece467bcc215d43.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wei|cophybench_benchmarking_physical_reasoning_from_conditional_video_observation"},"authorids":{"value":["~Fanyue_Wei1","~Kai_Xu7","~Yizhuo_Zhang7","~Pengzhan_Sun1","~Junbin_Xiao1","~Angela_Yao1"]},"authors":{"value":["Fanyue Wei","Kai Xu","Yizhuo Zhang","Pengzhan Sun","Junbin Xiao","Angela Yao"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a fluid simulation pipeline integrating numerical simulation, neural physics and generative control. This one provides better latency than existing physics-based methodologies by only employing numerical simulation when encountering complex fluid dynamics through an automatic fallback. In particular, a GNN-based neural simulator handles low-latency updates, while a fallback to the Material Point Method (MPM) retains accuracy when the dynamics is more complex. In addition, a diffusion-based generative controller enables interactive control by mapping user sketches to external force fields. Experiments show reduced latency (11-29%) and competitive physical fidelity across 2D and 3D settings."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- Can the fallback be learned or adaptive instead of rule-based? E.g. by having a confidence-based trigger instead of a fixed threshold.\n- I did not fully understand what influences the latency reduction (e.g. when to expect 10% rather than 30%). Is this dependent on the complexity of the dynamics?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The method is simple and intuitive, with the fallback mechanism ensuring that costly numerical simulation is only performed when the complexity of the fluid dynamics warrants it. While hybrid pipelines are sometimes under-appreciated in the literature, these often yield the best trade-offs.\n- The latency reduction is significant. Reducing ~10-30% latency while maintaining fidelity is impressive and of practical utility for interactive graphics applications.\n- Interactive control via sketches is visually compelling and, to the best of my knowledge, conceptually new."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- My main concern lies on the novelty of the proposed methodology with respect to PAC-Nerf [1]: although used for different objectives, both this work and Pac-NERF combine a learned neural surrogate with a physics-based simulator to get both efficiency and physical plausibility. This work should be mentioned and thoroughly discussed.\n- While visually compelling, the sketch interface seems more an application layer than a core technical contribution. It’s unclear to me if the diffusion-based control module is novel or simply applied to this domain.\n- The presentation needs some minor adjustment and refinement, with e.g. table 2 floating over section titles.\n- While I am not up to date with all the applicable methods, the set of considered baselines seems somewhat limited. This makes it hard to assess the actual benefits with respect to the state of the art.\n\nConsidering the weaknesses and the strengths, I am inclined to reject at this time. However, as this field is not my primary area of expertise, I am happy to revise my score if these concerns are adequately addressed in the rebuttal or not shared by the other reviewers.\n\n[1] Li, Xuan, et al. \"Pac-nerf: Physics augmented continuum neural radiance fields for geometry-agnostic system identification.\" ICLR 2023"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923843963,"tcdate":1761548721177,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13118/Reviewer_iDWG"],"signatures":["ICLR.cc/2026/Conference/Submission13118/Reviewer_iDWG"],"forum":"6vX0LH9Yt7","number":1,"license":"CC BY 4.0","cdate":1761548721177,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13118/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923843963,"domain":"ICLR.cc/2026/Conference","replyto":"6vX0LH9Yt7","id":"G1GGwDNoGJ","forumContent":{"TLDR":{"value":"We propose a neural physics system for real-time, interactive fluid simulations."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["fluid","simulation","animation","diffusion model"]},"supplementary_material":{"value":"/attachment/6763fe48e6f39912b7ef410d86e5b6ba6622f39c.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"We propose a neural physics system for real-time, interactive fluid simulations. Traditional physics-based methods, while accurate, are computationally intensive and suffer from latency issues. Recent machine‑learning methods reduce computational costs while preserving fidelity; yet most still fail to satisfy the latency constraints for real‑time use and lack support for interactive applications. To bridge this gap, we introduce a novel hybrid method that integrates numerical simulation, neural physics, and generative control. Our neural physics jointly pursues low-latency simulation and high physical fidelity by employing a fallback safeguard to classical numerical solvers.\nFurthermore, we develop a diffusion-based controller that is trained using a revserve modeling strategy to generate external dynamic force fields for fluid manipulation. Our system demonstrates robust performance across diverse 2D/3D scenarios, material types, and obstacle interactions, achieving real-time simulations at high frame rates (11~29% latency reduced) while enabling fluid control guided by\nuser-friendly freehand sketches. We present a significant step towards practical, controllable, and physically plausible fluid simulations for real-time interactive applications. We promise to release both models and data upon acceptance."},"_bibtex":{"value":"@misc{\nxu2026hybrid,\ntitle={Hybrid Neural-{MPM} for Interactive Fluid Simulations in Real-Time},\nauthor={Jingxuan Xu and Hong Huang and Chuhang Zou and Manolis Savva and Yunchao Wei and Wuyang Chen},\nyear={2026},\nurl={https://openreview.net/forum?id=6vX0LH9Yt7}\n}"},"title":{"value":"Hybrid Neural-MPM for Interactive Fluid Simulations in Real-Time"},"pdf":{"value":"/pdf/6b2fdec9d6548190e8d252dcc0273d12abd4b37c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"xu|hybrid_neuralmpm_for_interactive_fluid_simulations_in_realtime"},"authorids":{"value":["~Jingxuan_Xu1","~Hong_Huang5","~Chuhang_Zou1","~Manolis_Savva1","~Yunchao_Wei1","~Wuyang_Chen1"]},"authors":{"value":["Jingxuan Xu","Hong Huang","Chuhang Zou","Manolis Savva","Yunchao Wei","Wuyang Chen"]}},"version":2},{"content":{"venue":{"value":"Complex (1) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02466-5_92.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"zhao|synchronization_in_complex_networks_with_different_sort_of_communities"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Ming_Zhao:","~Tao_Zhou13","https://dblp.org/search/pid/api?q=author:Hui-Jie_Yang:","https://dblp.org/search/pid/api?q=author:Gang_Yan:","https://dblp.org/search/pid/api?q=author:Bing-Hong_Wang:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02466-5_92"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/ZhaoZYYW09,\n  author={Ming Zhao and Tao Zhou and Hui-Jie Yang and Gang Yan and Bing-Hong Wang},\n  title={Synchronization in Complex Networks with Different Sort of Communities},\n  year={2009},\n  cdate={1230768000000},\n  pages={924-933},\n  url={https://doi.org/10.1007/978-3-642-02466-5_92},\n  booktitle={Complex (1)},\n  crossref={conf/complex/2009-1}\n}\n"},"abstract":{"value":"In this paper, inspired by the idea that many real networks are composed by sorts of communities, we investigate the synchronization property of oscillators on such community networks. We identify the communities by two ways, one is by the structure of individual community and the other by the intrinsic frequencies probability density g(ω) of Kuramoto oscillators on different communities. For the two sorts of community networks, when the community structure is strong, only the oscillators on the same community synchronize. With the weakening of the community strength, an interesting phenomenon appears: although the global synchronization is not achieved, oscillators on the same sort of communities will synchronize independently. Global synchronization will appear with the further weakening of community structure."},"title":{"value":"Synchronization in Complex Networks with Different Sort of Communities"},"authors":{"value":["Ming Zhao","Tao Zhou","Hui-Jie Yang","Gang Yan","Bing-Hong Wang"]}},"tmdate":1767616583688,"pdate":1230768000000,"externalIds":["dblp:conf/complex/ZhaoZYYW09"],"tcdate":1767616531738,"writers":["~"],"signatures":["~Tao_Zhou13"],"forum":"jb6ldaornC","license":"CC BY-SA 4.0","number":715670,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1767616583688,"domain":"DBLP.org","id":"jb6ldaornC","version":2},{"content":{"summary":{"value":"This paper introduces a frequency-domain physics-guided framework to improve the physically plausible motion in video diffusion. This paper formulates rigid motions (translation, rotation, and scaling) within a unified spectral SIM(2) framework, and proposes corresponding differentiable frequency-domain losses. Empirically, the proposed method improves motion accuracy and temporal coherence across multiple backbones (Open-Sora, MVDIT, Hunyuan)."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"See the detailed comments in weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- This paper explores an important problem in video generation.\n- The SIM(2)-based spectral derivation unifies translation, rotation, and scaling within a mathematically sound framework.\n- The loss is architecture-agnostic and can be inserted into any diffusion model without modifying the backbone.\n- Evaluation spans three major video diffusion systems and includes multiple metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The innovation lies mainly in unifying them under the SIM(2) formulation.\n- The method only addresses translation, rotation, and scaling, which limits applicability to real-world complex scenes.\n- There are some related physics-constrained video generation works, such as [a], which should also be discussed. Also, except for comparing with the baseline models, it should compare with some related works.\n- [b] is a comprehensive physics generation benchmark designed to evaluate physical commonsense correctness in T2V generation. To validate the effectiveness of the proposed method, it is suggested to apply [b].\n- Since the proposed method is plugged into the baseline models, it should report the generation times.\n- How to set the temperature parameter τ? And how sensitive it is to τ?\n\n\n[a] MOTIONCRAFT: Physics-based Zero-Shot Video Generation\n[b] Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917208026,"tcdate":1761452918948,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4165/Reviewer_2TB5"],"signatures":["ICLR.cc/2026/Conference/Submission4165/Reviewer_2TB5"],"forum":"jhan3NJ5x1","number":1,"license":"CC BY 4.0","cdate":1761452918948,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4165/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917208026,"domain":"ICLR.cc/2026/Conference","replyto":"jhan3NJ5x1","id":"jwLoWU0o4p","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video generation","Diffusion model"]},"supplementary_material":{"value":"/attachment/61e5db89c3e0dcf46f52dfd534527c73711f043e.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Current video diffusion models generate visually compelling content but often violate \nbasic laws of physics, producing subtle artifacts like rubber-sheet deformations and \ninconsistent object motion. We introduce a frequency-domain physics prior that improves \nmotion plausibility without modifying model architectures. Our method decomposes common \nrigid motions (translation, rotation, scaling) into lightweight spectral losses, \nrequiring only 2.7% of frequency coefficients while preserving 97%+ of spectral energy. \nApplied to Open-Sora, MVDIT, and Hunyuan, our approach improves both motion accuracy and action recognition by ~11\\% on average on OpenVID-1M (relative), while maintaining visual quality. User studies show 74--83% preference for our physics-enhanced videos. It also reduces warping error by 22--37% (depending on the backbone) and improves temporal consistency scores. These results indicate that simple, global spectral cues are an effective drop-in regularizer for physically plausible motion in video diffusion."},"_bibtex":{"value":"@misc{\nanonymous2026physicsguided,\ntitle={Physics-Guided Motion Loss for Video Generation Model},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=jhan3NJ5x1}\n}"},"title":{"value":"Physics-Guided Motion Loss for Video Generation Model"},"pdf":{"value":"/pdf/5a7c6e0e967cec1df42675ec0f9003da99cdfd85.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"xue|physicsguided_motion_loss_for_video_generation_model"},"authorids":{"value":["~Bowen_Xue1","~Giuseppe_Claudio_Guarnera1","~Shuang_Zhao1","~Zahra_Montazeri1"]},"authors":{"value":["Bowen Xue","Giuseppe Claudio Guarnera","Shuang Zhao","Zahra Montazeri"]}},"version":2},{"content":{"summary":{"value":"This paper presents a physics-aware PLM reprogramming framework REPST tailored for spatio-temporal forecasting. The proposed REPST consists of a physics-aware spatio-temporal decomposer and a selective reprogrammed language model.  Experiment results confirm that the proposed framework unlocks the capabilities of PLMs to capture fine-grained spatio-temporal dynamics and achieves better performance than existing methods."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is well-written and easy to follow.\n2. The paper introduces a unique approach to enable PLMs to handle spatio-temporal data by using a physics-aware decomposer to disentangle complex spatio-temporal dynamics into components with rich physical semantics.\n3. Extensive experiments are conducted to validate the effectiveness of the proposal."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed decomposer is not clearly described. Since the decompostion is a common tool in time series analysis, there lacks a discussion of why the decomposed components are physics-aware. \n2. The usage of the reconstruction matrix is ambiguous. Further ablation studies are needed.\n3. Most Transformer-based baselines use 96 or longer input lengths. The paper only provides experimental results with an input length of 48, and the experiment under increasing the input length is missing."}},"nonreaders":[],"tmdate":1732591546788,"tcdate":1730119621691,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission732/Reviewer_EE3c"],"signatures":["ICLR.cc/2025/Conference/Submission732/Reviewer_EE3c"],"forum":"wCNuEA5MSv","number":1,"license":"CC BY 4.0","cdate":1730119621691,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission732/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732591546788,"domain":"ICLR.cc/2025/Conference","replyto":"wCNuEA5MSv","id":"LL3wmE1WBH","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["spatio-temporal forecasting","time series forecasting"]},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Spatio-temporal forecasting is pivotal in numerous real-world applications, including transportation planning, energy management, and climate monitoring. \nIn this work, we aim to harness the reasoning and generalization abilities of Pre-trained Language Models (PLMs) for more effective spatio-temporal forecasting, particularly in data-scarce scenarios. \nHowever, recent studies uncover that PLMs, which are primarily trained on textual data, often falter when tasked with modeling the intricate correlations inherent in numerical time series, thereby limiting their effectiveness in comprehending spatio-temporal data.\nTo bridge the gap, we propose REPST, a physics-aware PLM reprogramming framework tailored for spatio-temporal forecasting. \nSpecifically, we first propose a physics-aware decomposer that adaptively disentangles spatially correlated time series into interpretable sub-components, which facilitates PLM’s understanding of sophisticated spatio-temporal dynamics via a divide-and-conquer strategy.\nMoreover, we propose a selective discrete reprogramming scheme, which introduces an expanded spatio-temporal vocabulary space to project spatio-temporal series into discrete representations. This scheme minimizes the information loss during reprogramming and enriches the representations derived by PLMs.\nExtensive experiments on real-world datasets show that the proposed REPST outperforms twelve state-of-the-art baseline methods, particularly in data-scarce scenarios, highlighting the effectiveness and superior generalization capabilities of PLMs for spatio-temporal forecasting."},"_bibtex":{"value":"@misc{\nwang2025language,\ntitle={Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming},\nauthor={Hao Wang and Jindong Han and Wei Fan and Hao Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=wCNuEA5MSv}\n}"},"title":{"value":"Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming"},"pdf":{"value":"/pdf/cb7d921d45bfc48169f16c8f000ca68be9a42b3d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|language_model_empowered_spatiotemporal_forecasting_via_physicsaware_reprogramming"},"authorids":{"value":["~Hao_Wang92","~Jindong_Han1","~Wei_Fan6","~Hao_Liu17"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hao Wang","Jindong Han","Wei Fan","Hao Liu"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["vision-language model","generative video model","intuitive physics"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Vision-language models (VLMs) excel at interpreting what is visible, but can often struggle with future-dependent questions whose answers hinge on events that\nhave not yet occurred. Such questions require anticipating how a scene will unfold—something humans do naturally through mental simulation. We propose a\ntest-time approach that gives VLMs an analogous capability. To complement a\nVLM’s internal text-based reasoning, we use a generative video model to produce\nmultiple plausible future rollouts conditioned on the observed video. These “mental simulations” are fed back to the VLM as additional visual context before it answers. We evaluate this simulate-then-reason framework on three intuitive physics\nbenchmarks: Physion, CLEVRER, and real-world Physics-IQ videos. Across a\ndiverse set of open-source and proprietary VLMs, augmenting models with imagined futures can improve prediction accuracy over both observed-only baselines\nand strong test-time language reasoning baselines, suggesting that allocating test-time computation to explicit video-based mental simulation can enhance multi-modal reasoning about the physical world. Video examples are available on our\nproject page. (https://projectpagexxx.github.io/mental-simulation/)"},"_bibtex":{"value":"@inproceedings{\nanonymous2026mental,\ntitle={Mental Simulation for Vision-Language Models on Intuitive Physics},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=wd2lJCPwYB},\nnote={under review}\n}"},"title":{"value":"Mental Simulation for Vision-Language Models on Intuitive Physics"},"pdf":{"value":"/pdf/2e7e88c6bb234f1e80934b17d249e3b683468fee.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791226651959,"tcdate":1789142893682,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission14485/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission14485/Authors"],"forum":"wd2lJCPwYB","license":"CC BY 4.0","number":14485,"cdate":1789142893682,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission14485/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791226651959,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"wd2lJCPwYB","version":2},{"content":{"summary":{"value":"This paper introduces FEABench, a benchmark for evaluating large language models (LLMs) in physics, mathematics, and engineering tasks via finite element analysis (FEA), with COMSOL Multiphysics as the selected software. The authors point to limited research on LLMs handling complex numerical simulations crucial in these fields.\n\nFEABench includes two datasets: \n- FEABench Gold with 15 human-verified problems offering quantitative targets, and \n- FEABench Large with 200 parsed problems from COMSOL tutorials, which often require plot generation or multi-value computations rather than single, verifiable outcomes.\n\nThe study observes that while LLMs generate executable code, they face challenges like selecting appropriate physics interfaces and avoiding inaccuracies. To address this, the authors designed a multi-turn LLM agent system to iteratively refine solutions via COMSOL API interactions, feedback, and specialized sub-agents. This agent includes: an algorithmic call sequence to reduce errors, combined LLM and API feedback for consistency checks, and methods to sustain correction attempts despite tool call failures. The agent system consists of sub-agents, such as a Controller for solution generation, an Evaluator for feedback, a Corrector for solution adjustments, and a Tool Lookup Agent for retrieving information.\n\nDespite significant progress in execution accuracy, complete and correct solutions to benchmark problems are still to be achieved."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"I would appreciate if the authors could elaborate on the novelty of their agent design, especially in comparison to the extensive existing literature on LLM agent development. Specifically, clarifying how their approach differs in structure, feedback mechanisms, or task handling would help understand the potential algorithmic advancements introduced in this work."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Challenging Benchmark: FEABench introduces a novel and complex benchmark for assessing LLMs in physics, mathematics, and engineering tasks using FEA software, offering a way to evaluate these models in real-world problem-solving scenarios.\n- Evaluation: FEABench goes beyond correctness checks by employing a multi-dimensional evaluation strategy, providing a more detailed view of LLM capabilities.\n- Agent-Based Framework: The multi-turn agent approach illustrates the potential of combining LLMs with feedback systems and tool integration to tackle complex, iterative tasks more effectively. The agent reaches an executability score of 0.88 (88% of the generated code lines are executable). This is considerably higher than the single-shot approaches. The agent shows substantial improvements in interface factuality, interface recall, and feature recall, indicating a better understanding and use of physics concepts in the code. The agent successfully computes a \"Valid Target\" in two out of fifteen problems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited Success: While the agent framework in FEABench improves code executability, this is not sufficient for overall benchmark success. While the agent computes valid target values in two problems, only one of those values falls within 10% of the ground truth answer. This is a crucial metric for determining true problem-solving success. Table 5 shows that the agent-based approach does not lead to significant improvements in code similarity compared to single-shot prompts. The agent showed only modest improvement in accurately setting properties of physics features, and it failed to fully solve any FEABench problem. This raises questions about whether the agent's search strategy is well-suited for tackling the benchmark's challenges.\n- Absence of Agent Baselines: Although the paper proposes an agent-based approach, it does not provide a comparative analysis against existing (rich set of) LLM agent frameworks, making it difficult to assess improvements in code generation or overall performance. As a result, the primary contribution seems to be limited to the introduction of the FEABench dataset.\n- The iterative nature of the agent-based approach involves multiple cycles of code generation, execution, feedback analysis, and code correction. This process can be computationally expensive, particularly when dealing with complex FEA problems that require significant computational resources for simulation."}},"nonreaders":[],"tmdate":1732860177565,"tcdate":1730690434150,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12171/Reviewer_35U5"],"signatures":["ICLR.cc/2025/Conference/Submission12171/Reviewer_35U5"],"forum":"hDkLpu1E64","number":4,"license":"CC BY 4.0","cdate":1730690434150,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12171/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732860177565,"domain":"ICLR.cc/2025/Conference","replyto":"hDkLpu1E64","id":"l7kYMQS6xC","forumContent":{"TLDR":{"value":"How well can LLMs leverage FEA software to simulate and solve problems that require numerical analysis?"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["numerical analysis","finite element","benchmark","agents"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building precise simulations of the real world and invoking numerical solvers to answer quantitative problems is an essential requirement in engineering and science. We present FEABench, a benchmark to evaluate the ability of large language models (LLMs) and LLM agents to simulate and solve physics, mathematics and engineering problems using finite element analysis (FEA). We introduce a multipronged evaluation scheme to investigate the ability of LLMs to solve these problems by reasoning over natural language problem descriptions and operating COMSOL Multiphysics$^\\textregistered$, an FEA software, to compute the answers. In addition to testing state-of-the art-LLMs, we further design a language model agent equipped with the ability to interact with the software through its Application Programming Interface (API), examine its outputs and use tools to improve its solutions over multiple iterations. Our best performing strategy generates executable API calls 88\\% of the time. However, this benchmark still proves to be challenging enough that the LLMs and agents we tested were not able to completely and correctly solve any problem. LLMs that can successfully interact with and operate FEA software to solve problems such as those in our benchmark would significantly push the frontiers of their utility. Acquiring this capability would augment LLMs' reasoning skills with the precision of numerical solvers and advance the development of autonomous systems that can tackle complex problems in the real world."},"_bibtex":{"value":"@misc{\nmudur2025feabench,\ntitle={{FEAB}ench: Evaluating Language Models on Real World Physics Reasoning Ability},\nauthor={Nayantara Mudur and Hao Cui and Subhashini Venugopalan and Paul Raccuglia and Michael Brenner and Peter Christian Norgaard},\nyear={2025},\nurl={https://openreview.net/forum?id=hDkLpu1E64}\n}"},"title":{"value":"FEABench: Evaluating Language Models on Real World Physics Reasoning Ability"},"pdf":{"value":"/pdf/3e64111fb86b7cbb5ef6469de0f077b416722ed3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"mudur|feabench_evaluating_language_models_on_real_world_physics_reasoning_ability"},"authorids":{"value":["~Nayantara_Mudur1","~Hao_Cui3","~Subhashini_Venugopalan2","~Paul_Raccuglia1","~Michael_Brenner1","~Peter_Christian_Norgaard1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Nayantara Mudur","Hao Cui","Subhashini Venugopalan","Paul Raccuglia","Michael Brenner","Peter Christian Norgaard"]}},"version":2},{"content":{"summary":{"value":"The paper proposes APEX, a plug-and-play loop that (i) constructs a relational scene graph from two snapshots, (ii) applies a “difference-graph” attention module (a Graphormer variant) to identify salient interactions, (iii) enumerates candidate actions, (iv) performs forward simulation rollouts with a physics engine for each action, and (v) asks an LLM to select an action using textual rollout summaries. Experiments cover synthetic physics QA, a custom short-horizon Tetris setup, a dynamic obstacle-avoidance toy task, PHYRE, and a small “real-world” vignette."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"Can you precisely define (\\Delta G) for heterogeneous edge types and features. How are ((G_t, G_{t+\\Delta t})) fused in attention (encodings, positional terms, pairwise features)?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Pragmatic integration of a physics engine into an LLM loop; the overall system is easy to understand and replicate in spirit.\n\n- Clear problem motivation: teaching LLM agents to rely on external physics tools rather than internalizing fragile physical heuristics is a sensible direction.\n\n- Breadth of tasks (synthetic QA, games/toy control, PHYRE) demonstrates that the loop can be wired up across settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- limited methodological novelty\n- single-step lookahead: the algorithmic core is a 1-step brute-force evaluation of enumerated actions; claims and tables about multi-step/higher-horizon complexity are not matched by controlled, implemented evidence.\n- Tetris lacks standard hand-coded/Tetris-AI baselines (e.g., height/holes/bumpiness heuristics, MCTS with lookahead).\n- Obstacle avoidance omits classic MPC/DWA/A*/RRT with the same simulator and budget.\n- PHYRE relies on 10k uniformly random actions, effectively reducing to random search rather than demonstrating LLM-guided planning"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924232973,"tcdate":1762229720229,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13665/Reviewer_ZBxi"],"signatures":["ICLR.cc/2026/Conference/Submission13665/Reviewer_ZBxi"],"forum":"ROB3ALLKIX","number":4,"license":"CC BY 4.0","cdate":1762229720229,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13665/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924232973,"domain":"ICLR.cc/2026/Conference","replyto":"ROB3ALLKIX","id":"cuByninYji","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We build a framework that helps LLMs quantify when a cat will collide with them and assess the physical outcomes of different escape routes (e.g., whether they’ll crash into a table) by feeding them physics-based simulations."},"keywords":{"value":["Physics-Enhanced LLMs","Graph-Based Perception","Task Planning","Predictive Simulation"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Large Language Models (LLMs) demonstrate strong reasoning and task planning capabilities but remain fundamentally limited in physical interaction modeling. Existing approaches integrate perception via Vision-Language Models (VLMs) or adaptive decision-making through Reinforcement Learning (RL), but they fail to capture dynamic object interactions or require task-specific training, limiting their real-world applicability.\nWe introduce APEX (Anticipatory Physics-Enhanced Execution), a framework that equips LLMs with physics-driven foresight for real-time task planning. APEX constructs structured graphs to identify and model the most relevant dynamic interactions in the environment, providing LLMs with explicit physical state updates. Simultaneously, APEX provides low-latency forward simulations of physically feasible actions, allowing LLMs to select optimal strategies based on predictive outcomes rather than static observations.\nWe evaluate APEX on three benchmarks designed to assess perception, prediction, and decision-making: (1) Physics Reasoning Benchmark, testing causal inference and object motion prediction; (2) Tetris, evaluating whether physics-informed prediction enhances decision-making performance in long-horizon planning tasks; (3) Dynamic Obstacle Avoidance, assessing the immediate integration of perception and action feasibility analysis. APEX significantly outperforms standard LLMs and VLM-based models, demonstrating the necessity of explicit physics reasoning for bridging the gap between language-based intelligence and real-world task execution."},"_bibtex":{"value":"@misc{\nhuang2025apex,\ntitle={{APEX}: Empowering {LLM}s with Physics-Based Task Planning for Real-time Insight},\nauthor={Wanjing Huang and Weixiang Yan and Zhen Zhang and Ambuj Singh},\nyear={2025},\nurl={https://openreview.net/forum?id=ROB3ALLKIX}\n}"},"title":{"value":"APEX: Empowering LLMs with Physics-Based Task Planning for Real-time Insight"},"pdf":{"value":"/pdf/0529797af939423f17dcaccea10909babb2a5c4d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|apex_empowering_llms_with_physicsbased_task_planning_for_realtime_insight"},"authorids":{"value":["~Wanjing_Huang1","~Weixiang_Yan1","~Zhen_Zhang16","~Ambuj_Singh1"]},"authors":{"value":["Wanjing Huang","Weixiang Yan","Zhen Zhang","Ambuj Singh"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a pipeline to use LLMs for designing two-phase BCC/B2 superalloys. The paper starts from a thermo-calc–generated dataset, combining 200+ BCC compositions and ~90 B2 compositions into ~50k synthetic training instances. They first supervised finetuned several open LLMs (LLaMA-3.1, Gemma-2, OLMo-2) to reliably produce compositions in this structured schema. Then they sample new candidates from the SFT models, evaluate them with Thermo-Calc over a temperature range, and build a hierarchical, physics-based reward that prioritizes: (1) two-phase BCC+B2 presence, (2) BCC solidifying first, (3) B2 still present at room temperature, and (4) wide two-phase stability window. These rewards are converted into pairwise preferences and used to run DPO-style preference tuning."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Are all preference pairs derived from the same Thermo-Calc database/settings that were used to generate the SFT data? If so, can you provide any cross-tool or cross-database validation to show the model is not overfitting one CALPHAD implementation?\n\n2. Some base models (e.g., OLMo-2) do not benefit from physics-based DPO. Did you try alternative preference-learning setups ($\\beta$-DPO, listwise preferences, GRPO/PPO with the same reward) to stabilize training across models?\n\n3. In Table comparisons, API LLMs with good prompt engineering still look competitive. Do those API models produce samples that are judged better by Thermo-Calc because they encode extra “chemical intuition” not present in your synthetic SFT data? If yes, can high-reward API samples be distilled back into the local model?\n\n4. Since you claim the framework is general, can you show at least one non-BCC/B2 or non-superalloy example (even a small one) to demonstrate that replacing Thermo-Calc with another physics/simulation engine would still work?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Instead of materials generation in general, the paper targets “one-shot generation of a BCC composition, a B2 composition, and a B2 volume fraction” for a concrete, high-value family of alloys (BCC matrix + ordered B2 precipitates), which makes the problem easy to evaluate.\n\n2. Using physics-generated preferences (from Thermo-Calc) to do DPO on top of an SFT model is a reasonable and novel adaptation of recent preference-learning/LLM-alignment techniques to materials design.\n\n3. The paper actually prompts large API models and even uses prompt-optimization (e.g., DSPy/MIPRO-like) to show that generic LLMs tend to “hyper-fixate” on a narrow element set, whereas the proposed pipeline preserves diversity.\n\n4. The overall pipeline: synthetic SFT data from a simulator $\\rightarrow$ sampling $\\rightarrow$ physics-based reward $\\rightarrow$ pairwise DPO could be reused for other material design problems that have an automated simulator."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The same thermo-calc setup is used to (1) synthesize SFT data, (2) generate preference pairs, and (3) evaluate success. This makes it hard to tell whether the model learned transferable “materials knowledge” or just learned to speak to this particular CALPHAD database. Cross-simulator or literature-based sanity checks are missing.\n\n2.  Only some base models (e.g., LLaMA, Gemma) improve stably after physics-DPO, while others (e.g., OLMo-2) degrade, which suggests the pipeline is sensitive to model architecture or to off-distribution sampling. This weakens the “general framework” claim.\n\n3. The paper mostly reuses validity/coverage/novelty-style metrics from crystal/material generation, but the target here includes a continuous volume fraction and a thermo-mechanical notion of “two-phase stability over a temperature sweep.” A metric that directly measures the percentage of candidates that satisfy all four thermo-based sub-objectives would be more convincing.\n\n\n4. No experimental or literature recovery check. There is no demonstration that the model can rediscover known BCC/B2 alloy compositions from the literature, or that any predicted alloy is plausible under process constraints (cooling rate, aging). That makes the current contribution look more like a simulation-level template than an end-to-end materials design result."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762999983469,"tcdate":1761981274758,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20860/Reviewer_jX7K"],"signatures":["ICLR.cc/2026/Conference/Submission20860/Reviewer_jX7K"],"forum":"nEF9q1UmEZ","number":3,"license":"CC BY 4.0","cdate":1761981274758,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20860/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762999983469,"domain":"ICLR.cc/2026/Conference","replyto":"nEF9q1UmEZ","id":"2nuFFbVs7c","forumContent":{"TLDR":{"value":"We use preference learning to optimize local LMs to generate candidate compositions for BCC/B2 alloys"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Direct preference optimization","preference learning","materials science","alloys","language models"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"We apply preference learning to the task of language model generation of novel structural alloys. Where prior work focuses on generating stable inorganic crystals, our approach optimizes for the synthesizeability of a specific structural class: BCC/B2 superalloys, an underexplored family of materials with applications in extreme environments. Using three open-weight models (LLaMA-3.1, Gemma-2, and OLMo-2), we demonstrate that language models can be optimized for multiple design objectives using a single, unified reward signal through Direct Preference Optimization (DPO). Our reward signal is derived from thermodynamic phase calculations, offering a scientifically-grounded feedback for model tuning. To our knowledge, this is the first demonstration of preference-tuning a language model using physics-grounded feedback for targeted properties (in our case, BCC/B2 alloys). The resulting framework is general and adaptable to any design problem for which the design space is enumerable and simulation-based feedback is available."},"_bibtex":{"value":"@misc{\nghosh2025preference,\ntitle={Preference Learning from Physics-Based Feedback: Tuning Language Models to Design {BCC}/B2 Superalloys},\nauthor={Satanu Ghosh and Collin Holgate and Neal R Brodnik and Doug Downey and Samantha Daly and Tresa Pollock and Samuel Carton},\nyear={2025},\nurl={https://openreview.net/forum?id=nEF9q1UmEZ}\n}"},"title":{"value":"Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys"},"pdf":{"value":"/pdf/e7e19aee2f2570a808b1d3bdd813d60c2e51060f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ghosh|preference_learning_from_physicsbased_feedback_tuning_language_models_to_design_bccb2_superalloys"},"authorids":{"value":["~Satanu_Ghosh1","~Collin_Holgate1","~Neal_R_Brodnik1","~Doug_Downey1","~Samantha_Daly2","~Tresa_Pollock2","~Samuel_Carton1"]},"authors":{"value":["Satanu Ghosh","Collin Holgate","Neal R Brodnik","Doug Downey","Samantha Daly","Tresa Pollock","Samuel Carton"]}},"version":2},{"content":{"summary":{"value":"This paper presents a novel framework for generating 3D cinemagraphs from a set of multi-view static images. The key contribution is the integration of a physics-based simulation directly onto a 3D Gaussian Splatting (3D-GS) scene representation. Instead of relying on learned motion priors, the method first reconstructs a scene using 3D-GS. It then treats the individual Gaussians as a system of physical mass points. To manage complexity and enforce structural coherence, these Gaussians are clustered into \"SuperGaussians,\" which are approximated as rigid bodies. The animation is driven by a simulation that accounts for user-defined external forces (e.g., wind, spiral fields) and internal structural forces (elasticity and damping) that propagate through a sparse, locality-aware constraint graph. Finally, a trajectory blending technique is used to render a seamlessly looping video. The proposed method aims to produce physically plausible, interpretable, and highly controllable animations that surpass the realism of existing techniques."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"N/A"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The central contribution is non-trivial: merging a classical, interpretable physics simulation with a modern neural scene representation like 3D-GS. This approach moves away from black-box generative models for motion and toward a system grounded in first principles, which is a compelling research direction for controllable and physically-aware generation.\n- A direct benefit of the physics-based approach is that the resulting motion is both controllable and interpretable. Users can manipulate intuitive, high-level physical parameters like force fields, stiffness ($k$), and damping ($\\zeta$) to direct the animation, rather than navigating a complex latent space. The resulting motion can be understood through the lens of classical mechanics, which is a significant advantage over purely data-driven methods.\n- The clustering of primitives into \"SuperGaussians\" is a very intelligent way to manage the immense complexity of simulating millions of individual Gaussians. It provides a hierarchical structure that improves both computational efficiency and the structural coherence of the motion.\n- The use of a sparse, local \"constraint graph\" to model internal forces is a reasonable design choice. It reflects the local nature of real-world physical interactions and avoids the computational expense and potential instability of a fully-connected system."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The underlying physics is essentially a damped mass-spring system applied to rigid clusters. While effective for the subtle motions required for cinemagraphs, this is a very simplified model. The paper's own conclusion acknowledges that it cannot handle complex phenomena like exaggerated swinging or fluid-like dynamics. Furthermore, the framework does not appear to support crucial physical interactions such as collisions, fracture, or non-rigid deformations beyond simple elasticity, which limits its application to a narrow range of effects.\n- The method is demonstrated on scenes containing single, relatively isolated objects against clean backgrounds (e.g., NeRF synthetic data). It is unclear how the framework would scale to slightly larger, more complex, and cluttered real-world scenes. The SuperGaussian clustering and constraint graph construction could become significantly more challenging and potentially produce less meaningful results in scenes with many interacting or overlapping objects.\n- The paper claims strong user control, but the mechanism for applying forces seems to be at a high level (e.g., defining a global wind field). It is not clear how a user could apply a localized force to a specific semantic part of an object (e.g., pushing a single branch on the Ficus tree). This would presumably require an additional layer for semantic segmentation and selection of SuperGaussians, which is not discussed. There is no supplement demo shown at all, only text appendix."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918859892,"tcdate":1761983940108,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6483/Reviewer_Tp4o"],"signatures":["ICLR.cc/2026/Conference/Submission6483/Reviewer_Tp4o"],"forum":"byVg1nJRFt","number":4,"license":"CC BY 4.0","cdate":1761983940108,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6483/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918859892,"domain":"ICLR.cc/2026/Conference","replyto":"byVg1nJRFt","id":"3Gs1JtUtDs","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Cinemagraph"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"3D Cinemagraphs aim to generate visually compelling media by introducing subtle and continuous motion into otherwise static images. Recent efforts have explored this task through 3D reconstruction techniques, but they often fall short in delivering physically plausible and controllable animations. In this paper, we propose a novel physics-driven framework built upon 3D Gaussian Splatting (3DGS) to address these limitations. Given multi-view images and a user-specified force, our approach first reconstructs a 3D scene using 3DGS, then embeds the reconstruction into a physically consistent simulation environment. By modeling external and internal force fields and performing accurate force analysis within the reconstructed 3D space, we synthesize fine-grained and interpretable motion that aligns with physical intuition. Our method allows users to intuitively control motion effects via high-level physical parameters, achieving a delicate balance between realism and artistic flexibility. Extensive experiments under diverse force conditions demonstrate that our approach produces stable, interpretable, and visually appealing results, surpassing prior methods in both robustness and controllability."},"_bibtex":{"value":"@misc{\ncheng2026animating,\ntitle={Animating the Still: Physics-Based 3D Cinemagraph from Multi-View Images Using 3D Gaussian Splatting},\nauthor={Lechao Cheng and Jiyang Li and beihuang and Shengeng Tang and Zhangye Wang and Zhihui Yang and Jingxuan He},\nyear={2026},\nurl={https://openreview.net/forum?id=byVg1nJRFt}\n}"},"title":{"value":"Animating the Still: Physics-Based 3D Cinemagraph from Multi-View Images Using 3D Gaussian Splatting"},"pdf":{"value":"/pdf/3c56b5712ab511541dd90d173a1fb730d9ebd11e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"cheng|animating_the_still_physicsbased_3d_cinemagraph_from_multiview_images_using_3d_gaussian_splatting"},"authorids":{"value":["~Lechao_Cheng2","~Jiyang_Li1","~beihuang1","~Shengeng_Tang1","~Zhangye_Wang1","~Zhihui_Yang3","~Jingxuan_He2"]},"authors":{"value":["Lechao Cheng","Jiyang Li","beihuang","Shengeng Tang","Zhangye Wang","Zhihui Yang","Jingxuan He"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PRISM-PHYSICS, a framework for evaluating physics problems at the process level. Solutions are represented as directed acyclic graphs (DAGs) to capture dependencies between steps. The key innovation is the ancestor-closure scoring, which allocates partial credit based on intermediate steps. A rule-based symbolic equivalence checker ensures accurate comparison of formulas. Experimental results show that PRISM-PHYSICS provides detailed and reliable evaluations."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- PRISM-PHYSICS provides a large-scale, competition-level benchmark with carefully curated, DAG-structured solutions to complex physics problems.\n- A fully rule-based symbolic equivalence checker ensures consistent validation of diverse mathematical expressions, eliminating reliance on heuristic LLM scoring and offering a more reliable comparison across alternative formulations.\n- The ancestor-closure scoring policy allows for partial credit on intermediate steps, offering a more nuanced and fair assessment of student reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Does the system account for context-dependent variations in formulas? (e.g., solving a problem from both kinematics and dynamics perspectives, or analyzing it through momentum and energy considerations)\n- In physics problems, certain expressions may be contextually equivalent, but the strict analysis in this algorithm might overlook such context-dependent equivalence.   Does the current framework account for these context-sensitive variations?\n- If skips over intermediate, simpler steps during the solution process, would this result in incorrect evaluation by the proposed method?  \n- The summary of the experiment is not clear enough. It is hoped that the author will have a more clear and organized discussion of the experimental results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921361258,"tcdate":1761640894109,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9900/Reviewer_PRSU"],"signatures":["ICLR.cc/2026/Conference/Submission9900/Reviewer_PRSU"],"forum":"4PZMeopXzP","number":2,"license":"CC BY 4.0","cdate":1761640894109,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9900/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921361258,"domain":"ICLR.cc/2026/Conference","replyto":"4PZMeopXzP","id":"EKzfqWWsYD","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We present PRISM-Physics, a benchmark and a process-level evaluation framework that encodes physics solutions as DAGs and employs rule-based symbolic equivalence checking for reliable, fine-grained scoring."},"keywords":{"value":["Physics Reasoning","Process-Level Evaluation","Symbolic Equivalence","Scientific Problem Solving"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Benchmarks for competition-style reasoning have advanced evaluation in mathematics and programming, yet physics remains comparatively underexplored. Most existing physics benchmarks evaluate only final answers, which fail to capture reasoning processes, while recent stepwise methods rely on heuristic LLM-as-judge scoring or restrictive linear assumptions, limiting reliability and diagnostic validity.\nWe introduce PRISM-Physics, a process-level evaluation framework and benchmark for complex physics reasoning problems. Solutions are represented as directed acyclic graphs (DAGs) of formulas, explicitly encoding causal dependencies among intermediate steps to enable fine-grained, interpretable, and theoretically grounded scoring. \nWe prove the optimality of the DAG representation and the corresponding scoring policy. Combining with a fully rule-based method for symbolic formula equivalence matching that we developed, we ensure consistent validation across diverse formulations without heuristic judgments. Results show that our evaluation framework is more aligned with human experts' scoring. \nExperiments on state-of-the-art LLMs reveal persistent reasoning failures in physics, while step-level scoring offers both diagnostic insight and rich signals for later training. By combining structural rigor, theoretical guarantees, and symbolic validation, PRISM-Physics provides a principled foundation for advancing process-level evaluation and guiding the development of models with deeper scientific reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\nzhao2026prismphysics,\ntitle={{PRISM}-Physics: Causal {DAG}-Based Process Evaluation for Physics Reasoning},\nauthor={Wanjia Zhao and Qinwei Ma and Jingzhe Shi and Shirley Wu and Jiaqi Han and Yijia Xiao and Si-Yuan Chen and Xiao Luo and Ludwig Schmidt and James Zou},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=4PZMeopXzP}\n}"},"title":{"value":"PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning"},"pdf":{"value":"/pdf/95b751e95b88437e4484cd3de1a315d0a89884f4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhao|prismphysics_causal_dagbased_process_evaluation_for_physics_reasoning"},"authorids":{"value":["~Wanjia_Zhao1","~Qinwei_Ma1","~Jingzhe_Shi1","~Shirley_Wu1","~Jiaqi_Han2","~Yijia_Xiao1","~Si-Yuan_Chen1","~Xiao_Luo3","~Ludwig_Schmidt1","~James_Zou1"]},"authors":{"value":["Wanjia Zhao","Qinwei Ma","Jingzhe Shi","Shirley Wu","Jiaqi Han","Yijia Xiao","Si-Yuan Chen","Xiao Luo","Ludwig Schmidt","James Zou"]}},"version":2},{"content":{"summary":{"value":"This paper aims to integrate a spatio-temporal graph neural network with physics-aware causality for spatio-temporal modeling. The major contribution is the soft integration of physics equations with causality. Experiments over several synthetic and real-world datasets can verify the effectiveness of the proposed model."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The paper is well-written and easy to follow. Integrating spatio-temporal graph neural network is of great importance to many real-world applications.\n2. The paper conducts experiments over both synthetic and real-world datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My major concerns are:\n1. Insufficient related work. To the best of my knowledge, there is quite a large number of literature exploring the integration of physics law or causality into spatio-temporal graph neural networks [1,2,3,4,5,6]. For example, [1,2,6] employ neural ordinary differential equations to capture continuous ST dependencies. Ji et al. propose a physics-guided neural network for spatiotemporal modeling in traffic flows [3]. CaST designs a new framework for handling causality in spatio-temporal graphs [4]. However, this paper lacks a discussion on these studies and doesn't compare the proposed model with them either. What's the difference between them? Why should we use the proposed model? It would be good to survey more related publications before paper submission. What's more, the related work section should be included in the main body of the paper, instead of the appendix.\n2. The technical contribution of this work against existing approaches is not significant, which is clearly below the acceptance level of ICLR.\n3. The term \"causality\" in this paper is questionable. This causality is more similar to proximity in other spatio-temporal graph neural networks [7, 8], rather than the actual causality in causal inference. \n4. The learned causality in this paper lacks justification.\n5. The baselines used in this paper are weak and outdated. Please consider more recent baselines for comparison (see the above references). \n6. This paper lacks the experiment over one of the most popular tasks -- traffic forecasting, which is also driven by inherent physics laws.\n7. No source code for reproducing the results.\n8. No discussion on the model efficiency and model size.\n\nReference:\n\n[1] Fang, et al. \"Spatial-temporal graph ode networks for traffic flow forecasting.\" SIGKDD 2021.\n\n[2] Choi,et al. \"Graph neural controlled differential equations for traffic forecasting.\" AAAI 2022.\n\n[3] Ji et al. \"STDEN: Towards physics-guided neural networks for traffic flow prediction.\" AAAI 2021.\n\n[4] Xia et al. \"Deciphering Spatio-Temporal Graph Forecasting: A Causal Lens and Treatment.\" NeurIPS 2023.\n\n[5] Jia et al. \"Physics-guided recurrent graph model for predicting flow and temperature in river networks.\" SDM 2021.\n\n[6] Liang et al. \"Mixed-order relation-aware recurrent neural networks for spatio-temporal forecasting.\" TKDE 2022.\n\n[7] Wu et al. \"Graph wavenet for deep spatial-temporal graph modeling.\" IJCAI 2019.\n\n[8] Bai et al. \"Adaptive graph convolutional recurrent network for traffic forecasting.\" NeurIPS 2020."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Please reply to the questions in the weaknesses."},"rating":{"value":"3: reject, not good enough"},"details_of_ethics_concerns":{"value":"No"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636693845,"tcdate":1699168426325,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6312/Reviewer_pHbu"],"signatures":["ICLR.cc/2024/Conference/Submission6312/Reviewer_pHbu"],"forum":"2uHTuvDkLZ","number":4,"license":"CC BY 4.0","cdate":1699168426325,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6312/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636693845,"domain":"ICLR.cc/2024/Conference","replyto":"2uHTuvDkLZ","id":"OXuxaN1Vkv","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["physics-informed deep learning; causal learning; spatiotemporal learning"]},"supplementary_material":{"value":"/attachment/5873a933b20dc180f679318946973106e5a5d0fa.pdf"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Interpretable physics equations are widely recognized as valuable inductive biases for constructing robust spatiotemporal models. To harness these valuable pieces of knowledge, existing approaches often presuppose access to the exact underlying equations. However, such an assumption usually doesn't hold, especially in the context of real-world observations. Conversely, causality systematically captures the fundamental causal relations across space and time that are intrinsically present in physics dynamics. Nevertheless, causality is often ignored as a means of integrating prior physics knowledge. In this work, we propose a novel approach that effectively captures and leverages causality to integrate physics equations into spatiotemporal models, without assuming access to precise physics principles. \nSpecifically, we introduce a physics-aware spatiotemporal causal graph network (P-stCGN). Causal relationships are analytically derived from prior physics knowledge and serve as physics-aware causality labels. A causal module is introduced to learn causal weights from spatially close and temporally past observations to current observations via semi-supervised learning. Given the learned causal structure, a forecasting module is introduced to perform predictions guided by the cause-effect relations. Extensive experiments on time series data show that our semi-supervised causal learning approach is robust with noisy and limited data. Furthermore, our evaluations on real-world graph signals demonstrate superior forecasting performance, achieved by utilizing prior physics knowledge from a causal perspective."},"_bibtex":{"value":"@misc{\nseo2024physicsaware,\ntitle={Physics-aware Causal Graph Network for Spatiotemporal Modeling},\nauthor={Sungyong Seo and Zijun Cui and Sam Griesemer and Joshua Hikida and Yan Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=2uHTuvDkLZ}\n}"},"title":{"value":"Physics-aware Causal Graph Network for Spatiotemporal Modeling"},"pdf":{"value":"/pdf/b680101f5e19ac6f8f83c61e3ae0f38cbe0cd7bc.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"seo|physicsaware_causal_graph_network_for_spatiotemporal_modeling"},"authorids":{"value":["~Sungyong_Seo1","~Zijun_Cui1","~Sam_Griesemer1","joshua.hikida@gmail.com","~Yan_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sungyong Seo","Zijun Cui","Sam Griesemer","Joshua Hikida","Yan Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes *PhyMix*, a physics-guided framework for single-image 3D indoor scene generation. The authors introduce a unified *Physics Evaluator* consisting of four aspects and nine measurable physical constraints, and integrate its feedback into both training (implicit preference alignment using Scene-GRPO) and inference (explicit Test-Time Optimization). Experiments on 3D-FRONT show consistent improvements in physical plausibility and geometric fidelity, with qualitative results across various image domains."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"In Table 2, bolding the best results would make it easier for readers to compare methods."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The proposed Physics Evaluator provides a comprehensive and unified measurement of physical consistency, covering contact, stability, geometric priors, and deployability.\n2. The combination of implicit optimization (Scene-GRPO) and explicit refinement (TTO) is conceptually elegant and appears effective in improving physical plausibility.\n3. The method generalizes to multiple input domains (real, synthetic, cartoon, LLM-generated), showing robustness and practical applicability.\n4. The paper is overall well-written and should be easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The training pipeline depends on the Physics Evaluator, and some evaluation components (especially simulation-based stability $P_{sim}$ ) can be computationally expensive. The paper lacks a clear comparison of training/inference time and compute cost relative to baselines.\n2. The Physics Evaluator contains many hyperparameters (as discussed in the appendix). It is not clear whether these hyperparameters are object-category dependent, or how sensitive the evaluator is to different scene compositions. More justification on robustness across object types is needed.\n3. Qualitative comparisons in the main paper are limited. Given that physical consistency often manifests in motion or interaction, videos could better reflect physical plausibility. Currently, no supplementary video materials are provided."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915649654,"tcdate":1761718862209,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission968/Reviewer_5eD1"],"signatures":["ICLR.cc/2026/Conference/Submission968/Reviewer_5eD1"],"forum":"RK6j9cwSK4","number":2,"license":"CC BY 4.0","cdate":1761718862209,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission968/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915649654,"domain":"ICLR.cc/2026/Conference","replyto":"RK6j9cwSK4","id":"9FWKnOjRiH","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"PhyMix couples a Physics Evaluator with Scene-GRPO and test-time optimization to produce physically consistent single-image 3D indoor scenes."},"keywords":{"value":["3D Scene Generation","Indoor Scene Synthesis","Physics-Aware Optimization"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their reliability in robotics, embodied AI, and design. To examine this gap, we introduce a unified Physics Evaluator that measures four main aspects: contact, stability, geometric priors, and deployability, which are further decomposed into nine sub-constraints, establishing the first benchmark to measure physical consistency. Based on this evaluator, our analysis shows that state-of-the-art methods remain largely physics-unaware. To overcome this limitation, we further propose a framework that integrates feedback from the Physics Evaluator into both training and inference, enhancing the physical plausibility of generated scenes. Specifically, we propose PhyMix, which is composed of two complementary components: (i) implicit alignment via Scene-GRPO, a critic-free group-relative policy optimization that leverages the Physics Evaluator as a preference signal and biases sampling towards physically feasible layouts, and (ii) explicit refinement via a plug-and-play Test-Time Optimizer (TTO) that uses differentiable evaluator signals to correct residual violations during generation. Overall, our method unifies evaluation, reward shaping, and inference-time correction, producing 3D indoor scenes that are both visually faithful and physically plausible. Extensive evaluations on synthetic dataset confirm state-of-the-art performance in both visual fidelity and physical plausibility, and extensive qualitative examples on stylized and real-world images further showcase the method’s robustness."},"_bibtex":{"value":"@misc{\nwu2026phymix,\ntitle={PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit{\\textendash}Explicit Optimization},\nauthor={Dongli Wu and Jingyu Hu and Ka-Hei Hui and Xiaobao Wei and Zhengzhe Liu and Jianqiang Li},\nyear={2026},\nurl={https://openreview.net/forum?id=RK6j9cwSK4}\n}"},"title":{"value":"PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit–Explicit Optimization"},"pdf":{"value":"/pdf/a08871c55f476dfd4fca9f88ae466e84b0ec4927.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|phymix_towards_physically_consistent_singleimage_3d_indoor_scene_generation_with_implicitexplicit_optimization"},"authorids":{"value":["~Dongli_Wu1","~Jingyu_Hu4","~Ka-Hei_Hui1","~Xiaobao_Wei1","~Zhengzhe_Liu2","~Jianqiang_Li2"]},"authors":{"value":["Dongli Wu","Jingyu Hu","Ka-Hei Hui","Xiaobao Wei","Zhengzhe Liu","Jianqiang Li"]}},"version":2},{"content":{"venue":{"value":"PLoS Comput. Biol. 2020"},"pdf":{"value":"https://journals.plos.org/ploscompbiol/article/file?id=10.1371/journal.pcbi.1007730&type=printable"},"venueid":{"value":"dblp.org/journals/PLOSCB/2020"},"paperhash":{"value":"neupärtl|intuitive_physical_reasoning_about_objects_masses_transfers_to_a_visuomotor_decision_task_consistent_with_newtonian_physics"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Nils_Neupärtl:","https://dblp.org/search/pid/api?q=author:Fabian_Tatai:","~Constantin_A._Rothkopf1"]},"html":{"value":"https://doi.org/10.1371/journal.pcbi.1007730"},"_bibtex":{"value":"@article{DBLP:journals/ploscb/NeupartlTR20,\n  author={Nils Neupärtl and Fabian Tatai and Constantin A. Rothkopf},\n  title={Intuitive physical reasoning about objects' masses transfers to a visuomotor decision task consistent with Newtonian physics},\n  year={2020},\n  cdate={1577836800000},\n  journal={PLoS Comput. Biol.},\n  volume={16},\n  number={10},\n  url={https://doi.org/10.1371/journal.pcbi.1007730}\n}\n"},"abstract":{"value":"While interacting with objects during every-day activities, e.g. when sliding a glass on a counter top, people obtain constant feedback whether they are acting in accordance with physical laws. However, classical research on intuitive physics has revealed that people’s judgements systematically deviate from predictions of Newtonian physics. Recent research has explained at least some of these deviations not as consequence of misconceptions about physics but instead as the consequence of the probabilistic interaction between inevitable perceptual uncertainties and prior beliefs. How intuitive physical reasoning relates to visuomotor actions is much less known. Here, we present an experiment in which participants had to slide pucks under the influence of naturalistic friction in a simulated virtual environment. The puck was controlled by the duration of a button press, which needed to be scaled linearly with the puck’s mass and with the square-root of initial distance to reach a target. Over four phases of the experiment, uncertainties were manipulated by altering the availability of sensory feedback and providing different degrees of knowledge about the physical properties of pucks. A hierarchical Bayesian model of the visuomotor interaction task incorporating perceptual uncertainty and press-time variability found substantial evidence that subjects adjusted their button-presses so that the sliding was in accordance with Newtonian physics. After observing collisions between pucks, which were analyzed with a hierarchical Bayesian model of the perceptual observation task, subjects transferred the relative masses inferred perceptually to adjust subsequent sliding actions. Crucial in the modeling was the inclusion of a cost function, which quantitatively captures participants’ implicit sensitivity to errors due to their motor variability. Taken together, in the present experiment we find evidence that our participants transferred their intuitive physical reasoning to a subsequent visuomotor control task consistent with Newtonian physics and weighed potential outcomes with a cost functions based on their knowledge about their own variability."},"title":{"value":"Intuitive physical reasoning about objects' masses transfers to a visuomotor decision task consistent with Newtonian physics"},"authors":{"value":["Nils Neupärtl","Fabian Tatai","Constantin A. Rothkopf"]}},"tmdate":1741353526212,"pdate":1577836800000,"tcdate":1741353512592,"writers":["~"],"signatures":["~Constantin_Rothkopf1"],"forum":"qYefON5kag","license":"CC BY-SA 4.0","number":364885,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741353526212,"domain":"DBLP.org","id":"qYefON5kag","version":2},{"content":{"summary":{"value":"This work presents a physics-aware benchmark for 3D dynamic scene understanding. It features 17 simulated scenes with diverse materials—liquids, gases, rheological substances, and textiles—showcasing complex multi-body interactions generated by accurate physical solvers. Experiments reveal that existing DyNVS models achieve visual realism but fail to capture true physics, making Phys-Bench a key resource for developing physics-aware scene reconstruction."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Please refer to the weakness section."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This work focuses on the physics-aware 3D scene understanding, which is a significant problem. This benchmark is useful for evaluating model's physics understanding ability. \n2. The benchmark includes interactions of many different materials with modern simulation techniques."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Lack of visualizations. Can the author present videos of the benchmark?\n2. The benchmark only includes 17 scenes, which seems to be limited for evaluation. \n3. The proposed metrics are not well-justified. The author should provide more discussion about the selection of the metrics. Specifically, in TD, the authors don't explain how they match ground-truth points and reconstructed points. It's unclear how they compute this metric. Furthermore, both TD and AUOP are computed point-wise. This cannot well indicate the physical plausibility of the whole reconstructed 3D model. \n4. Justification of my current score:\n(1) Why not lower: this work is well-motivated, and I believe it's benchmarking a significant problem (physical plausibility) in the field of 3D reconstruction and understanding. \n(2) Why not higher: my main concern with this work is the limited scale of scenes. As a benchmark, it should provide a comprehensive evaluation for its target field. As for this paper, the authors do include different kinds of materials and different simulation techniques. However, the majority of this dataset is still scenes with several relatively simple objects and constrained spatial configurations. This limits its ability to fully test the generalization and robustness of learning-based methods, especially in more complex or cluttered environments. Moreover, it remains unclear how well the dataset covers real-world variability in lighting, geometry, and dynamic interactions, which are crucial for practical applications."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924696162,"tcdate":1761792044678,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14241/Reviewer_4CZi"],"signatures":["ICLR.cc/2026/Conference/Submission14241/Reviewer_4CZi"],"forum":"kwhk8o3k5O","number":1,"license":"CC BY 4.0","cdate":1761792044678,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14241/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924696162,"domain":"ICLR.cc/2026/Conference","replyto":"kwhk8o3k5O","id":"knOuiLwS0z","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["4D Gaussian Splatting","Physics","Dynamic Novel View Synthesis"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We introduce Phys-Bench, a novel physics-aware benchmark for 3D dynamic scene understanding. This benchmark is designed to evaluate methods for reconstructing 4D scenes and understanding underlying physics from given videos, with a main\nfocus on the Dynamic Novel View Synthesis (DyNVS) task. \nWhile existing algorithms and benchmarks primarily focus on photorealistic reconstruction, they largely overlook physics understanding. This neglect is a critical limitation, as a true understanding of dynamic scenes requires models to reason about physical interactions, not just appearance. \nOur benchmark provides complex dynamic scenarios with rich multi-object interactions, featuring realistic collisions and force exchanges that are faithfully generated to strictly adhere to physical laws. \nFurthermore, it contains a diverse range of physical materials, such as liquid, gas, rheological substances, and textiles, which move beyond the rigid bodies prevalent in existing benchmarks. \nTo enable quantitative evaluation, we provide essential ground-truth information such as 3D particle trajectories and physics parameters and propose two novel metrics tailored to assessing physical realism. \nWe further evaluate existing Dynamic Novel View synthesis and physics parameter estimation method on our benchmark and reveal their overlooked limitations in physics understanding and multi-body dynamics handling. \nWe believe Phys-Bench will serve as a crucial foundation for advancing research in dynamic view synthesis, physics-based scene understanding, and the integration of deep learning with physical simulation, ultimately enabling more faithful reconstruction and interpretation of complex 3D dynamic scenes."},"_bibtex":{"value":"@misc{\nkim2025physbench,\ntitle={Phys-Bench: A Physics-aware Benchmark with Multi-Body Interactions for 3D Dynamic Scene Understanding},\nauthor={Mijeong Kim and Gunhee Kim and Jungyoon Choi and Wonjae Roh and Bohyung Han},\nyear={2025},\nurl={https://openreview.net/forum?id=kwhk8o3k5O}\n}"},"title":{"value":"Phys-Bench: A Physics-aware Benchmark with Multi-Body Interactions for 3D Dynamic Scene Understanding"},"pdf":{"value":"/pdf/7397fc3721e1087578236428bbabf75295e20450.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kim|physbench_a_physicsaware_benchmark_with_multibody_interactions_for_3d_dynamic_scene_understanding"},"authorids":{"value":["~Mijeong_Kim1","~Gunhee_Kim4","~Jungyoon_Choi1","~Wonjae_Roh2","~Bohyung_Han1"]},"authors":{"value":["Mijeong Kim","Gunhee Kim","Jungyoon Choi","Wonjae Roh","Bohyung Han"]}},"version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_115.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"han|internalevolution_driven_growth_in_creationannihilation_cyclic_games"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xiao-Pu_Han:","https://dblp.org/search/pid/api?q=author:Luo-Luo_Jiang:","~Tao_Zhou13","https://dblp.org/search/pid/api?q=author:Bing-Hong_Wang:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_115"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/HanJZW09,\n  author={Xiao-Pu Han and Luo-Luo Jiang and Tao Zhou and Bing-Hong Wang},\n  title={Internal-Evolution Driven Growth in Creation-Annihilation Cyclic Games},\n  year={2009},\n  cdate={1230768000000},\n  pages={2377-2387},\n  url={https://doi.org/10.1007/978-3-642-02469-6_115},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"In this paper, the domain growth process in a novel kind of cyclic game is investigated by similation method. Different with the classical cyclic games, this process is called ”creation-annihilation process”, in which is just like the autocatalysis system. The results of numerical simulations show that the domain growth in such cyclic games with four or five states has a special feature: the growing domain usually has a stable boundary, and the growth is driven by the internal-evolution of the domain. Considering with the widespread of the cyclic autocatalysis in organism activities, such internal-evolution driven growth could be universal in many organism systems."},"title":{"value":"Internal-Evolution Driven Growth in Creation-Annihilation Cyclic Games"},"authors":{"value":["Xiao-Pu Han","Luo-Luo Jiang","Tao Zhou","Bing-Hong Wang"]}},"tmdate":1767616599864,"pdate":1230768000000,"externalIds":["dblp:conf/complex/HanJZW09"],"tcdate":1767616531666,"writers":["~"],"signatures":["~Tao_Zhou13"],"forum":"B7Q8kG7NbQ","license":"CC BY-SA 4.0","number":715696,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1767616599864,"domain":"DBLP.org","id":"B7Q8kG7NbQ","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_83.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"chen|collective_aggregation_pattern_dynamics_control_via_attractiverepulsive_function"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Michael_Z._Q._Chen:","https://dblp.org/search/pid/api?q=author:Zhao_Cheng:","https://dblp.org/search/pid/api?q=author:Hai-Tao_Zhang:","~Tao_Zhou13","https://dblp.org/search/pid/api?q=author:Ian_Postlethwaite:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_83"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/ChenCZZP09,\n  author={Michael Z. Q. Chen and Zhao Cheng and Hai-Tao Zhang and Tao Zhou and Ian Postlethwaite},\n  title={Collective Aggregation Pattern Dynamics Control via Attractive/Repulsive Function},\n  year={2009},\n  cdate={1230768000000},\n  pages={2064-2077},\n  url={https://doi.org/10.1007/978-3-642-02469-6_83},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"In the coordinated collective behaviors of biological swarms and flocks, the attractive/repulsive (A/R) functional link between each pair of particles plays an important role. By changing the slope of the A/R function, a dramatic transition between different aggregation patterns surfaces. With a high value of the slope, the particle aggregation shows a liquid-like pattern in which the outer particles are sparsely distributed while the inner ones densely. In addition, the particle density is reduced from the outside to the inside of each cluster. By comparison, when the slope decreases to a sufficiently low value, the particle aggregation exhibits a crystal-like pattern as the distance between each pair of neighboring particles remains constant. Remarkably, there is an obvious spinodal in the curve of particle-particle distance variance versus the slope, indicating a transition between liquid-like and crystal-like aggregation patterns. Significantly, this work may reveal some common mechanism behind the aggregation of physical particles and swarming of organisms in nature, and may find its potential engineering applications, for example, UAVs and multi-robot systems."},"title":{"value":"Collective Aggregation Pattern Dynamics Control via Attractive/Repulsive Function"},"authors":{"value":["Michael Z. Q. Chen","Zhao Cheng","Hai-Tao Zhang","Tao Zhou","Ian Postlethwaite"]}},"tmdate":1767616592740,"pdate":1230768000000,"externalIds":["dblp:conf/complex/ChenCZZP09"],"tcdate":1767616531873,"writers":["~"],"signatures":["~Tao_Zhou13"],"forum":"jwiZ115waa","license":"CC BY-SA 4.0","number":715681,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1767616592740,"domain":"DBLP.org","id":"jwiZ115waa","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_92.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"zhang|collective_behavior_coordination_and_aggregation_with_lowcost_communication"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Hai-Tao_Zhang:","https://dblp.org/search/pid/api?q=author:Michael_Z._Q._Chen:","~Tao_Zhou13","https://dblp.org/search/pid/api?q=author:Zhao_Cheng:","https://dblp.org/search/pid/api?q=author:Pin-Ze_Yu:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_92"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/ZhangCZCY09,\n  author={Hai-Tao Zhang and Michael Z. Q. Chen and Tao Zhou and Zhao Cheng and Pin-Ze Yu},\n  title={Collective Behavior Coordination and Aggregation with Low-Cost Communication},\n  year={2009},\n  cdate={1230768000000},\n  pages={2159-2170},\n  url={https://doi.org/10.1007/978-3-642-02469-6_92},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"An important natural phenomenon surfaces that satisfactory synchronization of self-driven particles can be achieved via remarkably reduced communication cost, especially for high density particle groups with low external noise. Statistical numerical evidence illustrates that a highly efficient manner is to distribute the communication messages as evenly as possible along the whole dynamic process, since it minimizes the communication redundancy. More surprisingly, it is discovered that there exists an abnormal region in the state diagram where moderately decreasing the communication cost can even improve the synchronization performance. Significantly, another interesting fact is found that low-cost communication can help the particles aggregate into synchronized clusters, which may be beneficial to explain the forming mechanism of individuals’ aggregation phenomena over biological flocks/swarms."},"title":{"value":"Collective Behavior Coordination and Aggregation with Low-Cost Communication"},"authors":{"value":["Hai-Tao Zhang","Michael Z. Q. Chen","Tao Zhou","Zhao Cheng","Pin-Ze Yu"]}},"tmdate":1767616562581,"pdate":1230768000000,"externalIds":["dblp:conf/complex/ZhangCZCY09"],"tcdate":1767616531772,"writers":["~"],"signatures":["~Tao_Zhou13"],"forum":"x5wjkddWmf","license":"CC BY-SA 4.0","number":715661,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1767616562581,"domain":"DBLP.org","id":"x5wjkddWmf","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_80.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"lawniczak|development_of_road_traffic_ca_model_of_4way_intersection_to_study_travel_time"},"authorids":{"value":["~Anna_T._Lawniczak1","https://dblp.org/search/pid/api?q=author:Bruno_N._Di_Stefano:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_80"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/LawniczakS09,\n  author={Anna T. Lawniczak and Bruno N. Di Stefano},\n  title={Development of Road Traffic CA Model of 4-Way Intersection to Study Travel Time},\n  year={2009},\n  cdate={1230768000000},\n  pages={2040-2049},\n  url={https://doi.org/10.1007/978-3-642-02469-6_80},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"We describe our development of a road traffic CA (Cellular Automata) model of the four most common types of 4-way intersection (Yield-controlled intersections, Stop-controlled intersections, Signal-controlled intersections, and Roundabout-based intersection). We developed this model to study how these four different types of 4-way intersection affect road traffic flow and congestion in general and “travel time” in particular. In this paper we describe the model and 4WayCA.exe, the traffic simulator software package in which the model has been implemented. We focus in particular on the model abstractions and on the simulator architecture."},"title":{"value":"Development of Road Traffic CA Model of 4-Way Intersection to Study Travel Time"},"authors":{"value":["Anna T. Lawniczak","Bruno N. Di Stefano"]}},"tmdate":1759275183061,"pdate":1230768000000,"externalIds":["dblp:conf/complex/LawniczakS09"],"tcdate":1759275140303,"writers":["~"],"signatures":["~Anna_T_Lawniczak1"],"forum":"6423ZtZs77","license":"CC BY-SA 4.0","number":632745,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1759275183061,"domain":"DBLP.org","id":"6423ZtZs77","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_57.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"lawniczak|entropy_based_detection_of_ddos_attacks_in_packet_switching_network_models"},"authorids":{"value":["~Anna_T._Lawniczak1","https://dblp.org/search/pid/api?q=author:Hao_Wu_0002:","https://dblp.org/search/pid/api?q=author:Bruno_N._Di_Stefano:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_57"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/LawniczakWS09,\n  author={Anna T. Lawniczak and Hao Wu and Bruno N. Di Stefano},\n  title={Entropy Based Detection of DDoS Attacks in Packet Switching Network Models},\n  year={2009},\n  cdate={1230768000000},\n  pages={1810-1822},\n  url={https://doi.org/10.1007/978-3-642-02469-6_57},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"Distributed denial-of-service (DDoS) attacks are network-wide attacks that cannot be detected or stopped easily. They affect “natural” spatio-temporal packet traffic patterns, i.e. “natural distributions” of packets passing through the routers. Thus, they affect “natural” information entropy profiles, a sort of “fingerprints”, of normal packet traffic. We study if by monitoring information entropy of packet traffic through selected routers one may detect DDoS attacks or anomalous packet traffic in packet switching network (PSN) models. Our simulations show that the considered DDoS attacks of “ping” type cause shifts in information entropy profiles of packet traffic monitored even at small sets of routers and that it is easier to detect these shifts if static routing is used instead of dynamic routing. Thus, network-wide monitoring of information entropy of packet traffic at properly selected routers may provide means for detecting DDoS attacks and other anomalous packet traffics."},"title":{"value":"Entropy Based Detection of DDoS Attacks in Packet Switching Network Models"},"authors":{"value":["Anna T. Lawniczak","Hao Wu","Bruno N. Di Stefano"]}},"tmdate":1759275157797,"pdate":1230768000000,"externalIds":["dblp:conf/complex/LawniczakWS09"],"tcdate":1759275140038,"writers":["~"],"signatures":["~Anna_T_Lawniczak1"],"forum":"hK0iwKzFKI","license":"CC BY-SA 4.0","number":632726,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1759275157797,"domain":"DBLP.org","id":"hK0iwKzFKI","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_37.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"geng|evolving_model_of_weighted_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xianmin_Geng:","https://dblp.org/search/pid/api?q=author:Hongwei_Zhou:","~Guanghui_Wen1"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_37"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/GengZW09,\n  author={Xianmin Geng and Hongwei Zhou and Guanghui Wen},\n  title={Evolving Model of Weighted Networks},\n  year={2009},\n  cdate={1230768000000},\n  pages={1575-1590},\n  url={https://doi.org/10.1007/978-3-642-02469-6_37},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"In this paper, in order to search the reason of the phenomena of power- law in the weighted networks, we present a general model for the growth of weighted networks that couples of new edges and vertices and the weights’ and intrinsic strengths’ dynamical evolution. This model is based on a simple weight and intrinsic strength driven dynamics and generates networks exhibiting the statistical properties observed in several real-world systems. Within this model we not only yields the scale-free behavior for the weight, strength and degree distributions, but also we give the analytical computation of the distributions of the weight, the strength and the degree .Simultaneity, by way of contrasting our results with those of the random model, we found the preferential attachment is necessary to the phenomena of scale-free of the strength and degree distributions. Finally, we found the analytical results are good consistent with those of numerical simulation. The conclusion from this model is helpful to the investigation of the topological role of weight and strength."},"title":{"value":"Evolving Model of Weighted Networks"},"authors":{"value":["Xianmin Geng","Hongwei Zhou","Guanghui Wen"]}},"tmdate":1749148025872,"pdate":1230768000000,"tcdate":1749147712574,"writers":["~"],"signatures":["~Guanghui_Wen1"],"forum":"raJtY0qABA","license":"CC BY-SA 4.0","number":557162,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1749148025872,"domain":"DBLP.org","id":"raJtY0qABA","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_37.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"geng|evolving_model_of_weighted_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xianmin_Geng:","https://dblp.org/search/pid/api?q=author:Hongwei_Zhou:","~Guanghui_Wen1"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_37"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/GengZW09,\n  author={Xianmin Geng and Hongwei Zhou and Guanghui Wen},\n  title={Evolving Model of Weighted Networks},\n  year={2009},\n  cdate={1230768000000},\n  pages={1575-1590},\n  url={https://doi.org/10.1007/978-3-642-02469-6_37},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"In this paper, in order to search the reason of the phenomena of power- law in the weighted networks, we present a general model for the growth of weighted networks that couples of new edges and vertices and the weights’ and intrinsic strengths’ dynamical evolution. This model is based on a simple weight and intrinsic strength driven dynamics and generates networks exhibiting the statistical properties observed in several real-world systems. Within this model we not only yields the scale-free behavior for the weight, strength and degree distributions, but also we give the analytical computation of the distributions of the weight, the strength and the degree .Simultaneity, by way of contrasting our results with those of the random model, we found the preferential attachment is necessary to the phenomena of scale-free of the strength and degree distributions. Finally, we found the analytical results are good consistent with those of numerical simulation. The conclusion from this model is helpful to the investigation of the topological role of weight and strength."},"title":{"value":"Evolving Model of Weighted Networks"},"authors":{"value":["Xianmin Geng","Hongwei Zhou","Guanghui Wen"]}},"tmdate":1749147837426,"pdate":1230768000000,"tcdate":1749147701108,"writers":["~"],"signatures":["~Guanghui_Wen1"],"forum":"59F8wc4JeS","license":"CC BY-SA 4.0","number":556772,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1749147837426,"domain":"DBLP.org","id":"59F8wc4JeS","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_60.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"ke|degree_distribution_of_a_twocomponent_growing_network"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jianhong_Ke:",""]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_60"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/KeC09,\n  author={Jianhong Ke and Xiaoshuang Chen},\n  title={Degree Distribution of a Two-Component Growing Network},\n  year={2009},\n  cdate={1230768000000},\n  pages={1838-1845},\n  url={https://doi.org/10.1007/978-3-642-02469-6_60},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"We propose a two-component growing network model which comprises two kinds of nodes. Such a network is constructed by introducing new nodes of either kind with no immediate links and creating new links between any two nodes. We then investigate the connectivity of the two-component growing network by means of the rate equation approach. For a network system with shifted linear connection rate kernels, the in-degree and out-degree distributions take power-law forms; while for a random growing network, the in-degree and out-degree distributions are both exponential. Moreover, the in-degree and out-degree distributions are correlated each other."},"title":{"value":"Degree Distribution of a Two-Component Growing Network"},"authors":{"value":["Jianhong Ke","Xiaoshuang Chen"]}},"tmdate":1745077794058,"pdate":1230768000000,"tcdate":1735882941793,"writers":["~"],"signatures":["~Xiaoshuang_Chen5"],"forum":"pOianppE4a","license":"CC BY-SA 4.0","number":260097,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1745077794058,"domain":"DBLP.org","id":"pOianppE4a","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_12.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"feng|a_new_bioinspired_approach_to_the_traveling_salesman_problem"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Francis_C._M._Lau_0001:","https://dblp.org/search/pid/api?q=author:Daqi_Gao:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_12"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/FengLG09b,\n  author={Xiang Feng and Francis C. M. Lau and Daqi Gao},\n  title={A New Bio-inspired Approach to the Traveling Salesman Problem},\n  year={2009},\n  cdate={1230768000000},\n  pages={1310-1321},\n  url={https://doi.org/10.1007/978-3-642-02469-6_12},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"The host-seeking behavior of mosquitoes is very interesting. In this paper, we propose a novel mosquito host-seeking algorithm (MHSA) as a new branch of biology-inspired algorithms for solving TSP problems. The MHSA is inspired by the host-seeking behavior of mosquitoes. We present the mathematical model, the algorithm, the motivation, and the biological model. The MHSA can work out the theoretical optimum solution, which is important and exciting, and we give the theoretical foundation and present experiment results that verify this fact."},"title":{"value":"A New Bio-inspired Approach to the Traveling Salesman Problem"},"authors":{"value":["Xiang Feng","Francis C. M. Lau","Daqi Gao"]}},"tmdate":1744105802561,"pdate":1230768000000,"tcdate":1744105521333,"writers":["~"],"signatures":["~Xiang_Feng2"],"forum":"UxJ1W4Gdj1","license":"CC BY-SA 4.0","number":382256,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1744105802561,"domain":"DBLP.org","id":"UxJ1W4Gdj1","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_16.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"chen|a_grid_resource_scheduling_algorithm_based_on_the_utility_optimization"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jiang_Chen:","~Jian_Peng5","https://dblp.org/search/pid/api?q=author:Xiaoyang_Cao:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_16"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/ChenPC09,\n  author={Jiang Chen and Jian Peng and Xiaoyang Cao},\n  title={A Grid Resource Scheduling Algorithm Based on the Utility Optimization},\n  year={2009},\n  cdate={1230768000000},\n  pages={1355-1362},\n  url={https://doi.org/10.1007/978-3-642-02469-6_16},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"To solve the problem of heterogeneity of user requirements in grid resource allocation, a grid resource scheduling algorithm based on utility function is proposed by analyzing the relationship between the executing time and cost and the user utility function, the theory of economics is used to solve the optimal problem of the user utility function. The result of experiment shows, when the system finished the same set of gridlets, the algorithm achieves better performance not only in cost than the algorithm based on the time optimization when they spent equal time, but also in time than the algorithm based on the cost optimization on the assumption that they consumed the equal quantity of cost."},"title":{"value":"A Grid Resource Scheduling Algorithm Based on the Utility Optimization"},"authors":{"value":["Jiang Chen","Jian Peng","Xiaoyang Cao"]}},"tmdate":1744081977817,"pdate":1230768000000,"tcdate":1744081947777,"writers":["~"],"signatures":["~jian_peng4"],"forum":"fpO35hisYz","license":"CC BY-SA 4.0","number":380190,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1744081977817,"domain":"DBLP.org","id":"fpO35hisYz","version":2},{"content":{"venue":{"value":"Complex (2) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02469-6_58.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"qu|enhancing_the_scalefree_networks_attack_tolerance"},"authorids":{"value":["~Zehui_Qu1","https://dblp.org/search/pid/api?q=author:Pu_Wang:","https://dblp.org/search/pid/api?q=author:Zhiguang_Qin:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02469-6_58"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/QuWQ09,\n  author={Zehui Qu and Pu Wang and Zhiguang Qin},\n  title={Enhancing the Scale-Free Network's Attack Tolerance},\n  year={2009},\n  cdate={1230768000000},\n  pages={1823-1826},\n  url={https://doi.org/10.1007/978-3-642-02469-6_58},\n  booktitle={Complex (2)},\n  crossref={conf/complex/2009-2}\n}\n"},"abstract":{"value":"Despite the large size of most communication systems such as the Internet and World Wide Web (WWW), there is a relatively short path between two nodes, revealing the networks’ small world characteristic which speeds the delivery of information and data. While these networks have a surprising error tolerance, their scale-free topology makes them fragile under intentional attack, leaving us a challenge on how to improve the networks’ robustness against attack without losing their small world merit. Here we try to enhance scale-free network’s tolerance under attack by using a method based on networks’ topology re-constructing."},"title":{"value":"Enhancing the Scale-Free Network's Attack Tolerance"},"authors":{"value":["Zehui Qu","Pu Wang","Zhiguang Qin"]}},"tmdate":1738847379183,"pdate":1230768000000,"tcdate":1738847371464,"writers":["~"],"signatures":["~Zehui_Qu1"],"forum":"CYnpdP6SkY","license":"CC BY-SA 4.0","number":307246,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1738847379183,"domain":"DBLP.org","id":"CYnpdP6SkY","version":2},{"content":{"summary":{"value":"This paper presents Lean4PHYS, a Lean4-based framework for formalizing college-level physics problems.   It includes PhysLib, a modular library with a systematic unit system and reusable theorems, and LeanPhysBench, a benchmark of 200 formalized physics problems.   The authors propose a pipeline to convert natural language physics questions into Lean4 proofs.   Experiments compare Lean-oriented provers with general-purpose LLMs on LeanPhysBench.   Results show LLMs outperform provers, with improved performance around 40.5% accuracy when using PhysLib context, highlighting the need for domain-specific knowledge in formal reasoning."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"see Weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper pioneers formal physics reasoning in Lean and introduces the first large-scale Lean4 physics benchmark covering topics from mechanics to modern physics. \n\n- PhysLib provides a modular, SI-based unit system and topic-structured theorems, ensuring dimensional consistency and extendability for accurate physical reasoning.\n\n- Comprehensive experiments show that PhysLib context consistently boosts model performance.  LLMs outperform specialized Lean provers, revealing that current Lean provers, trained for math, struggle with physics tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper lacks an explanation of the advantages of PhysLib's modular structure over organizing theorems by specific physics domains.  It also does not clarify how this hierarchical organization aids in retrieval and reasoning.\n- What defines the boundary between mathematical and physical problems, and why do Lean provers, which perform well in mathematics, fail to transfer their capabilities to the physics domain?\n- Would including non-competition problems, such as physics questions from middle school or high school exams, in the experiments provide a more comprehensive comparison?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919841693,"tcdate":1761640833727,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7797/Reviewer_VJbK"],"signatures":["ICLR.cc/2026/Conference/Submission7797/Reviewer_VJbK"],"forum":"wQ2jyFz18H","number":1,"license":"CC BY 4.0","cdate":1761640833727,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7797/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919841693,"domain":"ICLR.cc/2026/Conference","replyto":"wQ2jyFz18H","id":"HE7KZlDAN2","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"This paper presents Lean4PHYS, a reasoning framework for college-level physics problems in Lean4. It includes LeanPhysBench, the first benchmark in the field, and PhysLib, a community-driven repository that sets the foundation for the field."},"keywords":{"value":["Lean4","Reasoning","AIforScience"]},"supplementary_material":{"value":"/attachment/fbd9bc9c890ab7276d2401a43a0d60dbc1501c55.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present **Lean4PHYS**, a comprehensive reasoning framework for college-level physics problems in Lean4. To establish a solid foundation for formal reasoning in physics, **Lean4PHYS** launches *PhysLib*, a repository containing fundamental unit systems and essential theorems to formulate physics proofs in Lean4. It will be community-driven and long-term maintained. Lean4PHYS also includes *LeanPhysBench*, a college-level benchmark for evaluating LLMs' Lean4 formal physics reasoning capability. It contains 200 hand-crafted and peer-reviewed Lean4 theorem statements formalized from university textbooks and physics competition problems. Based on the *PhysLib* and *LeanPhysBench* we composed in **Lean4PHYS**, we perform exhaustive experiments of baseline results using major expert Math provers and state-of-the-art closed-source models, and provide an analysis of their performance. In the experiment, we identify that most expert provers do not outperform general models as they did in the math domain. This suggests potential overfitting to the math domain rather than learning formal reasoning for formal provers. We also conduct a comprehensive experiment showing that, with *PhysLib* in the context, LLMs' performance on *LeanPhysBench* increases by **11.90%** on average, proving the effectiveness of our repository in assisting LLMs in solving the Lean4 physics problem. To the best of our knowledge, we are the first study to provide a physics benchmark in Lean4."},"_bibtex":{"value":"@inproceedings{\nli2026leanphysics,\ntitle={Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4},\nauthor={Yuxin Li and Minghao LIU and Ruida WANG and JI WenZhao and Zhitao He and Rui Pan and Junming Huang and Tong Zhang and Yi R. Fung},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=wQ2jyFz18H}\n}"},"title":{"value":"Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4"},"pdf":{"value":"/pdf/71e0f9a2255dd3fbd4264072b0433afaf0786606.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|lean4physics_comprehensive_reasoning_framework_for_collegelevel_physics_in_lean4"},"authorids":{"value":["~Yuxin_Li13","~Minghao_LIU8","~Ruida_WANG1","~JI_WenZhao1","~Zhitao_He1","~Rui_Pan4","~Junming_Huang1","~Tong_Zhang2","~Yi_R._Fung1"]},"authors":{"value":["Yuxin Li","Minghao LIU","Ruida WANG","JI WenZhao","Zhitao He","Rui Pan","Junming Huang","Tong Zhang","Yi R. Fung"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a transformer-based architecture to scale neural operators to larger and more complex conditions involving spatiotemporal modeling. The novelty of this architecture is the use of the so-called latent rollout, which performs autoregressive modeling on the latent space, as opposed to the decoded space in other approaches. Authors claim that this architecture can operate without spatial structures, which is an inherent benefit of transformer-based approaches, and inspection of the latent-space in space-time. The authors show the usefulness of this approach across different physics-based problems. This is a meaningful contribution to design of NNs for physics applications, with potential broader applications in other ML domains. However, there are some questions that this reviewer hopes that the authors can address."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Authors should be commended on the discussion on memory complexity in Line 105 to 119. However, this could be summarized as a table for cleaner presentation. Can the authors also include time complexity information for a more transparent comparison of the trade-offs associated with the different architectures? (I know this has been somewhat benchmarked in Table 1, but it could be useful to have some more theoretical information)\n2. What kind of position encoding is used for the UPT? This is unclear in the paper. Does the type of position encoding matter for Lagrangian vs Eulerian.\n3. For a more comprehensive display of results, could the authors include accuracy metrics (e.g., mean error) and memory benefits to  Table 1?\n4. This has been briefly discussed in Appendix A. But for further clarification, is the present architecture approach limited to physics-related applications? Or are there opportunities for extending the present approach towards other autoregressive ML applications, such as video or language modeling?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. Paper is clearly written, and appendices help with self-completeness.\n2. Proposed method is shown to be better than existing popular architectures for physics-based modeling.\n3. Code is provided for reproducibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. A bit more analysis and information could be included to improve completeness. See Questions.\n2. Only test MSE has been used as an accuracy metric. Physics applications typically care about conservation of mass and performance near boundaries. However, the lack of further metrics is somewhat justified by the speedup and memory gains shown,"},"limitations":{"value":"Limitations are well-discussed in Appendix A."}},"nonreaders":[],"tmdate":1730879083110,"tcdate":1720374752924,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission6487/Reviewer_CUfq"],"signatures":["NeurIPS.cc/2024/Conference/Submission6487/Reviewer_CUfq"],"forum":"oUXiNX5KRm","number":3,"license":"CC BY 4.0","cdate":1720374752924,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission6487/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879083110,"domain":"NeurIPS.cc/2024/Conference","replyto":"oUXiNX5KRm","id":"tZk0Xyr9tW","forumContent":{"TLDR":{"value":"We introduce Universal Physics Transformers, an efficiently scalable neural operator framework to model a wide range of spatio-temporal problems – for Lagrangian and Eulerian discretization schemes."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["neural operator","computational fluid dynamics","Lagrangian simulations","transformers","latent space modeling"]},"supplementary_material":{"value":"/attachment/ffbee57cf57b024cc7564d74926efc865fcb2da6.zip"},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Neural operators, serving as physics surrogate models, have recently gained increased interest. With ever increasing problem complexity, the natural question arises: what is an efficient way to scale neural operators to larger and more complex simulations - most importantly by taking into account different types of simulation datasets. This is of special interest since, akin to their numerical counterparts, different techniques are used across applications, even if the underlying dynamics of the systems are similar. Whereas the flexibility of transformers has enabled unified architectures across domains, neural operators mostly follow a problem specific design, where GNNs are commonly used for Lagrangian simulations and grid-based models predominate Eulerian simulations. \n\nWe introduce Universal Physics Transformers (UPTs), an efficient and unified learning paradigm for a wide range of spatio-temporal problems. UPTs operate without grid- or particle-based latent structures, enabling flexibility and scalability across meshes and particles. UPTs efficiently propagate dynamics in the latent space, emphasized by inverse encoding and decoding techniques. Finally, UPTs allow for queries of the latent space representation at any point in space-time. We demonstrate diverse applicability and efficacy of UPTs in mesh-based fluid simulations, and steady-state Reynolds averaged Navier-Stokes simulations, and Lagrangian-based dynamics."},"_bibtex":{"value":"@inproceedings{\nalkin2024universal,\ntitle={Universal Physics Transformers: A Framework For Efficiently Scaling Neural Operators},\nauthor={Benedikt Alkin and Andreas F{\\\"u}rst and Simon Lucas Schmid and Lukas Gruber and Markus Holzleitner and Johannes Brandstetter},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=oUXiNX5KRm}\n}"},"title":{"value":"Universal Physics Transformers: A Framework For Efficiently Scaling Neural Operators"},"pdf":{"value":"/pdf/4ffed9700454b3329a91e1a3f4304eef2c75e090.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"alkin|universal_physics_transformers_a_framework_for_efficiently_scaling_neural_operators"},"authorids":{"value":["~Benedikt_Alkin1","~Andreas_Fürst1","~Simon_Lucas_Schmid1","~Lukas_Gruber2","~Markus_Holzleitner1","~Johannes_Brandstetter1"]},"authors":{"value":["Benedikt Alkin","Andreas Fürst","Simon Lucas Schmid","Lukas Gruber","Markus Holzleitner","Johannes Brandstetter"]}},"version":2},{"content":{"summary":{"value":"In this paper, the authors focus on the evaluation of text-to-video models. To this end, they propose a new benchmark as well as a new evaluation method. Named PhyGenBench, the dataset of prompts evaluate intuitive physics in subcontexts such as mechanics, optics, thermal dynamics, and material properties. Alongside this benchmark is PhyGenEval, an automated eval pipeline where a VLM is combined with GPT-4o to generate evaluative questions and answers. The authors compare PhyGenEval against human evaluations. They also perform some initial experiments with contemporary video models on their proposed benchmark."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"I believe PhyGenBench is an excellent contribution for the research community, but the PhyGenEval as described is problematic for the reasons listed above. Is it possible that PhyGenEval be described as a potential approach for automated evaluation to be iterated on in subsequent research, with PhyGenBench + human evaluation as the main contribution?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. Current generative models of video have serious issues with intuitive physics, and the research community needs a good benchmark to evaluate this capability. The proposed benchmark dataset can serve as a very important dataset for the community.\n\n2. Evaluating intuitive physics can be difficult, and PhyGenEval might be a promising method to automate evaluation without the need for human raters. \n\n3. The quantitative evaluation of current video models on the benchmark is a strong contribution and shows the need to improve these models in the realm of intuitive physics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While PhyGenEval is an interesting approach to automating evaluation on the benchmark, I assert that the approach has some issues and should not be adopted by the community at this moment as a standard evaluation; human raters should be used:\n\n1. The pipeline is not reliable enough, the PCA correlation results are only around .7 - .8\n\n2. The pipeline relies on proprietary models such as GPT-4o and may be difficult to reproduce with open models."}},"nonreaders":[],"tmdate":1731427583711,"tcdate":1730676877720,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2363/Reviewer_PPCZ"],"signatures":["ICLR.cc/2025/Conference/Submission2363/Reviewer_PPCZ"],"forum":"6rMHcLWxl4","number":4,"license":"CC BY 4.0","cdate":1730676877720,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2363/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427583711,"domain":"ICLR.cc/2025/Conference","replyto":"6rMHcLWxl4","id":"dpxenSOCnT","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["World Simulator","Physical Commonsense","Video Generation","Evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGenBench, a comprehensive \\textbf{Phy}sics \\textbf{Gen}eration \\textbf{Ben}chmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench, we propose a novel evaluation framework called PhyGenEval. This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through PhyGenBench and PhyGenEval, we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by PhyGenBench (e.g., dynamic scenarios). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We will release the data and codes at https://github.com/PhyGenBench/PhyGenBench"},"_bibtex":{"value":"@misc{\nmeng2025towards,\ntitle={Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation},\nauthor={Fanqing Meng and Jiaqi Liao and Xinyu Tan and Wenqi Shao and Quanfeng Lu and Kaipeng Zhang and Yu Cheng and Dianqi Li and Yu Qiao and Ping Luo},\nyear={2025},\nurl={https://openreview.net/forum?id=6rMHcLWxl4}\n}"},"title":{"value":"Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation"},"pdf":{"value":"/pdf/1814f0c3473ab9a04aca4edcd8aab3e678055bdd.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"meng|towards_world_simulator_crafting_physical_commonsensebased_benchmark_for_video_generation"},"authorids":{"value":["~Fanqing_Meng1","~Jiaqi_Liao2","~Xinyu_Tan1","~Wenqi_Shao2","~Quanfeng_Lu1","~Kaipeng_Zhang1","~Yu_Cheng1","~Dianqi_Li1","~Yu_Qiao1","~Ping_Luo2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fanqing Meng","Jiaqi Liao","Xinyu Tan","Wenqi Shao","Quanfeng Lu","Kaipeng Zhang","Yu Cheng","Dianqi Li","Yu Qiao","Ping Luo"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PhysVidBench, a new benchmark designed to evaluate physical commonsense understanding in T2V models. They argue that although recent T2V models demonstrate impressive visual realism, they frequently violate intuitive physics—showing implausible object motions, misuse of tools, or broken causal sequences. PhysVidBench is derived from the PIQA dataset and contains 383 base prompts plus 383 “upsampled” prompts, emphasizing tool use and affordance reasoning in everyday physical tasks. Evaluation is conducted via a three-stage pipeline: Generating physics-grounded yes/no questions; Producing dense video captions; Asking an LLM to answer the questions using only the captions. \nExperiments across major T2V models show that current systems achieve modest scores and fail most often in spatial reasoning and temporal dynamics. They further demonstrate human–model agreement and propose a lightweight prompt refinement loop that improves generation quality without retraining the model."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1.\tClear motivation and real-world relevance\nThe paper targets an underexplored yet essential capability: everyday physical commonsense. Unlike prior work focusing on abstract physics laws or motion smoothness, PhysVidBench tests realistic, goal-oriented interactions involving tool use, and material behavior.\n2.\tComprehensive and systematic evaluation\nThe benchmark spans seven reasoning dimensions (force, motion, affordance, material transformation, etc.) and incorporates a difficulty-based stratification. The analysis includes base vs. upsampled prompts, scale effects, and cross-model comparisons, offering a holistic diagnostic view.\n3.\tDiagnostic and iterative refinement utility\nThe proposed “error-guided prompt refinement” demonstrates that the same benchmark can be repurposed to guide model improvement. The iterative procedure yields consistent gains without model updates, proving its usefulness beyond static evaluation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tEvaluation of realism and ceiling\nAlthough the caption-based QA pipeline reduces hallucination, it evaluates textual rather than visual physical understanding. The final judgment depends on the only one captioner’s recall and Gemini’s internal physics priors. A single model may have deviations. The pipeline measures consistency within the text–caption–QA chain, not direct perception–reasoning, may introduce biases. \n2.\tLack of statistical significance reporting\nPerformance differences (e.g., base vs. upsampled prompts, difficulty tiers) are presented without confidence intervals or hypothesis testing. Including bootstrap-based confidence intervals or paired-sample tests would strengthen claims about robustness and consistency.\n3.\tLimited benchmark scale and coverage\nWhile high in quality, 383 prompts are relatively small compared to existing video benchmarks. The benchmark lacks broader coverage of outdoor or complex fluid–solid interactions, or other based datasets similar to PIQA.\n4.\tLack of granular failure analysis\nThe discussion identifies spatial and temporal reasoning as weak dimensions but does not quantify error categories (e.g., occlusion, contact, support relations, deformation). A detailed failure analysis would enhance the diagnostic value of the benchmark."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762930800265,"tcdate":1761558109649,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18831/Reviewer_8iig"],"signatures":["ICLR.cc/2026/Conference/Submission18831/Reviewer_8iig"],"forum":"RsletK8757","number":2,"license":"CC BY 4.0","cdate":1761558109649,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18831/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762930800265,"domain":"ICLR.cc/2026/Conference","replyto":"RsletK8757","id":"r5rhJE1PIi","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["benchmarks","video generation","physical commonsense","text-to-video","generative models","evaluation","multimodal reasoning","synthetic video","model assessment"]},"supplementary_material":{"value":"/attachment/3976b5069c0797816fd1af97929f4e2725c75d28.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in text-to-video (T2V) generation have enabled visually compelling outputs, but models still struggle with everyday physical commonsense, often producing videos that violate intuitive expectations of causality, object behavior, and tool use. We introduce PhysVidBench, a human-validated benchmark for assessing physical reasoning in T2V models. It comprises carefully curated prompts spanning seven dimensions of physical interaction, from material transformation to temporal dynamics, offering broad, multi-faceted coverage of scenarios where physical plausibility is critical. For each prompt, we generate videos using diverse state-of-the-art models, and evaluate them through a three-stage pipeline: grounded physics questions are derived from each prompt, generated videos are captioned with a vision–language model, and a language model answers the questions using only the captions. This strategy mitigates hallucination and produces scores that align closely with human judgments. Beyond evaluation, PhysVidBench also serves as a diagnostic tool, enabling feedback-driven refinement of model outputs. By emphasizing affordances and tool-mediated actions, areas often overlooked in existing benchmarks, PhysVidBench provides a structured, interpretable framework for assessing and improving everyday physical commonsense in T2V models."},"_bibtex":{"value":"@misc{\ntezcan2026can,\ntitle={Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models},\nauthor={Baris Sarper Tezcan and Enes Sanli and Erkut Erdem and Aykut Erdem},\nyear={2026},\nurl={https://openreview.net/forum?id=RsletK8757}\n}"},"title":{"value":"Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models"},"pdf":{"value":"/pdf/125450ebb628e517ed97d111d3871da1123d095f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"tezcan|can_your_model_separate_yolks_with_a_water_bottle_benchmarking_physical_commonsense_understanding_in_video_generation_models"},"authorids":{"value":["~Baris_Sarper_Tezcan1","~Enes_Sanli1","~Erkut_Erdem1","~Aykut_Erdem1"]},"authors":{"value":["Baris Sarper Tezcan","Enes Sanli","Erkut Erdem","Aykut Erdem"]}},"version":2},{"content":{"TLDR":{"value":"RL fine-tuning LLMs on synthetic data improves real-world multi-hop reasoning by teaching knowledge composition skills"},"venue":{"value":"ICLR 2026 Workshop VerifAI-2"},"keywords":{"value":["multi-hop reasoning","large language models","reinforcement learning","synthetic data"]},"abstract":{"value":"Reinforcement Learning (RL) has been shown to significantly boost reasoning capabilities of large language models (LLMs) in math, coding, and multi-hop reasoning tasks.\nHowever, RL fine-tuning requires abundant high-quality verifiable data, often sourced from human annotations, generated from frontier LLMs, or scored by LLM-based verifiers.\nAll three have considerable limitations: human-annotated datasets are small and expensive to curate, LLM-generated data is hallucination-prone and costly, and LLM-based verifiers are inaccurate and slow.\nIn this work, we investigate a cheaper alternative: RL fine-tuning on _rule-generated synthetic data_ for multi-hop reasoning tasks.\nWe discover that LLMs fine-tuned on synthetic data perform significantly better on popular real-world question-answering benchmarks, despite the synthetic data containing only fictional knowledge.\nOn stratifying performance by question difficulty, we find that synthetic data teaches LLMs to _compose knowledge_---a fundamental and generalizable reasoning skill.\nOur work highlights rule-generated synthetic reasoning data as a free and scalable resource to improve LLM reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\nkabra2026learning,\ntitle={Learning from Synthetic Data Improves Multi-hop Reasoning},\nauthor={Anmol Kabra and Yilun Yin and Albert Gong and Kamil{\\.{e}} Stankevi{\\v{c}}i{\\={u}}t{\\.{e}} and Dongyoung Go and Johann Lee and Katie Z Luo and Carla P Gomes and Kilian Q Weinberger},\nbooktitle={ICLR 2026 Workshop: VerifAI-2: The Second Workshop on AI Verification in the Wild},\nyear={2026},\nurl={https://openreview.net/forum?id=JEA9UcQncY}\n}"},"title":{"value":"Learning from Synthetic Data Improves Multi-hop Reasoning"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"venueid":{"value":"ICLR.cc/2026/Workshop/VerifAI-2"},"paperhash":{"value":"kabra|learning_from_synthetic_data_improves_multihop_reasoning"},"authorids":{"value":["~Anmol_Kabra1","~Yilun_Yin1","~Albert_Gong1","~Kamilė_Stankevičiūtė1","~Dongyoung_Go1","~Johann_Lee1","~Katie_Z_Luo1","~Carla_P_Gomes1","~Kilian_Q_Weinberger1"]},"Track":{"value":"long paper (up to 8 pages)"},"authors":{"value":["Anmol Kabra","Yilun Yin","Albert Gong","Kamilė Stankevičiūtė","Dongyoung Go","Johann Lee","Katie Z Luo","Carla P Gomes","Kilian Q Weinberger"]}},"tmdate":1773260207027,"pdate":1772424058776,"tcdate":1770316237201,"writers":["ICLR.cc/2026/Workshop/VerifAI-2","ICLR.cc/2026/Workshop/VerifAI-2/Submission31/Authors"],"signatures":["ICLR.cc/2026/Workshop/VerifAI-2/Submission31/Authors"],"forum":"JEA9UcQncY","license":"CC BY 4.0","number":31,"cdate":1770316237201,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/VerifAI-2/-/Submission","ICLR.cc/2026/Workshop/VerifAI-2/-/Post_Submission","ICLR.cc/2026/Workshop/VerifAI-2/-/Edit","ICLR.cc/2026/Workshop/VerifAI-2/Submission31/-/Camera_Ready"],"mdate":1773260207027,"odate":1772718030507,"domain":"ICLR.cc/2026/Workshop/VerifAI-2","id":"JEA9UcQncY","version":2},{"content":{"summary":{"value":"This paper focuses on Mesh Field Theory (MeshFT) and its neurbl version, MeshFT-Net, which represents an advance in mesh-based continuum physics. Ideas, such as the definition and separation of topological structure from metric and dissipative components, are novel. The authors define and formalize four minimal, yet critical, physical requirements: locality, permutation equivariance, orientation covariance, and energy balance/passivity. They demonstrate that the dynamics of mesh-based physics within these parameters allow local factorization into a port-Hamiltonian structure. The interconnection is determined exclusively by the mesh, leaving only the metric and dissipation components to be learned. Building on this insight, MeshFT-Net hardwires the signed incidence matrix as the fixed interconnection and learns positive-definite metrics and positive-semidefinite dissipative elements. MeshFT-Net is demonstrated to have near-zero energy drift and significantly improved physical fidelity. For the demonstration of this, authors use analytic plane-wave benchmarks, physics-consistency tests, and a real acoustic scattering dataset, \"The Well,. Authors compare their approach to  MeshGraphNet (MGN), MGN with a Hamiltonian penalty, and Hamiltonian neural networks."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. I think authors should narrow the \"mesh-based physics\" claims or add at least one non-linear / advection example to show the idea generalizes.\n2. Make comparison with GraphCON or other recent work focusing to graph simulators with corrected long-range information flow (qualitatively or quantitively). Also, there are some recent variants of MeshGraphNets, such as PI-MGNs (physics-informed MeshGraphNets), it is interesting to compare with them because paper speaks about data efficiency. \n3. It is interesting to see some list of counterexamples where main theorem about local reduction doesn't work.\n4. As I understand, Figures 2 and 4 provide data-size trends only for the analytic plane-wave task. I have not found it for example for “The Well” dataset. So, please either limit the “5× data efficiency” claim or add train-size analysis for additional datasets."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Demonstrated that mesh based physics, under simple physical assumptions, reduces to a port-Hamiltonian form where the interconnection is determined solely by mesh topology. \n2. MeshFT-Net hard-wires this interconnection and trains exclusively the metric and dissipation terms so that the updates are energy consistent. \n3. For 2D wave and acoustic tests, authors demonstrated near zero energy drift and more accurate wave speed and momentum conservation compared to MeshGraphNet and Hamiltonian baseline.. \n4. Similar or better accuracy is achieved with approximately five times less training data, and generalization occurs for different setups."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Other baselines compared remain classical (MGN, MGN-HP, HNN). Some recent work on graph simulators focus on oversmoothing or long-range dependencies. Additionally, there are some recent variants of MeshGraphNets, such as PI-MGNs (physics-informed MeshGraphNets). These baselines are not included and, as such, makes it harder to assess competitiveness.\n2. Authors talk about “mesh-based physics” broadly, but results are mainly on linear wave/acoustics.\n3. Probably, the \"5x data-efficiency\" claim needs additional verification and strict formulation (not rough)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358948121,"tcdate":1761838647295,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15580/Reviewer_QG3E"],"signatures":["ICLR.cc/2026/Conference/Submission15580/Reviewer_QG3E"],"forum":"dWtJXHZkFy","number":2,"license":"CC BY 4.0","cdate":1761838647295,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15580/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358948121,"domain":"ICLR.cc/2026/Conference","replyto":"dWtJXHZkFy","id":"1sc979xzRH","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"Mesh Field Theory disentangles mesh-based physics into a fixed topological wiring and learned metric operators, yielding stable long-horizon rollouts, high physical fidelity, strong data efficiency, and robust out-of-distribution performance."},"keywords":{"value":["Mesh Field Theory","Mesh-Based Physics","Port–Hamiltonian Dynamics","Structure-Preserving Simulation"]},"supplementary_material":{"value":"/attachment/e8695f4a30d3549a46fefe1578e58acef53f1720.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"We present Mesh Field Theory (MeshFT) and its neural realization, MeshFT-Net: a structure-preserving framework for mesh-based continuum physics that cleanly separates the physics’ topological structure from its metric structure. Imposing minimal physical principles (locality, permutation equivariance, orientation covariance, and energy balance/dissipation inequality), we prove a reduction theorem for mesh-based physics. Under these conditions, the physical dynamics admit a local factorization into a port–Hamiltonian form: the conservative interconnection is fixed uniquely by mesh topology, whereas metric effects enter only through constitutive relations and dissipation. This reduction clarifies what must be fixed and what should be learned, directly informing MeshFT-Net’s design.\nAcross evaluations on analytic and realistic datasets, physics-consistency tests, and out-of-distribution validation, MeshFT-Net achieves near-zero energy drift and strong physical fidelity—correct dispersion and momentum conservation—along with robust extrapolation and high data efficiency. By eliminating non-physical degrees of freedom and learning only metric-dependent structure, MeshFT provides a principled inductive bias for stable, faithful, and data-efficient physical simulation."},"_bibtex":{"value":"@misc{\nnoguchi2026mesh,\ntitle={Mesh Field Theory: Port{\\textendash}Hamiltonian Formulation of Mesh-Based Physics},\nauthor={Satoshi Noguchi and Yoshinobu Kawahara},\nyear={2026},\nurl={https://openreview.net/forum?id=dWtJXHZkFy}\n}"},"title":{"value":"Mesh Field Theory: Port–Hamiltonian Formulation of Mesh-Based Physics"},"pdf":{"value":"/pdf/4d34ac53bfc121e0f6691e6165be779c15bdc1a1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"noguchi|mesh_field_theory_porthamiltonian_formulation_of_meshbased_physics"},"authorids":{"value":["~Satoshi_Noguchi1","~Yoshinobu_Kawahara1"]},"authors":{"value":["Satoshi Noguchi","Yoshinobu Kawahara"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a reasoning framework named \"Deliberate-to-Intuitive\" (D2I), aimed at addressing the challenges of poor modality alignment and high training costs faced by Multimodal Large Language Models (MLLMs) in complex reasoning tasks, such as mathematical problems."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"Same as the Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"1. The main contribution of this paper is quite ingenious. A major bottleneck in current MLLM reasoning research is the high cost of obtaining high-quality, fine-grained reasoning-annotated data. The \"lightweight\" training paradigm proposed in this paper—which uses only rule-based format rewards without relying on any additional human annotations or content-level supervision—is used to enhance complex visual reasoning capabilities.\n2. The evaluation is not limited to in-domain datasets but also covers multiple out-of-domain math benchmarks and general MLLM benchmarks, strongly demonstrating the method's generalization ability.\n3. Rather than stopping at accuracy gains, the paper investigates why D2I works. The analyses of Pass@k, entropy distributions, and token-distribution shifts are persuasive; together they suggest that D2I encourages more exploratory, diverse, and flexible generation strategies, thereby outperforming D2D’s more rigid outputs.\n4. The paper is clearly written and easy to understand."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. An interesting finding of the paper is that different strategies perform differently on various benchmarks (e.g., JUS performs well on MathVerse, while PAR performs better on MathVista and MATH-Vision). The authors attribute this to the task types of the benchmarks (e.g., PAR helps in understanding \"complex visual layouts\"). However, this association is debatable. The core of D2I lies in intuitive reasoning, meaning these strategies are not used at test time. Why, then, would forcing the model to output coordinates (LOC) during training make it perform better on intuitive reasoning tasks that do not require coordinates?\n2. The case study in Figure 3 is very insightful, showing that $D2I_{jus}$ successfully identified \"vertical angles,\" while other models (including $D2I_{loc}$ and $D2I_{par}$) incorrectly identified them as \"adjacent angles\" or \"linear pairs\". However, this case only shows the successes and failures among D2I variants. To more forcefully support the core argument that D2I is superior to D2D (i.e., that D2D lacks flexibility), it is crucial to show the outputs of the D2D models. For example, if a $D2D_{par}$ model could generate the correct image parsing output but still produced the wrong final answer, it would strongly demonstrate the limitations of the D2D paradigm itself, rather than the model's learning ability.\n3. The core premise of this paper is that \"format rewards\" are more scalable than \"content rewards\". This is a reasonable assertion. However, the experiments lack an upper-bound baseline (even a theoretical one) that uses \"content rewards\"."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919029557,"tcdate":1761560006119,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6745/Reviewer_NXGx"],"signatures":["ICLR.cc/2026/Conference/Submission6745/Reviewer_NXGx"],"forum":"jxZ9IYlEsG","number":2,"license":"CC BY 4.0","cdate":1761560006119,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6745/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919029557,"domain":"ICLR.cc/2026/Conference","replyto":"jxZ9IYlEsG","id":"tL8oIDfm4H","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multimodal LLMs","LLM Reasoning","Reinforcement Learning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Reasoning is a key capability for large language models (LLMs), particularly when applied to complex tasks such as mathematical problem solving. \nHowever, multimodal reasoning research still requires further exploration of modality alignment and training costs. Many of these approaches rely on additional data annotation and relevant rule-based rewards to enhance the understanding and reasoning ability, which significantly increases training costs and limits scalability. To address these challenges, we propose the Deliberate-to-Intuitive reasoning framework (D2I) that improves the understanding and reasoning ability of multimodal LLMs (MLLMs) without extra annotations and complex rewards. Specifically, our method sets deliberate reasoning strategies to enhance modality alignment only through the rule-based format reward during training. \nWhile evaluating, the reasoning style shifts to intuitive, which removes deliberate reasoning strategies during training and implicitly reflects the model's acquired abilities in the response. D2I outperforms baselines across both in-domain and out-of-domain benchmarks. Our findings highlight the role of format reward in fostering transferable reasoning skills in MLLMs, and inspire directions for decoupling training-time reasoning depth from test-time response flexibility."},"_bibtex":{"value":"@misc{\nyu2026learning,\ntitle={Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal {LLM}s},\nauthor={Yahan Yu and Yuyang Dong and Masafumi Oyamada},\nyear={2026},\nurl={https://openreview.net/forum?id=jxZ9IYlEsG}\n}"},"title":{"value":"Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs"},"pdf":{"value":"/pdf/2e5cb55022d1bcc5df16db983e0516d8b21742f6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yu|learning_deliberately_acting_intuitively_unlocking_testtime_reasoning_in_multimodal_llms"},"authorids":{"value":["~Yahan_Yu1","~Yuyang_Dong1","~Masafumi_Oyamada1"]},"authors":{"value":["Yahan Yu","Yuyang Dong","Masafumi Oyamada"]}},"version":2},{"content":{"summary":{"value":"This paper introduces the General Physics Transformer (GPhyT) , a large-scale transformer model trained on a diverse 1.8 TB dataset of physics simulations. The authors' goal is to create a \"Physics Foundation Model\" (PFM) that, analogous to LLMs, follows a \"train once, deploy anywhere\" paradigm. The core idea is that a single model can learn to infer the governing physical dynamics (e.g., fluid-solid interaction, shock waves, convection) from a \"prompt\" of prior states, enabling in-context learning. The paper presents three key results: (1) superior performance against specialized architectures on seen physics; (2) zero-shot generalization to unseen boundary conditions and entirely new physical systems ; and (3) stable autoregressive rollouts for 50 timesteps."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The baseline comparison (Q1) uses FNO and UNet. Could the authors comment on why these were chosen over more modern neural operators? Do they believe GPhyT would maintain its performance gap against architectures specifically designed for greater generalization?\n2. The model is fixed-resolution , which is a major limitation compared to neural operators. What are the authors' thoughts on achieving resolution-invariance? Does the \"tubelet\" patching approach  fundamentally prevent this?\n3. How critical is the inclusion of explicit derivatives (dx, dy, dt) in the input ? Have the authors ablated this feature? How much does performance degrade if the model only receives the raw state fields (pressure, velocity, etc.)?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. A Compelling New Paradigm: The primary strength is the successful application of the in-context learning paradigm to physics simulation. The idea that a model can infer dynamics from a prompt of prior states , rather than being explicitly told the equations, is a significant shift from specialized solvers like PINNs or most Neural Operators.\n2. Strong Generalization Results: The paper's most convincing result is the zero-shot generalization to unseen boundary conditions (Table 2). The model achieves nearly identical MSE on \"open\" boundary variants as it does on \"known\" periodic/symmetric ones , demonstrating true generalization. The qualitative generalization to entirely new physics (Fig 4), while less accurate, is a powerful proof of concept.\n3. Large-Scale Data Curation: The 1.8 TB dataset is a significant contribution. More importantly, the data augmentation strategies—specifically variable time incrementsand per-dataset normalization—are intelligently designed to force the model to learn in-context inference rather than memorize specific scales or dynamics.\n4. Thoroughness and Transparency: The paper is rigorous, with detailed ablations on prompt size and integrator schemes . The authors are also very honest about the significant limitations , which grounds the paper's ambitious claims in reality."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Weak Baseline Comparisons: The \"up to 29x\" performance gain is against a standard FNO (from 2020) and a standard UNet. These architectures were not designed for the multi-physics, in-context inference task this paper proposes. A stronger comparison would involve more recent, advanced operator architectures. This overstates the \"breakthrough\" on known physics (Q1).\n2. Major Unsolved Limitations: The authors correctly identify the key limitations, but they are severe. The model is 2D onlyand, most critically, fixed-resolution. This is a significant step backward from the Neural Operator paradigm, which is built on discretization-invariance. A true PFM must be able to handle variable resolutions and 3D systems.\n3. Long-Term Stability is Still Low: While 50-timestep rollouts are demonstrated (Q3) , the error accumulation is non-trivial (Fig 5). The authors admit this \"falls considerably short of the precision exhibited by numerical solvers\" and is not ready for \"practical engineering applications\". This remains a key barrier.\n4. \"New Physics\" Generalization is Only Qualitative: The generalization to unseen physics (supersonic flow, turbulent layers) is visually impressive but shows very high quantitative error (Table 2). This is an exciting proof-of-concept for \"plausibility\" but not yet a robust, accurate generalization.\n5. Reliance on Explicit Derivatives: The model is given numerically computed first-order spatial and temporal derivatives (dx, dy, dt) as input channels . This is a strong inductive bias that likely helps significantly. It slightly undermines the narrative of learning from \"data alone\"  and feels more like a hybrid \"physics-informed\" input feature than a pure in-context inference from raw states."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920445165,"tcdate":1761998235582,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8602/Reviewer_PPrL"],"signatures":["ICLR.cc/2026/Conference/Submission8602/Reviewer_PPrL"],"forum":"q62POvqTLb","number":3,"license":"CC BY 4.0","cdate":1761998235582,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8602/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920445165,"domain":"ICLR.cc/2026/Conference","replyto":"q62POvqTLb","id":"w9AFDsKMIS","forumContent":{"TLDR":{"value":"We present GPhyT, a transformer trained on 1.8TB of diverse physics simulations that can zero-shot generalize to entirely new physical systems by inferring governing dynamics from context alone."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics Foundation Model","Multi-physics Learning","In-context Learning","Zero-shot Generalization","Scientific Machine Learning","Physics-Aware Machine Learning","Spatiotemporal Transformers"]},"supplementary_material":{"value":"/attachment/bb7557fc3f8da9f1e6e373f392add528e405a51d.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Foundation models have revolutionized natural language processing through a ``train once, deploy anywhere'' paradigm, where a single pre-trained model adapts to countless downstream tasks without retraining. Access to a Physics Foundation Model (PFM) would be transformative - democratizing access to high-fidelity simulations, accelerating scientific discovery, and eliminating the need for specialized solver development. Yet current physics-aware machine learning approaches remain fundamentally limited to single, narrow domains and require retraining for each new system. We present the General Physics Transformer (GPhyT), trained on 1.8 TB of diverse simulation data, that demonstrates foundation model capabilities are achievable for physics. Our key insight is that transformers can learn to infer governing dynamics from context, enabling a single model to simulate fluid-solid interactions, shock waves, thermal convection, and multi-phase dynamics without being told the underlying equations. GPhyT achieves three critical breakthroughs: (1) superior performance across multiple physics domains, outperforming specialized architectures by more than 7x, (2) plausible zero-shot generalization to entirely unseen physical systems through in-context learning, and (3) more stable long-term predictions through long-horizon rollouts. By establishing that a single model can learn generalizable physical principles from data alone, this work opens the path toward a universal PFM that could transform computational science and engineering."},"_bibtex":{"value":"@misc{\nwiesner2026towards,\ntitle={Towards a Physics Foundation Model},\nauthor={Florian Wiesner and Matthias Wessling and Stephen Baek},\nyear={2026},\nurl={https://openreview.net/forum?id=q62POvqTLb}\n}"},"title":{"value":"Towards a Physics Foundation Model"},"pdf":{"value":"/pdf/a4f22c125a4e4ee33cdd9e42a7b76196da8bf16f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wiesner|towards_a_physics_foundation_model"},"authorids":{"value":["~Florian_Wiesner1","~Matthias_Wessling1","~Stephen_Baek1"]},"authors":{"value":["Florian Wiesner","Matthias Wessling","Stephen Baek"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a novel framework that enables large language models (LLMs) to create and use their own tools by leveraging external reference materials (e.g. textbooks, knowledge documents). The approach is motivated by the observation that many complex tasks (causal reasoning, physics, chemistry, etc.) require domain knowledge that may not be encoded in the LLM’s parameters. Unlike prior tool-using paradigms, which rely purely on the model’s internal knowledge or a fixed set of APIs, REFTOOL “grounds” the tool creation process in external references."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.\tThe article's innovation seems extremely limited relative to existing work [1], essentially just pre-generating tools in a RAG-like manner. What do the authors believe is the innovative contribution of their research?\n[1] Cai, T., Wang, X., Ma, T., Chen, X., & Zhou, D. Large Language Models as Tool Makers. In The Twelfth International Conference on Learning Representations."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.\tEnabling LLMs to automatically create executable tools from unstructured text references is more efficient than traditional text retrieval. This structured generation approach (including pseudocode and code templates) is more effective than simple text retrieval. By transforming textbooks into a toolbox, REFTOOL empowers models to solve problems previously beyond their capabilities.\n2.\tREFTOOL achieves significant performance improvements on knowledge-intensive benchmarks like causal reasoning, physics, and chemistry, increasing accuracy by 12.3% compared to existing tool-generation methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThe method currently relies on GPT-4 (or similar large models) for tool generation and verification. While understandable (tool synthesis is difficult), this means the initial setup can be expensive and dependent on proprietary models. Without GPT-4 or similar high-performance models, the quality of the generated tools may decline. This introduces an unfairness in experimental baseline comparisons. The article does not deeply explore using open-source LLMs for tool creation; it would be worth knowing how a 70B open-source model performs on this task, even if some performance loss is expected.\n2.\tAlthough overall tool quality is high, the generated tools might be relatively unreliable in some highly complex domains (like chemistry). For instance, chemical reasoning may require extremely complex domain knowledge, and even with textbooks and GPT-4, generating completely correct code isn't guaranteed. The article notes that chemistry tools received relatively low scores in human evaluation. Therefore, for highly specialized or obscure knowledge, the method might require additional fine-tuning (perhaps incorporating symbolic systems or domain expert review).\n3.\tREFTOOL generates hundreds of tools per textbook (e.g., over 500 for physics). Storing and managing such a large toolbox could become overwhelming, especially when considering broader knowledge bases or multiple textbooks. Although hierarchical selection helps, it can become quite complex – for example, if a category still contains dozens of tools, providing all their descriptions to the LLM might exceed its context window limit."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923819771,"tcdate":1762159530881,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13096/Reviewer_M5bi"],"signatures":["ICLR.cc/2026/Conference/Submission13096/Reviewer_M5bi"],"forum":"mgTjEVzyXu","number":4,"license":"CC BY 4.0","cdate":1762159530881,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13096/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923819771,"domain":"ICLR.cc/2026/Conference","replyto":"mgTjEVzyXu","id":"AmydTHoUo4","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Tool Creation","Tool-Augmented Reasoning"]},"supplementary_material":{"value":"/attachment/e8147a9452a6831046071ea015c0ed42cf60dc37.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large Language Models (LLMs) can enhance their reasoning capabilities by using external tools. However, many tasks lack predefined tools. Prior works have explored instructing LLMs to generate tools on their own, but such approaches depend heavily on internal knowledge and struggle when tasks fall outside the model’s knowledge scope. To address this limitation, we propose RefTool, a reference-guided framework for automatic tool creation that leverages external materials, such as textbooks and knowledge snippets. RefTool consists of two modules: (1) tool creation, where LLMs generate executable tools from reference content, validate them using illustrative examples, and organize them hierarchically into a toolbox; and (2) tool utilization, where LLMs navigate the toolbox structure to select and apply the appropriate tools to solve problems. Experiments on causality, physics, and chemistry benchmarks demonstrate that RefTool outperforms existing tool-creation and domain-specific reasoning methods by 12.3% on average accuracy, while being cost-efficient and broadly generalizable to non-scientific tasks, e.g., extremely low-resource language translation. Analyses reveal that grounding tool creation in references produces accurate and faithful tools, and that the hierarchical structure facilitates effective tool selection. RefTool enables LLMs to overcome internal knowledge limitations, advancing generalizable reasoning in knowledge-intensive domains."},"_bibtex":{"value":"@inproceedings{\nliu2026reftool,\ntitle={RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning},\nauthor={Xiao Liu and Da Yin and Zirui Wu and Yansong Feng},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=mgTjEVzyXu}\n}"},"title":{"value":"RefTool: Reference-Guided Tool Creation for Knowledge-Intensive Reasoning"},"pdf":{"value":"/pdf/46655110545a337fed5938020925405c791813b5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|reftool_referenceguided_tool_creation_for_knowledgeintensive_reasoning"},"authorids":{"value":["~Xiao_Liu19","~Da_Yin2","~Zirui_Wu1","~Yansong_Feng1"]},"authors":{"value":["Xiao Liu","Da Yin","Zirui Wu","Yansong Feng"]}},"version":2},{"content":{"venue":{"value":"Complex 2012"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-03473-7_4.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2012"},"paperhash":{"value":"cardonarivera|largescale_conflicts_in_massively_multiplayer_online_games"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Rogelio_Cardona-Rivera:","~Kiran_Lakkaraju3","https://dblp.org/search/pid/api?q=author:Jonathan_Whetzel:","https://dblp.org/search/pid/api?q=author:Jeremy_R._Bernstein:"]},"html":{"value":"https://doi.org/10.1007/978-3-319-03473-7_4"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/Cardona-RiveraLWB12,\n  author={Rogelio Cardona-Rivera and Kiran Lakkaraju and Jonathan Whetzel and Jeremy R. Bernstein},\n  title={Large-Scale Conflicts in Massively Multiplayer Online Games},\n  year={2012},\n  cdate={1325376000000},\n  pages={40-51},\n  url={https://doi.org/10.1007/978-3-319-03473-7_4},\n  booktitle={Complex},\n  crossref={conf/complex/2012}\n}\n"},"abstract":{"value":"Complex systems are of interest to the scientific community due to their ubiquity and diversity in daily life. Popularity notwithstanding, the analysis of complex systems remains a difficult task, due to the problems in capturing high-volume data. Massively Multiplayer Online Games (MMOGs) have recently emerged as a tractable way to analyze complex system interactions, because these virtual environments are able to capture a great amount of data and at high-fidelity, often tracking the actions of many individuals at a time resolution of seconds. MMOGs have been used to study phenomena such as social networks and financial systems; our focus is to identify behaviors related to Large-Scale Conflict (LSC). In this paper, we review how one particular MMOG allows large-scale complex behavior to emerge and we draw parallels between virtual-world LSCs and real-world LSCs. The LSC-related behavior that we are interested in identifying deals with the conditions that lead a participant in the virtual environment (a game player) to engage and actively participate in a LSC, with the goal of informing an agent-based model that predicts when any one player is likely to engage in conflict. We identify virtual world behavioral analogues to real-world behavior of interest (i.e. insurgent behavior), and link the virtual behavior to a (previously derived) theoretical framework that analyzes the determinants of participation in the civil war of Sierra Leone. This framework identifies three general theories that collectively predict participation in civil war (a type of LSC); we operationalize one of the theories (Theory of Social Sanctions), and look at how insurgent behavior can occur as a function of community networks, which are assumed to impose social sanctions for non-participation in an LSC."},"title":{"value":"Large-Scale Conflicts in Massively Multiplayer Online Games"},"authors":{"value":["Rogelio Cardona-Rivera","Kiran Lakkaraju","Jonathan Whetzel","Jeremy R. Bernstein"]}},"tmdate":1728972544399,"pdate":1325376000000,"tcdate":1728972520974,"writers":["~"],"signatures":["~Kiran_Lakkaraju3"],"forum":"jy93luEvyp","license":"CC BY-SA 4.0","number":152217,"cdate":1325376000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1728972544399,"domain":"DBLP.org","id":"jy93luEvyp","version":2},{"content":{"summary":{"value":"This paper introduces a multiple physics pretraining approach for surrogate modeling, which learns general useful features across diverse physical tasks with a shared embedding and normalization strategy. The experiment results show the proposed MPP-pretrained model outperforms task-specific baselines on all pretraining sub-tasks and also show superior finetuning results on new physics tasks."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. Constructing a physics-based foundational model and exploring multiple task pretraining for computational physics tasks is both interesting and beneficial.\n2. The experiments demonstrate impressive results in both pretraining sub-tasks as well as substantial transfer potential for new tasks with low-data system.\n3. The authors have conducted several training strategies to perform the pretraining effectively."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. More attempts can be made to address the transferability problems between different types of physics equations. For instance, when discussing the Navier-Stokes (NS) equations, whether at high or low Reynolds numbers, the equation forms are quite similar, which allows for the investigation into whether a single model still possesses robust merged learning capabilities for more different classes of equations, such as diffusion equations and wave equations.\n\n2. What are the advantages and disadvantages of this new approach compared to the Finite Element Method (FEM) for systems with explicit control equations or empirical formulas? One of my biggest problem is whether the physics task should be treated as a purely data-driven problem, or if it ought to incorporate certain explicit priors or equation-based guidelines.\n\n3. For models without PDEs, how can we determine the similarity of multiple physics fields and whether they can be learned simultaneously? I’m concerned about the improved performances are due to the limited diversity of different physics tasks.\n\n4. In the appendix, experimental data from 1-step to 9-step show little change in the field, whether looking at ground truth or predicted solutions. If the selected time steps were longer, would the model still be able to accurately predict future changes?"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Overall, the purposes behind this work are valuable. However, there remain many questions that need to be addressed. Please consider to answer the above questions."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636222132,"tcdate":1698745067380,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2797/Reviewer_2tFL"],"signatures":["ICLR.cc/2024/Conference/Submission2797/Reviewer_2tFL"],"forum":"fH9eqpCcR3","number":5,"license":"CC BY 4.0","cdate":1698745067380,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2797/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636222132,"domain":"ICLR.cc/2024/Conference","replyto":"fH9eqpCcR3","id":"2gLOUHgBd8","forumContent":{"TLDR":{"value":"We develop approaches to enable autoregressive pretraining on multiple physical systems and show it can improve transfer performance across wide domain gaps."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["transfer learning","physics","pretraining","finetuning","surrogate models","spatiotemporal"]},"supplementary_material":{"value":"/attachment/1b15c36d7da4210bf437c0fbce1293449a229043.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We introduce multiple physics pretraining (MPP), an autoregressive task-agnostic pretraining approach for physical surrogate modeling. MPP involves training large surrogate models to predict the dynamics of multiple heterogeneous physical systems simultaneously by learning features that are broadly useful across diverse physical tasks. In order to learn effectively in this setting, we introduce a shared embedding and normalization strategy that projects the fields of multiple systems into a single shared embedding space. We validate the efficacy of our approach on both pretraining and downstream tasks over a broad fluid mechanics-oriented benchmark. We show that a single MPP-pretrained transformer is able to match or outperform task-specific baselines on all pretraining sub-tasks without the need for finetuning. For downstream tasks, we demonstrate that finetuning MPP-trained models results in more accurate predictions across multiple time-steps on new physics compared to training from scratch or finetuning pretrained video foundation models. We open-source our code and model weights trained at multiple scales for reproducibility and community experimentation. Video examples are included in the supplementary materials."},"_bibtex":{"value":"@misc{\nmccabe2024multiple,\ntitle={Multiple Physics Pretraining for Physical Surrogate Models},\nauthor={Michael McCabe and Bruno R{\\'e}galdo-Saint Blancard and Liam Holden Parker and Ruben Ohana and Miles Cranmer and Alberto Bietti and Michael Eickenberg and Siavash Golkar and Geraud Krawezik and Francois Lanusse and Mariel Pettee and Tiberiu Tesileanu and Kyunghyun Cho and Shirley Ho},\nyear={2024},\nurl={https://openreview.net/forum?id=fH9eqpCcR3}\n}"},"title":{"value":"Multiple Physics Pretraining for Physical Surrogate Models"},"pdf":{"value":"/pdf/8c0116933d8ac09056fd513515a6f906f3b9bb51.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"mccabe|multiple_physics_pretraining_for_physical_surrogate_models"},"authorids":{"value":["~Michael_McCabe2","~Bruno_Régaldo-Saint_Blancard1","~Liam_Holden_Parker1","~Ruben_Ohana1","~Miles_Cranmer2","~Alberto_Bietti1","~Michael_Eickenberg5","~Siavash_Golkar1","gkrawezik@flatironinstitute.org","~Francois_Lanusse2","~Mariel_Pettee1","~Tiberiu_Tesileanu1","~Kyunghyun_Cho1","~Shirley_Ho2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Michael McCabe","Bruno Régaldo-Saint Blancard","Liam Holden Parker","Ruben Ohana","Miles Cranmer","Alberto Bietti","Michael Eickenberg","Siavash Golkar","Geraud Krawezik","Francois Lanusse","Mariel Pettee","Tiberiu Tesileanu","Kyunghyun Cho","Shirley Ho"]}},"version":2},{"content":{"summary":{"value":"The paper first constructs a new physics reasoning benchmark, PhysicsQA, and then proposes to identify errors in reasoning and use reinforcement learning (RL) to refine.\nThe 3 main error types for physics problem solving  identified by the authors are \"Problem Miscomprehension, Incorrect Concept Application, and Calculation Errors.\"\nBy designing different modules to refine wrong reasoning steps, the authors leverage RL to tune small language models (SLMs).\nExperiments across several physics benchmarks have shown some improvements in the accuracy of SLMs."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"See Weakness"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The proposed PhysicsQA dataset might be useful for physics problem-solving with LLMs.\n2. The identified three main types of errors in physics reasoning might be useful for further development of stronger LLMs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"+ This paper brings PhysicsQA, but does not discuss more recent work (early 2025) in constructing benchmarks of physics reasoning.\n+ The reward design is questionable. I believe the reward used in the RL process can be easily hacked by just generating fewer steps. For example, the LLM can just put all steps into a \"single wrong step\".\n+ The experiments fail to cover reasoning models. I do not believe this kind of step-level refinement framework can work for reasoning models. Besides, it is hard to segment the whole reasoning process into steps.\n+ The method mainly relies on the three main error types. I wonder whether these error types are \"complete\", i.e., there might be many other error types. This kind of hand-crafted \"refinement\" module can not generalize to other unseen/undefined error types.\n+ The authors mainly focus on physics reasoning. In fact, the current trend treats physics reasoning along with other scientific reasoning/math reasoning as a whole -> \"reasoning tasks\". Therefore, more benchmarks in other domains should be used.\n+ The presentation is poor and confusing:\n  + The second paragraph of the introduction is really confusing. It seems like content that should appear in \"related work\". I can not see any motivation or differences from prior work in the current version of the introduction.\n  + The writing needs to be thoroughly improved for the \"methodology part\". The authors prefer to spend many words on how prior works solve certain challenges. From the writing, it seems all the techniques are already introduced in previous works, and this paper just puts them together.\n\n+ Others: many format issues\n  + All subsections are wrongly named, e.g., \"2.0.1\" -> \"2.1\"\n  + L154-L155: \"3\"-> \"Figure 3\"\n  + L409 and L418: broken references.\n  + Table captions should be put above the table.\n  + Table 2 exceeds the margin"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921836417,"tcdate":1760689460754,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10563/Reviewer_Zsnf"],"signatures":["ICLR.cc/2026/Conference/Submission10563/Reviewer_Zsnf"],"forum":"DiQ8d2kOxG","number":1,"license":"CC BY 4.0","cdate":1760689460754,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10563/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921836417,"domain":"ICLR.cc/2026/Conference","replyto":"DiQ8d2kOxG","id":"x3XhwbPlYz","forumContent":{"TLDR":{"value":"We introduce a refinement agent and use LoRA-based RLHF with a step-level reward model to improve reasoning using small LLM Models"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large language models","Physics Multi-step Reasoning","RLHF","LLM Agent"]},"supplementary_material":{"value":"/attachment/f77a87d6a789146f35431c9c30b602da96f147c8.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Large Language Models (LLMs) excel at many reasoning tasks but struggle with scientific domains like physics, which demand precise mathematical calculations alongside deep conceptual and factual understanding. In complex physics problem solving, LLMs commonly falter due to three core issues: misunderstanding the problem, incorrect application of concepts, and calculation mistakes. These challenges are more pronounced in small LLMs due to their limited capacity, making them more prone to failures. To address these limitations, we propose a modular reinforcement learning refinement framework tailored for small LLMs, integrating first step error localization, and correction through a Reinforcement Learning guided feedback mechanism. We also introduce PhysicsQA, a diverse benchmark of 370 physics problems designed to evaluate LLM reasoning across the aforementioned dimensions. Experimental results demonstrate improvements upto 10% in final answer accuracy reasoning using Small language models over existing approaches"},"_bibtex":{"value":"@misc{\njaiswal2026modular,\ntitle={Modular Refinement of Small Language Models for Physics Reasoning via Localized Error Feedback},\nauthor={Raj Jaiswal and Dhruv Jain and Rishabh Dhawan and Dhruvkumar Patel and Avinash Anand and Shin'ichi Satoh and Tanuja Ganu and Rajiv Ratn Shah and Erik Cambria and Zhengkui Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=DiQ8d2kOxG}\n}"},"title":{"value":"Modular Refinement of Small Language Models for Physics Reasoning via Localized Error Feedback"},"pdf":{"value":"/pdf/1cd18268e153ed7837fd6150db61d0942c901bf5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"jaiswal|modular_refinement_of_small_language_models_for_physics_reasoning_via_localized_error_feedback"},"authorids":{"value":["~Raj_Jaiswal1","~Dhruv_Jain2","~Rishabh_Dhawan1","~Dhruvkumar_Patel1","~Avinash_Anand1","~Shin'ichi_Satoh1","~Tanuja_Ganu1","~Rajiv_Ratn_Shah1","~Erik_Cambria1","~Zhengkui_Wang1"]},"authors":{"value":["Raj Jaiswal","Dhruv Jain","Rishabh Dhawan","Dhruvkumar Patel","Avinash Anand","Shin'ichi Satoh","Tanuja Ganu","Rajiv Ratn Shah","Erik Cambria","Zhengkui Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces HERON which is a framework to create efficient and robust long-horizon schedules for human robot collaboration by explicitly accounting for human uncertainty. The framework consists of three sequential modules operating in an iterative loop: \n\n1. Task Decompostion LLM: This translates natural language goal into a structured task graph of sub-tasks and dependencies.\n2. Physics-guided LLM: Augments each sub-task with an estimated execution time and an optimal agent assignment, reasoning over physical constraints and cost.\n3. MILP Optimizer: Generates an optimal schedule by solving a Mixed-Integer Linear Program that minimizes cost over makespan and workload distribution, while enforcing all temporal and resource constraints.\n\nA continuous Verifying Stage monitors execution, detects human uncertainty events, and triggers dynamic re-planning by re-entering the planning loop. Experiments across four complex kitchen tasks in the AI2-THOR simulation demonstrate that HERON achieves a 45% higher success rate and a 13% reduction in schedule timespan compared to existing LLM-based planning baselines (SMART-LLM, LiP-LLM, LLaMAR)"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Could the authors provide a simulation experiment where the $\\hat{t}_i$ values generated by the Physics-guided LLM are replaced with values drawn from a Uniform distribution or a simpler analytical model (like the ones used in the prompts, $t=s/(2a)$ for robot travel), and compare the resulting metrics? This would quantitatively isolate the benefit provided by the LLM's \"common-sense\" time estimation from the pure benefit of the MILP solver.\n2. How does the complexity of solving the MILP scale with the number of sub-tasks ($N$)? Since HERON performs dynamic re-planning based on perceived human behavior, the solver must return a new optimal schedule almost instantaneously. Could the authors provide re-planning time vs. $N$ plots to assure the real-time feasibility of the approach, or state the average and max computation times observed during dynamic re-planning?\n3. The paper mentions that VLM-based monitoring to automatically detect performance Variability ($\\xi_1$) lies beyond the current scope but is a future direction. Given that the human only provides feedback on actual completion time $t^{act}_i$ after the robot completes its first task, why not use a simpler form of VLM integration now? For instance, a VLM could visually monitor the human's progress on an assigned task (e.g., slicing a vegetable) to determine if they are currently engaged in the task, allowing for more proactive detection of a slow down $\\xi_1$ *before* the robot completes its step.\n4. While the $C(\\mathcal{S})$ objective function includes a workload distribution term ($\\lambda_{2} \\cdot \\sum \\sum x_{i,a} \\hat{t}_i(a)$), and the failure cases mention scenarios where the robot is idle, the resulting Balance (B) metric for HERON is often lower than LLaMAR's (e.g., Task 3: HERON is 0.56 vs. LLaMAR is 0.80). Please elaborate on the choice of weighting parameters $\\lambda_1$ and $\\lambda_2$. Were these chosen to strictly prioritize minimizing the makespan ($\\lambda_1$) over achieving a balanced workload ($\\lambda_2$), and would altering this trade-off significantly impact the overall success rate?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The framework is fundamentally designed to be resilient and uncertainty-aware, explicitly modeling and dynamically re-scheduling in response to three categories of human stochasticity: performance variability ($\\xi_1$), interruptions ($\\xi_2$), and dynamic goal changes ($\\xi_3$) which makes it more realistic.\n\n2. The overall quality of the evaluation is strong, rigorously comparing HERON against multiple relevant LLM-based planning baselines (SMART-LLM, LiP-LLM, LLaMAR) and validating performance across four task structures (parallel, sequential, hybrid).\n\n3. Ablation studies clearly demonstrate the necessity of both the Physics-guided LLM (although the validity of a physics-guided LLM in this paper is questionable, see weaknesses+questions) and the MILP Optimizer for sustaining high success rates and achieving temporal efficiency, confirming that dynamic planning alone is insufficient."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"major weaknesses:\n\n1. The paper acknowledges that the execution time estimates ($\\hat{t}_i$) generated by the physics-guided LLM were not quantitatively benchmarked against realistic human or robot measurements. Without this quantitative validation, the claim that the LLM generates \"realistic\" times remains qualitative. Since the entire MILP optimization relies on these $\\hat{t}_i$ values to minimize makespan, errors in the LLM's time estimates could lead the optimizer to generate a schedule that is optimal mathematically but suboptimal in reality.\n2. The qualitative failure modes note that the Task Decomposition LLM sometimes produces unnecessarily strict or redundant dependency constraints. This forces parallelizable tasks to execute sequentially, which is likely the core reason why the total timespan (TI) for HERON in some complex tasks (like Task 3) is still quite long compared to the makespan of the purely symbolic LLM baselines (though the Success Rate is higher). This highlights an unaddressed brittleness in the initial LLM parsing step."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921257002,"tcdate":1761858294121,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9765/Reviewer_Qpzk"],"signatures":["ICLR.cc/2026/Conference/Submission9765/Reviewer_Qpzk"],"forum":"QobJeymX6Z","number":1,"license":"CC BY 4.0","cdate":1761858294121,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9765/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921257002,"domain":"ICLR.cc/2026/Conference","replyto":"QobJeymX6Z","id":"W36xAi89Uw","forumContent":{"TLDR":{"value":"A framework that combines LLM-based task decomposition, physics-guided estimation, and MILP optimization to enable efficient and resilient human–robot collaboration under uncertainty."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Human-Robot Collaboration","Long-Horizon Planning","Task Scheduling"]},"supplementary_material":{"value":"/attachment/c0fa68634bf257dcd3b192ea6841103bcf745ee8.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"The integration of humans into long-horizon planning introduces unique challenges that extend beyond conventional robotic task planning. Unlike robots, humans exhibit inherent uncertainty in task execution, including variable performance, unexpected interruptions, and dynamic goal changes, all of which complicate efficient collaboration. To address these challenges, we propose Human-robot collaboration with Efficient and Resilient OptimizatioN for Long-horizon planning (HERON), a novel framework that combines large language models (LLMs), physics-guided reasoning, and optimization techniques. HERON leverages LLMs in two complementary roles: (i) decomposing natural language task descriptions into structured sub-tasks with agent assignments, and (ii) generating physics-guided execution time estimates and determining sub-task assignments for both human and robot agents based on physical constraints and complementarities. These outputs are incorporated into a mixed-integer linear programming scheduler, which dynamically re-schedules based on observed human uncertainties. This integration ensures that scheduling is not only feasible with respect to physical limitations but also robust to human unpredictability while maintaining efficiency in resource and time allocation. Experiments demonstrate that HERON enables resilient and adaptive human-robot collaboration, achieving more efficient scheduling and higher task success rates compared to existing LLM-based planning frameworks. Website at https://sites.google.com/view/heron-planner."},"_bibtex":{"value":"@misc{\nkim2026heron,\ntitle={{HERON}: Human-robot collaboration with Efficient and Resilient OptimizatioN for Long-horizon planning},\nauthor={Taehyeon Kim and Gyeongmin Kim and E. Cho Smith and Byung-Cheol Min},\nyear={2026},\nurl={https://openreview.net/forum?id=QobJeymX6Z}\n}"},"title":{"value":"HERON: Human-robot collaboration with Efficient and Resilient OptimizatioN for Long-horizon planning"},"pdf":{"value":"/pdf/50037fd37f8eac3749f9a003072abfa94ece03ef.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|heron_humanrobot_collaboration_with_efficient_and_resilient_optimization_for_longhorizon_planning"},"authorids":{"value":["~Taehyeon_Kim3","~Gyeongmin_Kim3","~E._Cho_Smith1","~Byung-Cheol_Min1"]},"authors":{"value":["Taehyeon Kim","Gyeongmin Kim","E. Cho Smith","Byung-Cheol Min"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Air-DualODE, a novel approach for air quality prediction that combines physics-based and data-driven methods using dual Neural ODEs. The physics branch implements a modified diffusion-advection equation with a correction term for open systems (BA-DAE). In contrast, the data-driven branch employs masked attention-based Neural ODEs to capture unknown dynamics. The two branches are temporally aligned using a decaying contrastive learning scheme and fused in latent space using GNN, demonstrating superior performance on city-scale (Beijing) and national-scale (KnowAir) datasets."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Q1. Given that real air pollution sources (industrial activity, vehicle emissions) and sinks (forests, lakes) exhibit complex non-linear relationships, why did you simplify the correction term $\\beta X$ as a linear term? What is the physical justification for setting $\\beta$'s range to $[−1, +∞)$?\n\nQ2. While you claim that the Physics branch explicitly models physical phenomena, how is this physical interpretability preserved when projecting into latent space?\n\nQ3. Regarding the temporal alignment process using Decay-TCL, how do the chosen values of $\\lambda_1 = 1$ and $\\lambda_2 =0.8$ guarantee physically meaningful alignment? What is the physical significance of using time-decaying weights?\n\nQ4. Why specifically choose Spatial-MSA in the Data-Driven branch? How does this align with a physics-informed approach?\n\nQ5. The authors justify GNN fusion based on 'distance-dependent influence', but isn't this characteristic already considered in the Physics branch?\n\nQ6. How can the authors justify the performance on the national-scale KnowAir dataset when case studies are limited to the Beijing dataset?\n\nQ7. How does the visualization of $\\beta$ values correspond to actual observed pollution source/sink data?\"\n\nQ8. Can authors perform sensitivity analysis for different ranges of $\\beta$ values?\n\nQ9. The authors seem to only consider the DOPRI5 ODE solver. Could they analyze performance and runtime differences when using simpler methods like Euler or RK4?\n\nQ10. The paper should reference and compare with recent work on climate modeling using diffusion and diffusion-advection equations in neural ODE frameworks [1,2]. Can authors clarify their position by analyzing similarities and differences in their approach to diffusion and advection?\n\nQ11. How does the intentional violation of conservation law in BA-DAE affect numerical stability, particularly for ODE solvers?\n\nQ12. How do you ensure that the BA-DAE in the Physics branch and Neural ODE in the Data-driven branch operate in the same state space?\n\nQ14. Is 'Physics-Informed' appropriate in the title? Would 'Physics-guided' or 'Physics-inspired' be more accurate, given that this might be confused with traditional PINN approaches?\n\nQ15. Can authors provide more details about the RNN used in the Coefficient Estimator?\n\nQ16. Figure 2 lacks several elements mentioned in the text, particularly α from equation 6. The relationship between Dynamics fusion and Section 3.4 equations needs clarification.\n\nQ17. Can authors provide visualizations or distribution analyses showing how Gdiff and Gadv change dynamically with wind speed and direction?\n\nQ18. While criticizing the computational cost of existing physics-based methods, how does your dual branch architecture with a complex fusion mechanism improve efficiency? Doesn't using two ODE solvers increase computational burden?\n\nQ19. Can the authors include the number of forward evaluations (NFE) comparisons in Table 2's ablation studies?\n\nQ20. The authors should cover related work on GNNs that redesign the diffusion equation [3,4] and its variations[5,6] using NODE. Can you discuss more about what the authors' methods have in common and what they differ from?\n\n> [1] Choi, Hwangyong, et al. \"Climate modeling with neural advection-diffusion equation.\" Knowledge and Information Systems 65.6 (2023): 2403-2427.\n> \n> [2] Hwang, Jeehyun, et al. \"Climate modeling with neural diffusion equations.\" 2021 IEEE International Conference on Data Mining (ICDM). IEEE, 2021.\n>\n> [3] Wang, Yifei, et al. \"Dissecting the diffusion process in linear graph convolutional networks.\" Advances in Neural Information Processing Systems 34 (2021): 5758-5769.\n>\n> [4] Chamberlain, Ben, et al. \"Grand: Graph neural diffusion.\" International conference on machine learning. PMLR, 2021.\n>\n> [5] Thorpe, Matthew, et al. \"GRAND++: Graph neural diffusion with a source term.\" ICLR (2022).\n>\n> [6] Choi, Jeongwhan, et al. \"Gread: Graph neural reaction-diffusion networks.\" International Conference on Machine Learning. PMLR, 2023."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses the limitations of pure physics-based and pure data-driven approaches by proposing a hybrid framework that attempts to leverage the advantages of both methods.\n\n2. The introduction of BA-DAE with a correction term represents an attempt to model open system dynamics, which is more realistic for air quality prediction than traditional closed system assumptions.\n\n3. The model achieves state-of-the-art performance across different spatial scales while maintaining some level of interpretability through its physics branch."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper oversimplifies complex air pollution dynamics using a linear correction term (βX) without proper theoretical justification, undermining its claim of accurate open system modeling.\n\n2. The approach loses physical interpretability when projecting to latent space and violates conservation laws, raising concerns about numerical stability and contradicting the paper's emphasis on physics-informed modeling.\n\n3. The computational efficiency claims are questionable as the dual branch architecture with multiple ODE solvers likely increases computational burden rather than reducing it.\n\n4. The experimental validation is limited, with case studies confined to Beijing data and lacking crucial analyses such as parameter sensitivity testing and solver comparisons.\n\n5. The technical documentation is incomplete, with key mathematical elements missing from figures and insufficient details about architectural choices, making reproducibility challenging."}},"nonreaders":[],"tmdate":1733143773073,"tcdate":1730720955681,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7746/Reviewer_6Tsm"],"signatures":["ICLR.cc/2025/Conference/Submission7746/Reviewer_6Tsm"],"forum":"kOJf7Dklyv","number":5,"license":"CC BY 4.0","cdate":1730720955681,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7746/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733143773073,"domain":"ICLR.cc/2025/Conference","replyto":"kOJf7Dklyv","id":"Ys1pxLTDIm","forumContent":{"TLDR":{"value":"A novel approach for physics-guided Neural ODE"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Air Quality Prediction; Physics-guided Deep Learning"]},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Air pollution significantly threatens human health and ecosystems, necessitating effective air quality prediction to inform public policy. Traditional approaches are generally categorized into physics-based and data-driven models. Physics-based models usually struggle with high computational demands and closed-system assumptions, while data-driven models may overlook essential physical dynamics, confusing the capturing of spatiotemporal correlations. Although some physics-guided approaches combine the strengths of both models, they often face a mismatch between explicit physical equations and implicit learned representations. To address these challenges, we propose Air-DualODE, a novel physics-guided approach that integrates dual branches of Neural ODEs for air quality prediction. The first branch applies open-system physical equations to capture spatiotemporal dependencies for learning physics dynamics, while the second branch identifies the dependencies not addressed by the first in a fully data-driven way. These dual representations are temporally aligned and fused to enhance prediction accuracy. Our experimental results demonstrate that Air-DualODE achieves state-of-the-art performance in predicting pollutant concentrations across various spatial scales, thereby offering a promising solution for real-world air quality challenges."},"_bibtex":{"value":"@inproceedings{\ntian2025air,\ntitle={Air Quality Prediction with Physics-Guided Dual Neural {ODE}s in Open Systems},\nauthor={Jindong Tian and Yuxuan Liang and Ronghui Xu and Peng Chen and Chenjuan Guo and Aoying Zhou and Lujia Pan and Zhongwen Rao and Bin Yang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=kOJf7Dklyv}\n}"},"title":{"value":"Air Quality Prediction with Physics-Guided Dual Neural ODEs in Open Systems"},"pdf":{"value":"/pdf/2dfd149abbbfe06c5b3c1f3e31a93bb4a9220042.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"tian|air_quality_prediction_with_physicsguided_dual_neural_odes_in_open_systems"},"authorids":{"value":["~Jindong_Tian1","~Yuxuan_Liang1","~Ronghui_Xu2","~Peng_Chen14","~Chenjuan_Guo1","~Aoying_Zhou1","~Lujia_Pan2","~Zhongwen_Rao1","~Bin_Yang4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jindong Tian","Yuxuan Liang","Ronghui Xu","Peng Chen","Chenjuan Guo","Aoying Zhou","Lujia Pan","Zhongwen Rao","Bin Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Physics Informed Neurally Constructed ODE Networks (PINECONEs), a pipeline to combine the Neural ODE family with physics-informed loss. The authors evaluate this framework on transport equations and Burger’s equations, compared with PINNs. The proposed method shows faster convergence and better accuracy when using first-order optimization methods."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- A framework is proposed by combining Neural ODE architectures and physics-informed loss. \n- This paper is easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The idea is not novel. There are already many works investigating the potential of combining neural differential equations with physics-informed loss [1,2,3]. \n\n- The baselines are not sufficient. The proposed method is only compared with standard PINNs. There are many variants of the PINN family, which show better performance [4,5,6]. To convince the readers, I think more baselines are expected.\n\n- The proposed method is only tested on 1D problems. There are many successful implementations of PINNs in 2D and 3D cases [4,5,6], but this paper only investigates 1D systems.\n\n---\n\n**Refs:**\n\n[1] Ji, Weiqi, et al. \"Stiff-pinn: Physics-informed neural network for stiff chemical kinetics.\" The Journal of Physical Chemistry A 125.36 (2021): 8098-8106.\n\n[2] Lai, Zhilu, et al. \"Structural identification with physics-informed neural ordinary differential equations.\" Journal of Sound and Vibration 508 (2021): 116196.\n\n[3] O'Leary, Jared, Joel A. Paulson, and Ali Mesbah. \"Stochastic physics-informed neural ordinary differential equations.\" Journal of Computational Physics 468 (2022): 111466.\n\n[4] Cho, Junwoo, et al. \"Separable Physics-Informed Neural Networks.\" arXiv preprint arXiv:2306.15969 (2023).\n\n[5] Wang, Sifan, Hanwen Wang, and Paris Perdikaris. \"Learning the solution operator of parametric partial differential equations with physics-informed DeepONets.\" Science advances 7.40 (2021): eabi8605.\n\n[6] Wang, Sifan, Shyam Sankaran, and Paris Perdikaris. \"Respecting causality is all you need for training physics-informed neural networks.\" arXiv preprint arXiv:2203.07404 (2022)."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Please see my concerns in **Weaknesses**."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637059117,"tcdate":1697841020287,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8482/Reviewer_cWov"],"signatures":["ICLR.cc/2024/Conference/Submission8482/Reviewer_cWov"],"forum":"TB5THwq1sq","number":1,"license":"CC BY 4.0","cdate":1697841020287,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8482/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637059117,"domain":"ICLR.cc/2024/Conference","replyto":"TB5THwq1sq","id":"qEguTaSqLy","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Scientific Machine Learning","Neural ODEs","PINNs","PDEs"]},"supplementary_material":{"value":"/attachment/45b469e8182ff7d4397b26dca2fa8a5593047a9f.pdf"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Recently, there has been a growing interest in using neural networks to approximate the solutions of partial differential equations (PDEs). Physics-informed neural networks (PINNs) have emerged as a promising framework for parameterizing PDE solutions using deep neural networks. However, PINNs often rely on memory-intensive optimizers to attain reasonable accuracy and can encounter training difficulties due to issues such as stiffness in the gradient flow of the loss. To address these challenges, we propose a novel network architecture that combines neural ordinary differential equations (ODEs) with physics-informed constraints in the loss function. In this approach, the dynamics within a neural ODE are expanded to include a system of ODEs whose solution provides the partial derivatives governing our PDE system. We call this architecture PINECONEs: physics-informed neurally constructed ODE networks. We evaluate the approach using simple but canonical PDEs from the literature to illustrate its potential. Our results show that training requires fewer iterations than previous approaches to achieve higher accuracy when using first-order optimization methods."},"_bibtex":{"value":"@misc{\nmckay2024physics,\ntitle={Physics Informed Neurally Constructed {ODE} Networks ({PIN}e{CON}es)},\nauthor={Maricela Best Mckay and Brian Wetton and Bhushan Gopaluni},\nyear={2024},\nurl={https://openreview.net/forum?id=TB5THwq1sq}\n}"},"title":{"value":"Physics Informed Neurally Constructed ODE Networks (PINeCONes)"},"pdf":{"value":"/pdf/eb690ac37ae7e839774e6c65bb6196df80c643ee.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"mckay|physics_informed_neurally_constructed_ode_networks_pinecones"},"authorids":{"value":["~Maricela_Best_Mckay1","~Brian_Wetton1","~Bhushan_Gopaluni1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Maricela Best Mckay","Brian Wetton","Bhushan Gopaluni"]}},"version":2},{"content":{"summary":{"value":"The paper presents SDT-Net, a deep learning framework for tracking space debris in complex skylight backgrounds, and introduces the Space Debris Tracking Dataset (SDTD), a large-scale synthetic dataset containing 18,040 video sequences with 62,562 frames and 250,000 synthetic debris instances. SDT-Net integrates feature enhancement, detection, and tracking modules to achieve high accuracy in cluttered, occluded, and dense debris environments. Evaluations on both synthetic and real Antarctic telescope data demonstrate strong performance, achieving a 73.2% MOTA score. The study highlights the potential of deep learning for real-time, transferable debris tracking and establishes SDTD as a benchmark for future research.Is this conversation helpful so far."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Here are several relevant prior works that the authors should cite, covering space-debris detection/tracking, multi-object tracking in astronomy, datasets, and physics-informed approaches:\n\n**Space-debris detection/tracking and optical observations**\n\nCament, L. et al., “Space Debris Tracking with the Poisson Labeled Multi-Bernoulli Multi-target Tracking Filter”, Sensors, 21(11):3684, 2021.\n\n“A Robust Vision-based Algorithm for Detecting and Classifying Small Orbital Debris” (NASA MSFC) – algorithm for small debris using optical detection. \nNASA Technical Reports Server\n\nNavya, M. et al., “Deep Learning-Based Space Debris Tracking and Mitigation”, J Electrical Systems, 20(1):606-611, 2024. \n\nZhou, D., Sun, G., Zhang, Z., Wu, L., “On Deep Recurrent Reinforcement Learning for Active Visual Tracking of Space Non-cooperative Objects”, arXiv:2212.14304, Dec 2022. \n\nRoll, D. S., Kurt, Z., Woo, W. L., “CosmosDSR – a methodology for automated detection and tracking of orbital debris using the Unscented Kalman Filter”, arXiv:2310.17158, Oct 2023. \n\n**Astronomical multi-object tracking / star-field object tracking**\n\nGuan, J., Cheng, H-Y., Wu, Y-P., Tian, C., Qi, J-Y., “Multi-target tracking for star sensor based on CenterTrack deep learning model”, Scientific Reports 15:37125 (2025). \n\n**Space-debris modelling / simulation and environment context**\n\nKim, et al., “Review of Space Debris Modeling Methods and Development Trends”, Journal of Astronautical Sciences, 41(4):209-… (2024) \n\nESA Space Debris Environment Report, https://sdup.esoc.esa.int/discosweb/statistics/, sdup.esoc.esa.int\n\n**Deep learning object/tracking methods in cluttered/low SNR astronomical settings**\n\nSDebrisNet: “SDebrisNet: A Spatial–Temporal Saliency Network for Space Debris”, Applied Sciences 13(8):4955 (2023). \n\n**Benchmarks/datasets for debris/satellite detection**\n\nThe authors should mention existing optical/space-object detection datasets, even if only for detection (not tracking) to position their contribution. For example, the Kaggle “Debris Detection Dataset” (optical images) – though limited. \n\n**Orbit/dynamics embedding into tracking**\n\nAlthough not directly DL-tracking, works that link vision tracking with orbit dynamics may strengthen the discussion. For example the PINN-based tracking after collision: “Tracking an Untracked Space Debris After an Inelastic Collision Using Physics Informed Neural Network”, arXiv:2307.09938 (2023)."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper proposes a deep-learning approach for space debris tracking in complex skylight backgrounds. The main contributions are:\n\n1.The authors introduce a novel dataset, the Space Debris Tracking Dataset (SDTD), created by an observation-based simulation scheme, drawing on astronomy images (from e.g. the Zwicky Transient Facility, ZTF) and synthetically imposing debris trajectories and backgrounds. The dataset reportedly includes 18,040 video sequences (≈ 62,562 frames) and ~250,000 synthetic debris instances. \nMoonlight\n\n2.They propose a network named SDT‑Net, which comprises a Region-of-Interest Feature Enhancement (RoIFE) module, a detection module and a tracking module (tracking by detection plus association across frames). The network is targeted at the multi-object tracking (MOT) task in astronomical / debris scenarios. \n\n3.They conduct experiments on their synthetic dataset and also evaluate transfer to real-world data: they claim a MOTA score (Multiple Object Tracking Accuracy) of ~73.2% (or ~70.6% in some versions) on a small real dataset collected at an Antarctic station. \n\n4.They argue that their dataset addresses the paucity of annotated debris-tracking data, and that SDT-Net exhibits robustness under dense debris, occlusion, and complex star-field backgrounds.\n\nThus the paper is an attempt to bring modern deep-MOT methods into the space-debris tracking domain, supported by a large synthetic benchmark."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the paper makes interesting advances, there are several concerns and weaknesses that the authors should address:\n\n1. Synthetic-to-real transfer gap / dataset realism\n\n1) Although the dataset is large and simulation-based, synthetic data may not fully replicate the statistical characteristics of real debris tracks, noise sources, background clutter, telescope artefacts, or imaging conditions (e.g., atmospheric scintillation, streak brightness variation, non-uniform PSF, variations in exposure times, sensor noise). The authors do test on a small real dataset, but the size is tiny (36 video sequences, ~2,228 frames) and limited to one station (Antarctic). This raises questions about generalisability to other sensors, orbital regimes, debris sizes, lighting conditions, star-field densities.\n\n2) The paper reports a single performance number on real data; more extensive evaluation across different observational setups would strengthen the claim of “strong transferability”.\n\n2. Dataset annotation / ground-truth fidelity and bias\n\n1) The synthetic generation process may introduce biases (e.g., debris speed, size, appearance, background variation) that favour their method, especially since the method is trained on the synthetic data. It’s unclear how well annotation errors, occlusion patterns, sensor artefacts, and false positives/negatives are handled.\n\n2) The real data annotations (astronomy experts) are limited in quantity; the annotation criteria, inter-annotator consistency, debris definitions (what qualifies as debris vs star/artefact) may affect reproducibility.\n\n3. Evaluation metrics and baseline comparisons\n\n1) The paper uses MOTA as a key metric; however, MOTA alone may not capture fine issues like ID-switches, fragmentation, false alarms in cluttered star fields, long-term track survival, or tracking latency (important for real-time/operational use).\n\n2) The baselines compared are relatively generic MOT methods (e.g., CenterTrack, OCSORT) rather than domain-specific methods tailored to astronomical debris or long-exposure streak detection. A stronger argument would include recent astronomy/space-debris tracking methods.\n\n3) The paper claims “state-of-the-art”, but many details about run-time, sensor input frame rate, false positive/false negative rates, resource usage (GPU/CPU) are missing. For an operational system, these are important.\n\n4. Scalability and real-time viability\n\n1) Space debris tracking in real operational settings often demands real-time or near-real-time performance, dealing with many debris objects, variable frame rates, large fields of view, and possibly resource-constrained platforms. The paper does not sufficiently discuss latency, computational load, or memory constraints.\n\n2) Dense debris scenarios (e.g., mega-constellations, low Earth orbit clutter) may stress the method beyond the distribution of synthetic data; how well does it scale beyond the densities in the dataset?\n\n5. Lack of orbital/physical modelling integration\n\n1) The method appears largely vision-based (image/video processing) without explicit incorporation of orbital dynamics, sensor geometry, debris kinematics, or space situational awareness (SSA) context (e.g., Two-Line Elements, orbital propagation). In many practical applications, combining image tracking with orbital dynamics yields more robust performance. The paper doesn’t show how their output could link to orbit prediction or catalogue maintenance.\n\n2) Without physics-based constraints (motion models, known debris motion patterns), the tracker may fail in ambiguous scenarios (e.g., overlapping tracks, rapid acceleration, non-linear motion), and the paper does not explore these limitations in depth.\n\n6. Generalisation to other observational platforms\n\n1) The dataset is constructed from ZTF images (ground-based optical telescope) and the real evaluation is from a single station. It is unclear how well the method would generalise to different sensors: e.g., space-based optical imagers, radar, different exposure times, spectral bands, or telescopes with different PSFs, different background noise levels, different orbital altitudes.\n\n2) The authors should discuss how the method would adapt to e.g., GEO, MEO, or LEO regimes, or to different sensors (infrared, radar) or daytime/nighttime imaging.\n\n7. Benchmark release and reproducibility\n\n1) The paper mentions that “dataset and code will be released soon”. Without immediate availability, reproducibility and community uptake may be limited. The authors should commit to making the dataset, annotations, evaluation scripts and code available under a clear license, and provide a leaderboard or standard evaluation split.\n\n2) If synthetic only, there is a risk that future users will duplicate their simulation bias. Clear documentation of simulation parameters, debris motion models, background modelling is needed.\n\n8. Limited real-world deployment discussion\n\n1) The paper could benefit from a deeper discussion of how this tracking method would integrate into operational debris tracking pipelines, what the false alarm risk is, how track continuity and object correlation across multiple passes/sensors would be handled, and what the end-to-end system implications are (e.g., collision avoidance, catalogue updating).\n\n2) It is also unclear how many frames per second, what field of view, and what detection sensitivity (size/magnitude of debris) the system supports; practical relevance to e.g., <10 cm debris tracking is not characterised."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918874373,"tcdate":1762465663165,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6509/Reviewer_9kga"],"signatures":["ICLR.cc/2026/Conference/Submission6509/Reviewer_9kga"],"forum":"j63W4sMjFE","number":4,"license":"CC BY 4.0","cdate":1762465663165,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6509/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918874373,"domain":"ICLR.cc/2026/Conference","replyto":"j63W4sMjFE","id":"bocoz5dggB","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["deep learning","AI for science","object tracking","dataset"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"With the rapid development of space exploration, space debris has attracted more attention due to its potential extreme threat, leading to the need for real-time and accurate debris tracking. However, existing methods are mainly based on traditional signal processing, which cannot effectively process the complex background and dense space debris. In this paper, we propose a deep learning-based Space Debris Tracking Network (SDT-Net) to achieve highly accurate debris tracking. SDT-Net effectively represents the feature of debris, enhancing the efficiency and stability of end-to-end model learning. To train and evaluate this model effectively, we also produce a large-scale dataset Space Debris Tracking Dataset (SDTD) by a novel observation-based data simulation scheme. SDTD contains 18,040 video sequences with a total of 62,562 frames and covers 250,000 synthetic space debris. Extensive experiments validate the effectiveness of our model and the challenging of our dataset. Furthermore, we test our model on real data from the Antarctic Station, achieving a MOTA score of 73.2%, which demonstrates its strong transferability to real-world scenarios."},"_bibtex":{"value":"@misc{\nzhuang2026high,\ntitle={High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset},\nauthor={Guohang Zhuang and Weixi Song and Jinyang Huang and chenwei yang and Wanli Ouyang and Yan Lu},\nyear={2026},\nurl={https://openreview.net/forum?id=j63W4sMjFE}\n}"},"title":{"value":"High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset"},"pdf":{"value":"/pdf/8ee8017238f5cd82f47ede42c9644e8f9740636b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhuang|high_performance_space_debris_tracking_in_complex_skylight_backgrounds_with_a_largescale_dataset"},"authorids":{"value":["~Guohang_Zhuang1","~Weixi_Song2","~Jinyang_Huang1","~chenwei_yang1","~Wanli_Ouyang1","~Yan_Lu10"]},"authors":{"value":["Guohang Zhuang","Weixi Song","Jinyang Huang","chenwei yang","Wanli Ouyang","Yan Lu"]}},"version":2},{"content":{"summary":{"value":"The authors introduce a interpretation of gradient decent based deep learning through the lens of thermodynamics. At first glance the idea is interesting and timely, however the evidence to support based on synthetic data it is very weak."},"correctness":{"value":"2: Fair"},"soundness":{"value":"2: Fair"},"strengths":{"value":"- It is an interesting perspective to reconcile deep learning with physics inspired learning.\n- Some analogues make intuitive sense (e.g. energy as loss, entropy as parameter entropy, and learning rate as temperature)"},"weaknesses":{"value":"- “Parameter entropy” computed as mean log-variance of weights is a questionable proxy for entropy\n-Experiments are only on synthetic data with small models. The claims of “phase transitions” and “entropy dissipation” are qualitative and rely on this questionable entropy proxy.\n- There are no details about the synthetic data."},"confidence":{"value":2},"rating":{"value":4}},"parentInvitations":"NLDL.org/2026/Abstracts_Track/-/Official_Review","nonreaders":[],"tmdate":1762345638263,"tcdate":1761914281188,"writers":["NLDL.org/2026/Abstracts_Track","NLDL.org/2026/Abstracts_Track/Submission43/Reviewer_zSNs"],"signatures":["NLDL.org/2026/Abstracts_Track/Submission43/Reviewer_zSNs"],"forum":"6aU0wfyIoz","number":1,"license":"CC BY 4.0","cdate":1761914281188,"readers":["everyone"],"invitations":["NLDL.org/2026/Abstracts_Track/Submission43/-/Official_Review","NLDL.org/2026/Abstracts_Track/-/Edit"],"mdate":1762345638263,"domain":"NLDL.org/2026/Abstracts_Track","replyto":"6aU0wfyIoz","id":"Enq6WgkE1t","forumContent":{},"version":2},{"content":{"summary":{"value":"This paper introduces a training-free, probabilistic framework for controllable visual generation by combining heterogeneous pre-trained \"expert\" models (e.g., generative models, discriminative VLMs, and physics simulator software). The core of the method is a novel sampling algorithm that draws from the combined \"Product of Experts\" distribution by interleaving three techniques: Annealed Importance Sampling (AIS) to gradually refine samples from noise, MCMC to ensure fidelity to the generative experts, and Sequential Monte Carlo (SMC) to resample and filter particles. The method is instantiated and evaluated on complex tasks, including graphics-engine-instructed image editing, physics-instructed video generation, and layout-controlled text-to-image synthesis. The results demonstrate superior controllability over baselines, effectively adhering to precise object poses and physics-based motion trajectories while maintaining high visual quality."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses section."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper provides a classic and mathematically sound algorithm for combining heterogeneous pre-trained models, including both generative and discriminative experts. The method can generate images and videos with high controllability without extra training.\n\n- The paper demonstrates many instantiations of this framework across different tasks (image editing, video generation, layout control) and expert types (flow models, autoregressive models, VLMs, physics engines), and the proposed method consistently outperforms monolithic baseline approaches.\n\n- The main paper and the appendix provide extensive implementation details, qualitative results, and ablations, which support its claims and reproducibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The method is inherently slow due to its iterative nature, requiring $T$ annealing steps, $L$ parallel particles, and $K$ MCMC steps per particle. This makes it less practical for real-world GenAI applications, with image generation taking ~4 minutes and video generation taking 5-30 minutes. In my understanding, the proposed method is more like a proof of concept instead of a feasible path for future visual generation frameworks.\n\n- The use of the physics engine can be \"ad-hoc.\" For instance, the autoregressive video task requires the user to provide a full RGB rendering from the simulator. If a user must **manually** create complex 3D renderings or animations in professional software like Blender to serve as a control signal, the practical cost and effort may be prohibitively high."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924547206,"tcdate":1761668841705,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14067/Reviewer_WhsG"],"signatures":["ICLR.cc/2026/Conference/Submission14067/Reviewer_WhsG"],"forum":"dTYbqgvZmc","number":1,"license":"CC BY 4.0","cdate":1761668841705,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14067/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924547206,"domain":"ICLR.cc/2026/Conference","replyto":"dTYbqgvZmc","id":"8LmBXDUG6K","forumContent":{"TLDR":{"value":"An inference-time sampling framework for image and video generation that compose knowledge from heterogeneous perceptual models."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["generative models","image generation","video generation"]},"supplementary_material":{"value":"/attachment/6315e54bc6734f58fa09865c6b27531f1dcf2b89.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Modern neural models capture rich priors and have complementary knowledge over shared data domains, e.g., images and videos. Integrating diverse knowledge from multiple sources—including visual generative models, visual language models, and sources with human-crafted knowledge such as graphics engines and physics simulators remains under-explored. We propose a probabilistic framework that combines information from these heterogeneous models, where expert models jointly shape a product distribution over outputs. To sample from this product distribution for controllable image/video synthesis tasks, we introduce an annealed  MCMC sampler in combination with SMC-style resampling to enable efficient inference-time model composition. Our framework empirically yields better controllability than monolithic methods and additionally provides flexible user interfaces for specifying visual generation goals."},"_bibtex":{"value":"@inproceedings{\nzhang2026product,\ntitle={Product of Experts for Visual Generation},\nauthor={Yunzhi Zhang and Carson Murtuza-Lanier and Zizhang Li and Yilun Du and Jiajun Wu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=dTYbqgvZmc}\n}"},"title":{"value":"Product of Experts for Visual Generation"},"pdf":{"value":"/pdf/042df847ea9eb76e6519dc134d3b69147dc38a68.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|product_of_experts_for_visual_generation"},"authorids":{"value":["~Yunzhi_Zhang1","~Carson_Murtuza-Lanier1","~Zizhang_Li1","~Yilun_Du1","~Jiajun_Wu1"]},"authors":{"value":["Yunzhi Zhang","Carson Murtuza-Lanier","Zizhang Li","Yilun Du","Jiajun Wu"]}},"version":2},{"content":{"summary":{"value":"This paper presents MAPS (Multi-Modal Scientific Reasoning with Physics Perception and Simulation), a novel framework designed to improve the performance of Multi-Modal Large Language Models (MLLMs) in expert-level scientific reasoning tasks, specifically in physical sciences. MAPS enhances the comprehension and analytical processes by integrating a Physics Perception Model (PPM) with a simulator for interpreting complex physical diagrams and quantitative reasoning."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"The paper demonstrates the effectiveness of MAPS using circuit analysis problems. Can the authors provide evidence or insights into how MAPS could be adapted for other physical sciences, such as mechanics or optics, which involve different types of diagrams and simulation languages?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"S1: The integration of perception and simulation for multi-modal scientific reasoning is well-conceived and leverages MLLM strengths while mitigating their weaknesses in handling complex diagrams.\n\nS2: Results from circuit analysis problems highlight a notable increase in accuracy, showcasing MAPS' ability to outperform current state-of-the-art methods.\n\nS3: The paper provides comprehensive explanations for data synthesis, PPM training, and the inference process, enhancing reproducibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1: The framework is tested primarily on circuit analysis, which may not fully capture its adaptability across different physical sciences.\n\nW2: The multi-step process involving diagram conversion, SL generation, and simulation may introduce cumulative errors, which could affect real-world applicability.\n\nW3: The reliance on synthetic data poses a challenge for real-world accuracy, as unseen or complex diagrams might not align with generated examples."}},"nonreaders":[],"tmdate":1731429289422,"tcdate":1730534982124,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11154/Reviewer_1DpV"],"signatures":["ICLR.cc/2025/Conference/Submission11154/Reviewer_1DpV"],"forum":"GR0y0F3Ipd","number":2,"license":"CC BY 4.0","cdate":1730534982124,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11154/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429289422,"domain":"ICLR.cc/2025/Conference","replyto":"GR0y0F3Ipd","id":"cxAGvn7not","forumContent":{"TLDR":{"value":"improving multi-modal scientific reasoning capability with physics perception model and simulation assistance"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-modal reasoning","scientific reasoning","physical simulation"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Pre-trained on extensive text and image corpora, current Multi-Modal Large Language Models (MLLM) have shown strong capabilities in general visual reasoning tasks. \nHowever, their performance is still lacking in physical domains that require understanding diagrams with complex physical structures and quantitative analysis based on multi-modal information. \nTo address this, we develop a new framework, named **M**ulti-Modal Scientific Re**A**soning with **P**hysics Perception and **S**imulation (**MAPS**) based on an MLLM. \nMAPS decomposes expert-level multi-modal reasoning task into physical diagram understanding via a Physical Perception Model (PPM) and reasoning with physical knowledge via a simulator. \nThe PPM module is obtained by fine-tuning a visual language model using carefully designed synthetic data with paired physical diagrams and corresponding simulation language descriptions. \nAt the inference stage, MAPS integrates the simulation language description of the input diagram provided by PPM and results obtained through a Chain-of-Simulation process with MLLM to derive the underlying rationale and the final answer. \nValidated using our collected college-level circuit analysis problems, MAPS significantly improves reasoning accuracy of MLLM and outperforms all existing models. \nThe results confirm MAPS offers a promising direction for enhancing multi-modal scientific reasoning ability of MLLMs. \nWe will release our code, model and dataset used for our experiments upon publishing of this paper."},"_bibtex":{"value":"@inproceedings{\nzhu2025maps,\ntitle={{MAPS}: Advancing Multi-Modal Reasoning in Expert-Level Physical Science},\nauthor={Erle Zhu and Yadi Liu and Zhe Zhang and Xujun Li and JinZhou and Xinjie Yu and Minlie Huang and Hongning Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=GR0y0F3Ipd}\n}"},"title":{"value":"MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science"},"pdf":{"value":"/pdf/c84516a8e4b9a68b710453218bcaa36dee327176.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhu|maps_advancing_multimodal_reasoning_in_expertlevel_physical_science"},"authorids":{"value":["~Erle_Zhu1","~Yadi_Liu1","~Zhe_Zhang24","~Xujun_Li2","~JinZhou1","~Xinjie_Yu1","~Minlie_Huang1","~Hongning_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Erle Zhu","Yadi Liu","Zhe Zhang","Xujun Li","JinZhou","Xinjie Yu","Minlie Huang","Hongning Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces LOCA, a framework for automatically cleaning scientific QA corpora through logical chain augmentation. LOCA reconstructs reasoning by inserting missing logical steps and separating each step into principles and derivations. Experiments on three physics benchmarks show that LOCA reduces corpus error rates from around 20% to below 2%."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. How sensitive are results to the number of review iterations (Ncorr, Nwrg)? How they chosen in the paper?\n2. In lines 38–39, the authors state that “our own expert analysis reveals that error rates in major benchmarks’ QA pairs can exceed 20%.” How is this “error” defined? What criteria or guidelines were provided to the experts? Does it include both incorrect answers (as shown in the appendix) and logically incomplete solutions?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"LOCA identifies an  issue (logical incompleteness of solution) in problem-solving benchmarks and proposes a solution to tackle this. \n\nAuthors show empirical results across PHYBench, PHYSICS, and ABench-Physics. It shows that LOCA can reduce the residual error rate compared to baselines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The framework mainly applies to problem-solving questions in derivation-heavy domains like physics or mathematics, while the title and introduction claim a broader scope across scientific domains. This seems inaccurate since corpora in other domains can contain other issues such as factual errors, unclear writing, or formatting problems. Clarifying the scope would make the paper more accurate.  \n- The definition of “error” appears to focus on logical incompleteness rather than final-answer correctness. In my opinion, it is somewhat unclear whether a solution with a correct final answer but missing intermediate steps should be considered erroneous. The paper would benefit from clarifying what kinds of reasoning flaws LOCA is designed to detect, and why such augmentation matters if a knowledgeable reader could easily fill in those steps.  \n- Evaluation benchmarks are limited to physics. Including results from other scientific fields, such as math, would strengthen the generalizability of LOCA.  \n- The iterative review loop depends on LLM judgment. However the paper lacks analysis of the reliability of these LLM-based reviewers. The paper would benefit from some human evaluation of review quality.\n- The dataset used for experiments is relatively small: only 100 questions per benchmark are sampled from larger datasets (e.g., PHYBench with 500 problems, PHYSICS with 1297).  \n- Hyperparameter choices (Ncorr, Nwrg) are not justified or analyzed for sensitivity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925121859,"tcdate":1761971642094,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14769/Reviewer_Vcy6"],"signatures":["ICLR.cc/2026/Conference/Submission14769/Reviewer_Vcy6"],"forum":"kdFjucrq7B","number":3,"license":"CC BY 4.0","cdate":1761971642094,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14769/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925121859,"domain":"ICLR.cc/2026/Conference","replyto":"kdFjucrq7B","id":"OnDPE9OcqM","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["scientific corpus cleaning","logical chain","AI for science","LLMs"]},"supplementary_material":{"value":"/attachment/9bd9e56c7096711b6b2d9ada26ff81d426ddc305.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"While Large Language Models (LLMs) excel in general domains, their reliability often falls short in scientific problem-solving. The advancement of scientific AI depends on large-scale, high-quality corpora. However, existing scientific question-answering (QA) datasets suffer from high error rates, frequently resulting from logical leaps and implicit reasoning within the answers. To address this issue, we introduce LOCA (Logical Chain Augmentation), a novel framework for automatically cleaning scientific corpora, implemented through an augment-and-review loop. At its core, LOCA enhances raw answers by completing missing logical steps and explicitly separating the underlying scientific principle from its subsequent derivation. By applying LOCA to challenging scientific corpora, we demonstrate that it can automatically filter noisy datasets, typically reducing the error rate from as high as 20\\% to below 2\\%. LOCA provides a scalable and effective methodology for creating high-quality scientific corpora, paving the way for more reliable training and evaluation of scientific AI."},"_bibtex":{"value":"@misc{\nfang2026loca,\ntitle={{LOCA}: Logical Chain Augmentation for Scientific Corpus Cleaning},\nauthor={Youle Fang and Dong-Shan Jian and Xiang Li and Ce Meng and Lingshi Meng and Chen-Xu Yan and Zhizhang Bian and Yan-Qing Ma},\nyear={2026},\nurl={https://openreview.net/forum?id=kdFjucrq7B}\n}"},"title":{"value":"LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning"},"pdf":{"value":"/pdf/6adf0ad32849f9012988103f077cd7de25b8d589.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"fang|loca_logical_chain_augmentation_for_scientific_corpus_cleaning"},"authorids":{"value":["~Youle_Fang1","~Dong-Shan_Jian1","~Xiang_Li161","~Ce_Meng1","~Lingshi_Meng3","~Chen-Xu_Yan1","~Zhizhang_Bian1","~Yan-Qing_Ma1"]},"authors":{"value":["Youle Fang","Dong-Shan Jian","Xiang Li","Ce Meng","Lingshi Meng","Chen-Xu Yan","Zhizhang Bian","Yan-Qing Ma"]}},"version":2},{"content":{"venue":{"value":"Nature Reviews Physics"},"pdf":{"value":"https://www.nature.com/articles/s42254-018-0005-3.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"battiston|taking_census_of_physics"},"html":{"value":"https://doi.org/10.1038/s42254-018-0005-3"},"abstract":{"value":"Over the past decades, the diversity of areas explored by physicists has exploded, encompassing new topics from biophysics and chemical physics to network science. However, it is unclear how these new subfields emerged from the traditional subject areas and how physicists explore them. To map out the evolution of physics subfields, here, we take an intellectual census of physics by studying physicists’ careers. We use a large-scale publication data set, identify the subfields of 135,877 physicists and quantify their heterogeneous birth, growth and migration patterns among research areas. We find that the majority of physicists began their careers in only three subfields, branching out to other areas at later career stages, with different rates and transition times. Furthermore, we analyse the productivity, impact and team sizes across different subfields, finding drastic changes attributable to the recent rise in large-scale collaborations. This detailed, longitudinal census of physics can inform resource allocation policies and provide students, editors and scientists with a broader view of the field’s internal dynamics. An analysis of the number of physicists and their career paths reveals the changing landscape of the physics subdisciplines, highlighting the connections between different fields and the effects of large collaborations."},"title":{"value":"Taking census of physics"},"authors":{"value":[{"fullname":"Federico Battiston","username":"https://orcid.org/orcid-search/search?searchQuery=Federico%20Battiston"},{"fullname":"Federico Musciotto","username":"https://orcid.org/orcid-search/search?searchQuery=Federico%20Musciotto"},{"fullname":"Dashun Wang","username":"https://orcid.org/orcid-search/search?searchQuery=Dashun%20Wang"},{"fullname":"Albert-László Barabási","username":"https://orcid.org/orcid-search/search?searchQuery=Albert-L%C3%A1szl%C3%B3%20Barab%C3%A1si"},{"fullname":"Michael Szell","username":"~Michael_Szell1"},{"fullname":"Roberta Sinatra","username":"https://orcid.org/orcid-search/search?searchQuery=Roberta%20Sinatra"}]}},"tmdate":1779472590284,"pdate":1546905600000,"externalIds":["doi:10.1038/s42254-018-0005-3"],"tcdate":1779472583569,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Michael_Szell1"],"forum":"ONztptEmdU","license":"CC BY-SA 4.0","number":71187,"cdate":1545246305152,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1779472590284,"domain":"OpenReview.net/Public_Article","id":"ONztptEmdU","version":2},{"content":{"summary":{"value":"This paper proposes a meta-learning approach to learning a synthetic model of the RL environment as a proxy for finding the optimal policy, instead of directly interacting with the real environment. The authors formulate the problem as a meta-learning problem and solve it using evolution strategy (ES). To overcome the difficulty of learning the dynamics of the actual environment, the authors propose that learning a synthetic *contextual bandit* (SCB) model is sufficient to achieve the goal and train the agent policy. They conduct experiments and provide ablation studies to analyze the choice and variants of the proposed method."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. Using a simpler model (synthetic contextual bandit, SCB) as a proxy is an interesting idea when environment is complex.\n2. Experiments on interpretability justify the choice of SCB as the proxy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"One of the main contributions of this paper is the use of a synthetic contextual bandit (SCB) as a proxy to the real environment. However, it is important to note that:\n\n* A contextual bandit (CB) can be converted to a Markov decision process (MDP), but not vice versa, because CB is stateless. This means that the SCB model may not be able to accurately capture the dynamics of more complex environments, such as Go, where state is essential for planning and decision-making.\n* It is also unclear how synthetic contextual bandit can work for partial observation environments, where the agent does not have access to all of the state information.\n\nIn other words, the SCB model may be able to learn to play simple games, such as Atari Breakout, where the state space is relatively small and statelessness is not a major issue. However, for more complex games, such as Go, where the state space is very large and statelessness is a major issue, the SCB model is unlikely to be able to learn to play at a high level. Additionally, it is unclear how the SCB model would perform in partial observation environments.\n\nIn addition to the limitations of the SCB model discussed above, there are several other issues with this paper:\n\n* One important paper is not cited nor discussed, which is closely related:\n\nFerreira et al. (2022) proposed a similar approach of learning synthetic environments and reward networks for reinforcement learning. It would be helpful to discuss the relationship between this work and the proposed method.\n\n* There is no comparison with the state-of-the-art results on the environments/tasks in this paper. It is understandable if the proposed method does not outperform the state of the art, as the inner policy optimization can be different. However, it would be interesting to see how the proposed method compares to other methods on the same tasks.\n\n* The motivation for using a synthetic environment is not entirely clear. In the first paragraph of the introduction, the authors mention the biological fact that organisms can learn from artificial stimuli. However, it is not clear how this relates to the use of synthetic environments in reinforcement learning. Additionally, the authors do not explain why learning or meta-learning RL policies is not sufficient.\n\n* The colors in Figure 5 should be revised so that synthetic curves have the same color, and the same for the real curves. This would make the figure easier to read and interpret.\n\nOverall, the use of a SCB as a proxy to the real environment is a promising approach, but it is important to be aware of its limitations. For more complex environments, such as Go, and partial observation environments, models like MCTS are still necessary for learning to play at a high level. The authors need to address the issues raised above in order to strengthen their work.\n\nReference\n\nFerreira, F., Nierhoff, T., Sälinger, A., & Hutter, F. *Learning Synthetic Environments and Reward Networks for Reinforcement Learning*. In ICLR 2022."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. Why do the authors choose to evaluate the proposed method on Brax environments, instead of Gym or other popular environments?\n\n2. The authors claim that limiting the episode length to one step in the synthetic environment does not qualitatively affect the performance of the agent (Section 4, last subsection; Section 5, beginning). However, I could not find any figure or table in the paper that shows this result. Can the authors please provide this information?\n\n3. The authors claim that \"even state-of-the-art RL algorithms such as PPO struggle with solving MountainCar\" (Section 1, second paragraph). However, I am not sure if this is true. Can the authors please provide citations to support this claim?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636839771,"tcdate":1698813783964,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission7108/Reviewer_XskH"],"signatures":["ICLR.cc/2024/Conference/Submission7108/Reviewer_XskH"],"forum":"VDkye4EKVe","number":3,"license":"CC BY 4.0","cdate":1698813783964,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission7108/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636839771,"domain":"ICLR.cc/2024/Conference","replyto":"VDkye4EKVe","id":"Gl3RuMKVCT","forumContent":{"TLDR":{"value":"We meta-optimize neural networks representing contextual bandits that rapidly train agents and broadly generalize."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Reinforcement Learning","Meta-Learning","Evolution Strategies","Environments"]},"supplementary_material":{"value":"/attachment/8d62a90cd9973786f748995dd7c6238e10656932.zip"},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Human agents often acquire skills under conditions that are significantly different from the context in which the skill is needed. For example, students prepare for an exam not by taking it, but by studying books or supplementary material. Can artificial agents benefit from training outside of their evaluation environment as well? In this project, we develop a novel meta-optimization framework to discover neural network-based synthetic environments.  We find that training contextual bandits suffices to train Reinforcement Learning agents that generalize well to their evaluation environment, eliminating the need to meta-learn a transition function. We show that the synthetic contextual bandits train Reinforcement Learning agents in a fraction of time steps and wall clock time, and generalize across hyperparameter settings and algorithms. Using our method in combination with a curriculum on the performance evaluation horizon, we are able to achieve competitive results on a number of challenging continuous control problems. Our approach opens a multitude of new research directions: Contextual bandits are easy to interpret, yielding insights into the tasks that are encoded by the evaluation environment. Additionally, we demonstrate that synthetic environments can be used in downstream meta-learning setups, derive a new policy from the differentiable reward function, and show that the synthetic environments generalize to entirely different optimization settings."},"_bibtex":{"value":"@misc{\nliesen2024discovering,\ntitle={Discovering Minimal Reinforcement Learning Environments},\nauthor={Jarek Luca Liesen and Chris Lu and Andrei Lupu and Jakob Nicolaus Foerster and Henning Sprekeler and Robert Tjarko Lange},\nyear={2024},\nurl={https://openreview.net/forum?id=VDkye4EKVe}\n}"},"title":{"value":"Discovering Minimal Reinforcement Learning Environments"},"pdf":{"value":"/pdf/bb34aea1e571e048f32abf23220142f6f76dfb44.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"liesen|discovering_minimal_reinforcement_learning_environments"},"authorids":{"value":["~Jarek_Luca_Liesen1","~Chris_Lu1","~Andrei_Lupu1","~Jakob_Nicolaus_Foerster1","~Henning_Sprekeler1","~Robert_Tjarko_Lange1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jarek Luca Liesen","Chris Lu","Andrei Lupu","Jakob Nicolaus Foerster","Henning Sprekeler","Robert Tjarko Lange"]}},"version":2},{"content":{"summary":{"value":"The authors present a physics-informed offline reinforcement learning framework for optimizing energy efficiency in data center cooling systems, addressing critical challenges like limited data and safety constraints. Using a graph neural network model that respects time-reversal symmetry, the framework enables efficient and robust policy learning from real-world operational data. The authors claimed that this method was successfully deployed in a large-scale commercial data center and achieved 14-18% energy savings over 1300 hours without violating safety constraints. The work demonstrates the potential of offline RL for complex, data-limited industrial applications and calls for a shift from simulation-based benchmarks to real-world problems for more practical and impactful RL research."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The proposed solution integrates a physics-informed dynamics model to accurately capture the complex thermal behavior within the server room, paired with a graph neural network that embeds domain knowledge to reduce data requirements. \n2. The authors claim that this approach produces well-structured and generalizable latent representations, facilitating a sample-efficient offline RL algorithm that maximizes the value function in latent space with appropriate regularization.\n3. The implementation includes a safety-aware reward function to ensure operational reliability.\n4. The premise of this paper is that as offline RL enables efficient policy learning from pre-collected data, eliminating the risks and costs associated with continuous interaction in safety-critical or resource-constrained environments, it is more effective and practical than online RL."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Claimed but Not Established: The paper asserts strong out-of-distribution (OOD) generalization capabilities and effectiveness with limited real-world data, but these claims are insufficiently substantiated.\n2. Lack of Industry Baseline: The real-world validation experiments show 14-18% energy savings in DC cooling without safety violations, but the paper is unable to present any well-defined industry practice baseline with similar objectives for comparison under similar constraints. The comparison lacks fairness. Also, a thorough optimization metric for a data center operation should factor in elements other than safety violations. \n3. Model Generalizability & Insufficient Benchmark Comparison: The method heavily relies on modeling, but the generalizability and robustness of the modeling technique remain unverified. The method's performance is not evaluated against established benchmarks, limiting the validation of its general effectiveness.\n4. Data and Experiment Limitations: The method's performance evaluation is constrained by the definitions of the experimental setup and the data distributions used in this study, which is not standardized. \n5. Minimal Algorithmic Novelty: The approach offers little innovation compared to existing methods, limiting its algorithmic contribution.\n6. Comparison with other Physics-informed modeling: The paper does not convincingly demonstrate how the proposed approach is superior to well-established physics simulation models bootstrapped with collected data that use online RL to train. The claimed higher sample efficiency with the physics-informed model is not unique, as similar benefits are observed with both online and offline RL approaches.\n7. Unclear Baseline Performance: There is an inadequate explanation for why aggressive baseline methods like CCA and CQL achieve lower energy consumption but fail to maintain critical thermal safety.\n\nIn summary, this paper would be better suited for a domain conference centered around physics modeling and the application of standard AI techniques for optimization."}},"nonreaders":[],"tmdate":1734401026675,"tcdate":1730719508451,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1285/Reviewer_TQKL"],"signatures":["ICLR.cc/2025/Conference/Submission1285/Reviewer_TQKL"],"forum":"W8xukd70cU","number":4,"license":"CC BY 4.0","cdate":1730719508451,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1285/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1734401026675,"domain":"ICLR.cc/2025/Conference","replyto":"W8xukd70cU","id":"NMkxhKshNW","forumContent":{"TLDR":{"value":"We developed and deployed a sample-efficient offline RL framework for energy efficiency optimization of data center's cooling system"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Offline Reinforcement learning","data center optimization","cooling system","energy saving"]},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The recent advances in information technology and artificial intelligence have fueled a rapid expansion of the data center (DC) industry worldwide, accompanied by an immense appetite for electricity to power the DCs. In a typical DC, around 30-40% of the energy is spent on the cooling system rather than on computer servers, posing a pressing need for developing new energy-saving optimization technologies for DC cooling systems. However, optimizing such real-world industrial systems faces numerous challenges, including but not limited to a lack of reliable simulation environments, limited historical data, and stringent safety and control robustness requirements. In this work, we present a novel physics-informed offline reinforcement learning (RL) framework for energy efficiency optimization of DC cooling systems. The proposed framework models the complex dynamical patterns and physical dependencies inside a server room using a purposely designed graph neural network architecture that is compliant with the fundamental time-reversal symmetry. Because of its well-behaved and generalizable state-action representations, the model enables sample-efficient and robust latent space offline policy learning using limited real-world operational data. Our framework has been successfully deployed and verified in a large-scale production DC for closed-loop control of its air-cooling units (ACUs). We conducted a total of 2000 hours of short and long-term experiments in the production DC environment. The results show that our method achieves 14-21% energy savings in the DC cooling system, without any violation of the safety or operational constraints. We have also conducted a comprehensive evaluation of our approach in a real-world DC testbed environment. Our results have demonstrated the significant potential of offline RL in solving a broad range of data-limited, safety-critical real-world industrial control problems."},"_bibtex":{"value":"@inproceedings{\nzhan2025data,\ntitle={Data Center Cooling System Optimization Using Offline Reinforcement Learning},\nauthor={Xianyuan Zhan and Xiangyu Zhu and Peng Cheng and Xiao Hu and Ziteng He and Hanfei Geng and Jichao Leng and Huiwen Zheng and Chenhui Liu and Tianshun Hong and Yan Liang and Yunxin Liu and Feng Zhao},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=W8xukd70cU}\n}"},"title":{"value":"Data Center Cooling System Optimization Using Offline Reinforcement Learning"},"pdf":{"value":"/pdf/b14590ff619ed3e0470e1dbf49699b24b14ac11f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhan|data_center_cooling_system_optimization_using_offline_reinforcement_learning"},"authorids":{"value":["~Xianyuan_Zhan1","~Xiangyu_Zhu4","~Peng_Cheng1","~Xiao_Hu7","~Ziteng_He1","~Hanfei_Geng1","~Jichao_Leng1","~Huiwen_Zheng2","~Chenhui_Liu1","~Tianshun_Hong1","~Yan_Liang6","~Yunxin_Liu2","~Feng_Zhao11"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xianyuan Zhan","Xiangyu Zhu","Peng Cheng","Xiao Hu","Ziteng He","Hanfei Geng","Jichao Leng","Huiwen Zheng","Chenhui Liu","Tianshun Hong","Yan Liang","Yunxin Liu","Feng Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper introduces FOLIAGE, a latent world model designed to understand and predict the dynamics of accretive surface growth—a phenomenon where surfaces grow by adding new material, such as in plants. The model features three core components: (1) An \"accretion-aware\" perception stack that fuses partial and multimodal sensor inputs (images, point clouds, meshes) into a compact latent state, using correspondence constraints and vertex age features to focus on newly grown regions. (2) An action-conditioned dynamics model that predicts the evolution of this latent state based on changes to material properties. (3) A physics-guided training paradigm that leverages privileged simulator information (e.g., per-vertex energies) at training time to shape the latent representation, while remaining solver-free at deployment. To support this research, the authors also introduce a new simulation platform (SURF-GARDEN) and an evaluation suite (SURF-BENCH), on which FOLIAGE demonstrates strong performance across various tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Could the authors better position their contribution for the ICLR audience? Are the core principles of FOLIAGE—such as age-based positional encoding or correspondence-driven multimodal fusion—applicable to more general and widely studied problems in dynamic 3D understanding, such as human motion analysis, long-term robotic interaction, or video prediction?\n2. The work relies entirely on the new SURF-GARDEN dataset. How do the authors see the proposed methods generalizing beyond this simulated environment? Were any experiments conducted on real-world data, or on other standard dynamic 3D datasets, even if they don't feature the specific \"accretive growth\" property?\n3. The physics-guided training is a key component. How much of the model's strong performance is attributable to having access to a high-fidelity simulator with privileged energy information, versus the novelty of the FOLIAGE architecture itself? Is it possible that a simpler model could perform similarly if given access to the same high-quality, physically-grounded data?\n4. Given the system's significant complexity, what is the single most important architectural or methodological insight that the broader community should take away from this work? If a researcher cannot replicate the entire SURF-GARDEN platform, what is the key component they could readily apply to their own problems?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. A major strength of this work is the introduction of the SURF-GARDEN platform and SURF-BENCH evaluation suite. This is a significant contribution that formalizes a challenging new problem domain, providing a high-quality dataset, a physically-based simulator, and a diverse set of tasks. This infrastructure is valuable for enabling future research in this specific area.\n2. The FOLIAGE model is technically sophisticated and thoughtfully designed. The approach of fusing heterogeneous sensor data via correspondence, using age features to encode temporal dynamics, and separating a deployable encoder from a privileged, physics-guided target encoder for training is a novel and powerful paradigm for learning from complex physical systems.\n3. The paper presents extensive experiments demonstrating that FOLIAGE consistently outperforms a wide range of strong baselines across the six core tasks defined in SURF-BENCH. The results are thorough, and the stress tests effectively showcase the robustness of the proposed model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The primary weakness of this work is its focus on \"accretive surface growth,\" a highly specialized sub-field of computational physics and biology. While the problem is scientifically interesting, its relevance and potential impact on the broader machine learning and representation learning community at ICLR is questionable. The paper does not make a compelling case for why this specific physical phenomenon should be of general interest, nor how the methods developed would apply to more mainstream problems in vision, graphics, or robotics.\n2. The paper's main achievement is the construction of a large, complex, and highly integrated system for solving one specific task. While impressive from an engineering perspective, it is difficult to distill a single, core machine learning method or principle that is broadly applicable. The individual components are clever adaptations of existing tools (GNNs, Transformers, EMA), but the overall contribution feels more like a strong application paper for a specialized venue (e.g., computer graphics, computational science) rather than a foundational methods paper for ICLR.\n3. The FOLIAGE system is exceptionally complex, involving a custom simulator, multiple specialized encoders for different modalities, a heterogeneous graph fusion mechanism, and a specific physics-guided training loop. This high degree of complexity, coupled with the reliance on a new custom dataset, may significantly limit the work's adoption, reproducibility, and the ability for other researchers to build upon its ideas."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925646371,"tcdate":1760786369219,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15361/Reviewer_Wk6Y"],"signatures":["ICLR.cc/2026/Conference/Submission15361/Reviewer_Wk6Y"],"forum":"eTXTOUrrhY","number":1,"license":"CC BY 4.0","cdate":1760786369219,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15361/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925646371,"domain":"ICLR.cc/2026/Conference","replyto":"eTXTOUrrhY","id":"ANa5dnx7Iq","forumContent":{"TLDR":{"value":"We present FOLIAGE, a latent world model for unbounded accretive surface evolution"},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Surface Growth","World Model","Multimodal"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Accretive surfaces grow by adding material and changing rest metrics, producing emergent, complex, and changing morphologies. We introduce FOLIAGE, a geometry-centric latent world model that infers a deployable state from heterogeneous, partial sensors and predicts its action-conditioned evolution. The perception stack aligns images, point clouds, and meshes through correspondence-constrained fusion and age features, then pools into global and young-region summaries that emphasize where change will occur next. Dynamics input act only on the latent, taking material coefficients and a horizon code to produce counterfactual roll-outs without entangling perception with control. Training-time physics guides representation via a target encoder that receives per-vertex energies and energy-gated message passing, while the deployable path relies solely on observable inputs. On the SURF-GARDEN data platform and the SURF-BENCH suite, FOLIAGE improves mesh topology classification by ~3 pp, reduces dense-correspondence geodesic error by ~10\\%, lifts cross-modal retrieval by ~25\\% mAP@100, increases growth-stage recognition by ~8 pp, lowers 5-step Chamfer by ~20\\%, and cuts inverse-material error by ~40\\% relative to strong baselines. Stress tests show graceful degradation under sensor loss, stable long-horizon roll-outs, and gains from train-only physics without test-time privileges. Code and datasets used in this study will be made publicly available upon publication to facilitate reproducibility and further research."},"_bibtex":{"value":"@misc{\nanonymous2026foliage,\ntitle={{FOLIAGE}: a Latent World Model for Accretive Surface Growth},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=eTXTOUrrhY}\n}"},"title":{"value":"FOLIAGE: a Latent World Model for Accretive Surface Growth"},"pdf":{"value":"/pdf/10b1072240dbcdc709a9e9ac2c4a5deccfedb375.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"liu|foliage_a_latent_world_model_for_accretive_surface_growth"},"authorids":{"value":["~Xiaoyi_Liu7","~Hao_Tang6"]},"authors":{"value":["Xiaoyi Liu","Hao Tang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes DIET-PATE, a variant of PATE that removes the need for public data drawn from the same distribution as the private dataset. The key idea is to (1) pretrain both teachers and student on large procedurally generated image datasets, and (2) apply data-free knowledge distillation with adapted batch-normalization statistics to reduce the covariate shift between synthetic and private data. The authors extend the approach to a distributed collaborative setting, DIET-CaPC, which integrates DIET-PATE into the CaPC framework and replaces expensive private inference on encrypted data with standard inference on synthetic queries, enabling the training of a single shared student. Experiments on MNIST, CIFAR-10, and TissueMNIST show that DIET-PATE approaches or exceeds standard PATE under comparable privacy budgets when in-distribution public data is absent, and that DIET-CaPC yields substantial efficiency gains and better privacy–utility trade-offs than CaPC."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please refer to 'Weaknesses'."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Addresses a real limitation of PATE and CaPC.\nThe paper tackles a well-known practical bottleneck of PATE: the dependence on public in-distribution data, which is often unrealistic in sensitive domains such as healthcare or finance. Likewise, DIET-CaPC directly targets CaPC’s main pain points—slow private inference and fragmented privacy budgets leading to only modest local improvements. The problem formulation and motivation are clear and compelling.\n\n2. Extension to distributed collaboration:\nDIET-CaPC is a meaningful extension, e.g., replacing encrypted private queries with synthetic queries removes the need for HE/MPC-based private inference and allows larger models compared to the CryptoNet-style networks used in CaPC. The protocol description (teachers, student party, privacy guardian) is clear, and the empirical comparison shows both speedups and more favorable privacy–utility trade-offs. \n\n3. Extensive evaluation and ablations: The experimental section is fairly thorough.\n\n4. Clarity and reproducibility:\nThe paper is overall well-written and easy to follow. Implementation details are described in the appendix with specificity to help reproduce the experiments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Incremental methodological novelty: It seems like DIET-PATE is an integration of PATE, programmatically generated pretraining, and data-free KD with BN “current statistics.” Each component is taken from prior work; the main novelty lies in combining them to remove PATE’s reliance on public in-distribution data and in porting this combination into CaPC. While this integration is practically useful, the methodological contribution may be seen as somewhat incremental for a top-tier venue, especially since there is no new privacy analysis, theoretical insight, or substantially new algorithmic mechanism beyond the pipeline design.\n\n2. Limited experimental scope and scalability:\nThe evaluation focuses on image classification tasks with relatively standard-sized models (ResNet10/18) and modest-resolution data (MNIST, CIFAR-10, TissueMNIST). It remains unclear how DIET-PATE behaves with larger architectures or more complex modalities (e.g., high-resolution images) where pretraining and synthetic generation might be more challenging. Given that the paper emphasizes improved scalability for DIET-CaPC, additional experiments or a more in-depth discussion of scaling bottlenecks would be valuable.  ￼\n\n3. Lack of security/robustness analysis:\nThe work is framed around privacy, but does not investigate robustness to adversarial manipulation of the synthetic data or the protocol. For example, what happens if the synthetic generation process is biased or partially corrupted, or if an adversary contributes malicious teachers in DIET-CaPC? Even a discussion of such risks and possible mitigations would enhance the paper’s practical relevance for sensitive domains.\n\n4. Intuitive design choices:\nIn DIET-PATE, teachers fine-tune only the last layer on private data, and the student also fine-tunes a small fraction of parameters; the paper shows empirically that this helps alignment and performance, but these choices are not deeply analyzed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915855722,"tcdate":1762705564516,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1685/Reviewer_kPqL"],"signatures":["ICLR.cc/2026/Conference/Submission1685/Reviewer_kPqL"],"forum":"Tqusxp1tXg","number":4,"license":"CC BY 4.0","cdate":1762705564516,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1685/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915855722,"domain":"ICLR.cc/2026/Conference","replyto":"Tqusxp1tXg","id":"yn82WYsdf0","forumContent":{"TLDR":{"value":"We present a distributed private learning approach that eliminates both PATE's requirement for in-distribution public data and CaPC's reliance on costly private inference, using programmatically generated data and data-free distillation."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["PATE","programmatically generated data","data free knowledge distillation","CaPC","privacy","differential privacy"]},"supplementary_material":{"value":"/attachment/ef1da9bb90a9fc7d53a3b1f55d778f5000422539.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"The PATE algorithm is one of the canonical approaches to private machine learning. It leverages a private dataset to label a public dataset, enabling knowledge transfer from teachers to a student model under differential privacy (DP) guarantees. However, PATE's reliance on public data from the same distribution as the private data poses a fundamental limitation, particularly in domains such as healthcare and finance, where in-distribution public data is typically unavailable. In this work, we propose DIET-PATE which overcomes this limitation. Therefore, it combines programmatically generated data and data-free knowledge distillation. Our experiments demonstrate that DIET-PATE closely matches the performance of standard PATE, despite the absence of in-distribution public data. Furthermore, we show that our approach seamlessly extends to distributed collaborative learning with CaPC. In this setting, only PATE-based learning can be used to provide DP guarantees, as teacher models are trained by different entities and exchange knowledge solely via labels. By eliminating the need for in-distribution data during knowledge transfer, our method removes CaPC’s reliance on private inference with encrypted data, substantially reducing computational overhead and, for the first time, enabling the use of more complex models and learning tasks. Moreover, leveraging programmatically generated data allows parties in CaPC to jointly train a global model, rather than just improving local ones, thereby achieving significantly higher utility. These advances extend the practicality of distributed private learning with PATE and CaPC to sensitive and complex domains."},"_bibtex":{"value":"@misc{\nmeintz2026distributed,\ntitle={Distributed {PATE} and Ca{PC} on a {DIET}: Private Knowledge Transfer without Public Data or Private Inference},\nauthor={Michel Meintz and Adam Dziedzic and Franziska Boenisch},\nyear={2026},\nurl={https://openreview.net/forum?id=Tqusxp1tXg}\n}"},"title":{"value":"Distributed PATE and CaPC on a DIET: Private Knowledge Transfer without Public Data or Private Inference"},"pdf":{"value":"/pdf/f0f27012983cab44e3b86eb29091f732f3fc9133.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"meintz|distributed_pate_and_capc_on_a_diet_private_knowledge_transfer_without_public_data_or_private_inference"},"authorids":{"value":["~Michel_Meintz1","~Adam_Dziedzic1","~Franziska_Boenisch2"]},"authors":{"value":["Michel Meintz","Adam Dziedzic","Franziska Boenisch"]}},"version":2},{"content":{"summary":{"value":"This paper proposes HGTFT, a transformer framework for forecasting in heterogeneous multi-domain physical systems. It fuses static and dynamic variables in a heterogeneous graph, integrates temporal attention with relation-specific aggregation. Evaluations show substantial improvements over baselines and strong zero/few-shot transfer across realistic multiphysics systems."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Table 2 shows large RCS improvements (e.g., 0.0158 → 0.0018 zero shot); could the authors how this metric generalizes to unseen physics domains?\n\n2. In Section 5.3, the weighting of the four loss terms (MSE, RCS, CRS, FDS) appears fixed; could the authors show how performance changes when these weights are learned or tuned, to verify robustness?\n\n3. The multi-instance normalization in Eq. (9) aggregates across percentile bounds. Could the authors explain how this compares against standard per-feature normalization across node types?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. the topic is timely and valuable.\n\n2. The paper clearly defines the new setting of heterogeneous graph forecasting in multi-domain physical systems, extending beyond conventional data.\n\n3. The proposed graph-temporal fusion with physics-aligned losses is technically well motivated and addresses both accuracy and physical consistency, which many data-driven models ignore\n\n4. Comprehensive experiments across synthetic and real-world datasets support the method’s claims"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The technical increment over existing graph transformers or physics-informed forecasting models is limited. This paper introduced: a heterogeneous graph encoder, a temporal transformer, and physics-inspired regularizers. However, each individual piece has been seen in earlier spatiotemporal or physics-informed learning work. The paper’s novelty lies more in integration and application to building energy systems than in a new architectural mechanism.\n\n2. The proposed physics-informed losses (RCS/CRS/FDS) are heuristic rather than derived from governing equations. The extent to which they enforce true physical constraints?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926586608,"tcdate":1761928969588,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16485/Reviewer_Bu3n"],"signatures":["ICLR.cc/2026/Conference/Submission16485/Reviewer_Bu3n"],"forum":"esM0zdV3NO","number":2,"license":"CC BY 4.0","cdate":1761928969588,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16485/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926586608,"domain":"ICLR.cc/2026/Conference","replyto":"esM0zdV3NO","id":"eEZKwK0JOL","forumContent":{"TLDR":{"value":"HGTFT, a pre-train and fine-tune framework that integrates heterogeneous spatiotemporal data with physics-informed constraints for accurate, physically consistent time series forecasting in Multi-Domain Physical Systems."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Heterogeneous Graph","Time Series Forecasting","Multiphysics","Physical Systems","Pre-training","Transformer"]},"supplementary_material":{"value":"/attachment/ca3774d7ae4892bac79634cf64764e0b3fc1e6dc.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Existing Transformer-based models effectively capture multivariate dependencies, while pre-trained large models achieve strong generalization but are often confined to single-object or single-physics settings. Spatial-temporal approaches leverage graph structures but fall short in modeling heterogeneous entities with diverse inter-variable interactions, and they often lack mechanisms to enforce physical consistency.\nTo address these challenges, we propose the Heterogeneous Graph Temporal Fusion Transformer (HGTFT), a pre-training and fine-tuning framework tailored for spatially and temporally structured physical environments. HGTFT tokenizes observation points and generates embeddings that capture both temporal patterns and spatial correlations, enabling the integration of heterogeneous static and dynamic information. \nWe further introduce optimized normalization and physics-informed loss functions that enhance predictive accuracy while improving physical plausibility. Applied to temperature, flow, and energy-related datasets in building environments, our approach demonstrates strong zero-shot generalization and achieves substantial accuracy gains through few-shot fine-tuning with domain-specific data."},"_bibtex":{"value":"@misc{\nsun2026heterogeneous,\ntitle={Heterogeneous Graph Temporal Fusion Transformer for Time Series Forecasting in Multi-Domain Physical Systems},\nauthor={Yifu Sun and Xin Li and Qi Shen and Qiang Dou and Yingqiu Qiu and Yuankai Zhao and Qianchuan Zhao},\nyear={2026},\nurl={https://openreview.net/forum?id=esM0zdV3NO}\n}"},"title":{"value":"Heterogeneous Graph Temporal Fusion Transformer for Time Series Forecasting in Multi-Domain Physical Systems"},"pdf":{"value":"/pdf/a883883d2e6fc778a14ab7211a7dda9de53e617e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sun|heterogeneous_graph_temporal_fusion_transformer_for_time_series_forecasting_in_multidomain_physical_systems"},"authorids":{"value":["~Yifu_Sun3","~Xin_Li94","~Qi_Shen4","~Qiang_Dou1","~Yingqiu_Qiu2","~Yuankai_Zhao1","~Qianchuan_Zhao1"]},"authors":{"value":["Yifu Sun","Xin Li","Qi Shen","Qiang Dou","Yingqiu Qiu","Yuankai Zhao","Qianchuan Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a new framework combining Large Language Models (LLMs) with physics engines and optimization algorithms for complex physical reasoning. This approach cab infer and simulate multi-object dynamics under physical interactions. The authors propose a new dataset, TraySim, where objects on a tray are impacted by an external force, simulating dynamic, inter-object interactions."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The authors may provide answers for my clarification questions about the method."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"This paper addresses an important question in AI by enabling large language models to perform complex physical reasoning, which is essential for real-world applications. The authors introduce a new framework, LLMPhy, that combines large language models with optimization algorithms and a physics simulator to estimate physical parameters and predict interactions. They also present a new dataset, TraySim, designed to benchmark models on multi-object physical reasoning tasks, making it a valuable resource for advancing AI’s capabilities in dynamic environments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"First, the paper lacks a discussion on its relationship with existing datasets like CoPhy https://arxiv.org/abs/1909.12000 and ComPhy https://comphyreasoning.github.io/, which also aim to infer physical properties of objects. These datasets offer benchmarks widely used in physical reasoning research, and a comparison of the proposed method on them would strengthen the validation of LLMPhy and clarify its distinct contributions.\n\nSecond, several established methods already tackle similar tasks, such as the approach presented in the ComPhy paper and methods like https://arxiv.org/abs/2012.08508 The absence of comparisons with these methods and relevant literature weakens the contextual positioning of LLMPhy within existing approaches.\n\nThird, the methodology lacks clarity, especially regarding critical details like how multiview images are utilized to reconstruct object shapes and the specific inputs provided to the LLM. For example, if the LLM samples optimization variables based on physics knowledge from its training data, is object category information also supplied to guide this process?\n\nLastly, the unclear explanation of how raw RGB video data is converted into simulation inputs makes it difficult to assess how this approach could apply in realistic scenarios beyond controlled simulations. This raises concerns about the generalizability of LLMPhy to real-world applications."}},"nonreaders":[],"tmdate":1731428017028,"tcdate":1730670431372,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12837/Reviewer_2zB8"],"signatures":["ICLR.cc/2025/Conference/Submission12837/Reviewer_2zB8"],"forum":"qGL6fE1lqd","number":2,"license":"CC BY 4.0","cdate":1730670431372,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12837/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428017028,"domain":"ICLR.cc/2025/Conference","replyto":"qGL6fE1lqd","id":"ZAHQmwfWCO","forumContent":{"TLDR":{"value":"Combining LLMs and physics engines for solving physical reasoning problems"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["large language models","physics simulators","physical reasoning"]},"supplementary_material":{"value":"/attachment/269b0fa498f32985b40c2d153c2c21b350427de2.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Physical reasoning is an important skill needed for robotic agents when operating in the real world. However, solving such reasoning problems often involves hypothesizing and reflecting over complex multi-body interactions under the effect of a multitude of physical forces and thus learning all such interactions poses a significant hurdle for state-of-the-art machine learning frameworks, including large language models (LLMs). To study this problem, we propose a new physical reasoning task and a dataset, dubbed TraySim. Our task involves predicting the dynamics of several objects on a tray that is given an external impact -- the domino effect of the ensued object interactions and their dynamics thus offering a challenging yet controlled setup, with the goal of reasoning being to infer the stability of the objects after the impact. To solve this complex physical reasoning task, we present LLMPhy, a zero-shot black-box optimization framework that leverages the physics knowledge and program synthesis abilities of LLMs, and synergizes these abilities with the world models built into modern physics engines. Specifically, LLMPhy uses an LLM to  generate code to iteratively estimate the physical hyperparameters of the system (friction, damping, layout, etc.) via an implicit analysis-by-synthesis approach using a (non-differentiable) simulator in the loop and uses the inferred parameters to imagine the dynamics of the scene towards solving the reasoning task.} To show the effectiveness of LLMPhy, we present experiments on our TraySim dataset to predict the steady-state poses of the objects. Our results show that the combination of the LLM and the physics engine leads to state-of-the-art zero-shot physical reasoning performance, while demonstrating superior convergence against standard black-box optimization methods and better estimation of the physical parameters. Further, we show that LLMPhy is capable of solving both continuous and discrete black-box optimization problems."},"_bibtex":{"value":"@misc{\ncherian2025llmphy,\ntitle={{LLMP}hy: Complex Physical Reasoning Using Large Language Models and World Models},\nauthor={Anoop Cherian and Radu Corcodel and Siddarth Jain and Diego Romeres},\nyear={2025},\nurl={https://openreview.net/forum?id=qGL6fE1lqd}\n}"},"title":{"value":"LLMPhy: Complex Physical Reasoning Using Large Language Models and World Models"},"pdf":{"value":"/pdf/1ac1d478c0fe0fe083ebaae5f4be44f529f96183.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"cherian|llmphy_complex_physical_reasoning_using_large_language_models_and_world_models"},"authorids":{"value":["~Anoop_Cherian1","~Radu_Corcodel1","~Siddarth_Jain2","~Diego_Romeres1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Anoop Cherian","Radu Corcodel","Siddarth Jain","Diego Romeres"]}},"version":2},{"content":{"summary":{"value":"This paper proposes MagniLearning, an adaptive weighting strategy for physics-informed neural PDE solvers. MagniLearning dynamically adjusts the weights assigned to spatial regions, temporal blocks, and knowledge-based constraints during training, aiming to prioritize regions and time intervals that most influence generalization."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"See Weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. **Comprehensive Framework**: The paper develops a unified strategy (MagniLearning) that adaptively reweights loss contributions from spatial regions, temporal segments, and knowledge/incomplete supervision, using structured leave-one-out risk estimators. This offers a principled, theoretically justified way to tackle long-standing challenges in neural PDE solvers—namely, poor generalization in underrepresented or hard-to-fit regions.\n2. **Theoretical Guarantees**: The work provides population risk bounds and convergence analyses (see Theorems 3.1 and 3.2, and proofs in Appendix A) that extend prior generalization results for PINNs and knowledge-augmented neural networks. These guarantees are stated with explicit dependencies on sample complexity, network width, and adaptive weighting parameters, which is valuable for researchers seeking theoretical insights into sample efficiency and over-parameterization regimes.\n3. **Detailed Mathematical Derivation**: All adaptive weighting formulas and risk definitions are spelled out with clear mathematical notation (e.g., Equations for LORO/LOTO/LOKO weighting, risk bounds, and Lemma A.6). The derivations show thoughtful treatment of weighting mechanisms and their statistical effects."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Experimental comparison and setup**: The experimental evaluation appears limited. It is essential to include comparisons with more strong and relevant baselines, such as PINNsformer [1] and RoPINN [2]. A more detailed experimental setup—following the standards set by these works—would help ensure a fair and convincing evaluation.\n2. **Broader Generalization and Robustness**: How would MagniLearning perform on more complex or higher-dimensional PDEs (e.g., Navier-Stokes) and under severe noise or imperfect physics constraints? Are there scalability limitations?\n3. **Ablation on Components**: Do LORO, LOKO, and LOTO contribute independently to overall performance, or is most of the gain explained by one of them? Please provide ablation or illustrative case studies per component.\n4. **Sensitivity Analysis**: How sensitive is MagniLearning’s performance to the key hyperparameters ($\\beta, \\lambda, \\kappa, \\epsilon$) and normalization choices? Would performance degrade for non-optimal choices, especially in more complex or high-dimensional domains?\n\n[1] PINNsFormer: A Transformer-Based Framework For Physics-Informed Neural Networks\n\n[2] RoPINN: Region Optimized Physics-Informed Neural Networks"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927929018,"tcdate":1761572471271,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18173/Reviewer_ECTK"],"signatures":["ICLR.cc/2026/Conference/Submission18173/Reviewer_ECTK"],"forum":"rE1ggCutRz","number":1,"license":"CC BY 4.0","cdate":1761572471271,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18173/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927929018,"domain":"ICLR.cc/2026/Conference","replyto":"rE1ggCutRz","id":"PBhA7F8kRn","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-Informed Neural Networks (PINNs)","Neural PDE solvers","Adaptive weighting","Generalization","Convergence acceleration"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Integrating domain knowledge into neural networks has advanced the development of Physics-Informed Neural Networks (PINNs), enabling solutions to partial differential equations across diverse applications. To enhance generalization in neural PDE solvers, we propose MagniLearning, an adaptive weighting strategy that dynamically adjusts the importance of spatial regions, knowledge components, and temporal segments during training. Our approach evaluates the impact of omitting each region or time block on model performance and assigns higher weights to the most influential data. This adaptive scheme accelerates convergence in neural PDE solvers by emphasizing the most informative regions and time segments, while enhancing robustness to noise and underrepresented physics. We formalize the method using an effective risk function that incorporates region- and time-dependent weights, and we provide theoretical guarantees for controlling the generalization error. Numerical experiments demonstrate that MagniLearning significantly improves both stability and accuracy."},"_bibtex":{"value":"@misc{\nfogh2026adaptive,\ntitle={Adaptive Spatial-Temporal Generalization for Physics-Informed Neural {PDE} Solvers},\nauthor={Fatemeh Fogh and Yufei Tang and Mohsen Ahmadi and Xingquan Zhu},\nyear={2026},\nurl={https://openreview.net/forum?id=rE1ggCutRz}\n}"},"title":{"value":"Adaptive Spatial-Temporal Generalization for Physics-Informed Neural PDE Solvers"},"pdf":{"value":"/pdf/9a603fabc6c0a0cbe2ae9c7bb60d57f96632be75.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"fogh|adaptive_spatialtemporal_generalization_for_physicsinformed_neural_pde_solvers"},"authorids":{"value":["~Fatemeh_Fogh1","~Yufei_Tang1","~Mohsen_Ahmadi1","~Xingquan_Zhu1"]},"authors":{"value":["Fatemeh Fogh","Yufei Tang","Mohsen Ahmadi","Xingquan Zhu"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2409.13886v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"jaiswal|learning_to_play_video_games_with_intuitive_physics_priors"},"authorids":{"value":["~Abhishek_Jaiswal1","~Nisheeth_Srivastava1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2409.13886"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2409-13886,\n  publtype={informal},\n  author={Abhishek Jaiswal and Nisheeth Srivastava},\n  title={Learning to Play Video Games with Intuitive Physics Priors},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2409.13886},\n  url={https://doi.org/10.48550/arXiv.2409.13886}\n}\n"},"abstract":{"value":"Video game playing is an extremely structured domain where algorithmic decision-making can be tested without adverse real-world consequences. While prevailing methods rely on image inputs to avoid the problem of hand-crafting state space representations, this approach systematically diverges from the way humans actually learn to play games. In this paper, we design object-based input representations that generalize well across a number of video games. Using these representations, we evaluate an agent's ability to learn games similar to an infant - with limited world experience, employing simple inductive biases derived from intuitive representations of physics from the real world. Using such biases, we construct an object category representation to be used by a Q-learning algorithm and assess how well it learns to play multiple games based on observed object affordances. Our results suggest that a human-like object interaction setup capably learns to play several video games, and demonstrates superior generalizability, particularly for unfamiliar objects. Further exploring such methods will allow machines to learn in a human-centric way, thus incorporating more human-like learning benefits."},"title":{"value":"Learning to Play Video Games with Intuitive Physics Priors"},"authors":{"value":["Abhishek Jaiswal","Nisheeth Srivastava"]}},"tmdate":1747395367330,"pdate":1704067200000,"tcdate":1747193162241,"writers":["~"],"signatures":["~Nisheeth_Srivastava1"],"forum":"V32D4AAUpS","license":"CC BY-SA 4.0","number":453144,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747395367330,"domain":"DBLP.org","id":"V32D4AAUpS","version":2},{"content":{"summary":{"value":"This paper proposes SONAR, a physics-constrained implicit neural representation (INR) for X-ray dark-field CT reconstruction. The method jointly models transmission, phase shift, and dark-field signals using a SIREN-based neural field, optimized directly in projection space under a Talbot–Lau interferometer forward model. A shot-wise training strategy with warm-start initialization enables stable phase retrieval across projections. Experiments on a human-scale DFCT prototype show improved noise suppression and reduced streak artifacts compared to classical sliding-window phase retrieval, leading to more coherent reconstructions. The approach is self-supervised and leverages physical constraints to reduce hallucination risk."},"review":{"value":"Overall, this paper presents a well-motivated and technically solid approach that integrates physics-based modeling with implicit neural representations for DFCT reconstruction. The formulation is elegant and aligns naturally with the underlying imaging physics.\n\nStrengths:\n- Strong integration of physics constraints via the Talbot–Lau forward model \n- Use of INRs is well-suited for continuous projection modeling \n- Self-supervised formulation avoids reliance on large training datasets \n- Shot-wise optimization with warm start is practical and effective \n- Clear qualitative improvements in artifact suppression and structural coherence \n\nWeaknesses:\n- The experimental evaluation is limited, primarily relying on a single dataset and qualitative comparisons. Additional quantitative analysis and broader validation would strengthen the conclusions \n\nOverall, the method is conceptually sound and promising, particularly for challenging DFCT reconstruction settings."},"strengths":{"value":"The paper introduces a well-designed physics-constrained INR framework tailored to dark-field CT, effectively combining implicit neural representations with a principled forward model. The self-supervised formulation is particularly appealing in this domain where training data is scarce. The shot-wise optimization strategy is practical and leverages inter-projection consistency. Qualitative results demonstrate clear improvements in noise reduction and artifact suppression, and the approach reduces the risk of hallucinations by enforcing physical consistency."},"weaknesses":{"value":"The experimental evaluation is somewhat limited, as it is primarily based on a single dataset and relies heavily on qualitative comparisons. While the improvements are visually convincing, additional quantitative metrics and validation across different acquisition settings or samples would further strengthen the evidence and generality of the approach."},"confidence":{"value":5},"rating":{"value":5},"justification_of_rating":{"value":"The paper presents a technically sound and well-motivated method with clear qualitative improvements and strong alignment with imaging physics. While the evaluation is somewhat limited, the contribution is meaningful and relevant, making it suitable for acceptance."},"title":{"value":"Physics-constrained neural representation for dark-field CT reconstruction"}},"parentInvitations":"MIDL.io/2026/Short_Papers/-/Official_Review","nonreaders":[],"tmdate":1778421299961,"tcdate":1777864443494,"writers":["MIDL.io/2026/Short_Papers","MIDL.io/2026/Short_Papers/Submission112/Reviewer_avuo"],"signatures":["MIDL.io/2026/Short_Papers/Submission112/Reviewer_avuo"],"forum":"a317VhxNvo","number":1,"license":"CC BY 4.0","cdate":1777864443494,"readers":["everyone"],"invitations":["MIDL.io/2026/Short_Papers/Submission112/-/Official_Review","MIDL.io/2026/Short_Papers/-/Edit"],"mdate":1778421299961,"domain":"MIDL.io/2026/Short_Papers","replyto":"a317VhxNvo","id":"pjp3nQoM0T","forumContent":{"TLDR":{"value":"SONAR is a neural representation adaptively trained under Talbot–Lau interferometer physics to suppress streak artifacts in X-ray dark-field CT."},"venue":{"value":"MIDL 2026 - Short Papers Poster"},"keywords":{"value":["Computed tomography","dark-field imaging","implicit neural representations"]},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Dark-field computed tomography (DFCT) enables functional lung imaging with small-angle X-ray scattering, but reconstructions are often degraded by streak artifacts. We propose SONAR (Shot-Optimized Neural Adaptive Representation), a projection-based implicit neural representation (INR) that jointly models transmission, phase shift, and dark-field signals across neighboring shots using a physics-based Talbot–Lau interferometer forward model. By leveraging adaptive per-projection optimization, SONAR effectively enables stabilized phase retrieval and suppresses streak artifacts for improved DFCT image quality on a grating-based human-scale prototype."},"_bibtex":{"value":"@inproceedings{\nfrey2026sonar,\ntitle={{SONAR}: A Physics-constrained Neural Representation for X-ray Dark-field {CT}},\nauthor={Daniel Frey and Theresa Hiu and Julian McGinnis and Tina Dorosti and Johannes Thalhammer and Sebastian Peterhansl and Zijin Huang and Franz Pfeiffer and Daniel Rueckert and Florian Schaff},\nbooktitle={Medical Imaging with Deep Learning - Short Papers},\nyear={2026},\nurl={https://openreview.net/forum?id=a317VhxNvo}\n}"},"title":{"value":"SONAR: A Physics-constrained Neural Representation for X-ray Dark-field CT"},"pdf":{"value":"/pdf/ce01d23b4b85ef46b789659edfb7d22b86cb1c17.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Short_Papers"},"paperhash":{"value":"frey|sonar_a_physicsconstrained_neural_representation_for_xray_darkfield_ct"},"authorids":{"value":["~Daniel_Frey1","~Theresa_Hiu1","~Julian_McGinnis1","~Tina_Dorosti1","~Johannes_Thalhammer1","sebastian.peterhansl@tum.de","~Zijin_Huang2","~Franz_Pfeiffer1","~Daniel_Rueckert2","~Florian_Schaff1"]},"registration":{"value":"Yes"},"authors":{"value":["Daniel Frey","Theresa Hiu","Julian McGinnis","Tina Dorosti","Johannes Thalhammer","Sebastian Peterhansl","Zijin Huang","Franz Pfeiffer","Daniel Rueckert","Florian Schaff"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"This paper introduces the Deep Complex Patio-Spectral Network (DCSNet), a fully complex-valued, token-based neural network developed for end-to-end foreground extraction and adaptable for image classification. Extensive experiments show that DCSNet surpasses current complex-valued approaches across various tasks with real and complex-valued data, achieving results on par with leading real-valued models."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please see the weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"This paper presents a novel complex-valued neural network, the Deep Complex Patio-Spectral Network (DCSNet), a fully complex-valued, token-based, end-to-end architecture designed for foreground extraction and adaptable to image classification tasks. Extensive experiments demonstrate that DCSNet outperforms existing complex-valued methods across diverse tasks involving both real and complex-valued data, achieving competitive results relative to state-of-the-art real-valued models. The paper is well-written, concise, and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Novelty: The novelty of this paper appears questionable. The authors claim that they propose the first token-based complex-valued network that maintains complex-valued information throughout. However, the fully Complex-valued Convolutional Network (FCCN) [1] also processes complex-valued data through the entire model. Could the authors clarify any differences between these two models? What advantages does DCSNet offer over FCCN?\n\n[1] Saurabh Yadav; Koteswar Rao Jerripothula, FCCNs: Fully Complex-valued Convolutional Networks using Complex-valued Color Model and Loss Function. ICCV 2023.\n\nComplex-valued Image Generation (R2C Method): The authors propose an R2C method for generating complex-valued images from real-valued images, presenting it as a novel complex-valued color transformation. However, methods such as quaternion representation, complex logarithmic transformation, and the Hilbert transform are well-established for generating complex-valued images. Could the authors specify the advantages of the R2C method over these alternatives?\n\nDCSNet Architecture:\n1. Fourier filters replace self-attention in DCSNet to retain information within the complex domain while preserving global context. How do Fourier filters achieve global information retention in this context, and why were they chosen?\n2. The paper briefly mentions dense tokens for image embedding but doesn’t fully explain their purpose. Are these tokens meant to capture pixel-level details (dense information) of the image? If so, why not use high-resolution Fourier filters as localized filters to capture this information directly?\n3. If large Fourier filters serve as global filters in the frequency domain while dense tokens capture image details, could a bank of wavelet filters offer a more effective solution? Wavelet filters with multiple resolutions could extract both global (large scale) and local (small scale) image features.\n\nResolution Tokens: The paper mentions multiple resolution tokens \\T_{i} for i \\in {0,1,2,3}. What was the reasoning for using exactly four resolutions? Would using more or fewer resolutions impact the results?\n\nTable 7 Clarification: In Table 7, there is a term \\calL_{isal}. Is this a typo? Please clarify its meaning if not."}},"nonreaders":[],"tmdate":1733128376563,"tcdate":1730292223401,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1276/Reviewer_QqAh"],"signatures":["ICLR.cc/2025/Conference/Submission1276/Reviewer_QqAh"],"forum":"9hmDl8fFDs","number":2,"license":"CC BY 4.0","cdate":1730292223401,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1276/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733128376563,"domain":"ICLR.cc/2025/Conference","replyto":"9hmDl8fFDs","id":"8GfV5XpQlN","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"A robust complex-valued approach in Spatio-spectral domain for multiple tasks on both real and complex data."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Deep Complex Newtworks","Complex-valued color transformation"]},"supplementary_material":{"value":"/attachment/d3c1bebc77a7f4edd23a402e6580595b34f472e4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex-valued neural networks have attracted growing attention for their ability to handle complex-valued data with enhanced representational capacity. However, their potential in computer vision remains relatively untapped. \nIn this paper, we introduce Deep Complex Spatio-Spectral Network (DCSNet), a fully complex-valued token-based, end-to-end neural network designed for binary segmentation tasks. Additionally, our DCSNet encoder can be used for image classification in the complex domain. We also propose an invertible real-to-complex (R2C) transform, which generates two complex-valued input channels, complex intensity and complex hue, while producing complex-valued images with distinct real and imaginary components.\nDCSNet operates in both spatial and spectral domains by leveraging complex-valued inputs and complex Fourier transform.\nAs a result, the complex-valued representation is maintained throughout DCSNet, and we avoid the information loss typically associated with Real$\\leftrightarrow$Complex transformations. Extensive experiments show that DCSNet surpasses existing complex-valued methods across various tasks on both real and complex-valued data and achieves competitive performance compared to existing real-valued methods, establishing a robust framework for handling both data types effectively."},"_bibtex":{"value":"@misc{\nyadav2025deep,\ntitle={Deep Complex Spatio-Spectral Networks with Complex Visual Inputs},\nauthor={Saurabh Yadav and Koteswar Rao Jerripothula},\nyear={2025},\nurl={https://openreview.net/forum?id=9hmDl8fFDs}\n}"},"title":{"value":"Deep Complex Spatio-Spectral Networks with Complex Visual Inputs"},"pdf":{"value":"/pdf/78685d9476ab42033341d680f3e952693b0cbc64.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yadav|deep_complex_spatiospectral_networks_with_complex_visual_inputs"},"authorids":{"value":["~Saurabh_Yadav2","~Koteswar_Rao_Jerripothula3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Saurabh Yadav","Koteswar Rao Jerripothula"]}},"version":2},{"content":{"summary":{"value":"Summary:  \nThis paper addresses the limitation of existing video generation models—their inability to adhere to physical laws despite producing visually realistic content—by proposingPhysMaster, a framework that enhances physics-awareness through learned physical representations and reinforcement learning (RL).  \n\nContributions:  \n（1）Novel Physical Representation Framework: PhysMaster introduces PhysEncoder to extract implicit physical knowledge from input images, bridging the gap between visual input and physical guidance—addressing the key challenge of translating visual cues to physics-aware generation.  \n（2）RL-Driven Optimization for Physical Representation: Unlike prior RL methods that fine-tune the entire video model, PhysMaster uses DPO to specifically optimize PhysEncoder. This avoids overfitting to specific physical scenarios and enhances generalizability across diverse phenomena.  \n（3）Efficient and Generalizable Pipeline: The three-stage pipeline (SFT + two DPO stages) balances effectiveness and efficiency. It works with small human-labeled datasets (500 preference pairs suffice) and generalizes from simple \"free-fall\" to 17 types of real-world physical events.  \n（4）SOTA Performance with Practical Value: PhysMaster outperforms existing models in both physical accuracy and inference speed (generating a 5-second video in 26 seconds), making it a plug-and-play solution for physics-aware video generation and potential applications like robotics training or scientific visualization."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Questions:    \n（1）Questions About PhysEncoder’s Learned Physical Knowledge  \nYou state PhysEncoder extracts physical priors from input images and PCA visualization shows grouping by external forces. But do you have evidence that PhysEncoder learns causal physical principles instead of just correlational visual patterns? For example, through counterfactual testing or feature attribution to prove it captures meaningful physical laws? This clarifies if PhysMaster advances physical understanding as claimed.  \n（2）Questions About Preference Data Scalability  \nYou note human annotation for DPO data is costly but dismiss AI evaluators without empirical validation. Have you tested physics-aware AI evaluators against human labels on a WISA-80K subset? What was the agreement rate? If not, why not use hybrid pipelines like AI pre-screening plus human validation to reduce effort? This addresses if the scalability bottleneck is unavoidable.  \n（3）Questions About Complex Physical Scenario Experiments  \nYour broader scenarios use WISA-80K but lack testing on dynamic multi-force or non-rigid interactions. Why were these excluded given they’re key for generalizability? For synthetic complex scenarios, have you compared generated motion to ground-truth physical quantities instead of just visual metrics? This checks if generalizability extends to challenging cases.  \n（4）Questions About Input Noise Robustness  \nAll experiments use high-quality inputs, but real-world inputs have noise. Have you tested how input corruption impacts PhysEncoder’s performance? If performance degrades, have you tried lightweight preprocessing like denoising? This assesses real-world utility as a plug-in solution."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Strengths:   \n（1）PhysMaster delivers originality through targeted innovations that address prior limitations: It framesphysical representation as a plug-in module (PhysEncoder) instead of embedding physics into the generation model, decoupling physical knowledge extraction from visual generation—solving the poor generalization of simulation-based methods and overfitting of end-to-end fine-tuning.  \n（2）The paper maintains high methodological and experimental quality: The three-stage pipeline logically addresses sequential challenges, using well-justified components.  \n（3）The paper is well-organized and accessible.  \n（4）PhysMaster has immediate and long-term value: Its plug-in design and efficiency make it practical for content creation and prototyping."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses:   \n（1）Over-Reliance on Human Annotation for Real-World Preference Data Limits Scalability. A core limitation of the work is its dependence on human annotators to construct preference datasets for DPO in real-world scenarios (e.g., WISA-80K). The paper acknowledges this is \"costly and time-consuming\", but it fails to address how this bottleneck restricts the framework’s scalability to more diverse physical phenomena (e.g., quantum mechanics, electromagnetism) or larger datasets. Worse, existing AI evaluators (e.g., VLMs) are dismissed as \"flawed in physics knowledge\" without any attempt to validate or improve them—a missed opportunity to reduce human effort.   \n（2）Lack of Analysis on PhysEncoder’s Physical Knowledge Granularity. The paper claims PhysEncoder extracts \"physical priors like relative positions and potential interactions\", but it provides no evidence ofwhat specific physical knowledgethe encoder actually learns.   \n（3）Experiments on Complex Physical Interactions Are Insufficient. The paper’s \"broader scenarios\" evaluation focuses on 17 physical events, but it neglectsdynamic, multi-object, or non-rigid interactions—the most challenging cases for physics-aware generation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916664683,"tcdate":1761894594036,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3318/Reviewer_aYxD"],"signatures":["ICLR.cc/2026/Conference/Submission3318/Reviewer_aYxD"],"forum":"CG2VPDZkwM","number":4,"license":"CC BY 4.0","cdate":1761894594036,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3318/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916664683,"domain":"ICLR.cc/2026/Conference","replyto":"CG2VPDZkwM","id":"U7F50CJve3","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["physics-aware video generation","representation learning","reinforcement learning"]},"supplementary_material":{"value":"/attachment/aa0b699477d9a37a094accff749767c9dba74a7d.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation models nowadays are capable of generating visually realistic videos, but often fail to adhere to physical laws, limiting their ability to generate physically plausible videos and serve as ''world models''. To address this issue, we propose PhysMaster, which captures physical knowledge as a representation for guiding video generation models to enhance their physics-awareness. Specifically, PhysMaster is based on the image-to-video task where the model is expected to predict physically plausible dynamics from the input image. Since the input image provides physical priors like relative positions and potential interactions of objects in the scenario, we devise PhysEncoder to encode physical information from it as an extra condition to inject physical knowledge into the video generation process. The lack of proper supervision on the model's physical performance beyond mere appearance motivates PhysEncoder to apply reinforcement learning with human feedback to physical representation learning, which leverages feedback from generation models to optimize physical representations with Direct Preference Optimization (DPO) in an end-to-end manner. PhysMaster provides a feasible solution for improving physics-awareness of PhysEncoder and thus of video generation, proving its ability on a simple proxy task and generalizability to wide-ranging physical scenarios. This implies that our PhysMaster, which unifies solutions for various physical processes via representation learning in the reinforcement learning paradigm, can act as a generic and plug-in solution for physics-aware video generation and broader applications."},"_bibtex":{"value":"@misc{\nji2025physmaster,\ntitle={PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning},\nauthor={Sihui Ji and Xi Chen and Xin Tao and Pengfei Wan and Hengshuang Zhao},\nyear={2025},\nurl={https://openreview.net/forum?id=CG2VPDZkwM}\n}"},"title":{"value":"PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning"},"pdf":{"value":"/pdf/c349ec6dad2b58c6a79e4570de82f1e8ca952db5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ji|physmaster_mastering_physical_representation_for_video_generation_via_reinforcement_learning"},"authorids":{"value":["~Sihui_Ji1","~Xi_Chen30","~Xin_Tao3","~Pengfei_Wan1","~Hengshuang_Zhao2"]},"authors":{"value":["Sihui Ji","Xi Chen","Xin Tao","Pengfei Wan","Hengshuang Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new knowledge-guided ML method, called PDED. PDED disentangles the data-based and physics-based modules, which are trained with labeled data and physics laws respectively. Ensembling learning and knowledge distillation are used to assemble those two representations. Superior performance has been shown to validate the effectiveness of the proposed method."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"- This paper proposes a new framework for combining physics laws and data, which aims to avoid the optimization issue in common physics-informed learning tasks. \n\n- This paper is well-written and well-structured."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- It seems that this method still has many hyperparameters to tune as shown in Eq. (14) and (15), though the goal of this paper is to mitigate the optimization issue in physics-informed learning. I am not sure of the magnitude differences between those loss terms in physics loss and data loss in  Eq. (14) and (15).\n\n- As shown in **Experiments settings** on Page 6, all of the hyper-parameters are set to 1. It seems they don’t have large variances in magnitudes. However, the test datasets focus on $x, t$ dimensions, which can be regarded as 1D PDEs in physics-informed learning. 1D PDEs are relatively easy to learn, and the optimization issues are not severe. I am wondering if the authors could test on more challenging 2D datasets in traffic modeling."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"- On Page 2, the authors claim that they discovered the optimization issue of physics-informed learning, where many existing research works have identified this problem [1-2]. I don’t think that is one of the contributions of this paper. \n\n- On Page 5, “As mentioned by Theorem 2.1, the effectiveness of ensemble teacher model is also determined by the the accuracy of individual student model.” There is a double “the” in this sentence. \n\n- This paper considers relative L2 errors, as commonly seen in physics-informed learning research. However, in the domain of traffic prediction, researchers also consider RMSE and MAPE to measure the errors from different perspectives. Is it possible to add more evaluation metrics since the traffic data has some rush hour phenomena, where extreme statistics are needed? \n\n- Regarding the noise setting in Section 3.4, is it common to set the noise ratio by using a small covariance of 0.01/0.04 in the Gaussian distribution? Can the authors add some references here?\n\n---\n**Refs:**\n\n[1] Krishnapriyan, Aditi, et al. \"Characterizing possible failure modes in physics-informed neural networks.\" Advances in Neural Information Processing Systems 34 (2021): 26548-26560.\n\n[2] Wang, Sifan, Yujun Teng, and Paris Perdikaris. \"Understanding and mitigating gradient flow pathologies in physics-informed neural networks.\" SIAM Journal on Scientific Computing 43.5 (2021): A3055-A3081."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636262535,"tcdate":1698115947197,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3151/Reviewer_m66B"],"signatures":["ICLR.cc/2024/Conference/Submission3151/Reviewer_m66B"],"forum":"GszBQ3ZTzk","number":1,"license":"CC BY 4.0","cdate":1698115947197,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3151/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636262535,"domain":"ICLR.cc/2024/Conference","replyto":"GszBQ3ZTzk","id":"vYvWGnGFRk","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed deep learning","Traffic state estimation","Knowledge distillation","Ensemble learning"]},"supplementary_material":{"value":"/attachment/a2e6b5b1436004e0907405012360cce043adef51.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Traditional physics-informed deep learning combines the data-driven methods with the model-based methods by incorporating physics loss as a constraint in total loss function in general, which aims to enforce the neural network to behave according to the physics property. However, this simple integration makes physical knowledge submerged in data information since data loss and physics loss could have large magnitude differences, conflicting directions of the gradients, and varying convergence rates so that the physics law may not work as expected and inhibits the model from working effectively furthermore, especially for traffic state estimation (TSE). To alleviate these issues, we propose a Physical knowledge combined Data information neural network with Ensemble Distillation framework (PDED) to first disentangle the data-driven model and physics-based model, and then reassemble them to take advantages of label information and physics property. Practically, we separately train data-driven model based on true labels and physics-based model according to physics laws. Then, we introduce the ensemble learning and knowledge distillation to assemble their representations of these two models for constructing a more competitive learnable online teacher model, which in turn distills knowledge to guide the update of them for learning richer knowledge to improve the performance of student models. Through extensive experiments on both synthetic dataset and real-world datasets, our model demonstrates better performance than the existing state-of-the-art methods."},"_bibtex":{"value":"@misc{\nfu2024pded,\ntitle={{PDED}: Revitalize physics laws submerged in data information for Traffic State Estimation},\nauthor={Yao Fu and Hong Zhao and Xiaoyu Cai and Ruiheng Yang and Weihao Jiang and Shiliang Pu},\nyear={2024},\nurl={https://openreview.net/forum?id=GszBQ3ZTzk}\n}"},"title":{"value":"PDED: Revitalize physics laws submerged in data information for Traffic State Estimation"},"pdf":{"value":"/pdf/dfcd70006ce024e9a689e1fc20180c31fce193e5.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"fu|pded_revitalize_physics_laws_submerged_in_data_information_for_traffic_state_estimation"},"authorids":{"value":["~Yao_Fu7","~Hong_Zhao5","~Xiaoyu_Cai1","~Ruiheng_Yang1","~Weihao_Jiang2","~Shiliang_Pu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yao Fu","Hong Zhao","Xiaoyu Cai","Ruiheng Yang","Weihao Jiang","Shiliang Pu"]}},"version":2},{"content":{"summary":{"value":"The paper proposed to enhance video generation model (I2V) 's obedience to physics rules by introducing an aditional module (named as PhysMaster) to learn physics information implicitly from I and then inject this condition by concatenaing the physical feature after the image latent from VAE. This module is based on DINO v2 and trained by DPO RL with preference data annotated by human. Training data used is from WISA-80K. For experiment, authors compared the RL-tuned I2V model (without any detail) to other general model as well as physics-enhanced model such as WISA/PhyT2V."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Only one addtional question: why RL here is necessary, comparing direct supervized finetuning on WISA-80K?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper is clearly written, and according to the results from Table 3, it achieved better performance than existing works w.r.t. physics metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed PhysMaster module (initialized with DINOv2) is designed to be responsible for physics information learning. However, it only took image as input, leaving text prompt untouched. With these settings, it's highly doubtful that the information encoded by PhysMaster module is really about physics. It's simple: how can we tell what physical rules would be involved with image given only?\n2. The proposed method is not training-free / plug-and-play for video generation model, this means we have to train the PhysMaster module as well as the I2V model (with lora in the paper). And since the whole model is finetuned with physics-centric data, it should definitely perform better on physics oriented metrics. This will make the contribution claims untenable."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916666575,"tcdate":1761735202485,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3318/Reviewer_sFF9"],"signatures":["ICLR.cc/2026/Conference/Submission3318/Reviewer_sFF9"],"forum":"CG2VPDZkwM","number":2,"license":"CC BY 4.0","cdate":1761735202485,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3318/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916666575,"domain":"ICLR.cc/2026/Conference","replyto":"CG2VPDZkwM","id":"Nzc1lZjYn8","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["physics-aware video generation","representation learning","reinforcement learning"]},"supplementary_material":{"value":"/attachment/aa0b699477d9a37a094accff749767c9dba74a7d.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation models nowadays are capable of generating visually realistic videos, but often fail to adhere to physical laws, limiting their ability to generate physically plausible videos and serve as ''world models''. To address this issue, we propose PhysMaster, which captures physical knowledge as a representation for guiding video generation models to enhance their physics-awareness. Specifically, PhysMaster is based on the image-to-video task where the model is expected to predict physically plausible dynamics from the input image. Since the input image provides physical priors like relative positions and potential interactions of objects in the scenario, we devise PhysEncoder to encode physical information from it as an extra condition to inject physical knowledge into the video generation process. The lack of proper supervision on the model's physical performance beyond mere appearance motivates PhysEncoder to apply reinforcement learning with human feedback to physical representation learning, which leverages feedback from generation models to optimize physical representations with Direct Preference Optimization (DPO) in an end-to-end manner. PhysMaster provides a feasible solution for improving physics-awareness of PhysEncoder and thus of video generation, proving its ability on a simple proxy task and generalizability to wide-ranging physical scenarios. This implies that our PhysMaster, which unifies solutions for various physical processes via representation learning in the reinforcement learning paradigm, can act as a generic and plug-in solution for physics-aware video generation and broader applications."},"_bibtex":{"value":"@misc{\nji2025physmaster,\ntitle={PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning},\nauthor={Sihui Ji and Xi Chen and Xin Tao and Pengfei Wan and Hengshuang Zhao},\nyear={2025},\nurl={https://openreview.net/forum?id=CG2VPDZkwM}\n}"},"title":{"value":"PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning"},"pdf":{"value":"/pdf/c349ec6dad2b58c6a79e4570de82f1e8ca952db5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ji|physmaster_mastering_physical_representation_for_video_generation_via_reinforcement_learning"},"authorids":{"value":["~Sihui_Ji1","~Xi_Chen30","~Xin_Tao3","~Pengfei_Wan1","~Hengshuang_Zhao2"]},"authors":{"value":["Sihui Ji","Xi Chen","Xin Tao","Pengfei Wan","Hengshuang Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LA-SR (Language Assistant for SR), a novel framework for unpaired real-world image super-resolution. The work is motivated by the difficulty of collecting paired real-world LR-HR data and the significant domain gap of synthetic degradations. The authors' key insight is that a single high-quality image naturally contains real-world LR patches (distant objects) and HR patches (near objects) due to varying depths of field, thus providing a source for unpaired, realistic training data."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. How does your data curation method handle scenes where depth does not strongly correlate with quality (e.g., motion blur on near objects, or sharp, distant textures)? Does this introduce significant label noise into your training set?\n2. How does the patch extraction algorithm perform on images with a shallow depth of field (e.g., portraits, documents)? Is there a risk of failing to extract a sufficient number or diversity of LR/HR patches from such images?\n3. Could you provide more details on the \"pre-defined quality texts\"? Was a richer vocabulary used beyond binary labels like {good}/{bad} to describe different quality dimensions (e.g., 'sharp', 'noisy', 'blurry')?\n4. The results in the appendix show that LA-SR does not always perform well on reference-based metrics like PSNR. Could you comment on this trade-off between perceptual quality and reconstruction fidelity within your framework?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.  The paper is exceptionally well-written and structured. Figure 1, in particular, provides a highly intuitive and concise overview of the entire framework, effectively communicating the core idea.\n2. The strategy of mining unpaired, real-world LR-HR patches from single images using depth is highly practical and scalable. It cleverly bypasses the need for specialized hardware or complex synthetic degradation pipelines, giving it significant real-world applicability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The framework's core assumption—that distance directly correlates with image quality ('far = low-res', 'near = high-res')—is an oversimplification of real-world imaging physics. This link can be broken by factors like motion blur on near objects, simple textures in the distance (e.g., a clear sky), or the camera's focal plane. Treating depth as the sole proxy for quality is a major limitation.\n2. The reliance on the depth-quality assumption may lead to sampling biases. For instance, the method might struggle to extract effective LR-HR pairs from images with a shallow depth of field (e.g., portraits, flat surfaces). This could result in a training dataset that lacks certain scene types, potentially harming the model's generalization.\n3. The use of binary, coarse-grained quality labels like {good} and {bad} is a simplistic form of supervision. This may not capture the full spectrum of real-world degradations (e.g., noise, compression artifacts) and could limit the model's ability to handle complex, unseen degradations beyond simple blurriness.\n4. While the paper focuses on perceptual quality, the reference-based metrics reported in the appendix (e.g., PSNR/SSIM) are not consistently strong compared to other methods. This suggests a significant trade-off where the model gains perceptual realism at the cost of reconstruction fidelity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919211462,"tcdate":1761886662749,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6995/Reviewer_PnCK"],"signatures":["ICLR.cc/2026/Conference/Submission6995/Reviewer_PnCK"],"forum":"1VckB5vsbW","number":3,"license":"CC BY 4.0","cdate":1761886662749,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6995/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919211462,"domain":"ICLR.cc/2026/Conference","replyto":"1VckB5vsbW","id":"bg3oWs0tVr","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Super-Resolution","Image restoration","Language","CLIP"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Single image super-resolution (SISR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs. \nTraining SR models typically requires paired HR–LR data, which is difficult to obtain in reality. As a result, most methods synthesize LR images by artificially degrading HR images with handcrafted kernels or camera ISP adjustments. However, these synthetic degradations fail to capture the complexity of real LR images, leading to poor generalization in practice. To address this, we observe that even within a single high-quality image, regions at different depths exhibit varying resolutions—where distant regions act as LR patches and closer ones as HR patches.  This allows the extraction of real, degradation-induced LR patches from real images. Since these LR patches lack paired HR counterparts, we propose LA-SR (Language Assistant for SR), a novel framework for unpaired SR. The key idea of LA-SR is to redefine unpaired SR in the \\emph{language space}, using vision-language models to bridge the LR–HR gap. LA-SR projects images into a semantic-rich space representing both content and quality, and applies two language-guided losses: linguistic-content loss to preserve semantic fidelity, and linguistic-quality loss to enhance perceptual realism. With this alignment, LA-SR effectively super-resolves real LR inputs, producing realistic outputs that overcome the limitations of synthetic-data-trained methods."},"_bibtex":{"value":"@misc{\npark2026languageassisted,\ntitle={Language-Assisted Super-Resolution from Real-World Low-Resolution Patches},\nauthor={JoonKyu Park and Kyoung Mu Lee},\nyear={2026},\nurl={https://openreview.net/forum?id=1VckB5vsbW}\n}"},"title":{"value":"Language-Assisted Super-Resolution from Real-World Low-Resolution Patches"},"pdf":{"value":"/pdf/900dbceebd78100db482e820d57e6dd810383bfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"park|languageassisted_superresolution_from_realworld_lowresolution_patches"},"authorids":{"value":["~JoonKyu_Park1","~Kyoung_Mu_Lee2"]},"authors":{"value":["JoonKyu Park","Kyoung Mu Lee"]}},"version":2},{"content":{"summary":{"value":"This paper proposes MIND to generate large-scale, high-quality synthetic math dialogues to pretrain LMs and enhance their mathematical reasoning abilities. MIND prompts a pretrained LLM (llama3-70B-instruct) to convert raw mathematical text into structured multi-turn conversations using diverse conversational styles. These synthetic dialogues break down complex problems step-by-step while injecting complementary explanations and reasoning. The authors generated 64B tokens of synthetic data using the 14B token OpenWebMath corpus. They conducted experiments with varying conversational styles and participants to assess impact on reasoning. Models pretrained on MIND-generated dialogues outperformed those trained on raw data in mathematical reasoning and general reasoning. In summary, MIND provides a way to upsample limited high-quality math data into structured dialogues that embed multi-hop reasoning."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- What is the impact of conversation length on model performance? Is there an optimal length?\n\n- How well does this approach generalize to other technical domains beyond mathematics?\n\n- The continued pretraining setup upsamples OpenWebMath but keeps CommonCrawl data constant. How much does the relative ratio of math vs general data impact results?\n\n- Will the data be released?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Originality: The paper presents an approach (MIND) to generate high-quality synthetic math dialogues that improve math reasoning in LMs. Compared to prior work on synthetic pretraining data that mostly rephrases raw text, MIND adds semantic variations and step-by-step reasoning that are crucial for complex math problem solving.\n\nQuality: The paper is thorough in evaluating MIND across multiple dimensions - testing different conversational styles, scaling behavior, and applicability to varying seed corpora. Ablations provide good insights, like the importance of knowledge gaps between dialogue participants. \n\nClarity: The paper is easy to read. Key aspects of the approach, experiments and results are clearly explained.\n\nSignificance: This work demonstrates the potential of structured conversational data to enhance reasoning in language models, especially when domain-specific high-quality data is limited."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Since the method uses the LLAMA3-70B-INSTRUCT model to generate conversations, it is unclear whether the improvements in downstream reasoning tasks come from the quality of the generated dialogues or are simply a result of model distillation from the powerful LLAMA3-70B model. The authors should investigate this and isolate the impact of the MIND-generated dialogues from the influence of the underlying LLM.\n\n- The experiments are based on a single in-house pretrained model checkpoint. It is possible that this model is not very well pretrained, making the improvements from synthetic dialogues appear more significant than they would be for a highly-optimized model. To demonstrate the generality of the MIND approach, the authors can experiment with multiple popular and high-quality pretrained models such as LLAMA-3, Mistral, GEMMA, etc. Consistent improvements across a range of strong baseline models would provide more convincing evidence for the effectiveness of the proposed method.\n\n- The paper focuses on math reasoning. It would be nice to see if MIND generalizes to other technical domains needing step-by-step reasoning, like physics, engineering, coding etc. Some preliminary experiments could help to see the broader applicability."}},"nonreaders":[],"tmdate":1731428819326,"tcdate":1730697220654,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11410/Reviewer_ikXv"],"signatures":["ICLR.cc/2025/Conference/Submission11410/Reviewer_ikXv"],"forum":"TuOTSAiHDn","number":2,"license":"CC BY 4.0","cdate":1730697220654,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11410/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428819326,"domain":"ICLR.cc/2025/Conference","replyto":"TuOTSAiHDn","id":"mU7sOaVdf9","forumContent":{"TLDR":{"value":"We propose a novel large-scale and diverse Math Informed syNthetic Dialogue (MIND) generation method that improves the mathematical reasoning ability of LLMs during pretraining."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["pretraining","mathematical reasoning","synthetic dialogue","LLM","reasoning"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The utility of synthetic data to enhance pretraining data quality and hence to improve downstream task accuracy has been widely explored in recent large language models (LLMs). Yet, these approaches fall inadequate in complex, multi-hop and mathematical reasoning tasks as the synthetic data typically fails to add complementary knowledge to the existing raw corpus. In this work, we propose a novel large-scale and diverse Math Informed syNthetic Dialogue (MIND) generation method that improves the mathematical reasoning ability of LLMs. Specifically, using MIND, we generate synthetic conversations based on OpenWebMath (OWM), resulting in a new math corpus, MIND-OWM. Our experiments with different conversational settings reveal that incorporating knowledge gaps between dialog participants is essential for generating high-quality math data. We further identify an effective way to format and integrate synthetic and raw data during pretraining to maximize the gain in mathematical reasoning, emphasizing the need to restructure raw data rather than use it as-is. Compared to pretraining just on raw data, a model pretrained on MIND-OWM shows significant boost in mathematical reasoning (GSM8K: +13.42%, MATH: +2.30%), including superior performance in specialized knowledge (MMLU: +4.55%, MMLU-STEM: +4.28%) and general\npurpose reasoning tasks (GENERAL REASONING: +2.51%)."},"_bibtex":{"value":"@inproceedings{\nakter2025mind,\ntitle={{MIND}: Math Informed syNthetic Dialogues for Pretraining {LLM}s},\nauthor={Syeda Nahida Akter and Shrimai Prabhumoye and John Kamalu and Sanjeev Satheesh and Eric Nyberg and Mostofa Patwary and Mohammad Shoeybi and Bryan Catanzaro},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=TuOTSAiHDn}\n}"},"title":{"value":"MIND: Math Informed syNthetic Dialogues for Pretraining LLMs"},"pdf":{"value":"/pdf/5313c64ed4eed9b848e009dd9d70799be396e7f8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"akter|mind_math_informed_synthetic_dialogues_for_pretraining_llms"},"authorids":{"value":["~Syeda_Nahida_Akter1","~Shrimai_Prabhumoye1","~John_Kamalu1","~Sanjeev_Satheesh3","~Eric_Nyberg1","~Mostofa_Patwary1","~Mohammad_Shoeybi1","~Bryan_Catanzaro1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Syeda Nahida Akter","Shrimai Prabhumoye","John Kamalu","Sanjeev Satheesh","Eric Nyberg","Mostofa Patwary","Mohammad Shoeybi","Bryan Catanzaro"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a way to make neural operators more universal by pretraining across multiple physics, and then fine tune on a new PDE. The suggestion is to keep a shared stack of kernel integral operator blocks and, at the fine tune time, freeze this backbone, while training only small input/output adapters. This cuts adaptation cost and tests, whether the backbone captures reusable physics. The authors build a multiphysics data pipeline that standardizes heterogeneous PDE datasets, enabling joint pretraining over advection, Burgers, Navier–Stokes, etc. The training scheme is: pretrain the full model on mixed physics, then fine tune only adapters for the target task. Fig. 2 illustrates the setup. The authors find that: (i) Pretraining a backbone and fine tuning only the small adapters improves error and shortens epochs versus training the whole model from scratch (table 1), (ii) When the fine tune task augments the equation with an extra term, the pretrained backbone plus adapters still wins (table 2), (iii) Pretrain on one set of PDEs , e.g. advection and Burgers, then fine tune on another, e.g. reaction–diffusion, yields lower error and faster epochs (table 3 and figure 3)."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Suggestions: (i) Report end to end efficiency and not just per epoch time. The tables list Avg. epoch time(s), but one cannot judge the total \nwall clock once pretraining is amortized. Add: total pretrain hours, fine tune hours, and amortization break even curves, i.e  accuracy vs. total time/energy, and include FLOPs and trained parameter counts per setting, (ii) Ablate the fine tuning strategy by adding controls: (1) unfreeze the backbone (full fine tune), (2) train adapters atop a randomly initialized frozen backbone, (3) partial unfreezing, e.g. top k operator blocks.This isolates how much gain comes from pretraining vs. fewer trainable parameters, (iii) Go beyond shared regular grids and same dimensionality by adding tests on irregular meshes, varying domains/BCs, and 3D cases to demonstrate geometry robust transfer."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"(i) Clear recipe: freeze backbone and train adapters. The pretrain to finetune scheme is explicit: fix the common integral operator stack and update only small input/output adapters, which highlights what transfers and cuts training cost. (ii) Well-designed evaluation: the authors test three distinct scenarios, new parameters, added physics/inputs, and cross PDE transfer, without changing the modeling recipe, (iii) Consistent empirical gains:  across tables, pretraining plus adapters improves accuracy and reduces average epoch time versus scratch, sometimes by large margins."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(i) Scope limited to same dimensionality and curated grids: the method is explicitly evaluated when pretrain and fine tune tasks have the same problem dimensionality, and the pipeline resamples all data to a shared fixed grid. That leaves open transfer across 2D to 3D, irregular meshes, or complex geometries/boundaries, (ii) Speedups reported per epoch, not end to end:  tables report Avg. epoch time(s) but not the total wall clock including pretraining. This makes it hard to judge the true efficiency once pretraining cost is amortized, (iii) Missing controls/ablations on fine tuning strategy: fine-tuning always freezes the operator backbone and there is no baseline that unfreezes the backbone after pretraining or trains adapters with a randomly initialized frozen backbone to isolate how much of the gain comes from pretraining vs. just fewer trainable parameters, (iv) Evaluation metrics are generic and not physics aware: the core metric used is range normalized MAE (NMAE) and MSE, and  there is no explicit assessment of conservation laws, stability under long rollouts, or other physics diagnostics, so it is unclear how models behave far off the train distribution."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919736888,"tcdate":1761578620845,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7672/Reviewer_xuEJ"],"signatures":["ICLR.cc/2026/Conference/Submission7672/Reviewer_xuEJ"],"forum":"utk1b1OSXN","number":1,"license":"CC BY 4.0","cdate":1761578620845,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7672/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919736888,"domain":"ICLR.cc/2026/Conference","replyto":"utk1b1OSXN","id":"XkYt0fZCDn","forumContent":{"TLDR":{"value":"There could be a physical pretrain of neural operator that could be used to solve various problems"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Neural Operator","Pretrain","Multiphysics","Adapters","Foundational models"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Although neural operators found common use in contemporary data-driven physical systems simulation, their training procedure remains computationally expensive and time-consuming. Some advances have been made with the study of downstream problems, where the model is trained on a simpler problem and later fine-tuned on a more challenging one to achieve better quality and lower over-all time costs. In this research we examine capabilities of transformer-based neural operator architectures, which were previously used only for particular problems solutions, in more generalized transfer learning. We evaluate performance of the transformer- and state space model-based neural operators on wide range of downstream PDE simulation problems, including extension of models to the out-of pretraining sample parameter values, addition of new variables into the dynamics, and transfer of operator, trained on datasets, composed of solutions of several differential equations. The results indicate the ability of neural operators of advanced architectures to be used to transfer knowledge between problems, involving partial differential equations."},"_bibtex":{"value":"@misc{\nmasliaev2025towards,\ntitle={Towards Universal Neural Operators through Multiphysics Pretraining},\nauthor={Mikhail Masliaev and Dmitry A. Gusarov and Ilya Markov and Alexander Hvatov},\nyear={2025},\nurl={https://openreview.net/forum?id=utk1b1OSXN}\n}"},"title":{"value":"Towards Universal Neural Operators through Multiphysics Pretraining"},"pdf":{"value":"/pdf/4d5adfb2c063c394f607c61676d96335d29cda0c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"masliaev|towards_universal_neural_operators_through_multiphysics_pretraining"},"authorids":{"value":["~Mikhail_Masliaev1","~Dmitry_A._Gusarov1","~Ilya_Markov2","~Alexander_Hvatov1"]},"authors":{"value":["Mikhail Masliaev","Dmitry A. Gusarov","Ilya Markov","Alexander Hvatov"]}},"version":2},{"content":{"summary":{"value":"This paper introduces SteelNet, a multimodal representation-learning framework for steel-mill monitoring that aims to stay robust when sensors fail and to surface actionable parameter-level attributions for defects. It pairs a new (partly synthetic) multimodal dataset with a model that uses a parameter encoder, self-attention, physics-informed constraints, modality dropout, and multi-task heads for defect class and “intensity” prediction, plus an attribution head to highlight root-cause variables."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"Q1. Figures / readability:\nCould you regenerate all figures with vector graphics, larger fonts, clear axis units, consistent color palettes, and readable legends—especially Figs. 5–9? (Also, please fix the caption typo “it’s mask” → “its mask” in Fig. 6.) \n\nQ2. Method clarity: fusion + physics constraint:\nHow exactly are image features fused with process-parameter embeddings, and what is the closed-form definition (with coefficients) of the physics-constraint term used in training? The ablation refers to “w/o Physics Constraints” but the constraint itself isn’t fully specified.\n\nQ3. Data realism and leakage control:\nGiven the reliance on synthetic parameters and static samples, can you provide any validation on real sensor time-series—or, at minimum, describe safeguards against leakage between the intensity model and the parameter generation pipeline? \n\nQ4. Baselines and statistical rigor:\nBeyond ablations, will you add head-to-head baselines (e.g., strong tabular models and established missing-modality methods), report mean±std over multiple seeds, and clarify the sensor-dropout protocol (failure sampling pattern, seeds, confidence intervals)?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"1. Integration of industrial requirements with AI design:\nThe paper effectively connects real industrial problems with AI methods. Instead of building a generic model, it designs the system to handle real issues in steel manufacturing, such as missing sensor data, few labeled samples, and the need for clear explanations of defect causes. This focus on practical challenges makes the work more useful and relevant for industry applications, not just academic research.\n\n2. Systematic ablation study:\nThe authors conduct an ablation analysis that clearly disentangles the effect of each model component — self-attention, modality-dropout, attribution loss, and physics constraints. \n\n3. Acknowledgment and discussion of limitations:\nThe paper clearly states its main limitations — including the use of synthetic data, simplified treatment of time, and absence of real factory deployment. This honest discussion shows that the authors understand the boundaries of their work and helps readers see what still needs to be tested in real conditions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Figures: low readability and polish (plus a minor mistake): \nSeveral plots are hard to read (small fonts, low contrast, dense legends) and look inconsistent across sections (e.g., architecture/importance figs; robustness and appendix plots). Clearer typography, consistent palettes, and vector graphics would help. Also fix the caption typo “it’s mask” → “its mask” in Fig. 6.\n\n2. System/algorithm description feels sketch-level rather than fully academic:\nWhile Fig. 2 and “Algorithm 1” give a high-level view, key pieces remain under-specified—especially the exact fusion with the image modality and the formal definition of the physics constraint (it appears as a black-box term PhysicsConstraint(z)). Please spell out the fusion pathway, units/normalization, and the constraint’s closed form (including any coefficients) so others can replicate precisely.\n\n3. Heavy reliance on synthetic parameters and static samples limits external validity:\nThe paper itself notes that process parameters are synthetically generated, time dynamics are omitted, and there is no validation on real sensor data. This makes generalization to real plants uncertain. A small real-sensor/time-series study (even a pilot) would strengthen the claims.\n\n4. Claims not fully supported as contributions:\nThe paper states it targets equipment availability and improves operational decision-making, but the experiments do not measure uptime/availability or control benefits; as written, that should not be counted as a demonstrated contribution. Also, the defect-intensity scoring (ResNet-34 trained on masks) is a useful dataset step, but not methodically novel—please frame it as dataset preparation rather than a core algorithmic contribution."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941650127,"tcdate":1761986604915,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21239/Reviewer_hZkH"],"signatures":["ICLR.cc/2026/Conference/Submission21239/Reviewer_hZkH"],"forum":"J9VRPrhwjM","number":4,"license":"CC BY 4.0","cdate":1761986604915,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21239/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941650127,"domain":"ICLR.cc/2026/Conference","replyto":"J9VRPrhwjM","id":"twa2TWlkl4","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multimodal Representation Learning","Cross-Modal Alignment","Steel Rolling Mills","Deep Learning","Casual Inference"]},"supplementary_material":{"value":"/attachment/9402892e9da3bc2b5e907cbf9c3819490d2d3e7a.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Steel rolling mills must continuously monitor various sensors like vibration probes, thermocouples, and flow meters to ensure safe operations and maintain product quality. However, the harsh industrial condition sometimes lead to sensor failures, unclear signals, and incomplete data. Traditional monitoring systems often struggle in these environments, makes it challenging to detect early signs of problems and predict failures. To this end, we propose SteelNet: A Multimodal Representation Learning Framework designed for robust learning from various industrial sensor data. SteelNet incorporates the cross-modal alignment and modality dropout strategies that enable consistent representation learning even when modalities are partially missing. The core problem being solved is improving equipment availability and optimizing process parameters. This framework allows for the early detection of critical events by combining information from multiple sensors, and effectively handling missing or data, which are common in industrial environments. By improving the reliability of anomaly detection and predictive insights, SteelNet not only strengthens fault tolerance but it also support better decision making in the process optimization. Although developed for steel rolling mills, it's applicability extends to real-world scenarios and other industry setups."},"_bibtex":{"value":"@misc{\nakash2026steelnet,\ntitle={SteelNet: Multimodal Representation Learning for Industrial Process Optimization},\nauthor={Hindesh Akash and Aman De Sarker and Gourav Sarkar and Darshan Sharma},\nyear={2026},\nurl={https://openreview.net/forum?id=J9VRPrhwjM}\n}"},"title":{"value":"SteelNet: Multimodal Representation Learning for Industrial Process Optimization"},"pdf":{"value":"/pdf/941f48bca38a5f22438f3984454eb05d2c8576b9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"akash|steelnet_multimodal_representation_learning_for_industrial_process_optimization"},"authorids":{"value":["~Hindesh_Akash1","~Aman_De_Sarker1","~Gourav_Sarkar2","~Darshan_Sharma1"]},"authors":{"value":["Hindesh Akash","Aman De Sarker","Gourav Sarkar","Darshan Sharma"]}},"version":2},{"content":{"summary":{"value":"The paper presents the Mesh-based Multi-Segment Graph Network (MMSGN), a model designed to simulate dynamic systems by combining local and long-range information exchanges within a physically aligned hierarchical mesh structure. By segmenting meshes based on physics-informed features, MMSGN efficiently captures complex dynamics and outperforms baseline methods in both prediction accuracy and mesh quality across several datasets. Through empirical analysis, the model demonstrates robust generalization capabilities and scalability, making it suitable for large-scale, complex simulations​."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. **Mesh Continuity**: While the authors use *Mesh Continuity* as the mesh quality benchmark, this metric may not be standard in all visual computing applications or mechanical engineering without proper citations. I would recommend using aspect ratio, or \"as-regular-as-possible\" metric to measure the uniformness for every element (triangle/tetrahedra) on the mesh, as it is a more common mesh quality metric in FEM literature. Furthermore, the sole use of Hausdorff distance for geometric fidelity could be limiting; adding Chamfer distance would provide a more balanced measure of mesh accuracy.\n2. **Underwhelming PE Impact**: Despite the geometric reasoning behind PE(e.g. understanding relative location between within segments), the benefit is marginal (around a 10% reduction in RMSE on the already low error). Could the authors elaborate on the observed impact of PE?\n3. **Selection of Optimal Number of Clusters**: The paper includes an empirical analysis in Appendix Table 4, showing that the number of clusters impacts accuracy (with only two data points along the # of clusters dimension shown in the table). Given the algorithm's similarity to K-means, selecting an optimal number of clusters and their initialization will likely be crucial for accuracy and convergence speed. Could the authors expand on their criteria for selecting the optimal number of clusters, and explain how initialization sensitivity is managed in the clustering process?\n\n\n**Misc**:\n- **Missing Error Metrics**: Figure 3 lacks error colormap ranges, making it difficult to assess error variability.\n- **Physics-Inspired Segmentation**: Consider exploring modal analysis in engineering and fracture modes in graphics for relevant physics-driven segmentation methods.\n\n\n**Rebuttal**:\n\nFirst off, I deeply apologize for missing the deadline to reply as I’ve been travelling. I have reviewed the revised manuscript and sincerely appreciate the additional statistics provided by the authors. I also agree with Reviewer YvcY that the statistical analysis convincingly demonstrates that segmentation is indeed highly useful. Kudos to the authors for addressing the issues raised in such a short amount of time.\n\nI believe my earlier discussion with the AC is relevant here, so I will share part of it for context.\n\n---\n\nHere, I want to raise my concerns. My expertise lies primarily in FEM for mesh-based methods and neural field/PINN approaches for physics + ML, which informed my review from the perspective of norms in mechanical engineering and computer graphics.\n\nWhile the authors provided additional statistics in their rebuttal, which I found convincing, they failed to adequately address or reinforce the concerns I raised:\n\n1. **Literature review**: The paper does not sufficiently engage with physics-inspired methods, particularly modal analysis, which has a long history in computational physics and mechanical engineering and has seen renewed interest in graphics (e.g., [[Benchekroun et al. 2023]](https://www.dgp.toronto.edu/projects/fast_complementary_dynamics_site/) and [[Sellan et al. 2022]](https://www.dgp.toronto.edu/projects/breaking-good/), and I am not suggesting to cite those papers, just want to highlight what **\"physics-informed features from geometry ONLY\"** *should* be claimed). Using a graph-based clustering algorithm for a physics problem with no physical grounding feels underwhelming and disconnected from established practices.\n\n2. **Clustering methodology**: The rebuttal essentially agrees that there is no physically motivated or elegant method to determine the optimal number of clusters for initial clustering [(reference)](https://openreview.net/forum?id=pzasy8KRWK&noteId=FGgmSZdP0U). This makes it hard for me to imagine this method being applied to any serious physics/engineering problems.\n\nAlthough the paper incorporates some good physics intuition, I struggle to view it as motivated by a desire to develop a true physics solver. Instead, it appears to explore clustering on graphs—a valid direction for follow-up work, but not convincingly framed here as a physics-based contribution.\n\n---\nI also found the exchange between the authors and Reviewer YvcY regarding whether the segmentation is \"physically-informed\" highly interesting [(the exchange can be found here).](https://openreview.net/forum?id=pzasy8KRWK&noteId=UlcjK7vO4H). However, while the geometric segmentation may have been effective for a *specific* simulation setup, this is a *correlation*, not *causation*. Furthermore, I question whether the method would generalize to the same geometric setup under different boundary conditions or initial conditions (again, sorry for the late reply, it is no longer possible for authors to provide more updates on this).\n\nTo avoid confusion, I strongly recommend refraining from describing the segmentation approach as \"physical.\" While this may seem like a matter of semantics, precise language is essential in technical writing to ensure clarity and avoid misrepresentation.\n\nGiven these considerations, I will not change my rating for the paper. That said, I greatly appreciate the professional conduct of the authors and the significant improvements made to the paper during the rebuttal process."},"rating":{"value":5},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Clarity of High-Level Concept**: The core idea—that deformations primarily remain local to the area of contact and propagate slowly across the entire structure so some clustered feature would be helpful—is intuitive and easy to follow. This perspective makes sense in scenarios where local interactions don’t significantly affect distant areas. I found the paper easy to follow. In addition, I think the paper provides good details on the setup of the experiments (including baselines).\n\n2. **Physics-guided Segmentation**: I like the idea of using segmentation/clustering to \"reduce feature space\" locally, it draws some interesting connections to modern CNN and ViT structure in CV, and reduced-order modelling in engineering. In addition, Positional Encoding (PE) and Segment Encoding (SE) definitions are grounded in physical intuition by constructions.\n\n3. **Performance**: Surprisingly fast compared to well-established baselines like MGN. However, I would like to learn some intuition behind this -- based on the neural network structure, I can't really tell where the performance boost comes from. Is it because the number of message-passing steps between nodes was reduced?"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Questionable Motivation in Localization**: The emphasis on local deformation may not always hold, especially in fields like computational physics or mechanical engineering, where elasticity often results in global, fast propagating deformation—particularly with low-stiffness materials. This raises concerns that the paper's foundation may not fully align with real-world mechanics.\n\n2. **Limitations in Segmentation Approach**: The paper’s reliance on *purely geometric* (and arguably *topological*) clustering for segmentation may not capture clusters that reflect true mechanical behavior. For instance, with a bird flapping its wings, both wings would oscillate at similar frequencies, we could naturally group them as a single cluster in modal analysis. However, the geometric-only approach used here might yield clusters highly dependent on the initial setup, potentially missing important physical interdependencies. A physics-inspired method, like modal analysis, could provide clusters with greater physical relevance and reduce sensitivity to initialization. This is my major concern with this submission, the current segmentation approach is too *geometric* and already exhibits a **strong** prior on how every deformation must be highly local.\n\n3. **Mesh Resolution Sensitivity and Segment Size Issues**: As number of segments grows, the number of finite elements within them decreases, leading to potential accuracy and performance downgrades. Appendix Table 4 reflects this, with errors increasing as segment size grows in some examples, but the paper lacks clarity on how it determines an optimal cluster count."}},"nonreaders":[],"tmdate":1733277958656,"tcdate":1730689229302,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3197/Reviewer_5KUk"],"signatures":["ICLR.cc/2025/Conference/Submission3197/Reviewer_5KUk"],"forum":"pzasy8KRWK","number":3,"license":"CC BY 4.0","cdate":1730689229302,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3197/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733277958656,"domain":"ICLR.cc/2025/Conference","replyto":"pzasy8KRWK","id":"q25WYvDk4P","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["dynamic system; physics simulation; solid mechanics; graph-based simulation;"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Dynamic systems evolve through complex interactions, where local events influence global behaviors, reflecting the interconnected nature of real-world phenomena. Simulating such systems demands models that effectively capture both local and long-range dynamics, while maintaining a balance between accuracy and computational efficiency. However, existing mesh-based Graph Neural Network (GNN) methods often struggle to achieve both high accuracy and efficiency, especially when dealing with large datasets, complex mesh structures, and extensive long-range effects. Inspired by how real-world dynamic systems operate, we present the Mesh-based Multi-Segment Graph Network (MMSGN), a novel framework designed to address these challenges by leveraging a physically aligned hierarchical information exchange mechanism. MMSGN combines micro-level local interactions with macro-level global exchanges, aligning the hierarchical mesh structure with the system’s physical properties to seamlessly capture both local and global dynamics. This approach enables precise modeling of complex behaviors while maintaining computational efficiency. We validate our model on multiple dynamic system datasets and compare it with several state-of-the-art methods. Our results demonstrate that MMSGN delivers superior accuracy and mesh quality, excels in managing long-range effects, and maintains high computational efficiency. Furthermore, MMSGN exhibits strong generalization capabilities, scaling effectively to larger physical domains. These advantages make MMSGN well-suited for simulating complex, large-scale dynamic systems across a variety of scenarios. Codes and data will be made publicly accessible upon acceptance."},"_bibtex":{"value":"@misc{\nlei2025physically,\ntitle={Physically Aligned Hierarchical Mesh-based Network for Dynamic System Simulation},\nauthor={Bo Lei and Victor M Castillo and Yeping Hu},\nyear={2025},\nurl={https://openreview.net/forum?id=pzasy8KRWK}\n}"},"title":{"value":"Physically Aligned Hierarchical Mesh-based Network for Dynamic System Simulation"},"pdf":{"value":"/pdf/779fe1e9cc7ae9291fd338fee97f9e4a14c314b4.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"lei|physically_aligned_hierarchical_meshbased_network_for_dynamic_system_simulation"},"authorids":{"value":["~Bo_Lei2","~Victor_M_Castillo1","~Yeping_Hu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Bo Lei","Victor M Castillo","Yeping Hu"]}},"version":2},{"content":{"summary":{"value":"This paper investigates whether using synthetic data in machine learning genuinely safeguards privacy, as often claimed. The evaluations is done on 4 different training paradigms-coreset selection, data distillation, data-free knowledge distillation and synthetic data generated from diffusion models. To test the privacy claims of these methods, the study uses membership inference attacks (MIAs), focusing on worst-case scenarios to rigorously assess privacy leakage. The paper also compares these methods to Differential Privacy Stochastic Gradient Descent (DPSGD), a technique known for providing formal privacy guarantees, and finds that DPSGD consistently outperforms synthetic data-based approaches in terms of the privacy-utility-efficiency balance. The findings reveal that none of the synthetic data techniques match DPSGD in safeguarding privacy effectively. Notably, the study also discovers that visual dissimilarity between synthetic and private data does not necessarily imply privacy protection, as even visually distinct synthetic data can leak information when model logits are similar. This highlights a risk that methods relying solely on visual or distributional differences may offer a false sense of privacy."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"Why do you consider coreset selection as synthetic data?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This broad approach offers a thorough understanding of various methodologies in synthetic data utilization and their impact on privacy.\nThe study juxtaposes synthetic data-based techniques with Differential Privacy-SGD (DPSGD) as a baseline, which helps readers contextualize the efficacy of synthetic data methods in privacy preservation compared to a gold-standard approach like DPSGD.\nThe study identifies instances where synthetic data, despite visual dissimilarity from private data, can still leak privacy information through logit similarities. This nuanced finding enhances the paper's depth by showing that visual similarity alone isn’t sufficient to evaluate privacy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The experiments focus on CIFAR-10 and specific models, such as ResNet-18, which may limit the generalizability of findings. The paper’s findings could vary across more complex datasets or architectures, and broader experiments could better represent the implications for privacy in diverse real-world scenarios​.\n\nTechniques like DPSGD are noted for efficiency, yet they are resource-intensive. The paper briefly mentions but does not deeply engage with the practical constraints of computational cost and scalability, which are critical factors for real-world implementation of privacy-preserving methods​.\n\nWhile the empirical evaluation is thorough, the paper lacks an in-depth theoretical framework to explain why certain synthetic data techniques lead to privacy leakage. A theoretical grounding could bolster the empirical findings and offer predictive insights for synthetic data privacy."}},"nonreaders":[],"tmdate":1733326265050,"tcdate":1730045526606,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission722/Reviewer_ybiF"],"signatures":["ICLR.cc/2025/Conference/Submission722/Reviewer_ybiF"],"forum":"C8niXBHjfO","number":2,"license":"CC BY 4.0","cdate":1730045526606,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission722/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733326265050,"domain":"ICLR.cc/2025/Conference","replyto":"C8niXBHjfO","id":"TY9OfcesFy","forumContent":{"TLDR":{"value":"We rigorously evaluate privacy leakage across various methods based on training with synthetic data, and none of these methods achieve a better trade-off than the differential privacy baselines."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["ML privacy","membership inference"]},"supplementary_material":{"value":"/attachment/6e1f58718b808847cde1bccb5ded6d7e22ee78c4.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"As synthetic data becomes increasingly popular in machine learning tasks, numerous methods---without formal differential privacy guarantees---use synthetic data for training. These methods often claim, either explicitly or implicitly, to protect the privacy of the original training data.\nIn this work, we explore four different training paradigms: coreset selection, dataset distillation, data-free knowledge distillation, and synthetic data generated from diffusion models. While all these methods utilize synthetic data for training, they lead to vastly different conclusions regarding privacy preservation. We caution that empirical approaches to preserving data privacy require careful and rigorous evaluation; otherwise, they risk providing a false sense of privacy."},"_bibtex":{"value":"@inproceedings{\nzhao2025does,\ntitle={Does Training with Synthetic Data Truly Protect Privacy?},\nauthor={Yunpeng Zhao and Jie Zhang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=C8niXBHjfO}\n}"},"title":{"value":"Does Training with Synthetic Data Truly Protect Privacy?"},"pdf":{"value":"/pdf/1a75577d2e1d1f3d2d69d02202f1c3559d7d8f3e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|does_training_with_synthetic_data_truly_protect_privacy"},"authorids":{"value":["~Yunpeng_Zhao2","~Jie_Zhang14"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yunpeng Zhao","Jie Zhang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes FLARE, a physics-informed GNN rewiring method for fluid simulation that selectively adds directional 2-hop edges based on instantaneous flow alignment. The authors show good performance over prior methods like PIORF and generic 2-hop rewiring."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"Q1. What is the rationale for using $T=0$ in the base FLARE experiments, and do you believe this sufficiently tests the physical alignment principle? What happens in the case of positive thresholds where $T>0$?\n\nQ2. More evidence is needed to support the claim that PIORF \"does not align with physical principles\" (line 237, footnote 2). The differences between FLARE and PIORF need to be discussed more clearly, and the velocity gradient (strain rate) aspect of PIORF should be addressed. Can you explain more precisely which aspects of PIORF are less physically straightforward?\n\nQ3. Why does FLARE underperform PIORF on Airfoil density with BSMS-GNN?\n\nQ4. Please provide a detailed computational cost comparison between baseline, PIORF, 2-HOP-ALL, and FLARE for each dataset. How much overhead does dynamic rewiring add?\n\nQ5. Your comparison confounds multiple factors (locality, directionality, selection criterion). Can you provide ablations that isolate each factor? Specifically, factors such as whether all connections other than inverse connections are directional, or the proportion of bidirectional connections?\n\nQ6. What are the detailed settings for 3-hop and 4-hop in the ablation study?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper demonstrates consistent and substantial performance improvements over PIORF and structural baselines across three diverse datasets (unsteady, steady, compressible/incompressible) and 3 architectures.\n2. The authors present a more straightforward solution than existing physics-based rewiring methods.\n3. The introduction clearly articulates what problems of existing methods the paper aims to solve."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed FLARE and its evaluation approach are insufficient to support the authors' claimed motivation of \"rigorous adherence to physical principles.\"\n\n2. The use of a zero threshold ($T=0$) raises questions about whether structural directionality (directional 2-hop) is the primary driver of performance rather than the alignment principle itself. Thus, the paper weakens the rigor of the \"physics-informed\" claim.\n\n3. There is insufficient fluid dynamics or GNN-theoretical (e.g., curvature, over-squashing) justification for determining 2-hop as the optimal locality. The paper presents only experimental results.\n\n4. The computational cost of dynamic graph reconstruction at every time step during rollout (inference time overhead) is not analyzed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925558516,"tcdate":1761981807808,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15260/Reviewer_PxA5"],"signatures":["ICLR.cc/2026/Conference/Submission15260/Reviewer_PxA5"],"forum":"izLvJEBkae","number":3,"license":"CC BY 4.0","cdate":1761981807808,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15260/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925558516,"domain":"ICLR.cc/2026/Conference","replyto":"izLvJEBkae","id":"cYNkSlm2DD","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed Graph Neural Networks; CFD; mesh; rewiring; unsteady flow"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"To overcome computation burden of traditional computational fluid dynamics (CFD) simulations, researchers have explored different architectures to develop physics-informed simulation methods. Among them, graph neural networks (GNN) are most suitable for adopting CFD meshes, which are extensively used in engineering and industrial applications. However, classical GNNs propagate information among neighbour nodes, which highly restrict information exchange within the network. To address this issue, graph rewiring methods have been developed for generic graph problems, but not particular for fluid simulation. PIORF, introducing edges connecting distant nodes, is the first graph rewiring method to do so, and previous experiments have demonstrated its effectiveness against state-of-the-art generic rewiring methods. Nevertheless, in this work, we found that simply connecting all 2-hop nodes can provide competitive performance with PIORF. This result raises three questions: 1) Is physics-informed rewiring really useful for improving flow predictions? 2) Should we consider just local connection, instead of connecting distant nodes? 3) Do we need to change the connections based on input flow for rollout simulations? By thoroughly adopting physical fluid principles, we propose a simple yet very efficient method, Flow Alignment Rewiring (FLARE) technique, which connects 2-hop nodes only when the node direction aligns with input flow direction. Hence, FLARE is a physics-informed local rewiring method, different from PIORF and well-aligned with fluid physics. Extensive numerical experiments on flows over a cylinder and single and tandem airfoil under different flow conditions and deep network architectures demonstrate that FLARE outperforms PIORF and various 2-hop rewiring approaches by a significant margin."},"_bibtex":{"value":"@misc{\nli2026graph,\ntitle={Graph Rewiring based on Flow Alignment for Improving Fluid Simulation},\nauthor={Zenong Li and Wei Xian Lim and Wai Lee Chan and Adams Wai-Kin Kong},\nyear={2026},\nurl={https://openreview.net/forum?id=izLvJEBkae}\n}"},"title":{"value":"Graph Rewiring based on Flow Alignment for Improving Fluid Simulation"},"pdf":{"value":"/pdf/fde6cfdb3e702fa96f577b9f6fb0cfb47669442b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|graph_rewiring_based_on_flow_alignment_for_improving_fluid_simulation"},"authorids":{"value":["~Zenong_Li1","~Wei_Xian_Lim2","~Wai_Lee_Chan1","~Adams_Wai-Kin_Kong1"]},"authors":{"value":["Zenong Li","Wei Xian Lim","Wai Lee Chan","Adams Wai-Kin Kong"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a physics-informed neural network for subglacial bed topography that couples two physics fidelities (shallow-ice approximation and a reduced-stokes surrogate) inside one PINN loss and adds a boundary-aware weak-form term that mixes Neumann traction and optional Dirichlet constraints at glacier margins. On a Greenland radar track dataset, the method reports test mse with $r^2$ and claims better physics residuals than single-fidelity or purely data-driven baselines."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please address the weaknesses above."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. The paper aims to address a relevant geoscience problem where labeled data are sparse and physical constraints matter, and motivates the need for physics-guided inversion rather than black-box regression.\n2. The multi-fidelity idea is sensible: use SIA for cheap broad constraints and a higher-order residual for added fidelity, with learned or fixed weights to balance them."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The “reduced-stokes” residual is specified as $r_{Stokes} = −ν\\Delta \\hat b − f$ with $ν=1$ and $f=0$, which collapses to a Poisson-like smoothness on the bed field rather than a demonstrably derived momentum balance tied to ice rheology or sliding. This risks being a hand-crafted regularizer rather than a true higher-fidelity physics term.\n2. The physical role of the predicted variable is unclear: the network maps surface features to bed elevation $\\hat b$, but the residuals are written directly on $\\hat b$ without showing how $\\hat b$ couples to velocity, thickness, or stresses in SIA/Stokes. \n3. Neumann boundary condition uses $g_N = 0$ by default, which is a strong assumption..\n4. The dataset split appears random 80/20 over track points; this can cause spatial leakage because nearby points on a flight line are strongly correlated. The text says the team ensured points were “not too similar,” but it lacks a rigorous spatial holdout protocol and distance thresholds.\n5. Metric reporting mixes “training units” and “physical units,” leading to confusing cross-model comparisons (e.g., random forest shows r²=0.987 yet huge mae/rmse due to unit scaling). \n6. The baseline shows a “weighted physics objective” of 0 while listing nonzero residuals; SIA and Stokes residuals are identical to two decimals in multiple rows. Can you explain this?\n7. The “uncertainty weighting” via Kendall log-variance is used for loss balancing, but no calibration, uncertainty evaluation, or learned weight trajectories are presented, so the “uncertainty” interpretation is not strong.\n8. Interpolated boundary labels may leak target information."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942562384,"tcdate":1761974231458,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23217/Reviewer_2Vt4"],"signatures":["ICLR.cc/2026/Conference/Submission23217/Reviewer_2Vt4"],"forum":"F9pbcENAXv","number":2,"license":"CC BY 4.0","cdate":1761974231458,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23217/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942562384,"domain":"ICLR.cc/2026/Conference","replyto":"F9pbcENAXv","id":"S94Ma2ARhm","forumContent":{"TLDR":{"value":"We introduce a multi-fidelity PINN framework that leverages boundary-aware losses for more accurate PDE-constrained learning tasks, demonstrated on ice-bed topography prediction."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-fidelity learning","PINNs","boundary-aware losses","Glacier bed inversion","Ice dynamics"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Predicting ice dynamics and sea-level rise requires an understanding of subglacial bedrock topography; however, inversion remains a challenging task in data-sparse regions where surface observations are limited. Some conventional machine learning methods face challenges in predicting subglacial topography due to heavy reliance on purely data correlations and cannot guarantee physical consistency, especially in data-sparse regions. Physics-Informed Neural Networks (PINNs) address this limitation by embedding partial differential equation (PDE) constraints into deep learning, enabling more physically consistent predictions. However, most existing PINN formulations depend on a single fidelity of physics, and soft boundary penalties can still compromise performance. We propose a multi-fidelity PINN framework for ice-bed topography prediction that advances beyond these limitations in two ways. First, we introduce multi-fidelity residual coupling, jointly enforcing the shallow-ice approximation (SIA) and reduced-Stokes equations within a single network. This coupling improves accuracy while maintaining physics consistency, achieving strong predictive performance (e.g., Test MSE = 0.028, and $R^2$ = 0.97). Second, we design a boundary-aware weak-form loss that supports traction/flux (Neumann) and optional Dirichlet constraints, allowing flexible enforcement of margin physics. Experiments show that hard Dirichlet enforcement over-constrains the model and reduces accuracy, while soft or selective enforcement preserves predictive quality. To our knowledge, this is the first Physics-Informed Neural Network (PINN) framework for predicting ice-bed topography that unifies multi-fidelity partial differential equation (PDE) residuals with configurable boundary-aware losses, providing a practical and extensible approach to physically plausible predictions in data-sparse regimes."},"_bibtex":{"value":"@misc{\ntabassum2026multifidelity,\ntitle={Multi-Fidelity Physics-Informed Neural Networks ({PINN}) with Boundary-Aware Losses for Ice-Bed Topography Prediction},\nauthor={Tartela Tabassum and Pavan Raj Ravi and Jianwu Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=F9pbcENAXv}\n}"},"title":{"value":"Multi-Fidelity Physics-Informed Neural Networks (PINN) with Boundary-Aware Losses for Ice-Bed Topography Prediction"},"pdf":{"value":"/pdf/0cf02073255bc7317e3fc6b5aa201ef578ed18b4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"tabassum|multifidelity_physicsinformed_neural_networks_pinn_with_boundaryaware_losses_for_icebed_topography_prediction"},"authorids":{"value":["~Tartela_Tabassum1","~Pavan_Raj_Ravi1","~Jianwu_Wang1"]},"authors":{"value":["Tartela Tabassum","Pavan Raj Ravi","Jianwu Wang"]}},"version":2},{"content":{"summary":{"value":"This work proposes ParFam, a simple regression method with a fixed and predefined structure, to tackle the symbolic regression problem. In ParFam, the function expression structure is directly specified by the user in advance, and then the coefficients are learned with a sparsity regularization from the observed data. In this way, the original symbolic regression problem can be reduced to a continuous optimization problem with respect to the coefficients.\n\nBased on the ground-truth problems from SRBench and the knowledge from the Cambridge Handbook of Physics Formulation, this work proposes a reasonable parametric expression structure to represent the physical formulas. Then, it uses a global continuous optimization method (basin-hopping algorithm) to find the optimal coefficients of the predefined expression. Experimental results show that ParFam can achieve promising performance on the symbolic regression problem for physics formulas."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"Symbolic regression (SR) is an important but difficult problem that can be found in various domains. The proposed ParFam method can achieve promising performance on SR for physics formulas in a straightforward way."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although I enjoy reading this paper and appreciate the explicit discussion on the limitations, I have some major concerns about ParFam.\n\n**1. Is It still Symbolic Regression?**\n\nTo my understanding, symbolic regression is a learning-based approach to find the mathematical expression of a function from the observed data, which includes two important components:\n\n- Learn the analytical function structure;\n\n- Optimize the coefficients (parameters) of the structure;\n\nThe former is unique for symbolic regression, which distinguishes it from the other regression problems. Symbolic regression is difficult and is currently shown to be HP-hard with formal proof [1]. I think this is the reason why an efficient (approximate) SR algorithm will be \"usually quite complicated\" as described in this work.   \n\nIn ParFam, however, the analytical structure learning step is totally bypassed with a predefined function structure. The original problem is hence reduced to sparse regression with a fixed structure, and the only goal is to find the optimal coefficients. Is it still symbolic regression?  \n\n**2. Strong Prior Knowledge on Physics are Required**\n\nTo achieve promising performance on SR problems for physics formulas, ParFam requires prior knowledge of all possible physics formulas, as from SRBench and the Cambridge Handbook of Physics Formulation. I think this prior knowledge is very strong and only specific to physics formulas, and is hard to be generalized for other SR problems in real-world applications.  \n\n**3. DL-ParFam**\n\nThe idea of DL-ParFam, a deep learning-based pretrain model for ParFam, is interesting. But it is currently more like a toy prototype, and only tested on very simple synthetic problems. To truly show the advantage of DL-ParFam over other pre-training-based SR methods, a concrete model design on real-world SR applications is required. \n\nIn DL-ParFam, the model only takes the function value y as input to predict the mask c for all parameters, and all information of the function input x is completely ignored. It is hard to believe this approach can provide a reasonably good prediction for real-world SR applications, especially those with complicated structures. \n\nTo build the pre-trained model, DL-ParFam requires the input data x to have the same dimension $m$, and the data should be sampled on the same grid across all different data sets. Can this requirement be easily satisfied for physical SR problems and other SR problems?\n\n**4. Experiments**\n\nSince ParFam has a strong prior knowledge of the physical formulas, it is expected it can have promising performance on the physics SR problems. Indeed, according to the results, ParFam even discards part of the observed data, and only requires a subset of 500-1000 data points for coefficient optimization. It is hard to imagine this procedure could work well for real-world SR problems.\n\nDL-ParFam is only tested on very simple synthetic problems. It is hard to judge its potential for solving real-world application problems with complicated structures."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Please see the weaknesses section. I am willing to adjust my rating if the issues in weaknesses are well addressed.\n\n[1] Symbolic Regression is NP-hard. TMLR 2022."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700567393536,"tcdate":1698759558730,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5667/Reviewer_LMH9"],"signatures":["ICLR.cc/2024/Conference/Submission5667/Reviewer_LMH9"],"forum":"5vXDQ65dzH","number":3,"license":"CC BY 4.0","cdate":1698759558730,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5667/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700567393536,"domain":"ICLR.cc/2024/Conference","replyto":"5vXDQ65dzH","id":"xHjMh4Lffj","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"TLDR":{"value":"We introduce the symbolic regression method ParFam, which achieves state-of-the-art by leveraging the structure of common physical laws, and show how to extend ParFam to improve the performance further using deep learning."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["symbolic regression","global optimization","deep learning"]},"primary_area":{"value":"general machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The problem of symbolic regression (SR) arises in many different applications, such as identifying physical laws or deriving mathematical equations describing the behavior of financial markets from given data. Various methods exist to address the problem of SR, often based on genetic programming. However, these methods are usually quite complicated and require a lot of hyperparameter tuning and computational resources. \nIn this paper, we present our new method ParFam that utilizes parametric families of suitable symbolic functions to translate the discrete symbolic regression problem into a continuous one, resulting in a more straightforward setup compared to current state-of-the-art methods. \nIn combination with a powerful global optimizer, this approach results in an effective method to tackle the problem of SR. \nFurthermore, it can be easily extended to more advanced algorithms, e.g., by adding a deep neural network to find good-fitting parametric families. \nWe prove the performance of ParFam with extensive numerical experiments based on the common SR benchmark suit SRBench, showing that we achieve state-of-the-art results. Our code can be found at https://anonymous.4open.science/r/parfam-90FC/README.md."},"_bibtex":{"value":"@misc{\nscholl2024parfam,\ntitle={ParFam - Symbolic Regression Based on Continuous Global Optimization},\nauthor={Philipp Scholl and Katharina Bieker and Hillary Hauger and Gitta Kutyniok},\nyear={2024},\nurl={https://openreview.net/forum?id=5vXDQ65dzH}\n}"},"title":{"value":"ParFam - Symbolic Regression Based on Continuous Global Optimization"},"pdf":{"value":"/pdf/ac908633e71d1dbb3761bd3911df5a9e2b618369.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"scholl|parfam_symbolic_regression_based_on_continuous_global_optimization"},"authorids":{"value":["~Philipp_Scholl2","~Katharina_Bieker1","~Hillary_Hauger3","~Gitta_Kutyniok2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Philipp Scholl","Katharina Bieker","Hillary Hauger","Gitta Kutyniok"]}},"version":2},{"content":{"summary":{"value":"This work focuses on tackling the problem of single-person motion estimation from a monocular video. Current approaches produce temporal artifacts such as jittering. Most approaches are entirely kinematic while others that combine physics, do it by re-simulating the kinematic inputs by using automatic PD controllers. These methods, however, require simplifying assumptions that compromise the motion realism.This method combines information from two sources to produce an optimal estimate of single-person 3D human motion. One source comes from an off-the-shelf kinematic motion estimation pipeline and the other from a differentiable physics formulation. These two \"measurements\" about the motion are combined in a Kalman-filter to generate an optimal output.  Here, the authors propose to selectively incorporate the physics models with the kinematics observations in an online setting, taking as inspiration the neural Kalman-filter. The method uses a meta-PD controller and a physics-based simulation step. The Kalman filter is realized via a recurrent neural network which aims to balance the kinematic inputs with the simulation. \n\nThe authors propose an end-to-end model for this purpose which is not trivial to accomplish. The method is capable of capturing accurate global trajectories and, at the same time, producing physically plausible human poses."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- L38: To be clear, what the authors refer to as a \"differentiable physics simulation\" is the use of rigid body dynamic equations or is it an actually differentiable simulator, e.g.,  Tiny Differentiable Simulator from PyBullet, similar to Gartner et. al?\n- L116: Authors create a proxy character (shown in appendix B). I wonder if it is possible to use the humanoid generated by SimPOE or KinPoly for the same end? This latter humanoid seems to be the most realistic representation of a 3D human body for simulation. I would like what the authors think about this and why they chose this specific form.\n- Sec. 3.2: How are the contacts modeled? Are they modeled directly with the GRU and data annotations or is there also a specific physics formulation to account for these contacts?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"* The paper is very well written and the experiments are well presented which makes the paper easy to follow. \n* Authors present extensive experiments comparing several SoTA methods and include in-domain and out-of-domain test data for these. \n* The setup and the method are sound. \n* I would say that this is the first paper that successfully combines information from kinematic estimates and a differentiable physics simulation step in an end-to-end manner. It is not trivial to refine kinematic estimates with physics simulation (or physics informed estimates) outside the RL framework. It seems that the neural Kalmal filter is a promising direction to bridge the gap between kinematics and physics. I believe that this work is of high significance for the field."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### **Presentation**\nThe qualitative results could be better presented as it is sometimes hard to have a good sense of the pose estimated by the kinematic approach (TRACE) both in Fig. 6 and in the supplementary .gif images. The way the poses and the original video are visualized can be improved. First, I would suggest making the kinematic skeleton more visible as it is “obscured” by OSDCap results. Even better, it would be nice to have SMPL visualizations of GT, kinematics and OSDCap as it is presented in Fig. 1. In my opinion, changing visualization styles within the paper can reduce the presentation quality. I also advise the authors to focus on video results with several examples. If possible I would like to see this as part of the rebuttal, if not, then this should be present for the camera ready version of the paper. \n\n### **Physics-based metrics**\nSec 4.3: It would be interesting to show the results for more physics-based metrics other than the Acceleration metric, for example, foot skating and ground penetration, which should be corrected by the physics formulation. As authors are modeling contacts and have physics-based losses (e.g., friction, velocity), it would be interesting to know these metrics in comparison with the SoTA or at least the baseline used (TRACE).\n\n### Minor\n- (Typo) Fig.2: 4th line: performs contains-->contains.\n- (Typo) L295: HMDCap -->OSDCap?\n\n### References\nI would advise the authors to include most recent papers that combine kinematic and physics estimates for pose/motion estimation, either in the introduction or related works. For example:\n- Ugrinovic et. al., “MultiPhys: Multi-Person Physics-aware 3D Motion Estimation” (CVPR 2024).\n- Zhang et. al. “PhysPT: Physics-aware Pretrained Transformer for Estimating Human\nDynamics from Monocular Videos” (CVPR 2024)."},"limitations":{"value":"Yes."}},"nonreaders":[],"tmdate":1730878639967,"tcdate":1720687950064,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission374/Reviewer_BVYQ"],"signatures":["NeurIPS.cc/2024/Conference/Submission374/Reviewer_BVYQ"],"forum":"RkOT8rAmRR","number":3,"license":"CC BY 4.0","cdate":1720687950064,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission374/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878639967,"domain":"NeurIPS.cc/2024/Conference","replyto":"RkOT8rAmRR","id":"BYpnxxKbdN","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["human motion","dynamics","optimal state","kalman filter","physics-based"]},"supplementary_material":{"value":"/attachment/b72ba1da11508232b4d347e8798c559c3b5203ba.zip"},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Human motion capture from monocular videos has made significant progress in recent years. However, modern approaches often produce temporal artifacts, e.g. in form of jittery motion and struggle to achieve smooth and physically plausible motions. Explicitly integrating physics, in form of internal forces and exterior torques, helps alleviating these artifacts. Current state-of-the-art approaches make use of an automatic PD controller to predict torques and reaction forces in order to re-simulate the input kinematics, i.e. the joint angles of a predefined skeleton. However, due to imperfect physical models, these methods often require simplifying assumptions and extensive preprocessing of the input kinematics to achieve good performance. To this end, we propose a novel method to selectively incorporate the physics models with the kinematics observations in an online setting, inspired by a neural Kalman-filtering approach. We develop a control loop as a meta-PD controller to predict internal joint torques and external reaction forces, followed by a physics-based motion simulation. A recurrent neural network is introduced to realize a Kalman filter that attentively balances the kinematics input and simulated motion, resulting in an optimal-state dynamics prediction. We show that this filtering step is crucial to provide an online supervision that helps balancing the shortcoming of the respective input motions, thus being important for not only capturing accurate global motion trajectories but also producing physically plausible human poses. The proposed approach excels in the physics-based human pose estimation task and demonstrates the physical plausibility of the predictive dynamics, compared to state of the art. The code is available on https://github.com/cuongle1206/OSDCap."},"_bibtex":{"value":"@inproceedings{\nle2024optimalstate,\ntitle={Optimal-state Dynamics Estimation for Physics-based Human Motion Capture from Videos},\nauthor={Cuong Le and Manon Kok and Viktor Johansson and Bastian Wandt},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=RkOT8rAmRR}\n}"},"title":{"value":"Optimal-state Dynamics Estimation for Physics-based Human Motion Capture from Videos"},"pdf":{"value":"/pdf/325366bc6a69db293281709cbf852252b3527c07.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"le|optimalstate_dynamics_estimation_for_physicsbased_human_motion_capture_from_videos"},"authorids":{"value":["~Cuong_Le1","~Viktor_Johansson1","~Manon_Kok1","~Bastian_Wandt2"]},"authors":{"value":["Cuong Le","Viktor Johansson","Manon Kok","Bastian Wandt"]}},"version":2},{"content":{"summary":{"value":"The paper presents a supervised framework that predicts physical properties of 3D scenes directly from visual inputs. The method, PIXIE, trains a feed-forward 3D U-Net on CLIP-distilled volumetric features to infer both discrete material types and continuous physical parameters from multi-view RGB images. These predictions can be coupled with Gaussian Splatting and simulated using the Material Point Method (MPM) to produce realistic physics-based animations. To support training, the authors introduce PIXIEVERSE, a dataset of 3D assets labeled with physical material annotations across 10 semantic categories. Experiments show that PIXIE achieves higher realism scores and is three orders of magnitude faster than test-time optimization baselines, while also generalizing zero-shot to real-world scenes despite being trained solely on synthetic data."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In real-world environments with cluttered backgrounds and potentially moving objects, how does the proposed method identify and isolate the dynamic regions relevant for physical simulation? \n2. The paper states that each Gaussian in the Gaussian Splatting model is treated as an MPM particle (Sec. 3.1), but this mapping might be uneven, as splats are not uniformly distributed and often concentrate near visible surfaces. Could this cause inconsistencies in material distribution or simulation stability? Have any corrective strategies been applied to mitigate these issues?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper introduces PIXIEVERSE, an open-source dataset of 1,624 3D assets annotated with physical material parameters, enabling future research.\n- The paper proposes the first supervised learning method that directly predicts both discrete material classes and continuous physical parameters (Young’s modulus, Poisson’s ratio, density) from 3D visual features, which enables faster inference than prior test-time optimization methods.\n- Extensive experiments on synthetic data and real-world data show the superior performance of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The PIXIEVERSE dataset relies heavily on semi-automatic annotations generated by vision-language models. Such labels may contain systematic biases or noise, and their accuracy is not quantitatively validated.\n- While Figure 4 reports a 2-second inference time, this does not account for the required NeRF/feature-field reconstruction step, which can be computationally expensive.\n- Although the paper claims zero-shot generalization on Spring-Gaus data, it omits comparisons against Spring-Gaus, instead asserting that \"no other baseline can generalize under this setting\".\n- Missing related work: It would be better if the author could compare with [1], which aims to estimate physical properties implicitly from videos.\n\n[1] Zhu X, Deng H, Yuan H, et al. Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D Video. ICLR 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923744881,"tcdate":1761641813291,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12998/Reviewer_wPyp"],"signatures":["ICLR.cc/2026/Conference/Submission12998/Reviewer_wPyp"],"forum":"PHUczJGCgc","number":2,"license":"CC BY 4.0","cdate":1761641813291,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12998/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923744881,"domain":"ICLR.cc/2026/Conference","replyto":"PHUczJGCgc","id":"aNoWYQ6lpy","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Computer vision","4D reconstruction","Physics learning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Inferring the physical properties of 3D scenes from visual information is a critical yet challenging task for creating interactive and realistic virtual worlds. While humans intuitively grasp material characteristics such as elasticity or stiffness, existing methods often rely on slow, per-scene optimization, limiting their generalizability and application. To address this problem, we introduce PIXIE, a novel method that trains a generalizable neural network to predict physical properties across multiple scenes from 3D visual features purely using supervised losses. Once trained, our feed-forward network can perform fast inference of plausible material fields, which coupled with a learned static scene representation like Gaussian Splatting enables realistic physics simulation under external forces. To facilitate this research, we also collected PIXIEVERSE, one of the largest known datasets of paired 3D assets and physic material annotations. Extensive evaluations demonstrate that PIXIE is about 1.46-4.39x better and orders of magnitude faster than test-time optimization methods. By leveraging pretrained visual features like CLIP, our method can also zero-shot generalize to real-world scenes despite only ever been trained on synthetic data. https://pixie-2026-12998.github.io/"},"_bibtex":{"value":"@misc{\nle2026pixie,\ntitle={Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels},\nauthor={Long Le and Ryan Lucas and Chen Wang and Chuhao Chen and Dinesh Jayaraman and Eric Eaton and Lingjie Liu},\nyear={2026},\nurl={https://openreview.net/forum?id=PHUczJGCgc}\n}"},"title":{"value":"Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels"},"pdf":{"value":"/pdf/67f200fa43c94eee29bb076510135bd2e6357bc2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"le|pixie_fast_and_generalizable_supervised_learning_of_3d_physics_from_pixels"},"authorids":{"value":["~Long_Le1","~Ryan_Lucas1","~Chen_Wang13","~Chuhao_Chen1","~Dinesh_Jayaraman2","~Eric_Eaton2","~Lingjie_Liu1"]},"authors":{"value":["Long Le","Ryan Lucas","Chen Wang","Chuhao Chen","Dinesh Jayaraman","Eric Eaton","Lingjie Liu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes AutoBio, a simulation benchmark designed specifically for laboratory tasks and environments. The authors provide a laboratory equipment asset generation pipeline, relevant physics plugins, rendering which supports e.g. transparent materials, and a data generation pipeline. The authors also benchmark VLA and IL baselines on tasks of varying levels of difficulty."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"- Are benchmarks on simulation performance available (e.g. simulation speed) to gauge evaluation speed, and potentially the applicability of online learning methods?\n- The authors note the potential multitask capability of the VLA models. Did authors run multitask experiments on the VLA and IL models?\n- Are there future plans to provide additional tasks (e.g. longer horizon, more subtasks, etc), leveraging the same physics and rendering framework?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Useful task suite in an area of robotic manipulation with few realistic benchmarks\n- Careful consideration of physics, rendering, assets, etc in the context of biology tasks\n- VLA and IL baselines are relevant and highlight weaknesses in more complex tasks\n- The presentation is clear and contributions well-explained"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Seeing as the realistic assets, physics, and rendering are a central focus, validation on a real robot setup (even on the simpler tasks) would support claims of realism\n- The paper notes VLAs may perform well as multi-task agents in the discussion, however this setting is not evaluated"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923976207,"tcdate":1761979475215,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13313/Reviewer_qBEk"],"signatures":["ICLR.cc/2026/Conference/Submission13313/Reviewer_qBEk"],"forum":"UUE6HEtjhu","number":4,"license":"CC BY 4.0","cdate":1761979475215,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13313/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923976207,"domain":"ICLR.cc/2026/Conference","replyto":"UUE6HEtjhu","id":"EkyMYUsEdL","forumContent":{"TLDR":{"value":"AutoBio offers a novel simulation and benchmark with biologically-grounded manipulation tasks for precise, multimodal robotic automation in digital biology laboratories."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["robotics","robot learning","vision language action model","biology experimental operation","AI for science"]},"supplementary_material":{"value":"/attachment/39a01109a6f1804a67f5f0bc0c89daefb6503bae.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Vision-language-action (VLA) models have shown promise as generalist robotic policies by jointly leveraging visual, linguistic, and proprioceptive modalities to generate action trajectories. While recent benchmarks have advanced VLA research in domestic tasks, professional science-oriented domains remain underexplored. We introduce AutoBio, a simulation framework and benchmark designed to evaluate robotic automation in biology laboratory environments—an application domain that combines structured protocols with demanding precision and multimodal interaction. AutoBio extends existing simulation capabilities through a pipeline for digitizing real-world laboratory instruments, specialized physics plugins for mechanisms ubiquitous in laboratory workflows, and a rendering stack that support dynamic instrument interfaces and transparent materials through physically based rendering. Our benchmark comprises biologically grounded tasks spanning three difficulty levels, enabling standardized evaluation of language-guided robotic manipulation in experimental protocols. We provide infrastructure for demonstration generation and seamless integration with VLA models. Baseline evaluations with SOTA VLA models reveal significant gaps in precision manipulation, visual reasoning, and instruction following in scientific workflows. By releasing AutoBio, we aim to catalyze research on generalist robotic systems for complex, high-precision, and multimodal professional environments."},"_bibtex":{"value":"@inproceedings{\nlan2026autobio,\ntitle={AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory},\nauthor={Zhiqian Lan and Yuxuan Jiang and Ruiqi Wang and Xuanbing Xie and Rongkui Zhang and Yicheng Zhu and LI PEIHANG and Tianshuo Yang and Tianxing Chen and Haoyu Gao and Xiaokang Yang and Xuelong Li and Hongyuan Zhang and Yao Mu and Ping Luo},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=UUE6HEtjhu}\n}"},"title":{"value":"AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory"},"pdf":{"value":"/pdf/0305ac07e8a1dd414215ee3c4ca7eaadd75b0b6a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lan|autobio_a_simulation_and_benchmark_for_robotic_automation_in_digital_biology_laboratory"},"authorids":{"value":["~Zhiqian_Lan1","~Yuxuan_Jiang1","~Ruiqi_Wang17","~Xuanbing_Xie1","~Rongkui_Zhang1","~Yicheng_Zhu2","~LI_PEIHANG1","~Tianshuo_Yang1","~Tianxing_Chen2","~Haoyu_Gao6","~Xiaokang_Yang1","~Xuelong_Li2","~Hongyuan_Zhang1","~Yao_Mu1","~Ping_Luo2"]},"authors":{"value":["Zhiqian Lan","Yuxuan Jiang","Ruiqi Wang","Xuanbing Xie","Rongkui Zhang","Yicheng Zhu","LI PEIHANG","Tianshuo Yang","Tianxing Chen","Haoyu Gao","Xiaokang Yang","Xuelong Li","Hongyuan Zhang","Yao Mu","Ping Luo"]}},"version":2},{"content":{"summary":{"value":"The paper proposes an approach to sample from the posterior distribution based on score-based diffusion models with a particular focus on inverse problems in physics. \n\nThe proposed approach, to my understanding, has two novel contributions:\n- they use reverse-physics simulations and augment them with score estimates that supposedly leads to superior performance in sampling,\n- they propose a multi-step training regime where scores are estimated in sequential manner for a given time-horizon, in contrary to the standard single-step score estimation.\n\nThe method was tested on different synthetic data experiments, and both the proposed approaches seem to produce meaningful improvements. "},"soundness":{"value":"3 good"},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"- How is the ground truth obtained in the case of SDEs?\n- In Section 3.4 (Fig. 6), why is the spectral error worse for SDEs compared to ODEs while it is the other way around in other experiments?"},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"strengths":{"value":"- The multi-step training of the score function seems to be effective in synthetic data applications.\n- The theoretical results, although not particularly novel, provide a complete picture of the proposed methods. \n- The experiments presented validate the proposed ideas well."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- One of the novelties claimed in the paper is the use of score-function to simply _refine_ the outputs of a reverse-physics simulator. Can't one simply view the reverse-physics simulator as (non-learned) _part of_ the parametric score model? In this case, the claim simply becomes that a physics-informed model to approximate the score is better than one that is oblivious to the physics? Isn't this an unsurprising statement? This is the underlying motivation behind physics-informed neural nets (PINNs), a relatively large area of research. Could the authors clarify?\n- In some of the experiments, the authors use the LPIPS metric to evaluate the quality of the solution. This does not make much sense as LPIPS is designed to evaluate the \"perceptual\" quality of the image, which has nothing to do with the physical accuracy (the metric one cares about)."},"limitations":{"value":"- Releasing the code could be crucial as there are many delicate details one might need to get right for the method to work."}},"nonreaders":[],"tmdate":1702410809901,"tcdate":1690557720617,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1882/Reviewer_SQFk"],"signatures":["NeurIPS.cc/2023/Conference/Submission1882/Reviewer_SQFk"],"forum":"2BpoGPSDCR","number":6,"license":"CC BY 4.0","cdate":1690557720617,"mdate":1702410809901,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1882/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"2BpoGPSDCR","id":"sSr3rt1TTH","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["inverse problems","diffusion models","learned corrections","score matching"]},"_bibtex":{"value":"@inproceedings{\nholzschuh2023solving,\ntitle={Solving Inverse Physics Problems with Score Matching},\nauthor={Benjamin Holzschuh and Simona Vegetti and Nils Thuerey},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=2BpoGPSDCR}\n}"},"title":{"value":"Solving Inverse Physics Problems with Score Matching"},"paperhash":{"value":"holzschuh|solving_inverse_physics_problems_with_score_matching"},"TLDR":{"value":"We propose a novel methodology for solving inverse problems that involve the temporal evolution of physical systems."},"abstract":{"value":"We propose to solve inverse problems involving the temporal evolution of physics systems by leveraging recent advances from diffusion models. \nOur method moves the system's current state backward in time step by step by combining an approximate inverse physics simulator and a learned correction function. \nA central insight of our work is that training the learned correction with a single-step loss is equivalent to a score matching objective, while recursively predicting longer parts of the trajectory during training relates to maximum likelihood training of a corresponding probability flow.\nWe highlight the advantages of our algorithm compared to standard denoising score matching and implicit score matching, as well as fully learned baselines for a wide range of inverse physics problems. The resulting inverse solver has excellent accuracy and temporal stability and, in contrast to other learned inverse solvers, allows for sampling the posterior of the solutions. Code and experiments are available at https://github.com/tum-pbs/SMDP."},"pdf":{"value":"/pdf/e14d868c93d4bb63690c02fd9c5fcb4f28eed65c.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Benjamin_Holzschuh1","~Simona_Vegetti1","~Nils_Thuerey1"]},"authors":{"value":["Benjamin Holzschuh","Simona Vegetti","Nils Thuerey"]}},"version":2},{"content":{"venue":{"value":"Intelligent Tutoring Systems 2002"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/3-540-47987-2_40.pdf"},"venueid":{"value":"dblp.org/conf/ITS/2002"},"paperhash":{"value":"vanlehn|minimally_invasive_tutoring_of_complex_physics_problem_solving"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Kurt_VanLehn:","~Collin_F._Lynch1","https://dblp.org/search/pid/api?q=author:Linwood_Taylor:","https://dblp.org/search/pid/api?q=author:Anders_Weinstein:","https://dblp.org/search/pid/api?q=author:Robert_Shelby:","https://dblp.org/search/pid/api?q=author:Kay_G._Schulze:","https://dblp.org/search/pid/api?q=author:Donald_Treacy:","https://dblp.org/search/pid/api?q=author:Mary_Wintersgill:"]},"html":{"value":"https://doi.org/10.1007/3-540-47987-2_40"},"_bibtex":{"value":"@inproceedings{DBLP:conf/its/VanLehnLTWSSTW02,\n  author={Kurt VanLehn and Collin F. Lynch and Linwood Taylor and Anders Weinstein and Robert Shelby and Kay G. Schulze and Donald Treacy and Mary Wintersgill},\n  title={Minimally Invasive Tutoring of Complex Physics Problem Solving},\n  year={2002},\n  cdate={1009843200000},\n  pages={367-376},\n  url={https://doi.org/10.1007/3-540-47987-2_40},\n  booktitle={Intelligent Tutoring Systems},\n  crossref={conf/its/2002}\n}\n"},"abstract":{"value":"Solving complex physics problems requires some kind of knowledge for selecting appropriate applications of physics principles. This knowledge is tacit, in that it is not explicitly taught in textbooks, existing tutoring systems or anywhere else. Experts seem to have acquired it via implicit learning and may not be aware of it. Andes is a coach for physics problem solving that has had good evaluations, but still does not teach complex problem solving as well as we would like. The conventional ITS approach to increasing its effectiveness requires teaching the tacit knowledge explicitly, and yet this would cause Andes to be more invasive. In particular, the textbooks and instructors would have to make space in an already packed curriculum for teaching the tacit knowledge. This paper discusses our attempts to teach the tacit knowledge without making Andes more invasive."},"title":{"value":"Minimally Invasive Tutoring of Complex Physics Problem Solving"},"authors":{"value":["Kurt VanLehn","Collin F. Lynch","Linwood Taylor","Anders Weinstein","Robert Shelby","Kay G. Schulze","Donald Treacy","Mary Wintersgill"]}},"tmdate":1718710555194,"pdate":1009843200000,"tcdate":1718710518843,"writers":["~"],"signatures":["~Collin_Lynch1"],"forum":"LTfqFU4LRX","license":"CC BY-SA 4.0","number":35745,"cdate":1009843200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1718710555194,"domain":"DBLP.org","id":"LTfqFU4LRX","version":2},{"content":{"summary":{"value":"The paper presents a scalable quantum control framework utilizing physics-informed reinforcement learning (RL) to address the challenge of optimizing control signals for complex quantum systems. The authors devise an RL algorithm that incorporates physical constraints to restrict the solution space, focusing on desired time scales of quantum state dynamics and realistic control signal limitations.  The method is evaluated on three quantum systems—multi-level Λ systems, Rydberg atoms, and superconducting transmons—demonstrating higher fidelities and robustness to perturbations compared to previous methods."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please see \"Weaknesses\" above."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- This manuscript introduces a novel framework that combines physics-informed constraints with reinforcement learning for quantum control. The use of physics-based constraints not only enhances solution quality but also ensures that the control signals are physically realistic and experimentally feasible.\n- The paper provides a thorough evaluation of the method across different quantum systems and under various conditions, which strengthens the credibility of the results."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The main concern is that I do not consider ICLR the appropriate venue for this paper, as it does not provide the necessary background information for readers unfamiliar with quantum control. It would be more suitable for the paper to be published in a physics journal rather than an AI conference. Even though as a paper for physics journal, it still lacks the self-contained nature and clarity necessary for a comprehensive understanding by the readers. For instance, the experiments presented in this paper appear to be related  to a pre-existing methodology known as 'STIRAP'. However, the manuscript fails to provide adequate explanation or details regarding this method, which could leave readers without a clear understanding of its role and significance in the study. \n\nFurthermore, I am not convinced that the results presented in this paper sufficiently substantiate the authors' claims. Specifically, the authors claim \"the applied constrained not only improves computational efficiency but also promotes the selection of more physically realistic control signals.\" However, the manuscript does not offer adequate numerical or theoretical evidence to validate this assertion."}},"nonreaders":[],"tmdate":1731428657619,"tcdate":1730453211421,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7692/Reviewer_LoSq"],"signatures":["ICLR.cc/2025/Conference/Submission7692/Reviewer_LoSq"],"forum":"YPvI7SofeZ","number":2,"license":"CC BY 4.0","cdate":1730453211421,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7692/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428657619,"domain":"ICLR.cc/2025/Conference","replyto":"YPvI7SofeZ","id":"xw9m9Leb8G","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["reinforcement learning","quantum computing","quantum control","quantum dynamics","control theory"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Quantum optimal control is concerned with the realisation of desired dynamics in quantum systems, serving as a linchpin for advancing quantum technologies and fundamental research. \nAnalytic approaches and standard optimisation algorithms do not yield satisfactory solutions for large quantum systems, and especially not for real world quantum systems which are open and noisy. \nWe devise a physics-informed Reinforcement Learning (RL) algorithm that restricts the space of possible solutions.\nWe incorporate priors about the desired time scales of the quantum state dynamics -- as well as realistic control signal limitations -- as constraints to the RL algorithm. \nThese physics-informed constraints additionally improve computational scalability by facilitating parallel optimisation. \nWe evaluate our method on three broadly relevant quantum systems (multi-level $\\Lambda$ system, Rydberg atom and superconducting transmon) and incorporate real-world complications, arising from dissipation and control signal perturbations. \nWe achieve both higher fidelities -- which exceed 0.999 across all systems --  and better robustness to time-dependent perturbations and experimental imperfections than previous methods. \nLastly, we demonstrate that incorporating multi-step feedback can yield solutions robust even to strong perturbations."},"_bibtex":{"value":"@misc{\nernst2025scaleable,\ntitle={Scaleable Quantum Control via Physics Constrained Reinforcement Learning},\nauthor={Jan Ole Ernst and Tim Franzmeyer and Aniket Chatterjee and Axel Kuhn},\nyear={2025},\nurl={https://openreview.net/forum?id=YPvI7SofeZ}\n}"},"title":{"value":"Scaleable Quantum Control via Physics Constrained Reinforcement Learning"},"pdf":{"value":"/pdf/6a8491f73a929f892cb314e02e6287998a50ab9e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ernst|scaleable_quantum_control_via_physics_constrained_reinforcement_learning"},"authorids":{"value":["~Jan_Ole_Ernst1","~Tim_Franzmeyer1","~Aniket_Chatterjee1","~Axel_Kuhn1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jan Ole Ernst","Tim Franzmeyer","Aniket Chatterjee","Axel Kuhn"]}},"version":2},{"content":{"summary":{"value":"The paper proposes NewtonBench, a benchmark for evaluating large language models (LLMs) on discovering physical laws through interactive experimentation. It introduces metaphysical shifts, i.e. a transformation method applied to canonical physics equations (for example, modifying exponents or couplings), to generate 108 shifted laws across 12 physics domains. Each law is instantiated in three configurations: 1) Vanilla Equation: the shifted law alone, 2) Simple System: the law embedded in a minimal experimental setup,\n3) Complex System: the law coupled with additional equations. Together these form 324 total tasks. The authors also provide a formal proof that all tasks are finitely solvable.\n\nLLMs interact with each environment by adjusting variables and observing outputs to infer the hidden governing law. Performance is evaluated by symbolic accuracy and robustness to noise. Eleven popular LLMs were tested, showing strong performance on simple cases but substantial drops for complex or noisy systems.\n\nMain contributions include:\n1. A large-scale, interactive benchmark for equation discovery.\n2. The metaphysical shift method for creating diverse, solvable tasks.\n3. A comprehensive evaluation of leading LLMs’ capabilities on the proposed benchmark."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper focus on interactive scientific discovery, which is an important and relatively underexplored area.\n2. It proposes an effective method for generating new equations and provides a formal proof of their finite solvability.\n3. The evaluation is thorough, covering a wide range of models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed \"metaphysical shifts\" fail to create a \"physically plausible universe\" and ultimately compromise the scientific relevance of the task. The core problem is that the physics descriptions become meaningless once the dimensional structure and proportional dependencies of core quantities are altered. In the laws altered by NewtonBench (such as Newton’s Law of Universal Gravitation, Coulomb’s Law, Heat Transfer, and Hooke’s Law), most are phenomenological in nature and contain at least one parameter or coefficient whose physical identity is defined by its role within that specific law. Examples include the gravitational constant, the permittivity, the thermal conductivity, or the spring constant. These quantities do not possess independent operational definitions outside the context of the laws that introduce them.\nTherefore, when the functional form of such a law is altered (for instance, by changing exponents or couplings), those parameters lose their original physical meaning, since their identity is inseparable from the structure of the law itself. The modified equations thus describe mathematically consistent systems but no longer preserve the physical semantics of the original quantities or interactions. This reframes the task from a genuine physics discovery challenge into interactive equation discovery within an abstract, non-physical framework.\n\n2. The so-called “model system discovery” in NewtonBench does not introduce a fundamentally different epistemic or physical level of reasoning. It merely increases equation compositionality (the number of nested or coupled expressions), which can always be flattened into a single symbolic equation. In real physics, a model system is defined by time-evolving state variables, interactions governed by symmetries or conservation laws, and differential equations that capture causality and dynamics. In contrast, the relations in NewtonBench reduce to static algebraic mappings between scalar quantities, making its “complex systems” only syntactically more complicated rather than fundamentally different, as they remain within the same layer of symbolic algebra."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919396539,"tcdate":1761600535650,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7266/Reviewer_9pai"],"signatures":["ICLR.cc/2026/Conference/Submission7266/Reviewer_9pai"],"forum":"Gk6umqW74m","number":2,"license":"CC BY 4.0","cdate":1761600535650,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7266/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919396539,"domain":"ICLR.cc/2026/Conference","replyto":"Gk6umqW74m","id":"9c7M648ITm","forumContent":{"TLDR":{"value":"We introduce NewtonBench, a benchmark for evaluating LLMs’ scientific law discovery via interactive, generalizable experiments."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["large language models","benchmark","virtual environment","generalization","agent","scientific law discovery"]},"supplementary_material":{"value":"/attachment/644261b3f64d1a719a04f6b1e6ebeac1a56e1865.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large language models (LLMs) are emerging as powerful tools for scientific law discovery, a foundational challenge in AI-driven science.\nHowever, existing benchmarks for this task suffer from a fundamental methodological trilemma, forcing a trade-off between scientific relevance, scalability, and resistance to memorization. Furthermore, they oversimplify discovery as static function fitting, failing to capture the authentic scientific process of uncovering embedded laws through the interactive exploration of complex model systems. To address these critical gaps, we introduce **NewtonBench**, a benchmark comprising 324 scientific law discovery tasks across 12 physics domains. Our design mitigates the evaluation trilemma by using counterfactual law shifts - systematic alterations of canonical laws - to generate a vast suite of problems that are scalable, scientifically relevant, and memorization-resistant.\nMoreover, we elevate the evaluation from static function fitting to interactive model discovery, requiring agents to experimentally probe simulated complex systems to uncover hidden principles. Our extensive evaluation of 11 state-of-the-art LLMs reveals a clear but fragile capability for discovery in frontier models: this ability degrades precipitously with increasing system complexity and exhibits extreme sensitivity to observational noise. Notably, we uncover a paradoxical effect of tool assistance: providing a code interpreter can hinder more capable models by inducing a premature shift from exploration to exploitation, causing them to satisfice on suboptimal solutions. These results demonstrate that robust, generalizable discovery in complex, interactive environments remains the core challenge for the future of automated science. By providing a scalable, robust, and scientifically authentic testbed, NewtonBench offers a crucial tool for measuring true progress and guiding the development of next-generation AI agents capable of genuine scientific discovery."},"_bibtex":{"value":"@inproceedings{\nzheng2026newtonbench,\ntitle={NewtonBench: Benchmarking Generalizable Scientific Law Discovery in {LLM} Agents},\nauthor={Tianshi Zheng and Kelvin Kiu Wai Tam and Newt Nguyen Kim Hue Nam and Baixuan Xu and Zhaowei Wang and Cheng Jiayang and Hong Ting Tsang and Weiqi Wang and Jiaxin Bai and Tianqing Fang and Yangqiu Song and Ginny Wong and Simon See},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Gk6umqW74m}\n}"},"title":{"value":"NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents"},"pdf":{"value":"/pdf/6375dd7eb877e02f27026fa1429638e0d73abf11.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zheng|newtonbench_benchmarking_generalizable_scientific_law_discovery_in_llm_agents"},"authorids":{"value":["~Tianshi_Zheng1","~Kelvin_Kiu_Wai_Tam1","~Newt_Nguyen_Kim_Hue_Nam1","~Baixuan_Xu1","~Zhaowei_Wang2","~Cheng_Jiayang1","~Hong_Ting_Tsang1","~Weiqi_Wang1","~Jiaxin_Bai1","~Tianqing_Fang1","~Yangqiu_Song1","~Ginny_Wong1","~Simon_See1"]},"authors":{"value":["Tianshi Zheng","Kelvin Kiu Wai Tam","Newt Nguyen Kim Hue Nam","Baixuan Xu","Zhaowei Wang","Cheng Jiayang","Hong Ting Tsang","Weiqi Wang","Jiaxin Bai","Tianqing Fang","Yangqiu Song","Ginny Wong","Simon See"]}},"version":2},{"content":{"summary":{"value":"This paper presents a new model that applies sinc interpolation to the previously proposed KAN model. It discusses the advantages of the proposed model for function fitting and solving PDEs in a Physics-informed manner."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"* What exactly is f(jh) in equation (6)? Is it simply the value of f at jh? Is equation (6) interpolating f(x) using those values? If so, do we need to know those values precisely to approximate f(x) with this formulation?\n* Could other problems with singularities be tested? Does the proposed method also show significant advantages and accurately approximate solutions in all cases where singular phenomena appear? Could it be applied to other boundary layer problems or higher-dimensional problems beyond two dimensions?\n* For line 140, it might be better to use the \\paragraph function for terms like “Convergence theorem” or “On a general interval (a, b)” to save space."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The study explores how the recently developed KAN model can better handle singular problems from the perspective of Physics-Informed Neural Networks (PINNs)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I believe the main goal of this paper is to improve existing PINN-based models through the development of sincKAN. However, as shown in Table 2, it's unclear if sincKAN demonstrates convincingly better performance than other existing models across various PDE problems. Clearly, as mentioned in the PI-KAN model, there are specific advantages and disadvantages to using KAN for solving PDEs, such as computational cost or training speed. If sincKAN does not consistently outperform across all datasets, as seen in Table 2, what advantages does sincKAN offer over using standard MLP or modified MLP to train Physics-informed Loss? While sincKAN appears to perform better on boundary layer problems, as shown in Table 3, the paper’s title and focus could have been more aligned with boundary layer issues if that is its primary strength. Additionally, are there recent studies related to PINNs specifically aimed at solving boundary layer problems? What would happen if epsilon increased to a larger value, such as 10^8? It might also help to include a basic explanation of KAN’s underlying mechanism for readers unfamiliar with the model, as certain concepts could be difficult to grasp without this context."}},"nonreaders":[],"tmdate":1731428289628,"tcdate":1730686382166,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10374/Reviewer_mr4i"],"signatures":["ICLR.cc/2025/Conference/Submission10374/Reviewer_mr4i"],"forum":"ihHeqPLRDk","number":4,"license":"CC BY 4.0","cdate":1730686382166,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10374/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428289628,"domain":"ICLR.cc/2025/Conference","replyto":"ihHeqPLRDk","id":"UAjUPVtGkQ","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed neural networks","Kolmogorov-Arnold Networks","Partial differential equations","Computational physics."]},"supplementary_material":{"value":"/attachment/78589ff6f8ba4e2c8cab881419682eb69f86cf4f.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this paper, we propose to use Sinc interpolation in the context of Kolmogorov-Arnold Networks, neural networks with learnable activation functions, which recently gained attention as alternatives to multilayer perceptron. Many different function representations have already been tried, but we show that Sinc interpolation proposes a viable alternative, since it is known in numerical analysis to represent well both smooth functions and functions with singularities. This is important not only for function approximation but also for the solutions of partial differential equations with physics-informed neural networks. Through a series of experiments, we show that SincKANs provide better results in almost all of the examples we have considered."},"_bibtex":{"value":"@misc{\nyu2025sinc,\ntitle={Sinc Kolmogorov-Arnold Network and Its Applications on Physics-informed Neural Networks},\nauthor={Tianchi Yu and JINGWEI QIU and Jiang Yang and Ivan Oseledets},\nyear={2025},\nurl={https://openreview.net/forum?id=ihHeqPLRDk}\n}"},"title":{"value":"Sinc Kolmogorov-Arnold Network and Its Applications on Physics-informed Neural Networks"},"pdf":{"value":"/pdf/b2625752041c98c9978af6d3f403718dc2e532ba.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yu|sinc_kolmogorovarnold_network_and_its_applications_on_physicsinformed_neural_networks"},"authorids":{"value":["~Tianchi_Yu1","~JINGWEI_QIU1","~Jiang_Yang1","~Ivan_Oseledets1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Tianchi Yu","JINGWEI QIU","Jiang Yang","Ivan Oseledets"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Workshop VerifAI-2"},"keywords":{"value":["model retraining; synthetic data; data verification"]},"abstract":{"value":"Synthetic data has been increasingly used to train frontier generative models. However, recent study raises key concerns that iteratively retraining a generative model on its self-generated synthetic data may keep deteriorating model performance, a phenomenon often coined model collapse. In this paper, we investigate ways to modify the synthetic retraining process to avoid model collapse, and even possibly help reverse the trend from collapse to improvement. Our key finding is that by injecting information through an external synthetic data verifier, whether a human or a better model, synthetic retraining will not cause model collapse. Specifically, we situate our theoretical analysis in the fundamental linear regression setting, showing that verifier-guided retraining can yield near-term improvements but ultimately drives the parameter estimate to the verifier's “knowledge center” in the long run. Our theory further predicts that, unless the verifier is perfectly reliable, these early gains will plateau and may even reverse. Indeed, our experiments across linear regression, Variational Autoencoders (VAEs) trained on MNIST, and SmolLM2-135M on the XSUM task confirm these theoretical insights."},"_bibtex":{"value":"@inproceedings{\nyi2026escaping,\ntitle={Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence},\nauthor={Bingji Yi and Qiyuan Liu and Yuwei Cheng and Haifeng Xu},\nbooktitle={ICLR 2026 Workshop: VerifAI-2: The Second Workshop on AI Verification in the Wild},\nyear={2026},\nurl={https://openreview.net/forum?id=dXH1zC8Goy}\n}"},"title":{"value":"Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"venueid":{"value":"ICLR.cc/2026/Workshop/VerifAI-2"},"paperhash":{"value":"yi|escaping_model_collapse_via_synthetic_data_verification_nearterm_improvements_and_longterm_convergence"},"authorids":{"value":["~Bingji_Yi1","~Qiyuan_Liu4","~Yuwei_Cheng2","~Haifeng_Xu1"]},"Track":{"value":"long paper (up to 8 pages)"},"authors":{"value":["Bingji Yi","Qiyuan Liu","Yuwei Cheng","Haifeng Xu"]}},"tmdate":1773260207793,"pdate":1772424057734,"tcdate":1769977117618,"writers":["ICLR.cc/2026/Workshop/VerifAI-2","ICLR.cc/2026/Workshop/VerifAI-2/Submission10/Authors"],"signatures":["ICLR.cc/2026/Workshop/VerifAI-2/Submission10/Authors"],"forum":"dXH1zC8Goy","license":"CC BY 4.0","number":10,"cdate":1769977117618,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/VerifAI-2/-/Submission","ICLR.cc/2026/Workshop/VerifAI-2/-/Post_Submission","ICLR.cc/2026/Workshop/VerifAI-2/-/Edit","ICLR.cc/2026/Workshop/VerifAI-2/Submission10/-/Camera_Ready"],"mdate":1773260207793,"odate":1772718030066,"domain":"ICLR.cc/2026/Workshop/VerifAI-2","id":"dXH1zC8Goy","version":2},{"content":{"summary":{"value":"This paper proposes STAN (Spatio-Temporal Attention Network), a reinforcement learning-based policy network combined with attention mechanisms for continuous low-thrust spacecraft collision avoidance in complex, multi-debris orbital environments. The work addresses key limitations in prior methods, including the inability to handle variable numbers of debris, and insufficient integration of physics-informed risk metrics into learning-based control."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"The authors can address the experimental limitaions concerns on choosing shorter time frame for experiments, why didn't consider the smaller pertubations in the orbital dynamics, circular initialization of debris orbits and sensor noise in the modelling scenrios."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"ST-Attention design combines self-attention with physically-informed bias to prioritize debris threats effectively. It addresses the gap in standard attention mechanisms that lack domain knowledge.\n\nArchitecture - Strengths:\n1. Combines learned and domain knowledge through fusing of self-attention for complex interactions with explicit physics-informed bias features, capturing both debris correlations and collision risks.\n\n2. Scalable to arbitrary debris counts through encoding features and using attention allows the model to handle variable numbers of debris objects and multi-threat scenarios.\n\n3. Learnable weighting setup let the network adjust the importance of physical features relative to learned embeddings, improving adaptability to different orbital contexts or debris densities."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Architecture - Limitations:\n\n1. Physics-aware attention bias incorporates the distance of closest approach (DCA) and time to closest approach (TCA) as a learnable bias, the model ensures attention is grounded in domain-relevant risk indicators, not just learned embeddings. Model may be more biased towards immediate critical threats than long-range operational settings such as mission fuel consumption.\n\n2. Uniform broadcast of bias across all columns to form 𝑁×𝑁. This assumes the physical importance of debris i affects all pairwise interactions equally, which may ignore interaction-specific relationships between debris pairs (e.g., cross-collision influence).\n\n3. Mean pooling across debris dimension reduces to a single vector 𝑓 and  both heads rely on it. May lose fine-grained per-debris information, limiting the decoder’s ability to make nuanced, individual maneuvers for specific threats.\n\n4. Reward penalizes collision probability but does not explicitly enforce safety constraints. Policy may occasionally select risky maneuvers if they increase cumulative reward. Does not explicitly account for sensor noise or uncertainty in debris position, which could make the reward function misleading in real operational settings.\n\n5. Some terms may overlap in effect, e.g., minimum distance softness (p_s) and collision probability (p_c) are related; summing both may overweight certain safety aspects.\n\nExperiments - Limitations:\n\n1. Current approach considers Two-body orbital dynamics only and it neglects higher-order perturbations such as J2 (Earth oblateness), atmospheric drag, solar radiation pressure, and third-body effects (Moon, Sun).\n\n2. Multistage collisions are simulated progressively, but each scenario cover only a short timeframe (e.g., 1.2 hours in experiments). The simulation is for 1.2-hour mission window and does not test long-term maneuver planning, cumulative collision risk, or fuel optimization over multiple orbital periods. Policy generalization to extended missions or sequential conjunctions is unclear.\n\n3. Spacecraft near-circular orbit, debris initialized along same orbit. Unrealistic debris distribution case, the real debris is in multiple orbital planes, inclinations, and eccentricities.\n\n4. Exact positions and velocities of debris are used. Ignores sensor noise, tracking errors, and uncertainty in debris catalog."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920420154,"tcdate":1761838171746,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8566/Reviewer_eUN8"],"signatures":["ICLR.cc/2026/Conference/Submission8566/Reviewer_eUN8"],"forum":"aZs6DkGM2I","number":2,"license":"CC BY 4.0","cdate":1761838171746,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8566/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920420154,"domain":"ICLR.cc/2026/Conference","replyto":"aZs6DkGM2I","id":"XQmmHYe75b","forumContent":{"TLDR":{"value":"We propose STAN, a spatio-temporal attention network for real-time low-thrust avoidance of multistage space debris collisions using deep reinforcement learning."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Space Debris Collision Avoidance","Deep Reinforcement Learning","Spatio-Temporal Attention","Continuous Low-Thrust Control","Policy Network Design"]},"supplementary_material":{"value":"/attachment/1a3a7b0f9fbd5927ca9f72aeab75a707522e3ea9.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"The rapid expansion of space missions has led to an exponential increase in space debris, posing severe threats to spacecraft. Existing approaches struggle to handle multistage collision risks in cluttered orbital environments, and the use of continuous low-thrust propulsion further complicates avoidance planning. To address these challenges, we propose the Spatio-Temporal Attention Network (**STAN**), which employs novel Spatio-Temporal Attention (**ST-Attention**) layers in place of conventional attention mechanisms. STAN encodes satellite-debris pairs and integrates time and distance into attention weight computation, enabling the model to generate context-aware low-thrust maneuvers. The model is trained using deep reinforcement learning across four representative multistage collision scenarios, jointly optimizing collision probability, fuel consumption, and orbital deviation. Experimental results show that STAN outperforms baseline methods in safety performance, fuel efficiency, and orbit preservation."},"_bibtex":{"value":"@misc{\nyang2026stan,\ntitle={{STAN}: A Spatio-Temporal Attention Network for Space Debris Multistage Collision Avoidance},\nauthor={Liwen Yang and Xue Bai and Xiaoyi Wang and Ming Xu},\nyear={2026},\nurl={https://openreview.net/forum?id=aZs6DkGM2I}\n}"},"title":{"value":"STAN: A Spatio-Temporal Attention Network for Space Debris Multistage Collision Avoidance"},"pdf":{"value":"/pdf/706d4e60cd5e28f3485ea7f7a132a3dc192f4593.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yang|stan_a_spatiotemporal_attention_network_for_space_debris_multistage_collision_avoidance"},"authorids":{"value":["~Liwen_Yang1","~Xue_Bai3","~Xiaoyi_Wang10","~Ming_Xu14"]},"authors":{"value":["Liwen Yang","Xue Bai","Xiaoyi Wang","Ming Xu"]}},"version":2},{"content":{"summary":{"value":"This work presents a system named “Physics-Integrated Gaussian Splatting,” aiming to synthesize \ncontrollable and physically inspired fire and combustion effects in real-world 3D scenes.  The method integrates several coordinated modules to form a complete pipeline from real-scene  reconstruction to fire simulation:\n 1) a multimodal large model predicts material types and combustion-related parameters in 2D projection \nspace and back-project them to 3D Gaussian\n 2) the 3D scene is voxelized, where solid and air regions are distinguished by Gaussian density—simplified\n fire simulations are applied to air voxels, while thermal diffusion and charring are modeled for solids, with \nintuitive user controls such as ignition position and airflow\n 3) a unified rendering framework generates multiple visual effects, including flames, smoke, charring, and \nindirect illumination.\n \nThe work does not introduce a new rendering theory or physical model but integrates existing techniques\n 3D Gaussian Splatting, simplified fluid simulation, and multimodal reasoning—into a complete interactive\n pipeline, producing visually realistic fire results.\n\n Experiments on multiple real and synthetic scenes demonstrate visually realistic and controllable fire \ngeneration, and the results appear to outperform existing generation-based approaches"},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- How does the system perform on uncommon materials, composite materials, or objects with unusual textures?\n\n- Could the authors clarify or quantify how much these physical simplifications affect the realism of the generated fire?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"This paper presents an integrated system combining 3D Gaussian Splatting, simplified physical simulation, and multimodal reasoning to construct a physically-informed and controllable fire generation pipeline. \n\nThis integration and simplification provide a certain degree of novelty in application, making complex fire simulation more practical and easier to operate. The system is well-designed, with modules for material prediction, fire simulation, and rendering fire results. Experiments on multiple real and synthetic scenes demonstrate that the generated fire is visually plausible and controllable, outperforming existing generation-based methods. The user interaction design further enhances the system’s operability. \n\nThe paper provides clear and understandable descriptions of the system modules  and workflow, and the authors indicate that the source code will be released in the future, which facilitates follow-up \nresearch. Overall, this work offers a complete and practical solution for interactive fire generation, with utility for computer graphics, visual effects, and virtual environment fire simulation"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the integration offers practical value, the method has limited theoretical and technical novelty.  Moreover, the approach heavily depends on the MLLM’s 2D material inference capability, and its performance on uncommon materials, composite materials, or extreme fire conditions remains unexplored. \nAdditionally, the paper does not provide comparisons between the simplified physical simulation and a full \nphysics-based simulation, making it difficult to justify the acceptability of the simplifications in practice. Finally\n While FieryGS accounts for multiple effects of fire combustion, it does not explicitly capture or evaluate the \ndynamic lighting effects generated by fire, which contribute to the perceived motion and liveliness of the \nscen"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921144747,"tcdate":1762170236059,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9601/Reviewer_EZSh"],"signatures":["ICLR.cc/2026/Conference/Submission9601/Reviewer_EZSh"],"forum":"ziKFH7whvy","number":4,"license":"CC BY 4.0","cdate":1762170236059,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9601/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921144747,"domain":"ICLR.cc/2026/Conference","replyto":"ziKFH7whvy","id":"gDGQKRPIiU","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["3D Gaussian Splatting","Physics Simulation","Combustion Simulation","Novel View Synthesis"]},"supplementary_material":{"value":"/attachment/7a397bad1665e2554d7033b8bf2d44fb2f67df67.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We consider the problem of synthesizing photorealistic, physically plausible combustion effects in in-the-wild 3D scenes. Traditional CFD and graphics pipelines can produce realistic fire effects but rely on handcrafted geometry, expert-tuned parameters, and labor-intensive workflows, limiting their scalability to the real world. Recent scene modeling advances like 3D Gaussian Splatting (3DGS) enable high-fidelity real-world scene reconstruction, yet lack physical grounding for combustion. To bridge this gap, we propose FieryGS, a physically-based framework that integrates physically-accurate and user-controllable combustion simulation and rendering within the 3DGS pipeline, enabling realistic fire synthesis for real scenes. Our approach tightly couples three key modules: (1) multimodal large-language-model-based physical material reasoning, (2) efficient volumetric combustion simulation, and (3) a unified renderer for fire and 3DGS. By unifying reconstruction, physical reasoning, simulation, and rendering, FieryGS removes manual tuning and automatically generates realistic, controllable fire dynamics consistent with scene geometry and materials. Our framework supports complex combustion phenomena—including flame propagation, smoke dispersion, and surface carbonization—with precise user control over fire intensity, airflow, ignition location and other combustion parameters. Evaluated on diverse indoor and outdoor scenes, FieryGS outperforms all comparative baselines in visual realism, physical fidelity, and controllability."},"_bibtex":{"value":"@inproceedings{\nshen2026fierygs,\ntitle={Fiery{GS}: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting},\nauthor={Qianfan Shen and Ningxiao Tao and Qiyu Dai and Tianle Chen and Minghan Qin and Yongjie Zhang and Mengyu Chu and Wenzheng Chen and Baoquan Chen},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ziKFH7whvy}\n}"},"title":{"value":"FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting"},"pdf":{"value":"/pdf/4dc4a8361c1a8896d2fba2be944cba84986d0a9d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"shen|fierygs_inthewild_fire_synthesis_with_physicsintegrated_gaussian_splatting"},"authorids":{"value":["~Qianfan_Shen1","~Ningxiao_Tao1","~Qiyu_Dai1","~Tianle_Chen4","~Minghan_Qin1","~Yongjie_Zhang3","~Mengyu_Chu3","~Wenzheng_Chen1","~Baoquan_Chen1"]},"authors":{"value":["Qianfan Shen","Ningxiao Tao","Qiyu Dai","Tianle Chen","Minghan Qin","Yongjie Zhang","Mengyu Chu","Wenzheng Chen","Baoquan Chen"]}},"version":2},{"content":{"TLDR":{"value":"A violation-of-expectation surprise score is not physics evidence on its own. I decompose it on V-JEPA 2 and V-JEPA 2.1, move it with a knob that carries no physics, and specify the six quantities that must be reported beside it."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["intuitive physics","violation of expectation","evaluation methodology","video world models","latent-predictive models","V-JEPA","representation diagnostics","measurement validity"]},"supplementary_material":{"value":"/attachment/b24e5f04b1738f0a1f0e287b002d63e3f1f48d33.zip"},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Latent-predictive video models are said to acquire intuitive physics when they are more surprised by an impossible event than by a possible one (Garrido et al., 2025). That surprise is a single scalar, the difference in prediction error between two otherwise identical clips. The benchmark that defines that score chose the measure deliberately, naming one alternative protocol (Weihs et al., 2022) and reporting that the alternative had not yielded evidence of understanding (Bordes et al., 2025). Nothing has established what the difference contains. We decompose the squared form exactly into magnitude, direction and a selection term, and show it is not a fixed property of the model. We measure it on V-JEPA 2 (Assran et al., 2025) and V-JEPA 2.1 (Mur-Labadia et al., 2026). Both are publicly released latent-predictive encoders, and only V-JEPA 2.1 exposes intermediate depths to read at. Reading V-JEPA 2.1 at the four encoder depths its predictor was trained to read, with nothing attached on top, the reported accuracy spans 10.18 percentage points with the model, the clips and the readout all unchanged. Across the two encoders every term keeps its sign and none keeps its size, so the two standard normalisations disagree about which model is more direction-driven. A motion-only observer touching neither model shows that most or more of what the score detects is available from movement alone. We then move the score using a regulariser carrying no physics. An untrained random projection preserves accuracy at 88.39% against the frozen encoder's 89.55%. Training a small network in its place, under the coefficients published with VICReg (Bardes et al., 2022), drives it to 46.57%, a point estimate below chance whose 95% confidence interval contains chance. Raising one of those three coefficients recovers 94.62%. A violation-of-expectation score is not readable as physics evidence until six quantities are reported beside it. We specify those six and show what each one catches."},"_bibtex":{"value":"@inproceedings{\nanonymous2026does,\ntitle={Does Surprise Measure Intuitive Physics? The Score Depends on How the Model Is Read},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KMwOmfu3bA},\nnote={under review}\n}"},"title":{"value":"Does Surprise Measure Intuitive Physics? The Score Depends on How the Model Is Read"},"pdf":{"value":"/pdf/541ad33a35104d894d92c0deb1439702a27ba2a4.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791234583190,"tcdate":1789782545957,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission51807/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission51807/Authors"],"forum":"KMwOmfu3bA","license":"CC BY 4.0","number":51807,"cdate":1789782545957,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission51807/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791234583190,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"KMwOmfu3bA","version":2},{"content":{"summary":{"value":"The paper proposes a deletion‑based framework to probe how much LLMs depend on their CoT when solving physics problems. By systematically removing portions of the generated reasoning and measuring changes in answer accuracy, final answer length, and information overlap, the authors study three open‑source models (Magistral, Phi‑4 and Qwen‑A3B) on three physics benchmarks. Experiments reveal that explicit reasoning prompts improve performance but the CoT can be removed without dramatically hurting accuracy, as models \"cram\" reconstructed steps into the final answer. They conclude that current accuracy‑only evaluations are insufficient and calls for metrics that assess the faithfulness of reasoning."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. How does the deletion framework differ from or extend prior CoT‑evaluation methods (e.g., perturbation-based evaluations)? What is novel beyond applying it to physics tasks.\n\n2. Is this framework suitable for other domains, like mathematics or commonsense reasoning?\n\n3. Can you provide more analysis on what kinds of information models \"cram\" into the final answer when reasoning is removed? Are they recalling memorized formulas or recomputing reasoning?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The work tackles an important question about whether CoT explanations genuinely reflect model reasoning, which is crucial for using LLMs in scientific domains.\n\n- The deletion strategy is clearly described and measures multiple downstream effects, such as accuracy, answer length, lexical and frequency overlap. This provides a structured way to examine reliance on intermediate reasoning.\n\n- The experiments cover three different benchmarks of physics domain and multiple LLMs, the authors explore effects of prompt explicitness and different deletion strategies, with the analysis is carefully presented and supported by figures."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Prior research has already highlighted the gap between answer accuracy and CoT faithfulness and proposed evaluation frameworks.  For instance, Nguyen et al. [1] introduce discriminative and generative evaluations that showed LLMs may reach correct answers through incorrect reasoning, and Barez et al. [2] argue that CoT is not, by itself, an adequate explanation. The deletion framework is a more like a straightforward application of such idea to physics domain and does not reveal its novelty or specifity.\n\n- The analysis and experimental obversations are surface‑level, lacking of in-depth exploration of LLMs' internal activation or behavioural pattern. Thus, it cannot wwell explain why models can reconstruct missing steps or whether they use memorised templates versus genuine reasoning.\n\n- The experiments consider end‑of‑scratchpad truncation, random deletion and removal of annotated physics tokens. More nuanced manipulations, such as deleting specific reasoning types, shuffling steps, may yield deeper insight into what information is truly required.\n\n[1] Nguyen et al., Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs, 2024.\n\n[2] Barez et al., Chain-of-Thought Is Not Explainability, 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917423076,"tcdate":1761979855030,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4526/Reviewer_iLnu"],"signatures":["ICLR.cc/2026/Conference/Submission4526/Reviewer_iLnu"],"forum":"GiItKTlJIB","number":4,"license":"CC BY 4.0","cdate":1761979855030,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4526/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917423076,"domain":"ICLR.cc/2026/Conference","replyto":"GiItKTlJIB","id":"1ZfO851Ikv","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"LLMs can solve physics problems by patching gaps in heavily deleted CoT reasoning traces, but without true faithfulness."},"keywords":{"value":["chain-of-thought","reasoning","evaluation"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Reasoning-focused language models are increasingly applied to AI for science, but evaluation has not kept pace: benchmarks largely measure end-task accuracy while ignoring whether models genuinely depend on their own reasoning traces. This gap is critical in domains like physics problem solving, where equations, units, and structured terminology make reasoning reliability both essential and testable. We introduce a systematic deletion framework that intercepts chain-of-thought (CoT) mid-generation, removes tokens, and measures downstream effects. Applied to three open-source models—Magistral, Phi-4, and Qwen-A3B—across multiple physics benchmarks, our method shows that models remain accurate under heavy deletions (40–60\\%) by “cramming” reconstructed steps into final answers. Overlap analyses reveal that deleted equations and facts often reappear, but inconsistently across strategies, exposing shallow and opportunistic reliance on CoT. These findings underscore that current accuracy-based evaluations are insufficient for scientific domains, and point toward the need for methods that assess reasoning faithfulness as a core requirement for advancing AI for science."},"_bibtex":{"value":"@misc{\nhawthorne2025how,\ntitle={How Much Chain-of-Thought Do {LLM}s Really Need for Physics?},\nauthor={Anton Gonzalvez Hawthorne and Daanish Shabbir and Isabelle Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=GiItKTlJIB}\n}"},"title":{"value":"How Much Chain-of-Thought Do LLMs Really Need for Physics?"},"pdf":{"value":"/pdf/5b25875679498c6ba2d54f32ef752eebe9c53647.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hawthorne|how_much_chainofthought_do_llms_really_need_for_physics"},"authorids":{"value":["~Anton_Gonzalvez_Hawthorne1","~Daanish_Shabbir1","~Isabelle_Lee1"]},"authors":{"value":["Anton Gonzalvez Hawthorne","Daanish Shabbir","Isabelle Lee"]}},"version":2},{"content":{"summary":{"value":"The paper introduces CrystalSeg, a physics-guided, GPU-accelerated pipeline that (i) programmatically generates CAD-based 3D crystal/loop/liquor scenes, (ii) renders phase-contrast projections via multislice wave propagation under the Fresnel approximation with a pragmatic detector model (Gaussian PSF blur, Poisson shot noise, Gaussian read noise, and column-wise gain to induce ring artifacts), (iii) reconstructs volumes by FBP, and (iv) trains nnU-Net to segment crystal, mother liquor, and mounting loop. Training on Real+Syn outperforms OnlyReal and OnlySyn. Physics-guided simulation also matches real projections/reconstructions better than absorption-only and a diffusion “style-transfer” tool. The authors position this as enabling fully automated ray-tracing absorption correction, markedly reducing manual effort and addressing a practical bottleneck in long-wavelength crystallography."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"· Generalization: How does performance vary across beamlines/energies, detector PSFs, and significantly different crystal morphologies? Any cross-site test sets?\n· Synthetic–real gap: Do you measure distribution shift between synthetic and real reconstructions (e.g., edge statistics or learned feature distance), and how does it correlate with segmentation error?\n· Simulator ablation: Can you quantify segmentation gains from (i) phase contrast vs. absorption-only, (ii) PSF/noise, and (iii) ring-artifact modeling?\n· Baselines: How does Real+Syn compare with (i) Real-only + heavy augmentation and Real-only + self-training to isolate the marginal benefit of simulation, and (ii) competitive 3D segmentation baselines."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Long-wavelength crystallography needs accurate per-voxel segmentation for absorption correction, manual work is a bottleneck. The work is well-scoped to this pain point and ties to validation.\n· Physically grounded simulator that models phase contrast, δ/β materials, detector response, and common artifacts, yielding more realistic projections/reconstructions than absorption-only surrogates.\n· Physics-guided projections/recons exceed absorption-only in SSIM/PSNR for two samples, and synthetic data substantially boosts segmentation when mixed with real data.\n· Include task-level validation, demonstrated downstream utility for automated absorption correction, indicating impact beyond proxy segmentation metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"· Training uses synthetic datasets with only 5 real test sets. Beamlines, energy and sample-morphology diversity are not systematically evaluated. Generalization beyond the reported setting is unclear.\n· Results hinge on nnU-Net only, no comparisons to other strong 3D segmentation baselines, nor to real-only + heavy augmentation / self-training controls.\n· There is no ablation study for phase-contrast, PSF/noise, or ring-artifact modeling. The “Nano Banana” comparison is illustrative but not a physics baseline, absorption-only is the only true ablation.\n· The contribution asserts the “first fully automated” solution for absorption correction; evidence is compelling for two case studies but lacks tests across varied samples/beamlines.\n· Downstream crystallographic validation is limited, only two case studies are shown.\n· The paper states variation in refractive indices and randomization, but the sampled ranges and their match to beamline conditions are not tabulated. Ring-artifact modeling via column gain may not capture diverse ring etiologies.\n· The manuscript claims hours to seconds reduction, but does not report generation/training/inference timings on stated hardware.\n· Minor wording issue, contributions list says “ray-racing” instead of “ray-tracing.”"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921361289,"tcdate":1761672286274,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9902/Reviewer_d2Em"],"signatures":["ICLR.cc/2026/Conference/Submission9902/Reviewer_d2Em"],"forum":"aK7knflHF6","number":2,"license":"CC BY 4.0","cdate":1761672286274,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9902/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921361289,"domain":"ICLR.cc/2026/Conference","replyto":"aK7knflHF6","id":"OmS8RqeGIP","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["AI4S","segmentation","synthetic dataset","tomography_reconstruction"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Automated 3D segmentation of tomographic volumes is a critical bottleneck in long-wavelength X-ray crystallography, a technique crucial for drug development and validating structural models from systems like AlphaFold3. This segmentation is a prerequisite for ray-tracing absorption correction, which is necessary for data processing in X-ray crystallography experiments. However, it is currently performed manually by experts, which is a process that is slow, costly, and prevents full automation of the scientific pipeline. The primary barrier to automation is the prohibitive expense and difficulty of collecting annotated segmentation data.\nTo address this data scarcity problem, we present **CrystalSeg**, a novel, GPU-accelerated simulation and segmentation pipeline. It generates vast amounts of annotated data by simulating synchrotron X-ray tomography images and their corresponding reconstructed 3D volumes. We demonstrate that segmentation networks trained on CrystalSeg's synthetic data achieve dramatic performance gains over models trained on limited real data, with **improvements of 29.2\\% in Recall, 30.5\\% in IoU, and 24.9\\% in F1 score** for finding the crystal.\nCrystalSeg effectively reduces the expert labor required for segmentation from hours to minutes. More importantly, it enables, for the first time, a fully automated solution for ray-tracing absorption correction in long-wavelength crystallography, making this advanced structural biology technique more scalable and accessible."},"_bibtex":{"value":"@misc{\nlu2025crystalseg,\ntitle={CrystalSeg: Automating Synchrotron Tomographic Reconstruction Segmentation for Crystallography with Physically Guided Simulations},\nauthor={Yishun Lu and Shuyang Sun and Ramona Duman and Armin Wagner and Philip Torr and Wesley Armour},\nyear={2025},\nurl={https://openreview.net/forum?id=aK7knflHF6}\n}"},"title":{"value":"CrystalSeg: Automating Synchrotron Tomographic Reconstruction Segmentation for Crystallography with Physically Guided Simulations"},"pdf":{"value":"/pdf/b1c5058d0adf5f692961b1be5faeab410bd39dad.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"lu|crystalseg_automating_synchrotron_tomographic_reconstruction_segmentation_for_crystallography_with_physically_guided_simulations"},"authorids":{"value":["~Yishun_Lu1","~Shuyang_Sun1","~Ramona_Duman1","~Armin_Wagner1","~Philip_Torr1","~Wesley_Armour1"]},"authors":{"value":["Yishun Lu","Shuyang Sun","Ramona Duman","Armin Wagner","Philip Torr","Wesley Armour"]}},"version":2},{"content":{"summary":{"value":"This paper tries to use a series of transformer models to fit the hidden physics rules in a dynamic reconstruction framework from given observed multi-view videos. The main focus of their pipeline is to model how the implicit forces are propogated between particles in different scale. By achieving this, they show good future prediction performance on three datasets."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. Do you include the background in your experiments? Only foreground objects are shown in the qualitative results. \n2. The visual metrics, such as PSNRs, SSIMs, LPIPSs, in Spring-Gaus synthetic dataset is extremely low. Even the best 14.080 PSNR is far worse than expectation, can the authors explain the too low PSNRs? And the scores for spring-gaus is too low compared to its original paper. Can the authors explain this?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. The paper has a detailed explanation in implementation. \n2. Thorough ablation studies on each modules, giving readers’ an in-detail understanding about the ability for each modules. \n3. The research problem is important, because enabling models to understand physics is beneficial for many downstream tasks such as robotics and the world model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although this paper achieves great performance in given datasets compared to the baselines, there are strong weaknesses of the proposed methods. \n1. Unlike spring-gs and pac-nerf fitting parameters for simulation pipelines, this paper uses transformers to implicitly learn the propagation rule of forces. However, this learned propagation rule is only fitted per-scene-wise, which means it could only overfit to observed forces. How the model can be generalized to novel environment setting and unobserved external forces propagation is NOT evaluated and shown. Although the future prediction to some extent requires this ability, the future parts in the dataset is quite simple, and the configurations of the objects are not `new' for the network, it is still not clear whether the model really learns physics. Therefore, in order to really support the arguments that the model learns the physics, experiments about resimulation with different initial object configurations and differnet environment settings are important. \n2. The baselines and related works are simple. There are missing future prediction methods to discuss, such as FreeGave (Li et al, CVPR2025), GaussianPrediction (Zhao et al, SIGGRAPH2024). As for the baselines to directly compare with, GIC (Cai et al, NeurIPS 2024 oral) should be included. \n3. The presentation is too tedious. There are too many efforts put in the implementation part (the definition for every single mlps and transformers), making it extremely hard for people to understand what is the main purpose. After reading all the equations with a great effort, I finally find most of the equations are not important. I strongly request the authors to revise the method part. \n\nI’m open to increase my score if the authors can address my concerns above."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920286728,"tcdate":1761877478596,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8380/Reviewer_8NkT"],"signatures":["ICLR.cc/2026/Conference/Submission8380/Reviewer_8NkT"],"forum":"fclyfPF3fJ","number":4,"license":"CC BY 4.0","cdate":1761877478596,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8380/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920286728,"domain":"ICLR.cc/2026/Conference","replyto":"fclyfPF3fJ","id":"HTaFsR7Fj2","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Modeling Object Dynamics","Spatial Completion","Temporal Aggregation","Particle Graph Transformer"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Increasing interaction demands with dynamic objects require accurate modeling of their dynamics and precise prediction of motion trajectories from limited observations. Existing approaches rely on the coordinates of downsampled Key Points as the feature basis and model their interactions within local neighborhoods, resulting in the loss of fine-grained details and homogenized particle representations. In this work, we propose DyG$^2$T, a dynamics modeling framework that leverages spatiotemporally completed particle representations for multi-scale force propagation. Spatially, each Key Point enriches fine-grained edge features and spatial geometry by aggregating position information from corresponding raw particles and relative coordinates from neighboring Key Points. Temporally, after supplementing Key Points with inter-frame relative motion offsets via Motion Align Net, the Temporal Attention is applied to aggregate Key Point features across adjacent frames, preserving the dynamic evolution patterns of particles. For comprehensive interactive modeling, a Particle Graph Transformer establishes multi-scale force propagation paths from contact-near to distant Key Points, preserving discriminative long-range dependencies critical for accurate trajectory modeling. Experiments on synthetic and real-world datasets demonstrate that DyG$^2$T achieves accurate trajectory decoding, strong cross-object and real-world generalization."},"_bibtex":{"value":"@misc{\nwang2025dygt,\ntitle={DyG\\${\\textasciicircum}2\\$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer},\nauthor={Yansong Wang and Zhaobo Qi and Yundong Sun and Xinyan Liu and Beichen Zhang and Shuhui Wang and Weigang Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=fclyfPF3fJ}\n}"},"title":{"value":"DyG$^2$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer"},"pdf":{"value":"/pdf/5e0477efed40487b6c5eef4e2bcf50004862e44c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|dyg^2t_modeling_object_dynamics_with_3d_gaussian_temporalspatial_particle_graph_transformer"},"authorids":{"value":["~Yansong_Wang5","~Zhaobo_Qi1","~Yundong_Sun1","~Xinyan_Liu1","~Beichen_Zhang2","~Shuhui_Wang1","~Weigang_Zhang1"]},"authors":{"value":["Yansong Wang","Zhaobo Qi","Yundong Sun","Xinyan Liu","Beichen Zhang","Shuhui Wang","Weigang Zhang"]}},"version":2},{"content":{"summary":{"value":"Summary: The paper presents a new approach to learning to simulate physical systems using Graph Neural Networks (GNNs). Traditional GNNs rely on fixed, manually designed hierarchies / meshes, which fail to adapt to the evolving dynamics in physical simulations. The authors propose the Dynamic Hierarchical Message Passing (DHMP) model, which introduces dynamic, context-aware and data-driven hierarchies.\n\nKey innovations of DHMP include:\n* Anisotropic Message Passing: Facilitates direction-specific message propagation, allowing better representation of physical processes.\n* Differentiable Node Selection: This component allows for learning adaptable, multi-scale graph structures that evolve over time.\n\nDHMP outperforms existing methods, achieving an average of 22.7% improvement in five classic physics simulation datasets. It effectively models both local and long-range dependencies in time-varying, mesh-based systems."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I would like the authors to address the potential limitations that I have listed."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The Dynamic Hierarchical Message Passing (DHMP) model adapts its graph structure dynamically, effectively captures long-range dependencies and handles unseen mesh structures, making it a strong solution for complex physics simulations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Some potential limitations:\n\n* Novelty of multi-scale graph neural networks by differentiable node selection: This work \"Multiresolution equivariant graph variational autoencoder\" by Truong Son Hy and Risi Kondor (https://iopscience.iop.org/article/10.1088/2632-2153/acc0d8) has already proposed a similar idea using Gumbel-Softmax for node sampling to construct an adaptive hierarchy.\n\n* Increased Complexity: Theoretically, the dynamic adaptation of hierarchies and anisotropic message passing introduces additional computational overhead, making it more complex and potentially slower than fixed-hierarchy models. Could you please analyse the time complexity and the space complexity of your model and compare with other baselines? In the Appendix, Table 13 includes comparison with other baselines in terms of training cost, inference time and number of parameters that suggest the computational overhead of this work is not significant. \n\n* Stability of Differentiable Node Selection (DiffSELECT): This learning mechanism might face instability challenges during training, especially in highly dynamic or chaotic systems, which could lead to less reliable performance in certain scenarios. The key component / function is the Gumbel-Softmax in the node sampling / selection (see Equation 6 in Section 3.3). However, the Gumbel-Softmax is sensitive with its temperature hyper-parameter \\tau (see PyTorch instruction: https://pytorch.org/docs/stable/generated/torch.nn.functional.gumbel_softmax.html). How can you select the temperature hyper-parameter? Is it the same for every scenario?\n\nIt would be great if the authors can try on some turbulence datasets to showcase the stability of method.\n\n* Limitation in Generalization: While DHMP performs well on some specific physics simulation datasets, its generalization to other domains or non-mesh-based applications may require further modification or tuning, limiting its broader applicability. Do you have any plan to apply your model into other domains or non-mesh-based applications? \n\nI suggest the authors to check PDEBench benchmark: https://arxiv.org/abs/2210.07182"}},"nonreaders":[],"tmdate":1731427988521,"tcdate":1730145362085,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5636/Reviewer_rDem"],"signatures":["ICLR.cc/2025/Conference/Submission5636/Reviewer_rDem"],"forum":"r8t6OsLP2s","number":2,"license":"CC BY 4.0","cdate":1730145362085,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5636/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427988521,"domain":"ICLR.cc/2025/Conference","replyto":"r8t6OsLP2s","id":"C0WxpI0Xhf","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics Simulation","Message Passing Networks"]},"supplementary_material":{"value":"/attachment/84cbd23b6d9c15f4dcc187e1bb93d339014d5e15.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Graph neural networks have emerged as a powerful tool for large-scale mesh-based physics simulation. Existing approaches primarily employ hierarchical, multi-scale message passing to capture long-range dependencies within the graph. However, these graph hierarchies are typically fixed and manually designed, which do not adapt to the evolving dynamics present in complex physical systems. In this paper, we introduce a novel neural network named DHMP, which learns **D**ynamic **H**ierarchies for **M**essage **P**assing networks through a differentiable node selection method. The key component is the *anisotropic* message passing mechanism, which operates at both intra-level and inter-level interactions. Unlike existing methods, it first supports directionally non-uniform aggregation of dynamic features between adjacent nodes within each graph hierarchy. Second, it determines node selection probabilities for the next hierarchy according to different physical contexts, thereby creating more flexible message shortcuts for learning remote node relations. Our experiments demonstrate the effectiveness of DHMP, achieving $22.7$\\% improvement on average compared to recent fixed-hierarchy message passing networks across five classic physics simulation datasets."},"_bibtex":{"value":"@misc{\ndeng2025discovering,\ntitle={Discovering Message Passing Hierarchies for Mesh-Based Physics Simulation},\nauthor={Huayu Deng and Xiangming Zhu and Yunbo Wang and Xiaokang Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=r8t6OsLP2s}\n}"},"title":{"value":"Discovering Message Passing Hierarchies for Mesh-Based Physics Simulation"},"pdf":{"value":"/pdf/4358239cefc024c026977740a2ac39236803d519.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"deng|discovering_message_passing_hierarchies_for_meshbased_physics_simulation"},"authorids":{"value":["~Huayu_Deng1","~Xiangming_Zhu2","~Yunbo_Wang2","~Xiaokang_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Huayu Deng","Xiangming Zhu","Yunbo Wang","Xiaokang Yang"]}},"version":2},{"content":{"summary":{"value":"This paper discusses the compromise in gated linear RNNs, which remove the dependency of the gate projection on the previous hidden state to maintain linear complexity. While it's intuitive that removing this dependency leads to a performance drop, the paper provides a new perspective by analyzing the distribution shift of gate values. The authors uncover that after removing this dependency, the first layer of the model shows a significant distribution shift in its gate values, while subsequent layers exhibit a more moderate shift. Building on this, the paper introduces a trick that applies Gumbel-Softmax initialization to the first layer, forcing the model to output gate values close to either 0 or 1, thus enhancing performance. The experiments verify that this trick is effective for the synthetic recall tasks designed by the authors. However, for more complex language tasks, the trick is not as effective."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. The method excels on copying tasks but shows almost no improvement on WikiText-103 language modeling . Does this imply the method fails to generalize to complex, real-world tasks?\n2. The analysis uses a minimal model. How do we know this distribution shift problem even exists in more complex SOTA architectures like Mamba or GLA, which were not tested?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper investigates the intuitive performance drop caused by removing the gate's dependency on $h_{t-1}$. It empirically finds that this removal causes a \"distribution shift\" in the gate values. The discovery that $h_{t-1}$ is the primary driver for pushing gate values towards 0 or 1 in nonlinear RNNs is useful and inspirational.\n2. The solution of applying Gumbel-Softmax initialization only to the first layer based on the diagnosis that this is where the shift is most clear, which is a targeted approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the Gumbel-Softmax initialization shows outstanding performance on synthetic tasks, it demonstrates almost no improvement on the more important, complex real-world task of WikiText-103 language modeling (Table 1).\n2. All analyses and experiments are conducted on a minimal gated linear RNN or a simple 6-layer version. The paper completely lacks validation on any modern, state-of-the-art Linear RNN architecture (e.g., Mamba, RWKV). And the research can be expanded to modern architectures that adopt the matrix-formed memory (e.g., Gated Deltanet, TTT).\n3. The benchmark is relatively insufficient. Some basic widely used language tasks such PIQA, ARC are not included. This may due to that model is too small that cannot handle even a bit complex tasks. However, the results on small models may not be persuasive."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926481873,"tcdate":1761570248458,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16349/Reviewer_wM4S"],"signatures":["ICLR.cc/2026/Conference/Submission16349/Reviewer_wM4S"],"forum":"yLA9l9hykz","number":1,"license":"CC BY 4.0","cdate":1761570248458,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16349/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926481873,"domain":"ICLR.cc/2026/Conference","replyto":"yLA9l9hykz","id":"GCOFtnovg1","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Recurrent Neural Network","Gate mechanism"]},"supplementary_material":{"value":"/attachment/eabf0830918567b555bcceb1949f51417a439308.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Linear Recurrent Neural Networks (RNNs) have attracted attention for their memory and computational efficiency.\nIn particular, gated linear RNNs enable nonlinear transformations through gating mechanisms while still maintaining linear time complexity by removing hidden states from them.\nHowever, the impact of the gate mechanisms and such removal of hidden states from them remains unexplored.\nHere we empirically investigate the impact of these gating mechanisms and find that gate values near zero or one highly depend on hidden states, leading to unintended distribution shifts of gate values when hidden states are removed in gated linear RNNs.\nBased on our findings, we propose an algorithm to mitigate the distribution shifts, which empirically improves performance on long-sequence modeling tasks."},"_bibtex":{"value":"@misc{\niwamori2026towards,\ntitle={Towards Understanding Gated Linear Recurrent Neural Networks},\nauthor={Toshiya Iwamori and Mahito Sugiyama},\nyear={2026},\nurl={https://openreview.net/forum?id=yLA9l9hykz}\n}"},"title":{"value":"Towards Understanding Gated Linear Recurrent Neural Networks"},"pdf":{"value":"/pdf/9b9edb2b974c5293bac768c2ba3eb824740d18c8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"iwamori|towards_understanding_gated_linear_recurrent_neural_networks"},"authorids":{"value":["~Toshiya_Iwamori1","~Mahito_Sugiyama1"]},"authors":{"value":["Toshiya Iwamori","Mahito Sugiyama"]}},"version":2},{"content":{"summary":{"value":"This paper presents M²F-PINN, a Multi-Scale Frequency-domain Physics-Informed Neural Network designed for ocean current forecasting. The model integrates multi-frequency feature embeddings, momentum-equation-based physical constraints, and a 3D Swin Transformer backbone to capture multi-scale dynamics. Specifically, it introduces Fourier mapping to address spectral bias in the model, incorporates PDE residuals from zonal and meridional momentum equations as physics loss terms, and learns spatiotemporal evolution via a Transformer-based autoregressive framework. Experiments conducted on the GLORYS12 reanalysis dataset (2005–2008) show that the proposed model outperforms several baselines in both prediction accuracy and physical consistency."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Since the GLORYS12 dataset is a reanalysis data rather than a pure numerical model output, it may not strictly obey the PDEs used for the physics-informed loss. Could the authors clarify whether the GLORYS fields are consistent with the momentum equations? If inconsistencies exist, how are they handled during optimization, and how does this affect model stability and interpretability?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"(1) Multi-scale representation via frequency embeddings. The use of Gaussian Fourier features helps alleviate spectral bias and improves learning of high-frequency dynamics, a common limitation of PINNs and spatiotemporal neural operators.\n(2) Across all evaluation metrics, M²F-PINN achieves higher accuracy and better stability compared to state-of-the-art models such as XiHe, and WenHai. The results suggest that the multi-scale frequency-domain learning strategy, combined with physics-informed constraints, effectively enhances both prediction accuracy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) Experimental setup simplicity and comparison fairness. The experimental design focuses on a relatively simplified ocean forecasting setup. In contrast, baseline models such as XiHe and WenHai were originally developed for more complex, multi-variable, and higher-resolution scenarios. Using them under a simplified setting may reduce the fairness and persuasive power of the comparison, as these models are not optimized for such reduced configurations. However, the simplified experimental configuration in this paper makes the comparison less convincing.\n(2) Limited novelty in methodological design. The integration of frequency-domain representations (Fourier features), NTK-inspired spectral components, and PDE-based physics losses is reasonable and well-motivated but not conceptually new. Similar combinations have appeared in recent physics-informed operator learning and PINNs literature. As a result, the innovation of M²F-PINN lies more in assembling existing techniques than in introducing a fundamentally new modeling principle or architecture.\n(3) Minor typo. In line 48, the referenced model name “AI-GMOS” is a typo — it should be “AI-GOMS.”"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931448106,"tcdate":1761468451393,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19568/Reviewer_ZXEn"],"signatures":["ICLR.cc/2026/Conference/Submission19568/Reviewer_ZXEn"],"forum":"4z7tMSzQET","number":1,"license":"CC BY 4.0","cdate":1761468451393,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19568/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931448106,"domain":"ICLR.cc/2026/Conference","replyto":"4z7tMSzQET","id":"LjZaCzzeXC","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"This work introduces  a multi-scale, frequency-domain, physics-informed neural network for ocean forecasting that enhances physical interpretability and effectively captures frequency information."},"keywords":{"value":["physics-informed neural networks(PINN)","multi-scale Fourier feature","ocean forecasting"]},"supplementary_material":{"value":"/attachment/0f7263a6ae9549b8db445a92f0dd4b9d255ff90f.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics‐informed neural networks (PINNs) embed physical laws into data-driven learning and are becoming increasingly influential in climate and ocean forecasting. Yet effectively capturing multi-scale variability across high and low frequencies while maintaining training stablility and ensuring convergence remains challenging for conventional PINNs. We introduce M$^2$F-PINN, a novel Transformer-based multi-scale frequency-domain multi-PINN algorithm designed to 1) mitigate spectral bias via Fourier representation learning, and 2) analyze multi-scale characteristics through frequency-domain modeling, and 3) incorporate physics priors using multiple PINNs. M$^2$F-PINN leverages multi-scale Fourier networks to learn spectral components and multi-scale interactions, and employs a 3D Swin Transformer in an autoregressive setting to capture spatiotemporal regularities. The advantages of M$^2$F-PINN include: 1) adaptively learns multi-scale frequency components to enhance the modeling of multi-scale dynamics; 2) jointly estimates physical coefficients within the PINN modules, refining representations of physical processes; 3) preserves the Transformer framework, enabling compatibility with diverse architectures and structural decoupling; 4) extensive experiments on real-world ocean datasets show that M$^2$F-PINN outperforms deep-learning baselines and competitive ocean models (e.g., XiHe, WenHai) in predicting ocean current fields, achieving superior performance across multiple time horizons."},"_bibtex":{"value":"@misc{\nlinfei2026mfpinn,\ntitle={M{\\texttwosuperior}F-{PINN}: A Multi-Scale Frequency-Domain Multi-Physics-Informed Neural Network for Ocean Forecasting},\nauthor={Cao Linfei and Jiachen Yang and Meng Xi and Jingyi He and Fei Gao and Jiabao Wen and Zhijing Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=4z7tMSzQET}\n}"},"title":{"value":"M²F-PINN: A Multi-Scale Frequency-Domain Multi-Physics-Informed Neural Network for Ocean Forecasting"},"pdf":{"value":"/pdf/12031d2723c890bb5102cc1a2a3540f3957dd963.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"linfei|mfpinn_a_multiscale_frequencydomain_multiphysicsinformed_neural_network_for_ocean_forecasting"},"authorids":{"value":["~Cao_Linfei1","~Jiachen_Yang2","~Meng_Xi2","~Jingyi_He2","~Fei_Gao29","~Jiabao_Wen2","~Zhijing_Wang2"]},"authors":{"value":["Cao Linfei","Jiachen Yang","Meng Xi","Jingyi He","Fei Gao","Jiabao Wen","Zhijing Wang"]}},"version":2},{"content":{"TLDR":{"value":"Synthetic MuJoCo and PhiFlow scenes alone, with no human labels and under $30 of hosted-API credits, lift Qwen3-VL-30B on PhysBench by +6.9pp (+20.1pp Scene) without degrading prior textbook-physics ability on ScienceQA."},"venue":{"value":"AI4Physics"},"pdf":{"value":"/pdf/0583bc04ef33ffec1091e3c3595c048578b722bc.pdf","readers":["everyone"]},"keywords":{"value":["vision-language models","physical reasoning","synthetic data","simulation-to-real","MuJoCo","PhiFlow","GRPO","parameter-efficient fine-tuning"]},"venueid":{"value":"ICML.cc/2026/Workshop/AI4Physics"},"paperhash":{"value":"r|synthetic_physics_as_supervision_learning_realworld_physical_reasoning_in_visionlanguage_models"},"authorids":{"value":["~Swastik_R1","~Natesha_B_V1"]},"abstract":{"value":"Vision-language models (VLMs) remain unreliable on visually grounded physics reasoning in real-world media, and human-labelled physics supervision is expensive. We study whether *synthetic* simulator scenes alone can supply an effective training signal. We fine-tune Qwen3-VL-30B on 14,597 rigid-body and fluid scenes with free-text answers read directly from simulator state, using no human annotation and no local GPU. On PhysBench Test (n=9,786), this synthetic-only SFT lifts accuracy from 40.7% to 47.6% (+6.9pp), improves 27/39 subtasks, and yields a +20.1pp gain on the Scene domain. A follow-up GRPO stage with simulator-verifiable rewards preserves this aggregate gain (+7.0pp over baseline) while adding targeted improvements on the un-trained `general:relationships` domain, and a data-fidelity probe at both stages shows the gains track the physical quality of the synthetic signal. On a vision-essential ScienceQA-Physics probe, prior textbook-physics ability is *not* significantly degraded, so the model remains a generalist VLM. The full pipeline costs under \\$30 in hosted-API credits, providing evidence that synthetic physics is a practical and reproducible supervision signal for grounded visual reasoning in VLMs."},"_bibtex":{"value":"@inproceedings{\nr2026synthetic,\ntitle={Synthetic Physics as Supervision: Learning Real-World Physical Reasoning in Vision-Language Models},\nauthor={Swastik R and Natesha B V},\nbooktitle={ICML 2026 Workshop on AI for Physics},\nyear={2026},\nurl={https://openreview.net/forum?id=ItlBE5738K}\n}"},"title":{"value":"Synthetic Physics as Supervision: Learning Real-World Physical Reasoning in Vision-Language Models"},"authors":{"value":["Swastik R","Natesha B V"]}},"tmdate":1784576772400,"pdate":1780110847196,"tcdate":1778078374052,"writers":["ICML.cc/2026/Workshop/AI4Physics","ICML.cc/2026/Workshop/AI4Physics/Submission76/Authors"],"signatures":["ICML.cc/2026/Workshop/AI4Physics/Submission76/Authors"],"forum":"ItlBE5738K","license":"CC BY 4.0","number":76,"cdate":1778078374052,"readers":["everyone"],"invitations":["ICML.cc/2026/Workshop/AI4Physics/-/Submission","ICML.cc/2026/Workshop/AI4Physics/-/Post_Submission","ICML.cc/2026/Workshop/AI4Physics/-/Edit","ICML.cc/2026/Workshop/AI4Physics/Submission76/-/Camera_Ready_Revision"],"mdate":1784576772400,"odate":1782846806976,"domain":"ICML.cc/2026/Workshop/AI4Physics","id":"ItlBE5738K","version":2},{"content":{"summary":{"value":"This work presents a more accessible physics simulation by introducing the language model so that the physics solver can be interacted by using the test, called text2PDE. The proposed method is verified cylinder flow and buoyancy-driven flow although they cannot demonstrate enough the claimed effect."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. The reviewer has some doubts about the selected generative model. For such prediction and simulation tasks, diffusion models are not always effective due to their SDE features. Flow matching, which is based on ODEs, appears more suitable for PDE solvers. Could the authors discuss the application of both models for PDE solvers and explain why they chose diffusion models over flow matching?\n2. In line 229, \"we find that GANs and perceptual guidance can in certain cases improve reconstruction performance, but for simplicity, we omit them in our main results.\" The author said GANs can improve the performance, but why you do not apply GANs?\n3. In Figure 5, the middle two figures do not show turbulent flow, making this characterization inaccurate. Specifically, the second figure clearly shows laminar wake flow, and while the third figure shows vortices, it has not yet reached turbulent conditions."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The developed framework encodes and decodes arbitrary PDE data into a latent space to generate unstructured physics solutions.\n2. The language model is applied in the framework to make the PDE solver more accessible.\n3. The paper applies two experiments to demonstrate the superior performance of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The reviewer raises doubts about the novelty and contribution of the work, seemingly only using the encoding of the language model as a condition to generate solutions of PDEs.\n2. Why do the two experiments adopt different baselines and lack baselines such as FNO for cylinder flow?\n3. From line 047 to 050, while this statement may seem obvious, it's not a scientific issue. It is necessary to have physics and numerical knowledge to build PDE solvers. At the same time, approaches based on FNO, DeepOnet and similar numerically-driven methods can learn accurate PDE solvers without requiring physics knowledge.\n4. From line 053, this work deviates from the challenges described above, and there are concerns about potential overclaiming. Based on the author's two experiments, the proposed model only works in specific PDEs, spatiotemporal dimensions, and simple scenarios where the required physics knowledge is quite basic. In other words, the authors failed to address a crucial question: for complex PDE systems, how to represent the system using language, when mathematical formulas are clearly more suitable, while it is the main challenge the reviewer understand from the previous paragraph.\n5. Spatial and temporal resolution is crucial for PDE solvers. How does this work handle different spatial resolutions? Specifically, how does this method distinguish between initial conditions with different resolutions within the same system?\n6. Could you provide more details about the autoencoder in Sec. 3.1? This part is quite abstract and difficult for readers to understand intuitively. Please briefly explain the autoencoder's architecture, input/output structure, and training methodology. Additionally, during the diffusion model training, is the encoder trained simultaneously or are its parameters fixed?\n7. In line 235, the decoder appears to be CNN-based, which imposes constraints on the spatial dimensions of the output solutions. Furthermore, interpolation operations can lead to inaccuracies in the results, even when solutions on regular grids are accurate.\n8. In the method section, the authors have placed too many crucial details in the appendix, like Appendix C.2, B.1, B.2, D, making it impossible to understand the method without referring to the appendix.\n9. The selected parameters for cylinder flow are overly simplistic, only covering laminar flow cases. Generally, the Kármán vortex street phenomenon is more common in cylinder flows, yet the authors did not include this scenario.\n10. All data used is 2D, and it's unclear whether this model can handle 1D and 3D experiments. The lack of such experiments significantly reduces the persuasiveness of the work.\n11. Without knowing the scale of the data, reporting only L1 loss makes evaluation difficult. Could authors additionally report the relative L2 loss?\n12. As the authors acknowledge, the inference time of diffusion models limits their application, which is confirmed by the statistical results. The authors should consider using inference acceleration methods such as DDIM and verify that the method remains effective under DDIM sampling."}},"nonreaders":[],"tmdate":1731428180172,"tcdate":1730271483697,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4592/Reviewer_zsUD"],"signatures":["ICLR.cc/2025/Conference/Submission4592/Reviewer_zsUD"],"forum":"Nb3a8aUGfj","number":1,"license":"CC BY 4.0","cdate":1730271483697,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4592/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428180172,"domain":"ICLR.cc/2025/Conference","replyto":"Nb3a8aUGfj","id":"gXI89kutQm","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"We develop novel methods for text-conditioned physics simulation that are accurate and efficient."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["AI4Science","PDE","Neural Operator","Latent Diffusion","Text2PDE"]},"supplementary_material":{"value":"/attachment/04393117b5086d1a03a72c313369c9fab8480be3.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advances in deep learning have inspired numerous works on data-driven solutions to partial differential equation (PDE) problems. These neural PDE solvers can often be much faster than their numerical counterparts; however, each presents its unique limitations and generally balances training cost, numerical accuracy, and ease of applicability to different problem setups. To address these limitations, we introduce several methods to apply latent diffusion models to physics simulation. Firstly, we introduce a mesh autoencoder to compress arbitrarily discretized PDE data, allowing for efficient diffusion training across various physics. Furthermore, we investigate full spatiotemporal solution generation to mitigate autoregressive error accumulation. Lastly, we investigate conditioning on initial physical quantities, as well as conditioning solely on a text prompt to introduce text2PDE generation. We show that language can be a compact, interpretable, and accurate modality for generating physics simulations, paving the way for more usable and accessible PDE solvers. Through experiments on both uniform and structured grids, we show that the proposed approach is competitive with current neural PDE solvers in both accuracy and efficiency, with promising scaling behavior up to $\\sim$3 billion parameters. By introducing a scalable, accurate, and usable physics simulator, we hope to bring neural PDE solvers closer to practical use."},"_bibtex":{"value":"@inproceedings{\nzhou2025textpde,\ntitle={Text2{PDE}: Latent Diffusion Models for Accessible Physics Simulation},\nauthor={Anthony Zhou and Zijie Li and Michael Schneier and John R Buchanan Jr and Amir Barati Farimani},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Nb3a8aUGfj}\n}"},"title":{"value":"Text2PDE: Latent Diffusion Models for Accessible Physics Simulation"},"pdf":{"value":"/pdf/8bec6717db33fa179f192c141cc2a48e28c648d8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhou|text2pde_latent_diffusion_models_for_accessible_physics_simulation"},"authorids":{"value":["~Anthony_Zhou1","~Zijie_Li2","~Michael_Schneier1","~John_R_Buchanan_Jr1","~Amir_Barati_Farimani2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Anthony Zhou","Zijie Li","Michael Schneier","John R Buchanan Jr","Amir Barati Farimani"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2506.13777v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"xu|a_survey_of_physicsinformed_ai_for_complex_urban_systems"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:En_Xu:","https://dblp.org/search/pid/api?q=author:Huandong_Wang:","https://dblp.org/search/pid/api?q=author:Yunke_Zhang:","~Sibo_Li1","https://dblp.org/search/pid/api?q=author:Yinzhou_Tang:","https://dblp.org/search/pid/api?q=author:Zhilun_Zhou:","https://dblp.org/search/pid/api?q=author:Yuming_Lin_0003:","https://dblp.org/search/pid/api?q=author:Yuan_Yuan_0032:","https://dblp.org/search/pid/api?q=author:Xiaochen_Fan:","","~Yong_Li7"]},"html":{"value":"https://doi.org/10.48550/arXiv.2506.13777"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2506-13777,\n  publtype={informal},\n  author={En Xu and Huandong Wang and Yunke Zhang and Sibo Li and Yinzhou Tang and Zhilun Zhou and Yuming Lin and Yuan Yuan and Xiaochen Fan and Jingtao Ding and Yong Li},\n  title={A Survey of Physics-Informed AI for Complex Urban Systems},\n  year={2025},\n  month={June},\n  cdate={1748736000000},\n  journal={CoRR},\n  volume={abs/2506.13777},\n  url={https://doi.org/10.48550/arXiv.2506.13777}\n}\n"},"abstract":{"value":"Urban systems are typical examples of complex systems, where the integration of physics-based modeling with artificial intelligence (AI) presents a promising paradigm for enhancing predictive accuracy, interpretability, and decision-making. In this context, AI excels at capturing complex, nonlinear relationships, while physics-based models ensure consistency with real-world laws and provide interpretable insights. We provide a comprehensive review of physics-informed AI methods in urban applications. The proposed taxonomy categorizes existing approaches into three paradigms - Physics-Integrated AI, Physics-AI Hybrid Ensemble, and AI-Integrated Physics - and further details seven representative methods. This classification clarifies the varying degrees and directions of physics-AI integration, guiding the selection and development of appropriate methods based on application needs and data availability. We systematically examine their applications across eight key urban domains: energy, environment, economy, transportation, information, public services, emergency management, and the urban system as a whole. Our analysis highlights how these methodologies leverage physical laws and data-driven models to address urban challenges, enhancing system reliability, efficiency, and adaptability. By synthesizing existing methodologies and their urban applications, we identify critical gaps and outline future research directions, paving the way toward next-generation intelligent urban system modeling."},"title":{"value":"A Survey of Physics-Informed AI for Complex Urban Systems"},"authors":{"value":["En Xu","Huandong Wang","Yunke Zhang","Sibo Li","Yinzhou Tang","Zhilun Zhou","Yuming Lin","Yuan Yuan","Xiaochen Fan","Jingtao Ding","Yong Li"]}},"tmdate":1769357834589,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2506-13777"],"tcdate":1768608292931,"writers":["~"],"signatures":["~Yong_Li7"],"forum":"jwsoa1WZ1R","license":"CC BY-SA 4.0","number":769219,"cdate":1748736000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769357834589,"domain":"DBLP.org","id":"jwsoa1WZ1R","version":2},{"content":{"venue":{"value":"Crossref"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-032-04965-0_16.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"chalcroft|domainagnostic_stroke_lesion_segmentation_using_physicsconstrained_synthetic_data"},"html":{"value":"https://doi.org/10.1007/978-3-032-04965-0_16"},"abstract":{"value":"Segmenting stroke lesions in MRI is challenging due to diverse acquisition protocols that limit model generalisability. In this work, we introduce two physics-constrained approaches to generate synthetic quantitative MRI (qMRI) images that improve segmentation robustness across heterogeneous domains. Our first method, qATLAS, trains a neural network to estimate qMRI maps from standard MPRAGE images, enabling the simulation of varied MRI sequences with realistic tissue contrasts. The second method, qSynth, synthesises qMRI maps directly from tissue labels using label-conditioned Gaussian mixture models, ensuring physical plausibility. Extensive experiments on multiple out-of-domain datasets show that both methods outperform a baseline UNet, with qSynth notably surpassing previous synthetic data approaches. These results highlight the promise of integrating MRI physics into synthetic data generation for robust, generalisable stroke lesion segmentation. Code is available at https://github.com/liamchalcroft/qsynth."},"title":{"value":"Domain-Agnostic Stroke Lesion Segmentation Using Physics-Constrained Synthetic Data"},"authors":{"value":[{"fullname":"Liam Chalcroft","username":"~Liam_Chalcroft1"},{"fullname":"Jenny Crinion","username":"https://orcid.org/orcid-search/search?searchQuery=Jenny%20Crinion"},{"fullname":"Cathy J. Price","username":"https://orcid.org/orcid-search/search?searchQuery=Cathy%20J.%20Price"},{"fullname":"John Ashburner","username":"https://orcid.org/orcid-search/search?searchQuery=John%20Ashburner"}]}},"tmdate":1777813222190,"pdate":1767225600000,"externalIds":["doi:10.1007/978-3-032-04965-0_16"],"tcdate":1777813209484,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Liam_Chalcroft1"],"forum":"0OESPQzX8F","license":"CC BY-SA 4.0","number":65778,"cdate":1758193848378,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1777813222190,"domain":"OpenReview.net/Public_Article","id":"0OESPQzX8F","version":2},{"content":{"summary":{"value":"This paper proposes STANCE, an image-to-video generation method that improves motion coherence and physical plausibility. The key ideas are: 1 Sparse-to-Dense motion cues: converting instance-level user inputs (arrows, masks, mass) into dense 2.5D control fields. 2. Dense RoPE: selecting active motion tokens and tagging them with rotary embeddings to prevent control collapse. 3. Joint RGB + auxiliary map generation (segmentation or depth) to stabilize spatio-temporal consistency. The method is built on CogVideoX and trained on ~200k simulated rigid-body scenes. Experiments show improvements on a Physics-IQ metric and qualitative control results."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Does your method work for non-rigid motion as well?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.Clear practical focus on controllable object motion\nThe design targets rigid-body interactions and physical plausibility, which is aligned with real-world applications (AR/VR, robotics, creative tools).\n\n2.Simple modular ideas with observable gains\nSparse→dense cues and Dense-RoPE are easy to implement and demonstrate measurable improvements in control fidelity and Physics-IQ.\n\n3.Reasonable dataset and ablation studies\nThe authors curated a specialized Kubric-style dataset and ablated control injection and joint auxiliary supervision."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.Dataset scope is narrow and biased toward synthetic rigid-body scenes\nThe method is primarily validated on artificial Kubric-like collisions, which limits generalization claims. Real-world evaluations are limited to simple tabletop toys and lack diverse environments or camera motions.\n\n2.Limited quantitative evidence and baselines\nQuantitative evaluation relies mainly on Physics-IQ and FVD, without human studies or perceptual realism metrics. Baselines include SG-I2V, DragAnything, MoFA-Video, etc., but more recent state-of-the-art controllable video frameworks are missing (e.g., VACE frameworks).\n\n3.Technical novelty is incremental\nThe contributions largely combine known ingredients—instance masks, flow-derived cues, and RoPE tagging—into a pipeline. While practical, the approach feels like an engineering refinement of existing image/video control pipelines. The method does not propose fundamentally new control paradigms or physical modeling insights."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916463517,"tcdate":1761976677336,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2961/Reviewer_kddo"],"signatures":["ICLR.cc/2026/Conference/Submission2961/Reviewer_kddo"],"forum":"FwtKMYHov7","number":3,"license":"CC BY 4.0","cdate":1761976677336,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2961/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916463517,"domain":"ICLR.cc/2026/Conference","replyto":"FwtKMYHov7","id":"jBSqPVb3yz","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Generation","Generative Model"]},"supplementary_material":{"value":"/attachment/415a700d351d8bd244411f2900432bea49e53253.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation has recently made striking visual progress, but maintaining coherent object motion and interactions remains difficult. We trace two practical bottlenecks: (i) human-provided motion hints (e.g., small 2D maps) often collapse to too few effective tokens after encoding, weakening guidance; and (ii) optimizing for appearance and motion in a single head can favor texture over temporal consistency. We present STANCE, an image-to-video framework that addresses both issues with two simple components.\nFirst, we introduce Instance Cues—a pixel-aligned control signal that turns sparse, user-editable hints into a dense 2.5D (camera-relative) motion field by averaging per-instance flow and augmenting with monocular depth over the instance mask. This reduces depth ambiguity compared to 2D drag/arrow inputs while remaining easy to user. Second, we preserve the salience of these cues in token space with Dense RoPE, which tags a small set of motion tokens (anchored on the first frame) with time-addressable rotary embeddings. Paired with joint RGB + auxiliary-map prediction (segmentation or depth), our model anchors structure while RGB handles appearance, stabilizing optimization and improving temporal coherence without requiring per-frame trajectory scripts."},"_bibtex":{"value":"@misc{\nanonymous2026stance,\ntitle={{STANCE}: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=FwtKMYHov7}\n}"},"title":{"value":"STANCE: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding"},"pdf":{"value":"/pdf/dcb00196385b99164d59c430fb8613633f2432c0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"chen|stance_motion_coherent_video_generation_via_sparsetodense_anchored_encoding"},"authorids":{"value":["~ZhiFei_Chen1","~Tianshuo_Xu1","~Leyi_Wu1","~Luozhou_Wang2","~Dongyu_Yan1","~Zihan_You2","~Wenting_Luo1","~Guo_Zhang2","~Ying-Cong_Chen1"]},"authors":{"value":["ZhiFei Chen","Tianshuo Xu","Leyi Wu","Luozhou Wang","Dongyu Yan","Zihan You","Wenting Luo","Guo Zhang","Ying-Cong Chen"]}},"version":2},{"content":{"summary":{"value":"The paper proposes PhyMAGIC, a training-free framework for generating physically consistent motion from a single image. It combines a pre-trained image-to-video diffusion model, LLM-based confidence reasoning, and a differentiable physics simulator to create 3D assets suitable for simulation without fine-tuning or supervision. By iteratively refining motion prompts with simulation feedback, PhyMAGIC produces realistic and physically plausible dynamics, outperforming existing video generators and physics-aware methods in both physical consistency and visual quality."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"please refer to the weeknesses."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The writing is clear.\n\n2. The integration of LLMs, generation models, and physics simulators is technically compelling.\n\n3. The problem addressed is interesting and meaningful."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tWhat is the ultimate goal of this work — generating 3D or generating video? My understanding is that rendering videos from 3D Gaussians is the final objective, with the video primarily serving as a source of physical information. In that case, is a video generation model truly necessary? Could the physical quantities instead be provided directly by an LLM’s prior knowledge or manually specified by humans? Would this affect the final results? Overall, the connection between video generation and 3D rendering feels quite disconnected.\n\n2.\tPrior works such as DreamPhysics, PhyDreamer, and Physics3D also combine video generation with physics simulators. Could you elaborate more specifically on what the distinct contributions of this work are compared to them? Overall, I think the framework is quite similar.\n\n3.\tSince the Trellis model is used, the system can only generate single-object videos. It cannot handle scenes with backgrounds or objects of higher complexity and realism. For instance, PhysCtrl（NeurIPS25）, which also integrates a simulator with video generation, can produce more complex and natural video results.\n\n4.\tThe rendered outputs suffer from the inherent granularity of Gaussians, leading to a noticeable lack of sharpness and fine details — for example, in the wolf case, most detail is lost in later frames. Moreover, as no video visualizations are provided, it is difficult to fully assess the generation quality; my observations are based solely on the few provided images."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921679194,"tcdate":1761814367239,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10350/Reviewer_ZtRm"],"signatures":["ICLR.cc/2026/Conference/Submission10350/Reviewer_ZtRm"],"forum":"nruZar3Aaz","number":2,"license":"CC BY 4.0","cdate":1761814367239,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10350/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921679194,"domain":"ICLR.cc/2026/Conference","replyto":"nruZar3Aaz","id":"qAAu1A9pK8","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["3D Dynamic Generation","Phyical Priors","Physical Simulation","LLM Reasoning","MPM Simulation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in 3D content generation have amplified demand for dynamic models that are both visually realistic and physically consistent. However, state-of-the-art video diffusion models frequently produce implausible results such as momentum violations and object interpenetrations. Existing physics-aware approaches often rely on task-specific fine-tuning or supervised data, which limits their scalability and applicability. To address the challenge, we present PhyMAGIC, a training-free framework that generates physically consistent motion from a single image. PhyMAGIC integrates a pre-trained image-to-video diffusion model, confidence-guided reasoning via large language models (LLMs), and a differentiable physics simulator to produce 3D assets ready for downstream physical simulation without fine-tuning or manual supervision. By iteratively refining motion prompts using LLM-derived confidence scores and leveraging simulation feedback, PhyMAGIC steers generation toward physically consistent dynamics. Comprehensive experiments demonstrate that PhyMAGIC outperforms state-of-the-art video generators and physics-aware baselines, enhancing physical property inference and motion–text alignment while maintaining visual fidelity."},"_bibtex":{"value":"@misc{\nmeng2025phymagic,\ntitle={Phy{MAGIC}: Physical Motion-Aware Generative Inference with Confidence-guided {LLM}},\nauthor={Siwei Meng and Yawei Luo and Ping Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=nruZar3Aaz}\n}"},"title":{"value":"PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM"},"pdf":{"value":"/pdf/af1594b19fb1335afb9429816fc8f2f5c4f00a74.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"meng|phymagic_physical_motionaware_generative_inference_with_confidenceguided_llm"},"authorids":{"value":["~Siwei_Meng1","~Yawei_Luo3","~Ping_Liu1"]},"authors":{"value":["Siwei Meng","Yawei Luo","Ping Liu"]}},"version":2},{"content":{"summary":{"value":"This position paper argues that modern biology should be understood as a direct outgrowth of physics, both in conceptual foundations and methodological approaches. The authors contend that many breakthroughs in biology—particularly in molecular biology, biophysics, and systems biology—derive from physical principles and tools originally developed in physics. They advocate for a deeper integration of physics-based thinking into biological research, especially in the era of large-scale data, complex systems modeling, and AI-driven discovery. The paper reviews historical examples, outlines conceptual parallels between the disciplines, and identifies opportunities for cross-disciplinary training and research frameworks. The authors propose that embracing physics-style modeling and inference can accelerate biological discovery and improve the rigor of biological sciences."},"agreement":{"value":4},"code_of_conduct_acknowledgement":{"value":"Yes"},"confidence":{"value":4},"questions":{"value":"1. Can the authors provide quantitative or bibliometric evidence showing trends in the adoption of physics-inspired methods in biological research?\n2 How do the authors envision AI acting as a bridge between physics and biology in practical collaborative projects?\n3. Are there specific subfields in biology where the physics-based approach has met resistance, and if so, how might these challenges be addressed?"},"rating":{"value":7},"author_identification":{"value":"No."},"discussion":{"value":3},"ethics":{"value":["NO or VERY MINOR ethics concerns only"]},"significance":{"value":3},"presentation":{"value":3},"thoroughness":{"value":4},"alternative_position":{"value":"Yes, and alternative positions are well-considered and addressed by the argument"},"strengths":{"value":"- The central thesis is clearly articulated, with multiple historical and contemporary examples supporting the link between physics and biology.\n- The paper is well-structured, progressing logically from conceptual framing to historical context and forward-looking recommendations.\n- The emphasis on training, methodology transfer, and the potential for AI to strengthen cross-disciplinary research is timely.\n- The writing style is clear and accessible, making it approachable for a broad NeurIPS audience."},"weaknesses":{"value":"- The paper’s examples, while compelling, are weighted toward molecular and systems biology; broader coverage of other biological subfields could improve generality.\n- Although the argument is persuasive, quantitative evidence showing the direct impact of physics-derived approaches in recent biological breakthroughs is limited.\n- The discussion of AI integration is relatively brief and could be expanded with specific scenarios or case studies.\n- The call to action could be made more actionable, e.g., outlining concrete steps for the NeurIPS community to engage with biological research."},"context":{"value":3},"position":{"value":"Yes, the paper argues for or against a position related to machine learning."},"support":{"value":3}},"parentInvitations":"NeurIPS.cc/2025/Position_Paper_Track/-/Official_Review","nonreaders":[],"tmdate":1761707010347,"tcdate":1754706088692,"writers":["NeurIPS.cc/2025/Position_Paper_Track","NeurIPS.cc/2025/Position_Paper_Track/Submission455/Reviewer_86jt"],"signatures":["NeurIPS.cc/2025/Position_Paper_Track/Submission455/Reviewer_86jt"],"forum":"HzGZVYi8fK","number":3,"license":"CC BY 4.0","cdate":1754706088692,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Position_Paper_Track/Submission455/-/Official_Review","NeurIPS.cc/2025/Position_Paper_Track/-/Edit"],"mdate":1761707010347,"domain":"NeurIPS.cc/2025/Position_Paper_Track","replyto":"HzGZVYi8fK","id":"oFGh3DANou","forumContent":{"TLDR":{"value":"We argue that Physics-Informed ML must evolve to meet the unique challenges of biological modeling, and that this evolution represents not a limitation, but a major opportunity, giving rise to Biology-Informed Machine Learning (BIML)."},"venue":{"value":"NeurIPS 2025 Position Paper Track"},"keywords":{"value":["Physics-Informed Machine Learning","Scientific Machine Learning","Computational Biology","Probabilistic methods"]},"abstract":{"value":"Physics-Informed Machine Learning (PIML) has successfully integrated mechanistic understanding into machine learning, particularly in domains governed by well-known physical laws.\nThis success has motivated efforts to apply PIML to biology, a field rich in dynamical systems but shaped by different constraints.\nBiological modeling, however, presents unique challenges: multi-faceted and uncertain prior knowledge, heterogeneous and noisy data, partial observability, and complex, high-dimensional networks.\n\\textbf{In this position paper, we argue that these challenges should not be seen as obstacles to PIML, but as catalysts for its evolution. We propose Biology-Informed Machine Learning (BIML): a principled extension of PIML that retains its structural grounding while adapting to the practical realities of biology.}\nRather than replacing PIML, BIML retools its methods to operate under softer, probabilistic forms of prior knowledge.\nWe outline four foundational pillars as a roadmap for this transition: uncertainty quantification, contextualization, constrained latent structure inference, and scalability.\nFoundation Models and Large Language Models will be key enablers, bridging human expertise with computational modeling.\nWe conclude with concrete recommendations to build the BIML ecosystem and channel PIML-inspired innovation toward challenges of high scientific and societal relevance."},"_bibtex":{"value":"@inproceedings{\nmartinelli2025position,\ntitle={Position: Biology is the Challenge Physics-Informed {ML} Needs to Evolve},\nauthor={Julien Martinelli},\nbooktitle={The Thirty-Ninth Annual Conference on Neural Information Processing Systems Position Paper Track},\nyear={2025},\nurl={https://openreview.net/forum?id=HzGZVYi8fK}\n}"},"title":{"value":"Position: Biology is the Challenge Physics-Informed ML Needs to Evolve"},"pdf":{"value":"/pdf/dbf168623ba9aab185715b03ce580a4776f6eb65.pdf"},"lay_summary":{"value":"Physics-Informed Machine Learning (PIML) has successfully integrated mechanistic understanding into machine learning, particularly in domains governed by well-known physical laws.\nThis success has motivated efforts to apply PIML to biology, a field rich in dynamical systems but shaped by different constraints.\nBiological modeling, however, presents unique challenges: multi-faceted and uncertain prior knowledge, heterogeneous and noisy data, partial observability, and complex, high-dimensional networks.\nIn this position paper, we argue that these challenges should not be seen as obstacles to PIML, but as catalysts for its evolution. We propose Biology-Informed Machine Learning (BIML): a principled extension of PIML that retains its structural grounding while adapting to the practical realities of biology.\nRather than replacing PIML, BIML retools its methods to operate under softer, probabilistic forms of prior knowledge.\nWe outline four foundational pillars as a roadmap for this transition: uncertainty quantification, contextualization, constrained latent structure inference, and scalability.\nFoundation Models and Large Language Models will be key enablers, bridging human expertise with computational modeling.\nWe conclude with concrete recommendations to build the BIML ecosystem and channel PIML-inspired innovation toward challenges of high scientific and societal relevance."},"venueid":{"value":"NeurIPS.cc/2025/Position_Paper_Track"},"paperhash":{"value":"martinelli|position_biology_is_the_challenge_physicsinformed_ml_needs_to_evolve"},"authorids":{"value":["~Julien_Martinelli1"]},"authors":{"value":["Julien Martinelli"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a hybrid simulation approach that integrates neural networks with physics-based models to enhance accuracy and efficiency. It introduces an implicit gray-box model, where deep neural networks (DNNs) and physical equations share state variables. This implicit integration captures complex coupled interactions and reduces training data requirements. The effectiveness of this approach is demonstrated through simulations of steady-state and transient behaviors in power systems."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"1. Since the NN-based model is used for computing Newton-Raphson and integrated with the physics-based model for internal optimization, the accuracy can be better (Figure 3 and Table 1). How about the performance of the traditional physics-based model in Figure 3 and Table 1? \n2. How to deal with noisy data with the proposed method? How to effectively separate noise and real hidden physics? This will strongly influence the optimization of NN."},"rating":{"value":1},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"This paper presents an implicit hybrid model method for physics-based models and NN-based models. The NN-based models can help extract sensitivity terms and help with the convergence."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the motivation of this paper is good, it is hard to know whether the proposed method is effective in more general and challenging problems. Only the power system example is not sufficient. More challenging and 3D transient examples are needed. Strong and clear examples with enough evidence are required.\n2. This paper claims to focus on large-scale systems, but there are no descriptions of the degrees of freedom of the demonstration example. \n3. The literature review is not comprehensive. Only PINN (line 123) is mentioned in the paper. A comprehensive literature is needed, such as Fourier neural operator, DeepONet, JAX-CFD, and other physics-informed machine learning methods.\n4. Typo: Lin369, there is no Figure 11."}},"nonreaders":[],"tmdate":1731427938351,"tcdate":1730696695611,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4821/Reviewer_MDmX"],"signatures":["ICLR.cc/2025/Conference/Submission4821/Reviewer_MDmX"],"forum":"sSWiZr8QU7","number":4,"license":"CC BY 4.0","cdate":1730696695611,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4821/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427938351,"domain":"ICLR.cc/2025/Conference","replyto":"sSWiZr8QU7","id":"Q1exPCixnT","forumContent":{"TLDR":{"value":"We present a new simulation paradigm that directly integrates DNNs with numerical engines of physics-based solvers to enable simulation of a fully implicit gray box modeling."},"venue":{"value":"ICLR 2025 Conference Desk Rejected Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["gray box modeling","simulation","neural networks"]},"supplementary_material":{"value":"/attachment/d46e857012ac5443346dbdb14f6d311a7a70729f.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Simulation is vital for scientific and engineering disciplines, as it enables the prediction and design of physical systems. However, the computational challenges inherent to large-scale simulations often arise from complex device models featuring high degrees of nonlinearities or hidden physical behaviors not captured by first principles. Gray-box models that combine deep neural networks (DNNs) with physics-based models have been proposed to address the computational challenges in modeling complex physical systems. A well-crafted gray box model capitalizes on the interpretability and accuracy of a physical model while incorporating deep neural networks to capture hidden physical behaviors and mitigate computational load associated with highly nonlinear components. Previously, gray box models have been constructed by defining an explicit combination of physics-based and black-box models to represent the behavior of sub-systems; however this alone cannot represent the coupled interactions that define the behavior of the entire physical system. We, therefore, explore an implicit gray box model, where both DNNs (trained on measurement and simulated data) and physical equations share a common set of state-variables. While this approach captures coupled interactions at the boundary of data-driven and physics-based models, simulating the implicit gray box model remains an open-ended problem. In this work, we introduce a new hybrid simulation that directly integrates DNNs into the numerical solvers of simulation engines to fully simulate implicit gray box models of large physical systems. This is accomplished by backpropagating through the DNN to calculate specific Jacobian values during each iteration of the numerical method. The hybrid simulation of implicit gray-box models improves the accuracy and runtime compared to full physics-based simulation and enables reusable DNN models with lower data requirements for training. For demonstration, we explore the advantages of this approach as compared to physics-based, black box, and other gray box methods for simulating the steady-state and electromagnetic transient behavior of power systems."},"_bibtex":{"value":"@misc{\nagarwal2024a,\ntitle={A Hybrid Simulation of {DNN}-based Gray Box Models},\nauthor={Aayushya Agarwal and Yihan Ruan and Lawrence Pileggi},\nyear={2024},\nurl={https://openreview.net/forum?id=sSWiZr8QU7}\n}"},"title":{"value":"A Hybrid Simulation of DNN-based Gray Box Models"},"pdf":{"value":"/pdf/91be2685b8516df0dec43e00cb865f1ad6f4029a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"agarwal|a_hybrid_simulation_of_dnnbased_gray_box_models"},"authorids":{"value":["~Aayushya_Agarwal1","~Yihan_Ruan1","~Lawrence_Pileggi1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Aayushya Agarwal","Yihan Ruan","Lawrence Pileggi"]}},"version":2},{"content":{"summary":{"value":"The authors propose the Pseudo Physics-Informed Neural Operator (PPI-NO), which couples the existing concepts of physics discovery and neural operator learning. In particular, a surrogate partial differential equation (PDE) representation is learned from data using a neural network. Afterwards, the neural-network-PDE model is used as a regularizer to refine the training of the neural operator. The authors claim that the coupling helps the neural operator learn effectively in the low data limit."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. l083. You define f(x) as the source function. I believe neural operators go beyond simply source functions to solution mapping.\n2. l087. \\mathbb{F} and \\mathbb{U} are not defined.\n3. l147. How order of derivatives should be chosen?\n4. Eq. (5). Why generate N' samples in the second term? Instead, why can we not use the available N samples from the first term?\n5. In section 4, important literature in this area are missing. For example, SNO [1], CNO [2], LNO [3], and PIWNO [4] are not reviewed.\n6. l301. Why are the same derivatives not used across all the examples? How are they chosen?\n7. l302. For the SIF example, why are polynomials of the derivatives not used? \n8. l311. What do the iterations denote?\n9. l317. For the SIF example, 400-600 training samples are used. Obtaining such a training set using high-fidelity crack simulations is very costly. This completely defeats the purpose of the proposed framework.\n10. Table 1. Why does the error in DONet-Darcy, DONet-Poisson, and DONet-Advection examples increase with the increase in training data?\n11. In Table 2. The decrease in error in the case of PPI-NO is very marginal. This indicates that the incorporation of rudimentary physics is ineffective in complex problems like the SIF prediction. \n12. l413. Should the baseline comparison be moved to an ablation study in the given setup? Otherwise, the comparison for physics accuracy should be made with dedicated physics discovery algorithms like PINN-SR [5].\n\n\n[1] Fanaskov, Vladimir Sergeevich, and Ivan V. Oseledets. \"Spectral neural operators.\" Doklady Mathematics. Vol. 108. No. Suppl 2. Moscow: Pleiades Publishing, 2023.\n\n[2] Raonic, Bogdan, et al. \"Convolutional neural operators for robust and accurate learning of PDEs.\" Advances in Neural Information Processing Systems 36 (2024).\n\n[3] Cao, Qianying, Somdatta Goswami, and George Em Karniadakis. \"Laplace neural operator for solving differential equations.\" Nature Machine Intelligence 6.6 (2024): 631-640.\n\n[4] Navaneeth, N., Tapas Tripura, and Souvik Chakraborty. \"Physics informed WNO.\" Computer Methods in Applied Mechanics and Engineering 418 (2024): 116546.\n\n[5] Chen, Zhao, Yang Liu, and Hao Sun. \"Physics-informed learning of governing equations from scarce data.\" Nature communications 12.1 (2021): 6136."},"rating":{"value":3},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper is written clearly and has appropriate results to support the authors' claim. Further, the paper proposes the integration of the non-trivial concepts of physics discovery and neural operator learning, which is an important problem."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The idea of coupling physics discovery with NN has been explored earlier. For e.g., see PINN-SR [1].\n\n2. The basic idea of the manuscript is problematic. The discovered \"pseudo\" physics is not exact and hence is of much lower-fidelity (and is unlikely to generalize). The data available is of higher fidelity. Therefore a composite loss function where one term is of higher fidelity and the other is of lower fidelity will, in theory, stop the model from generalization. This fact has been previously pointed on in [2] and as a remedy transfer learning was proposed. \n\n3. Even by incorporating rudimentary physics information, a significant decrease in error is not observed in Table 1 (which is not totally unexpected given the point above). In the results of the DONet-Darcy flow, DONet-Diffusion, and all Poisson and advection equations, the reduction in error is minimal, which makes the contribution of the discovered physics marginal.\n\n4. Like any other basis function-based physics-discovery algorithms, this framework also requires careful selection of the derivatives, which limits the proposed framework's applicability. It is evident in Table 1. Even when the training data is increased, the relative error increases instead of decreasing in some cases. This may be due to faulty physics identification. I will also add that since the exact terms are not known, using a L2 loss in generally not preferred (as with L2 error, even those terms that are supposed to be absent will have non-zero weights). This contributes to the error in equation discovery and hence, the accuracy of the overall method.\n\n5. Important aspects like the effect of incorporating physics on the zero-shot prediction on super- and sub-resolutions, as well as generalization to out-of-distribution input, have not been studied. These are required to gauge the strength of the proposed framework correctly.\n\n[1] Chen, Zhao, Yang Liu, and Hao Sun. \"Physics-informed learning of governing equations from scarce data.\" Nature communications 12.1 (2021): 6136.\n[2] Chakraborty S. Transfer learning based multi-fidelity physics informed deep neural network. Journal of Computational Physics. 2021 Feb 1;426:109942."}},"nonreaders":[],"tmdate":1731428278630,"tcdate":1730520150208,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4999/Reviewer_T4BB"],"signatures":["ICLR.cc/2025/Conference/Submission4999/Reviewer_T4BB"],"forum":"CrmUKllBKs","number":4,"license":"CC BY 4.0","cdate":1730520150208,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4999/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428278630,"domain":"ICLR.cc/2025/Conference","replyto":"CrmUKllBKs","id":"54g0akh3II","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Pseudo Physics","Data-Driven Physics Discovery","PDEs","Neural Operator","AI for science","Scientific Machine Learning"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in operator learning are transforming the landscape of computational physics and engineering, especially alongside the rapidly evolving field of physics-informed machine learning. The convergence of these areas offers\nexciting opportunities for innovative research and applications. However, merging\nthese two realms often demands deep expertise and explicit knowledge of physical systems, which may be challenging or even impractical in relatively complex applications. To address this limitation, we propose a novel framework: Pseudo\nPhysics-Informed Neural Operator (PPI-NO). In this framework, we construct a\nsurrogate physics system for the target system using partial differential equations\n(PDEs) derived from simple, rudimentary physics knowledge, such as basic differential operators. We then couple the surrogate system with the neural operator model, utilizing an alternating update and learning process to iteratively enhance\nthe model’s predictive power. While the physics derived via PPI-NO may not mirror the ground-truth underlying physical laws — hence the term “pseudo physics” — this approach significantly enhances the accuracy of current operator learning\nmodels, particularly in data scarce scenarios. Through extensive evaluations across\nfive benchmark operator learning tasks and an application in fatigue modeling,\nPPI-NO consistently outperforms competing methods by a significant margin. The\nsuccess of PPI-NO may introduce a new paradigm in physics-informed machine\nlearning, one that requires minimal physics knowledge and opens the door to\nbroader applications in data-driven physics learning and simulations."},"_bibtex":{"value":"@misc{\nchen2025pseudo,\ntitle={Pseudo Physics-Informed Neural Operators},\nauthor={Keyan Chen and Yile Li and Da Long and WEI W. XING and Jacob Hochhalter and Shandian Zhe},\nyear={2025},\nurl={https://openreview.net/forum?id=CrmUKllBKs}\n}"},"title":{"value":"Pseudo Physics-Informed Neural Operators"},"pdf":{"value":"/pdf/864b77c21caf8f31310746c5d9b464fe5feadfa1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|pseudo_physicsinformed_neural_operators"},"authorids":{"value":["~Keyan_Chen3","~Yile_Li1","~Da_Long1","~WEI_W._XING1","~Jacob_Hochhalter1","~Shandian_Zhe1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Keyan Chen","Yile Li","Da Long","WEI W. XING","Jacob Hochhalter","Shandian Zhe"]}},"version":2},{"content":{"summary":{"value":"A new approach for performing ILP using neural network parameterization is proposed. The main contribution is an approach capable of learning a broader set of rules than what is currently possible with state-of-the-art neural ILP methods. The proposed method relies on evaluating/verifying SAT on tree-like FOL structures by adapting a factor-graph message passing approach similar to existing methods. The novel aspect is “folding” which is to use this to learn more complex FOL structures by introducing constraints for merging similar structures. \n\nFurther, to learn using neural methods, each of the discrete operations in the evaluation and learning of FOL structures is encoded as a differentiable operation on a continuous tensor representation (similar to TensorLog). Experiments are performed on 3 datasets (one synthetic and 2 other standard ones) and comparisons with an existing state-of-the-art neural ILP method show promising results."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"- Learning more complex structures from neural ILPs seems like a significant contribution\n- The idea of using constraints for merging to form complex structures from trees and encoding them with neural nets seems interesting"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The experiments dot not adequately show the impact of the proposed approach. Specifically, there is a single synthetic example (community) on which the complex rule learning outcome is demonstrated. It seems like the other compared approach fails here. There are 2 other benchmarks, but it seems like the proposed approach is not necessary here. \n- The paper leverages existing approaches (e.g. TensorLog, message-passing, etc.) so it was hard to understand the novel contributions of the paper."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Can there be a more comprehensive evaluation done to show that i) complex rules are required for real-world cases and ii) existing methods fail for such cases while the proposed method can effectively learn such rules.\n\nOne of the aspects shown in the experiments is also learning time. How do more complex structures affect this?\n\nIf the novel contributions were better highlighted it would be useful to evaluate significance of the proposed method."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637049592,"tcdate":1698689189302,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8424/Reviewer_Y7bE"],"signatures":["ICLR.cc/2024/Conference/Submission8424/Reviewer_Y7bE"],"forum":"p6hIAEHwSp","number":2,"license":"CC BY 4.0","cdate":1698689189302,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8424/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637049592,"domain":"ICLR.cc/2024/Conference","replyto":"p6hIAEHwSp","id":"y8OTHPqtg3","forumContent":{"TLDR":{"value":"This paper extends differentiable backwards chaining inductive logic programming techniques to support subgraph-shaped rules."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["inductive logic programming","subgraph rules","gradient-based"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Differentiable inductive logic programming techniques have proven effective at learning logic rules from noisy datasets; however, existing algorithms incur pernicious trade-offs between rule expressivity and scalability to large problems. Forward-chaining ILP algorithms can learn arbitrary rules, but their memory requirements scale exponentially with problem size. Backwards-chaining ILP algorithms address this limitation but do so with loss of generality by imposing the restrictive constraint that rules must be expressible as ensembles of independent chain-like Horn clauses. In this paper we present FUSE-ILP, a technique that relaxes this chain-like constraint and enables the differentiable evaluation of a restricted class of subgraph-like rules. Our method extends TensorLog-inspired backwards-chaining ILP techniques with branch masking and leaf grouping, which enable tree-like rule evaluation and “folding” of these trees into subgraphs. We demonstrate that this formulation allows our algorithm to learn more expressive rules than previous backwards-chaining algorithms while retaining a similar computational cost."},"_bibtex":{"value":"@misc{\njohnson2024efficient,\ntitle={Efficient Subgraph Rule Induction via Tree Folding in Differentiable Logic Programming},\nauthor={Blair Johnson and Faramarz Fekri and James Clayton Kerce},\nyear={2024},\nurl={https://openreview.net/forum?id=p6hIAEHwSp}\n}"},"title":{"value":"Efficient Subgraph Rule Induction via Tree Folding in Differentiable Logic Programming"},"pdf":{"value":"/pdf/7a54af3d55039a2f6aed6ee3e188313fd5e1e04c.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"johnson|efficient_subgraph_rule_induction_via_tree_folding_in_differentiable_logic_programming"},"authorids":{"value":["~Blair_Johnson1","~Faramarz_Fekri1","~James_Clayton_Kerce1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Blair Johnson","Faramarz Fekri","James Clayton Kerce"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the challenge of denoising magnetometer data for magnetic anomaly navigation, a critical capability for environments where GPS signals are unavailable or compromised. The authors propose a novel approach that embeds two physics-based constraints directly into neural network architectures. The paper also emphasizes continuous-time dynamics and long-term memory for handling irregularly sampled magnetometer data. To address data scarcity, the authors develop synthetic datasets using the World Magnetic Model combined with time-series conditional GANs. Experimental results demonstrate that their physics-aware constraints improve predictive accuracy and physical plausibility across multiple architectures (CNNs, MLPs, LTCs, and Contiformers)."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See Weaknesses."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses an important practical problem in navigation where GPS is unavailable or compromised, which has significant real-world applications in both civilian and military contexts.\n2. The implementation of divergence-free fields through vector potential representation is mathematically sound and directly enforces a key physical constraint."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper combines several existing techniques (physics-informed neural networks, equivariant networks, continuous-time models). Thus, this paper has limited novelty and contribution.\n2. The paper claims to \"outperform the state-of-the-art\" but does not compare against recent relevant methods that address similar problems. Therefore, the experiment is limited.\n3. The paper mentions using \"Flt1005, Flt1006\" datasets but provides no information about their size, diversity, or how they were acquired.\n4. The performance improvements could equally be attributed to the network architecture choices rather than the physics constraints themselves.\n5. The authors should experiment on real datasets, not synthetic datasets."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927705383,"tcdate":1761988217696,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17880/Reviewer_Lzc1"],"signatures":["ICLR.cc/2026/Conference/Submission17880/Reviewer_Lzc1"],"forum":"WlZZcV67uz","number":3,"license":"CC BY 4.0","cdate":1761988217696,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17880/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927705383,"domain":"ICLR.cc/2026/Conference","replyto":"WlZZcV67uz","id":"ZH3TP2dQBA","forumContent":{"TLDR":{"value":"Closed-form solution for Contiformer Using Damped Harmonic Oscillators"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["irregular time-series","physics-aware neural networks","magnetic navigation"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Physics Aware Neural Networks : Denoising for Magnetic Navigation\n\nMagnetic-anomaly navigation, which leverages small-scale local variations in the Earth's magnetic field, has emerged as a promising alternative for environments where GPS signals are unavailable or compromised. Airborne systems face a fundamental challenge in extracting the necessary geomagnetic field data: the aircraft itself induces magnetic noise, which needs to be removed. While the classical Tolles-Lawson model addresses this, it inadequately handles stochastically corrupted magnetic data needed for operational navigation.\nTo handle stochastic noise, we propose a novel approach using two physics-based constraints: divergence-free vector fields and E(3)-equivariance. Our constraints guarantee that the generated magnetic field obeys Maxwell's equations, and ensure that outputs change appropriately with sensor position/orientation. The divergence-free constraint is implemented by defining a neural network outputting vector potential A, with the magnetic field constructed as its curl. For E(3)-equivariance, we use tensor products of geometric tensors representable using spherical harmonics with well-known rotational transformations. By enforcing physical consistency and constraining the space of admissible field functions, our formulation acts as an implicit regularizer, improving spatiotemporal performance. \nWe conduct ablation studies evaluating these constraints' individual and combined effects across CNNs, MLPs, Liquid Time Constant Models, and Contiformers. We note that continuous-time dynamics and long-term memory are critical for modelling magnetic time-series data; the Contiformer architecture, which inherently possesses both, surpasses state-of-the-art methods in our experiments. To handle data scarcity, we develop synthetic datasets by utilising the World Magnetic Model (WMM) in conjunction with time-series conditional GANs, generating realistic and temporally consistent magnetic field sequences spanning various trajectory patterns and environmental scenarios. Our experiments demonstrate that embedding these constraints significantly improves predictive accuracy and physical plausibility outperforming the state-of-the-art across both classical and unconstrained deep learning approaches."},"_bibtex":{"value":"@misc{\ngupta2026physics,\ntitle={Physics Aware Neural Networks : Denoising for Magnetic Navigation},\nauthor={Debayan Gupta and ARITRA DAS and Yashas Shende and Arghya Pathak and Reva Laxmi Chauhan and Muskaan Chugh},\nyear={2026},\nurl={https://openreview.net/forum?id=WlZZcV67uz}\n}"},"title":{"value":"Physics Aware Neural Networks : Denoising for Magnetic Navigation"},"pdf":{"value":"/pdf/82fdfd6b65a14f8c1eae9fe2777d1f50e4ee7e7b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"gupta|physics_aware_neural_networks_denoising_for_magnetic_navigation"},"authorids":{"value":["~Debayan_Gupta1","~ARITRA_DAS2","~Yashas_Shende1","~Arghya_Pathak1","~Reva_Laxmi_Chauhan1","~Muskaan_Chugh1"]},"authors":{"value":["Debayan Gupta","ARITRA DAS","Yashas Shende","Arghya Pathak","Reva Laxmi Chauhan","Muskaan Chugh"]}},"version":2},{"content":{"summary":{"value":"This paper introduced a Hessian-Free Natural Gradient Descent (HF-NGD) framework for Physics-Informed Machine Learning including PINNs and neural operators. The method is matrix-free, with additional memory cost for iterations. The method achieves state-of-the-art performance on different benchmarks, 1~2 orders improvement compared to SGD and LBFGS , and over a 2-orders improvement on neural operators. This paper is the first to apply low-rank preconditioning in Gauss-Newton methods for PINN and neural operator optimization, making contributions to the community of physics-informed machine learning."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"see in the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The originality comes from the combination of preconditioning technique, the Matrix-free Hessian approximation, PCG and line search, making an explicit experimental improvement on the prediction accuracy. The comparison to UGNNG and DGNNG shows the importance of preconditioning.\n2. The paper is well-written and easy to follow, both theoretical and experimental richly.\n3. The paper makes contribution to the second-order optimization techniques on physics-informed machine learning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The contributions should be clarified more clearly compared to existing second order optimization techniques. The title is “Hessian free …” but there are works without explicitly forming the Hessian matrix, and the improvements of this paper may mainly comes from the preconditioning?\n\n2.Since the method is more like a combination of mathematical techniques, the ablation study should be reported, to show the independent importance of preconditioning, matrix-free computation cost, and so on.\n\n3.Section 3.2 and 3.3 are mathematically expressed independently, and the relation to Hessian matrix and physics-informed machine learning should be clarified more clearly to make the paper more readable."}},"nonreaders":[],"tmdate":1731428990147,"tcdate":1729392378989,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8803/Reviewer_7CZX"],"signatures":["ICLR.cc/2025/Conference/Submission8803/Reviewer_7CZX"],"forum":"Oqk1Ui6m0n","number":1,"license":"CC BY 4.0","cdate":1729392378989,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8803/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428990147,"domain":"ICLR.cc/2025/Conference","replyto":"Oqk1Ui6m0n","id":"XKieNV58Co","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"We propose a scalable Hessian-Free Natural Gradient Descent with matrix-free Hessian approximations and preconditioning, achieving state-of-the-art performance for large neural networks in Physics-Informed Machine Learning."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["PINNs","Gauss-Newton","Function-space optimization","Hessian-Free Optimization","Second-order Optimizers"]},"supplementary_material":{"value":"/attachment/04b1ec3cca72bdecc3f98a8b1727773aeccac4ff.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Physics-Informed Machine Learning (PIML) methods, such as Physics-Informed Neural Networks (PINNs), are notoriously difficult to optimize. Recent advances utilizing second-order optimization techniques, including natural gradient and Gauss-Newton methods, have significantly improved training accuracy over first-order methods. However, these approaches are computationally prohibitive, as they require evaluating, storing, and inverting large curvature matrices, limiting scalability to small networks. To overcome this limitation, we introduce a Hessian-Free Natural Gradient Descent framework that employs a matrix-free approximation of the Hessian. This approach circumvents the need for explicitly constructing the Hessian matrix and incorporates a novel preconditioning scheme that significantly enhances convergence rates. Our method enables scaling to large neural networks with  up to a million of parameters. Empirically, we demonstrate that our approach outperforms state-of-the-art optimizers, such as LBFGS and Adam, achieving orders-of-magnitudes accuracy improvements across various benchmark PDE problems."},"_bibtex":{"value":"@misc{\njnini2024hessianfree,\ntitle={Hessian-Free Natural Gradient Descent for Physics Informed Machine Learning},\nauthor={Anas Jnini and Flavio Vella},\nyear={2024},\nurl={https://openreview.net/forum?id=Oqk1Ui6m0n}\n}"},"title":{"value":"Hessian-Free Natural Gradient Descent for Physics Informed Machine Learning"},"pdf":{"value":"/pdf/2ec601a985178e5a24fcdcf801a8c56ef2ceabc1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"jnini|hessianfree_natural_gradient_descent_for_physics_informed_machine_learning"},"authorids":{"value":["~Anas_Jnini1","~Flavio_Vella1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Anas Jnini","Flavio Vella"]}},"version":2},{"content":{"summary":{"value":"This paper identifies a core gap in MLLM's ability to do human-like visual reasoning, which they term as _visual knowledge_. This refers to intuitive principles which humans use freely to understand the world, like intuitive physics and social cues. Based on this, the paper proposes two main contributions: \n1. A new benchmark, **VKBench**, which is curated from a set of existing datasets (IntPhys 2, PACS, VSI-Bench, VLEP, Social-IQ 2.0, RexTime) to benchmark MLLM's visual knowledge across world and human-centric axes. It consists of 1,680 multiple-choice questions across 1,249 videos, covering eight distinct types of visual knowledge. It was carefully constructed to avoid audio and linguistic biases. \n2. A new dataset, **VKQA**, of visual knowledge video examples, as well as method, **Video-VK+**, demonstrating that visual knowledge can be taught to these models. The authors utilize a \"See-Think-Answer\" format with RL to enforce the model to first visually process the input before making deductions. This is done on Qwen-2.5-VL-7B Instruction backbone, and this generally improves the models by 4.58% on average on their suite of benchmarks.\n\nThe authors benchmark many state of the art models on **VKBench**, and find that models still fall behind human performance (15% at best), especially at Intuitive Physics and Spatial Awareness."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. It's not quite clear to me from the text if the QAs are also pulled from the existing datasets, or if only the annotations are used to then synthesize new QAs. Clarification on this would be helpful. \n\nI think the dataset needs to go under significant revisions for this to be helpful for the community. While there clearly has been a large amount of detailed effort invested, I don't think that this benchmark is of critical value given the already immense number of video benchmarks which exist. It seems to be in a niche which is already covered by other benchmarks (given that the questions are collected from other popular datasets), and even the filtering done does not have validation to show that this provides a meaningful improvement over alternatives. Therefore I recommend a reject based on my current review of the paper."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"**Well-curated benchmark methodology**: VKBench was collected from existing datasets, so while there is no new contribution of base data on this front, there was consideration to how biased questions could be to audio and language, which are important problems present in recent benchmarks. There was also human validation done on the questions, in addition to shuffling the choices. \n\n**Simultaneous proposal of problem, benchmark, and method**: The authors not only codify a problem, but they also propose a benchmark, dataset, and method to solve it. This is a fair undertaking which shows that the overarching paper was thought out well in advance. \n\n**Nice figures and writing**: The figures are colorful and illustrative, and supplement the main text very well. Overall the writing of the paper is clear apart from minor details not described."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Benchmark is not difficult nor prescriptive**: While the benchmark claims to evaluate visual knowledge, it's not clear to me what information I gain by benchmarking my model on VKBench. First, benchmarks currently introduced where SoTA models achieve 71% accuracy is not helpful, as I suspect such a gap can be closed quite quickly (especially as it seems to be implied that this is a knowledge issue from the methods section). Given that random chance is so high for many questions, I find it surprising the authors didn't add in more options or leave the questions as short answers. This to me also points to how the current division of visual knowledge axes is too coarse, as models are somewhat uniform across them, and I can't tell where a model specifically lacks when it performs poorly. I would like more feedback on where models tend to fail, or _how they do_, rather than attributing it to being a knowledge or processing issue. \n\nRather than looking at Pearson correlations within the benchmark itself, you should compare to other benchmarks to see how well being good at X task correlates with having a strong world simulator in your model, or social understanding (of which there exist other benchmarks already). This would provide stronger validation that this benchmark is indeed prescriptive. \n\n**Weak method contributions beyond GRPO-Zero**: From what I can tell, it seems that the majority of the contribution of the method lies with GRPO(-Zero), which is prior work. For instance, MVBench, Video-MME, and MMVU are very similar between GRPO-Zero and Video-VK+, while I suspect VKBench may be more knowledge-based as a benchmark, which is why it improves better when incorporating extra data from VKQA. In fact, the See-Think-Answer SFT does not improve VKBench at all, while only providing modest contributions for the rest of the benchmarks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915777102,"tcdate":1762122624837,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1472/Reviewer_XciW"],"signatures":["ICLR.cc/2026/Conference/Submission1472/Reviewer_XciW"],"forum":"P798W8Ag7L","number":4,"license":"CC BY 4.0","cdate":1762122624837,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1472/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915777102,"domain":"ICLR.cc/2026/Conference","replyto":"P798W8Ag7L","id":"U8VEkHIb2d","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multimodal Large Language Model","Video Large Language Model","Visual Knowledge","Benchmark","Datasets"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical and social principles. This capability, which we term visual knowledge, forms a bridge between perception and reasoning, yet remains an underexplored gap in current systems.\nTo systematically measure this capability, we present VKBench, a comprehensive video benchmark featuring 1,680 questions in 1,249 videos, covering eight core types of visual knowledge spanning both world-centric (e.g., intuitive physics) and human-centric (e.g., subjective intentions). Results show that leading models still fall short of human performance, with particularly notable gaps in world-centric visual knowledge.\nTo bridge this gap, we introduce VKQA, a new dataset, and Video-VK+, a baseline model that explicitly incorporates visual knowledge into MLLMs. Video-VK+ follows a structured See–Think–Answer format and adopts reinforcement learning with visual knowledge reward. This approach improves performance on VKBench by 3.7% and surpasses existing models on multiple video benchmarks.\nOur findings highlight visual knowledge as a key component for developing more robust and generalizable MLLMs that can not only see but also truly understand our world."},"_bibtex":{"value":"@misc{\njiang2025benchmarking,\ntitle={Benchmarking Visual Knowledge in Multimodal Large Language Models},\nauthor={Tianxiang Jiang and Sheng Xia and Yicheng Xu and Linquan Wu and Xiangyu Zeng and Limin Wang and Yu Qiao and Yi Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=P798W8Ag7L}\n}"},"title":{"value":"Benchmarking Visual Knowledge in Multimodal Large Language Models"},"pdf":{"value":"/pdf/186de1c1275cf653df49719895560101bf3f1938.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"jiang|benchmarking_visual_knowledge_in_multimodal_large_language_models"},"authorids":{"value":["~Tianxiang_Jiang1","~Sheng_Xia1","~Yicheng_Xu1","~Linquan_Wu1","~Xiangyu_Zeng4","~Limin_Wang1","~Yu_Qiao1","~Yi_Wang19"]},"authors":{"value":["Tianxiang Jiang","Sheng Xia","Yicheng Xu","Linquan Wu","Xiangyu Zeng","Limin Wang","Yu Qiao","Yi Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Directed Cellular Sheaves and a corresponding Directed Sheaf Laplacian, enabling Sheaf Neural Networks to incorporate edge directionality, a missing capability in current SNNs. The authors prove key spectral properties and show that this formulation recovers classical sheaf Laplacians and magnetic Laplacians as special cases. They further propose DSNN, demonstrating consistent gains on both real-world graphs and synthetic directional SBM settings."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. Is $q$ learned or tuned per dataset? If tuned, how stable is performance across $q$ values?\n\n2. Does DSNN maintain benefits in settings where directional edges are sparse or only weakly informative?\n\n3. How does performance degrade if a portion of edge directions are flipped or randomized?\n\n4. Why the random split is adopted for the node classification instead of the widely-used public splits? Performance on small homophilic and heterophilic benchmarks can vary noticeably with random seeds, so it would be useful to justify this choice and clarify whether public splits are also tested."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"**1. Principled directional sheaf formulation**\n\nThe paper introduces directed cellular sheaves and a corresponding directed sheaf Laplacian, providing the first rigorous sheaf-theoretic framework for directed graphs and addressing a clear limitation of existing SNNs.\n\n**2. Solid theoretical foundation**\n\nThe authors prove Hermiticity, PSD spectrum bounds, and show that the proposed operator recovers classical sheaf Laplacians and magnetic Laplacians as special cases, demonstrating a sound and unifying mathematical design.\n\n**3. Comprehensive experimental results**\n\nAcross both real-world and synthetic benchmarks, the model outperforms existing SNNs and competitive direction-aware GNNs, with especially strong results in heterophilic and direction-dominated settings, validating the benefits of directional sheaf modeling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Limited intuition for the directional mechanism**\n\nWhile the mathematical construction is provided, the paper provides limited high-level insight into **how and why** the complex restriction maps enhance directional information flow in practice. The introduction of the complex phase feels algebraically motivated rather than guided by an intuitive model of directional propagation.\n\n**2. Scope of experimental evaluation**\n\nThe evaluation focuses primarily on small-to-medium-scale datasets. There is no demonstration on larger real-world directed benchmarks (e.g., OGB-ArXiv, and arxiv-year). \n\n**3. Ablations could be deeper**\n\n3.1. The effect of stalk dimension $d$\n\n3.2. Sensitivity to direction sparsity or unreliable edge orientation (i.e., direction noise)\n\n3.3. Effect of learning vs. fixing the phase $q$\n\n**4. Writing clarity**\n\nThe definition and construction of the directed cellular sheaf are mathematically sound but presented in a dense, notation-heavy manner. Adding intuitive explanations, intermediate steps, and conceptual guidance (e.g., how complex phases encode directional flow at a high level) would make the framework more accessible and easier to follow for a broader audience beyond sheaf specialists."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917906435,"tcdate":1761839119005,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5143/Reviewer_47m4"],"signatures":["ICLR.cc/2026/Conference/Submission5143/Reviewer_47m4"],"forum":"iDiiETH7Qv","number":3,"license":"CC BY 4.0","cdate":1761839119005,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5143/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917906435,"domain":"ICLR.cc/2026/Conference","replyto":"iDiiETH7Qv","id":"fOcvN1yF8V","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["directed sheaf neural network","directed graphs","directed cellular sheaves"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Sheaf Neural Networks (SNNs) are a powerful algebraic-topology generalization of Graph Neural Networks (GNNs), and have been shown to significantly improve our ability to model complex relational data. While the GNN literature proved that incorporating directionality can substantially boost performance in many real-world applications, no SNNs approaches are known with such a capability. To address this limitation, we introduce the Directed Cellular Sheaf, a generalized cellular sheaf designed to explicitly account for edge orientations. Building on it, we define a corresponding sheaf Laplacian, the Directed Sheaf Laplacian $L^{\\widetilde{\\mathcal{F}}}$, which exploits the sheaf's structure to capture both the graph’s topology and its directions. $L^{\\widetilde{\\mathcal{F}}}$ serves as the backbone of the Directed Sheaf Neural Network (DSNN), the first SNN model to embed a directional bias into its architecture. Extensive experiments on twelve real-world benchmarks show that DSNN consistently outperforms many baseline methods. The source\ncode can be found at https://github.com/hakanaktas0/DSNN."},"_bibtex":{"value":"@inproceedings{\nfiorini2026sheaves,\ntitle={Sheaves Reloaded: A Direction Awakening},\nauthor={Stefano Fiorini and Hakan Aktas and Iulia Duta and Pietro Morerio and Alessio Del Bue and Pietro Lio and Stefano Coniglio},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=iDiiETH7Qv}\n}"},"title":{"value":"Sheaves Reloaded: A Direction Awakening"},"pdf":{"value":"/pdf/57e845fda6d56b487ec365a89e5b7211f789a44c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"fiorini|sheaves_reloaded_a_direction_awakening"},"authorids":{"value":["~Stefano_Fiorini1","~Hakan_Aktas2","~Iulia_Duta1","~Pietro_Morerio1","~Alessio_Del_Bue2","~Pietro_Lio1","~Stefano_Coniglio2"]},"authors":{"value":["Stefano Fiorini","Hakan Aktas","Iulia Duta","Pietro Morerio","Alessio Del Bue","Pietro Lio","Stefano Coniglio"]}},"version":2},{"content":{"venue":{"value":"BIOSTEC (2) 2024"},"venueid":{"value":"dblp.org/conf/BIOSTEC/2024"},"paperhash":{"value":"santangelo|synthcheck_a_dashboard_for_synthetic_data_quality_assessment"},"authorids":{"value":["~Gabriele_Santangelo1","","",""]},"html":{"value":"https://doi.org/10.5220/0012558700003657"},"_bibtex":{"value":"@inproceedings{DBLP:conf/biostec/SantangeloNBD24,\n  author={Gabriele Santangelo and Giovanna Nicora and Riccardo Bellazzi and Arianna Dagliati},\n  title={SynthCheck: A Dashboard for Synthetic Data Quality Assessment},\n  year={2024},\n  cdate={1704067200000},\n  pages={246-256},\n  url={https://doi.org/10.5220/0012558700003657},\n  booktitle={BIOSTEC (2)},\n  crossref={conf/biostec/2024-2}\n}\n"},"abstract":{"value":"In recent years, synthetic data generation has become a topic of growing interest, especially in healthcare, where they can support the development of robust Artificial Intelligence (AI) tools. Additionally, synthetic data offer advantages such as easier sharing and consultation compared to original data, which are subject to patient privacy laws that have become increasingly rigorous in recent years. To ensure a safe use of synthetic data, it is necessary to assess their quality. Synthetic data quality evaluation is based on three properties: resemblance, utility, and privacy, that can be measured using different statistical approaches. Automatic evaluation of synthetic data quality can foster their safe usage within medical AI systems. For this reason, we have developed a dashboard application, in which users can perform a comprehensive quality assessment of their synthetic data. This is achieved through a user-friendly interface, providing easy access and intuitive functionalities"},"title":{"value":"SynthCheck: A Dashboard for Synthetic Data Quality Assessment"},"authors":{"value":["Gabriele Santangelo","Giovanna Nicora","Riccardo Bellazzi","Arianna Dagliati"]}},"tmdate":1772592485607,"pdate":1735603200000,"externalIds":["dblp:conf/biostec/SantangeloNBD24"],"tcdate":1772592479642,"writers":["~"],"signatures":["~Gabriele_Santangelo1"],"forum":"5sBoCqQKqt","license":"CC BY-SA 4.0","number":847327,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772592485607,"domain":"DBLP.org","id":"5sBoCqQKqt","version":2},{"content":{"venue":{"value":"Complex (1) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02466-5_3.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"feng|optimization_using_a_new_bioinspired_approach"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Francis_C._M._Lau_0001:","https://dblp.org/search/pid/api?q=author:Daqi_Gao:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02466-5_3"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/FengLG09,\n  author={Xiang Feng and Francis C. M. Lau and Daqi Gao},\n  title={Optimization Using a New Bio-inspired Approach},\n  year={2009},\n  cdate={1230768000000},\n  pages={39-51},\n  url={https://doi.org/10.1007/978-3-642-02466-5_3},\n  booktitle={Complex (1)},\n  crossref={conf/complex/2009-1}\n}\n"},"abstract":{"value":"There is growing interest in bio(logy)-inspired approaches that are inspired by the principles of biology and that can solve difficult problems. In this paper, we propose a new computational algorithm that is inspired by molecular mechanics for the solution of complex problems. There is a deep and useful connection between mechanics mechanics and combinatorial optimization. This connection exposes new information and allows an unfamiliar perspective on traditional optimization problems and approaches. The alternative of molecular mechanics algorithm (MMA) to traditional approaches has the advantages of inherent parallelism and the ability to deal with a variety of complicated social interactions, autonomous behaviors and multiple objectives."},"title":{"value":"Optimization Using a New Bio-inspired Approach"},"authors":{"value":["Xiang Feng","Francis C. M. Lau","Daqi Gao"]}},"tmdate":1744105802920,"pdate":1230768000000,"tcdate":1744105521742,"writers":["~"],"signatures":["~Xiang_Feng2"],"forum":"7jkTxwxJVt","license":"CC BY-SA 4.0","number":382270,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1744105802920,"domain":"DBLP.org","id":"7jkTxwxJVt","version":2},{"content":{"TLDR":{"value":"This paper introduces two low-level tasks to test intuitive physics understanding and proposes Scene Dynamic Field, a method to integrate visual representation from physics simulators to MLLMs while showcasing generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multi-modal LLM","Intuitive Physics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. \nIn this work, we investigate the first step of physical reasoning, i.e., **intuitive physics understanding**, revealing substantial limitations in understanding the dynamics of continuum objects. \nTo isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. \nTo address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. \nSDF substantially improves performance, achieving up to $20.7\\%$ gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field."},"_bibtex":{"value":"@inproceedings{\nli2026beyond,\ntitle={Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models},\nauthor={Nanxi Li and Xiang Wang and Yuanjie Chen and Haode Zhang and Hong Li and Yong-Lu Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Ax02eR2c3d}\n}"},"title":{"value":"Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models"},"pdf":{"value":"/pdf/3b199b74bbdbddb8166b5e799fc00ddb71a413ae.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|beyond_static_vision_scene_dynamic_field_unlocks_intuitive_physics_understanding_in_multimodal_large_language_models"},"authorids":{"value":["~Nanxi_Li1","~Xiang_Wang19","~Yuanjie_Chen2","~Haode_Zhang2","~Hong_Li5","~Yong-Lu_Li1"]},"authors":{"value":["Nanxi Li","Xiang Wang","Yuanjie Chen","Haode Zhang","Hong Li","Yong-Lu Li"]}},"tmdate":1775877009051,"pdate":1769435910907,"tcdate":1758034327732,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7741/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission7741/Authors"],"forum":"Ax02eR2c3d","license":"CC BY 4.0","number":7741,"cdate":1758034327732,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission7741/-/Full_Submission","ICLR.cc/2026/Conference/Submission7741/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission7741/-/Camera_Ready_Revision"],"mdate":1775877009051,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"Ax02eR2c3d","version":2},{"content":{"summary":{"value":"This paper systematically investigates how Transformer language models learn and process complex recursive language structures by designing a set of synthetic context-free grammars (CFGs) with long-range dependencies and local ambiguities. The main contributions are: (1) introducing a multi-head linear probing method to analyze internal representations; (2) developing new visualization and quantification tools to understand attention patterns; (3) demonstrating that GPT models can learn complex CFGs in a dynamic programming-like manner and revealing the critical role of boundary-based attention in handling long-range dependencies; (4) validating the model's effectiveness and robustness in handling such complex language structures. These findings offer new insights into the inner workings of large language models."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"The paper suggests that Transformer models learn CFGs by emulating dynamic programming (DP) algorithms. Would it be possible to test this hypothesis more directly, perhaps by comparing the model’s internal representations to the states of known DP algorithms? The current evidence feels somewhat indirect.\n\nThe paper also notes BERT’s weaker performance in predicting deep non-terminal (NT) structures but doesn’t delve into the reasons behind this. While attributing it to the locality of the MLM task seems plausible, more direct evidence would strengthen this claim.\n\nCould the authors discuss potential limitations? Acknowledging the method’s constraints would be helpful for future research directions."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"The highlight of this paper is its novel approach to studying large language models by using carefully designed synthetic CFGs to explore model mechanisms. This method is innovative because it offers a controlled, quantifiable experimental environment, making complex model behaviors interpretable. By demonstrating how CFGs can be used to investigate the learning processes of models, this work provides a powerful analytical framework for future researchers. Combining theoretical analysis with experimental design, this approach not only aids in understanding existing models but also inspires new directions for probing the inner workings of even larger language models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper mainly focuses on studying canonical CFGs, and the exploration of non-canonical CFGs is not sufficiently in-depth.\n\nThe research results are primarily based on specific synthetic CFGs, and their generalizability needs to be further verified."}},"nonreaders":[],"tmdate":1731429193273,"tcdate":1730721627054,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13213/Reviewer_Mj12"],"signatures":["ICLR.cc/2025/Conference/Submission13213/Reviewer_Mj12"],"forum":"J6qrIjTzoM","number":4,"license":"CC BY 4.0","cdate":1730721627054,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13213/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429193273,"domain":"ICLR.cc/2025/Conference","replyto":"J6qrIjTzoM","id":"sTVfvKNJ7S","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["generative language models","interpretability","induction head","inner workings"]},"supplementary_material":{"value":"/attachment/8c1c0e398b6a253a61554122b9ccc7090686bdf9.zip"},"primary_area":{"value":"interpretability and explainable AI"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Transformer-based language models are effective but complex, and understanding their inner workings is a significant challenge. Previous research has primarily explored how these models handle simple tasks like name copying or selection, and we extend this by investigating how these models grasp complex, recursive language structures defined by context-free grammars (CFGs). We introduce a family of synthetic CFGs that produce hierarchical rules, capable of generating lengthy sentences (e.g., hundreds of tokens) that are locally ambiguous and require dynamic programming to parse. Despite this complexity, we demonstrate that generative models like GPT can accurately learn this CFG language and generate sentences based on it. We explore the model's internals, revealing that its hidden states precisely capture the structure of CFGs, and its attention patterns resemble the information passing in a dynamic programming algorithm."},"_bibtex":{"value":"@misc{\nallen-zhu2025interpretability,\ntitle={Interpretability of Language Models for Learning Hierarchical Structures},\nauthor={Zeyuan Allen-Zhu and Yuanzhi Li},\nyear={2025},\nurl={https://openreview.net/forum?id=J6qrIjTzoM}\n}"},"title":{"value":"Interpretability of Language Models for Learning Hierarchical Structures"},"pdf":{"value":"/pdf/5048fd3734b9eb68785aef352d991a6fe73893d1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"allenzhu|interpretability_of_language_models_for_learning_hierarchical_structures"},"authorids":{"value":["~Zeyuan_Allen-Zhu1","~Yuanzhi_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyuan Allen-Zhu","Yuanzhi Li"]}},"version":2},{"content":{"venue":{"value":"CoRR 2020"},"pdf":{"value":"http://arxiv.org/pdf/2011.07193v2"},"venueid":{"value":"dblp.org/journals/CORR/2020"},"paperhash":{"value":"ota|towards_humanlevel_learning_of_complex_physical_puzzles"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Kei_Ota:","~Devesh_K._Jha1","https://dblp.org/search/pid/api?q=author:Diego_Romeres:","https://dblp.org/search/pid/api?q=author:Jeroen_van_Baar:","https://dblp.org/search/pid/api?q=author:Kevin_A._Smith:","https://dblp.org/search/pid/api?q=author:Takayuki_Semitsu:","https://dblp.org/search/pid/api?q=author:Tomoaki_Oiki:","https://dblp.org/search/pid/api?q=author:Alan_Sullivan:","https://dblp.org/search/pid/api?q=author:Daniel_Nikovski:","https://dblp.org/search/pid/api?q=author:Joshua_B._Tenenbaum:"]},"html":{"value":"https://arxiv.org/abs/2011.07193"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2011-07193,\n  publtype={informal},\n  author={Kei Ota and Devesh K. Jha and Diego Romeres and Jeroen van Baar and Kevin A. Smith and Takayuki Semitsu and Tomoaki Oiki and Alan Sullivan and Daniel Nikovski and Joshua B. Tenenbaum},\n  title={Towards Human-Level Learning of Complex Physical Puzzles},\n  year={2020},\n  cdate={1577836800000},\n  journal={CoRR},\n  volume={abs/2011.07193},\n  url={https://arxiv.org/abs/2011.07193}\n}\n"},"abstract":{"value":"Humans quickly solve tasks in novel systems with complex dynamics, without requiring much interaction. While deep reinforcement learning algorithms have achieved tremendous success in many complex tasks, these algorithms need a large number of samples to learn meaningful policies. In this paper, we present a task for navigating a marble to the center of a circular maze. While this system is very intuitive and easy for humans to solve, it can be very difficult and inefficient for standard reinforcement learning algorithms to learn meaningful policies. We present a model that learns to move a marble in the complex environment within minutes of interacting with the real system. Learning consists of initializing a physics engine with parameters estimated using data from the real system. The error in the physics engine is then corrected using Gaussian process regression, which is used to model the residual between real observations and physics engine simulations. The physics engine augmented with the residual model is then used to control the marble in the maze environment using a model-predictive feedback over a receding horizon. To the best of our knowledge, this is the first time that a hybrid model consisting of a full physics engine along with a statistical function approximator has been used to control a complex physical system in real-time using nonlinear model-predictive control (NMPC)."},"title":{"value":"Towards Human-Level Learning of Complex Physical Puzzles"},"authors":{"value":["Kei Ota","Devesh K. Jha","Diego Romeres","Jeroen van Baar","Kevin A. Smith","Takayuki Semitsu","Tomoaki Oiki","Alan Sullivan","Daniel Nikovski","Joshua B. Tenenbaum"]}},"tmdate":1717809170614,"pdate":1577836800000,"tcdate":1717809142535,"writers":["~"],"signatures":["~Devesh_K._Jha1"],"forum":"RljeNVDQ5T","license":"CC BY-SA 4.0","number":25251,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1717809170614,"domain":"DBLP.org","id":"RljeNVDQ5T","version":2},{"content":{"venue":{"value":"Computer Physics Communications"},"pdf":{"value":"https://www.sciencedirect.com/science/article/pii/S0010465525001523/pdfft?md5=089e3a516a9309db6ec6f5f209cc6ba5&pid=1-s2.0-S0010465525001523-main.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"chandra|generalizable_models_of_magnetic_hysteresis_via_physicsaware_recurrent_neural_networks"},"html":{"value":"https://doi.org/10.1016/j.cpc.2025.109650"},"abstract":{"value":"Hysteresis is a ubiquitous phenomenon in magnetic materials; its modeling and identification are crucial for understanding and optimizing the behavior of electrical machines. Such machines often operate under uncertain conditions, necessitating modeling methods that can generalize across unobserved scenarios. Traditional recurrent neural architectures struggle to generalize hysteresis patterns beyond their training domains. This paper mitigates the generalization challenge by introducing a physics-aware recurrent neural network approach to model and generalize the hysteresis manifesting in sequentiality and history-dependence. The proposed method leverages ordinary differential equations (ODEs) governing the phenomenological hysteresis models to update hidden recurrent states. The effectiveness of the proposed method is evaluated by predicting generalized scenarios, including first-order reversal curves and minor loops. The results demonstrate robust generalization to previously untrained regions, even with noisy data, an essential feature that hysteresis models must have. The results highlight the advantages of integrating physics-based ODEs into recurrent architectures, including superior performance over traditional methods in capturing the complex, nonlinear hysteresis behaviors in magnetic materials. The codes and data related to the paper are at github.com/chandratue/HystRNN."},"title":{"value":"Generalizable models of magnetic hysteresis via physics-aware recurrent neural networks"},"authors":{"value":[{"fullname":"Abhishek Chandra","username":"~Abhishek_Chandra2"},{"fullname":"Taniya Kapoor"},{"fullname":"Bram Daniels"},{"fullname":"Mitrofan Curti"},{"fullname":"Koen Tiels"},{"fullname":"Daniel M. Tartakovsky"},{"fullname":"Elena A. Lomonova"}]}},"tmdate":1789090292569,"pdate":1756684800000,"externalIds":["doi:10.1016/j.cpc.2025.109650"],"tcdate":1763760244707,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Abhishek_Chandra2"],"forum":"pPvpTOkIPD","license":"CC BY-SA 4.0","number":17446,"cdate":1746489242401,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789090292569,"domain":"OpenReview.net/Public_Article","id":"pPvpTOkIPD","version":2},{"content":{"summary":{"value":"The paper proposes a Lorentz equivariant transformer (L-GATr) based on geometric algebra for high energy physics. It generalizes the Geometric Algebra Transformer (GATr) from $E(3)$ equivariance to the Lorentz group. The proposed transformer is then developed into a generative model based on Riemannian flow matching for particle data. L-GATr is evaluated on several high energy physics tasks, including quantum field theory amplitude surrogates, top tagging, and generative modeling for event reconstruction."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. What are $y_m, y_p, \\eta$, and $\\phi$ in (4)?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is well-written and easy to follow. In addition, the problem is well-motivated.\n2. The proposed L-GATr generalizes GATr from $E(3)$ to the Lorentz group.\n3. As stated by the authors, the Lorentz-equivariant flow matching proposed in section 3.2 is the first generative model proposed for particle physics. \n4. Compared to graph-based Lorentz equivariant networks, the transformer architecture is more efficient and scalable. \n5. The proposed method is shown to be more data efficient than the baselines in both amplitude surrogates and generative modeling experiments. \n6. The ability to scale with the data is also verified in several experiments. \n7. The benefits of Riemannian flow matching compared to the Euclidean version are demonstrated in the experiments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. It is a bit unclear to me what has been modified for Lorentz equivariance in the transformer framework. Specifically, (1) and (2) look the same as (4) and (5) in [13]. It’ll be great if the authors can state the changes explicitly. \n2. The current presentation of the experiments can be a bit hard to understand for people without a physics background. Adding some basic introduction to the problems can strengthen the paper. \n3. In the top tagging experiment, the performance of the proposed method is marginally worse than the baseline method. \n4. Although the proposed method is claimed to support symmetry-breaking data, its effect is not well studied in the experiments."},"limitations":{"value":"As mentioned in the paper, the proposed L-GATr has additional computational overhead compared to traditional transformers. Secondly, even though the framework allows additional inputs to address symmetry breaking issues, the effect of such an approach is not well studied. It is unclear how well the proposed method can handle symmetry breaking inputs."}},"nonreaders":[],"tmdate":1730878888655,"tcdate":1720797608520,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission3969/Reviewer_w7CH"],"signatures":["NeurIPS.cc/2024/Conference/Submission3969/Reviewer_w7CH"],"forum":"X34GKv8sYT","number":2,"license":"CC BY 4.0","cdate":1720797608520,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission3969/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878888655,"domain":"NeurIPS.cc/2024/Conference","replyto":"X34GKv8sYT","id":"LHZuDUgr8k","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"A Lorentz-equivariant Transformer architecture plus Lorentz-equivariant flow matching for high-energy physics"},"keywords":{"value":["Geometric deep learning","equivariance","Lorentz symmetry","Transformer","flow matching","high-energy physics","particle physics"]},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Extracting scientific understanding from particle-physics experiments requires solving diverse learning problems with high precision and good data efficiency. We propose the Lorentz Geometric Algebra Transformer (L-GATr), a new multi-purpose architecture for high-energy physics. L-GATr represents high-energy data in a geometric algebra over four-dimensional space-time and is equivariant under Lorentz transformations, the symmetry group of relativistic kinematics. At the same time, the architecture is a Transformer, which makes it versatile and scalable to large systems. L-GATr is first demonstrated on regression and classification tasks from particle physics. We then construct the first Lorentz-equivariant generative model: a continuous normalizing flow based on an L-GATr network, trained with Riemannian flow matching. Across our experiments, L-GATr is on par with or outperforms strong domain-specific baselines."},"_bibtex":{"value":"@inproceedings{\nspinner2024lorentzequivariant,\ntitle={Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics},\nauthor={Jonas Spinner and Victor Breso Pla and Pim De Haan and Tilman Plehn and Jesse Thaler and Johann Brehmer},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=X34GKv8sYT}\n}"},"title":{"value":"Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics"},"pdf":{"value":"/pdf/8ddc6739ac0f4a9405f3368dba8b88248133a915.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"spinner|lorentzequivariant_geometric_algebra_transformers_for_highenergy_physics"},"authorids":{"value":["~Jonas_Spinner1","~Victor_Breso_Pla1","~Pim_De_Haan1","~Tilman_Plehn1","~Jesse_Thaler1","~Johann_Brehmer1"]},"authors":{"value":["Jonas Spinner","Victor Breso Pla","Pim De Haan","Tilman Plehn","Jesse Thaler","Johann Brehmer"]}},"version":2},{"content":{"summary":{"value":"This paper proposes ThinkPlace, a vision-language-guided video editing framework that achieves physically and visually coherent custom subject integration without explicit physics simulation.\nThe system introduces three core components (1) a VLM-based Chain-of-Thought (CoT) reasoning module (2) a Spatial Direct Preference Optimization (Spatial-DPO) post-training stage; and (3) an iterative corrective editing loop. Built on VACE/Wan2.1 diffusion transformers and Qwen-VL2.5, ThinkPlace achieves improved physical plausibility and video realism on 200 custom-integration test videos."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the weakness section."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The Interaction CoT reasoning explicitly analyzes scene physics, lighting, and motion, generating structured spatial and semantic guidance before synthesis. \n\n- The Spatial-DPO and VLM-based preference evaluation remove the need for human annotation, a practical and scalable design validated by gains in PC/PR (physical realism) metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although ThinkPlace leads on all metrics, several improvements are numerically small and may fall within measurement variance. Therefore, (1) it is not clear how significant is the improvement of the method, and (2) Reporting standard deviations or significance tests would help substantiate the claimed superiority, especially for perceptual and identity metrics.\n\n- In this paper, the experiments focus on short, curated physics-aware scenes. It is unclear how the model behaves under open-domain or high-motion videos, e.g., handheld or complex lighting conditions. Including stress-test examples (hand-held camera, low-light, or multi-object scenes) would demonstrate robustness.\n\n- The entire system’s success depends on VLM inference accuracy—errors in bounding-box localization or physical reasoning could cascade. It should better quantify reasoning reliability, for example by reporting bounding-box IoU against manual annotations or showing failure cases where VLM misguidance causes errors (e.g., incorrect occlusion handling).\n\n- The iterative VLM post-evaluation adds 2–3 refinement cycles per video (Sec. 4.4), which can be computationally expensive. \n\n- While the qualitative figures are provided, user studies or pairwise preference tests could more convincingly validate perceived realism, especially since much of the novelty lies in subjective physical plausibility."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916459618,"tcdate":1761940703610,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2957/Reviewer_mqDX"],"signatures":["ICLR.cc/2026/Conference/Submission2957/Reviewer_mqDX"],"forum":"uYsok31soR","number":5,"license":"CC BY 4.0","cdate":1761940703610,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2957/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916459618,"domain":"ICLR.cc/2026/Conference","replyto":"uYsok31soR","id":"tWUehxsJy2","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Custom Subject Integration","Vision-Language Models","Direct Preference Optimization"]},"supplementary_material":{"value":"/attachment/62c023ab57d433efecb98ef6f27860c21d7c4c61.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Contemporary video editing methods have achieved remarkable visual fidelity for custom subject integration, yet they fundamentally lack the capability to model causally realistic interactions between inserted objects and their environments. This limitation results in physically implausible editing outcomes, violating basic physical laws. \nIn this work, we present ThinkPlace, an end-to-end framework that addresses these challenges by leveraging Vision-Language Models (VLM) as a reasoning brain to guide physically-aware video editing without explicit physics simulation. Our approach introduces three key innovations: First, we develop a VLM-guided chain-of-thought reasoning pipeline that generates environment-aware guidance tokens while providing physically plausible editing regions for the downstream video diffusion model. Second, we introduce a Spatial Direct Preference Optimization post-training which also employs VLM for enhancing visual naturalness of editing results.\nThird, we leverage VLM for post-evaluation, triggering corrective refinement cycles that progressively improves integration quality.\nExtensive experiments demonstrate ThinkPlace achieves physically-coherent custom subject integration compared with State-of-the-art solutions. Our work represents a significant step toward bridging the gap between visual quality and physical realism in video editing applications."},"_bibtex":{"value":"@misc{\ngu2025think,\ntitle={Think Before You Place: Chain-of-Thought Video Editing for Environment-Aware Custom Subject Integration},\nauthor={Bohai Gu and Taiyi Wu and Dazhao Du and Alan Zhao and Song Guo},\nyear={2025},\nurl={https://openreview.net/forum?id=uYsok31soR}\n}"},"title":{"value":"Think Before You Place: Chain-of-Thought Video Editing for Environment-Aware Custom Subject Integration"},"pdf":{"value":"/pdf/0f2e4ec94596276ed5feeb83de6681f49c47d2f8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"gu|think_before_you_place_chainofthought_video_editing_for_environmentaware_custom_subject_integration"},"authorids":{"value":["~Bohai_Gu1","~Taiyi_Wu1","~Dazhao_Du2","~Alan_Zhao1","~Song_Guo5"]},"authors":{"value":["Bohai Gu","Taiyi Wu","Dazhao Du","Alan Zhao","Song Guo"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a physics-encoded inverse modeling approach to estimate Arctic snow depth from reanalysis inputs. A sequence model (RNN/attention + MLP head) is coupled to a linearized hydrostatic relation and regularized with a contrastive term built from noise-perturbed views of the inputs. On a central-Arctic, daily ERA-style dataset, the method beats several neural baselines and tracks seasonal dynamics."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Please add baselines from (i) PINNs/physics-loss regularized nets; (ii) a simple hydrostatic-only regression. This will help isolate the gain from your inverse head vs. generic physics regularization.\n\n2. For ablations, please compare between removing each loss term (prediction, physics, contrastive) and each architectural block (attention, inverse-param head).\n\n3. What independent observations of snow depth (e.g., ICESat/ICESat-2 freeboard-derived products, Operation IceBridge, in-situ buoys) can you use to validate the learned model and rule out circularity with Eq. (1)? If none, please justify why the proxy target suffices. Else, please train on one region/period, test OOD on another.\n\n4. Please precisely define the target space and explain how your architecture or loss enforces a surjective inverse (or provide a counterexample/limitation). Any theoretical or empirical evidence (e.g., coverage metrics over plausible parameter sets)?\n\n5. Clarify why calling the objective “supervised” is appropriate when positives are noise-perturbed views; discuss sensitivity to $\\tau$, batch size, and negative sampling strategy."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. **Physics-guided objective:** Encoding a (linearized) hydrostatic relation into the learning target promotes physically plausible estimates rather than unconstrained regression.\n\n2. **Signal in data-scarce regimes:** The contrastive regularizer appears to stabilize training when labels/coverage are limited."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Foundation-model baselines for Earth/physics tasks:** There’s no comparison to ClimaX [1] (a pre-trained foundation model for weather/climate that adapts to heterogeneous variables and scales), which already shows strong transfer on ERA/CMIP tasks after lightweight fine-tuning. Even a small ClimaX adapter fine-tuned to snow-relevant channels would be an informative baseline, or a complementary pretrain-then-physics-tune strategy.\n\n2. **External truth & circularity.** The target appears (from the description) to be derived from the same hydrostatic relation the model encodes. That risks **learning to the proxy**, which basically means evaluating the model on the same training set. Independent validation against satellite/freeboard-derived snow products, OIB tracks, or buoy stations is critical to establish real-world value, not just self-consistency.\n\n3. **Physics-informed peers beyond vanilla nets.** The baselines don’t include modern physics-aware methods (PINNs [3], neural operators with constraints, hybrid data-assimilation losses). Given ClimaX [1] and PhysiX [2] show strong data-driven physics fidelity at scale, this weakens the claims of state-of-the-art for this work.\n\n4. **Supervised contrastive clarity.** Using noise-perturbed views as positives is data-augmentation/self-supervised flavored; it would help to justify the “supervised” label and analyze sensitivity to temperature, batch size, and negative sampling.\n\n\n[1] Nguyen, Tung, et al. \"Climax: A foundation model for weather and climate.\" arXiv preprint arXiv:2301.10343 (2023).\n\n[2] Nguyen, Tung, et al. \"PhysiX: A Foundation Model for Physics Simulations.\" arXiv preprint arXiv:2506.17774 (2025).\n\n[3] Raissi, Maziar, Paris Perdikaris, and George E. Karniadakis. \"Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations.\" Journal of Computational physics 378 (2019): 686-707."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931632724,"tcdate":1761844498672,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19786/Reviewer_t4ue"],"signatures":["ICLR.cc/2026/Conference/Submission19786/Reviewer_t4ue"],"forum":"49JcR9oeoo","number":1,"license":"CC BY 4.0","cdate":1761844498672,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19786/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931632724,"domain":"ICLR.cc/2026/Conference","replyto":"49JcR9oeoo","id":"DeFt3BAATP","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Inverse modeling","Supervised representation learning","Surjective mapping","Arctic snow depth","Knowledge guidance"]},"supplementary_material":{"value":"/attachment/bf47f34f6a8ea7e7f0ff7790c09f265da7a14919.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"The accurate estimation of Arctic snow depth ($h_s$) remains a critical time-varying inverse problem due to the extreme scarcity and noise inherent in associated sea ice parameters. Existing process-based and data-driven models are either highly sensitive to sparse data or lack the physical interpretability required for climate-critical applications. To address this gap, we introduce PhysE-Inv, a novel framework that integrates a sophisticated sequential architecture, an LSTM Encoder-Decoder with Multi-head Attention and physics-guided contrastive learning, with physics-guided inference.Our core innovation lies in a surjective, physics-constrained inversion methodology. This methodology first leverages the hydrostatic balance forward model as a target-formulation proxy, enabling effective learning in the absence of direct $h_s$ ground truth; second, it uses reconstruction physics regularization over a latent space to dynamically discover hidden physical parameters from noisy, incomplete time-series input. Evaluated against state-of-the-art baselines, PhysE-Inv significantly improves prediction performance, reducing error by 20\\% while demonstrating superior physical consistency and resilience to data sparsity compared to empirical methods. This approach pioneers a path for noise-tolerant, interpretable inverse modeling, with wide applicability in geospatial and cryospheric domains."},"_bibtex":{"value":"@misc{\nsampath2026physerlinv,\ntitle={Phys{ERL}-Inv: A Physics-Encoded Inverse Modeling Approach for Arctic Snow Depth Prediction},\nauthor={Akila Sampath and Vandana Janeja and Jianwu Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=49JcR9oeoo}\n}"},"title":{"value":"PhysERL-Inv: A Physics-Encoded Inverse Modeling Approach for Arctic Snow Depth Prediction"},"pdf":{"value":"/pdf/7871ac9fc35149e3f88a3a37b1d92d3464bce152.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sampath|physerlinv_a_physicsencoded_inverse_modeling_approach_for_arctic_snow_depth_prediction"},"authorids":{"value":["~Akila_Sampath1","~Vandana_Janeja1","~Jianwu_Wang1"]},"authors":{"value":["Akila Sampath","Vandana Janeja","Jianwu Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a symmetric invariant data structure, the Symbolic Graph, replacing expression trees in symbolic regressions. It also develops a GNN to encode physics constraints and guide the MCTS algorithm. Numerical experiments on two benchmarks and one synthetic dataset show the performance improvement of the proposed method."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Why were the NNGP and SPL baselines not compared in the real-world experiment (Table 3)?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The presentation, visualization, and the design of real-world experiments.\n2. The idea of handling symmetric invariance by a new data structure."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The necessity of GNN. It seems that the proposed symbolic graph(SG) can be viewed as a tree with more branches and special nodes (e.g. summation, directed minus). If this is true, then the SG (with trivial modifications) can also be processed by existing non-graph neural networks designed for expression trees.\n2. The necessity of physics-constrained NN. It seems that the NN is self-trained to learn the physics constraints from hand-crafted rules. But If the rules are simple and explicit, then one can bypass the self-training and directly encode the rules into the MCTS process. If the rules are complex or unknown, then how to train the network to learn them?"}},"nonreaders":[],"tmdate":1731428943096,"tcdate":1730474213255,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7682/Reviewer_4ixR"],"signatures":["ICLR.cc/2025/Conference/Submission7682/Reviewer_4ixR"],"forum":"Ia17iAtr0P","number":1,"license":"CC BY 4.0","cdate":1730474213255,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7682/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428943096,"domain":"ICLR.cc/2025/Conference","replyto":"Ia17iAtr0P","id":"Mb9Bkmo9OX","forumContent":{"TLDR":{"value":"Physics-constrained symbolic regression, Monte-Carlo tree search with symmetric invariant representations and graph neural networks."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Symbolic Regression","Physics-constrained","Graph Neural Network","Reinforcement Learning","Monte-Carlo Tree Search","Expression Tree","Automated Feature Engineering","Symbolic Graph"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"As data-driven scientific discovery increasingly demands explainable over ‘black-box’ machine learning (ML) methods, Symbolic Regression (SR) that derives analytical expressions can help identify key functional dependencies in complex systems. However, traditional SR methods often suffer from (a) inefficient exploration due to their inability to compress the search space of equivalent expressions, and (b) non-physical solutions that violate fundamental physics constraints. We here introduce a symmetric invariant representation of candidate analytical expressions using a Symbolic Graph (SG), on which the Symbolic Graph Neural Network (SGNN) encodes operators, symmetries,   constraints and constant fitting knowledge. We further develop reinforcement learning (RL) algorithms with Monte-Carlo Tree Search (MCTS) on our SGNN for SR. Such a physics-constrained graph symbolic regression (PCGSR) method effectively compresses the search space for efficient SR. Experiments on synthetic and real-world scientific datasets demonstrate the efficiency and accuracy of our PCGSR in discovering underlying expressions and adhering to physical laws, yielding physically meaningful solutions."},"_bibtex":{"value":"@misc{\nxiang2025physicsconstrained,\ntitle={Physics-constrained Graph Symbolic Regression},\nauthor={Ziyu Xiang and Kenna Ashen and Xiaofeng Qian and Xiaoning Qian},\nyear={2025},\nurl={https://openreview.net/forum?id=Ia17iAtr0P}\n}"},"title":{"value":"Physics-constrained Graph Symbolic Regression"},"pdf":{"value":"/pdf/f49a5930438a1b85984252300a705d7a4a0f259c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"xiang|physicsconstrained_graph_symbolic_regression"},"authorids":{"value":["~Ziyu_Xiang1","~Kenna_Ashen1","~Xiaofeng_Qian1","~Xiaoning_Qian2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ziyu Xiang","Kenna Ashen","Xiaofeng Qian","Xiaoning Qian"]}},"version":2},{"content":{"summary":{"value":"The paper introduces CryoGEM, an innovative method combining physics-based cryo-EM simulation with unpaired noise translation via contrastive learning to generate high-quality synthetic cryo-EM datasets. The approach significantly improves the visual quality of generated images and enhances downstream tasks like particle picking and pose estimation, leading to better 3D reconstructions."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Is there any reference indicating whether the Gaussian noise distribution accurately represents the actual physical noise?\n\nIn practice, it is relatively easy to obtain a large number of observed samples of the target image in transmission images. Even if the proposed approach enhances the results, can we still easily access more samples of the target particle with minimal effort?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1 Extensive experiments demonstrate that CryoGEM produces high-quality synthetic cryo-EM images that significantly outperform existing methods like CycleGAN and CUT. The visual quality of the generated images is notably superior, preserving structural details and realistic noise patterns.\n\n2. The synthetic datasets generated by CryoGEM improve the performance of downstream tasks, such as particle picking and pose estimation. The paper reports substantial improvements in these tasks, leading to better resolution in the final 3D reconstructions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The physics-based simulation in CryoGEM relies on a coarse result as an input. This requirement can be a significant limitation in scenarios where obtaining a reliable coarse result is challenging, such as with very small or highly dynamic molecules."},"limitations":{"value":"See Weakness"}},"nonreaders":[],"tmdate":1730879086098,"tcdate":1720473501636,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission6516/Reviewer_D477"],"signatures":["NeurIPS.cc/2024/Conference/Submission6516/Reviewer_D477"],"forum":"edOZifvwMi","number":2,"license":"CC BY 4.0","cdate":1720473501636,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission6516/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879086098,"domain":"NeurIPS.cc/2024/Conference","replyto":"edOZifvwMi","id":"vz1BiqLSVf","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Image Synthesis","Contrastive Learning","Cryo-EM"]},"primary_area":{"value":"generative_models"},"abstract":{"value":"In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of proteins, such as the SARS-COV-2 spike protein. To achieve high-resolution reconstruction, a comprehensive data processing pipeline has been adopted. However, its performance is still limited as it lacks high-quality annotated datasets for training. To address this, we introduce physics-informed generative cryo-electron microscopy (CryoGEM), which for the first time integrates physics-based cryo-EM simulation with a generative unpaired noise translation to generate physically correct synthetic cryo-EM datasets with realistic noises. Initially, CryoGEM simulates the cryo-EM imaging process based on a virtual specimen. To generate realistic noises, we leverage an unpaired noise translation via contrastive learning with a novel mask-guided sampling scheme. Extensive experiments show that CryoGEM is capable of generating authentic cryo-EM images. The generated dataset can be used as training data for particle picking and pose estimation models, eventually improving the reconstruction resolution."},"_bibtex":{"value":"@inproceedings{\nzhang2024cryogem,\ntitle={Cryo{GEM}: Physics-Informed Generative Cryo-Electron Microscopy},\nauthor={Jiakai Zhang and Qihe Chen and Yan Zeng and Wenyuan Gao and Xuming He and Zhijie Liu and Jingyi Yu},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=edOZifvwMi}\n}"},"title":{"value":"CryoGEM: Physics-Informed Generative Cryo-Electron Microscopy"},"pdf":{"value":"/pdf/28c9b2115b82c3166b55e72a5bbf65dc9b1611fc.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"zhang|cryogem_physicsinformed_generative_cryoelectron_microscopy"},"authorids":{"value":["~Jiakai_Zhang3","~Qihe_Chen2","~Yan_Zeng3","~Wenyuan_Gao1","~Xuming_He3","~Zhijie_Liu3","~Jingyi_Yu5"]},"authors":{"value":["Jiakai Zhang","Qihe Chen","Yan Zeng","Wenyuan Gao","Xuming He","Zhijie Liu","Jingyi Yu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a novel neural operator framework (COMPOL) for multi-physics PDE problems. COMPOL include a recurrent and an attention based aggregation maechanics to model interdependency across physics process. The experiments show superiors accuacy across multiple new PDE datasets."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. It is an interesting question how the number of physics field correlated with the improvement of performance. Intuitive the aggregation is helpful when there is more physics coupled. For example, a standard NS equation has x-velocity, y-velocity, and p-pressure. How many physics field will the COMPOL model be helpful?\n2. Most of the datasets studied are newly generated. It will be interested to have an existing dataset, especially on weather foreacst. In ERA5 there are over 90 physics fields (temperature, velocity, humidity, etc). The model should show improve of performance.\n\nMinor suggestion: please make the main figure large (full linewidth) and add figures for each dataset."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper considers a relatively new problem setting of multi-physics modeling.  \n- It provides several new datasets consist of multiple physics fields and variables. \n- The architecture is flexible and agnostic to the specific inner neural operator choices."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The recurrent module and attention mechanism have been widely studied in ML and PDE community. Especially, the inter-physics attention has been applied in [1]. It will be helpful to discuss and compare with the previous work.\n2. Using recurrent structure and attention usually increase the runtime significantly, it will be helpful to report the runtime and plot convergence of the cost accuracy tradeoff (at what runtime, the model get what error rate). This will be helpful to justify the extra runtime and memory usage.\n\n\n[1] Rahman, Md Ashiqur, et al. \"Pretraining codomain attention neural operators for solving multiphysics pdes.\" Advances in Neural Information Processing Systems 37 (2024): 104035-104064."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918366219,"tcdate":1761838596210,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5942/Reviewer_A2xK"],"signatures":["ICLR.cc/2026/Conference/Submission5942/Reviewer_A2xK"],"forum":"k3DrCkpCok","number":3,"license":"CC BY 4.0","cdate":1761838596210,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5942/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918366219,"domain":"ICLR.cc/2026/Conference","replyto":"k3DrCkpCok","id":"eZPl2MdfYm","forumContent":{"TLDR":{"value":"A versatile multi-physics operator learning framwork"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["operator learning","physical simulations","coupled and multi-physics simulations"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Multi-physics simulations play an essential role in accurately modeling complex interactions across diverse scientific and engineering domains. Although neural operators, especially the Fourier Neural Operator (FNO), have significantly improved computational efficiency, they often fail to effectively capture intricate correlations inherent in coupled physical processes. To address this limitation, we introduce COMPOL, a novel coupled multi-physics operator learning framework. COMPOL extends conventional operator architectures by incorporating sophisticated recurrent and attention-based aggregation mechanisms, effectively modeling interdependencies among interacting physical processes within latent feature spaces. Our approach is architecture-agnostic and seamlessly integrates into various neural operator frameworks that involve latent space transformations. Extensive experiments on diverse benchmarks—including biological reaction-diffusion systems, pattern-forming chemical reactions, multiphase geological flows, and thermo-hydro-mechanical processes — demonstrate that COMPOL consistently achieves superior predictive accuracy compared to state-of-the-art methods."},"_bibtex":{"value":"@misc{\nsun2026compol,\ntitle={{COMPOL}: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations},\nauthor={Yifei Sun and Tao Wang and Junqi Qu and Yushun Dong and Hewei Tang and Shibo Li},\nyear={2026},\nurl={https://openreview.net/forum?id=k3DrCkpCok}\n}"},"title":{"value":"COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations"},"pdf":{"value":"/pdf/786ba3fc88eeb538ba6f9fce819f955aa4e43702.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"sun|compol_a_unified_neural_operator_framework_for_scalable_multiphysics_simulations"},"authorids":{"value":["~Yifei_Sun11","~Tao_Wang65","~Junqi_Qu1","~Yushun_Dong1","~Hewei_Tang1","~Shibo_Li1"]},"authors":{"value":["Yifei Sun","Tao Wang","Junqi Qu","Yushun Dong","Hewei Tang","Shibo Li"]}},"version":2},{"content":{"summary":{"value":"This paper conducts an in-depth study on the problem of visual information redundancy in vision-language large models (VLLMs), and points out that visual information redundancy is one of the important reasons for the poor performance of the model in complex visual tasks, such as fine-grained object recognition and spatial reasoning. The author constructed a synthetic dataset, quantified the complexity of tasks, and found that complex tasks (such as object counting) require more specialized visual tokens, have lower redundancy, and are sensitive to compression. However, simple tasks (such as color recognition) are not sensitive to redundancy and can even tolerate up to 99% token discard. Through fine-tuning experiments on the model, the author found that fine-tuning mainly altered the text representation of the model, while the visual representation changed relatively little. Moreover, different types of tasks (spatial reasoning vs. object localization) affected the internal representation of the model in different ways. Based on these findings, the authors proposed compression strategies and training suggestions for VLLMs, namely, appropriately compressing information in the early layers, carefully compressing in the middle layers, reducing the compression ratio for complex tasks, significantly compressing for simple tasks, and paying more attention to the updates of text and multimodal projection layers during fine-tuning."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- How to specifically configure the vision token compression on the task with different complexities? And, will the configuration setting be quite different among different types of models?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- This paper proposes and comprehensively applies multiple quantitative indicators (such as Gini coefficient, stable rank, participation rate, etc.) to systematically analyze visual information redundancy from the two levels of token norm and matrix rank, surpassing previous studies that only focused on attention distribution and providing a more comprehensive tool for understanding the internal visual information processing of VLLMs.\n\n- The experiments precisely controls variables through the construction of synthetic datasets, the negative correlation between task complexity (such as the number of objects and the difficulty of spatial reasoning) and the degree of visual information redundancy was clearly verified for the first time, providing direct evidence for explaining the performance bottleneck of VLLMs in complex visual tasks"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Some findings, such as \"there is a connection between task complexity and visual compression\", are similar with the conclusions given in previous works like PDrop[1]. \n\n- Fine-tuning experiments are only based on simplified subsets of COCO and GQA (such as objects with only \"left-right\" relationships), and more complex spatial relationships (such as spatial reasoning in ERQA) have not been tested, which may underestimate the model's redundant performance in real complex tasks.\n\n- The experiment mainly uses syntheti5c data of simple geometric shapes (fixed color/shape/size), lacking complex factors such as texture, occlusion, and lighting changes in real images, which may lead to insufficient generalization of the conclusion in real scenes.\n\n[1] Xing, et al. Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy reduction. CVPR, 202"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931246494,"tcdate":1762097432889,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19294/Reviewer_QipL"],"signatures":["ICLR.cc/2026/Conference/Submission19294/Reviewer_QipL"],"forum":"jvkukNvnny","number":4,"license":"CC BY 4.0","cdate":1762097432889,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19294/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931246494,"domain":"ICLR.cc/2026/Conference","replyto":"jvkukNvnny","id":"vILhLPNLbH","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["vision language modeling","large language models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Vision capabilities in vision large language models (VLLMs) have consistently lagged behind their linguistic capabilities. In particular, numerous benchmark studies have demonstrated that VLLMs struggle when fine-grained visual information or spatial reasoning is required. However, we do not yet understand exactly why VLLMs struggle so much with these tasks relative to others. Some works have focused on visual redundancy as an explanation, where high-level visual information is uniformly spread across numerous tokens and specific, fine-grained visual information is discarded. In this work, we investigate this premise in greater detail, seeking to better understand exactly how various types of visual information are processed by the model and what types of visual information are discarded. To do so, we introduce a simple synthetic benchmark dataset that is specifically constructed to probe various visual features, along with a set of metrics for measuring visual redundancy, allowing us to better understand the nuances of their relationship. Then, we explore fine-tuning VLLMs on a number of complex visual tasks to better understand how redundancy and compression change based upon the complexity of the data that a model is trained on. We find that there is a connection between task complexity and visual compression, implying that having a sufficient ratio of high complexity visual data is crucial for altering the way that VLLMs distribute their visual representation and consequently improving their performance on complex visual tasks. We hope that this work will provide valuable insights for training the next generation of VLLMs."},"_bibtex":{"value":"@misc{\nhannan2026seeing,\ntitle={Seeing Beyond Redundancy: Task Complexity's Role in Vision Token Specialization in {VLLM}s},\nauthor={Darryl Hannan and John Cooper and Joshua Dylan White and Yijing Watkins},\nyear={2026},\nurl={https://openreview.net/forum?id=jvkukNvnny}\n}"},"title":{"value":"Seeing Beyond Redundancy: Task Complexity's Role in Vision Token Specialization in VLLMs"},"pdf":{"value":"/pdf/8dd6f31460fd364411640ad2f63b27f0a26b1c34.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"hannan|seeing_beyond_redundancy_task_complexitys_role_in_vision_token_specialization_in_vllms"},"authorids":{"value":["~Darryl_Hannan1","~John_Cooper2","~Joshua_Dylan_White1","~Yijing_Watkins1"]},"authors":{"value":["Darryl Hannan","John Cooper","Joshua Dylan White","Yijing Watkins"]}},"version":2},{"content":{"summary":{"value":"This paper proposes DA-Transformer, a physics-informed Transformer model designed for spatiotemporal air quality forecasting. It explicitly integrates diffusion and advection mechanisms, key physical processes governing pollutant transport, directly into the Transformer encoder layers. Diffusion and Advection operators are differentiable and jointly learned with the model weights. The authors demonstrate the effectiveness of their approach across three diverse real-world datasets (Beijing, UK, California), showing consistent improvements over classical, deep sequence-based, Transformer-based, and graph-based baselines. The paper includes some ablation studies to assess the impact of the physics-informed modules and learnable positional encoding."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Could the authors provide a formal complexity analysis of the DA-Transformer as well as the other data-driven baselines (e.g., STGCN[3], AirFormer[4], AirPhyNet[1] and Air-DualODE[2])?\n\n2. The authors compared several time-series models, but LSTM and Informer is already a relatively early work in Time Series Prediction. Could the authors include a comparison with more recent methods such as PatchTST[5]?\n\n3. Could the authors provide a sensitivity analysis of some key hyperparameters?\n\n4. Considering that sudden changes are crucial in air quality prediction, could the authors include corresponding comparison results in Table 1? This is a common practice in other air quality prediction studies.\n\n\n[1]. AirPhyNet: Harnessing Physics-Guided Neural Networks for Air Quality Prediction.\n\n[2]. Air Quality Prediction with Physics-Guided Dual Neural ODEs in Open Systems.\n\n[3]. Spatio-temporal graph convolutional networks: a deep learning framework for traffic forecasting.\n\n[4]. Airformer: Predicting nationwide air quality in china with transformers.\n\n[5]. A time series is worth 64 words: Long-term forecasting with transformers."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. Instead of using external solvers or GNN-based PDE approximations, the paper proposes a clean integration of discrete diffusion/advection within Transformer layers.\n\n2. Physics modules (e.g., temperature-based diffusion coefficients, wind-driven advection) are well grounded in the paper.\n\n3. Multiple datasets with varying environmental regimes enhance credibility of generalizability claims."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The introduction section fails to summarize the core contributions of the paper and is unnecessarily verbose.\n\n2. The authors should maintain consistent terminology throughout the paper—specifically, they should decide between using “physics-informed” (e.g., lines 20, 84) and “physics-guided” (e.g., line 455), as the two terms are not conceptually equivalent.\n\n3. The paper includes insufficient baselines for spatiotemporal models.\n\n4. The novelty of the paper is quite limited, as many of the proposed methods are adapted from existing work (e.g., learnable physics modules are commonly used in AirPhyNet[1] and Air-DualODE[2]), and therefore should not be claimed as original contributions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943481902,"tcdate":1761806365052,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission25577/Reviewer_32GY"],"signatures":["ICLR.cc/2026/Conference/Submission25577/Reviewer_32GY"],"forum":"PLO1gjCMk5","number":4,"license":"CC BY 4.0","cdate":1761806365052,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission25577/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943481902,"domain":"ICLR.cc/2026/Conference","replyto":"PLO1gjCMk5","id":"arlR6677Cq","forumContent":{"TLDR":{"value":"A physics-informed Transformer that learns temperature-conditioned diffusion and wind-driven advection to improve long-horizon PM2.5 forecasting across regions."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Spatiotemporal Forecasting","Air Quality","Physics-informed Learning","Transformers","Diffusion and Advection"]},"supplementary_material":{"value":"/attachment/127ec3d3ecc63d58a3d88e1354dd42e816827af4.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Air pollution is a major concern for public health and the environment globally, which highlights the need for effective monitoring and predictive modeling to mitigate its impact. Although data-driven models have shown promising results in air quality prediction, they still struggle to model the underlying physical mechanisms of pollutant dispersion, where diffusion governs small-scale spreading and advection drives large-scale directional transport. To address this limitation, we propose the Diffusion-Advection Transformer (DA-Transformer), a novel physics-informed architecture. Specifically, the model integrates the two key physical mechanisms by embedding diffusion and advection as differential equation-based components. These physics-informed modules are incorporated into a Transformer framework to enable the model to better capture pollutant transport dynamics, such as local diffusion-driven smoothing and wind-induced directional propagation in air quality data. Experiments on three real-world datasets demonstrate that DA-Transformer consistently outperforms baseline models in $\\mathrm{PM}_{2.5}$ concentration prediction and achieves substantial gains over its variants that exclude diffusion and advection in their model design."},"_bibtex":{"value":"@misc{\nanonymous2026diffusionadvection,\ntitle={Diffusion-Advection Transformer for Air Quality Prediction},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=PLO1gjCMk5}\n}"},"title":{"value":"Diffusion-Advection Transformer for Air Quality Prediction"},"pdf":{"value":"/pdf/e159a98d7df32c21638799cb787ba60fd05cb848.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"zhang|diffusionadvection_transformer_for_air_quality_prediction"},"authorids":{"value":["~Luyang_Zhang3","~Chunbo_Luo2","~Geyong_Min1"]},"authors":{"value":["Luyang Zhang","Chunbo Luo","Geyong Min"]}},"version":2},{"content":{"summary":{"value":"This study investigates a novel complex-valued deep CNN designed for foreground extraction. It proposes a new method for encoding RGB images as complex values, an end-to-end token-based architecture that maintains the complex representation throughout, and an improved training pipeline. The authors demonstrate superior performance on a variety of complex-valued image benchmarks."},"soundness":{"value":4},"confidence":{"value":2},"questions":{"value":"See weaknesses above."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"I am unfamiliar with the literature on complex-valued neural networks, and defer to the opinion of more reviewers. As a neophyte to this field, I found the manuscript overall to be very readable, well-written, interesting, and convincing. The novel encoding, architecture, and training pipeline seem to work well, and produce a very capable model for handling complex-valued inputs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"This may be my ascribed to my naivety for the field, but I'm unsure how to interpret the benchmark results. While DCSNet certainly seems to outperform other complex-valued neural networks, the margin of victory is often in the range of 1-5 percent. It is difficult to tell whether this represents a fundamental advance, or a marginal improvement. Further, Table 2 seems to indicate that complex-valued neural networks in general frequently fail to outperform their real-valued counterparts. How should these results be interpreted with respect to the broader viability of complex-valued neural networks?"}},"nonreaders":[],"tmdate":1731427401976,"tcdate":1730213020086,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1276/Reviewer_B2bV"],"signatures":["ICLR.cc/2025/Conference/Submission1276/Reviewer_B2bV"],"forum":"9hmDl8fFDs","number":1,"license":"CC BY 4.0","cdate":1730213020086,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1276/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427401976,"domain":"ICLR.cc/2025/Conference","replyto":"9hmDl8fFDs","id":"4nilEgoqUB","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"A robust complex-valued approach in Spatio-spectral domain for multiple tasks on both real and complex data."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Deep Complex Newtworks","Complex-valued color transformation"]},"supplementary_material":{"value":"/attachment/d3c1bebc77a7f4edd23a402e6580595b34f472e4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex-valued neural networks have attracted growing attention for their ability to handle complex-valued data with enhanced representational capacity. However, their potential in computer vision remains relatively untapped. \nIn this paper, we introduce Deep Complex Spatio-Spectral Network (DCSNet), a fully complex-valued token-based, end-to-end neural network designed for binary segmentation tasks. Additionally, our DCSNet encoder can be used for image classification in the complex domain. We also propose an invertible real-to-complex (R2C) transform, which generates two complex-valued input channels, complex intensity and complex hue, while producing complex-valued images with distinct real and imaginary components.\nDCSNet operates in both spatial and spectral domains by leveraging complex-valued inputs and complex Fourier transform.\nAs a result, the complex-valued representation is maintained throughout DCSNet, and we avoid the information loss typically associated with Real$\\leftrightarrow$Complex transformations. Extensive experiments show that DCSNet surpasses existing complex-valued methods across various tasks on both real and complex-valued data and achieves competitive performance compared to existing real-valued methods, establishing a robust framework for handling both data types effectively."},"_bibtex":{"value":"@misc{\nyadav2025deep,\ntitle={Deep Complex Spatio-Spectral Networks with Complex Visual Inputs},\nauthor={Saurabh Yadav and Koteswar Rao Jerripothula},\nyear={2025},\nurl={https://openreview.net/forum?id=9hmDl8fFDs}\n}"},"title":{"value":"Deep Complex Spatio-Spectral Networks with Complex Visual Inputs"},"pdf":{"value":"/pdf/78685d9476ab42033341d680f3e952693b0cbc64.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yadav|deep_complex_spatiospectral_networks_with_complex_visual_inputs"},"authorids":{"value":["~Saurabh_Yadav2","~Koteswar_Rao_Jerripothula3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Saurabh Yadav","Koteswar Rao Jerripothula"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel physics-guided approach, DA-Transformer, for modeling air quality data. Specifically, it integrates the diffusion–advection phenomenon of pollutant transport into the Transformer architecture through a first-order discretization of the corresponding partial differential equations. Experiments are conducted on three representative datasets, including one newly constructed by the authors. The comparative results demonstrate that the proposed method outperforms the baselines."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.Could the authors explain why the ablation results vary across datasets when the diffusion–advection modules are removed? In particular, why are the performance differences on the UK and California datasets relatively small—are there other influencing factors specific to these regions?\n\n2.MAPE is also an important metric for spatiotemporal forecasting; achieving improvements only in MAE and RMSE is not sufficiently convincing.\n\n3.Could author provide more results on other spatiotemporal prediction baselines like Airformer[1], PM2.5GNN[2]? \n\n4.Could the authors provide evaluation results for sudden change scenarios in Table 1? The definition of sudden changes can be referenced from AirFormer[1].\n\n[1]. Yuxuan Liang, et al. AirFormer: Predicting Nationwide Air Quality in China with Transformers. AAAI 2023\n\n[2]. Shuo Wang, et al. Pm2. 5-gnn: A domain knowledge enhanced graph neural network for pm2. 5 forecasting. ACM SIGSPATIAL 2020"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"1.A novel physics-guided approach based on the Transformer architecture is proposed, which differs from existing physics-guided methods in air quality prediction.\n\n2.The Fig. 1 clearly illustrates the proposed method.\n\n3.The DA-Transformer is evaluated on three distinct datasets, each representing a unique environmental setting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The authors should clarify their motivation and the challenges of existing physics-guided approaches in the introduction. Currently, this paper gives the impression that the main goal is simply to incorporate physics-guided modules into a Transformer architecture, rather than to advance the modeling paradigm within the physics-guided domain.\n\n2.The DA-Transformer introduces the discrete diffusion-advection equation in the latent space; however, the latent representation H is derived from pollutant and meteorological data. This raises a serious concern: does such a representation truly satisfy the physical assumptions of the diffusion-advection equation? This is a serious issue that undermines the physical validity of the model.\n\n3.No evidence is given that such surrogate station-order-based derivatives approximate true spatial gradients. This could lead to physically incorrect modeling in irregular sensor layouts.\n\n4.Although the authors define a discretized solution scheme to embed PDEs into the Transformer architecture, they only adopt a first-order discretization. While this choice is simple and computationally efficient, it raises concerns about potential cumulative errors compared to ODE-solver-based methods. It is similar to using Euler’s method to solve differential equations.\n\n5.The related work section is insufficient, and the baselines listed in the main experiment table are quite limited. For example, AirFormer, a data-driven model specifically designed for air pollutant, is not mentioned or compared.\n\n6.There is an inconsistency between the problem statement and the description in the methodology section. The output dimension of DA-transformer is confusing.\n\n7.The novelty of this work is limited, as it primarily focuses on incorporating physics modules into the Transformer architecture without introducing fundamentally new modeling techniques.\n\n8.The introduction does not effectively highlight the main contributions and is overly wordy."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943482261,"tcdate":1761663141108,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission25577/Reviewer_RSdU"],"signatures":["ICLR.cc/2026/Conference/Submission25577/Reviewer_RSdU"],"forum":"PLO1gjCMk5","number":3,"license":"CC BY 4.0","cdate":1761663141108,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission25577/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943482261,"domain":"ICLR.cc/2026/Conference","replyto":"PLO1gjCMk5","id":"NrNSbUTINB","forumContent":{"TLDR":{"value":"A physics-informed Transformer that learns temperature-conditioned diffusion and wind-driven advection to improve long-horizon PM2.5 forecasting across regions."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Spatiotemporal Forecasting","Air Quality","Physics-informed Learning","Transformers","Diffusion and Advection"]},"supplementary_material":{"value":"/attachment/127ec3d3ecc63d58a3d88e1354dd42e816827af4.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Air pollution is a major concern for public health and the environment globally, which highlights the need for effective monitoring and predictive modeling to mitigate its impact. Although data-driven models have shown promising results in air quality prediction, they still struggle to model the underlying physical mechanisms of pollutant dispersion, where diffusion governs small-scale spreading and advection drives large-scale directional transport. To address this limitation, we propose the Diffusion-Advection Transformer (DA-Transformer), a novel physics-informed architecture. Specifically, the model integrates the two key physical mechanisms by embedding diffusion and advection as differential equation-based components. These physics-informed modules are incorporated into a Transformer framework to enable the model to better capture pollutant transport dynamics, such as local diffusion-driven smoothing and wind-induced directional propagation in air quality data. Experiments on three real-world datasets demonstrate that DA-Transformer consistently outperforms baseline models in $\\mathrm{PM}_{2.5}$ concentration prediction and achieves substantial gains over its variants that exclude diffusion and advection in their model design."},"_bibtex":{"value":"@misc{\nanonymous2026diffusionadvection,\ntitle={Diffusion-Advection Transformer for Air Quality Prediction},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=PLO1gjCMk5}\n}"},"title":{"value":"Diffusion-Advection Transformer for Air Quality Prediction"},"pdf":{"value":"/pdf/e159a98d7df32c21638799cb787ba60fd05cb848.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"zhang|diffusionadvection_transformer_for_air_quality_prediction"},"authorids":{"value":["~Luyang_Zhang3","~Chunbo_Luo2","~Geyong_Min1"]},"authors":{"value":["Luyang Zhang","Chunbo Luo","Geyong Min"]}},"version":2},{"content":{"venue":{"value":"Complex (1) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02466-5_9.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"sheridan|on_scalefree_prior_distributions_and_their_applicability_in_largescale_network_inference_with_gaussian_graphical_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Paul_Sheridan:","https://dblp.org/search/pid/api?q=author:Takeshi_Kamimura:","~Hidetoshi_Shimodaira1"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02466-5_9"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/SheridanKS09,\n  author={Paul Sheridan and Takeshi Kamimura and Hidetoshi Shimodaira},\n  title={On Scale-Free Prior Distributions and Their Applicability in Large-Scale Network Inference with Gaussian Graphical Models},\n  year={2009},\n  cdate={1230768000000},\n  pages={110-117},\n  url={https://doi.org/10.1007/978-3-642-02466-5_9},\n  booktitle={Complex (1)},\n  crossref={conf/complex/2009-1}\n}\n"},"abstract":{"value":"This paper concerns the specification, and performance, of scale-free prior distributions with a view toward large-scale network inference from small-sample data sets. We devise three scale-free priors and implement them in the framework of Gaussian graphical models. Gaussian graphical models are used in gene network inference where high-throughput data describing a large number of variables with comparatively few samples are frequently analyzed by practitioners. And, although there is a consensus that many such networks are scale-free, the modus operandi is to assign a random network prior. Simulations demonstrate that the scale-free priors outperform the random network prior at recovering scale-free trees with degree exponents near 2, such as are characteristic of many real-world systems. On the other hand, the random network prior compares favorably at recovering scale-free trees characterized by larger degree exponents."},"title":{"value":"On Scale-Free Prior Distributions and Their Applicability in Large-Scale Network Inference with Gaussian Graphical Models"},"authors":{"value":["Paul Sheridan","Takeshi Kamimura","Hidetoshi Shimodaira"]}},"tmdate":1733720740530,"pdate":1230768000000,"tcdate":1733720734967,"writers":["~"],"signatures":["~Hidetoshi_Shimodaira1"],"forum":"hhpevFMSNw","license":"CC BY-SA 4.0","number":254271,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1733720740530,"domain":"DBLP.org","id":"hhpevFMSNw","version":2},{"content":{"summary":{"value":"This manuscript proposes a benchmark to evaluate physical realism in image editing. The authors consider physical effects beyond simple instruction completion. They have collected a physics-aware benchmark, PICABench, to evaluate the performance of existing image editing models. Meanwhile, the authors also utilize a VLM-as-a-judge (PICAEval) to evaluate the edited images. Beyond the benchmark, the authors also propose constructing a physics-aware dataset (PICA-100K) from videos. Experimental results demonstrate that existing models still struggle to achieve physical realism. Furthermore, the proposed dataset is shown to enhance the physical realism of a fine-tuned base model."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This manuscript considers the role of physics in the image editing task, which is an interesting and rarely explored area of research. The proposed benchmark is well-constructed and provides a comprehensive analysis of existing models regarding the physical realism of their edits. In addition, the proposed PICA-100K dataset is a clever approach to learn physics from synthetic videos."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The primary weakness is that the proposed solution, fine-tuning on PICA-100K, yields very small gains. While this is a positive result and better than the baseline model, the small margin doesn't present PICA-100K as a definitive solution. Moreover, the final fine-tuned model still underperforms other top-tier models (e.g., GPT-Image-1, Seedream 4.0, Qwen-Image-Edit) in the overall benchmark, as shown in Table 1.\n\n2. To better demonstrate the dataset's effectiveness, the authors should try fine-tuning other base models (besides Flux.1 Kontext) and report their performance improvements.\n\n3. Meanwhile, the concept of constructing image editing pairs from synthetic videos is not entirely novel and has been explored in previous works (e.g., Frame2Frame[1], ByteMorph[2]). The reviewer recommends the authors explicitly to describe the differences and contributions of their data generation pipeline compared to these prior works.\n\n[1] Pathways on the Image Manifold: Image Editing via Video Generation\n\n[2] ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915921671,"tcdate":1761983762288,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1871/Reviewer_Ytia"],"signatures":["ICLR.cc/2026/Conference/Submission1871/Reviewer_Ytia"],"forum":"AWxI5xnuZB","number":4,"license":"CC BY 4.0","cdate":1761983762288,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1871/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915921671,"domain":"ICLR.cc/2026/Conference","replyto":"AWxI5xnuZB","id":"r5zaDnUtEr","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["image edit; benchmark; dataset"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Image editing has achieved remarkable progress recently. Modern editing models could already follow complex instructions to manipulate the original content. However, beyond completing the editing instructions, the accompanying physical effects are the key to the generation realism. For example, removing an object should also remove its shadow, reflections, and interactions with nearby objects. Unfortunately, existing models and benchmarks mainly focus on instruction completion but overlook these physical effects. So, at this moment, how far are we from physically realistic image editing? To answer this, we introduce PICABench, which systematically evaluates physical realism across eight sub-dimension(spanning optics, mechanics, and state transitions) for most of the common editing operations(add, remove, attribute change, etc). We further propose the PICAEval, a reliable evaluation protocol that uses VLM-as-a-judge with per-case, region-level human annotations and questions. Beyond benchmarking, we also explore effective solutions by learning physics from videos and construct a training dataset PICA-100K.After evaluating most of the mainstream models, we observe that physical realism remains a challenging problem with large rooms to explore. We hope that our benchmark and proposed solutions can serve as a foundation for future work moving from naive content editing toward physically consistent realism."},"_bibtex":{"value":"@inproceedings{\npu2026picabench,\ntitle={{PICAB}ench: How Far are We from Physical Realistic  Image Editing?},\nauthor={Yuandong Pu and Le Zhuo and Songhao Han and Jinbo Xing and Kaiwen Zhu and Shuo Cao and Bin Fu and Si Liu and Hongsheng Li and Yu Qiao and Wenlong Zhang and Xi Chen and Yihao Liu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=AWxI5xnuZB}\n}"},"title":{"value":"PICABench: How Far are We from Physical Realistic  Image Editing?"},"pdf":{"value":"/pdf/3181431c47dfc2d9c17b62f29f6f70bdf43c06ed.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"pu|picabench_how_far_are_we_from_physical_realistic_image_editing"},"authorids":{"value":["~Yuandong_Pu1","~Le_Zhuo2","~Songhao_Han1","~Jinbo_Xing1","~Kaiwen_Zhu2","~Shuo_Cao2","~Bin_Fu1","~Si_Liu5","~Hongsheng_Li3","~Yu_Qiao1","~Wenlong_Zhang3","~Xi_Chen30","~Yihao_Liu1"]},"authors":{"value":["Yuandong Pu","Le Zhuo","Songhao Han","Jinbo Xing","Kaiwen Zhu","Shuo Cao","Bin Fu","Si Liu","Hongsheng Li","Yu Qiao","Wenlong Zhang","Xi Chen","Yihao Liu"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2403.11237v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"zhang|force_dataset_and_method_for_intuitive_physics_guided_humanobject_interaction"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xiaohan_Zhang:","https://dblp.org/search/pid/api?q=author:Bharat_Lal_Bhatnagar:","https://dblp.org/search/pid/api?q=author:Sebastian_Starke:","https://dblp.org/search/pid/api?q=author:Ilya_Petrov:","~Vladimir_Guzov1","~Helisa_Dhamo1","https://dblp.org/search/pid/api?q=author:Eduardo_Pérez-Pellitero:","https://dblp.org/search/pid/api?q=author:Gerard_Pons-Moll:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2403.11237"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2403-11237,\n  publtype={informal},\n  author={Xiaohan Zhang and Bharat Lal Bhatnagar and Sebastian Starke and Ilya Petrov and Vladimir Guzov and Helisa Dhamo and Eduardo Pérez-Pellitero and Gerard Pons-Moll},\n  title={FORCE: Dataset and Method for Intuitive Physics Guided Human-object Interaction},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2403.11237},\n  url={https://doi.org/10.48550/arXiv.2403.11237}\n}\n"},"abstract":{"value":"Interactions between human and objects are influenced not only by the object's pose and shape, but also by physical attributes such as object mass and surface friction. They introduce important motion nuances that are essential for diversity and realism. Despite advancements in recent kinematics-based methods, this aspect has been overlooked. Generating nuanced human motion presents two challenges. First, it is non-trivial to learn from multi-modal human and object information derived from both the physical and non-physical attributes. Second, there exists no dataset capturing nuanced human interactions with objects of varying physical properties, hampering model development. This work addresses the gap by introducing the FORCE model, a kinematic approach for synthesizing diverse, nuanced human-object interactions by modeling physical attributes. Our key insight is that human motion is dictated by the interrelation between the force exerted by the human and the perceived resistance. Guided by a novel intuitive physics encoding, the model captures the interplay between human force and resistance. Experiments also demonstrate incorporating human force facilitates learning multi-class motion. Accompanying our model, we contribute the FORCE dataset. It features diverse, different-styled motion through interactions with varying resistances."},"title":{"value":"FORCE: Dataset and Method for Intuitive Physics Guided Human-object Interaction"},"authors":{"value":["Xiaohan Zhang","Bharat Lal Bhatnagar","Sebastian Starke","Ilya Petrov","Vladimir Guzov","Helisa Dhamo","Eduardo Pérez-Pellitero","Gerard Pons-Moll"]}},"tmdate":1731528000004,"pdate":1704067200000,"tcdate":1731331471475,"writers":["~"],"signatures":["~Vladimir_Guzov1"],"forum":"OUD5MPLauH","license":"CC BY-SA 4.0","number":181568,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731528000004,"domain":"DBLP.org","id":"OUD5MPLauH","version":2},{"content":{"summary":{"value":"This paper proposes PINFDIT, a novel framework aiming to be a general-purpose model for time series analysis. It combines a Transformer architecture with a Diffusion framework to perform a variety of tasks, including forecasting, imputation, anomaly detection, and synthetic data generation. The core contributions of the paper are (1) a comprehensive masking strategy to handle imperfect time series data (missing values, multi-resolution) and (2) a novel \"energy-based physics-informed sampling\" technique that allows for the injection of physical laws at the inference stage without retraining or architectural modification.\n\nThis work tackles a very ambitious and important problem. The attempt to unify a pre-trained, general-purpose model with domain-specific physical knowledge is highly timely. The use of calibrated Langevin dynamics for correction at inference time is a practical and clever idea that could overcome the limitations of existing approaches like PINNs.\n\nHowever, there are several significant weaknesses regarding the paper's core claims, experimental fairness, and methodological clarity. Without addressing these, it is difficult to evaluate the generality and true contribution of the proposed framework."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.  **(Re: W1)** Could you please provide a detailed explanation of how the partial differential terms and the gradient $\\nabla K$ in Eq. (5) and Eq. (8) are computed for discrete time series data $x^{tar}$?\n2.  **(Re: W3)** What was the specific reason for \"bypassing pretraining\" in the anomaly detection task? Does this imply that the pre-trained general-purpose model may be unsuitable for certain downstream tasks?\n3.  **(Re: W3)** For a fair comparison in the imputation experiments (Tables 4/15), can you provide results for PINFDIT trained from scratch (without pre-training)?\n4.  **(Re: W4)** Can you show what happens to performance when the physics-injection module (Alg. 1, lines 6-8) is applied with an intentionally incorrect physical law, or to a non-physical dataset like NASDAQ?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.  **Novel and Practical Physics-Injection Method:** Unlike traditional Physics-Informed Machine Learning (e.g., PINNs) which requires including PDE residuals in the training loss and retraining for specific problems, the proposed \"model-editing-free\" inference-time correction is highly practical. The ability to maintain the flexibility of a pre-trained general model while \"plugging in\" domain knowledge as needed is a major advantage.\n2.  **Comprehensive Problem Definition and Task Scope:** The paper has an ambitious goal of solving a wide range of time series tasks (forecasting, imputation, anomaly detection, generation) with a single model. It is also commendable that it directly addresses real-world data complexities such as missing values, multi-resolution sampling, and irregular intervals.\n3.  **Extensive Experimental Validation:** The authors have conducted a vast number of experiments to demonstrate their model's performance. This includes PDE simulator-based forecasting, practical forecasting on real-world data (climate, healthcare, finance), generative and imputation tasks, zero-shot performance, and detailed ablation studies. Achieving state-of-the-art (SOTA) performance on many benchmarks is impressive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  **Ambiguity in the Implementation of Physics Injection:**\n    * The concrete implementation of the physics-injection mechanism, one of the paper's core contributions, is critically unclear.\n    * Physical laws are quantified by an energy function $K(x^{tar};F)$ (the squared PDE residual).\n    * For correction during inference, the gradient of this energy function ($\\nabla K(x_{j}^{tar};F)$) is required (see Algorithm 1).\n    * **Critical Question:** $x^{tar}$ is discrete time series data. How are partial differential terms like $\\frac{\\partial x^{tar}}{\\partial t}$ and $\\frac{\\partial x^{tar}}{\\partial u_{i}}$ computed? If finite differences were used, there is no discussion of the associated discretization error, stability, or sensitivity to the choice of scheme. This information is essential for reproducing the methodology and verifying its validity.\n    * The authors claim they do not require a \"differentiable simulator\", but it appears they *do* require a \"differentiable (physical) residual function.\" The distinction and constraints are not adequately explained.\n\n2.  **Disconnect in \"General-Purpose\" and \"Foundation Model\" Claims:**\n    * The paper positions PINFDIT as a \"proto-foundation model\".\n    * However, the model's operation contradicts this claim. The model's pre-training (using the Chronos dataset) is entirely physics-agnostic and learns only statistical patterns.\n    * The physics injection is merely a post-hoc correction step (Algorithm 1, lines 6-8) that operates *outside* the pre-trained model during inference. This is less a \"Physics-Informed\" Transformer and more a \"General-Purpose Transformer\" combined with a \"Physics-Based Corrector Plugin.\" While this modularity is a strength, the title \"Physics-Informed\" is misleading as the model itself does not *learn* or *internalize* the physics.\n\n3.  **Unfair Experimental Comparisons and Lack of Consistency:**\n    * **Imputation (Tables 4 & 15):** The authors state, \"All baseline models are trained in a full-shot setting, while PINFDIT leverages a pre-trained foundation model, fine-tuning it on realistic datasets\". This is a **clearly unfair comparison**. PINFDIT benefits from pre-training on massive external data (Chronos, ~5B time points), while baselines like PatchTST and TimesNet are trained from scratch only on the task data (e.g., ETTh1). It is impossible to distinguish if PINFDIT's superiority comes from its architecture or simply from pre-training. A fair comparison would require reporting results for PINFDIT trained from scratch on the same data.\n    * **Anomaly Detection (Table 14):** The authors state they \"opted to bypass pretraining\" for this task and introduced an additional pre-processing step (Spectral Residue). This severely undermines the core narrative of a single foundation model for all tasks. If the pre-trained model is actually detrimental for anomaly detection (as it \"may inadvertently overfit by reconstructing anomalies\"), this clearly exposes a limitation of the proposed general-purpose model.\n\n4.  **Lack of Robustness Analysis for the Physics-Injection Plugin:**\n    * The paper only presents positive cases where physics injection improves performance (e.g., PDE simulations, ERA5 climate forecasting).\n    * There are no failure-case or sensitivity analyses. What happens if an incorrect physical law is injected (e.g., applying Burgers' equation to non-fluid data)? What happens if it's applied to data with no physical basis (e.g., the NASDAQ financial dataset)? Does performance degrade? Proving the validity of this modular approach requires such robustness checks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942813811,"tcdate":1761714393522,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23805/Reviewer_EgRK"],"signatures":["ICLR.cc/2026/Conference/Submission23805/Reviewer_EgRK"],"forum":"EphTlUJ4XN","number":1,"license":"CC BY 4.0","cdate":1761714393522,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23805/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942813811,"domain":"ICLR.cc/2026/Conference","replyto":"EphTlUJ4XN","id":"Pu33yhAQoA","forumContent":{"TLDR":{"value":"Physics-Guided Inference in Time Series Diffusion Transformers"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion; Transformer; Time Series; Physics Informed Machine Learning;Physics-Guided Inference in Time Series Diffusion Transformers"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Time series analysis underpins scientific advances. While specialized models have advanced various time series tasks, scientific domains face unique challenges: limited samples with complex physical dynamics, missing observations, multi-resolution sampling, and requirements for physical consistency. With the increasing demands on generative modeling capabilities, we introduce PINFDiT, a diffusion transformer-based model with physics injection during inference. Our approach combines a transformer backbone for capturing temporal dependencies with a comprehensive masking strategy that addresses imperfect data. The diffusion framework enables high-quality sample generation with inherent generative capability. In addition, our model-free physics-guided correction steers generated samples toward physically consistent solutions using calibrated Langevin dynamics, which balances distribution fidelity and physical law adherence without architectural modifications or retraining. Our evaluation demonstrates PINFDiT's effectiveness across multivariate forecasting with imperfect data, physics knowledge incorporation in data-limited scenarios, zero-shot and fine-tuning performance across diverse domains, establishing it as a proto-foundation model that bridges the gap between general-purpose and domain-specific models."},"_bibtex":{"value":"@inproceedings{\ncao2026pinfdit,\ntitle={{PINFD}iT: Energy-Based Physics-Informed Diffusion Transformers for General-purpose Time Series Tasks},\nauthor={Defu Cao and Wen Ye and Yizhou Zhang and Sam Griesemer and Yan Liu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EphTlUJ4XN}\n}"},"title":{"value":"PINFDiT: Energy-Based Physics-Informed Diffusion Transformers for General-purpose Time Series Tasks"},"pdf":{"value":"/pdf/db4837ca9639380f0ee731e255ce49e3ec012290.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"cao|pinfdit_energybased_physicsinformed_diffusion_transformers_for_generalpurpose_time_series_tasks"},"authorids":{"value":["~Defu_Cao1","~Wen_Ye3","~Yizhou_Zhang3","~Sam_Griesemer1","~Yan_Liu1"]},"authors":{"value":["Defu Cao","Wen Ye","Yizhou Zhang","Sam Griesemer","Yan Liu"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"venue":{"value":"DT4H 2026 Poster"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"keywords":{"value":["High-frequency ultrasound","Synthetic generation","Skin layer segmentation","SLEB","k-Wave"]},"abstract":{"value":"High-frequency ultrasound (HFUS) enables noninvasive visualization of superficial skin structures, but automated skin-layer analysis is limited by the scarcity of densely annotated data. Existing real HFUS datasets commonly provide annotations for superficial targets\nsuch as the epidermis and subepidermal low-echogenic band (SLEB), while dense labels for deeper structures such as dermis, subcutaneous tissue, fascia, and muscle are rarely available. We propose a physics-guided synthetic HFUS generation framework for skin layer segmentation. The framework constructs multilayer acoustic skin phantoms, assigns layer dependent acoustic properties, and uses k-Wave simulation to generate paired synthetic HFUS images, dense layer masks, and simulation metadata. To evaluate whether the generated data provide transferable supervision, we use it for downstream segmentation pretraining and fine-tune the models on real Mendeley HFUS data. Synthetic pretraining followed by real fine-tuning achieved real-domain performance comparable to real-only training and improved mean Dice/IoU in three of four evaluated trainable architectures. These results suggest that physics-guided synthetic HFUS images contain transferable anatomical and textural cues for real-domain skin layer segmentation, although further reduction of the synthetic-real appearance gap is needed to enable greater gains. The code and data are available at: https://github.com/Finn-02/synthetic-hfus-skin-layer-segmentation."},"_bibtex":{"value":"@inproceedings{\nju2026physicsguided,\ntitle={Physics-Guided Synthetic High-Frequency Ultrasound Generation for Skin Layer Segmentation},\nauthor={Junkyung ju and Kyungho Yoon and Minwoo Shin},\nbooktitle={2nd International Workshop on Digital Twin for Healthcare},\nyear={2026},\nurl={https://openreview.net/forum?id=vVvAcwZvMS}\n}"},"title":{"value":"Physics-Guided Synthetic High-Frequency Ultrasound Generation for Skin Layer Segmentation"},"original_pdf":{"value":"/pdf/cdb320e78988813e911505bb132b3fe4d6777a88.pdf"},"venueid":{"value":"MICCAI.org/2026/Workshop/DT4H"},"paperhash":{"value":"ju|physicsguided_synthetic_highfrequency_ultrasound_generation_for_skin_layer_segmentation"},"authorids":{"value":["~Junkyung_ju1","~Kyungho_Yoon1","~Minwoo_Shin2"]},"Contribution":{"value":"We propose a physics-guided synthetic HFUS generation framework that generates paired images, dense multilayer labels, and simulation metadata. Experiments demonstrate its transferability for real-domain skin segmentation."},"authors":{"value":["Junkyung ju","Kyungho Yoon","Minwoo Shin"]}},"tmdate":1787950137345,"pdate":1787825867030,"tcdate":1781876700840,"writers":["MICCAI.org/2026/Workshop/DT4H","MICCAI.org/2026/Workshop/DT4H/Submission15/Authors"],"signatures":["MICCAI.org/2026/Workshop/DT4H/Submission15/Authors"],"forum":"vVvAcwZvMS","license":"CC BY 4.0","number":15,"cdate":1781876700840,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/DT4H/-/Submission","MICCAI.org/2026/Workshop/DT4H/-/Submission_Change_Before_Bidding","MICCAI.org/2026/Workshop/DT4H/-/Submission_Change_Before_Reviewing","MICCAI.org/2026/Workshop/DT4H/Submission15/-/Camera_Ready_Revision","MICCAI.org/2026/Workshop/DT4H/-/Edit","MICCAI.org/2026/Workshop/DT4H/-/Submission_Release"],"mdate":1787950137345,"odate":1787825867030,"domain":"MICCAI.org/2026/Workshop/DT4H","id":"vVvAcwZvMS","version":2},{"content":{"summary":{"value":"The authors propose diffusion-based inverse problem solvers involving the temporal evolution of physics systems. The method utilize a combination of score function and an inverse physics simulator, which corresponds to reverse of drift term in diffusion models, to moves the system’s state backward in time. They demonstrate the effectiveness of their method in a wide range of inverse physics problems."},"soundness":{"value":"4 excellent"},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"In the evaluation of this methods, is it acceptable to not compare it with conventional method such as finite element methods?"},"rating":{"value":"4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"4 excellent"},"contribution":{"value":"2 fair"},"strengths":{"value":"The authors propose multi-step loss to capture long-range dependence in physics system, which has potential implication for general diffusion models.\nThe experiment are conducted extensively."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The method is modification of diffusion model with adoption of inverse physic simulator in order to address nonlinear drift term of the physics system. In this regard, its novelty is limited and its applicability may be restricted to certain scenarios. \n\nMinor/errata\n\nHats are missing in equation (4)\n"},"limitations":{"value":"It is suggested to include the future direction of this work."}},"nonreaders":[],"tmdate":1702410810146,"tcdate":1688713157597,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1882/Reviewer_og5h"],"signatures":["NeurIPS.cc/2023/Conference/Submission1882/Reviewer_og5h"],"forum":"2BpoGPSDCR","number":3,"license":"CC BY 4.0","cdate":1688713157597,"mdate":1702410810146,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1882/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"2BpoGPSDCR","id":"LgIju3mFlh","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["inverse problems","diffusion models","learned corrections","score matching"]},"_bibtex":{"value":"@inproceedings{\nholzschuh2023solving,\ntitle={Solving Inverse Physics Problems with Score Matching},\nauthor={Benjamin Holzschuh and Simona Vegetti and Nils Thuerey},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=2BpoGPSDCR}\n}"},"title":{"value":"Solving Inverse Physics Problems with Score Matching"},"paperhash":{"value":"holzschuh|solving_inverse_physics_problems_with_score_matching"},"TLDR":{"value":"We propose a novel methodology for solving inverse problems that involve the temporal evolution of physical systems."},"abstract":{"value":"We propose to solve inverse problems involving the temporal evolution of physics systems by leveraging recent advances from diffusion models. \nOur method moves the system's current state backward in time step by step by combining an approximate inverse physics simulator and a learned correction function. \nA central insight of our work is that training the learned correction with a single-step loss is equivalent to a score matching objective, while recursively predicting longer parts of the trajectory during training relates to maximum likelihood training of a corresponding probability flow.\nWe highlight the advantages of our algorithm compared to standard denoising score matching and implicit score matching, as well as fully learned baselines for a wide range of inverse physics problems. The resulting inverse solver has excellent accuracy and temporal stability and, in contrast to other learned inverse solvers, allows for sampling the posterior of the solutions. Code and experiments are available at https://github.com/tum-pbs/SMDP."},"pdf":{"value":"/pdf/e14d868c93d4bb63690c02fd9c5fcb4f28eed65c.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Benjamin_Holzschuh1","~Simona_Vegetti1","~Nils_Thuerey1"]},"authors":{"value":["Benjamin Holzschuh","Simona Vegetti","Nils Thuerey"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Great GATsBi, a hybrid multimodal framework for bicycle trajectory forecasting. It integrates physics-based modeling (capturing vehicle-like dynamics via constant velocity, constant acceleration, kinematics, and extended Kalman filtering) and social-based modeling (capturing pedestrian-like social interactions via Graph Attention Networks, GATs). The framework also incorporates an anticipation mechanism and perception decay inspired by psychological and social science insights. It is evaluated on a self-collected controlled mass cycling dataset (circular track, varying traffic density) and generalized to pedestrian datasets (ETH, HOTEL). The paper claims Great GATsBi outperforms physics-based and social-based baselines, addressing the gap of neglecting bicycles’ dual behavioral nature in prior trajectory forecasting work."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the weaknesses for my main concerns.\nIn addition, I would like the authors to clarify the following:\n1. Could the authors further clarify how the proposed anticipation mechanism avoids the potential circular reasoning problem identified in the weaknesses?\n2. Could the authors provide an ablation study on the physics model ensemble (e.g., using different combinations of the four models) to conclusively demonstrate the necessity of including all of them?\n3. Could the missing citation on line 68 be added, and could the resolution of Figure 1 be improved?\n4. Could the authors specify the neighbor selection strategy for the social graph and potentially include an ablation study on its impact?\n5. Why were bicycle-specific baselines excluded from comparisons? If pedestrian baselines (SocialLSTM/Social-BiGAT) were adapted to bicycle dynamics (e.g., adjusting for speed), would their performance narrow the gap with Great GATsBi?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The model develops a hybrid framework that combines physics-based and social-based modeling to effectively match bicycles’ dual behavioral traits of vehicle-like dynamics and pedestrian-like flexibility.\n2. The social module innovatively integrates psychological and social science insights (neighbor trajectory anticipation and perception decay) to make social interaction modeling more realistic.\n3. A high-quality controlled mass cycling dataset is built to avoid external interferences, providing reliable support for verifying the model’s performance in bicycle dynamics and social interactions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Risk of Circular Logic in the Core Innovation: The \"anticipation mechanism\" in the social module requires predicting the future trajectories of neighboring agents (using a simple const.v model, mentioned in ilne 239) to serve as input for forecasting the ego agent's future. This creates a potential circular argument: predicting agent A's future relies on first predicting agent B's future, which is itself a challenging prediction problem. The model sidesteps this fundamental issue rather than solving it. If the const.v predictions for neighbors are unreliable, this \"anticipated\" input may introduce noise rather than beneficial information.\n2. Unclear Motivation for Physics Models: The physics module ensembles four models, but their fusion is performed opaquely through LSTM encoding and concatenation. The paper fails to justify why this specific combination of models is necessary and sufficient, nor does it provide significance analysis to demonstrate each model's unique contribution. Notably, since the simplest const.v model performs best among the individual baselines (Table 1), the motivation for including the more complex and poorer-performing kinematic and xkalman models is questionable. This appears more like model stacking than a deliberate design. An ablation study comparing different subsets (e.g., 2 or 3 models) is needed to substantiate that using all four is optimal.\n3. Writing: Line 68 appears to have a missing citation.\n4. Figures: Figure 1 is blurring, and the overlaid trajectories are difficult to discern.\n5. Social Graph Construction: Line 284 states \"at most five neighbors at a distance below 20m are considered,\" but the specific selection strategy (e.g., the five closest? random selection?) is not specified. This strategy can significantly impact the results and should be discussed or ablated.\n6. Lack of Novelty: The core methodology primarily combines existing techniques: GATs for social modeling (from Social-BiGAT), physics model ensembling, and multimodal output (GMM)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918347748,"tcdate":1762068043755,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5917/Reviewer_HCK9"],"signatures":["ICLR.cc/2026/Conference/Submission5917/Reviewer_HCK9"],"forum":"7RyvWhdiqp","number":4,"license":"CC BY 4.0","cdate":1762068043755,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5917/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918347748,"domain":"ICLR.cc/2026/Conference","replyto":"7RyvWhdiqp","id":"uuuovyWFxB","forumContent":{"TLDR":{"value":"We present the Great GATsBi, a domain-knowledge-based, hybrid, multimodal trajectory prediction framework for bicycles."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Trajectory Forecasting","Behavioral Modeling","Physics-informed Machine Learning","Bicycles","Autonomous Driving"]},"supplementary_material":{"value":"/attachment/5906e91492901afdf8ee22f4b802ec7dc64808be.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Accurate prediction of road user movement is increasingly required by many applications ranging from advanced driver assistance systems to autonomous driving, and especially crucial for road safety. Even though most traffic accident facilities account to bicycles, they have received little attention, as previous work focused mainly on pedestrians and motorized vehicles. In this work, we present the Great GATsBi, a domain-knowledge-based, hybrid, multimodal trajectory prediction framework for bicycles. The model incorporates both physics-based modeling (inspired by motorized vehicles) and social-based modeling (inspired by pedestrian movements) to explicitly account for the dual nature of bicycle movement. The social interactions are modeled with a graph attention network, and include decayed historical, but also anticipated, future trajectory data of a bicycles neighborhood, following recent insights from psychological and social studies. The results indicate that the proposed ensemble of physics models - performing well in the short-term predictions - and social models - performing well in the long-term predictions - exceeds state-of-the-art performance. We also conducted a controlled mass-cycling experiment to demonstrate the framework's performance when forecasting bicycle trajectories and modeling social interactions with road users."},"_bibtex":{"value":"@misc{\nriehl2026great,\ntitle={Great {GAT}sBi: Hybrid, Multimodal, Trajectory Forecasting for Bicycles using Anticipation Mechanism},\nauthor={Kevin Riehl and Shaimaa K. El-Baklish and Anastasios Kouvelas and Michail A. Makridis},\nyear={2026},\nurl={https://openreview.net/forum?id=7RyvWhdiqp}\n}"},"title":{"value":"Great GATsBi: Hybrid, Multimodal, Trajectory Forecasting for Bicycles using Anticipation Mechanism"},"pdf":{"value":"/pdf/79cd8d22d3f5a87819a17cf521089ee50965b290.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"riehl|great_gatsbi_hybrid_multimodal_trajectory_forecasting_for_bicycles_using_anticipation_mechanism"},"authorids":{"value":["~Kevin_Riehl1","~Shaimaa_K._El-Baklish1","~Anastasios_Kouvelas1","~Michail_A._Makridis1"]},"authors":{"value":["Kevin Riehl","Shaimaa K. El-Baklish","Anastasios Kouvelas","Michail A. Makridis"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a physics-in-the-loop scheme for the design of architectural structures. The authors here use a neural network(MLP) to learn the mapping from the desired structural shape to intermediate mechanical properties (bar stiffnesses) and then use a physics model to generate a physically feasible shape approximating the target.  The authors then apply this method to the design of masonry shells and cable towers,  comparing it with two neural network baselines: one trained to produce physically feasible shapes and the other trained to produce feasible shapes while also ensuring mechanical stability. Results on both these case studies show the proposed approach outperforming both these baselines and being on par with numerical optimization."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Overall, the paper is technically sound and proposes an interesting integration of ML and physics. However, there are some concerns as follows:\n\n1. The proposed hybrid approach requires running the physics model during inference, likely increasing training costs. While the authors report inference time, the training time is not provided. Given the limited representational capacity of the network, it must be retrained whenever the design is re-parameterized (a likely scenario during conceptualization). Reporting training times would enable a better assessment of the method’s real-world viability; if training significantly exceeds the duration of several optimizations, direct optimization might be preferable.\n\n2. Similarly, providing training time metrics for the other baselines would help clarify trade-offs and be useful in scenarios where slight accuracy losses are acceptable for performance gains.\n\n3. From Fig.9 , the optimization initialized with the proposed method converges quickly as compared to the other initializations. This leads to an interesting question of how the NN initialized optimizations would perform. This approach would have the benefit of outputting a guaranteed local minima while avoiding the additional physics overhead during training. Including these results would enhance the paper's contribution to the community.\n\n4. The case studies considered here have relatively low DOFs, and the physics relies on a linear FDM. The authors could discuss the viability of this approach for structures where linear FDM is inapplicable or for dynamic scenarios. In such cases—and even when FDM is applicable but the structure has much higher DOFs—would this approach remain feasible?\n\n5. Furthermore, the variation in MLP and PINN inference times is puzzling. Since the input sizes remain constant, one would expect the inference times for the fully trained networks to be similar. Could the authors comment on this discrepancy?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The idea of combining physics with ML for architectural design is promising, as it removes the need for generating costly training labels. Instead, the model can learn by integrating the physics model with inexpensive loss functions.\n2. The out-of-distribution performance is also interesting, as it potentially reduces the need for extensive variability in the training data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The case studies provided are relatively simplistic and do not reflect real-world applications. The physics model used is also quite basic, limiting the method's applicability in practical scenarios.\n2. Additionally, as the authors acknowledge, even when trained, the proposed parameterization lacks flexibility and requires retraining whenever the design representation changes."}},"nonreaders":[],"tmdate":1731428823255,"tcdate":1730783683819,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11555/Reviewer_NVMt"],"signatures":["ICLR.cc/2025/Conference/Submission11555/Reviewer_NVMt"],"forum":"Tpjq66xwTq","number":4,"license":"CC BY 4.0","cdate":1730783683819,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11555/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428823255,"domain":"ICLR.cc/2025/Conference","replyto":"Tpjq66xwTq","id":"h39cLoAY6R","forumContent":{"TLDR":{"value":"We couple neural networks with a differentiable mechanics simulator to accelerate the solution of shape-matching problems for mechanical design."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Differentiable physics","mechanical design","physics-in-the-loop neural networks","inverse problems","architectural structures"]},"supplementary_material":{"value":"/attachment/2784b8e0c84714c2f2410c81d51183fec7ccdff1.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Designing mechanically efficient geometry for architectural structures like shells, towers, and bridges, is an expensive iterative process.\nExisting techniques for solving such inverse problems rely on traditional optimization methods, which are slow and computationally expensive, limiting iteration speed and design exploration.\nNeural networks would seem to offer a solution via data-driven amortized optimization, but they often require extensive fine-tuning and cannot ensure that important design criteria, such as mechanical integrity, are met.\nIn this work, we combine neural networks with a differentiable mechanics simulator to develop a model that accelerates the solution of shape approximation problems for architectural structures represented as bar systems.\nThis model explicitly guarantees compliance with mechanical constraints while generating designs that closely match target geometries.\nWe validate our approach in two tasks, the design of masonry shells and cable-net towers.\nOur model achieves better accuracy and generalization than fully neural alternatives, and comparable accuracy to direct optimization but in real time, enabling fast and reliable design exploration.\nWe further demonstrate its advantages by integrating it into 3D modeling software and fabricating a physical prototype.\nOur work opens up new opportunities for accelerated mechanical design enhanced by neural networks for the built environment."},"_bibtex":{"value":"@inproceedings{\npastrana2025realtime,\ntitle={Real-time design of architectural structures with differentiable mechanics and neural networks},\nauthor={Rafael Pastrana and Eder Medina and Isabel M. de Oliveira and Sigrid Adriaenssens and Ryan P Adams},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Tpjq66xwTq}\n}"},"title":{"value":"Real-time design of architectural structures with differentiable mechanics and neural networks"},"pdf":{"value":"/pdf/9e8177c8e24077c2b8e64f036a1f189ed28f3c7f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"pastrana|realtime_design_of_architectural_structures_with_differentiable_mechanics_and_neural_networks"},"authorids":{"value":["~Rafael_Pastrana1","~Eder_Medina1","~Isabel_M._de_Oliveira1","~Sigrid_Adriaenssens1","~Ryan_P_Adams1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rafael Pastrana","Eder Medina","Isabel M. de Oliveira","Sigrid Adriaenssens","Ryan P Adams"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Lean4PHYS, which is a Lean4-based framework for formal physics. Lean4PHYS includes PhysLib (a repository of a physics unit system and commonly used theorems) and LeanPhysBench (a benchmark of 200 hand-crafted theorems from high school competitions to elementary college level)."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- Did the authors attempt to compare the accuracy of the models in this Lean 4-based setup vs natural setup (i.e., asking the model the question in natural language)?\n    - This would be important to check if the bottleneck is in the physics understanding or in the Lean 4 code understanding\n- I would expect PhysLib to be used by the model in a tool-calling fashion such that we do not need to present all the available concepts to the model in the context window. Is that the case?\n- L377-379: “models with weaker in-context learning perform reatively badly on this level of problems. It is because they cannot infer the new out-of-distribution syntax or unit-handling rules from context.” → This seems to be an overclaiming since none of the experiments are checking the in-context capabilities. Not to mention, we cannot confidently say that this data is OOD because we do not have access to the pretraining data of the models. Am I understanding the sentence properly?"},"rating":{"value":2},"details_of_ethics_concerns":{"value":"The paper uses copyrighted materials from publishers such as Pearson and Science Press. The authors noted that they \"reformulated and rephrased\" the materials; however, I am not certain that this sufficiently addresses the copyright limitations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"The first framework for physics problems in Lean4 for LLM. This may be of significant interest to the community that studies LLM applications to physics. The PhysLib library and LeanPhysBench test dataset can be a useful artefact for future studies."},"flag_for_ethics_review":{"value":["Yes, Legal compliance (e.g., GDPR, copyright, terms of use, web crawling policies)"]},"weaknesses":{"value":"- Copyright infringement.\n    - **Due to this issue, I decide to assign a very low score to the paper albeit the important contribution. However, I am very open to changing my score once this issue is clarified/addressed.**\n    - The authors mentioned that “rather than copying the questions verbatim, we reformulated and rephrased them based on the underlying physics ideas\", however, I am not certain that this is sufficient. The key issue hinges on \"substantial similarity\" and whether the original work's creative expression has been copied, even in a modified form. Given that the underlying physics idea of the questions are copied (perhaps to the point that there exists a one-to-one mapping between the textbooks and the dataset questions), this seems to constitute substantial similarity. The authors may need to ask for **explicit permission** from the publishers.\n- Statistical robustness\n    - Given that the evaluations were done with non-zero temperature, the authors should report the stochasticity of the results (e.g., standard error).\n- Lack of implementation elaboration\n    - How do the authors present PhysLib to the model? Would it fit into the context window?\n    - For non-experts, it is challenging to understand what the task looks like, particularly because the prompt asks the model to “complete the following Lean4 code” instead of the commonly known question-answering setup."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358100461,"tcdate":1761862920991,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7797/Reviewer_6F2Q"],"signatures":["ICLR.cc/2026/Conference/Submission7797/Reviewer_6F2Q"],"forum":"wQ2jyFz18H","number":3,"license":"CC BY 4.0","cdate":1761862920991,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7797/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358100461,"domain":"ICLR.cc/2026/Conference","replyto":"wQ2jyFz18H","id":"C5VFMjlGNr","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"This paper presents Lean4PHYS, a reasoning framework for college-level physics problems in Lean4. It includes LeanPhysBench, the first benchmark in the field, and PhysLib, a community-driven repository that sets the foundation for the field."},"keywords":{"value":["Lean4","Reasoning","AIforScience"]},"supplementary_material":{"value":"/attachment/fbd9bc9c890ab7276d2401a43a0d60dbc1501c55.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present **Lean4PHYS**, a comprehensive reasoning framework for college-level physics problems in Lean4. To establish a solid foundation for formal reasoning in physics, **Lean4PHYS** launches *PhysLib*, a repository containing fundamental unit systems and essential theorems to formulate physics proofs in Lean4. It will be community-driven and long-term maintained. Lean4PHYS also includes *LeanPhysBench*, a college-level benchmark for evaluating LLMs' Lean4 formal physics reasoning capability. It contains 200 hand-crafted and peer-reviewed Lean4 theorem statements formalized from university textbooks and physics competition problems. Based on the *PhysLib* and *LeanPhysBench* we composed in **Lean4PHYS**, we perform exhaustive experiments of baseline results using major expert Math provers and state-of-the-art closed-source models, and provide an analysis of their performance. In the experiment, we identify that most expert provers do not outperform general models as they did in the math domain. This suggests potential overfitting to the math domain rather than learning formal reasoning for formal provers. We also conduct a comprehensive experiment showing that, with *PhysLib* in the context, LLMs' performance on *LeanPhysBench* increases by **11.90%** on average, proving the effectiveness of our repository in assisting LLMs in solving the Lean4 physics problem. To the best of our knowledge, we are the first study to provide a physics benchmark in Lean4."},"_bibtex":{"value":"@inproceedings{\nli2026leanphysics,\ntitle={Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4},\nauthor={Yuxin Li and Minghao LIU and Ruida WANG and JI WenZhao and Zhitao He and Rui Pan and Junming Huang and Tong Zhang and Yi R. Fung},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=wQ2jyFz18H}\n}"},"title":{"value":"Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4"},"pdf":{"value":"/pdf/71e0f9a2255dd3fbd4264072b0433afaf0786606.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|lean4physics_comprehensive_reasoning_framework_for_collegelevel_physics_in_lean4"},"authorids":{"value":["~Yuxin_Li13","~Minghao_LIU8","~Ruida_WANG1","~JI_WenZhao1","~Zhitao_He1","~Rui_Pan4","~Junming_Huang1","~Tong_Zhang2","~Yi_R._Fung1"]},"authors":{"value":["Yuxin Li","Minghao LIU","Ruida WANG","JI WenZhao","Zhitao He","Rui Pan","Junming Huang","Tong Zhang","Yi R. Fung"]}},"version":2},{"content":{"summary":{"value":"The paper presents a novel physics-based human motion capture method that is physically explainable, conforming to the PD control theory and rigid body dynamics. The key designs of the method involve an integration of a kinematic Kalman filter and Newtonian equation-based physics simulation, and learnable Kalman gains, PD gains, external forces and robot inertia biases. The method outperforms existing kinematics-based and physics-based motion capture methods on keypoint accuracies."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Typos:\n\n* In Figure 2, \"$q_{t|t-1}$\" -> \"$q_{t+1|t}$\"\n* In the caption of Figure 2, \"performs\" is redundant\n* In Equation 6, \"$\\sum_c^2$\" -> \"$\\sum_{c=1}^2$\"\n* In line 295, \"HMDCap\" -> \"OSDCap\"\n* In line 478, \"the\" is redundant"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"(1) The method is physically explainable without unrealistic approximations of the control process and the robot dynamics.\n\n(2) Under the paradigm of using physics simulation to capture human motion, the method provides novel insights about which physical properties should be modeled by neural networks.\n\n(3) The method is superior to previous kinematics-based and physics-based methods in the accuracy of joint predictions.\n\n(4) The writing is clear and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) The contact modeling only considers foot-ground contact, ignoring full-body contact that commonly appears in human-object and human-scene interaction scenarios. Besides, the contact on each foot is represented as a force vector on a pre-defined contact point, ignoring changes in the contact point and the resultant torque of the contact.\n\n(2) The method updates the inertia matrix $M$ online. However, the inertia matrix is the attribute of the robot and should be fixed values during the whole motion capture process for better physical interpretability.\n\n(3) To fully examine the generalizability of the proposed method, existing physics-based methods should also be compared on datasets Fit3D and SportsPose."},"limitations":{"value":"One limitation is that the human dynamic model is formulated as a connection of circles and cylinders, which neglects the modeling of geometric details and wearings of humans."}},"nonreaders":[],"tmdate":1730878640107,"tcdate":1720611772186,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission374/Reviewer_QEqx"],"signatures":["NeurIPS.cc/2024/Conference/Submission374/Reviewer_QEqx"],"forum":"RkOT8rAmRR","number":2,"license":"CC BY 4.0","cdate":1720611772186,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission374/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878640107,"domain":"NeurIPS.cc/2024/Conference","replyto":"RkOT8rAmRR","id":"0bJc7wio71","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["human motion","dynamics","optimal state","kalman filter","physics-based"]},"supplementary_material":{"value":"/attachment/b72ba1da11508232b4d347e8798c559c3b5203ba.zip"},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Human motion capture from monocular videos has made significant progress in recent years. However, modern approaches often produce temporal artifacts, e.g. in form of jittery motion and struggle to achieve smooth and physically plausible motions. Explicitly integrating physics, in form of internal forces and exterior torques, helps alleviating these artifacts. Current state-of-the-art approaches make use of an automatic PD controller to predict torques and reaction forces in order to re-simulate the input kinematics, i.e. the joint angles of a predefined skeleton. However, due to imperfect physical models, these methods often require simplifying assumptions and extensive preprocessing of the input kinematics to achieve good performance. To this end, we propose a novel method to selectively incorporate the physics models with the kinematics observations in an online setting, inspired by a neural Kalman-filtering approach. We develop a control loop as a meta-PD controller to predict internal joint torques and external reaction forces, followed by a physics-based motion simulation. A recurrent neural network is introduced to realize a Kalman filter that attentively balances the kinematics input and simulated motion, resulting in an optimal-state dynamics prediction. We show that this filtering step is crucial to provide an online supervision that helps balancing the shortcoming of the respective input motions, thus being important for not only capturing accurate global motion trajectories but also producing physically plausible human poses. The proposed approach excels in the physics-based human pose estimation task and demonstrates the physical plausibility of the predictive dynamics, compared to state of the art. The code is available on https://github.com/cuongle1206/OSDCap."},"_bibtex":{"value":"@inproceedings{\nle2024optimalstate,\ntitle={Optimal-state Dynamics Estimation for Physics-based Human Motion Capture from Videos},\nauthor={Cuong Le and Manon Kok and Viktor Johansson and Bastian Wandt},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=RkOT8rAmRR}\n}"},"title":{"value":"Optimal-state Dynamics Estimation for Physics-based Human Motion Capture from Videos"},"pdf":{"value":"/pdf/325366bc6a69db293281709cbf852252b3527c07.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"le|optimalstate_dynamics_estimation_for_physicsbased_human_motion_capture_from_videos"},"authorids":{"value":["~Cuong_Le1","~Viktor_Johansson1","~Manon_Kok1","~Bastian_Wandt2"]},"authors":{"value":["Cuong Le","Viktor Johansson","Manon Kok","Bastian Wandt"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a multiphysics training scheme that jointly learns from full PDE simulations and their decomposed “basic forms” the authors argue that this injects “fundamental physics knowledge” into neural operators (NOs) and improves data efficiency and OOD generalization.  The key contribution lies in identifying and leveraging \"fundamental physics knowledge\" through decomposed basic PDE forms.  This has not  been explored extensively in the neural operator literature.  \nThe authors target two central SciML issues, 1. data hunger and 2. poor OOD transfer, for operator learning across 1D/2D/3D PDEs (Diffusion-Reaction, Navier–Stokes, Kuramoto–Sivashinsky, plus ScalarFlow). \nFormulations of PDEs and “basic forms” are standard and correctly specified\nThe paper is generally well written with helpful overview figures (Fig. 3 pipeline; Fig. 4 gallery of PDEs/basic forms) and plots tying simulation cost to nRMSE. Implementation, data splits, and training schedules are placed in appendices. \nMinor typos remain but do not impede readability.\nCentral claims are supported with experiments and results that are a bit light on content\nThe validation of physics (central theme) is light given no explicit checks on mass/energy conservations.  Another issue is the heuristic treatment of the fundamental physics term."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See weakness section , addressing those would be good\nrecommend to:\nBetter justify the decomposition principles\nProvide theoretical analysis or at least intuition for why the approach works\nCompare with more baselines\nDiscuss limitations and failure cases more thoroughly"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"no ethical issues"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The authors target two central SciML issues, 1. data hunger and 2. poor OOD transfer, for operator learning across 1D/2D/3D PDEs (Diffusion-Reaction, Navier–Stokes, Kuramoto–Sivashinsky, plus ScalarFlow). \nThe key contribution lies in identifying and leveraging \"fundamental physics knowledge\" through decomposed basic PDE forms.  This has not  been explored extensively in the neural operator literature.  \nProposed benefits such as:\nData efficiency, Long horizon stability , OOD generalization , are all desirable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The term \"fundamental physics knowledge\" is somewhat vague and could be better defined\nSection 3.1 could be more systematic in explaining the decomposition principles\nSome notation inconsistencies (e.g., switching between v and u for solutions\n\nMissing error bars in main results (added later in appendix)\nthere is limited statistical analysis\nThe ScalarFlow experiment (Section 4.5) is somewhat disconnected and brief\nClaims about \"fundamental physics knowledge\" being key are not fully validated (could be just multi-task regularization)\nThere is a lack of theoretical insight, no formal explanation of why the approach works beyond intuition is provided.\nThe decomposition rules in Section 3.1.1 seem ad-hoc without principled justification\n\nEvaluation is very basic, results only compare against vanilla baseline and spatiotemporal downsampling\nThere is no comparison/discussion with other data-efficient methods or recent foundation models\nReal-world evaluation is very limited (only ScalarFlow)\nInconsistent terminology: \"Fundamental physics knowledge\" vs \"basic forms\" used interchangeably\nMissing details: How are the mixture ratios exactly determined? Training time comparisons?\nLimited discussion: When would this approach fail? What about PDEs that don't decompose nicely?\nPresentation issues: Some figures (especially in appendix) are too small to read clearly\nScalability concerns: All experiments on relatively small-scale problems\n\nGrammer-Typos\nline 483: and outha ha h-of-distribution generalization."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925071492,"tcdate":1761788075973,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14708/Reviewer_kRiv"],"signatures":["ICLR.cc/2026/Conference/Submission14708/Reviewer_kRiv"],"forum":"mJiPqOzc3O","number":2,"license":"CC BY 4.0","cdate":1761788075973,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14708/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925071492,"domain":"ICLR.cc/2026/Conference","replyto":"mJiPqOzc3O","id":"g7p98K39ls","forumContent":{"TLDR":{"value":"We propose to incorporate fundamental physics knowledge into learning neural operators to enhance its data efficiency, long-term consistency, and OOD generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Neural Operator","PDE","Fundamental Physics Knowledge"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Recent advances in scientific machine learning (SciML) have enabled neural operators (NOs) to serve as powerful surrogates for modeling the dynamic evolution of physical systems governed by partial differential equations (PDEs). While existing approaches focus primarily on learning simulations from the target PDE, they often overlook more fundamental physical principles underlying these equations. Inspired by how numerical solvers are compatible with simulations of different settings of PDEs, we propose a multiphysics training framework that jointly learns from both the original PDEs and their simplified basic forms. Our framework enhances data efficiency, reduces predictive errors, and improves out-of-distribution (OOD) generalization, particularly in scenarios involving shifts of physical parameters and synthetic-to-real transfer. Our method is architecture-agnostic and demonstrates consistent improvements in normalized root mean square error (nRMSE) across a wide range of 1D/2D/3D PDE problems. Through extensive experiments, we show that explicit incorporation of fundamental physics knowledge significantly strengthens the generalization ability of neural operators.\nWe will release models and codes at https://sites.google.com/view/sciml-fundemental-pde."},"_bibtex":{"value":"@inproceedings{\nma2026learning,\ntitle={Learning Data-Efficient and Generalizable Neural Operators via Fundamental Physics Knowledge},\nauthor={Siying Ma and Mehrdad Momeni Zadeh and Mauricio Soroco and Wuyang Chen and Jiguo Cao and Vijay Ganesh},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=mJiPqOzc3O}\n}"},"title":{"value":"Learning Data-Efficient and Generalizable Neural Operators via Fundamental Physics Knowledge"},"pdf":{"value":"/pdf/27f1b69d2b552cb0e5d0a96e3231fc4148675ae2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ma|learning_dataefficient_and_generalizable_neural_operators_via_fundamental_physics_knowledge"},"authorids":{"value":["~Siying_Ma1","~Mehrdad_Momeni_Zadeh1","~Mauricio_Soroco1","~Wuyang_Chen1","~Jiguo_Cao1","~Vijay_Ganesh1"]},"authors":{"value":["Siying Ma","Mehrdad Momeni Zadeh","Mauricio Soroco","Wuyang Chen","Jiguo Cao","Vijay Ganesh"]}},"version":2},{"content":{"summary":{"value":"This paper proposes to feed complex-valued deep learning networks with coefficients extracted of Complex-valued Scattering Representations (CSR) of input data when dealing with either complex or real input samples in a supervised learning scenario. CSR consists on the application of a sequence of layers composed by convolutions with Morlet filters with learnable parameters followed by a novel learnable activation function. The complex-valued coefficients obtained at each CSR layer are subsequently fed into traditional complex-valued deep learning models. Empirical findings demonstrate that the introduction of CSR can lead to improved classification performance, particularly when dealing with a limited number of training data samples."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The main strength of this study lies in its provision of an effective decomposition of a real- or complex-valued signal. This decomposition enables successful classification by deep complex-valued neural networks, which have gained prominence in recent literature due to advances in manifold geometry and group theory.\n\n- Another notable contribution is the introduction of a novel complex-valued activation function equipped with learnable parameters, which has demonstrated empirical improvements in classification performance.\n\n- The experimental results presented in this work convincingly showcase the superiority of the proposed approach when compared to conventional complex-valued deep neural networks without using the proposed CSR."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Main issues\n\n-\tAlthough there is an intuitive understanding of why the suggested CSR offers improved data representation for classification, the paper falls short in terms of providing a solid theoretical foundation to comprehensively explain its underlying principles and limitations.\n\n-\tThe contributions outlined in the paper signify an evolutionary enhancement of previously introduced complex-valued deep neural networks.\n\nMinor issues\n\n-\tBefore eq. (3): “w is the learnable parameters” ->  “w is the vector of learnable parameters”"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"I notice that a particular linear transformation of real-valued input data is applied to convert it into a complex-valued input (sliding color encoding proposed in Singhal et al., 2022a). I have a few questions regarding this selection:\n\n1.\tGiven the numerous possible methods to transform real-valued data into complex-valued inputs, why was this particular transformation chosen as the preferred option?\n\n2.\tIs there a conceptual or empirical rationale behind the adoption of this specific transformation?\n\n3.\tHave any attempts been made to introduce a trainable linear transformation, allowing it to be adjusted during training? Do you believe that this approach might offer any advantages?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636436794,"tcdate":1697751647358,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4587/Reviewer_G9gH"],"signatures":["ICLR.cc/2024/Conference/Submission4587/Reviewer_G9gH"],"forum":"z9ySIS1inA","number":2,"license":"CC BY 4.0","cdate":1697751647358,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4587/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636436794,"domain":"ICLR.cc/2024/Conference","replyto":"z9ySIS1inA","id":"novUq6zhii","forumContent":{"TLDR":{"value":"A Novel and Universal Complex-valued Representation for Complex-valued Deep Learning"},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Complex-valued Deep Learning","Scattering Representations","Representation Learning","Training with limited-labled data"]},"supplementary_material":{"value":"/attachment/c8a8892c8d20f2f83c5975cab8c82e2d4760bfff.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Complex-valued deep learning has made significant progress with manifold geometry and group theory. It delivers leaner and better classifiers with novel complex-valued layer functions and network architectures, not only on naturally complex-valued data such as Magnetic Resonance imaging (MRI) but also on real-valued data such as RGB or multi-spectral images. However, current complex-valued representations for complex-valued and real-valued inputs are rudimentary, focusing on channel characteristics (e.g., sliding encoding) without capturing spatial and spatial-frequency properties of the input data. We propose Complex-valued Scattering Representations (CSR) as universal complex-valued representations and integrate them into complex-valued deep learning networks.  To obtain CSR, We construct filters based on complex-valued Morlet wavelets with tunable parameters and develop learnable high-dimensional complex-valued ReLU as the non-linear activation function.  By incorporating these novel components into complex-valued models, our models significantly outperform real-valued counterparts and existing complex-valued models on RGB, multi-spectral image (MSI), and MRI patch classification tasks, especially under limited labeled training data settings, greatly enhancing complex-valued networks on a broader range of applications."},"_bibtex":{"value":"@misc{\nwang2024complexvalued,\ntitle={Complex-valued Scattering Representations},\nauthor={Ke Wang and Utkarsh Singhal and Michael Lustig and Stella X. Yu},\nyear={2024},\nurl={https://openreview.net/forum?id=z9ySIS1inA}\n}"},"title":{"value":"Complex-valued Scattering Representations"},"pdf":{"value":"/pdf/e25a1b9fe25f9d17ea0545c45fdc1c379c459088.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|complexvalued_scattering_representations"},"authorids":{"value":["~Ke_Wang8","~Utkarsh_Singhal1","~Michael_Lustig2","~Stella_X._Yu2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ke Wang","Utkarsh Singhal","Michael Lustig","Stella X. Yu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PRISM-Physics, a large-scale physics reasoning benchmark with a proposed DAG-based evaluation protocol that addresses the limitations of existing LLM-as-judge scoring methods. The evaluation framework includes a fully rule-based symbolic formula equivalence checker to ensure consistent validation across diverse mathematical formulations, thereby eliminating reliance on subjective judgments. In the experiments, the paper investigates a diverse set of leading LLMs on PRISM-Physics and demonstrates the superiority of the proposed evaluation protocol compared to the LLM-as-judge method."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses 1, 2."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The idea of using DAG to judge the correctness of the final answer and intermediate steps is reasonable, and I agree that using the LLM-as-judge method to evaluate the correctness of physics problems is challenging and prone to errors.\n2. The theoretical analysis part of the paper is solid.\n3. The experiment is comprehensive and convincing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. My main concern is that, although PRISM-Physics can conduct rule-based judgments to determine the correctness of the final answer and intermediate steps using a DAG, the construction of the DAG still heavily relies on LLM-based extraction and rewriting. Compared to the existing LLM-as-judge method, the uncertainty introduced by LLMs seems to have merely shifted from the judgment stage to the preprocessing stage.\n2. Another concern lies in the scalability and additional computational cost of the proposed evaluation protocol. Compared to existing benchmarks, PRISM-Physics requires an annotated DAG in addition to the final answer for each question in order to perform a more rigorous evaluation. Thus, the scalability of the proposed protocol appears limited. If we aim to extend this rigorous protocol to other existing benchmarks, what additional requirements would those questions need to meet? Furthermore, if we intend to construct a DAG for a new physics problem, how much extra computational cost would this introduce in the preprocessing stage?\n3. Typo: in Line 362: \"zero-shot **COTzheg** prompts\". In Table 1, it would be better to retain the same number of digits after the decimal point and to bold the best results.\n4. The results in Figure 8 (Appendix E.2) are difficult to understand. The authors should at least explain the meaning of each rectangle in the text and clarify whether the difference shown represents \"multimodal – text\" or \"text – multimodal\"."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921360718,"tcdate":1761790096247,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9900/Reviewer_4ABk"],"signatures":["ICLR.cc/2026/Conference/Submission9900/Reviewer_4ABk"],"forum":"4PZMeopXzP","number":3,"license":"CC BY 4.0","cdate":1761790096247,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9900/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921360718,"domain":"ICLR.cc/2026/Conference","replyto":"4PZMeopXzP","id":"umUVtNSp82","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We present PRISM-Physics, a benchmark and a process-level evaluation framework that encodes physics solutions as DAGs and employs rule-based symbolic equivalence checking for reliable, fine-grained scoring."},"keywords":{"value":["Physics Reasoning","Process-Level Evaluation","Symbolic Equivalence","Scientific Problem Solving"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Benchmarks for competition-style reasoning have advanced evaluation in mathematics and programming, yet physics remains comparatively underexplored. Most existing physics benchmarks evaluate only final answers, which fail to capture reasoning processes, while recent stepwise methods rely on heuristic LLM-as-judge scoring or restrictive linear assumptions, limiting reliability and diagnostic validity.\nWe introduce PRISM-Physics, a process-level evaluation framework and benchmark for complex physics reasoning problems. Solutions are represented as directed acyclic graphs (DAGs) of formulas, explicitly encoding causal dependencies among intermediate steps to enable fine-grained, interpretable, and theoretically grounded scoring. \nWe prove the optimality of the DAG representation and the corresponding scoring policy. Combining with a fully rule-based method for symbolic formula equivalence matching that we developed, we ensure consistent validation across diverse formulations without heuristic judgments. Results show that our evaluation framework is more aligned with human experts' scoring. \nExperiments on state-of-the-art LLMs reveal persistent reasoning failures in physics, while step-level scoring offers both diagnostic insight and rich signals for later training. By combining structural rigor, theoretical guarantees, and symbolic validation, PRISM-Physics provides a principled foundation for advancing process-level evaluation and guiding the development of models with deeper scientific reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\nzhao2026prismphysics,\ntitle={{PRISM}-Physics: Causal {DAG}-Based Process Evaluation for Physics Reasoning},\nauthor={Wanjia Zhao and Qinwei Ma and Jingzhe Shi and Shirley Wu and Jiaqi Han and Yijia Xiao and Si-Yuan Chen and Xiao Luo and Ludwig Schmidt and James Zou},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=4PZMeopXzP}\n}"},"title":{"value":"PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning"},"pdf":{"value":"/pdf/95b751e95b88437e4484cd3de1a315d0a89884f4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhao|prismphysics_causal_dagbased_process_evaluation_for_physics_reasoning"},"authorids":{"value":["~Wanjia_Zhao1","~Qinwei_Ma1","~Jingzhe_Shi1","~Shirley_Wu1","~Jiaqi_Han2","~Yijia_Xiao1","~Si-Yuan_Chen1","~Xiao_Luo3","~Ludwig_Schmidt1","~James_Zou1"]},"authors":{"value":["Wanjia Zhao","Qinwei Ma","Jingzhe Shi","Shirley Wu","Jiaqi Han","Yijia Xiao","Si-Yuan Chen","Xiao Luo","Ludwig Schmidt","James Zou"]}},"version":2},{"content":{"summary":{"value":"Large language models have achieved strong performance on many natural language processing tasks, but they struggle with complex reasoning tasks. Many existing studies use synthetic data to train large language models to improve their reasoning ability. However, these studies often lack disciplinary guidance. \n\nTo address this issue, this paper uses some unlabeled documents to generate synthetic data, including books and web data. Specifically, the process is divided into three parts: Data Extraction and Preprocessing, Code Synthesis, and Filtering and Output. The authors conduct an in-depth analysis of the synthetic data, examining data difficulty, diversity, and disciplinary distribution. \n\nFinally, they train the model using the synthetic data and find that it leads to better training results."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"+ The paper tests multiple-choice tasks. Does it also perform well on open-ended questions, such as GSM8k or MATH500?\n\n+ Why does combining web data and book data lead to better results? What are the differences and characteristics of these two types of data when merged?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"+ This paper proposes a new method for data synthesis that can generate useful data for the model from unlabeled documents.\n\n+ This paper conducts an in-depth analysis of the data, including data difficulty, data diversity, and disciplinary distribution, helping readers better understand the synthetic data and providing more information for future research.\n\n+ The synthetic data in this paper can effectively improve the model's reasoning ability and proves to be effective across tests in multiple disciplines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"+ Previous studies have also generated training data from unlabeled data (such as DeepSeek-Math, JiuZhang3.0, and MAmmoTH 2.0). Even without using complex logic extraction and logic retrieval processes, their results are more significant than existing methods.\n\n+ The process of generating synthetic data uses very large models, which incurs high costs. Does this indicate that the method has limitations in practical applications?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915749413,"tcdate":1760604512966,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1366/Reviewer_wEhN"],"signatures":["ICLR.cc/2026/Conference/Submission1366/Reviewer_wEhN"],"forum":"SQVxBJhIrK","number":1,"license":"CC BY 4.0","cdate":1760604512966,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1366/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915749413,"domain":"ICLR.cc/2026/Conference","replyto":"SQVxBJhIrK","id":"V4FeYZSbGt","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Large Language Models","Data Synthesis","Synthetic Data","Reasoning","Post-Training","Supervised Fine-Tuning"]},"supplementary_material":{"value":"/attachment/295edaaa8d16fab031e4c61c54a7efd46c00b736.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large language models (LLMs) perform strongly on many language tasks but still struggle with complex multi-step reasoning across disciplines. Existing reasoning datasets often lack disciplinary breadth, reasoning depth, and diversity, as well as guiding principles for question synthesis. We propose DESIGNER: a DESIGN-logic-guidEd Reasoning data synthesis pipeline that leverages naturally available, extensive raw documents to generate multidisciplinary questions. The central insight is the notion of Design Logic, a form of reusable meta-knowledge that encapsulates the structured process human experts use to transform knowledge into complex exam questions, enabling LLMs to generate new questions with the same complex reasoning patterns from entirely different source texts with explicit control over difficulty, diversity, and question types. We use LLMs to reverse-engineer and abstract over 120,000 Design Logics from existing questions across various disciplines. By designing a two-stage retrieve-and-generate mechanism to match these Design Logics with raw corpus, we synthesized two large-scale reasoning datasets that span 75 disciplines: DLR-Book (3.04 million questions from the book corpus) and DLR-Web (1.66 million questions from the web corpus). Data analysis indicates that the questions synthesized by our method exhibit greater difficulty and diversity compared to those in the baseline datasets. Supervised fine-tuning (SFT) on Qwen3 and Llama3 with our data substantially improves multidisciplinary reasoning and outperforms baseline datasets. Notably, by applying SFT on the base versions of these models using only our data, we even surpass their official final models that have undergone the full post-training."},"_bibtex":{"value":"@inproceedings{\nliu2026designer,\ntitle={{DESIGNER}: Design-Logic-Guided Multidisciplinary Data Synthesis for {LLM} Reasoning},\nauthor={Weize Liu and Yongchi Zhao and Yijia Luo and Mingyu Xu and Jiaheng Liu and Yanan Li and Xiguo Hu and ZhiqiBai and Yuchi Xu and Wenbo Su and Bo Zheng},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=SQVxBJhIrK}\n}"},"title":{"value":"DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning"},"pdf":{"value":"/pdf/f8a592c0240f08acdfa3e01977a7f32b90276ba8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|designer_designlogicguided_multidisciplinary_data_synthesis_for_llm_reasoning"},"authorids":{"value":["~Weize_Liu1","~Yongchi_Zhao1","~Yijia_Luo1","~Mingyu_Xu3","~Jiaheng_Liu1","~Yanan_Li8","~Xiguo_Hu1","~ZhiqiBai1","~Yuchi_Xu1","~Wenbo_Su2","~Bo_Zheng5"]},"authors":{"value":["Weize Liu","Yongchi Zhao","Yijia Luo","Mingyu Xu","Jiaheng Liu","Yanan Li","Xiguo Hu","ZhiqiBai","Yuchi Xu","Wenbo Su","Bo Zheng"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a benchmark that progressively increases task difficulty from basic synthetic retrieval to complex multi-step reasoning across three domains. It is used to evaluate 33 long-context models, and results uncover the limitations in current long-context capabilities that challenge existing claims of solved long-context understanding."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"What are the challenges/difficulties in constructing the benchmark? What can be learned from the evaluation results to improve the performance of the models?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"+ a new benchmark that progressively increases task difficulty from basic synthetic retrieval to complex multi-step reasoning across three domains\n+ comprehensive evaluation using 33 long-context models to uncover the limitations in current long-context capabilities"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- while the benchmark is good, the technical contributions are limited\n- the findings reported in the evaluation are less insightful"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917431443,"tcdate":1761997747333,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4545/Reviewer_MZgD"],"signatures":["ICLR.cc/2026/Conference/Submission4545/Reviewer_MZgD"],"forum":"GzzyuhP5Kz","number":3,"license":"CC BY 4.0","cdate":1761997747333,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4545/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917431443,"domain":"ICLR.cc/2026/Conference","replyto":"GzzyuhP5Kz","id":"rGyBtzliI0","forumContent":{"TLDR":{"value":"We propose a novel synthetic benchmark RULERv2 to evaluate long-context language models"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Long-context","Evaluation","Benchmark","Synthetic"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in long-context language models have spurred development of diverse benchmarks that often test multiple skills simultaneously, making it difficult to identify specific failure modes. To address this, we introduce RULERv2, a benchmark with systematic difficulty progression from basic synthetic retrieval to complex multi-step reasoning across three domains: multi-key NIAH, multi-value NIAH, and multi-doc QA.We conduct a large-scale evaluation of leading models, including seven closed-source and 26 open-weight models. Our findings reveal a notable performance gap between the two. Critically, we demonstrate that all models, including those claiming million-token context windows, exhibit performance degradation with increasing length, highlighting an unresolved challenge. Our analysis shows that explicit decomposition into a retrieve-then-solve strategy outperforms the implicit, single-step approach, and chain-of-thought reasoning enables models to discover effective decomposition autonomously. Finally, we find that even top-performing open-weight models struggle with fundamental retrieval and copying tasks, leading to degraded performance on more complex problems."},"_bibtex":{"value":"@misc{\nhsieh2026rulerv,\ntitle={{RULER}v2: From Basic Retrieval to Complex Reasoning, A Bottom-Up Benchmark for Long-Context Evaluation},\nauthor={Cheng-Ping Hsieh and Faisal Ladhak and Krishna C Puvvada and Boris Ginsburg},\nyear={2026},\nurl={https://openreview.net/forum?id=GzzyuhP5Kz}\n}"},"title":{"value":"RULERv2: From Basic Retrieval to Complex Reasoning, A Bottom-Up Benchmark for Long-Context Evaluation"},"pdf":{"value":"/pdf/df541773b0f7bcdd79bdb5da4411a7706761ff94.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"hsieh|rulerv2_from_basic_retrieval_to_complex_reasoning_a_bottomup_benchmark_for_longcontext_evaluation"},"authorids":{"value":["~Cheng-Ping_Hsieh1","~Faisal_Ladhak2","~Krishna_C_Puvvada1","~Boris_Ginsburg1"]},"authors":{"value":["Cheng-Ping Hsieh","Faisal Ladhak","Krishna C Puvvada","Boris Ginsburg"]}},"version":2},{"content":{"summary":{"value":"The paper introduces OmniChat, a novel spoken dialogue system enhanced by synthetic data for handling diverse scenarios. The key contributions include:\n1. ShareChatX - a large-scale synthetic spoken dialogue dataset covering various scenarios including emotional dialogues, audio events, and music contexts\n2. OmniChat - a multi-turn spoken dialogue system with a heterogeneous feature fusion module (Mix-Former) for optimizing feature selection across different dialogue contexts\n3. Comprehensive analysis of synthetic data usage in training spoken dialogue systems, including optimal ratios between synthetic and real data\n4. State-of-the-art performance achieved on the DailyTalk dataset and other complex dialogue scenarios"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Could you please share the ablation studies on different components of the Mix-Former architecture?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper addresses a critical challenge in spoken dialogue systems by leveraging synthetic data to overcome the scarcity of large-scale, high-quality spoken dialogue datasets.\n2. The proposed Mix-Former module effectively integrates multiple expert features (speech, emotion, beat) to handle diverse dialogue scenarios.\n3. Comprehensive experimental analysis provides valuable insights into optimal training strategies, including the ideal balance between synthetic and real data.\n4. The work demonstrates significant practical impact through state-of-the-art performance on real-world datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks detailed comparison with some recent baseline methods in spoken dialogue systems, particularly in terms of model size and computational requirements.\n2. The evaluation metrics could be more comprehensive, especially for measuring the quality of generated speech beyond just content and emotion accuracy.\n3. The methodology for ensuring quality control in synthetic data generation could be explained more thoroughly."}},"nonreaders":[],"tmdate":1731427733582,"tcdate":1730641580897,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3090/Reviewer_ns19"],"signatures":["ICLR.cc/2025/Conference/Submission3090/Reviewer_ns19"],"forum":"cVgOIjcNoQ","number":2,"license":"CC BY 4.0","cdate":1730641580897,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3090/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427733582,"domain":"ICLR.cc/2025/Conference","replyto":"cVgOIjcNoQ","id":"cCdr7zmrRw","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"Enhancing Spoken Dialogue Systems with Scalable Synthetic Data"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Spoken Dialogue System","Synthetic Data","Multi-modal Large Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these systems still struggle to handle the full complexity of real-world conversations, including audio events, musical contexts, and emotional expressions, mainly because current dialogue datasets are constrained in both scale and scenario diversity. In this paper, we propose leveraging synthetic data to enhance the dialogue models across diverse scenarios. We introduce **ShareChatX**, the first comprehensive, large-scale dataset for spoken dialogue that spans diverse scenarios. Based on this dataset, we introduce **OmniChat**, a multi-turn dialogue system with a heterogeneous feature fusion module, designed to optimize feature selection in different dialogue contexts. In addition, we explored critical aspects of training dialogue systems using synthetic data. Through comprehensive experimentation, we determined the ideal balance between synthetic and real data, achieving state-of-the-art results on the real-world dialogue dataset DailyTalk. We also highlight the crucial importance of synthetic data in tackling diverse, complex dialogue scenarios, especially those involving audio and music. For more details, please visit our demo page at \\url{https://sharechatx.github.io/}."},"_bibtex":{"value":"@misc{\ncheng2024omnichat,\ntitle={OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios},\nauthor={Xize Cheng and Dongjie Fu and Xiaoda Yang and Minghui Fang and Ruofan Hu and Jingyu Lu and Bai Jionghao and Zehan Wang and Shengpeng Ji and Rongjie Huang and Linjun Li and Yu Chen and Tao Jin and Zhou Zhao},\nyear={2024},\nurl={https://openreview.net/forum?id=cVgOIjcNoQ}\n}"},"title":{"value":"OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios"},"pdf":{"value":"/pdf/18710de70969375ea1129596e69146d27b8f844c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"cheng|omnichat_enhancing_spoken_dialogue_systems_with_scalable_synthetic_data_for_diverse_scenarios"},"authorids":{"value":["~Xize_Cheng1","~Dongjie_Fu1","~Xiaoda_Yang1","~Minghui_Fang1","~Ruofan_Hu2","~Jingyu_Lu1","~Bai_Jionghao2","~Zehan_Wang2","~Shengpeng_Ji1","~Rongjie_Huang1","~Linjun_Li2","~Yu_Chen29","~Tao_Jin2","~Zhou_Zhao3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xize Cheng","Dongjie Fu","Xiaoda Yang","Minghui Fang","Ruofan Hu","Jingyu Lu","Bai Jionghao","Zehan Wang","Shengpeng Ji","Rongjie Huang","Linjun Li","Yu Chen","Tao Jin","Zhou Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a comprehensive framework named Lean4PHYS, designed for formal reasoning on university-level physics problems using the Lean4 proof assistant. The framework consists of two core components: PhysLib, a community-driven library that provides a foundational unit system and commonly used theorems for formal physics reasoning; and LeanPhysBench, a benchmark dataset of 200 problems manually constructed and formalized from university textbooks and physics competitions. Based on this framework, the authors evaluate the performance of several mainstream large language models (including both general-purpose models and those specialized in Lean mathematical proofs). The experimental results show that general-purpose LLMs generally outperform math-specialized models on physics reasoning tasks, revealing a potential overfitting issue of the latter to the mathematics domain. Furthermore, the study demonstrates that using PhysLib as contextual information significantly improves the performance of all tested models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1.  The experimental results indicate that providing PhysLib as context to the models significantly improves their performance. Given that PhysLib, as a foundational library, could be quite large, the paper seems to lack a specific description of how it was effectively integrated into the model's prompt context window. Did the authors provide the entire library's content, or was some form of retrieval mechanism used to select relevant theorems and definitions? Clarifying this implementation detail is crucial for the reproducibility and understanding of the results."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"*   **Significant Contribution to the Research Community:** This paper contributes two extremely valuable resources: **PhysLib**, a modular and extensible foundational library for physics, and **LeanPhysBench**, the first benchmark dedicated to evaluating formal reasoning capabilities in physics. These two achievements provide a solid infrastructure and a fair evaluation standard for subsequent researchers to enter and work in this field, which will undoubtedly promote the development of the entire community.\n*   **Exhaustive Experiments and Deep Insights:** The paper's experimental design is very comprehensive, not only testing multiple top-tier general-purpose LLMs but also comparing them with several Lean-specialized models that excel in mathematics. The results reveal an important finding that \"specialized models have limited cross-domain (from math to physics) generalization ability,\" prompting deep reflection on model generalization and domain overfitting. At the same time, the experiments clearly quantify the effectiveness of the PhysLib library in assisting models with physics reasoning, proving its design value."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"*   **Insufficient Discussion of Related Work:** The \"Related Work\" section mentions Lean's application in physics and other non-mathematical fields but fails to deeply discuss the differences and connections with some directly related works. For example, the paper mentions the `PhysLean` [Tooby-Smith & contributors (2024)] project but only briefly describes it as \"theorem-specific, small-scale, and non-modular.\" Considering that `PhysLean` also aims to formalize physics in Lean4, the authors should have elaborated more on the fundamental differences and specific advantages of Lean4PHYS in terms of design philosophy, implementation methods (such as the construction of the unit system), coverage, and modular design compared to `PhysLean`. Adding such in-depth comparative analysis would better highlight the uniqueness and irreplaceable contribution of this work."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919841236,"tcdate":1761729154624,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7797/Reviewer_SHVy"],"signatures":["ICLR.cc/2026/Conference/Submission7797/Reviewer_SHVy"],"forum":"wQ2jyFz18H","number":2,"license":"CC BY 4.0","cdate":1761729154624,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7797/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919841236,"domain":"ICLR.cc/2026/Conference","replyto":"wQ2jyFz18H","id":"HPaYTuYLd2","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"This paper presents Lean4PHYS, a reasoning framework for college-level physics problems in Lean4. It includes LeanPhysBench, the first benchmark in the field, and PhysLib, a community-driven repository that sets the foundation for the field."},"keywords":{"value":["Lean4","Reasoning","AIforScience"]},"supplementary_material":{"value":"/attachment/fbd9bc9c890ab7276d2401a43a0d60dbc1501c55.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present **Lean4PHYS**, a comprehensive reasoning framework for college-level physics problems in Lean4. To establish a solid foundation for formal reasoning in physics, **Lean4PHYS** launches *PhysLib*, a repository containing fundamental unit systems and essential theorems to formulate physics proofs in Lean4. It will be community-driven and long-term maintained. Lean4PHYS also includes *LeanPhysBench*, a college-level benchmark for evaluating LLMs' Lean4 formal physics reasoning capability. It contains 200 hand-crafted and peer-reviewed Lean4 theorem statements formalized from university textbooks and physics competition problems. Based on the *PhysLib* and *LeanPhysBench* we composed in **Lean4PHYS**, we perform exhaustive experiments of baseline results using major expert Math provers and state-of-the-art closed-source models, and provide an analysis of their performance. In the experiment, we identify that most expert provers do not outperform general models as they did in the math domain. This suggests potential overfitting to the math domain rather than learning formal reasoning for formal provers. We also conduct a comprehensive experiment showing that, with *PhysLib* in the context, LLMs' performance on *LeanPhysBench* increases by **11.90%** on average, proving the effectiveness of our repository in assisting LLMs in solving the Lean4 physics problem. To the best of our knowledge, we are the first study to provide a physics benchmark in Lean4."},"_bibtex":{"value":"@inproceedings{\nli2026leanphysics,\ntitle={Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4},\nauthor={Yuxin Li and Minghao LIU and Ruida WANG and JI WenZhao and Zhitao He and Rui Pan and Junming Huang and Tong Zhang and Yi R. Fung},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=wQ2jyFz18H}\n}"},"title":{"value":"Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4"},"pdf":{"value":"/pdf/71e0f9a2255dd3fbd4264072b0433afaf0786606.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|lean4physics_comprehensive_reasoning_framework_for_collegelevel_physics_in_lean4"},"authorids":{"value":["~Yuxin_Li13","~Minghao_LIU8","~Ruida_WANG1","~JI_WenZhao1","~Zhitao_He1","~Rui_Pan4","~Junming_Huang1","~Tong_Zhang2","~Yi_R._Fung1"]},"authors":{"value":["Yuxin Li","Minghao LIU","Ruida WANG","JI WenZhao","Zhitao He","Rui Pan","Junming Huang","Tong Zhang","Yi R. Fung"]}},"version":2},{"content":{"TLDR":{"value":"Higher synthetic data novelties do not translate to downstream utility: in pore-pressure prediction, simple interpolation resamplers outperform complex physics and data-driven generators."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["pore pressure prediction","well logs","tabular data","synthetic data generation","physics-informed generative models","CTGAN","SMOTE","benchmark","TSTR","distributional fidelity","data scarcity","petrophysics"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Pore pressure prediction is essential for safe well planning, yet machine learning models are limited by the scarcity of real well data. Synthetic data generation is a common approach, including interpolation methods, physics-informed generators, and data-driven models. To our knowledge, no prior work systematically compares these approaches under a matched hyperparameter-tuning budget using both distributional and downstream evaluation. We introduced WellBench, a multi-basin benchmark, and a Physics-Optimized Forward Model (POFM), a physics-based scientific simulator for synthetic well-log generation that integrates petrophysical relationships and is calibrated to real well-log distributions using Tree-structured Parzen Estimation (TPE). We compared POFM against a TPE-tuned CTGAN, SMOTE, and Smoothed Bootstrap across four regions. POFM achieved the highest Distance-to-Closest-Record (DCR), indicating the greatest separation from real training records, while the interpolation generators best reproduced the real marginals but yielded the lowest DCR. Across seven downstream models, only the interpolation generators achieved positive blind-well Train-Synthetic-Test-Real ($R^2$) (up to $\\approx 0.98$), while CTGAN and POFM yielded negative $R^2$. Under a depth-held-out split, Train-Real-Test-Real (TRTR) failed at every real-data fraction, highlighting the challenge posed by limited real-data coverage. Yet TSTR with interpolation-based generators outperformed TRTR at every fraction. These results reveal a fundamental tradeoff between distributional fidelity and downstream utility. They also show that interpolation-based generators can outperform complex synthetic-data approaches when real data is scarce."},"_bibtex":{"value":"@inproceedings{\nanonymous2026wellbench,\ntitle={WellBench: An Open-Source Benchmark for{\\textbackslash}{\\textbackslash}Synthetic Well Log Generation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=r3SqIM6iat},\nnote={under review}\n}"},"title":{"value":"WellBench: An Open-Source Benchmark for\\\\Synthetic Well Log Generation"},"pdf":{"value":"/pdf/25dd1382fa592c710431d50587bfbedcd7598264.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791226736379,"tcdate":1789172026188,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission14668/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission14668/Authors"],"forum":"r3SqIM6iat","license":"CC BY 4.0","number":14668,"cdate":1789172026188,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Edit","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission14668/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing"],"mdate":1791226736379,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"r3SqIM6iat","version":2},{"content":{"comment":{"value":"> The algorithm might rely on high-quality and multi-view RGBD videos. In the experiments, the background is clean and the objects are relatively simple. There are not so many tests on real-world data where the depth values are noisy and the view angles are sparse.\n\nOur experiments are in a simplified setting, but we find that in this setting VPD is robust to having few views and using depth predicted by a convnet instead of ground truth. Fig. 6b shows the performance of the model trained with varying numbers of conditioning views, and Fig. 6c shows that it performs nearly as well with predicted depth as with ground truth. The Deformables datasets have only four camera views, which is at least feasible to implement in the real world.\n\nApplying VPD to real-world data is a high priority piece of future work for us, though we were unable to do so for this submission due to a lack of pre-existing datasets with multi-view RGB-D data.\n\n> What's the memory and time consumption to train this pipeline?\n\nThis model is memory-intensive, especially for the graph network, as the graphs contain ~30,000 points and ~100,000 edges. During training, with 6 recursive rollout steps and a backward pass, memory consumption is above 30GB, and training the whole model end to end takes 2-3 days. For training, the NeRF rendering speed is not a major bottleneck, since the model only renders a random subset of pixels for each training step, but it could be improved by skipping rendering in regions with no nearby points. Improving the efficiency of the model will be a key direction for future work.\n\n> In Figure 8, while the PSNR stays the same between 1024->2048 and 4096->8192 while takes big jumps between 2048->4096 and 8192→16384.\n\nThanks for pointing this out. We have re-run this evaluation with more samples and updated Figure 8 in the paper.\n\n> This method seems to treat the same points in different views as different points.\n\nIt would be interesting to investigate methods for explicitly merging and de-duplicating points which are very near to one another. So far this happens only implicitly in the model, which learns to cope with the non-uniform point sampling distribution. Note that this affects the set of inputs to the model but not the loss, which is computed on pixels.\n\n[1] Guan, Shanyan, et al. \"Neurofluid: Fluid dynamics grounding with particle-driven neural radiance fields.\" International Conference on Machine Learning. PMLR, 2022.\n\n[2] Xue, Haotian, et al. \"3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.\n\n[3] Pfaff, T., Fortunato, M., Sanchez-Gonzalez, A., & Battaglia, P. W. (2020). Learning mesh-based simulation with graph networks. International Conference on Learning Representations."},"title":{"value":"Reply 2 / 2"}},"tmdate":1700669058527,"tcdate":1700669058527,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3677/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission3677/Authors"],"forum":"4rBEgZCubP","number":3,"license":"CC BY 4.0","cdate":1700669058527,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3677/-/Official_Comment"],"mdate":1700669058527,"domain":"ICLR.cc/2024/Conference","replyto":"9kIUS6MABr","id":"49hYqcOs5T","forumContent":{"TLDR":{"value":"Learned dynamics models that combine 3D representations and ray-based rendering."},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["simulation","dynamics","nerf","particle dynamics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Realistic simulation is critical for applications ranging from robotics to animation. Traditional analytic simulators sometimes struggle to capture sufficiently realistic simulation which can lead to problems including the well known \"sim-to-real\" gap in robotics. Learned simulators have emerged as an alternative for better capturing real-world physical dynamics, but require access to privileged ground truth physics information such as precise object geometry or particle tracks. Here we propose a method for learning simulators directly from observations. Visual Particle Dynamics (VPD) jointly learns a latent particle-based representation of 3D scenes, a neural simulator of the latent particle dynamics, and a renderer that can produce images of the scene from arbitrary views. VPD learns end to end from posed RGB-D videos and does not require access to privileged information. Unlike existing 2D video prediction models, we show that VPD's 3D structure enables scene editing and long-term predictions. These results pave the way for downstream applications ranging from video editing to robotic planning."},"_bibtex":{"value":"@inproceedings{\nwhitney2024learning,\ntitle={Learning 3D Particle-based Simulators from {RGB}-D Videos},\nauthor={William F Whitney and Tatiana Lopez-Guevara and Tobias Pfaff and Yulia Rubanova and Thomas Kipf and Kim Stachenfeld and Kelsey R Allen},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=4rBEgZCubP}\n}"},"title":{"value":"Learning 3D Particle-based Simulators from RGB-D Videos"},"pdf":{"value":"/pdf/c084f0fff026d69efbf43d593934b7e30a668247.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"whitney|learning_3d_particlebased_simulators_from_rgbd_videos"},"authorids":{"value":["~William_F_Whitney1","~Tatiana_Lopez-Guevara1","~Tobias_Pfaff1","~Yulia_Rubanova2","~Thomas_Kipf2","~Kim_Stachenfeld1","~Kelsey_R_Allen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["William F Whitney","Tatiana Lopez-Guevara","Tobias Pfaff","Yulia Rubanova","Thomas Kipf","Kim Stachenfeld","Kelsey R Allen"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Fun3D, a physics-informed text-to-3D generation method aimed at producing physically plausible 3D shapes based on text prompts. Existing text-to-3D models primarily focus on visual and geometric accuracy but lack physical realism, which limits practical applications. Fun3D addresses this gap by integrating physics, specifically solid mechanics, into the generative process. It uses a two-stage framework: an initial 3D shape is generated via 3D diffusion models and then optimized through a differentiable physics layer. This layer utilizes a mix of geometry and physics constraints, leveraging finite element method (FEM) data to improve stability and load-bearing capacity. Experiments demonstrate that Fun3D yields more physically robust shapes compared to baseline models like Diffusion-SDF, making it suitable for engineering and other real-world applications."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Please address the concern about the weakness.\n2. I am curious about the applicability of Fun3D to other types of objects where some of the parts are soft and some of the parts are solid. For example, in animals in Figure 6, the strength of the animal and the stress on them do not necessarily depend on the geometry we see on the outside. They are often more related to the structure of their bones and muscles. So why is Figure 6 shown or discussed in this paper, and why the proposed method can help animal generation to have better physical properties?\n3. If some artist or architect wants to build something that is against the analysis of FEM, how do you balance the strength of the generated object and the design proposed by them?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. By incorporating solid mechanics and FEM-inspired optimization into 3D generation, the paper advances text-to-3D methods beyond visual realism, aiming for physical feasibility, which is valuable for applications requiring structural integrity.\n\n2. The use of a neural network-based differentiable physics layer allows the system to be trained end-to-end, optimizing geometry while maintaining physics constraints.\n\n3. The topic is important to real-world applications if we want to use a generative model to help produce solid objects."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Comparisons: The authors only compare their method to Diffusion-SDF, neglecting recent advancements like SDFusion (CVPR 2023) and LucidDreamer (CVPR 2024), which utilize different geometry extraction and physics-informed components. Including these would provide a fuller assessment of the method’s capabilities.\n\n2. Evaluation Metric Bias: Physical strength is evaluated using FEM, which is also an integral component of Fun3D’s training. This could bias results in favor of the proposed model. Additional evaluation metrics, such as load capacity or material distribution uniformity, could provide a more unbiased assessment of physical properties."}},"nonreaders":[],"tmdate":1732426814451,"tcdate":1730620314582,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9022/Reviewer_abDh"],"signatures":["ICLR.cc/2025/Conference/Submission9022/Reviewer_abDh"],"forum":"6SMeOas0JX","number":2,"license":"CC BY 4.0","cdate":1730620314582,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9022/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732426814451,"domain":"ICLR.cc/2025/Conference","replyto":"6SMeOas0JX","id":"UApLeWfy1l","forumContent":{"TLDR":{"value":"Looks great, functions better"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["3D shape generation","Functional 3D model","Physics perception","Differentiable physics layer","Solid mechanics"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-to-3D shape generation has shown great promise in generating novel 3D content based on given text prompts. However, existing generative methods mostly focus on geometric or visual plausibility while ignoring function for the generated 3D shapes. This greatly hinders the practicality of generated 3D shapes in real-world applications. In this work, we propose Fun3D, a physics driven functional text-to-3D shape generation method. By analyzing the solid mechanics of generated 3D shapes, we reveal that the 3D shapes generated by existing text-to-3D generation methods are impractical for real-world applications as the generated 3D shapes do not conform to the laws of physics. To this end, we leverage 3D diffusion models to provide 3D shape priors and design a data-driven differentiable physics layer to optimize 3D shape priors with solid mechanics. This allows us to optimize geometry efficiently and learn physics information about 3D shapes at the same time. Experimental results demonstrate that our method can consider both geometric plausibility and functional requirement, further bridging 3D virtual modeling and physical worlds."},"_bibtex":{"value":"@misc{\nxu2025looks,\ntitle={Looks Great, Functions Better: Physics Compliance Text-to-3D Shape Generation},\nauthor={Qingshan Xu and Jiao Liu and Melvin Wong and Caishun Chen and Yew-Soon Ong},\nyear={2025},\nurl={https://openreview.net/forum?id=6SMeOas0JX}\n}"},"title":{"value":"Looks Great, Functions Better: Physics Compliance Text-to-3D Shape Generation"},"pdf":{"value":"/pdf/485f4f11c3bf39c2443a11179a743b357c2f1bb0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"xu|looks_great_functions_better_physics_compliance_textto3d_shape_generation"},"authorids":{"value":["~Qingshan_Xu1","~Jiao_Liu3","~Melvin_Wong1","~Caishun_Chen1","~Yew-Soon_Ong1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Qingshan Xu","Jiao Liu","Melvin Wong","Caishun Chen","Yew-Soon Ong"]}},"version":2},{"content":{"summary":{"value":"The present manuscript proposes to leverage finite-element techniques to learn the solution to parametric PDEs using neural networks. Minimizing a classical Galerkin approximation, the method trains a neural network to predict the coefficients of a nodal FEM basis for a given PDE parameter, where the mesh and nodal basis generation is performed using FEniCS.\nMoreover, building upon rich literature on FEM methods, corresponding theoretical guarantees are developed. Finally, numerical experiments on complex geometries show that the proposed approach can outperform supervised and unsupervised approaches based on DeepONets."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The development of hybrid methods, combining existing FEM solvers with deep learning, is a promising research direction.\n- On complex geometries, the method performs significantly better than existing physics-informed neural operators."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) The stated contributions seem to \"oversell\" the method:\n    - As discussed later, all physics-informed neural operators do not require any training data.\n    - Most neural operators can deal with any form of PDE data (forcing, coefficients, boundary conditions, initial conditions). \n    - In contrast, the proposed method appears not to work on *any* form of PDE data but only on data given in a parametrized form, i.e., by a finite-dimensional parameter vector. In particular, one cannot use (discretizations of) arbitrary input functions, which is possible for, e.g., FNO. \n\n2) Further numerical results are needed:\n    - The hyperparameters of the baselines seem not to have been optimized for the considered problems.\n    - It would be good to present comparisons on problems that have been considered in the PIDeepONet or PINO papers since these methods have not been optimized for complex geometries. In this context, other approaches have been suggested; see, e.g., https://arxiv.org/pdf/2207.05209.pdf.\n    - For the current problems, it would be more suitable to see comparisons against graph-based neural PDE solvers.\n    - It is essential to compare runtimes of the considered methods.\n\n3) Motivation of the work and comparisons with classical FEM methods:\n    - It seems that the proposed approach is merely learning a surrogate model for solving the linear/linearized system of equations arising in FEM. It still requires carefully choosing basis functions and meshes and assembling stiffness matrices (i.e., in the specific case of the present work, it is heavily relying on FEniCS). While current operator learning methods can not yet achieve the same accuracies as specialized numerical solvers, they are more universal and do not need to be adapted to specific PDEs.\n   - Considering training cost, what is the advantage of the proposed approach to just solving the linear system in (6) with a suitable solver? There is a single comparison of the runtime of FEONet and FEM in the appendix, but this seems to be a very critical point. Especially given that the achieved accuracy of FEONet seems to be orders of magnitude worse than the FEM solver. In the case of varying forcing functions, it seems that one could reuse the inverse of the stiffness matrix and just compute a single matrix-vector product to arrive at the solution (which could also be batched).\n    - There should be more information in the main text on how the bilinear form $B$ is computed for varying PDE data."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"What is the advantage of minimizing (9) instead of solving the FEM system and directly regressing the optimal coefficients?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636051413,"tcdate":1698820180951,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1251/Reviewer_jxCk"],"signatures":["ICLR.cc/2024/Conference/Submission1251/Reviewer_jxCk"],"forum":"wwJJUamHVp","number":2,"license":"CC BY 4.0","cdate":1698820180951,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1251/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636051413,"domain":"ICLR.cc/2024/Conference","replyto":"wwJJUamHVp","id":"kegBzHIW5m","forumContent":{"TLDR":{"value":"We proposed a novel approach for solving parametric PDEs based on the finite element methods, which is called Finite Element Operator Network (FEONet)."},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Scientific machine learning","finite element methods","physics-informed operator learning","parametric partial differential equations"]},"supplementary_material":{"value":"/attachment/200f5cdd1371a5aa07c99437bcd402868b2e224c.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Partial differential equations (PDEs) underlie our understanding and prediction of natural phenomena across numerous fields, including physics, engineering, and finance. However, solving parametric PDEs is a complex task that necessitates efficient numerical methods. In this paper, we propose a novel approach for solving parametric PDEs using a Finite Element Operator Network (FEONet). Our proposed method leverages the power of deep learning in conjunction with traditional numerical methods, specifically the finite element method, to solve parametric PDEs in the absence of any paired input-output training data. We demonstrate the effectiveness of our approach on several benchmark problems and show that it outperforms existing state-of-the-art methods in terms of accuracy, generalization, and computational flexibility. Our FEONet framework shows potential for application in various fields where PDEs play a crucial role in modeling complex domains with diverse boundary conditions and singular behavior. Furthermore, we provide theoretical convergence analysis to support our approach, utilizing finite element approximation in numerical analysis."},"_bibtex":{"value":"@misc{\nlee2024finite,\ntitle={Finite Element Operator Learning for Solving Parametric {PDE}s without Labeled Data},\nauthor={Jae Yong Lee and Seungchan Ko and Youngjoon Hong},\nyear={2024},\nurl={https://openreview.net/forum?id=wwJJUamHVp}\n}"},"title":{"value":"Finite Element Operator Learning for Solving Parametric PDEs without Labeled Data"},"pdf":{"value":"/pdf/973ff82866889d54203f7bea39625f261cb55371.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"lee|finite_element_operator_learning_for_solving_parametric_pdes_without_labeled_data"},"authorids":{"value":["~Jae_Yong_Lee2","~Seungchan_Ko1","~Youngjoon_Hong2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jae Yong Lee","Seungchan Ko","Youngjoon Hong"]}},"version":2},{"content":{"venue":{"value":"CoRR 2021"},"pdf":{"value":"https://arxiv.org/pdf/2105.07426v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"gaikwad|curiositydriven_intuitive_physics_learning"},"html":{"value":"https://arxiv.org/abs/2105.07426"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2105-07426,\n  publtype={informal},\n  author={Tejas Gaikwad and Romi Banerjee},\n  title={Curiosity-driven Intuitive Physics Learning},\n  year={2021},\n  cdate={1609459200000},\n  journal={CoRR},\n  volume={abs/2105.07426},\n  url={https://arxiv.org/abs/2105.07426}\n}\n"},"abstract":{"value":"Biological infants are naturally curious and try to comprehend their physical surroundings by interacting, in myriad multisensory ways, with different objects - primarily macroscopic solid objects - around them. Through their various interactions, they build hypotheses and predictions, and eventually learn, infer and understand the nature of the physical characteristics and behavior of these objects. Inspired thus, we propose a model for curiosity-driven learning and inference for real-world AI agents. This model is based on the arousal of curiosity, deriving from observations along discontinuities in the fundamental macroscopic solid-body physics parameters, i.e., shape constancy, spatial-temporal continuity, and object permanence. We use the term body-budget to represent the perceived fundamental properties of solid objects. The model aims to support the emulation of learning from scratch followed by substantiation through experience, irrespective of domain, in real-world AI agents."},"title":{"value":"Curiosity-driven Intuitive Physics Learning"},"authors":{"value":[{"fullname":"Tejas Gaikwad","username":""},{"fullname":"Romi Banerjee","username":"~Romi_Banerjee1"}]}},"tmdate":1786938676675,"pdate":1640908800000,"externalIds":["dblp:journals/corr/abs-2105-07426"],"tcdate":1786938668108,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Romi_Banerjee1"],"forum":"xlYZDD8wQo","license":"CC BY-SA 4.0","number":133378,"cdate":1609459200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1786938676675,"domain":"OpenReview.net/Public_Article","id":"xlYZDD8wQo","version":2},{"content":{"summary":{"value":"The paper addresses the task of generating physically plausible environments. To tackle challenges in both spatial arrangement and physics, the authors propose the PhyScensis framework, which leverages a large language model (LMM) to generate predicates and employs a physics engine as the physics solver. Their framework also incorporates feedback from the physics engine back to the LMM for further refinement, resulting in realistic layouts and physically stable scenes. Experimental results demonstrate superior performance compared to previous methods."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"* How are objects selected—only at the category level, or is there a more detailed retrieval?\n* Are there predefined rules for selection, or is it random? For example, in the “table for 4” case, why are all plates the same? Is this constraint imposed by the LLM?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"* The overall system, which integrates LLM-based predicates, a physics-based solver, a geometry-based spatial solver, and feedback to the LLM, is well-designed. This results in layouts that are both reasonable and physically stable.\n* Physics-plausible scene generation is an interesting and important direction, particularly for large-scale scene generation.\n* The experiments are thorough, including ablations and additional evaluations on downstream robotics tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* It is unclear what text prompts are used in the test set for all methods. How many prompts are there, and how diverse are they?\n* There is no discussion of failure cases, particularly regarding physics. What are the limitations of the current predefined predicates?\n* Regarding the LLM, it is unclear how it determines object sizes and how it selects objects from the candidate object set."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921012969,"tcdate":1761947360546,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9407/Reviewer_MLGL"],"signatures":["ICLR.cc/2026/Conference/Submission9407/Reviewer_MLGL"],"forum":"aCVfhY4Qen","number":3,"license":"CC BY 4.0","cdate":1761947360546,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9407/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921012969,"domain":"ICLR.cc/2026/Conference","replyto":"aCVfhY4Qen","id":"bZY3RF3O4a","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Physical Scene Generation"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Automatically generating interactive 3D environments is crucial for scaling up robotic data collection in simulation. While prior work has primarily focused on 3D asset placement, it often overlooks the physical relationships between objects (e.g., contact, support, balance, and containment), which are essential for creating complex and realistic manipulation scenarios such as tabletop arrangements, shelf organization, or box packing. Compared to classical 3D layout generation, producing complex physical scenes introduces additional challenges: (a) higher object density and complexity (e.g., a small shelf may hold dozens of books), (b) richer supporting relationships and compact spatial layouts, and (c) the need to accurately model both spatial placement and physical properties.\nTo address these challenges, we propose PhyScensis, an LLM agent-based framework powered by a physics engine, to produce physically plausible scene configurations with high complexity.\nSpecifically, our framework consists of three main components: an LLM agent iteratively proposes assets with spatial and physical predicates; a solver, equipped with a physics engine, realizes these predicates into a 3D scene; and feedback from the solver informs the agent to refine and enrich the configuration. \nMoreover, our framework preserves strong controllability over fine-grained textual descriptions and numerical parameters (e.g., relative positions, scene stability), enabled through probabilistic programming for stability and a complementary heuristic that jointly regulates stability and spatial relations.\nExperimental results show that our method outperforms prior approaches in scene complexity, visual quality, and physical accuracy, offering a unified pipeline for generating complex physical scene layouts for robotic manipulation."},"_bibtex":{"value":"@inproceedings{\nwang2026physcensis,\ntitle={PhyScensis: Physics-Augmented {LLM} Agents for Complex Physical Scene Arrangement},\nauthor={Yian Wang and Han Yang and Minghao Guo and Xiaowen Qiu and Tsun-Hsuan Wang and Wojciech Matusik and Joshua B. Tenenbaum and Chuang Gan},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=aCVfhY4Qen}\n}"},"title":{"value":"PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement"},"pdf":{"value":"/pdf/a80ca2e2f82bcb6b1458663143e7981305c1250c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|physcensis_physicsaugmented_llm_agents_for_complex_physical_scene_arrangement"},"authorids":{"value":["~Yian_Wang1","~Han_Yang4","~Minghao_Guo1","~Xiaowen_Qiu1","~Tsun-Hsuan_Wang2","~Wojciech_Matusik2","~Joshua_B._Tenenbaum1","~Chuang_Gan1"]},"authors":{"value":["Yian Wang","Han Yang","Minghao Guo","Xiaowen Qiu","Tsun-Hsuan Wang","Wojciech Matusik","Joshua B. Tenenbaum","Chuang Gan"]}},"version":2},{"content":{"summary":{"value":"This works proposes ElastoGen, a hybrid model for 4D elastodynamics. It incorporates physics priors and can be embedded in a larger, differentiable deep learning model for end-to-end 4D generation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What are the limitations of this method compared to PhysDreamer?\n\n2. Does the method need full access to a mesh of the objects, or can this be learned from data?\n\n3. There are numerous typo's throughout the manuscript, even in the title of a section ('synamics') and in the acronyms ('NerualMTL')."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. This work successfully incorporates relevant physics priors in the dynamical modeling of soft materials. Incorporating stronger priors into generative models is a relevant idea."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The related work section on Generative models is extremely broad and this work is not well positioned. Ultimately, this work is compared to Zhang et al. 2024, but this is not even mentioned or discussed in the related work.\nThe concept of *4D* is not adequately explained. The long lists of references do not help to situate this work or understand the relevant context.\n\n2. In terms of differentiable physics modelling, I think the authors should be aware of the related research line commonly called *differentiable physics* and discuss their method compared to those approaches:\n\nDegrave, Jonas, et al. \"A differentiable physics engine for deep learning in robotics.\" Frontiers in neurorobotics 13 (2019): 6.\n\nde Avila Belbute-Peres, F., Smith, K., Allen, K., Tenenbaum, J., & Kolter, J. Z. (2018). End-to-end differentiable physics for learning and control. Advances in neural information processing systems, 31.\n\nHu, Y., Anderson, L., Li, T. M., Sun, Q., Carr, N., Ragan-Kelley, J., & Durand, F. (2020, January). DiffTaichi: Differentiable Programming for Physical Simulation. In International Conference on Learning Representations.\n\n3. The experimental validation is very limited, both quantitatively and qualitatively. It is very hard to estimate the value of this work on the experimental aspect, which is very important for this kind of work."}},"nonreaders":[],"tmdate":1731427749384,"tcdate":1730405650504,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3147/Reviewer_ksPC"],"signatures":["ICLR.cc/2025/Conference/Submission3147/Reviewer_ksPC"],"forum":"j50c2tkQUu","number":2,"license":"CC BY 4.0","cdate":1730405650504,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3147/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427749384,"domain":"ICLR.cc/2025/Conference","replyto":"j50c2tkQUu","id":"7OsOUn8afp","forumContent":{"TLDR":{"value":"We present ElastoGen, a knowledge-driven model that generates physically accurate and coherent 4D elastodynamics."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["generative model","machine learning","neural network architectures"]},"supplementary_material":{"value":"/attachment/d663f217ebed123028262859ca69a298fd4e6448.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present ElastoGen, a knowledge-driven model that generates physically accurate and coherent 4D elastodynamics. Instead of relying on petabyte-scale data-driven learning, ElastoGen leverages the principles of physics-in-the-loop and learns from established physical knowledge, such as partial differential equations and their numerical solutions. The core idea of ElastoGen is converting the global differential operator, corresponding to the nonlinear elastodynamic equations, into iterative local convolution-like operations, which naturally fit modern neural networks. Each network module is specifically designed to support this goal rather than functioning as a black box. As a result, ElastoGen is exceptionally lightweight in terms of both training requirements and network scale. Additionally, due to its alignment with physical procedures, ElastoGen efficiently generates accurate dynamics for a wide range of hyperelastic materials and can be easily integrated with upstream and downstream deep modules to enable end-to-end 4D generation."},"_bibtex":{"value":"@misc{\nfeng2024elastogen,\ntitle={ElastoGen: 4D Generaetive Elastodynamics},\nauthor={Yutao Feng and Yintong Shang and Xiang Feng and Lei Lan and Shandian Zhe and Tianjia Shao and Hongzhi Wu and Kun Zhou and Hao Su and Chenfanfu Jiang and Yin Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=j50c2tkQUu}\n}"},"title":{"value":"ElastoGen: 4D Generaetive Elastodynamics"},"pdf":{"value":"/pdf/9fa9241ebde9ae6fc4ab85819eb32a48a6087153.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"feng|elastogen_4d_generaetive_elastodynamics"},"authorids":{"value":["~Yutao_Feng1","~Yintong_Shang1","~Xiang_Feng1","~Lei_Lan1","~Shandian_Zhe1","~Tianjia_Shao1","~Hongzhi_Wu1","~Kun_Zhou1","~Hao_Su1","~Chenfanfu_Jiang3","~Yin_Yang4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yutao Feng","Yintong Shang","Xiang Feng","Lei Lan","Shandian Zhe","Tianjia Shao","Hongzhi Wu","Kun Zhou","Hao Su","Chenfanfu Jiang","Yin Yang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces PhysFlow, a generative model for protein backbone structures. The authors' core motivation is that standard diffusion and flow models use a \"noising\" process that is not physically realistic, ignoring principles like topological integrity and steric clashes. To address this, they propose a physics-inspired, non-linear \"unfolding\" process grounded in Hamiltonian dynamics. This forward process is designed to unfold a protein into a secondary structure (like a $\\beta$-sheet) while explicitly preserving bonds and using a Coulomb-like repulsion term to prevent residue collisions.This physics-driven process is integrated into the flow-matching paradigm on $SE(3)$ to learn the distribution of protein backbones. The model also incorporates sequence information, enabling it to perform both unconditional backbone generation and sequence-conditioned folding. The authors claim that PhysFlow achieves state-of-the-art performance in unconditional generation, producing more designable and novel structures, and accurately folds monomer sequences."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"See above."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The paper's motivation to incorporate more physics into the generative forward process is a novel and interesting research direction. The idea of replacing a generic noise process with a more structured unfolding trajectory that respects physical constraints like collision avoidance is intuitive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper proposes a complex, non-linear, second-order Hamiltonian system for the forward unfolding process but completely fails to discuss or justify the reverse generative process. It is a major theoretical gap to assume that a standard, first-order flow-matching objective is sufficient to learn the true time-reversal of such a complex non-linear dynamic system, which is required to satisfy the Continuity Equation or the Fokker-Planck equation. The original reverse processes for diffusion models or flow matching generally cannot be directly applied here without theoretical modification.\n2. The model is explicitly designed to unfold to a target distribution $p_0$ defined as a linear $\\beta$-sheet, yet the empirically generated samples are biased toward the $\\alpha$-helix structure (82.2%). This contradiction fundamentally weakens the paper's core claim, suggesting that the physics-guided forward process is either not being reversed correctly or has little impact on the final generated distribution.\n3. The paper's motivation, that a \"physically plausible\" forward trajectory is necessary, is questionable. The power of modern generative models originates from the strong generalization ability of the denoising network, which requires the forward process to be stochastic and highly noisy to enable learning recovery from arbitrary, highly corrupted states. Restricting the model to a single, well-defined forward path may hurt its ability to generalize to novel states not encountered on this narrow trajectory.\n4. The proposed algorithm requires an expensive pre-processing step to simulate and store the forward trajectory for every protein. This prevents on-the-fly training and introduces a critical scalability bottleneck, especially since the forward simulation takes over one second per protein (as shown in Figure 4), severely limiting its application to larger datasets like AlphaFold DB.\n5. The evaluation is undermined by non-standard and weak comparisons. (1) The authors compare against some models trained on shorter sequences (e.g., up to 256 aa) while generating samples up to 300 aa, leading to an unfair comparison. (2) They exclude stronger baselines (like Proteina and Genie2) with the excuse of using different datasets, even though this paper uses a self-curated dataset as well. (3) The comparison for the folding task (Table 3) is made against co-design models (MultiFlow, FoldFlow-2) rather than specialized, state-of-the-art folding models, making the claim of strong folding performance unconvincing.\n6. The reported \"Diversity Cluster Ratio\" is affected by the designability value. For a clear assessment of structural variety, the authors should report the absolute number of designable clusters instead, which is a more robust metric for diversity.\n7. The paper lacks definitions for key terms, specifically \"linear diffusion\" in the first paragraph of Section 3 and \"unfolding flow\" (Line 187).\n8. The choice to set the target state $p_0$ from a beta-sheet distribution is arbitrary and unexplained. Ideally, the model should be able to generate any secondary structure, so the selection of a single predefined distribution for the \"unfolded\" state needs strong justification.\n9. The paper is lacking crucial hyperparameter configurations for the forward simulation, including $\\gamma$ (drag coefficient), $\\sigma_\\beta^2$, $\\mu_\\beta$ (for the target distribution), and $\\sigma_v^2$, hindering the reproducibility of the work.\n10.  In Table 1, the categorization of existing models into DDPM, CFM, OT, or DSM is an oversimplification, as these algorithms are well-known to be deeply related and can be unified under the Stochastic Interpolant framework.\n11. Equation (3) is wrong, as in flow matching, there is only one ODE. The reverse process simply simulates the same ODE in the reverse time direction."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916325905,"tcdate":1761552410457,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2670/Reviewer_uDYi"],"signatures":["ICLR.cc/2026/Conference/Submission2670/Reviewer_uDYi"],"forum":"qEzgXEBLIH","number":2,"license":"CC BY 4.0","cdate":1761552410457,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2670/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916325905,"domain":"ICLR.cc/2026/Conference","replyto":"qEzgXEBLIH","id":"9V3Lm1Xb9Y","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"We propose a novel physics-informed generative model for protein backbone structure generation using flow matching"},"keywords":{"value":["Protein Structure Generative Models","Structure prediction","Physics-informed generative model","Flow Matching"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Protein structure prediction and folding are fundamental to understanding biology, with recent deep learning advances reshaping the field. Diffusion-based generative models have revolutionized protein design, enabling the creation of novel proteins. However, these methods often neglect the intrinsic physical realism of proteins, driven by noising dynamics that lack grounding in physical principles. To address this, we first introduce a physically motivated non-linear noising process, grounded in classical physics, that unfolds proteins into secondary structures (e.g., $\\alpha$-helices, linear $\\beta$-sheets) while preserving structural integrity—maintaining bonds and preventing collisions. We then integrate this process with the flow-matching paradigm on $\\mathrm{SE(3)}$ to model the invariant distribution of protein backbones with high fidelity, incorporating sequence information to enable sequence-conditioned folding and expand the generative capabilities of our model. Experimental results demonstrate state-of-the-art performance in unconditional protein generation, producing more designable and novel protein structures while accurately folding monomer sequences into precise protein conformations."},"_bibtex":{"value":"@misc{\nverma2026let,\ntitle={Let Physics Guide Your Protein Flows: Topology-aware Unfolding and Generation},\nauthor={Yogesh Verma and Markus Heinonen and Vikas K Garg},\nyear={2026},\nurl={https://openreview.net/forum?id=qEzgXEBLIH}\n}"},"title":{"value":"Let Physics Guide Your Protein Flows: Topology-aware Unfolding and Generation"},"pdf":{"value":"/pdf/762b9d6d31e2214ed703dfdeacc3ef5e8c753988.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"verma|let_physics_guide_your_protein_flows_topologyaware_unfolding_and_generation"},"authorids":{"value":["~Yogesh_Verma1","~Markus_Heinonen1","~Vikas_K_Garg1"]},"authors":{"value":["Yogesh Verma","Markus Heinonen","Vikas K Garg"]}},"version":2},{"content":{"summary":{"value":"The paper proposes PhyGenBench and PhyGenEval. PhyGenBench is a benchmark with about 160 text prompts used to evaluate models' video generation ability on physics-related text prompts. PhyGenEval is an evaluation framework of PhyGenBench, used to automatically assess the video quality of physics laws, via GPT-prompted questions and VLM perception."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weakness above."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The topic is novel and interesting. Evaluating physics in AI-generated videos is really important. This paper is the first one on this topic as far as I know.\n\n- Experiments show that PhyGenEval is closer to human value."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The PhyGenBench is a dataset with 160 text prompts. As a comparison, for the works mentioned in this paper, VideoPhy has 688 prompts with 36.5k human annotations, and DEVIL has more than 800 prompts. Only 160 text prompts may not represent the full complexity of physics law.\n\n- In PhyGenEval, the overall score is set on a four-point scale, but even the top-performing video generation model scores only 0.5 on average. That means the model gets a 0 score in more than half of the test cases. This suggests that the evaluation metric might be overly strict, potentially limiting its effectiveness in distinguishing between models. Such stringent scoring could reduce the benchmark’s ability to accurately reflect model performance differences.\n\n- Since the topic is related to evaluating the physics in generative models, I think it is better to add some discussion on physical reasoning benchmarks in related works, which has been a heated debate topic, such as SuperCLEVR-Physics[1], ContPhy[2], Physion[3] and so on.\n\n[1] Wang, X., Ma, W., Wang, A., Chen, S., Kortylewski, A., & Yuille, A. (2024). Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering. ArXiv. https://arxiv.org/abs/2406.00622\n\n[2] Zheng, Z., Yan, X., Chen, Z., Wang, J., Lim, Q. Z., Tenenbaum, J. B., & Gan, C. (2024). ContPhy: Continuum Physical Concept Learning and Reasoning from Videos. ArXiv. https://arxiv.org/abs/2402.06119\n\n[3] Bear, D. M., Wang, E., Mrowca, D., Binder, F. J., Tung, H., Pramod, R. T., Holdaway, C., Tao, S., Smith, K., Sun, F., Kanwisher, N., Tenenbaum, J. B., Yamins, D. L., & Fan, J. E. (2021). Physion: Evaluating Physical Prediction from Vision in Humans and Machines. ArXiv. https://arxiv.org/abs/2106.08261"}},"nonreaders":[],"tmdate":1731427584100,"tcdate":1730115336100,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2363/Reviewer_eSnR"],"signatures":["ICLR.cc/2025/Conference/Submission2363/Reviewer_eSnR"],"forum":"6rMHcLWxl4","number":1,"license":"CC BY 4.0","cdate":1730115336100,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2363/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427584100,"domain":"ICLR.cc/2025/Conference","replyto":"6rMHcLWxl4","id":"QL3IDfJsLo","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["World Simulator","Physical Commonsense","Video Generation","Evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGenBench, a comprehensive \\textbf{Phy}sics \\textbf{Gen}eration \\textbf{Ben}chmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench, we propose a novel evaluation framework called PhyGenEval. This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through PhyGenBench and PhyGenEval, we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by PhyGenBench (e.g., dynamic scenarios). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We will release the data and codes at https://github.com/PhyGenBench/PhyGenBench"},"_bibtex":{"value":"@misc{\nmeng2025towards,\ntitle={Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation},\nauthor={Fanqing Meng and Jiaqi Liao and Xinyu Tan and Wenqi Shao and Quanfeng Lu and Kaipeng Zhang and Yu Cheng and Dianqi Li and Yu Qiao and Ping Luo},\nyear={2025},\nurl={https://openreview.net/forum?id=6rMHcLWxl4}\n}"},"title":{"value":"Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation"},"pdf":{"value":"/pdf/1814f0c3473ab9a04aca4edcd8aab3e678055bdd.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"meng|towards_world_simulator_crafting_physical_commonsensebased_benchmark_for_video_generation"},"authorids":{"value":["~Fanqing_Meng1","~Jiaqi_Liao2","~Xinyu_Tan1","~Wenqi_Shao2","~Quanfeng_Lu1","~Kaipeng_Zhang1","~Yu_Cheng1","~Dianqi_Li1","~Yu_Qiao1","~Ping_Luo2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fanqing Meng","Jiaqi Liao","Xinyu Tan","Wenqi Shao","Quanfeng Lu","Kaipeng Zhang","Yu Cheng","Dianqi Li","Yu Qiao","Ping Luo"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Velocity-Regularized Adam (VRAdam), a new optimizer for deep neural network training that introduces a physics-inspired regularization mechanism. Based on the classical Adam optimizer, VRAdam incorporates a quartic kinetic energy term to dynamically regulate the effective learning rate based on the velocity (momentum) of parameter updates. The resulting learning rate shrinks automatically in high-velocity regimes, reducing oscillations and improving convergence stability. The paper provides a theoretical framework, proving uniform exponential stability via Lyapunov analysis for stochastic non-convex objectives. Extensive experiments on CIFAR-10, WikiText-2, GridWorld, and GPT-2 fine-tuning show that VRAdam achieves faster convergence, smoother training curves, and better generalization compared to AdamW and other optimizers."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. See weakness.\n\n2. The author choose the quartic kinetic term NRQCD system as T(v), is there any other choice? Can you report the ablation study since it deserves to be the key insights of the improvement against AdamW.\n\n3. How sensitive is VRAdam to the choice of the velocity penalizer β₃? Does it generalize well across tasks without tuning?\n\n4. Does the quartic kinetic term introduce any bias that could affect convergence to flatter minima or generalization in practice? Can it be integrated with other techiques to improve generalization like in ‘Improving Generalization of Deep Neural Networks by Optimum Shifting , AAAI25’?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The introduction of a quartic kinetic energy term as a stabilizing mechanism is a fresh and interesting physics-based perspective on optimizer design.\n\n2. The analogy between optimization trajectories and particle dynamics adds an intuitive understanding of how the method moves away from instability near the edge of stability.\n\n3. The paper rigorously proves global uniform exponential stability and convergence under mild conditions, supported by clear mathematical derivations.\n\n4. VRAdam consistently outperforms AdamW, RAdam, RMSProp, and SGD across diverse tasks (CNNs, Transformers, GFlowNets, and LLMs)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The presentation is not friendly to readers not familiar with optimization techniques using Langevin dynamics.\n\n2. Lots of hyperparameters are used. Although $\\beta_3$ is claimed to be robust, the practical sensitivity of VRAdam to its hyperparameters $like (\\alpha_0,\\alpha_1,\\beta_3)$ is not fully explored. Are they same with Adam or will be influenced by $\\beta_3$ ?\n\n3. While AdamW(2017) is a strong baseline, the study omits comparisons with newer optimizers (e.g., LION, AdaHessian), which are relevant for modern deep learning tasks.\n\n4. The heavy use of physical analogies (NRQCD, Lagrangians) brings difficulties to understand the motivation and improvement of the algorithm for readers unfamiliar with physics.\n\n5. The experiments only report the validation and test loss, the improvement against AdamW seems modest, whether VRAdam improves the task’s accuracy is not reported. Besides, no clear ablation on the contribution of the quartic term versus standard momentum damping—this would help isolate the true effect of velocity regularization."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941752902,"tcdate":1761021258287,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21408/Reviewer_uYkK"],"signatures":["ICLR.cc/2026/Conference/Submission21408/Reviewer_uYkK"],"forum":"6BhduwrCp3","number":1,"license":"CC BY 4.0","cdate":1761021258287,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21408/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941752902,"domain":"ICLR.cc/2026/Conference","replyto":"6BhduwrCp3","id":"YtfPNtleeV","forumContent":{"TLDR":{"value":"We introduce Velocity-Regularized Adam (**VRAdam**), a velocity penalizing optimizer for deep neural networks."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Optimization in deep learning","physics-inspired","edge of stability"]},"supplementary_material":{"value":"/attachment/468e70955b1c75bdc4a0a86debc3c34f97ba565d.zip"},"primary_area":{"value":"optimization"},"abstract":{"value":"We introduce Velocity-Regularized Adam (VRAdam), a physics-inspired optimizer for training deep neural networks that draws on ideas from quartic terms for kinetic energy with its stabilizing effects on various system dynamics. \nPrevious algorithms, including the ubiquitous Adam, operate at the so-called adaptive edge of stability regime during training, leading to rapid oscillations and slowed convergence of loss.\nHowever, VRAdam adds a higher order penalty on the learning rate based on the velocity such that the algorithm automatically slows down whenever weight updates become large. In practice, we observe that the effective dynamic learning rate shrinks in high-velocity regimes, and damping oscillations. By combining this velocity‑based regularizer for global damping with Adam’s per‑parameter scaling, we create a powerful hybrid optimizer. For this optimizer, we provide rigorous theoretical analysis of operation at the edge of stability from a physical and control perspective for the momentum. Furthermore, we derive convergence bounds with the rate $\\mathcal{O}(\\ln(N)/\\sqrt{N})$ for a stochastic non‑convex objective under mild assumptions. We demonstrate that VRAdam exceeds the performance against standard optimizers including AdamW. We benchmark various tasks such as image classification, language modeling, and generative modeling using diverse architectures and training methodologies including Convolutional Neural Networks (CNNs), Transformers, and GFlowNets."},"_bibtex":{"value":"@inproceedings{\nvaidhyanathan2026a,\ntitle={A Physics-Inspired Optimizer: Velocity Regularized Adam},\nauthor={Pranav Vaidhyanathan and Lucas Schorling and Natalia Ares and Michael A Osborne},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6BhduwrCp3}\n}"},"title":{"value":"A Physics-Inspired Optimizer: Velocity Regularized Adam"},"pdf":{"value":"/pdf/6faf88b351b5b4e9db79349ea4e560efd40f82d0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"vaidhyanathan|a_physicsinspired_optimizer_velocity_regularized_adam"},"authorids":{"value":["~Pranav_Vaidhyanathan1","~Lucas_Schorling1","~Natalia_Ares1","~Michael_A_Osborne1"]},"authors":{"value":["Pranav Vaidhyanathan","Lucas Schorling","Natalia Ares","Michael A Osborne"]}},"version":2},{"content":{"summary":{"value":"This paper introduces FLARE (Flow Alignment Rewiring) for improving GNN-based CFD simulations.\nClassical GNNs operating on CFD meshes suffer from limited information propagation (over-squashing) and physics-misaligned connectivity, since message passing is confined to mesh adjacency that is unrelated to flow direction.\nThe proposed method rewires the mesh dynamically based on local flow alignment:\n* only 2-hop local connections are considered (for locality);\n* edges are directional (for unidirectional transport);\n* new edges are added only when the velocity aligns with the added edge.\n\nThis yields a direction-aware, physics-consistent connectivity pattern that adapts during rollout as the predicted velocity field changes.\nExperiments on three datasets—CylinderFlow, Airfoil, and Tandem-Airfoil-Cruise and across three backbone architectures (MeshGraphNet, BSMS-GNN, Transolver+) show that FLARE outperforms both the prior physics-informed rewiring method PIORF and 2-hop variants."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* What are computational expenses for your model?\n* Can your approach be extended to multiphase flows or multi-physics simulations?\n* How large is the runtime overhead of recomputing the rewired graph at each timestep?\n* Your alignment score is dimensional, so how will you handle the cases of extremely slow flows in some points or extremely strong flows?\n* Could the same idea be applied beyond fluids (e.g., heat diffusion, elasticity)?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"* the proposed approach is based on the first principles of fluid mechanics—locality, directionality, and flow alignment\n* model-agnostic rewiring approach that can plug into any message-passing GNN or hybrid GNN-Transformer architecture without altering its core equations\n* during rollout, rewiring is based on predicted velocities, so it's dynamically adjusted\n* ablations provided: direction reversal, hop distance\n* clear answers to guiding questions: The experiments directly address the three motivating questions—confirming that (1) physics-informed rewiring is beneficial, (2) local directional links suffice, and (3) dynamic flow-based updates matter."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* the proposed rewiring way doesn't depend on fluid velocity magnitude\n* the paper is mainly empirical and proposes mostly an engineering solution\n* dynamic rewiring at each step may add runtime cost, but no timing or complexity study is reported\n* limited physical validation metrics: The study reports RMSE/MSE only, some physics-informed metrics would improve the results"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925557862,"tcdate":1761988647028,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15260/Reviewer_UERp"],"signatures":["ICLR.cc/2026/Conference/Submission15260/Reviewer_UERp"],"forum":"izLvJEBkae","number":4,"license":"CC BY 4.0","cdate":1761988647028,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15260/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925557862,"domain":"ICLR.cc/2026/Conference","replyto":"izLvJEBkae","id":"RJmBtSYowI","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed Graph Neural Networks; CFD; mesh; rewiring; unsteady flow"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"To overcome computation burden of traditional computational fluid dynamics (CFD) simulations, researchers have explored different architectures to develop physics-informed simulation methods. Among them, graph neural networks (GNN) are most suitable for adopting CFD meshes, which are extensively used in engineering and industrial applications. However, classical GNNs propagate information among neighbour nodes, which highly restrict information exchange within the network. To address this issue, graph rewiring methods have been developed for generic graph problems, but not particular for fluid simulation. PIORF, introducing edges connecting distant nodes, is the first graph rewiring method to do so, and previous experiments have demonstrated its effectiveness against state-of-the-art generic rewiring methods. Nevertheless, in this work, we found that simply connecting all 2-hop nodes can provide competitive performance with PIORF. This result raises three questions: 1) Is physics-informed rewiring really useful for improving flow predictions? 2) Should we consider just local connection, instead of connecting distant nodes? 3) Do we need to change the connections based on input flow for rollout simulations? By thoroughly adopting physical fluid principles, we propose a simple yet very efficient method, Flow Alignment Rewiring (FLARE) technique, which connects 2-hop nodes only when the node direction aligns with input flow direction. Hence, FLARE is a physics-informed local rewiring method, different from PIORF and well-aligned with fluid physics. Extensive numerical experiments on flows over a cylinder and single and tandem airfoil under different flow conditions and deep network architectures demonstrate that FLARE outperforms PIORF and various 2-hop rewiring approaches by a significant margin."},"_bibtex":{"value":"@misc{\nli2026graph,\ntitle={Graph Rewiring based on Flow Alignment for Improving Fluid Simulation},\nauthor={Zenong Li and Wei Xian Lim and Wai Lee Chan and Adams Wai-Kin Kong},\nyear={2026},\nurl={https://openreview.net/forum?id=izLvJEBkae}\n}"},"title":{"value":"Graph Rewiring based on Flow Alignment for Improving Fluid Simulation"},"pdf":{"value":"/pdf/fde6cfdb3e702fa96f577b9f6fb0cfb47669442b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|graph_rewiring_based_on_flow_alignment_for_improving_fluid_simulation"},"authorids":{"value":["~Zenong_Li1","~Wei_Xian_Lim2","~Wai_Lee_Chan1","~Adams_Wai-Kin_Kong1"]},"authors":{"value":["Zenong Li","Wei Xian Lim","Wai Lee Chan","Adams Wai-Kin Kong"]}},"version":2},{"content":{"summary":{"value":"### Summary\n\nThis paper presents a new method (EDF) for dataset distillation targeting improved performance in complex scenarios like ImageNet-1K subsets. EDF integrates Grad-CAM activation maps to focus on high-activation, discriminative regions within synthetic data, unlike previous methods that treat all image pixels equally. It also introduces the *Common Pattern Dropout* module to remove low-loss supervision signals, which often contain non-discriminative common patterns. Additionally, the authors contribute a new benchmark, *Complex Dataset Distillation (Comp-DD)*, to help the community evaluate distillation methods in complex scenarios. Experimental results show EDF outperforming state-of-the-art (SOTA) methods in challenging datasets and achieving \"lossless\" performance on some subsets."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"None"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- **Novelty**: EDF’s use of Grad-CAM to target high-activation areas in synthetic images is innovative, particularly for complex datasets where discriminative features are limited to small regions. The *Common Pattern Dropout* module’s filtering of low-loss signals is a creative solution in an attempt to minimize non-discriminative features, improving the overall quality of synthetic data.\n\n- **Benchmark Contribution**: The creation of the Comp-DD benchmark provides a valuable tool for future work, establishing a way to test dataset distillation methods on complex scenarios.\n\n- **Strong Results**: EDF achieves SOTA performance across a variety of complex dataset subsets, with substantial gains in accuracy and the ability to achieve near-lossless performance in certain settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Inaccuracy of Grad-CAM Definition in Section 2.2**: According to the original Grad-CAM paper, $M^c$ should be a gradient-weighted sum of all feature maps in the last convolutional layer. Therefore, the symbol $l$ in Equation (3) should denote the $l$-th feature map instead of the $l$-th convolutional layer.\n\n**Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D.** (2017). *Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization.* In *Proceedings of the IEEE International Conference on Computer Vision* (pp. 618-626). doi:[10.1109/ICCV.2017.74](https://doi.org/10.1109/ICCV.2017.74)\n\n- **Dependency on Hyperparameters**: EDF’s performance depends heavily on various hyperparameters (dropout ratio $\\alpha$, Grad-CAM update frequency $K$, enhancement factor $\\beta$), requiring fine-tuning for a particular dataset, which could limit usability in practical applications without extensive experimentation.\n\n- **Generalizability to Non-Complex Scenarios**: EDF shows significant improvements on complex datasets; however, its benefits on simpler datasets (like CIFAR) are less clear. It’s valuable to show EDF’s generalizability across datasets of various complexity.\n\n- **Limited Complexity of Comp-DD**: ImageNet is a curated dataset, with most object-related images in a portrait style. Consequently, the subsets chosen from ImageNet in Comp-DD are likely also portrait-oriented with larger background regions. This setup may limit the true complexity of the benchmark. To further increase complexity, incorporating zoomed-out, cropped objects with relevant labels from uncurated datasets such as COCO, Objects365, or SA-1B could provide a more challenging and diverse set of scenarios, better suited to the intended purpose.\n\n- **Model Bias in Generating Activation Maps**: It is well-known that neural networks can capture biased or non-generalizable features, as discussed in the Grad-CAM paper. If the activation maps focus on such less-generalizable regions, this could undermine the effectiveness of the *Discriminative Area Enhancement (DAE)* module. However, it is unclear how the proposed method addresses this issue of potential bias in the activation maps. Clarification on how EDF mitigates or adapts to these biased features would strengthen the approach.\n\n- **Clarification on Low-Loss Regions and Discriminative Features**: The paper assumes that low-loss regions correspond to common, less-discriminative features. Although the paper demonstrates activation area shifts with different loss levels during distillation (Fig. 2a), additional qualitative results—such as visual analysis with superposed images and highlighted activation regions—would help support this assumption. It is important to note that a smaller activation region does not necessarily indicate less-discriminative features.\n\nRecommendation:\nOverall, my recommendation for this paper is borderline reject. However, the score could be improved if the concerns outlined above are properly addressed."}},"nonreaders":[],"tmdate":1731427341678,"tcdate":1730675976316,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission929/Reviewer_GwUt"],"signatures":["ICLR.cc/2025/Conference/Submission929/Reviewer_GwUt"],"forum":"SDV7Y6Dhx9","number":3,"license":"CC BY 4.0","cdate":1730675976316,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission929/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427341678,"domain":"ICLR.cc/2025/Conference","replyto":"SDV7Y6Dhx9","id":"RioSeAiGum","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["dataset distillation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Dataset distillation has demonstrated strong performance on simple datasets like CIFAR, MNIST, and TinyImageNet but struggles to achieve similar results in more complex scenarios. \nIn this paper, we propose a novel approach that \\textbf{e}mphasizes the \\textbf{d}iscriminative \\textbf{f}eatures (obtained by Grad-CAM) for dataset distillation, called \\textbf{EDF}.\nOur approach is inspired by a key observation: in simple datasets, high-activation areas typically occupy most of the image, whereas in complex scenarios, the size of these areas is much smaller.\nUnlike previous methods that treat all pixels equally when synthesizing images, EDF uses Grad-CAM activation maps to enhance high-activation areas.\nFrom a supervision perspective, we downplay supervision signals that have lower losses, as they contain common patterns.\nAdditionally, to help the DD community better explore complex scenarios, we build the Complex Dataset Distillation (Comp-DD) benchmark by meticulously selecting sixteen subsets, eight easy and eight hard, from ImageNet-1K.\nNotably, EDF consistently outperforms SOTA results in complex scenarios, such as ImageNet-1K subsets.\nHopefully, more researchers will be inspired and encouraged to enhance the practicality and efficacy of DD. \nOur code and benchmark will be made public."},"_bibtex":{"value":"@misc{\nwang2024emphasizing,\ntitle={Emphasizing Discriminative Features for Dataset Distillation in Complex Scenarios},\nauthor={Kai Wang and Zekai Li and Zhi-Qi Cheng and Samir Khaki and Ahmad Sajedi and Shanmukha Ramakrishna Vedantam and Konstantinos N Plataniotis and Alexander G Hauptmann and Yang You},\nyear={2024},\nurl={https://openreview.net/forum?id=SDV7Y6Dhx9}\n}"},"title":{"value":"Emphasizing Discriminative Features for Dataset Distillation in Complex Scenarios"},"pdf":{"value":"/pdf/51a3d875a6c2106cf4c4fad147811067e4216779.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|emphasizing_discriminative_features_for_dataset_distillation_in_complex_scenarios"},"authorids":{"value":["~Kai_Wang8","~Zekai_Li2","~Zhi-Qi_Cheng1","~Samir_Khaki1","~Ahmad_Sajedi2","~Shanmukha_Ramakrishna_Vedantam1","~Konstantinos_N_Plataniotis1","~Alexander_G_Hauptmann1","~Yang_You1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kai Wang","Zekai Li","Zhi-Qi Cheng","Samir Khaki","Ahmad Sajedi","Shanmukha Ramakrishna Vedantam","Konstantinos N Plataniotis","Alexander G Hauptmann","Yang You"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a method for generating synthetic microservice call graphs using large language models (LLMs). It attempts to address the difficulty of obtaining real-world traces by training LLMs to produce hierarchical and constraint-abiding call graphs recursively, breaking down the complex task into simpler sub-tasks. The authors also employ instruction tuning to align model outputs with specific trace features, in order to enhance the model's ability to generate valid and realistic traces. The paper evaluates the proposed approach by substituting for real-world data in system management tasks and adapt to downstream tasks like predicting trace features and infilling missing data."},"soundness":{"value":1},"confidence":{"value":3},"questions":{"value":"1. The authors should further elaborate the necessity of synthetic trace generation in the microservices domain, especially when ample tracking data is typically available in microservices-based architectures? Additionally, if privacy concerns limit the use of real-world tracking data, why your synthetic trace generation method still need real-world data for training?\n\n2. It is recommended that the authors compare the proposed method with existing trace generation methods or simulation systems, such as TrainTicket, particularly in terms of resource consumption and effectiveness, to demonstrate the advantages or unique features of their approach.\n\n3. The authors are suggested to define the \"accuracy\" metric used in their experiments and explain how \"valid following the initial instructions\" is quantified. Furthermore, it is suggested to include additional validation metrics that capture the complexity of microservice interactions, such as instance response times and trace branching.\n\n4.The authors are advised to provide more experimental details, including whether the test sets used for comparing real and synthetic data are the same, and whether the inaccurate data is include for training TraceVAE. Additionally, it is recommended to include comparative experiments with existing trace generation methods to prove the effectiveness of the proposed method.\n\n5.Please add more experiments about, what are the implications that accuracy declines when the number of edges exceeds 5 or the depth exceeds 2? For example, given the presence of a certain proportion of inaccurate data in the generated traces, the authors should discuss how to manage these data, especially when using them as training data for anomaly detection or root cause analysis, to avoid affecting model performance."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper introduces an approach to generating synthetic microservice call graphs using large language models (LLMs), which is an interesting idea for system workload tracing. It attempts  to handle complex and arbitrary hierarchical structures and implicit constraints within microservice call graphs through a recursive generation method. The paper highlights the potential of synthetic traces to replace real-world data in optimizing and tuning system management tasks, offering significant advantages in terms of privacy and data availability. The method of this paper is of a certain level of innovation and practical value."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. I have serious doubts about the motivation of this paper: Do we really need synthetic traces in the microservices trace domain?\n    1）None of the three \"synthetic trace generation\" methods mentioned by the authors are for generating microservices traces: (Bergsma et al., 2021) is for generating cloud workloads, and (Jiang et al., 2023; Yin et al., 2022) are for producing network traces.\n\n   2）The authors only made a brief statement about the motivation, \"Obtaining real-world traces is often hindered by privacy concerns and their general unavailability,\" without providing any arguments or details. Moreover, the authors' method also requires \"1.36 million microservice call graph samples\" for training and validation, which is contradictory.\n\n   3）In fact, in enterprises with microservices as the basic architecture, trace data is abundant and easily obtainable. There is no issue of insufficient data. If privacy prevents the use of this data, then the authors' method would also be inoperable.\n\n   4）Furthermore, there are many microservices simulation systems, such as TrainTicket (X. Zhou, X. Peng et al., “Fault analysis and debugging of microservice systems: Industrial survey, benchmark system, and empirical study,” TSE’18, 2018). These systems also consume far less resources than the \"4xA100\" used in this paper.\n\n2. The authors' experiments are insufficient and fail to prove the effectiveness.\n    1) As a key metric for measuring the quality of generated traces, \"Accuracy,\" the authors did not elaborate on its definition or how to count the \"valid following the initial instructions.\" The authors also did not introduce the \"initial instructions.\"\n\n    2) We know that the calling relationships between microservices in real environments are very complex, so to verify the accuracy of the generated traces, many aspects need to be validated, not just \"Distribution of Popular Calls\" and \"Heavy-hitter Prediction.\" For example: response delays of instances, branching of traces, etc.\n\n   3) Figure 3 shows that only when the number of edges is less than 5 or the depth is less than 2 can the accuracy of the generated traces be guaranteed to be high. Otherwise, a certain proportion of inaccurate data will appear. I doult that if this inaccurate data is used as training data for anomaly detection or root cause analysis, it will severely affect the model's performance. For example, for anomaly detection, the authors only mentioned that TraceVAE performs similarly with real and synthetic data, suggesting that the authors provide more experimental details (such as whether the test sets are the same) and compare more algorithms.\n\n   4) When comparing effects, it is not very meaningful for the authors to always compare with untrained LLMs. It is recommended to add comparisons with existing trace generation methods. For example, which has a greater overhead and better effect between the TrainTicket simulation system and this paper's generation method."}},"nonreaders":[],"tmdate":1731428417337,"tcdate":1730684963586,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12738/Reviewer_KLHo"],"signatures":["ICLR.cc/2025/Conference/Submission12738/Reviewer_KLHo"],"forum":"f9GURUHZQo","number":3,"license":"CC BY 4.0","cdate":1730684963586,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12738/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428417337,"domain":"ICLR.cc/2025/Conference","replyto":"f9GURUHZQo","id":"ouMIzOWLaM","forumContent":{"TLDR":{"value":"We train a language model to generate synthetic computer system traces, specifically microservice call graphs."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic data","synthetic trace","microservice","large language model","machine learning for systems"]},"supplementary_material":{"value":"/attachment/a3ce86bc05387e93c1d4440cda9c2150d9103a3b.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Computer system workload traces, which record hardware or software events during application execution, are essential for understanding the behavior of complex systems and managing their processing and memory resources. However, obtaining real-world traces can be challenging due to the significant collection overheads in performance and privacy concerns that arise in proprietary systems. As a result, synthetic trace generation is considered a promising alternative to using traces collected in real-world production deployments. This paper proposes to train a large language model (LLM) to generate synthetic workload traces, specifically microservice call graphs. To capture complex and arbitrary hierarchical structures and implicit constraints in such traces, we fine-tune LLMs to generate each layer recursively, making call graph generation a sequence of easier steps. To further enforce learning constraints in traces and generate uncommon situations, we apply additional instruction tuning steps to align our model with the desired trace features. Our evaluation results show that our model can generate diverse realistic traces under various conditions and outperform existing methods in accuracy and validity. We show that our synthetically generated traces can effectively substitute real-world data in optimizing or tuning systems management tasks. We also show that our model can be adapted to perform key downstream trace-related tasks, specifically, predicting key trace features and infilling missing data given partial traces."},"_bibtex":{"value":"@misc{\nkim2025large,\ntitle={Large Language Models as Realistic Microservice Trace Generators},\nauthor={Donghyun Kim and Sriram Ravula and Taemin Ha and Alex Dimakis and Daehyeok Kim and Aditya Akella},\nyear={2025},\nurl={https://openreview.net/forum?id=f9GURUHZQo}\n}"},"title":{"value":"Large Language Models as Realistic Microservice Trace Generators"},"pdf":{"value":"/pdf/4b031907890b6c901f041bd1fb2a704a57090bad.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"kim|large_language_models_as_realistic_microservice_trace_generators"},"authorids":{"value":["~Donghyun_Kim12","~Sriram_Ravula1","~Taemin_Ha1","~Alex_Dimakis1","~Daehyeok_Kim1","~Aditya_Akella1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Donghyun Kim","Sriram Ravula","Taemin Ha","Alex Dimakis","Daehyeok Kim","Aditya Akella"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PhyWorldBench, a comprehensive benchmark for evaluating physical realism in text-to-video generation models. The benchmark contains 1,050 prompts spanning 10 main physics categories (fundamental, composite, and anti-physics), with each category divided into 5 subcategories and 7 scenarios. The authors evaluate 12 state-of-the-art models (5 proprietary, 7 open-source) and propose CAP (Context-Aware Prompt), a method using MLLMs for automated evaluation. Results show that even the best models (Pika 2.0 achieving 26.2% success rate) struggle significantly with physical realism, particularly with complex interactions and anti-physics scenarios."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- The paper mentions models sometimes \"rationalize\" physics violations by generating static scenes. Could you quantify how often this happens across models? This seems like an important failure mode worth analyzing systematically.\n- Have you considered releasing a \"difficulty rating\" for each prompt based on model performance? This could help researchers identify particularly challenging scenarios to focus on."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Important and timely problem. The paper tackles a crucial gap in bridging video generation and physical reasoning. The motivation is clear and aligns well with current research trends.\n- Comprehensive experimental design. The evaluation is thorough - 12 models tested, 12,600 videos generated, large-scale human annotation via MTurk. They evaluate across a diverse set of tasks, object falling, rolling, fluid dynamics, etc. \n- Practical automated evaluator. The CAP method achieves 80.3% ROC-AUC for semantic adherence and 75.1% for physical commonsense (Table 2), which is pretty solid for zero-shot evaluation. \n- Clear organization and presentation. The hierarchical organization is logical and makes the benchmark easy to understand and extend. The tables and plots are easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- No end-to-end pipeline for new models. The paper doesn’t provide a unified automatic evaluation framework for new video generation models. While CAP is proposed as an automatic evaluator, there's no clear standalone pipeline or released code/API for researchers to evaluate their own models. The paper mentions \"we will open-source our codebase\" (reproducibility statement) but it's unclear if this includes an easy-to-use evaluation script. For a benchmark paper, providing a simple evaluate interface would greatly increase adoption. The current workflow seems to require manual video generation then evaluation, which is cumbersome. For practical use, an automated submission or leaderboard system would make the benchmark much more usable.\n- Metric definitions design lack justification. A few metrics (like the physics consistency score) are described in high-level terms but not fully formalized. It’s hard to know how reproducible they are from the text alone, e.g., whether they use learned physical estimators or ground-truth simulations. The Yes/No binary evaluation (Section 3.1) might be too coarse-grained - a partial physics violation might deserve a score between 0 and 1. The paper acknowledges models often \"rationalize\" actions rather than fail outright (Section 4.4), suggesting a more nuanced scoring could capture important phenomena.\n- Limited real-world coverage. Most of the benchmark focuses on synthetic data (e.g., MuJoCo or Unity scenes). It would be great to see more real videos or robotic interactions to test generalization.\n- The scope of the benchmark is somehow overlapped with previous works, e.g., the inclusion of an \"anti-physics\" category is similar as “counterfactual prompts” in [1]. \n\n[1] T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation. 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920346862,"tcdate":1761031520386,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8464/Reviewer_FCTE"],"signatures":["ICLR.cc/2026/Conference/Submission8464/Reviewer_FCTE"],"forum":"rlZeILv3fm","number":2,"license":"CC BY 4.0","cdate":1761031520386,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8464/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920346862,"domain":"ICLR.cc/2026/Conference","replyto":"rlZeILv3fm","id":"tgrWH9jead","forumContent":{"venue":{"value":"ICLR 2026 Oral"},"TLDR":{"value":"Large-scale, multidimensional video generation for physics"},"keywords":{"value":["Video Generation","Video Evaluation"]},"supplementary_material":{"value":"/attachment/bd8cdaa60661c40d3e703b3dd09ff6689653ba72.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents $PhyWorldBench$\n, a comprehensive benchmark designed to evaluate video generation models based on their adherence to the laws of physics. The benchmark covers multiple levels of physical phenomena, ranging from fundamental principles like object motion and energy conservation to more complex scenarios involving rigid body interactions and human or animal motion. Additionally, we introduce a novel \"Anti-Physics\" category, where prompts intentionally violate real-world physics, enabling the assessment of whether models can follow such instructions while maintaining logical consistency. Besides large-scale human evaluation, we also design a simple yet effective method that could utilize current MLLM to evaluate the physics realism in a zero-shot fashion. We evaluate 10 state-of-the-art text-to-video generation models, including five open-source and five proprietary models, with a detailed comparison and analysis. we identify pivotal challenges models face in adhering to real-world physics. Through systematic testing of their outputs across 1,050 curated prompts—spanning fundamental, composite, and anti-physics scenarios—we identify pivotal challenges these models face in adhering to real-world physics. We then rigorously examine their performance on diverse physical phenomena with varying prompt types, deriving targeted recommendations for crafting prompts that enhance fidelity to physical principles."},"_bibtex":{"value":"@inproceedings{\ngu2026phyworldbench,\ntitle={\\$PhyWorldBench\\$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models},\nauthor={Jing Gu and Xian Liu and Yu Zeng and Ashwin Nagarajan and Fangrui Zhu and Daniel Hong and Yue Fan and Qianqi Yan and Kaiwen Zhou and Ming-Yu Liu and Xin Eric Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=rlZeILv3fm}\n}"},"title":{"value":"$PhyWorldBench$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models"},"pdf":{"value":"/pdf/6138b5d05836a6d9ee27ae5fc6f3bbbe3667ff02.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gu|phyworldbench_a_comprehensive_evaluation_of_physical_realism_in_texttovideo_models"},"authorids":{"value":["~Jing_Gu2","~Xian_Liu1","~Yu_Zeng1","~Ashwin_Nagarajan1","~Fangrui_Zhu1","~Daniel_Hong1","~Yue_Fan3","~Qianqi_Yan1","~Kaiwen_Zhou3","~Ming-Yu_Liu1","~Xin_Eric_Wang2"]},"authors":{"value":["Jing Gu","Xian Liu","Yu Zeng","Ashwin Nagarajan","Fangrui Zhu","Daniel Hong","Yue Fan","Qianqi Yan","Kaiwen Zhou","Ming-Yu Liu","Xin Eric Wang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2606.28971v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"cui|selfevolving_agentic_image_restoration_via_deliberate_planning_and_intuitive_execution"},"html":{"value":"https://doi.org/10.48550/arXiv.2606.28971"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2606-28971,\n  publtype={informal},\n  author={Shuang Cui and Fan Ji and Guanglong Sun and Yufei Guo and Xiongxin Tang and Jiangmeng Li and Fanjiang Xu},\n  title={Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution},\n  year={2026},\n  month={June},\n  cdate={1780272000000},\n  journal={CoRR},\n  volume={abs/2606.28971},\n  url={https://doi.org/10.48550/arXiv.2606.28971}\n}\n"},"abstract":{"value":"Real-world image restoration (IR) remains challenging due to complex and coupled degradations. While recent agentic IR frameworks leverage Large Language Models for flexible tool planning, they face two critical limitations. First, from a search scheme perspective, excessive reliance on greedy strategies fails to balance exploration and exploitation. Second, existing agentic systems underutilize information, exhibiting episodic amnesia. To address these challenges, we propose \\textbf{Self-Evolving Agentic Image Restoration (SEAR)}, which formulates restoration as a sequential decision-making problem. Inspired by the dual-process theory, SEAR comprises an Intuitive Executor and a Deliberate Planner, respectively following the fast-thinking \\textit{System 1} and slow-thinking \\textit{System 2} principles. The Deliberate Planner employs Pruning-Aware Monte Carlo Tree Search for long-horizon reasoning, utilizing a hybrid no-reference reward and a Multimodal Large Language Model (MLLM)-based tournament to prevent metric exploitation. Complementarily, the Intuitive Executor leverages a self-evolving episodic memory indexed by degradation-aware state fingerprints. This mechanism distills expensive search trajectories into adaptive expertise, overcoming episodic amnesia while progressively amortizing cold-start exploration costs through memory reuse. Extensive experiments on synthetic and real-world benchmarks demonstrate its strong perceptual and quantitative performance."},"title":{"value":"Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution"},"authors":{"value":[{"fullname":"Shuang Cui","username":""},{"fullname":"Fan Ji","username":""},{"fullname":"Guanglong Sun","username":""},{"fullname":"Yufei Guo","username":"~Yufei_Guo1"},{"fullname":"Xiongxin Tang","username":""},{"fullname":"Jiangmeng Li","username":""},{"fullname":"Fanjiang Xu","username":""}]}},"tmdate":1784037798753,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2606-28971"],"tcdate":1784037794412,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Yufei_Guo1"],"forum":"02DjjkWwpU","license":"CC BY-SA 4.0","number":54705,"cdate":1780272000000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784037798753,"domain":"OpenReview.net/Public_Article","id":"02DjjkWwpU","version":2},{"content":{"summary":{"value":"This paper proposes a new physics-informed fluid predictor, named PINP. PINP firstly estimates the underlying pressure and velocity filed from observed fluid, which is constrained by a discretized physics loss. Then it employs an interpolation formalization of integral for future prediction, where an additional correction network is presented to reduce the error of discretized PDE predictor. Experimentally, PINP performs well in 2D and 3D flows and weather prediction tasks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.\tAbout implementation of baselines.\n\nIn NowCastNet, ensuring eidetic prediction results is one significant contribution of this paper. However, as shown in Figure 8, its prediction is quite blurry. I am wondering how the authors experimented with this baseline.\n\nBesides, in the supplementary materials, the prediction results of LSM and FNO appear strange periodic shakes. Actually, I think a well-trained deep model will not make such weird predictions. Did the authors carefully tune these two baselines?\n\n2.\tAbout spatial generalization.\n\nWhy PINP can achieve spatial generalization? Can the authors provide some intuitive explanations?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"This paper is overall well-written.\n\nThe idea of incorporating physics loss into fluid prediction is reasonable.\n\nThe authors have provided comprehensive experiments to verify the effectiveness of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThe title is kind of overclaimed. \n\n\nSince this paper is tailored to fluid prediction, I think “physics-informed fluid predictor” is more suitable. Otherwise, it is a little bit overclaimed, where there are extensive prediction tasks that PINP cannot solve, such as rigid body movement (controlled by classical physics) or magnetic field (governed by electromagnetism).\n\n2.\tA series of technical designs are underexplored or not well supported.\n\n(1)\tPINP adopts the discretized PDE loss for physics constraint, which may bring serious approximation error. The current design is based on the assumption that the differential operator can be approximated by spatial or temporal difference, which cannot be satisfied, especially in low-resolution data. Note that I am not saying that being physics-informed is a bad idea. The canonical physics-informed neural works employ the auto differential in neural works for approximation, which is much more precise than the discretization in PINP.\n\n(2)\tI cannot figure out that why additionally predicting the pressure field can boost the performance. As shown in Figure 2 (b), the predicted pressure field is only used in physical constraint loss, which cannot affect the future prediction process. This means that predicting the pressure field is just to fit the physical loss, which brings a new meaningless task. According to my experience, I think this design can only bring extra load to the model instead of benefiting the prediction. Besides, as shown in Figure 9(a), removing physical constraints will not bring a serious decrease. Further, How about keeping the second equation in Eq.(12) but removing the pressure-related one? I believe that the benefit of physics loss is mainly brought by the incompressible term loss. \n\n(3)\tThe design of the correction network is also weird. As formalized in Eq.(10), the inputs and outputs of the correction network are both expected to be close to the ground truth. Under this constraint, why correction network is necessary? (Minor: Eq.(10) may have a typo, where the comma should be “-”).\n\n(4)\tAbout the temporal loss. I am curious about how likely is this loss function to work. Some statistical results on how many times this loss is non-zero are expected.\n\nGoing further from (2), I doubt that the prediction of pressure field is useless in the current design, which is listed as one of the main contributions w.r.t. other papers. I think compared with Helmfluid, the advantage of PNIP lies in the physical loss, which can provide a more direct and explicit constraint to the velocity field.\n\nIn summary, I think there are many unsupported designs in the proposed method, which may affect the claim of the main contribution of this paper.\n\n3.\tAbout the efficiency. \n\nI am curious about the training overload. Since the calculation of loss in Eq.(16) may also cause extra computation costs than other baselines."}},"nonreaders":[],"tmdate":1732589304271,"tcdate":1730008439512,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3043/Reviewer_3QEU"],"signatures":["ICLR.cc/2025/Conference/Submission3043/Reviewer_3QEU"],"forum":"vAuodZOQEZ","number":1,"license":"CC BY 4.0","cdate":1730008439512,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3043/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732589304271,"domain":"ICLR.cc/2025/Conference","replyto":"vAuodZOQEZ","id":"Xpsa4bLaJj","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"We integrated PDEs into the network and loss function, enabling the prediction of observable physical quantities and the inference of future latent physical quantities as interpretation."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Fluid dynamics","Spatiotemporal prediction","Physics-informed learning"]},"supplementary_material":{"value":"/attachment/ec4cb85553651f092bb0fa29bd389e39c1a363ee.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Accurately predicting fluid dynamics and evolution has been a long-standing challenge in physical sciences. Conventional deep learning methods often rely on the nonlinear modeling capabilities of neural networks to establish mappings between past and future states, overlooking the fluid dynamics, or only modeling the velocity field, neglecting the coupling of multiple physical quantities. In this paper, we propose a new physics-informed learning approach that incorporates coupled physical quantities into the prediction process to assist with forecasting. Central to our method lies in the discretization of physical equations, which are directly integrated into the model architecture and loss function. This integration enables the model to provide robust, long-term future predictions. By incorporating physical equations, our model demonstrates temporal extrapolation and spatial generalization capabilities. Experimental results show that our approach achieves the state-of-the-art performance in spatiotemporal prediction across both numerical simulations and real-world extreme-precipitation nowcasting benchmarks."},"_bibtex":{"value":"@inproceedings{\nchen2025pinp,\ntitle={{PINP}: Physics-Informed Neural Predictor with latent estimation of fluid flows},\nauthor={Huaguan Chen and Yang Liu and Hao Sun},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=vAuodZOQEZ}\n}"},"title":{"value":"PINP: Physics-Informed Neural Predictor with latent estimation of fluid flows"},"pdf":{"value":"/pdf/d89b44f8442093480a12ba08c016520fc392eddf.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"chen|pinp_physicsinformed_neural_predictor_with_latent_estimation_of_fluid_flows"},"authorids":{"value":["~Huaguan_Chen1","~Yang_Liu52","~Hao_Sun4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Huaguan Chen","Yang Liu","Hao Sun"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a decoder-only Transformer framework, named DoPformer, for physics-informed neural networks (PINNs). The key idea is to simplify the PINNsFormer encoder–decoder design by using a lightweight decoder-only structure that retains temporal coupling through self-attention, while reducing the number of parameters. To enhance spectral fidelity and efficiency, the model integrates two optional modules: a Fourier neural-operator branch (DoPformer+NO) and a feed-forward block based on the Kolmogorov–Arnold network (KAN) (DoPformer+KAN). The paper evaluates these variants on canonical PDEs and reports that the models achieve competitive or better accuracy with fewer trainable parameters."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How is the Fourier operator integrated within the physics-informed loss? If the neural operator branch relies on paired mappings, how does it remain consistent with the PINN formulation?\n\n2. What happens if the encoder block is retained (i.e., a hybrid of PINNsFormer with KAN)? Would it improve accuracy by being parameter-efficient?\n\n3. It is interesting to explore a Chebyshev-based KAN formulation [2] and report its comparative performance.\n\n4. How would the proposed model handle complex or irregular geometries, where token formation and windowing become nontrivial?\n\n5. Please provide computational time comparisons to demonstrate the claimed efficiency.\n\n6. How is initialization handled for the KAN variant, and are there recommended basis configurations or hyperparameters that influence performance?\n\n7. Please explain why the Navier–Stokes case for DoPformer+KAN shows a sharp rise in rMAE and rRMSE? Is this a typo or due to instability?\n\n8. The paper mentions a drastic reduction in parameters when replacing MLPs with KANs. Please elaborate on whether the baseline MLP has already overfitted the problem and how fair the comparison is across architectures?\n\n[2] Shukla, Khemraj, et al. \"A comprehensive and FAIR comparison between MLP and KAN representations for differential equations and operator networks.\" Computer Methods in Applied Mechanics and Engineering 431 (2024): 117290."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The model achieves accuracy with a smaller number of parameters, showing potential for light-weight physics-informed architectures.\n\n2. The decoder-only design simplifies the architecture while maintaining competitive performance, making it practical for resource-limited PDE simulation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The integration of a Fourier neural operator within a physics-informed framework is conceptually confusing. Neural operators typically rely on supervised input–output mappings, while PINNs operate in a semi-supervised or unsupervised setting. This ambiguity requires clarification, particularly regarding how neural operators operate without paired data.\n\n2. The work does not address scalability issues. Memory consumption and out-of-memory (OOM) behavior common in PINNsFormer and PINNMamba are not discussed or benchmarked, which is crucial for higher-dimensional or more complex PDEs. Please see the last table in the appendix of PINNsMamba paper [1] for a detailed list of problems where PINNsFormer and PINNMamba face the issue of OOM. \n\n3. The range of problems tested remains limited. Highly oscillatory or strongly nonlinear PDEs are missing, making it unclear whether the Fourier or KAN augmentations are truly beneficial beyond canonical cases.\n\n4. Important architectural and implementation details are underspecified. For instance, initialization schemes for KAN, specific basis functions, and exact parameterization strategies are not described, making reproducibility challenging.\n\n5. The presentation could be improved. Some tables show abrupt jumps in errors (e.g., rMAE and rRMSE for the Navier–Stokes case with DoPformer+KAN) without explanation. Moreover, the discussion of why parameter reduction is so drastic compared to MLPs is insufficient.\n\n[1] Xu, Chenhui, et al. \"Sub-Sequential Physics-Informed Learning with State Space Model.\" Forty-second International Conference on Machine Learning."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931519916,"tcdate":1761992733748,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19675/Reviewer_frNH"],"signatures":["ICLR.cc/2026/Conference/Submission19675/Reviewer_frNH"],"forum":"B5BpwOHPlW","number":3,"license":"CC BY 4.0","cdate":1761992733748,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19675/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931519916,"domain":"ICLR.cc/2026/Conference","replyto":"B5BpwOHPlW","id":"dPrBOMJ9Zx","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Physics Informed Neural Network","PINN","Neural Operators"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Physics-Informed Neural Networks (PINNs) approximate PDE solutions by embedding physical constraints into training, yet MLP-based backbones often suffer from instability and loss of fidelity on long horizons. Recent sequence models (e.g., Transformers) alleviate some of these issues, but their encoder–decoder design adds parameters and memory pressure with limited benefit for autoregressive pseudo-sequences.\nWe introduce \\textbf{DoPformer}, a \\emph{decoder-only} Transformer tailored to physics-informed learning. DoPformer consumes short spatio–temporal pseudo-sequences, uses multi-head self-attention with WaveAct activations, and applies a sequential physics loss across the window. Removing the encoder and cross-attention yields a lighter model while preserving long-range temporal coupling through self-attention.\nTo further boost spectral accuracy, we explore two optional modules: (i) a Fourier \\emph{neural-operator} branch (\\textit{DoPformer+NO}) that improves oscillatory regimes and long-horizon rollouts; and (ii) a compact \\emph{KAN}-based feed-forward replacement (\\textit{DoPformer+KAN}) that drastically reduces parameters while maintaining strong accuracy.\nAcross convection, reaction, wave, and 2D Navier–Stokes equations, DoPformer consistently improves PINN accuracy and stability; the NO and KAN variants deliver additional gains depending on stiffness and spectral content. Our numerical results show that on these benchmarks DoPformer attains state-of-the-art accuracy among physics-informed models while using substantially fewer parameters."},"_bibtex":{"value":"@misc{\nbuzaev2025decoder,\ntitle={Decoder Only Transformer for Physics Informed Neural Networks},\nauthor={Fedor Buzaev and Andrei Ermakov and Dmitry Efremenko and Daria Pugacheva and Mariia Ivanova and Denis Derkach},\nyear={2025},\nurl={https://openreview.net/forum?id=B5BpwOHPlW}\n}"},"title":{"value":"Decoder Only Transformer for Physics Informed Neural Networks"},"pdf":{"value":"/pdf/378167ec8283d0e6a6688ce474c642f46e313c89.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"buzaev|decoder_only_transformer_for_physics_informed_neural_networks"},"authorids":{"value":["~Fedor_Buzaev1","~Andrei_Ermakov1","~Dmitry_Efremenko1","~Daria_Pugacheva1","~Mariia_Ivanova1","~Denis_Derkach1"]},"authors":{"value":["Fedor Buzaev","Andrei Ermakov","Dmitry Efremenko","Daria Pugacheva","Mariia Ivanova","Denis Derkach"]}},"version":2},{"content":{"summary":{"value":"This paper introduces DORIC (Domain-Universal, ODE-Regularized, Interpretable-Concept Transformer), a novel framework for time-series forecasting that combines explainable concept bottlenecks with physics-informed regularization. The model routes multivariate input through five interpretable latent concepts before predicting via a driven–damped ODE head. Unlike prior Transformers that optimize efficiency or frequency decomposition, DORIC emphasizes scientific plausibility and explainability. Experiments across six benchmarks (Electricity, Traffic, Weather, Illness, Exchange Rate, and ETT) show DORIC achieves the lowest MSE/MAE in most metrics."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. Can the authors provide quantitative or qualitative evidence (e.g., visualizations, case studies, or human evaluations) showing that the learned latent concepts correspond to their intended meanings and remain interpretable after training?\n2. Can the authors clarify why a driven–damped ODE was chosen as the universal physical prior across all datasets, and provide evidence—empirical or theoretical—that the learned ODE parameters correspond to meaningful dynamics rather than serving only as a generic regularizer?\n3. How are the ODE coefficients ($\\beta$, $\\gamma$) initialized and constrained? Are they shared across datasets?\n4. Can the method handle irregularly sampled or non-stationary time-series data without retraining?\n5. Are there any computational trade-offs compared to other Transformers (e.g., inference latency)?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper’s main strength is its innovative integration of concept bottlenecks and physics-informed residuals within a Transformer, effectively combining interpretability with physical plausibility. It presents a comprehensive evaluation across six diverse datasets and strong baselines, supported by detailed ablation studies that clearly show each component’s contribution. The authors emphasize interpretability through five structured latent concepts and provide theoretical grounding via expressiveness and convergence analyses. Additionally, the appendix enhances reproducibility by including implementation details and pseudo-code."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper’s central claim of interpretability is not convincingly demonstrated. While the architecture enforces a five-concept bottleneck, there is no theoretical proof or empirical validation that these learned concepts retain their intended meanings. Evidence is limited to internal correlations, without visual, human, or domain-level verification of interpretability.\n2. The paper introduces a driven–damped ODE as the core of its “physics-informed” design, but the justification is mostly heuristic. The ODE form is applied uniformly across unrelated domains without evidence that such dynamics are meaningful or empirically valid, and the learned coefficients are never analyzed. As a result, the physics component functions more as a generic smoothness prior than a genuinely grounded physical model.\n3. Figures lack axis explanations such as Figure 2 and Figure 3, making it hard to interpret visual differences quantitatively.\n4. There are some grammar errors:\nLine 149-150: “time series data first enters …” → plural mismatch; should be “data first enter …” or “the time-series signal first enters …”.\nLine 150-151: “prediction..” → double period; correct to “prediction.”\nLine 372-373: “Quantitative performance :” → remove space before colon.\nLine 323: “For he detailed theorem setting …” → missing “t”; should read “For the detailed theorem setting …”."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762947991314,"tcdate":1762947991314,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10601/Reviewer_FQ9K"],"signatures":["ICLR.cc/2026/Conference/Submission10601/Reviewer_FQ9K"],"forum":"oy4fc9h9oT","number":4,"license":"CC BY 4.0","cdate":1762947991314,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10601/-/Official_Review"],"mdate":1762947991314,"domain":"ICLR.cc/2026/Conference","replyto":"oy4fc9h9oT","id":"08oGragWHN","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Cross-domain generalization","Differentiable physics residuals","Physics-regularized forecasting","Time-series Transformer"]},"supplementary_material":{"value":"/attachment/6ac4d75b8a07056fcc57c632202974a94e8ff356.zip"},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Accurate, explainable and physically credible forecasting remains a persistent challenge for multivariate time-series with domain-varying statistical properties. \nWe propose DORIC, a Domain-Universal, ODE-Regularized, Interpretable-Concept Transformer for Time-Series Forecasting that generates predictions through five self-supervised, domain-agnostic concepts while enforcing differentiable residuals grounded in \nfirst-principles constraints. The concepts are softly regressed toward analytic statistics of the raw signal, and a driven–damped ODE head couples these concepts to the forecast as a shared, mean-reverting dynamical template across datasets. Unlike prior efficiency-focused Transformers, such as Informer(sparse attention) or FEDformer(frequency priors), DORIC combines latent explainability with explicit scientific constraints, while preserving the attention mechanism’s capacity to model long-range dependencies.  \nWe evaluate DORIC on six publicly-available datasets and it achieves the lowest error in eight of twelve MSE/MAE metrics.  \nCompared with TimeMixer, DORIC outperforms it on four datasets while maintaining strong interpretability. \nInterpretability analyses show that the learned concepts remain strongly aligned with their analytic targets, physics residuals stay relative to the signal scale, and the learned ODE coefficients follow domain-consistent patterns.\nAblation studies reveal complementary contributions: removing the physics residual increases average MSE from 0.328 to 0.547, eliminating concept alignment raises it to 0.698, and replacing the shared encoder with disjoint concept heads results in a 76% increase."},"_bibtex":{"value":"@misc{\nma2026signals,\ntitle={Signals, Concepts, and Laws: Toward Universal, Explainable Time-Series Forecasting},\nauthor={Hongwei Ma and Junbin Gao and Minh-Ngoc Tran},\nyear={2026},\nurl={https://openreview.net/forum?id=oy4fc9h9oT}\n}"},"title":{"value":"Signals, Concepts, and Laws: Toward Universal, Explainable Time-Series Forecasting"},"pdf":{"value":"/pdf/a508b4e475adef3862f73335d07c1eee297e1d44.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"ma|signals_concepts_and_laws_toward_universal_explainable_timeseries_forecasting"},"authorids":{"value":["~Hongwei_Ma1","~Junbin_Gao1","~Minh-Ngoc_Tran1"]},"authors":{"value":["Hongwei Ma","Junbin Gao","Minh-Ngoc Tran"]}},"version":2},{"content":{"summary":{"value":"This work attempts to provide regularization to deep operator network training by adding a \"pseudo-physics\" component when there is no knowledge of the PDE to inform training. A comprehensive experimental study with ablation is provided."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"The approach to learn the physics resembles that in Section IV-B of Zhang et al. \"Deep Learning and Symbolic Regression for\nDiscovering Parametric Equations\". The authors should give that reference and compare their approach to theirs.\n\nWhy do the authors use the acronym \"DONet\" for \"DeepONet\"? The latter is the term widely used in the literature.\n\nOn page 2, the discretized versions of u and f aren't \"collocation points\". That refers to points where a PDE residual is minimized.\n\nStill page 2, the efficiency of FNO does not reside in performing the convolution in the frequency domain, per se, but in learning the parameters in the frequency domain.\n\nOn page 3, it's not clear what the authors mean by having more data by decomposing the 128x128 input in 16,384 points. This is still the same amount of data.\n\nOn page 4, the authors say that the convolution layer is used to compensate for errors in the discretization of the derivatives. How?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"This is a method that attempts to provide physics regularization to the training of deep operator networks, which are usually trained only from data.\n\nA comprehensive ablation study is provided."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"This is a bootstrapping approach where one attempts to learn the physics and then use it to improve training of the operator network over a data-drive baseline. It is not clear how the pseudo-physics constraint helps achieve a better solution. This could happen simply by additional training of the operator network. There is no firm rationale for how this should work.\n\nNo comparison is made with the physics-informed neural operator using the correct physics. The authors' method should give a solution with accuracy between the data-driven and the physics-informed cases, but we do not know how much improvement is made unless we can see what the accuracy of the fully physics-informed operator network is.\n\nThere is some incorrect terminology (see Questions) and incorrect technical statements. For example, in Section 3.1, the authors state that the PDE solution can be obtained through integration of Green's function, but this is only true for linear PDEs.\n\nThe authors assess the additional number of parameters in their model, which is small, but nothing is said about the additional training and inference time incurred. The latter is important because there is a lot of iterative training and refinement in the proposed method."}},"nonreaders":[],"tmdate":1731428278095,"tcdate":1730656992141,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4999/Reviewer_MPgk"],"signatures":["ICLR.cc/2025/Conference/Submission4999/Reviewer_MPgk"],"forum":"CrmUKllBKs","number":6,"license":"CC BY 4.0","cdate":1730656992141,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4999/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428278095,"domain":"ICLR.cc/2025/Conference","replyto":"CrmUKllBKs","id":"RIa1LlKBfn","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Pseudo Physics","Data-Driven Physics Discovery","PDEs","Neural Operator","AI for science","Scientific Machine Learning"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in operator learning are transforming the landscape of computational physics and engineering, especially alongside the rapidly evolving field of physics-informed machine learning. The convergence of these areas offers\nexciting opportunities for innovative research and applications. However, merging\nthese two realms often demands deep expertise and explicit knowledge of physical systems, which may be challenging or even impractical in relatively complex applications. To address this limitation, we propose a novel framework: Pseudo\nPhysics-Informed Neural Operator (PPI-NO). In this framework, we construct a\nsurrogate physics system for the target system using partial differential equations\n(PDEs) derived from simple, rudimentary physics knowledge, such as basic differential operators. We then couple the surrogate system with the neural operator model, utilizing an alternating update and learning process to iteratively enhance\nthe model’s predictive power. While the physics derived via PPI-NO may not mirror the ground-truth underlying physical laws — hence the term “pseudo physics” — this approach significantly enhances the accuracy of current operator learning\nmodels, particularly in data scarce scenarios. Through extensive evaluations across\nfive benchmark operator learning tasks and an application in fatigue modeling,\nPPI-NO consistently outperforms competing methods by a significant margin. The\nsuccess of PPI-NO may introduce a new paradigm in physics-informed machine\nlearning, one that requires minimal physics knowledge and opens the door to\nbroader applications in data-driven physics learning and simulations."},"_bibtex":{"value":"@misc{\nchen2025pseudo,\ntitle={Pseudo Physics-Informed Neural Operators},\nauthor={Keyan Chen and Yile Li and Da Long and WEI W. XING and Jacob Hochhalter and Shandian Zhe},\nyear={2025},\nurl={https://openreview.net/forum?id=CrmUKllBKs}\n}"},"title":{"value":"Pseudo Physics-Informed Neural Operators"},"pdf":{"value":"/pdf/864b77c21caf8f31310746c5d9b464fe5feadfa1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|pseudo_physicsinformed_neural_operators"},"authorids":{"value":["~Keyan_Chen3","~Yile_Li1","~Da_Long1","~WEI_W._XING1","~Jacob_Hochhalter1","~Shandian_Zhe1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Keyan Chen","Yile Li","Da Long","WEI W. XING","Jacob Hochhalter","Shandian Zhe"]}},"version":2},{"content":{"summary":{"value":"Motivated by the convergence discrepancy of the loss terms in physics-informed training of neural networks, the paper proposes a new activation function, IRELU, and a label rescaling method, while abandoning common normalization techniques."},"presentation":{"value":"1 poor"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The paper tries to address an important challenge with PINNs, concerning the discrepancy between loss terms in physics-informed training. \n- The direction taken by authors in focusing on the units of derivatives and labels is interesting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- As also pointed out by the authors, the proposed IRELU activation has limited usability in physics-informed models, where one might need derivatives of an arbitrary order w.r.t. inputs, while derivatives for IRELU are $0$ for third and higher order derivatives. Even for second order PDEs, the effects of a constant second derivative of $1$ in IRELU (for $x>0$) need more attention and study. \n- Preventing vanishing and exploding gradients is one major characteristic of RELU. The gradient propagation of IRELU is not studied, though. As in your experiments, physics-informed training usually involves training with the solution data as well (IC, BC, etc.), where the first order derivative of IRELU ($=x$ for $x>0$) appears in the optimization. This is concerning for exploding gradients, especially since the paper also abandons normalization.\n- The notion of the 'difficulty of learning derivative-constrained info based on the loss scaling term' is inaccurate. While convergence discrepancy of different loss terms is known to happen in physics-informed training, the convergence rate is shown to be in favor of the residual loss (derivative-constrained term) in some cases [1].\n- Other works have studied activation functions in the physics-informed setting before [2, 3]. Lack of review and comparison with such works is surprising. Moreover, the authors do mention the adaptive loss scaling methods, but, there is again no comparison with those methods.\n\n\n[1] Wang, S., Yu, X., & Perdikaris, P. (2020). When and why PINNs fail to train: A neural tangent kernel perspective. ArXiv. /abs/2007.14527\n\n[2] Sitzmann, V., Martel, J., Bergman, A., Lindell, D., & Wetzstein, G. (2020). Implicit neural representations with periodic activation functions. Advances in neural information processing systems, 33, 7462-7473.\n\n[3] Jagtap, A. D., Kawaguchi, K., & Karniadakis, G. E. (2020). Adaptive activation functions accelerate convergence in deep and physics-informed neural networks. Journal of Computational Physics, 404, 109136.\n\n### Minor Comments\n- There are a few grammatical errors, and the readability can also be improved.\n- In Sec 3.2, Results, the references to Fig. 2 seem to be meant for Fig. 1."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. A more in-depth study of the proposed activation function would be really helpful. Authors may want to explain what characteristics IRELU shares with RELU and how it addresses the issues with polynomial activation functions.\n2. The presentation and readability can be greatly improved by adding more plots instead of tables; Especially, in the experiments and ablation study to show how the proposed methods contribute to addressing the loss discrepancy. Also, the plotting style in Figures 1a and 1b is not informative and rather confusing. \n3. Section 4.2 is very limited in justifying the proposed rescaling method and how it improves the learning of the derivative-constrained info. I would appreciate more details and insights regarding the choice of $C$ and why normalization methods are discouraged.\n4. As mentioned in the Weaknesses, comparison with other activation functions that are designed for or tested with PINNs is crucial."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636678828,"tcdate":1698994474866,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6220/Reviewer_WsvB"],"signatures":["ICLR.cc/2024/Conference/Submission6220/Reviewer_WsvB"],"forum":"knl4kGCagT","number":3,"license":"CC BY 4.0","cdate":1698994474866,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6220/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636678828,"domain":"ICLR.cc/2024/Conference","replyto":"knl4kGCagT","id":"JjjP93VSD3","forumContent":{"TLDR":{"value":"We propose methods to improve training of derivative-constrained NNs commonly found in physics-based applications."},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Scientific Machine learning","Physics-informed neural networks","Derivative-constrained"]},"supplementary_material":{"value":"/attachment/cdfef5f9cfbbc2ed062f59527ffd044b8389bfa4.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We refer to the setting where the (partial) derivatives of a neural network’s (NN’s)\npredictions with respect to its inputs are used as additional training signal as a\nderivative-constrained (DC) NN. This situation is common in physics-informed\nsettings in the natural sciences. We propose an integrated RELU (IReLU) acti-\nvation function to improve training of DC NNs. We also investigate denormal-\nization and label rescaling to help stabilize DC training. We evaluate our meth-\nods on physics-informed settings including quantum chemistry and Scientific Ma-\nchine Learning (SciML) tasks. We demonstrate that existing architectures with\nactivations replaced with IReLU activations combined with denormalization/label\nrescaling better incorporate training signal provided by derivative constraints."},"_bibtex":{"value":"@misc{\nlo2024on,\ntitle={{ON} {TRAINING} {DERIVATIVE}-{CONSTRAINED} {NEURAL} {NETWORKS}},\nauthor={Kai Chieh Lo and Daniel Huang},\nyear={2024},\nurl={https://openreview.net/forum?id=knl4kGCagT}\n}"},"title":{"value":"ON TRAINING DERIVATIVE-CONSTRAINED NEURAL NETWORKS"},"pdf":{"value":"/pdf/418906c84cfb646bbf036152d08c8d9f701a3a53.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"lo|on_training_derivativeconstrained_neural_networks"},"authorids":{"value":["~Kai_Chieh_Lo1","~Daniel_Huang3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kai Chieh Lo","Daniel Huang"]}},"version":2},{"content":{"correctness":{"value":"Is correct."},"summary_and_contributions":{"value":"This paper constructs a simulation framework for the popular social media platform Reddit using LLM agents seeded with synthetic personal profiles and generates SynthPAI dataset.\nContributions:\n1. A personalized LLM agents-based comment thread simulation framework for producing high-fidelity, diverse, and PAI research-aiding synthetic comments. \n2. A curated and hand-labeled synthetic PAI dataset, SynthPAI, of over 7800 comments, creating the first open and privacy-preserving dataset for PAI research. \n3. A public release of there framework and dataset SynthPAI and an extensive experimental evaluation showing that SynthPAI is diverse, realistic, and enables insightful PAI research."},"confidence":{"value":3},"documentation":{"value":"The dataset documentation is clear."},"rating":{"value":7},"title":{"value":"Review for \"A Synthetic Dataset for Personal Attribute Inference\""},"ethics":{"value":"No ethics concerns."},"clarity":{"value":"This paper is well written."},"review":{"value":"Pros:\n1. The use of LLM agents to generate synthetic data mitigates privacy concerns and provides a novel approach to dataset creation for PAI research.\n2. The synthetic comments are almost indistinguishable from real ones, as demonstrated by human studies, ensuring the dataset's utility.\n3. Extensive validation through comparisons with real-world data and across multiple state-of-the-art LLMs reinforces the dataset's reliability.\n4. The public release of the dataset and framework supports further research and development in PAI and privacy-preserving methodologies.\n\nCons:\n1. While the dataset includes a diverse range of attributes, it could be expanded to include more complex or nuanced personal information.\n2. The quality and diversity of the synthetic data are heavily dependent on the capabilities of the LLMs used, which may introduce inherent biases.\n3. Despite the automation, the need for manual verification and labeling of the generated comments is labor-intensive and may not scale efficiently."},"strengths":{"value":"1. By using synthetic data, the authors effectively address privacy concerns that hinder research using real personal data.\n2. The dataset includes a wide range of topics and demographic attributes, making it representative of real-world scenarios.\n3. The comprehensive validation through human studies and replication of prior experiments ensures the dataset's credibility and applicability.\n4. The framework allows for the scalable generation of synthetic data, which can be adapted for various research needs and extended with additional attributes."},"flag_for_ethics_review":{"value":"2: No, there are no or only very minor ethics concerns"},"relation_to_prior_work":{"value":"This paper discusses the relation with prior work."},"opportunities_for_improvement":{"value":"1. It’s better to apply the framework on more social media platform. Future work could include a broader and more nuanced set of personal attributes to enhance the dataset's applicability.\n2. Improving the accuracy of automated labeling could reduce the manual effort required and increase scalability.\n3. Addressing and mitigating biases inherent in the LLMs used to generate the synthetic data could further improve the dataset's quality and fairness.\n4. Incorporating features such as upvotes or comment metadata from real-world platforms could enhance the realism and utility of the synthetic dataset."},"additional_feedback":{"value":"N/A"},"limitations":{"value":"The authors adequately address the limitations and future works."}},"nonreaders":[],"tmdate":1731500655345,"tcdate":1719878906196,"writers":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission307/Reviewer_RgP5"],"signatures":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission307/Reviewer_RgP5"],"forum":"1nqfIQIQBf","number":1,"license":"CC BY 4.0","cdate":1719878906196,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission307/-/Official_Review","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1731500655345,"domain":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","replyto":"1nqfIQIQBf","id":"dhzF4s6nmd","forumContent":{"TLDR":{"value":"We build an LLM-agent simulation framework for Reddit to generate synthetic data to advance inference-based privacy research."},"venue":{"value":"NeurIPS 2024 Track Datasets and Benchmarks Poster"},"pdf":{"value":"/pdf/9eda6f7930bb32b3cac7bf0f90459d4c360a52be.pdf"},"keywords":{"value":["privacy","synthetic data","large language models","social media"]},"venueid":{"value":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track"},"paperhash":{"value":"yukhymenko|a_synthetic_dataset_for_personal_attribute_inference"},"authorids":{"value":["~Hanna_Yukhymenko1","~Robin_Staab1","~Mark_Vero1","~Martin_Vechev1"]},"abstract":{"value":"Recently powerful Large Language Models (LLMs) have become easily accessible to hundreds of millions of users world-wide. However, their strong capabilities and vast world knowledge do not come without associated privacy risks. In this work, we focus on the emerging privacy threat LLMs pose – the ability to accurately infer personal information from online texts. Despite the growing importance of LLM-based author profiling, research in this area has been hampered by a lack of suitable public datasets, largely due to ethical and privacy concerns associated with real personal data. We take two steps to address this problem: (i) we construct a simulation framework for the popular social media platform Reddit using LLM agents seeded with synthetic personal profiles; (ii) using this framework, we generate *SynthPAI*, a diverse synthetic dataset of over 7800 comments manually labeled for personal attributes. We validate our dataset with a human study showing that humans barely outperform random guessing on the task of distinguishing our synthetic comments from real ones. Further, we verify that our dataset enables meaningful personal attribute inference research by showing across 18 state-of-the-art LLMs that our synthetic comments allow us to draw the same conclusions as real-world data. Combined, our experimental results, dataset and pipeline form a strong basis for future privacy-preserving research geared towards understanding and mitigating inference-based privacy threats that LLMs pose."},"_bibtex":{"value":"@inproceedings{\nyukhymenko2024a,\ntitle={A Synthetic Dataset for Personal Attribute Inference},\nauthor={Hanna Yukhymenko and Robin Staab and Mark Vero and Martin Vechev},\nbooktitle={The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track},\nyear={2024},\nurl={https://openreview.net/forum?id=1nqfIQIQBf}\n}"},"title":{"value":"A Synthetic Dataset for Personal Attribute Inference"},"authors":{"value":["Hanna Yukhymenko","Robin Staab","Mark Vero","Martin Vechev"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"TLDR":{"value":"Physics-based character control with dexterous human-object interaction via synthetic video imitation"},"keywords":{"value":["Character Animation","Human-Object Interaction","Video Imitation"]},"supplementary_material":{"value":"/attachment/3c9e19068957ff2229955e3098ebbad188a770f3.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Physics-based character control for dexterous Human-Object Interaction (HOI) relies heavily on motion-capture data, which limits its applicability beyond captured scenarios. Meanwhile, video generative models can synthesize realistic HOI videos for unseen objects and scenarios directly from text prompts, but 3D interactions learned from such synthetic data remain kinematic without ensuring physical plausibility. Imitating these videos with physics-based control is a promising direction, yet unreliable object tracking and hand-object misalignment make it difficult to obtain accurate imitation targets. To address these challenges, we present DeVI (Dexterous Video Imitation), a novel framework that learns physics-based character control for dexterous HOI by imitating text-conditioned synthetic videos. To bypass unreliable 3D object motion reconstruction, we introduce hybrid imitation targets that combine reconstructed 3D human motion with tracked 2D object trajectories. Our Visual HOI Alignment further refines the human reference to align with both the generated video and the initial 3D object configuration, making it suitable for physical interaction. Extensive experiments show that DeVI's hybrid representation is more effective than those used by the baselines and that Visual HOI Alignment improves reconstruction quality and imitation success. We further demonstrate physically plausible functional manipulation from text instructions, including target-aware interactions and articulated object manipulation, without requiring motion-capture demonstrations."},"_bibtex":{"value":"@inproceedings{\nanonymous2026devi,\ntitle={De{VI}: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=dqrBs2AyAC},\nnote={under review}\n}"},"title":{"value":"DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation"},"pdf":{"value":"/pdf/01586aa34218818a5b0408b0955a7c1f2afa3464.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791230343555,"tcdate":1789555553853,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission24765/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission24765/Authors"],"forum":"dqrBs2AyAC","license":"CC BY 4.0","number":24765,"cdate":1789555553853,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission24765/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791230343555,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"dqrBs2AyAC","version":2},{"content":{"summary":{"value":"This paper presents an optimization framework named PitStop for optimizing physics-informed objective functions. In the current paradigm, most work optimizes Physics-informes loss, which corresponds to the supervision loss of the temporal derivative of the governing physics model. By optimizing this derivative-based loss, the existing methods aim at attaining the optimal Supervision loss. However, this work points out that the optimal points of these two different loss functions are not always the same, and based on this observation, proposes PitStop as an alternative for the existing classical gradient-based optimization methods like Gradient Descent and Gauss-Newton methods. To analyze the properties of the PitStop, the authors use the lens of linear fixed-point iterations. With this interpretation, the authors claim that while PitStop could converge to the worse fixed point than the Gradient Descent in terms of the physics-informed loss, it eventually converges to the better Supervision loss faster than these methods. The authors provide experimental results on one toy example, harmonic oscillator, Burger's equation, and Navier-Stokes equation, and show that PitStop converges faster to the better solution than the conventional methods."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"See above."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The motivation is good, and authors provided a thorough theoretical analysis of their approach. \n- With simple experiments, the authors effectively show that the current approach to minimize the Physics-informed loss does not always align with the final goal to minimize the Supervision loss. They also sho"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The overall description was hard to follow. More intuitive explanation about why PitStop works would be appreciated.\n- The overall description was hard to follow, making it challenge to reproduce the results\n- Even though the authors gave detailed definitions and analysis of their approach, it is unclear how we can implement it. Seeing the equation 7 and 8, I feel like we can reproduce the results by only cutting the gradient flow across different time steps, but I'm not sure if I understand it correctly. It would be helpful if the authors provide an explicit algorithm (or pseudo code).\n- The experimental results are not convincing that this approach is an overall better approach than existing methods.  For example,  in   Figure 4, I believe (d) GD gives better result than PitStop."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924522237,"tcdate":1761895246376,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14035/Reviewer_WCm7"],"signatures":["ICLR.cc/2026/Conference/Submission14035/Reviewer_WCm7"],"forum":"3yOpgyOcyd","number":4,"license":"CC BY 4.0","cdate":1761895246376,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14035/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924522237,"domain":"ICLR.cc/2026/Conference","replyto":"3yOpgyOcyd","id":"LdF7EjKcKQ","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Optimization","Physics-Informed Training"]},"supplementary_material":{"value":"/attachment/4dddd8ed4b74348f3a93c3480e25ca585a0a9745.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics-informed learning offers a powerful approach for modeling physical systems by enforcing governing equations directly within the training process. However, optimizing such models remains inherently challenging, especially for large systems, due to the ill-conditioned nature of these residual-based loss functions. In this paper, we critically examine the limitations of classical optimization techniques by developing a comprehensive theoretical framework for physics-informed setups, including insights on convergence guarantees, convergence speed and fixed points. Next, we introduce PitStop, a novel optimization method for physics-informed training based on gradient stopping, which overcomes the limitations of classical methods by backpropagating feedback differently from the standard chain rule of calculus. The method is motivated and mathematically analyzed in our theoretical framework, incurs no additional computational cost compared to standard gradients, and achieves superior results in our experiments. Our work paves the way for more scalable and reliable physics-informed model training by fundamentally rethinking optimization paradigms."},"_bibtex":{"value":"@misc{\nschnell2026pitstop,\ntitle={PitStop: Physics-Informed Training with Gradient Stopping},\nauthor={Patrick Schnell and Nils Thuerey},\nyear={2026},\nurl={https://openreview.net/forum?id=3yOpgyOcyd}\n}"},"title":{"value":"PitStop: Physics-Informed Training with Gradient Stopping"},"pdf":{"value":"/pdf/6cbe4dc946e836592e8c8d339c89a9d54a085ffc.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"schnell|pitstop_physicsinformed_training_with_gradient_stopping"},"authorids":{"value":["~Patrick_Schnell1","~Nils_Thuerey1"]},"authors":{"value":["Patrick Schnell","Nils Thuerey"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a cross-modal representation learning framework named CARL, designed to address the issue that optimizing purely statistical objectives can disrupt underlying causal structures. CARL jointly optimizes three structure-preserving losses—Conditional Independence Preservation, Markov Boundary Preservation, and Monotonic Alignment Consistency—to ensure that the learned representation space retains the causal structure of the original data. The authors validate the approach on synthetic datasets and a real-world Human Phenotype Project (HPP) dataset, and provide theoretical guarantees showing that causal queries remain identifiable in the representation space."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1 Have the authors considered evaluating CARL on larger-scale cross-modal tasks (e.g., CLIP-style vision-language alignment) to assess generalization?\n\n2 CARL trains multiple independent predictors with cross-validation. Has its scalability on large-scale data been evaluated?\n\n3 For highly imbalanced modality-information densities (e.g., image vs. text), can CARL still effectively prevent the lower-density modality from being overshadowed?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1 First systematic treatment of cross-modal causal structure preservation, formalizing the CSP principle and the three core challenges.\n\n2 Introduces the ε-CSP definition, a attainability/consistency theorem, and a theorem for preserving identifiability of causal queries, providing rigorous guarantees.\n\n3 The three losses are well motivated and complementary, balancing conditional independence with information retention.\n\n4 Synthetic and real data jointly verify effectiveness, robustness, and interpretability.\n\n5 Successfully recovers known medical causal pathways on the HPP dataset, showcasing potential in complex biomedical scenarios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1 Although the appendix contains detailed proofs, the main text could better explain some theoretical results (e.g., an intuitive reading of the error bounds) to improve readability.\n\n2 While experiments span synthetic and real data, they do not include larger-scale cross-modal benchmarks (e.g., vision-language tasks), limiting the demonstration of generalization.\n\n3 In the DUAL configuration the method avoids using both image modalities simultaneously, which may limit information utilization in certain practical settings."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926942957,"tcdate":1762796247286,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16917/Reviewer_DeFu"],"signatures":["ICLR.cc/2026/Conference/Submission16917/Reviewer_DeFu"],"forum":"I43IOiimO6","number":4,"license":"CC BY 4.0","cdate":1762796247286,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16917/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926942957,"domain":"ICLR.cc/2026/Conference","replyto":"I43IOiimO6","id":"SZUUSLtvS2","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Structure-preserving Constraints","Representation Learning"]},"supplementary_material":{"value":"/attachment/429e88ccc8b5eb0ec3f760c75fd2f1613f860ae4.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Cross-modal representation learning is fundamental for extracting structured information from multimodal data to enable semantic understanding and reasoning. However, current methods optimize statistical objectives without explicit causal constraints, where nonlinear mappings can introduce spurious dependencies or eliminate critical mediators, leading to representation-induced structural drift that undermines the reliability of causal inference. Therefore, establishing theoretical guarantees for causal invariance in cross-modal representation learning remains a foundational challenge. To this end, we propose Causal Alignment and Representation Learning (CARL), which explicitly embeds causal structure preservation constraints into cross-modal alignment objectives. Specifically, CARL introduces a multi-consistency loss architecture that jointly optimizes conditional independence preservation and information bottleneck regularization to balance cross-modal compression with critical variable retention, ensuring low-density modalities are not masked by high-density reconstruction demands. We further incorporate monotonic alignment consistency loss to establish correspondence between semantic similarity and representation distance through Spearman correlation, and Markov boundary preservation loss to maintain identifiability conditions including backdoor, frontdoor, and instrumental variable criteria in the shared representation space. In synthetic experiments with known causal ground truth, CARL achieves state-of-the-art performance in preserving conditional independence patterns and maintaining causal query identifiability under structural uncertainty. Real-world validation on Human Phenotype Project data reveals that CARL successfully preserves causal structures between fundus vascular representations and cardiovascular events, demonstrating its capacity for reliable cross-modal causal inference in complex biomedical applications."},"_bibtex":{"value":"@inproceedings{\nli2026carl,\ntitle={{CARL}: Preserving Causal Structure in Representation Learning},\nauthor={Yulong Li and Xiwei Liu and Feilong Tang and Zhixiang Lu and Ming Hu and Yichen Li and Haochen Xue and Peixin Guo and Jionglong Su and Yutong Xie and Eran Segal and Imran Razzak},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=I43IOiimO6}\n}"},"title":{"value":"CARL: Preserving Causal Structure in Representation Learning"},"pdf":{"value":"/pdf/0a69c1508caeab4368226b7153992f49c27f86d4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|carl_preserving_causal_structure_in_representation_learning"},"authorids":{"value":["~Yulong_Li7","~Xiwei_Liu2","~Feilong_Tang2","~Zhixiang_Lu1","~Ming_Hu4","~Yichen_Li5","~Haochen_Xue1","~Peixin_Guo2","~Jionglong_Su1","~Yutong_Xie4","~Eran_Segal1","~Imran_Razzak2"]},"authors":{"value":["Yulong Li","Xiwei Liu","Feilong Tang","Zhixiang Lu","Ming Hu","Yichen Li","Haochen Xue","Peixin Guo","Jionglong Su","Yutong Xie","Eran Segal","Imran Razzak"]}},"version":2},{"content":{"summary":{"value":"This paper explores the mathematical capabilities of the LLaMA-2 7B model, suggesting that small language models can achieve high performance in mathematical reasoning when fine-tuned with synthetic data generated by GPT-4 Turbo. The LLaMA-2 7B model achieves 82.4% accuracy on the GSM8K benchmark and 40.1% on the MATH benchmark, outperforming some larger models, including earlier versions of GPT-4. By scaling up the supervised fine-tuning (SFT) data with GPT-4-generated questions, the authors claim that the synthetic data performs comparably to real data. The approach purportedly avoids performance saturation and shows competitive results across benchmarks."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"weaknesses"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"Strengths:\nScaling Approach: The paper demonstrates that scaling synthetic data can effectively improve performance for smaller models, which challenges the assumption that only large-scale models excel in complex reasoning tasks.\nHigh Benchmark Performance: The LLaMA-2 7B model's performance on GSM8K and MATH benchmarks is commendable and suggests that synthetic data can be a viable alternative to real data in specific contexts.\nExperiment Setup: The paper includes various data scaling scenarios and compares performance against state-of-the-art models, adding insight into the practical applications of synthetic data for improving mathematical capabilities in language models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1 Training Data Overlap Concerns: There’s a significant possibility that LLaMA-2's pretraining data includes or paraphrases the benchmarks used in evaluation, especially given the unknowns around its training corpus. This overlap could mean that the results simply \"unlock\" capabilities that were already in the model due to exposure during pretraining. Testing on a model where pretraining data is known and does not include the benchmarks would provide a more credible assessment.\n2 Limited Benchmark Selection: The focus on GSM8K and MATH restricts evaluation to high school and grade school mathematics. More challenging, diverse benchmarks like MMLU (especially college-level sections) and OlympiadBench could provide a stronger test of the model's robustness. Success on OlympiadBench (IMO) [3], for example, would be a stronger indication of broader mathematical reasoning abilities.\n3 Overstated Claims of Generality: The paper frequently implies that its findings apply to \"common 7B LMs,\" yet the experiments focus solely on LLaMA-2 7B. Generalizing these results to other 7B models without further testing is premature.\n4 No Discussion on Synthetic Data Risks: The paper lacks a discussion on \"model collapse\" from synthetic data, an issue documented in the literature [2]. While using GPT-4 may mitigate some risks, the absence of model degradation across such extensive data scaling is unusual. This aspect needs further exploration to explain the improvement consistency.\n5 Large Data Volume Requirements: The use of 1024 generations per problem (yielding up to 76M examples) raises questions about the efficiency and feasibility of this approach. Such large data requirements, combined with the potential overlap between synthetic and benchmark data, make the setup less compelling.\n6 Reliance on GPT-4 as Judge: The use of GPT-4 to evaluate generated responses without validating its correlation with human evaluations weakens the reliability of the reported improvements. A smaller-scale human evaluation or citation demonstrating alignment between GPT-4 judgments and human judgments would improve credibility.\n7 Power Law Claim without Evidence: The claim of \"power law scaling\" is unsupported by a functional form or detailed analysis. Without a clearly defined functional relationship, the observed improvements should be framed as scaling trends, not power laws.\nReservations - Writing Concerns:\n8 Misleading Introduction on Emergence: The introduction on emergence theory is confusing (It’s known emergence is caused by metrics now and in addition this is not how we grade people) and does not add value to the main narrative. It distracts from the paper's core contributions and should be revised or strongly suggested its removed.\n9 Lack of Quantitative Detail in the Abstract: Statements like \"significant improvement\" are vague without absolute or relative performance changes. Quantitative results and implications should be included in the abstract to clearly communicate the impact of the findings.\n10 Clarity and Readability: The paper would benefit from a thorough review by a native English speaker. There are multiple areas where phrasing is unclear or lacks precision, detracting from the scientific rigor of the presentation.\n11 Captions and Figures: The captions do not provide sufficient context or summaries of the figures' main takeaways. Including summary values (like percentages) in the captions or directly in the figures would make it easier for readers to interpret the results.\n12 “Table 2 shows the results.” This type of sentence raises major concerns about the writing of the whole paper."}},"nonreaders":[],"tmdate":1731428411071,"tcdate":1730693322013,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11934/Reviewer_D8Z5"],"signatures":["ICLR.cc/2025/Conference/Submission11934/Reviewer_D8Z5"],"forum":"fL8sds4naU","number":3,"license":"CC BY 4.0","cdate":1730693322013,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11934/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428411071,"domain":"ICLR.cc/2025/Conference","replyto":"fL8sds4naU","id":"QOUTfMbkcI","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"Our method effectively increases the scale and diversity of SFT data, which can stimulate the mathematical capabilities of common general pretrained LLMs."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large language model","Math capabilities","Synthetic data","Alignment"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"It was once believed that mathematical capabilities in language models required either large model scales or extensive math-related data pre-training. However, this paper demonstrates that the small-scale LLaMA-2 7B model already possesses strong mathematical potential. This is evidenced by its impressive scores of 97.6% on GSM8K benchmark and 70% on MATH benchmark, achieved by selecting the oracle response from 1024 generations. Equipped GPT-4 Turbo as an additional verification, LLaMA-2 7B also achieves 91.8% accuracy on GSM8K benchmark. This indicates that the primary issue within current models is the difficulty in consistently eliciting the inherent mathematical capabilities. We find that scaling up synthetic SFT data, which proves to be nearly as effective as real data, can significantly enhance the reliability of generating correct answers. Surprisingly, even with approximately one million samples, we observe no clear performance saturation. And our method is more efficient with large data scale than previous works. This approach achieves an accuracy of 82.4% on GSM8K and 40.1% on MATH using LLaMA-2 7B model, surpassing GPT-3.5 Turbo. Our 70B model even exceeds an early version of GPT-4 on MATH and out-of-domain Hungarian National High School Math Exam. These results demonstrate our method significantly elicits the general mathematical capabilities of language models. Also, we provide insights into scaling behaviors across different reasoning complexities."},"_bibtex":{"value":"@misc{\nli2024common,\ntitle={Common 7B Language Models Already Possess Strong Math Capabilities},\nauthor={Chen Li and Weiqi Wang and Jingcheng Hu and Yuxing Wei and Han Hu and Zheng Zhang and Houwen Peng and Nanning Zheng},\nyear={2024},\nurl={https://openreview.net/forum?id=fL8sds4naU}\n}"},"title":{"value":"Common 7B Language Models Already Possess Strong Math Capabilities"},"pdf":{"value":"/pdf/fce275767d90e74e35ac012738dbeb6b93317c28.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|common_7b_language_models_already_possess_strong_math_capabilities"},"authorids":{"value":["~Chen_Li29","~Weiqi_Wang3","~Jingcheng_Hu1","~Yuxing_Wei1","~Han_Hu1","~Zheng_Zhang4","~Houwen_Peng2","~Nanning_Zheng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chen Li","Weiqi Wang","Jingcheng Hu","Yuxing Wei","Han Hu","Zheng Zhang","Houwen Peng","Nanning Zheng"]}},"version":2},{"content":{"summary":{"value":"This paper proposes UrbanGraph, a framework for urban microclimate prediction that integrates physics-informed, dynamic, and heterogeneous graph neural networks. The key idea is to explicitly encode time-varying physical processes, such as shading, vegetation evapotranspiration, and convective diffusion, into the topology of a dynamic heterogeneous graph. This graph structure is then processed by a spatio-temporal model combining a Relational Graph Convolutional Network (RGCN) for spatial dependencies and an LSTM for temporal evolution. Moreover, the authors also curate the UMC4/12 dataset, a high-resolution, physics-based simulation benchmark. Empirical experiments demonstrate that UrbanGraph outperforms various strong baselines."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How expensive is rebuilding 3 physics-driven edge sets every hour for a large city-level grid (say 50k–100k nodes)? Or have you considered or tested any strategies to manage this complexity?\n2. You report on six target variables (UTCI, PET, AT, MRT, WS, RH). Are they trained with a single shared UrbanGraph and separate heads, or trained separately per target (Eq. (1) suggests separate mappings)? If separate, could multitask training further boost R² through shared spatial embeddings?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The problem is well-motivated.\n2. The proposed idea of encoding observational/structural bias directly into the topology, i.e., physics-informed edges rather than physics-informed losses, is principled and technically sound.\n3. The dataset contribution is valuable to the community.\n4. The paper is well written."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The method is evaluated on a CFD-style simulator (ENVI-met) under several urban configurations. This is fine for an anonymized submission, but the key claim is physics-informed generalization to realistic urban microclimates. Without any real-world/field-sensor validation, it’s hard to tell if the hand-crafted dynamic edge rules are robust to noisy or incomplete inputs. Would it be possible to deploy on real city data (even small-scale)?\n2. The model uses an LSTM as the temporal evolution module. They compare to GRU and Transformer, but the argument for LSTM as the final choice is mostly empirical. Given that the graph itself is dynamic, a temporal attention or cross-time graph operation might better exploit the structured changes in edges."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927648404,"tcdate":1761876696107,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17804/Reviewer_jFx8"],"signatures":["ICLR.cc/2026/Conference/Submission17804/Reviewer_jFx8"],"forum":"ckjNF94cIi","number":3,"license":"CC BY 4.0","cdate":1761876696107,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17804/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927648404,"domain":"ICLR.cc/2026/Conference","replyto":"ckjNF94cIi","id":"BeOAzPHwKh","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Spatio-Temporal Graph","Heterogeneous Graph","Dynamic Graph","Physics-Informed ML","Urban Microclimate"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"With rapid urbanization, predicting urban microclimates has become critical, as it affects building energy demand and public health risks. However, existing generative and homogeneous graph approaches fall short in capturing physical consistency, spatial dependencies, and temporal variability. \\revise{To address this, we introduce UrbanGraph, a framework founded on a novel structure-based inductive bias. Unlike implicit graph learning, UrbanGraph transforms physical first principles into a dynamic causal topology, explicitly encoding time-varying causalities (e.g., shading and convection) directly into the graph structure to ensure physical consistency and data efficiency. Results show that UrbanGraph achieves state-of-the-art performance across all baselines. Specifically, the use of explicit causal pruning significantly reduces the model's floating-point operations (FLOPs) by 73.8\\% and increases training speed by 21\\% compared to implicit graphs. Our contribution includes the first high-resolution benchmark for spatio-temporal microclimate modeling, and a generalizable explicit topological encoding paradigm applicable to urban spatio-temporal dynamics governed by known physical equations."},"_bibtex":{"value":"@inproceedings{\nxin2026urbangraph,\ntitle={UrbanGraph: Physics-Informed Spatio-Temporal Dynamic Heterogeneous Graphs for Urban Microclimate Prediction},\nauthor={Weilin Xin and Chenyu Huang and Peilin Li and Jing Zhong and Jiawei Yao},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ckjNF94cIi}\n}"},"title":{"value":"UrbanGraph: Physics-Informed Spatio-Temporal Dynamic Heterogeneous Graphs for Urban Microclimate Prediction"},"pdf":{"value":"/pdf/ce906d47a515a333d8ac768dcd795f73b586e3da.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"xin|urbangraph_physicsinformed_spatiotemporal_dynamic_heterogeneous_graphs_for_urban_microclimate_prediction"},"authorids":{"value":["~Weilin_Xin1","~Chenyu_Huang7","~Peilin_Li3","~Jing_Zhong1","~Jiawei_Yao7"]},"authors":{"value":["Weilin Xin","Chenyu Huang","Peilin Li","Jing Zhong","Jiawei Yao"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a physics-AI hybrid modeling framework for fine-grained weather forecast. They propose to adaptively tune a PDE kernel together with a neural network as the encoder. Following the Euler time stepping, the PDE kernel can perform a fine-grained temporal forecast, which act as a physics-guided modeling part."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"I encourage the authors to address my concerns listed in the weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The combination of AI and physics is crucial and novel for weather forecast.\n2. The fine-grained weather forecast is of interest to nowcasting and temporal downscaling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Some opportunities to improve:\n\n1. The experimental details are way from sufficient. What are the hyperparameters you use, except for learning rate? How do you divide validation and test set? How do you divide input and label? What are the inputs and outputs to the model? Their sizes? What are the datasets' statistics, as there is none introduced? What is the time cost or number of parameters for your model compared to other models? How do you obtain the results in the tables, since there are no error bars? Are these all based on a one-time run? Are they statistically significant, and what are the p values, etc.? Table 3 does not demonstrate that the proposed model is the best model. Is there a specific reason why 120-min forecast is especially good? Any pseudo algorithm for understanding? Is there code for understanding and checking?\n\n2. The experimental results are both not good enough and not analyzed well enough. It looks strange from Figure 4 that FourcastNet is much worse than ECMWF-IFS even though FourcastNet should have been better as reported in the original paper. There needs discussion to justify, or maybe it is due to experimental setting. It is hard to know due to lack of experimental details pointed out above. The nowcast results are basically suggesting that the proposed model is no better than previous models, except for 120 min that is a less realistic setting for real life. Figure 6 is revealing very minimum information about the comparison between models. All the errors look alike, and it is hard to know which model's prediction is better without ground truth. I am not convinced by the explanation why the physics weight is decaying within each hour: unless your neural network counterpart's prediction does not change much, the ratio between physics and AI does not directly reflect how much they change. It could be that the AI part is predicting/contributing less, although the weight is surging. It does not mean anything. \n\n3. I want to highlight this problem using a separate paragraph. The ablation study is strange: it seems the physics part is not always helping. This might relate to how you use the PDE kernel and derive the PDEs. Your appendix A needs to cite references and name the equation names you are using. For example, how do you determine the coefficient for those constants in the equations? Shouldn't that be learned, since you cannot tell how big is the friction coefficient, let alone the term, for example? Moreover, Eq. 14 seems incorrect. I cannot recall such an equation in fluid dynamics. It is the continuity equation if you change p to z. However, pressure level slash geopotential slash height are not strictly the same thing. I feel concerned that they are assumed to be the same without any justification."},"limitations":{"value":"Limitations are discussed."}},"nonreaders":[],"tmdate":1730878929173,"tcdate":1719893385957,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission4475/Reviewer_jiN5"],"signatures":["NeurIPS.cc/2024/Conference/Submission4475/Reviewer_jiN5"],"forum":"ioAlzcELTf","number":1,"license":"CC BY 4.0","cdate":1719893385957,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission4475/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878929173,"domain":"NeurIPS.cc/2024/Conference","replyto":"ioAlzcELTf","id":"t7FY2WbMre","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["weather forecast","physics-AI hybrid model","partial differential equation","machine learning"]},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Data-driven artificial intelligence (AI) models have made significant advancements in weather forecasting, particularly in medium-range and nowcasting. However, most data-driven weather forecasting models are black-box systems that focus on learning data mapping rather than fine-grained physical evolution in the time dimension. Consequently, the limitations in the temporal scale of datasets prevent these models from forecasting at finer time scales. This paper proposes a physics-AI hybrid model (i.e., WeatherGFT) which generalizes weather forecasts to finer-grained temporal scales beyond training dataset. Specifically, we employ a carefully designed PDE kernel to simulate physical evolution on a small time scale (e.g., 300 seconds) and use a parallel neural networks with a learnable router for bias correction. Furthermore, we introduce a lead time-aware training framework to promote the generalization of the model at different lead times. The weight analysis of physics-AI modules indicates that physics conducts major evolution while AI performs corrections adaptively. Extensive experiments show that WeatherGFT trained on an hourly dataset, effectively generalizes forecasts across multiple time scales, including 30-minute, which is even smaller than the dataset's temporal resolution."},"_bibtex":{"value":"@inproceedings{\nxu2024generalizing,\ntitle={Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-{AI} Hybrid Modeling},\nauthor={Wanghan Xu and Fenghua Ling and Wenlong Zhang and Tao Han and Hao Chen and Wanli Ouyang and LEI BAI},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=ioAlzcELTf}\n}"},"title":{"value":"Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling"},"pdf":{"value":"/pdf/b8b7c1d96b59d20bcb7dfbdb207746fc2568d879.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"xu|generalizing_weather_forecast_to_finegrained_temporal_scales_via_physicsai_hybrid_modeling"},"authorids":{"value":["~Wanghan_Xu1","~Fenghua_Ling1","~Wenlong_Zhang3","~Tao_Han4","~Hao_Chen14","~Wanli_Ouyang1","~LEI_BAI1"]},"authors":{"value":["Wanghan Xu","Fenghua Ling","Wenlong Zhang","Tao Han","Hao Chen","Wanli Ouyang","LEI BAI"]}},"version":2},{"content":{"summary":{"value":"IFIN is a physics-guided network that interleaves differentiable forward operators with learnable inverse modules at each stage, jointly exploiting measurement and image domains to keep physical consistency and enrich features. With a physics-guided kernel adaptation for PSF mismatch, it achieves state-of-the-art lensless imaging reconstruction and improved robustness to noise and model errors, especially under severe blur."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.The paper claims to be physics-inspired, but the analysis is superficial, providing only a brief explanation of the forward and inverse processes in lensless imaging.\n2.It is unclear how the learnable PSF is obtained—the paper should specify the initialization, optimization strategy, and physical constraints involved.\n3.In Figures 3–5 and Table 1, the qualitative comparisons do not include recent methods from the past two years.\n4.For in-the-wild lensless imaging, there is no quantitative comparison. Additionally, it is unclear how many images are included in this dataset and how many images are contained in the proposed SV Lensless dataset."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"Interleaving differentiable forward operators with learnable inverse modules enforces physical consistency while enriching representations in both measurement and image domains.\n﻿The physics-guided kernel adaptation mitigates PSF mismatch/incompleteness, enabling constrained blind deconvolution and reducing sensitivity to model errors.\n﻿Demonstrates state-of-the-art reconstruction quality on challenging lensless benchmarks (including a new dataset)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Physics grounding is insufficient. Although the paper claims to be physics-inspired, it does not substantiate the physical modeling in depth (assumptions, operator derivations, constraints, or validation against instrument physics).\nRequires accurate, differentiable operators; robustness claims under strong model misspecification (nonlinear aberrations, spatially variant PSFs, misalignment) aren’t quantified.\nBlind PSF refinement can be ill-posed; constraints, regularizers, and failure modes (e.g., texture transfer, PSF–image ambiguity) are not discussed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926363883,"tcdate":1761578840307,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16203/Reviewer_eE1v"],"signatures":["ICLR.cc/2026/Conference/Submission16203/Reviewer_eE1v"],"forum":"q1TpQ6guwX","number":2,"license":"CC BY 4.0","cdate":1761578840307,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16203/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926363883,"domain":"ICLR.cc/2026/Conference","replyto":"q1TpQ6guwX","id":"opHPDadVgU","forumContent":{"TLDR":{"value":"We propose IFIN, the network couples forward physics and learned inverse at every layer with learnable calibration-free PSF, and shows state-of-the-art lensless imaging results under spatially varying blur and noise."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Computational Imaging","Lensless Imaging","Physics-guided Learning","Inverse Problem"]},"supplementary_material":{"value":"/attachment/9d05fbb0955cc21e2c4d9132124a9b7ae3b31195.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Inverse modeling plays a central role across computational optical imaging problems, including microscopy, imaging through scattering media, and lensless cameras, where the forward model often manifests as a severe blur. Discrepancies between the model and the actual imaging process further aggravate the ill-posed nature of the inverse problem. Physics-enabled methods that integrate analytical forward models with data-driven networks have been explored, but most incorporate physics only in a one-sided manner—either operating purely in the measurement space or only after inversion—thereby discarding complementary cues and reducing robustness to calibration errors.\nHere, we propose the Integrated Forward–Inverse Network (IFIN), a physics-guided deep neural network that interleaves differentiable forward operators with learnable inverse modules at every stage of the hierarchy. This design preserves physical consistency while shaping richer feature representations by jointly leveraging information from both measurement and image domains. A physics-guided kernel adaptation further compensates for inaccurate or unavailable PSF calibration, dynamically refining the kernel for blind deconvolution under system constraints.\nIFIN is especially effective when measurements are severely blurred by large point-spread functions, where conventional CNN-based inversion is limited by local receptive fields and underutilizes the measurement signal. On challenging lensless imaging benchmarks—including our newly introduced dataset, IFIN achieves state-of-the-art reconstruction quality and improved robustness under noise and model mismatch."},"_bibtex":{"value":"@misc{\nbae2026integrated,\ntitle={Integrated Forward{\\textendash}Inverse Network for Physics-Guided Image Reconstruction},\nauthor={Donggeon Bae and Jaewoo Jung and Yong Guk Kang and Kyung Chul Lee and Taeyoung Kim and Joonsik Park and Sangjun Byun and Jongho Kim and Nakkyu Baek and Hyeonyong Lee and Kyunghoon Jung and Seung Ah Lee},\nyear={2026},\nurl={https://openreview.net/forum?id=q1TpQ6guwX}\n}"},"title":{"value":"Integrated Forward–Inverse Network for Physics-Guided Image Reconstruction"},"pdf":{"value":"/pdf/0606a689890b9911fc407f24347161e12f8d2e59.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"bae|integrated_forwardinverse_network_for_physicsguided_image_reconstruction"},"authorids":{"value":["~Donggeon_Bae1","~Jaewoo_Jung1","~Yong_Guk_Kang1","~Kyung_Chul_Lee1","~Taeyoung_Kim11","~Joonsik_Park1","~Sangjun_Byun1","~Jongho_Kim3","~Nakkyu_Baek1","~Hyeonyong_Lee1","~Kyunghoon_Jung2","~Seung_Ah_Lee1"]},"authors":{"value":["Donggeon Bae","Jaewoo Jung","Yong Guk Kang","Kyung Chul Lee","Taeyoung Kim","Joonsik Park","Sangjun Byun","Jongho Kim","Nakkyu Baek","Hyeonyong Lee","Kyunghoon Jung","Seung Ah Lee"]}},"version":2},{"content":{"summary":{"value":"This paper addresses a crucial and often overlooked limitation in single-image 3D indoor scene generation: the lack of physical plausibility. While many recent methods achieve high visual fidelity, the resulting 3D scenes often contain obvious physical errors—such as floating objects, collisions, or unstable arrangements—making them unreliable for downstream applications like robotics and embodied AI. The authors propose a two-fold solution: a comprehensive Physics Evaluator and a novel framework, PhyMix, which integrates physics-based guidance into both the training and inference stages. The results demonstrate a significant advancement in generating scenes that are both visually faithful and physically consistent."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1. Is Eq.(3)  eps-prediction? Flow matching normally optimizes the conditional velocity field.\n\n2. Why the negative FM loss related to likelihood proxy?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"• Systematic Benchmarking: The introduction of the unified Physics Evaluator is arguably the most significant contribution. It provides the field with a long-overdue, systematic, and comprehensive set of nine metrics for physical plausibility. This moves the community past ad hoc collision or grounding checks toward a holistic assessment.\n• Elegant Technical Solution: The implicit-explicit optimization strategy (Scene-GRPO + TTO) is a theoretically elegant and effective method for handling the dual challenge of integrating both non-differentiable and differentiable constraints into a diffusion-based generative pipeline. The ablation studies confirm the necessity and complementarity of both components.\n• Strong Empirical Results and Validation: The performance gains are compelling, showing the method raises the overall physical score by +20.2% relative to the strongest baseline (MIDI). Crucially, the authors validate their metrics with a perceptual user study (MOS and Pairwise Preference), confirming that the quantitative scores align strongly with human judgment of physical plausibility"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Scene-GRPO is an application of the established flow-GRPO/GRPO preference learning paradigm, borrowing its framework directly from the LLM .\n\n2. The current Physics Evaluator relies on simple physical approximations (e.g., center-of-mass checks for static stability) that may fail to capture subtle or fine-grained edge cases, such as an object barely balancing on a thin edge, or the long-term effects of complex weight distribution and material properties. Furthermore, the reliance on these simplified physical metrics—and the design of the nine corresponding constraints—leans too heavily on hand-crafted engineering for the loss design. A more scientifically rigorous approach would involve integrating a sophisticated, general-purpose physics simulation engine for evaluation and differentiable loss, rather than relying on a custom set of rules derived from simplified geometric priors."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915649243,"tcdate":1761992132172,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission968/Reviewer_uokm"],"signatures":["ICLR.cc/2026/Conference/Submission968/Reviewer_uokm"],"forum":"RK6j9cwSK4","number":4,"license":"CC BY 4.0","cdate":1761992132172,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission968/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915649243,"domain":"ICLR.cc/2026/Conference","replyto":"RK6j9cwSK4","id":"UdOLbqIFsd","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"PhyMix couples a Physics Evaluator with Scene-GRPO and test-time optimization to produce physically consistent single-image 3D indoor scenes."},"keywords":{"value":["3D Scene Generation","Indoor Scene Synthesis","Physics-Aware Optimization"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Existing single-image 3D indoor scene generators often produce results that look visually plausible but fail to obey real-world physics, limiting their reliability in robotics, embodied AI, and design. To examine this gap, we introduce a unified Physics Evaluator that measures four main aspects: contact, stability, geometric priors, and deployability, which are further decomposed into nine sub-constraints, establishing the first benchmark to measure physical consistency. Based on this evaluator, our analysis shows that state-of-the-art methods remain largely physics-unaware. To overcome this limitation, we further propose a framework that integrates feedback from the Physics Evaluator into both training and inference, enhancing the physical plausibility of generated scenes. Specifically, we propose PhyMix, which is composed of two complementary components: (i) implicit alignment via Scene-GRPO, a critic-free group-relative policy optimization that leverages the Physics Evaluator as a preference signal and biases sampling towards physically feasible layouts, and (ii) explicit refinement via a plug-and-play Test-Time Optimizer (TTO) that uses differentiable evaluator signals to correct residual violations during generation. Overall, our method unifies evaluation, reward shaping, and inference-time correction, producing 3D indoor scenes that are both visually faithful and physically plausible. Extensive evaluations on synthetic dataset confirm state-of-the-art performance in both visual fidelity and physical plausibility, and extensive qualitative examples on stylized and real-world images further showcase the method’s robustness."},"_bibtex":{"value":"@misc{\nwu2026phymix,\ntitle={PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit{\\textendash}Explicit Optimization},\nauthor={Dongli Wu and Jingyu Hu and Ka-Hei Hui and Xiaobao Wei and Zhengzhe Liu and Jianqiang Li},\nyear={2026},\nurl={https://openreview.net/forum?id=RK6j9cwSK4}\n}"},"title":{"value":"PhyMix: Towards Physically Consistent Single-Image 3D Indoor Scene Generation with Implicit–Explicit Optimization"},"pdf":{"value":"/pdf/a08871c55f476dfd4fca9f88ae466e84b0ec4927.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|phymix_towards_physically_consistent_singleimage_3d_indoor_scene_generation_with_implicitexplicit_optimization"},"authorids":{"value":["~Dongli_Wu1","~Jingyu_Hu4","~Ka-Hei_Hui1","~Xiaobao_Wei1","~Zhengzhe_Liu2","~Jianqiang_Li2"]},"authors":{"value":["Dongli Wu","Jingyu Hu","Ka-Hei Hui","Xiaobao Wei","Zhengzhe Liu","Jianqiang Li"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to merely hundreds or thousands of examples, PhysInOne provides 2 million videos across 153,810 dynamic 3D scenes, covering 71 basic physical phenomena in mechanics, optics, fluid dynamics, and magnetism. Distinct from previous works, our scenes feature multiobject interactions against complex backgrounds, with comprehensive ground-truth annotations including 3D geometry, semantics, dynamic motion, physical properties, and text descriptions. We demonstrate PhysInOne’s efficacy across four emerging applications: physics-aware video generation, long-/short-term future frame prediction, physical property estimation, and motion transfer. Experiments show that fine-tuning foundation models on PhysInOne significantly enhances physical plausibility, while also exposing critical gaps in modeling complex physical dynamics and estimating intrinsic properties. As the largest dataset of its kind, orders of magnitude beyond prior works, PhysInOne establishes a new benchmark for advancing physics-grounded world models in generation, simulation, and embodied AI."},"_bibtex":{"value":"@inproceedings{\nzhou2026physinone,\ntitle={PhysInOne: Visual Physics Learning and Reasoning in One Suite},\nauthor={Siyuan Zhou and Hejun Wang and Hu Cheng and Jinxi Li and Dongsheng Wang and Junwei Jiang and Yixiao Jin and Jiayue Huang and Shiwei Mao and Shangjia Liu and Yafei Yang and Hongkang Song and Shenxing Wei and Zihui Zhang and Bing Wang and Zhihua Wang and Chuhang Zou and Bo Yang},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=OvbeHLZaSH}\n}"},"title":{"value":"PhysInOne: Visual Physics Learning and Reasoning in One Suite"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Zhou_PhysInOne_Visual_Physics_Learning_and_Reasoning_in_One_Suite_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"zhou|physinone_visual_physics_learning_and_reasoning_in_one_suite"},"authorids":{"value":["~Siyuan_Zhou3","~Hejun_Wang1","~Hu_Cheng2","~Jinxi_Li2","~Dongsheng_Wang12","~Junwei_Jiang1","~Yixiao_Jin2","~Jiayue_Huang1","~Shiwei_Mao1","~Shangjia_Liu1","~Yafei_YANG1","~Hongkang_Song1","~Shenxing_Wei2","~Zihui_Zhang2","~Bing_WANG8","~Zhihua_Wang9","~Chuhang_Zou1","~Bo_Yang7"]},"authors":{"value":["Siyuan Zhou","Hejun Wang","Hu Cheng","Jinxi Li","Dongsheng Wang","Junwei Jiang","Yixiao Jin","Jiayue Huang","Shiwei Mao","Shangjia Liu","Yafei Yang","Hongkang Song","Shenxing Wei","Zihui Zhang","Bing Wang","Zhihua Wang","Chuhang Zou","Bo Yang"]}},"tmdate":1789656674536,"pdate":1789656438489,"tcdate":1765220587613,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission32399/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission32399/Authors"],"forum":"OvbeHLZaSH","license":"CC BY 4.0","number":32399,"cdate":1765220587613,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission32399/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission32399/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656674536,"odate":1789656438489,"domain":"thecvf.com/CVPR/2026/Conference","id":"OvbeHLZaSH","version":2},{"content":{"venue":{"value":"Journal of Physics: Conference Series"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"batalo|statistical_moments_for_simulation_calibration_with_modelbridge"},"authorids":{"value":["~Bojan_Batalo1","~Lincon_Souza1","~Keisuke_Yamazaki1"]},"html":{"value":"https://iopscience.iop.org/article/10.1088/1742-6596/2701/1/012047"},"abstract":{"value":"Computer simulations are actively used for analyzing complex phenomena, especially in fields where access to their real-world counterparts is not feasible, such as physics, chemistry, material science, and others. The key to executing successful simulations is making sure that the parameters of a simulator reflect real-world scenarios, a tedious and error-prone effort addressed through simulation calibration. Recently, several methods have been proposed to automatize this task by learning from previously calibrated simulations using a model-bridge paradigm: a complex simulation is replaced by a simpler surrogate model, which can then be bridged to the calibrated simulation parameters. However, designing the surrogate model is a non-trivial problem involving trade-offs between simplicity of representation, interpretability and calibration accuracy, as well as the complexity of the bridge model required to map the surrogate to calibrated parameters. Further, while effective, such approaches can be non-intuitive for practitioners due to their distance from the simulation. In this paper, we view a simulation as a distribution of output variables, which can be easily represented by statistical moments. This yields a very simple and interpretable surrogate that can be bridged to calibrated parameters with a simple linear regression. We show that our method outperforms previous approaches, in terms of calibration accuracy and time, through experiments on simulations of turbulent flow dynamics and synthetic signals."},"title":{"value":"Statistical moments for simulation calibration with model-bridge"},"authors":{"value":["Bojan Batalo","Lincon Souza","Keisuke Yamazaki"]}},"tmdate":1751517022018,"pdate":1708873200000,"tcdate":1751517022018,"writers":["~Bojan_Batalo1","~Lincon_Souza1","~Keisuke_Yamazaki1"],"signatures":["~Bojan_Batalo1"],"forum":"YgwfB3SINo","license":"CC BY-NC 4.0","number":37537,"cdate":1751517022018,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1751517022018,"domain":"OpenReview.net/Archive","id":"YgwfB3SINo","version":2},{"content":{"summary":{"value":"This paper proposes an innovative framework named MAPS, which relies on multimodal large language models (MLLMs). During the inference phase, it integrates the simulated language descriptions of input diagrams generated by PPM and data obtained through the Chain-of-Simulation process, aiming to address complex structural and quantitative analysis issues in the field of physics. Through empirical studies on university-level circuit analysis problems, the authors demonstrate the significant effect of MAPS in enhancing the inference accuracy of MLLMs, outperforming all existing models. These results reveal the potential of MAPS in strengthening the multimodal scientific reasoning capabilities of MLLMs, providing a promising direction for the development of this field."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"Incorrect spelling and expression:\n- Line 243 - diagram diagram -> diagram\n- Appendix A.2, Figure 5, Row2 - Chnese -> Chinese\n\nSome Confusions:\nThe research idea proposed in this paper is highly innovative, but during my reading, I encountered several questions that I hope the authors can provide answers to:\n- Table 1 presents the conversion effectiveness of PPM on the synthetic dataset ppm-syn-lprc-test and the real-world dataset SimpleCircuitEval. It is noted that the recognition accuracy of Label-Type diagrams is lower than that of Numerical-Type diagrams, which seems contrary to common belief since label-type diagrams are generally considered easier to identify than numerical-type diagrams. What is the reason for this difference?\n- The test dataset SimpleCircuitEval comprises only 79 questions. Is this sample size sufficient to support a convincing evaluation result? Does it adequately represent the universality of LPRC question types?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"The advantages of this paper are mainly reflected in the following aspects:\n- This paper proposes an innovative process framework that can combine physical perception models (PPMs) with simulator outcomes to infer answers to physical problems. The framework integrates the understanding of physical diagrams with the reasoning of physical knowledge, and its effectiveness has been validated through experiments;\n- The paper designs and introduces a synthetic dataset named ppm-syn-lprc, which is used for fine-tuning the visual-language model PPM. The generation process of the synthetic dataset is elaborated in detail, and the scalability of this method is demonstrated. For instance, this approach can be applied to other fields such as mechanical systems, providing clear directions and ideas for improving the model's reasoning capabilities in other areas of physics;\n- The structure of the paper is well-organized and logical, making it easy for readers to understand and follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Writing:\n- There are some typos in the article, and it is recommended that the author carefully proofread and corrected them to enhance the professionalism and readability of the paper.\n\nExperimental Design:\n- Generalization Issues: The results of this study have only been validated on the GPT-4V model, which may not be sufficient to demonstrate the applicability of the framework to other model architectures. It is suggested that the authors extend the evaluation of the framework to different model architectures to strengthen the generalizability and credibility of the experimental conclusions.\n- Evaluation of Synthetic Data Effectiveness: The paper mentions using CogVLM-17B as the base model for PPM and fine-tuning it on the synthetic dataset. To clarify whether the improvement in the model’s ability to understand physical diagrams is due to the inherent capabilities of the base model or the introduction of synthetic data, the authors are advised to add comparative experiments before and after fine-tuning. This will help verify the effectiveness of synthetic data and reinforce the persuasiveness of the research findings."}},"nonreaders":[],"tmdate":1732960264197,"tcdate":1730709975145,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11154/Reviewer_xCVY"],"signatures":["ICLR.cc/2025/Conference/Submission11154/Reviewer_xCVY"],"forum":"GR0y0F3Ipd","number":4,"license":"CC BY 4.0","cdate":1730709975145,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11154/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732960264197,"domain":"ICLR.cc/2025/Conference","replyto":"GR0y0F3Ipd","id":"VhX3fOLhWI","forumContent":{"TLDR":{"value":"improving multi-modal scientific reasoning capability with physics perception model and simulation assistance"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-modal reasoning","scientific reasoning","physical simulation"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Pre-trained on extensive text and image corpora, current Multi-Modal Large Language Models (MLLM) have shown strong capabilities in general visual reasoning tasks. \nHowever, their performance is still lacking in physical domains that require understanding diagrams with complex physical structures and quantitative analysis based on multi-modal information. \nTo address this, we develop a new framework, named **M**ulti-Modal Scientific Re**A**soning with **P**hysics Perception and **S**imulation (**MAPS**) based on an MLLM. \nMAPS decomposes expert-level multi-modal reasoning task into physical diagram understanding via a Physical Perception Model (PPM) and reasoning with physical knowledge via a simulator. \nThe PPM module is obtained by fine-tuning a visual language model using carefully designed synthetic data with paired physical diagrams and corresponding simulation language descriptions. \nAt the inference stage, MAPS integrates the simulation language description of the input diagram provided by PPM and results obtained through a Chain-of-Simulation process with MLLM to derive the underlying rationale and the final answer. \nValidated using our collected college-level circuit analysis problems, MAPS significantly improves reasoning accuracy of MLLM and outperforms all existing models. \nThe results confirm MAPS offers a promising direction for enhancing multi-modal scientific reasoning ability of MLLMs. \nWe will release our code, model and dataset used for our experiments upon publishing of this paper."},"_bibtex":{"value":"@inproceedings{\nzhu2025maps,\ntitle={{MAPS}: Advancing Multi-Modal Reasoning in Expert-Level Physical Science},\nauthor={Erle Zhu and Yadi Liu and Zhe Zhang and Xujun Li and JinZhou and Xinjie Yu and Minlie Huang and Hongning Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=GR0y0F3Ipd}\n}"},"title":{"value":"MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science"},"pdf":{"value":"/pdf/c84516a8e4b9a68b710453218bcaa36dee327176.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhu|maps_advancing_multimodal_reasoning_in_expertlevel_physical_science"},"authorids":{"value":["~Erle_Zhu1","~Yadi_Liu1","~Zhe_Zhang24","~Xujun_Li2","~JinZhou1","~Xinjie_Yu1","~Minlie_Huang1","~Hongning_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Erle Zhu","Yadi Liu","Zhe Zhang","Xujun Li","JinZhou","Xinjie Yu","Minlie Huang","Hongning Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a unified pipeline that jointly generates shadows and performs object relighting, achieving physics-based control. It also introduces the ShadRel dataset for coupled light transport. The experimental setup is comprehensive, and the visual results are excellent for both synthetic and real-world cases."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"Please refer to the weakness."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- This paper presents the first generative method for the joint synthesis of shadows and object relighting from a single 2D RGB image. A key innovation is the introduction of a control mechanism conditioned on monocular depth, which enables the generation of continuous and coherent shadow and relighting effects in direct response to continuous variations in illumination.\n- This paper introduces ShadRel, a large-scale synthetic dataset developed for the tasks of shadow generation and object relighting. The dataset is notable for its extensive volume and the inclusion of diverse and complex lighting scenarios, rendering it a valuable asset for training and evaluating models in these domains.\n- This paper introduces a novel light-aware occlusion representation, LGI map. The methodology for acquiring this map is clearly articulated, and both theoretical analysis and experimental results compellingly demonstrate that the LGI map enables the generation of results with physically-based shadow and illumination constraints. This representation effectively encodes light-aware occlusion from monocular depth, providing a physics-inspired prior that constrains generative models to produce more realistic and consistent outputs.\n- The paper is well-supported by thorough experimentation, clear theoretical explanations, and compelling visual results."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The visual results presented in the paper are compelling. However, the experiments primarily showcase objects with relatively simple geometric structures. To more rigorously assess the robustness and generalization capabilities of the proposed method, I would encourage the authors to include results on more structurally complex objects. For instance, objects with fine-grained details, intricate parts, or significant self-occlusion (e.g., a bicycle, a detailed sculpture, or a potted plant) would serve as more challenging test cases. Demonstrating high-fidelity shadow and relighting effects on such objects would provide stronger evidence of the model's effectiveness and further strengthen the paper's contributions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925921343,"tcdate":1761968570500,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15667/Reviewer_YF9H"],"signatures":["ICLR.cc/2026/Conference/Submission15667/Reviewer_YF9H"],"forum":"yrVAA0czRz","number":4,"license":"CC BY 4.0","cdate":1761968570500,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15667/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925921343,"domain":"ICLR.cc/2026/Conference","replyto":"yrVAA0czRz","id":"WHacwsURkK","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Shadow Generation","Relight","Bridge matching"]},"supplementary_material":{"value":"/attachment/d85be80c26bbc80412d5e4b93e9fc862b08dcad3.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We propose Light–Geometry Interaction (LGI) maps, a novel representation that encodes light-aware occlusion from monocular depth. Unlike ray tracing, which requires full 3D reconstruction, LGI captures essential light–shadow interactions reliably and accurately, computed from off-the-shelf 2.5D depth map predictions. LGI explicitly ties illumination direction to geometry, providing a physics-inspired prior that constrains generative models. Without such prior,  these models often produce floating shadows, inconsistent illumination, and implausible shadow geometry. Building on this representation, we propose a unified pipeline for joint shadow generation and relighting-unlike prior methods that treat them as disjoint tasks-capturing the intrinsic coupling of illumination and shadowing essential for modeling indirect effects. By embedding LGI into a bridge-matching generative backbone, we reduce ambiguity and enforce physically consistent light–shadow reasoning. To enable effective training, we curated the first large-scale benchmark dataset for joint shadow and relighting, covering reflections, transparency, and complex interreflections. Experiments show significant gains in realism and consistency across synthetic and real images. LGI thus bridges geometry-inspired rendering with generative modeling, enabling efficient, physically consistent shadow generation and relighting."},"_bibtex":{"value":"@inproceedings{\nwang2026joint,\ntitle={Joint Shadow Generation and Relighting via Light-Geometry Interaction Maps},\nauthor={Shan Wang and Peixia Li and Chenchen Xu and Ziang Cheng and Jiayu Yang and Hongdong Li and Pulak Purkait},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=yrVAA0czRz}\n}"},"title":{"value":"Joint Shadow Generation and Relighting via Light-Geometry Interaction Maps"},"pdf":{"value":"/pdf/2fece7ce99f2c1676671d9376517363fc28a44da.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|joint_shadow_generation_and_relighting_via_lightgeometry_interaction_maps"},"authorids":{"value":["~Shan_Wang2","~Peixia_Li2","~Chenchen_Xu1","~Ziang_Cheng1","~Jiayu_Yang1","~Hongdong_Li1","~Pulak_Purkait1"]},"authors":{"value":["Shan Wang","Peixia Li","Chenchen Xu","Ziang Cheng","Jiayu Yang","Hongdong Li","Pulak Purkait"]}},"version":2},{"content":{"review":{"value":"The paper proposes GeoPT, a unified pretrained model for general physics simulation based on lifted geometric pretraining, which augments geometry with synthetic data without the requirement of physics labels. Pretrained on over one million data points, GeoPT achieves superior performance across multiple benchmarks from different domains, significantly accelerating the convergence speed. \n\nPros:\n- The proposed method shows the pathway for a unified pretraining approach for general physics simulation, significantly reducing the requirement of physical labels for training models.\n- The model manifests strong scalability with respect to model size and data diversity, highlighting the potential of GeoPT in handling more complex simulations.\n- The paper is well written, with the core idea clearly explained.\n\nCons:\n- The synthetic dynamics are sampled uniformly and are not constrained by conservation laws, PDE structure, or physical boundary conditions. Therefore, it may not faithfully reflect the structure of real physical dynamics."},"confidence":{"value":4},"rating":{"value":9},"title":{"value":"Review of GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training"}},"parentInvitations":"ICLR.cc/2026/Workshop/FM4Science/-/Official_Review","nonreaders":[],"tmdate":1772549369765,"tcdate":1771986970799,"writers":["ICLR.cc/2026/Workshop/FM4Science","ICLR.cc/2026/Workshop/FM4Science/Submission77/Reviewer_w8qP"],"signatures":["ICLR.cc/2026/Workshop/FM4Science/Submission77/Reviewer_w8qP"],"forum":"N9qIqvanBj","number":3,"license":"CC BY 4.0","cdate":1771986970799,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/FM4Science/Submission77/-/Official_Review","ICLR.cc/2026/Workshop/FM4Science/-/Edit"],"mdate":1772549369765,"domain":"ICLR.cc/2026/Workshop/FM4Science","replyto":"N9qIqvanBj","id":"UKdCfOiHkz","forumContent":{"TLDR":{"value":"This paper presents GeoPT as a unified model pre-trained from large-scale geometries for general physics simulation."},"venue":{"value":"ICLR 2026 Workshop FM4Science Poster"},"pdf":{"value":"/pdf/3df9a842bf6bf180192840eca36de1f7eda780d3.pdf"},"keywords":{"value":["Neural Simulation","PDE Solving","Self-Supervised Pre-Training"]},"venueid":{"value":"ICLR.cc/2026/Workshop/FM4Science"},"paperhash":{"value":"wu|geopt_scaling_physics_simulation_via_lifted_geometric_pretraining"},"authorids":{"value":["~Haixu_Wu1","~Minghao_Guo1","~Zongyi_Li1","~Zhiyang_Dou1","~Mingsheng_Long5","~Kaiming_He2","~Wojciech_Matusik2"]},"abstract":{"value":"Neural simulators promise efficient surrogates for physics simulation, but scaling them is bottlenecked by the prohibitive cost of generating high-fidelity training data. Pre-training on abundant off-the-shelf geometries offers a natural alternative, yet faces a fundamental gap: supervision on static geometry alone ignores dynamics and can lead to negative transfer on physics tasks. We present GeoPT, a unified pre-trained model for general physics simulation based on lifted geometric pre-training. The core idea is to augment geometry with synthetic dynamics, enabling dynamics-aware self-supervision without physics labels. Pre-trained on over one million samples, GeoPT consistently improves industrial-fidelity benchmarks spanning fluid mechanics for cars, aircraft, and ships, and solid mechanics in crash simulation, reducing labeled data requirements by 20-60% and accelerating convergence by 2$\\times$. These results show that lifting with synthetic dynamics bridges the geometry-physics gap, unlocking a scalable path for neural simulation and potentially beyond. Code is available at https://github.com/Physics-Scaling/GeoPT."},"_bibtex":{"value":"@inproceedings{\nwu2026geopt,\ntitle={Geo{PT}: Scaling Physics Simulation via Lifted Geometric Pre-Training},\nauthor={Haixu Wu and Minghao Guo and Zongyi Li and Zhiyang Dou and Mingsheng Long and Kaiming He and Wojciech Matusik},\nbooktitle={ICLR 2026 Workshop on Foundation Models for Science: Real-World Impact and Science-First Design},\nyear={2026},\nurl={https://openreview.net/forum?id=N9qIqvanBj}\n}"},"title":{"value":"GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training"},"authors":{"value":["Haixu Wu","Minghao Guo","Zongyi Li","Zhiyang Dou","Mingsheng Long","Kaiming He","Wojciech Matusik"]}},"version":2},{"content":{"summary":{"value":"The authors propose a Gaussian Splatting-inspired approach to ultrasound channel fitting and subsequent strain rate estimation. This is achieved by initializing a set of point scatters, which is further parameterized by a per-frame intensity. Given newly measured RF data, the method subsequently optimizes point scatter locations and intensities using a physics-inspired model. For strain rate computation, an affine transform is fitted between scatters in the first and last frames. However, the method is only validated against standard block-matching speckle tracking, while simulated data appears to be generated with a forward model that closely matches the one used for reconstruction, possibly biasing the method's performance."},"clarity_and_organization":{"value":"Fair"},"best_paper_award":{"value":"No"},"justification_for_rating":{"value":"This is a promising research direction for inverse ultrasound imaging. Preliminary results are encouraging, however, the boundary between previous work on differentiable off-grid ultrasound imaging and the proposed method is obscure. The proposed strain rate stage is also rather unspecified, especially given that it is the final purpose of the method. I would support acceptance if the contribution is reframed more precisely, the methods section is clarified appropriately, and a comparison to a fairer baseline is included."},"strengths":{"value":"- **Novel formulation:** the physics-inspired formulation for the reconstruction of ultrasound data is interesting, and compellingly combines Gaussian Splatting with an ultrasound model.\n- **Qualitatively positive:** The three figures support the idea of a more smooth, noise-robust estimation of strain rate in cardiovascular data."},"weaknesses":{"value":"- **Method framing:** The method is presented as an off-grid point-scatter approach analogous to splatting, but most of the underlying physics model seems to be inherited from INFER [1]. Methodological novelty therefore mostly lies in joint tracking across frames, which is absolutely novel, but should be presented fairly. The final strain rate computation is a conventional affine fit + Gaussian smoothing. The final strain rate map is also smoothed out substantially, which weakens the claim regarding localized strain rate estimation.\n- **Potential forward model mismatch:** Simulated measures are generated with Zea, while the proposed method follows a closely related ultrasound model. This may result in the method being evaluated under near perfect conditions, biasing the result. It would be interesting to evaluate performance under e.g. channel noise, mismatch in transmit waveform, or data generated by an independent simulator.\n- **Baseline does not operate under the same conditions:** While the proposed method is allowed to use the raw channel data, block matching only receives the resulting images, and the experiments therefore confounded the tracking method with input information. An RF-domain or phase-based baseline would strengthen the evaluation. Furthermore, a comparison with framewise INFER [1] followed by deformable registration would demonstrate the added benefit of jointly modelling the trajectory as proposed here.\n- **Lacking overview Figure:** A general overview Figure describing the methodological contribution would strengthen the paper. Section 2 currently moves directly into tensor and index definitions without an initial methodological overview, making it more difficult to follow.\n\n[1] Van de Schaft, V., Nolan, O., & Van Sloun, R. J. (2025). Off-Grid Ultrasound Imaging by Stochastic Optimization. IEEE Transactions on Ultrasonics, Ferroelectrics, and Frequency Control."},"confidence":{"value":3},"rating":{"value":3},"title":{"value":"A novel physics-informed model for ultrasound acquisition modelling"},"llm_acknowledgment":{"value":["Yes"]},"additional_comments":{"value":"- $\\Phi$ does not seem to be defined anywhere in Equations 1-3, while being listed as an optimizable parameter in Section 2: Optimization."}},"parentInvitations":"MICCAI.org/2026/Workshop/Off-Grid/-/Official_Review","nonreaders":[],"tmdate":1785502711230,"tcdate":1784316095848,"writers":["MICCAI.org/2026/Workshop/Off-Grid","MICCAI.org/2026/Workshop/Off-Grid/Submission24/Reviewer_7iFj"],"signatures":["MICCAI.org/2026/Workshop/Off-Grid/Submission24/Reviewer_7iFj"],"forum":"DwdQQ9g0Wz","number":1,"license":"CC BY 4.0","cdate":1784316095848,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/Off-Grid/Submission24/-/Official_Review","MICCAI.org/2026/Workshop/Off-Grid/-/Edit"],"mdate":1785502711230,"domain":"MICCAI.org/2026/Workshop/Off-Grid","replyto":"DwdQQ9g0Wz","id":"AxyQlggoDZ","forumContent":{"TLDR":{"value":"We estimate the strain rate from ultrasound data by inverting a physical model of acoustic scattering"},"venue":{"value":"Off-Grid 2026 Poster"},"code":{"value":"https://anonymous.4open.science/r/strain-off-grid-CC28"},"keywords":{"value":["Ultrasound","Strain Rate Estimation","Cardiac","Model-Based"]},"formatting_agreement":{"value":["Yes"]},"abstract":{"value":"Ultrasound Strain rate estimation provides a quantitative measure of cardiac muscle function, which makes it an important tool for the diagnosis and monitoring of heart disease. The most common approach for strain rate estimation, speckle tracking through block matching, assumes the matching block undergoes a rigid translation between frames, which is not true in strained tissue. We present a method for estimating the strain rate by tracking off-grid point scatterers between transmissions in the channel data. Our method attempts to reconstruct the observed signals as a sum of point sources, akin to Gaussian Splatting. We evaluate our method on 2D simulated phantom data and show we obtain lower error in the estimated strain rate maps and are better able to detect an ischemic region in the myocardium. We also show that our method is able to recover strain rate maps from in-vivo data. All code is available at https://anonymous.4open.science/r/strain-off-grid-CC28."},"_bibtex":{"value":"@inproceedings{\nschaft2026strain,\ntitle={Strain Rate Estimation from Ultrasound Channel Data},\nauthor={Vincent van de Schaft and Ruud van Sloun},\nbooktitle={Off-Grid: 1st Workshop on Continuous Representations and Grid-Free Methods in Medical Imaging},\nyear={2026},\nurl={https://openreview.net/forum?id=DwdQQ9g0Wz}\n}"},"title":{"value":"Strain Rate Estimation from Ultrasound Channel Data"},"originality_agreement":{"value":["Yes"]},"pdf":{"value":"/pdf/8a34603ce79ef050c65d02ee5559fdaac052689d.pdf"},"supplementary_material_agreement":{"value":["Yes"]},"venueid":{"value":"MICCAI.org/2026/Workshop/Off-Grid"},"paperhash":{"value":"schaft|strain_rate_estimation_from_ultrasound_channel_data"},"authorids":{"value":["~Vincent_van_de_Schaft1","~Ruud_van_Sloun2"]},"authors":{"value":["Vincent van de Schaft","Ruud van Sloun"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a method to generate synthetic long documents for model training. It aims to solve the problem of standard method that may concatenate unrelated short document to form a long document, and the similarity-based methods that concatenate too similar documents leading to the lack of diversity. The proposed method - Quest - selects the documents to be concatenated by the common keywords they would have in the generated queries to which they may answer. This is intended to concatenate documents that may be on a common topic (keywords), without being too similar.\nThe experiments on several datasets show that the synthetic data generated by the proposed method can lead to some improvements in the test tasks (on average)."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"Have you tried simpler methods to select documents to be concatenated, e.g. based on controlled similarity measure, so that the documents are within some range of similarity? Would this achieve the same goal as the proposed method?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The problem observed in the previous methods to generate synthetic data is interesting: the documents to be concatenated should be similar to some extent, but not too similar.\n2. The experiments show that the method can result in improvement on some tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper proposes a quite complex algorithm using doc2query, keyword indexing, then sampling through keywords. The paper does not motivate why these steps are necessary. One could imagine that much simpler methods would be able to lead to synthetic data that are to some extent similar, as is required by the authors. For example, if it is observed that the method based on KNN lead to concatenating too similar documents, it would be possible to select documents to be concatenated that are similar to some extent, but not the most similar. Could this simpler method lead to similar results?\n2. The paper tends to over-claim the advantage of the proposed method. For example, it is said \"Table 1 compares the Longbench results, showing that Quest consistently outperforms other methods across model sizes, ...\". Looking at Table 1, one can see that the proposed method outperforms the others on average, but underperform them on several datasets. The advantage of the method is not so clear. This over-claim appear in other observations as well (e.g. Table 3 shows that Pythia obtains the best performance on 4 datasets, while Quest only on 2). Globally, if one consider all the datasets, the demonstration that the proposed method is better is very weak.\n3. Some of the figures are unclear. Fig. 1 is not well explained. It intends to show the perfect performance of Quest on a specific task. However, the figure does not provide very useful information in addition to say it is perfect. One would also like to see some other methods in comparison, and to understand better the task itself. In Fig. 2, the dotted lines are not explained. One can guess later that they correspond to the performance of the model in some task.\n4. Fig. 5 compares the results using the documents of some length in the original dataset, and the synthetic documents. It shows that the latter perform slightly better. Do these data have equivalent sizes? If the synthetic data are much more than the real subset of data, the observation would not be surprising. However, it would be difficult to conclude that the synthetic data are better than the real data, only they are more."}},"nonreaders":[],"tmdate":1731427948805,"tcdate":1730785638926,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6964/Reviewer_ki5C"],"signatures":["ICLR.cc/2025/Conference/Submission6964/Reviewer_ki5C"],"forum":"sAYnDWaGd5","number":4,"license":"CC BY 4.0","cdate":1730785638926,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6964/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427948805,"domain":"ICLR.cc/2025/Conference","replyto":"sAYnDWaGd5","id":"aBGrcYocr2","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["longcontext","pre-training","scaling"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in large language models (LLMs) have highlighted the importance of extending context lengths for handling complex tasks. While traditional methods for training on long contexts often use filtered long documents, these approaches lead to domain imbalances, limiting model performance. To address this, techniques like random document concatenation (Standard) and similarity-based methods (KNN, ICLM) have been developed. However, they either sacrifice semantic coherence or diversity. To balance both aspects, we introduce Quest, a query-centric data synthesis method aggregating semantically relevant yet diverse documents. Quest uses a generative model to predict potential queries for each document, grouping documents with similar queries and keywords. Extensive experiments demonstrate Quest's superior performance on long-context tasks, achieving remarkable results with context lengths of up to 1M tokens and confirming its scalability across various model sizes."},"_bibtex":{"value":"@inproceedings{\ngao2025quest,\ntitle={Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model},\nauthor={Chaochen Gao and Xing W and Qi Fu and Songlin Hu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=sAYnDWaGd5}\n}"},"title":{"value":"Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model"},"pdf":{"value":"/pdf/2756cbc55bcf8b00533b450650cce52020323ce5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"gao|quest_querycentric_data_synthesis_approach_for_longcontext_scaling_of_large_language_model"},"authorids":{"value":["~Chaochen_Gao1","~Xing_W1","~Qi_Fu2","~Songlin_Hu2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chaochen Gao","Xing W","Qi Fu","Songlin Hu"]}},"version":2},{"content":{"summary":{"value":"To overcome the limitation that the standard formulation of Vision Transformers (ViTs) uses one-dimensional positional encodings which disrupt the intrinsic two-dimensional spatial geometry of images, the paper proposes Weierstrass Positional Encoding for Vision Transformers (WePE), a mathematically principled method that preserves the intrinsic spatial structure of images through complex-domain mapping based on the Weierstrass elliptic function. By encoding 2D coordinates on the complex plane and exploiting the doubly periodic properties of the function, WePE provides a continuous, resolution-agnostic, and geometrically consistent positional representation. Extensive experiments demonstrate that this approach consistently improves performance over conventional positional encodings while introducing negligible computational or memory overhead."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Boundary Conditions & Aspect Ratios. How does the method behave for non-square inputs and unusual aspect ratios? Provide results where height/width change substantially and discuss any need to retune lattice parameters.\n\nFailure Cases & Diagnostics. The paper would benefit from explicit failure analyses (qualitative heatmaps, attention distance histograms, error vs. spatial separation). Where does the method underperform, and why?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"originality:\n This paper proposes a novel definition and solution to address the long-standing challenge of positional encoding in Vision Transformers (ViTs). By introducing a Weierstrass elliptic positional encoding (WePE) framework grounded in complex analysis, it provides a fresh, mathematically principled perspective on preserving 2D spatial geometry. Unlike heuristic sinusoidal or rotary encodings, WePE employs the doubly periodic Weierstrass elliptic function and its derivative to construct a compact, continuous, and resolution-invariant spatial representation. \nsignificance:\n The positional encoding problem in ViTs is fundamental to numerous vision tasks—ranging from visual grounding and visual reasoning to detection and segmentation. The proposed WePE framework provides a general, plug-and-play solution with provable geometric properties, distance-decay behavior, and strong empirical gains across benchmarks. It could potentially broadly influence both theoretical and applied research communities, setting a new direction for geometry-aware representation learning in vision transformers."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper is technically rich but somewhat difficult to follow, particularly for readers who are not deeply familiar with complex analysis or elliptic function theory. While the mathematical rigor is commendable, the exposition occasionally prioritizes formal derivations over intuitive explanations. As a result, the connection between the underlying mathematical properties (e.g., periodicity, addition law, distance decay) and their concrete implications for Vision Transformer performance is not always clearly articulated."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931387115,"tcdate":1761985800126,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19480/Reviewer_DryE"],"signatures":["ICLR.cc/2026/Conference/Submission19480/Reviewer_DryE"],"forum":"lVmnN8g6lW","number":3,"license":"CC BY 4.0","cdate":1761985800126,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19480/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931387115,"domain":"ICLR.cc/2026/Conference","replyto":"lVmnN8g6lW","id":"744aBW260t","forumContent":{"TLDR":{"value":"We propose WePE for ViTs. By exploiting the doubly periodicity of the Weierstrass elliptic function, it avoids flattening 2D images into 1D sequences and preserves the natural geometric inductive bias disrupted by conventional encodings."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Weierstrass Elliptic Function","Positional Encoding","Vision Transformers","Double periodicity"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Vision Transformers (ViTs) have demonstrated remarkable success in computer vision tasks.  However, their reliance on learnable one-dimensional positional encoding  disrupts the inherent two-dimensional spatial structure of images due to patch flattening. Existing positional encoding approaches lack geometric constraints and fail to preserve a monotonic correspondence between Euclidean spatial distances and sequential index distances, thereby limiting the model's capacity to leverage spatial proximity priors effectively. Recognizing that periodicity is particularly beneficial for positional encoding, we propose Weierstrass elliptic Positional Encoding (WePE), a mathematically principled approach that encodes two-dimensional coordinates in the complex domain. This method maps the normalized two-dimensional patch coordinates onto the complex plane and constructs a compact four-dimensional positional feature based on the Weierstrass elliptic function $\\wp(z)$ and its derivative. The doubly periodic property of $\\wp(z)$ enables a principled encoding of 2D positional information, while their intrinsic lattice structure aligns naturally with the geometric regularities of patch grids in images. Their nonlinear geometric characteristics enable faithful modeling of spatial distance relationships,  while the associated algebraic addition formula allows relative positional information between arbitrary patch pairs to be derived directly from their absolute encodings. WePE is a plug-and-play, resolution-agnostic positional module that integrates seamlessly with existing ViTs. Extensive experiments demonstrate that WePE delivers consistent performance gains in most scenarios, while its implementation with precomputed lookup tables ensures that these improvements incur no noticeable computational or memory overhead. In addition, several analyses and ablation studies bring further confirmation to the effectiveness of our method."},"_bibtex":{"value":"@misc{\nanonymous2026weierstrass,\ntitle={Weierstrass Positional Encoding for Vision Transformers},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=lVmnN8g6lW}\n}"},"title":{"value":"Weierstrass Positional Encoding for Vision Transformers"},"pdf":{"value":"/pdf/9499698f8b41d25eb5f7c9a9056a754d6d5f82be.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"xin|weierstrass_positional_encoding_for_vision_transformers"},"authorids":{"value":["~Zhihang_Xin1","~Rui_Wang14","~Xitong_Hu1","~Xiaojun_Wu2"]},"authors":{"value":["Zhihang Xin","Rui Wang","Xitong Hu","Xiaojun Wu"]}},"version":2},{"content":{"venue":{"value":"IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/4609443/10330207/10415179.pdf"},"venueid":{"value":"dblp.org/journals/STAEORS/2024"},"paperhash":{"value":"liu|c^2n^2_complexvalued_contourlet_neural_network"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Mengkun_Liu:","~Licheng_Jiao2","~Xu_Liu5","~Lingling_Li1","https://dblp.org/search/pid/api?q=author:Fang_Liu_0001:","https://dblp.org/search/pid/api?q=author:Shuyuan_Yang:","https://dblp.org/search/pid/api?q=author:Yuwei_Guo_0001:","https://dblp.org/search/pid/api?q=author:Puhua_Chen:"]},"html":{"value":"https://doi.org/10.1109/JSTARS.2024.3358846"},"_bibtex":{"value":"@article{DBLP:journals/staeors/LiuJLLLYGC24,\n  author={Mengkun Liu and Licheng Jiao and Xu Liu and Lingling Li and Fang Liu and Shuyuan Yang and Yuwei Guo and Puhua Chen},\n  title={$C^{2}N^{2}$: Complex-Valued Contourlet Neural Network},\n  year={2024},\n  cdate={1704067200000},\n  journal={IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.},\n  volume={17},\n  pages={4478-4491},\n  url={https://doi.org/10.1109/JSTARS.2024.3358846}\n}\n"},"abstract":{"value":"Complex-valued convolutional neural networks (CV-CNN) have recently gained recognition in feature representation learning. It implements the repeated application of the operations in convolution, local average pooling, and the absolute value of the resulting vectors. However, it is only conducted in the complex spatial domain, and lacks effective representation of directionality, singularity, and regularity in the complex spectral domain for anomaly detection of images. This is the key to feature learning representation of high-order singularity. To solve this problem, a complex-valued contourlet neural network (C $^{2}$ N $^{2}$ ) is proposed in this article. It is novel in this sense that, different from the CV-CNN in the spatial domain, the spectral stream of C $^{2}$ N $^{2}$ can enhance the multiresolution sparse representation of nonsubsampled contourlet (NSCT) with multiscales and multidirections for images. Furthermore, the spectral feature integration module is proposed to capture the statistical properties of the NSCT coefficients. It is shown that the proposed network can improve the distinguishability of feature learning and classification ability in theoretical analysis and experiments on three benchmark datasets (Flevoland, Xi'an, and Germany) compared with developed methods. Polarimetric synthetic aperture radar image classification is widely used in the fields of agriculture, forestry, and military. It must be emphasized that there is potential in effective feature learning representation and the generalization capability of C $^{2}$ N $^{2}$ in deep learning, recognition, and interpretation."},"title":{"value":"$C^{2}N^{2}$: Complex-Valued Contourlet Neural Network"},"authors":{"value":["Mengkun Liu","Licheng Jiao","Xu Liu","Lingling Li","Fang Liu","Shuyuan Yang","Yuwei Guo","Puhua Chen"]}},"tmdate":1774445948567,"pdate":1704067200000,"tcdate":1727748386092,"writers":["~"],"signatures":["~Licheng_Jiao2"],"forum":"HclqVMAAYV","license":"CC BY-SA 4.0","number":126148,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1774445948567,"domain":"DBLP.org","id":"HclqVMAAYV","version":2},{"content":{"summary":{"value":"Existing single-models fail to solve Physics Olympiad tasks because it requires both complex reasoning and multimodal understanding. In this paper, the authors propose PHYSICSMINIONS, a framework of three connected studios: Visual, Logic, and Review, which refine solutions through repeated feedback. This coevolutionary process unites structured perception, reasoning, and verification, enabling the first open-source gold medal and near human-level results in the latest International Physics Olympiad."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"In this paper, the authors mention that the dual-verifier design assumes that consistent verification implies correctness. I'm curious what mechanisms exist to detect or prevent \"false consensus\", where both verifiers reinforce the same error?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper clearly identifies the limitations of the single model architecture in Physics Olympiad tasks (fail to handle both complex reasoning and multimodal understanding).\n2. They evaluate on a real and challenging benchmark to ensure fair and representative evaluation and significantly improve the performance (i.e. approaching human-expert level).\n3. The proposed system includes three components and delegate the tasks. This coevolutionary multi-agent system allows agents to cooperate, reflect, and verify each other's outputs, overcoming the limits of single model reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. I think it would be good to report the computation cost since the coevolution loop likely adds some iterations for the communications. I want to see a comparison between the single model and multi agent system.\n2. If I understand correctly, all agents use the same LLM, I think the system lacks heterogeneity and diversity of reasoning. LLM from different model family could have different aspect of reasoning, visualization abilities. What's the performance to combine different LLMs? I'm curious if this can further improve the results.\n3. Related work section does not cover the research on MLLM-based chart and figure understanding. I think the authors could add an section to talk about the visualization ability of current LLM on understanding the chart and figure (for example, CharXiv https://arxiv.org/abs/2406.18521, an evaluation benchmark)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915744717,"tcdate":1761976612444,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1350/Reviewer_KQyp"],"signatures":["ICLR.cc/2026/Conference/Submission1350/Reviewer_KQyp"],"forum":"kipQYpoZf1","number":2,"license":"CC BY 4.0","cdate":1761976612444,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1350/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915744717,"domain":"ICLR.cc/2026/Conference","replyto":"kipQYpoZf1","id":"EL9UIUNs6j","forumContent":{"TLDR":{"value":"PhysicsMinions is a coevolutionary multimodal multi-agent system that achieves gold-medal performance in latest physics Olympiads, significantly outperforming single-model baselines and advancing open-source models to human-expert levels."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics Olympiad","multi-agent system","coevolutionary framework","multimodal model"]},"supplementary_material":{"value":"/attachment/33ae425d5b96c718e9f850d263c500e27f92b666.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics is central to understanding and shaping the real world, and the ability to solve physics problems is a key indicator of real-world physical intelligence. Physics Olympiads, renowned as the crown of competitive physics, provide a rigorous testbed requiring complex reasoning and deep multimodal understanding, yet they remain largely underexplored in AI research. Existing approaches are predominantly single-model based, and open-source MLLMs rarely reach gold-medal-level performance. To address this gap, we propose PhysicsMinions, a coevolutionary multi-agent system for Physics Olympiad. Its architecture features three synergistic studios: a Visual Studio to interpret diagrams, a Logic Studio to formulate solutions, and a Review Studio to perform dual-stage verification. The system coevolves through an iterative refinement loop where feedback from the Review Studio continuously guides the Logic Studio, enabling the system to self-correct and converge towards the ground truth. Evaluated on the HiPhO benchmark spanning 7 latest physics Olympiads, PhysicsMinions delivers three major breakthroughs: (i) Strong generalization: it consistently improves both open-source and closed-source models of different sizes, delivering clear benefits over their single-model baselines; (ii) Historic breakthroughs: it elevates open-source models from only 1–2 to 6 gold medals across 7 Olympiads, achieving the first-ever open-source gold medal in the latest International Physics Olympiad (IPhO) under the average-score metric; and (iii) Scaling to human expert: it further advances the open-source Pass@32 score to 26.8/30 points on the latest IPhO, ranking 4$^\\text{th}$ of 406 contestants and far surpassing the top single-model score of 22.7 (ranked 22$^\\text{nd}$). Generally, PhysicsMinions offers a generalizable framework for Olympiad-level problem solving, with the potential to extend across disciplines."},"_bibtex":{"value":"@misc{\nyu2026physicsminions,\ntitle={PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System},\nauthor={Fangchen Yu and Junchi Yao and Ziyi Wang and Haiyuan Wan and Youling Huang and Bo Zhang and Shuyue Hu and Dongzhan Zhou and Ning Ding and Ganqu Cui and LEI BAI and Wanli Ouyang and Peng Ye},\nyear={2026},\nurl={https://openreview.net/forum?id=kipQYpoZf1}\n}"},"title":{"value":"PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System"},"pdf":{"value":"/pdf/9b0cafdc54384f883a43e35363ab5ceebd7ea9e5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yu|physicsminions_winning_gold_medals_in_the_latest_physics_olympiads_with_a_coevolutionary_multimodal_multiagent_system"},"authorids":{"value":["~Fangchen_Yu1","~Junchi_Yao1","~Ziyi_Wang33","~Haiyuan_Wan1","~Youling_Huang2","~Bo_Zhang17","~Shuyue_Hu1","~Dongzhan_Zhou1","~Ning_Ding5","~Ganqu_Cui1","~LEI_BAI1","~Wanli_Ouyang1","~Peng_Ye4"]},"authors":{"value":["Fangchen Yu","Junchi Yao","Ziyi Wang","Haiyuan Wan","Youling Huang","Bo Zhang","Shuyue Hu","Dongzhan Zhou","Ning Ding","Ganqu Cui","LEI BAI","Wanli Ouyang","Peng Ye"]}},"version":2},{"content":{"summary":{"value":"The paper introduces PHYSHANDI, a physics-based framework for reconstructing and simulating 3D hand-deformable object interactions from sparse-view RGB-D videos. The technical contributions include: (1) Dense hand modeling using the MANO parametric hand model. (2) Deformable object simulation via a spring-mass system, where object deformations are driven by forces from reconstructed hand motions. (3) A three-stage optimization pipeline. In experimetns, the method outperforms the state-of-the-art baseline (PhysTwin) in reconstruction accuracy, future prediction, and generalization to unseen interactions, particularly in scenarios with dense hand-object contacts."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"The main flaw is lack of experiments on hand reconstruction quality."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.\tThe novel dataset DENSEHDI focuses on dense hand-object contacts, addressing a gap in existing benchmarks. \n\n2.\tThis paper achieves the dense 3D reconstruction of both hands and deformable objects simultaneously, ensuring physical plausibility and coherence between hands and object dynamics. \n\n3.\tThe inverse physics design help improve accuracy in sparse-view settings, especially the single-view cases."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tPhysics-based simulation and optimization may require significant computational resources. Computational cost should be clarified. \n\n2.\tThis paper lacks experiments directly evaluating the accuracy of hand reconstructions. The title of the paper gives hand and object the equal position, but only evaluate objects. The claim in the Abstract “such simulation of object deformations can, in turn, refine and improve hand reconstruction via inverse physics” is not fully supported. \n\n3.\tThe quantitative results do not significantly outperform previous PhysTwin method in Table. 1, especially for  metrics such as SSIM. Why?\n\n4. It seems that the model assumes fixed hand-object contact topology within a sequence, limiting its applicability to interactions with dynamic contact changes (e.g., sliding or rolling). \n\n5. Previous paper \"InteractionFusion: Real-time Reconstruction of Hand Poses and Deformable Objects in Hand-object Interactions, ACM SIGGRAPH 2019.\" should be cited. This is \"another kind\" of hand-deformable object interactions\". In comparison with this paper, this manuscript is more like a dynamic object simulation or cloth simulation paper with a hand-refinement module, which targets refining the object dynamics at the end."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916678034,"tcdate":1761993700032,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3336/Reviewer_rNHX"],"signatures":["ICLR.cc/2026/Conference/Submission3336/Reviewer_rNHX"],"forum":"BTJmEqVUDF","number":4,"license":"CC BY 4.0","cdate":1761993700032,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3336/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916678034,"domain":"ICLR.cc/2026/Conference","replyto":"BTJmEqVUDF","id":"oCaSmdUZXO","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Interacting hand-object reconstruction"]},"supplementary_material":{"value":"/attachment/43fbd41b3de1234cb4be79a917ce3b4bc16c768c.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"While existing methods for reconstructing hand–object interactions have made impressive progress, they either focus on rigid or part-wise rigid objects—limiting their ability to model real-world objects (e.g., cloth, stuffed animals) that exhibit highly non-rigid deformations—or model deformable objects without full 3D hand reconstruction. To bridge this gap, we present PhysHandi (Physics-based Reconstruction of Hand and Deformable Object Interactions), a framework that enables full 3D reconstruction of both interacting hands and non-rigid objects. Our key idea is to physically simulate object deformations driven by forces induced from densely reconstructed 3D hand motions, ensuring that the reconstructed object dynamics are both physically plausible and coherent with the interacting hand movements. Furthermore, we demonstrate that such simulation of object deformations can, in turn, refine and improve hand reconstruction via inverse\nphysics. In experiments, PhysHandi outperforms the state-of-the-art baseline across reconstruction, future prediction, and generalization to unseen interactions."},"_bibtex":{"value":"@misc{\nlee2026physhandi,\ntitle={PhysHandi: Physics-Based Reconstruction of Hand-Deformable Object Interactions},\nauthor={Jihyun Lee and Changmin Lee and Donghwan Kim and Tae-Kyun Kim},\nyear={2026},\nurl={https://openreview.net/forum?id=BTJmEqVUDF}\n}"},"title":{"value":"PhysHandi: Physics-Based Reconstruction of Hand-Deformable Object Interactions"},"pdf":{"value":"/pdf/7813122df6fcafe4c3ac29f8ab0317c614d92d6d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"lee|physhandi_physicsbased_reconstruction_of_handdeformable_object_interactions"},"authorids":{"value":["~Jihyun_Lee3","~Changmin_Lee2","~Donghwan_Kim6","~Tae-Kyun_Kim2"]},"authors":{"value":["Jihyun Lee","Changmin Lee","Donghwan Kim","Tae-Kyun Kim"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a way for Image Generation Models to generate k steps in the future for a given input video. The challenge is getting the physics correct when simulating the next k steps. The proposed method is inspired by human mental simulation (where a trajectory is simulated in mind to predict the future state) and Chain of Thought prompting in LLMs (where the model is asked to answer step-by-step). The method works by giving the model input frames and asking the model to simulate small time-steps. Implicitly, the model de-renders, simulates the next step using transition dynamics, and re-renders to give the next frame. They find the physics of the trajectory stay consistent with a small delta t."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- The number of intuitive physics studies comparing machine learning models and humans' surprise rating to the plausibility of the scene. Are the authors planning on exploring this avenue to see if the mental simulation hypothesis still holds? \n- Have you considered testing whether Chain-of-Time generalizes beyond intuitive physics to intuitive psychology? (model vs human rating for plausibility rating of psychology scenes) \n- Have the authors explored combining their method with an external physics engine?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- shows strong results on 2d motion and gravity scene\n- There is a partial success in more complex simulations, like fluids and a bouncing ball\n- The method works at inference time and works with existing models like GPT-4o"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The mechanism is implicit (de-render, transition based on world transition matrix, rendering) and difficult to test\n- In 3D scenes, the early error seems to compound, making it difficult to simulate longer time-steps\n- Not much comparison with other existing methods (Video or World-Models) for generating physically plausible images generation \n- The generalization seems limited to very simple scenes and breaks when applied to more complex physics problems (fluid, bouncing)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942553712,"tcdate":1761880317169,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23193/Reviewer_nzrx"],"signatures":["ICLR.cc/2026/Conference/Submission23193/Reviewer_nzrx"],"forum":"f6ugrBWs3K","number":2,"license":"CC BY 4.0","cdate":1761880317169,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23193/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942553712,"domain":"ICLR.cc/2026/Conference","replyto":"f6ugrBWs3K","id":"eBIcyiJUnG","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-modal Language Models","Spatial and Temporal Perception","Image Generation","Physical Reasoning"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"We propose a novel method to improve the physical simulation ability of vision-language models. This Chain-of-Time simulation is motivated by in-context reasoning in machine learning, and mental simulation in humans. The method involves generating a series of intermediate images during a simulation. Chain of Time is used at inference time and requires no additional fine-tuning for performance benefits. We apply the Chain-of-Time method to synthetic and real-world domains, including 2-D graphics simulations and natural 3-D videos. These domains test a variety of particular physical properties, including velocity, acceleration, fluid dynamics, and conservation of momentum. We found that using Chain-of-Time simulation substantially improves the performance of state-of-the-art Image Generation Model. Beyond examining performance, we also analyze the specific states of the world simulated by an image model at each time step, which sheds light on the dynamics underlying these simulations. This analysis reveals insights that are hidden from traditional evaluations of physical reasoning, including cases where an Image Generation Model is able to simulate physical properties that unfold over time, such as velocity, gravity, and collisions domain well. Our analysis also highlights particular cases where the Image Generation Model struggles to infer particular physical parameters from input images, despite being capable of simulating relevant physical processes."},"_bibtex":{"value":"@misc{\nwang2026chain,\ntitle={Chain of Time: In-Context Physical Simulation with Image Generation Models},\nauthor={YingQiao Wang and Eric Bigelow and Boyi Li and Tomer Ullman},\nyear={2026},\nurl={https://openreview.net/forum?id=f6ugrBWs3K}\n}"},"title":{"value":"Chain of Time: In-Context Physical Simulation with Image Generation Models"},"pdf":{"value":"/pdf/5d0ef4077bf57397482b6d12a29a564f7f8e8459.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|chain_of_time_incontext_physical_simulation_with_image_generation_models"},"authorids":{"value":["~YingQiao_Wang1","~Eric_Bigelow1","~Boyi_Li1","~Tomer_Ullman1"]},"authors":{"value":["YingQiao Wang","Eric Bigelow","Boyi Li","Tomer Ullman"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios, overlooking shape complexities and assembly trajectories in industrial assemblies. We introduce AssemblyBench, a synthetic dataset of 2,789 industrial objects with multimodal instruction manuals, corresponding 3D part models, and part assembly trajectories. We also propose a transformer-based model, AssemblyDyno, which uses the instructional manual and the 3D shape of each part to jointly predict assembly order and part assembly trajectories. AssemblyDyno outperforms prior works in both assembly pose estimation and trajectory feasibility, where the latter is evaluated by our physics-based simulations."},"_bibtex":{"value":"@inproceedings{\nli2026assemblybench,\ntitle={AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects},\nauthor={Danrui Li and Jiahao Zhang and Bernhard Egger and Moitreya Chatterjee and Suhas Lohit and Tim K. Marks and Anoop Cherian},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=tcRurT2RT3}\n}"},"title":{"value":"AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Li_AssemblyBench_Physics-Aware_Assembly_of_Complex_Industrial_Objects_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"li|assemblybench_physicsaware_assembly_of_complex_industrial_objects"},"authorids":{"value":["~Danrui_Li1","~Jiahao_Zhang3","~Bernhard_Egger1","~Moitreya_Chatterjee1","~Suhas_Lohit1","~Tim_K._Marks1","~Anoop_Cherian1"]},"authors":{"value":["Danrui Li","Jiahao Zhang","Bernhard Egger","Moitreya Chatterjee","Suhas Lohit","Tim K. Marks","Anoop Cherian"]}},"tmdate":1789656565885,"pdate":1789656437770,"tcdate":1765220277433,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission31178/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission31178/Authors"],"forum":"tcRurT2RT3","license":"CC BY 4.0","number":31178,"cdate":1765220277433,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission31178/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission31178/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656565885,"odate":1789656437770,"domain":"thecvf.com/CVPR/2026/Conference","id":"tcRurT2RT3","version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2605.12845v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"li|assemblybench_physicsaware_assembly_of_complex_industrial_objects"},"html":{"value":"https://doi.org/10.48550/arXiv.2605.12845"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2605-12845,\n  publtype={informal},\n  author={Danrui Li and Jiahao Zhang and Bernhard Egger and Moitreya Chatterjee and Suhas Lohit and Tim K. Marks and Anoop Cherian},\n  title={AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects},\n  year={2026},\n  month={May},\n  cdate={1777593600000},\n  journal={CoRR},\n  volume={abs/2605.12845},\n  url={https://doi.org/10.48550/arXiv.2605.12845}\n}\n"},"abstract":{"value":"Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios, overlooking shape complexities and assembly trajectories in industrial assemblies. We introduce AssemblyBench, a synthetic dataset of 2,789 industrial objects with multimodal instruction manuals, corresponding 3D part models, and part assembly trajectories. We also propose a transformer-based model, AssemblyDyno, which uses the instructional manual and the 3D shape of each part to jointly predict assembly order and part assembly trajectories. AssemblyDyno outperforms prior works in both assembly pose estimation and trajectory feasibility, where the latter is evaluated by our physics-based simulations."},"title":{"value":"AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects"},"authors":{"value":[{"fullname":"Danrui Li","username":"~Danrui_Li1"},{"fullname":"Jiahao Zhang","username":""},{"fullname":"Bernhard Egger","username":""},{"fullname":"Moitreya Chatterjee","username":""},{"fullname":"Suhas Lohit","username":""},{"fullname":"Tim K. Marks","username":""},{"fullname":"Anoop Cherian","username":""}]}},"tmdate":1784392326622,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2605-12845"],"tcdate":1784392322732,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Danrui_Li1"],"forum":"XKRQaf9r1V","license":"CC BY-SA 4.0","number":63453,"cdate":1777593600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784392326622,"domain":"OpenReview.net/Public_Article","id":"XKRQaf9r1V","version":2},{"content":{"summary":{"value":"LOCA proposes an agent-based workflow for quality filtering of scientific QA data (in practice only physics QA is used). The method consists of three main components:\n1. Logic-chain enhancement: Complete the implicit reasoning steps in the original answer and decompose each step into a pair: (Principle, Derivation).\n2. Iterative review: Two reviewers (agents) independently verify the principle and derivation of each step. A QA pair is accepted if it passes three consecutive review rounds; conversely, if it accumulates five failures, it is discarded.\n3. Final check: The enhanced final answer must remain identical to the original answer.\nExperiments demonstrate a significant reduction in residual error rate on multiple physics QA datasets."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Consider validating whether the cleaned dataset further improves downstream model training quality.\n- Consider evaluating variants that incorporate multiple distinct models into the enhancement/review loop to avoid single-model bias."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The idea of augmenting hidden reasoning chains is novel and interesting.\n- On three physics QA datasets the method shows strong empirical improvement.\n- The (Principle, Derivation) pairing naturally matches reasoning patterns in scientific problems and provides a clear structure for human auditing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although the title claims to target scientific corpora, the evaluation is limited to physics, leaving generalization to other domains unverified.\n- Using a single LLM for both enhancement and reviewing may introduce self-bias and may partially explain the advantage observed under evaluation with Gemini 2.5 Pro.\n- The contribution is more of a engineering design rather than algorithmic innovation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925122693,"tcdate":1761289049442,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14769/Reviewer_dSxn"],"signatures":["ICLR.cc/2026/Conference/Submission14769/Reviewer_dSxn"],"forum":"kdFjucrq7B","number":1,"license":"CC BY 4.0","cdate":1761289049442,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14769/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925122693,"domain":"ICLR.cc/2026/Conference","replyto":"kdFjucrq7B","id":"uXkCp1SaVN","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["scientific corpus cleaning","logical chain","AI for science","LLMs"]},"supplementary_material":{"value":"/attachment/9bd9e56c7096711b6b2d9ada26ff81d426ddc305.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"While Large Language Models (LLMs) excel in general domains, their reliability often falls short in scientific problem-solving. The advancement of scientific AI depends on large-scale, high-quality corpora. However, existing scientific question-answering (QA) datasets suffer from high error rates, frequently resulting from logical leaps and implicit reasoning within the answers. To address this issue, we introduce LOCA (Logical Chain Augmentation), a novel framework for automatically cleaning scientific corpora, implemented through an augment-and-review loop. At its core, LOCA enhances raw answers by completing missing logical steps and explicitly separating the underlying scientific principle from its subsequent derivation. By applying LOCA to challenging scientific corpora, we demonstrate that it can automatically filter noisy datasets, typically reducing the error rate from as high as 20\\% to below 2\\%. LOCA provides a scalable and effective methodology for creating high-quality scientific corpora, paving the way for more reliable training and evaluation of scientific AI."},"_bibtex":{"value":"@misc{\nfang2026loca,\ntitle={{LOCA}: Logical Chain Augmentation for Scientific Corpus Cleaning},\nauthor={Youle Fang and Dong-Shan Jian and Xiang Li and Ce Meng and Lingshi Meng and Chen-Xu Yan and Zhizhang Bian and Yan-Qing Ma},\nyear={2026},\nurl={https://openreview.net/forum?id=kdFjucrq7B}\n}"},"title":{"value":"LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning"},"pdf":{"value":"/pdf/6adf0ad32849f9012988103f077cd7de25b8d589.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"fang|loca_logical_chain_augmentation_for_scientific_corpus_cleaning"},"authorids":{"value":["~Youle_Fang1","~Dong-Shan_Jian1","~Xiang_Li161","~Ce_Meng1","~Lingshi_Meng3","~Chen-Xu_Yan1","~Zhizhang_Bian1","~Yan-Qing_Ma1"]},"authors":{"value":["Youle Fang","Dong-Shan Jian","Xiang Li","Ce Meng","Lingshi Meng","Chen-Xu Yan","Zhizhang Bian","Yan-Qing Ma"]}},"version":2},{"content":{"summary":{"value":"The paper proposes REPST, a novel framework for spatio-temporal forecasting that leverages Pre-trained Language Models (PLMs), traditionally used for text, by adapting them for numerical time series analysis. Recognizing the limitations of PLMs in modeling complex spatio-temporal correlations, REPST introduces two key components: a physics-aware decomposer that breaks down spatially correlated time series into interpretable sub-components, enhancing PLM understanding through a divide-and-conquer strategy; and a selective discrete reprogramming scheme that expands the spatio-temporal vocabulary, minimizing information loss and enriching PLM representations. Experiments on real-world datasets show that REPST outperforms twelve existing methods, demonstrating strong performance, especially in data-scarce settings, and unlocking the potential of PLMs for spatio-temporal tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* Is the physics-aware component primarily derived from DMD modes? If so, how does this differ from other decomposition methods like PCA or eigenvectors, which can also capture patterns in data?\n\n* How does DMD, which primarily captures temporal embeddings, contribute to enhancing spatial information? I’m still unclear about how DMD facilitates better spatial representation in the context of spatio-temporal dynamics.\n\n* Could autoencoders, which also offer non-linear embeddings, serve as an alternative to the DMD for capturing dynamic information? Would such embeddings also be considered physics-aware in this context as you also have some augmentation from the data?\n\n* How exactly does the authors’ approach to “reprogramming” the PLM differ from simply changing input structures? Were there any modifications to the PLM architecture or retraining steps involved?\n\n* Have the baseline models been tested with the same augmented data as REPST? If not, how would the results compare under such conditions?\n\n* How do the authors substantiate their claim of improved reasoning capabilities in the model, especially in terms of spatial reasoning? Is there specific evidence beyond improved zero-shot performance?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"* I think the authors’ approach to leveraging PLMs for spatio-temporal forecasting is innovative, especially considering the usual challenges these models face with numerical time series. By adapting PLMs for spatio-temporal data, they explore an intriguing direction that could have broader implications for forecasting tasks in various fields.\n\n* The proposed REPST framework’s integration of the physics-aware decomposer and selective discrete reprogramming is a creative attempt to enrich PLM comprehension of complex spatio-temporal patterns. This combination appears to facilitate a more structured understanding of the input data, which is promising for zero-shot and few-shot learning.\n\n* The forecasting results presented in Table 1 are interesting, showing that REPST outperforms state-of-the-art baselines, especially in data-scarce scenarios. This suggests that the framework could be beneficial in practical applications where limited data is a significant constraint."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* I’m not entirely convinced by the claim of “physics-aware” decomposition. The use of the Dynamic Mode Decomposition (DMD) model seems more like pattern extraction from the input data X rather than incorporating actual physical laws. DMD is primarily designed for linear systems and is mainly focused on capturing patterns. If the authors are labeling it as “physics-aware,” I wonder why they didn’t opt for eigenvectors or simpler methods like PCA, which could also provide interpretable components. This makes me question whether the use of DMD here truly justifies the physics-informed label. Please clarify your reasoning behind using DMD over other methods like PCA or eigenvectors. Additionally, please provide more evidence or examples of how your decomposition method incorporates physical principles beyond pattern extraction.\n\n* The terminology of “reprogramming” feels overstated to me. Typically, reprogramming would imply substantial changes to the PLM’s architecture or layers. Based on Figure 2, however, it seems that the base architecture of ChatGPT2 is not significantly modified, apart from the input transformation to align with spatio-temporal dynamics. I would appreciate a clearer explanation of what changes were made to the PLM, and whether these modifications involved retraining or fine-tuning beyond input alignment. Please provide a more detailed explanation of the changes made to the PLM architecture, if any, and clarify whether any retraining or fine-tuning was involved beyond input alignment.\n\n* The results in Table 1 are intriguing, but I wonder if there is an inconsistency in how the baseline models were evaluated. It appears that REPST benefits from augmented data generated through the DMD process, while the baselines might have used only the original data. I believe a fair comparison would require running the baselines on the same augmented data to better understand the performance gap. Please clarify whether the baseline models were evaluated using the same augmented data as REPST, and if not, provide results of baselines run on the augmented data for a fair comparison.\n\n* I find the claims about reasoning and generalizability improvements through PLM usage to be a bit unclear. The authors emphasize generalization, demonstrated by zero-shot performance improvements, but I don’t see a convincing demonstration of enhanced reasoning capabilities, particularly in spatial dimensions. If the improved reasoning is attributed to the DMD-based decomposition, it seems weak, as DMD primarily enhances temporal embedding rather than spatial reasoning. Please provide specific examples or analyses that demonstrate enhanced reasoning capabilities, particularly in spatial dimensions. Also, please clarify how DMD contributes to spatial reasoning, if at all."}},"nonreaders":[],"tmdate":1731427308829,"tcdate":1730144449196,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission732/Reviewer_Xa2y"],"signatures":["ICLR.cc/2025/Conference/Submission732/Reviewer_Xa2y"],"forum":"wCNuEA5MSv","number":2,"license":"CC BY 4.0","cdate":1730144449196,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission732/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427308829,"domain":"ICLR.cc/2025/Conference","replyto":"wCNuEA5MSv","id":"30QXO60fIb","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["spatio-temporal forecasting","time series forecasting"]},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Spatio-temporal forecasting is pivotal in numerous real-world applications, including transportation planning, energy management, and climate monitoring. \nIn this work, we aim to harness the reasoning and generalization abilities of Pre-trained Language Models (PLMs) for more effective spatio-temporal forecasting, particularly in data-scarce scenarios. \nHowever, recent studies uncover that PLMs, which are primarily trained on textual data, often falter when tasked with modeling the intricate correlations inherent in numerical time series, thereby limiting their effectiveness in comprehending spatio-temporal data.\nTo bridge the gap, we propose REPST, a physics-aware PLM reprogramming framework tailored for spatio-temporal forecasting. \nSpecifically, we first propose a physics-aware decomposer that adaptively disentangles spatially correlated time series into interpretable sub-components, which facilitates PLM’s understanding of sophisticated spatio-temporal dynamics via a divide-and-conquer strategy.\nMoreover, we propose a selective discrete reprogramming scheme, which introduces an expanded spatio-temporal vocabulary space to project spatio-temporal series into discrete representations. This scheme minimizes the information loss during reprogramming and enriches the representations derived by PLMs.\nExtensive experiments on real-world datasets show that the proposed REPST outperforms twelve state-of-the-art baseline methods, particularly in data-scarce scenarios, highlighting the effectiveness and superior generalization capabilities of PLMs for spatio-temporal forecasting."},"_bibtex":{"value":"@misc{\nwang2025language,\ntitle={Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming},\nauthor={Hao Wang and Jindong Han and Wei Fan and Hao Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=wCNuEA5MSv}\n}"},"title":{"value":"Language Model Empowered Spatio-Temporal Forecasting via Physics-Aware Reprogramming"},"pdf":{"value":"/pdf/cb7d921d45bfc48169f16c8f000ca68be9a42b3d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|language_model_empowered_spatiotemporal_forecasting_via_physicsaware_reprogramming"},"authorids":{"value":["~Hao_Wang92","~Jindong_Han1","~Wei_Fan6","~Hao_Liu17"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hao Wang","Jindong Han","Wei Fan","Hao Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes DGNet, a discrete Green network framework for solving spatiotemporal PDEs with a particular focus on explicit modeling of source terms.\nUnlike conventional neural PDE solvers that implicitly mix the source and system state, DGNet explicitly decouples the system evolution from the source response, inspired by the Green’s function formalism.\nThe model constructs a graph-based discrete operator that preserves the superposition principle. The operator $\\mathcal{L}$ is composed of a physics-based component $\\mathcal{L}{\\text{physics}}$ and a neural correction component $\\mathcal{L}{\\text{NN}}$, where the former computes spatial derivatives such as gradients and Laplacians numerically, and the latter leverages a message passing neural network (MPNN) to correct approximation errors.\nDGNet further incorporates Crank–Nicolson time integration to ensure stability in temporal evolution. Across multiple PDE benchmarks, including irregular mesh scenarios and novel source terms, DGNet achieves state-of-the-art accuracy and demonstrates strong robustness compared to existing operator-learning methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- In the ablation study, does “w/o $\\mathcal{L}_{\\text{NN}}$” correspond to solving the PDE purely via numerical methods (i.e., without any learned correction)?\n- How exactly is $\\mathcal{L}_{\\text{physics}}$ computed in practice—are the gradients and Laplacians incorporated as feature inputs to the MPNN, or directly as numerical differential operators?\n- The paper emphasizes strong generalization to unseen source terms, but given the small test sizes (e.g., 3–20), can the authors provide additional experiments or statistical evidence to support this claim?\n- How does DGNet compare to PHYMPGN (ICLR 2025) under identical experimental settings? The differences in reported performance seem unexpectedly large.\n- Could larger-scale experiments (more trajectories or longer temporal horizons) be conducted to further validate generalization?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper addresses a longstanding limitation in neural PDE solvers—their inability to generalize to unseen source terms—by explicitly incorporating the Green’s function concept into a learnable, graph-based framework.\nThis formulation is conceptually elegant: by separating the effect of the source term from the system dynamics, DGNet captures the response structure in a principled and interpretable manner.\nThe combination of physics-based discretization and neural correction strikes a strong balance between numerical fidelity and data-driven flexibility. In particular, the hybrid operator $\\mathcal{L} = \\mathcal{L}{\\text{physics}} + \\mathcal{L}{\\text{NN}}$ effectively merges computational stability with adaptability, while maintaining computational efficiency.\nEmpirical results show exceptionally high performance—often outperforming existing baselines by large margins—and stability on unseen forcing terms.\nOverall, DGNet’s design provides a conceptually grounded and practically effective approach for learning PDE dynamics on irregular meshes and under novel source conditions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Despite its strong results, the paper suffers from several clarity and interpretability issues that limit its accessibility.\nFirst, the presentation of the operator $\\mathcal{L}$ and its components lacks precision. Although equations define $\\mathcal{L}{\\text{physics}}$ and $\\mathcal{L}{\\text{NN}}$, it remains unclear how $\\mathcal{L}{\\text{physics}}$ is numerically computed and integrated into the overall update rule—whether gradients and Laplacians are used merely as input features or as direct numerical operators.\nSimilarly, the interaction between $\\mathcal{L}{\\text{NN}}$ (the MPNN correction) and the source term $f$ is only superficially discussed, leaving readers uncertain about how the model actually combines physical and learned dynamics to advance the PDE state $u$.\nAlthough the abstract emphasizes source-term generalization, the main text provides limited explanation of how the method ensures this capability beyond architectural intuition.\nSecond, the neural component is relatively minimal. Apart from the residual MPNN correction, the rest of the solver heavily relies on standard numerical discretizations. While this hybrid structure is defensible, the role of the NN component could be elaborated to justify the “learning” aspect of the framework.\nThird, the evaluation protocol raises concerns about robustness. Reported test sets are extremely small (3–20 samples), making it difficult to confidently assess generalization claims. The performance gains may partially stem from limited sampling rather than true model generalization.\nFinally, after the numerical solution step, the exposition becomes sparse, with several mathematical derivations (Eqs. 5–7) presented without intuitive interpretation or discussion."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931072746,"tcdate":1761974877442,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19031/Reviewer_HxRT"],"signatures":["ICLR.cc/2026/Conference/Submission19031/Reviewer_HxRT"],"forum":"EJ8HnNTEAv","number":3,"license":"CC BY 4.0","cdate":1761974877442,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19031/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931072746,"domain":"ICLR.cc/2026/Conference","replyto":"EJ8HnNTEAv","id":"YYtHA4EKYG","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"A data-efficient method for learning spatiotemporal PDEs that generalizes to unseen source terms."},"keywords":{"value":["Partial Differential Equations","Data-efficient learning","Graph Neural Networks","Physics-Informed Machine Learning","Generalization"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Spatiotemporal partial differential equations (PDEs) underpin a wide range of scientific and engineering applications. Neural PDE solvers offer a promising alternative to classical numerical methods. However, existing approaches typically require large numbers of training trajectories, while high-fidelity PDE data are expensive to generate. Under limited data, their performance degrades substantially, highlighting their low data efficiency. A key reason is that PDE dynamics embody strong structural inductive biases that are not explicitly encoded in neural architectures, forcing models to learn fundamental physical structure from data. A particularly salient manifestation of this inefficiency is poor generalization to unseen source terms. In this work, we revisit Green’s function theory—a cornerstone of PDE theory—as a principled source of structural inductive bias for PDE learning. Based on this insight, we propose DGNet, a discrete Green network for data-efficient learning of spatiotemporal PDEs. The key idea is to transform the Green’s function into a graph-based discrete formulation, and embed the superposition principle into the hybrid physics–neural architecture which reduces the burden of learning physical priors from data, thereby improving sample efficiency. Across diverse spatiotemporal PDE scenarios, DGNet consistently achieves state-of-the-art accuracy using only tens of training trajectories. Moreover, it exhibits robust zero-shot generalization to unseen source terms, serving as a stress test that highlights its data-efficient structural design."},"_bibtex":{"value":"@inproceedings{\ntan2026dgnet,\ntitle={{DGN}et: Discrete Green Networks for Data-Efficient Learning of Spatiotemporal {PDE}s},\nauthor={Yingjie Tan and Quanming Yao and Yaqing Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EJ8HnNTEAv}\n}"},"title":{"value":"DGNet: Discrete Green Networks for Data-Efficient Learning of Spatiotemporal PDEs"},"pdf":{"value":"/pdf/107ff4674b22e4b0b19c7d4c071741bba67df063.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"tan|dgnet_discrete_green_networks_for_dataefficient_learning_of_spatiotemporal_pdes"},"authorids":{"value":["~Yingjie_Tan3","~Quanming_Yao3","~Yaqing_Wang2"]},"authors":{"value":["Yingjie Tan","Quanming Yao","Yaqing Wang"]}},"version":2},{"content":{"comment":{"value":"Thanks for your response.\n\nSo, I gather the physics in this work is more leaning toward intuitive physics or cognitive physics. I accept narrowing the scope, but I strongly advise making it clearer. Because when this work becomes a popular benchmark, and expose to a wider audience, they might think researchers in the AI field are not rigorous."},"title":{"value":"Reply to part 2"}},"tmdate":1700598465546,"tcdate":1700598465546,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6857/Reviewer_P1S1"],"signatures":["ICLR.cc/2024/Conference/Submission6857/Reviewer_P1S1"],"forum":"pNlntv7A9X","number":18,"license":"CC BY 4.0","cdate":1700598465546,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6857/-/Official_Comment"],"mdate":1700598465546,"domain":"ICLR.cc/2024/Conference","replyto":"5NMiRqKpgm","id":"vE8uvUomOt","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Neuro-symbolic Visual Reasoning","Physical Reasoning"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We introduce the Soft-Body Physical Dataset (SOPHY), a novel benchmark for evaluating machine models in physical reasoning across diverse scenarios for soft bodies. The SOPHY is specifically designed to be complementary with existing physical reasoning benchmarks by encompassing diverse physical property inferences for soft bodies like physical parameters such as mass and density across dynamic situations and predicting corresponding dynamics. This comprehensive dataset enables the development and assessment of AI models with human-like visual reasoning abilities in understanding both rigid objects and soft objects’ visual attributes, physical properties, and dynamics while devising goal-oriented solutions. We evaluated a range of AI models and found that they still struggle to achieve satisfactory performance, which shows that current AI models still lack physical commonsense for soft objects and illustrates the value of the proposed dataset. We hope the SOPHY fosters advancements in AI perception and reasoning in diverse physical environments, bridging the gap between human and machine intelligence in the physical world."},"_bibtex":{"value":"@misc{\nzheng2024softphy,\ntitle={SoftPhy: Soft-Body Physical Concept  Learning  and Reasoning from Videos},\nauthor={Zhicheng Zheng and Xin Yan and Zhenfang Chen and Jingzhou Wang and Qin Zhi Eddie Lim and Joshua B. Tenenbaum and Chuang Gan},\nyear={2024},\nurl={https://openreview.net/forum?id=pNlntv7A9X}\n}"},"title":{"value":"SoftPhy: Soft-Body Physical Concept  Learning  and Reasoning from Videos"},"pdf":{"value":"/pdf/e38976beee0ed6da591f0664691f8962cdefc0d3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"zheng|softphy_softbody_physical_concept_learning_and_reasoning_from_videos"},"authorids":{"value":["~Zhicheng_Zheng2","~Xin_Yan3","~Zhenfang_Chen1","~Jingzhou_Wang3","~Qin_Zhi_Eddie_Lim1","~Joshua_B._Tenenbaum1","~Chuang_Gan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhicheng Zheng","Xin Yan","Zhenfang Chen","Jingzhou Wang","Qin Zhi Eddie Lim","Joshua B. Tenenbaum","Chuang Gan"]}},"version":2},{"content":{"summary":{"value":"This paper introduces BuildArena, a physics-aligned interactive benchmark for evaluating LLMs in language-driven engineering construction. The benchmark enable models to construct and test 3D structures under physical constraints using natural language. \nThe paper evaluates eight major LLMs across physics-based construction tasks, reporting success rates, performance indicators, and token-cost analyses. Results show that while models exhibit rudimentary 3D construction skills and creative strategies, they still fail at compositional precision, hierarchical assembly, and spatial conflict resolution."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Whether the paper use specific prompt across different LLMs?\n\n- How model perform when using multi-round CoT?\n\n- How do BuildArena tasks compare to human performance under textual-only conditions?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- BuildArena pioneers the integration of language, physics simulation, and 3D assembly for LLM benchmarking.\n\n- The paper is well-structured, systematically progressing.\n\n- BuildArena establishes a foundational benchmark for evaluating LLMs’ physics-grounded reasoning and interactive 3D construction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The benchmark currently performs single-round evaluation. There is no closed-loop feedback integrating simulation outcomes into iterative model improvement or self-correction.\n\n- The 3D Spatial Geometric Computation Library, though impressive, mirrors only a subset of Besiege's physics primitives. As the authors admit, limited module diversity constrains object complexity and realism.\n\n- The evaluation metrics capture outcome quality but miss process-level score, i.e., whether models optimize design efficiency, robustness trade-offs, or use feedback adaptively.\n\n- Since tasks involve 3D spatial reasoning, a natural baseline would include VLMs."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916359678,"tcdate":1762285016314,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2752/Reviewer_6rvr"],"signatures":["ICLR.cc/2026/Conference/Submission2752/Reviewer_6rvr"],"forum":"oml3PWSYcc","number":4,"license":"CC BY 4.0","cdate":1762285016314,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2752/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916359678,"domain":"ICLR.cc/2026/Conference","replyto":"oml3PWSYcc","id":"aYSsMMXRju","forumContent":{"TLDR":{"value":"We provide BuildArena, a physics‑aligned interactive benchmark that tests the engineering construction capabilities of frontier LLMs."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Engineering construction","LLM","benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Engineering construction automation aims to transform natural language specifications into physically viable structures, requiring complex integrated reasoning under strict physical constraints. While modern LLMs possess broad knowledge and strong reasoning capabilities that make them promising candidates for this domain, their construction competencies remain largely unevaluated. To address this gap, we introduce BuildArena, the first physics-aligned interactive benchmark designed for language-driven engineering construction. It takes a first step towards engineering automation using LLMs. Technically, it contributes to the community in two aspects: (1) an extendable task design strategy spanning static and dynamic mechanics across multiple difficulty tiers; (2) a 3D Spatial Geometric Computation Library for supporting construction based on language instructions. On eight frontier LLMs, BuildArena comprehensively evaluates their capabilities for language-driven and physics-grounded construction automation. We release the code at https://anonymous.4open.science/r/BuildArena-9B7B/ to benefit construction automation in engineering applications."},"_bibtex":{"value":"@misc{\nanonymous2026buildarena,\ntitle={BuildArena: A Physics\\nobreakdash-Aligned Interactive Benchmark of {LLM}s for Engineering Construction},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=oml3PWSYcc}\n}"},"title":{"value":"BuildArena: A Physics‑Aligned Interactive Benchmark of LLMs for Engineering Construction"},"pdf":{"value":"/pdf/26c6163a776c85a6701b6693a31ffcf1c6edc130.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"xia|buildarena_a_physicsaligned_interactive_benchmark_of_llms_for_engineering_construction"},"authorids":{"value":["~Tian_Xia15","~Tianrun_Gao3","~Wenhao_Deng2","~Long_Wei1","~Xiaowei_Qian3","~Jiang_Yixian1","~Chenglei_Yu1","~Tailin_Wu1"]},"authors":{"value":["Tian Xia","Tianrun Gao","Wenhao Deng","Long Wei","Xiaowei Qian","Jiang Yixian","Chenglei Yu","Tailin Wu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces SPARK, a physics-guided augmentation framework for modeling dynamical systems that overcomes the limitations of traditional numerical and data-driven methods. By incorporating a unique compression and augmentation plugin, along with an attention mechanism and Fourier-enhanced graph ODE, SPARK improves model generalization and robustness, especially in data-scarce situations and distribution shifts. Experimental results highlight SPARK's strong performance in accurately predicting complex spatiotemporal dynamics, particularly in challenging cases like sea ice evolution, effectively capturing intricate physical phenomena."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In line 163, it needs references for those methods which simply concatenate boundary information with node features.\n2. What does boundary information refer to? Give some examples please.\n3. In abstract, what's the meaning of \"stable data distribution\"? Provide explanations about it and why does it can cause ineffectiveness of data scarcity and distribution shifts."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. By incorporating boundary information and physical parameters, SPARK enhances the model's ability to generalize across different physical scenarios, which is crucial for real-world applications.\n2. The paper provides extensive experimental results across various benchmark datasets, demonstrating SPARK's superior performance compared to existing models, particularly in handling out-of-distribution scenarios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The symbols and formulas appear to be somewhat disorganized, which makes it difficult for readers to understand the meaning.  Clear definitions and a more structured presentation of the equations would greatly enhance the paper's accessibility and overall readability.\n2. The lack of novelty. This paper claims to be the first to use physics-guided compression and augmentation. But there has been a paper [1] doing like this. The techniques of the two papers are very similar, including : (1) using VQ-VAE to compress information (2) augmenting training set by the top-K discrete embeddings.\n3. The proposed methodology, may be complex to implement in practice. The paper could provide more guidance or examples on how to effectively apply SPARK in different contexts.\n\n[1] Wu, Hao, et al. \"BeamVQ: Aligning Space-Time Forecasting Model via Self-training on Physics-aware Metrics.\" arXiv preprint arXiv:2405.17051 (2024)."}},"nonreaders":[],"tmdate":1731427676763,"tcdate":1730537189234,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2890/Reviewer_bjTD"],"signatures":["ICLR.cc/2025/Conference/Submission2890/Reviewer_bjTD"],"forum":"BZQmpsuW7D","number":2,"license":"CC BY 4.0","cdate":1730537189234,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2890/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427676763,"domain":"ICLR.cc/2025/Conference","replyto":"BZQmpsuW7D","id":"duclUCOIh7","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["dynamical system","augmentation"]},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In dynamical system modeling, traditional numerical methods have a solid theoretical foundation but are limited by high computational costs and sensitivity to initial conditions. Current data-driven approaches use deep learning models to capture complex spatiotemporal features, but they rely heavily on large amounts of data and assume a stable data distribution, making them ineffective against data scarcity and distribution shifts. To address these challenges, we propose SPARK, a physics-guided quantized augmentation plugin. SPARK integrates boundary information and physical parameters, using a reconstruction autoencoder to build a physics-rich discrete memory bank for data compression. It then enhances selected samples for downstream tasks with this pre-trained memory bank. SPARK then utilizes an attention mechanism to model historical observations and combines fourier-enhanced graph ODE to efficiently predict long-term dynamical systems, enhancing robustness and adaptability to complex physical environments. Extensive experiments on benchmark datasets show that our approach significantly outperforms various baseline methods in handling distribution shifts and data scarcity."},"_bibtex":{"value":"@misc{\nxu2025spark,\ntitle={{SPARK}: Physics-Guided Quantitative Augmentation for Dynamical System Modeling},\nauthor={Fan Xu and Penghao Zhao and Zhipeng Xu and XINLIANG ZHOU and Xinping Yi and Qingsong Wen and Hao Wu and Kun Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=BZQmpsuW7D}\n}"},"title":{"value":"SPARK: Physics-Guided Quantitative Augmentation for Dynamical System Modeling"},"pdf":{"value":"/pdf/45e2355f4b317c1e910b7d82a4f4cbd6beba1662.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"xu|spark_physicsguided_quantitative_augmentation_for_dynamical_system_modeling"},"authorids":{"value":["~Fan_Xu5","~Penghao_Zhao1","~Zhipeng_Xu1","~XINLIANG_ZHOU1","~Xinping_Yi1","~Qingsong_Wen2","~Hao_Wu39","~Kun_Wang15"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fan Xu","Penghao Zhao","Zhipeng Xu","XINLIANG ZHOU","Xinping Yi","Qingsong Wen","Hao Wu","Kun Wang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2505.01736v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"wan|pesanet_physicsencoded_spectral_attention_network_for_simulating_pdegoverned_complex_systems"},"html":{"value":"https://doi.org/10.48550/arXiv.2505.01736"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2505-01736,\n  publtype={informal},\n  author={Han Wan and Rui Zhang and Qi Wang and Yang Liu and Hao Sun},\n  title={PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems},\n  year={2025},\n  month={May},\n  cdate={1746057600000},\n  journal={CoRR},\n  volume={abs/2505.01736},\n  url={https://doi.org/10.48550/arXiv.2505.01736}\n}\n"},"abstract":{"value":"Accurately modeling and forecasting complex systems governed by partial differential equations (PDEs) is crucial in various scientific and engineering domains. However, traditional numerical methods struggle in real-world scenarios due to incomplete or unknown physical laws. Meanwhile, machine learning approaches often fail to generalize effectively when faced with scarce observational data and the challenge of capturing local and global features. To this end, we propose the Physics-encoded Spectral Attention Network (PeSANet), which integrates local and global information to forecast complex systems with limited data and incomplete physical priors. The model consists of two key components: a physics-encoded block that uses hard constraints to approximate local differential operators from limited data, and a spectral-enhanced block that captures long-range global dependencies in the frequency domain. Specifically, we introduce a novel spectral attention mechanism to model inter-spectrum relationships and learn long-range spatial features. Experimental results demonstrate that PeSANet outperforms existing methods across all metrics, particularly in long-term forecasting accuracy, providing a promising solution for simulating complex systems with limited data and incomplete physics."},"title":{"value":"PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems"},"authors":{"value":[{"fullname":"Han Wan","username":""},{"fullname":"Rui Zhang","username":""},{"fullname":"Qi Wang","username":""},{"fullname":"Yang Liu","username":"~Yang_Liu104"},{"fullname":"Hao Sun","username":""}]}},"tmdate":1784019463427,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2505-01736"],"tcdate":1784019454512,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Yang_Aron_Liu1"],"forum":"lYsmQuPm9C","license":"CC BY-SA 4.0","number":54169,"cdate":1746057600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784019463427,"domain":"OpenReview.net/Public_Article","id":"lYsmQuPm9C","version":2},{"content":{"venue":{"value":"IJCAI 2025"},"pdf":{"value":"https://www.ijcai.org/proceedings/2025/0862.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"wan|pesanet_physicsencoded_spectral_attention_network_for_simulating_pdegoverned_complex_systems"},"html":{"value":"https://doi.org/10.24963/ijcai.2025/862"},"_bibtex":{"value":"@inproceedings{DBLP:conf/ijcai/WanZWLS25,\n  author={Han Wan and Rui Zhang and Qi Wang and Yang Liu and Hao Sun},\n  title={PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems},\n  year={2025},\n  cdate={1735689600000},\n  pages={7751-7759},\n  url={https://doi.org/10.24963/ijcai.2025/862},\n  booktitle={IJCAI},\n  crossref={conf/ijcai/2025}\n}\n"},"abstract":{"value":"Accurately modeling and forecasting complex systems governed by partial differential equations (PDEs) is crucial in various scientific and engineering domains. However, traditional numerical methods struggle in real-world scenarios due to incomplete or unknown physical laws. Meanwhile, machine learning approaches often fail to generalize effectively when faced with scarce observational data and the challenge of capturing local and global features. To this end, we propose the Physics-encoded Spectral Attention Network (PeSANet), which integrates local and global information to forecast complex systems with limited data and incomplete physical priors. The model consists of two key components: a physics-encoded block that uses hard constraints to approximate local differential operators from limited data, and a spectral-enhanced block that captures long-range global dependencies in the frequency domain. Specifically, we introduce a novel spectral attention mechanism to model inter-spectrum relationships and learn long-range spatial features. Experimental results demonstrate that PeSANet outperforms existing methods across all metrics, particularly in long-term forecasting accuracy, providing a promising solution for simulating complex systems with limited data and incomplete physics."},"title":{"value":"PeSANet: Physics-encoded Spectral Attention Network for Simulating PDE-Governed Complex Systems"},"authors":{"value":[{"fullname":"Han Wan","username":""},{"fullname":"Rui Zhang","username":""},{"fullname":"Qi Wang","username":""},{"fullname":"Yang Liu","username":"~Yang_Liu104"},{"fullname":"Hao Sun","username":""}]}},"tmdate":1784019461594,"pdate":1767139200000,"externalIds":["dblp:conf/ijcai/WanZWLS25"],"tcdate":1784019454500,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Yang_Aron_Liu1"],"forum":"Z7k4WyDwkh","license":"CC BY-SA 4.0","number":54167,"cdate":1735689600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784019461594,"domain":"OpenReview.net/Public_Article","id":"Z7k4WyDwkh","version":2},{"content":{"venue":{"value":"CogSci 2023"},"pdf":{"value":"https://escholarship.org/content/qt3hq021qs/qt3hq021qs.pdf?t=rxyatc"},"venueid":{"value":"dblp.org/conf/COGSCI/2023"},"paperhash":{"value":"chen|just_in_time_representations_for_mental_simulation_in_intuitive_physics"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Tony_Chen:","~Kelsey_R._Allen1","https://dblp.org/search/pid/api?q=author:Samuel_J._Cheyette:","https://dblp.org/search/pid/api?q=author:Josh_Tenenbaum_0001:","https://dblp.org/search/pid/api?q=author:Kevin_A._Smith:"]},"html":{"value":"https://escholarship.org/uc/item/3hq021qs"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cogsci/ChenAC0S23,\n  author={Tony Chen and Kelsey R. Allen and Samuel J. Cheyette and Josh Tenenbaum and Kevin A. Smith},\n  title={\"Just In Time\" Representations for Mental Simulation in Intuitive Physics},\n  year={2023},\n  cdate={1672531200000},\n  url={https://escholarship.org/uc/item/3hq021qs},\n  booktitle={CogSci},\n  crossref={conf/cogsci/2023}\n}\n"},"abstract":{"value":"Author(s): Chen, Tony; Allen, Kelsey R; Cheyette, Samuel J.; Tenenbaum, Josh; Smith, Kevin A | Abstract: Many models of intuitive physical reasoning posit some kind of mental simulation mechanism, yet everyday environments frequently contain far more objects than people could plau- sibly represent with their limited cognitive capacity. What determines which objects are actually included in our repre- sentations? We asked participants to predict how a ball will bounce through a complex field of obstacles, and probed work- ing memory for objects in the scene that were more and less likely to be relevant to the ball’s trajectory. We evaluate differ- ent accounts of relevance and find that successful object mem- ory is best predicted by how frequently a ball’s trajectory is expected to contact that object under a probabilistic simulation model. This suggests that people construct representations for mental simulation efficiently and dynamically, on the fly, by adding objects “just in time”: only when they are expected to become relevant for the next stage of simulation."},"title":{"value":"\"Just In Time\" Representations for Mental Simulation in Intuitive Physics"},"authors":{"value":["Tony Chen","Kelsey R. Allen","Samuel J. Cheyette","Josh Tenenbaum","Kevin A. Smith"]}},"tmdate":1727914026132,"pdate":1672531200000,"tcdate":1727914009075,"writers":["~"],"signatures":["~Kelsey_R_Allen1"],"forum":"NlBVfS3lun","license":"CC BY-SA 4.0","number":137770,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727914026132,"domain":"DBLP.org","id":"NlBVfS3lun","version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["Neural Simulation","PDE Solving","Self-Supervised Pre-Training"]},"_bibtex":{"value":"@inproceedings{\nwu2026geopt,\ntitle={Geo{PT}: Scaling Physics Simulation via Lifted Geometric Pre-Training},\nauthor={Haixu Wu and Minghao Guo and Zongyi Li and Zhiyang Dou and Mingsheng Long and Kaiming He and Wojciech Matusik},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=cVO0eX8vLQ}\n}"},"title":{"value":"GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training"},"paperhash":{"value":"wu|geopt_scaling_physics_simulation_via_lifted_geometric_pretraining"},"originally_submitted_PDF":{"value":"/pdf/1cee686b2929d2462f3aa9cc5e5c310e2f31807b.pdf"},"TLDR":{"value":"This paper presents GeoPT as a unified model pre-trained from large-scale geometries for general physics simulation."},"primary_area":{"value":"deep_learning->selfsupervised_learning"},"abstract":{"value":"Neural simulators promise efficient surrogates for physics simulation, but scaling them is bottlenecked by the prohibitive cost of generating high-fidelity training data. Pre-training on abundant off-the-shelf geometries offers a natural alternative, yet faces a fundamental gap: supervision on static geometry alone ignores dynamics and can lead to negative transfer on physics tasks. We present GeoPT, a unified pre-trained model for general physics simulation based on lifted geometric pre-training. The core idea is to augment geometry with synthetic dynamics, enabling dynamics-aware self-supervision without physics labels. Pre-trained on over one million samples, GeoPT consistently improves industrial-fidelity benchmarks spanning fluid mechanics for cars, aircraft, and ships, and solid mechanics in crash simulation, reducing labeled data requirements by 20-60% and accelerating convergence by 2$\\times$. These results show that lifting with synthetic dynamics bridges the geometry-physics gap, unlocking a scalable path for neural simulation and potentially beyond. Code is available at https://github.com/Physics-Scaling/GeoPT."},"link_to_code":{"value":"https://github.com/Physics-Scaling/GeoPT"},"pdf":{"value":"/pdf/44655a3a462b8690affca72fd8e1b1291e582765.pdf"},"lay_summary":{"value":"Physics simulation is a foundation of science and engineering, but it is often very expensive to compute. Neural simulators can make predictions much faster, but building foundation models for physics usually requires large amounts of costly simulation data. This paper focuses on this missing piece in constructing a physics foundation model.\n\nWe introduce GeoPT, a pre-trained model that learns from abundant 3D geometry data instead. Since static shapes alone do not contain physical dynamics, GeoPT adds synthetic dynamics to geometry during pre-training, forming a new lifted pre-training paradigm. This helps the model learn patterns that are useful for downstream physics prediction.\n\nOn industrial-scale tasks, including fluid simulation and crash simulation, GeoPT improves accuracy, reduces the need for labeled simulation data by up to 60%, and speeds up training. This suggests a scalable path toward a general-purpose physics foundation model."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Haixu_Wu1","~Minghao_Guo1","~Zongyi_Li1","~Zhiyang_Dou1","~Mingsheng_Long5","~Kaiming_He2","~Wojciech_Matusik2"]},"authors":{"value":["Haixu Wu","Minghao Guo","Zongyi Li","Zhiyang Dou","Mingsheng Long","Kaiming He","Wojciech Matusik"]}},"tmdate":1790067273697,"pdate":1777576172025,"tcdate":1768002111271,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission2468/Authors"],"signatures":["ICML.cc/2026/Conference/Submission2468/Authors"],"forum":"cVO0eX8vLQ","license":"CC BY 4.0","number":2468,"cdate":1768002111271,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission2468/-/Full_Submission","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/Submission2468/-/Camera_Ready_Revision"],"mdate":1790067273697,"odate":1782341901997,"domain":"ICML.cc/2026/Conference","id":"cVO0eX8vLQ","version":2},{"content":{"TLDR":{"value":"This paper presents GeoPT as a unified model pre-trained from large-scale geometries for general physics simulation."},"venue":{"value":"ICLR 2026 Workshop FM4Science Poster"},"pdf":{"value":"/pdf/3df9a842bf6bf180192840eca36de1f7eda780d3.pdf"},"keywords":{"value":["Neural Simulation","PDE Solving","Self-Supervised Pre-Training"]},"venueid":{"value":"ICLR.cc/2026/Workshop/FM4Science"},"paperhash":{"value":"wu|geopt_scaling_physics_simulation_via_lifted_geometric_pretraining"},"authorids":{"value":["~Haixu_Wu1","~Minghao_Guo1","~Zongyi_Li1","~Zhiyang_Dou1","~Mingsheng_Long5","~Kaiming_He2","~Wojciech_Matusik2"]},"abstract":{"value":"Neural simulators promise efficient surrogates for physics simulation, but scaling them is bottlenecked by the prohibitive cost of generating high-fidelity training data. Pre-training on abundant off-the-shelf geometries offers a natural alternative, yet faces a fundamental gap: supervision on static geometry alone ignores dynamics and can lead to negative transfer on physics tasks. We present GeoPT, a unified pre-trained model for general physics simulation based on lifted geometric pre-training. The core idea is to augment geometry with synthetic dynamics, enabling dynamics-aware self-supervision without physics labels. Pre-trained on over one million samples, GeoPT consistently improves industrial-fidelity benchmarks spanning fluid mechanics for cars, aircraft, and ships, and solid mechanics in crash simulation, reducing labeled data requirements by 20-60% and accelerating convergence by 2$\\times$. These results show that lifting with synthetic dynamics bridges the geometry-physics gap, unlocking a scalable path for neural simulation and potentially beyond. Code is available at https://github.com/Physics-Scaling/GeoPT."},"_bibtex":{"value":"@inproceedings{\nwu2026geopt,\ntitle={Geo{PT}: Scaling Physics Simulation via Lifted Geometric Pre-Training},\nauthor={Haixu Wu and Minghao Guo and Zongyi Li and Zhiyang Dou and Mingsheng Long and Kaiming He and Wojciech Matusik},\nbooktitle={ICLR 2026 Workshop on Foundation Models for Science: Real-World Impact and Science-First Design},\nyear={2026},\nurl={https://openreview.net/forum?id=N9qIqvanBj}\n}"},"title":{"value":"GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training"},"authors":{"value":["Haixu Wu","Minghao Guo","Zongyi Li","Zhiyang Dou","Mingsheng Long","Kaiming He","Wojciech Matusik"]}},"tmdate":1777215753338,"pdate":1772546524981,"tcdate":1770495670933,"writers":["ICLR.cc/2026/Workshop/FM4Science","ICLR.cc/2026/Workshop/FM4Science/Submission77/Authors"],"signatures":["ICLR.cc/2026/Workshop/FM4Science/Submission77/Authors"],"forum":"N9qIqvanBj","license":"CC BY 4.0","number":77,"cdate":1770495670933,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/FM4Science/-/Submission","ICLR.cc/2026/Workshop/FM4Science/-/Post_Submission","ICLR.cc/2026/Workshop/FM4Science/-/Edit","ICLR.cc/2026/Workshop/FM4Science/Submission77/-/Revision"],"mdate":1777215753338,"odate":1772546524981,"domain":"ICLR.cc/2026/Workshop/FM4Science","id":"N9qIqvanBj","version":2},{"content":{"summary":{"value":"This paper introduces P-ALIGN, a framework that applies self-alignment concepts from large language models to physical dynamical system modeling. The key innovation is enabling dynamical system models to iteratively improve through self-discovery and physics-aware curation of training data. The framework aims to enhance both statistical accuracy and physical consistency of predictions and is validated by empirical results."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.\tCan you clarify the steps in self-discovery section (page 4): how the sub regions/representative vectors are calculated? equation 6 and 8 seems conflicting, is Z’_t calculated by matrix operation or projection to nearest anchor point?\n\n2.\tCan you show how the physics rewards function is defined for these specific examples?\n\n3.\tCan you show how the detailed structure of these encoder/decoders?\n\n4.\tCan you elaborate on problem formulation? Are you mapping from a single step observation to a trajectory (X_t->X_{t+1},X_{t+2..})? If this is the case, the setting is questionable. In practice, a single step might not contain all the information needed to predict the future (i.e. there could exist high-order temporal effect such as acceleration). And in real-world the past observations are available, it is almost free lunch to use that for massively improving the forecasting results. See the section 2.1 in this paper (https://proceedings.neurips.cc/paper_files/paper/2022/hash/87f476af4053961667c2c08e9f4b850e-Abstract-Conference.html) for the standard problem setting in dynamical system learning.\n\n5.\tIn line 337, the author claims ‘The filtered hypothesis space H′ is smaller’. Can you clarify why it is smaller and smaller compared to which baseline?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1.The design and adaptation of self-alignment concept from LLM to physics system is novel.\n\n2. the designed framework is architecture agnostic and can be used in many scenarios.\n\n3. the paper explains the reduced generalization error upper bound \n\n4. There shows an impressive and consistent improvement across different scenarios"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tMany of the implementation details are not clear or missing, see my question 1,2,3\n\n2.\tThe method requires prior knowledge of physics constraints, which is not necessarily available in real-world scenarios. Also, this could be an unfair comparison with vanilla baseline, since the latter could be easily improved by adding a regularization term to penalize the physics violation. \n\n3.\tMy major concern is the problem setting in this paper. For dynamical/physics system forecasting specially with these PDE like inputs, using multiple past observation to predict the future is a commonly used and very effective way to improve the forecasting. While this paper seems not using these basics from dynamical systems, therefore the experiments could be less legit. see my question 4."}},"nonreaders":[],"tmdate":1731427722495,"tcdate":1730143138594,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3054/Reviewer_BMrU"],"signatures":["ICLR.cc/2025/Conference/Submission3054/Reviewer_BMrU"],"forum":"AgTSjXh7vl","number":3,"license":"CC BY 4.0","cdate":1730143138594,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3054/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427722495,"domain":"ICLR.cc/2025/Conference","replyto":"AgTSjXh7vl","id":"BNLWVqqcd2","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["dynamic systems modeling","physical consistency."]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Deep learning has emerged as the new paradigm in modeling complex physical dynamical systems. Nevertheless, data-driven methods learn patterns by optimizing statistical metrics, tend to overlook the adherence to physical laws. Previous work have attempted to incorporate physical constraints into neural networks, but they often face limitations due to lack of flexibility or optimization challenges. In this paper, we propose a novel framework, Physics-aware Self-Alignment (P-Align), to enhance the physical consistency of dynamical systems modeling.  P-Align enables dynamical system models to provides physics-aware rewards, which makes self-alignment of dynamical system models possible. Comprehensive experiments show that \\method{} not only gave an average statistical skill score boost of more than 32% for ten backbones on five datasets, but also significantly enhances physics-aware metrics. All of our source codes will be released via GitHub."},"_bibtex":{"value":"@misc{\nxu2024palign,\ntitle={P-Align: Self-Alignment in Physical Dynamical System Modeling},\nauthor={Zhipeng Xu and Fan Xu and Hanbin Wang and XINLIANG ZHOU and Lilan Peng and Qingsong Wen and Kun Wang and Hao Wu},\nyear={2024},\nurl={https://openreview.net/forum?id=AgTSjXh7vl}\n}"},"title":{"value":"P-Align: Self-Alignment in Physical Dynamical System Modeling"},"pdf":{"value":"/pdf/981c03a0c4b975e74db34a898eaf4d2fe0e44e3d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"xu|palign_selfalignment_in_physical_dynamical_system_modeling"},"authorids":{"value":["~Zhipeng_Xu1","~Fan_Xu5","~Hanbin_Wang1","~XINLIANG_ZHOU1","~Lilan_Peng1","~Qingsong_Wen2","~Kun_Wang15","~Hao_Wu39"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhipeng Xu","Fan Xu","Hanbin Wang","XINLIANG ZHOU","Lilan Peng","Qingsong Wen","Kun Wang","Hao Wu"]}},"version":2},{"content":{"summary":{"value":"In complex systems with minimal information of the underlying physics, it is difficult to model physics based losses. To overcome this issue, the authors propose a surrogate model that learns the inverse mapping between the solution at discrete points, its derivatives and the source term at the corresponding discrete points. This model effectively serves as the “teacher model” for a neural operator framework that learns the solution from the source term. The derivatives of the solution are computed based on numerical differences. It is pseudo physics informed because the operator is trained using the surrogate model rather than loss functions and residuals defined over the actual PDE."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.\tWhen the source terms and the boundary conditions are known, the PDEs can be estimated using Monte-Carlo Walk-on-Spheres (WOS). Neural Walk-on-spheres trains neural networks based on WOS estimates. How does the error rate of the Surrogate model compare against random-walks that accumulate the source term over the green function? \n\n2.\tWhat is the justification for using convolution neural networks as the surrogate model, to capture neighborhood information? A radius based graph neural network is discretization agnostic and works especially well in sparse settings. \n\n3.\tThe surrogate model is not discretization agnostic. The functions sampled would have to be the same discretization as it was trained on. Which would mean that the neural operator model can predict any sparse distribution of points, but the second loss term (i.e. the surrogate model) has to be a fixed discretization. This seems like a bottleneck. Were there reasons for not making the second model a neural operator. Perhaps using [1] would be a good way to ensure operator learning through the entire pipeline. \n\n[1] Wang, Tian, and Chuang Wang. \"Latent Neural Operator for Solving Forward and Inverse PDE Problems.\" arXiv preprint arXiv:2406.03923 (2024)."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Originality: The use of the inverse model as the ground truth in cases where the data is sparse and the governing PDE is unknown is quite promising because in a lot of applied settings, it is not always the case that a governing PDE is known. \n\nQuality: The intuition of the paper is quite clear. There are thorough experiments on standard benchmark datasets. The authors perform several ablations to substantiate their claims. The figures clearly indicate the message the authors are trying to convey. \n\nSignificance: This is a novel idea that builds upon the Physics-informed ML literature, combining inverse-PDE estimates into the learning pipeline as an alternative to physics based residuals and losses."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"They use the neighborhood information captured within the convolution layers as a way to compensate for errors in numerical differences. A graph neural operator would be both discretization agnostic and would be better for capturing neighborhood information. \n\nThe training dataset seems quite low. It is not clear whether 5 examples indicate 5 instances of the same PDE with different co-efficients or whether it’s 5 different sparse representations, with the same co-efficients. \n\nThe property of neural operator is that it’s discretization agnostic. The authors don’t mention what discretizations they tested. 128x128 grid is not indicative of the discretization, but rather the resolution. By this I mean that this setting could be a set of densely located 128x128 points in a very small area within a large mesh or a set of 128x128 sparse points spread over the entire mesh. \n\nWhile comparing against data driven FNO models is a good baseline, the authors propose this architecture as a substitute for Physics informed ML. Therefore, it would be appropriate to show how this scales against PINNs and PINOs. \n\nIn the FNO paper, the models were trained on training sets with 1000 instances. However, the authors here use a significantly smaller training dataset. Could it be possible that the failure scenarios shown in Figure 3. are because the FNO models require a larger training set to converge? Perhaps a more fair comparison would be to train both the FNO model and the PPI-FNO model on the larger dataset. \nIt seems unreasonable to think that a system is so sparse that the training dataset only has 5 instances. Moreover, it is not clear whether sparsity refers to the size of the training dataset or the number of points within the mesh (sparse discretization)."}},"nonreaders":[],"tmdate":1731428278750,"tcdate":1730131844900,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4999/Reviewer_VSib"],"signatures":["ICLR.cc/2025/Conference/Submission4999/Reviewer_VSib"],"forum":"CrmUKllBKs","number":3,"license":"CC BY 4.0","cdate":1730131844900,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4999/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428278750,"domain":"ICLR.cc/2025/Conference","replyto":"CrmUKllBKs","id":"CWtN045Q4n","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Pseudo Physics","Data-Driven Physics Discovery","PDEs","Neural Operator","AI for science","Scientific Machine Learning"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in operator learning are transforming the landscape of computational physics and engineering, especially alongside the rapidly evolving field of physics-informed machine learning. The convergence of these areas offers\nexciting opportunities for innovative research and applications. However, merging\nthese two realms often demands deep expertise and explicit knowledge of physical systems, which may be challenging or even impractical in relatively complex applications. To address this limitation, we propose a novel framework: Pseudo\nPhysics-Informed Neural Operator (PPI-NO). In this framework, we construct a\nsurrogate physics system for the target system using partial differential equations\n(PDEs) derived from simple, rudimentary physics knowledge, such as basic differential operators. We then couple the surrogate system with the neural operator model, utilizing an alternating update and learning process to iteratively enhance\nthe model’s predictive power. While the physics derived via PPI-NO may not mirror the ground-truth underlying physical laws — hence the term “pseudo physics” — this approach significantly enhances the accuracy of current operator learning\nmodels, particularly in data scarce scenarios. Through extensive evaluations across\nfive benchmark operator learning tasks and an application in fatigue modeling,\nPPI-NO consistently outperforms competing methods by a significant margin. The\nsuccess of PPI-NO may introduce a new paradigm in physics-informed machine\nlearning, one that requires minimal physics knowledge and opens the door to\nbroader applications in data-driven physics learning and simulations."},"_bibtex":{"value":"@misc{\nchen2025pseudo,\ntitle={Pseudo Physics-Informed Neural Operators},\nauthor={Keyan Chen and Yile Li and Da Long and WEI W. XING and Jacob Hochhalter and Shandian Zhe},\nyear={2025},\nurl={https://openreview.net/forum?id=CrmUKllBKs}\n}"},"title":{"value":"Pseudo Physics-Informed Neural Operators"},"pdf":{"value":"/pdf/864b77c21caf8f31310746c5d9b464fe5feadfa1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|pseudo_physicsinformed_neural_operators"},"authorids":{"value":["~Keyan_Chen3","~Yile_Li1","~Da_Long1","~WEI_W._XING1","~Jacob_Hochhalter1","~Shandian_Zhe1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Keyan Chen","Yile Li","Da Long","WEI W. XING","Jacob Hochhalter","Shandian Zhe"]}},"version":2},{"content":{"summary":{"value":"This paper presents \"GenCANN,\" a framework where an LLM is prompted to generate and iteratively refine a physics-constrained neural network (CANN) for material modeling. The system is evaluated on three mechanics datasets, where it is shown to match or exceed the performance of \"human-designed\" baselines.\n\nWhile the goal of automating scientific modeling is relevant, the paper in its current form is not suitable for publication. Its central claims are invalidated by a combination of critical factors: 1) The problem formulation is a \"toy-level\" task that is orders of magnitude simpler than established code-generation/agentic (eg, MLE bench) benchmarks, making the LLM's success unsurprising. 2) The experimental methodology is fatally flawed, primarily by failing to account for the stochasticity of LLM outputs and by using confounded, unfair baselines. 3) The work is fundamentally irreproducible as submitted, lacking the necessary artifacts (prompts and conversation logs) to verify its claims."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Did the authors run the 3-hop refinement experiment only once for each dataset? If not, what are the mean and standard deviation of the R2 scores over (e.g.) 10 or 20 independent trials? Without this, the results in Table 1 are meaningless.\n\n- How can the authors justify comparing a 3-hop search algorithm (GenCANN) against a single, static, published baseline (CANN)? A fair baseline would be a human expert given the same 3-hop refinement process, or a standard hyperparameter optimization (e.g., random search) run with a similar budget.\n\n- Given that Figure 6 clearly demonstrates that R2 hides catastrophic model failures at the boundaries, why was this aggregate metric chosen as the sole feedback signal for the LLM, rather than a more robust, physics-aware metric (e.g., max relative error)?\n\n- Will the authors provide the full, unedited prompts and complete conversation logs for all refinement trials to allow for verification of the generative process?"},"rating":{"value":0},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The paper addresses an important and popular application area: the use of LLMs to automate and lower the barrier to entry for complex constitutive modeling tasks.\n\nThe proposed framework, which combines LLM-based code generation with a physics-informed constitutive model ML framework (CANN), is clearly presented."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The premise of using LLMs as code-generating/agents for scientific tasks is not novel; in fact, the paper's own literature review cites numerous recent examples. The specific task delegated to the LLM is a simple, \"fill-in-the-blanks\" hyperparameter selection for a small regression model within a predefined, human-authored code skeleton. This is far simpler than rigorous, established benchmarks like SWE-bench (which requires fixing real-world GitHub issues) or MLE-bench (which involves end-to-end ML competitions). Given that LLMs are known to struggle on these complex benchmarks, their success on this paper's highly constrained, simple regression task is entirely expected and provides no new insight.  \n\n- The core method involves an LLM in a stochastic 3-hop refinement loop, where the \"best-performing version is kept\". However, the results in Table 1 are presented as single, deterministic $R^2$ scores. The paper makes no mention of if running this 3-hop-trials multiple repeats, nor does it report any mean, standard deviation, or variance for its results. This is a critical omission. As presented, the results are not scientific findings; they are single, unreplicated cherry-picks or \"lucky runs.\"  \n\n- The comparison to a \"human-designed CANN\" is not clearly stated. The paper states that these baselines are static, pre-existing models taken from prior literature (e.g., \"the best CANN reported in the literature (Pierre et al., 2023)\" and \"the initial CANN publication (Linka et al., 2021)\"). This is an \"apples-to-oranges\" comparison. The GenCANN is a 3-hop search algorithm allowed to optimize for the specific dataset, while the \"human\" baseline is just a single, static point from previous results.  \n\n- The paper's own analysis reveals the confounding variable that likely explains its results: the LLM-generated models are simply massively larger than the baselines. For the Skin dataset, the paper admits the GenCANN's (128-128-64-32 layers) superior accuracy \"likely reflects its larger architecture\" compared to the baseline's (12-12 layers). The finding that \"a much larger network achieves a better R2 score on a regression task\" can be trivially found by human too (without nearly no effort) and does not show any advantage of \"LLM-auto-design\".\n\n- The paper's own data shows that $R^2$ is a flawed and non-robust feedback signal. For the synthetic rubber dataset, Table 1 reports a \"perfect\" R2=1.00 for both the GenCANN and the baseline. However, Figure 6(c) clearly shows the exact same baseline model failing catastrophically (high relative error) at the x-axis boundaries. The aggregate $R^2$ score completely hides this critical failure. Using this highly-compressed metric to guide the LLM's refinement does not make sense-how did it lead to a better results in Figure 6, while the $R^2$ is also 1 for a poor architecture? \n\n- The work is unverifiable. The authors do not provide the exact prompts or the full conversation logs from the 3-hop refinement process. The \"Exemplary GENCANN IMPLEMENTATION\" shows only the final product, not the process, and the \"Reproducibility Statement\" links only to code repositories, not these critical generative artifacts."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915667172,"tcdate":1761987561302,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1061/Reviewer_aqh8"],"signatures":["ICLR.cc/2026/Conference/Submission1061/Reviewer_aqh8"],"forum":"hBoFfrg3HJ","number":4,"license":"CC BY 4.0","cdate":1761987561302,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1061/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915667172,"domain":"ICLR.cc/2026/Conference","replyto":"hBoFfrg3HJ","id":"qeFXYsK6Vh","forumContent":{"TLDR":{"value":"We show that LLMs can automatically generate physics-constrained neural networks, reducing the expertise needed for constitutive modeling in solid mechanics."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large Language Models (LLMs)","Physics-constrained neural networks","Constitutive Artificial Neural Networks (CANNs)","Automated model generation","Data-driven solid mechanics"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Large language model (LLM)-based agentic frameworks increasingly adopt the paradigm of dynamically generating task-specific agents. We suggest that not only agents but also specialized software modules for scientific and engineering tasks can be generated on demand. We demonstrate this concept in the field of solid mechanics. There, so-called constitutive models are required to describe the relationship between mechanical stress and body deformation. Constitutive models are essential for both the scientific understanding and industrial application of materials. However, even recent data-driven methods of constitutive modeling, such as constitutive artificial neural networks (CANNs), still require substantial expert knowledge and human labor. We present a framework in which an LLM generates a CANN on demand, tailored to a given material class and dataset provided by the user. The framework covers LLM-based architecture selection, integration of physical constraints, and complete code generation. Evaluation on three benchmark problems demonstrates that LLM-generated CANNs achieve accuracy comparable to or greater than manually engineered counterparts, while also exhibiting reliable generalization to unseen loading scenarios and extrapolation to large deformations. These findings indicate that LLM-based generation of physics-constrained neural networks can substantially reduce the expertise required for constitutive modeling and represent a step toward practical end-to-end automation."},"_bibtex":{"value":"@misc{\ntacke2026automating,\ntitle={Automating modeling in mechanics: {LLM}s as designers of physics-constrained neural networks for constitutive modeling of materials},\nauthor={Marius Tacke and Matthias Busch and Kian Abdolazizi and Jonas F. Eichinger and Kevin Linka and Christian J Cyron and Roland Aydin},\nyear={2026},\nurl={https://openreview.net/forum?id=hBoFfrg3HJ}\n}"},"title":{"value":"Automating modeling in mechanics: LLMs as designers of physics-constrained neural networks for constitutive modeling of materials"},"pdf":{"value":"/pdf/b33d2b7c279c0c6d84e2749c8af4200fbee70bcd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"tacke|automating_modeling_in_mechanics_llms_as_designers_of_physicsconstrained_neural_networks_for_constitutive_modeling_of_materials"},"authorids":{"value":["~Marius_Tacke1","~Matthias_Busch1","~Kian_Abdolazizi1","~Jonas_F._Eichinger1","~Kevin_Linka1","~Christian_J_Cyron1","~Roland_Aydin1"]},"authors":{"value":["Marius Tacke","Matthias Busch","Kian Abdolazizi","Jonas F. Eichinger","Kevin Linka","Christian J Cyron","Roland Aydin"]}},"version":2},{"content":{"summary":{"value":"The manuscript describes an application of the conditional diffusion model for the simulation of complex physical system, turbulent flows. The authors performed an extensive numerical studies using a range of solvers and a few different flow geometries. It is shown that overall the diffusion model outperforms supervised approaches."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"The authors performed large-scale simulations and extensive numerical experiments to investigate the conditional diffusion model for the physics problems. The result seems to suggest an advantage of the diffusion model in physics simulations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While it is interesting to see the capability of the diffusion model in learning physics problems, the study does not go beyond a relatively straightforward application of the conditional diffusion model, which does not align well with the scope of ICLR. The authors used simple evaluation metrics, which may miss important characteristics of physics problems."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"1. One of the most important characteristics of the physics problem is the conservation law. If not the mass conservation constraint, the computation becomes just a very simple matrix vector multiplications. What's the divergence-free error of the diffusion model and how it compares with the computational physics model?\n\n2. MSE error may not be the best metric to investigate the physics problem. For example, it will be helpful to compare the power spectrum to see if the nonlinear energy transfer is correctly represented in the diffusion model. For the turbulent problems considered, there are well defined metrics that give better representation of the physics. The authors need to compare those metrics, instead of simple MSE.\n\n3. What does it mean to have a posterior sampling? While the diffusion model can sample from the probability distribution, the problem itself is deterministic. It does not make a sense, simply because the diffusion model can generate a sample from a distribution, suddenly the authors arguing that they can sample from a posterior distribution when the problem setup is deterministic. If the authors consider the primitive variables as random variables, the problem formulation has also be properly stated and changed.\n\n4. Again from the comment 3, the paper lacks a proper problem formulation."},"rating":{"value":"5: marginally below the acceptance threshold"},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635947085,"tcdate":1698260318049,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission216/Reviewer_1NQ6"],"signatures":["ICLR.cc/2024/Conference/Submission216/Reviewer_1NQ6"],"forum":"1hhja8ZxcP","number":1,"license":"CC BY 4.0","cdate":1698260318049,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission216/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635947085,"domain":"ICLR.cc/2024/Conference","replyto":"1hhja8ZxcP","id":"3PbMl6CZwT","forumContent":{"TLDR":{"value":"We employ autoregressive diffusion models for the simulation of turbulent flows, and find that they yield excellent accuracy and rollout stability."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["turbulent flow","PDEs","numerical simulation","diffusion models","autoregressive models"]},"supplementary_material":{"value":"/attachment/00bf9b95aea315c609939fe20bbf16083220ea6c.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Simulating turbulent flows is crucial for a wide range of applications, and machine learning-based solvers are gaining increasing relevance. However, achieving stability when generalizing to longer rollout horizons remains a persistent challenge for learned PDE solvers. We address this challenge by introducing a fully data-driven fluid solver that utilizes an autoregressive rollout based on conditional diffusion models. We show that this approach offers clear advantages in terms of rollout stability compared to other learned baselines. Remarkably, these improvements in stability are achieved without compromising the quality of generated samples, and our model successfully generalizes to flow parameters beyond the training regime. Additionally, the probabilistic nature of the diffusion approach allows for inferring predictions that align with the statistics of the underlying physics. We quantitatively and qualitatively evaluate the performance of our method on a range of challenging scenarios, including incompressible and transonic flows, as well as isotropic turbulence."},"_bibtex":{"value":"@misc{\nkohl2024turbulent,\ntitle={Turbulent Flow Simulation using Autoregressive Conditional Diffusion Models},\nauthor={Georg Kohl and Liwei Chen and Nils Thuerey},\nyear={2024},\nurl={https://openreview.net/forum?id=1hhja8ZxcP}\n}"},"title":{"value":"Turbulent Flow Simulation using Autoregressive Conditional Diffusion Models"},"pdf":{"value":"/pdf/e389cadc04229cd70935a93668fcf9086b51a828.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"kohl|turbulent_flow_simulation_using_autoregressive_conditional_diffusion_models"},"authorids":{"value":["~Georg_Kohl1","~Liwei_Chen2","~Nils_Thuerey1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Georg Kohl","Liwei Chen","Nils Thuerey"]}},"version":2},{"content":{"summary":{"value":"Paper proposes a simulation system for building HVAC systems which can be used to train reinforcement learning based control agents. The simulation uses physics-based approach for simulation, overcoming the limitations of ML based system that do not generalize to out-of-distribution data. Traditional physics-based simulations are too slow and computationally expensive, which make RL agents difficult to train. The simulator is calibrated to real-world data using hyper-parameter tuning of physics equations. An RL agent is trained to show the feasibility of the simulation system."},"presentation":{"value":"1 poor"},"contribution":{"value":"1 poor"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The problem statement is clear, and is a well-known problem in this domain\n- The idea makes sense, using fast physics-based simulations overcomes the limitations of prior simulation systems like EnergyPlus\n- A simulator of a real world building that can be created in just 3 hours is appealing"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Paper claims that the simulations will generalize to out-of-distribution data as it is based on physics, however no evidence has been provided to show the generalization. As a result, it is unclear if the RL agent is going to learn a meaningful policy.\n- The drift in error accumulates over time, and has been shown in the evaluation period of 6 hours. RL agents require simulation for at least a year, it is unclear how the simulation system can be practically useful. \n- No details of the RL agent simulation is provided. What is the episode length? What were the weather conditions? What were the occupancy conditions? \n- The main contribution of the paper is based on physics equations of heat transfer. However, they have not been explained adequately. Precise description of the equations used is required to assess the fidelity of the physics simulations. \n- Figure 3 indicates that the control system in the simulation has not been properly modeled. The control system should capture the transition from day to night mode. \n- It is unclear if an error of 0.83 degrees is meaningful. First, it's unclear if the unit is Celsius or Fahrenheit. Second, it is unclear how much the indoor temperature drifts during a 6 hour period. If the range of temperature is only 2 degrees, then an error of 0.83 is >40%. \n- Unclear why certain aspects like radiative gains were ignored in the simulator."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Weaknesses above summarize the questions I have for the authors. Brief summary below\n- What are the details of the RL simulation system? How do you know if the simulation is accurate? I would have expected a controlled experiment where the HVAC settings are changed and the corresponding measurements are matched against the simulations. \n- What are the equations used in the simulator?\n- What is the range of temperature values in the train and test dataset?\n- How long do you run the simulation while training the RL agent? Is the error accumulation sufficiently low to train an agent during that period?"},"rating":{"value":"1: strong reject"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636632694,"tcdate":1698901958918,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5942/Reviewer_P7Wf"],"signatures":["ICLR.cc/2024/Conference/Submission5942/Reviewer_P7Wf"],"forum":"5XUlfPcQnG","number":3,"license":"CC BY 4.0","cdate":1698901958918,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5942/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636632694,"domain":"ICLR.cc/2024/Conference","replyto":"5XUlfPcQnG","id":"cplZNmbIgv","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"TLDR":{"value":"Customizable Simulator for training RL agent to optimize an HCAV system of a commercial building."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["HVAC","Reinforcement Learning","Simulation"]},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Modern commercial Heating, Ventilation, and Air Conditioning (HVAC) systems form a complex and interconnected thermodynamic system with the building and outside weather conditions, and current setpoint control policies are not fully optimized for minimizing energy use and carbon emission. Given a suitable training environment, a Reinforcement Learning (RL) model is able to improve upon these policies, but training such a model, especially in a way that scales to thousands of buildings, presents many practical challenges. To address these challenges, we propose a novel simulation based approach, where a customized simulator is used to train the agent for each building. Our simulator is lightweight and calibrated with recorded data from the building to achieve sufficient fidelity. On a two-story, 68,000 square foot building, with 127 devices, we were able to calibrate our simulator to have just over half a degree of drift from the real world over a 6 hour period. We train an RL agent on this simulator and demonstrate that our agent is able to learn an improved policy. This approach is an important step toward having a real-world Reinforcement Learning control system that can be scaled to many buildings, allowing for greater efficiency and resulting in reduced energy consumption and carbon emissions."},"_bibtex":{"value":"@misc{\ngoldfeder2024a,\ntitle={A Calibrated Simulation for Offline Training of Reinforcement Learning Agents to Optimize Energy and Emission in Office Buildings},\nauthor={Judah Goldfeder and John Sipple},\nyear={2024},\nurl={https://openreview.net/forum?id=5XUlfPcQnG}\n}"},"title":{"value":"A Calibrated Simulation for Offline Training of Reinforcement Learning Agents to Optimize Energy and Emission in Office Buildings"},"pdf":{"value":"/pdf/bdfc2ab6a1b4a3683e5ecf3eb6413ea673300d81.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"goldfeder|a_calibrated_simulation_for_offline_training_of_reinforcement_learning_agents_to_optimize_energy_and_emission_in_office_buildings"},"authorids":{"value":["~Judah_Goldfeder1","sipple@google.com"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Judah Goldfeder","John Sipple"]}},"version":2},{"content":{"summary":{"value":"The paper tackles a core mismatch in physics-guided diffusion: prior methods enforce PDE constraints on the posterior mean at noisy states, which doesn’t guarantee the final samples obey the physics (a Jensen-gap issue). It proposes Physics-Informed Distillation (PIDDM): train a standard teacher diffusion model, then distill a one-step student and penalize the PDE residual directly on the student’s outputs; at test time, a tiny latent-noise refinement can further reduce residuals. The same student supports forward, inverse, and partial-reconstruction by optimizing its latent under data masks plus the PDE residual. On Darcy/Poisson/Burgers, PIDDM achieves lower PDE error and competitive or better MSE than prior physics-aware baselines, while using ~1 NFE. The key contribution is a simple recipe that enforces physics on final samples to avoid the Jensen gap and cutting compute."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"### Questions/Confusions\n- **On \"bypassing Jensen's gap via the true posterior $p(x_0|x_t)$\" claim**. In section 4.2, the paper claims its improvement arises from *\"enforcing PDE constraints on the true posterior $x_0 \\sim p(x_0|x_t)$, which is fundamentally more accurate than on the posterior mean $\\mathbb{E}(x_0|x_t)$, thereby theoretically bypassing Jensen's Gap\"*. However, the method described in Algorithm 1 and Equation (8) applies the PDE loss to the output of the student model . This output is a single, deterministic function of ϵ, trained to approximate samples from the teacher. Is it not a significant overstatement to call this \"enforcing constraints on the true posterior\"? The method appears to constrain a single, learned approximation of a sample, not the true (and intractable) posterior distribution. Could the authors please clarify this discrepancy in the theoretical framing?\n\n### Errors in the tables.\nIn Table 1: ECI, instead of PIDDM-1, is the second best in terms of SMSE for Poisson equation.   Diffusion PDE, instead of PIDDM-1, is the second best in terms of MMSE and SMSE for Burger's equation. \n\n### Clarity on Residual Implementation\nSection 2.1 mentions that finite difference methods are commonly used to approximate the differential operators, and Appendix C provides the governing PDE equations. However, the specific details on the implementation of the physics residual operator used in loss functions, such as the exact finite difference stencils. order of accuracy, handling of initial conditions/boundary conditions are missing. While readers may find the details in the code provided, could the authors include theses important implementation details directly in the appendix? \n\n### Minor typos\n- Line 351-352: Burger's equation -> Burgers' equation \n- Table 1: Burger -> Burgers"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"- Originality. This paper takes a simple but fresh angle: stop enforcing physics on noisy states on posterior means and put the PDE loss on the actual final samples. Doing this via teacher -> one step student distillation with PDE loss is a clean combo. It is not a brand-new primitive but a well-targeted rethink of where the constraints belongs.\n\n- Quality. The paper supports its claims across some PDE benchmarks, and comprarisons are made against some of the relevant and competitive baselines. The results consistently demonstrate PIDDM's effectiveness. Extensive ablation studies further investigate key design choices.\n\n- Clarity. This paper clearly articulates and empirically shows the issue of Jensen's Gap. The motivation is strong. The objective and algorithms are readable.\n\n- Significance. It enables one-step unconditional generation (NFE=1), and the same distilled student model can be adapted via inference-time optimization to handle different (but inherently the same under the paper's framing) problems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### Unjustified Complexity for Downstream Tasks Compared to Standard Inverse Problem Methods\nThe paper's approach to downstream tasks uses a complex, iterative optimization that appears potentially redundant, relatively fragile, and poorly justified given the problem's inherent structure and existing solution paradigms.\n- By modeling the joint filed $x=(u,a)$, all downstream tasks inherently become inverse problems: estimating the unknown components of $x$ given partial observations defined by mask $M$ and governing physics $\\mathcal{R}(x)=0$. Whether the known part is $a$ (forward) $x$ (inverse), or a subset of both (reconstuction) is merely a difference in the mask M. \n- For the primary downstream tasks, the method (Algorithm 3) optimizes $\\varepsilon$ in the latent space that has the same dimensionality as the field, and backpropagates through both a student model distilled from a pre-trained teacher model, and the discrete residual implemented implemented in a differentiable way. \n- Since the method already requires a differentiable implementation of the physics residual $\\mathcal{R}(x)$, one could directly solve these problems via standard gradient-based optimization in the data space. The objective would simply minimize a combination of the data mismatch term and the physics residual term\n$$\\min_x ||(x-x^\\prime)\\odot M||^2 + \\lambda ||\\mathcal{R}(x\\odot(1-M)+x^\\prime\\odot M)||^2.$$\nThis avoids the need for any model (no teacher, no student) and optimizes directly on the quantity of interest. The paper does not provide theoretical argument or empirical comparison to demonstrate that Algorithm 3 offers advantages (e.g., better convergence, finding better minima, improved sample quality) over the simpler, direct optimization approach. The significant overhead of training the diffusion and student model seems entirely unjustified without demonstrating superiority over simpler optimization baselines.\n- The authors acknowlege that the work  assumes access to a well-trained teacher model, but even assuming a perfectly trained teacher and student, optimizing through the high-dimensional latent space via the complex, learned student inherently complicates the optimization landscape.\n- More importantly, direct optimization relies primarily on the physics model and observed data without concerns about out-of-distribution generalization. In contrast, PIDDM relies heavily on the student model. Given the infinite dimensionality of PDE solution spaces and infinite choices of parameters/initial conditions/boundary conditions, it is highly likely that the teacher model (and thus the student) will be trained on insufficient or non-reprensentative data. If the specific problem instance falls outside the distribution learned by the teacher/student, the student model could actively harm the optimization. \n\n- **This questions the utility of the proposed framework for anything beyond unconditional generation.**  \n\n### Unclear statistical significance\nThe paper compares PIDDM against several baselines using metrics like MMSE, SMSE, FPD, and PDE error, reporting single numerical values for each method and dataset (Tables 1, 2, 3, 4). However, as these are generative models, their outputs are inherently stochastic, depending on factors like the initial noise seed, and potentially the optimization path in Algorithm 3. Without error bars, it is impossible to assess the statistical significance of the reported differences between methods. This is particularly problematic where the performance metrics between PIDDM and certain baselines are very close. The authors are encouraged to report uncertainty calculated over multiple independent runs for all reported metrics to strengthen the paper's claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923040064,"tcdate":1761151700100,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12068/Reviewer_ZQ1z"],"signatures":["ICLR.cc/2026/Conference/Submission12068/Reviewer_ZQ1z"],"forum":"hW7P3x9W8A","number":1,"license":"CC BY 4.0","cdate":1761151700100,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12068/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923040064,"domain":"ICLR.cc/2026/Conference","replyto":"hW7P3x9W8A","id":"LnY31If10J","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Diffusion","Physical Sciences"]},"supplementary_material":{"value":"/attachment/bdf75a4d6c308d3a5a9ea981e5401828c7091e06.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Modeling physical systems in a generative manner offers several advantages, including the ability to handle partial observations, generate diverse solutions, and address both forward and inverse problems. Recently, diffusion models have gained increasing attention in the modeling of physical systems, particularly those governed by partial differential equations (PDEs). However, diffusion models only access noisy data $\\boldsymbol{x}_t$ at intermediate steps, making it infeasible to directly enforce constraints on the clean sample $\\boldsymbol{x}_0$ at each noisy level. As a workaround, constraints are typically applied to the expectation of clean samples $\\mathbb{E}[\\boldsymbol{x}_0|\\boldsymbol{x}_t]$, which is estimated using the learned score network. However, imposing PDE constraints on the expectation does not strictly represent the one on the true clean data, known as Jensen's Gap. This gap creates a trade-off: enforcing PDE constraints may come at the cost of reduced accuracy in generative modeling. To address this, we propose a simple yet effective post-hoc distillation approach, where PDE constraints are not injected directly into the diffusion process, but instead enforced during a post-hoc distillation stage. We term our method as Physics-Informed Distillation of Diffusion Models (PIDDM). This distillation not only facilitates single-step generation with improved PDE satisfaction, but also support both forward and inverse problem solving and reconstruction from randomly partial observation. Extensive experiments across various PDE benchmarks demonstrate that PIDDM significantly both improves PDE satisfaction and generative modeling over several recent and competitive baselines, such as PIDM, DiffusionPDE, and ECI-sampling, while achieving lower computational overhead and avoiding extensive hyperparameter tuning. Our approach can shed light on more efficient and effective strategies for incorporating physical constraints into diffusion models."},"_bibtex":{"value":"@misc{\nzhang2025physicsinformed,\ntitle={Physics-Informed Distillation of Diffusion Models for {PDE}-Constrained Generation},\nauthor={Yi Zhang and Difan Zou},\nyear={2025},\nurl={https://openreview.net/forum?id=hW7P3x9W8A}\n}"},"title":{"value":"Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation"},"pdf":{"value":"/pdf/17d887d71f2708c8581d2d7408e3bf48c01acea0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|physicsinformed_distillation_of_diffusion_models_for_pdeconstrained_generation"},"authorids":{"value":["~Yi_Zhang94","~Difan_Zou1"]},"authors":{"value":["Yi Zhang","Difan Zou"]}},"version":2},{"content":{"summary":{"value":"The paper presents m-PhOeNIX (Multi-Physics Operator Network for In-Context Learning), a model that combines local wavelet experts and context gates to enable multi-task and sequential learning for various physics-driven PDEs. The proposed framework aims to allow the model to learn new PDE systems without requiring extensive re-training, preventing catastrophic forgetting. However, significant limitations in theoretical rigor, computational efficiency, and a lack of sufficient validation experiments reduce the overall impact of the work."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See Weakness."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The idea is interesting to use local wavelet experts and context gates to create a flexible framework for capturing multi-scale features across multiple physics systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The theoretical foundation of m-PhOeNIX is insufficiently developed. The authors need to provide a clearer theoretical rationale and motivation for combining the existing architectures. The paper also lacks a formal analysis of why Daubechies wavelets were chosen over other types of wavelets. \n\n2. Wavelet-based operations are typically more computationally expensive than FFT, especially for high-resolution data or real-time applications. However, the paper does not provide any benchmarks on runtime or memory usage. This information is essential to evaluate the model’s practical viability in large-scale scientific applications.\n\n3. The use of multiple wavelet experts and context gates adds significant model complexity, which scales up with the number of experts and task diversity. This may introduce memory overheads, making m-PhOeNIX less scalable for high-dimensional PDEs.\n\n4. Boundary conditions significantly impact the solutions to PDEs, yet m-PhOeNIX does not explore how it would handle varying boundary conditions across tasks. A strategy for managing changing boundary conditions could be promising.\n\n5. The model is designed to handle cases where the underlying PDEs are unknown or complex, making it potentially valuable for real-world applications. However, the model has been tested only on synthetic data, limiting its demonstrated applicability. It is essential to test the model on real-world datasets and more complex, higher-dimensional systems. Such validation would provide stronger evidence for the model's claimed adaptability and robustness in handling diverse and realistic physical phenomena."}},"nonreaders":[],"tmdate":1731427859383,"tcdate":1730709058264,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7512/Reviewer_aFVW"],"signatures":["ICLR.cc/2025/Conference/Submission7512/Reviewer_aFVW"],"forum":"ubUTIlAH0m","number":3,"license":"CC BY 4.0","cdate":1730709058264,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7512/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427859383,"domain":"ICLR.cc/2025/Conference","replyto":"ubUTIlAH0m","id":"6YqGpnAlXC","forumContent":{"TLDR":{"value":"This framework does simultaneous and sequential learning of solution operators of multiple heterogeneous physical systems without catastrophic forgetting."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multi-physics operator learning","neural operator","catastrophic forgetting","continual learning","wavelet"]},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We propose a multi-physics operator network for simultaneous and sequential learning of solution operators of multiple heterogeneous parametric partial differential equations. Existing neural operators are adept at learning the solution operator of only a single physical system, and adapting to new physical equations requires training a new surrogate model from scratch with physics-specific intensive hyperparameter tuning. The proposed multi-physics neural operator leverages the recent advancements in wavelet-based kernel integral-induced neural operator modeling and instantiates a memory-based ensembling strategy for projecting heterogeneous physical systems into a common shared feature space. The local channel-level ensembling is supported by context gates, which not only utilize the shared features to embed the features of multiple heterogeneous physical systems into the network parameters but also allow the multi-physics operator to learn new solution operators by transferring knowledge sequentially; this allows the proposed model to continually learn without forgetting. We illustrate the efficacy of our algorithm by simultaneously and sequentially learning six complex time-dependent solution operators of six physical systems. The inference results on the simultaneous and sequentially trained models depict the ability to infer previously seen physical systems without fine-tuning and catastrophic forgetting, indicating the characteristics of a foundation model. The framework also demonstrates the super-resolution property and generalization to out-of-distribution input conditions."},"_bibtex":{"value":"@misc{\ntripura2024multiphysics,\ntitle={Multi-Physics Operator Network for In-context learning (m-PhOe{NIX})},\nauthor={Tapas Tripura and Souvik Chakraborty},\nyear={2024},\nurl={https://openreview.net/forum?id=ubUTIlAH0m}\n}"},"title":{"value":"Multi-Physics Operator Network for In-context learning (m-PhOeNIX)"},"pdf":{"value":"/pdf/579d1c318f1d6b3fb010607dd3bf0e71ff37c5f9.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"tripura|multiphysics_operator_network_for_incontext_learning_mphoenix"},"authorids":{"value":["~Tapas_Tripura1","~Souvik_Chakraborty2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Tapas Tripura","Souvik Chakraborty"]}},"version":2},{"content":{"summary":{"value":"This paper proposes using a trained language model rather than a general-purpose language model as a user simulator to generate a synthetic conversation dataset. The dataset is then used to train pre-trained language models. The authors train an assistant model called PlatoLM on the synthetic conversation data generated by the trained user simulator. They show that PlatoLM outperforms models trained on synthetic conversations produced by a general-purpose language model."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. This paper demonstrates the efficacy of training a user simulator model for generating synthetic training data to improve language models. The approach of training a user simulator could be broadly applied across domains when curating datasets to train language models.\n2. The comprehensive experiments present promising results when training language models with synthetic conversation datasets produced by the proposed approach of using a trained user simulator model. The trained models outperform those trained on synthetic data generated by a general-purpose language model.\n3. The authors curate a high-quality, human-like multi-turn conversation dataset using the trained user simulator model. The dataset will be open-sourced."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The proposed approach of training a user simulator model to generate synthetic training data, while logical, may lack sufficient novelty. Using a trained language model as a user simulator aligns with prior work on conversational agents and data augmentation. The straightforward nature of training a user simulator model makes the technique intuitive, but also means the work is incremental."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. In section 5.3, what could be the possible reason for the unstable performance increase when scaling up training samples\n2. A minor typo in section 3.2.1, ChaTGPT should be ChatGPT"},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700716844969,"tcdate":1698812638963,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission7381/Reviewer_6WiD"],"signatures":["ICLR.cc/2024/Conference/Submission7381/Reviewer_6WiD"],"forum":"9nddtu94uX","number":4,"license":"CC BY 4.0","cdate":1698812638963,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission7381/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700716844969,"domain":"ICLR.cc/2024/Conference","replyto":"9nddtu94uX","id":"q7HN518ft2","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large Language Model","User Simulation","Human Computer Interaction"]},"supplementary_material":{"value":"/attachment/65555c22ee2606bad7996d1fe101f345329957b8.zip"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The unparalleled performance of closed-sourced ChatGPT has sparked efforts towards its democratization, with notable strides made by leveraging real user and ChatGPT conversations, as evidenced by Vicuna. However, due to challenges in gathering conversations involving human participation, current endeavors like Baize and UltraChat aim to automatically generate conversational data. They primarily rely on ChatGPT conducting roleplay to simulate human behaviors based on instructions rather than genuine learning from humans, resulting in limited scope, diminished diversity, and an absence of genuine multi-round conversational dynamics. To address the above issues, we target human questions extracted from genuine human-machine conversations as a learning goal and train a user simulator called Socratic to produce a high-quality human-centric synthetic conversation dataset. Subsequently, this dataset was used to train our assistant model, named PlatoLM. PlatoLM achieves the SOTA performance among 7B  models (including  LLaMA-2-7B-chat and Vicuna-7B) in both Vicuna-Bench and pairwise comparison in MT-Bench; the effectiveness of PlatoLM is also evidenced by manual evaluation."},"_bibtex":{"value":"@misc{\nanonymous2024platolm,\ntitle={Plato{LM}: Teaching {LLM}s  via a Socratic  Questioning User Simulator},\nauthor={Anonymous},\nyear={2024},\nurl={https://openreview.net/forum?id=9nddtu94uX}\n}"},"title":{"value":"PlatoLM: Teaching LLMs  via a Socratic  Questioning User Simulator"},"pdf":{"value":"/pdf/b84fcdc29b25ccef28d006dc9a10875ca09b1216.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"kong|platolm_teaching_llms_via_a_socratic_questioning_user_simulator"},"authorids":{"value":["~Chuyi_Kong1","~Yaxin_FAN2","~Xiang_Wan1","~Feng_Jiang4","~Benyou_Wang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chuyi Kong","Yaxin FAN","Xiang Wan","Feng Jiang","Benyou Wang"]}},"version":2},{"content":{"summary":{"value":"This work focuses on the logicality of the scientific reasoning process.\nThe authors propose three new metrics to assess the logicality of an LLM’s reasoning process: Logical Fidelity, Causal Connection, and Inferential Progress.\nThey construct an 80K logicality-related SFT dataset and a 864-example in-domain test set.\nExperiments have shown that their constructed training dataset can effectively improve LLM logicality in physics reasoning and the final task performances."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How do you judge the answer correctness of those open-ended questions on PHYSLOGIC? Are you relying on the rule-based method?\n2. I am curious about the range of the proposed three metrics.\n3. How to set the coefficients in the formula in L225?\n4. How to get the Nexus and Weight for each question? Is the Nexus unique for a certain question? (Maybe we can combine the two steps in Nexus into a \"bigger\" one)\n5. In your metrics, you mainly use the cosine similarity of text embeddings. Have you ever tried NLI models, like [1-2].\n6. I am also curious whether Logical Nexus is specific to physics reasoning? It seems that we could also define \"Nexus\" or extract it using your prompt for other domains (math, other STEM domains).\n\n\n[1] Golovneva, O., Chen, M., Poff, S., Corredor, M., Zettlemoyer, L., Fazel-Zarandi, M., & Celikyilmaz, A. (2022). Roscoe: A suite of metrics for scoring step-by-step reasoning. arXiv preprint arXiv:2212.07919.\n\n[2] Xu, X., Diao, S., Yang, C., & Wang, Y. (2024). Can We Verify Step by Step for Incorrect Answer Detection?. arXiv preprint arXiv:2402.10528."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The idea to check logicality from different perspectives is interesting and proven useful for SFT.\n2. The proposed PHYSLOGIC benchmark might be useful for related domains.\n3. The experiments are comprehensive, and the whole paper is full of details.\n4. The collected SFT dataset in itself can be useful for enhancing the physics reasoning ability of LLMs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This paper focuses exclusively on physics reasoning. It might be better to discuss other possible related domains, for example, math and code. I recommend that the authors at least evaluate their trained models as well as baselines on other popular benchmarks (e.g., AIME25, LiveCodeBench) to see the generalizability of their method.\n2. As the PHYSLOGIC is collected from arXiv papers. It is possible that PHYSLOGIC has a data contamination issue.\n3. Lack of an ablation study on the different percentiles for the Logical-Distillation (L-D) method. It seems that the authors use 50% directly.\n4. The proposed metrics heavily rely on \"reasoning step segmentation.\" How to segment such long responses from large reasoning models into \"reasoning steps\"? I think there is no consensus on this problem, and I believe a different segmentation strategy will influence these metrics.\n5. I recommend that the authors spend more text explaining the definition of the new metrics, especially Causal Connection. It is not very intuitive to see the meaning from the formula.\n6. It is better to report the new proposed three metrics on the training data of different methods.\n7. I am uncertain about the reliability of the new metrics. You know, many works try to define some process-level metrics, but many of these are not that reliable."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917290792,"tcdate":1761047981781,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4313/Reviewer_Sq3b"],"signatures":["ICLR.cc/2026/Conference/Submission4313/Reviewer_Sq3b"],"forum":"QxBqTm4B3H","number":1,"license":"CC BY 4.0","cdate":1761047981781,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4313/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917290792,"domain":"ICLR.cc/2026/Conference","replyto":"QxBqTm4B3H","id":"axv4ZdNH6w","forumContent":{"TLDR":{"value":"We designed an algorithm to assess the logicality of LLMs' scientific reasoning and leveraged it to construct an SFT dataset for physics reasoning that exhibits high logicality."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["scientific logicality","logicality assessment","LLM physics reasoning"]},"supplementary_material":{"value":"/attachment/6ea8c23dbc308e7133cf680fd08b812841d5e867.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"With the continuous advancement of reasoning abilities in Large Language Models (LLMs), their application to scientific reasoning tasks has gained significant research attention. Current research primarily emphasizes boosting LLMs' performances on scientific QA benchmarks by training on larger, more comprehensive datasets with extended reasoning chains. However, these approaches neglect the essence of scientific reasoning process -- logicality, which is the rational foundation to ensure the validity of reasoning steps leading to reliable conclusions. In this work, we make the first systematic investigation into the internal logicality underlying LLM scientific reasoning, and develop a scientific logicality enriched methodology, including a set of assessment criteria and data sampling methods for logicality-guided training, to improve the logical faithfulness as well as task performance. Further, we take physics, characterized by its diverse logical structures and formalisms, as an exemplar discipline to practise the above methodology. For data construction, we extract scientific problems from academic literature and sample a high-quality dataset exhibiting strong logicality. Experiments based on three different backbone LLMs reveal that: 1) the training data we constructed can effectively improve the scientific logicality in LLM reasoning; and 2) the enriched scientific logicality plays a critical role in solving scientific problems."},"_bibtex":{"value":"@misc{\nyu2026scientific,\ntitle={Scientific logicality enriched methodology for {LLM} reasoning: A practice in physics},\nauthor={Zhaoxin Yu and Nan Xu and Kun Chen and Jiahao Zhao and Lei Wang and Wenji Mao},\nyear={2026},\nurl={https://openreview.net/forum?id=QxBqTm4B3H}\n}"},"title":{"value":"Scientific logicality enriched methodology for LLM reasoning: A practice in physics"},"pdf":{"value":"/pdf/9817c4107cc3386b18dd7468c6080e9418ef0138.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yu|scientific_logicality_enriched_methodology_for_llm_reasoning_a_practice_in_physics"},"authorids":{"value":["~Zhaoxin_Yu1","~Nan_Xu4","~Kun_Chen9","~Jiahao_Zhao1","~Lei_Wang85","~Wenji_Mao1"]},"authors":{"value":["Zhaoxin Yu","Nan Xu","Kun Chen","Jiahao Zhao","Lei Wang","Wenji Mao"]}},"version":2},{"content":{"summary":{"value":"This paper introduces DiffuseBot, a physics-augmented diffusion model designed for generating and optimizing the morphologies and control mechanisms of soft robots. DiffuseBot aims to bridge the gap between virtually generated content and physical utility in the domain of soft robotics. Firstly, it combines the diffusion process with a physical simulation that serves as a performance certificate, thereby ensuring the feasibility and effectiveness of the generated designs. Secondly, it details a co-design procedure that simultaneously optimizes the physical design and control of the soft robots, leveraging insights from differentiable simulation. The paper validates the efficacy of this approach by presenting a variety of both simulated and physically fabricated robots, along with their diverse capabilities."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. In general, the paper is well written, with only minor flaws. Even those unfamiliar with soft robot design will find the paper easy to comprehend.\n\n2. Although diffusion models are expressive and powerful, their performance for tasks dealing with physical tasks often falls short. Thus, injecting a physics prior or 'physics-augmented diffusion model' is crucial. I think the method proposed in this paper is interesting and promising. \n\n3. The evaluation is comprehensive and thoughtful. The physical robot is impressive.\n"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Overall, I did not identify any major weaknesses in the paper, but here are a few points that could strengthen it:\n\n1. While the writing is generally clear, certain sections could benefit from clearer exposition, such as:\n*  The section on diffusion as co-design is not very intuitive, especially for audiences not familiar with soft robot design. Specifically, it should be clearer how gradient-based optimization benefits robot design and what exactly line 152's \"synergy\" means.\n* It would be helpful if the authors clarify that the \"condition\" in this work actually refers to text.\n2. The robot's actuator and stiffness seem oversimplified, having only constant stiffness. Given that the gradient of $\\Psi_{act}$ is almost zero, it appears that the actuator and stiffness are solely determined by the geometry.\n3. A similar idea of tuning in the embedding space is proposed in[1]. A discussion and connection to this existing work could be interesting.  \n4. In general, the method the paper uses to inject a physics prior into the generation process could be applicable to more general scenarios. Works like Diffuser[2] or Decision Diffuser[3] generate state sequences with diffusion models, but the generated states can sometimes be physically implausible. A deeper discussion about the potential of the method could make the paper stronger.\n\n[1] Gal, Rinon, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H. Bermano, Gal Chechik and Daniel Cohen-Or. “An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.”, ICLR, 2023.\n\n[2] Janner, Michael, Yilun Du, Joshua B. Tenenbaum and Sergey Levine. “Planning with Diffusion for Flexible Behavior Synthesis.”, ICML, 2022. \n\n[3] Ajay, Anurag, Yilun Du, Abhi Gupta, Joshua B. Tenenbaum, T. Jaakkola and Pulkit Agrawal. “Is Conditional Generative Modeling all you need for Decision-Making?” ICLR, 2023. "},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"\n1. I do not fully understand how the k-means clustering is performed for actuator and stiffness generation. Specifically, what kind of feature is used for clustering?\n\n2. In line 86, which structural biases are you referring to ?\n"},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"see weakness."}},"nonreaders":[],"tmdate":1702411286621,"tcdate":1689063372607,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission10463/Reviewer_zr6A"],"signatures":["NeurIPS.cc/2023/Conference/Submission10463/Reviewer_zr6A"],"forum":"1zo4iioUEs","number":3,"license":"CC BY 4.0","cdate":1689063372607,"mdate":1702411286621,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission10463/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"1zo4iioUEs","id":"u78CXxuHHt","forumContent":{"venue":{"value":"NeurIPS 2023 oral"},"keywords":{"value":["soft robot","diffusion model","co-design"]},"supplementary_material":{"value":"/attachment/ffc986d9683f1056512609ca36966932ba11a6bd.pdf"},"_bibtex":{"value":"@inproceedings{\nwang2023diffusebot,\ntitle={DiffuseBot: Breeding Soft Robots With Physics-Augmented Generative Diffusion Models},\nauthor={Tsun-Hsuan Wang and Juntian Zheng and Pingchuan Ma and Yilun Du and Byungchul Kim and Andrew Everett Spielberg and Joshua B. Tenenbaum and Chuang Gan and Daniela Rus},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=1zo4iioUEs}\n}"},"title":{"value":"DiffuseBot: Breeding Soft Robots With Physics-Augmented Generative Diffusion Models"},"paperhash":{"value":"wang|diffusebot_breeding_soft_robots_with_physicsaugmented_generative_diffusion_models"},"TLDR":{"value":"Use diffusion models augmented by physics-based simulation to breed soft robots"},"abstract":{"value":"Nature evolves creatures with a high complexity of morphological and behavioral intelligence, meanwhile computational methods lag in approaching that diversity and efficacy.  Co-optimization of artificial creatures' morphology and control in silico shows promise for applications in physical soft robotics and virtual character creation; such approaches, however, require developing new learning algorithms that can reason about function atop pure structure. In this paper, we present DiffuseBot, a physics-augmented diffusion model that generates soft robot morphologies capable of excelling in a wide spectrum of tasks. \\name bridges the gap between virtually generated content and physical utility by (i) augmenting the diffusion process with a physical dynamical simulation which provides a certificate of performance, and (ii) introducing a co-design procedure that jointly optimizes physical design and control by leveraging information about physical sensitivities from differentiable simulation.  We showcase a range of simulated and fabricated robots along with their capabilities. Check our website: https://diffusebot.github.io/"},"pdf":{"value":"/pdf/99219790f12d9993715f2e6a2eb03395ccb6b38e.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Tsun-Hsuan_Wang2","~Juntian_Zheng1","~Pingchuan_Ma3","~Yilun_Du1","~Byungchul_Kim1","~Andrew_Everett_Spielberg1","~Joshua_B._Tenenbaum1","~Chuang_Gan1","~Daniela_Rus1"]},"authors":{"value":["Tsun-Hsuan Wang","Juntian Zheng","Pingchuan Ma","Yilun Du","Byungchul Kim","Andrew Everett Spielberg","Joshua B. Tenenbaum","Chuang Gan","Daniela Rus"]}},"version":2},{"content":{"summary":{"value":"**Summary:**  \nThis paper presents STANCE, a controllable image-to-video generation framework that aims to improve physical and temporal coherence in diffusion-based video generation. The authors identify two key issues in prior methods—loss of signal density after encoding sparse controls and entanglement of appearance and motion supervision—and propose two simple but effective solutions: Instance Cues, which expand sparse user-editable arrows and masks into dense, camera-relative 2.5D motion fields, and Dense RoPE, which assigns spatially addressable rotary embeddings to selected motion tokens, keeping them spatially anchored during generation. The model jointly predicts RGB and an auxiliary structural map (depth or segmentation), which acts as a consistency witness. Experiments on a large synthetic dataset and real-world examples show that STANCE yields more coherent physical interactions and consistent motion, outperforming multiple baselines such as VLIPP, MoFA-Video, and MotionPro."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"see the weakness"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"**Strengths:**  \n- Clearly identifies key weaknesses in existing controllable video generation methods: sparse tokenization and entangled training objectives.  \n- Proposes **Instance Cues** that make user input dense, interpretable, and camera-relative, improving control precision.  \n- The **Dense RoPE** mechanism elegantly addresses loss of spatial anchoring after tokenization, preserving effective motion tokens even for small or thin objects.  \n- Joint RGB + structural stream training is simple yet empirically effective in stabilizing geometry and improving contact plausibility.  \n- Comprehensive experiments, including synthetic and real-world captured videos, demonstrate consistent improvements in physical coherence (reported with Physics-IQ metric).  \n- The model allows intuitive editing (direction, speed, mass, ∆z) with visually consistent outcomes and realistic cause–effect relations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses:  \n1. While practical, the technical novelty is moderate—Instance Cues and Dense RoPE extend existing token-density and positional embedding ideas rather than introducing entirely new principles.  \n2. Real-world validations are limited in scale; demonstrations mostly show toy cases with one or two rigid objects, leaving uncertainty for complex multi-agent or deformable dynamics.  \n3. The method depends heavily on precomputed instance masks and monocular depth estimation, which may constrain general applicability outside clean lab setups.  \n4. It remains unclear how robust the model is when user inputs deviate from the training distributions (e.g., unrealistic mass or velocity values).  \n5. The auxiliary head’s role is primarily empirical; there is little theoretical or analytical discussion explaining why it particularly improves temporal stability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916463204,"tcdate":1762095185031,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2961/Reviewer_jb9S"],"signatures":["ICLR.cc/2026/Conference/Submission2961/Reviewer_jb9S"],"forum":"FwtKMYHov7","number":4,"license":"CC BY 4.0","cdate":1762095185031,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2961/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916463204,"domain":"ICLR.cc/2026/Conference","replyto":"FwtKMYHov7","id":"tZECOKXTxy","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Generation","Generative Model"]},"supplementary_material":{"value":"/attachment/415a700d351d8bd244411f2900432bea49e53253.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation has recently made striking visual progress, but maintaining coherent object motion and interactions remains difficult. We trace two practical bottlenecks: (i) human-provided motion hints (e.g., small 2D maps) often collapse to too few effective tokens after encoding, weakening guidance; and (ii) optimizing for appearance and motion in a single head can favor texture over temporal consistency. We present STANCE, an image-to-video framework that addresses both issues with two simple components.\nFirst, we introduce Instance Cues—a pixel-aligned control signal that turns sparse, user-editable hints into a dense 2.5D (camera-relative) motion field by averaging per-instance flow and augmenting with monocular depth over the instance mask. This reduces depth ambiguity compared to 2D drag/arrow inputs while remaining easy to user. Second, we preserve the salience of these cues in token space with Dense RoPE, which tags a small set of motion tokens (anchored on the first frame) with time-addressable rotary embeddings. Paired with joint RGB + auxiliary-map prediction (segmentation or depth), our model anchors structure while RGB handles appearance, stabilizing optimization and improving temporal coherence without requiring per-frame trajectory scripts."},"_bibtex":{"value":"@misc{\nanonymous2026stance,\ntitle={{STANCE}: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=FwtKMYHov7}\n}"},"title":{"value":"STANCE: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding"},"pdf":{"value":"/pdf/dcb00196385b99164d59c430fb8613633f2432c0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"chen|stance_motion_coherent_video_generation_via_sparsetodense_anchored_encoding"},"authorids":{"value":["~ZhiFei_Chen1","~Tianshuo_Xu1","~Leyi_Wu1","~Luozhou_Wang2","~Dongyu_Yan1","~Zihan_You2","~Wenting_Luo1","~Guo_Zhang2","~Ying-Cong_Chen1"]},"authors":{"value":["ZhiFei Chen","Tianshuo Xu","Leyi Wu","Luozhou Wang","Dongyu Yan","Zihan You","Wenting Luo","Guo Zhang","Ying-Cong Chen"]}},"version":2},{"content":{"venue":{"value":"Crossref"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-032-05573-6_5.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"nikoshin|from_tissuemimicking_phantoms_to_physicsbased_scans_synthetic_oct_for_fewshot_foundation_model_training"},"html":{"value":"https://doi.org/10.1007/978-3-032-05573-6_5"},"abstract":{"value":"The past decade has seen a substantial increase in demand for high-quality, controllable synthetic data for training medical imaging foundation models. Current synthetic Optical Coherence Tomography (OCT) data generation methods often face a trade-off between physical realism and biological congruence. To address these limitations, we propose a novel pipeline that synergizes data-driven analysis with a physics-based simulation. Our method first analyzes real OCT scans using a zero-shot Segment Anything Model (SAM) to extract realistic parameters of anatomical structures. Subsequently, these parameters are used to generate varied, biologically congruent structural masks. We then simulate a realistic distribution of optical scatterers within these masks and perform a virtual OCT scan that replicates the physics of light-tissue interaction. This physics-based generation pipeline is fundamentally different from physics-informed neural networks. We demonstrate that a MedSAM model trained exclusively on our synthetic data achieves Dice scores of 96% (Noise), 83% (Epidermis), and 97% (Dermis) on unseen real scans. This performance closely matches that of a model engaged in few-shot learning on real data (98%, 85%, and 97% respectively), validating the efficacy of our approach for training models in data-scarce scenarios."},"title":{"value":"From Tissue-Mimicking Phantoms to Physics-Based Scans: Synthetic OCT for Few-Shot Foundation Model Training"},"authors":{"value":[{"fullname":"Denis Nikoshin","username":"https://orcid.org/orcid-search/search?searchQuery=Denis%20Nikoshin"},{"fullname":"Daniil Mikhailenko","username":"https://orcid.org/orcid-search/search?searchQuery=Daniil%20Mikhailenko"},{"fullname":"Alexander Sovetsky","username":"https://orcid.org/orcid-search/search?searchQuery=Alexander%20Sovetsky"},{"fullname":"Alexander Matveyev","username":"~Alexander_Matveyev1"},{"fullname":"Vladimir Zaitsev","username":"https://orcid.org/orcid-search/search?searchQuery=Vladimir%20Zaitsev"},{"fullname":"Lev Matveev","username":"https://orcid.org/orcid-search/search?searchQuery=Lev%20Matveev"}]}},"tmdate":1784725583406,"pdate":1767225600000,"externalIds":["doi:10.1007/978-3-032-05573-6_5"],"tcdate":1784725577774,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Alexander_Matveyev1"],"forum":"bRSlMsPR4w","license":"CC BY-SA 4.0","number":86789,"cdate":1758359392610,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784725583406,"domain":"OpenReview.net/Public_Article","id":"bRSlMsPR4w","version":2},{"content":{"venue":{"value":"SASHIMI@MICCAI 2025 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-032-05573-6_5.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"nikoshin|from_tissuemimicking_phantoms_to_physicsbased_scans_synthetic_oct_for_fewshot_foundation_model_training"},"html":{"value":"https://doi.org/10.1007/978-3-032-05573-6_5"},"_bibtex":{"value":"@inproceedings{DBLP:conf/miccai/NikoshinMSMZM25,\n  author={Denis Nikoshin and Daniil Mikhailenko and Alexander Sovetsky and Alexander L. Matveyev and Vladimir Y. Zaitsev and Lev A. Matveev},\n  title={From Tissue-Mimicking Phantoms to Physics-Based Scans: Synthetic OCT for Few-Shot Foundation Model Training},\n  year={2025},\n  cdate={1735689600000},\n  pages={44-51},\n  url={https://doi.org/10.1007/978-3-032-05573-6_5},\n  booktitle={SASHIMI@MICCAI 2025},\n  crossref={conf/miccai/2025sashimi}\n}\n"},"abstract":{"value":"The past decade has seen a substantial increase in demand for high-quality, controllable synthetic data for training medical imaging foundation models. Current synthetic Optical Coherence Tomography (OCT) data generation methods often face a trade-off between physical realism and biological congruence. To address these limitations, we propose a novel pipeline that synergizes data-driven analysis with a physics-based simulation. Our method first analyzes real OCT scans using a zero-shot Segment Anything Model (SAM) to extract realistic parameters of anatomical structures. Subsequently, these parameters are used to generate varied, biologically congruent structural masks. We then simulate a realistic distribution of optical scatterers within these masks and perform a virtual OCT scan that replicates the physics of light-tissue interaction. This physics-based generation pipeline is fundamentally different from physics-informed neural networks. We demonstrate that a MedSAM model trained exclusively on our synthetic data achieves Dice scores of 96% (Noise), 83% (Epidermis), and 97% (Dermis) on unseen real scans. This performance closely matches that of a model engaged in few-shot learning on real data (98%, 85%, and 97% respectively), validating the efficacy of our approach for training models in data-scarce scenarios."},"title":{"value":"From Tissue-Mimicking Phantoms to Physics-Based Scans: Synthetic OCT for Few-Shot Foundation Model Training"},"authors":{"value":[{"fullname":"Denis Nikoshin","username":""},{"fullname":"Daniil Mikhailenko","username":""},{"fullname":"Alexander Sovetsky","username":""},{"fullname":"Alexander L. Matveyev","username":""},{"fullname":"Vladimir Y. Zaitsev","username":""},{"fullname":"Lev A. Matveev","username":"~Lev_A._Matveev1"}]}},"tmdate":1784549898980,"pdate":1767139200000,"externalIds":["dblp:conf/miccai/NikoshinMSMZM25"],"tcdate":1784549895422,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Lev_Matveev1"],"forum":"2omtfsIxhX","license":"CC BY-SA 4.0","number":73135,"cdate":1735689600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784549898980,"domain":"OpenReview.net/Public_Article","id":"2omtfsIxhX","version":2},{"content":{"summary":{"value":"The paper presents PHUMA, a large-scale physically grounded humanoid motion dataset designed for stable imitation learning in simulation. It addresses the problem that motion data extracted from Internet videos, such as Humanoid-X, often contain artifacts like floating, penetration, and joint-limit violations that degrade physics-based policy training. PHUMA introduces a physics-aware curation process that filters motions for contact and balance consistency, and a physics-constrained retargeting algorithm, PhySINK, that enforces non-floating, non-penetration, and non-skating constraints. The resulting dataset contains about 76k motion clips covering 73 hours of locomotion data. Experiments on the Unitree G1 and H1-2 humanoids in Isaac Gym show that policies trained on PHUMA achieve substantially higher imitation success and physical stability than those trained on LaFAN1, AMASS, or Humanoid-X."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"The imitation results are reported on both Unitree G1 and H1-2. Are the same PHUMA motions directly retargeted to each robot, or are robot-specific shape/scale adjustments applied?\n\nAre there qualitative examples where PhySINK still fails?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper focuses on physically grounded humanoid motion, addressing a gap in large-scale imitation datasets where stability and contact consistency are often ignored.\n\n2. The proposed PhySINK pipeline is technically solid and produces cleaner motion data with fewer artifacts than Humanoid-X.\n\n3. The experiments are comprehensive within simulation, demonstrating clear quantitative improvements on multiple humanoid platforms and providing a useful dataset that can benefit future physics-based imitation learning research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the paper presents a clean and useful dataset, its novelty appears somewhat limited relative to recent works that also improve over Humanoid-X through physics-aware retargeting. Methods such as ASAP (RSS 2025), KunfuBot (NeurIPS 2025), and GMR have already introduced substantial innovations in motion retargeting and data cleaning—ASAP integrates RL-based physical simulation during retargeting, KunfuBot applies extensive filtering for realistic contacts, and GMR focuses specifically on retargeting fidelity—whereas PHUMA mainly improves data quality through curation and constraints. Compared with those works, the contribution here lies primarily in dataset refinement rather than methodological advance. \n\n2. In addition, the dataset only covers locomotion, leaving out other interaction behaviors, and all results are limited to simulation without any hardware validation. There is no qualitative simulation videos for reference, which makes it difficult to assess the realism of the resulting motions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921772404,"tcdate":1761943690861,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10482/Reviewer_1eFX"],"signatures":["ICLR.cc/2026/Conference/Submission10482/Reviewer_1eFX"],"forum":"DDogB73gS4","number":3,"license":"CC BY 4.0","cdate":1761943690861,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10482/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921772404,"domain":"ICLR.cc/2026/Conference","replyto":"DDogB73gS4","id":"JRkCKEDwy1","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["humanoid","retarget","dataset"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Motion imitation is a promising approach for humanoid locomotion, enabling agents to acquire humanlike behaviors. Existing methods typically rely on high-quality motion capture datasets such as AMASS, but these are scarce and expensive, limiting scalability and diversity. Recent studies attempt to scale data collection by converting large-scale internet videos, exemplified by Humanoid-X. However, they often introduce physical artifacts such as floating, penetration, and foot skating, which hinder stable imitation.\nIn response, we introduce \\textbf{PHUMA}, a \\textbf{P}hysically-grounded \\textbf{HUMA}noid locomotion dataset that leverages human video at scale, while addressing physical artifacts through careful data curation and physics-constrained retargeting.\nPHUMA enforces joint limits, ensures ground contact, and eliminates foot skating, producing motions that are both large-scale and physically reliable.\nWe evaluated PHUMA in two sets of conditions: (i) imitation of unseen motion from self-recorded test videos and (ii) path following with pelvis-only guidance. In both cases, PHUMA-trained policies outperform Humanoid-X and AMASS, achieving significant gains in imitating diverse motions. PHUMA will be publicly released to support future research."},"_bibtex":{"value":"@misc{\nlee2026phuma,\ntitle={{PHUMA}: Physically-Grounded Humanoid Locomotion Dataset},\nauthor={Kyungmin Lee and Sibeen Kim and Minho Park and Hyunseung Kim and Dongyoon Hwang and Hojoon Lee and Jaegul Choo},\nyear={2026},\nurl={https://openreview.net/forum?id=DDogB73gS4}\n}"},"title":{"value":"PHUMA: Physically-Grounded Humanoid Locomotion Dataset"},"pdf":{"value":"/pdf/a36747cac87a488230dfe8cb6d5c0aa5a6db82d7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"lee|phuma_physicallygrounded_humanoid_locomotion_dataset"},"authorids":{"value":["~Kyungmin_Lee2","~Sibeen_Kim1","~Minho_Park3","~Hyunseung_Kim1","~Dongyoon_Hwang1","~Hojoon_Lee1","~Jaegul_Choo1"]},"authors":{"value":["Kyungmin Lee","Sibeen Kim","Minho Park","Hyunseung Kim","Dongyoon Hwang","Hojoon Lee","Jaegul Choo"]}},"version":2},{"content":{"summary":{"value":"The paper seeks to investigate how current video world models understand physics. It examines this question from multiple angles using a synthetic “toy ball” dataset and the IntPhys benchmark, together with hierarchical probing, subspace geometry analysis, and attention ablation experiments. The main contribution is the identification of a **Physics Emergence Zone (PEZ)** located at roughly one-third of the network depth, where directed motion information and the ability to distinguish physically possible from impossible events emerge simultaneously. This provides a valuable heuristic for understanding hierarchical specialization in video Transformers (see Fig. 1 and Sec. 4.2)."},"presentation":{"value":3},"significance":{"value":3},"code_of_conduct_acknowledgement":{"value":"Affirmed."},"soundness":{"value":3},"compliance_with_LLM_reviewing_policy":{"value":"Affirmed."},"key_questions_for_authors":{"value":"1. Did the experiments vary properties of the video data such as frame rate and resolution? If low-resolution videos were used, would the depth at which the PEZ appears remain the same?\n2. The paper explicitly attributes the “sawtooth” pattern in iterative null-space projection to “paired (e.g., sine-cosine) feature encoding” (see Fig. 4c and Sec. 7.2). Could you provide either a mathematical derivation or a controlled synthetic experiment to verify that sinusoidal pairs are the unique or most likely cause of this specific degradation curve?\n3. Do CNN-based networks exhibit a similar phenomenon?"},"overall_recommendation":{"value":5},"confidence":{"value":4},"originality":{"value":3},"strengths_and_weaknesses":{"value":"## Strengths\n1. The paper clearly defines the **Physics Emergence Zone** and validates its consistency across different architectures and model scales; it also offers a refined perspective on the separation between scalar and vector motion representations.\n2. It combines multiple analytical tools, including subspace geometry and attention head analysis, which helps avoid the limitations of relying on a single method.\n3. By measuring the principal angles between the IntPhys and Direction subspaces, the paper finds that these subspaces are nearly orthogonal, effectively ruling out the possibility that higher-level physical judgments are made by reusing a compact set of motion variables.\n\n## Weaknesses\n1. The analysis is limited to Transformer-based architectures and methods. It remains unclear whether similar conclusions would hold for CNN-based models.\n2. The study does not appear to explore other types of motion, such as harmonic motion or circular motion."},"limitations":{"value":"Yes."}},"parentInvitations":"ICML.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1782345510114,"tcdate":1773302126566,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission17789/Reviewer_zUMQ"],"signatures":["ICML.cc/2026/Conference/Submission17789/Reviewer_zUMQ"],"forum":"aijGVmEG9Y","number":2,"license":"CC BY 4.0","cdate":1773302126566,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/Submission17789/-/Official_Review","ICML.cc/2026/Conference/-/Edit"],"mdate":1782345510114,"domain":"ICML.cc/2026/Conference","replyto":"aijGVmEG9Y","id":"XZtzIJUlX7","forumContent":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["mechanistic interpretability","video interpretability","world model interpretability","video models","video world models"]},"_bibtex":{"value":"@inproceedings{\njoseph2026interpreting,\ntitle={Interpreting Physics in Video World Models},\nauthor={Sonia Joseph and Quentin Garrido and Randall Balestriero and Matthew Kowal and Thomas Fel and Shahab Bakhtiari and Blake Aaron Richards and Michael Rabbat},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=aijGVmEG9Y}\n}"},"title":{"value":"Interpreting Physics in Video World Models"},"paperhash":{"value":"joseph|interpreting_physics_in_video_world_models"},"originally_submitted_PDF":{"value":"/pdf/c723ed60771151bdc50fd69ae2df2d6cd296f2b6.pdf"},"TLDR":{"value":"Interpretability shows that intuitive physics and motion direction in video world models emerge mid-network without compact, reusable state."},"primary_area":{"value":"deep_learning->other_representation_learning"},"abstract":{"value":"A long-standing question in physical reasoning is whether video models rely on factorized physical state variables, or on task-specific distributed representations. We present the first mechanistic interpretability study of physical variables inside large-scale video encoders, combining layerwise probing, subspace geometry, patch-level decoding, and targeted attention ablations to characterize where and how physical information is orga- nized. Across architectures, we identify a sharp intermediate-depth transition, the Physics Emergence Zone, at which physical variables become linearly accessible. Scalar speed and acceleration are available from early layers, whereas motion direction emerges only at the Physics Emergence Zone, mirroring the V1 to MT motion hierarchy in primate visual cortex. Direction is encoded as a circular high-dimensional population code: dozens of orthogonal probe dimensions must be steered jointly to change the decoded direction, orders of magnitude more than the low-dimensional steering interventions seen in language models. These findings argue against compact physics- engine state variables and support distributed, hierarchically-organized, “brain-like” representations that are nonetheless sufficient for making physical predictions."},"pdf":{"value":"/pdf/2b60051caf32fc825d23bf1a91a00b518f3d5799.pdf"},"lay_summary":{"value":"We investigated how AI video models understand physics. Rather than storing simple variables like a traditional physics engine, they appear to build distributed, brain-like representations spread across many neurons. We found a specific “Physics Emergence Zone” in the middle of the network where physical understanding suddenly becomes accessible, with more complex concepts like motion direction emerging later than simpler concepts like speed. These findings suggest that video models reason about the physical world using hierarchical representations that resemble aspects of biological vision."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Sonia_Joseph1","~Quentin_Garrido1","~Randall_Balestriero1","~Matthew_Kowal1","~Thomas_Fel2","~Shahab_Bakhtiari1","~Blake_Aaron_Richards1","~Michael_Rabbat1"]},"authors":{"value":["Sonia Joseph","Quentin Garrido","Randall Balestriero","Matthew Kowal","Thomas Fel","Shahab Bakhtiari","Blake Aaron Richards","Michael Rabbat"]}},"version":2},{"content":{"comment":{"value":"Thank you for the comprehensive rebuttal, which addresses most of my earlier concerns. I still strongly encourage the authors to expand IntPhys 2 beyond rigid-body dynamics. Including deformable objects, fluids, or granular media would significantly broaden the benchmark’s applicability and deepen its scientific value.\n\nRegarding the proposed avenues for improvement: could the authors report results when the input videos are down-sampled to a lower spatial resolution (e.g., 128×128) so that a longer temporal context (e.g., 48 or more frames) can be fed to the models? I am particularly interested in whether this trade-off (lower fidelity for longer context) yields measurable gains on the IntPhys 2 tasks."}},"parentInvitations":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Official_Comment","tmdate":1761797341278,"tcdate":1754195056990,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_PW3t"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/Reviewer_PW3t"],"forum":"Xpf5x3mLvn","number":3,"license":"CC BY 4.0","cdate":1754195056990,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1018/-/Official_Comment","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1761797341278,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","replyto":"ZFPwO7yD5g","id":"yteG963lb4","forumContent":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Video benchmark","Intuitive Physic","Evaluation"]},"supplementary_material":{"value":"/attachment/4932fdafe9cb6b17072f67b2caf246775bce6feb.zip"},"_bibtex":{"value":"@misc{\nbordes2026intphys,\ntitle={IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments},\nauthor={Florian Bordes and Quentin Garrido and Justine T Kao and Adina Williams and Michael Rabbat and Emmanuel Dupoux},\nyear={2026},\nurl={https://openreview.net/forum?id=Xpf5x3mLvn}\n}"},"title":{"value":"IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments"},"paperhash":{"value":"bordes|intphys_2_benchmarking_intuitive_physics_understanding_in_complex_synthetic_environments"},"dataset_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh.zip"},"TLDR":{"value":"We introduced a new intuitive physic benchmark IntPhys 2, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments"},"primary_area":{"value":"datasets_&_benchmarks_for_computer_vision"},"abstract":{"value":"We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to  macroscopic objects: Permanence, Immutability, Spatio-Temporal Continuity, and Solidity. These conditions are inspired by research into intuitive physical understanding emerging during early childhood. IntPhys 2 offers a comprehensive suite of tests, based on the violation of expectation framework, that challenge models to differentiate between possible and impossible events within controlled and diverse virtual environments. Alongside the benchmark, we provide performance evaluations of several state-of-the-art models. Our findings indicate that while these models demonstrate basic visual understanding, they face significant challenges in grasping intuitive physics across the four principles in complex scenes, with most models performing at chance levels (50\\%), in stark contrast to human performance, which achieves near-perfect accuracy. This underscores the gap between current models and human-like intuitive physics understanding, highlighting the need for advancements in model architectures and training methodologies."},"pdf":{"value":"/pdf/c29b03435c4329f0422348326d7ca2b5d606960a.pdf"},"croissant_file":{"value":"/attachment/54b155200f4140df0470921a91b0093300c98b6a.json"},"code_URL":{"value":"https://dl.fbaipublicfiles.com/IntPhys2/SW50UGh5czJEYXRh_code.zip"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"authorids":{"value":["~Florian_Bordes1","~Quentin_Garrido1","~Justine_T_Kao1","~Adina_Williams1","~Michael_Rabbat1","~Emmanuel_Dupoux1"]},"authors":{"value":["Florian Bordes","Quentin Garrido","Justine T Kao","Adina Williams","Michael Rabbat","Emmanuel Dupoux"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a modular refinement framework for physics reasoning in small language models (SLMs). The method performs SFT warm-up followed by an agent-guided RL loop with error localization and targeted feedback. The authors also introduce PhysicsQA, a curated benchmark of 370 high-school-level physics questions with verified CoT traces."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Topical relevance — Physics reasoning remains a challenging and important domain for LLMs;\n\nClear decomposition of failure modes — The taxonomy (miscomprehension / conceptual error / calculation) provides a structured interpretation of LLM errors, aligning method design with observed failure types.\n\nConsistent improvements on small models — On multiple benchmarks, the proposed refinement gives measurable gains over CoT/RAG/FT/DPO on 1B–3B open-source models, showing practical merit under constrained budgets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Lack of comparison with existing physics reasoning benchmarks\nThe paper introduces PhysicsQA as its main benchmark but does not compare or position the proposed method against existing, more challenging physics-reasoning-centered benchmarks, such as Physreason: A comprehensive benchmark towards physics-based reasoning.\nGiven that Physreason explicitly targets physics reasoning under structured evaluation, the absence of comparison or even discussion makes it unclear whether the proposed framework actually advances physics reasoning, or merely overfits to a custom dataset.\n\nMissing evaluation under strong “deliberate thinking” modes of frontier models\nRecent “o-series / R1-style” models (e.g., OpenAI o3-mini(o3-medium), DeepSeek-R1, Gemini-1.5-Pro “thinking mode”) have shown substantial gains on reasoning-heavy tasks simply by changing the inference-time policy (deliberate sampling / longer reasoning / reflection).\nThe paper does not evaluate how the proposed method compares under identical think-mode inference, nor whether the gains persist if SLM baselines also use deliberation. This omission weakens the claimed contribution.\n\nThe “small model deficit” argument is not convincing under current scaling dynamics\nThe motivation heavily relies on “SLMs struggle, thus SLM refinement is necessary”. But with current and near-term scaling trends (e.g., 7B on consumer hardware with KV-cache offloading, quantization, speculative decoding)，the gap between 3B and 7B inference cost is shrinking rapidly. In this context, the paper does not justify why narrowing the reasoning gap of 3B models is more important than (i) studying scaling laws for physics reasoning, or (ii) using slightly larger, but still affordable, 7B class models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921834917,"tcdate":1761377106976,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10563/Reviewer_3ydd"],"signatures":["ICLR.cc/2026/Conference/Submission10563/Reviewer_3ydd"],"forum":"DiQ8d2kOxG","number":2,"license":"CC BY 4.0","cdate":1761377106976,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10563/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921834917,"domain":"ICLR.cc/2026/Conference","replyto":"DiQ8d2kOxG","id":"hRih8pfyqy","forumContent":{"TLDR":{"value":"We introduce a refinement agent and use LoRA-based RLHF with a step-level reward model to improve reasoning using small LLM Models"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large language models","Physics Multi-step Reasoning","RLHF","LLM Agent"]},"supplementary_material":{"value":"/attachment/f77a87d6a789146f35431c9c30b602da96f147c8.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Large Language Models (LLMs) excel at many reasoning tasks but struggle with scientific domains like physics, which demand precise mathematical calculations alongside deep conceptual and factual understanding. In complex physics problem solving, LLMs commonly falter due to three core issues: misunderstanding the problem, incorrect application of concepts, and calculation mistakes. These challenges are more pronounced in small LLMs due to their limited capacity, making them more prone to failures. To address these limitations, we propose a modular reinforcement learning refinement framework tailored for small LLMs, integrating first step error localization, and correction through a Reinforcement Learning guided feedback mechanism. We also introduce PhysicsQA, a diverse benchmark of 370 physics problems designed to evaluate LLM reasoning across the aforementioned dimensions. Experimental results demonstrate improvements upto 10% in final answer accuracy reasoning using Small language models over existing approaches"},"_bibtex":{"value":"@misc{\njaiswal2026modular,\ntitle={Modular Refinement of Small Language Models for Physics Reasoning via Localized Error Feedback},\nauthor={Raj Jaiswal and Dhruv Jain and Rishabh Dhawan and Dhruvkumar Patel and Avinash Anand and Shin'ichi Satoh and Tanuja Ganu and Rajiv Ratn Shah and Erik Cambria and Zhengkui Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=DiQ8d2kOxG}\n}"},"title":{"value":"Modular Refinement of Small Language Models for Physics Reasoning via Localized Error Feedback"},"pdf":{"value":"/pdf/1cd18268e153ed7837fd6150db61d0942c901bf5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"jaiswal|modular_refinement_of_small_language_models_for_physics_reasoning_via_localized_error_feedback"},"authorids":{"value":["~Raj_Jaiswal1","~Dhruv_Jain2","~Rishabh_Dhawan1","~Dhruvkumar_Patel1","~Avinash_Anand1","~Shin'ichi_Satoh1","~Tanuja_Ganu1","~Rajiv_Ratn_Shah1","~Erik_Cambria1","~Zhengkui_Wang1"]},"authors":{"value":["Raj Jaiswal","Dhruv Jain","Rishabh Dhawan","Dhruvkumar Patel","Avinash Anand","Shin'ichi Satoh","Tanuja Ganu","Rajiv Ratn Shah","Erik Cambria","Zhengkui Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces PhyRL, a Physics-Guided Reinforcement Learning framework that improves the robustness and out-of-distribution generalization of spatio-temporal Graph Neural Networks. A Physics Reward Oracle provides rewards based on physical consistency, robustness, and uncertainty reliability, guiding the model via reinforcement learning rather than standard loss minimization. Experiments on N-body, Navier–Stokes, and sea surface temperature tasks show that PhyRL significantly enhances accuracy, stability, and uncertainty calibration, offering a new reward-driven approach to physics-informed modeling."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. In line 192, you mention computing the consistency loss by integrating the squared PDE residual over the spatio-temporal domain, with differential operators approximated on the graph. Could you provide the explicit computational formula or an example of how this residual is implemented on discrete graph structures?\n2. The reward terms (e.g., R = $\\exp(-\\lambda v)$ ) all adopt an exponential transformation. What is the rationale behind this specific functional form? Have you compared it with other mappings such as linear or logistic functions?\n3. In line 203, when defining the robustness reward, how is the perturbation noise intensity ($\\sigma$) determined? Is it fixed across datasets or tuned per task?\n4. The uncertainty reward relies on Monte Carlo Dropout, which requires dropout layers. How would this be applied to architectures like Transformers or Graph Neural Operators that lack dropout modules?\n5. During Physics-Guided Tree Search, are the multiple candidate trajectories generated through MC Dropout sampling? If so, how many samples are drawn at each step, and how sensitive is the search result to this number?\n6. There are typos in Figure 2 (“\\textbf{Ours}”), which should be corrected in the final version."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper presents a novel perspective by reframing physics-informed learning as a reinforcement learning problem, where physical laws are not imposed as soft constraints but embedded as verifiable reward signals.\n- The work addresses the limitation of current spatio-temporal GNNs—poor OOD generalization—and introduces a broadly applicable framework for improving the physical fidelity and reliability of data-driven dynamical models. \n- The paper is well-organized and clearly written, with intuitive figures (e.g., Fig. 1) that effectively illustrate the dual-phase framework."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Key implementation and training information are missing. The GNN architecture, baseline configurations, and RL training details (e.g., learning rate, steps, reward curves) are not provided. Important hyperparameters ($w_c, w_r, w_u, \\lambda_c, \\lambda_r, \\lambda_u$) and their sensitivity are also unreported, making the results less convincible.\n- The framework’s computational cost is not discussed. Reinforcement learning loops can be significantly more expensive than supervised training, and the paper provides no analysis of training efficiency or scalability.\n- Missing related work[1], which also integrates reinforcement learning into GNN-based physical simulations through adaptive remeshing.\n- Although the paper claims to release code, the provided anonymous link only contains a README without actual implementation. \n\n[1] Learning Controllable Adaptive Simulation for Multi-resolution Physics.\nTailin Wu, Takashi Maruyama, Qingqing Zhao, Gordon Wetzstein, Jure Leskovec. ICLR 2023"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942721558,"tcdate":1760608025934,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23578/Reviewer_v2Jd"],"signatures":["ICLR.cc/2026/Conference/Submission23578/Reviewer_v2Jd"],"forum":"jEgWebcmUc","number":1,"license":"CC BY 4.0","cdate":1760608025934,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23578/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942721558,"domain":"ICLR.cc/2026/Conference","replyto":"jEgWebcmUc","id":"djSrkfNk0M","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["spatio-temporal; out-of-distribution"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Modeling and predicting spatio-temporal dynamical systems are pivotal in numerous scientific and engineering domains, yet their inherent complexity and the stringent requirement for out-of-distribution (OOD) generalization pose significant challenges. While existing models based on Graph Neural Networks (GNNs) excel at capturing spatio-temporal dependencies, they often exhibit insufficient robustness and inaccurate uncertainty estimation when confronted with unseen or perturbed dynamical patterns.\nTo address this challenge, this paper proposes a novel training framework inspired by Direct Preference Optimization (DPO) to enhance the OOD generalization capabilities of Multi-scale Spatio-Temporal Graph (MSTG) models for dynamical systems. At the core of our approach is the construction of an automated \"physics preference oracle\" that leverages uncertainty estimation and system perturbations to generate paired trajectory preference data. Specifically, the model generates multiple candidate future trajectories by applying perturbations to the input or leveraging its inherent stochasticity. The oracle then automatically evaluates these trajectories and identifies \"preferred\" and \"dispreferred\" outcomes based on metrics such as physical consistency, robustness against perturbations, and the reliability of the predicted uncertainty. Using this preference dataset, we introduce a DPO-style loss function to directly optimize the MSTG model, encouraging it to favor predictions that are more consistent with physical laws, more resilient to perturbations, and provide reliable uncertainty estimates (i.e., high uncertainty) in OOD scenarios.\nThis method aims to elevate dynamical system modeling from mere data fitting to learning and internalizing the intrinsic physical properties and robust behaviors of the system. Experiments demonstrate that our proposed framework significantly improves the OOD generalization, prediction accuracy, and quality of uncertainty quantification for MSTG models in complex spatio-temporal modeling and prediction tasks. This research offers a new perspective on leveraging DPO-like reinforcement learning paradigms to tackle fundamental challenges in scientific computing."},"_bibtex":{"value":"@misc{\nzeng2026reinforcing,\ntitle={Reinforcing Spatio-Temporal Graph Neural Networks with a Physics Reward Oracle},\nauthor={Hansheng Zeng and Yuqi Li and Chuanguang Yang and Weilun Feng and Zeyu Dong and Yao Lu and Yingli Tian and Hao Wu},\nyear={2026},\nurl={https://openreview.net/forum?id=jEgWebcmUc}\n}"},"title":{"value":"Reinforcing Spatio-Temporal Graph Neural Networks with a Physics Reward Oracle"},"pdf":{"value":"/pdf/b782e29976bfa559857d02bf479aa3c9b893522e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zeng|reinforcing_spatiotemporal_graph_neural_networks_with_a_physics_reward_oracle"},"authorids":{"value":["~Hansheng_Zeng1","~Yuqi_Li7","~Chuanguang_Yang1","~Weilun_Feng2","~Zeyu_Dong2","~Yao_Lu15","~Yingli_Tian1","~Hao_Wu39"]},"authors":{"value":["Hansheng Zeng","Yuqi Li","Chuanguang Yang","Weilun Feng","Zeyu Dong","Yao Lu","Yingli Tian","Hao Wu"]}},"version":2},{"content":{"venue":{"value":"European Journal of Physics"},"pdf":{"value":"https://iopscience.iop.org/article/10.1088/1361-6404/ace748/pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"xie|making_exciton_physics_easy_and_affordable"},"html":{"value":"https://doi.org/10.1088/1361-6404/ace748"},"abstract":{"value":"Making exciton physics easy and affordable, Xie, Yong, Ersu, Gulsum, Pucher, Thomas, Kuriakose, Sruthi, Zhang, Wenliang, Al-Enizi, Abdullah M, Albrithen, Hamad A H, Nafady, Ayman, Bratschitsch, Rudolf, Island, Joshua O, Castellanos-Gomez, Andres"},"title":{"value":"Making exciton physics easy and affordable"},"authors":{"value":[{"fullname":"Yong Xie","username":"~Yong_Xie5"},{"fullname":"Gulsum Ersu","username":"https://orcid.org/orcid-search/search?searchQuery=Gulsum%20Ersu"},{"fullname":"Thomas Pucher","username":"https://orcid.org/orcid-search/search?searchQuery=Thomas%20Pucher"},{"fullname":"Sruthi Kuriakose","username":"https://orcid.org/orcid-search/search?searchQuery=Sruthi%20Kuriakose"},{"fullname":"Wenliang Zhang","username":"https://orcid.org/orcid-search/search?searchQuery=Wenliang%20Zhang"},{"fullname":"Abdullah M Al-Enizi","username":"https://orcid.org/orcid-search/search?searchQuery=Abdullah%20M%20Al-Enizi"},{"fullname":"Hamad A H Albrithen","username":"https://orcid.org/orcid-search/search?searchQuery=Hamad%20A%20H%20Albrithen"},{"fullname":"Ayman Nafady","username":"https://orcid.org/orcid-search/search?searchQuery=Ayman%20Nafady"},{"fullname":"Rudolf Bratschitsch","username":"https://orcid.org/orcid-search/search?searchQuery=Rudolf%20Bratschitsch"},{"fullname":"Joshua O Island","username":"https://orcid.org/orcid-search/search?searchQuery=Joshua%20O%20Island"},{"fullname":"Andres Castellanos-Gomez","username":"https://orcid.org/orcid-search/search?searchQuery=Andres%20Castellanos-Gomez"}]}},"tmdate":1777903806205,"pdate":1693526400000,"externalIds":["doi:10.1088/1361-6404/ace748"],"tcdate":1777903794054,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Yong_Xie5"],"forum":"Pyb63AtuwZ","license":"CC BY-SA 4.0","number":66408,"cdate":1689288634767,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1777903806205,"domain":"OpenReview.net/Public_Article","id":"Pyb63AtuwZ","version":2},{"content":{"summary":{"value":"This paper proposes a physics-based noise model for EMCCD cameras. The statistical model includes some typical noise components for EMCCD sensors, and a calibration method is proposed for adaptation this noise model on each sensor. Through careful noise modeling and calibration, the authors synthesize realistic EMCCD noise data for training, and effectively improve the learning of deep denoisiers in both macroscopic testset and microscopic testset."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Why ELD presents banding patterns in Fig. 7?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper introduces the first EMCCD denoising method utilizing physics-based noise modeling method.\n- The overall writing of this paper is good and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- This paper proposes the first noise modeling method for EMCCD sensors, and there are indeed some new adaptations on this sensor type. However, the main idea borrows many contributions from the similar task of CMOS noise modeling, and seems to be a EMCCD-version of ELD [1]. Specifically, the entire pipeline, i.e., physics-based noise modeling ->  calibration -> synthesis -> denoise pipeline is the same with ELD. The noise components and calibration process are also similar with ELD. In addition, the modeling of FPN and pre-processing operation comes from PMN [2] .\n- For Fig. 7, why ELD presents banding patterns, even after calibration using the target device? ELD calibrates row noise using bias frames, and the variance for row noise would be close to zero on sensors without obvious banding patterns if correctly calibrated. I wonder why ELD still causes such row patterns on EMCCD sensors.\n- There should be more comparisons with sota methods, for both noise modeling and self-supervised denoising methods. For example, [3] proposes a general noise modeling method which uses poisson sampling for signal-dependent noise and GAN for signal-independent noise. I think [3] can also handle EMCCD sensors. Stronger baselines for self-supervised methods are also recommended to compare [4].\n- I concern that it is not rigorous to use SID clean images to synthesize noisy pairs for training. Different from EMCCD sensors, SID dataset uses Sony cameras with CMOS sensors. Each sensor type has its own unique recipe for generating RAW data; even using clean images from one type of CMOS sensor to generate synthetic noisy pairs and then testing on real data from a different CMOS sensor can lead to negative effects, not to mention EMCCD data. Therefore, I believe that SID clean data is not a suitable choice for this application.\n- Section 2.3 is not necessary since no deep denoiser architecture is proposed.\n\n\n[1] Physics-based Noise Modeling for Extreme Low-light Photography. TPAMI 2021\n[2] Learnability Enhancement for Low-Light Raw Image Denoising- A Data Perspective. TPAMI 2023\n[3] Towards General Low-Light Raw Noise Synthesis and Modeling. ICCV 2023\n[4] Exploring Efficient Asymmetric Blind-Spots for Self-Supervised Denoising in Real-World Scenarios. CVPR 2024"}},"nonreaders":[],"tmdate":1731427816243,"tcdate":1730674373006,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4427/Reviewer_e1uJ"],"signatures":["ICLR.cc/2025/Conference/Submission4427/Reviewer_e1uJ"],"forum":"vmulbBDCan","number":3,"license":"CC BY 4.0","cdate":1730674373006,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4427/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427816243,"domain":"ICLR.cc/2025/Conference","replyto":"vmulbBDCan","id":"ns6eFh9wSU","forumContent":{"TLDR":{"value":"A novel noise model and calibration procedure for EMCCD, synthesizing authentic training data for a neural network to achieve state-of-the-art EMCCD denoising performance."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["EMCCD","physics-based noise modeling","deep high-sensitivity imaging","fluorescence microscopy image denoising"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Electron-multiplying charge-coupled device (EMCCD) has been instrumental in sensitive observations under low-light situations including astronomy, material science, and biology. \nDespite its ingenious designs to enhance target signals overcoming read-out circuit noises, produced images are not completely noise free, which could still cast a cloud on desired experiment outcomes, especially in fluorescence microscopy.\nExisting studies on EMCCD's noise model have been focusing on statistical characteristics in theory, yet unable to incorporate latest advancements in the field of computational photography, where physics-based noise models are utilized to guide deep learning processes, creating adaptive denoising algorithms for ordinary image sensors.\nStill, those models are not directly applicable to EMCCD.\nIn this paper, we intend to pioneer EMCCD denoising by introducing a systematic study on physics-based noise model calibration procedures for an EMCCD camera, accurately estimating statistical features of observable noise components in experiments, which are then utilized to generate substantial amount of authentic training samples for one of the most recent neural networks.\nA first real-world test image dataset for EMCCD is captured, containing both images of ordinary daily scenes and those of microscopic contents.\nBenchmarking upon the testset and authentic microscopic images, we demonstrate distinct advantages of our model against previous methods for EMCCD and physics-based noise modeling, forging a promising new path for EMCCD denoising."},"_bibtex":{"value":"@inproceedings{\njiang2025revolutionizing,\ntitle={Revolutionizing {EMCCD} Denoising through a Novel Physics-Based Learning Framework for Noise Modeling},\nauthor={Haiyang Jiang and Tetsuichi Wazawa and Imari Sato and Takeharu Nagai and Yinqiang Zheng},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=vmulbBDCan}\n}"},"title":{"value":"Revolutionizing EMCCD Denoising through a Novel Physics-Based Learning Framework for Noise Modeling"},"pdf":{"value":"/pdf/b0ed39c76ba54d2f078d1263ff3edaac3b15cfdb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"jiang|revolutionizing_emccd_denoising_through_a_novel_physicsbased_learning_framework_for_noise_modeling"},"authorids":{"value":["~Haiyang_Jiang1","~Tetsuichi_Wazawa1","~Imari_Sato1","~Takeharu_Nagai1","~Yinqiang_Zheng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haiyang Jiang","Tetsuichi Wazawa","Imari Sato","Takeharu Nagai","Yinqiang Zheng"]}},"version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["llm","reinforcement learning","physics simulator","llm reasoning","math"]},"_bibtex":{"value":"@inproceedings{\nprabhudesai2026simreason,\ntitle={Sim2Reason: Solving Physics Olympiad via Reinforcement Learning on Physics Simulators},\nauthor={Mihir Prabhudesai and Aryan Satpathy and Yangmin Li and Zheyang Qin and Nikash Bhardwaj and Amir Zadeh and Chuan Li and Katerina Fragkiadaki and Deepak Pathak},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=rQ7eobSCZZ}\n}"},"title":{"value":"Sim2Reason: Solving Physics Olympiad via Reinforcement Learning on Physics Simulators"},"paperhash":{"value":"prabhudesai|sim2reason_solving_physics_olympiad_via_reinforcement_learning_on_physics_simulators"},"originally_submitted_PDF":{"value":"/pdf/3b4ac9ac141ae6b43086d1ef2ee9f0ebfdba73e6.pdf"},"TLDR":{"value":"Training LLMs with RL on synthetic QA data generated from physics simulators yields strong zero-shot gains on real-world physics olympiad-level benchmarks."},"primary_area":{"value":"deep_learning->large_language_models"},"abstract":{"value":"We have witnessed remarkable advances in LLM reasoning capabilities with the advent of DeepSeek-R1. However, much of this progress has been fueled by the abundance of internet question–answer (QA) pairs—a major bottleneck going forward, since such data is limited in scale and concentrated mainly in domains like mathematics. In contrast, other sciences such as physics lack large-scale QA datasets to effectively train reasoning-capable models. In this work, we show that physics simulators can serve as a powerful alternative source of supervision for training LLMs for physical reasoning. We generate random scenes in physics engines,  create synthetic question–answer pairs from simulated interactions, and train LLMs using reinforcement learning on this synthetic data. Our models exhibit zero-shot sim-to-real transfer to real-world physics benchmarks: for example, training solely on synthetic simulated data improves performance on IPhO (International Physics Olympiad) problems by 5-10 percentage points across model sizes. These results demonstrate that physics simulators can act as scalable data generators, enabling LLMs to acquire deep physical reasoning skills beyond the limitations of internet-scale QA data. Code available at: https://sim2reason.github.io/."},"link_to_code":{"value":"https://sim2reason.github.io/"},"pdf":{"value":"/pdf/7dc6af4e63ab1eddbcafb49d28592674d6a629a2.pdf"},"lay_summary":{"value":"What if AI learned physics the way Newton did –  by experiencing it?\n\nWe built Sim2Reason: train LLMs inside virtual worlds governed by real physics laws, zero human annotation.\n\nResult: +5–10% improvement on International Physics Olympiad, zero-shot.\n\nThe core bottleneck: human-annotated textbook QA data is scarce and narrow. Less than 1% of DeepSeek-R1's 800K training pairs touch STEM. We sidestep this entirely — physics simulators already encode the laws of nature. So we just... learn from them directly.\n\nMore details along with video explanation could be found at: https://sim2reason.github.io/."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Mihir_Prabhudesai1","~Aryan_Satpathy2","~Yangmin_Li1","~Zheyang_Qin1","~Nikash_Bhardwaj1","~Amir_Zadeh1","~Chuan_Li4","~Katerina_Fragkiadaki1","~Deepak_Pathak1"]},"authors":{"value":["Mihir Prabhudesai","Aryan Satpathy","Yangmin Li","Zheyang Qin","Nikash Bhardwaj","Amir Zadeh","Chuan Li","Katerina Fragkiadaki","Deepak Pathak"]}},"tmdate":1790067428048,"pdate":1777576836188,"tcdate":1769206230366,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission28696/Authors"],"signatures":["ICML.cc/2026/Conference/Submission28696/Authors"],"forum":"rQ7eobSCZZ","license":"CC BY 4.0","number":28696,"cdate":1769206230366,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission28696/-/Full_Submission","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/Submission28696/-/Camera_Ready_Revision"],"mdate":1790067428048,"odate":1782342021091,"domain":"ICML.cc/2026/Conference","id":"rQ7eobSCZZ","version":2},{"content":{"summary":{"value":"The paper presents **STEAM**, a method for generating synthetic medical data to support causal inference while addressing privacy limitations. STEAM is designed to replicate critical aspects of real data, including covariates, treatment assignments, and outcome distributions, to enable accurate analysis of treatment effects.\n\nTo evaluate synthetic data quality, the authors introduce new metrics that assess how well the generated data supports causal inference, addressing gaps in traditional evaluation methods. STEAM also incorporates differential privacy to enhance security. Empirical results suggest that STEAM performs well in complex, high-dimensional scenarios, making it suitable for applications in healthcare and other fields requiring secure, synthetic data for causal analysis."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"The $U_{PEHE}$ metric is based on evaluating a family of CATE estimators. Could the authors clarify what size this family should ideally be for reliable estimation, and outline the computational cost involved? Additional context on this aspect would be helpful to assess the feasibility of $U_{PEHE}$ across different applications.\n\nHow might the STEAM approach be extended or adapted to support arbitrary causal graphs? Any insights on this would provide useful context for understanding the broader applicability of the method to more complex causal structures."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.**Novelty**\n  \nThe paper presents a novel approach to synthetic data generation that prioritizes causal relationships, often overlooked in traditional methods. **STEAM** addresses this by modeling covariates, treatment assignments, and outcomes, while ensuring usual Differential Privacy concerns are easily transferable.\n\n2.**Quality**\n  \nThe methodology is thorough, with clear desiderata and structured evaluation metrics. Empirical results show STEAM’s effectiveness, especially in complex, high-dimensional scenarios.\n\n3.**Clarity**\n\nThe paper is well-organized, clearly guiding the reader through the problem, methodology, and results. Each component and metric is explained with clarity.\n\n4.**Significance**\n\nBy enabling privacy-preserving, causally accurate synthetic data, this work has broad applications, particularly in healthcare and other sensitive fields. The proposed metrics and STEAM framework make synthetic data more viable for impactful, real-world research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited Accessibility of Metrics**  \n   The paper would benefit from including reminders of key equations, particularly for Jensen-Shannon Divergence (JSD), $P_\\alpha$, and $R_\\beta$. This addition would improve accessibility for non-expert readers, helping them better understand and apply the proposed evaluation metrics.\n\n2. **Lack of Comparison with Relevant Causal Generative Models**  \n   The evaluation does not include comparisons with established causal generative models that handle interventional and counterfactual data, such as DCM [1], VACA [2], and DoWhy-GCM [3]. These models account for treatment, outcome, and counterfactuals and are capable of generating data with similar causal structures. Benchmarking STEAM against these methods could provide a more comprehensive view of its relative performance and potential advantages in synthetic data generation for causal inference.\n\n3. **Omission of Closely Related Work**  \n   The paper does not sufficiently reference the reasearch area of causal generative models. Inclusing mentions of these works would strengthen the contextual background, positioning STEAM within the landscape of existing work and clarifying its contributions to the field.\n\nReferences:\n1. Blöbaum, P., Götz, P., Budhathoki, K., Mastakouri, A. A., & Janzing, D. (2024). *DoWhy-GCM: An extension of DoWhy for causal inference in graphical causal models*. [arXiv:2206.06821](https://arxiv.org/abs/2206.06821).\n2. Sanchez-Martin, P., Rateike, M., & Valera, I. (2021). *VACA: Design of Variational Graph Autoencoders for Interventional and Counterfactual Queries*. [arXiv:2110.14690](https://arxiv.org/abs/2110.14690).\n3. Chao, P., Blöbaum, P., Patel, S., & Kasiviswanathan, S. P. (2024). *Modeling Causal Mechanisms with Diffusion Models for Interventional and Counterfactual Queries*. [arXiv:2302.00860](https://arxiv.org/abs/2302.00860)."}},"nonreaders":[],"tmdate":1732527317339,"tcdate":1730714938544,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3845/Reviewer_9kc4"],"signatures":["ICLR.cc/2025/Conference/Submission3845/Reviewer_9kc4"],"forum":"lTldTFWbJ8","number":2,"license":"CC BY 4.0","cdate":1730714938544,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3845/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732527317339,"domain":"ICLR.cc/2025/Conference","replyto":"lTldTFWbJ8","id":"mGbHnCMdfe","forumContent":{"TLDR":{"value":"We identify problems with the standard synthetic data generation pipeline when producing data containing treatments, and we propose a set of metrics and a generation method which perform better in this setting."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic data","evaluation","generative models","metrics","treatment effect analysis"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Causal inference on medical data, such as estimation of treatment effects, is crucial to ensure the efficacy and safety of medical interventions. However, privacy concerns frequently limit access to the patient data necessary for such analyses. Generative models can produce synthetic data that preserves privacy and closely approximates the real data distribution, yet existing  methods are not designed for data containing treatments and the specific challenges their downstream use pose. With our work we establish a set of desiderata that synthetic data containing treatments should satisfy: preservation of (i) the covariate distribution, (ii) the treatment assignment mechanism, and (iii) the outcome generation mechanism. Based on these desiderata, we propose a principled set of evaluation metrics to assess such synthetic data. Finally, we present STEAM: a novel method for generating Synthetic data for Treatment Effect Analysis in Medicine. STEAM mimics the data-generating process of real-world data containing treatments, and can ensure differential privacy. We empirically demonstrate that STEAM achieves state-of-the-art performance across our metrics as compared to existing generative models, particularly as the complexity of the generative task increases."},"_bibtex":{"value":"@misc{\namad2025generation,\ntitle={Generation and Evaluation of Synthetic Data Containing Treatments},\nauthor={Harry Amad and Zhaozhi Qian and Dennis Frauen and Julianna Piskorz and Stefan Feuerriegel and Mihaela van der Schaar},\nyear={2025},\nurl={https://openreview.net/forum?id=lTldTFWbJ8}\n}"},"title":{"value":"Generation and Evaluation of Synthetic Data Containing Treatments"},"pdf":{"value":"/pdf/b9ae540d4531db13df09b4ee8e9e7d7794e3cc30.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"amad|generation_and_evaluation_of_synthetic_data_containing_treatments"},"authorids":{"value":["~Harry_Amad1","~Zhaozhi_Qian1","~Dennis_Frauen1","~Julianna_Piskorz1","~Stefan_Feuerriegel1","~Mihaela_van_der_Schaar2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Harry Amad","Zhaozhi Qian","Dennis Frauen","Julianna Piskorz","Stefan Feuerriegel","Mihaela van der Schaar"]}},"version":2},{"content":{"summary":{"value":"This paper proposes RareFlow, a physics-aware dual-conditioning super-resolution framework designed for cross-sensor and out-of-distribution remote sensing imagery. The method integrates a dual-conditioning architecture combining Gated ControlNet and text-guided semantic conditioning, an uncertainty-gated control mechanism using Monte Carlo dropout and a physics-aware loss. Experiments demonstrate that RareFlow significantly improves perceptual realism and physical consistency over SOTA baselines such as SeeSR, AdcSR, and ZoomLDM."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper presents a well-motivated and technically innovative framework that effectively addresses the issue of generating physically inconsistent details in out-of-distribution scientific imagery. RareFlow’s dual-conditioning architecture successfully balances structural fidelity and semantic guidance, while its uncertainty-gated control mechanism adaptively suppresses hallucinated features under high uncertainty. The introduction of a physics-aware loss formulation, which integrates spectral alignment, radiometric consistency, and perceptual quality, provides a principled bridge between visual realism and physical correctness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.While the paper repeatedly emphasizes its “physics-aware”, the methodology relies primarily on heuristic loss formulations rather than an explicit physical modeling framework. No clear theoretical connection is established between the proposed losses and physics. As a result, the term “physics-aware” feels more empirical than principled, potentially overstating the scientific rigor of the approach.\n2.Although most backbone components are frozen, the overall training pipeline still feels fragmented and under-specified, relying on several pretrained modules (VAE, SD3, ControlNet) whose interconnections are not fully transparent.\n3.Key hyperparameters (e.g., λ-weights for each loss) are not justified or tuned systematically."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918663850,"tcdate":1762085751070,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6375/Reviewer_jpJK"],"signatures":["ICLR.cc/2026/Conference/Submission6375/Reviewer_jpJK"],"forum":"9iAWnRGTAU","number":3,"license":"CC BY 4.0","cdate":1762085751070,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6375/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918663850,"domain":"ICLR.cc/2026/Conference","replyto":"9iAWnRGTAU","id":"dwkfdHjAS2","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We present RareFlow, a physics-aware SR framework designed for OOD robustness."},"keywords":{"value":["satellite images","remote sensing","diffusion models"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Super-resolution (SR) for remote sensing imagery often fails under out-of-distribution (OOD) conditions, such as rare geomorphic features captured by diverse sensors, producing visually plausible but physically inaccurate results.We present RareFlow, a physics-aware SR framework designed for OOD robustness. RareFlow's core is a dual-conditioning architecture. A Gated ControlNet preserves fine-grained geometric fidelity from the low-resolution input, while textual prompts provide semantic guidance for synthesizing complex features. To ensure physically sound outputs, we introduce a multifaceted loss function that enforces both spectral and radiometric consistency with sensor properties. Furthermore, the framework quantifies its own predictive uncertainty by employing a stochastic forward pass approach; the resulting output variance directly identifies unfamiliar inputs, mitigating feature hallucination.We validate RareFlow on a new, curated benchmark of multi-sensor satellite imagery. In blind evaluations, geophysical experts rated our model's outputs as approaching the fidelity of ground truth imagery, significantly outperforming state-of-the-art baselines. This qualitative superiority is corroborated by quantitative gains in perceptual metrics, including a nearly 40\\% reduction in FID. RareFlow provides a robust framework for high-fidelity synthesis in data-scarce scientific domains and offers a new paradigm for controlled generation under severe domain shift."},"_bibtex":{"value":"@misc{\nfallah2025rareflow,\ntitle={RareFlow: Physics-Aware Flow-Matching for Cross-Sensor Super-Resolution of Rare-Earth Features},\nauthor={Forouzan Fallah and Wenwen Li and Chia-Yu Hsu and Hyunho Lee and Yezhou Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=9iAWnRGTAU}\n}"},"title":{"value":"RareFlow: Physics-Aware Flow-Matching for Cross-Sensor Super-Resolution of Rare-Earth Features"},"pdf":{"value":"/pdf/8109a4af23fac37fb408d7f598f88ffd279e1e02.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"fallah|rareflow_physicsaware_flowmatching_for_crosssensor_superresolution_of_rareearth_features"},"authorids":{"value":["~Forouzan_Fallah1","~Wenwen_Li1","~Chia-Yu_Hsu2","~Hyunho_Lee5","~Yezhou_Yang1"]},"authors":{"value":["Forouzan Fallah","Wenwen Li","Chia-Yu Hsu","Hyunho Lee","Yezhou Yang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces the Synthetic Dataset Quality Estimation (SYNQUE) framework, which addresses the problem of ranking synthetic datasets by their expected real-world task performance using only a small set of unlabelled real samples. The authors propose two types of proxy metrics to estimate dataset quality without training task-specific models. The first category includes representation-based proxies that adapt traditional distributional and diversity measures such as MMD, PAD, and MAUVE to an embedding space. The second, called LENS (LLM-Evaluated Normalized Score), is an LLM-based approach that employs rubric-guided reasoning to compare synthetic and real data directly in natural language. Experiments across diverse domains, including sentiment analysis, Text2SQL, image classification, and web navigation, show that these proxies correlate well with downstream task performance."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* Could you provide more discussion or analysis on how sensitive the representation-based proxies are to the choice of encoder? For example, have you compared different embedding models (e.g., smaller Qwen variants or open-source encoders) to assess whether the correlations remain stable?\n\n\nOverall, the paper is conceptually interesting, but the justification of key assumptions and the empirical analysis could be strengthened to make the contribution more convincing. I encourage the authors to address the above concerns and clarify these points in their rebuttal."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"* The main contribution of the paper is the introduction of a novel and practically relevant problem setting, Synthetic Dataset Quality Estimation (SYNQUE), which focuses on evaluating and ranking synthetic datasets using only unlabelled real data. By formulating the problem without using labelled data, the approach substantially reduces computational cost, avoiding repeated model training while still providing informative quality estimates.\n\n* The paper proposes several representation-based proxy metrics that estimate data quality through measures of diversity and distributional alignment, without requiring labelled real samples or downstream model training.\n\n* The authors further present LENS, an LLM-based evaluation framework that introduces principled debiasing strategies to mitigate order bias, label bias, and score bias, leading to more consistent and interpretable LLM-based judgments.\n\n* It addresses a timely and important challenge in understanding and quantifying the quality of synthetic data, which is increasingly critical for large-scale model development.\n\n* The paper provides comprehensive experimental validation across diverse domains, showing that the proposed proxies correlate strongly with downstream task performance, with LENS achieving the most reliable results on complex and long-horizon tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* **Framing and originality could be clarified.**  \n  While the problem setting is interesting, the distinction between synthetic dataset quality estimation and general dataset quality estimation remains under-specified. The paper does not fully explain what makes evaluating synthetic data uniquely challenging beyond the label-free constraint. A more explicit comparison with recent works such as [1], [2], and [3] would help position SYNQUE within the broader landscape of data selection and importance estimation. In particular, [2] also studies quality estimation for LLM-generated data and highlights the gap between synthetic and real data, which could provide valuable context for the present work.\n\n* **Limited theoretical motivation for proxy metrics.**  \n  The paper provides intuitive explanations for why metrics such as MMD², PAD, and MAUVE may correlate with downstream performance, but lacks theoretical justification or formal conditions under which these proxies should succeed or fail. Similarly, the rationale for why LENS’s rubric-based reasoning captures true data–distribution similarity remains largely empirical.\n\n* **Quality–diversity balance is not explicitly addressed.**  \n  Each proposed proxy primarily captures either data quality (alignment) or diversity (coverage), but the paper does not present a unified formulation that balances the two. Prior work [1, 2] has shown that maintaining both quality and diversity is crucial for effective data selection, especially for synthetic data. Discussing or analysing this trade-off would strengthen the contribution.\n\n* **Label-quality robustness is not evaluated.**  \n  Although the framework avoids using labels for real data, it still depends on synthetic data labels that may be noisy or hallucinatory. The method implicitly assumes that these synthetic labels are of reasonable quality and aligned with the downstream task, since all proposed proxies rely solely on input features rather than label consistency. It would be valuable to discuss the performance gap between the proposed label-free setting and small labelled baselines, as well as the impact of varying levels of label noise in the synthetic data.\n\n* **Inter-dataset dependency is not discussed.**  \n  The framework assumes that each synthetic dataset can be evaluated independently, yet in practice, different synthetic datasets may share overlapping samples or originate from similar generative prompts or models. Such dependencies could bias the proxy correlations and ranking results, especially when datasets are not mutually independent. It would be useful to discuss how SYNQUE handles or mitigates inter-dataset overlap and whether the proposed proxies remain valid under such dependencies.\n\n[1] *Harnessing Diversity for Important Data Selection in Pretraining Large Language Models* (ICLR 2025)\n\n[2] *Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification* (ICLR 2025)\n\n[3] *Most Influential Subset Selection: Challenges, Promises, and Beyond* (NeurIPS 2024)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925820988,"tcdate":1761490881381,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15544/Reviewer_mUYC"],"signatures":["ICLR.cc/2026/Conference/Submission15544/Reviewer_mUYC"],"forum":"QAOLaVXiLg","number":1,"license":"CC BY 4.0","cdate":1761490881381,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15544/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925820988,"domain":"ICLR.cc/2026/Conference","replyto":"QAOLaVXiLg","id":"awo2kURzfS","forumContent":{"TLDR":{"value":"We introduce the SynQuE problem: estimating quality of synthetic datasets anchored on real unannotated data alongside baselines and our novel LENS metric."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Evaluation","Synthetic Data"]},"supplementary_material":{"value":"/attachment/75bbe9f91946deb8f844ebdfdf7c7732cb00f92e.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"We introduce and formalize the Synthetic Dataset Quality Estimation (SYNQUE) problem: ranking synthetic datasets by their expected real-world task performance using only limited unannotated real data.\nThis addresses a critical and open challenge where data is scarce due to collection costs or privacy constraints.\nWe establish the first comprehensive benchmarks for this problem by introducing and evaluating proxy metrics that choose synthetic data for training to maximize task performance on real data.\nWe introduce the first proxy metrics for SYNQUE by adapting distribution and diversity-based distance measures to our context via embedding models.\nTo address the shortcomings of these metrics on complex planning tasks, we propose SYNQUE, a novel proxy that leverages large language model reasoning.\nOur results show that SYNQUE proxies correlate with real task performance across diverse tasks, including sentiment analysis, Text2SQL, web navigation, and image classification, with LENS consistently outperforming others on complex tasks by capturing nuanced characteristics.\nFor instance, on text-to-SQL parsing, training on the top-3 synthetic datasets selected via SYNQUE proxies can raise accuracy from 30.4\\% to 38.4 (+8.1)\\% on average compared to selecting data indiscriminately.\nThis work establishes SYNQUE as a practical framework for synthetic data selection under real-data scarcity and motivates future research on foundation model-based data characterization and fine-grained data selection."},"_bibtex":{"value":"@misc{\nchen2026synque,\ntitle={{SYNQUE}: Estimating Synthetic Dataset Quality Without Annotations},\nauthor={Arthur Chen and Victor Zhong},\nyear={2026},\nurl={https://openreview.net/forum?id=QAOLaVXiLg}\n}"},"title":{"value":"SYNQUE: Estimating Synthetic Dataset Quality Without Annotations"},"pdf":{"value":"/pdf/75ff8dca1a28785d783c9597fed79fc19be7181c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"chen|synque_estimating_synthetic_dataset_quality_without_annotations"},"authorids":{"value":["~Arthur_Chen2","~Victor_Zhong1"]},"authors":{"value":["Arthur Chen","Victor Zhong"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a physics-informed autoregressive network (PIANO) for solving time-dependent PDEs by conditioning future predictions on previously computed states. The approach reformulates PINNs to include autoregression, aiming to improve temporal stability and accuracy. The authors present theoretical analysis suggesting that standard PINNs are temporally unstable and validate their method on canonical PDE benchmarks and a weather forecasting task. The results indicate performance gains compared to existing non-autoregressive PINN variants."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How does PIANO differ fundamentally from existing autoregressive or recurrent PINN formulations such as those in [1, 2, 3, 4, 5]?\n\n2. Comparisons against other autoregressive physics-informed frameworks, rather than only non-autoregressive baselines, would be beneficial.\n\n3. How does the model perform on strongly nonlinear or chaotic PDEs?\n\n4. Is the method limited to uniform Cartesian grids, or can it be generalized to irregular domains and complex geometries?\n\n5. It would be beneficial to clarify how boundary and initial conditions are integrated into the loss, and why certain terms (such as the initial condition) appear to be omitted.\n\n6. What is the computational overhead of the autoregressive rollout compared to traditional PINNs or neural operators?\n\n7. Given the similarities with neural ODEs or recurrent sequence models, please discuss the conceptual differences that can improve temporal stability.\n\n8. Has the impact of gradient accumulation over time on training stability and efficiency been examined for long sequences?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses the issue of temporal instability in PINNs by proposing a physics-informed autoregressive formulation.\n\n2. It presents both theoretical analysis and empirical demonstrations across PDE benchmarks and real-world weather forecasting.\n\n3. The proposed architecture can potentially be integrated into various PINN setups."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The literature review does not adequately position PIANO among previous autoregressive PINN variants. Several related works using similar recurrent or sequence-based formulations are missing, which weakens the novelty claim.\n\n2. Comparisons are mostly against non-autoregressive baselines. Including autoregressive approaches such as [1, 2, 3, 4, 5] would provide a fairer benchmark.\n\n3. While the PDE benchmarks are relevant, the method has not been validated on more complex, nonlinear, or higher-order PDEs, which makes it difficult to ascertain its robustness and generality.\n\n4. It is unclear whether the proposed model can handle irregular or non-Cartesian geometries, as all experiments appear to be conducted on fixed, uniform grids.\n\n5. The connection of PIANO to neural ODE and recurrent neural PDE solvers is evident, but this similarity is not discussed, giving the impression of overlap rather than a clear methodological distinction.\n\n6. The paper is difficult to read, as the authors introduce multiple terminologies (PIEL, PIANO), which can be confusing.\n\n7. The paper does not clearly specify what exactly is minimized, whether it is energy, the PDE residual, or another form of physics-constrained objective, and how boundary and initial conditions are incorporated.\n\n8. Additional clarification on the autoregressive loss design, gradient stability over long rollouts, and computational scaling would help substantiate the claimed advantages.\n\n[1] Lippe, Phillip, et al. \"Pde-refiner: Achieving accurate long rollouts with neural pde solvers.\" Advances in Neural Information Processing Systems 36 (2023): 67398-67433.\n\n[2] Kapoor, Taniya, et al. \"Neural oscillators for generalization of physics-informed machine learning.\" Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 38. No. 12. 2024.\n\n[3] Bergamin, Federico, et al. \"Guided autoregressive diffusion models with applications to PDE simulation.\" ICLR 2024 Workshop on AI4DifferentialEquations In Science. 2024.\n\n[4] Michałowska, Katarzyna, et al. \"Neural operator learning for long-time integration in dynamical systems with recurrent neural networks.\" 2024 International Joint Conference on Neural Networks (IJCNN). IEEE, 2024.\n\n[5] Koehler, Felix, Simon Niedermayr, and Nils Thuerey. \"APEBench: A benchmark for autoregressive neural emulators of PDEs.\" Advances in Neural Information Processing Systems 37 (2024): 120252-120310."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762939028810,"tcdate":1761989419824,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20928/Reviewer_d19m"],"signatures":["ICLR.cc/2026/Conference/Submission20928/Reviewer_d19m"],"forum":"9y2IyqaWxs","number":5,"license":"CC BY 4.0","cdate":1761989419824,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20928/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762939028810,"domain":"ICLR.cc/2026/Conference","replyto":"9y2IyqaWxs","id":"OYjmyQVbfk","forumContent":{"TLDR":{"value":"PIANO is a novel physics-informed autoregressive framework for time-dependent PDEs that conditions each prediction on prior states, improving stability and accuracy over existing methods."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Physics-Informed Neural Networks","Autoregressive Networks"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Solving time-dependent partial differential equations (PDEs) is fundamental to modeling critical phenomena across science and engineering. Physics-Informed Neural Networks (PINNs) solve PDEs using deep learning. However, PINNs perform pointwise predictions that neglect the autoregressive property of dynamical systems, leading to instabilities and inaccurate predictions. We introduce Physics-Informed Autoregressive Networks (PIANO)---a framework that redesigns PINNs to model dynamical systems. PIANO operates autoregressively, explicitly conditioning future predictions on the past. It is trained through a self-supervised rollout mechanism while enforcing physical constraints. We present a rigorous theoretical analysis demonstrating that PINNs suffer from temporal instability, while PIANO achieves stability through autoregressive modeling. Extensive experiments on challenging time-dependent PDEs demonstrate that PIANO achieves state-of-the-art performance, significantly improving accuracy and stability over existing methods. We further show that PIANO outperforms existing methods in weather forecasting."},"_bibtex":{"value":"@misc{\nnagda2026piano,\ntitle={{PIANO}: Physics-Informed Autoregressive Networks},\nauthor={Mayank Nagda and Jephte Abijuru and Phil Ostheimer and Stephan Mandt and Marius Kloft and Sophie Fellenz},\nyear={2026},\nurl={https://openreview.net/forum?id=9y2IyqaWxs}\n}"},"title":{"value":"PIANO: Physics-Informed Autoregressive Networks"},"pdf":{"value":"/pdf/a50c7f3c76f822e9166a0518472ea5af202eedcf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"nagda|piano_physicsinformed_autoregressive_networks"},"authorids":{"value":["~Mayank_Nagda1","~Jephte_Abijuru1","~Phil_Ostheimer1","~Stephan_Mandt1","~Marius_Kloft1","~Sophie_Fellenz1"]},"authors":{"value":["Mayank Nagda","Jephte Abijuru","Phil Ostheimer","Stephan Mandt","Marius Kloft","Sophie Fellenz"]}},"version":2},{"content":{"summary":{"value":"This manuscript challenges gradient descent optimization techniques. Motivated by physics, the authors present VRAdam, which automatically controls the learning rate based on the momentum. Experiments on CIFAR, Wikitext, GFlowNet, and GPT-2 demonstrate improvements."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses above. My score is based on the assumption that all typos are corrected in the revised manuscript."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The motivation, especially the physics-inspired optimizer, is interesting. Indeed, the momentum optimizer itself is derived from Newtonian dynamics.\n- The theoretical analysis looks solid and promising for the improved convergence.\n- Source code is available, which eases deployment in practice."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The norm of $v_t$ of the VRAdam looks to compute the global norm, but it may be appropriate to use the parameter-wise norm to allow parameter-wise learning rate control. This choice is not sufficiently discussed.\n- VRAdam brings three hyperparameters of \\alpha_0, \\alpha_1, and \\beta_3. I think these additional hyperparameters make it difficult to adopt the VRAdam in practice.\n- Accuracy of 80% for ResNet-32 with CIFAR-10 is a weak baseline.\n- Table 1 is not convincing enough. These results should be supplemented with more quantitative and qualitative analysis.\n- The experimental results, such as Table 2, are focused on the final validation loss. Is it possible to demonstrate other indices, such as practical ones? I think certain practitioners may want to capture the performance more practically, but the value of loss is difficult to understand on an absolute scale. It is also difficult to understand whether it corresponds to sufficient convergence or is still far from convergence.\n- LLM results were only trained for 2 epochs, which I think is insufficient for convergence.\n- I think Eq. 36 would be -m/4 + O(\\lambda), not -3m \\lambda/4 + O(\\lambda^2). Is it possible to provide an exact derivation?\n- Writing should be improved.\n    - Kingma & Ba (2017) → Kingma & Ba (2015)\n    - “the, global” → “the global”\n    - “d” → “(d)” for the caption of Figure 2.\n    - “physical inspired” → “physics-inspired” at Line 70.\n    - “Note, that” → “Note that” at Line 784.\n- The manuscript writes to compute velocity norm, whereas the source code computes gradient norm by default (normgrad=True). To be compatible with the description in this manuscript, normgrad=False may be correct for default."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941751985,"tcdate":1761813058885,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21408/Reviewer_SPPq"],"signatures":["ICLR.cc/2026/Conference/Submission21408/Reviewer_SPPq"],"forum":"6BhduwrCp3","number":4,"license":"CC BY 4.0","cdate":1761813058885,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21408/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941751985,"domain":"ICLR.cc/2026/Conference","replyto":"6BhduwrCp3","id":"AZMlmX1MRj","forumContent":{"TLDR":{"value":"We introduce Velocity-Regularized Adam (**VRAdam**), a velocity penalizing optimizer for deep neural networks."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Optimization in deep learning","physics-inspired","edge of stability"]},"supplementary_material":{"value":"/attachment/468e70955b1c75bdc4a0a86debc3c34f97ba565d.zip"},"primary_area":{"value":"optimization"},"abstract":{"value":"We introduce Velocity-Regularized Adam (VRAdam), a physics-inspired optimizer for training deep neural networks that draws on ideas from quartic terms for kinetic energy with its stabilizing effects on various system dynamics. \nPrevious algorithms, including the ubiquitous Adam, operate at the so-called adaptive edge of stability regime during training, leading to rapid oscillations and slowed convergence of loss.\nHowever, VRAdam adds a higher order penalty on the learning rate based on the velocity such that the algorithm automatically slows down whenever weight updates become large. In practice, we observe that the effective dynamic learning rate shrinks in high-velocity regimes, and damping oscillations. By combining this velocity‑based regularizer for global damping with Adam’s per‑parameter scaling, we create a powerful hybrid optimizer. For this optimizer, we provide rigorous theoretical analysis of operation at the edge of stability from a physical and control perspective for the momentum. Furthermore, we derive convergence bounds with the rate $\\mathcal{O}(\\ln(N)/\\sqrt{N})$ for a stochastic non‑convex objective under mild assumptions. We demonstrate that VRAdam exceeds the performance against standard optimizers including AdamW. We benchmark various tasks such as image classification, language modeling, and generative modeling using diverse architectures and training methodologies including Convolutional Neural Networks (CNNs), Transformers, and GFlowNets."},"_bibtex":{"value":"@inproceedings{\nvaidhyanathan2026a,\ntitle={A Physics-Inspired Optimizer: Velocity Regularized Adam},\nauthor={Pranav Vaidhyanathan and Lucas Schorling and Natalia Ares and Michael A Osborne},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6BhduwrCp3}\n}"},"title":{"value":"A Physics-Inspired Optimizer: Velocity Regularized Adam"},"pdf":{"value":"/pdf/6faf88b351b5b4e9db79349ea4e560efd40f82d0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"vaidhyanathan|a_physicsinspired_optimizer_velocity_regularized_adam"},"authorids":{"value":["~Pranav_Vaidhyanathan1","~Lucas_Schorling1","~Natalia_Ares1","~Michael_A_Osborne1"]},"authors":{"value":["Pranav Vaidhyanathan","Lucas Schorling","Natalia Ares","Michael A Osborne"]}},"version":2},{"content":{"summary":{"value":"This paper presents InPhyRe, a benchmark designed to evaluate the inductive physical reasoning of Large Multimodal Models. By constructing synthetic visual scenarios that intentionally violate universal physical laws, InPhyRe provides a method to quantitatively measure a model's ability to infer and apply novel physical principles from in-context visual examples. The benchmark is used to evaluate over LMMs on their parametric knowledge, inductive reasoning, and susceptibility to language bias. The results show that current models struggle significantly on InPhyRe, demonstrating weak inductive capabilities and a strong reliance on language cues over genuine visual understanding."},"soundness":{"value":1},"confidence":{"value":2},"questions":{"value":"The authors argue that \"inductive physical reasoning, is a hallmark of intelligence that humans develop at a very young age.\" Previous experiments on inductive and physical reasoning in infants were mostly conducted in real-world environments and did not involve scenarios that violate real-world physics. Could the authors provide further human studies to support this claim?\n\nI'll be glad to increase my score if the concerns above are properly resolved."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Clear definition and writing. The paper is clearly written and very effectively identifies the current shortcoming of current LLM/VLMs in inductive physical reasoning.\n2. The authors conducted comprehensive experiments on curent LMMs.\n3. I think the paper overall challenges the in-context viusal reasoning capabilities of LMMs, pushing the community to develop more adversarial and out-of-distribution benchmarks to probe the true limits of these models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. My biggest concern is the positioning of this paper. The authors select examples that violate physical principles in order to avoid being covered by the LLM's pretrained parametric knowledge. However, I think the model's weak performance in such an environment only indicates that it is dominated by a strong physics prior, not necessarily that its inductive reasoning ability is weak. To the best of my knowledge, experiments related to the violation-of-expectation paradigm in psychology [1] have shown that human infants, when encountering phenomena that appear to violate their expectations, also tend to assume their internal world model is correct, rather than overturning their established mental simulation model based on a few sampled trajectories.\n\n2. There is a huge gap between the benchmark's design and the real world applicability. Real-world scenarios, such as autonomous driving on diverse ground materials mentioned, are still physically consistent and will not feature impossible events like violation of object permanence.\n\n3. I think the dataset design is a bit simple. The dataset only includes basic geometric shapes on a flat surface, meaning the model does not need to handle any real-world visual challenges such as lighting and shadows, occlusion, complex backgrounds, or material reflections. Have the authors considered including more diverse object materials (e.g., soft bodies/cloth) and richer shapes?\n\n4. The paper lacks further experiments to investigate the cause of this phenomenon, such as visualizing specific attention maps or applying linear probing to inner representations. \n\n[1] Piloto LS, et al. Intuitive physics learning in a deep-learning model inspired by developmental psychology. Nature Human Behaviour, 2022."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924868503,"tcdate":1760683423805,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14468/Reviewer_q4tW"],"signatures":["ICLR.cc/2026/Conference/Submission14468/Reviewer_q4tW"],"forum":"IIrPoZ28dN","number":2,"license":"CC BY 4.0","cdate":1760683423805,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14468/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924868503,"domain":"ICLR.cc/2026/Conference","replyto":"IIrPoZ28dN","id":"IvGTPyrkDR","forumContent":{"TLDR":{"value":"We evaluate whether large multimodal models can infer and apply unseen physical laws from demonstration samples."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["large multimodal models","physical reasoning","inductive reasoning"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large multimodal models (LMMs) encode universal physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potential collision event from visual input. However, since parametric knowledge includes only the physical laws seen during training, it is insufficient for reasoning when the inference scenario follows physical laws unseen during training. In contrast, humans can adapt their physical reasoning to unseen physical environments with only a few visual examples. This inductive physical reasoning ability is indispensable for LMMs if they are to replace human agents in safety-critical applications. Despite its importance, existing visual benchmarks evaluate only the parametric knowledge in LMMs, and not inductive physical reasoning. To this end, we propose InPhyRe, the first visual question answering benchmark to measure inductive physical reasoning in LMMs. InPhyRe evaluates LMMs on their ability to predict the outcome of collision events in algorithmically generated synthetic videos. By inspecting over 13 open-source and proprietary LMMs, InPhyRe informs us that (1) LMMs struggle to apply their limited parametric knowledge about universal physical laws to reasoning, (2) inductive physical reasoning in LMMs is weak when inference scenarios obey physical laws unseen during training, and (3) inductive physical reasoning in LMMs suffers from language bias and largely ignores the visual inputs, questioning the trustworthiness of LMMs regarding visual inputs."},"_bibtex":{"value":"@misc{\nsreekumar2026inphyre,\ntitle={InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning},\nauthor={Gautam Sreekumar and Vishnu Boddeti},\nyear={2026},\nurl={https://openreview.net/forum?id=IIrPoZ28dN}\n}"},"title":{"value":"InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning"},"pdf":{"value":"/pdf/b87a0e49daf5f5f05caff235d97e823b17b5dbfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sreekumar|inphyre_discovers_large_multimodal_models_struggle_in_inductive_physical_reasoning"},"authorids":{"value":["~Gautam_Sreekumar1","~Vishnu_Boddeti1"]},"authors":{"value":["Gautam Sreekumar","Vishnu Boddeti"]}},"version":2},{"content":{"venue":{"value":"SIGGRAPH 2024"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"yao|moconvq_unified_physicsbased_motion_control_via_scalable_discrete_representations"},"authorids":{"value":["~Heyuan_Yao1","~Zhenhua_Song1","yuyangzhou2002@gmail.com","~Tenglong_Ao1","~Baoquan_Chen1","~Libin_Liu1"]},"html":{"value":"https://arxiv.org/pdf/2310.10198"},"abstract":{"value":"In this work, we present MoConVQ, a novel unified framework for physics-based motion control leveraging scalable discrete representations. Building upon vector quantized variational autoencoders (VQ-VAE) and model-based reinforcement learning, our approach effectively learns motion embeddings from a large, unstructured dataset spanning tens of hours of motion examples. The resultant motion representation not only captures diverse motion skills but also offers a robust and intuitive interface for various applications. We demonstrate the versatility of MoConVQ through several applications: universal tracking control from various motion sources, interactive character control with latent motion representations using supervised learning, physics-based motion generation from natural language descriptions using the GPT framework, and, most interestingly, seamless integration with large language models (LLMs) with in-context learning to tackle complex and abstract tasks."},"title":{"value":"MoConVQ: Unified physics-based motion control via scalable discrete representations"},"authors":{"value":["Heyuan Yao","Zhenhua Song","Yuyang Zhou","Tenglong Ao","Baoquan Chen","Libin Liu"]}},"tmdate":1790500635909,"pdate":1721318400000,"tcdate":1790500635909,"writers":["~Heyuan_Yao1","~Zhenhua_Song1","yuyangzhou2002@gmail.com","~Tenglong_Ao1","~Baoquan_Chen1","~Libin_Liu1"],"signatures":["~Libin_Liu1"],"forum":"ECN9BFvYRp","license":"arXiv.org perpetual, non-exclusive license","number":54949,"cdate":1790500635909,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1790500635909,"domain":"OpenReview.net/Archive","id":"ECN9BFvYRp","version":2},{"content":{"venue":{"value":"ACM Trans. Graph. 2024"},"venueid":{"value":"dblp.org/journals/TOG/2024"},"paperhash":{"value":"yao|moconvq_unified_physicsbased_motion_control_via_scalable_discrete_representations"},"authorids":{"value":["~Heyuan_Yao1","https://dblp.org/search/pid/api?q=author:Zhenhua_Song:","https://dblp.org/search/pid/api?q=author:Yuyang_Zhou:","https://dblp.org/search/pid/api?q=author:Tenglong_Ao:","https://dblp.org/search/pid/api?q=author:Baoquan_Chen:","https://dblp.org/search/pid/api?q=author:Libin_Liu_0002:"]},"html":{"value":"https://doi.org/10.1145/3658137"},"_bibtex":{"value":"@article{DBLP:journals/tog/YaoSZACL24,\n  author={Heyuan Yao and Zhenhua Song and Yuyang Zhou and Tenglong Ao and Baoquan Chen and Libin Liu},\n  title={MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete Representations},\n  year={2024},\n  month={July},\n  cdate={1719792000000},\n  journal={ACM Trans. Graph.},\n  volume={43},\n  number={4},\n  pages={144:1-144:21},\n  url={https://doi.org/10.1145/3658137}\n}\n"},"abstract":{"value":"In this work, we present MoConVQ, a novel unified framework for physics-based motion control leveraging scalable discrete representations. Building upon vector quantized variational autoencoders (VQ-VAE) and model-based reinforcement learning, our approach effectively learns motion embeddings from a large, unstructured dataset spanning tens of hours of motion examples. The resultant motion representation not only captures diverse motion skills but also offers a robust and intuitive interface for various applications. We demonstrate the versatility of MoConVQ through several applications: universal tracking control from various motion sources, interactive character control with latent motion representations using supervised learning, physics-based motion generation from natural language descriptions using the GPT framework, and, most interestingly, seamless integration with large language models (LLMs) with in-context learning to tackle complex and abstract tasks."},"title":{"value":"MoConVQ: Unified Physics-Based Motion Control via Scalable Discrete Representations"},"authors":{"value":["Heyuan Yao","Zhenhua Song","Yuyang Zhou","Tenglong Ao","Baoquan Chen","Libin Liu"]}},"tmdate":1741156732993,"pdate":1704067200000,"tcdate":1741156729334,"writers":["~"],"signatures":["~Heyuan_Yao1"],"forum":"UQQbJbd3vx","license":"CC BY-SA 4.0","number":353136,"cdate":1719792000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741156732993,"domain":"DBLP.org","id":"UQQbJbd3vx","version":2},{"content":{"summary":{"value":"odelbench proposes an end-to-end benchmark to assess the capability of AI systems in \nextracting executable physics-based models from scientific papers. Each task is provided with \npaper context and data; systems must implement a physics model in Python, fit parameters \nunder physical constraints, and produce metrics MSE, R^2, and a comparison plot. Evaluation \nleverages expert gold models to derive a hierarchical weighted binary rubric covering physics \ncorrectness, completeness (artifact and executability), and reproduction quality (fit). Baselines \nwith general LLMs show that code can often run, while adherence to physics or constraints and \nreproducibility often fail, which motivates this benchmark as a standardized way of tracking \nscientific modeling capability."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. How will the rubric accept multiple valid methods without false negatives? Is there a \nsystem to handle this case? \n2. How consistent are your graders? Can you share how often different judges agree when \nscoring (even a simple % would help)? \n3. If there are multiple correct ways to model the same system, how do you avoid \npenalizing a valid but different solution? \n4. Does adding a short “planning step” (write the physics assumptions first, code second) \nimprove scores?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"On soundness, the problem is well-motivated; the framework (inputs, required artifacts, rubric, scoring) is \nspecified clearly. The rubric design and constraint-aware fitting are reasonable and grounded in \ndomain principles. Baseline results (with distributions and variability) support the central claims \nabout current LLM limitations. Threats to validity (judge variance, domain breadth) are \nacknowledged with concrete mitigation/roadmap. \n\nOn presentation, writing is clear; the pipeline/rubric figures communicate the workflow; the ring-resonator \nexample makes the abstraction level concrete. Prior work is positioned well (PaperBench, \nHumanEval/SWE-bench, PINNs, LLM-as-judge). \n\nOn contribution, the benchmark targets an important, under-served evaluation capability: literature-to-model \nwith physical constraints and reproducibility. Using gold models to auto-derive rubrics is a \nuseful, reproducible idea."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"On soundness, baselines are limited in diversity and ablations (e.g., planning vs. no-planning, different optimizers). \n\nOn presentation, the paper could be improved by a tighter, tabular summary of rubric categories/weights across several tasks and a short “failure gallery” with side-by-side artifacts. \n\nOn contribution, significance is currently bottlenecked by domain breadth (20 \nphotonics tasks) and limited baseline analysis, but the design is extensible and the contribution \nis likely valuable to ICLR.\n\nMore specifically,\n\n- Domain scope. The initial dataset is narrow (photonics, 20 tasks); generalization to \nother physics/engineering areas is mentioned but I don’t see it in the writing. \n- Judge reliability. LLM-as-judge is used as a yes/no framework; while variance is \nacknowledged, it’s not clear how we can track the variance of an LLM’s output across \nmany runs. \n- Baselines: Limited baseline diversity and missing ablations (e.g., planning step, \nconstraint reparameterizations, optimizer/backends, data digitization noise). \n- Multiple-valid-solutions. Many physics problems admit non-unique but valid \nsolutions; there isn’t anything included for field aware answers."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920984807,"tcdate":1761966163729,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9367/Reviewer_BZeV"],"signatures":["ICLR.cc/2026/Conference/Submission9367/Reviewer_BZeV"],"forum":"Bw9LCBz9KW","number":3,"license":"CC BY 4.0","cdate":1761966163729,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9367/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920984807,"domain":"ICLR.cc/2026/Conference","replyto":"Bw9LCBz9KW","id":"PvWEmM58mc","forumContent":{"TLDR":{"value":"ModelBench is a benchmark for testing whether AI systems can read physics papers and produce executable, physics-based models."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Scientific AI benchmarks","Physics","LLM-as-judge","Rubric-based evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We introduce **ModelBench**, a benchmark for evaluating whether AI systems can extract\nexecutable physics-based models from scientific literature. ModelBench couples\n(i) gold-standard reference models,\n(ii) a hierarchical, weighted binary rubric covering physics correctness, code quality, and reproduction quality, and\n(iii) a judge protocol that produces pass/fail scores at rubric leaves.\nUnlike code-generation benchmarks that test function-level correctness, ModelBench\ntargets the end-to-end task of reconstructing physically grounded models from incomplete and underspecified scientific descriptions.\nWe release the benchmark specification, rubric generator and judge prompts,\nand an initial set of 20 gold models within the field of photonic integrated circuits, alongside scripts for fully reproducible evaluation.\nCandidate systems are required to produce a Python implementation of the model,\na plot of the fitted results, and evaluate MSE and $R^2$ metric of the fit.\nUsing general-purpose LLMs as neutral baselines, we report aggregate scores and case studies that reveal common failure modes\n(e.g., constraint violations, phenomenological overfitting) and show how rubric structure aids in diagnostic evaluation.\nWe discuss limitations (judge variance, dataset breadth, implicit-knowledge gaps) and outline a roadmap to expand domains,\ntighten constraint checking, and support multiple valid solutions. ModelBench provides a transparent platform\nfor tracking scientific modeling capabilities in AI under physical and empirical constraints."},"_bibtex":{"value":"@misc{\nschoolkate2026modelbench,\ntitle={ModelBench: A Benchmark for Extracting Executable, Physics-Based Models from Scientific Literature},\nauthor={Pim Schoolkate and Patrick Huembeli and Frank Sch{\\\"a}fer and Krystian Nowakowski and Carlos Arribalzaga Jov{\\'e} and Frank Koppens and Dirk Englund and Jacob M. Taylor},\nyear={2026},\nurl={https://openreview.net/forum?id=Bw9LCBz9KW}\n}"},"title":{"value":"ModelBench: A Benchmark for Extracting Executable, Physics-Based Models from Scientific Literature"},"pdf":{"value":"/pdf/76d73ba2e8b90fc035df3773654a2a2f446bc09b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"schoolkate|modelbench_a_benchmark_for_extracting_executable_physicsbased_models_from_scientific_literature"},"authorids":{"value":["~Pim_Schoolkate1","~Patrick_Huembeli1","~Frank_Schäfer1","~Krystian_Nowakowski1","~Carlos_Arribalzaga_Jové1","~Frank_Koppens1","~Dirk_Englund1","~Jacob_M._Taylor1"]},"authors":{"value":["Pim Schoolkate","Patrick Huembeli","Frank Schäfer","Krystian Nowakowski","Carlos Arribalzaga Jové","Frank Koppens","Dirk Englund","Jacob M. Taylor"]}},"version":2},{"content":{"summary":{"value":"### Summary\nThis paper proposes a physics-guided image enhancement framework integrating a differentiable camera response model (CRM) with deep generative learning. It decouples the camera response function (CRF) and brightness transformation function (BTF), introduces a dual-branch contrastive autoencoder (CAE) for domain-invariant feature extraction, and an adaptive feature distribution matching (AFDM) module for differentiable eCDF alignment. Experiments on underwater and low-light datasets show improved generalization and real-world robotic deployment.\n\n### Strengths\n- Novel integration of physical modeling (CRM) and deep networks.  \n- Dual-branch contrastive learning enhances robustness to domain shifts.  \n- AFDM offers a differentiable alternative to AdaIN/histogram matching.  \n- Comprehensive comparisons with strong baselines.\n\n### Weaknesses\n- Conceptual novelty is limited; CRF–BTF modeling and ADMM optimization follow prior work (e.g., LECARM).  \n- The “differentiable physics” part is loosely coupled with learning; no end-to-end integration.  \n- Writing and figures are dense and hard to follow.  \n- Evaluations rely mainly on no-reference metrics without perceptual or user studies.  \n- Focused on applications rather than learning theory—less aligned with ICLR scope."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How is the CRF calibration integrated during training or inference? Is it pre-computed or updated jointly with BTF?  \n2. What is the nature of the “guide image” used for AFDM—does it come from another domain, or is it sampled within the same dataset?  \n3. Can the authors report perceptual metrics (e.g., LPIPS, PSNR) or user study results to strengthen claims about visual quality?  \n4. How does the method generalize across *unseen* domains (e.g., training on low-light and testing on haze)?  \n5. What is the computational complexity compared to strong baselines such as Diff-Retinex++ or WFI2-Net?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Physics-inspired framework:**\nThe idea of combining radiometric camera modeling with deep generative learning is meaningful and provides some interpretability rarely seen in image enhancement research.\n\n2. **Dual-branch contrastive learning:**\nThe proposed CAE design with two symmetrical branches improves robustness to domain shifts and is theoretically analyzed (Eq. 8–9) for stability.\n\n3. **Differentiable eCDF alignment:**\nThe AFDM module introduces an elegant differentiable alternative to AdaIN or histogram matching, aligning distributions across domains.\n\n4. **Comprehensive experiments:**\nThe paper presents extensive visual and quantitative results on multiple domains (underwater, low-light, haze), and real-world robotic deployment adds practical value."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited novelty:**  \n   The contributions appear incremental. The CRF–BTF decoupling and ADMM-based CRF estimation follow previous works such as LECARM (Ren et al., TCSVT’18). Dual-branch contrastive learning and feature matching extend standard paradigms rather than introduce fundamentally new learning concepts.\n\n2. **Weak physics–learning integration:**  \n   Although termed “differentiable physics,” the physical CRF calibration is treated as an offline optimization rather than an end-to-end differentiable module. The connection between physical modeling and learned BTF remains loosely coupled.\n\n3. **Clarity and presentation issues:**  \n   The paper is dense, with unclear notation (e.g., reusing *f, g, E, P*), verbose equations, and missing intuition. The forward pipeline is not clearly explained—particularly how CRF calibration influences BTF training and inference. Figures are complex but not explanatory.\n\n4. **Empirical rigor:**  \n   Evaluation relies mainly on no-reference metrics (UIQM, UCIQE, CCF), which are noisy and limited in reflecting perceptual quality. No LPIPS/PSNR results, runtime comparison, or user study are provided. Statistical significance of the reported +1.226 UIQM improvement is unclear.\n\n5. **Language quality:**  \n   The writing contains numerous grammatical errors and awkward phrasing, making it difficult to follow in several sections (especially §2.2–2.3). The overall readability is below the ICLR standard.\n\n6. **Reproducibility concerns:**  No code or config reference is provided."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916675961,"tcdate":1761851326876,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3333/Reviewer_v49C"],"signatures":["ICLR.cc/2026/Conference/Submission3333/Reviewer_v49C"],"forum":"3NbMnPh7vM","number":2,"license":"CC BY 4.0","cdate":1761851326876,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3333/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916675961,"domain":"ICLR.cc/2026/Conference","replyto":"3NbMnPh7vM","id":"HsGXG5KXUO","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["camera response model","camera response function","brightness transformation function","dual branch contrastive encoder"]},"supplementary_material":{"value":"/attachment/660ccad137ee6911a53f5155f75bede5e2812362.pdf"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Visual perception in the wild have demonstrated transformative potential across a wide range of applications, spanning from planetary exploration to deep-sea monitoring missions. However, a fundamental challenge remains in enabling visual perception enhancement that can explicitly extract rules and support interactive, precise manipulation in unknown, dynamic environments—particularly under conditions of large scale data absence, heterogeneous data distribution, and without the supervision of annotated images. Our approach introduces a differentiable physics framework that unifies the camera response model (CRM) with deep learning to achieve visual perception enhancement under multiple degradation conditions. Specifically, grounded in fundamental principles of radiation physics, we formulate the camera response function (CRF) calibration as a constrained optimization problem. Then we reconstruct the brightness transformation function (BTF) in traditional CRM as a multi-scale generative network, completely decoupling it from the CRF. Meanwhile, we design a dual-branch contrastive encoder that enables the BTF to regulate the irradiance enhancement process through multi-scale exposure distributions learned from guide images. This offers a flexible BTF interface supporting stable and controllable domain generalization for image enhancement. Through comprehensive experiments, our method significantly advances domain generalization capabilities in adaptive image enhancement, outperforming specialized counterparts by margins of +1.226 (UIQM) averaged across challenging unseen underwater domains."},"_bibtex":{"value":"@misc{\nmao2026mastering,\ntitle={Mastering Domain Shift Image Enhancement Via Differentiable Physics},\nauthor={Ruiqi Mao and Rongxin Cui},\nyear={2026},\nurl={https://openreview.net/forum?id=3NbMnPh7vM}\n}"},"title":{"value":"Mastering Domain Shift Image Enhancement Via Differentiable Physics"},"pdf":{"value":"/pdf/00b30a2ef42d23d0c6ecb6959a6f12f9246198ad.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"mao|mastering_domain_shift_image_enhancement_via_differentiable_physics"},"authorids":{"value":["~Ruiqi_Mao1","~Rongxin_Cui1"]},"authors":{"value":["Ruiqi Mao","Rongxin Cui"]}},"version":2},{"content":{"venue":{"value":"Complex & Intelligent Systems"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s40747-025-02194-z.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"jung|physicsinformed_neural_network_and_momentum_contrastive_learning_for_battery_state_of_health_estimation"},"html":{"value":"https://doi.org/10.1007/s40747-025-02194-z"},"abstract":{"value":"Estimating the State of health (SoH) of lithium-ion batteries is essential for ensuring their safe and efficient operation across various applications. Traditional approaches often struggle to balance accuracy, physical consistency and data efficiency. This paper proposes a novel combination model of Physics-Informed Neural Network and Momentum Contrastive Learning for Battery State of Health Estimation that associates the interpretability of physics-based model with the representational power of contrastive learning. Our innovation lies in developing a unified optimization strategy that carefully balances an estimation physics-informed architecture and the power of contrastive learning. To specifically improve the physics-informed network, we leverage a shared feature encoder to improve representation learning for accurate SoH estimation. For contrastive learning, we design a physics-guided data augmentation strategy with a shared encoder, which generates realistic variations of battery degradation patterns and a momentum encoder architecture, which stabilizes the learning process. Extensive experiments on the NASA lithium-ion battery datasets demonstrate that our model achieves superior performance over state-of-the-art baselines such CNN, BPINN, Informer and XGBoost-ARIMA, achieving a mean absolute error (MAE) average of 0.095% and a root mean squared error (RMSE) average of 0.117% across all batteries. The associations of physics constraints with contrastive learning improve prediction accuracy and enhance model generalization across different battery types and operating conditions, addressing key limitations in existing battery health estimation approaches."},"title":{"value":"Physics-informed neural network and momentum contrastive learning for battery state of health estimation"},"authors":{"value":[{"fullname":"Jiwoo Jung"},{"fullname":"Yipene Cedric Francois Bassole","username":"~Yipene_Cedric_Francois_Bassole1"},{"fullname":"Yunsick Sung","username":"~Yunsick_Sung1"}]}},"tmdate":1789636142600,"pdate":1765929600000,"externalIds":["doi:10.1007/s40747-025-02194-z"],"tcdate":1767701888186,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Yunsick_Sung1"],"forum":"Ak8qMjPRUT","license":"CC BY-SA 4.0","number":22607,"cdate":1765969850709,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1789636142600,"domain":"OpenReview.net/Public_Article","id":"Ak8qMjPRUT","version":2},{"content":{"summary":{"value":"This paper proposes a post-training framework that fine-tunes flow-matching generators using randomized weak-form PDE residuals and a joint latent-parameter pathway, so the model produces physics-consistent fields while simultaneously inferring hidden coefficients; experiments on canonical PDE tasks indicate reduced residuals with limited impact on sample diversity.\n\nContributions.\n1) Post-training physics alignment: turns an already trained flow-matching model into a physics-respecting generator by minimizing weak-form residuals with compact test functions, avoiding unstable high-order derivatives and limiting drift from the base distribution.\n2) Joint state–parameter generation: augments the generator with a learned evolution for latent physical parameters and an inverse predictor, and fine-tunes both under an adjoint-matching objective (with a scaled memoryless schedule) to couple solutions and parameters.\n3) Practical control and coverage: demonstrates denoising, sparse-observation guidance, and boundary-condition adaptation, and exposes simple knobs to trade off constraint strength versus fidelity/diversity, with lightweight fine-tuning overhead."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1) Diversity vs. PDE correctness. For well-posed forward/inverse PDEs, please justify when output diversity is desirable; otherwise, replace or augment SSIM-based diversity with task metrics (solution error L2/H1, weak/strong residual distributions, boundary-violation rates) and, for partial-observation settings, include posterior calibration (coverage vs. nominal).\n\n2) Scope of test problems. Add at least one oscillatory elliptic case (Poisson/Helmholtz) and one basic incompressible flow (e.g., lid-driven cavity or cylinder shedding); if new runs are infeasible, provide higher-resolution or 3D variants or a brief scaling analysis (compute, stability bottlenecks).\n\n3) Baselines and recovery metrics. Include matched-compute head-to-head with (i) training-time physics-regularized flow matching, (ii) inference-time projection/ECI, and (iii) a classical PDE-constrained inversion baseline. Report solution error, boundary violations, weak/strong residuals, latent-parameter MAE/RMSE, and wall-clock.\n\n4) Ablations for method choices. Provide a small ablation comparing weak vs. strong residuals (stability, final residuals) and sensitivity to test-function sampling; show how the scaled memoryless parameter kappa affects stability, residual reduction, and drift from the base distribution."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1) Proposes a post-training route to impose physics on pretrained flow-matching models via weak-form PDE residuals, coupled with joint latent-parameter evolution for inverse problems without paired labels; also introduces a scaled memoryless noise schedule within adjoint matching.\n\n2) Grounds the method in adjoint matching and implements randomized local test functions for stable weak residuals; experiments span Darcy denoising, sparse-observation guidance, linear-elasticity boundary adaptation, and a small natural-image recoloring case, with ablations showing a residual–diversity trade-off.\n\n3) Clearly states goals and contributions, provides a pipeline diagram, derives the weak forms and test-function design, includes a full training algorithm and detailed dataset/backbone specs, and offers a reproducibility statement."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) Diversity objective may be misaligned for PDE solvers. For well-posed forward/inverse PDEs the target is a single solution; promoting output “diversity” is not desirable, and when partial observations make the task ill-posed, diversity stems from the problem, not the pipeline. The paper treats diversity as a knob/metric (SSIM-based) and studies its trade-off against residuals (Fig.\\ 3), which can conflict with PDE goals.\n\n2) Test problems are not comprehensive. Evaluations focus on Darcy denoising, sparse-obs guidance, and a linear-elasticity boundary change, plus a small image recoloring demo; there is no coverage of more challenging PDEs such as Poisson, Navier–Stokes, or Helmholtz, nor larger-scale or multi-physics settings.\n\n3) Limited baselines and quantitative recovery metrics. Beyond an ECI comparison in the elasticity case, there is no systematic head-to-head with alternative physics-constrained generative methods, and the paper provides little quantitative evaluation of latent-parameter recovery accuracy or real-data tests (most results are residual reductions and visuals)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358627137,"tcdate":1761982203504,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17809/Reviewer_Pud7"],"signatures":["ICLR.cc/2026/Conference/Submission17809/Reviewer_Pud7"],"forum":"khBHJz2wcV","number":2,"license":"CC BY 4.0","cdate":1761982203504,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17809/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358627137,"domain":"ICLR.cc/2026/Conference","replyto":"khBHJz2wcV","id":"zOIoPgVHjN","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We fine-tune pretrained flow-matching models using weak-form PDE residual rewards to generate physically consistent fields and infer latent parameters for inverse problems without paired solution–parameter training data."},"keywords":{"value":["Generative Modeling","Physics‑Informed Machine Learning","Inverse Problems","Parameter Identification"]},"supplementary_material":{"value":"/attachment/31af82a68d8c2d764654675d29c5798cbeb8796c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We present a framework for fine-tuning flow-matching generative models to enforce physical constraints and solve inverse problems in scientific systems. Starting from a model trained on low-fidelity or observational data, we apply a differentiable post-training procedure that minimizes weak-form residuals of governing partial differential equations (PDEs), promoting physical consistency and adherence to boundary conditions without distorting the underlying learned distribution. To infer unknown physical inputs, such as source terms, material parameters, or boundary data, we augment the generative process with a learnable latent parameter predictor and propose a joint optimization strategy. The resulting model produces physically valid field solutions alongside plausible estimates of hidden parameters, effectively addressing ill-posed inverse problems in a data-driven yet physics-aware manner. We validate our method on canonical PDE problems, demonstrating improved satisfaction of physical constraints and accurate recovery of latent coefficients. Further, we confirm cross-domain utility through fine-tuning of natural-image models. Our approach bridges generative modelling and scientific inference, opening new avenues for simulation-augmented discovery and data-efficient modelling of physical systems."},"_bibtex":{"value":"@inproceedings{\ntauberschmidt2026physicsconstrained,\ntitle={Physics-Constrained Fine-Tuning of Flow-Matching Models for Generation and Inverse Problems},\nauthor={Jan Tauberschmidt and Sophie Fellenz and Sebastian Josef Vollmer and Andrew B. Duncan},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=khBHJz2wcV}\n}"},"title":{"value":"Physics-Constrained Fine-Tuning of Flow-Matching Models for Generation and Inverse Problems"},"pdf":{"value":"/pdf/bd101e2dc591436d01faae477b1916554f68cba9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"tauberschmidt|physicsconstrained_finetuning_of_flowmatching_models_for_generation_and_inverse_problems"},"authorids":{"value":["~Jan_Tauberschmidt1","~Sophie_Fellenz1","~Sebastian_Josef_Vollmer1","~Andrew_B._Duncan1"]},"authors":{"value":["Jan Tauberschmidt","Sophie Fellenz","Sebastian Josef Vollmer","Andrew B. Duncan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a specific structure to encode physics prior to the training and use Euler/RK for time stepping to achieve good generalization capability under a data-scarce situation."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"The paper explicitly takes into account the physics of the system when designing the system, yielding better generalization capability compare to baselines like FNO"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I am a bit confused with the experimental setting. I really like the argument of baking more physics prior to the model. However, it seems that during the training, the model is still trained with a large-scale dataset - where one needs up to 10^6 times to generate this dataset."},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. I am curious any thoughts on why FNO performs so badly even with the full dataset for training? This is different from what I generally get from various literature. \n2.  I am curious why different padding strategy corresponds to boundary condition. How does it help enforce the boundary condition?\n3. how could it generalize to mesh base simulation with adaptive resolution?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636209854,"tcdate":1699241398240,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2688/Reviewer_La8N"],"signatures":["ICLR.cc/2024/Conference/Submission2688/Reviewer_La8N"],"forum":"UHIKtKzTj7","number":4,"license":"CC BY 4.0","cdate":1699241398240,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2688/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636209854,"domain":"ICLR.cc/2024/Conference","replyto":"UHIKtKzTj7","id":"Sma0CY2TZX","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Process systems modeling","Physics-informed machine learning","Temporal-spatial stepping method","Out-of-sample generalizability."]},"supplementary_material":{"value":"/attachment/bba5ee17f6691d592c3e248133053c5555cbae92.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Process systems, which play a fundamental role in various scientific and engineering fields, often rely on computational models to capture their complex temporal-spatial dynamics. However, due to limited insights into the intricate physical principles, these models can be imprecise or inapplicable, coupled with a significant computational demand exacerbating inefficiencies. To address these challenges, we propose a physics-aware proxy model (PAPM) to explicitly incorporate partial prior mechanistic knowledge, including conservation and constitutive relations. Additionally, to enhance the inductive biases about strict physical laws and broaden the applicability scope, we introduce a holistic temporal and spatial stepping method (TSSM) aligned with the distinct equation characteristics of different process systems, resulting in better out-of-sample generalization. We systematically compare state-of-the-art pure data-driven models and physics-aware models, spanning five two-dimensional non-trivial benchmarks in nine generalization tasks. Notably, PAPM achieves an average absolute performance improvement of 6.4%, while requiring fewer FLOPs, and only 1% of the parameters compared to the prior leading method, PPNN. Through such analysis, the structural design and specialized spatio-temporal modeling schemes (i.e., TSSM) of PAPM exhibit not only the most balanced trade-off between accuracy and computational efficiency among all methods evaluated, but also an impressive out-of-sample generalization."},"_bibtex":{"value":"@misc{\nliu2024papm,\ntitle={{PAPM}: A Physics-aware Proxy Model for Process Systems},\nauthor={Pengwei Liu and Zhongkai Hao and Xingyu Ren and Hangjie Yuan and Dong Ni},\nyear={2024},\nurl={https://openreview.net/forum?id=UHIKtKzTj7}\n}"},"title":{"value":"PAPM: A Physics-aware Proxy Model for Process Systems"},"pdf":{"value":"/pdf/b922d6af06adf7a8c43b4720ecc76728df0e52a9.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|papm_a_physicsaware_proxy_model_for_process_systems"},"authorids":{"value":["~Pengwei_Liu1","~Zhongkai_Hao1","~Xingyu_Ren2","~Hangjie_Yuan1","~Dong_Ni3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pengwei Liu","Zhongkai Hao","Xingyu Ren","Hangjie Yuan","Dong Ni"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a comprehensive benchmark designed to evaluate methods estimating heterogeneous treatment effects (HTEs) from (right-censored) survival data. Estimating HTEs in survival analysis is vital for applications such as precision medicine but remains challenging due to censoring, unobserved counterfactuals, and complex modeling assumptions. This paper addresses these challenges by providing a unified categorization of survival HTE methods into three families—outcome imputation, direct-survival CATE models, and survival meta-learners—along with modular implementations of their variants. It also includes synthetic datasets with known ground-truth HTEs that systematically vary causal assumptions (e.g., randomization, confounding, positivity, informative censoring) and survival dynamics (e.g., Cox, AFT, Poisson models with varying censoring rates). Additionally, it offers multiple semi-synthetic datasets and two real-world datasets."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"-\tIf ground-truth hazard functions are available, it would be more meaningful to evaluate HTE performance with respect to the hazard functions themselves. Since many of the evaluated methods are capable of directly estimating conditional hazards from data, it is unclear why these performance results are not reported in the paper.\n-\tRegarding Weakness 2, it remains unclear how the benchmark results can be practically leveraged when applying the methods to new datasets or models."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"-\tAddresses a Critical Need: The paper tackles a significant gap in the causal inference and survival analysis literature -- the lack of standardized evaluation practices for HTE estimation with survival data. \n-\tThe benchmark incorporates multiple data types (synthetic, semi-synthetic, real). The synthetic data explores a wide range of crucial factors: different causal assumption violations (i.e., confounding, positivity, informative censoring) and commonly used survival/censoring distributions (Cox, AFT, Poisson)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Major comments:\n\n-\tLimited synthetic data generation processes: The synthetic datasets are restricted to three survival time generation mechanisms—Cox, AFT, and Poisson—which may not capture more complex, non-parametric scenarios. For instance, SurvITE defines discrete hazards and samples time-to-event or censoring outcomes based on those hazards. Also, the validity and representativeness of such synthetic data remain unclear. In particular, how treatment effects, selection biases, and covariate shifts influence the generated time-to-event outcomes is not well articulated, as the datasets are primarily described in terms of their underlying hazard functions.\n- Lack of actionable insights from the benchmark: The benchmark results do not provide sufficient insight into how model design decisions should be made based on the findings. As a result, the analysis feels largely descriptive rather than being guided by principled intuitions or hypotheses about model behavior. The paper could better highlight how specific model characteristics or data properties influence performance, offering clearer guidance for model selection and application.\n-\tThe paper introduces a category of \"Direct-Survival CATE Models\", but, deep survival models that directly estimate the conditional average treatment effect (CATE) (such as SurvITE) on survival outcomes are not included. \n-\tAlthough the paper claims to provide a unified framework for handling various survival models, it does not address models that directly predict time-to-event outcomes (e.g., Chafuwa et al., 2021). It is unclear how such models would be integrated into the proposed benchmark.\n-\tEstimates of hazard or survival functions can vary substantially depending on the chosen time horizon, yet the benchmark does not clearly specify or analyze how this dependency affects the evaluation results.\n\nMinor comments:\n\n-\tVisualizations (e.g., conditional hazard and survival function plots) for synthetic datasets under different causal configurations would help clarify and validate the underlying data generation processes."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924258843,"tcdate":1761576017892,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13710/Reviewer_FuYz"],"signatures":["ICLR.cc/2026/Conference/Submission13710/Reviewer_FuYz"],"forum":"qG6O3jMkCj","number":2,"license":"CC BY 4.0","cdate":1761576017892,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13710/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924258843,"domain":"ICLR.cc/2026/Conference","replyto":"qG6O3jMkCj","id":"mrpcJUcMaP","forumContent":{"TLDR":{"value":"We present SurvHTE-Bench, a comprehensive causal inference benchmark to evaluate methods that estimate heterogeneous treatment effects from censored survival data, enabling rigorous, fair, and reproducible comparison across diverse causal scenarios."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Causal Inference","Survival Analysis","Treatment Effect","Datasets and Benchmarks","CATE"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as precision medicine and individualized policy-making. Yet, the survival analysis setting poses unique challenges for HTE estimation due to censoring, unobserved counterfactuals, and complex identification assumptions. Despite recent advances, from causal survival forests to survival meta-learners and outcome imputation approaches, evaluation practices remain fragmented and inconsistent. We introduce SurvHTE‐Bench, the first comprehensive benchmark for HTE estimation with censored outcomes. The benchmark spans (i) a modular suite of synthetic datasets with known ground truth, systematically varying causal assumptions and survival dynamics, (ii) semi-synthetic datasets that pair real-world covariates with simulated treatments and outcomes, and (iii) real-world datasets from a twin study (with known ground truth) and from an HIV clinical trial. Across synthetic, semi-synthetic, and real-world settings, we provide the first rigorous comparison of survival HTE methods under diverse conditions and realistic assumption violations. SurvHTE‐Bench establishes a foundation for fair, reproducible, and extensible evaluation of causal survival methods."},"_bibtex":{"value":"@inproceedings{\nnoroozizadeh2026survhtebench,\ntitle={Surv{HTE}-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis},\nauthor={Shahriar Noroozizadeh and Xiaobin Shen and Jeremy Weiss and George H. Chen},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=qG6O3jMkCj}\n}"},"title":{"value":"SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis"},"pdf":{"value":"/pdf/50449a9c07ecf9a08caac37e6b186bd641bc7a48.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"noroozizadeh|survhtebench_a_benchmark_for_heterogeneous_treatment_effect_estimation_in_survival_analysis"},"authorids":{"value":["~Shahriar_Noroozizadeh1","~Xiaobin_Shen1","~Jeremy_Weiss1","~George_H._Chen1"]},"authors":{"value":["Shahriar Noroozizadeh","Xiaobin Shen","Jeremy Weiss","George H. Chen"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewers for their comments. Below, we list the strengths of our work as identified by the reviewers:\n\n1. Broad, comprehensive, and extensive evaluation [unjk, q4tW, WEJ2]\n2. Interesting findings [nZYb]\n3. “Well-organized”, “clearly written” presentation [nZYb, q4tW, WEJ2]\n\nTo address concerns regarding task validity, novelty, and rigor, we have added **four major experiments** in the revised PDF:\n\n1. **Human Baseline (App. E.9):** We evaluated 10 human subjects. Even with only one* demonstration and no QA pairs, humans achieved **~90% accuracy** on most scenarios. This definitely proves the task is well-defined and visually solvable.\n2. **All-Frames Evaluation (App. E.10):** We tested providing all video frames during evaluation. Performance did not improve (and often degraded), confirming that the bottleneck is **reasoning**, not perception/motion estimation.\n3. **Attention Analysis (App. F.1):** We visualized attention maps (e.g., Gemma3-12B), revealing that LMMs attend to image tokens **an order of magnitude less** than text tokens. This provides a structural explanation for the language bias we observed.\n4. **Prompt Robustness (App. E.11):** We analyzed sensitivity to prompt perturbations. Results show a low coefficient of variation for larger models, confirming experimental stability.\n\nThe concerns raised by reviewers stem from a **misunderstanding of our evaluation goal**. We briefly state our goal and its importance, along with an example below.\n\n**Physical reasoning** is the task of predicting the outcome of a physical event based on initial observations. For example, given the dashcam footage from a car, we wish to predict whether a collision will occur or not, whether the car will enter the opposite lane, and so on. Note that physical reasoning is a **normative task**, i.e., predictions are about events yet to happen from the initial conditions based on some physical rules.\n\n**InPhyRe’s Goal**: Evaluate whether LMMs can infer and apply physical laws from demonstration samples when the underlying rules were not observed during training.\n\n**Comparison of InPhyRe with other datasets**\n\nPCB is Physics Context Builder (V. Balazadeh et al., CVPR 2025).\n\n| Benchmark | Task | Example query | Do physical conditions/laws change between training and testing? | Require on-the-fly physical condition/law inference from demos/test sample? |\n| --- | --- | --- | --- | --- |\n| CLEVRER | Factual and counterfactual physical reasoning | “What shape is the object that collides with the cyan cylinder?” | No | No |\n| ComPhy | General physical reasoning requiring latent property reasoning | “Which event would happen if the grey cylinder were not magnetic?” | No | Yes |\n| CoPhy | Counterfactual physical reasoning | “what happens if we push this ball to the right?” | No | No (models are trained to do counterfactual reasoning) |\n| PhysBench | General physical reasoning about property, dynamics, relations, etc. | “Which object will the cart hit first?” “What is the object closest to the teacup in the Figure?” | No | No |\n| IntPhys (v1 and v2) | Physical plausibility prediction | Predict the next frame based on the initial frames of some event | Yes | No |\n| Physion | Object contact prediction | The initial frames of a video showing falling dominoes | No | No |\n| Physion++ | Object contact prediction through physical property understanding | The initial frames of a video showing falling dominoes | No | No |\n| ContPhy | Physical property and dynamics prediction | “Is the density of the orange fluid greater than that of the green fluid?” ”What can we do to guide most of the orange fluid into cyan cylinder?” | No | No |\n| PCB | Train smaller VLMs to support models during inference | Not an evaluation benchmark. PCBs provide scene narration to support the inference model. | No | No |\n| **InPhyRe** | Infer physical laws from demo samples and apply them to the evaluation sample. | “Will the velocity of the green cube increase after collision with the red sphere?” | **Yes** | **Yes** |\n\n**Inductive physical reasoning (IPR) vs counterfactual physical reasoning (CPR)**: The physical laws do not change between training and testing in CPR. Compared to CPR, IPR measures a stronger form of generalization. E.g., ComPhy does not check whether the LMM actually understood the underlying physics or made a guess based on a memorized training set.\n\n**IPR vs intuitive physical reasoning**: The goal in intuitive physical reasoning is to evaluate physical reasoning in models through its ability to detect violated physics. **Violation detection does not require understanding the differing rules**. In contrast, IPR requires inferring the underlying rules from demos."},"title":{"value":"Global Response - Clarification of our objective (part 1/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763728508375,"tcdate":1763728508375,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14468/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission14468/Authors"],"forum":"IIrPoZ28dN","number":1,"license":"CC BY 4.0","cdate":1763728508375,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14468/-/Official_Comment"],"mdate":1763728508375,"domain":"ICLR.cc/2026/Conference","replyto":"IIrPoZ28dN","id":"eRR0uXWLuT","forumContent":{"TLDR":{"value":"We evaluate whether large multimodal models can infer and apply unseen physical laws from demonstration samples."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["large multimodal models","physical reasoning","inductive reasoning"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large multimodal models (LMMs) encode universal physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potential collision event from visual input. However, since parametric knowledge includes only the physical laws seen during training, it is insufficient for reasoning when the inference scenario follows physical laws unseen during training. In contrast, humans can adapt their physical reasoning to unseen physical environments with only a few visual examples. This inductive physical reasoning ability is indispensable for LMMs if they are to replace human agents in safety-critical applications. Despite its importance, existing visual benchmarks evaluate only the parametric knowledge in LMMs, and not inductive physical reasoning. To this end, we propose InPhyRe, the first visual question answering benchmark to measure inductive physical reasoning in LMMs. InPhyRe evaluates LMMs on their ability to predict the outcome of collision events in algorithmically generated synthetic videos. By inspecting over 13 open-source and proprietary LMMs, InPhyRe informs us that (1) LMMs struggle to apply their limited parametric knowledge about universal physical laws to reasoning, (2) inductive physical reasoning in LMMs is weak when inference scenarios obey physical laws unseen during training, and (3) inductive physical reasoning in LMMs suffers from language bias and largely ignores the visual inputs, questioning the trustworthiness of LMMs regarding visual inputs."},"_bibtex":{"value":"@misc{\nsreekumar2026inphyre,\ntitle={InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning},\nauthor={Gautam Sreekumar and Vishnu Boddeti},\nyear={2026},\nurl={https://openreview.net/forum?id=IIrPoZ28dN}\n}"},"title":{"value":"InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning"},"pdf":{"value":"/pdf/b87a0e49daf5f5f05caff235d97e823b17b5dbfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sreekumar|inphyre_discovers_large_multimodal_models_struggle_in_inductive_physical_reasoning"},"authorids":{"value":["~Gautam_Sreekumar1","~Vishnu_Boddeti1"]},"authors":{"value":["Gautam Sreekumar","Vishnu Boddeti"]}},"version":2},{"content":{"summary":{"value":"To address the challenge of interpretability, generalizability and long-horizon predictive performance, the author proposed the PDE-Diffusion model. The model incorporate physics-based priors, two stage diffusion model to tackle the mentioned challenges. It was also tested on extensive datasets."},"presentation":{"value":"1 poor"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"Originality: The author propose to embed a physics-based prior into the sampling process for diffusion model to enhance the learning result. \n\nQuality and clarify: The methodology part misses details about the physics-based priors. The result part is confusing, indicating the proposed model cannot beat the SOTA models.\n\nSignificance: The author claims that the physics-embedded framework is more interoperable and generalizable, meanwhile mitigate the problem of temporal coherence."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The methodology part is not clear in cases of how to embed the physics prior $\\epsilon_I$ and $\\epsilon_B$ during the reverse sampling process. The result doens't support the claim where in many cases, the PDE-Diffusion result is much worse than FNO. The reviewer also has major concern about the writing style. It seems not rigorous and scientific on delivering the message. Details can be found in the questions part."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"1. Why not put the algorithm 3 in the main manuscript? The innovative part physics-conditioning is in algorithm 3 while algorithm 1 and 2 are well-known algorithm. Moreover, I didn't see the equation on how to deal with the $\\epsilon_I$ and $\\epsilon_B$ in either main manuscript or the appendix. How do you embed the physics prior into the model?\n\n2. The result is confusing. In Table 2 and 3, the result for the diffusion models are exactly the same. However, the metrics are evaluated for different datasets. How is it possible? Moreover, the bold font seems to be used pretty arbitrary, not always the best result. In fact, the proposed model's performance is worse than FNO in may cases. In that case, the result doesn't support the authors' claim that the proposed structure is better than SOTA.\n\n3. The writing is poor, seems not to be proofread. To name a few. In Table 2, for FNO, there is a blank without numbers. In page 7 section 5.3, \"has a decrease of x%\" is obviously not a completed part. The authors should fill the exact number instead of a placeholder. In page 13, \"which is reasonable when reasonable in the ...\" is obviously a wrong sentence."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637165356,"tcdate":1698797822638,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission9255/Reviewer_p93J"],"signatures":["ICLR.cc/2024/Conference/Submission9255/Reviewer_p93J"],"forum":"3sOE3MFepx","number":3,"license":"CC BY 4.0","cdate":1698797822638,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission9255/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637165356,"domain":"ICLR.cc/2024/Conference","replyto":"3sOE3MFepx","id":"jCtjBiec8B","forumContent":{"TLDR":{"value":"In this paper we propose to solve PDE with a physical guided diffusion model whose reverse process is conditioned by initial/boundary and PDE guidance."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["AI for science","PDE","diffusion model","generative model"]},"supplementary_material":{"value":"/attachment/7fecff3c4ea127c6363e0c6652a5afb4cd969e6c.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Solving partial differential equations (PDEs) is crucial in various disciplines, and their resolution often necessitates the use of computationally intensive numerical methods as well as specialized domain expertise. While data-driven approaches have emerged as promising alternatives, they encounter limitations in terms of generalizability, interpretability, and long-horizon predictive performance, as well as issues related to temporal incoherence. To address these challenges, we introduce the PDE-Diffusion, a two-stage model with three distinctive features: (i) the incorporation of physics-based priors to enhance model interpretability and generalization, (ii) a two-stage diffusion model that efficiently handles physical field forecasting without requiring multi-frame inputs, and (iii) the assimilation of PDE-informed constraints to ensure temporal coherence while producing high-quality predictive results. We conduct extensive experiments to evaluate PDE-Diffusion's capabilities using the PDEBench dataset and two of our newly proposed datasets. The results indicate that PDE-Diffusion delivers state-of-the-art performance in all cases."},"_bibtex":{"value":"@misc{\ngao2024pdediffusion,\ntitle={{PDE}-Diffusion: Physic guided diffusion model for solving partial derivative equations},\nauthor={Chonghan Gao and Haoyi Zhou and wen xin gong and QING PO WU WU and Tianyu Chen and Qian Yu and Shanghang Zhang and Jianxin Li},\nyear={2024},\nurl={https://openreview.net/forum?id=3sOE3MFepx}\n}"},"title":{"value":"PDE-Diffusion: Physic guided diffusion model for solving partial derivative equations"},"pdf":{"value":"/pdf/20350cdee6651c3f540fbc1cb3c8e77f26cccdcc.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"gao|pdediffusion_physic_guided_diffusion_model_for_solving_partial_derivative_equations"},"authorids":{"value":["~Chonghan_Gao1","~Haoyi_Zhou1","~wen_xin_gong1","~QING_PO_WU_WU1","~Tianyu_Chen1","~Qian_Yu4","~Shanghang_Zhang4","~Jianxin_Li3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chonghan Gao","Haoyi Zhou","wen xin gong","QING PO WU WU","Tianyu Chen","Qian Yu","Shanghang Zhang","Jianxin Li"]}},"version":2},{"content":{"summary":{"value":"This paper proposes EnvSocial-Diff, a diffusion-based crowd simulator that augments social-physics (SFM) with (1) structured environmental conditioning (obstacles, objects of interest, lighting, etc.) and (2) an Individual–Group Interaction (IGI) module that mixes pairwise and group-conformity signals via a GNN. On GC and UCY, the method reports consistent gains over SFM/PCS/SPDiff and other baselines across various metrics, with ablations for each factor."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- How sensitive is the method to the accuracy and granularity of scene annotations (objects of interest, lighting, obstacle maps)? For example, if these are noisy, incomplete, or generated automatically from imperfect scene-understanding systems, how does performance degrade?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Many prior pedestrian prediction models are purely data-driven and overlook structured environmental and social factors. There are important factors in pedestrian simulation. This work tries to bridge the social-force ideas with modern generative modeling and scene modeling, which is refreshing.\n    - Accurate modelling the prior scene information will be essential for future crowd simulation works. \n\n- The decomposition into destination force + diffusion refinement is clean and easy to follow. It preserves interpretability while still benefiting from generative diversity.\n- The group similarity and alignment logic, aggregated via GNN, is conceptually simple but grounded in real social-navigation dynamics. The ablations suggest it brings in meaningful improvement."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The **demo video** is very difficult to interpret. This is currently the biggest presentation gap. As it stands, it is hard to tell what is happening, which agents belong to which group, or how environment cues influence behavior. Since one of the main claims is improved realism and responsiveness to context, the qualitative visualization should make these effects obvious. Overlays, legends, visual callouts, and side-by-side comparisons would help a lot.\n- The framework layers several components (diffusion, semantic encoders, GNN, physics prior). While individually reasonable, it would be good to comment on compute cost and whether a simpler scene-encoder and diffusion baseline might achieve similar gains.\n- GC and UCY are standard dataset in crowd simulation, but they are relatively small scenes. Since the paper emphasizes complex social and environmental structure, a brief discussion or qualitative example in a denser or more varied environment would help support the generality claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921948007,"tcdate":1761971715853,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10714/Reviewer_ZL51"],"signatures":["ICLR.cc/2026/Conference/Submission10714/Reviewer_ZL51"],"forum":"2XBAm3Dbnt","number":3,"license":"CC BY 4.0","cdate":1761971715853,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10714/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921948007,"domain":"ICLR.cc/2026/Conference","replyto":"2XBAm3Dbnt","id":"Zliq8UNuwD","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"A diffusion-based crowd simulation model with environmental conditioning and individual-group interaction."},"keywords":{"value":["Crowd simulation","Social physics force","Diffusion model"]},"supplementary_material":{"value":"/attachment/b5e0a6a16bde36ba1f6b943f605e484f6315ea8e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Modeling realistic pedestrian trajectories requires accounting for both social interactions and environmental context, yet most existing approaches largely emphasize social dynamics. We propose EnvSocial-Diff: a diffusion-based crowd simulation model informed by social physics and augmented with environmental conditioning and individual-group interaction. Our structured environmental conditioning module explicitly encodes obstacles, objects of interest, and lighting levels, providing interpretable signals that capture scene constraints and attractors. In parallel, the individual-group interaction module goes beyond individual-level modeling by capturing both fine-grained interpersonal relations and group-level conformity through a graph-based design. Experiments on multiple benchmark datasets demonstrate that EnvSocial-Diff outperforms the latest state-of-the-art methods, underscoring the importance of explicit environmental conditioning and multi-level social interaction for realistic crowd simulation."},"_bibtex":{"value":"@inproceedings{\nzhao2026envsocialdiff,\ntitle={EnvSocial-Diff: A Diffusion-Based Crowd Simulation Model with Environmental Conditioning and Individual-Group Interaction},\nauthor={Bingxue Zhao and Qi Zhang and Hui Huang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=2XBAm3Dbnt}\n}"},"title":{"value":"EnvSocial-Diff: A Diffusion-Based Crowd Simulation Model with Environmental Conditioning and Individual-Group Interaction"},"pdf":{"value":"/pdf/6be61c23c2e7def47a9666339c640097edb54588.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhao|envsocialdiff_a_diffusionbased_crowd_simulation_model_with_environmental_conditioning_and_individualgroup_interaction"},"authorids":{"value":["~Bingxue_Zhao1","~Qi_Zhang11","~Hui_Huang3"]},"authors":{"value":["Bingxue Zhao","Qi Zhang","Hui Huang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes an architecture called Physics-Informed Deep Inverse Operator Networks (PI-DIONs), which can learn the solution operator of PDE-based inverse problems without labeled training data. The architecture  of PI-DIONs is based on DeepONet, and trained with both the physics-infomred loss and data reconstruction loss. The stability estimates established in the inverse problem literature are extended to the operator learning framework. Experiments are conducted to demonstrate the effectiveness of PI-DIONs in learning the solution operators of the inverse problems without the need for labeled data."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. DeepONet and FNO are used for forward problems traditionally, how did they deal with inverse problems in your experiments?\n2. How is the labeled training target f mentioned in line 399 used?  The loss for target f is absent in line 152."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The integration of physics-informed losses into an inverse problem framework based on operator learning is novel, and in principle PI-DIONs can solve the inverse problems (at least in scenarios mentioned in experiments) fast and without the need for labeled data.\n2. Theoretical analysis of the stability estimates is provided."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Line 243, \"where the term ∥f − f^\\star∥L2(Ωm) in the righthand side\", there is no such term there. Please clarify the equation in line 242 and include all terms on the right-hand side of the equation. \n2. It seems that the input to the reconstruction and inverse branch networks is fixed in shape, corresponding to the partial measurement with given geometry. The observed data in PINNs can have variable count and locations. Please discuss how PI-DIONs might be adapted to handle variable measurement geometries and if there are any limitations on the types of measurement setups it can handle. \n3. In the experiments, PI-DIONs are compared with purely data-driven DeepONet and FNO, which both did not take physics information into account. If possible, please include comparisons with PINNs in the experiments, since both your PI-DIONs and PINNs are physics-informed methods for inverse problems.\n4. The simultaneous training of physics-informed losses for 1000 samples is a difficult task (similar to train 1000 PINNs simultaneously). I am curious about the training difficulties encountered. Please provide specific details on training time, hardware used, and any convergence challenges encountered. If possible, please also include an ablation study on the effect of sample size on PI-DIONs' performance since smaller sample size may lead to easier optimization.\n5. The theoretical analysis on stability estimate is extended from existing key results that considered the single element case. \n6. Please provide a clear definition of u in line 152 and describe its relationship with partial measurement. In line 456, it is better to write \"f(x,y) = 100x(1 − x)y(1 − y) \", so does line 450. \n\nConsidering the above weaknesses, I give a score of 3 to the current version of this paper."}},"nonreaders":[],"tmdate":1732961374695,"tcdate":1730345855472,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6116/Reviewer_3Fuq"],"signatures":["ICLR.cc/2025/Conference/Submission6116/Reviewer_3Fuq"],"forum":"0FxnSZJPmh","number":2,"license":"CC BY 4.0","cdate":1730345855472,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6116/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732961374695,"domain":"ICLR.cc/2025/Conference","replyto":"0FxnSZJPmh","id":"CEg1cvslMi","forumContent":{"TLDR":{"value":"We propose a novel architecture called Physics-Informed Deep Inverse Operator Networks (PI-DIONs), which can learn the solution operator of PDE-based inverse problems without any labeled training data."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Inverse Problems","Stability","Operator Learning","Physics-Informed Machine Learning"]},"supplementary_material":{"value":"/attachment/fbb8719feb5bf7c131012ed11421beed954dc966.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Inverse problems involving partial differential equations (PDEs) can be seen as discovering a mapping from measurement data to unknown quantities, often framed within an operator learning approach. However, existing methods typically rely on large amounts of labeled training data, which is impractical for most real-world applications. Moreover, these supervised models may fail to capture the underlying physical principles accurately. To address these limitations, we propose a novel architecture called Physics-Informed Deep Inverse Operator Networks (PI-DIONs), which can learn the solution operator of PDE-based inverse problems without any labeled training data. We extend the stability estimates established in the inverse problem literature to the operator learning framework, thereby providing a robust theoretical foundation for our method. These estimates guarantee that the proposed model, trained on a finite sample and grid, generalizes effectively across the entire domain and function space. Extensive experiments are conducted to demonstrate that PI-DIONs can effectively and accurately learn the solution operators of the inverse problems without the need for labeled data."},"_bibtex":{"value":"@inproceedings{\ncho2025physicsinformed,\ntitle={Physics-Informed Deep Inverse Operator Networks for Solving {PDE} Inverse Problems},\nauthor={Sung Woong Cho and Hwijae Son},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=0FxnSZJPmh}\n}"},"title":{"value":"Physics-Informed Deep Inverse Operator Networks for Solving PDE Inverse Problems"},"pdf":{"value":"/pdf/b56f166dad491faab7b756b572988dae7aa557cc.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"cho|physicsinformed_deep_inverse_operator_networks_for_solving_pde_inverse_problems"},"authorids":{"value":["~Sung_Woong_Cho1","~Hwijae_Son1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sung Woong Cho","Hwijae Son"]}},"version":2},{"content":{"summary":{"value":"The existing literature offers limited theoretical exploration of conditional class probabilities-based algorithms for classification in complex scenarios, particularly concerning the cross-entropy loss. In this paper, authors investigated the convergence rates of algorithms based on conditional class probabilities estimation using kernel logistic regression in complex classification scenarios, such as long-tailed learning, domain adaptation, and transfer learning."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please explain the 3rd weakness mentioned above."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"No."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. A new oracle inequality for kernel logistic regression is established that holds with high probability.\n\n2. The approximation error of kernel logistic regression w.r.t. the cross-entropy loss is derived.\n\n3. The optimal convergence rates w.r.t. the cross-entropy loss for complex variants of classification problems is established."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. As stated in the 3rd paragraph, the main motivation is \"The existing literature offers limited theoretical exploration of conditional class probabilities-based algorithms for classification in complex scenarios, particularly concerning the cross-entropy loss.” It can be better to show the connection of the findings in this paper (i.e., kernel logistic regression for complex classification scenarios) and existing findings for standard classification scenario.\n\n2. It can be better to show the connection of the three contributions in introduction. In the current version, it is even hard to find if there is some connection them with the title of this paper.\n\n3. The main novelty is to restrict applications into three kinds of scenarios, i.e., long-tailed learning, domain adaptation, and transfer learning, which are called complex classification scenarios. I know all of them, but I don't know why they are put together in this paper (especially for long-tailed learning). Is there any logic to that? Why does the author focus on these three scenarios? The current version feels pieced together."}},"nonreaders":[],"tmdate":1731428329484,"tcdate":1730020406656,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5162/Reviewer_ygYV"],"signatures":["ICLR.cc/2025/Conference/Submission5162/Reviewer_ygYV"],"forum":"WlhVRh2rQ0","number":2,"license":"CC BY 4.0","cdate":1730020406656,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5162/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428329484,"domain":"ICLR.cc/2025/Conference","replyto":"WlhVRh2rQ0","id":"Dq7gl4eRM1","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["complex classification scenarios","long-tailed learning","domain adaptation","transfer learning","kernel methods","logistic regression","learning theory"]},"primary_area":{"value":"learning theory"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex classification scenarios, including long-tailed learning, domain adaptation, and transfer learning, present substantial challenges for traditional algorithms. Conditional class probability (CCP) predictions have recently become critical components of many state-of-the-art algorithms designed to address these challenging scenarios. Among kernel methods, kernel logistic regression (KLR) is distinguished by its effectiveness in predicting CCPs through the minimization of the cross-entropy (CE) loss. Despite the empirical success of CCP-based approaches, the theoretical understanding of their performance, particularly regarding the CE loss, remains limited. In this paper, we bridge this gap by demonstrating that KLR-based algorithms achieve minimax optimal convergence rates for the CE loss under mild assumptions in these complex tasks, thereby establishing their theoretical efficiency in such demanding contexts."},"_bibtex":{"value":"@inproceedings{\nwen2025optimal,\ntitle={Optimal Learning of Kernel Logistic Regression for Complex Classification Scenarios},\nauthor={Hongwei Wen and Annika Betken and Hanyuan Hang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=WlhVRh2rQ0}\n}"},"title":{"value":"Optimal Learning of Kernel Logistic Regression for Complex Classification Scenarios"},"pdf":{"value":"/pdf/8297274cc7adba6bc8e2ae70ffdeb5e760ecd863.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"wen|optimal_learning_of_kernel_logistic_regression_for_complex_classification_scenarios"},"authorids":{"value":["~Hongwei_Wen1","~Annika_Betken1","~Hanyuan_Hang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hongwei Wen","Annika Betken","Hanyuan Hang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces AGG-RL, a novel framework for sound source localization (SSL). It aims to achieve SSL with geometric invariance and grid flexibility by jointly learning audio-geometric representations and grid representations in a shared latent space. To achieve this, the approach proposes a learnable non-uniform discrete Fourier transform (LNuDFT) that assigns frequency bins based on physical informativeness, and a relative microphone position encoding (rMPE) aligned with TDOA. Experiments were conducted on synthetic and real datasets, demonstrating improved performance over baselines."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Experiment (v) still outperforms the proposed one on the Dynamic-U data and performs on par with the proposed one for NAO robot, but significantly degrades on the Eigenmike. Please expand on this discrepancy?\n\n2. In L70, “Both components facilitate the extraction of spatial representations with physics-based inductive biases”: Why are LNuDFT and rMPE considered as imposing (physics-based) inductive biases? To me, these are perceived as a process that introduces and makes learnable parameters for feature extraction more flexible, and I see it as a process that relieves the inductive bias.\n\n3. Can the LNuDFT initialization or update scheme get stuck in poor local minima, e.g., if the initial frequency allocation is far from optimal? Have the authors tried more physically motivated or data-driven initializations beyond logit-based mapping?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- **Originality**: The integration of physics-informed design (LNuDFT, rMPE) with a unified latent space for audio, geometry, and grid is original and is well-motivated by known limitations of geometry-specific and fixed-grid methods in SSL.\n- **Quality**: The experimental design is thorough: ablation studies, varied datasets (including real recordings), and consistent baselines provide credible evidence for claims of generalization and robustness. The efficiency analysis (parameter count, FLOPs) is also useful.\n- **Clarity**: The paper is mostly well written, with helpful architectural diagrams and clear mathematical formulations. The method and its motivation are explained coherently, and appendices include implementation details and sample code links.\n- **Significance**: The experimental results demonstrate that this framework provides substantial improvements to SSL in real-world environments where array and grid conditions may vary."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **(Lack of justification for the claim on being physics-informed)** The paper mentions LNuDFT and rMPE as two main components proposed as ‘physics-informed components’ in the abstract, but the reason why these are ‘physics-informed’ was not explained.\n    - Although Appendix A.1 supports that LNuDFT's ‘trainable manner’ emphasizes informative phase regions/causes to some extent, but the way of drawing the claim that 'it emphasizes informative regions' was limited to the common sense about speech signals. (For example, would the frequency response of LNuDFT parameters become different for the other signals, e.g., bird chirps or ambient wind noise etc?)\n    - Similarly, for rMPE, it was not demonstrated that using this method actually captures “inter-channel time differences” more effectively.\n\n    Performing DFT non-uniformly and encoding information relatively can be done independently of physics; connecting it to physical phenomena seems to require more detailed justification.\n\n2. **(Difficulty in result analysis)** While it is commendable that comparisons were made on both real and synthetic datasets, the fact that all microphone arrays in the synthetic data were selected as dynamic microphones complicates result analysis.\n    - Particularly, examining the results in Table 3 reveals a noticeable trend of performance differences, but it is unclear whether this stems from differences between real and synthetic data or from differences in dynamic microphones. The authors also did not draw a clear conclusion on this point."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921791060,"tcdate":1762533115287,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10505/Reviewer_aRq1"],"signatures":["ICLR.cc/2026/Conference/Submission10505/Reviewer_aRq1"],"forum":"bWXpJFesLS","number":3,"license":"CC BY 4.0","cdate":1762533115287,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10505/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921791060,"domain":"ICLR.cc/2026/Conference","replyto":"bWXpJFesLS","id":"inIORM9Ln3","forumContent":{"TLDR":{"value":"This paper proposes audio-geometry-grid representation learning for grid-flexible and geometry-invariant sound source localization, leveraging learnable non-uniform discrete Fourier transform and relative microphone positional encoding."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Sound Source Localization","Geometry-Invariant","Grid-Flexible","Representation Learning","Physics-Informed Design","Learnable Non-uniform DFT","Relative Microphone Positional Encoding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Sound source localization (SSL) is a fundamental task in spatial audio understanding, yet most deep neural network-based methods are constrained by fixed array geometries and predefined directional grids, limiting generalizability and scalability. To address these issues, we propose _audio-geometry-grid representation learning_ (AGG-RL), a novel framework that jointly learns audio-geometry and grid representations in a shared latent space, enabling both geometry-invariant and grid-flexible SSL. Moreover, to enhance generalizability and interpretability, we introduce two physics-informed components: a _learnable non-uniform discrete Fourier transform_ (LNuDFT), which optimizes the dense allocation of frequency bins in a non-uniform manner to emphasize informative phase regions, and a _relative microphone positional encoding_ (rMPE), which encodes relative microphone coordinates in accordance with the nature of inter-channel time differences. Experiments on synthetic and real datasets demonstrate that AGG-RL achieves superior performance, particularly under unseen conditions. The results highlight the potential of representation learning with physics-informed design towards a universal solution for spatial acoustic scene understanding across diverse scenarios."},"_bibtex":{"value":"@inproceedings{\nbaek2026physicsinformed,\ntitle={Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization},\nauthor={Min-Sang Baek and Gyeong-Su Kim and Donghyun Kim and Joon-Hyuk Chang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=bWXpJFesLS}\n}"},"title":{"value":"Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization"},"pdf":{"value":"/pdf/dcd11d5d38c6a0bf65ded06ef46bf42c740dab7c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"baek|physicsinformed_audiogeometrygrid_representation_learning_for_universal_sound_source_localization"},"authorids":{"value":["~Min-Sang_Baek1","~Gyeong-Su_Kim1","~Donghyun_Kim20","~Joon-Hyuk_Chang1"]},"authors":{"value":["Min-Sang Baek","Gyeong-Su Kim","Donghyun Kim","Joon-Hyuk Chang"]}},"version":2},{"content":{"venue":{"value":"SciPost Physics"},"pdf":{"value":"/pdf/9ef873a6c81a2d1c9c4217e9da3e62b3c286557e.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"rocamonde|testing_new_physics_models_with_global_comparisons_to_collider_measurements_the_contur_toolkit"},"authorids":{"value":["~Juan_Rocamonde1"]},"abstract":{"value":"Measurements at particle collider experiments, even if primarily aimed at understanding Standard Model processes, can have a high degree of model independence, and implicitly contain information about potential contributions from physics beyond the Standard Model. The Contur package allows users to benefit from the hundreds of measurements preserved in the Rivet library to test new models against the bank of LHC measurements to date. This method has proven to be very effective in several recent publications from the Contur team, but ultimately, for this approach to be successful, the authors believe that the Contur tool needs to be accessible to the wider high energy physics community. As such, this manual accompanies the first user-facing version: Contur v2. It describes the design choices that have been made, as well as detailing pitfalls and common issues to avoid. The authors hope that with the help of this documentation, external groups will be able to run their own Contur studies, for example when proposing a new model, or pitching a new search."},"title":{"value":"Testing new physics models with global comparisons to collider measurements: the Contur toolkit"},"authors":{"value":["Juan Rocamonde"]}},"tmdate":1776814390334,"pdate":1621461600000,"tcdate":1776814390334,"writers":["~Juan_Rocamonde1"],"signatures":["~Juan_Rocamonde1"],"forum":"jvNYdRLxwh","license":"CC BY 4.0","number":47262,"cdate":1776814390334,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1776814390334,"domain":"OpenReview.net/Archive","id":"jvNYdRLxwh","version":2},{"content":{"venue":{"value":"Nuclear Physics B"},"pdf":{"value":"/pdf/c2cda42619572b97ede47b9a4b78b59ffda46fc3.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"pal|multilepton_probes_of_new_physics_and_leptonuniversality_in_topquark_interactions"},"authorids":{"value":["~Kuntal_Pal1","jose.wudka@ucr.edu","yoavafik@gmail.com","shaouly@physics.technion.ac.il","adlersoni@gmail.com"]},"html":{"value":"https://doi.org/10.1016/j.nuclphysb.2022.115849"},"abstract":{"value":"We explore the sensitivity to new physics (NP) in the associated production of top-quarks with leptons pp→tt¯ℓ+ℓ−, which leads to the multi-leptons signals pp→nℓ+𝚓𝚎𝚝𝚜+⧸ET, where n=2,3,4. The NP is parameterized via 4-Fermi effective tt¯ℓ+ℓ− contact interactions of various types, which are generated by multi-TeV heavy scalar, vector or tensor exchanges in tt¯→ℓ+ℓ−; we focus on the case of ℓ=e,μ. We match the 4-Fermi ttℓℓ terms to the SMEFT operators and also give examples of specific underlying heavy physics that can generate such terms. Analysis of the SM signals and corresponding backgrounds shows that the di-lepton and tri-lepton channels are much better probes of the effective tt¯ℓ+ℓ− 4-Fermi terms than the four-lepton one at the 13 TeV LHC. Therefore, the best sensitivity is obtained in the di- and tri-lepton channels, for which the dominant background pp→tt¯ and pp→WZ, respectively, can be essentially eliminated after applying the 2ℓ and 3ℓ selections and a sufficiently high invariant mass selection for the opposite sign same flavor (OSSF) lepton-pair. We explore two cases: lepton flavor universal (LFU) NP where the ttee and ttμμ contact interactions are of same size and LFU violating (LFUV) NP, where the scale of the ttμμ terms is assumed to be much lower. We show that in both cases it is possible to obtain new 95\\% CL bounds on the scale of the ttℓℓ contact interactions at the level Λ≳2−3 TeV, which are considerably tighter than the current bounds on these 4-Fermi terms."},"title":{"value":"Multi-lepton probes of new physics and lepton-universality in top-quark interactions"},"authors":{"value":["Kuntal Pal","Jose Wudka","Yoav Afik","Shaouly Bar-Shalom","Amarjit Soni"]}},"tmdate":1756657588247,"pdate":1656658800000,"tcdate":1756657588247,"writers":["~Kuntal_Pal1","jose.wudka@ucr.edu","yoavafik@gmail.com","shaouly@physics.technion.ac.il","adlersoni@gmail.com"],"signatures":["~Kuntal_Pal1"],"forum":"57Pt6HKaf3","license":"CC BY 4.0","number":38688,"cdate":1756657588247,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1756657588247,"domain":"OpenReview.net/Archive","id":"57Pt6HKaf3","version":2},{"content":{"summary":{"value":"This paper develops a rigorous statistical-learning theory for physics-informed learning under dependent (mixing) data. It studies regularized empirical risk minimization where the regularizer encodes known physical laws through a linear differential operator, and the data arise from a stochastic dynamical system $X_{t+1}=f_{\\star}(X_t)+W_t$. Using tools from Sobolev-space analysis, the small-ball method, and martingale offset complexity, the authors prove complexity-dependent excess-risk bounds showing that when the physical prior is aligned with the ground-truth dynamics (i.e., PDE residual of $f_{\\star}$ is nearly zero), the learning rate acclerates from the traditional Sobolev minimax rate $O(T^{-d})$ to the fast i.i.d optimal rate $O(1/T)$, even when samples are correlated. A simple unicycle-dynamics experiment empirically confirms the predicted speed-up, demonstrating that properly aligned physics-based regularization can provably improve sample efficiency in learning dynamical systems."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"In relation to the weaknesses stated above, please see my questions/comments below:\n\n1. The paper presents a two-term bound combining the $T^{-d}$ and $T^{-1}$ rates, but it is unclear how to interpret the intermediate regime. Could the authors clarify whether the transition between the two regimes is smooth or abrupt, and provide any threshold conditions under which the fast term becomes dominant? In addition, it'd be great if the authors could shed some light on the following: how robust this behavior is to imperfect knowledge alignment? For example, if $\\|\\|D\\left(f_{\\star}\\right)\\|\\|_{L^2}$ is small but nonzero, does the convergence rate degrade gradually (and remain faster than $T^{-d}$ ) or does it collapse sharply to the slower regime?\n\n3. Since $\\Psi(f)=\\|\\|D(f)\\|\\|_{L^2}^2$ may involve higher-order derivatives, do the authors foresee computational or numerical challenges when scaling beyond low-dimensional synthetic problems? Any guidance on discretization or numerical stability when implementing this regularizer in practice would be helpful.\n\n3. The analysis assumes that $D$ is a linear and elliptic operator, which can be restrictive. Could the authors comment on whether any part of the analysis may extend to mildly nonlinear or non-elliptic operators, or if linear ellipticity is fundamentally required for the proof techniques used?\n\n4. The paper imposes $s \\geq 2 d_X$ to ensure the burn-in term vanishes. Is this threshold believed to be intrinsic to the problem, or could the two-phase rate behavior persist under weaker smoothness assumptions (e.g., $s>d_X$ or $s>3 d_X / 2$ ) with possibly different constants?\n\n5. Could the authors provide any preliminary numerical results or intuition on how large $T$ must be for the asymptotic behavior to manifest in practice? Even a brief discussion of the finite-sample regime would help readers assess when the theoretical rates become observable."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper addresses an important topic at the intersection of physics-informed learning and statistical learning under dependence. The authors aim to provide a theoretical foundation for when incorporating physical structure can mitigate the challenges of temporally correlated data. In this context, the paper offers several notable positive aspects:\n\n1. The central idea of using elliptic differential operators to encode physical priors and recover i.i.d.-like learning rates (under suitable alignment) despite sample correlation is novel and insightful.\n\n2. The theoretical analysis is mathematically sound. The assumptions are clearly stated, and the main theorems follow logically from the lemmas and proof techniques provided. The paper carefully extends existing theory to accommodate a Sobolev-based physics regularizer in the presence of Markovian data, which is nontrivial. While I did not look into the proofs in detail and some proof components rely on established techniques, the overall combination is coherent and technically competent.\n\n3. The results contribute to a better theoretical understanding of physics-informed learning, particularly in settings where data dependence is unavoidable, such as system identification and scientific modeling. Showing that incorporating correct physical structure can mitigate the adverse effects of temporal correlation is a useful insight, although demonstrated under idealized assumptions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the paper makes a meaningful theoretical contribution and is clearly written, several limitations temper its overall impact and practical relevance. Most of these relate to the idealized nature of the assumptions and the gap between the theory and empirical applicability. The following points highlight areas where the work could be strengthened or where additional clarification or experimentation would improve the contribution:\n\n\n1. The fast-rate $O(1 / T)$ convergence is achieved only under an idealized knowledge alignment condition, namely when $\\|\\|D\\left(f_{\\star}\\right)\\|\\|_{L^2} \\approx 0$. In practice, the physical operator $D$ is rarely known with such precision, and even modest mismatch can revert the rate to the slower $O\\left(T^{-d}\\right)$ regime. The paper does not provide a quantitative robustness analysis that would clarify how sensitive the rates are to partial or imperfect alignment, which limits the applicability of the theoretical claims in realistic settings.\n\n2. The theoretical guarantees depend on $\\lambda_T$ choices tied to latent problem quantities (e.g., $\\Psi(f_{\\star}), \\sigma_W^2$ ). While the authors note that cross-validation could be used in principle, the paper offers limited practical guidance or empirical validation for tuning $\\lambda_T$ in realistic settings.\n\n3. The single toy experiment (unicycle) is supportive but narrow (one setting, small MLP, no real-world data/baselines), so the robustness of the phase transition across architectures/noise regimes remains unclear. Even in this controlled synthetic setting, implementing the physics-informed regularization term $\\Psi(f)= \\|\\|D(f)\\|\\|_{L^2}^2$ may require evaluating higher-order derivatives and Sobolev norms, which can be computationally demanding in higher dimensions. In addition, the paper does not discuss how discretization, numerical differentiation, or instability in PDE solvers or neural approximators would affect the performance or validity of the theoretical bounds, leaving a gap between the continuous theory and practical implementation.\n\n4. The framework relies on $D$ being a known, linear elliptic operator. Many modern scientific machine learning applications involve unknown or nonlinear physics, or operators that must be learned jointly with the model (e.g., operator learning, neural PDE surrogates, and PINNs). As the current analysis does not extend to such settings, it is unclear how the insights would generalize to applications where the governing equations are only partially known or inherently nonlinear."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922857285,"tcdate":1761982540457,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11838/Reviewer_6PJM"],"signatures":["ICLR.cc/2026/Conference/Submission11838/Reviewer_6PJM"],"forum":"IvLVPbeoRx","number":2,"license":"CC BY 4.0","cdate":1761982540457,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11838/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922857285,"domain":"ICLR.cc/2026/Conference","replyto":"IvLVPbeoRx","id":"Ms1clVzy5v","forumContent":{"TLDR":{"value":"We prove that adding correct prior domain knowledge to nonparametric learning with dependent data speeds up learning"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["learning with dependent data","physics-informed machine learning","convergence rates","complexity-dependent bounds"]},"supplementary_material":{"value":"/attachment/a00585da11ea98b0c0c19b8e1c5d1c3a2b58963e.pdf"},"primary_area":{"value":"learning theory"},"abstract":{"value":"A major challenge in physics-informed machine learning is to understand how the incorporation of prior domain knowledge affects learning rates when data are dependent. Focusing on empirical risk minimization with physics-informed regularization, we derive complexity-dependent bounds on the excess risk in probability and in expectation. We prove that, when the physical prior information is aligned, the learning rate improves from the (slow) Sobolev minimax rate to the (fast) optimal i.i.d. one without sample-size deflation due to data dependence."},"_bibtex":{"value":"@inproceedings{\nscampicchio2026physicsinformed,\ntitle={Physics-informed learning under mixing: How physical knowledge speeds up learning},\nauthor={Anna Scampicchio and Leonardo Felipe Toso and Rahel Rickenbach and James Anderson and Melanie Zeilinger},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=IvLVPbeoRx}\n}"},"title":{"value":"Physics-informed learning under mixing: How physical knowledge speeds up learning"},"pdf":{"value":"/pdf/9917146046b56820383565485dace6196f5b94c5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"scampicchio|physicsinformed_learning_under_mixing_how_physical_knowledge_speeds_up_learning"},"authorids":{"value":["~Anna_Scampicchio1","~Leonardo_Felipe_Toso1","~Rahel_Rickenbach1","~James_Anderson6","~Melanie_Zeilinger1"]},"authors":{"value":["Anna Scampicchio","Leonardo Felipe Toso","Rahel Rickenbach","James Anderson","Melanie Zeilinger"]}},"version":2},{"content":{"summary":{"value":"This paper proposes SINDy-SHRED, which is a hybrid network that combines deep neural networks and physics discovery. The pipeline uses sampling methods to obtain low-dimensional sensor measurements, GRUs for learning low-dim dynamics, SINDy for incorporating physics, and a shallow decoder for reconstruction. The proposed method is flexible, easy, efficient, and robots. This method has demonstrated effectiveness across many synthetic and real-world datasets."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- This paper mentioned laptop-level computing. However, I didn’t find the computing configuration in the paper. How does this paper validate this statement? Also, this paper shows the training time in Table 1 without introducing the computing platforms the authors used.\n\n- For the implementation of SINDy loss, do you train the entire measured time sequence as one batch? Suppose you have a sensor measurement s_1 with a time duration of [0,1000]. Do GRU layers take 1000 steps as input without batches? Otherwise, why do you compute the SINDy loss?\n\n- I assume the learned latent dynamics are not sufficiently precise in terms of discovered equations. How do you measure the error propagation of latent discovered physics?\n\n- On Page 4, why does this paper claim that the proposed method is able to avoid spectral bias issues without an encoder? The spectral bias comes from the neural network itself, not just encoders. More discussions are appreciated. \n\n- How do you select the basis functions in the SINDy library? If the latent space is smooth, then you do not need to consider too many basis functions. What are the basis function libraries for the tested cases?\n\n- The authors claim the GRU produces a smoother latent space than LSTM. Do you have an ablation study on that or visualizations?\n\n- I think the superiority of the proposed method comes from the incorporation of the discovered latent physics, as shown in Figure 10. It would be good to have an ablation study to train the same pipeline without SINDy loss to give readers a better sense. \n\n- As shown in Figure 14, how does this paper obtain the original latent space?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper is interesting due to the incorporation of potential physics into the machine learning models for learning long-term dynamics. \n\n- The proposed method shows good flexibility, efficiency, and robustness.\n\n- The experiments are extensive across synthetic and real-world dynamical systems. \n\n- The paper is well-written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- More clarifications on the methodology itself and the experimental setups are appreciated. Please see many questions below.\n\n- One minor issue should be fixed. This paper does not use consistent citation formats. For example, on Page 6, the authors use “NOAA (Reynolds et al., 2002)” and “Williams et al. (2024).”"}},"nonreaders":[],"tmdate":1731427277549,"tcdate":1730154135301,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission554/Reviewer_55op"],"signatures":["ICLR.cc/2025/Conference/Submission554/Reviewer_55op"],"forum":"iyGkoWP6nA","number":2,"license":"CC BY 4.0","cdate":1730154135301,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission554/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427277549,"domain":"ICLR.cc/2025/Conference","replyto":"iyGkoWP6nA","id":"74vUquyelm","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"We introduce SINDy-SHRED, an architecture that simultaneously solves the sensing and model identification problems, achieving state-of-the-art performance for accurate and stable long-term predictions."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Sata-driven modeling","scientific machine learning","sparse identification of nonlinear dynamics","AI in dynamic systems","spatiotemporal modeling","PDEs"]},"supplementary_material":{"value":"/attachment/5ee5f36d842f8c4a7d38c8b25c095ded8b7567ed.pdf"},"primary_area":{"value":"learning on time series and dynamical systems"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Spatio-temporal modeling of real-world data is a challenging problem as a result of inherent high-dimensionality, noisy observations, and expensive data collection procedures. In this paper, we present Sparse Identification of Nonlinear Dynamics with SHallow Recurrent Decoder networks (SINDy-SHRED) to jointly solve the sensing and model identification problems with simple implementation, efficient computation, and robust performance. SINDy-SHRED utilizes Gated Recurrent Units (GRUs) to model the temporal sequence of sensor measurements along with a shallow decoder network to reconstruct the full spatio-temporal field from the latent state space using only a few available sensors. Our proposed algorithm in\u0002troduces a SINDy-based regularization. Beginning with an arbitrary latent state space, the dynamics of the latent space progressively converges to a SINDy-class functional, provided the projection remains within the set. We conduct a system\u0002atic experimental study including synthetic PDE data, real-world sensor measure\u0002ments for sea surface temperature, and direct video data. With no explicit encoder, SINDy-SHRED allows for efficient training with minimal hyperparameter tuning and laptop-level computing. SINDy-SHRED demonstrates robust generalization in a variety of applications with minimal to no hyperparameter adjustments. Additionally, the interpretable SINDy model of latent state dynamics enables accurate long-term video predictions, achieving state-of-the-art performance and outperforming all baseline methods considered, including Convolutional LSTM, PredRNN, ResNet, and SimVP."},"_bibtex":{"value":"@misc{\ngao2024sparse,\ntitle={Sparse identification of nonlinear dynamics with Shallow Recurrent Decoder Networks},\nauthor={Liyao Gao and Jan P. Williams and J. Nathan Kutz},\nyear={2024},\nurl={https://openreview.net/forum?id=iyGkoWP6nA}\n}"},"title":{"value":"Sparse identification of nonlinear dynamics with Shallow Recurrent Decoder Networks"},"pdf":{"value":"/pdf/eb5bb625ca6a446dca906ab63913740482974798.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"gao|sparse_identification_of_nonlinear_dynamics_with_shallow_recurrent_decoder_networks"},"authorids":{"value":["~Liyao_Gao2","~Jan_P._Williams1","~J._Nathan_Kutz1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Liyao Gao","Jan P. Williams","J. Nathan Kutz"]}},"version":2},{"content":{"comment":{"value":"We provide a summary of the reviewers' concerns and our responses here. Our work introduces a **new research question:** Can LMMs adapt to novel unseen scenarios by inferring underlying physical laws from a few demonstration examples? We call this “inductive physical reasoning.” Such a capability will significantly expand the deployment scope and effectiveness of LMMs into novel scenarios by simply relying on demonstration samples using this “inductive physical reasoning” rather than the prevailing solution of expensive fine-tuning. The reviewers’ negative response appears to stem from a misunderstanding of our research question and how it differs from the existing works (see the table in the global response).\n\nOur research insights **identify generalization limits and language bias of LMMs** in physical reasoning. We believe our work will inspire the research community to develop solutions to address this limitation.\n\n---\n\n**Short summary**: Our work identified a new drawback in physical reasoning using LMMs, namely, **inductive physical reasoning**, and provided a benchmark to evaluate inductive physical reasoning in LMMs. Many of the reviewers’ concerns:\n\n1. were due to **incorrectly understanding inductive physical reasoning as similar to existing research directions** such as counterfactual physical reasoning (see table in the global response for the differences), and\n2. required **clarifications about how we arrived at our conclusions**.\n\nWe directly address the reviewers’ concerns **through existing or additional experiments** (listed below).\n\n---\n\n**List of experiments added during rebuttal**:\n\nE1. **Human Baseline (App. E.9):** Even with only one demonstration sample and no QA pairs, ten human subjects achieved **~90% mean accuracy** on most scenarios. This definitively proves the task is well-defined and visually solvable.\n\nE2. **All-Frames Evaluation (App. E.10):** Performance did not improve (and often degraded) when all evaluation frames were provided. This confirms that the bottleneck is **reasoning**, not perception/motion estimation.\n\nE3. **Attention Analysis (App. F.1):** Visualizing attention maps in Gemma3-12B revealed that LMMs attended to image tokens **an order of magnitude less** than text tokens. This provides a structural explanation for the language bias we observed.\n\nE4. **Prompt Robustness (App. E.11):** We analyzed the sensitivity of evaluation accuracy to prompt perturbations. Results show a low coefficient of variation for larger models, confirming experimental stability.\n\nE5. **Visually more complex scene (App. E.8)**: LMMs performed similarly or worse on InPhyRe with visually more complex and diverse scenarios (e.g., background objects, changing camera pose, illumination direction). This suggests that performance can only worsen with further scene complexity.\n\n---\n\n---\n\n**Questions asked by the reviewers and our responses**:\n\n1. [unjk] **Is inductive physical reasoning the same as counterfactual/intuitive physical reasoning?** No. In counterfactual/intuitive physical reasoning, the evaluated model does not have to infer any unseen physical laws. The objective of inductive physical reasoning has never been considered in prior works (see table in global response).\n2. [unjk] **Are existing benchmarks sufficient?** No. To measure inductive physical reasoning, (1) the test samples must follow physical laws different from the train samples, and (2) demo samples with consistently different physical laws must be available for the LMM to infer the underlying laws on-the-fly. No such benchmark exists.\n3. [nZYb, q4tW, WEJ2] **Aren’t the observations due to a lack of training data with impossible physics?** If the observed poor performance was only due to impossible physics, then LMMs would have performed well in our SB scenario that didn’t violate any true physics. This is evidence that LMMs indeed have a problem with reasoning from demo samples.\n4. [nZYb, q4tW, WEJ2] **Will observations from proposed InPhyRe generalize to more complex scenes?** Our observations were made using (a) physically well-understood scenarios, (b) without confounding visual factors. In a more complex scene, performance can only get worse (Exp. E5).\n5. [q4tW] **Is the broader goal of InPhyRe to improve LMM reasoning on impossible physics?** No, impossible physics is our methodology, not the end goal. Our broader goal is adaptive physical reasoning using LMMs.\n6. [unjk] **Is the observed language bias due to template copying?** If the template copying hypothesis were true, then LMMs would have performed well in irregular scenarios with question-answer pairs in demos by directly matching the answer templates from demo and evaluation samples."},"title":{"value":"Rebuttal Summary for AC (part 1/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1770917940904,"tcdate":1764552392550,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14468/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission14468/Authors"],"forum":"IIrPoZ28dN","number":10,"license":"CC BY 4.0","cdate":1764552392550,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14468/-/Official_Comment","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770917940904,"domain":"ICLR.cc/2026/Conference","replyto":"IIrPoZ28dN","id":"D3zxRW1lWA","forumContent":{"TLDR":{"value":"We evaluate whether large multimodal models can infer and apply unseen physical laws from demonstration samples."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["large multimodal models","physical reasoning","inductive reasoning"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large multimodal models (LMMs) encode universal physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potential collision event from visual input. However, since parametric knowledge includes only the physical laws seen during training, it is insufficient for reasoning when the inference scenario follows physical laws unseen during training. In contrast, humans can adapt their physical reasoning to unseen physical environments with only a few visual examples. This inductive physical reasoning ability is indispensable for LMMs if they are to replace human agents in safety-critical applications. Despite its importance, existing visual benchmarks evaluate only the parametric knowledge in LMMs, and not inductive physical reasoning. To this end, we propose InPhyRe, the first visual question answering benchmark to measure inductive physical reasoning in LMMs. InPhyRe evaluates LMMs on their ability to predict the outcome of collision events in algorithmically generated synthetic videos. By inspecting over 13 open-source and proprietary LMMs, InPhyRe informs us that (1) LMMs struggle to apply their limited parametric knowledge about universal physical laws to reasoning, (2) inductive physical reasoning in LMMs is weak when inference scenarios obey physical laws unseen during training, and (3) inductive physical reasoning in LMMs suffers from language bias and largely ignores the visual inputs, questioning the trustworthiness of LMMs regarding visual inputs."},"_bibtex":{"value":"@misc{\nsreekumar2026inphyre,\ntitle={InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning},\nauthor={Gautam Sreekumar and Vishnu Boddeti},\nyear={2026},\nurl={https://openreview.net/forum?id=IIrPoZ28dN}\n}"},"title":{"value":"InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning"},"pdf":{"value":"/pdf/b87a0e49daf5f5f05caff235d97e823b17b5dbfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sreekumar|inphyre_discovers_large_multimodal_models_struggle_in_inductive_physical_reasoning"},"authorids":{"value":["~Gautam_Sreekumar1","~Vishnu_Boddeti1"]},"authors":{"value":["Gautam Sreekumar","Vishnu Boddeti"]}},"version":2},{"content":{"summary":{"value":"This paper studies domain shift caused by image-quality degradations in medical imaging, focusing on the gap between well-curated training data and noisier real-world clinical acquisitions. The authors introduce synthetic, physics-informed corruption pipelines for chest X-rays and dermoscopy, covering radiograph-specific degradations (e.g., acquisition geometry, anti-scatter grid artefacts, beam energy, collimation, detector performance, focal spot, tube current) and dermoscopy degradations (e.g., blur, colour reproduction, focal-plane effects, resolution). They then benchmark several segmentation models (U-Net, Swin-Unet, SAM) and classification models (ResNet-18, ConvNeXt-tiny, ViT-B/16) trained on (i) clean data, (ii) standard augmentations (blur/noise/brightness), and (iii) the proposed corruptions. They include auxiliary tasks for predicting corruption severity and corruption type."},"justification_of_final_rating":{"value":"I would like to thank the authors for their work on the rebuttal. The addition of in-depth discussions about the corruption setups and reliance on Table 1 improves the paper, but still, the results remain weak."},"confidence":{"value":4},"final_rating":{"value":3},"justification_of_the_preliminary_rating":{"value":"I recommend a weak reject due to the limited strength of the evidence supporting the main robustness claims: the experimental section reports a single metric per task (Dice for segmentation, accuracy for classification), which is insufficient to substantiate robustness under distribution shift (e.g., no severity-wise robustness curves, worst-case performance, fairness across subgroups, calibration/uncertainty, or variability across seeds/splits). In addition, Figures 2–5 mainly show training with one corruption type at a time, which is informative but does not directly answer whether physics-based corruption yields robustness under realistic mixtures of corruptions and severities. The paper contains a promising and interesting methodological contribution, but would need stronger and clearer robustness evaluation (additional metrics, severity-based analysis, mixed-corruption training/testing, and uncertainty/variance reporting) to be published. I doubt these additional experiments could be conducted during the rebuttal and thus recommend that the authors consider reworking the paper before re-submitting it."},"confidentiality_llm_acknowledgment":{"value":"Yes"},"strengths":{"value":"* The paper addresses a timely and important challenge: image-quality shift is a realistic deployment bottleneck in medical AI, and the paper tackles it directly.\n* Using physics-based degradations is a stronger and more clinically aligned approach than purely generic perturbations, and the paper provides substantial methodological detail. \n* The manuscript positions itself clearly among prior corruption/augmentation and physics-driven generalization literature. \n* Evaluating both segmentation and classification, including a foundation model (SAM), is valuable for readers trying to understand whether robustness trends generalize."},"weaknesses":{"value":"* Robustness is not fully operationalized in the evaluation: the experiments primarily report a single metric per task (Dice for segmentation, accuracy for classification). For robustness claims, additional metrics are usually necessary (e.g., calibration/uncertainty, worst-case performance, stability across severity levels, and class-imbalance–aware metrics). \n* The protocol lacks clarity: what distribution is the model evaluated on? From Figures 2–5, it is not always clear whether performance corresponds to (a) testing on clean images, (b) testing on the same corruption used in training, or (c) testing across corruption types/severities (including unseen ones). This makes it hard to interpret the robustness conclusions. \n* The results presentation also makes robustness conclusions difficult to interpret. The barplots show training with one corruption type at a time, which is helpful to evaluate the potential of each type of corruption, but does not answer the key question: does physics-based corruption-aware training improve performance under a realistic mixture of corruptions and severities? \n* The results are inconsistent across tasks and models, which weakens actionable takeaways. The paper highlights variability, but the current presentation does not distill which physics-based corruptions are reliably beneficial, nor when they outperform standard augmentation (and under which evaluation conditions). \n* Table 1 is underused. The auxiliary “corruption severity/type prediction” task is an interesting addition, yet it is only briefly discussed. The near-zero performance for some settings (e.g., “No Corruption” for certain models) deserves explanation, and the connection to the main robustness story could be strengthened."},"detailed_comments":{"value":"1. Clarify the train/test corruption setup for every figure/table by stating explicitly which corruptions/severities are used in training AND/OR in testing.\n\n2. Report robustness as a function of severity (curves, not only bars). For each model: performance vs. severity level (1–4), and summary metrics such as worst-case performance, calibration, or fairness metrics across image qualities.\n\n3. Add uncertainty/stability reporting. Re-run with multiple seeds and/or splits and provide confidence intervals. This is particularly important given the dataset size and the variability already observed.\n\n4. Include a “mixed-corruption” training baseline. In addition to “single corruption type” training, include a setting in which the training data contains a realistic mixture of corruption types and severities: this would align more directly with the paper’s robustness goal. Relatedly, adding a selection step to justify the choice of this \"mixture of corruption\" could help clarify which type of corruption is most effective.\n\n5. Improve reliance on Table 1 in the main narrative. In particular, explain: why some models fail on “no corruption” or particular corruption types, and whether corruption prediction correlates with downstream robustness."},"questions_to_address_in_the_rebuttal":{"value":"Main focus points: \n- Clarify the methodology and implementation for the train/test corruption setup. \n- Add robustness metrics."},"preliminary_rating":{"value":2}},"parentInvitations":"MIDL.io/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1771530042015,"tcdate":1767783337838,"writers":["MIDL.io/2026/Conference","MIDL.io/2026/Conference/Submission214/Reviewer_Dz8m"],"signatures":["MIDL.io/2026/Conference/Submission214/Reviewer_Dz8m"],"forum":"fPXKcRF9fb","number":1,"license":"CC BY 4.0","cdate":1767783337838,"readers":["everyone"],"invitations":["MIDL.io/2026/Conference/Submission214/-/Official_Review","MIDL.io/2026/Conference/-/Edit"],"mdate":1771530042015,"domain":"MIDL.io/2026/Conference","replyto":"fPXKcRF9fb","id":"7sqCV2yaye","forumContent":{},"version":2},{"content":{"summary":{"value":"This paper provides a theoretical analysis of why SSM needs to be parameterized by complex numbers instead of real numbers. It shows that there exist complex LTI systems that could not be well-approximated by real systems of comparable size. Moreover, it proves that certain dynamics cannot be approximated by real LTI systems unless it has an exponentially large state-space dimension or large system matrices; yet, using complex LTI systems resolves this issue."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. You formulate the problem using the discrete LTI system. Of course, in SSMs, the parameters are from the continuous-time LTI systems. While this does not change the basic question you are exploring because the real axis in the left half-plane gets mapped to the real axis in the unit disk under virtually all discretization schemes, would Theorem 2 be changed if you take the discretization into account?\n2. The paper studies two cases: $\\mathbf{A}$, $\\mathbf{B}$, and $\\mathbf{C}$ are either all restricted to real or all allowed to be complex. Intuitively, however, the important thing is that $\\mathbf{A}$ has to be complex. Have you looked into the case where $\\mathbf{A}$ is complex and $\\mathbf{B}$ and $\\mathbf{C}$ are real? In that case, which world would it fall into?\n3. In section 3.3, instead of giving two examples of dynamics that are poorly approximated by real systems, maybe there could be a discussion of what you called forward difference. It would be helpful to relate Theorem 2 to some general (and easily interpretable) properties of the dynamics."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The theoretical statements are precisely made. The sketch of the proofs are helpful for understanding the paper.\n2. The comparison between real and complex is fairly thorough, encompassing different perspectives."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. My main concern is about the contribution of this work. While it is true that many ML models use real parameterizations, the diagonal matrix $\\mathbf{A}$ comes from diagonalizing a general state matrix. Therefore, unless one puts restrictions on the matrix to be diagonalized (e.g. Hermitian), it is natural to assume that $\\mathbf{A}$ should be complex-valued. Showing why a real parameterization does not work well sounds like a slightly artificial problem and adds relatively little to the SSM community.\n2. The experiments could not strongly corroborate the theory. In addition to showing the performance, maybe some synthetic experiments would be helpful to show the exponential gap in Theorem 2 and Proposition 3."},"limitations":{"value":"None."}},"nonreaders":[],"tmdate":1730879946242,"tcdate":1718228066803,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission18021/Reviewer_UNm3"],"signatures":["NeurIPS.cc/2024/Conference/Submission18021/Reviewer_UNm3"],"forum":"h15RyEj151","number":1,"license":"CC BY 4.0","cdate":1718228066803,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission18021/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879946242,"domain":"NeurIPS.cc/2024/Conference","replyto":"h15RyEj151","id":"A3eet4IOWT","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"We establish formal gaps between real and complex SSMs in terms of expressiveness and practical learnability"},"keywords":{"value":["Neural Networks","Theory","Structured State Space Models","Mamba","S4","Complex parametrization"]},"primary_area":{"value":"learning_theory"},"abstract":{"value":"Structured state space models (SSMs), the core engine behind prominent neural networks such as S4 and Mamba, are linear dynamical systems adhering to a specified structure, most notably diagonal. In contrast to typical neural network modules, whose parameterizations are real, SSMs often use complex parameterizations. Theoretically explaining the benefits of complex parameterizations for SSMs is an open problem. The current paper takes a step towards its resolution, by establishing formal gaps between real and complex diagonal SSMs. Firstly, we prove that while a moderate dimension suffices in order for a complex SSM to express all mappings of a real SSM, a much higher dimension is needed for a real SSM to express mappings of a complex SSM. Secondly, we prove that even if the dimension of a real SSM is high enough to express a given mapping, typically, doing so requires the parameters of the real SSM to hold exponentially large values, which cannot be learned in practice. In contrast, a complex SSM can express any given mapping with moderate parameter values. Experiments corroborate our theory, and suggest a potential extension of the theory that accounts for selectivity, a new architectural feature yielding state of the art performance."},"_bibtex":{"value":"@inproceedings{\nran-milo2024provable,\ntitle={Provable Benefits of Complex Parameterizations for Structured State Space Models},\nauthor={Yuval Ran-Milo and Eden Lumbroso and Edo Cohen-Karlik and Raja Giryes and Amir Globerson and Nadav Cohen},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=h15RyEj151}\n}"},"title":{"value":"Provable Benefits of Complex Parameterizations for Structured State Space Models"},"pdf":{"value":"/pdf/36066871b1ea5ba0d5c2126254596aa0019d538c.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"ranmilo|provable_benefits_of_complex_parameterizations_for_structured_state_space_models"},"authorids":{"value":["~Yuval_Ran-Milo1","~Eden_Lumbroso1","~Edo_Cohen-Karlik1","~Raja_Giryes1","~Amir_Globerson1","~Nadav_Cohen1"]},"authors":{"value":["Yuval Ran-Milo","Eden Lumbroso","Edo Cohen-Karlik","Raja Giryes","Amir Globerson","Nadav Cohen"]}},"version":2},{"content":{"summary":{"value":"\nThis paper tackles the task of illumination-aware conditional image repainting. Given an input image and a set of conditions, the proposed method aims to inpaint / re-generate a certain region based on the input conditions. This can be used to achieve functionalities such as object insertion and image composition. Compared to prior works, this paper is with the goal of injecting physics-based illumination information into the image generation process. \n\nIn a high-level, instead of formulating this task as a simple image-to-image translation in 2D image space, this work aims to introduce explicit physics-based rendering in 3D into a 2D neural renderer. This can be achieved by incorporating physics-based rendering buffers. To enable training and evaluation of the method, the authors also curate a photorealistic synthetic dataset with material and lighting conditions. \n\nThe results of the proposed method is qualitative visualized and quantitatively evaluated. A user study is included to compare the photorealism of the edited results. The proposed method can significantly outperform baselines wrt lighting effects. \n"},"soundness":{"value":"3 good"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Please see weaknesses section above. "},"rating":{"value":"7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"4 excellent"},"contribution":{"value":"4 excellent"},"strengths":{"value":"In general I find this paper with a sufficient amount of workload and technically solid. \n\nOriginality: \n\n- The task definition is well motivated. The analysis on why we need 3D information in conditional generation is generally informative and convincing.\n- The proposed method is sensible and novel. Despite a complicated pipeline, it presents a smart approach to inject physics-based rendering process into a 2D neural renderer. \n\nQuality: \n\n- The qualitative and quantitative results outperform baselines and achieves SOTA. \n\n\nClarity: \n\n- This paper is well written and easy to follow. \n- The descriptions on method details in main paper and supp are thorough. \n\n\nSignificance: \n\n- This paper proposes a carefully designed approach for illumination-aware image generation, which has not been extensively explored in recent generative models.\n- The proposed dataset can be beneficial for future research works. \n\n"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"In general I do not find critical concerns of this paper but have some questions to further elicit insights: \n\n- The proposed lighting representation is a slightly modified version of prior works. How does the parametric light representation (in Eq.9) compare to prior sky models [22, 23, 32, 63]? \n- The model is trained on synthetic data, which can be a concern when the ultimate goal is to apply on real-world imagery. How well does it work on real-world images, and how to measure the domain gap? \n- What is the core advantage of generative repainting compared to fully physics-based lighting estimation methods such as SOLID-Net? \n\nThe motivation of conditional image repainting is still a relatively small scope. The authors could consider including discussion of these works in related works. For explicit lighting estimation, a line of work estimates 3D lighting volume: \n- Wang et al. Neural Light Field Estimation for Street Scenes with Differentiable Virtual Object Insertion \n- Li et al. Spatiotemporally Consistent HDR Indoor Lighting Estimation \n\nIn a similar spirit to this paper, many works in relighting and neural rendering also combine neural modules with PBR. For example, \n- Philip et al. Multi-view Relighting using a Geometry-Aware Network \n- Pandey et al. Total Relighting: Learning to Relight Portraits for Background Replacement \n\n"},"limitations":{"value":"The limitations and failure cases are discussed in paper and supp. "}},"nonreaders":[],"tmdate":1702410835813,"tcdate":1688717507537,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission2461/Reviewer_Ex3g"],"signatures":["NeurIPS.cc/2023/Conference/Submission2461/Reviewer_Ex3g"],"forum":"9UxUTGCteW","number":4,"license":"CC BY 4.0","cdate":1688717507537,"mdate":1702410835813,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission2461/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"9UxUTGCteW","id":"POzCTt17gt","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Illumination","Image Generation","Conditional Image Repainting"]},"supplementary_material":{"value":"/attachment/4f98f8ca7375343ee5170308585780d582a7d96d.pdf"},"_bibtex":{"value":"@inproceedings{\ntang2023luminaire,\ntitle={Lumin{AIR}e: Illumination-Aware Conditional Image Repainting for Lighting-Realistic Generation},\nauthor={Jiajun Tang and Haofeng Zhong and Shuchen Weng and Boxin Shi},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=9UxUTGCteW}\n}"},"title":{"value":"LuminAIRe: Illumination-Aware Conditional Image Repainting for Lighting-Realistic Generation"},"paperhash":{"value":"tang|luminaire_illuminationaware_conditional_image_repainting_for_lightingrealistic_generation"},"TLDR":{"value":"integrating explicit illumination constraints into current image geration pipeline with new dataset proposed."},"abstract":{"value":"We present the ilLumination-Aware conditional Image Repainting (LuminAIRe) task to address the unrealistic lighting effects in recent conditional image repainting (CIR) methods. The environment lighting and 3D geometry conditions are explicitly estimated from given background images and parsing masks using a parametric lighting representation and learning-based priors. These 3D conditions are then converted into illumination images through the proposed physically-based illumination rendering and illumination attention module. With the injection of illumination images, physically-correct lighting information is fed into the lighting-realistic generation process and repainted images with harmonized lighting effects in both foreground and background regions can be acquired, whose superiority over the results of state-of-the-art methods is confirmed through extensive experiments. For facilitating and validating the LuminAIRe task, a new dataset Car-LuminAIRe with lighting annotations and rich appearance variants is collected."},"pdf":{"value":"/pdf/bb9aa3d889bc285de59d81db06a1a33a910389dd.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Jiajun_Tang2","~Haofeng_Zhong1","~Shuchen_Weng1","~Boxin_Shi3"]},"authors":{"value":["Jiajun Tang","Haofeng Zhong","Shuchen Weng","Boxin Shi"]}},"version":2},{"content":{"summary":{"value":"This paper tackles the problem of generating text-conditioned character animation that is physics-based. It proposes a two-stage approach that first generates a kinematic motion conditioned on text using diffusion, and then tracks this motion with a physics-based motion VAE in the learned latent space. Experiments on the common KIT-ML and HumanML3D benchmarks show improved performance over prior work in physics-based motion from text, and show the ability to specify goal waypoints for the motion to hit while following human text prompts."},"soundness":{"value":"2 fair"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Overall, I think a general physics-based text-to-motion model is an important and novel direction, and the hierarchical approach of diffusion planning with physics-based tracker could be a strong baseline going forward. But I’m mainly concerned that the quality of the output from the high-level diffusion planner is compromising the comparison to DReCon and therefore the need for a latent motion VAE tracker has not been fully justified. I would really like to see how InsActor performs when plugging in a SOTA diffusion model like MDM [33] as the planner. Moreover, an evaluation of standalone tracking performance would make the comparison between the two tracking approaches (latent vs state-based) much more clear. \n\nSome other comments and suggestions that didn’t have an influence on my rating: \n* The related work (Sec 2) is missing relevant physics-based human animation methods and a discussion of why they are difficult to scale up to the general text-to-motion task. E.g. [Peng et al., ASE: Large-Scale Reusable Adversarial Skill Embeddings for Physically Simulated Characters, SIGGRAPH 2022] [Won et al., A Scalable Approach to Control Diverse Behaviors for Physically Simulated Characters, SIGGRAPH 2020], etc..\n* Concurrent work PhysDiff [Yuan et al., arxiv 2022] is an alternative approach to adding physicality to text-to-motion diffusion and could be discussed in future revisions of the paper. Also related, concurrent work Trace and Pace [Rempe et al., CVPR 2023] gives controllability over physics based characters with guidance of a diffusion planner.\n\n===================== After Rebuttal ============================\n\nAfter considering other reviews and discussions with authors and between reviewers, I have decided to slightly raise my score and am leaning towards accept. I think the paper lays out a compelling kinematic diffusion + physics-based tracking idea that can serve as a baseline and inspire improvement in each component of the system including differentiable simulation, motion diffusion, and physics-based tracking. The evaluations show that this hierarchical approach works better than state+action diffusion, and that tracking with a latent model is more robust than direct state-based tracking for planned motions from MotionDiffuse.\n\nHowever, I am still quite concerned about the qualitative results and would really encourage the authors to update the paper text to discuss these qualitative issues such that future work can pursue important directions (e.g., the choice of Brax and differentiable simulation in general rather than RL, and the noisy plans from MotionDiffuse especially in the waypoint setting). It would also be good to clarify why in Table 2 of the rebuttal doc, motion quality (FID) drops significantly from kinematic “Planner” (b) output to full physics-based InsActor (c), indicating that adding physics-based tracking is not necessarily improving motion realism despite it being physically constrained (unlike in PhysDiff). I encourage the authors to show some video results of the planner output vs InsActor on regular text-to-motion (not the waypoint setting) to demonstrate the difference and potential advantages/disadvantages of using the latent tracking technique vs RL and more reliable simulators as in DReCon. \n\n=========================================================="},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"strengths":{"value":"Physics-driven text-to-motion is an important problem and conditioning on free-form text input has not really been tackled in the literature, so the paper is novel in that respect. Diffusion has shown promising results recently, but these are all kinematic.\n\nThe proposed idea of using high-level diffusion followed by physics-based tracking is simple and solid. It would make a good first baseline for future work in this area. \n\nTechnically, InsActor uses a physics-based motion VAE to track kinematic motion which is novel, and it uses differentiable physics to train rather than a learned world model or RL. If this approach really does work to track general motions in the HumanML3D dataset, it would contribute an alternative to recent RL approaches that can be difficult to generalize. \n\nThe proposed method is evaluated on HumanML3D and KIT, which are the most relevant benchmarks for the text-to-motion task. Also, the DReCon baseline in Tables 1 and 2 is an important baseline that uses the alternative state-based approach to tracking.\n\nThe supplementary video and figures are visually pleasing, and I appreciate that the supp video is extensive and shows many results. The shown demo is also cool, and demonstrates fast generation capabilities (relative to other diffusion approaches)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"To better motivate the need for physics in text-to-motion, there should be a comparison between the proposed InsActor and a state-of-the-art kinematic diffusion model like MDM [33] added to Tables 1 and 2. Currently some numbers in these tables, e.g. FID and diversity, are worse than those reported in MDM, and it’s not clear why that’s the case since the high-level planner in InsActor is very similar. Does this high-level planner perform worse than MDM? Or does the physics-based motion tracking somehow have a large effect on motion quality and diversity? The high-level policy ablation reported in Table 4 takes a step in this direction, but it is trained on rollouts from the motion VAE instead of directly on mocap data as done in MDM and other text-to-motion diffusion models.\n\nLooking at video results at 0:46 and 01:54, I don’t think the high-level diffusion planner is on par with recent models like MDM. Both with and without waypoint guidance there are some significant artifacts like jittering and skating, some of which seem to be affecting the final motion from InsActor (e.g. some noisy popping of limbs and unnatural sliding). I understand this planner is not necessarily the main contribution, but I think poor kinematic motions from the planner undermines the comparison to the target-state tracking policy DReCon, which may perform better when operating on, e.g., outputs from MDM that better reflect realistic motion. Since the low-level motion VAE model for tracking (Sec 4.2) is a key contribution of the work, it’s very important to justify that it is necessary by showing that the DReCon baseline is still inferior when operating on more reasonable kinematic inputs.\n\nThe methods Sec 4 is missing some details that could improve understanding and reproducibility: \n* The tasks states described in L92 are all local, so how is the global root trajectory modeled in the diffusion Sec 4.1? \n* L154: if the pose state is in the local frame, how is inpainting performed to ensure motion meets a global target waypoint?\n* What is the architecture of the diffusion model? Is it using 1D convolutions as in Diffuser or a transformer as in other human motion diffusion models? $\\mu$ and $\\Sigma$ in Eqn 2 are never defined. In general, I’m wondering why not use a SOTA motion diffusion model out-of-the-box for this high-level planning component?\n* L185: is the low-level motion VAE trained directly on outputs of the diffusion model or on mocap data from the dataset? If on mocap, why is the encoder (Eqn 4) expected to produce reasonable results when operating on noisy and unrealistic pose transitions? \n* Similarly, Sec 5.3 shows robustness to perturbations from boxes, but is this kind of perturbation seen in training of the low-level policy too? If not, how does this robustness arise without using RL for training (i.e. without some exploration). \n\nAn evaluation on the low-level tracking component by itself would be very helpful. E.g. reporting tracking errors for the latent policy from InsActor compared to DReCon for both motion-captured and diffusion-generated motions. The current metrics in Tables 1 and 2 were designed for kinematic text-to-motion models, and I would think are mostly influenced by the diffusion planner which is the same for InsActor and DReCon, so a tracking-only evaluation could help parse the difference in performance. There are also open-source RL physics-based trackers that may be worth considering, e.g. [Luo et al., Dynamics-Regulated Kinematic Policy for Egocentric Pose Estimation, NeurIPS 2021]."},"limitations":{"value":"Limitations are sufficiently discussed in Sec 6."}},"nonreaders":[],"tmdate":1702410785681,"tcdate":1688421403631,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1407/Reviewer_s1C4"],"signatures":["NeurIPS.cc/2023/Conference/Submission1407/Reviewer_s1C4"],"forum":"hXevuspQnX","number":2,"license":"CC BY 4.0","cdate":1688421403631,"mdate":1702410785681,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1407/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"hXevuspQnX","id":"DkDJVwQSMg","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Physics-based Animation; Human Motion Generation"]},"supplementary_material":{"value":"/attachment/a49a89bbc360367c9de8d7d999d042f7be8de7ed.zip"},"_bibtex":{"value":"@inproceedings{\nren2023insactor,\ntitle={InsActor: Instruction-driven Physics-based Characters},\nauthor={Jiawei Ren and Mingyuan Zhang and Cunjun Yu and Xiao Ma and Liang Pan and Ziwei Liu},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=hXevuspQnX}\n}"},"title":{"value":"InsActor: Instruction-driven Physics-based Characters"},"paperhash":{"value":"ren|insactor_instructiondriven_physicsbased_characters"},"TLDR":{"value":"We generate motions for physically-based characters from human instructions."},"abstract":{"value":"Generating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animations that reflect high-level human instructions remains a difficult problem due to the complexity of physical environments and the richness of human language. \nIn this paper, we present $\\textbf{InsActor}$, a principled generative framework that leverages recent advancements in diffusion-based human motion models to produce instruction-driven animations of physics-based characters.\nOur framework empowers InsActor to capture complex relationships between high-level human instructions and character motions by employing diffusion policies for flexibly conditioned motion planning.\nTo overcome invalid states and infeasible state transitions in planned motions, InsActor discovers low-level skills and maps plans to latent skill sequences in a compact latent space. \nExtensive experiments demonstrate that InsActor achieves state-of-the-art results on various tasks, including instruction-driven motion generation and instruction-driven waypoint heading. Notably, the ability of InsActor to generate physically simulated animations using high-level human instructions makes it a valuable tool, particularly in executing long-horizon tasks with a rich set of instructions. Our project page is available at [jiawei-ren.github.io/projects/insactor/index.html](https://jiawei-ren.github.io/projects/insactor/index.html)"},"pdf":{"value":"/pdf/451a0328421292fd32d77f3ee70ec51fbec5546f.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Jiawei_Ren1","~Mingyuan_Zhang1","~Cunjun_Yu1","~Xiao_Ma2","~Liang_Pan2","~Ziwei_Liu1"]},"authors":{"value":["Jiawei Ren","Mingyuan Zhang","Cunjun Yu","Xiao Ma","Liang Pan","Ziwei Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes RareFlow which is designed OOD-robust super-resolution in remote sensing, especially rare geomorphic features across heterogeneous sensors. The method combines flow-matching with dual conditioning: a Gated ControlNet to preserve fine-grained geometry from the LR input and text prompts to steer complex feature synthesis. Training uses physics-aware losses to enforce spectral/radiometric consistency with sensor properties, and a stochastic forward-pass uncertainty estimates unfamiliar inputs to mitigate hallucinations. A curated cross-sensor benchmark shows large qualitative gains and almost 40% FID reduction over SOTA baselines."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"Check the weakness."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1) They designed losses tied to sensor spectra which helps keeping outputs physically plausible, not just visually sharp. This also ensure that the large-scale radiometric information are preserved and the hallucination is prevented.\n\n2) A gated ControlNet (for structure) plus text prompts (for semantics) jointly handle appearance and geometry under distribution shift.\n\n3) They got a strong performance on benchmark compared to SOTA methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) How sensitive are the results to prompt wordings? Can you do some ablation on this? \n\n2) Do the physics losses transfer to unseen sensors and bands without retraining? Can you provide any zero-shot evaluation?\n\n3) Do the authors have any intuition why PSNR is low comparatively compared to other metrics in Table 3?\n\n4)  Do the author have some ablations on different parts of the losses. Which components (spectral vs. radiometric) drive more gains? Provide loss term ablations."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918663369,"tcdate":1762125346326,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6375/Reviewer_U7Ez"],"signatures":["ICLR.cc/2026/Conference/Submission6375/Reviewer_U7Ez"],"forum":"9iAWnRGTAU","number":4,"license":"CC BY 4.0","cdate":1762125346326,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6375/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918663369,"domain":"ICLR.cc/2026/Conference","replyto":"9iAWnRGTAU","id":"GQupSLyQax","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We present RareFlow, a physics-aware SR framework designed for OOD robustness."},"keywords":{"value":["satellite images","remote sensing","diffusion models"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Super-resolution (SR) for remote sensing imagery often fails under out-of-distribution (OOD) conditions, such as rare geomorphic features captured by diverse sensors, producing visually plausible but physically inaccurate results.We present RareFlow, a physics-aware SR framework designed for OOD robustness. RareFlow's core is a dual-conditioning architecture. A Gated ControlNet preserves fine-grained geometric fidelity from the low-resolution input, while textual prompts provide semantic guidance for synthesizing complex features. To ensure physically sound outputs, we introduce a multifaceted loss function that enforces both spectral and radiometric consistency with sensor properties. Furthermore, the framework quantifies its own predictive uncertainty by employing a stochastic forward pass approach; the resulting output variance directly identifies unfamiliar inputs, mitigating feature hallucination.We validate RareFlow on a new, curated benchmark of multi-sensor satellite imagery. In blind evaluations, geophysical experts rated our model's outputs as approaching the fidelity of ground truth imagery, significantly outperforming state-of-the-art baselines. This qualitative superiority is corroborated by quantitative gains in perceptual metrics, including a nearly 40\\% reduction in FID. RareFlow provides a robust framework for high-fidelity synthesis in data-scarce scientific domains and offers a new paradigm for controlled generation under severe domain shift."},"_bibtex":{"value":"@misc{\nfallah2025rareflow,\ntitle={RareFlow: Physics-Aware Flow-Matching for Cross-Sensor Super-Resolution of Rare-Earth Features},\nauthor={Forouzan Fallah and Wenwen Li and Chia-Yu Hsu and Hyunho Lee and Yezhou Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=9iAWnRGTAU}\n}"},"title":{"value":"RareFlow: Physics-Aware Flow-Matching for Cross-Sensor Super-Resolution of Rare-Earth Features"},"pdf":{"value":"/pdf/8109a4af23fac37fb408d7f598f88ffd279e1e02.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"fallah|rareflow_physicsaware_flowmatching_for_crosssensor_superresolution_of_rareearth_features"},"authorids":{"value":["~Forouzan_Fallah1","~Wenwen_Li1","~Chia-Yu_Hsu2","~Hyunho_Lee5","~Yezhou_Yang1"]},"authors":{"value":["Forouzan Fallah","Wenwen Li","Chia-Yu Hsu","Hyunho Lee","Yezhou Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes PCDM (physics-constrained diffusion model), an inverse problem solver that leverages diffusion model as plug-and-play prior. PCDM uses the idea of variable splitting and proposes to solve the underlying optimization problem with implicit diffusion model regularization. The authors demonstrate its application in full-waveform inversion, data assimilation, and topology optimization."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. I'm a bit surprised at how well the Opt w/o diff baseline can recover the large structure of the ground truth, as shown in Figure 2 and Table 1. This contrasts with traditional FWI literature findings [1] and my own experimental validation on OpenFWI dataset. I'm curious how the authors implement the FWI problem and the corresponding baselines. More specifically, \n\t1. What is exactly the Opt w/o diff baseline in Table 1? Is that the Adam optimizer? What initialization strategy was employed? What are the specific hyperparameters used to report the results? \n\t3. Why are the residuals of InversionNet and VelocityGAN omitted from Table 1?  \n\t4. Given that OpenFWI paper does not provide the gradient implementation of the forward model, how did the authors implement the gradient? \n2. What are the hyperparameter selection criteria across compared methods? \n3. Is there any supplementary material or code to facilitate the reproducibility?\n\n[1] : Virieux, Jean, and Stéphane Operto. \"An overview of full-waveform inversion in exploration geophysics.\" _Geophysics_ 74.6 (2009): WCC1-WCC26."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Applying plug-and-play diffusion model methods to physics-constrained inverse problems is relatively new to the diffusion model community. \n2. The paper is generally easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed PCDM appears to be mathematically equivalent to a special case of the algorithm in Li et al. [1] (specifically, the case using Tweedie's formula). The claim of algorithmic novelty is questionable (line 100). \n2. The \"physics-constrained\" aspect really comes from the inverse problem itself instead of the novel algorithmic design. Most existing gradient-based plug-and-play diffusion model methods can incorporate that physics loss such as DiffPIR [2], DPS,DAPS [3], RED-diff [4], [5]. These methods are not compared or  discussed in the paper. \n3. The experimental comparison excludes many recent and relevant algorithms. For example DiffPIR [2] and DAPS [3], RED-diff [4].  \n4. Reproducibility concerns: important experimental and implementation details are insufficiently documented. See more concrete questions in the next section.\n5. There is a lack of ablation studies on important algorithm design parameters, such as the number of likelihood steps per iteration, the optimization threshold $t_s$, and sensitivity to the optimizer configurations. \n\n\n[1] : Li, Xiang, et al. \"Decoupled data consistency with diffusion purification for image restoration.\" _arXiv preprint arXiv:2403.06054_ (2024).\n[2] : Zhu, Yuanzhi, et al. \"Denoising Diffusion Models for Plug-and-Play Image Restoration.\" _arXiv preprint arXiv:2305.08995_ (2023).\n[3] : Zhang, Bingliang, et al. \"Improving diffusion inverse problem solving with decoupled noise annealing.\" _arXiv preprint arXiv:2407.01521_ (2024).\n[4] : Mardani, Morteza, et al. \"A Variational Perspective on Solving Inverse Problems with Diffusion Models.\" _The Twelfth International Conference on Learning Representations_.\n[5] : Peng, Xinyu, et al. \"Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance.\" _Forty-first International Conference on Machine Learning_. 2024."}},"nonreaders":[],"tmdate":1731429382175,"tcdate":1730352926517,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12020/Reviewer_DxjV"],"signatures":["ICLR.cc/2025/Conference/Submission12020/Reviewer_DxjV"],"forum":"Da3j02cHe0","number":2,"license":"CC BY 4.0","cdate":1730352926517,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12020/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429382175,"domain":"ICLR.cc/2025/Conference","replyto":"Da3j02cHe0","id":"fDb3r4NHHV","forumContent":{"TLDR":{"value":"We propose a novel framework for solving physics-constrained inverse problems by integrating physics constraints and diffusion models."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["physics-constraints inverse problem","diffusion model","PDE","generative modeling"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Solving inverse problems in scientific and engineering domains often involves complex, nonlinear forward physics and ill-posed conditions. \nRecent advancements in diffusion model have shown promise for general inverse problems, yet their application to scientific domains remains less explored and is hindered by the complexity and high non-linearity of physics constraints. We present a physics-constrained diffusion model (PCDM) designed to solve inverse problems in scientific and engineering domains by efficiently integrating pre-trained diffusion models and physics-constrained objectives.\nWe leverage accelerated diffusion sampling to enable a practical generation process while strictly adhering to physics constraints by solving optimization problems at each timestep. By decoupling the likelihood optimization from the reverse diffusion steps, we ensure that the solutions remain physically consistent, even when employing fewer sampling steps.\nWe validate our method on a wide range of challenging physics-constrained inverse problems, including data assimilation, topology optimization, and full-waveform inversion. Experimental results show that our approach significantly outperforms existing methods in efficiency and precision, making it practical for real-world applications."},"_bibtex":{"value":"@misc{\nlee2025efficient,\ntitle={Efficient Physics-Constrained Diffusion Models for Solving Inverse Problems},\nauthor={Seungjun Lee and Shinjae Yoo},\nyear={2025},\nurl={https://openreview.net/forum?id=Da3j02cHe0}\n}"},"title":{"value":"Efficient Physics-Constrained Diffusion Models for Solving Inverse Problems"},"pdf":{"value":"/pdf/4ebfdbdf76dc91e6a3c9f8d62ac65169c876ccf7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"lee|efficient_physicsconstrained_diffusion_models_for_solving_inverse_problems"},"authorids":{"value":["~Seungjun_Lee1","~Shinjae_Yoo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Seungjun Lee","Shinjae Yoo"]}},"version":2},{"content":{"summary":{"value":"This paper introduces several ideas to boost the efficiency of marginal motion prediction: (1) represent all input entities as polylines without global pose attributes, (2) use transformer architectures but limit attention to K nearest neighbors, (3) directly use relative pose in transformer computations, (4) apply full self-attention only to map tokens, which can be cached during online inference, (5) obtain traffic light and agent features hierarchically, with cross-attention, and (6) use a final cross-attention block for all agent-anchor pairs to directly decode trajectories without any clustering or ensembling. All of these ideas are intuitive and an ablation study discusses some of their individual contributions. The final model obtains reasonable performance on WOMD and Argoverse 2 while scaling to dense traffic much more feasibly than one of the existing SoTA methods, Wayformer."},"soundness":{"value":"3 good"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. Is the training time of HTPR similar to Wayformer for the same number of training epochs?\n2. How is the KNARPE operation implemented in practice? Do you still compute and mask a dense attention matrix, or implement custom kernels to only compute attention where needed?\n3. How important is the post-processing described in L254-257? Are these techniques commonly applied by methods on these leaderboards?\n4. Could you please elaborate on L261-262, what does sampling 25% and 50% mean in this context?\n5. Would it be possible to compare the inference time (Fig. 4) to a GNN method with relative pose encodings?\n\nMinor:\n\n1. Have you tried adding the blue dots from Fig. 1b to Fig. 1a as well? This could make it clearer to understand which agents are being used in Fig. 1b.\n2. Could Fig. 5 be simplified, in particular by removing the striking colors for the map elements? An alternative option would be to add a legend describing all colors.\n\nUpdate:\n\nThank you for the detailed responses to all questions. The rebuttal addresses all of my concerns, and I would like to maintain my positive rating."},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"strengths":{"value":"The key contribution of this work lies in clearly highlighting of some problematic practices which are still commonly used in most research on motion forecasting in autonomous driving (heavy emphasis on the offline setting), and bringing efficiency for online inference to the forefront. The ideas presented to improve efficiency are not all new, but in combination interesting and well-motivated. Despite the large number of complex technical concepts covered in the draft, the presentation is clear and it is possible to follow and understand all components."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While all the proposed ideas are simple and intuitive, putting them all together yields a complex architecture with a large space of design choices and hyper-parameters. This model trains for 10 days, despite the efficient vectorized input and hierarchical architecture focused on efficiency. The runtime analysis only presents a comparison to an agent-centric baseline Wayformer, which fails to provide evidence for whether the proposed model is efficient among relative pose based forecasting methods (e.g., no evidence for the claim made in L092 that GNNs are more demanding than transformers in the online inference setting)."},"limitations":{"value":"Limitations are discussed in Section 5."}},"nonreaders":[],"tmdate":1702411150670,"tcdate":1688665819933,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission8148/Reviewer_Gjvi"],"signatures":["NeurIPS.cc/2023/Conference/Submission8148/Reviewer_Gjvi"],"forum":"YcmGuwdLoU","number":4,"license":"CC BY 4.0","cdate":1688665819933,"mdate":1702411150670,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission8148/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"YcmGuwdLoU","id":"jXT0GtWcJP","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Motion Prediction","Autonomous Driving","Transformer"]},"supplementary_material":{"value":"/attachment/a50195cd1394de25f2d169613ac2bdc3956afea4.pdf"},"_bibtex":{"value":"@inproceedings{\nzhang2023realtime,\ntitle={Real-Time Motion Prediction via Heterogeneous Polyline Transformer with Relative Pose Encoding},\nauthor={Zhejun Zhang and Alexander Liniger and Christos Sakaridis and Fisher Yu and Luc Van Gool},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=YcmGuwdLoU}\n}"},"title":{"value":"Real-Time Motion Prediction via Heterogeneous Polyline Transformer with Relative Pose Encoding"},"paperhash":{"value":"zhang|realtime_motion_prediction_via_heterogeneous_polyline_transformer_with_relative_pose_encoding"},"TLDR":{"value":"We use hierarchical Transformers with pairwise-relative representation to realize accurate and efficient motion prediction in autonomous driving."},"abstract":{"value":"The real-world deployment of an autonomous driving system requires its components to run on-board and in real-time, including the motion prediction module that predicts the future trajectories of surrounding traffic participants. Existing agent-centric methods have demonstrated outstanding performance on public benchmarks. However, they suffer from high computational overhead and poor scalability as the number of agents to be predicted increases. To address this problem, we introduce the K-nearest neighbor attention with relative pose encoding (KNARPE), a novel attention mechanism allowing the pairwise-relative representation to be used by Transformers. Then, based on KNARPE we present the Heterogeneous Polyline Transformer with Relative pose encoding (HPTR), a hierarchical framework enabling asynchronous token update during the online inference. By sharing contexts among agents and reusing the unchanged contexts, our approach is as efficient as scene-centric methods, while performing on par with state-of-the-art agent-centric methods. Experiments on Waymo and Argoverse-2 datasets show that HPTR achieves superior performance among end-to-end methods that do not apply expensive post-processing or model ensembling. The code is available at https://github.com/zhejz/HPTR."},"pdf":{"value":"/pdf/8ffb36ed1a34f61b9c70fc5dc0630245574a2ccd.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Zhejun_Zhang1","~Alexander_Liniger1","~Christos_Sakaridis1","~Fisher_Yu2","~Luc_Van_Gool1"]},"authors":{"value":["Zhejun Zhang","Alexander Liniger","Christos Sakaridis","Fisher Yu","Luc Van Gool"]}},"version":2},{"content":{"summary":{"value":"This paper proposes \"AdS-GNN,\" a new graph neural network that is equivariant to conformal transformations (translations, rotations, scale transformations, and special conformal transformations). To achieve this equivariance, the authors, inspired by insights from the AdS/CFT correspondence in physics, introduce a method to lift input data from flat Euclidean space $\\mathbb{R}^d$ to Anti-de Sitter (AdS) space $AdS_{d+1}$, which has one additional dimension. On AdS space, the conformal transformations of the original space manifest as isometric transformations (distance-preserving transformations). Therefore, the paper efficiently achieves conformal equivariance by constructing a message-passing GNN that utilizes the proper distance of AdS space as an invariant. The proposed method was evaluated on tasks from computer vision (e.g., SuperPixel MNIST) and statistical physics (e.g., 2D/3D Ising models). Particularly in the physics tasks, it was confirmed to show high generalization performance even under extrapolation (OOD) scenarios, such as scale transformations or changes in system size (number of points). Furthermore, the interpretability of the model was also demonstrated, as it was able to extract a physical universal quantity—the scaling dimension of the Ising model—from the trained network with high precision."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"**Regarding generality:** While I understand the method's primary contribution lies in physics tasks, what advantages do the authors believe this conformal equivariance approach offers over existing equivariant GNNs (e.g., those with rotational or translational equivariance) in domains outside of physics, such as general computer vision or robotics? I would also like to ask for the authors' insights into why the method failed to achieve SOTA performance on SuperPixel MNIST. Are there specific reasons to consider, such as the approximation error from the lifting procedure, the nature of the point cloud data, or a potential lack of expressive power in the architecture?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"**Novelty of the Idea**: The application of the profound idea of AdS/CFT correspondence from physics to the context of geometric deep learning, thereby constructing a GNN architecture with conformal equivariance, is highly original and commendable.\n\n**High Affinity with Physics Tasks**: Due to its design background, the proposed method is an excellent fit for tasks in statistical physics where conformal symmetry plays a dominant role (e.g., the Ising model near its critical point).\n\n**Excellent Generalization and Interpretability**: In experiments on physics tasks, AdS-GNN demonstrated performance superior to existing equivariant GNNs (like EGNN). Its robustness to extrapolation tasks, such as changing the system size (number of points), is particularly noteworthy. Furthermore, the fact that the model can automatically learn and extract a physically meaningful universal quantity—the scaling dimension—from data and recover its true value with high precision is a testament to the model's high interpretability and a significant contribution."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Lack of Generality and Performance**: The main contribution of this method appears to be limited to physics tasks. As shown by the experimental results (Table 1), in standard image (point cloud) classification tasks like SuperPixel MNIST, the performance does not reach that of existing SOTA methods (e.g., PONITA), and the method's superiority in general-purpose benchmarks has not been demonstrated.\n\n**Limited Applicability**: This method requires prior knowledge that the target data or task possesses \"conformal symmetry.\" Its application to many general machine learning tasks where such strong symmetry does not exist or is unknown is difficult, and the method's utility is inherently restricted.\n\n**Insufficient Appeal to the ICLR Community**: As a result of the above two points, the paper's contribution feels strongly directed primarily at the physics community. There is a lack of discussion or evidence regarding what new possibilities this conformal equivariance approach could bring to the broader ICLR audience (including computer vision, reinforcement learning, natural language processing, etc.)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924343677,"tcdate":1761976373741,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13818/Reviewer_HgBi"],"signatures":["ICLR.cc/2026/Conference/Submission13818/Reviewer_HgBi"],"forum":"EIyvsL5Cue","number":3,"license":"CC BY 4.0","cdate":1761976373741,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13818/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924343677,"domain":"ICLR.cc/2026/Conference","replyto":"EIyvsL5Cue","id":"xpeVZNPVXs","forumContent":{"TLDR":{"value":"We suggest a framework for building conformal group equivariant models that are consistent under angle preserving transformations which include translations, rotations, reflections, scaling and special conformal transformation."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["equivariance; conformal group; scale equivariance; ising model"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Conformal symmetries, i.e.\\ coordinate transformations that preserve angles, play a key role in many fields, including physics, mathematics, computer vision and (geometric) machine learning. Here we build a neural network that is equivariant under general conformal transformations. To achieve this, we lift data from flat Euclidean space to Anti de Sitter (AdS) space. This allows us to exploit a known correspondence between conformal transformations of flat space and isometric transformations on the Anti de Sitter space. We then build upon the fact that such isometric transformations have been extensively studied on general geometries in the geometric deep learning literature. In particular, we employ message-passing layers conditioned on the proper distance, yielding a computationally efficient framework. We validate our model on tasks from computer vision and statistical physics, demonstrating strong performance, improved generalization capacities, and the ability to extract conformal data such as scaling dimensions from the trained network."},"_bibtex":{"value":"@inproceedings{\nzhdanov2026adsgnn,\ntitle={AdS-{GNN} - a Conformally Equivariant Graph Neural Network},\nauthor={Maksim Zhdanov and Nabil Iqbal and Erik J Bekkers and Patrick Forr{\\'e}},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EIyvsL5Cue}\n}"},"title":{"value":"AdS-GNN - a Conformally Equivariant Graph Neural Network"},"pdf":{"value":"/pdf/b47c2fac11a8229e55aebee668d5b700e87aca2a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhdanov|adsgnn_a_conformally_equivariant_graph_neural_network"},"authorids":{"value":["~Maksim_Zhdanov1","~Nabil_Iqbal1","~Erik_J_Bekkers1","~Patrick_Forré1"]},"authors":{"value":["Maksim Zhdanov","Nabil Iqbal","Erik J Bekkers","Patrick Forré"]}},"version":2},{"content":{"summary":{"value":"The authors propose using a physics-informed neural surrogate based on a sinusoidal representation network that reproduces Laplacian perfusion physics to mitigate the high computational costs associated with finite element methods in studying cerebral perfusion and simulating blood flow patterns in stroke."},"correctness":{"value":"4: Excellent"},"soundness":{"value":"4: Excellent"},"strengths":{"value":"The abstract is well written and results are shown on a very challenging real life dataset.\nData-driven approaches to mitigate the computational costs associated with physics modeling are highly relevant for many applications, hence this contribution will be interesting to a broader audience beyond the application in the focus of this study."},"weaknesses":{"value":"The authors do not mention that code for the method is being released which would be very valuable in terms of reproducibility and future research in physics-informed ML."},"confidence":{"value":4},"rating":{"value":5}},"parentInvitations":"NLDL.org/2026/Abstracts_Track/-/Official_Review","nonreaders":[],"tmdate":1762345636019,"tcdate":1761637076538,"writers":["NLDL.org/2026/Abstracts_Track","NLDL.org/2026/Abstracts_Track/Submission33/Reviewer_LN8Q"],"signatures":["NLDL.org/2026/Abstracts_Track/Submission33/Reviewer_LN8Q"],"forum":"b16RymJRX3","number":1,"license":"CC BY 4.0","cdate":1761637076538,"readers":["everyone"],"invitations":["NLDL.org/2026/Abstracts_Track/Submission33/-/Official_Review","NLDL.org/2026/Abstracts_Track/-/Edit"],"mdate":1762345636019,"domain":"NLDL.org/2026/Abstracts_Track","replyto":"b16RymJRX3","id":"d8N8rG6TCP","forumContent":{},"version":2},{"content":{"summary":{"value":"The authors identified the limitations of non-Hodge aware learners on simplicial complex (SC) data and proposed a convolutional structure that 1) decomposes the upper and lower k-Laplacian and 2) takes the inter-simplicial couplings into account. The paper has presented a justification for the performance by analyzing the Dirichlet energy and oversmoothing. Additionally, they provided theoretical perturbations bound to study the robustness of the proposed convolution layer. The claims are supported by experiments on synthetic and real SC datasets."},"soundness":{"value":"2 fair"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. Related to Weakness #2, why the Hodge Laplacian smoothing [31] not Hodge-aware? I think it can also learn from the different subspaces of the Hodge Laplacian, just not independently. Maybe adding some definition/citation as per #2 will clarify it a bit. \n1. [Minor language usage suggestion] Consider rewrite L24-L25 to improve clarity; for instance, you might rewrite it as something like “A SC can be informally viewed as an extension of a graph. For example, one of the simplest SC (SC_2) can be constructed from a graph by inducing some triangles over the edge set.”.\n1. [Minor language usage suggestion] There is an extra e.g., in L27\n1. [Minor language usage suggestion] Consider breaking L27-29 into multiple sentences to improve clarity.\n1. [Minor notation issue] I would consider changing the notation of the $\\mathbf B$ matrix in L259 to reduce confusion."},"rating":{"value":"4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"strengths":{"value":"1. [Originality] Incorporate the well-known Hodge theorem into the learning task on simplicial complex. The Hodge theorem provides a good intuition and explanation for the learning of the simplicial signal on SC. \n1. [Quality] Empirical examples on the synthetic datasets on Dirichlet energy and stability bound to support the theoretical claims. \n1. [Clarity] Great overview of the simplicial complex and Hodge decomposition. The authors also provided a motivation/justification for why the proposed layer works with Dirichlet energy minimization. \n1. [Significance] Being able to learn the simplicial signal in different Hodge subspaces is an important task."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. If this framework needs to be applied to graph only having edges (i.e., SC of order 1), one usually can apply something like clique-complex (or any other methods to fill in the Simplicios) from that graph. In this case, $n_k$ is generally large (worst case $n_k = \\mathcal O(n^k)$), resulting in a huge $L_k$ matrix. How practical is it to use the proposed method under this scenario?\n1. Can you provide a definition/discussion or citation of Hodge-aware? It is not clear to me where it is defined throughout the manuscript. I can get some high-level ideas by reading Theorem 7, but I think it would be nice if you could explicitly call it out at the beginning (e.g., in introduction or background).\n1. The discussion for preventing “over-smoothing” in Section 3 is great, it provides some high-level motivations of the choices you made. However, I am not sure if that it can support the claims.  Specifically, to really prevent “over-smoothing” of the Dirichlet energy, shouldn’t we bound the $D(x_k^{\\ell+1})$ in other way around, i.e., with a lower bound rather than an upper bound? If we can show that $D(x_k^{\\ell+1})$ can be lower bounded, the claim can be more convincing. \n1. Consider adding some high-level intuition on what harmonic flow is using the edge space example (e.g., flow cycling around global topoplogical structures); this will give readers having no background in Hodge decomposition a better understanding of what a “harmonic flow” is.\n1. [Typo] L186 there is typo/grammatical issue, do you mean “$\\tilde{h}_k = \\text{diag}(...)$ is the frequency response of $\\mathbb H_k$”?\n"},"limitations":{"value":"The paper discusses some of its limitations and requirements/assumptions. No significant social impact is identified from this work. \n"}},"nonreaders":[],"tmdate":1702411094978,"tcdate":1688688596836,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission7111/Reviewer_XgBj"],"signatures":["NeurIPS.cc/2023/Conference/Submission7111/Reviewer_XgBj"],"forum":"QSJKrO1Qpy","number":2,"license":"CC BY 4.0","cdate":1688688596836,"mdate":1702411094978,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission7111/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"QSJKrO1Qpy","id":"cq16UDFlKh","forumContent":{"venue":{"value":"Submitted to NeurIPS 2023"},"keywords":{"value":["hodge decomposition","simplicial complexes","spectral simplicial theory","simplicial neural network","stability"]},"supplementary_material":{"value":"/attachment/2f34bd8cadecf97dba76f29874c45578aec0d467.zip"},"_bibtex":{"value":"@misc{\nyang2023hodgeaware,\ntitle={Hodge-Aware Learning on Simplicial Complexes},\nauthor={Maosheng Yang and Elvin Isufi},\nyear={2023},\nurl={https://openreview.net/forum?id=QSJKrO1Qpy}\n}"},"title":{"value":"Hodge-Aware Learning on Simplicial Complexes"},"paperhash":{"value":"yang|hodgeaware_learning_on_simplicial_complexes"},"abstract":{"value":"  Neural networks on simplicial complexes (SCs) can learn from data residing on simplices such as nodes, edges, triangles, etc. \n  However, existing works often overlook the Hodge theory that decomposes simplicial data into three orthogonal characteristic subspaces, such as the identifiable gradient, curl and harmonic components of edge flows.\n  In this paper, we aim to incorporate this data inductive bias into learning on SCs. \n  Particularly, we present a general convolutional architecture \n  which respects the three key principles of uncoupling the lower and upper simplicial adjacencies, accounting for the inter-simplicial couplings, and performing higher-order convolutions. \n  To understand these principles, we first use Dirichlet energy minimizations on SCs to interpret their effects on mitigating the simplicial oversmoothing. \n  Then, through the lens of spectral simplicial theory,\n  we show the three principles promote the Hodge-aware learning of this architecture, in the sense that the three Hodge subspaces are invariant under its learnable functions and the learning in two nontrivial subspaces are independent and expressive.\n  To further investigate the learning ability of this architecture, we also study it is stable against small perturbations on simplicial connections.\n  Finally, we experimentally validate the three principles by comparing with methods that either violate or do not respect them.\n  Overall, this paper bridges learning on SCs with the Hodge decomposition, highlighting its importance for rational and effective learning from simplicial data."},"pdf":{"value":"/pdf/ad4f57e2f1ac83c0662a24f3f4daf85a5dc7389a.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference/Rejected_Submission"},"authorids":{"value":["~Maosheng_Yang1","~Elvin_Isufi1"]},"authors":{"value":["Maosheng Yang","Elvin Isufi"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (2) 2023"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-53499-7_1.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2023"},"paperhash":{"value":"park|identifying_wellconnected_communities_in_realworld_and_synthetic_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Minhyuk_Park:","~Yasamin_Tabatabaee1","https://dblp.org/search/pid/api?q=author:Vikram_Ramavarapu:","https://dblp.org/search/pid/api?q=author:Baqiao_Liu:","https://dblp.org/search/pid/api?q=author:Vidya_Kamath_Pailodi:","https://dblp.org/search/pid/api?q=author:Rajiv_Ramachandran:","https://dblp.org/search/pid/api?q=author:Dmitriy_Korobskiy:","https://dblp.org/search/pid/api?q=author:Fábio_Ayres:","https://dblp.org/search/pid/api?q=author:George_Chacko:","https://dblp.org/search/pid/api?q=author:Tandy_J._Warnow:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-53499-7_1"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/ParkTRLPRKACW23,\n  author={Minhyuk Park and Yasamin Tabatabaee and Vikram Ramavarapu and Baqiao Liu and Vidya Kamath Pailodi and Rajiv Ramachandran and Dmitriy Korobskiy and Fábio Ayres and George Chacko and Tandy J. Warnow},\n  title={Identifying Well-Connected Communities in Real-World and Synthetic Networks},\n  year={2023},\n  cdate={1672531200000},\n  pages={3-14},\n  url={https://doi.org/10.1007/978-3-031-53499-7_1},\n  booktitle={COMPLEX NETWORKS (2)},\n  crossref={conf/complexnetworks/2023-2}\n}\n"},"abstract":{"value":"Integral to the problem of detecting communities through graph clustering is the expectation that they are “well-connected”. Surprisingly, we find that the output of multiple clustering approaches–the Leiden algorithm with either the Constant Potts Model or modularity as quality function, Iterative K-Core Clustering, Infomap, and Markov Clustering–include communities that fail even a mild requirement for well-connectedness. As a remediation strategy, we have developed the “Connectivity Modifier” (CM), which iteratively removes small edge cuts and re-clusters until all communities detected are well-connected. Results from real-world networks with up to 75,025,194 nodes illustrate how CM enables additional insights into community structure within networks, while results on synthetic networks show that the CM algorithm improves accuracy in recovering true communities. Our study also raises questions about the “clusterability” of networks and mathematical models of community structure."},"title":{"value":"Identifying Well-Connected Communities in Real-World and Synthetic Networks"},"authors":{"value":["Minhyuk Park","Yasamin Tabatabaee","Vikram Ramavarapu","Baqiao Liu","Vidya Kamath Pailodi","Rajiv Ramachandran","Dmitriy Korobskiy","Fábio Ayres","George Chacko","Tandy J. Warnow"]}},"tmdate":1730169140917,"pdate":1672531200000,"tcdate":1730169137367,"writers":["~"],"signatures":["~Yasamin_Tabatabaee1"],"forum":"K28ry0VMyW","license":"CC BY-SA 4.0","number":162434,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1730169140917,"domain":"DBLP.org","id":"K28ry0VMyW","version":2},{"content":{"summary":{"value":"The paper presents a novel paradigm for constructing world models that serve as explicit representations of real-world environments and their dynamics. By integrating advances in real-time photorealism, such as Gaussian Splatting, with physics simulators, the authors propose a system capable of generating new data for imitation learning. Additionally, the paper demonstrates the application of this model in real-world scenarios, showing how data collected from the world model can be used to train robots via imitation learning, with promising results when transferring learned behaviors to real-world tasks.\n\n**Strengths:**\n\n1. The paper introduces an innovative approach by leveraging world models to generate robotic data for imitation learning, which is a contribution to the field.\n2. The experiments are detailed, covering both simulation environments and real-world robot demonstrations, providing a robust evaluation of the approach.\n3. A creative method for augmenting data used in imitation learning is introduced, which could lead to improved learning efficiency.\n\n**Weaknesses:**\n\n1. The absence of publicly available source code limits the reproducibility of the results. It is suggested to release the code during the rebuttal stage.\n2. Some figures in the paper need improvement, as the text in several instances is too small to read clearly.\n3. The predictions demonstrated in the paper are limited to simple tasks and physics environments, and future work should focus on extending these predictions to more challenging tasks and complex physical simulations.\n\nIn conclusion, the paper presents a compelling framework that blends world modeling with imitation learning, but there are areas for improvement, particularly in terms of figure clarity, task complexity, and providing source code for reproducibility."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Could you please show some performance results in more complex physical environments and challenging tasks? Even if they were unsuccessful, it would be helpful to see such results, even though they are not included in the paper."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The paper introduces an innovative approach by leveraging world models to generate robotic data for imitation learning, which is a contribution to the field.\n2. The experiments are detailed, covering both simulation environments and real-world robot demonstrations, providing a robust evaluation of the approach.\n3. A creative method for augmenting data used in imitation learning is introduced, which could lead to improved learning efficiency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The absence of publicly available source code limits the reproducibility of the results. It is suggested to release the code during the rebuttal stage.\n2. Some figures in the paper need improvement, as the text in several instances is too small to read clearly.\n3. The predictions demonstrated in the paper are limited to simple tasks and physics environments, and future work should focus on extending these predictions to more challenging tasks and complex physical simulations."}},"nonreaders":[],"tmdate":1731429309021,"tcdate":1729451851072,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9572/Reviewer_Mne8"],"signatures":["ICLR.cc/2025/Conference/Submission9572/Reviewer_Mne8"],"forum":"3RSLW9YSgk","number":1,"license":"CC BY 4.0","cdate":1729451851072,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9572/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429309021,"domain":"ICLR.cc/2025/Conference","replyto":"3RSLW9YSgk","id":"oKXoMqvhkR","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["World model;  Imagination; Imitation Learning; Gaussian Splatting; Compositional; Physics-informed; Object-centric;"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"A world model provides an agent with a representation of its environment, enabling it to predict the causal consequences of its actions. Current world models typically cannot directly and explicitly imitate the actual environment in front of a robot, often resulting in unrealistic behaviors and hallucinations that make them unsuitable for real-world robotics applications. \nTo overcome those challenges, we propose to rethink robot world models as learnable digital twins. We introduce DreMa, a new approach for constructing digital twins automatically using learned explicit representations of the real world and its dynamics, bridging the gap between traditional digital twins and world models.\nDreMa replicates the observed world and its structure by integrating Gaussian Splatting and physics simulators, allowing robots to imagine novel configurations of objects and to predict the future consequences of robot actions thanks to its compositionality.\nWe leverage this capability to generate new data for imitation learning by applying equivariant transformations to a small set of demonstrations. Our evaluations across various settings demonstrate significant improvements in accuracy and robustness by incrementing actions and object distributions, reducing the data needed to learn a policy and improving the generalization of the agents. \nAs a highlight, we show that a real Franka Emika Panda robot, powered by DreMa’s imagination, can successfully learn novel physical tasks from just a single example per task variation (one-shot policy learning).\nOur project page can be found in: https://dreamtomanipulate.github.io/."},"_bibtex":{"value":"@inproceedings{\nbarcellona2025dream,\ntitle={Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination},\nauthor={Leonardo Barcellona and Andrii Zadaianchuk and Davide Allegro and Samuele Papa and Stefano Ghidoni and Efstratios Gavves},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=3RSLW9YSgk}\n}"},"title":{"value":"Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination"},"pdf":{"value":"/pdf/09c7c86828171ae29d694536b94c26a0c887d7f3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"barcellona|dream_to_manipulate_compositional_world_models_empowering_robot_imitation_learning_with_imagination"},"authorids":{"value":["~Leonardo_Barcellona1","~Andrii_Zadaianchuk1","~Davide_Allegro1","~Samuele_Papa1","~Stefano_Ghidoni1","~Efstratios_Gavves1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Leonardo Barcellona","Andrii Zadaianchuk","Davide Allegro","Samuele Papa","Stefano Ghidoni","Efstratios Gavves"]}},"version":2},{"content":{"summary":{"value":"The authors propose a novel pipeline for creating large amounts of instruction data with varying levels of complexity using LLM. The proposed approach starts from a set of seed instructions and first uses Evol-Instruct, a suit of prompt templates that can make LLMs such as ChatGPT to rewrite seed instructions into more complex ones. Then all generated data are used for fine-tuning open-source LLMs such as LLaMA. The authors used the proposed method to collect a set of instruction data and fine-tuned LLaMA into \"WizardLM\", and conduct a suit of evaluation on various benchmarks. Experimental results show some improvement over representative baselines including Alpaca and Vicuna models."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The idea of using LLMs to rewrite and synthesize more complex instructions is interesting and intuitive.\n2. The paper is overall well-written and easy to follow. \n3. The authors evaluate WizardLM on a wide range of datasets/benchmarks and the experimental results look promising."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The technical contribution of the proposed method is not very significant because compared to self-instruct, it is only adding the command for LLM to generate more complex instruction. The success of the proposed method largely depend on the abilities of powerful LLMs such as ChatGPT. While interesting and intuitive, I'm not sure the technical contribution of the manuscript is suitable for conferences such as ICLR.\n\n2. The authors compare WizardLM with Vicuna by using the same amount of generated instructions. However, the cost or token consumption used for collecting the datasets (especially compared with Alpaca) should also be controlled or at least mentioned. \n\n3. The manuscript lacks comparisons with other methods for instruction data generation methods such as Baize, CAMEL, etc. For benchmark results, the role of the seed instructions is very important. I assume the seed data is very different, which could be a major cause of the performance difference. Therefore, more detailed ablation study is required to make the results more convincing.\n\n4. Also, it would be very helpful to test the proposed data generation methods using other LLMs (e.g., LLaMA) instead of ChatGPT to better understand the proposed data synthesis pipeline.\n\n4."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Please see the above weakness section for questions."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1701747070176,"tcdate":1699190668458,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6631/Reviewer_JU8z"],"signatures":["ICLR.cc/2024/Conference/Submission6631/Reviewer_JU8z"],"forum":"CfXh93NDgH","number":4,"license":"CC BY 4.0","cdate":1699190668458,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6631/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1701747070176,"domain":"ICLR.cc/2024/Conference","replyto":"CfXh93NDgH","id":"k3GRtuH3mH","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large Language Model","Instruction Fine-tuning"]},"supplementary_material":{"value":"/attachment/8b061fd1e3b9512d743627cf2a25daab1314afc5.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Training large language models (LLMs) with open-domain instruction following data brings colossal success. However, manually creating such instruction data is very time-consuming and labor-intensive. Moreover, humans may struggle to produce high-complexity instructions. In this paper, we show an avenue for creating large amounts of instruction data with varying levels of complexity using LLM instead of humans. Starting with an initial set of instructions, we use our proposed Evol-Instruct to rewrite them step by step into more complex instructions. Then, we mix all generated instruction data to fine-tune LLaMA. We call the resulting model WizardLM. Both automatic and human evaluations consistently indicate that WizardLM outperforms baselines such as Alpaca (trained from Self-Instruct) and Vicuna (trained from human-created instructions). The experimental results demonstrate that the quality of instruction-following dataset crafted by Evol-Instruct can significantly improve the performance of LLMs."},"_bibtex":{"value":"@inproceedings{\nxu2024wizardlm,\ntitle={Wizard{LM}: Empowering Large Pre-Trained Language Models to Follow Complex Instructions},\nauthor={Can Xu and Qingfeng Sun and Kai Zheng and Xiubo Geng and Pu Zhao and Jiazhan Feng and Chongyang Tao and Qingwei Lin and Daxin Jiang},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=CfXh93NDgH}\n}"},"title":{"value":"WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions"},"pdf":{"value":"/pdf/4760b6872282dc467e4ea0d825a5a186c2b18b66.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"xu|wizardlm_empowering_large_pretrained_language_models_to_follow_complex_instructions"},"authorids":{"value":["~Can_Xu2","~Qingfeng_Sun1","~Kai_Zheng8","~Xiubo_Geng2","~Pu_Zhao3","~Jiazhan_Feng1","~Chongyang_Tao1","~Qingwei_Lin1","~Daxin_Jiang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Can Xu","Qingfeng Sun","Kai Zheng","Xiubo Geng","Pu Zhao","Jiazhan Feng","Chongyang Tao","Qingwei Lin","Daxin Jiang"]}},"version":2},{"content":{"summary":{"value":"The paper presents a conformally equivariant graph neural network through showing how to lift data from Euclidean space to Anti de Sitter (AdS) space. Distance metrics in AdS space can then be used in conjunction with invariant message passing to create neural networks that are equivariant to the conformal group. Experimental results on computer vision tasks (SuperPixel MNIST and shape segmentation) and physics tasks (2d/3d Ising model and N-body simulation) are presented. In particular, for the physics tasks, AdS-GNN seems to require less data and generalizes better. Additionally, it recovers the correct conformal dimension, illustrating interpretability."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"What is the computational cost of these models compared to other invariant models (e.g. EGNN)? Does the lifting procedure incur significant computational costs?\n\nAnother scientific domain of interest could be fluid dynamics. At sufficiently high Reynolds number, the statistics of turbulent motions in the so-called “inertial range” become universal and exhibit scale invariance properties [2]. I would be interested to see the performance of these models on turbulence modeling tasks.\n\nIn what settings would conformal equivariance be useful for images? It seems like it would be more useful in dynamical systems/more-physics motivated problems as shown in the physics task experiments. Is there further motivation for why one would want robustness to conformal transformations in computer vision tasks?\n\nWould it be possible to have non-scalar features in message passing? Were there any experiments done extending the model to include non-scalar features?\n\nHow is the regulator $z_0$ chosen, and how does this impact the resulting lifting procedure? Are there other ways that one could lift data into AdS/why was this way in particular chosen? If the lifting procedure breaks certain conformal transformations, which transformations does it break? Figure 3 shows one broken transformation, but what do these “special” transformations correspond to physically, and why would this not be a problem that these are broken? Is it accurate to call the model conformally equivariant if this is the case (maybe approximately conformally equivariant)? Correct me if I'm wrong, but currently it seems conformally invariant rather than equivariant?\n\nWere any comparisons done to pre-existing scale invariant models? This would be interesting to include with comparison to the rotationally/translationally invariant models.\n\n[1] Batatia et al. MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields (2022).\n\n[2] Pope, Stephen. Turbulent Flows. (Cambridge University Press, 2000)."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper presents an interesting new framework for conformal group equivariance, using ideas from theoretical physics (and building open pre-existing work for constructing group equivariant convolutional layers). It is well-organized with experiments in multiple sub-domains, and provides a self contained introduction to the conformal group. I particularly like the figure at the top of pg. 3 showing possible transformations under the conformal group. Experimental results show that AdS-GNN maintains scale invariance (e.g. SuperPixel MNIST). The physics tasks are good illustrations of the utility of AdS-GNN, showing that AdS-GNN outperforms other models (EGNN and MPNN) for predicting N-point correlation functions for the 2D Ising model and requires less training data. I found it particularly interesting as well that the learned values of the conformal dimensions match the ground truth, and it would be interesting to explore this point further in more complex physics datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I am not sure of the usefulness of the image experiments. I wouldn’t expect image datasets to require conformal equivariance/to me the orientation of the image would probably matter the most. However, the AdS-GNN model is invariant, so this orientation is not taken into account. It does not seem to be persuasive that one would use AdS-GNN in a computer vision setting.\n\nI found the part about embedding points in AdS somewhat confusing, and the choice of regular $z_0$ is unclear. I am concerned that this would impact the dataset/break certain symmetries (see questions).  I think further clarification is needed in this section, perhaps an additional figure showing another embedded shape (or embeddings of circles, triangles as in the shape analysis data) would be helpful.\n\nAdS is unable to handle orientation and relies on invariant descriptors. This seems to be a significant limitation, for both image datasets and physics datasets, in light of other work on equivariant neural networks with message passing of higher-order tensorial features (e.g. [1]). \n\nThe shape segmentation task is quite toy, the authors could consider exploring a more realistic dataset such as ModelNet.\n\nOverall, most of the experiments seem to test scale invariance, but there are pre-existing scale invariant models (from my understanding). It may be good to benchmark against these pre-existing scale invariant models and see what conformal invariance actually gives us."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359965532,"tcdate":1761921941479,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13818/Reviewer_jAnW"],"signatures":["ICLR.cc/2026/Conference/Submission13818/Reviewer_jAnW"],"forum":"EIyvsL5Cue","number":2,"license":"CC BY 4.0","cdate":1761921941479,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13818/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359965532,"domain":"ICLR.cc/2026/Conference","replyto":"EIyvsL5Cue","id":"JHYs9GVdJ7","forumContent":{"TLDR":{"value":"We suggest a framework for building conformal group equivariant models that are consistent under angle preserving transformations which include translations, rotations, reflections, scaling and special conformal transformation."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["equivariance; conformal group; scale equivariance; ising model"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Conformal symmetries, i.e.\\ coordinate transformations that preserve angles, play a key role in many fields, including physics, mathematics, computer vision and (geometric) machine learning. Here we build a neural network that is equivariant under general conformal transformations. To achieve this, we lift data from flat Euclidean space to Anti de Sitter (AdS) space. This allows us to exploit a known correspondence between conformal transformations of flat space and isometric transformations on the Anti de Sitter space. We then build upon the fact that such isometric transformations have been extensively studied on general geometries in the geometric deep learning literature. In particular, we employ message-passing layers conditioned on the proper distance, yielding a computationally efficient framework. We validate our model on tasks from computer vision and statistical physics, demonstrating strong performance, improved generalization capacities, and the ability to extract conformal data such as scaling dimensions from the trained network."},"_bibtex":{"value":"@inproceedings{\nzhdanov2026adsgnn,\ntitle={AdS-{GNN} - a Conformally Equivariant Graph Neural Network},\nauthor={Maksim Zhdanov and Nabil Iqbal and Erik J Bekkers and Patrick Forr{\\'e}},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EIyvsL5Cue}\n}"},"title":{"value":"AdS-GNN - a Conformally Equivariant Graph Neural Network"},"pdf":{"value":"/pdf/b47c2fac11a8229e55aebee668d5b700e87aca2a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhdanov|adsgnn_a_conformally_equivariant_graph_neural_network"},"authorids":{"value":["~Maksim_Zhdanov1","~Nabil_Iqbal1","~Erik_J_Bekkers1","~Patrick_Forré1"]},"authors":{"value":["Maksim Zhdanov","Nabil Iqbal","Erik J Bekkers","Patrick Forré"]}},"version":2},{"content":{"summary":{"value":"This paper studies in-context learning (ICL) for tabular data, building upon TabPFN. The authors showed that fine-tuning leads to better downstream performance, and better pre-training with more complex data (forest-based) helps the fine-tuning performance. The authors proposed a novel way to create less realistic but more complex pre-training data. The resulting pre-trained model (including both original and new pre-train data), TabForestPFN, when fine-tuned, achieves the best performance across a wide range of tabular datasets and tasks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"I do not have further questions besides the weaknesses listed above. It would be great if the authors could address the weaknesses during the rebuttal. I'm open to adjusting my ratings based on the authors' responses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"S1. The paper identifies that fine-tuning is still needed to achieve the best performance on each downstream task.\n\nS2. It proposes a novel, seemingly simple-to-implement way to create more complex pre-training data. The idea to incorporate decision trees in data generation is neat, as it encourages the neural net model to mimic Tree-based models' predictions.\n\nS3. Expensive experiments are conducted."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1. The motivations of this paper are not very clear and convincing. First, it is intuitive that fine-tuning could create more complex boundaries and improve the downstream performance --- the fine-tuning process needs to fit the downstream data. Second, it is intuitive that when the downstream data are sufficient (large context), then fine-tuning would improve more. (Otherwise, it may overfit.) Third, I'm not entirely convinced that pre-training from scratch cannot generate complex boundaries. Basically, if one over-fits the model to the training data, it should create a complex enough boundary. Fourth, I didn't fully capture \"We find that fine-tuned ICL-transformers, in contrast, are able to create these complex decision boundaries similar to tree-based methods.\" Again, it is intuitive that fine-tuning would create more complex boundaries; do ICL-transformers play an important role in this statement? \n\nW2. It is a bit hard to digest what each step in Algorithm 1 does. Can the authors provide a visualization, for example, a series of sub-figures where each of them corresponds to the outcome of each step in Algorithm 1?\n\nW3. Experimental results on WhyTrees show good signs of TabForestPFN (or just TabForest) but Table 3 (on TabZilla) seems to suggest no benefit of TabForestPFN over TabPFN if both are fine-tuned. Can the authors provide more details about these? \n\nW4. I'm not sure if I completely buy the argument that TabPFN achieves better ZSL performance because its training data are more realistic. If the downstream tasks (realistic data) do require more complex boundaries, isn't it possible to create realistic but complex data?\n\nW5. The analyses lack sufficient insights. For example, the way section 5.6 is written seems to be too intuitive: fine-tuning excels when there are sufficient samples; a larger context (i.e., more training data for the downstream tasks) improves the performance. I am curious why at larger support sizes (Figure 7 (b)), fine-tuning shows more benefits than zero-shot. It would be great if the authors could provide more insights rather than superficially summarizing the observations."}},"nonreaders":[],"tmdate":1731427955943,"tcdate":1730826881665,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3840/Reviewer_dstQ"],"signatures":["ICLR.cc/2025/Conference/Submission3840/Reviewer_dstQ"],"forum":"pE0UM18TQh","number":3,"license":"CC BY 4.0","cdate":1730826881665,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3840/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427955943,"domain":"ICLR.cc/2025/Conference","replyto":"pE0UM18TQh","id":"bH5kRgXALg","forumContent":{"TLDR":{"value":"We introduce TabForestPFN, a fine-tuned tabular in-context learning transformer with improvements based on increasing the complexity of pretraining datasets."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["tabular classification","tabular in-context learning transformer","fine-tuning"]},"supplementary_material":{"value":"/attachment/5e319ef3d4cdabb45414290901a3e72f0a763b80.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The recently introduced TabPFN pretrains an In-Context Learning (ICL) transformer on synthetic data to perform tabular data classification. In this work, we extend TabPFN to the fine-tuning setting, resulting in a significant performance boost. We also discover that fine-tuning enables ICL-transformers to create complex decision boundaries, a property regular neural networks do not have. Based on this observation, we propose to pretrain ICL-transformers on a new forest dataset generator which creates datasets that are unrealistic, but have complex decision boundaries. TabForest, the ICL-transformer pretrained on this dataset generator, shows better fine-tuning performance when pretrained on more complex datasets. Additionally, TabForest outperforms TabPFN on some real-world datasets when fine-tuning, despite having lower zero-shot performance due to the unrealistic nature of the pretraining datasets. By combining both dataset generators, we create TabForestPFN, an ICL-transformer that achieves excellent fine-tuning performance and good zero-shot performance."},"_bibtex":{"value":"@misc{\nbreejen2024finetuned,\ntitle={Fine-tuned In-Context Learning Transformers are Excellent Tabular Data Classifiers.},\nauthor={Felix den Breejen and Sangmin Bae and Stephen Cha and Se-Young Yun},\nyear={2024},\nurl={https://openreview.net/forum?id=pE0UM18TQh}\n}"},"title":{"value":"Fine-tuned In-Context Learning Transformers are Excellent Tabular Data Classifiers."},"pdf":{"value":"/pdf/128e9c89a9461a2bce674deefb9db1a2bc2fdc6b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"breejen|finetuned_incontext_learning_transformers_are_excellent_tabular_data_classifiers"},"authorids":{"value":["~Felix_den_Breejen1","~Sangmin_Bae1","~Stephen_Cha1","~Se-Young_Yun1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Felix den Breejen","Sangmin Bae","Stephen Cha","Se-Young Yun"]}},"version":2},{"content":{"summary":{"value":"This paper presents a dark matter detection dataset and benchmark from a dark matter direct detection experiment. The core problem is about time series denoising. The baseline approaches they present use very basic ML algorithms. The physics experiment is interesting but this paper is clearly a physics paper, written with lots of technical language about the physics experiment, and without any strong argument for where this fits into the landscape of machine learning, or what machine learning people would use this for, and thus is inappropriate for ICLR."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Questions:\n- What specific ML research questions would this dataset help enable?\n\nSuggestions:\n- I suggest you consider working with someone from the machine learning community to understand how to make such work more impactful for a broader audience."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- Interesting physics problem that I have not seen explored at all in ML venues.\n- The evaluation criteria is interesting and I also don't think this specific metric has been explored much.\n- Potentially an interesting problem for a benchmark."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- This paper is fundamentally a dark matter instrumentation paper that was submitted to an ML venue. It is not appropriate for consideration in this conference. The authors should work with someone in the field of machine learning, or even someone in the physics and machine learning community, so they learn how to frame contributions so that they would be useful for the broader ICLR audience. Even as a physicist myself, I found the paper difficult to understand or to see how to repurpose the dataset for my own work at the intersection.\n- It is written heavily with physics terminology with little effort to make it accessible to the audience of the conference.\n- Not clear how this work could help advance ML methods.\n- Uses basic off-the-shelf machine learning architectures as baselines, without any technical innovations.\n- Basic analysis: \"we observed that the FC Net model achieved the best performance with a denoising score of 6.43\" - without any further general insights presented. No further analysis of why this was true, or what specifically the ML community should pursue in addressing the problem posed by this paper."}},"nonreaders":[],"tmdate":1731428066554,"tcdate":1731405072381,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10765/Reviewer_Wx5G"],"signatures":["ICLR.cc/2025/Conference/Submission10765/Reviewer_Wx5G"],"forum":"p2QAOORDoG","number":4,"license":"CC BY 4.0","cdate":1731405072381,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10765/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428066554,"domain":"ICLR.cc/2025/Conference","replyto":"p2QAOORDoG","id":"Pg4V4Pvn8Z","forumContent":{"TLDR":{"value":"TIDMAD is the first dataset and benchmark from a dark matter physics experiment, providing ultra-long time series data and comprehensive tools that enable machine learning models to directly advance the fundamental physics search for dark matter."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["benchmark","dataset","denoising","public dataset","dark matter","physics","time series","ultra-long time series"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Dark matter makes up approximately 85\\% of total matter in our universe, yet it has never been directly observed in any laboratory on Earth. The origin of dark matter is one of the most important questions in contemporary physics, and a convincing detection of dark matter would be a Nobel-Prize-level breakthrough in fundamental science. The ABRACADABRA experiment was specifically designed to search for dark matter. Although it has not yet made a discovery, ABRACADABRA has produced several dark matter search results widely endorsed by the physics community. The experiment generates ultra-long time-series data at a rate of 10 million samples per second, where the dark matter signal would manifest itself as a sinusoidal oscillation mode within the ultra-long time series. In this paper, we present the TIDMAD --- a comprehensive data release from the ABRACADABRA experiment including three key components: an ultra-long time series dataset divided into training, validation, and science subsets; a carefully-designed denoising score for direct model benchmarking; and a complete analysis framework which produces a community-standard dark matter search result suitable for publication as a physics paper. This data release enables core AI algorithms to extract the signal and produce real physics results thereby advancing fundamental science. The data downloading and associated analysis scripts are available at https://anonymous.4open.science/r/TIDMAD."},"_bibtex":{"value":"@misc{\nfry2024tidmad,\ntitle={{TIDMAD}: Time Series Dataset for Discovering Dark Matter with {AI} Denoising},\nauthor={J. T. Fry and Xinyi Hope Fu and Kaliro{\\\"e} Mabelle West Pappas and Zhenghao Fu and Lindley Winslow and Aobo Li},\nyear={2024},\nurl={https://openreview.net/forum?id=p2QAOORDoG}\n}"},"title":{"value":"TIDMAD: Time Series Dataset for Discovering Dark Matter with AI Denoising"},"pdf":{"value":"/pdf/e3ed31c8edd1624646499922afc7af3176d3f49a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"fry|tidmad_time_series_dataset_for_discovering_dark_matter_with_ai_denoising"},"authorids":{"value":["~J._T._Fry1","~Xinyi_Hope_Fu1","~Kaliroë_Mabelle_West_Pappas1","~Zhenghao_Fu1","~Lindley_Winslow1","~Aobo_Li2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["J. T. Fry","Xinyi Hope Fu","Kaliroë Mabelle West Pappas","Zhenghao Fu","Lindley Winslow","Aobo Li"]}},"version":2},{"content":{"summary":{"value":"This paper proposes the Discrete Volterra Network (DiVo), which integrates Volterra series with deep learning for time series forecasting. The method reformulates continuous Volterra integrals into discrete, learnable coefficient matrices through Kronecker-powered polynomial expansions. To address practical challenges, DiVo introduces multi-channel mechanisms for time-varying systems and redundancy-aware sparsification combining fixed masking with low-rank decomposition. Experiments on synthetic chaotic systems and real-world datasets show improvements over baseline methods."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. Given exponential parameter growth, how does DiVo scale beyond toy problems? What are the memory requirements for realistic input dimensions and sequence lengths?\n\n2. Why do well-established methods perform so poorly? Were hyperparameters properly tuned? Can these results be reproduced independently?\n\n3.  How severely does limiting to order k=2-3 constrain the model's ability to capture real nonlinear dynamics compared to neural networks?\n\n4. How does performance degrade on noisy, high-dimensional, or non-stationary real-world data that doesn't follow idealized dynamical systems?\n\n5.  What guidelines exist for selecting polynomial order and channel count without extensive cross-validation, which may be computationally prohibitive?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. The discrete Volterra reparameterization provides a principled mathematical foundation by connecting classical nonlinear system theory with modern neural architectures.\n\n2. DiVo demonstrates impressive results on synthetic chaotic datasets, suggesting it can capture complex nonlinear dynamics effectively.\n\n3.  The learned coefficient matrices offer direct insights into feature interactions and temporal dependencies, with validation showing recovery of known system dynamics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The exponential growth of Kronecker products with polynomial order creates prohibitive computational costs. Even with sparsification, the method is limited to very low orders (k≤3), severely constraining expressiveness for complex real-world systems.\n\n2. The dramatic underperformance of established deep learning methods (LSTM achieving 0.834 MSE vs DiVo's 0.305 on ETT) suggests serious implementation issues or unfair comparisons. No details are provided about hyperparameter tuning or multiple runs for baselines.\n\n3. Most evaluated datasets are small-scale (thousands of samples) and low-dimensional. The method's performance on modern large-scale time series applications remains undemonstrated and questionable given computational constraints.\n\n4. The paper lacks convergence analysis for the optimization procedure, approximation error bounds for truncated series, and theoretical justification for the multi-channel extension. The \"rethinking nonlinear dynamics\" claim is overstated for what is essentially polynomial feature engineering.\n\n5. The core contribution reduces to discretizing Volterra series and learning polynomial coefficients - concepts well-established in system identification. The engineering improvements (masking, low-rank decomposition) are incremental.\n\n6. Missing statistical significance tests, limited dataset diversity, unclear baseline implementations, and cherry-picked results undermine the empirical claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925761480,"tcdate":1761820540893,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15469/Reviewer_Krqz"],"signatures":["ICLR.cc/2026/Conference/Submission15469/Reviewer_Krqz"],"forum":"i4Mph3pvwr","number":1,"license":"CC BY 4.0","cdate":1761820540893,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15469/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925761480,"domain":"ICLR.cc/2026/Conference","replyto":"i4Mph3pvwr","id":"POl1VED3b1","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"DiVo is a compact, interpretable deep learning model that integrates Volterra series to capture nonlinear dynamics and memory in time series, outperforming traditional models in accuracy and efficiency."},"keywords":{"value":["Nonlinear Dynamics system","Discrete Volterra series","Time series"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Deep learning models have achieved remarkable success in modeling complex time series data, yet their black-box nature limits interpretability and explicit representation of intrinsic dynamic structures such as nonlinear interactions and memory effects. Observing the inherent compatibility of Volterra series' polynomial integral kernels with GPU-accelerated deep learning frameworks, we propose the Discrete Volterra Network (DiVo), a novel deep learning model family integrating Volterra series to explicitly learn dynamic characteristics from time series data. Specifically, DiVo computes discrete Volterra coefficient matrices via polynomial expansions, converting nonlinear time series modeling into linear polynomial coefficient learning. To address practical challenges, we introduce adaptive channel selection to remove strict dependence on time-invariant sequences, and propose a redundancy-aware sparsification strategy that combines fixed masking of Volterra features with sparsified low-rank decomposition to eliminate redundancy in both the feature and parameter spaces, yielding a compact model representation.Extensive experiments on diverse real-world datasets show DiVo significantly outperforms traditional deep models in prediction accuracy, interpretability, and parameter efficiency."},"_bibtex":{"value":"@misc{\nhuang2026rethinking,\ntitle={Rethinking Nonlinear Dynamics in Deep Time Series Models},\nauthor={Yixue Huang and Hao Li and Jiaming Fan and Jiadi Li and Ding Wanli and Ying Jiang and Meikang Qiu and Hao Jiang},\nyear={2026},\nurl={https://openreview.net/forum?id=i4Mph3pvwr}\n}"},"title":{"value":"Rethinking Nonlinear Dynamics in Deep Time Series Models"},"pdf":{"value":"/pdf/2b46ba9690b119402ddd0ba0d18b7ec3b6d2d8a9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"huang|rethinking_nonlinear_dynamics_in_deep_time_series_models"},"authorids":{"value":["~Yixue_Huang2","~Hao_Li55","~Jiaming_Fan2","~Jiadi_Li1","~Ding_Wanli1","~Ying_Jiang1","~Meikang_Qiu2","~Hao_Jiang18"]},"authors":{"value":["Yixue Huang","Hao Li","Jiaming Fan","Jiadi Li","Ding Wanli","Ying Jiang","Meikang Qiu","Hao Jiang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2505.21879v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"liu|symbolic_foundation_regressor_on_complex_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Weiting_Liu:","~Jiaxu_Cui1","https://dblp.org/search/pid/api?q=author:Jiao_Hu:","~En_Wang1","https://dblp.org/search/pid/api?q=author:Bo_Yang_0002:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2505.21879"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2505-21879,\n  publtype={informal},\n  author={Weiting Liu and Jiaxu Cui and Jiao Hu and En Wang and Bo Yang},\n  title={Symbolic Foundation Regressor on Complex Networks},\n  year={2025},\n  month={May},\n  cdate={1746057600000},\n  journal={CoRR},\n  volume={abs/2505.21879},\n  url={https://doi.org/10.48550/arXiv.2505.21879}\n}\n"},"abstract":{"value":"In science, we are interested not only in forecasting but also in understanding how predictions are made, specifically what the interpretable underlying model looks like. Data-driven machine learning technology can significantly streamline the complex and time-consuming traditional manual process of discovering scientific laws, helping us gain insights into fundamental issues in modern science. In this work, we introduce a pre-trained symbolic foundation regressor that can effectively compress complex data with numerous interacting variables while producing interpretable physical representations. Our model has been rigorously tested on non-network symbolic regression, symbolic regression on complex networks, and the inference of network dynamics across various domains, including physics, biochemistry, ecology, and epidemiology. The results indicate a remarkable improvement in equation inference efficiency, being three times more effective than baseline approaches while maintaining accurate predictions. Furthermore, we apply our model to uncover more intuitive laws of interaction transmission from global epidemic outbreak data, achieving optimal data fitting. This model extends the application boundary of pre-trained symbolic regression models to complex networks, and we believe it provides a foundational solution for revealing the hidden mechanisms behind changes in complex phenomena, enhancing interpretability, and inspiring further scientific discoveries."},"title":{"value":"Symbolic Foundation Regressor on Complex Networks"},"authors":{"value":["Weiting Liu","Jiaxu Cui","Jiao Hu","En Wang","Bo Yang"]}},"tmdate":1772935213654,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2505-21879"],"tcdate":1757310435391,"writers":["~"],"signatures":["~Jiaxu_Cui1"],"forum":"DyvsMnHBrW","license":"CC BY-SA 4.0","number":621104,"cdate":1746057600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1772935213654,"domain":"DBLP.org","id":"DyvsMnHBrW","version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2312.17135v1"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"ren|insactor_instructiondriven_physicsbased_characters"},"authorids":{"value":["~Jiawei_Ren1","https://dblp.org/search/pid/api?q=author:Mingyuan_Zhang:","https://dblp.org/search/pid/api?q=author:Cunjun_Yu:","~Xiao_Ma2","~Liang_Pan2","https://dblp.org/search/pid/api?q=author:Ziwei_Liu_0002:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2312.17135"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2312-17135,\n  publtype={informal},\n  author={Jiawei Ren and Mingyuan Zhang and Cunjun Yu and Xiao Ma and Liang Pan and Ziwei Liu},\n  title={InsActor: Instruction-driven Physics-based Characters},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2312.17135},\n  url={https://doi.org/10.48550/arXiv.2312.17135}\n}\n"},"abstract":{"value":"Generating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animations that reflect high-level human instructions remains a difficult problem due to the complexity of physical environments and the richness of human language. In this paper, we present InsActor, a principled generative framework that leverages recent advancements in diffusion-based human motion models to produce instruction-driven animations of physics-based characters. Our framework empowers InsActor to capture complex relationships between high-level human instructions and character motions by employing diffusion policies for flexibly conditioned motion planning. To overcome invalid states and infeasible state transitions in planned motions, InsActor discovers low-level skills and maps plans to latent skill sequences in a compact latent space. Extensive experiments demonstrate that InsActor achieves state-of-the-art results on various tasks, including instruction-driven motion generation and instruction-driven waypoint heading. Notably, the ability of InsActor to generate physically simulated animations using high-level human instructions makes it a valuable tool, particularly in executing long-horizon tasks with a rich set of instructions."},"title":{"value":"InsActor: Instruction-driven Physics-based Characters"},"authors":{"value":["Jiawei Ren","Mingyuan Zhang","Cunjun Yu","Xiao Ma","Liang Pan","Ziwei Liu"]}},"tmdate":1743245844638,"pdate":1672531200000,"tcdate":1727692171549,"writers":["~"],"signatures":["~Liang_Pan2"],"forum":"ftUS4hCrnN","license":"CC BY-SA 4.0","number":109938,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1743245844638,"domain":"DBLP.org","id":"ftUS4hCrnN","version":2},{"content":{"summary":{"value":"The paper mainly presents a synthetic data generation protocol for axiom learning in transformers. Through training a GPT-2 like transformer on the synthetic demonstrations, experiments show good generalization ability when evaluating it on a larger network. The paper is well formatted and the axiom generation procedure is well explained."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"There are some minor questions that need further clarification\n1. Line 317 \"We train a decoder-based 67 million parameter model based on GPT-2’s architecture\". Does it mean the model in the paper is trained from scratch?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Teaching LLM to reason is an important topic.\n2. Using high-quality synthetic data for pertaining LLM is important and the explanations of the synthetic axiom generation are clear."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is no proper explanation for why training the transformer on the given symbolic expressions will generalize better and even outperform the SoTA GPT-4 model. If my understanding is correct, the author only changes the dataset for pertaining and still uses negative log-likelihood (SFT) as the loss function to train the transformer.  Is it possible the model overfits the symbolic expression dataset the author has provided?\n2. The author attempts to investigate how different embedding techniques, i.e., SPE, LPE, and NoPE, will affect the causal reasoning ability of LLM. As far as I know, most open-source LLMs like llama, they are all using RoPE [1] for positional embedding which is a kind of relative embedding suitable for format learning. I am curious why the author did not mention RoPE in this paper. Moreover, the evaluation of positional encoding is only limited to the final performance and I could not find any deep insight into why SPE will perform better. [2] is a good reference paper for probing different LLM layers and testing their effectiveness.\n\n[1] Jianlin Su et al. RoFormer: Enhanced Transformer with Rotary Position Embedding\n\n[2] Tian Ye et al. Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process"}},"nonreaders":[],"tmdate":1731427896794,"tcdate":1729628843783,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13074/Reviewer_kNHW"],"signatures":["ICLR.cc/2025/Conference/Submission13074/Reviewer_kNHW"],"forum":"te30nmLaFf","number":1,"license":"CC BY 4.0","cdate":1729628843783,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13074/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427896794,"domain":"ICLR.cc/2025/Conference","replyto":"te30nmLaFf","id":"UVL6JUnaum","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Causal Axioms","Transformers","Generalization","LLMs"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"For text-based AI systems to interact in the real world, causal reasoning is an essential skill. Since active interventions are costly to execute, we study to what extent an agent can learn  causal reasoning from symbolic demonstrations of causal axioms. Specifically, we consider an axiomatic training setup where an agent learns from multiple demonstrations of a causal axiom (or rule), rather than incorporating the axiom as an inductive bias or inferring it from data values. A key question is whether the agent would learn to generalize from the axiom demonstrations to new scenarios. For example, if a transformer model is trained on demonstrations of the causal transitivity axiom over small graphs, would it generalize to applying the transitivity axiom over large graphs? \nOur results, based on a novel axiomatic training scheme, indicate that such generalization is possible. For the transitivity axiom, we find that a 67 million parameter transformer model, when trained on linear causal chains (along with some noisy variations) can generalize well to new kinds of graphs, including longer causal chains, causal chains with reversed order, and graphs with branching; even when it is not explicitly trained for such settings. We extend axiomatic training to a harder task of inferring causation from correlation statements and find similar generalization. On both tasks, our model performs at par (or even better) than many larger language models such as GPT-4, Gemini Pro, and Phi-3. Overall, our axiomatic training framework provides a new paradigm of learning causal reasoning in language models that can be extended to arbitrary axioms, as long as sufficient demonstrations can be generated."},"_bibtex":{"value":"@misc{\nvashishtha2025teaching,\ntitle={Teaching Transformers Causal Reasoning through Axiomatic Training},\nauthor={Aniket Vashishtha and Abhinav Kumar and Atharva Pandey and Abbavaram Gowtham Reddy and Vineeth N. Balasubramanian and Amit Sharma},\nyear={2025},\nurl={https://openreview.net/forum?id=te30nmLaFf}\n}"},"title":{"value":"Teaching Transformers Causal Reasoning through Axiomatic Training"},"pdf":{"value":"/pdf/61061bb223af7ae12635feed01dad10feb1cd1d7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"vashishtha|teaching_transformers_causal_reasoning_through_axiomatic_training"},"authorids":{"value":["~Aniket_Vashishtha1","~Abhinav_Kumar3","~Atharva_Pandey1","~Abbavaram_Gowtham_Reddy1","~Vineeth_N._Balasubramanian2","~Amit_Sharma3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Aniket Vashishtha","Abhinav Kumar","Atharva Pandey","Abbavaram Gowtham Reddy","Vineeth N. Balasubramanian","Amit Sharma"]}},"version":2},{"content":{"summary":{"value":"The paper focuses on the research of synthetic data for object re-identification. Although using the name of object re-identification, it only focuses on persons and vehicles. The main contribution of the paper is the proposed Alice benchmarks, which include both synthetic and real-world data. Specifically, an application is training a model with fully-labeled synthetic data and then tuning it on the unlabeled real-world data. A significance of the paper is its real-world data do not have pre-define “cluster” architecture, which is more practical. The authors try to regard it as a successor \n of the famous Market benchmark."},"soundness":{"value":"2 fair"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Please see the weakness.,"},"rating":{"value":"6: marginally above the acceptance threshold"},"details_of_ethics_concerns":{"value":"The authors discuss the legitimate in data-labeling, however, the main step that incurs the privacy problem is capturing data. How can your data captured in Figure 2 protects the privacy of persons and vehicles. As far as I concern, the passerbys cannot know their cars or themselves are trained for AI models. A license from local government is not enough, it seems to be the human right matters;"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"strengths":{"value":"1. Focusing on synthetic data is good; which does not invade privacy;\n2. Good writing and easy to follow;\n3. Some insightful discussions in DISCUSSION AND FUTURE WORK"},"flag_for_ethics_review":{"value":["Yes, Responsible research practice (e.g., human subjects, data release)"]},"weaknesses":{"value":"Although with these strengths, I think the paper is still far away from the acceptance of ICLR 2024.\n1. The contribution is not enough. I fully understand the importance of the benchmark. But as far as I am concerned, the paper’s biggest significance is the “un-clustered” data architecture. Some other contributions as benchmarking existing works and discussion. They cannot be considered as significance due to (1) no significance contribution beyond other works. (2) no novel method is proposed to customize the new-labeled synthetic to real dataset.\n2. The domain adaptation seems not suitable for synthetic to real datasets. The gain of synthetic data is easy and straightforward. Therefore, we should generate more diversed synthetic data to achieve domain generalization rather than domain adaptation. Please refer to ClonePerson and RandPerson for this kind of research;\n3. The authors discuss the legitimate in data-labeling, however, the main step that incurs the privacy problem is capturing data. How can your data captured in Figure 2 protects the privacy of persons and vehicles. As far as I concern, the passerbys cannot know their cars or themselves are trained for AI models. A license from local government is not enough, it seems to be the human right matters;\n4. The authors try to provide some understanding, however, these understandings are not well-defined by mathematics. Instead, they are some intuitive ones without proof. Some understanding in theory is expected."}},"nonreaders":[],"tmdate":1700194049890,"tcdate":1697187748377,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4743/Reviewer_6iq6"],"signatures":["ICLR.cc/2024/Conference/Submission4743/Reviewer_6iq6"],"forum":"vkkHqoerLV","number":1,"license":"CC BY 4.0","cdate":1697187748377,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4743/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700194049890,"domain":"ICLR.cc/2024/Conference","replyto":"vkkHqoerLV","id":"I3RxY7zWkA","forumContent":{"TLDR":{"value":"This paper introduces Alice benchmarks for \"synthetic to real\" domain adaptive object re-identification."},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Synthetic Data","Object re-ID","Benchmarks"]},"supplementary_material":{"value":"/attachment/7f089240669968080aecb950b4db3b05b819493e.pdf"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"For object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap between synthetic source and real-world target. To facilitate developing more new approaches in learning from synthetic data, we introduce the Alice benchmarks, large-scale datasets providing benchmarks as well as evaluation protocols to the research community. Within the Alice benchmarks, two object re-ID tasks are offered: person and vehicle re-ID. We collected and annotated two challenging real-world target datasets: AlicePerson and AliceVehicle, captured under various illuminations, image resolutions, etc. As an important feature of our real target, the clusterability of its training set is not manually guaranteed to make it closer to a real domain adaptation test scenario. Correspondingly, we reuse existing PersonX and VehicleX as synthetic source domains. The primary goal is to train models from synthetic data that can work effectively in the real world. In this paper, we detail the settings of Alice benchmarks, provide an analysis of existing commonly-used domain adaptation methods, and discuss some interesting future directions. An online server has been set up for the community to evaluate methods conveniently and fairly. Datasets and the online server details are available at https://sites.google.com/view/alice-benchmarks."},"_bibtex":{"value":"@inproceedings{\nsun2024alice,\ntitle={Alice Benchmarks: Connecting Real World Re-Identification with the Synthetic},\nauthor={Xiaoxiao Sun and Yue Yao and Shengjin Wang and Hongdong Li and Liang Zheng},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=vkkHqoerLV}\n}"},"title":{"value":"Alice Benchmarks: Connecting Real World Re-Identification with the Synthetic"},"pdf":{"value":"/pdf/b35efd7b4482c9cea9bc862179b522b5782d15d3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"sun|alice_benchmarks_connecting_real_world_reidentification_with_the_synthetic"},"authorids":{"value":["~Xiaoxiao_Sun1","~Yue_Yao1","~Shengjin_Wang1","~Hongdong_Li1","~Liang_Zheng4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiaoxiao Sun","Yue Yao","Shengjin Wang","Hongdong Li","Liang Zheng"]}},"version":2},{"content":{"summary":{"value":"This paper presents DiffPhy, a framework aimed at improving the physical realism of text-to-video (T2V) diffusion models. While existing models excel at generating visually high-quality videos, they often ignore physical laws such as gravity, force, and motion consistency. DiffPhy addresses this gap by introducing a physics-aware fine-tuning paradigm that integrates reasoning from Large Language Models (LLMs) and Multimodal LLMs (MLLMs). The LLM first performs chain-of-thought reasoning on text prompts to infer relevant physical attributes, phenomena, and enhanced contextual descriptions. The MLLM then verifies whether the generated video aligns with these inferred rules, producing differentiable supervision signals that guide the diffusion model’s updates. The training combines three main objectives—physical phenomena loss, physical commonsense loss, and semantic consistency loss—and further employs an attention-injection mechanism to correct physically implausible generations.\n\nTo support this learning process, the authors construct a new dataset, HQ-Phy, consisting of roughly 8,000 real-world videos emphasizing physical interactions and realistic motion. Through extensive experiments on VideoPhy2 and PhyGenBench benchmarks, DiffPhy demonstrates measurable gains over several open and closed-source baselines, including Wan 2.1-14B, Kling, and CogVideoX. Both automated and human evaluations show that the proposed method produces videos with higher semantic alignment and stronger physical plausibility. Overall, the paper provides a technically coherent and empirically validated approach for enhancing diffusion-based video generation with explicit physical reasoning, offering a meaningful step toward bridging the gap between visual fidelity and physical correctness, though some improvements remain moderate in magnitude."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"N/A"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1. The paper is clearly written and logically organized. Each component of the proposed framework—LLM reasoning, MLLM verification, loss formulation, and failure-aware refinement—is explained in a step-by-step manner, supported by intuitive figures (e.g., Figure 2 and Figure 7). The motivation for combining symbolic reasoning with diffusion training is easy to follow, making the work accessible to both machine learning and vision audiences.\n2. By infusing physical rules into video diffusion, the paper addresses one of the most pressing limitations of current T2V systems—the lack of physical realism and commonsense consistency. This direction has high potential impact not only for video generation but also for downstream domains such as robotics simulation, digital content creation, and physics-based reasoning benchmarks. The work can be viewed as a meaningful step toward unifying generative AI with structured world modeling. Overall, the contribution is both timely and relevant to the evolving landscape of physics-informed generative models.\n3. The introduction of the HQ-Phy dataset represents an additional and meaningful contribution. By curating approximately 8 000 real-world videos covering diverse physical interactions—such as gravity-driven motion, collisions, and fluid dynamics—the authors address a key limitation of existing benchmarks, which are often synthetic or too small to support effective fine-tuning. HQ-Phy provides valuable training material for future research in physics-aware video generation and helps bridge the current data gap in real-world physical phenomena. This dataset substantially strengthens the paper’s practical impact and long-term significance to the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The discussion in Section 2 (Video Physics Reasoning) is relatively narrow. It mainly contrasts simulator-based and representation-learning approaches but overlooks a growing line of research that leverages post-training or preference optimization techniques to enhance physical reasoning in generative models. Recent works such as [1-3] are highly relevant and should be discussed. These studies demonstrate that physics awareness can also be introduced during post-training, offering an important comparative context for DiffPhy.\n2. The method is only evaluated on a single backbone (Wan 2.1-14B). Although this model is a strong open-source baseline, improvements demonstrated on one architecture do not fully establish the generality or robustness of the proposed framework. To make the empirical evidence more convincing, DiffPhy should be applied to additional diffusion backbones to verify that the approach generalizes across different architectures and data. \n3. In lines 221–223, the authors state that they decode the predicted latent clip $x_t\\in R^{m\\times c\\times h \\times w}$ at a sampled timestep $t$ into pixel space for MLLM evaluation. However, for larger $t$, these intermediate latents are typically highly noisy and lack meaningful semantic structure. It remains unclear how an MLLM can reliably evaluate alignment or physical correctness from such noisy decoded clips. Without additional clarification or ablation showing the sensitivity of the evaluation to timestep noise, this procedure may introduce unstable or unreliable supervision signals.\n4. This paper strongly depends on physical reasoning ability of LLM and MLLM. Therefore, the entire framework implicitly assumes that the underlying LLM and MLLM possess strong physical reasoning and evaluation capabilities. However, the paper does not include any diagnostic experiments or analysis to validate this assumption. If these models fail to accurately reason about or judge physical phenomena, the provided feedback could be misleading, thereby undermining the claimed improvements. A more systematic assessment of how the quality of the LLM/MLLM affects the overall performance would strengthen the paper’s empirical credibility.\n\n[1] Ipo: Iterative preference optimization for text-to-video generation\n\n[2] Pisa experiments: Exploring physics post-training for video diffusion models by watching stuff drop\n\n[3] RDPO: Real Data Preference Optimization for Physics Consistency Video Generation"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918559428,"tcdate":1761192649320,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6229/Reviewer_YtuE"],"signatures":["ICLR.cc/2026/Conference/Submission6229/Reviewer_YtuE"],"forum":"lPKsPBstHg","number":1,"license":"CC BY 4.0","cdate":1761192649320,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6229/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918559428,"domain":"ICLR.cc/2026/Conference","replyto":"lPKsPBstHg","id":"nXRE8iCc4Y","forumContent":{"TLDR":{"value":"We propose DiffPhy, a generic framework that enables physically-correct and semantically coherent video generation by fine-tuning a pre-trained video diffusion model."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Diffusion","Video Generation","Physical Commensense"]},"supplementary_material":{"value":"/attachment/dd73e8deab01feb18bdae06f4f6ecc4dca24d57c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions, interactions, and dynamics introduce great difficulties when learning physics from data. In this work, we propose DiffPhy, a generic framework that enables physically-correct and photo-realistic video generation by fine-tuning a pre-trained video diffusion model. Our method leverages large language models (LLMs) to infer rich physical context from the text prompt. To incorporate this context into the video diffusion model, we use a multimodal large language model (MLLM) to verify intermediate latent variables against the inferred physical rules, guiding the model’s gradient updates accordingly. MLLM’s textual output is transformed into continuous signals. We then formulate a set of training objectives that jointly ensure physical accuracy and semantic alignment with the input text.  Additionally, failure facts of physical phenomena are corrected via attention injection. We also establish a high-quality physical video dataset containing diverse phyiscal actions and events to facilitate effective finetuning. Extensive experiments on public benchmarks demonstrate that DiffPhy is able to produce state-of-the-art results across diverse physics-related scenarios. Code and data will be made available post-review."},"_bibtex":{"value":"@misc{\nzhang2026think,\ntitle={Think Before You Diffuse: Infusing Physical Rules into Video Diffusion},\nauthor={Ke Zhang and Cihan Xiao and Jiacong Xu and Yiqun Mei and Vishal M. Patel},\nyear={2026},\nurl={https://openreview.net/forum?id=lPKsPBstHg}\n}"},"title":{"value":"Think Before You Diffuse: Infusing Physical Rules into Video Diffusion"},"pdf":{"value":"/pdf/abed66725f92f4403a0e6585ebdac7432376d20d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|think_before_you_diffuse_infusing_physical_rules_into_video_diffusion"},"authorids":{"value":["~Ke_Zhang17","~Cihan_Xiao1","~Jiacong_Xu1","~Yiqun_Mei1","~Vishal_M._Patel1"]},"authors":{"value":["Ke Zhang","Cihan Xiao","Jiacong Xu","Yiqun Mei","Vishal M. Patel"]}},"version":2},{"content":{"summary":{"value":"The paper proposes to use improve Multi-Modal LLM's reasoning capability by fine-tuning with domain-specific formal languages and using external physics simulator."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"None."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. There is some novelty in augmenting MLLM with the formal languages describing the circuits and using an external physics simulator.\n2. The proposed method works well in improving the performance.\n3. The paper is written clearly and easy to follow"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper only proposed a solution for a very specific domain of physical reasoning, namely circuit analysis. This implies that for every other sub-domain, we need to use another set of formal language and physics simulator, which is not scalable. This also implies that this philosophy cannot be extended to physical reasoning problems that don't have any formal languages describing it, which reduces its potential in the frontiers of science research."}},"nonreaders":[],"tmdate":1731429289381,"tcdate":1729937364172,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11154/Reviewer_b9qd"],"signatures":["ICLR.cc/2025/Conference/Submission11154/Reviewer_b9qd"],"forum":"GR0y0F3Ipd","number":1,"license":"CC BY 4.0","cdate":1729937364172,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11154/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429289381,"domain":"ICLR.cc/2025/Conference","replyto":"GR0y0F3Ipd","id":"Mr4mbKP8h1","forumContent":{"TLDR":{"value":"improving multi-modal scientific reasoning capability with physics perception model and simulation assistance"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-modal reasoning","scientific reasoning","physical simulation"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Pre-trained on extensive text and image corpora, current Multi-Modal Large Language Models (MLLM) have shown strong capabilities in general visual reasoning tasks. \nHowever, their performance is still lacking in physical domains that require understanding diagrams with complex physical structures and quantitative analysis based on multi-modal information. \nTo address this, we develop a new framework, named **M**ulti-Modal Scientific Re**A**soning with **P**hysics Perception and **S**imulation (**MAPS**) based on an MLLM. \nMAPS decomposes expert-level multi-modal reasoning task into physical diagram understanding via a Physical Perception Model (PPM) and reasoning with physical knowledge via a simulator. \nThe PPM module is obtained by fine-tuning a visual language model using carefully designed synthetic data with paired physical diagrams and corresponding simulation language descriptions. \nAt the inference stage, MAPS integrates the simulation language description of the input diagram provided by PPM and results obtained through a Chain-of-Simulation process with MLLM to derive the underlying rationale and the final answer. \nValidated using our collected college-level circuit analysis problems, MAPS significantly improves reasoning accuracy of MLLM and outperforms all existing models. \nThe results confirm MAPS offers a promising direction for enhancing multi-modal scientific reasoning ability of MLLMs. \nWe will release our code, model and dataset used for our experiments upon publishing of this paper."},"_bibtex":{"value":"@inproceedings{\nzhu2025maps,\ntitle={{MAPS}: Advancing Multi-Modal Reasoning in Expert-Level Physical Science},\nauthor={Erle Zhu and Yadi Liu and Zhe Zhang and Xujun Li and JinZhou and Xinjie Yu and Minlie Huang and Hongning Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=GR0y0F3Ipd}\n}"},"title":{"value":"MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science"},"pdf":{"value":"/pdf/c84516a8e4b9a68b710453218bcaa36dee327176.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhu|maps_advancing_multimodal_reasoning_in_expertlevel_physical_science"},"authorids":{"value":["~Erle_Zhu1","~Yadi_Liu1","~Zhe_Zhang24","~Xujun_Li2","~JinZhou1","~Xinjie_Yu1","~Minlie_Huang1","~Hongning_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Erle Zhu","Yadi Liu","Zhe Zhang","Xujun Li","JinZhou","Xinjie Yu","Minlie Huang","Hongning Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel approach to modeling physical dynamics by casting it as a next-step geometric graph prediction problem. Instead of representing physical systems as sequences of state vectors, the authors model each timestep as a graph and use geometric GNNs to predict the evolution of these graphs over time. The approach aims to integrate structural inductive biases with learned dynamics, leveraging the flexibility of GNNs in representing complex interactions and spatial relationships. The method is evaluated on several physical simulation tasks and is compared against existing neural and physics-based baselines."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Pls refer to weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The idea of formulating physical dynamics as a geometric graph prediction problem is conceptually interesting, and could potentially offer a unified framework for structured dynamical modeling.\n\n2. The paper leverages geometric deep learning techniques, such as equivariant GNNs, which are well-suited to modeling the symmetries inherent in physical systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the overall framework is promising, the technical details are somewhat underdeveloped. Key components (e.g., how graphs are constructed at each timestep, how node/edge features evolve) are described at a high level and lack sufficient mathematical clarity or justification.\n\n2. The experimental results look weak. Only a small number of benchmarks are used, and comparisons with strong recent baselines (e.g., learned simulators like GNS or differentiable physics engines) are missing.\n\n3. The model’s ability to generalize to systems with different numbers of components or interaction types is not convincingly demonstrated, which is a critical capability for physical simulation models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920798224,"tcdate":1761460544900,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9097/Reviewer_isjy"],"signatures":["ICLR.cc/2026/Conference/Submission9097/Reviewer_isjy"],"forum":"BhYqfdHLA0","number":2,"license":"CC BY 4.0","cdate":1761460544900,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9097/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920798224,"domain":"ICLR.cc/2026/Conference","replyto":"BhYqfdHLA0","id":"ivsbqoZTSQ","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Equivariance","Spatio-Temporal Transformer"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Physical dynamics simulation serves as a foundational component in scientific computing and AI applications. This paper presents a novel approach that redefines the problem as autoregressive prediction of spatiotemporal graph sequences. Built upon the expressivity of Transformer, we propose an Equivariant Spatiotemporal Transformer (EST), extending conventional Transformers with specialized equivariant spatiotemporal blocks. These blocks systematically alternate between spatial and temporal modules, rigorously maintaining E(3) symmetries throughout the process. Moreover, the design incorporates a novel Temporal Difference Graph (TDG) module derived from frame-wise variations, effectively modeling global dynamics and addressing cumulative errors in autoregressive predictions. Unlike traditional graph neural networks, our EST can process variable-length historical sequences and mitigate the persistent challenge of error accumulation in autoregressive processes. Comprehensive evaluations across multiscale physical systems (molecular-, protein-, and macroscopic-scale) demonstrate that our method achieves state-of-the-art performance, thereby showcasing its robust and versatile dynamics simulation capabilities."},"_bibtex":{"value":"@misc{\nli2026physical,\ntitle={Physical Dynamics as Next Geometric Graph Prediction},\nauthor={Zongzhao Li and Jiacheng Cen and Liming Wu and Hao Sun and Ruihua Song and Hangyu Mao and zhangfuzheng and Di ZHANG and Wenbing Huang},\nyear={2026},\nurl={https://openreview.net/forum?id=BhYqfdHLA0}\n}"},"title":{"value":"Physical Dynamics as Next Geometric Graph Prediction"},"pdf":{"value":"/pdf/af459679e2db3ef79a9ece873a4118b4dd335c29.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|physical_dynamics_as_next_geometric_graph_prediction"},"authorids":{"value":["~Zongzhao_Li2","~Jiacheng_Cen1","~Liming_Wu1","~Hao_Sun4","~Ruihua_Song1","~Hangyu_Mao2","~zhangfuzheng1","~Di_ZHANG3","~Wenbing_Huang1"]},"authors":{"value":["Zongzhao Li","Jiacheng Cen","Liming Wu","Hao Sun","Ruihua Song","Hangyu Mao","zhangfuzheng","Di ZHANG","Wenbing Huang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces DiffLM, a novel framework for controllable synthetic data generation using large language models (LLMs). DiffLM addresses the challenges of LLM-based data synthesis, such as limited understanding of target data distributions and complex prompt engineering, particularly for structured data. The key innovation lies in decoupling the learning of data distributions from the LLM's training objectives. This is achieved by employing a variational autoencoder (VAE) to learn a latent representation of the real data, which is then used to guide the LLM's generation process. To enhance the quality of the learned latent space, a diffusion model is incorporated, mitigating the limitations of traditional VAEs in capturing complex distributions. Furthermore, a soft prompt injection module seamlessly integrates the learned latent information into the LLM decoding process without retraining, preserving the LLM's inherent knowledge and reasoning abilities. The authors evaluate DiffLM on seven real-world structured datasets spanning tabular, code, and tool data. Their experiments demonstrate that DiffLM generates high-quality synthetic data, achieving comparable or even superior performance to existing methods on downstream tasks, and even surpassing the performance of real data in certain cases. The proposed framework offers a flexible and robust approach for controllable data synthesis, paving the way for wider adoption of LLMs in various data generation scenarios."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please see comments above"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- DiffLM introduces a combination of VAEs, diffusion models, and LLMs for synthetic data generation. The incorporation of a latent diffusion model within the VAE framework, coupled with the soft-prompt injection mechanism, distinguishes DiffLM from prior work that primarily focuses on either fine-tuning LLMs or using simpler latent variable models for text generation.\n- The paper provides a thorough evaluation across a diverse set of structured datasets and tasks (tabular, code, and tool generation), demonstrating the robustness and adaptability of the proposed framework.\n- The observation that synthetic data generated by DiffLM can outperform real data on certain downstream tasks is  compelling and suggests a potential for knowledge enhancement through synthetic data.\n- The proposed DiffLM framework addresses a significant challenge in leveraging LLMs for data synthesis, namely, controlling the generated data's structure and distribution. The ability to generate high-quality synthetic data for structured formats has broad implications for various applications, including data augmentation, privacy preservation, and software testing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- This work suggests that the proposed approach can achieve controllable synthetic data generation, but the model didnt use any disentangled latent variables learning or other grounding techniques to achieve such controllability, can the authors clarify what exactly the controllability refers to?\n- In experiments section the works shows that the synthetic data can outperform real data for continued pre-training, which is intriguing and interesting, it would be more helpful to provide more analysis and empirical insights on why and how this is made possible.\n- How does the interpolation of latent space look like? How and where did the diffusion model improve the issues of regular VAE model as described in the main text?\n- While automated metrics like DCR and downstream task performance are useful, they do not fully capture the nuances of data quality, especially for complex structured data. Incorporating human evaluation, particularly for code and tool generation, would provide a more holistic assessment of the generated data's usability and realism."}},"nonreaders":[],"tmdate":1731428406043,"tcdate":1731186040396,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10707/Reviewer_hoMx"],"signatures":["ICLR.cc/2025/Conference/Submission10707/Reviewer_hoMx"],"forum":"fRmfDqZ2yq","number":4,"license":"CC BY 4.0","cdate":1731186040396,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10707/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428406043,"domain":"ICLR.cc/2025/Conference","replyto":"fRmfDqZ2yq","id":"zlgul1FCdw","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"DiffLM combines VAEs and diffusion models with LLMs to generate high-quality synthetic data via a novel latent feature injection method."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic data generation","diffusion models","language model"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in large language models (LLMs) have significantly enhanced their knowledge and generative capabilities, leading to a surge of interest in leveraging LLMs for high-quality data synthesis. However, synthetic data generation via prompting LLMs remains challenging due to LLMs' limited understanding of target data distributions and the complexity of prompt engineering, especially for structured formatted data. To address these issues, we introduce DiffLM, a controllable data synthesis framework based on variational autoencoder (VAE), which further (1) leverages diffusion models to reserve more information of original distribution and format structure in the learned latent distribution and (2) decouples the learning of target distribution knowledge from the LLM's generative objectives via a plug-and-play latent feature injection module. As we observed significant discrepancies between the VAE's latent representations and the real data distribution, the latent diffusion module is introduced into our framework to learn a fully expressive latent distribution. Evaluations on seven real-world datasets with structured formatted data (i.e., Tabular, Code and Tool data) demonstrate that DiffLM generates high-quality data, with performance on downstream tasks surpassing that of real data by 2\\%–7\\% in certain cases. Data and code will be released upon acceptance."},"_bibtex":{"value":"@misc{\nzhou2025difflm,\ntitle={Diff{LM}: Controllable Synthetic Data Generation via Diffusion Language Models},\nauthor={Ying Zhou and Xinyao Wang and Yulei Niu and Yaojie Shen and Lexin Tang and Fan Chen and Ben He and Le Sun and Longyin Wen},\nyear={2025},\nurl={https://openreview.net/forum?id=fRmfDqZ2yq}\n}"},"title":{"value":"DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models"},"pdf":{"value":"/pdf/211932b9eeac41ea5e65e45aa150f5399b9b18fb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhou|difflm_controllable_synthetic_data_generation_via_diffusion_language_models"},"authorids":{"value":["~Ying_Zhou5","~Xinyao_Wang2","~Yulei_Niu1","~Yaojie_Shen1","~Lexin_Tang1","~Fan_Chen5","~Ben_He1","~Le_Sun1","~Longyin_Wen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ying Zhou","Xinyao Wang","Yulei Niu","Yaojie Shen","Lexin Tang","Fan Chen","Ben He","Le Sun","Longyin Wen"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Physical Scene Generation"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Automatically generating interactive 3D environments is crucial for scaling up robotic data collection in simulation. While prior work has primarily focused on 3D asset placement, it often overlooks the physical relationships between objects (e.g., contact, support, balance, and containment), which are essential for creating complex and realistic manipulation scenarios such as tabletop arrangements, shelf organization, or box packing. Compared to classical 3D layout generation, producing complex physical scenes introduces additional challenges: (a) higher object density and complexity (e.g., a small shelf may hold dozens of books), (b) richer supporting relationships and compact spatial layouts, and (c) the need to accurately model both spatial placement and physical properties.\nTo address these challenges, we propose PhyScensis, an LLM agent-based framework powered by a physics engine, to produce physically plausible scene configurations with high complexity.\nSpecifically, our framework consists of three main components: an LLM agent iteratively proposes assets with spatial and physical predicates; a solver, equipped with a physics engine, realizes these predicates into a 3D scene; and feedback from the solver informs the agent to refine and enrich the configuration. \nMoreover, our framework preserves strong controllability over fine-grained textual descriptions and numerical parameters (e.g., relative positions, scene stability), enabled through probabilistic programming for stability and a complementary heuristic that jointly regulates stability and spatial relations.\nExperimental results show that our method outperforms prior approaches in scene complexity, visual quality, and physical accuracy, offering a unified pipeline for generating complex physical scene layouts for robotic manipulation."},"_bibtex":{"value":"@inproceedings{\nwang2026physcensis,\ntitle={PhyScensis: Physics-Augmented {LLM} Agents for Complex Physical Scene Arrangement},\nauthor={Yian Wang and Han Yang and Minghao Guo and Xiaowen Qiu and Tsun-Hsuan Wang and Wojciech Matusik and Joshua B. Tenenbaum and Chuang Gan},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=aCVfhY4Qen}\n}"},"title":{"value":"PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement"},"pdf":{"value":"/pdf/a80ca2e2f82bcb6b1458663143e7981305c1250c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|physcensis_physicsaugmented_llm_agents_for_complex_physical_scene_arrangement"},"authorids":{"value":["~Yian_Wang1","~Han_Yang4","~Minghao_Guo1","~Xiaowen_Qiu1","~Tsun-Hsuan_Wang2","~Wojciech_Matusik2","~Joshua_B._Tenenbaum1","~Chuang_Gan1"]},"authors":{"value":["Yian Wang","Han Yang","Minghao Guo","Xiaowen Qiu","Tsun-Hsuan Wang","Wojciech Matusik","Joshua B. Tenenbaum","Chuang Gan"]}},"tmdate":1775877045773,"pdate":1769435976857,"tcdate":1758121408401,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9407/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission9407/Authors"],"forum":"aCVfhY4Qen","license":"CC BY 4.0","number":9407,"cdate":1758121408401,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission9407/-/Full_Submission","ICLR.cc/2026/Conference/Submission9407/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission9407/-/Camera_Ready_Revision"],"mdate":1775877045773,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"aCVfhY4Qen","version":2},{"content":{"summary":{"value":"This paper addresses the fundamental question of predicting training data influence on target benchmarks in multimodal large language models (MLLMs) before any training takes place. The authors conduct a comprehensive empirical analysis using 14 vision-language datasets across 7 diverse tasks and make several counter-intuitive discoveries that challenge conventional wisdom about data selection.\nKey Contributions\n1. Counter-intuitive Empirical Findings: The paper argues that intuitive task similarity is an unreliable predictor of cross-task generalization. \n2. Dataset-level vs. Task-level Influence: The authors demonstrate that data influence operates at the individual dataset level rather than broad task categories.\n3. DATAPROPHET Metric: The paper introduces a training-free, interpretable metric that combines three components:\n    - Multimodal Perplexity: Measures source data difficulty relative to target\n    - Cross-dataset Similarity: Captures alignment in questions, answers, and images using MLLM embeddings\n    - Source Dataset Diversity: Quantifies question coverage using clustering-based silhouette coefficient and entropy"},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"**Contradictory Evidence in Task Category Claims**\nThere appears to be a contradiction between your claim that \"datasets from the same task category do not necessarily help each other the most\" (Figure 1b) and the actual results in Figure 2. Taking ChartQA as an example from the Chart understanding task family: the highest improvement comes from Chart2Text (+16.04%), which is indeed from the same task category as defined in Section 2.1. Spatial reasoning tasks like Open-Spatial (+12.46%) and CLEVR(R) (+9.19%) show lower improvements than the same-category Chart2Text. This pattern appears to support intuitive same-task-category transfer rather than contradicting it. Could you:\n  - Clarify how you define \"task category\" boundaries for this specific analysis?\n  - Explain why the ChartQA example doesn't contradict your main claims about counter-intuitive transfer patterns?\n  - Provide a more systematic analysis of when same-category transfer does vs. doesn't dominate?\n\n**Fundamental Flaws in Cross-Task Transfer Analysis**\n\nYour Observations 2 and 3 in Section 2.2 are based on comparing relative improvements across different target datasets from the same source dataset (same row, different columns in Figure 2). However, this analytical approach suffers from critical confounding factors that invalidate your conclusions:\n\nCore Problem: You compare relative improvements across tasks with fundamentally different baseline difficulties and improvement potentials. The tasks themselves have different intrinsic characteristics that affect their \"improvability,\" making cross-task comparisons of relative gains meaningless.\n\n Specific Evidence from Your Data: For Observation 2, you claim: \"after training on OCR-VQA data, the relative performance gain achieved on ScreenQA (OCR, 17.88%) is lower than that achieved on GeomVerse (map understanding,21.74%).\"\n   However, examining Figure 2's data reveals confounding factors:\n  - GeomVerse shows consistently higher average relative gains (16.89%) compared to ScreenQA (11.14%) across ALL source datasets\n  - GeomVerse's self-improvement ($\\Delta_{s→s}$) is 55.71% vs ScreenQA's 30.38%\n  - This suggests GeomVerse is simply more \"improvable\" as a benchmark, regardless of source dataset\n\n Questions:\n\n  1. Confounding Control: How do you distinguish between genuine cross-task transfer effects and task-intrinsic improvability differences? Your current analysis cannot separate these factors.\n  2. Baseline Normalization: Have you considered normalizing improvements by task-specific baselines or maximum achievable gains? Without such normalization, comparing raw relative\n  improvements across different tasks is scientifically invalid.\n  3. Alternative Explanations: How do you rule out that the observed patterns are due to:\n    - Different evaluation metric sensitivities\n    - Varying task saturation points\n    - Benchmark design artifacts\n    - Different training dynamics needed by task types\n  4. Causal Claims: Your claim that transfer depends on \"individual datasets\" rather than \"task categories\" requires showing that dataset-specific factors (beyond task characteristics) drive\n  the observed patterns. How do you establish this causal relationship?\n\n  The Same Issue Applies to Observation 3: Your examples of \"text-rich tasks influencing vision-centric ones more than text-rich ones\" likely reflect the same confounding - vision-centric\n  tasks may simply have more room for improvement rather than indicating genuine cross-modal transfer superiority.\n\n  Impact on Paper's Validity: This analytical flaw undermines your central claims about dataset-specific vs. task-specific influence. Without proper controls for task-intrinsic factors, your\n  conclusions about \"counter-intuitive\" transfer patterns may be statistical artifacts rather than genuine insights.\n\n  Suggested Resolution: Re-analyze your data using improvement metrics that account for task-specific characteristics, or restrict comparisons to tasks with matched baseline difficulties and\n  improvement potentials."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"**Originality**\n\n- The paper introduces a research question - predicting cross-dataset influence in MLLMs before any training occurs. While data selection has been extensively studied, the specific focus on training-free prediction of multimodal data influence represents a clear departure from existing gradient-based or proxy-model approaches.\n\n- Systematic Empirical Investigation: The comprehensive 14×14 influence matrix analysis is unprecedented in scope for multimodal settings. This systematic approach to mapping cross-task transfer patterns fills an important gap in understanding MLLM behavior.\n\n**Quality**\n- The paper demonstrates experimental breadth by selecting 14 diverse vision-language datasets spanning 7 task families (OCR, chart understanding, document understanding, general VQA, spatial reasoning, counting, and map reasoning). This coverage, with 2 datasets per task type, enables robust cross-task transfer analysis. The systematic 14×14 experimental matrix (196 training-evaluation combinations) provides empirical evidence for understanding how different source datasets influence performance across various target benchmarks.\n- Rigorous Experimental Design: The controlled experimental setup is well-designed with consistent hyperparameters, fixed compute budgets (20K samples each), and standardized evaluation protocols across all 14 datasets. \n\n**Clarity**\n- Well-Structured Presentation: The paper follows a logical progression from motivation → empirical analysis → method development → validation. The three-part takeaway in Figure 1 effectively communicates the main idea.\n\n- Effective Visualization: Figure 2's heatmap clearly illustrates the nature of cross-dataset influence. The color-coding and task groupings make patterns easily interpretable.\n\n**Significance**\nThe paper challenges the widespread assumption that \"similar tasks help similar tasks more.\" The systematic influence analysis framework and the 14-dataset benchmark provide a reference for comparative studies."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Single Model Architecture**: All experiments are conducted exclusively on InternVL3-2B, which severely limits the generalizability of findings. Different MLLM architectures (e.g., GPT-4V, LLaVA, BLIP families) may exhibit fundamentally different cross-task transfer patterns due to varying pretraining objectives, data distributions, and architectural choices. The authors provide no evidence that DATAPROPHET's effectiveness extends beyond this single model family, making it unclear whether the discovered \"counter-intuitive\" patterns are universal phenomena or model-specific artifacts.\n\n**Ill-Defined Problem Formulation and Claims**\n- Vague \"Counter-Intuitive\" Claims: The paper's central claims about counter-intuitive findings are problematic due to poorly defined baselines. In the Introduction, the authors state that \"humans may intuitively assume that training the model on OCR task data will help its performance on chart tasks more than spatial reasoning tasks, since OCR and chart both require extracting text and numbers in an image.\" However, \"intuitive\" is not a well-defined, measurable concept. The authors provide no systematic survey of expert opinions, formal definition of intuitive similarity, or principled baseline for what constitutes \"expected\" transfer patterns. This makes their counter-intuitive claims essentially unfalsifiable and scientifically questionable.\n\n- Undefined Task Categories: The paper repeatedly refers to findings being \"dependent on individual datasets\" rather than \"task category,\" but \"task category\" itself lacks clear definition. The authors don't explain what constitutes a task boundary, how fine-grained the categorization should be, or what criteria distinguish one task category from another. This conceptual ambiguity undermines the central thesis about dataset-level vs. task-level influence. This subjective categorization makes it impossible to assess whether the reported transfer patterns reflect genuine task relationships or merely artifacts of the chosen taxonomy.\n\n**Limited Long-term Training Analysis**: All experiments use single-epoch training, but production MLLM training typically involves multiple epochs and complex scheduling. The influence patterns observed in short training runs may not persist during extended training."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915671254,"tcdate":1760877250207,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1078/Reviewer_LeY5"],"signatures":["ICLR.cc/2026/Conference/Submission1078/Reviewer_LeY5"],"forum":"iYMZKz5BGz","number":1,"license":"CC BY 4.0","cdate":1760877250207,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1078/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915671254,"domain":"ICLR.cc/2026/Conference","replyto":"iYMZKz5BGz","id":"xLwwTz7BmE","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["MLLM"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Conventional wisdom in selecting supervision data for multimodal large language models (MLLMs) is to prioritize datasets that are intuitively similar to the target task (e.g. text-rich v.s. vision-centric). However, it remains unclear how reliably such similarity translates into improved performance on the test benchmarks. \nIn this paper, we take the first step to study the problem in MLLMs: can we predict a training data's influence on a target benchmark even before any training takes place?\nTo answer this question, we first conduct an in-depth analysis using 14 vision-language datasets covering 7 diverse tasks. Our analysis shows that intuitive task\nsimilarity is unreliable in predicting task generalizability, and that transfer depends on the specific dataset rather than the broader task category. \nWe propose DATAPROPHET, a training-free, simple yet effective metric based on multimodal perplexity, similarity, and data diversity. Our experiments demonstrate that the influence rankings for different supervision datasets derived from DATAPROPHET is strongly-correlated with rankings based on the actual performance increase after training, with a Kendall’s $\\tau$ correlation coefficient of 86.0\\%. Moreover, we show that DATAPROPHET can help select better supervision data, achieving up to 6.9\\% improvement in average over uniform selection, 1.4\\% over SoTA training-based baseline, and 0.2\\% higher than oracle experiment performance-based selection. Our code and data will be released."},"_bibtex":{"value":"@inproceedings{\nqi2026demystifying,\ntitle={Demystifying Supervision Data Generalization in Multimodal {LM}s},\nauthor={Xuan Qi and Luxi He and Dan Roth and Xingyu Fu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=iYMZKz5BGz}\n}"},"title":{"value":"Demystifying Supervision Data Generalization in Multimodal LMs"},"pdf":{"value":"/pdf/0b260577fc26a0d4698609336828a91ec288059e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"qi|demystifying_supervision_data_generalization_in_multimodal_lms"},"authorids":{"value":["~Xuan_Qi3","~Luxi_He1","~Dan_Roth3","~Xingyu_Fu1"]},"authors":{"value":["Xuan Qi","Luxi He","Dan Roth","Xingyu Fu"]}},"version":2},{"content":{"venue":{"value":"CoRR 2021"},"pdf":{"value":"https://arxiv.org/pdf/2110.06023v1"},"venueid":{"value":"dblp.org/journals/CORR/2021"},"paperhash":{"value":"battiston|the_physics_of_higherorder_interactions_in_complex_systems"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Federico_Battiston:","https://dblp.org/search/pid/api?q=author:Enrico_Amico:","https://dblp.org/search/pid/api?q=author:Alain_Barrat:","https://dblp.org/search/pid/api?q=author:Ginestra_Bianconi:","https://dblp.org/search/pid/api?q=author:Guilherme_Ferraz_de_Arruda:","https://dblp.org/search/pid/api?q=author:Benedetta_Franceschiello:","https://dblp.org/search/pid/api?q=author:Iacopo_Iacopini:","https://dblp.org/search/pid/api?q=author:Sonia_Kéfi:","https://dblp.org/search/pid/api?q=author:Vito_Latora:","https://dblp.org/search/pid/api?q=author:Yamir_Moreno:","https://dblp.org/search/pid/api?q=author:Micah_M._Murray:","https://dblp.org/search/pid/api?q=author:Tiago_P._Peixoto:","~Francesco_Vaccarino1","https://dblp.org/search/pid/api?q=author:Giovanni_Petri:"]},"html":{"value":"https://arxiv.org/abs/2110.06023"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2110-06023,\n  publtype={informal},\n  author={Federico Battiston and Enrico Amico and Alain Barrat and Ginestra Bianconi and Guilherme Ferraz de Arruda and Benedetta Franceschiello and Iacopo Iacopini and Sonia Kéfi and Vito Latora and Yamir Moreno and Micah M. Murray and Tiago P. Peixoto and Francesco Vaccarino and Giovanni Petri},\n  title={The physics of higher-order interactions in complex systems},\n  year={2021},\n  cdate={1609459200000},\n  journal={CoRR},\n  volume={abs/2110.06023},\n  url={https://arxiv.org/abs/2110.06023}\n}\n"},"abstract":{"value":"Complex networks have become the main paradigm for modelling the dynamics of interacting systems. However, networks are intrinsically limited to describing pairwise interactions, whereas real-world systems are often characterized by higher-order interactions involving groups of three or more units. Higher-order structures, such as hypergraphs and simplicial complexes, are therefore a better tool to map the real organization of many social, biological and man-made systems. Here, we highlight recent evidence of collective behaviours induced by higher-order interactions, and we outline three key challenges for the physics of higher-order systems."},"title":{"value":"The physics of higher-order interactions in complex systems"},"authors":{"value":["Federico Battiston","Enrico Amico","Alain Barrat","Ginestra Bianconi","Guilherme Ferraz de Arruda","Benedetta Franceschiello","Iacopo Iacopini","Sonia Kéfi","Vito Latora","Yamir Moreno","Micah M. Murray","Tiago P. Peixoto","Francesco Vaccarino","Giovanni Petri"]}},"tmdate":1769343402897,"pdate":1640908800000,"externalIds":["dblp:journals/corr/abs-2110-06023"],"tcdate":1769343382912,"writers":["~"],"signatures":["~Francesco_Vaccarino1"],"forum":"bombWVvO8S","license":"CC BY-SA 4.0","number":797242,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769343402897,"domain":"DBLP.org","id":"bombWVvO8S","version":2},{"content":{"summary":{"value":"This paper introduces Physics-Inspired Neural Compression (PINC), a framework for compressing high-dimensional plasma turbulence simulations while preserving essential physical properties. The authors address the challenge of storing large gyrokinetic simulation data by proposing a single evaluation pipeline that measures both spatial and temporal turbulence characteristics. They investigate two neural compression paradigms (autoencoders and neural implicit field) and augment them with physics-informed loss terms derived from gyrokinetic integrals and turbulence spectra. The resulting PINC models achieve great compression rates while maintaining key physical quantities. The paper provides detailed quantitative and qualitative analyses, demonstrating that PINC significantly improves physics preservation compared to traditional compression methods."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"- How sensitive is the proposed physics-inspired loss formulation to the specific weighting or scaling of the individual components?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- The paper addresses an underexplored yet highly relevant problem: data compression for large-scale physics simulations, rather than accelerating the simulations themselves. This is a practical and impactful direction, as it targets a major bottleneck in scientific computing: data storage and accessibility.\n\n- The proposed Physics-Inspired Neural Compression (PINC) is conceptually well-motivated, bridging neural compression and physics-informed learning in a principled way. The inclusion of physically meaningful loss terms for gyrokinetics demonstrates a strong understanding of the domain.\n\n- The experimental evaluation is extensive and well-structured, including comparisons with traditional compression methods, ablations of individual loss components, and scaling analyses across compression ratios.\n\n- The results are compelling, showing that the proposed approach achieves extreme compression (up to 70,000x) while maintaining physically relevant quantities, a feat that existing methods fail to achieve.\n\n- The paper is clearly written and well organized, with solid theoretical grounding and reproducibility details (including code, configurations, and dataset description).\n\n- The authors provide a balanced discussion of limitations and outline meaningful future directions, such as incorporating temporal consistency and extending the approach to other domains.\n\nOverall, the work sets a new benchmark for physics-preserving compression and highlights the potential of neural networks as a viable alternative to traditional methods in high-dimensional scientific data management."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Paper is really strong, minor potential issues: \n\n- While the proposed PINC framework is convincing, it remains domain-specific, with the physics-informed losses tailored to gyrokinetic equations. It is unclear how easily the method generalizes to other scientific domains (e.g., fluid dynamics, astrophysics). \n\n- Are there insights from this study that could inform the design of neural operators or surrogate models, possibly using compressed representations as priors or initialization?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927980651,"tcdate":1761988176684,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18255/Reviewer_Dmk1"],"signatures":["ICLR.cc/2026/Conference/Submission18255/Reviewer_Dmk1"],"forum":"fixbsplpdw","number":3,"license":"CC BY 4.0","cdate":1761988176684,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18255/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927980651,"domain":"ICLR.cc/2026/Conference","replyto":"fixbsplpdw","id":"lQauM9UWlC","forumContent":{"TLDR":{"value":"Neural compression methods enable extreme compression of plasma turbulence simulation data while maintaining low reconstruction error and preserving key physical characteristics."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics-inspired","turbulence","plasma","neural compression","autoencoders","neural fields"]},"supplementary_material":{"value":"/attachment/0814372529d7d0e61c5f895f3e8f4b70b3af92df.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"High-fidelity scientific simulations are now producing unprecedented amounts of data, creating a storage and analysis bottleneck. A single simulation can generate tremendous data volumes, often forcing researchers to discard valuable information. A prime example of this is plasma turbulence described by the Gyrokinetic equations: nonlinear, multiscale, and 5D in phase space. They represent one of the most computationally demanding frontiers of modern science, with runs taking weeks and resulting in tens of terabytes of data dumps.\nThe increasing storage demands underscore the importance of compression, however, compressed snapshots might not preserve essential physical characteristics after reconstruction. To assess whether such characteristics are captured, we propose a spatiotemporal evaluation pipeline which accounts for structural phenomena and multi-scale transient fluctuations. Indeed, we find that various compression techniques lack preservation of temporal turbulence characteristics. Therefore, we explore Physics-Informed Neural Compression (PINC), which incorporates physics-informed losses tailored to gyrokinetics and enables extreme compressions of over 100000x. This direction provides a viable and scalable solution to the prohibitive storage demands of gyrokinetics, enabling post-hoc analyses that were previously infeasible."},"_bibtex":{"value":"@misc{\ngalletti2026physicspreserving,\ntitle={Physics-Preserving Compression of High-Dimensional Plasma Turbulence Simulations},\nauthor={Gianluca Galletti and Gerald Gutenbrunner and Fabian Paischer and Sandeep Suresh Cranganore and William Hornsby and Naomi Carey and Lorenzo Zanisi and Stanislas Pamela and Johannes Brandstetter},\nyear={2026},\nurl={https://openreview.net/forum?id=fixbsplpdw}\n}"},"title":{"value":"Physics-Preserving Compression of High-Dimensional Plasma Turbulence Simulations"},"pdf":{"value":"/pdf/88f0fd80bf4066acbbabbaac93b2beb84178ea60.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"galletti|physicspreserving_compression_of_highdimensional_plasma_turbulence_simulations"},"authorids":{"value":["~Gianluca_Galletti1","~Gerald_Gutenbrunner1","~Fabian_Paischer1","~Sandeep_Suresh_Cranganore1","~William_Hornsby1","~Naomi_Carey1","~Lorenzo_Zanisi1","~Stanislas_Pamela1","~Johannes_Brandstetter1"]},"authors":{"value":["Gianluca Galletti","Gerald Gutenbrunner","Fabian Paischer","Sandeep Suresh Cranganore","William Hornsby","Naomi Carey","Lorenzo Zanisi","Stanislas Pamela","Johannes Brandstetter"]}},"version":2},{"content":{"summary":{"value":"This paper presents M2F-PINN, a Multi-scale Frequency-domain Multi-Physics-Informed Neural Network for large-scale ocean current forecasting. The authors aim to address two known issues in ocean forecasting models: (1) deep models’ difficulty in learning both low- and high-frequency ocean dynamics due to spectral bias, and (2) the lack of physical consistency in purely data-driven ocean models. M2F-PINN propose several crucial components: multi-scale Fourier feature embeddings, a 3D Swin Transformer backbone for spatiotemporal feature extraction, and multiple Physics-Informed Neural Network (PINN) modules. The model is trained on the GLORYS12 reanalysis dataset (global ocean, 2005–2008) and evaluated against CNNs, RNNs, MeshGraphNets, Fourier Neural Operators, and state-of-the-art ocean models. M2F-PINN achieves the best results across 1–60-day prediction horizons, outperforming strong baselines and showing improved long-term physical consistency."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- How exactly are forecasts generated? Is the model trained in an autoregressive fashion (iterative multi-step rollout) or direct multi-horizon regression? If autoregressive, how is error accumulation handled?\n\n- Given that only U and V momentum equations are included, how can the model capture coupled dynamics driven by temperature and salinity gradients? Would incorporating the full Navier–Stokes or continuity equations improve realism?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper systematically combines frequency-domain embeddings, physics-informed losses, and Transformer architectures in a coherent way. Although each component exists independently, their integration into an end-to-end framework is well-executed.\n\n- The NTK-based spectral analysis provides a clear explanation of how Fourier embeddings reshape learning rates across frequency bands, offering some theoretical insight beyond empirical results."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The proposed components—multi-scale Fourier features, PINN-based physical losses, and uncertainty-weighted multi-task optimization—are all mature and widely used in physics-informed and climate modeling literature (e.g., FNO, ClimODE, LangYa, frequency-domain PINNs). The contribution is largely an engineering combination rather than a fundamentally new algorithmic advance. The claimed innovation in “frequency-domain multi-PINN” lacks distinctive technical depth beyond prior works. \n\n- The physics component only implements two simplified momentum equations for (U, V), omitting key coupled variables such as temperature, salinity, and density that govern realistic ocean circulation. As a result, the “multi-physics” claim is overstated—the scope is narrow and its real-world interpretability limited.\n\n- CNN and RNN serve as trivial references and do not represent current forecasting standards. Comparisons to XiHe and WenHai are insufficiently documented—there is no clarification whether they were re-trained under the same data regime or results are copied from publications. Missing comparisons with more recent AI-based earth system models (e.g., GraphCast, Fuxi, Pangu) weaken the empirical credibility.\n\n- Despite the title emphasizing forecasting, the paper fails to specify how prediction is carried out.The architecture uses a Swin Transformer encoder-decoder, but it is unclear whether forecasts are autoregressive (iteratively rolling forward) or direct multi-step predictions.The “Algorithm 1” description is ambiguous: while labeled “autoregressive,” it lacks any explicit temporal unrolling or teacher-forcing details. Without a precise description of forecast horizon handling, it is hard to interpret the reported 1-, 7-, and 30-day results or to assess generalization stability.\n\n- The paper asserts that Fourier embeddings capture “multi-scale oceanic structures,” yet provides no frequency-space visualizations or power-spectrum analysis to substantiate this claim. Likewise, the physical residuals (PIC) are introduced, but no concrete examples of physically consistent predictions (e.g., energy or vorticity preservation) are shown."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931446825,"tcdate":1762061035613,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19568/Reviewer_JjqD"],"signatures":["ICLR.cc/2026/Conference/Submission19568/Reviewer_JjqD"],"forum":"4z7tMSzQET","number":4,"license":"CC BY 4.0","cdate":1762061035613,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19568/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931446825,"domain":"ICLR.cc/2026/Conference","replyto":"4z7tMSzQET","id":"XLlZthvz49","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"This work introduces  a multi-scale, frequency-domain, physics-informed neural network for ocean forecasting that enhances physical interpretability and effectively captures frequency information."},"keywords":{"value":["physics-informed neural networks(PINN)","multi-scale Fourier feature","ocean forecasting"]},"supplementary_material":{"value":"/attachment/0f7263a6ae9549b8db445a92f0dd4b9d255ff90f.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics‐informed neural networks (PINNs) embed physical laws into data-driven learning and are becoming increasingly influential in climate and ocean forecasting. Yet effectively capturing multi-scale variability across high and low frequencies while maintaining training stablility and ensuring convergence remains challenging for conventional PINNs. We introduce M$^2$F-PINN, a novel Transformer-based multi-scale frequency-domain multi-PINN algorithm designed to 1) mitigate spectral bias via Fourier representation learning, and 2) analyze multi-scale characteristics through frequency-domain modeling, and 3) incorporate physics priors using multiple PINNs. M$^2$F-PINN leverages multi-scale Fourier networks to learn spectral components and multi-scale interactions, and employs a 3D Swin Transformer in an autoregressive setting to capture spatiotemporal regularities. The advantages of M$^2$F-PINN include: 1) adaptively learns multi-scale frequency components to enhance the modeling of multi-scale dynamics; 2) jointly estimates physical coefficients within the PINN modules, refining representations of physical processes; 3) preserves the Transformer framework, enabling compatibility with diverse architectures and structural decoupling; 4) extensive experiments on real-world ocean datasets show that M$^2$F-PINN outperforms deep-learning baselines and competitive ocean models (e.g., XiHe, WenHai) in predicting ocean current fields, achieving superior performance across multiple time horizons."},"_bibtex":{"value":"@misc{\nlinfei2026mfpinn,\ntitle={M{\\texttwosuperior}F-{PINN}: A Multi-Scale Frequency-Domain Multi-Physics-Informed Neural Network for Ocean Forecasting},\nauthor={Cao Linfei and Jiachen Yang and Meng Xi and Jingyi He and Fei Gao and Jiabao Wen and Zhijing Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=4z7tMSzQET}\n}"},"title":{"value":"M²F-PINN: A Multi-Scale Frequency-Domain Multi-Physics-Informed Neural Network for Ocean Forecasting"},"pdf":{"value":"/pdf/12031d2723c890bb5102cc1a2a3540f3957dd963.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"linfei|mfpinn_a_multiscale_frequencydomain_multiphysicsinformed_neural_network_for_ocean_forecasting"},"authorids":{"value":["~Cao_Linfei1","~Jiachen_Yang2","~Meng_Xi2","~Jingyi_He2","~Fei_Gao29","~Jiabao_Wen2","~Zhijing_Wang2"]},"authors":{"value":["Cao Linfei","Jiachen Yang","Meng Xi","Jingyi He","Fei Gao","Jiabao Wen","Zhijing Wang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a hybrid-physics architecture and training method that achieves better physical consistency and generalizability than traditional simulators and neural models for elastic physics rollout. The modular design for elastic physics is inspired by the FEM method."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Do you plan to release the code for your model training and simulations?\n- What are the costs in terms of flops and time of the NMF model inference vs the simulator used to generate ground truth trajectories?\n- Since the neural constitutive model requires the $B_e$ matrix and $V_e$ values to work, do you provide these in your test scenarios to the model? Is the $B_e$ term easy to derive in some way or should it also be learned?\n- How do you determine when to stop the separate training phase for each module? Do you train for a fixed number of steps or use some other condition for stopping?\n- You use a regularization constant of 0.1 for the volume loss term you add during joint training. How did you select this value? What happens if you increase it to make the physical constraint stronger? Does the overall trajectory become more physically plausible or does your model fail to learn due to the difficulty of satisfying this constraint?\n- There are few typos I noticed in the paper that can be corrected:\n  - Line 121: \"Examples include learned data-driven discretization stencils (Bar-Sinai et al., 2019) and learned residual\ndynamics on top of analytical dynamics to account for unmodeled (Yin et al., 2021).\" Maybe you want to write unmodeled dynamics?\n  - Line 211: \"Still start from classical FEM.\" Rewrite this to be a more clear introductory sentence.\n  - Line 424: \"With out the elaborative physical-aligned architecture and specialized training strategy in NMP, we cannot release the unique benefits of modularization.\" I think you mean to write \"we cannot realize\" instead of \"we cannot release\"."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The experiments provide good empirical evidence that their NMF model has better long term rollout performance for both seen and unseen scenarios in comparison to the baselines.\n- Their NMF model allows the addition of soft physics constraints in a fairly straightforward way.\n- The paper provides a thorough explanation of their modular model design and reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The architecture is specialized for elastic physics, which will not generalize to other equations and setttings. However, I acknowledge the idea of replacing components of a simulator with neural modules is an interesting idea.\n- The paper proposes a way to replace parts of an FEM method with neural networks, but this is not straightforward to apply to other types of solver schemes such as finite difference methods and spectral methods.\n- The strain-displacement matrix is required a priori to use the NMF model."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917052277,"tcdate":1761926034192,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3825/Reviewer_ZLNW"],"signatures":["ICLR.cc/2026/Conference/Submission3825/Reviewer_ZLNW"],"forum":"qYVa1obqTZ","number":2,"license":"CC BY 4.0","cdate":1761926034192,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3825/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917052277,"domain":"ICLR.cc/2026/Conference","replyto":"qYVa1obqTZ","id":"o4vAX1vFNg","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"This paper presents Neural Modular Physics, a thorough modular framework for elastic simulation, enabling better generalization and stable long-horizon simulation."},"keywords":{"value":["Neural Simulation","Neural Modular Networks","Elastic Dynamics","learned simulation"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Learning-based methods have made significant progress in physics simulation, typically approximating dynamics with a monolithic end-to-end optimized neural network. Although these models offer an effective way to simulation, they may lose essential features compared to traditional numerical simulators, such as physical interpretability and reliability. Drawing inspiration from classical simulators that operate in a modular fashion, this paper presents Neural Modular Physics (NMP) for elastic simulation, which combines the approximation capacity of neural networks with the physical reliability of traditional simulators. Beyond the previous monolithic learning paradigm, NMP enables direct supervision of intermediate quantities and physical constraints by decomposing elastic dynamics into physically meaningful neural modules connected through intermediate physical quantities. With a specialized architecture and training strategy, our method transforms the numerical computation flow into a modular neural simulator, achieving improved physical consistency and generalizability. Experimentally, NMP demonstrates superior generalization to unseen initial conditions and resolutions, stable long-horizon simulation, better preservation of physical properties compared to other neural simulators, and greater feasibility in scenarios with unknown underlying dynamics than traditional simulators."},"_bibtex":{"value":"@misc{\nli2026neural,\ntitle={Neural Modular Physics for Elastic Simulation},\nauthor={Yifei Li and Haixu Wu and Zeyi Xu and Tuur Stuyck and Wojciech Matusik},\nyear={2026},\nurl={https://openreview.net/forum?id=qYVa1obqTZ}\n}"},"title":{"value":"Neural Modular Physics for Elastic Simulation"},"pdf":{"value":"/pdf/72ba4587211cfe2a7de32067a758b4ec539f8603.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|neural_modular_physics_for_elastic_simulation"},"authorids":{"value":["~Yifei_Li7","~Haixu_Wu1","~Zeyi_Xu2","~Tuur_Stuyck1","~Wojciech_Matusik2"]},"authors":{"value":["Yifei Li","Haixu Wu","Zeyi Xu","Tuur Stuyck","Wojciech Matusik"]}},"version":2},{"content":{"summary":{"value":"This work studies contrastive learning in three settings: 1) gaussian data with linear model, 2) physics datasets non-linear model, 3) image datasets with non-linear model. The work develops theory in the linear setting, suggesting that while the training loss may converge monotonically, the geometry of the representation space undergoes discrete phase transitions. Numerical experiments provided in all considered settings.\n\n## Detailed Summary\n\n### Linear setting:\n* learn invariance to additive augmentation vectors, and end up with representations that span the space orthogonal to the additive augmentation vectors d… these vectors are called targets t\n* assume infinitesimal step-size\n* show that sampling augmentations from the space spanned by d according to some continuous distribution with support orthogonal to the target vectors is equivalent in this setting to sampling augmentations from a discrete uniform distribution over a set of basis vectors spanning the space orthogonal to the predefined targets t\n* Prop3: the weight matrix converges to the matrix whose columns form a basis for the space spanned by the targets\n* define a cosine metric between matrices, that is used to measure how close two matrices are to one another\n* argue that while the weight matrix converges to the target matrix monotonically, it’s repulsion from the space spanned by the augmentation basis vectors is non-monotonic with the cosine measure\n* No phase transition observed in numerical experiments\n\n### Physics dataset\n* consider the Kepler dataset, where each sample consists of a video of a point moving along a 3d trajectory, which can be described with three “latent” variables\n* train a contrastive model where positive samples come from frames in the same trajectory and negatives come from different trajectories, parametrized by a different set of latent variables\n* show than when sampling positives at short temporal horizons (i.e., “weak” augmentations), the model first learns a type of shortcut solution that exhibits a certain geometry in latent space, and then phase transitions suddenly into another geometry that captures the underlying factors of variation\n\n### Imagenet/CIFAR10\n* consider supervised contrastive learning, where \"strength\" of augmentations are computed by sampling images from the same class with varying visual similarity\n* measure the degree of representation clustering during training with adjusted Rank Index (ARI) and adjusted Mutual Information (AMI) and find that the rate at which these metrics improve depend on the \"strength\" of the augmentations"},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"### Significance\n* This work studies an interesting problem targeting an improved understanding of the training dynamics of contrastive learning\n\n### Quality\n* First seek to motivate claims with a theoretical explanation in the linear settings\n* Consider both small-scale (linear models with gaussian data / physics data) and larger-scale (ImageNet) problem settings\n\n### Clarity\n* Figures in the setting of the physics dataset are interesting and engaging"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### Theory\n* The theoretical intuitions derived from the theory are quite hand-wavy, even for describing how representations evolve in the toy setting. For instance, to argue how a phase could be stable, it is stated that the exponential terms must dominate the expression for the matrix cosine, which requires $cos_t \\approx c^1_i / \\sqrt{c^2_i}$. But, in fact, these quantities depend on a complex ratio of matrix power series, which themselves depend on the eigendecomposition of the matrix B, which itself is computed from a Cholesky factorization of the expected outer product of the additive data augmentation vectors, and this of course depends on the distribution from which augmentation vectors are sampled. When would such a criteria ever be satisfied in practice and how would you know by looking at your data or your data augmentations?\n* Why do you not observe phase transformations in the linear setting? How would you re-design your problem based on your theory to observe phase transitions?\n\n### Physics Dataset\n* It is not clear to me whether observations in the physics dataset originate from the rough theoretical motivation developed in the linear setting.\n* It is also not clear to me how hyper-parameter considerations (learning rate/schedules/etc.) affect any observed behaviours in terms of the geometry of the representation space.\n\n### Image Datasets\n* Quite frankly, I do not see any significant phase transition based on the ARI metric or the AMI metric.\n* The low-dimensional clusterings of the image representations for supervised contrastive learning show that supervised contrastive learning clusters images form the same class together during training, however, this has already been observed in the literature, and I'm not sure that there are any new interesting intuitions about how the geometry of the representation space evolves.\n* Not clear to me how the theory in the linear setting has contributed to explaining the progressive image clustering in this supervised contrastive learning setting."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"* While you observe a phase transition in the Physics Dataset, how can you verify that this is due to the intuition developed in the theory section?\n\n* How robust are these observations to changes in hyper-parameters, such as learning rate magnitude or schedules. What if you reverted to the infinitesimal learning rate assumption made in the linear setting?\n\n* How does the AMI/ARI metric relate to the convergence to the “target” matrix versus the potentially non-monotonic convergence to the B matrix in the linear setting?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637145496,"tcdate":1698810302906,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission9102/Reviewer_7pQx"],"signatures":["ICLR.cc/2024/Conference/Submission9102/Reviewer_7pQx"],"forum":"dAqH7CfHjL","number":3,"license":"CC BY 4.0","cdate":1698810302906,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission9102/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637145496,"domain":"ICLR.cc/2024/Conference","replyto":"dAqH7CfHjL","id":"H6duqYvj9P","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["representation learning","training dynamics","contrastive learning"]},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"How do self-supervised models actually train? We study the training dynamics of contrastive learning in three settings: a theoretical linear setting, on a low-dimensional physics-inspired dataset, and on full-fledged computer vision datasets including ImageNet. In all three settings, we show the existence of *phases*, i.e. locally stable or metastable representations, and of *phase transitions*, wherein a model rapidly and unexpectedly switches between different phases. Geometrically motivated metrics are developed to measure phase transitions. Finally, we show that phase transitions can be sped up with more robust augmentations. Code and visualizations will be made public upon publication."},"_bibtex":{"value":"@misc{\ncy2024phase,\ntitle={Phase Transitions in Contrastive Learning},\nauthor={Ali Cy and Anugrah Chemparathy and Michael Han and Rumen Dangovski and Peter Y. Lu and Charlotte Loh and Marin Soljacic},\nyear={2024},\nurl={https://openreview.net/forum?id=dAqH7CfHjL}\n}"},"title":{"value":"Phase Transitions in Contrastive Learning"},"pdf":{"value":"/pdf/5a8349b80474ac3fbffb875ce7f9090c1c0ec43c.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"cy|phase_transitions_in_contrastive_learning"},"authorids":{"value":["~Ali_Cy1","~Anugrah_Chemparathy1","~Michael_Han1","~Rumen_Dangovski1","~Peter_Y._Lu1","~Charlotte_Loh1","~Marin_Soljacic1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ali Cy","Anugrah Chemparathy","Michael Han","Rumen Dangovski","Peter Y. Lu","Charlotte Loh","Marin Soljacic"]}},"version":2},{"content":{"comment":{"value":"**Q3. [Figure 4: I would be interested in a deeper analysis of, for example, the differences between humans and DPI-Net; while on some tasks, the accuracy seems similar, the correlation is low, indicating they do well in different domains? Would it be possible to highlight properties/trials in which either human does better than DPI-Net and vice versa?]** \\\nGood suggestion! We agree that a comprehensive analysis of the variations between human and DPI-Net performance is essential. Our current findings suggest that while the accuracy levels may be comparable, the correlation between their performances is low, indicating their proficiency in distinct domains. We have conducted a more thorough analysis of the data for highlighting instances where either humans or DPI-Net outperform the other. \n\nThe representative videos are uploaded to our [webpage](https://dingmyu.github.io/physion_v2/). We can see that humans tend to have trouble dealing with unusual camera angles, contrasting physical properties (e.g., two objects with one very heavy and one very light, or very large elasticity), and long-term predictions. But DPI-Net can handle these scenarios better, especially for trials in mass and deformability. Instead models may fail in simple situations (due to lack of distance perception and intuitive physics) that are easier for humans to understand.\n\nIn our revised manuscript, we will incorporate a detailed breakdown of the trials or properties in which humans exhibit better performance compared to baseline models, as well as cases where each model demonstrates superiority over human performance. We believe that this analysis will provide valuable insights into the capabilities and limitations of different methods and contribute to a more comprehensive physical scene understanding.\n\n\n**Q4. [In the discussion, the authors say that none of the models match human performance, but it seems like DPI-Net does even better, so that sentence is not quite correct?]** \\\nThanks for pointing this out, we'll rephrase the sentence. Their overall performances are comparable.\n\n\n**Q5. [Table 1 has a column 'Diverse Phenomena'. This term is never referred to/explained in the main text. The same applies to few-shot reasoning. As of now, these are the two entries in the table discriminating physion++ from IntPhys. It would therefore be very valuable to explain these or refer to them in the main text, highlighting how physion++ is different.]** \\\nWe apologize for any confusion this may have caused and will make it clearer in our revision. Regarding 'Diverse Phenomena', we refer to whether the dataset covers diverse scenarios. As shown in Figure 1, our dataset has 9 scenarios, including rigid bodies, fluids, soft bodies, objects of various shapes, and various physical properties. Previous physical reasoning datasets such as Comphy and CLEVRER only have single or two scenarios with few object shapes.\n\nAs for 'Few-shot Reasoning', we mean that our dataset allows physical properties to be judged from a few reference video frames (inference phase in Figure 2). Completing the OCP task based on few-shot reference videos can also be understood as property inference. We'll make it clearer.\n\n\\\nThanks again for your time and effort! For any other questions, please feel free to let us know during the rebuttal window.\n\nBest, \\\nAuthors"},"title":{"value":"Rebuttal by Authors (Part 2/2)"}},"tmdate":1704277176386,"tcdate":1692757048932,"writers":["NeurIPS.cc/2023/Track/Datasets_and_Benchmarks","NeurIPS.cc/2023/Track/Datasets_and_Benchmarks/Submission735/Authors"],"signatures":["NeurIPS.cc/2023/Track/Datasets_and_Benchmarks/Submission735/Authors"],"forum":"5Exz7eaBXH","number":7,"license":"CC BY 4.0","cdate":1692757048932,"mdate":1704277176386,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Track/Datasets_and_Benchmarks/Submission735/-/Official_Comment","NeurIPS.cc/2023/Track/Datasets_and_Benchmarks/-/Edit"],"domain":"NeurIPS.cc/2023/Track/Datasets_and_Benchmarks","replyto":"gPxEOH1B3z","id":"gJEF66ada6","forumContent":{"venue":{"value":"NeurIPS 2023 Datasets and Benchmarks Poster"},"keywords":{"value":["intuitive physics","visual dynamics learning","physical property inference"]},"supplementary_material":{"value":"/attachment/c7626bc687f48e0c3db966b5ccadc8667a9c7813.pdf"},"abstract":{"value":"General physical scene understanding requires more than simply localizing and recognizing objects -- it requires knowledge that objects can have different latent properties (e.g., mass or elasticity), and that those properties affect the outcome of physical events. While there has been great progress in physical and video prediction models in recent years, benchmarks to test their performance typically do not require an understanding that objects have individual physical properties, or at best test only those properties that are directly observable (e.g., size or color). This work proposes a novel dataset and benchmark, termed Physion++, that rigorously evaluates visual physical prediction in artificial systems under circumstances where those predictions rely on accurate estimates of the latent physical properties of objects in the scene. Specifically, we test scenarios where accurate prediction relies on estimates of properties such as mass, friction, elasticity, and deformability, and where the values of those properties can only be inferred by observing how objects move and interact with other objects or fluids. We evaluate the performance of a number of state-of-the-art prediction models that span a variety of levels of learning vs. built-in knowledge, and compare that performance to a set of human predictions. We find that models that have been trained using standard regimes and datasets do not spontaneously learn to make inferences about latent properties, but also that models that encode objectness and physical states tend to make better predictions. However, there is still a huge gap between all models and human performance, and all models' predictions correlate poorly with those made by humans, suggesting that no state-of-the-art model is learning to make physical predictions in a human-like way. These results show that current deep learning models that succeed in some settings nevertheless fail to achieve human-level physical prediction in other cases, especially those where latent property inference is required. Project page: https://dingmyu.github.io/physion_v2/"},"_bibtex":{"value":"@inproceedings{\ntung2023physion,\ntitle={Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties},\nauthor={Hsiao-Yu Tung and Mingyu Ding and Zhenfang Chen and Daniel Bear and Chuang Gan and Joshua B. Tenenbaum and Daniel LK Yamins and Judith E Fan and Kevin A. Smith},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track},\nyear={2023},\nurl={https://openreview.net/forum?id=5Exz7eaBXH}\n}"},"title":{"value":"Physion++: Evaluating Physical Scene Understanding that Requires Online Inference of Different Physical Properties"},"pdf":{"value":"/pdf/38bc1aa97fbdd30eaa5cc429f54b89c1c6651450.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Track/Datasets_and_Benchmarks"}},"version":2},{"content":{"summary":{"value":"This paper proposes **MH-pFedHNDD** (Model-Heterogeneous personalized Federated learning framework based on HyperNetworks with Data Distillation), which aims to address both model and data heterogeneity in personalized federated learning. The key innovation is integrating data distillation into a hypernetwork-based framework. The method involves two phases: (1) a distillation phase where clients generate synthetic data using *Contrastive Condensation Loss* to create compact, discriminative synthetic datasets that are aggregated on the server; (2) a training phase where clients train personalized models generated by the hypernetwork using both local and synthetic data, guided by a *Reg Loss* with universum negatives. Experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate improvements over existing methods."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"### Q1. Distillation phase design and frequency\n\nSeveral design choices regarding the distillation phase need clarification:\n- **One-time vs periodic distillation**: Why is data distillation performed only once at the beginning? Intuitively, as the hypernetwork improves during training, it should generate better distillation models, leading to higher-quality synthetic data. Have you experimented with:\n  - Periodic re-distillation (e.g., every K rounds)?\n  - Adaptive re-distillation based on performance metrics?\n  - What is the accuracy vs. distillation frequency trade-off?\n- **Step 7 explanation**: What is Step 7 in the Distillation Phase shown in Figure 1? The text describes Steps 1-6 but Step 7 is clearly marked in the figure without any explanation.\n- **Sensitivity to initialization**: How sensitive is the method to the quality of the initial distillation model and the architecture of the distillation model?\n\n### Q2. Computational and communication costs\n\nThis is critical for practical federated learning deployment but is completely missing from the paper:\n- **Computational overhead**:\n  - How much time does the distillation phase take compared to the total training time?\n  - What is the per-round training time with synthetic data vs. MH-pFedHN (without synthetic data)?\n  - How many iterations/epochs are needed for distillation optimization?\n- **Communication costs**:\n  - What is the size of synthetic data $S_i$ uploaded by each client?\n  - How does total communication cost (synthetic data + model updates) compare with MH-pFedHN and MH-pFedHNGD?\n  - Is there a trade-off between IPC and communication cost?\n- **Wall-clock time**: Can you provide convergence curves showing:\n  - Accuracy vs. wall-clock time (not just communication rounds)\n  - Time to reach certain accuracy thresholds\n  - Comparison with baselines in terms of real training time\n\n### Q3. Experimental Settings and Hyperparameter Sensitivity\n- Only IPC (Images Per Class) is studied in Table 4, but critical hyperparameters $\\lambda_{CC}$, $\\lambda_{S}$, $\\lambda_{Reg}$ are not systematically analyzed. Could author explain how these hyperparameters were chosen?\n- How the number of communication rounds, local epoch, and local distillation iterations decided? Why MH-pFedHNDD performs 3000 local distillation while other methods only perform 1000 local distillation iterations?\n\n### Q4. Synthetic data quality and characteristics\n\nTo better understand what the method learns:\n- **Privacy**: If a client has very few data samples for a specific class (e.g. 50), would distilled data similar to real data? Is the proposed method safe from membership attack?\n- **Visualization**: Can you provide visualizations of generated synthetic images for different clients and classes?\n- **IPC trade-offs**: \n  - What is the optimal IPC value across different settings?\n  - Why does IPC=50 perform worse than IPC=25 (76.92% vs 77.50% at α=0.02)?\n  - Is there overfitting when IPC is too large?\n  - Is large IPC leakage more privacy?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- **Well-motivated design choices**: The two-phase framework with carefully designed loss functions (Contrastive Condensation Loss for synthetic data generation and Reg Loss with universum negatives for training) is well-motivated. The visualizations in Figure 2 effectively demonstrate how these components improve feature alignment and generalization.\n\n- **Comprehensive experimental evaluation**: The paper conducts extensive experiments across multiple datasets (CIFAR-10, CIFAR-100, Tiny-ImageNet), various non-IID settings ($\\alpha$ = 0.02, 0.05, 0.1), and both homogeneous and heterogeneous model scenarios. The generalization experiments (Section 4.2) and ablation studies provide good evidence for the method's effectiveness.\n\n- **Consistent performance improvements on complex datasets**: MH-pFedHNDD shows consistent improvements over strong baselines, particularly on more complex datasets (CIFAR-100 and Tiny-ImageNet), demonstrating its effectiveness in challenging scenarios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Limited novelty in technical components**: While the integration is novel, the individual technical components are largely borrowed from existing work without significant modifications:\n  - The hypernetwork architecture directly follows MH-pFedHN (Zhang et al., 2025) without changes\n  - $\\mathcal{L}_{CC}$ is essentially SupConLoss (Khosla et al., 2020) applied to the context of synthetic data generation\n  - $\\mathcal{L}_{Reg}$ is directly based on UniConLoss (Han et al., 2022)\n  - The data distillation method uses standard Distribution Matching (Zhao & Bilen, 2023)\n\n- **Computational and communication overhead not analyzed**: This is a critical omission for practical federated learning:\n  - No report of computational cost for the distillation phase, which requires training a separate distillation model and optimizing synthetic data\n  - Communication costs not compared with baselines - while synthetic data may be compressed, the initial distillation phase and ongoing exchange add overhead\n  - No wall-clock time comparisons or convergence speed analysis (only communication rounds shown)\n  - Without this analysis, it's unclear whether the accuracy improvements justify the additional costs\n\n- **Writing quality and clarity issues**:\n  - Grammatical errors and inconsistent capitalization (e.g., \"personalized Federated learning framework based on HyperNetworks with Data Distillation\")\n  - **Step 7 in Figure 1 (Distillation Phase)** is shown in the figure but never explained in the text, leaving the framework description incomplete\n  - The \"Difference\" row in Tables 1–2 is unclear: it appears to report MH‑pFedHNDD’s gain over the second‑best method (or its gap to the best), but this should be stated explicitly in the caption or text"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925496078,"tcdate":1761929580185,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15193/Reviewer_8vK2"],"signatures":["ICLR.cc/2026/Conference/Submission15193/Reviewer_8vK2"],"forum":"etHBN33ny5","number":4,"license":"CC BY 4.0","cdate":1761929580185,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15193/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925496078,"domain":"ICLR.cc/2026/Conference","replyto":"etHBN33ny5","id":"fJEj7kvqD0","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We are the first to propose a data-driven perspective  for model heterogeneity in pFL. Our MH-pFedHNDD is the first to integrate data distillation into the pFL hypernetwork as well as better balance personalization and generalization."},"keywords":{"value":["Hypernetwork","Personalized Federated Learning","Data Distillation"]},"supplementary_material":{"value":"/attachment/a4ade09ebb19aed96027b73ca9f0656986d29aa0.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Personalized federated learning (pFL) aims to provide each client with a customized model based on global knowledge. However, in highly heterogeneous scenarios, pFL often struggles to obtain effective global information and faces a trade-off between personalization and generalization, which can degrade overall generalization performance.\nTo address this issue, we propose a **M**odel-**H**eterogeneous **p**ersonalized **Fed**erated learning framework based on **H**yper**N**etworks with **D**ata **D**istillation, **MH-pFedHNDD**, which, for the first time, incorporates data distillation into a hypernetwork-based federated learning framework, introducing a data-driven perspective to tackle this problem.\nWe design two effective regularization terms:\n(1) Contrastive Condensation Loss, which encourages the latent embeddings of synthetic data to be more compact and closely aligned with the local data of clients used as anchors;\n(2) Reg Loss, which integrates the latent embeddings of all clients’ synthetic data as anchors to guide the optimization direction for generalization, thereby enhancing each client’s personalized optimization performance on its local data along with the use of universum negatives.\nBy leveraging synthetic data distilled with more robust global information, our method enhances local training on clients, is the first to alleviate the imbalance between commonality and personalization for hypernetworks, and improves the performance and generalization of the hypernetwork. \nExtensive experiments under various settings demonstrate the effectiveness of our MH-pFedHNDD in personalized federated learning. Our code is available at \\url{https://anonymous.4open.science/r/MH-pFedHNDD}."},"_bibtex":{"value":"@misc{\nzhang2025exploring,\ntitle={Exploring Hypernetwork to Enhance Model Heterogeneous Personalized Federated Learning with Data Distillation},\nauthor={Chen Zhang and Husheng Li and Xiang Liu and Linshan Jiang and Danxin Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=etHBN33ny5}\n}"},"title":{"value":"Exploring Hypernetwork to Enhance Model Heterogeneous Personalized Federated Learning with Data Distillation"},"pdf":{"value":"/pdf/567b156e565ec94fa34cece2acb80d0c3bf957af.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|exploring_hypernetwork_to_enhance_model_heterogeneous_personalized_federated_learning_with_data_distillation"},"authorids":{"value":["~Chen_Zhang27","~Husheng_Li4","~Xiang_Liu15","~Linshan_Jiang1","~Danxin_Wang1"]},"authors":{"value":["Chen Zhang","Husheng Li","Xiang Liu","Linshan Jiang","Danxin Wang"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["benchmark","human-centric reasoning","3D spatial reasoning","physical reasoning","multimodal large language models","structured representations"]},"supplementary_material":{"value":"/attachment/cc8712e4192c5d776d79c4d4924e93552d682dcb.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Understanding how another person can move and interact with the physical world is essential to human coordination and collaboration. Yet even a simple question—can someone bend down to pick up an object without hitting a nearby table or losing balance?—requires reasoning about articulated motion, collision, and physical stability that remains challenging for multimodal large language models.\nTo study Spatio-Physics in humAn-centric Reasoning, we introduce SPARKBench, a benchmark with 5,291 questions over 1,442 images.\nIt covers two macro-level categories: 3D Spatial reasoning about pose, orientation, and relations, and Spatio-Physics reasoning about collision, articulation and kinematics, and intuitive physics. We benchmark 12 vision-language models, finding that all of these models struggle to judge body orientation and movement limits while scoring lower on counterfactual than on descriptive questions.\nTo understand how to improve human-centric spatio-physics reasoning, we compare three representations (Image, Graph, and Code) and two states, estimated from the image or taken from the rendered scene.\nWe learn two best lessons for human-centric spatio-physics reasoning:\n(1) Structured body information can improve spatial judgments, but physical reasoning also needs reliable physical state relevant to the question.\n(2) Computing the judgment from that state can improve physical reasoning when they use reliable scene geometry and physical properties.\nSPARKBench with the best lessons provides a basis for identifying model weaknesses and designing better representations and computations for human-centric spatio-physics reasoning."},"_bibtex":{"value":"@inproceedings{\nanonymous2026sparkbench,\ntitle={{SPARKB}ench: Best Lessons from Benchmarking Human-Centric Spatio-Physics Reasoning},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KtBQgF9qZv},\nnote={under review}\n}"},"title":{"value":"SPARKBench: Best Lessons from Benchmarking Human-Centric Spatio-Physics Reasoning"},"pdf":{"value":"/pdf/e26b24306a015b2054281b4651a2a092d34a0a2a.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791221979393,"tcdate":1788288429295,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission7566/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission7566/Authors"],"forum":"KtBQgF9qZv","license":"CC BY 4.0","number":7566,"cdate":1788288429295,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission7566/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791221979393,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"KtBQgF9qZv","version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (2) 2018"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-05414-4_4.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"fornaia|a_general_powerful_graph_pattern_matching_system_for_data_analysis"},"html":{"value":"https://doi.org/10.1007/978-3-030-05414-4_4"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/FornaiaMPT18,\n  author={Andrea Fornaia and Misael Mongiovì and Giuseppe Pappalardo and Emiliano Tramontana},\n  title={A General Powerful Graph Pattern Matching System for Data Analysis},\n  year={2018},\n  cdate={1514764800000},\n  pages={40-53},\n  url={https://doi.org/10.1007/978-3-030-05414-4_4},\n  booktitle={COMPLEX NETWORKS (2)},\n  crossref={conf/complexnetworks/2018-2}\n}\n"},"abstract":{"value":"Graph pattern matching is a powerful mechanism for searching on network data. Most of the graph pattern matching tools available are based on subgraph isomorphism, i.e. finding a one-to-one correspondence between nodes of a query graph and nodes of a target graph. Often this approach is not flexible enough, since it does not allow the query to represent sets of nodes of any size that share the same profile of connectivity. We propose a novel powerful graph matching approach that overcomes the existing limits and allows the user to define complex patterns in a simple and intuitive way. In our approach, queries are expressed as graphs, whose nodes and edges specify structural constraints and filtering criteria. We show that, despite its simplicity, the proposed approach can solve a large variety of practical problems."},"title":{"value":"A General Powerful Graph Pattern Matching System for Data Analysis"},"authors":{"value":[{"fullname":"Andrea Fornaia","username":""},{"fullname":"Misael Mongiovì","username":"~Misael_Mongiovì2"},{"fullname":"Giuseppe Pappalardo","username":""},{"fullname":"Emiliano Tramontana","username":""}]}},"tmdate":1784964769719,"pdate":1546214400000,"externalIds":["dblp:conf/complexnetworks/FornaiaMPT18"],"tcdate":1784964760591,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Misael_Mongiovì2"],"forum":"YtJGIvxM3e","license":"CC BY-SA 4.0","number":103959,"cdate":1514764800000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784964769719,"domain":"OpenReview.net/Public_Article","id":"YtJGIvxM3e","version":2},{"content":{"venue":{"value":"J. Complex. 1995"},"pdf":{"value":"https://www.sciencedirect.com/science/article/pii/S0885064X85710175/pdf?md5=905d0829acfa1190c9ea7f6309503d75&pid=1-s2.0-S0885064X85710175-main.pdf"},"venueid":{"value":"dblp.org/journals/JC/1995"},"paperhash":{"value":"edelman|on_the_determinant_of_a_uniformly_distributed_complex_matrix"},"authorids":{"value":["~Alan_Edelman1"]},"html":{"value":"https://doi.org/10.1006/jcom.1995.1017"},"_bibtex":{"value":"@article{DBLP:journals/jc/Edelman95,\n  author={Alan Edelman},\n  title={On the Determinant of a Uniformly Distributed Complex Matrix},\n  year={1995},\n  cdate={788918400000},\n  journal={J. Complex.},\n  volume={11},\n  number={3},\n  pages={352-357},\n  url={https://doi.org/10.1006/jcom.1995.1017}\n}\n"},"abstract":{"value":"We derive the joint density for the singular values of a random complex matrix A uniformly distributed on ||A||F = 1 This joint density allows us to obtain the conditional expectation of det(AHA) = |det A|2 given the smallest singular value. This result has been used by Shub and Smale in their analysis of the complexity of Bezout′s theorem."},"title":{"value":"On the Determinant of a Uniformly Distributed Complex Matrix"},"authors":{"value":["Alan Edelman"]}},"tmdate":1747082511610,"pdate":788918400000,"tcdate":1747081675480,"writers":["~"],"signatures":["~Alan_Edelman1"],"forum":"4EXIbDywpW","license":"CC BY-SA 4.0","number":423059,"cdate":788918400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747082511610,"domain":"DBLP.org","id":"4EXIbDywpW","version":2},{"content":{"summary":{"value":"This paper proposes to add synthetic anomalies to diversify the collected anomalies under the semi-supervised setting. Authors connect anomaly detection with binary classification and introduces synthetic anomalies to mitigate two issues in semi-supervised AD: false negative modeling and insufficient regularity of learning. Some theoretical analyses are provided to justify the effectiveness of incorporating synthetic anomalies. Experiments across tabular, image, and text datasets  demonstrate the applicability of the proposed framework."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How sensitive is the framework to the way synthetic anomalies are generated? Would a more structured generator improve performance?\n2. It is possible to generalize the theoretical guarantees to broader architectures or activation functions?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1.  Some theoretical analyses are conducted on the effectiveness of introducing synthetic anomalies for semi-supervised anomaly detection.\n\n2. The experiments span diverse modalities (tabular, image, text) and multiple AD methods, showing general applicability of the “synthetic anomaly” principle."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Formulating anomaly task as binary classification is fundamentally inappropriate. Since the type of anomaly is uncountable, anomaly detection is usually formulated as one-class classification to model the distribution of normal data or to learn the pattern of them. Using binary classifier may learn a unreliable decision boundary.\n \n2. The theoretical analysis of convergence is narrow to the network using ReLU as activation function. Extending the theoretical guarantees to broader architectures or activation functions would significantly strengthen the generality and impact of the results.\n\n3. The novelty of this paper is weak. While the theoretical framing is elegant, the core idea is adding synthetic anomalies, which is not new. The main contribution lies in extending this idea to a semi-supervised setting, which feels incremental and does not substantially push the frontier of anomaly detection research.\n\n4. Synthetic anomaly generation is overly simplistic. The use of uniformly random noise as synthetic anomalies is questionable, especially for complex or high-dimensional data. This weakens the practical significance of the framework and may not generalize to high-dimensional or structured data. There is no comparison with more informative or adaptive anomaly generation methods.\n\n5. The presented ablation resembles a sensitivity analysis rather than a comprehensive investigation.\n\n6. The writing of this work is terrible and should be significantly improvoed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931301458,"tcdate":1761968850473,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19369/Reviewer_SFnp"],"signatures":["ICLR.cc/2026/Conference/Submission19369/Reviewer_SFnp"],"forum":"W7QcymYxXK","number":2,"license":"CC BY 4.0","cdate":1761968850473,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19369/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931301458,"domain":"ICLR.cc/2026/Conference","replyto":"W7QcymYxXK","id":"LC8eM1Vx8w","forumContent":{"TLDR":{"value":"We generalize unsupervised anomaly detection to the semi-supervised setting by showing that synthetic anomalies — previously used in unsupervised AD — remain provably and empirically beneficial with limited labeled anomalies."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Statistical Learning Theory; Anomaly Detection; Supervised Learning; Classification; Synthetic Data"]},"supplementary_material":{"value":"/attachment/d8ab4b83e3de28506794a4b0fbfb6ec6444931f3.zip"},"primary_area":{"value":"learning theory"},"abstract":{"value":"Anomaly detection (AD) is a critical task across domains such as cybersecurity and healthcare. In the unsupervised setting, an effective and theoretically-grounded principle is to train classifiers to distinguish normal data from (synthetic) anomalies. We extend this principle to semi-supervised AD, where training data also include a limited labeled subset of anomalies possibly present in test time. We propose a theoretically-grounded and empirically effective framework for semi-supervised AD that combines known and synthetic anomalies during training. To analyze semi-supervised AD, we introduce the first mathematical formulation of semi-supervised AD, which generalizes unsupervised AD. Here, we show that synthetic anomalies enable (i) better anomaly modeling in low-density regions and (ii) optimal convergence guarantees for neural network classifiers — the first theoretical result for semi-supervised AD. We empirically validate our framework on five diverse benchmarks, observing consistent performance gains. These improvements also extend beyond our theoretical framework to other classification-based AD methods, validating the generalizability of the synthetic anomaly principle in AD."},"_bibtex":{"value":"@misc{\nzhou2026bridging,\ntitle={Bridging Unsupervised and Semi-Supervised Anomaly Detection: A Provable and Practical Framework with Synthetic Anomalies},\nauthor={Tian-Yi Zhou and Matthew Lau and Xiangchi Yuan and Jizhou Chen and Xiaoming Huo and Wenke Lee},\nyear={2026},\nurl={https://openreview.net/forum?id=W7QcymYxXK}\n}"},"title":{"value":"Bridging Unsupervised and Semi-Supervised Anomaly Detection: A Provable and Practical Framework with Synthetic Anomalies"},"pdf":{"value":"/pdf/965d51a98e663af8cb10cad0ece815e5df12f3b2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhou|bridging_unsupervised_and_semisupervised_anomaly_detection_a_provable_and_practical_framework_with_synthetic_anomalies"},"authorids":{"value":["~Tian-Yi_Zhou1","~Matthew_Lau1","~Xiangchi_Yuan1","~Jizhou_Chen1","~Xiaoming_Huo1","~Wenke_Lee1"]},"authors":{"value":["Tian-Yi Zhou","Matthew Lau","Xiangchi Yuan","Jizhou Chen","Xiaoming Huo","Wenke Lee"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the core bottlenecks of physical consistency and parameter controllability in state-of-the-art text-to-video (T2V) generation. Existing models rely on appearance-level motion memorization without understanding underlying physical laws, leading to unrealistic motions (e.g., upward-falling objects) and poor generalization. To solve this, the authors propose NewtonGen, a hybrid framework integrating data-driven T2V generation with physics-informed Neural Newtonian Dynamics (NND). NND leverages unified neural ordinary differential equations (ODEs) combined with a residual MLP to learn latent dynamics from small-scale \"physics-clean\" data (generated via a custom simulator). It models 9-dimensional latent physical states (position, velocity, rotation, size, etc.) to predict physics-consistent trajectories. During inference, NND’s state predictions and scene prompts guide a motion-controlled T2V model (Go-with-the-Flow) to generate videos."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please follow weakness."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Directly addresses two critical limitations of state-of-the-art text-to-video (T2V) models—physical inconsistency (e.g., upward-falling objects, abrupt velocity changes) and lack of parameter controllability—which scaling alone cannot resolve. This fills a key research gap, as current models only learn motion distributions from visual appearance rather than underlying physical laws.\n\n2. NND’s design (physics-informed linear neural ODEs + residual MLP) unifies diverse Newtonian motions (12 types tested, including uniform velocity, parabolic motion, rotation, deformation, and 3D motion) into a single framework. It efficiently learns latent dynamics from a small amount of \"physics-clean\" data, avoiding over-reliance on large-scale datasets.\n\n3. Compares against 5 representative baselines (closed-source: Sora, Veo3; open-source: CogVideoX-5B, Wan2.2; physics-based: PhysT2V) across 12 motion types, ensuring fairness by standardizing generation settings.\n\n4. Separates physical dynamics reasoning (NND, for state prediction) from video generation (motion-controlled model like Go-with-the-Flow). This design retains the visual quality of existing T2V models while injecting physical constraints and enabling easy integration with other T2V frameworks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Relies on continuous ODEs and thus cannot handle discrete physical events (e.g., collisions, rebounds, explosions). This restricts its applicability to real-world scenarios involving interactive or abrupt physical interactions.\n2.  NND is trained exclusively on data from a custom Python physics simulator (designed to avoid motion blur, noise, or background distractions). Performance on real-world videos (with clutter, occlusions, or motion blur) is unproven, as no experiments on real datasets are reported.\n3. While NewtonGen outperforms baselines on physical consistency (PIS), it does not quantitatively evaluate visual quality (e.g., texture realism, shadow dynamics, background coherence) against competitors. This makes it unclear if physical consistency comes at the cost of visual fidelity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915566771,"tcdate":1761837225262,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission615/Reviewer_UcEn"],"signatures":["ICLR.cc/2026/Conference/Submission615/Reviewer_UcEn"],"forum":"rJ6N6sunaU","number":2,"license":"CC BY 4.0","cdate":1761837225262,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission615/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915566771,"domain":"ICLR.cc/2026/Conference","replyto":"rJ6N6sunaU","id":"9jX0T4kXxD","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"NewtonGen leverages Neural Newtonian Dynamics to learn general latent dynamics, enabling physically aware and controllable Text-to-Video generation."},"keywords":{"value":["Generative Models","Video Generation"]},"supplementary_material":{"value":"/attachment/9cc815abed32c4037e71149b52140fdc9597336b.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"A primary bottleneck in large-scale text-to-video generation today is physical consistency and controllability. Despite recent advances, state-of-the-art models often produce unrealistic motions, such as objects falling upward, or abrupt changes in velocity and direction. Moreover, these models lack precise parameter control, struggling to generate physically consistent dynamics under different initial conditions. We argue that this fundamental limitation stems from current models learning motion distributions solely from appearance, while lacking an understanding of the underlying dynamics. In this work, we propose NewtonGen, a framework that integrates data-driven synthesis with learnable physical principles. At its core lies trainable Neural Newtonian Dynamics (NND), which can model and predict a variety of Newtonian motions, thereby injecting latent dynamical constraints into the video generation process. By jointly leveraging data priors and dynamical guidance, NewtonGen enables physically consistent video synthesis with precise parameter control.  All data and code are available at https://github.com/pandayuanyu/NewtonGen."},"_bibtex":{"value":"@inproceedings{\nyuan2026newtongen,\ntitle={NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics},\nauthor={Yu Yuan and Xijun Wang and Tharindu Wickremasinghe and Zeeshan Nadir and Bole Ma and Stanley H. Chan},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=rJ6N6sunaU}\n}"},"title":{"value":"NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics"},"pdf":{"value":"/pdf/77f56149d8717bd2f3e32f8d954800e0aebadb4a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|newtongen_physicsconsistent_and_controllable_texttovideo_generation_via_neural_newtonian_dynamics"},"authorids":{"value":["~Yu_Yuan4","~Xijun_Wang4","~Tharindu_Wickremasinghe1","~Zeeshan_Nadir1","~Bole_Ma2","~Stanley_H._Chan2"]},"authors":{"value":["Yu Yuan","Xijun Wang","Tharindu Wickremasinghe","Zeeshan Nadir","Bole Ma","Stanley H. Chan"]}},"version":2},{"content":{"summary":{"value":"This work explored how synthetic data affects machine learning model. They measured the distribution gap via Mahalanobis distance in feature space. They investigated about (1) the effect of synthetic data in terms of distribution gap, (2) the importance of real data, (3) the importance of sim2real transformation quality, and (4) the characteristics of the synthetic data pool."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"2 fair"},"strengths":{"value":"- Synthetic data plays an important role in the field where data acquisition is difficult. For example, synthetic data is mainly used in optical flow [A]. Since synthetic data is widely used, research is needed to understand how it is used. This work conducted it.\n- Based on Sim2Real transformation, they conducted the experiments of synthetic data in terms of the size of the real dataset, the size of the synthetic dataset, data selection methods (random, PTL), and synthetic pool comparison.\n\n[A] FlowNet: Learning Optical Flow with Convolutional Networks, ICCV 2015\n\n[B] Fake it till you make it: face analysis in the wild using synthetic data alone, ICCV, 2021"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The authors expanded the domain gap used in PTL to the dataset comparison. The expansion is incremental.\n- In Sec. 5.1, the authors commented, “…, particularly the impact of synthetic images more significant when the amount of real data is small”. However, in Fig. 2(a), |Performance of VisDrone w/ synth - Performance of VisDrone w/o synth| is larger in 200 samples.\n- In Table 2, the more real data authors use the smaller the domain gap. For example, the distribution gap is changed to 35.1 (20) → 31.7 (50) → 23.2 (100) → 19.8 (200) in the 50% column of Okutama. Why is VisDrone used more, and has the distribution gap decreased?\n- The authors commented, “random selection is generally more effective in reducing the distribution gap using synthetic data than PTL”. It seems that the measure of the distribution gap used is not sufficient to delve into Sim2Real transformation alone. To delve into the synthetic data, it is better to provide the various views using another distribution measure, the number of parameters used in learning, etc.\n- PTL shows a higher performance than random selection despite the high distribution gap, which means the model is biased toward the reference dataset. The authors said the quality of Sim2Real transformation is important. What is the quality authors said?"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636134047,"tcdate":1698576304884,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2025/Reviewer_Eyf8"],"signatures":["ICLR.cc/2024/Conference/Submission2025/Reviewer_Eyf8"],"forum":"yW0hLmwq4f","number":2,"license":"CC BY 4.0","cdate":1698576304884,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2025/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636134047,"domain":"ICLR.cc/2024/Conference","replyto":"yW0hLmwq4f","id":"sGoqR0atfv","forumContent":{"TLDR":{"value":"A comprehensive study to understand the sim2real transformation mechanism and the impact of using synthetic data in training."},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Synthetic data","sim2real transformation","distribution gap"]},"supplementary_material":{"value":"/attachment/a3580bb32fd4656fcd0a5962aabc95a89cee4a3a.pdf"},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Synthetic data has the distinct advantage of building a large-scale labeled dataset for almost free. Still, it should be carefully integrated into learning; otherwise, the expected performance gains are difficult to achieve. The biggest hurdle for synthetic data to achieve increased training performance is the domain gap with the (real) test data. As a common solution to deal with the domain gap, the sim2real transformation is used, and its quality is affected by three factors: i) the real data serving as a reference when calculating the domain gap, ii) the synthetic data chosen to avoid the transformation quality degradation, and iii) the synthetic data pool from which the synthetic data is selected. In this paper, we investigate the impact of these factors on maximizing the effectiveness of synthetic data in training in terms of improving learning performance and acquiring domain generalization ability--two main benefits expected of using synthetic data. As an evaluation metric for the second benefit, we introduce a method for measuring the distribution gap between two datasets, which is derived as the normalized sum of the Mahalanobis distances of all test data. As a result, we have discovered several important findings that have never been investigated or have been used previously without accurate understanding. We expect that these findings can break the current trend of either naively using or being hesitant to use synthetic data in machine learning due to the lack of understanding, leading to more appropriate use in future research."},"_bibtex":{"value":"@misc{\nlee2024delving,\ntitle={Delving Deep into Sim2Real Transformation: Maximizing Impact of Synthetic Data in Training},\nauthor={Hyungtae Lee and Yan Zhang and Yi-Ting Shen and Heesung Kwon and Shuvra Shikhar Bhattacharyya},\nyear={2024},\nurl={https://openreview.net/forum?id=yW0hLmwq4f}\n}"},"title":{"value":"Delving Deep into Sim2Real Transformation: Maximizing Impact of Synthetic Data in Training"},"pdf":{"value":"/pdf/b4188d0ab9c4c8a2ce8bd4bf0e839d0ae9318e07.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"lee|delving_deep_into_sim2real_transformation_maximizing_impact_of_synthetic_data_in_training"},"authorids":{"value":["~Hyungtae_Lee3","~Yan_Zhang26","~Yi-Ting_Shen1","~Heesung_Kwon2","~Shuvra_Shikhar_Bhattacharyya1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hyungtae Lee","Yan Zhang","Yi-Ting Shen","Heesung Kwon","Shuvra Shikhar Bhattacharyya"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewer for recognizing the PECS metric and the layered mechanistic analysis. We address the concerns below.\n\n**(1, 2a) Cross-architecture mechanism and family diversity.** \n\nWe extended probing and activation steering to **Gemma 4 E4B** (Google) and **LLaVA-NeXT-Video-7B** (Meta), selected for (i) open weights (required for hidden-state access), (ii) matched scale (7–8B), and (iii) organizational distinctness from Qwen3-VL-8B (Alibaba). The three families share no pre or post-training pipeline, ruling out a within-family artifact.\n\n**Paper claims tested (all three replicate across both new families):**\n1. **VLMs fail to express epistemic restraint** (low PECS).\n2. **VLMs encode the void/control distinction** (linear probing).\n3. **Single-layer activation steering causally controls abstention**.\n\n*Claim 1 (replicated): VLMs fail to express epistemic restraint.* Both new models score low PECS under guided prompting: **Gemma PECS = 0.224 ± 0.006**, **LLaVA PECS = 0.116 ± 0.005**. Both well below saturation.\n\n*Claim 2 (replicated): VLMs encode the void/control distinction.* The distinction is linearly decodable in both new models and transfers across datasets, including on confabulated-only samples (Cf→Cf, inputs where the model failed to abstain). Replicating the paper's Table 3 (LR-probe cross-dataset transfer AUROC, best layer) as **AUC / Cf→Cf**:\n\n| Transfer | Qwen3-VL-8B (paper) | Gemma 4 E4B | LLaVA-NeXT-Video-7B |\n|---|---|---|---|\n| Cross-modality (vis↔text) | 0.888 / 0.902 | 0.879 / 0.839 | 0.920 / 0.903 |\n| Cross-domain (occl↔chaotic) | 0.851 / 0.858 | 0.850 / 0.710 | 0.791 / 0.725 |\n| Cross-both (diff. domain + void) | 0.760 / 0.748 | 0.797 / 0.689 | 0.710 / 0.709 |\n\n*Claim 3 (replicated, with per-family variation): single-layer activation steering causally controls abstention.* Replicating the paper's Table 4 (single-layer void-direction steering, Guided/Guided regime, Ctrl(+α) ↑ / Void(−α) ↓ abstention rate):\n\n| Family | Ctrl(+α) | Void(−α) |\n|---|---|---|\n| Qwen3-VL-8B (paper, L20) | 11→60 | 67→27 |\n| Gemma 4 E4B (L35) | 25→57 | 78→65 |\n| LLaVA-NeXT-Video-7B (L28) | 16→50 | 52→28 |\n\n**(3i) The interaction of prompt regime with activation steering is per-family.** The paper (§5.5 Finding 1) observed that under standard-prompt inference, +α-induced abstention is gated. We investigated this interaction in both new families. Ctrl(+α) abstention rate, 2×2 direction × inference matrix (Direction = prompt the steering vector was extracted under; Inference = prompt during steered generation):\n\n| Direction / Inference | Std / Std | Std / Guided | Guided / Std | Guided / Guided |\n|---|---|---|---|---|\n| Qwen | 0→1 | 11→54 | 0→1 | 11→60 |\n| Gemma | 3→17 | 25→39 | 3→20 | 25→57 |\n| LLaVA | 13→50 | 16→50 | 13→51 | 16→50 |\n\nThe two Std-inference columns isolate the interaction. On **Qwen**, standard inference cleanly gates +α (≤1%). On **Gemma**, standard inference partially gates +α (17–20%). On **LLaVA**, standard inference does not prevent activation steering (≥50%).\n\n**Summary of claims.** All three claims replicate across the new families, confirming the headline thesis (encode but fail to express). The §5.5 Finding 1 inference-gate sub-claim under Claim 3 is per-family, and we will scope it accordingly in the revision.\n\n**(2b) Chaotic taxonomy vs. temporal truncation.** Truncation alone is not sufficient for abstention to be the correct outcome. The distinction is clearest when chaos and truncation are varied independently:\n\n1. **No chaos + truncation** (e.g. IntPhys, CLEVRER): the model should **infer** the outcome, not an abstention case.\n2. **Chaos + no truncation** (our **control**): fully resolvable; models **answer** it (Seesaw control accuracy 82.9%, Table 25).\n3. **Chaos + truncation** (our **chaotic void**): the model should **abstain**.\n\nOnly case (3) implies abstention. This is different from (1) by chaos and (2) by truncation, so the taxonomy is not merely \"missing evidence due to truncation.\" Our results supports this: the chaotic void direction is nearly orthogonal to the occlusion direction (cos = −0.05; Table 24) and domain-specific; a temporal \"missing-evidence\" signal would instead align with occlusion's. \n\nOn the sampling concern (25 frames at 5 FPS): in control scenes we verify the resolved end-state is sampled, and in test scenes we verify the video is truncated before the outcome resolves.\n\n**(3) Differentiation from IntPhys 2.** We did not cite IntPhys 2 (we cited the original IntPhys), and thank the reviewer for raising it. The key difference is the protocol: IntPhys 2 assumes a determinate \"yes/no\" answer, whereas we require models to **both** answer correctly under the **control** condition and abstain under the **occluded / chaotic / ill-posed** conditions: so the overlap lies entirely in our control pairs. In short, IntPhys 2 measures *what* models predict; TRAPSBench measures *whether* they should not predict at all. We will include IntPhys2 in our discussion."},"title":{"value":"Thank you for your detailed review. We hope to address 1. Core claims on representation x output replicated across Gemma + Llava; 2. chaotic taxonomy -> truncation + chaos = abstention. Truncation is insufficient 3. differentiation from IntPhys2 - we focus on abstention, not correctness"}},"parentInvitations":"colmweb.org/COLM/2026/Conference/-/Official_Comment","tmdate":1787538549803,"tcdate":1780270431012,"writers":["colmweb.org/COLM/2026/Conference","colmweb.org/COLM/2026/Conference/Submission285/Authors"],"signatures":["colmweb.org/COLM/2026/Conference/Submission285/Authors"],"forum":"o5cFzLUOf0","number":3,"license":"CC BY 4.0","cdate":1780270431012,"readers":["everyone"],"invitations":["colmweb.org/COLM/2026/Conference/Submission285/-/Official_Comment","colmweb.org/COLM/2026/Conference/-/Edit"],"mdate":1787538549803,"domain":"colmweb.org/COLM/2026/Conference","replyto":"GQ1NJCRrFG","id":"K1y27HxUza","forumContent":{"TLDR":{"value":"We introduce TRAPSBench: matched answerable/unanswerable physics videos. Across 16 VLMs, epistemic restraint is poor; probing and steering in three open-weight families reveal an internal uncertainty signal that is not reliably expressed."},"venue":{"value":"COLM 2026"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the COLM Code of Ethics on https://colmweb.org/CoE.html"},"LLM_usage":{"value":"In accordance with the Policy on the use of Large Language Models in the Call for Papers https://colmweb.org/cfp.html, I certify that my submission discloses any substantive use of LLMs in the research and paper-writing process."},"keywords":{"value":["vision-language models","epistemic uncertainty","abstention","physical reasoning","activation steering"]},"abstract":{"value":"When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undeterminable from the visual evidence. Furthermore, we introduce Penalized Epistemic Calibration Score (PECS), a new robust metric that requires models to both answer correctly when the outcome is knowable, and abstain when the outcome is not. Across 16 VLMs spanning five families, spontaneous restraint is poor: the best PECS is 0.292. The bottleneck is expression, not perception: linear probes decode answerability from hidden states at up to 0.91 AUROC across physics domains; steering a single-layer void direction causally induces or suppresses abstention. Our results replicate across three open-weight families (Qwen, Gemma, LLaVA). The failure is also more pronounced in visual than textual uncertainty: models detect textual impossibility about 4x more readily than missing visual evidence. Closing this representation--output gap likely requires output-stage interventions."},"_bibtex":{"value":"@inproceedings{\npramono2026trapsbench,\ntitle={{TRAPSB}ench: Vision-Language Models Encode but Fail to Express Epistemic Restraint},\nauthor={Fnu pramono and John Cai and Sourabh Kulkarni},\nbooktitle={Third Conference on Language Modeling},\nyear={2026},\nurl={https://openreview.net/forum?id=o5cFzLUOf0}\n}"},"title":{"value":"TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint"},"pdf":{"value":"/pdf/9ec7e4bcede6cc8741505be3be57fb97e3c46759.pdf"},"venueid":{"value":"colmweb.org/COLM/2026/Conference"},"paperhash":{"value":"pramono|trapsbench_visionlanguage_models_encode_but_fail_to_express_epistemic_restraint"},"authorids":{"value":["~Fnu_pramono1","~John_Cai1","~Sourabh_Kulkarni1"]},"author_guide":{"value":"I certify that this submission complies with the submission instructions as described on https://colmweb.org/AuthorGuide.html"},"authors":{"value":["Fnu pramono","John Cai","Sourabh Kulkarni"]}},"version":2},{"content":{"summary":{"value":"This paper introduces an ambitious and conceptually interesting framework for generating photonic band diagrams (BDs) from 3D photonic crystal (PhC) structures using transformer-based latent diffusion models. The approach represents a clear methodological advance over prior deep learning surrogates that relied on CNNs or VAEs, as it leverages transformers to capture long-range electromagnetic couplings and diffusion models to synthesize fine-grained spectral details. The proposed Material-to-Context (M2C) encoder, ViT-based BD autoencoder, and conditional latent diffusion model together form a scalable architecture with the potential to drastically reduce the computational cost of BD prediction—reportedly achieving up to 900× speedup compared to rigorous coupled-wave analysis (RCWA) simulations.\n\nHowever, the experimental evidence is less convincing than the conceptual innovation. The reported quantitative metrics (Dice ≈ 0.37 on the small dataset and ≈ 0.23 on the large one) indicate limited fidelity, and the generated BDs, while visually plausible, often miss important spectral features. In photonic design, such small spectral mismatches can correspond to large physical deviations, so it remains unclear whether the model’s outputs are reliable enough for practical use. The paper acknowledges these limitations but does not offer concrete solutions, such as uncertainty quantification, physics-informed regularization, or calibration, to mitigate or measure these discrepancies."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Could you provide quantitative comparisons against standard baselines such as CNNs, VAEs, or U-Nets used in previous photonic band prediction works? This would clarify whether the diffusion-based approach offers a real advantage beyond architectural novelty.\n\nHow well does the model generalize to more complex 3D photonic structures beyond the synthetic stacked-layer dataset? Have you tested any out-of-distribution geometries or permittivity ranges to support your scalability claim?\n\nThe reported Dice and mAP scores are relatively low. Could you analyze where the errors come from—for example, from the transformer encoding, latent diffusion stage, or reconstruction process?\n\nGiven that small spectral deviations can have large physical impacts, have you considered adding uncertainty estimation or confidence measures to assess the reliability of generated band diagrams?\n\nYou mention physics-informed strategies as future work. Could you outline more concretely how such priors or constraints might be incorporated into your diffusion pipeline?\n\nFinally, do you view this model primarily as a fast surrogate for exploration or as a physically accurate predictor for design optimization? Clarifying this distinction would help position the contribution more precisely."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper’s main strength lies in its novel application of latent diffusion models to the problem of photonic band diagram (BD) generation, a domain where traditional solvers like RCWA are computationally intensive. The authors creatively combine transformer encoders and diffusion-based generative modeling, demonstrating how recent advances in generative AI can be repurposed for physics-driven simulation tasks. This cross-disciplinary approach is original and timely, bridging modern deep learning architectures with complex photonic design problems.\n\nConceptually, using a diffusion process to synthesize physically meaningful spectra represents a significant methodological innovation, offering the potential to replace thousands of numerical Maxwell solves with a single generative inference step. The proposed pipeline, especially the Material-to-Context (M2C) transformer, is thoughtfully designed and technically sound, highlighting an insightful understanding of how self-attention mechanisms can capture multi-layer optical coupling effects."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the conceptual innovation is clear, the paper’s experimental and analytical depth falls short of its stated ambition. The most significant limitation is the lack of comparison against strong baselines. The authors mention prior CNN-, VAE-, and U-Net–based models for photonic band or dispersion prediction but do not provide direct quantitative comparisons to these architectures. Without such baselines, it is difficult to assess whether the proposed diffusion-transformer framework offers a tangible performance advantage, or whether the observed results could be matched by simpler models.\n\n2. The paper also provides limited discussion on generalization. Although the authors claim that the framework can be scaled to arbitrary 3D photonic structures, all experiments are confined to synthetic, highly regular datasets built from stacked holey and uniform layers. There is no evidence that the model can handle more complex geometries, continuous permittivity variations, or realistic fabrication noise—conditions essential for true generalizability.\n\n3. Another weakness lies in the lack of critical analysis of the model’s underperformance. Reported metrics (Dice ≈ 0.37 on the small dataset, ≈ 0.23 on the large one) and visual comparison are substantially lower than what would constitute reliable physical prediction, yet the paper does not explore the causes. For instance, it remains unclear whether errors arise from the diffusion process, the latent space compression, or the transformer conditioning. A deeper ablation study or error decomposition would have clarified the model’s limitations and helped guide future improvements.\n\n4. Finally, while the authors briefly mention possible remedies—such as larger encoders, physics-informed priors, or ensemble sampling—these ideas are presented only as speculation rather than experimentally supported strategies. Overall, the work would benefit greatly from stronger empirical validation, explicit baseline comparison, and systematic error analysis, which would make its claims of physical relevance more convincing and actionable."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926836206,"tcdate":1761988699285,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16805/Reviewer_zhRf"],"signatures":["ICLR.cc/2026/Conference/Submission16805/Reviewer_zhRf"],"forum":"hLpJjDmFlQ","number":2,"license":"CC BY 4.0","cdate":1761988699285,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16805/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926836206,"domain":"ICLR.cc/2026/Conference","replyto":"hLpJjDmFlQ","id":"fkVo7DyvS2","forumContent":{"TLDR":{"value":"We propose a new method using latent diffusion models and transformers to generate band diagrams from 3D photonic crystals."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Photonic crystals","band diagrams","diffusion","transformers","surrogate models"]},"supplementary_material":{"value":"/attachment/441580fd9c5ae7aa704bc63803600b13d347a93f.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Photonic crystals enable fine control over light propagation at the nanoscale, and thus play a central role in the development of photonic and quantum technologies. Photonic band diagrams (BDs) are a key tool to investigate light propagation into such inhomogeneous structured materials. However, computing BDs requires solving Maxwell’s equations across many configurations, making it numerically expensive, especially when embedded in optimization loops for inverse design techniques, for example. To address this challenge, we introduce the first approach for BD generation based on diffusion models, with the capacity to later generalize and scale to arbitrary three-dimensional structures. This preliminary study couples a transformer encoder, which extracts contextual embeddings from the input structure, with a latent diffusion model to generate the corresponding BD. In addition, we provide insights into why transformers and diffusion models are well suited to capture the complex interference and scattering phenomena inherent to photonics. This cross-disciplinary approach is bridging modern deep learning architectures with complex photonic design problems, paving the way for new surrogate modeling strategies in this domain."},"_bibtex":{"value":"@misc{\ndelchevalerie2026towards,\ntitle={Towards Photonic Band Diagram Generation with Transformer-Latent Diffusion Models},\nauthor={Valentin Delchevalerie and Nicolas Roy and Arnaud Bougaham and Alexandre Mayer and Benoit Frenay and Micha{\\\"e}l Lobet},\nyear={2026},\nurl={https://openreview.net/forum?id=hLpJjDmFlQ}\n}"},"title":{"value":"Towards Photonic Band Diagram Generation with Transformer-Latent Diffusion Models"},"pdf":{"value":"/pdf/7b37cd9874a62413c6a6baea7e66e322ef56503b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"delchevalerie|towards_photonic_band_diagram_generation_with_transformerlatent_diffusion_models"},"authorids":{"value":["~Valentin_Delchevalerie1","~Nicolas_Roy1","~Arnaud_Bougaham1","~Alexandre_Mayer1","~Benoit_Frenay1","~Michaël_Lobet1"]},"authors":{"value":["Valentin Delchevalerie","Nicolas Roy","Arnaud Bougaham","Alexandre Mayer","Benoit Frenay","Michaël Lobet"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PRISM-PHYSICS, a benchmark and a process-level evaluation framework that encodes physics solutions as DAGs and employs rule-based symbolic equivalence checking for reliable, fine-grained scoring."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Refer to the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. A large-scale benchmark of competition-level physics problems with carefully curated, DAG-structured solutions. \n2. A DAG-based scoring policy that explicitly models causal dependencies among formulas, enabling fine-grained and interpretable process-level evaluation.\n3. A fully rule-based symbolic formula equivalence checker to reliably validate diverse mathematical expressions, ensuring consistent comparison across alternative formulations and eliminating reliance on heuristic LLM-as-judge scoring."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. My main concern is that Figures 1 through 5 are very unclear, and even when enlarged twice, they are still hard to read. These figures should ideally be the most direct representation of the data analysis in this study, PRISM-PHYSICS. I hope the authors can improve the clarity of these figures.\n2. Figure 1 appears on page 2, but there is no corresponding content on the first two pages. Is its placement here inappropriate? Additionally, Figure 1 is not referenced anywhere in the text. The same issue applies to Figure 2.\n3. I feel that the originality of the article is somewhat limited. Could the authors provide an explanation of the connection between the challenges presented in this study and the research methods? At the moment, the challenges and methods do not seem to align well.\n4. Regarding the analysis of the experimental section, I hope the authors can summarize it more clearly. The current summary of the experiments is not very clear.\n5. How was the step level notation done? I would appreciate clarification on this point."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360378315,"tcdate":1761359121975,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9900/Reviewer_4yRy"],"signatures":["ICLR.cc/2026/Conference/Submission9900/Reviewer_4yRy"],"forum":"4PZMeopXzP","number":1,"license":"CC BY 4.0","cdate":1761359121975,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9900/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360378315,"domain":"ICLR.cc/2026/Conference","replyto":"4PZMeopXzP","id":"gKfT4qqLPF","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We present PRISM-Physics, a benchmark and a process-level evaluation framework that encodes physics solutions as DAGs and employs rule-based symbolic equivalence checking for reliable, fine-grained scoring."},"keywords":{"value":["Physics Reasoning","Process-Level Evaluation","Symbolic Equivalence","Scientific Problem Solving"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Benchmarks for competition-style reasoning have advanced evaluation in mathematics and programming, yet physics remains comparatively underexplored. Most existing physics benchmarks evaluate only final answers, which fail to capture reasoning processes, while recent stepwise methods rely on heuristic LLM-as-judge scoring or restrictive linear assumptions, limiting reliability and diagnostic validity.\nWe introduce PRISM-Physics, a process-level evaluation framework and benchmark for complex physics reasoning problems. Solutions are represented as directed acyclic graphs (DAGs) of formulas, explicitly encoding causal dependencies among intermediate steps to enable fine-grained, interpretable, and theoretically grounded scoring. \nWe prove the optimality of the DAG representation and the corresponding scoring policy. Combining with a fully rule-based method for symbolic formula equivalence matching that we developed, we ensure consistent validation across diverse formulations without heuristic judgments. Results show that our evaluation framework is more aligned with human experts' scoring. \nExperiments on state-of-the-art LLMs reveal persistent reasoning failures in physics, while step-level scoring offers both diagnostic insight and rich signals for later training. By combining structural rigor, theoretical guarantees, and symbolic validation, PRISM-Physics provides a principled foundation for advancing process-level evaluation and guiding the development of models with deeper scientific reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\nzhao2026prismphysics,\ntitle={{PRISM}-Physics: Causal {DAG}-Based Process Evaluation for Physics Reasoning},\nauthor={Wanjia Zhao and Qinwei Ma and Jingzhe Shi and Shirley Wu and Jiaqi Han and Yijia Xiao and Si-Yuan Chen and Xiao Luo and Ludwig Schmidt and James Zou},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=4PZMeopXzP}\n}"},"title":{"value":"PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning"},"pdf":{"value":"/pdf/95b751e95b88437e4484cd3de1a315d0a89884f4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhao|prismphysics_causal_dagbased_process_evaluation_for_physics_reasoning"},"authorids":{"value":["~Wanjia_Zhao1","~Qinwei_Ma1","~Jingzhe_Shi1","~Shirley_Wu1","~Jiaqi_Han2","~Yijia_Xiao1","~Si-Yuan_Chen1","~Xiao_Luo3","~Ludwig_Schmidt1","~James_Zou1"]},"authors":{"value":["Wanjia Zhao","Qinwei Ma","Jingzhe Shi","Shirley Wu","Jiaqi Han","Yijia Xiao","Si-Yuan Chen","Xiao Luo","Ludwig Schmidt","James Zou"]}},"version":2},{"content":{"TLDR":{"value":"Synthetic spacecraft imagery alone fails to transfer to real observations, but hybrid training with web imagery improves detection and yields more semantically aligned internal representations."},"venue":{"value":"CVPR 2026 Workshop SynData4CV"},"pdf":{"value":"/pdf/422d473936a8c9c0ff4137cdc76232786f7a6ef8.pdf"},"keywords":{"value":["Synthetic data generation","Sim-to-real transfer","Spacecraft perception","Object detection","Representation learning","Space domain vision","Digital twin environment"]},"venueid":{"value":"thecvf.com/CVPR/2026/Workshop/SynData4CV"},"paperhash":{"value":"issitt|assessing_the_predictive_value_of_physicsgrounded_synthetic_data_for_computer_vision_in_space_environments","readers":["everyone"]},"authorids":{"readers":["everyone"],"value":["~Arianna_Issitt1","~Emily_Happy1","~Elijah_Clark1","~Mackenzie_J._Meni1","~Ryan_T._White1"]},"abstract":{"value":"Synthetic imagery is widely used to train space-domain vision systems for space applications because real orbital imagery is scarce and expensive to collect, yet it remains unclear which aspects of physical realism in synthetic data generation influence sim-to-real generalization. In this work, we investigate the predictive value of physics-grounded synthetic data using a Digital Twin Environment (DTE) that generates spacecraft imagery under controlled orbital geometry while varying illumination and surface reflectance modeling. Rather than optimizing downstream models, we fix the learning pipeline and use detection performance as a probe of data quality. Models trained under different synthetic regimes are evaluated on a web-scraped spacecraft imagery dataset and on a real laboratory dataset of an unseen spacecraft mockup under spacelike lighting. Synthetic-only training fails to transfer to real imagery, achieving near-zero detection and segmentation accuracy on the laboratory dataset. However, when synthetic imagery is used to augment web-collected training imagery, performance improves substantially, with the best model reaching 0.724 mAP@0.5 on the laboratory dataset, improving upon previously reported performance on this dataset (0.394 mAP@0.5). To interpret these differences, we analyze internal feature representations using PEEK-based information statistics and find that synthetic-only training concentrates representational variance near the backbone terminus, whereas models trained with real-data involvement maintain more distributed deep representations. Rotation estimation experiments using Structure-from-Motion reconstruction further show that physics-grounded rendering is necessary when synthetic imagery is intended to emulate real-world observations, as heuristic appearance models fail to preserve feature correspondences required by motion-estimation pipelines. These results provide insight into how synthetic data influences representation learning and offer guidance for designing synthetic data pipelines for space-domain vision systems."},"title":{"value":"Assessing the Predictive Value of Physics-Grounded Synthetic Data for Computer Vision in Space Environments"},"authors":{"readers":["everyone"],"value":["Arianna Issitt","Emily Happy","Elijah Clark","Mackenzie J. Meni","Ryan T. White"]}},"tmdate":1782159693105,"tcdate":1773363994977,"writers":["thecvf.com/CVPR/2026/Workshop/SynData4CV","thecvf.com/CVPR/2026/Workshop/SynData4CV/Submission36/Authors"],"signatures":["thecvf.com/CVPR/2026/Workshop/SynData4CV/Submission36/Authors"],"forum":"CklXJ7AiAb","license":"CC BY 4.0","number":36,"cdate":1773363994977,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Workshop/SynData4CV/-/Submission","thecvf.com/CVPR/2026/Workshop/SynData4CV/-/Post_Submission","thecvf.com/CVPR/2026/Workshop/SynData4CV/-/Edit"],"mdate":1782159693105,"domain":"thecvf.com/CVPR/2026/Workshop/SynData4CV","id":"CklXJ7AiAb","version":2},{"content":{"summary":{"value":"This paper proposes a method for generating synthetic microservice workload traces by fine-tuning large language models (LLMs) to replicate complex, hierarchical microservice call graphs. Using a recursive generation approach, the model creates layers of the graph step-by-step, preserving structural constraints to ensure realistic trace outputs. Additional instruction tuning enhances the model's ability to follow specific trace requirements and generate uncommon scenarios, making the synthetic traces suitable as substitutes for real-world data in downstream tasks like anomaly detection. The method is shown to outperform existing generative approaches, offering a promising solution for environments where real trace data is limited or privacy-sensitive."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Given that the recursive approach discards previously generated layers, how does this affect the model's performance in scenarios where long-term dependencies across multiple layers are critical? Could the authors provide additional insights or results on the effects of layer depth and sequence length on trace validity?\n\nThe paper mentions the use of hand-crafted templates for instruction tuning. Have the authors considered experimenting with automatically generated or dynamically adapted instructions to better match a broader range of conditions? Could this approach improve model generalization?\n\nHave the authors considered or conducted tests of the model in real-world environments? While the model performs well on simulated data, deploying it in live production systems might reveal additional insights about practical constraints, resource usage, and unexpected errors.\n\nThe results suggest a performance drop with increased trace complexity. Can the authors clarify the main factors contributing to this decline? Is it due to limited model capacity, recursive generation limitations, or instruction tuning?\n\nInstruction tuning, especially with user-specific attributes, might raise concerns about inadvertently learning or generating sensitive patterns. Did the authors take any measures to ensure privacy during training? Are there risks of exposing specific information, especially if used in sensitive environments?\n\nThe paper focuses on microservice call graphs, but this method may be adaptable to other hierarchical data. Do the authors foresee limitations in applying this approach to other domains, such as healthcare workflows or complex IoT interactions?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"Innovative Use of LLMs for Trace Generation: The proposed use of large language models (LLMs) for generating realistic microservice workload traces, specifically microservice call graphs, is promising, leveraging recent advancements in language models to solve challenges related to data availability and privacy in system performance analysis.\n\nRecursive Generation of Complex Graph Structures: The proposed method addresses the hierarchical and recursive nature of call graphs, which many synthetic data generation techniques struggle to capture. By breaking down call graph generation into recursive layers, the model can better manage complexity and maintain structural constraints, which is essential for realistic trace generation.\n\nInstruction Tuning for Enhanced Trace Validity: The authors incorporate instruction tuning, adding intermediate reasoning steps that enforce trace constraints. This strengthens the model's ability to follow user specifications and generate traces that respect the structural dependencies within microservice architectures, such as start and finish times in hierarchical call graphs.\n\nDemonstrated Practical Application: The paper shows that synthetically generated traces can effectively replace real traces in downstream tasks, including microservice management tasks like critical component extraction and anomaly detection. This is a valuable contribution, as it shows that synthetic data can potentially reduce reliance on sensitive real-world data while still maintaining performance.\n\nComprehensive Evaluation and Benchmarks: The authors conduct thorough experiments that compare the recursive and instruction-tuned model against baselines, such as probabilistic models and other generative methods like GANs and VAEs. They provide performance metrics across various tasks (e.g., trace generation, infilling, and downstream prediction tasks), giving a clear view of the model's strengths and limitations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited Exploration of Long-Term Dependencies: While the recursive approach is effective in handling hierarchical structures, it does not retain previously generated layers or edges in memory. This may limit the model's ability to capture long-term dependencies or sequential behaviors across more complex, larger traces, which could be important in certain applications involving extensive trace histories.\n\nReliance on Manually Constructed Instruction Templates: The use of hand-crafted templates for instruction tuning may restrict the model’s adaptability and efficiency. Generating diverse instruction formats, possibly with automated methods, might improve the model’s generalization and reduce the manual effort required to customize instructions for different scenarios.\n\nPerformance Trade-Offs in High Complexity Scenarios: The model demonstrates a drop in accuracy with increasing complexity, particularly with greater numbers of edges and layers in the call graphs. Although the recursive approach and instruction tuning mitigate some of this complexity, the performance could suffer in very large or deeply nested call graphs, limiting its effectiveness for the most demanding or detailed traces.\n\nSynthetic Data Quality Compared to Real Data: Despite the high accuracy and validity of synthetic traces for downstream tasks, the paper notes a slight performance drop when using synthetic traces compared to real data in certain tasks. This suggests that, while synthetic traces are a viable alternative, they may still not fully replicate the nuances of real-world data.\n\nLack of Real-World Deployment Evidence: While the model is tested in a simulated environment and evaluated on trace data from a known dataset (Alibaba v2022), there is limited evidence of its performance in a live production environment. Real-world deployment could reveal additional challenges or performance limitations, especially under more varied or unforeseen conditions.\n\nPotential Privacy Risks with Instruction-Tuned Models: The instruction tuning approach, while effective, could raise privacy concerns if used with highly specific instructions or in contexts where sensitive information may inadvertently be reflected in the generation outputs."}},"nonreaders":[],"tmdate":1731428417182,"tcdate":1730415559907,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12738/Reviewer_Yg3a"],"signatures":["ICLR.cc/2025/Conference/Submission12738/Reviewer_Yg3a"],"forum":"f9GURUHZQo","number":1,"license":"CC BY 4.0","cdate":1730415559907,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12738/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428417182,"domain":"ICLR.cc/2025/Conference","replyto":"f9GURUHZQo","id":"1htjkjHMKx","forumContent":{"TLDR":{"value":"We train a language model to generate synthetic computer system traces, specifically microservice call graphs."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic data","synthetic trace","microservice","large language model","machine learning for systems"]},"supplementary_material":{"value":"/attachment/a3ce86bc05387e93c1d4440cda9c2150d9103a3b.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Computer system workload traces, which record hardware or software events during application execution, are essential for understanding the behavior of complex systems and managing their processing and memory resources. However, obtaining real-world traces can be challenging due to the significant collection overheads in performance and privacy concerns that arise in proprietary systems. As a result, synthetic trace generation is considered a promising alternative to using traces collected in real-world production deployments. This paper proposes to train a large language model (LLM) to generate synthetic workload traces, specifically microservice call graphs. To capture complex and arbitrary hierarchical structures and implicit constraints in such traces, we fine-tune LLMs to generate each layer recursively, making call graph generation a sequence of easier steps. To further enforce learning constraints in traces and generate uncommon situations, we apply additional instruction tuning steps to align our model with the desired trace features. Our evaluation results show that our model can generate diverse realistic traces under various conditions and outperform existing methods in accuracy and validity. We show that our synthetically generated traces can effectively substitute real-world data in optimizing or tuning systems management tasks. We also show that our model can be adapted to perform key downstream trace-related tasks, specifically, predicting key trace features and infilling missing data given partial traces."},"_bibtex":{"value":"@misc{\nkim2025large,\ntitle={Large Language Models as Realistic Microservice Trace Generators},\nauthor={Donghyun Kim and Sriram Ravula and Taemin Ha and Alex Dimakis and Daehyeok Kim and Aditya Akella},\nyear={2025},\nurl={https://openreview.net/forum?id=f9GURUHZQo}\n}"},"title":{"value":"Large Language Models as Realistic Microservice Trace Generators"},"pdf":{"value":"/pdf/4b031907890b6c901f041bd1fb2a704a57090bad.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"kim|large_language_models_as_realistic_microservice_trace_generators"},"authorids":{"value":["~Donghyun_Kim12","~Sriram_Ravula1","~Taemin_Ha1","~Alex_Dimakis1","~Daehyeok_Kim1","~Aditya_Akella1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Donghyun Kim","Sriram Ravula","Taemin Ha","Alex Dimakis","Daehyeok Kim","Aditya Akella"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PhyGenBench, a new benchmark for evaluating the physical commonsense capabilities of Text-to-Video (T2V) models, particularly their understanding of intuitive physics. It includes prompts across 27 physical laws within four domains: mechanics, optics, thermal, and material properties. To evaluate performs on this benchmark, the authors propose PhyGenEval, a hierarchical evaluation framework using advanced vision-language models (VLMs) and large language models. Experimental results reveal that current T2V models lack robust physical commonsense, underscoring the gap between these models and true world simulators."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- What is \"the final score is calculated as 0 according to 4.2\" (line 292)? Is the example in Figure receive 0 after this physical commonsense evaluation?\n\n- It seems like the entire evaluation rely on closed sourced LLM: GPT-4o. If in the future, GPT-4o becomes unavailable, how should people compare results?\n\n- some typos such as we pue more detailed (line 410), Appendix C.2 (line 418)"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"- This paper clearly have great novelty. It focus on intuitive physics is unique and addresses an important gap in T2V evaluation.\n- PhyGenEval's three tiered framework (key phenomena detection, order verification, and overall naturalness) thoroughly assesses physical realism.\n- By getting more attention on the gap in physical commonsense, the benchmark provides great insights on how to improve video generation models to become a real world simulator."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper includes extensive comparisons to demonstrate PhyGenEval’s effectiveness, suggesting that a two-stage evaluation strategy may align more closely with human judgments for both InternVideo2 and GPT-4-o. Line 965 also notes that alternative open-source models achieve a high correlation coefficient with human evaluations. However, it appears that the main results rely on a specific version of GPT-4-o, which is not explicitly mentioned. As a benchmark, would future users need to evaluate all baselines and methods on updated versions of GPT-4-o to ensure fair comparisons? While the paper suggests that evaluation costs are minimal, I am concerned that this reliance on a specific model version may affect consistency. Have the authors considered using other LVLMs in place of GPT-4-o?\n- Certain T2I models may perform poorly on specific prompts. I am not fully convinced that the proposed evaluation method can robustly handle these lower-quality videos.\n- The issue of hallucination in large language models (LLMs) does not appear to be addressed in the evaluation protocol, potentially impacting the reliability of the benchmark. It would be beneficial if the authors considered this factor in their assessment framework.\n\n- The author promised more human evaluation results in Appendix C.2 but this result seems under Appendix C.1. The writing seems to be confusing. Also between line 899 and 905, I believe the annotation should be done more rigorously. I am expecting carefully validate results from human annotators or I think the results can be noisy. I think showing the instructions to the human annotators can be particularly helpful."}},"nonreaders":[],"tmdate":1731427583962,"tcdate":1730602697625,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2363/Reviewer_QELL"],"signatures":["ICLR.cc/2025/Conference/Submission2363/Reviewer_QELL"],"forum":"6rMHcLWxl4","number":2,"license":"CC BY 4.0","cdate":1730602697625,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2363/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427583962,"domain":"ICLR.cc/2025/Conference","replyto":"6rMHcLWxl4","id":"vx1rOjGZXG","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["World Simulator","Physical Commonsense","Video Generation","Evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGenBench, a comprehensive \\textbf{Phy}sics \\textbf{Gen}eration \\textbf{Ben}chmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench, we propose a novel evaluation framework called PhyGenEval. This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through PhyGenBench and PhyGenEval, we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by PhyGenBench (e.g., dynamic scenarios). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We will release the data and codes at https://github.com/PhyGenBench/PhyGenBench"},"_bibtex":{"value":"@misc{\nmeng2025towards,\ntitle={Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation},\nauthor={Fanqing Meng and Jiaqi Liao and Xinyu Tan and Wenqi Shao and Quanfeng Lu and Kaipeng Zhang and Yu Cheng and Dianqi Li and Yu Qiao and Ping Luo},\nyear={2025},\nurl={https://openreview.net/forum?id=6rMHcLWxl4}\n}"},"title":{"value":"Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation"},"pdf":{"value":"/pdf/1814f0c3473ab9a04aca4edcd8aab3e678055bdd.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"meng|towards_world_simulator_crafting_physical_commonsensebased_benchmark_for_video_generation"},"authorids":{"value":["~Fanqing_Meng1","~Jiaqi_Liao2","~Xinyu_Tan1","~Wenqi_Shao2","~Quanfeng_Lu1","~Kaipeng_Zhang1","~Yu_Cheng1","~Dianqi_Li1","~Yu_Qiao1","~Ping_Luo2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fanqing Meng","Jiaqi Liao","Xinyu Tan","Wenqi Shao","Quanfeng Lu","Kaipeng Zhang","Yu Cheng","Dianqi Li","Yu Qiao","Ping Luo"]}},"version":2},{"content":{"summary":{"value":"The paper proposes AWML, a unified framework that (i) learns structured latent world models with modular priors, (ii) generates certified counterfactual examples by recombining learned modules, (iii) accepts only those synthetics whose uncertainty is below a calibrated threshold, and (iv) decomposes priors into transferable and mutable components to enable adaptive transfer across environments. The authors provide a suite of finite‑sample generalization, modular amplification, and certified augmentation bounds (Theorems 3.1–3.6), a transfer bound (Theorem 3.7) and a greedy exploration guarantee (Theorem 3.8). Empirically, AWML is evaluated on a synthetic modular AR(1) environment and on the Uganda LSMS survey for a binary electrification prediction task. The experiments aim to validate modular amplification, uncertainty filtering and the overall AWML pipeline."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"# Questions\n\n## Module Estimation\nHow do you estimate the per‑module total variation bounds $\\delta_m$ in practice?\nWhat diagnostics do you use to ensure that the factorization assumption holds?\n\n## Calibration Constant $L$\nWhat estimator did you use for the calibration constant in Theorem 3.6?\nDid you validate that the estimated $L$ is close to the true bound on held‑out data?\n\n## Hyperparameter Selection\nDo you have a principled way to choose the acceptance threshold $u$ beyond cross‑validation?\n\n## Scalability\nWhat is the computational cost (time, memory) of modular recombination in high‑dimensional domains?\nDo you use any pruning or sampling strategies to keep the synthetic pool tractable?\n\n## Baselines\nHave you compared AWML against meta‑learning (e.g., MAML), self‑supervised pretraining, or active learning baselines on the LSMS dataset? Can you provide also a qualitative discussion of expected performance differences?\n\n## Active Exploration\nThe paper claims support for “information‑gain driven acquisition”.\nDid you evaluate the greedy exploration guarantee (Theorem 3.8) on any real or synthetic tasks?\n\n## Real‑World Domains\nBeyond LSMS, have you tested AWML on domains where structured world models are natural (robotics, physical simulation)?\nIf so, can you share results? How do you envision the method scaling to such settings?\n\n## Uncertainty Estimation\nWhich uncertainty estimator did you use (ensemble variance, conformal scores, etc.) in LSMS?\nHow robust is the filtering to mis‑calibrated uncertainty scores?\n\n## Future Work\nDo you plan to relax the modular independence assumption (e.g., via hierarchical or graph‑structured modules)?\nHow would you integrate AWML with continual learning or online adaptation?\n\n# Suggestions\n1. Introduce a high‑level diagram of the AWML pipeline early in the paper. Show how factual data → encoder $\\phi$ → modular dynamics $p_\\theta^{(m)}$ → synthetic generation via recombination → uncertainty filter → augmented dataset → downstream predictor.\n\n2. Simplify notation in the main text. Move heavy definitions (e.g., $\\mathcal{H}_P$, Rademacher complexity) to a notation table or appendix; use more descriptive variable names.\n\n3. Include a comparison table with at least three baseline data‑efficient methods (MAML, self‑supervised pretraining, active learning) on the LSMS dataset. Even if the baselines are simple, they provide a sanity check.\n\n4. Show runtime and memory for generating $B_{\\max}$ synthetic samples in a realistic setting (e.g., 1000‑dimensional image features). This will help readers assess practicality.\n\n5. Clarify the role of the “adaptive transfer”: explain how transferable vs. mutable priors are identified and used in practice."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"## Novel combination of ideas\nThe paper unifies several research threads: structured latent world models, counterfactual data augmentation, uncertainty‑aware acceptance, and adaptive transfer in a single coherent framework.\n\n##Finite‑sample theory\nThe authors derive explicit bounds that relate modular estimation error, synthetic sample size, and bias, providing a principled way to trade off variance reduction against augmentation bias.\n\n## Certified counterfactuals\nBy explicitly bounding the total variation between synthetic and factual distributions, the method offers an auditable safety guarantee that is rarely addressed in data‑augmentation work.\n\n## Empirical validation\nSynthetic experiments demonstrate the expected scaling of error with effective sample size; LSMS results show tangible AUC gains in a low‑label regime.\n\n## Clear structure\nThe manuscript is logically organized (intro → contributions → theory → experiments → appendix)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"## Unverified Assumptions\nThe modular factorization and per‑module TV bounds $\\delta_m$ are essential to the theoretical guarantees, yet authors do not report how these bounds are estimated or validated.  Synthetic experiments vary $N_{\\text{eff}}$ while holding $\\delta_m$ fixed, and no sensitivity analysis is provided.\n\n## Scalability\nThe algorithmic loop in Appendix A.11 suggests drawing up to $B_{\\max}$ synthetic samples and filtering them, yet the paper does not report runtime or memory usage.  In high‑dimensional domains (images), modular recombination would explode combinatorially.\n\n## Empirical baseline comparison\nThe experiments compare AWML only against a factual‑only baseline. No state‑of‑the‑art data‑efficient methods (meta‑learning, self‑supervised pretraining, active learning) are benchmarked.\n\n## Parameter tuning\nKey hyperparameters (number of modules $M$, parent sets $pa(m)$, threshold $u$, calibration constant $L$) are not systematically studied. The paper offers limited guidance on how to choose them in practice.\n\n## Complexity of exposition\nThe notation is heavy, and the main narrative often jumps between equations without intuitive explanations. A diagram of the AWML pipeline would help readers grasp the flow.\n\n## Real‑world impact\nBeyond LSMS, there is no demonstration on other realistic domains (robotics, medical imaging) where structured world models and counterfactuals are crucial.\n\n## Uncertainty estimation\nThe calibration constant $L$ is treated as a black‑box estimator, but the paper does not provide experimental evidence that the estimated $L$ is close to the true bound.\n\n## Active exploration\nThe claim that AWML supports “information‑gain driven acquisition” is mentioned but not evaluated."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922431770,"tcdate":1761760050978,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11282/Reviewer_zZf9"],"signatures":["ICLR.cc/2026/Conference/Submission11282/Reviewer_zZf9"],"forum":"IOKftyz5iP","number":2,"license":"CC BY 4.0","cdate":1761760050978,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11282/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922431770,"domain":"ICLR.cc/2026/Conference","replyto":"IOKftyz5iP","id":"ojoW9sd8ZC","forumContent":{"TLDR":{"value":"AWML is a framework that certifiably augments data via modular recombination and uncertainty filtering, yielding provable sample-efficiency gains—safely “making more data” when real data are scarce."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Adaptive world models","Structured Latent Dynamics","Modular causal modeling","Neural operators","Counterfactual recombination"]},"supplementary_material":{"value":"/attachment/60418fd34d72950b8d5da2ab3864661b9b6835e9.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"We propose Adaptive World Models for Data-Efficient Learning (AWML), a framework that combines structured latent world models, counterfactual augmentation, and calibrated uncertainty filtering to improve learning under scarce labels. AWML learns modular latent dynamics from domain priors, generates counterfactuals by recombining learned modules, and accepts synthetic samples only when a calibrated uncertainty score verifies their reliability.\n\nOur analysis shows that modular recombination improves estimation rates while limiting bias and that calibrated filtering controls the deviation between synthetic and factual distributions. Together, these results yield a clear excess-risk guarantee that captures the bias–variance trade-off through the acceptance threshold and effective sample size.\n\nAWML is instantiated with neural-operator backbones, modular causal blocks, and safeguards such as ensemble calibration and diagnostic audit flags.\n\nEmpirically, AWML reduces error in synthetic autoregressive tasks and achieves substantial gains in real low-label conditions; on the Uganda LSMS 2019 survey, AUC rises from 0.8797 to 0.9402 with only 25 labels.\n\nOverall, AWML provides a principled and practical approach to safe, data-efficient learning."},"_bibtex":{"value":"@misc{\nkatende2025adaptive,\ntitle={Adaptive World Models for Data-Efficient Learning ({AWML})},\nauthor={Ronald Katende},\nyear={2025},\nurl={https://openreview.net/forum?id=IOKftyz5iP}\n}"},"title":{"value":"Adaptive World Models for Data-Efficient Learning (AWML)"},"pdf":{"value":"/pdf/e208b57d83e7b1ca1fe91fc1a7a308c238e77821.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"katende|adaptive_world_models_for_dataefficient_learning_awml"},"authorids":{"value":["~Ronald_Katende1"]},"authors":{"value":["Ronald Katende"]}},"version":2},{"content":{"summary":{"value":"This work tackles the task of spatiotemporal forecasting for PDEs governed by different unknown physics; each trajectory is thus described by a specific equation or set of coefficients. To build a neural solver able to generalize to trajectories described by different physics, the authors advocate the use of in-context learning to adapt a cheap neural ODE solver via a transformed-based hypernetwork. The hypernetwork uses preceding T successive states for adapting the forecaster network to each unknown physics. The method is evaluated on a wide range of datasets and performs better or is competitive with existing baselines. It is also has been adapted to new physics via fine-tuning."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"To me, when I think of ICL in text, arbitrary sized examples can be given to adapt the language model to a new task. Can your method be applied to arbitrary sized sequences, i.e., can the hypernetwork adapt the integrated network given N past states, where N can vary at inference ?\n\nIn the umap visualisation, at epoch 0, the points seem to be already very well clustered. How do you explain that?\n\nI am bit concerned with the choices of the datasets. First, qualitatively, it is hard to see very big changes from t=0 to t=32. But more importantly, the designed task is to adapt the network to different unknown physics. It is not clearly stated in the paper. How much different physics are present ? Are the pde coefficients used during training the same during inference or test is done on new coefficients?\n\nConcerning the NRMSE results, you presented results at specific time-steps. Do the results correspond to NRMSE for that specic frame at time T or for average score from time 0 to T? It would have been nice to include an average score over the full trajectory."},"rating":{"value":3},"details_of_ethics_concerns":{"value":"No ethics review needed."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper is well written and easy to follow throughout and the motivations well explained.\n\nThe tackled task is important. Most of the times, machine learning approaches for solving PDE consider that a large number of trajectories are available for one single PDE equation with fixed parameters. In practical scenarios, multiple and unknown physics are expected, leading to very different dynamical behaviors, which need to be captured by a single neural network.\n\nThe forecaster network is more interpretable than existing data-driven approaches via a neural ODE using a simple CNN. The authors connect it to classical numerical methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Novelty: \n- the technical contribution is relatively low. Adapting neural ODE-solver to different physics has already been explored. The use of hypernetworks for adapting neural networks has been well studied in the meta-learning litterature. it has also been used for dynamical systems to condition a neural-ode like solver [1]. The differences with [1] is 1) that environments are supposed known, each environment describing a specific physic, while here, the environments would correspond to the first T states. 2) the use of ICL against weights updates.\n- icl works on quantized tokens in text, its use for continuous vectors and physical data is not straightforward. This has not been really well introduced in the paper. The authors notably justify their approach is close to [2], which employs a transformer approach for learning multiple physics using past states as input. It does not imply that [2] has ICL properties to me, as stated in the paper. \n\nEvaluation:\n- empirical results: while the method is competitive to AViT, it should have been compared to more related works. The method is presented as a meta-learning framework and comparison with respect to these approaches are thus important and should be included, to justify the strength of ICL compared to existing meta-learning strategies [1, 3, 4, 5, 6].\n\nChoice of title:\n- the title is not really appropriate: It sounds like the network is able to adapt to new physics (new equations) without fine-tuning, only by explicitly giving a new context, but the network needs to be fine-tuned as some layers of the network are specific to each dataset. I don't really understand why the authors decided to train a network on different PDEs simultaneously for leveraging the capacities of ICL, if fine-tuning is still needed for unseen PDEs.\n\n[1] CoDA - Kirchmeyer et al., Generalizing to New Physical Systems via Context-Informed Dynamics Model. ICML, 2022.\n\n[2] MPP - McCabe et al., Multiple Physics Pretraining for Physical Surrogate Models. 2024, https://arxiv.org/pdf/2310.02994.\n\n[3] DyAd - Wang et al. 2022, Meta-Learning Dynamics Forecasting Using Task Inference, NeurIPS, 2022.\n\n[4] CAVIA - Zintgraf et al., Fast Context Adaptation via Meta-Learning. ICML, 2019.\n\n[5] CAMEL - Blanke et al., Interpretable Meta-Learning of Physical Systems, ICLR, 2024.\n\n[6] FOCA - Park et al., First-order Context-based Adaptation for Generalizing to New Dynamical Systems, 2023, https://openreview.net/pdf?id=AW0i0lOhzqJ."}},"nonreaders":[],"tmdate":1731428386354,"tcdate":1730472838242,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12097/Reviewer_f5UB"],"signatures":["ICLR.cc/2025/Conference/Submission12097/Reviewer_f5UB"],"forum":"fzZfju8y0g","number":3,"license":"CC BY 4.0","cdate":1730472838242,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12097/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428386354,"domain":"ICLR.cc/2025/Conference","replyto":"fzZfju8y0g","id":"djZBmgQVv2","forumContent":{"TLDR":{"value":"We propose \"in-context neural PDE,\" which learns from previous states the parameters to feed into a neural solver in order to predict the next state of a PDE."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Spatio-temporal prediction","PDEs","in-context learning","neural solvers"]},"supplementary_material":{"value":"/attachment/8cffca9e734cf65f0f6cf7446c87b3a090428fa0.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We address the problem of predicting the next state of a dynamical system governed by *unknown* temporal partial differential equations (PDEs) using limited time-lapse data. While transformers offer a natural solution to this task through in-context learning, the inductive bias of temporal PDEs suggests a more tailored and effective approach. Specifically, when the underlying temporal PDE is fully known, classical numerical solvers can evolve the state with only a few parameters. Building on this observation, we introduce a large transformer-based hypernetwork that processes successive states to generate parameters for a much smaller neural ODE-like solver, which then predicts the next state through time integration. This framework, termed as *in-context neural PDE*, decouples parameter estimation from state prediction, offering closer alignment with classical numerical methods for improved interpretability while preserving the in-context learning capabilities of transformers. \nNumerical experiments on diverse physical datasets demonstrate that our method outperforms standard transformer-based models, reducing sample complexity and improving generalization, making it an efficient and scalable approach for spatiotemporal prediction in complex physical systems."},"_bibtex":{"value":"@misc{\nmorel2025incontext,\ntitle={In-Context Neural {PDE}: Learning to Adapt a Neural Solver to Different Physics},\nauthor={Rudy Morel and Jiequn Han and Edouard Oyallon},\nyear={2025},\nurl={https://openreview.net/forum?id=fzZfju8y0g}\n}"},"title":{"value":"In-Context Neural PDE: Learning to Adapt a Neural Solver to Different Physics"},"pdf":{"value":"/pdf/c3f4ab893f824ac81b801af2514bb6410864d961.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"morel|incontext_neural_pde_learning_to_adapt_a_neural_solver_to_different_physics"},"authorids":{"value":["~Rudy_Morel1","~Jiequn_Han1","~Edouard_Oyallon1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rudy Morel","Jiequn Han","Edouard Oyallon"]}},"version":2},{"content":{"summary":{"value":"The paper TempoPFN addresses the task of zero-shot time series forecasting, which is an importance task in the time series domain, and recently getting popularity. The paper has 2 main contributions. It trains linear RNNs, recently proposed in literature, only on synthetic data and evaluate on a exhaustive benchmark GIFT-Eval. The paper proposes a pipeline to generate a diverse set of synthetic data."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See weaknesses."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1. Presentation of the paper is good.\n2. The synthetic data generation pipeline is pretty exhaustive, and covers many types of synthetic data.\n3. The author(s) promise to open source the codes and pipelines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"However, the paper has a few weaknesses.\n1. The novelty of the paper is limited. First, linear RNNs are not new, they are adopted from the literature. Training models purely on synthetic data is not new; the paper just creates a more diverse set of synthetic data. Given the idea of TabPFN or ForecastPFN, the proposed work can be tried quite trivially, without much complications.\n2. The performance of the model is not up to the mark and marginally better than TabPFN-TS in CRPS and (from 0.544 to 0.536), but worse in MASE (0.771 to 0.797) after including significantly more and diverse set of synthetic data. This questions the efficacy and strength of the proposed work. The method is significantly worse than other methods that use real data alongwith synthetic data. \n3. \"We developed several new generators to fill gaps in existing approaches and capture specific temporal behaviors.\" -- There are infinite ways to generate time series, why only these types? Is there any real motivation behind the time series generation?\n4. Comment: The paper title is TEMPOPFN where PFN stands for Prior Data Fitted Networks. The paper should discuss PFNs (at least briefly) for the readers who are unaware about it. I did not find any mention of PFNs in the methodology."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926740244,"tcdate":1761927741059,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16687/Reviewer_8sXE"],"signatures":["ICLR.cc/2026/Conference/Submission16687/Reviewer_8sXE"],"forum":"WcEbBJeqQ0","number":4,"license":"CC BY 4.0","cdate":1761927741059,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16687/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926740244,"domain":"ICLR.cc/2026/Conference","replyto":"WcEbBJeqQ0","id":"buqBqJxfwE","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["time series forecasting","RNNs","synthetic data","mamba","linear rnn"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Foundation models for zero-shot time series forecasting face challenges in efficient long-horizon prediction and reproducibility, with existing synthetic-only approaches underperforming on challenging benchmarks. This paper presents TempoPFN, a univariate time series foundation model based on linear Recurrent Neural Networks (RNNs) pre-trained exclusively on synthetic data. The model uses a GatedDeltaProduct architecture with state-weaving for fully parallelizable training across sequence lengths, eliminating the need for windowing or summarization techniques while maintaining robust temporal state-tracking. Our comprehensive synthetic data pipeline unifies diverse generators including stochastic differential equations, Gaussian processes, and audio synthesis with novel augmentations such as time-varying TSMixup, differentiation, and integration. In zero-shot evaluations on the Gift-Eval benchmark, TempoPFN achieves state-of-the-art performance, matching models trained on real-world data while being significantly more efficient than existing baselines. We open-source our complete data generation pipeline and training code."},"_bibtex":{"value":"@misc{\nmoroshan2026tempopfn,\ntitle={Tempo{PFN}: Synthetic Pre-training of Linear {RNN}s for Zero-shot Time Series Forecasting},\nauthor={Vladyslav Moroshan and Julien Siems and Arber Zela and Timur Carstensen and Frank Hutter},\nyear={2026},\nurl={https://openreview.net/forum?id=WcEbBJeqQ0}\n}"},"title":{"value":"TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting"},"pdf":{"value":"/pdf/c1874d8823a787844f901b1d1aba9a394536051f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"moroshan|tempopfn_synthetic_pretraining_of_linear_rnns_for_zeroshot_time_series_forecasting"},"authorids":{"value":["~Vladyslav_Moroshan1","~Julien_Siems1","~Arber_Zela1","~Timur_Carstensen1","~Frank_Hutter1"]},"authors":{"value":["Vladyslav Moroshan","Julien Siems","Arber Zela","Timur Carstensen","Frank Hutter"]}},"version":2},{"content":{"summary":{"value":"This paper presents HGS, a method for simplifying the structure of mechanistic neural ODEs. The authors point out that the mechanistic models often used may be too complex for the small datasets we typically have in fields like healthcare, which can lead to overfitting. Their proposed solution includes three steps: first to modify the model's graph structure based on its topology by collapsing cycles and adding shortcuts, and then to use regularization to learn which connections to prune. The method is evaluated on synthetic data and a real world glucose forecasting task."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"Please address the weaknesses above."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The work is well motivated. Trying to bridge the gap between complex mechanistic models built by experts and data-driven methods is a critical research area, and the paper does a good job of framing the problem.\n2. The idea to first perform a structural simplification of the graph before applying a data-driven pruning is quite clever. It's a nice way to inject some domain agnostic heuristics to constrain the learning problem.\n3. The set of experiments seems comprehensive. The authors have benchmarked their method against a good range of strong baselines, and the ablation study clearly shows that each part of their pipeline contributes to the final result."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My main concerns are with the real-world application, and I would need the authors to address these points before I could reconsider my score.\n\n1. My primary concern is how the model handles patient variability. The T1DEXI dataset includes 105 different people, and it's a physiological fact that glucose dynamics differ significantly between individuals. From my reading of Equation 1, the MNODE learns a single set of dynamics for everyone, i.e., an \"average patient\" model. This seems to sidestep an important challenge in this domain. Could you clarify if this is indeed the case, and if so, explain the rationale behind this choice? To be clear, this has been discussed extensively in recent years ([1-4] for example).\n2. The paper mentions that the cross validation splits were created from a \"random permutation\" of the 342 time series. To avoid data leakage, it's standard practice in clinical ML to split data at the patient level (i.e., all data from a single person stays in one fold). Could you confirm whether your evaluation followed this practice? If not, the reported performance might not accurately reflect generalization to new, unseen patients.\n3. The paper uses the term MNODE, but it seems that the functional forms of the UVA-Padova model are discarded, with only the graph structure being retained. The actual dynamics are then learned by MLPs. This is a very weak form of mechanistic prior. It would be helpful to discuss the trade-offs here.\n4. It's worth noting that in the synthetic experiments, the TCN model starts to outperform HGS on the RMSE metric as the sample size grows to N=1000. While HGS remains more robust in terms of Peak RMSE, this suggests that its strong inductive bias might become a disadvantage when more data is available. Could you comment on this trade-off?\n\n[1] Generative ODE Modeling with Known Unknowns\n\n[2] Physics-Integrated Variational Autoencoders for Robust and Interpretable Generative Modeling\n\n[3] Learning Physics Constrained Dynamics Using Autoencoders\n\n[4] CONFIDE: Contextual Finite Difference Modelling of PDEs"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359504868,"tcdate":1760968622117,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13562/Reviewer_5nbh"],"signatures":["ICLR.cc/2026/Conference/Submission13562/Reviewer_5nbh"],"forum":"QBzFrjEF59","number":1,"license":"CC BY 4.0","cdate":1760968622117,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13562/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359504868,"domain":"ICLR.cc/2026/Conference","replyto":"QBzFrjEF59","id":"Oj5gt8R9Q4","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Predictive Sparsity","Hybrid Neural ODE","Group LASSO","Glucose Prediction"]},"supplementary_material":{"value":"/attachment/15a9c79c1213533ad31355922d0f0c51a56cf5b2.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Hybrid neural ordinary differential equations (neural ODEs) integrate mechanistic models with neural ODEs, offering strong inductive bias and flexibility, and are particularly advantageous in data-scarce healthcare settings. However, excessive latent states and interactions from mechanistic models can lead to training inefficiency and over-fitting, limiting practical effectiveness of hybrid neural ODEs. In response, we propose a new hybrid pipeline for automatic state selection and structure optimization in mechanistic neural ODEs, combining domain-informed graph modifications with data-driven regularization to sparsify the model for improving predictive performance and stability while retaining mechanistic plausibility. Experiments on synthetic and real-world data show improved predictive performance and robustness with desired sparsity, establishing an effective solution for hybrid model reduction in healthcare applications."},"_bibtex":{"value":"@inproceedings{\nzou2026automatic,\ntitle={Automatic and Structure-Aware Sparsification of Hybrid Neural {ODE}s with Application to Glucose Prediction},\nauthor={Bob Junyi Zou and Lu Tian},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=QBzFrjEF59}\n}"},"title":{"value":"Automatic and Structure-Aware Sparsification of Hybrid Neural ODEs with Application to Glucose Prediction"},"pdf":{"value":"/pdf/29fb984a8d9a1cba8595f511bb42e2de6204479b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zou|automatic_and_structureaware_sparsification_of_hybrid_neural_odes_with_application_to_glucose_prediction"},"authorids":{"value":["~Bob_Junyi_Zou1","~Lu_Tian4"]},"authors":{"value":["Bob Junyi Zou","Lu Tian"]}},"version":2},{"content":{"venue":{"value":"NeurIPS 2023"},"pdf":{"value":"https://proceedings.neurips.cc/paper_files/paper/2023/file/bc943cd038a5531d5433b1431c822c01-Paper-Conference.pdf"},"venueid":{"value":"dblp.org/conf/NIPS/2023"},"paperhash":{"value":"ren|insactor_instructiondriven_physicsbased_characters"},"authorids":{"value":["~Jiawei_Ren1","https://dblp.org/search/pid/api?q=author:Mingyuan_Zhang:","https://dblp.org/search/pid/api?q=author:Cunjun_Yu:","~Xiao_Ma2","~Liang_Pan2","https://dblp.org/search/pid/api?q=author:Ziwei_Liu_0002:"]},"html":{"value":"http://papers.nips.cc/paper_files/paper/2023/hash/bc943cd038a5531d5433b1431c822c01-Abstract-Conference.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/nips/RenZYMPL23,\n  author={Jiawei Ren and Mingyuan Zhang and Cunjun Yu and Xiao Ma and Liang Pan and Ziwei Liu},\n  title={InsActor: Instruction-driven Physics-based Characters},\n  year={2023},\n  cdate={1672531200000},\n  url={http://papers.nips.cc/paper_files/paper/2023/hash/bc943cd038a5531d5433b1431c822c01-Abstract-Conference.html},\n  booktitle={NeurIPS},\n  crossref={conf/nips/2023}\n}\n"},"abstract":{"value":"Generating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animations that reflect high-level human instructions remains a difficult problem due to the complexity of physical environments and the richness of human language. In this paper, we present $\\textbf{InsActor}$, a principled generative framework that leverages recent advancements in diffusion-based human motion models to produce instruction-driven animations of physics-based characters.Our framework empowers InsActor to capture complex relationships between high-level human instructions and character motions by employing diffusion policies for flexibly conditioned motion planning.To overcome invalid states and infeasible state transitions in planned motions, InsActor discovers low-level skills and maps plans to latent skill sequences in a compact latent space. Extensive experiments demonstrate that InsActor achieves state-of-the-art results on various tasks, including instruction-driven motion generation and instruction-driven waypoint heading. Notably, the ability of InsActor to generate physically simulated animations using high-level human instructions makes it a valuable tool, particularly in executing long-horizon tasks with a rich set of instructions. Our project page is available at [jiawei-ren.github.io/projects/insactor/index.html](https://jiawei-ren.github.io/projects/insactor/index.html)"},"title":{"value":"InsActor: Instruction-driven Physics-based Characters"},"authors":{"value":["Jiawei Ren","Mingyuan Zhang","Cunjun Yu","Xiao Ma","Liang Pan","Ziwei Liu"]}},"tmdate":1743245844634,"pdate":1672531200000,"tcdate":1727692171557,"writers":["~"],"signatures":["~Liang_Pan2"],"forum":"Mip6pmPZ47","license":"CC BY-SA 4.0","number":109941,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1743245844634,"domain":"DBLP.org","id":"Mip6pmPZ47","version":2},{"content":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Physics-based Animation; Human Motion Generation"]},"supplementary_material":{"value":"/attachment/a49a89bbc360367c9de8d7d999d042f7be8de7ed.zip"},"_bibtex":{"value":"@inproceedings{\nren2023insactor,\ntitle={InsActor: Instruction-driven Physics-based Characters},\nauthor={Jiawei Ren and Mingyuan Zhang and Cunjun Yu and Xiao Ma and Liang Pan and Ziwei Liu},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=hXevuspQnX}\n}"},"title":{"value":"InsActor: Instruction-driven Physics-based Characters"},"paperhash":{"value":"ren|insactor_instructiondriven_physicsbased_characters"},"TLDR":{"value":"We generate motions for physically-based characters from human instructions."},"abstract":{"value":"Generating animation of physics-based characters with intuitive control has long been a desirable task with numerous applications. However, generating physically simulated animations that reflect high-level human instructions remains a difficult problem due to the complexity of physical environments and the richness of human language. \nIn this paper, we present $\\textbf{InsActor}$, a principled generative framework that leverages recent advancements in diffusion-based human motion models to produce instruction-driven animations of physics-based characters.\nOur framework empowers InsActor to capture complex relationships between high-level human instructions and character motions by employing diffusion policies for flexibly conditioned motion planning.\nTo overcome invalid states and infeasible state transitions in planned motions, InsActor discovers low-level skills and maps plans to latent skill sequences in a compact latent space. \nExtensive experiments demonstrate that InsActor achieves state-of-the-art results on various tasks, including instruction-driven motion generation and instruction-driven waypoint heading. Notably, the ability of InsActor to generate physically simulated animations using high-level human instructions makes it a valuable tool, particularly in executing long-horizon tasks with a rich set of instructions. Our project page is available at [jiawei-ren.github.io/projects/insactor/index.html](https://jiawei-ren.github.io/projects/insactor/index.html)"},"pdf":{"value":"/pdf/451a0328421292fd32d77f3ee70ec51fbec5546f.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Jiawei_Ren1","~Mingyuan_Zhang1","~Cunjun_Yu1","~Xiao_Ma2","~Liang_Pan2","~Ziwei_Liu1"]},"authors":{"value":["Jiawei Ren","Mingyuan Zhang","Cunjun Yu","Xiao Ma","Liang Pan","Ziwei Liu"]}},"tmdate":1698949699415,"pdate":1695325757960,"tcdate":1683025789249,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1407/Authors"],"signatures":["NeurIPS.cc/2023/Conference/Submission1407/Authors"],"forum":"hXevuspQnX","number":1407,"cdate":1683025789249,"mdate":1698949699415,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/-/Submission","NeurIPS.cc/2023/Conference/-/Post_Submission","NeurIPS.cc/2023/Conference/Submission1407/-/Revision","NeurIPS.cc/2023/Conference/Submission1407/-/Supplementary_Material_Revision","NeurIPS.cc/2023/Conference/-/Edit","NeurIPS.cc/2023/Conference/Submission1407/-/Camera_Ready_Revision"],"odate":1698949699399,"domain":"NeurIPS.cc/2023/Conference","id":"hXevuspQnX","version":2},{"content":{"summary":{"value":"The paper proposes GOLA (Graph-based Operator Learning with Attention), an operator-learning framework for PDEs that works on irregular domains and sparse, nonuniform samples. GOLA first embeds inputs with a learnable Fourier encoder that projects function values at arbitrary coordinates into a spectral basis with complex-valued modes, then performs attention-enhanced message passing on a proximity graph built from spatial points to capture both local and global dependencies. Experiments on four 2D PDE families—Darcy Flow, Advection, Eikonal, and Nonlinear Diffusion—show consistent gains over baselines (DeepONet, AFNO, GKN)."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. **Stronger, more representative experiments.**\n   Could you add evaluations on *truly* irregular domains (unstructured/anisotropic meshes and complex boundaries), e.g., airfoil/cylinder flows or CAD-like geometries, and compare against more recent GNN/INR/Transformer-based operators designed for such settings? This would directly address external validity and positioning (see Weaknesses).\n\n2. **Fourier encoding on irregular meshes.**\n   Please clarify why the proposed Fourier encoder is applicable on nonuniform point sets:\n\n   * Is the encoding purely a function of coordinates (i.e., independent of grid regularity), or does its validity rely on the fact that your samples come from an underlying uniform lattice?\n   * How does aliasing/spectral leakage behave on truly unstructured meshes where Nyquist-style guarantees don’t hold?\n   * Can the approach extend to 3D unstructured meshes (e.g., tetrahedral/hexahedral) and curved boundaries, and what modifications (basis choice, graph construction, positional encodings) would be required?"},"rating":{"value":2},"details_of_ethics_concerns":{"value":"No ethics concerns."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1. **Learnable Fourier encoder on spatial graphs.**\n   Turning inputs into a complex Fourier basis lets the model capture long-range/global interactions with few coefficients, while graph message passing refines local structure. Because of the learned frequencies, the model can generalize across resolutions (the spectral representation is not tied to a fixed grid). This also reduces aliasing artifacts compared with naïve coordinate MLPs and gives a compact, physics-plausible feature space for operator learning.\n\n2. **Clear and sufficiently detailed method presentation.**\n   The paper lays out the pipeline cleanly—Fourier encoding → graph construction → attention-augmented message passing → decoding—with consistent notation and design choices explained (edge features, attention rationale, training objective). The ablations and component descriptions make it straightforward to reproduce and to understand which parts drive gains.\n\n3. **Helpful visualizations.**\n   The figures (architecture block diagram, graph construction sketches, and qualitative field reconstructions) make the workflow and design choices easier to follow and provide intuitive evidence for how attention and spectral features influence predictions. These plots also highlight resolution/generalization behavior and error patterns, aiding interpretability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Outdated or incomplete comparisons on irregular-domain PDE learning.**\n   The related work and experimental baselines underrepresent recent approaches across **GNN** [1], **implicit neural representations (INR/SIREN/coordinate networks)** [2], and **Transformer-style neural operators** [3] tailored to unstructured meshes or scattered points. Without head-to-head evaluations against stronger and more recent models (e.g., graph/mesh operators with positional encodings, attention-based operators on point clouds), it’s hard to substantiate GOLA’s claimed advantage. This limits the paper’s positioning and makes the empirical novelty less compelling.\n\n2. **“Irregular domains” are synthetically derived from uniform grids, weakening the motivation.**\n   The paper’s “irregular sampling” is obtained by subsampling a uniform lattice, which **does not reflect** the practical challenges that motivate irregular discretizations: complex or curved boundaries, anisotropic resolution to capture fine features, mesh adaptivity, or true unstructured meshes common in CFD and geometry-rich PDEs (e.g., airfoils, cylinder flows, ShapeNet-like CAD geometries). As a result, the experiments do not test boundary handling, mesh heterogeneity, or topology changes—key factors for demonstrating real-world utility. The focus on **only 2D** further narrows external validity; many operator-learning applications (fluid/solid mechanics, climate/ocean) require robust performance on 3D unstructured meshes.\n\n\n[1] Li, Zhihao, et al. \"Harnessing scale and physics: A multi-graph neural operator framework for pdes on arbitrary geometries.\" Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1. 2025.\n\n[2] Wang, Honghui, Shiji Song, and Gao Huang. \"GridMix: Exploring Spatial Modulation for Neural Fields in PDE Modeling.\" The Thirteenth International Conference on Learning Representations. 2025.\n\n[3] Wu, Haixu, et al. \"Transolver: A Fast Transformer Solver for PDEs on General Geometries.\" International Conference on Machine Learning. PMLR, 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921144268,"tcdate":1761272707328,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9600/Reviewer_Q32t"],"signatures":["ICLR.cc/2026/Conference/Submission9600/Reviewer_Q32t"],"forum":"fJZCNGRxFF","number":2,"license":"CC BY 4.0","cdate":1761272707328,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9600/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921144268,"domain":"ICLR.cc/2026/Conference","replyto":"fJZCNGRxFF","id":"BNEUXv3JFE","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["PDEs","Neural Operator","AI for science","Scientific Machine Learning","Operator Learning","Graph Neural Networks"]},"supplementary_material":{"value":"/attachment/acccde82db7b4d4864b616d6e7850ed5b3df6c48.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Operator learning seeks to approximate mappings from input functions to output solutions, particularly in the context of partial differential equations (PDEs). While recent advances such as DeepONet and Fourier Neural Operator (FNO) have demonstrated strong performance, they often rely on regular grid discretizations, limiting their applicability to complex or irregular domains. In this work, we propose a Graph-based Operator Learning with Attention (GOLA) framework that addresses this limitation by constructing graphs from irregularly sampled spatial points and leveraging attention-enhanced Graph Neural Netwoks (GNNs) to model spatial dependencies with global information. To improve the expressive capacity, we introduce a Fourier-based encoder that projects input functions into a frequency space using learnable complex coefficients, allowing for flexible embeddings even with sparse or nonuniform samples. We evaluated our approach across a range of 2D PDEs, including Darcy Flow, Advection, Eikonal, and Nonlinear Diffusion, under varying sampling densities. Our method consistently outperforms baselines, particularly in data-scarce regimes, demonstrating strong generalization and efficiency on irregular domains."},"_bibtex":{"value":"@misc{\nli2026graphbased,\ntitle={Graph-Based Operator Learning from Limited Data on Irregular Domains},\nauthor={Yile Li and Shandian Zhe},\nyear={2026},\nurl={https://openreview.net/forum?id=fJZCNGRxFF}\n}"},"title":{"value":"Graph-Based Operator Learning from Limited Data on Irregular Domains"},"pdf":{"value":"/pdf/3cf28682c933ff5c711ee2e700f990b717073d9b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|graphbased_operator_learning_from_limited_data_on_irregular_domains"},"authorids":{"value":["~Yile_Li1","~Shandian_Zhe1"]},"authors":{"value":["Yile Li","Shandian Zhe"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a physics-informed neural network for radial phase retrieval that combines optical priors with a manifold-based design. The model enforces strict radial symmetry by projecting inputs onto a quotient manifold, which removes global and angular phase ambiguities. It uses the nonlinear Schrödinger equation (NLSE) to model high-frequency wave propagation and the transport-of-intensity equation (TIE) to capture low-frequency behavior. These physics-based components are integrated within a learned neural framework. This resulted in more accurate and stable phase reconstructions than prior approaches."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Can the authors provide more intuition of the main ideas? The manuscript is difficult to follow, and it would help to see more discussion and diagrams of the manifold projection, the dual PDE branches, and the overall data flow.\n\n2. How sensitive is the model to deviations from perfect radial symmetry? For example, would mild astigmatism in the input significantly degrade performance?\n\n3. Are the generalization claims really valid? Simple extrapolation from 1-3 rings to 4-9 rings is not sufficient to claim generalizability."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"* Physics-informed architecture: Integrates optical physics (NLSE and TIE) directly into the neural network design, grounding the model in physical principles rather than purely data-driven fitting.\n\n* Compared to other approaches, such as U-Net and FNO, the proposed method achieves more accurate and stable phase reconstructions\n* The method has a strong theoretical grounding, and this helps with the interpretability \n\n* The method performances has some generalizablity and appears to be robust to noise."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The paper is difficult to follow. Main ideas and intuitions are not clearly stated. Several sections rely on dense mathematical notation (e.g., Hankel transforms, Lipschitz bounds, quotient manifolds) with limited intuitive explanation or guidance for the reader. The connection between the theoretical guarantees and the implemented network is unclear, and the paper quickly dives into architectural details (Figure 1) without first establishing sufficient conceptual context.\n\n* The results are shown on synthetic images only. There is no validation on real experiments or physical hardware setups. Such experiments could help determine whether the strong symmetry assumption holds. The method relies on strict radial symmetry that may not be valid under optical aberrations or misalignment.\n\n* How does the method compare against classical methods such as Gerchberg–Saxton or multi-plane TIE solvers? Only a few learned method results are shown. Are these the only methods available for comparison? I suggest the authors expand their comparison baseline to include well-established classical approaches.\n\n* Despite the complexity of the proposed method, in terms of core reconstruction, the results are very close to those of a U-Net. I question if the added complexity is justified by the improvements.\n\n* Some of the ablation results are unexpected, and the takeaways from each experiment are not clearly stated. For example, in Table 6, the performance on clean and noisy measurements is almost identical. The authors should discuss why this occurs and specify the noise level used in the experiments. In Table 5, I notice negative delta values, which indicate improvement when certain modules are removed. Does this mean that those components may not be beneficial? These points are not discussed or analyzed in the paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925477904,"tcdate":1761897906970,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15167/Reviewer_411b"],"signatures":["ICLR.cc/2026/Conference/Submission15167/Reviewer_411b"],"forum":"jS3EKPSaAR","number":2,"license":"CC BY 4.0","cdate":1761897906970,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15167/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925477904,"domain":"ICLR.cc/2026/Conference","replyto":"jS3EKPSaAR","id":"e3lz5ySCQe","forumContent":{"TLDR":{"value":"a pde based physic informed neural network that performs well in generalization"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["phase retrieval;PDE networks; outer-ring extrapolation; inverse problems"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Phase retrieval from intensity-only measurements is severely ill-posed due to global-gauge and rotational symmetries. We consider outer-ring generalization: training with supervision from only a few inner rings and testing the model’s ability to reconstruct a broader set of unseen outer rings. We introduce a physics-informed hybrid network that combines (i) radial priors encoded by a smooth exponentiated spline and a \\emph{monotone} outer-radius booster, (ii) two differentiable PDE branches---a Strang-split Kerr--NLSE pathway for high-frequency synthesis and a TIE-based low-pass pathway for coarse structure---and (iii) a strict radial projection enforcing output symmetry, together with a radius-dependent $\\alpha$-fusion. Across the tested configurations, when trained only on a few rings (1-3), our model reconstructs more rings(4-9) than conventional methods, and achieves better stability in peak\npositions and amplitude calibration under out-of-distribution settings. This provides some inspiration for enhancing the generalization of physics-informed neural networks when applied to optical inverse problems. Ablations isolate the contribution of the alpha fusion, PDE coupling, and monotone\nboosting. We will release pseudo-code to facilitate reproducibility."},"_bibtex":{"value":"@misc{\nyao2026physicsinformed,\ntitle={{PHYSICS}-{INFORMED} {RADIAL} {PHASE} {RETRIEVAL} {NEURAL} {NETWORK} {WITH} {HYBRID} {DEEP} {PRIORS} {AND} {DUAL} {PDE}},\nauthor={ZIYONG YAO},\nyear={2026},\nurl={https://openreview.net/forum?id=jS3EKPSaAR}\n}"},"title":{"value":"PHYSICS-INFORMED RADIAL PHASE RETRIEVAL NEURAL NETWORK WITH HYBRID DEEP PRIORS AND DUAL PDE"},"pdf":{"value":"/pdf/d0ee395178c85a98450cfb5c748c70d7faeb7393.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yao|physicsinformed_radial_phase_retrieval_neural_network_with_hybrid_deep_priors_and_dual_pde"},"authorids":{"value":["~ZIYONG_YAO1"]},"authors":{"value":["ZIYONG YAO"]}},"version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Feature Selection","FDR Control","Hypothesis Testing; Gaussian Mirror; Boosting Power"]},"primary_area":{"value":"probabilistic_methods"},"abstract":{"value":"Recent advances in false discovery rate (FDR)-controlled feature selection methods have improved reliability by effectively limiting false positives, making them well-suited for complex applications. A popular FDR-controlled framework called data splitting uses the \"mirror statistics\" to select features. However, we find that the unit variance assumption on mirror statistics could potentially limit the feature selection power. To address this, we generalize the mirror statistics in the Gaussian mirror framework and introduce a new approach called \"generalized Gaussian mirror\" ($\\text{G}^2\\text{M}$), which adaptively learns the variance and forms new test statistics.  We demonstrate both theoretically and empirically that the proposed test statistics achieve higher power than those of Gaussian mirror and data splitting. Comparisons with other FDR-controlled frameworks on synthetic, semi-synthetic, and real datasets highlight the superior performance of the $\\text{G}^2\\text{M}$ method in achieving higher power while maintaining FDR control. These findings suggest the potential for the $\\text{G}^2\\text{M}$ method for practical applications in real-world problems. Code is available in https://github.com/skyve2012/G2M."},"_bibtex":{"value":"@inproceedings{\nshen2025textgtextm,\ntitle={\\${\\textbackslash}text\\{G\\}{\\textasciicircum}2{\\textbackslash}text\\{M\\}\\$: A Generalized Gaussian Mirror Method to Boost Feature Selection Power},\nauthor={Hongyu Shen and Zhizhen Zhao},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=cldPfIoRiA}\n}"},"title":{"value":"$\\text{G}^2\\text{M}$: A Generalized Gaussian Mirror Method to Boost Feature Selection Power"},"pdf":{"value":"/pdf/6fd4edbf133fc27608a82bfe04ff7c54d6171fbf.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"shen|\\textg^2\\textm_a_generalized_gaussian_mirror_method_to_boost_feature_selection_power"},"authorids":{"value":["~Hongyu_Shen1","~Zhizhen_Zhao2"]},"authors":{"value":["Hongyu Shen","Zhizhen Zhao"]}},"tmdate":1783627333664,"pdate":1758217203466,"tcdate":1746932868093,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission18904/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission18904/Authors"],"forum":"cldPfIoRiA","license":"CC BY-NC 4.0","number":18904,"cdate":1746932868093,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission18904/-/Full_Submission","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission18904/-/Camera_Ready_Revision"],"mdate":1783627333664,"odate":1761704972492,"domain":"NeurIPS.cc/2025/Conference","id":"cldPfIoRiA","version":2},{"content":{"comment":{"value":"> The theoretical result also tells little information about what happens when the minimum of the objective function is not exactly achieved. Does one have error bounds from the optimization gap?\n\nWe have strengthened the theoretical results to address cases where the objective is not minimized exactly. In **Proposition 3**, we show that the minimizer of the Monte-Carlo estimator for the joint DiffCLF and DSM objective converges in probability to the true minimizer. Additionally, we establish asymptotic normality for this convergence, providing a quantitative characterization of the discrepancy due to finite-sample.\n\n---\n\n> The numerical validation of the proposed method is also rather weak, as it is only tested on synthetic Gaussian mixtures; which is far from practical setting of generative models.\n\nWe agree that Gaussian mixtures are synthetic and do not reflect the full complexity of practical generative modeling. Therefore, we conduct additional experiments on complex molecular systems, which showcase the effectiveness and scalability of the proposed method:\n - see **Figure 4, Figure 10, and L469-L474** for complex molecular systems, \n - and **Table 2 and L512-L521** for solvation free energy difference estimation on alanine dipeptide. \n\nHowever, we would like to first emphasize three points. \n - (i) Although Gaussian mixtures form the core benchmark, the experiments cover a broad range of tasks from energy estimation to recalibration via composition. \n - (ii) The synthetic setting enables precise, quantitative evaluation, which is essential for isolating the behavior of the objective itself. \n - (iii) **Section 5** shows that even in this simplified setup, DiffCLF significantly outperforms existing methods, highlighting both the limitations of previous objectives and the robustness of our approach; on more complex tasks, DiffCLF should therefore at least match, and often improve upon, existing baselines.\n\nTo further demonstrate applicability beyond synthetic data, we additionally reproduced the energy-learning experiments of **[1]** and show that DiffCLF estimates energies substantially more accurately than DSM on challenging molecular systems such as the alanine dipeptide and the protein Chignolin. In particular, we train energy-based diffusion model with DSM only and DSM+DiffCLF, and simulate Langevin dynamics using the learned energy at t=0 to sample from the target distribution. In other words, we obtain machine learning force field models by training energy-based diffusion models. **Figure 4** and **Figure 10** show that our method could simulate samples very similar to the equilibrium ones, demonstrating that it almost recovers the true energy landscape while simply training with DSM doesn’t. More specifically, our method requires much less computational budget compared to **[1]**.\n\nOn the other hand, we also demonstrate experiments on free energy difference estimation. In particular, we use the loss in **[2]** jointly with our diffusive classification loss to train energy-based models, and follow the same settings in **[2]** to estimate solvation free energy difference of alanine dipeptides through Thermodynamics Integration. The results (see **Table 2**) show that our method could improve the accuracy of free energy difference estimation, which demonstrates the effectiveness of our loss.\n\n---\n\n[1] Michael Plainer, Hao Wu, Leon Klein, Stephan Günnemann, & Frank Noé. (2025). Consistent Sampling and Simulation: Molecular Dynamics with Energy-Based Diffusion Models.\n\n[2] Máté, Bálint, François Fleuret, and Tristan Bereau. \"Solvation free energies from neural thermodynamic integration.\" The Journal of Chemical Physics 162.12 (2025)."},"title":{"value":"Response to Reviewer 8TVA [2/2]"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764113263675,"tcdate":1764113263675,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14499/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission14499/Authors"],"forum":"kpwGlSY1He","number":10,"license":"CC BY 4.0","cdate":1764113263675,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14499/-/Official_Comment"],"mdate":1764113263675,"domain":"ICLR.cc/2026/Conference","replyto":"qpdn7yfbIT","id":"D15ZbW5IRu","forumContent":{"TLDR":{"value":"We introduce a Diffusive Classification (DiffCLF) loss to learn energy-based generative models."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Energy-based Models","Diffusion Models","Stochastic Interpolants"]},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"abstract":{"value":"Score-based generative models have recently achieved remarkable success. While they are usually parameterized by the score, an alternative way is to use a series of time-dependent energy-based models (EBMs), where the score is obtained from the negative input-gradient of the energy. Crucially, EBMs can be leveraged not only for generation, but also for tasks such as compositional sampling or model recalibration via Monte Carlo methods. However, training EBMs remains challenging. Direct maximum likelihood is computationally prohibitive due to the need for nested sampling, while score matching, though efficient, suffers from mode blindness. To address these issues, we introduce the *Diffusive Classification* (DiffCLF) objective, a simple method that avoids blindness while remaining computationally efficient. DiffCLF reframes EBM learning as a supervised classification problem across noise levels, and can be seamlessly combined with standard score-based objectives. We validate the effectiveness of DiffCLF by comparing the estimated energies against ground truth in analytical Gaussian mixture cases, and by applying the trained models to tasks such as model composition and recalibration. Our results show that DiffCLF enables EBMs with higher fidelity and broader applicability than existing approaches."},"_bibtex":{"value":"@misc{\ngrenioux2026a,\ntitle={A Diffusive Classification Loss for Learning Energy-based Generative Models},\nauthor={Louis Grenioux and RuiKang OuYang and Jos{\\'e} Miguel Hern{\\'a}ndez-Lobato},\nyear={2026},\nurl={https://openreview.net/forum?id=kpwGlSY1He}\n}"},"title":{"value":"A Diffusive Classification Loss for Learning Energy-based Generative Models"},"pdf":{"value":"/pdf/7d1ac509df4762e85214d9655d4779b2d942c496.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"grenioux|a_diffusive_classification_loss_for_learning_energybased_generative_models"},"authorids":{"value":["~Louis_Grenioux1","~RuiKang_OuYang1","~José_Miguel_Hernández-Lobato1"]},"authors":{"value":["Louis Grenioux","RuiKang OuYang","José Miguel Hernández-Lobato"]}},"version":2},{"content":{"summary":{"value":"solving complex, multi-step, long-horizon tasks using prompting or fine-tuning poses various challenges. The paper presents a sound approach to enhance the capabilities of LLMs through self-improvement on web automation task. They propose evaluation techniques and effective use of in-distribution and out-of-distribution synthetic data, showcasing some performance gains."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"While the study is promising, its impact would be further enhanced by addressing the concerns regarding the selection of certain variables used in the experiments and evals and model diversity. \n\n### Suggestions for Improvement:\n- **Expand Hyperparameter Exploration:** Providing a sensitivity analysis or a justification for the selection of hyperparameters used in generating synthetic data would enhance the transparency and reproducibility of the results.\n- **Broader Agent Evaluation:** Including diverse model architectures in the experiments could help validate the proposed self-improvement techniques across a range of LLMs, potentially increasing the impact and applicability of the findings.\n\nI would be happy to increase the score post author responses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"### strenghts:\n1. **Capabilities through Self-Improvement:** The paper demos how LLMs can extend their capabilities through self-improvement techniques, particularly in the context of complex, long-horizon web agent tasks. This ability to acquire new capabilities while largely retaining existing ones is notable.\n2. **Appropriate Eval Metrics:** The introduction of novel metrics, and scores to evaluate the quality of trajectories, adds depth to the evaluation process. These metrics provide a nuanced view of the models' performance, moving beyond simple task completion metrics and enabling a more detailed assessment of the agent's robustness and capabilities. the additions seem appropriate  enough to capture both success, relevance, and efficiency of trajectoriesl\n3. **Synthetic Data Utilization:** The use of both in-domain and out-of-domain synthetic data for fine-tuning is commendable. This approach not only addresses the challenge of data scarcity but also demonstrates a systematic mechanism for enhancing agent model generalization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Good to address:\n1. **Hyperparameter Selection:** The paper lacks a clear justification for the choice of hyperparameters in the synthetic data generation process and generating new objectives. e.g. using 4 or 2 few-shot samples, temperature value, and perhaps just one sentence on why 0.7 cosine similarity used is better. Any missing details may lead readers to question the replicability and robustness of the results. A more thorough analysis or rationale for these choices would strengthen the paper.\n2. **Limited Model / Results:** The study's reliance on a single or fewer models limits its generalizability. Expanding the experiments to include a subset of models on appropriate agent frameworks that came out recently, such as smaller or different architectures, could provide insights into how these techniques scale or vary across different model sizes and types. I am thinking aloud if the results section can be tightened and elaborated more. But I will wait to see if my other reviewers has some feedback or ideas on this front."}},"nonreaders":[],"tmdate":1731428244446,"tcdate":1730580038268,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13182/Reviewer_xuoK"],"signatures":["ICLR.cc/2025/Conference/Submission13182/Reviewer_xuoK"],"forum":"jwME4SY0an","number":3,"license":"CC BY 4.0","cdate":1730580038268,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13182/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428244446,"domain":"ICLR.cc/2025/Conference","replyto":"jwME4SY0an","id":"EHYxLBMHBq","forumContent":{"TLDR":{"value":"In this work, we fine-tune large language models on synthetic data to self-improve at web agent tasks and evaluate on the WebArena benchmark, achieving a 31% improvement over the base model through our self-improvement procedure."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["llm","llms","synthetic data","web agents","agents","self-improvement","unsupervised learning"]},"supplementary_material":{"value":"/attachment/206f86832242309efcba12c73509c26ce9b74604.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Training models to act as agents that can effectively navigate and perform actions in a complex environment, such as a web browser, has typically been challenging due to lack of training data. Large language models (LLMs) have recently demonstrated some capability to navigate novel environments as agents in a zero-shot or few-shot fashion, purely guided by natural language instructions as prompts. Recent research has also demonstrated LLMs have the capability to exceed their base performance through self-improvement, i.e. fine-tuning on data generated by the model itself. In this work, we explore the extent to which LLMs can self-improve their performance as agents in long-horizon tasks in a complex environment using the WebArena benchmark. In WebArena, an agent must autonomously navigate and perform actions on web pages to achieve a specified objective. We explore fine-tuning on three distinct synthetic training data mixtures and achieve a 31% improvement in task completion rate over the base model on the WebArena benchmark through a self-improvement procedure. We additionally contribute novel evaluation metrics for assessing the performance, robustness, capabilities, and quality of trajectories of our fine-tuned agent models to a greater degree than simple, aggregate-level benchmark scores currently used to measure self-improvement."},"_bibtex":{"value":"@misc{\npatel2025large,\ntitle={Large Language Models Can Self-Improve At Web Agent Tasks},\nauthor={Ajay Patel and Markus Hofmarcher and Claudiu Leoveanu-Condrei and Marius-Constantin Dinu and Chris Callison-Burch and Sepp Hochreiter},\nyear={2025},\nurl={https://openreview.net/forum?id=jwME4SY0an}\n}"},"title":{"value":"Large Language Models Can Self-Improve At Web Agent Tasks"},"pdf":{"value":"/pdf/4d2920a9588bb50604035b663bde9663b71df2e2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"patel|large_language_models_can_selfimprove_at_web_agent_tasks"},"authorids":{"value":["~Ajay_Patel2","~Markus_Hofmarcher1","~Claudiu_Leoveanu-Condrei1","~Marius-Constantin_Dinu1","~Chris_Callison-Burch1","~Sepp_Hochreiter1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ajay Patel","Markus Hofmarcher","Claudiu Leoveanu-Condrei","Marius-Constantin Dinu","Chris Callison-Burch","Sepp Hochreiter"]}},"version":2},{"content":{"venue":{"value":"ICML 2025 poster"},"keywords":{"value":["Physics-enhanced Machine Learning","State Space Model","Long-term Dynamics Forecasting","Dynamical Systems"]},"_bibtex":{"value":"@inproceedings{\nwang2025a,\ntitle={A Generalizable Physics-Enhanced State Space Model for Long-Term Dynamics Forecasting in Complex Environments},\nauthor={Yuchen Wang and Hongjue Zhao and Haohong Lin and Enze Xu and Lifang He and Huajie Shao},\nbooktitle={Forty-second International Conference on Machine Learning},\nyear={2025},\nurl={https://openreview.net/forum?id=9NrUIaH1sx}\n}"},"title":{"value":"A Generalizable Physics-Enhanced State Space Model for Long-Term Dynamics Forecasting in Complex Environments"},"paperhash":{"value":"wang|a_generalizable_physicsenhanced_state_space_model_for_longterm_dynamics_forecasting_in_complex_environments"},"TLDR":{"value":"We propose Phy-SSM, a general-purpose framework that integrates partial physics knowledge into state space models (SSMs) for long-term dynamics forecasting."},"primary_area":{"value":"deep_learning->sequential_models_time_series"},"abstract":{"value":"This work aims to address the problem of long-term dynamic forecasting in complex environments where data are noisy and irregularly sampled. While recent studies have introduced some methods to improve prediction performance, these approaches still face a significant challenge in handling long-term extrapolation tasks under such complex scenarios. To overcome this challenge, we propose Phy-SSM, a general-purpose framework that integrates partial physics knowledge into state space models (SSMs) for long-term dynamics forecasting in complex environments. Our motivation is that SSMs can effectively capture long-range dependencies in sequential data and model continuous dynamical systems, while the incorporation of physics knowledge improves generalization ability. The key challenge lies in how to seamlessly incorporate partially known physics into SSMs. To achieve this, we decompose partially known system dynamics into known and unknown state matrices, which are integrated into a Phy-SSM unit. To further enhance long-term prediction performance, we introduce a physics state regularization term to make the estimated latent states align with system dynamics. Besides, we theoretically analyze the uniqueness of the solutions for our method. Extensive experiments on three real-world applications, including vehicle motion prediction, drone state prediction, and COVID-19 epidemiology forecasting, demonstrate the superior performance of Phy-SSM over the baselines in both long-term interpolation and extrapolation tasks. The source code will be publicly available upon publication."},"link_to_code":{"value":"https://github.com/511205787/Phy_SSM-ICML2025"},"pdf":{"value":"/pdf/1502e6e8ed21903934f784df699d9c56ad12a7b0.pdf"},"lay_summary":{"value":"Predicting how systems evolve over time — such as the motion of cars and drones, or the spread of disease — is challenging in real-world settings, especially when data is noisy and our understanding of the underlying physics is incomplete. We wondered: can we still make accurate predictions using only partial knowledge of the physical system?\n\nTo tackle this, we developed Phy-SSM — a method that seamlessly combines partially known physics with a deep state space model, a type of AI that excels at learning from time series data. Our key innovation is to explicitly separate the known and unknown parts of the physical system, then merge them in a unified framework that helps the model learn the true system dynamics. Surprisingly, we found that even incomplete physics knowledge can significantly improve the model’s ability to generalize for long-term predictions on unseen data.\n\nPhy-SSM allows users to easily incorporate partial physics knowledge into different systems to make more accurate forecasts. Our research highlights the potential of combining traditional physics with modern AI to tackle real-world prediction challenges."},"venueid":{"value":"ICML.cc/2025/Conference"},"authorids":{"value":["~Yuchen_Wang6","~Hongjue_Zhao1","~Haohong_Lin1","~Enze_Xu1","~Lifang_He1","~Huajie_Shao1"]},"authors":{"value":["Yuchen Wang","Hongjue Zhao","Haohong Lin","Enze Xu","Lifang He","Huajie Shao"]}},"tmdate":1753292992901,"pdate":1746105447335,"tcdate":1737603397370,"writers":["ICML.cc/2025/Conference","ICML.cc/2025/Conference/Submission8797/Authors"],"signatures":["ICML.cc/2025/Conference/Submission8797/Authors"],"forum":"9NrUIaH1sx","license":"CC BY 4.0","number":8797,"cdate":1737603397370,"readers":["everyone"],"invitations":["ICML.cc/2025/Conference/-/Submission","ICML.cc/2025/Conference/-/Post_Submission","ICML.cc/2025/Conference/Submission8797/-/Full_Submission","ICML.cc/2025/Conference/-/Edit","ICML.cc/2025/Conference/Submission8797/-/Camera_Ready_Revision"],"mdate":1753292992901,"odate":1750231275835,"domain":"ICML.cc/2025/Conference","id":"9NrUIaH1sx","version":2},{"content":{"summary":{"value":"This paper proposes to use LLM-generated synthetic data to augment the training of stance classification models. It also proposes a synthetic data-based active learning method that uses synthetic data to facilitate the selection of unlabelled data for human annotation. Experiments are conducted on the German subset of the X-stance dataset (with the help of machine translation). The results demonstrate that including synthetic data in training can improve stance prediction. The synthetic data-based active learning method, however, is not clearly better than a random selection-based baseline active learning method."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- It would be very helpful to explain why only a German dataset is used for the experiments. Also, if German text is used, have the authors considered using a different LLM that has good German language processing capabilities for the experiments?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The proposed method is sound. I do not see any major issue with the method.\n- Although the idea of using synthetic data to augment models is not entirely new, it probably has not been widely explored for stance prediction.\n- The authors conducted extensive experiments to evaluate the method, including varying the size of the synthetic dataset, comparing with meaningful baselines, and the further experiments that compare with a LLM zero-shot baseline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The experiments are conducted using a German dataset, but translation into and back from English is used in order for the method to work (probably because of limited German language understanding and generation capabilities of the Mistral model that is used?) There is no explanation of why the authors do not evaluate the method using an English dataset.\n- The novelty and impact of the work is still limited. (1) Using synthetic data to augment models is not new. Although applying the idea to stance prediction might be new, it is one of many NLP tasks. The way synthetic data is generated and used during training in this paper is also standard, hence there is limited technical contribution. (2) The idea of using synthetic data for active learning is very interesting and is novel based on my knowledge. However, its effectiveness is limited based on the experiments. Therefore, overall, although the work is very solid in general, its novelty and impact may not meet the standard of this conference.\n- There is room for improvement in terms of presentation. In particular, the active learning method proposed can benefit from first presenting an overview of the high-level intuition behind the method before describing the method itself."}},"nonreaders":[],"tmdate":1731812261049,"tcdate":1730607261202,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6567/Reviewer_cUfQ"],"signatures":["ICLR.cc/2025/Conference/Submission6567/Reviewer_cUfQ"],"forum":"ws5phQki00","number":2,"license":"CC BY 4.0","cdate":1730607261202,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6567/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731812261049,"domain":"ICLR.cc/2025/Conference","replyto":"ws5phQki00","id":"nPD3AtwS6E","forumContent":{"TLDR":{"value":"We study and show how to leverage LLM-generated synthetic data for stance detection in online discussions, which is a challenging stance detection task because of the broad range of debate questions."},"venue":{"value":"ICLR 2025 Spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["large language models","stance detection","data augmentation","active learning","online political discussions"]},"supplementary_material":{"value":"/attachment/a99830444732d29e57db39e7ad67fec6f8bd38ae.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Stance detection holds great potential to improve online political discussions through its deployment in discussion platforms for purposes such as content moderation, topic summarisation or to facilitate more balanced discussions. Typically, transformer-based models are employed directly for stance detection, requiring vast amounts of data. However, the wide variety of debate topics in online political discussions makes data collection particularly challenging. LLMs have revived stance detection, but their online deployment in online political discussions faces challenges like inconsistent outputs, biases, and vulnerability to adversarial attacks. We show how LLM-generated synthetic data can improve stance detection for online political discussions by using reliable traditional stance detection models for online deployment, while leveraging the text generation capabilities of LLMs for synthetic data generation in a secure offline environment. To achieve this, (i) we generate synthetic data for specific debate questions by prompting a Mistral-7B model and show that fine-tuning with the generated synthetic data can substantially improve the performance of stance detection, while remaining interpretable and aligned with real world data. (ii) Using the synthetic data as a reference, we can improve performance even further by identifying the most informative samples in an unlabelled dataset, i.e., those samples which the stance detection model is most uncertain about and can benefit from the most. By fine-tuning with both synthetic data and the most informative samples, we surpass the performance of the baseline model that is fine-tuned on all true labels, while labelling considerably less data."},"_bibtex":{"value":"@inproceedings{\nwagner2025the,\ntitle={The Power of {LLM}-Generated Synthetic Data for Stance Detection in Online Political Discussions},\nauthor={Stefan Sylvius Wagner and Maike Behrendt and Marc Ziegele and Stefan Harmeling},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=ws5phQki00}\n}"},"title":{"value":"The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions"},"pdf":{"value":"/pdf/5b56a56d8a30ca36a3662831fa3394f290a4450f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"wagner|the_power_of_llmgenerated_synthetic_data_for_stance_detection_in_online_political_discussions"},"authorids":{"value":["~Stefan_Sylvius_Wagner1","~Maike_Behrendt1","~Marc_Ziegele1","~Stefan_Harmeling1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Stefan Sylvius Wagner","Maike Behrendt","Marc Ziegele","Stefan Harmeling"]}},"version":2},{"content":{"summary":{"value":"The authors address the inefficiency of complex Monte Carlo approximations for Entropy Search (ES)-based acquisition functions in Bayesian Optimization (BO), which cause numerical errors and require cumbersome hand-crafted implementations. It proposes a two-stage amortization strategy using Prior-data Fitted Networks (PFNs): a base PFN is first trained to condition on information about the optima, and then the α-PFN is trained to predict the expected information gain using the information gains computed by the base PFN, enabling approximation in a single forward pass.\n\nKey contributions:\n1) It innovatively applies PFNs to amortize the approximation of ES variants (Predictive Entropy Search, Max-value Entropy Search, Joint Entropy Search), replacing costly sampling-based methods with a fast single forward pass and achieving speedups of at least 12 times (up to 41.2 times for 8-dimensional problems); \n2) It supports fully Bayesian Gaussian Process models by integrating hyperparameter uncertainty, a capability rarely seen in existing ES implementations;\n 3) Empirically, it matches the performance of state-of-the-art ES methods on synthetic Gaussian Process benchmarks and real-world hyperparameter optimization tasks from LC-Bench while maintaining high computational efficiency."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Supplement high-dimensional performance decomposition experiments: Count α-PFN’s time in \"acquisition function evaluation\", \"data preprocessing,\" and \"model inference\" for 10-50D synthetic tasks, compare with traditional ES. Analyze high-dimensional feature importance via attention visualization; add mutual information-based feature selection if redundant features exist.\n\n2. Optimize synthetic trajectory parameters: Use grid search (instead of random sampling) to find optimal ranges based on real-task (e.g., LC-Bench) historical queries. Introduce DTW distance to verify trajectory consistency, improving PFN training data authenticity and model generalization.\n\n3. Explore hyperparameter sampling optimization: Test 5/10/20/50 sampling times’ impact on α-PFN. Fix 10 times if performance is comparable to more samplings. Try gradient-based methods (e.g., HMC) to enhance hyperparameter coverage and fully Bayesian model accuracy.\n\n4. Improve out-of-distribution domain shift: Add domain adversarial training or augment GP samples like Levy function for tasks (e.g., Levy 4D). Supplement \"domain shift degree-model performance\" curves to clarify α-PFN’s adaptation boundary."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"This paper proposes a two-stageα-PFN framework, which uses PFN amortized approximation to break through the limitation that traditional Entropy Search (ES)-based acquisition functions rely on complex Monte Carlo sampling. It is compatible with PES/MES/JES variants and supports fully Bayesian Gaussian Processes. Experiments on synthetic data and LC-Bench tasks verify that its performance is comparable to that of mainstream ES methods, with a speedup of at least 12× (reaching 41.2× for 8-dimensional tasks). The paper has a clear logical structure and explicit details, providing a new paradigm for the research on efficient acquisition functions in Bayesian Optimization and reducing the computational cost of practical applications."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Insufficient generalization in high-dimensional scenarios: The model is only trained up to 6 dimensions and validated on 8-dimensional tasks, failing to cover higher dimensions (e.g., 50+ dimensions), and the reasons for performance degradation in high dimensions are not analyzed. It is recommended to supplement synthetic experiments for 10-50 dimensions, analyze the feature utilization efficiency of PFN, and conduct corresponding optimizations.\n\n2. Limited coverage of real-world scenarios: The validation is only conducted on LC-Bench hyperparameter optimization tasks, without involving complex real-world scenarios such as multi-objective optimization and noisy data. It is recommended to expand to multi-objective Bayesian Optimization (BO) tasks, validate on real noisy datasets, and supplement noise-resistant strategies.\n\n3. Unoptimized PFN architecture: A fixed 6-layer Transformer architecture is adopted, with no exploration of the impact of different numbers of layers or attention heads, nor any attempts at model lightweighting. It is recommended to conduct architecture ablation experiments to select the optimal configuration and attempt model distillation.\n\n4. Incomplete baseline comparisons: Only traditional ES variants and Log EI are used for comparison, with no inclusion of emerging methods (e.g., BORE). It is recommended to supplement such baselines to clarify the technical positioning of α-PFN."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926689732,"tcdate":1761877154330,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16620/Reviewer_WKw4"],"signatures":["ICLR.cc/2026/Conference/Submission16620/Reviewer_WKw4"],"forum":"AfzjIpcUjD","number":3,"license":"CC BY 4.0","cdate":1761877154330,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16620/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926689732,"domain":"ICLR.cc/2026/Conference","replyto":"AfzjIpcUjD","id":"dQI0ctlliM","forumContent":{"TLDR":{"value":"We use the framework of Prior-data Fitted Networks (PFNs) to develop the α-PFN transformer, that learns to approximate the entropy search acquisition function in a single forward pass for fast Bayesian Optimization."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["prior fitted network","Bayesian Optimization","entropy search","transformer","metalearning","information-theoretic acquisition functions","in-context learning"]},"supplementary_material":{"value":"/attachment/b8fbb9d63d07710e5042afadd3b645f350fd5f19.zip"},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"abstract":{"value":"Information‐theoretic acquisition functions such as Entropy Search (ES) offer a principled exploration–exploitation framework for Bayesian optimization (BO). However, their practical implementation relies on complicated and slow approximations, i.e., a Monte Carlo estimation of the information gain. This complexity can introduce numerical errors and requires specialized, hand-crafted implementations.\nWe propose a two‐stage amortization strategy that learns to approximate entropy search-based acquisition functions using Prior‐data Fitted Networks (PFNs) in a single forward pass. A first PFN is trained to be conditioned on information about the optima; second, the $\\alpha$‐PFN is trained to predict the expected information gain by training on information gains measured with the first PFN. The $\\alpha$-PFN offers a scalable and learnable approximation, which replaces the complex approximations with a single forward pass per candidate, enabling rapid and extensible acquisition evaluation. Empirically, our approach is competitive with state‐of‐the‐art entropy search implementations on synthetic and real‐world benchmarks while accelerating the different entropy search variants by over at least a factor of 12x, with the largest speed ups around 30x for the highest 8 dimensional problems."},"_bibtex":{"value":"@misc{\nviering2026alphapfn,\ntitle={\\${\\textbackslash}alpha\\$-{PFN}: Fast Entropy Search via In-Context Learning},\nauthor={Tom Julian Viering and Steven Adriaensen and Herilalaina Rakotoarison and Samuel M{\\\"u}ller and Carl Hvarfner and Frank Hutter and Eytan Bakshy},\nyear={2026},\nurl={https://openreview.net/forum?id=AfzjIpcUjD}\n}"},"title":{"value":"$\\alpha$-PFN: Fast Entropy Search via In-Context Learning"},"pdf":{"value":"/pdf/63320b0fd19e22c28eb65444355157514ea05128.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"viering|\\alphapfn_fast_entropy_search_via_incontext_learning"},"authorids":{"value":["~Tom_Julian_Viering1","~Steven_Adriaensen1","~Herilalaina_Rakotoarison1","~Samuel_Müller1","~Carl_Hvarfner1","~Frank_Hutter1","~Eytan_Bakshy1"]},"authors":{"value":["Tom Julian Viering","Steven Adriaensen","Herilalaina Rakotoarison","Samuel Müller","Carl Hvarfner","Frank Hutter","Eytan Bakshy"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a multiple physics pretraining approach for surrogate modeling, which learns general useful features across diverse physical tasks with a shared embedding and normalization strategy. The experiment results show the proposed MPP-pretrained model outperforms task-specific baselines on all pretraining sub-tasks and also show superior finetuning results on new physics tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Does MPP need to be trained in a context of very similar physical backgrounds (e.g., SWE, DiffRe2D and CNS)? How can we determine the similarity of multiple physics fields and whether they can be learned simultaneously?\n2. Single MPP can learn the dynamics for multiple classes of physical behavior. If the physical equations are vastly different and include different forms of dynamic processes, can one model still perform well? For example, can a single model handle Newtonian fluids and non-Newtonian fluids, rigid body dynamics or elastodynamics?\n3. Do baseline methods use the same normalized training loss as MPP? Why do you evaluate using Normalized Mean Squared Error (NMSE) instead of Mean Squared Error (MSE) during evaluation?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The idea of constructing a large pre-trained base model for physical simulations is promising. The experiments are performed for diverse physical systems. The results show that large surrogate models outperform strong baselines, and even models with a relatively few parameters can learn such diverse physics evolutions and perform competitively."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The validation of the approach is primarily conducted on fluid mechanics-oriented benchmarks. While this is a solid start, the applicability of the approach to other domains of physics remains to be demonstrated extensively.\n- The proposed model can handle simulations on structured meshes. However, simulations on unstructured mesh are under exploration."},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1730878818243,"tcdate":1720760653052,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission2983/Reviewer_mqdr"],"signatures":["NeurIPS.cc/2024/Conference/Submission2983/Reviewer_mqdr"],"forum":"DKSI3bULiZ","number":2,"license":"CC BY 4.0","cdate":1720760653052,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission2983/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878818243,"domain":"NeurIPS.cc/2024/Conference","replyto":"DKSI3bULiZ","id":"ENppGnTIYl","forumContent":{"TLDR":{"value":"We develop approaches to enable autoregressive pretraining on multiple physical systems and show it can improve transfer performance across domain gaps."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["transfer learning","physics","pretraining","finetuning","surrogate models","spatiotemporal"]},"supplementary_material":{"value":"/attachment/3bda966ab5ebc7a1ef9f0d7091eec3a790198a1f.zip"},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"We introduce multiple physics pretraining (MPP), an autoregressive task-agnostic pretraining approach for physical surrogate modeling of spatiotemporal systems with transformers. In MPP, rather than training one model on a specific physical system, we train a backbone model to predict the dynamics of multiple heterogeneous physical systems simultaneously in order to learn features that are broadly useful across systems and facilitate transfer. In order to learn effectively in this setting, we introduce a shared embedding and normalization strategy that projects the fields of multiple systems into a shared embedding space. We validate the efficacy of our approach on both pretraining and downstream tasks over a broad fluid mechanics-oriented benchmark. We show that a single MPP-pretrained transformer is able to match or outperform task-specific baselines on all pretraining sub-tasks without the need for finetuning. For downstream tasks, we demonstrate that finetuning MPP-trained models results in more accurate predictions across multiple time-steps on systems with previously unseen physical components or higher dimensional systems compared to training from scratch or finetuning pretrained video foundation models. We open-source our code and model weights trained at multiple scales for reproducibility."},"_bibtex":{"value":"@inproceedings{\nmccabe2024multiple,\ntitle={Multiple Physics Pretraining for Spatiotemporal Surrogate Models},\nauthor={Michael McCabe and Bruno R{\\'e}galdo-Saint Blancard and Liam Holden Parker and Ruben Ohana and Miles Cranmer and Alberto Bietti and Michael Eickenberg and Siavash Golkar and Geraud Krawezik and Francois Lanusse and Mariel Pettee and Tiberiu Tesileanu and Kyunghyun Cho and Shirley Ho},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=DKSI3bULiZ}\n}"},"title":{"value":"Multiple Physics Pretraining for Spatiotemporal Surrogate Models"},"pdf":{"value":"/pdf/f17e2dec513bf41ef3a4a0158ce2876724a18c38.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"mccabe|multiple_physics_pretraining_for_spatiotemporal_surrogate_models"},"authorids":{"value":["~Michael_McCabe2","~Bruno_Régaldo-Saint_Blancard1","~Liam_Holden_Parker1","~Ruben_Ohana1","~Miles_Cranmer2","~Alberto_Bietti1","~Michael_Eickenberg5","~Siavash_Golkar1","~Geraud_Krawezik1","~Francois_Lanusse2","~Mariel_Pettee1","~Tiberiu_Tesileanu1","~Kyunghyun_Cho1","~Shirley_Ho2"]},"authors":{"value":["Michael McCabe","Bruno Régaldo-Saint Blancard","Liam Holden Parker","Ruben Ohana","Miles Cranmer","Alberto Bietti","Michael Eickenberg","Siavash Golkar","Geraud Krawezik","Francois Lanusse","Mariel Pettee","Tiberiu Tesileanu","Kyunghyun Cho","Shirley Ho"]}},"version":2},{"content":{"summary":{"value":"This paper discusses a hypothesis testing scenario, where evaluating the likelihood function is intractable and hence the standard likelihood ratio test (LRT) becomes infeasible. The paper considers an alternative approach, where a neural classifier is trained based on simulated data under different hypotheses. In particular, it considers a solution where multiple instances are assigned to one common label (called MIL). The paper examines this idea in particle physics, for detecting deviations from Standard Model in collision experiments. By experimentation on synthetic data, it is shown that the proposed method is superior to the combination of decisions on individual instances (ensemble methods). This observation is further theoretically justified by arguments involving Fisher information and Cramer-Rao bound."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"After understanding the problem of interest, I am surprised about the use of LRT, in this context. The reason is that one of the alternatives is presented as a composite hypothesis (theta\\neq 0). A consequence of applying LRT is that the alternative values of theta (e.g. \\theta_1) must be selected beforehand. How can this be done? And are not tests such as GLRT more suitable for this scenario?\n\nIf I understand it correctly, the training procedure does not explicitly bias the logits toward the individual likelihood ratios. Is it possible to guarantee that they are estimates of LRs? Indeed, Appendix C shows that they are biased in nature. And if biased, how is the presented theory based on CRB relevant to them?\n\nAs a minor comment in (11), do you mean by “+ 0” a higher order term?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"I am not familiar with physics literature and hence cannot assess the significance of the paper within this field. From a general statistics perspective, especially in the context of data fusion, the contribution of the paper is a lightweight method, based on LRT, which fuses multiple instances at a feature level rather than at a decision level. Feature-level fusion is known to be superior to decision-level fusion, especially in low-SNR regimes, but is generally considered a complex task."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I wonder how novel or substantial contribution is. As already mentioned, the fact that feature-level fusion is superior to the decision level is well-known and intuitive. The use of NNs for estimating likelihood ratios is not entirely new and is extensively discussed in the context of neural ratio estimation (NRE). From the perspective of MIL, the paper considers a simplified scenario, which to me is a repetition of the\nstandard point estimation theory with multiple observations. A major part of the theoretical discussions, e.g. the vanishing of ML error and growth of FI with O(\\sqrt{N}), can be found in multiple classical sources.\n\nAnother drawback of the suggested approach is that for deployment, it requires an ensemble of independent observations of similar size to the ones used for training. This can be a limitation in practice.\n\nThe presentation of the paper can also be improved. It is sometimes difficult to understand the motivation behind the concepts introduced. For example, I am not familiar with the notion of effective Fisher information, and it is not clear to me what it implies. Some notations remain unexplained too. For example, in line 198 e_ij seems to refer to the elements of e_i, but this is not defined. Moreover, \\theta_SM and \\theta_SMEFT are not properly introduced."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926083373,"tcdate":1761922730501,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15863/Reviewer_PHoD"],"signatures":["ICLR.cc/2026/Conference/Submission15863/Reviewer_PHoD"],"forum":"PLva6Rol4W","number":2,"license":"CC BY 4.0","cdate":1761922730501,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15863/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926083373,"domain":"ICLR.cc/2026/Conference","replyto":"PLva6Rol4W","id":"Uzuy7w1t89","forumContent":{"TLDR":{"value":"We introduce an information-theoretic view of Multiple Instance Learning for its use in parameter estimation problems in High-Energy Physics, and demonstrate improved performance."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multiple Instance Learning","High-Energy Physics","Hypothesis Testing","Fisher Information","Parameter Estimation","Simulation-based Inference"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"In this work, we introduce a new information-theoretic perspective on Multiple Instance Learning (MIL) for parameter estimation with i.i.d. data, and show that MIL can outperform single-instance learners in low-signal regimes. Prior work \\citep{nachman_learning_2021} argued that per-instance methods are often sufficient, but this conclusion presumes enough per-instance signal to train near-optimal classifiers. We demonstrate that even state-of-the-art per-instance models can fail to reach optimal classifier performance in challenging low-signal regimes, whereas MIL can mitigate this sub-optimality. As a concrete application, we constrain Wilson coefficients of the Standard Model Effective Field Theory (SMEFT) using kinematic information from subatomic particle collision events at the Large Hadron Collider (LHC). In experiments, we observe that under specific modeling and weak signal conditions, pooling instances can increase the effective Fisher information compared to single-instance approaches."},"_bibtex":{"value":"@misc{\nazakli2026increasing,\ntitle={Increasing Information Extraction in Low-Signal Regimes via Multiple Instance Learning},\nauthor={Atakan AZAKLI and Bernd STELZER},\nyear={2026},\nurl={https://openreview.net/forum?id=PLva6Rol4W}\n}"},"title":{"value":"Increasing Information Extraction in Low-Signal Regimes via Multiple Instance Learning"},"pdf":{"value":"/pdf/d2052c9157811a19186311763d656a1d44c875a0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"azakli|increasing_information_extraction_in_lowsignal_regimes_via_multiple_instance_learning"},"authorids":{"value":["~Atakan_AZAKLI1","~Bernd_STELZER1"]},"authors":{"value":["Atakan AZAKLI","Bernd STELZER"]}},"version":2},{"content":{"summary":{"value":"This paper addresses a fundamental challenge in training Physics-Informed Neural Networks (PINNs): the instability and poor performance caused by conflicting gradients from the multiple loss terms (e.g., PDE residuals and boundary conditions). The authors propose a novel solution by reformulating PINN training as a nonconvex-strongly concave saddle-point problem (SPP). The authors provide a solid theoretical foundation for their proposed Bregman Gradient Descent Ascent (BGDA) algorithm, proving its convergence to a stationary point even in this non-Euclidean geometry setting. Extensive experiments on the comprehensive PINNacle benchmark demonstrate that AdaptiveBGDA significantly outperforms a wide range of state-of-the-art optimizers and weighting schemes across different PDE problems."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. See weakness above.\n\n2. Could you give insights on how practitioners should go about selecting the regularization parameter λ and the initial learning rates for a new problem?\n\n3. The experiments on high dimension is simply on d=5 dimensions, which is very weak indeed. Could you report results on higher dimensions like d=100?\n\n4. The theory and experiments consider a small number of loss terms (M). How would the method scale and perform if M were very large, for instance, in problems with a very large number of boundary condition constraints or coupled multi-physics systems?\n\n5. Assumption 1 needs the Lipschitz continuous for the loss function, which is widely used in the optimization literature. However, since the PINN loss is highly non-convex and ill-conditioned, does assumption 1 really hold for the PINN loss? How about the convergence"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The saddle-point reformulation is a principled and theoretically grounded approach to the known problem of loss imbalance in PINNs..\n\n2. The results are extensive, covering wide variety of PDEs (Poisson, Heat, Navier-Stokes, etc.) and challenging features (complex geometry, multiple domains, long time intervals). \n\n3. The proposed AdaptiveBGDA algorithm is a concrete and practical contribution that can be readily adopted by other researchers and practitioners to improve the stability and performance of their PINN models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although rich experiments are conducted based on PINNacle, the relative error values in Table 3 are relatively bad and did not reflect the state of the art results, and the improvement seems modest. For example, the Poisson and Heat equations are simple and can be easily solved by traditional numerical methods with high accuracy, however, the errors stays around $\\mathcal{O}(10^{-2})$, which are too large. For the complex chaotic KS equation, the improvements are from 9.57E-1 to 9.53E-1, however, for the same solution, work [a] in the year 2023 has reported a much better result 1.61E-1.\n[a] AN EXPERT’S GUIDE TO TRAINING PHYSICS-INFORMED NEURAL NETWORKS, 2023.\n\n2. The experiments in the main text seem to use a standard PINN architecture.  However, vanilla PINN architectures are hard to present the sota results[b]. While Appendix A shows results on more advanced architectures, a deeper discussion on how the saddle-point optimization interacts with specific architectural choices (e.g., Fourier feature networks, transformers) would be valuable.\n[b] Causality-enhanced Discreted Physics-informed Neural Networks for Predicting\nEvolutionary Equations,2024\n\n3. While the paper shows a runtime speedup, the saddle-point formulation inherently introduces additional computational overhead per iteration due to the proximal step for updating π\n\n4. The method introduces new hyperparameters (λ for the regularizer, step sizes γ_θ and γ_π). Although the paper uses fixed values across all benchmarks to show robustness, optimal performance in new, highly complex problems might still require tuning."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919419706,"tcdate":1761020972100,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7302/Reviewer_xUeL"],"signatures":["ICLR.cc/2026/Conference/Submission7302/Reviewer_xUeL"],"forum":"EQNp3sFrY3","number":1,"license":"CC BY 4.0","cdate":1761020972100,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7302/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919419706,"domain":"ICLR.cc/2026/Conference","replyto":"EQNp3sFrY3","id":"Mpajqfx3gd","forumContent":{"TLDR":{"value":"Adaptive weighting of losses corresponding to equations and boundary conditions during PINN training"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Physics-informed neural networks","Multi-task learning","Saddle-point problems","Scientific machine learning"]},"primary_area":{"value":"optimization"},"abstract":{"value":"Physics-informed neural networks (PINNs) have gained prominence in recent years and are now effectively used in a number of applications. However, their performance remains unstable due to the complex landscape of the loss function. To address this issue, we reformulate PINN training as a nonconvex-strongly concave saddle-point problem. After establishing the theoretical foundation for this approach, we conduct an extensive experimental study, evaluating its effectiveness across various tasks and architectures. Our results demonstrate that the proposed method outperforms the current state-of-the-art techniques."},"_bibtex":{"value":"@inproceedings{\nbylinkin2026enhancing,\ntitle={Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation},\nauthor={Dmitry Bylinkin and Mikhail Aleksandrov and Savelii Chezhegov and Aleksandr Beznosikov},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EQNp3sFrY3}\n}"},"title":{"value":"Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation"},"pdf":{"value":"/pdf/ce0ff129724ea008d7d711a49805886c6f706d9b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"bylinkin|enhancing_stability_of_physicsinformed_neural_network_training_through_saddlepoint_reformulation"},"authorids":{"value":["~Dmitry_Bylinkin1","~Mikhail_Aleksandrov1","~Savelii_Chezhegov1","~Aleksandr_Beznosikov1"]},"authors":{"value":["Dmitry Bylinkin","Mikhail Aleksandrov","Savelii Chezhegov","Aleksandr Beznosikov"]}},"version":2},{"content":{"venue":{"value":"Reliability Engineering & System Safety"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"hajiha|a_physicsregularized_datadriven_approach_for_health_prognostics_of_complex_engineered_systems_with_dependent_health_states"},"authorids":{"value":["mhajiha@uark.edu","~Xiao_Liu37","~Young_M_Lee1","rxm991@miami.edu"]},"html":{"value":"https://www.sciencedirect.com/science/article/pii/S0951832022003106"},"abstract":{"value":"Advances in sensing technology enable the monitoring of critical operating parameters of complex engineering systems. However, having sensor measurements does not necessarily imply that one has observed the true system health states, which are often hidden and need to be estimated from observable sensor signals. This paper proposes a physics-regularized data-driven approach for the health prognostics of complex engineered systems with multiple hidden and dependent health states. The framework consists of a data layer and a physics layer. The data layer captures the statistically-correlated temporal dynamics of hidden system states (such as degradation), while the physics layer imposes regularizations among observed system operating parameters and system health states through system working principles and governing physics. The proposed approach addresses some common challenges arising from the health prognostics of complex engineered systems, including the integration of engineering domain knowledge and sensor data streams, the estimation of hidden system health states from monitored system operation parameters, and the statistical dependency among the temporal dynamics of multiple system state variables. A case study based on a real dataset is presented to illustrate the proposed physical–statistical approach. It is shown that the interpretability of data-driven system prognostics can be significantly strengthened if a solid connection is established between sensor data and system physics."},"title":{"value":"A physics-regularized data-driven approach for health prognostics of complex engineered systems with dependent health states"},"authors":{"value":["Mohammadmahdi Hajiha","Xiao Liu","Young M Lee","Moghaddass Ramin"]}},"tmdate":1716328524983,"pdate":1656024915373,"tcdate":1716328524983,"writers":["mhajiha@uark.edu","~Xiao_Liu37","~Young_M_Lee1","rxm991@miami.edu"],"signatures":["~Xiao_Liu37"],"forum":"StyLWR6sXl","license":"CC BY 4.0","number":26394,"cdate":1716328524983,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1716328524983,"domain":"OpenReview.net/Archive","id":"StyLWR6sXl","version":2},{"content":{"comment":{"value":"We sincerely thank the reviewer for their insightful feedback and constructive suggestions. We appreciate the reviewer's positive feedback that our motivation \"correctly focuses on a simpler, core problem\" by separating physical motion perception from complex reasoning. We are also glad the reviewer found that our representation SDF \"helps the model learn the idea of dynamics, letting it generalize.\" We have provided detailed responses below to address the reviewer's questions and concerns.\n\n### Q1:  **Regarding the simplicity of the SDF representation and comparison to alternatives:**\n\nThe reviewer correctly observes that our Scene Dynamic Field (SDF) is a concise representation, encoding projected velocity magnitude into a single color channel. This was a deliberate design choice and is central to our paper's core objective: to train Vision-Language Models (VLMs) for intuitive physical understanding.\n\n1. **The Goal is Abstract Intuitive Understanding, Not Fine-Grained Reconstruction:** Our primary goal is to enhance the *intuitive physics understanding* of MLLMs for high-level reasoning and planning. As we state in Section 5, we aim to \"appropriately abstract\"  representations from simulators. The SDF is designed to be an intermediate visual prompt that is \"perceptually calibrated\" and optimized for VLM consumption, rather than a high-fidelity, complex data field.\n\n2. **VLM Suitability and Scalability:** As the reviewer suggests, representations like full 3D velocity vectors would contain more data. However, we argue that current VLMs are not well-equipped to directly interpret such fine-grained, raw vector fields. This would require massive, specialized training for the VLM to develop the necessary 3D spatial understanding, which runs counter to our goal of a \"cost-efficient approach\". 3D vectors could be challenging for current VLMs to leverage directly. As noted in recent work (e.g., [1]Fan et al., 2025; [2] Zheng et al., 2024), enabling VLMs to gain a true 3D spatial and vector-based understanding typically requires massive-scale, specialized effort. Our focus, in contrast, is to inject intuitive physical dynamics so that a VLM can readily consume and integrate into its reasoning process, making our abstract representation a more suitable choice for this goal.\n\n   [1] Fan, Zhiwen et al. VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction. *ArXiv* abs/2505.20279 (2025)\n\n   [2] Zheng, D., Huang, S., & Wang, L. (2024). Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding. CVPR 2025\n\n3. **On Alternative Representations (e.g., Optical Flow):** To directly address the reviewer's excellent suggestion, we conducted an additional experiment comparing our SDF to other representations. We directly evaluate SDF's effectiveness by comparing across four prompt conditions using the same question answering protocol:\n\n(1) **w/o Visual Prompt**: directly QA without any visual prompt.\n\n(2) **w/ Optical Flow**: add the reconstructed optical flow in prompt for QA.\n\n(3) **w/ Depth**: add depth estimation of the last frame in prompt for QA.\n\n(4) **w/ SDF(ours)**: add our proposed SDF as visual prompt for QA.\n\n| Visual Prompt | NFS Acc | TCV Acc |\n| ------------- | ------- | ------- |\n| No            |  29.1   |  57.5   |\n| Optical Flow  |  27.0   |  65.1   |\n| Depth         |  32.4   |  60.1   |\n| SDF (ours)    |  41.2   |  68.9   |\n\nThese results indicate that optical flow is less effective on NFS where reconstruction noise limits its discriminative value. In contrast, SDF achieves the best performance across both tasks, showing that clean velocity from the simulator yields a more reliable and more generalizable visual prompt.\n\nOur simulator-generated SDF, while simpler, provides a *cleaner* and *more reliable* signal of the underlying motion. It bypasses the noise of 2D reconstruction and directly encodes the \"ground truth\" dynamics from the physics engine. This reliability is crucial for building a robust intuitive understanding. We include the experiment and an example figure in the revised supplementary on page 16 with blue highlight. We thank the reviewer again for prompting this valuable addition."},"title":{"value":"Rebuttal to Reviewer tuv4 (Part 1/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763626890496,"tcdate":1763626890496,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7741/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission7741/Authors"],"forum":"Ax02eR2c3d","number":1,"license":"CC BY 4.0","cdate":1763626890496,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7741/-/Official_Comment"],"mdate":1763626890496,"domain":"ICLR.cc/2026/Conference","replyto":"El4N02cPEa","id":"P6ztPR5hR3","forumContent":{"TLDR":{"value":"This paper introduces two low-level tasks to test intuitive physics understanding and proposes Scene Dynamic Field, a method to integrate visual representation from physics simulators to MLLMs while showcasing generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multi-modal LLM","Intuitive Physics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. \nIn this work, we investigate the first step of physical reasoning, i.e., **intuitive physics understanding**, revealing substantial limitations in understanding the dynamics of continuum objects. \nTo isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. \nTo address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. \nSDF substantially improves performance, achieving up to $20.7\\%$ gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field."},"_bibtex":{"value":"@inproceedings{\nli2026beyond,\ntitle={Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models},\nauthor={Nanxi Li and Xiang Wang and Yuanjie Chen and Haode Zhang and Hong Li and Yong-Lu Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Ax02eR2c3d}\n}"},"title":{"value":"Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models"},"pdf":{"value":"/pdf/3b199b74bbdbddb8166b5e799fc00ddb71a413ae.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|beyond_static_vision_scene_dynamic_field_unlocks_intuitive_physics_understanding_in_multimodal_large_language_models"},"authorids":{"value":["~Nanxi_Li1","~Xiang_Wang19","~Yuanjie_Chen2","~Haode_Zhang2","~Hong_Li5","~Yong-Lu_Li1"]},"authors":{"value":["Nanxi Li","Xiang Wang","Yuanjie Chen","Haode Zhang","Hong Li","Yong-Lu Li"]}},"version":2},{"content":{"summary":{"value":"This paper presents a synthetic data generation method that aims to provide a simpler alternative to established methods like SurvivalGAN, but it suffers from significant theoretical and practical flaws that undermine its contribution. The method lacks a clear causal motivation, relies on questionable assumptions, and introduces privacy risks. Its empirical results are unimpressive, and the supposed benefits over existing methods are either marginal or unsubstantiated. Overall, while the paper proposes a straightforward approach, the methodology lacks the rigor and depth needed to be a meaningful advancement in synthetic data generation for survival analysis.\n\nIn summary, while the paper presents a straightforward approach to synthetic data generation, it is marred by theoretical inconsistencies, privacy concerns, and underwhelming empirical results. The lack of alignment between assumptions and methodology, combined with the inadequate validation and privacy assessment, suggests that the approach is neither a substantial nor a practical improvement over existing methods."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"1. Could the authors clarify how their generation process aligns with the stated assumption of conditional independence between observed and censoring times given the covariates? Additionally, could they provide theoretical or empirical evidence demonstrating that their method preserves this property? This clarification may help address the apparent inconsistency between their assumptions and implementation, and potentially refine the methodological foundation of the paper.\n2. Could the authors conduct additional experiments to better highlight the strengths of their method, highlighting all the various metrics, including those in SurvivalGAN? Additionally, could they provide a more in-depth analysis of the trade-offs between performance gains and computational costs, especially given that the large language model (LLM) used for generation shows only marginal improvements over non-LLM methods while incurring significantly higher generation times? \n3. Could the authors provide a more detailed discussion of how their method compares to SurvivalGAN in terms of balancing predictions across different racial groups, particularly given that SurvivalGAN appears to achieve more consistent results across subgroups?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The paper introduces a synthetic data generation approach that prioritizes simplicity and ease of implementation, offering a straightforward alternative to more complex methods like SurvivalGAN."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Below is a list of weaknesses.\n1. The sequence of modeling in this paper—generating from p(e), p(t∣e), and p(x∣t,e)—is poorly motivated and lacks a sound theoretical basis. Conditioning covariates on both time and event does not align with standard causal structures, where covariates would typically precede and influence both event and time. This choice suggests a misunderstanding of causal dependencies, undermining the model's credibility. Without a more thoughtful theoretical foundation, the proposed structure appears arbitrary and counterintuitive.\n\n2. Additionally, the generation method is overly simplistic, involving direct sampling from the empirical distribution for p(t) and p(e) and applying any conditional generative model for p(x∣t,e). While this approach may seem appealing due to its simplicity, it fails to capitalize on more sophisticated methods that could enhance performance or generalizability. Direct sampling from the empirical distribution introduces serious privacy concerns. Each synthetic sample directly mirrors the time and event distributions of real data, making this approach antithetical to the purpose of synthetic data, which is to offer privacy-preserving alternatives to real data. Although the authors employ the DCR metric to assess privacy, this measure lacks theoretical guarantees and is incompatible with robust privacy frameworks like differential privacy. Furthermore, while the authors acknowledge that kernel density estimation could reduce privacy risks, they do not explore this option experimentally, leaving questions about the practical utility and safety of their approach.\n\n3. The authors’ assumption that \"observed and censoring times are conditionally independent given the covariates\" is also problematic. Their generation process contradicts this assumption. If t⊥⊥e∣x, then they should be able to generate from both marginals independently. For instance, if the conditional generator of x given t and e is constant, then t and e will not be conditionally independent, thus breaking a fundamental assumption in survival analysis. This inconsistency between theoretical assumptions and practical implementation significantly weakens the methodological foundation of the paper.\n\n4. Building on the above, this assumption that observed and censoring times are conditionally independent given the covariates, raises significant concerns in real-world settings, where this assumption is often unrealistic, as censoring times can be linked to factors that influence survival times even when conditioned on covariates. For instance, in medical studies, a patient’s decision to exit a study (censoring) might be directly related to their declining health—a factor that also impacts survival—thus violating the assumption of conditional independence. Additionally, this assumption does not account for informative censoring, where censoring itself is indicative of an imminent event, such as a severe health deterioration. Ignoring this dependency can lead to substantial bias, as informative censoring demands adjustments beyond the simple conditional independence framework. Furthermore, the assumption introduces potential biases in models that involve complex covariate interactions, as it overlooks latent factors affecting both censoring and survival. This is especially problematic in high-dimensional settings, where covariate interactions are often intricate and less observable. By assuming independence, the model also loses flexibility in capturing dependencies within the joint distribution of censoring and survival times, leading to suboptimal performance when censoring mechanisms interact with the event in ways that covariates alone cannot account for. Finally, the assumption is difficult to validate, as it is practically impossible to confirm whether censored individuals would have different survival times. This lack of verifiability undermines the robustness of the analysis and may lead to overly optimistic model performance if the assumption does not hold in the underlying data. Together, these issues suggest that this assumption, while convenient, significantly limits the reliability and practical applicability of the method presented in the paper.\n\n5. The empirical results presented in this paper do little to substantiate the method’s efficacy. The improvements obtained through conditional sampling are marginal, with slim gains across covariate quality metrics and Brier scores, which calls into question the practical impact of this approach. For a method that claims to advance synthetic data generation, these results are underwhelming, as they demonstrate limited progress over simpler baselines. Additionally, the performance of the large language model (LLM) used for generation does not show a substantial advantage over non-LLM methods, and its generation time is many orders of magnitude slower. The time cost far outweighs any minor gains in performance, making this approach inefficient and unsuitable for practical use.\n\n6. Experiment 4.3, which highlights the supposed benefits of the proposed method for subgroup analysis, is also unclear. The authors do not adequately define what constitutes good performance for subgroup analysis, leaving the results ambiguous. Interestingly, the findings suggest that SurvivalGAN achieves more balanced downstream predictions across different racial groups, which could be seen as a point in favor of SurvivalGAN rather than the proposed approach. \n\n7. Importantly and totally ignored by this paper, SurvivalGAN offers a variety of performance metrics specifically tailored to synthetic data validation, whereas this paper relies on less appropriate metrics and a flawed validation mechanism. This lack of rigor in validation only further undermines the credibility of the proposed method."}},"nonreaders":[],"tmdate":1731428412769,"tcdate":1730551512131,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8040/Reviewer_pobY"],"signatures":["ICLR.cc/2025/Conference/Submission8040/Reviewer_pobY"],"forum":"fHqwCsDK1z","number":3,"license":"CC BY 4.0","cdate":1730551512131,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8040/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428412769,"domain":"ICLR.cc/2025/Conference","replyto":"fHqwCsDK1z","id":"KK7xtv5Rjc","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Survival Data Generation","Survival Analysis","Tabular Data Generation","Generative Modeling"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Synthetic data generation holds considerable promise, offering avenues to enhance privacy, fairness, and data accessibility. Despite the availability of various methods for generating synthetic tabular data, challenges persist, particularly in specialized applications such as survival analysis. One significant obstacle in survival data generation is censoring, which manifests as not knowing the precise timing of observed (target) events for certain instances. Existing methods face difficulties in accurately reproducing the real distribution of event times for both observed (uncensored) events and censored events, i.e., the generated event-time distributions do not accurately match the underlying distributions of the real data. So motivated, we propose a simple paradigm to produce synthetic survival data by generating covariates conditioned on event times (and censoring indicators), thus allowing one to reuse existing conditional generative models for tabular data without significant computational overhead, and without making assumptions about the (usually unknown) generation mechanism underlying censoring. We evaluate this method via extensive experiments on real-world datasets. Our methodology outperforms multiple competitive baselines at generating survival data, while improving the performance of downstream survival models trained on it and tested on real data. Importantly, our approach achieves these improvements without compromising patient privacy, offering a balanced solution for synthetic survival data generation."},"_bibtex":{"value":"@misc{\nashhad2025conditioning,\ntitle={Conditioning on Time is All You Need for Synthetic Survival Data Generation},\nauthor={Mohd Ashhad and Ricardo Henao},\nyear={2025},\nurl={https://openreview.net/forum?id=fHqwCsDK1z}\n}"},"title":{"value":"Conditioning on Time is All You Need for Synthetic Survival Data Generation"},"pdf":{"value":"/pdf/55551aed9063f70cd713bccd9fe0c77b3ff171c8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ashhad|conditioning_on_time_is_all_you_need_for_synthetic_survival_data_generation"},"authorids":{"value":["~Mohd_Ashhad1","~Ricardo_Henao1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mohd Ashhad","Ricardo Henao"]}},"version":2},{"content":{"summary":{"value":"This paper assesses whether physics-informed constraints improve near-shore chlorophyll-a (as a HAB proxy) forecasting in the California Current. It compares a ConvLSTM, a TFT, and a physics-guided ConvLSTM with a soft advection–diffusion loss, trained on ~4-km, 8-day composites with multiple dynamic/static drivers."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- 8-day target is under-motivated. Why 8? Is it inherited from MODIS composites or operational cadence? Would be good to compare with 7 or 10? How does that impact model performance?\n- The “binary floor flag” is mentioned but how/where it’s used isn’t obvious?\n- The paper has minimal baselines for comparison. Fourier Neural Operator / AFNO have been shown to work on ocean and climate data, why not use those to compare as baselines? Or use HABnet?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Clear hypothesis and goal of adding a physics loss to help forecast coastal chlorophyll/HABs.\n    \n- Physics guidance seems to help the most at near-term horizons and in spatial structure"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Performance of all models look the same and hardly enough gains quantitatively.\n    \n- Would be good to report uncertainty or statistical measure to further compare.\n    \n- Although the authors have provided qualitative analysis but it hardly compares all the baselines or shows sufficient evidence to understand why and how one is better.\n    \n- The paper is written in a sort of crytic manner from teh perspective of an ML reader, assuming much domain knowledge. It would be good to have an appendix section detailing all the introductory domain information, explaining datasets, ecological features and domain terms in detail.\n    \n- The paper has citations missing for the datasets mentioned. Such as ERA5/MODIS/CMEMS.\n    \n- In Fig. 2 the three models appear nearly indistinguishable; please quantify significance and show uncertainty bars.\n    \n- 8-day target is under-motivated. Why 8? Is it inherited from MODIS composites or operational cadence? Would be good to compare with 7 or 10? How does that impact model performance?\n    \n- The paper has minimal baselines for comparison. Fourier Neural Operator / AFNO have been shown to work on ocean and climate data, why not use those to compare as baselines? Or use HABnet?\n    \n- I am unsure of the novelty of this work, please elaborate on that. Right now it seems an applied work to me. While novelty is and should not be the main criteria, however this paper also lack insights of value. This work can benefit with more experiments and deeper discussions between various methods.\n    \n- The “binary floor flag” is mentioned but how/where it’s used isn’t obvious?\n    \n- Can you show more ablations such as: (i) raw vs monthly-anomaly features, (ii) dynamic vs static drivers, (iii) different λ schedules. Analyzing which ecological features impact the most or which are not of much use?\n    \nMinor\n    \n- Move domain background (drivers, sensors, acronyms) to a succinct “Data & Drivers” subsection with 3–4 line blurbs + references; explicitly cite ERA5, MODIS, GLORYS/CMEMS.\n    \n- Put full architecture and training hyperparams in Appendix: layer specs, kernels, stencils, normalization, masking, loss weights, schedulers, early stopping.\n    \n- Fig. 2 left should be color (grayscale makes curves indistinguishable), right panel is blurry even when zoomed.\n    \n- Include the structure-diagnostic plots (your Figs. 3–4 type) for all three models, not just the PINN, so it is clear if physics helps/hurts.\n    \n- In multi-column examples, clearly label which column is GT."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916524447,"tcdate":1761981569238,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3047/Reviewer_CrZH"],"signatures":["ICLR.cc/2026/Conference/Submission3047/Reviewer_CrZH"],"forum":"JS6H6R60I6","number":4,"license":"CC BY 4.0","cdate":1761981569238,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3047/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916524447,"domain":"ICLR.cc/2026/Conference","replyto":"JS6H6R60I6","id":"QH4wEEHAwQ","forumContent":{"TLDR":{"value":"This study develops physics-informed deep learning models that combine data to more accurately forecast harmful algal blooms along California’s coast, improving spatial fidelity and ecological insight for management and public health."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["harmful algal blooms","spatiotemporal modeling","physics-informed neural networks","ocean forecasting","remote sensing"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Harmful algal blooms (HABs) are increasing in frequency, duration, and extent along the California coast, driven by climate variability, nutrient enrichment, and complex physical–biogeochemical interactions. Forecasting HAB development and spread remains a challenge, especially in Eastern Boundary Upwelling Systems where advection, stratification, and episodic river inputs strongly shape bloom dynamics. Existing approaches often trade physical realism for statistical flexibility, limiting generalization across bloom regimes. We present a physics-informed deep learning framework for nearshore chlorophyll-a forecasting in the California Current System, integrating multi-sensor satellite products, atmospheric and ocean reanalysis fields, and static geospatial predictors. Three architectures are evaluated: a convolutional long short-term memory network (ConvLSTM), a Temporal Fusion Transformer (TFT), and a physics-informed ConvLSTM (PINN) incorporating the two-dimensional advection–diffusion equation as a soft training constraint. A multi-year, 4 km-resolution dataset (2003–2021) is processed via a tailored feature engineering pipeline with quality-controlled gap-filling, rolling statistics, lagged predictors, and climatology-based anomalies. Models are assessed with strict spatiotemporal cross-validation, emphasizing spatial fidelity, bloom footprint representation, and predictor interpretability. Post-hoc explainability analyses identify key environmental drivers consistent with known upwelling–bloom linkages in the region. We present comparative skill assessments, spatial bias analyses, and predictor attribution results, highlighting the advantages and trade-offs of adding physical constraints to coastal HAB forecasting models. This work delivers a scalable and transferable methodology with direct implications for ecosystem management, fisheries, and public health."},"_bibtex":{"value":"@misc{\nmohanty2026physicsguided,\ntitle={Physics-Guided Neural Forecasts of Nearshore Harmful Algal Blooms in the California Current System},\nauthor={Yashnil Mohanty},\nyear={2026},\nurl={https://openreview.net/forum?id=JS6H6R60I6}\n}"},"title":{"value":"Physics-Guided Neural Forecasts of Nearshore Harmful Algal Blooms in the California Current System"},"pdf":{"value":"/pdf/c08cc651a6ddeba4fd8a92f3b258a53d17814ed0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"mohanty|physicsguided_neural_forecasts_of_nearshore_harmful_algal_blooms_in_the_california_current_system"},"authorids":{"value":["~Yashnil_Mohanty1"]},"authors":{"value":["Yashnil Mohanty"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LogicSR, a unified benchmark for discovering logical expressions from data. It bridges the gap between symbolic regression, which targets continuous functions, and logic synthesis, which assumes complete, noiseless specifications. LogicSR combines real-world datasets from digital circuits and biological networks with a scalable synthetic generator that creates diverse Boolean formulas under noise and incompleteness. The authors evaluate 14 methods across four paradigms (logic synthesis, symbolic regression, neural, and LLM-based). Results show that traditional logic synthesis excels on small tasks but fails to generalize, symbolic and neural methods handle larger scales at higher cost, and current LLMs struggle with complex logical reasoning. LogicSR thus establishes a rigorous foundation for cross-domain evaluation in logical discovery."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- How would LogicSR extend to non-Boolean (e.g., multi-valued or fuzzy logic) functions?\n\n- How would LogicSR scale to more complex reasoning tasks with implications?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- Logical discovery is becoming crucial for interpretable and neuro-symbolic AI, yet lacks a unified benchmark.\nLogicSR clearly fills this gap, making the paper’s contribution broadly valuable.\n\n- The inclusion of both real-world and synthetic datasets, with controlled levels of noise and data incompleteness, enables evaluation under realistic conditions, which was missing in prior symbolic regression or logic synthesis work.\n\n- The study connects traditionally isolated communities (EDA, symbolic regression, neuro-symbolic AI, LLMs).\nThe quantitative analysis across scales, operators, and noise levels gives an excellent overview of capability boundaries.\n\n- The paper details generation algorithms, metrics, and reproducibility measures (open-source plan, parameter documentation), making the benchmark credible and replicable.\n\n- The observation that LLMs (even GPT-o3) fail on complex Boolean reasoning is both surprising and informative for the broader research community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although logical discovery is interesting in general, the paper could be strengthened by demonstrating concrete downstream use cases, such as how the discovered logical expressions could enhance interpretability in scientific modeling, improve circuit design efficiency, or support symbolic reasoning in neuro-symbolic systems. Without such examples, the broader practical impact remains abstract, and elaborated discussions are encouraged.\n\n- The benchmark currently excludes XOR, NAND, implication, or higher-arity operators, which limits expressiveness and generalization analysis.\n\n- The LLMs are only tested with single-prompt fitting tasks. Multi-step prompting or program-of-thought reasoning (which could improve symbolic regression) is not considered, leaving the evaluation somewhat incomplete.\n\n- The two-stage synthetic data generator is central to the paper, but the influence of its parameters (priority decay, layer weighting, merging strategy) is not systematically analyzed.\n\n- While some limitations can be inferred, they are not explicitly stated. An elaborated discussion on limitations or failure cases of the proposed approach would strengthen the paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919215992,"tcdate":1761865066560,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7001/Reviewer_12eU"],"signatures":["ICLR.cc/2026/Conference/Submission7001/Reviewer_12eU"],"forum":"8ixcRzGuff","number":3,"license":"CC BY 4.0","cdate":1761865066560,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7001/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919215992,"domain":"ICLR.cc/2026/Conference","replyto":"8ixcRzGuff","id":"uNjeuNnbye","forumContent":{"TLDR":{"value":"LogicSR is a unified benchmark for logical symbolic regression with real-world and synthetic datasets. It evaluates 14 methods on their conciseness, accuracy and efficiency, revealing current limitations and directions for neuro-symbolic AI."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Symbolic Regression","Logical Reasoning","Neuro-symbolic Learning","Benchmark Dataset","Boolean Expressions"]},"supplementary_material":{"value":"/attachment/a1169543d88070d9a92e64a93c156463051de602.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Discovering underlying logical expressions from data is a critical task for interpretable AI and scientific discovery, yet it remains poorly served by existing research infrastructure. The field of Symbolic Regression (SR) primarily focuses on continuous mathematical functions, while Logic Synthesis (LS) is designed for exact, noise-free specifications, not for learning from incomplete or noisy data. This leaves a crucial gap for evaluating algorithms that can learn generalizable logical rules in realistic scenarios. To address this, we introduce LogicSR, a large-scale and comprehensive benchmark for logical symbolic regression. LogicSR is built from two sources: real-world problems from digital circuits and biological networks, and a novel synthetic data generator capable of producing a diverse set of complex logical formulas at scale. We use LogicSR to conduct a rigorous evaluation of 17 algorithms, spanning classical logic solvers, modern machine learning models, and Large Language Models (LLMs). Our findings reveal that the logical modeling capabilities and generalization robustness of these algorithms significantly depend on task scale and logical complexity, with current cutting-edge LLMs showing limited complex logical reasoning ability. LogicSR provides a robust foundation to benchmark progress, unify evaluation across disparate fields, and steer the future development of powerful neuro-symbolic systems."},"_bibtex":{"value":"@misc{\nanonymous2026logicsr,\ntitle={Logic{SR}: A Unified Benchmark for Logical Discovery from Data},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=8ixcRzGuff}\n}"},"title":{"value":"LogicSR: A Unified Benchmark for Logical Discovery from Data"},"pdf":{"value":"/pdf/67697adbac3b3a1e6e7bbadb5de5a27d446714e5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"zhang|logicsr_a_unified_benchmark_for_logical_discovery_from_data"},"authorids":{"value":["~Zimeng_Zhang1","~Xin_Zheng21","~Feifei_Zhang4","~Yunxin_Liu2","~Yuanchun_Li1"]},"authors":{"value":["Zimeng Zhang","Xin Zheng","Feifei Zhang","Yunxin Liu","Yuanchun Li"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewer for their detailed feedback. \n\n> determine the limitations that exist in similar proposals in the literature\n\nWe would like to point that we discuss prior benchmarks like PhysBench, Intphys Rlbench, OpenAI Gym, Planbench and Minerva  in the related work section. We also extensively discuss various drawbacks of prior benchmarks like Kinetix, Xland and Minecraft (see last paragraph of related work section) that are fixed by BuilderBench.\n\nIn Appendix C and Table 1 we have added a comparison of BuilderBench with many (including the ones mentioned by the reviewer) prior benchmarks. We will summarize the main highlight below:\n\n* Task-suite:  We have provided a task suite which\nrequires skills such as reasoning about commutativity and associativity of pick and place ordering, maximizing overhangs, packing problems, intuitive physics, counterweights, buttresses, and\ntemporary scaffolding. **Building such tasks is non trivial and essential for evaluating zero shot reasoning from scratch.**\n\n* Fast interaction **(3x-25x)** : Training agents to solve these tasks will presumably require large amounts of interaction. Hence fast simulators are necessary. Even if our tasks are replicated in other hardware accelarated simulators, scaling the number of objects in a scene greatly slows the simulation down. This is a well known problem. We significantly reduced this speed drop using a combination of MuJoCo's cpu threading and end to end jitting, similar to EnvPool XLA. To show this, we have added a simple speed test of BuilderBench with ManiSkill, Genesis and MuJoCo MJX when simulating 10 cubes and a robot:\n\n| Simulator  | Simulation Speed with 10 cubes (Frames per Second) |\n|---------|----------|\n| BuilderBench     |  312461  |\n| MuJoCo MJX     | 12376   |\n| Genesis | 18466 |\n| Maniskill | 98084 |\n\nThese comparisons clearly highlight the limitations that exist in similar proposals in the literature shows the need for BuilderBench.\n\n> Craftium\n\nWe thank the reviewer for pointing this work out. While we have cited Minecraft, we weren't aware of the Craftium project. We have added this both to the related work section and Appendix C and Table 1. We note several points that make BuilderBench particularly appealing for research on open-ended exploration, zero shot reasoning, and learning purely from trial and error\n\n* Task-suite: We have provided a task suite which requires skills such as reasoning about commutativity and associativity of pick and place ordering, maximizing overhangs, packing problems, intuitive physics, counterweights, buttresses, and temporary scaffolding. Building such tasks is non trivial and essential for evaluating zero shot reasoning from scratch. \n\n* Distinct Physics Domain: BuilderBench focuses on tasks requiring continuous physics reasoning (e.g., maximizing overhangs, utilizing counterweights, creating temporary scaffolding). These tasks are impossible to set up in the discrete, block-based world of Craftium.\n\n* Solving tasks whose solutions are unkown to us: In BuilderBench, we provide some stable final structures which we don't know how to stably build. Such tasks are not framed in Craftium. These are exactly the tasks we want machine learning models to solve. BuilderBench provides a concrete experimental instantiation of such tasks.\n\n* Most tasks in minecraft / craftium are programmable -- that is, one could write a program to solve them (killing monsters, searching for diamonds etc.). But there are many tasks in BuilderBench who's solutions do not seem to have simple programs (high Kolmogorov complexity). For instance scaffolding or generating counterweights. \n\n* Learning purely from trial and error: Training agents to solve such complex tasks will presumably require large amounts of interaction. Hence fast simulators are necessary. Craftium is not hardware accelerated. Based on Figure 7 in the Craftium paper, the BuilderBench simulator is ~400 times faster. This boost will drastically increase the speed of algorithmic iteration.\n\n**We believe the new detailed comparisons in the revised manuscript and the clear differentiation from Craftium, directly address the reviewer's concerns regarding the limitations of prior work and the novelty of our contribution. We are confident that BuilderBench provides a necessary and distinct benchmark for open-ended exploration and zero shot learning.**"},"title":{"value":"Reply by Authors"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763605129448,"tcdate":1763605080660,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19521/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission19521/Authors"],"forum":"oGbzk8xuVT","number":8,"license":"CC BY 4.0","cdate":1763605080660,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19521/-/Official_Comment"],"mdate":1763605129448,"domain":"ICLR.cc/2026/Conference","replyto":"dIgVYFMNJO","id":"AmNScadYdn","forumContent":{"TLDR":{"value":"BuilderBench is benchmark for research towards generalist agents that learn to solve diverse and complex tasks via interaction with a fast and open-ended environment"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["reinforcement learning","unsupervised environment design","open-endedness","benchmark","generalist agents."]},"supplementary_material":{"value":"/attachment/ff2e2a6d87ae6798eabaac805bf0928ccf0587c0.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"Today's AI models learn primarily through mimicry and sharpening, so it is not surprising that they struggle to solve problems beyond the limits set by existing data. To solve novel problems, agents should acquire skills for exploring and learning through experience. Finding a scalable learning mechanism for developing agents that learn through interaction remains a major open problem. In this work, we introduce BuilderBench, a benchmark to accelerate research into agent pre-training that centers open-ended exploration. BuilderBench requires agents to learn how to build any structure using blocks. BuilderBench is equipped with $(1)$ a hardware accelerated simulator of a robotic agent interacting with various physical blocks, and $(2)$ a task-suite with over 42 diverse target structures that are carefully curated to test an understanding of physics, mathematics, and long-horizon planning. During training, agents have to explore and learn general principles about the environment without any external supervision. During evaluation, agents have to build the unseen target structures from the task suite. Solving these tasks requires a sort of \\emph{embodied reasoning} that is not reflected in words but rather in actions, experimenting with different strategies and piecing them together. Our experiments show that many of these tasks challenge the current iteration of algorithms. Hence, we also provide a ``training wheels'' protocol, in which agents are trained and evaluated to build a single target structure from the task suite. Finally, we provide single-file implementations of six different algorithms as a reference point for researchers."},"_bibtex":{"value":"@misc{\nghugare2026builderbench,\ntitle={BuilderBench -- A benchmark for generalist agents},\nauthor={Raj Ghugare and Catherine Ji and Kathryn Wantlin and Jin Schofield and Benjamin Eysenbach},\nyear={2026},\nurl={https://openreview.net/forum?id=oGbzk8xuVT}\n}"},"title":{"value":"BuilderBench -- A benchmark for generalist agents"},"pdf":{"value":"/pdf/ff3d667b80238409f15a1da01b0e63557797e4b5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"ghugare|builderbench_a_benchmark_for_generalist_agents"},"authorids":{"value":["~Raj_Ghugare1","~Catherine_Ji1","~Kathryn_Wantlin1","~Jin_Schofield1","~Benjamin_Eysenbach1"]},"authors":{"value":["Raj Ghugare","Catherine Ji","Kathryn Wantlin","Jin Schofield","Benjamin Eysenbach"]}},"version":2},{"content":{"summary":{"value":"This paper presents a physics-inspired neural network for estimating the location of an acoustical source using the deformations of optical fibres as input. This is an innovative approach, and I have not seen similar approaches before. The physical basis is presented clearly and with appropriate detail. The text and grammar are mostly good, but variables, acronyms and figures are not thoroughly described. The balance between thorough physical modelling and brief experiments feels awkward."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"All my questions have been resolved in the revision."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Contributions of the article are very clearly specified\n- The connection between physics and neural networks in the application scenario is a strong contribution"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Major issues:\nFirst, the correspondence between the real-world location of fibre optic cables and the Halton distribution is weak at best. Second, we probably do not have the accurate location of a fibre optic cable. Yet both the distribution and accurate location of measurement points are required for accurate results. Thus, the correspondence between the experiment and real-world scenarios is weak. More realistic scenarios would have helped.\n- This paper claims to be about using fibre optic cables as microphones, but it contains only experiments with piece-wise linear and random arrays. The only connection to fibre optics is the underlying physics model. Experiments closer to fibre optics scenarios would help a lot.\n- The balance of the paper is skewed since the physics side is strong, but experiments are brief.\n\nMinor issues:\n- All text, and especially equations in Figure 2, are very small.\n- Interpretation of symbols in Fig 3-4 are not explained."}},"nonreaders":[],"tmdate":1732796213016,"tcdate":1730659476722,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11052/Reviewer_gyvL"],"signatures":["ICLR.cc/2025/Conference/Submission11052/Reviewer_gyvL"],"forum":"S2WUJUETyc","number":4,"license":"CC BY 4.0","cdate":1730659476722,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11052/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732796213016,"domain":"ICLR.cc/2025/Conference","replyto":"S2WUJUETyc","id":"YHNwv83gN7","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["distributed acoustic sensing","physics informed neural networks","room acoustics","sound source localization"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Distributed Acoustic Sensing (DAS) is an emerging technology that transforms standard optical fibers into dense arrays of acoustic sensors, offering unprecedented opportunities for smart city applications, indoor monitoring of human activity, and surveillance without compromising privacy. In this paper, we integrate DAS with Physics-Informed Neural Networks (PINNs) for indoor sound source localization. By embedding the acoustic wave equation and impedance boundary conditions into the neural network architecture, we exploit physical laws to guide the learning process, improving accuracy and generalization. We propose two strategies for real-time sound source localization using DAS data. The first strategy involves training the PINN on all available data simultaneously, while the second strategy incrementally feeds data over time, simulating real-time data acquisition. Using real indoor DAS measurements, we demonstrate the effectiveness of our approach in deciphering complex room acoustics and accurately inferring sound source locations under both strategies. Our framework provides a novel solution for real-time indoor positioning and human activity surveillance, offering significant advantages over traditional camera-based systems by preserving individual privacy."},"_bibtex":{"value":"@misc{\ngu2025integrating,\ntitle={Integrating Distributed Acoustic Sensing and {PINN} Frameworks for Enhanced Indoor Sound Source Localization},\nauthor={Chen Gu and Zhuoyu Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=S2WUJUETyc}\n}"},"title":{"value":"Integrating Distributed Acoustic Sensing and PINN Frameworks for Enhanced Indoor Sound Source Localization"},"pdf":{"value":"/pdf/f562aa2bbe114949d59d94e1600379022031ad74.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"gu|integrating_distributed_acoustic_sensing_and_pinn_frameworks_for_enhanced_indoor_sound_source_localization"},"authorids":{"value":["~Chen_Gu1","~Zhuoyu_Chen2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chen Gu","Zhuoyu Chen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes an optimization method for control parameter search in physics-based models. The method relies on the categorical reparametrization of the control parameters through Gumbel-Softmax, simultaneously enabling non-deterministic search across the control parameter space, and  improving the robustness of gradient descent under the presence of noise.\nThe method is evaluated on a two-pool proton exchange model using an MLP, evaluated under varying noise levels and dataset sizes. Compared to a \"conventional\" approach (that does not rely on the Gumbel-Softmax reparametrization), the proposed method demonstrates better robustness to noise and improved accuracy,"},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- Can you clarify what are the characteristics of the  \"conventional\" method used as a baseline?\n- Could you clarify how does this method extend to other use cases (physics-based models and beyond)?\n- What do you see as the main contribution of the paper? Is it the overall optimization methodology or its applicability to the considered application?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"This paper presents a concise method for parameter search in noisy environments, demonstrating its potential on a complex physical system simulator. The simplicity of the proposed approach contributes to its robustness, making it promising for applications that demand accurate parameter estimation under noise and under a large search space.\n\nNotably, the method simultaneously optimizes the parameters of a neural network, which estimates system properties, alongside the control parameters defining the physical system. This integrated approach appears well-suited to the challenges in the targeted application domain, addressing both noise and the inherent difficulty of optimizing control parameters in complex systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper lacks clear contextualization relative to prior work, omitting key references to research on soft labels and reparameterization in search problems—well-studied subjects in ML and optimization. \n- Validation is limited to a single physics-based model, narrowing its scope, with unclear implications for generalization to related systems. The exclusive use of simulated data in experiments further restricts claims of noise robustness. \n- Additionally, the method is only compared to an undefined \"conventional\" approach, with no comparisons to other ML-based methods in the literature, making its contribution to the field unclear."}},"nonreaders":[],"tmdate":1731427530158,"tcdate":1730755721900,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2036/Reviewer_zeR2"],"signatures":["ICLR.cc/2025/Conference/Submission2036/Reviewer_zeR2"],"forum":"m9BiWVTJDx","number":3,"license":"CC BY 4.0","cdate":1730755721900,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2036/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427530158,"domain":"ICLR.cc/2025/Conference","replyto":"m9BiWVTJDx","id":"gH6f31aUGo","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["categorical reparameterization","Gumbel Softmax","Optimization","System properties","physics model"]},"primary_area":{"value":"optimization"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Optimizing control parameters is crucial to estimate reliable tissue characteristics in quantitative MRI. Basically, multiple hardware parameters are simultaneously controlled to generate a signal from MRI system. Repetitive acquisitions with different control parameter combinations create distinct signal modulations and then tissue characteristics are deduced from prior knowledge of physics-based relationship among modulated signals, control parameters, and tissue characteristics. The choice of control parameters, which determines the attribute of signal modulation, directly impacts the inverse problem in tissue characteristic estimation. Thus, the multidimensional control parameter optimization remains an open research topic in MRI field for accurate analysis of tissue characteristics. Typically, optimal parameters are determined by iteratively updating sets of control parameters to maximize the estimation accuracy of the tissue characteristics. However, the conventional optimization process is restricted to explore only the vicinity of control parameters at the current iteration. Therefore, it could highly depend on initialization and current parameters, which might lead to inefficient search especially when noise is present in the system. In this work, to mitigate this limitation, we propose a novel Gumbel-Softmax-based optimization scheme that enables a probabilistic search across an expanding set of all candidates for each control parameter using categorical reparameterization. As a case study, the proposed method is employed to find optimal control parameters for quantitative MRI. We demonstrate that our Gumbel-Softmax-based optimization simultaneously explores the entire range of control parameters from early iterations and outperforms the conventional optimization approach on accuracy of MR tissue characteristic estimation and repeatability of optimization, especially under noisy environments."},"_bibtex":{"value":"@misc{\nkang2025a,\ntitle={A Probabilistic Approach to Optimizing Hardware Control Parameters in System Property Estimation using Gumbel-Softmax},\nauthor={Beomgu Kang and Hyunseok Seo},\nyear={2025},\nurl={https://openreview.net/forum?id=m9BiWVTJDx}\n}"},"title":{"value":"A Probabilistic Approach to Optimizing Hardware Control Parameters in System Property Estimation using Gumbel-Softmax"},"pdf":{"value":"/pdf/5a14859f66ea6c98843528dc320ebf3d2780557f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"kang|a_probabilistic_approach_to_optimizing_hardware_control_parameters_in_system_property_estimation_using_gumbelsoftmax"},"authorids":{"value":["~Beomgu_Kang1","~Hyunseok_Seo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Beomgu Kang","Hyunseok Seo"]}},"version":2},{"content":{"venue":{"value":"ICDSP 2022"},"venueid":{"value":"dblp.org/conf/ICDSP2/2022"},"paperhash":{"value":"cheng|speaker_adaption_with_intuitive_prosodic_features_for_statistical_parametric_speech_synthesis"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Pengyu_Cheng:","~Zhen-Hua_Ling1"]},"html":{"value":"https://doi.org/10.1145/3529570.3529602"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icdsp2/ChengL22,\n  author={Pengyu Cheng and Zhen-Hua Ling},\n  title={Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis},\n  year={2022},\n  cdate={1640995200000},\n  pages={187-193},\n  url={https://doi.org/10.1145/3529570.3529602},\n  booktitle={ICDSP},\n  crossref={conf/icdsp2/2022}\n}\n"},"abstract":{"value":"In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, speech rate and energy considering that they are directly related with the overall prosodic characteristics of different speakers. The intuitive prosodic features are extracted at utterance-level or speaker-level, and are further integrated into the existing speaker-encoding-based and speaker-embedding-based adaptation frameworks respectively. The acoustic models are sequence-to-sequence ones based on Tacotron2. Intuitive prosodic features are concatenated with text encoder outputs and speaker vectors for decoding acoustic features. Experimental results have demonstrated that our proposed methods can achieve better objective and subjective performance than the baseline methods without intuitive prosodic features. Besides, the proposed speaker adaption method with utterance-level prosodic features has achieved the best similarity of synthetic speech among all compared methods."},"title":{"value":"Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis"},"authors":{"value":["Pengyu Cheng","Zhen-Hua Ling"]}},"tmdate":1727705570975,"pdate":1640995200000,"tcdate":1727703358142,"writers":["~"],"signatures":["~Zhen-Hua_Ling1"],"forum":"5HNW2unAiC","license":"CC BY-SA 4.0","number":115608,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727705570975,"domain":"DBLP.org","id":"5HNW2unAiC","version":2},{"content":{"venue":{"value":"CoRR 2022"},"pdf":{"value":"http://arxiv.org/pdf/2203.00951v1"},"venueid":{"value":"dblp.org/journals/CORR/2022"},"paperhash":{"value":"cheng|speaker_adaption_with_intuitive_prosodic_features_for_statistical_parametric_speech_synthesis"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Pengyu_Cheng:","~Zhen-Hua_Ling1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2203.00951"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2203-00951,\n  publtype={informal},\n  author={Pengyu Cheng and Zhen-Hua Ling},\n  title={Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis},\n  year={2022},\n  cdate={1640995200000},\n  journal={CoRR},\n  volume={abs/2203.00951},\n  url={https://doi.org/10.48550/arXiv.2203.00951}\n}\n"},"abstract":{"value":"In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, speech rate and energy considering that they are directly related with the overall prosodic characteristics of different speakers. The intuitive prosodic features are extracted at utterance-level or speaker-level, and are further integrated into the existing speaker-encoding-based and speaker-embedding-based adaptation frameworks respectively. The acoustic models are sequence-to-sequence ones based on Tacotron2. Intuitive prosodic features are concatenated with text encoder outputs and speaker vectors for decoding acoustic features.Experimental results have demonstrated that our proposed methods can achieve better objective and subjective performance than the baseline methods without intuitive prosodic features. Besides, the proposed speaker adaption method with utterance-level prosodic features has achieved the best similarity of synthetic speech among all compared methods."},"title":{"value":"Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis"},"authors":{"value":["Pengyu Cheng","Zhen-Hua Ling"]}},"tmdate":1727705570351,"pdate":1640995200000,"tcdate":1727703358142,"writers":["~"],"signatures":["~Zhen-Hua_Ling1"],"forum":"2YWmbb4uRc","license":"CC BY-SA 4.0","number":115607,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727705570351,"domain":"DBLP.org","id":"2YWmbb4uRc","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"https://arxiv.org/pdf/2412.00087v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"wang|onion_physicsinformed_deep_learning_model_for_line_integral_diagnostics_across_fusion_devices"},"authorids":{"value":["","","","","","","","","","","~Changqing_Zou1",""]},"html":{"value":"https://doi.org/10.48550/arXiv.2412.00087"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2412-00087,\n  publtype={informal},\n  author={Cong Wang and Weizhe Yang and Haiping Wang and Renjie Yang and Jing Li and Zhijun Wang and Xinyao Yu and Yixiong Wei and Xianli Huang and Zhaoyang Liu and Changqing Zou and Zhifeng Zhao},\n  title={ONION: Physics-Informed Deep Learning Model for Line Integral Diagnostics Across Fusion Devices},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2412.00087},\n  url={https://doi.org/10.48550/arXiv.2412.00087}\n}\n"},"abstract":{"value":"Rapid reconstruction of 2D plasma profiles from line-integral measurements is important in nuclear fusion. This paper introduces a physics-informed model architecture called Onion, that can enhance the performance of models and be adapted to various backbone networks. The model under Onion incorporates physical information by a multiplication process and applies the physics-informed loss function according to the principle of line integration. Prediction results demonstrate that the additional input of physical information improves the deep learning model's ability, leading to a reduction in the average relative error E_1 between the reconstruction profiles and the target profiles by approximately 0.84x10^(-2) on synthetic datasets and about 0.06x10^(-2) on experimental datasets. Furthermore, the implementation of the Softplus activation function in the final two fully connected layers improves model performance. This enhancement results in a reduction in the E_1 by approximately 1.06x10^(-2) on synthetic datasets and about 0.11x10^(-2) on experimental datasets. The incorporation of the physics-informed loss function has been shown to correct the model's predictions, bringing the back-projections closer to the actual inputs and reducing the errors associated with inversion algorithms. Besides, we have developed a synthetic data model to generate customized line-integral diagnostic datasets and have also collected soft x-ray diagnostic datasets from EAST and HL-2A. This study achieves reductions in reconstruction errors, and accelerates the development of surrogate models in fusion research."},"title":{"value":"ONION: Physics-Informed Deep Learning Model for Line Integral Diagnostics Across Fusion Devices"},"authors":{"value":["Cong Wang","Weizhe Yang","Haiping Wang","Renjie Yang","Jing Li","Zhijun Wang","Xinyao Yu","Yixiong Wei","Xianli Huang","Zhaoyang Liu","Changqing Zou","Zhifeng Zhao"]}},"tmdate":1772987594232,"pdate":1735603200000,"externalIds":["dblp:journals/corr/abs-2412-00087"],"tcdate":1772987587737,"writers":["~"],"signatures":["~Changqing_Zou2"],"forum":"NMlb3R13g3","license":"CC BY-SA 4.0","number":849891,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772987594232,"domain":"DBLP.org","id":"NMlb3R13g3","version":2},{"content":{"summary":{"value":"This paper provides a unified framework inspired the bracket-based dynamical system to analysis the oversmoothing problem in GNN. The past work may leverage the opposite physics concept such as the reversible processes, irreversible process and therefore it is not clear how such concept help to design GNNs, while the framework in this paper gives a deeper understanding. Leveraging the data-driven exterior calculus, this paper constructs four novel architectures which span the both reversibility and irreversibility spectrum using geometric brackets as a means of parameterizing dynamics abstractly without empirically assuming a physical model. Interestingly, it can reinterpretate the message-passing  and GAT as the fluxes and conservation balances of physics simulator.  It also generalize the attention mechanism which extends the nodal feature to higher order cliques and provide a unified evaluation of dissipation.  Emperically, the author compare the architectures derived from this framework with classical GNN,e.g., GAT, GDE over several benchmarks."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. This paper is well-written and offers valuable insights into the construction of graph neural networks. As far as I am aware, the idea presented in this paper is innovative. However, since I lack a background in physics, I would appreciate hearing the suggestions of other reviewers who are familiar with bracket-based dynamical systems.\n\n2. The experimental result in damped double pendulum and MuJoCO dynamics looks good.\n\n3. The author made the code available, promoting reproducibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This paper may be a little hard to understand for the reader without physics background.\n\n2. It seems that the algorithm proposed by the author does not outperform the baselines in node classification problem."},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"I am curious about the precise computational complexity of the architectures proposed by the author, specifically in the context of N nodes or N neighbors. It would be beneficial to have concrete values rather than just the order of complexity. Additionally, it would be valuable to understand the advantages of these architectures in comparison to traditional graph neural networks (GNNs).\n\nRegarding the node classification experiment, the algorithm's performance is generally comparable to the baseline, although it does appear weaker in some cases. However, it showcases excellent performance in physical systems like the double pendulum and MuJoCo dynamics. Could this be attributed to the notion that physics-inspired frameworks are inherently more suitable for describing and modeling physical systems rather than social networks?\n"},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1702411450057,"tcdate":1688621871349,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission13826/Reviewer_dWFZ"],"signatures":["NeurIPS.cc/2023/Conference/Submission13826/Reviewer_dWFZ"],"forum":"4SoTUaTK8N","number":1,"license":"CC BY 4.0","cdate":1688621871349,"mdate":1702411450057,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission13826/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"4SoTUaTK8N","id":"vCjAN8NSZP","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["graph neural networks","structure preserving machine learning","neural ordinary differential equations","hamiltonian dynamics","metriplectic dynamics"]},"supplementary_material":{"value":"/attachment/2060a3f3fa2523cc0b79e36d00df7c48c79e733b.pdf"},"_bibtex":{"value":"@inproceedings{\ngruber2023reversible,\ntitle={Reversible and irreversible bracket-based dynamics for deep graph neural networks},\nauthor={Anthony Gruber and Kookjin Lee and Nathaniel Trask},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=4SoTUaTK8N}\n}"},"title":{"value":"Reversible and irreversible bracket-based dynamics for deep graph neural networks"},"paperhash":{"value":"gruber|reversible_and_irreversible_bracketbased_dynamics_for_deep_graph_neural_networks"},"TLDR":{"value":"Novel bracket-inspired GNN architectures are versatile and establish role of reversibility/irreversibility in stability of deep GNNs."},"abstract":{"value":"Recent works have shown that physics-inspired architectures allow the training of deep graph neural networks (GNNs) without oversmoothing. The role of these physics is unclear, however, with successful examples of both reversible (e.g., Hamiltonian) and irreversible (e.g., diffusion) phenomena producing comparable results despite diametrically opposed mechanisms, and further complications arising due to empirical departures from mathematical theory. This work presents a series of novel GNN architectures based upon structure-preserving bracket-based dynamical systems, which are provably guaranteed to either conserve energy or generate positive dissipation with increasing depth.  It is shown that the theoretically principled framework employed here allows for inherently explainable constructions, which contextualize departures from theory in current architectures and better elucidate the roles of reversibility and irreversibility in network performance. Code is available at the Github repository \\url{https://github.com/natrask/BracketGraphs}."},"pdf":{"value":"/pdf/2710528ccf5111376c98f9cdac5a6e5ac0fec1c5.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["adgrube@sandia.gov","~Kookjin_Lee1","~Nathaniel_Trask2"]},"authors":{"value":["Anthony Gruber","Kookjin Lee","Nathaniel Trask"]}},"version":2},{"content":{"summary":{"value":"The paper aims to address two issues in game generation: lack automated evaluation metrics and struggle with complex content. Specifically, this paper proposes AVR‑Eval, a relative, pairwise metric that compares two pieces of web‑based multimedia (games or animations) using audio‑visual recordings (AVRs) as inputs to an omni‑modal judge, followed by a text‑model review step. An ablation shows that multi‑round description → comparison and a text‑model review notably reduce failure modes. Besides, this paper proposes AVR‑Agent. In the first stage, the coding model selects which assets to use to produce the desired content given\nthe original description. In the second stage, the coding model is asked to generate the content based on the original description, chosen assets, general guidelines, and evaluation criteria. In the third stage, the content is improved over multiple steps including content description and feedback for the content. Empirical results demonstrate the effectiveness of the Agent and Eval."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"N/A"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. Target an interesting goal in achieving automated game design.\n2. Each component in the AVR-Eval or the AVR-Agent is evaluated carefully to demonstrate its effectiveness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Overall the paper is very engineering for designing pipeline and prompts for AVR-Eval and AVR-Agent, and lack main technical contribution.\n2. Not much related work being discussed in the paper so it's hard to place the paper in existing literature. \n3. While AVR‑Eval is intuitive and the ablation is convincing, there is no study of alignment with human raters.\n4. The benchmark uses five game and animations. Many results may not carry to richer game loops, content pipelines, or larger engine use (e.g., asset streaming, physics edge cases, level generation at scale). \n5. The writing should be improved to better separate the discussion about approach and results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931362303,"tcdate":1761984822307,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19440/Reviewer_bTx2"],"signatures":["ICLR.cc/2026/Conference/Submission19440/Reviewer_bTx2"],"forum":"IrGJvFKuX2","number":2,"license":"CC BY 4.0","cdate":1761984822307,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19440/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931362303,"domain":"ICLR.cc/2026/Conference","replyto":"IrGJvFKuX2","id":"lNTIqZ5fw8","forumContent":{"TLDR":{"value":"New metric for multimedia evaluation and multi-agent framework for video game generation"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video-game","llms","multi-agent","agent","animations"]},"supplementary_material":{"value":"/attachment/f6bed286ed5d90b4bdd88e0260c5f7eb97fea704.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Generating novel video games is  a challenging problem. Large Language Models (LLMs) can generate games and animations, but lack automated evaluation metrics and struggle with complex content. To tackle these issues, we built a new metric and multi-agent system. First, we propose AVR-Eval, a metric for multimedia content where a model compares the Audio-Visual Recordings (AVRs) of two contents and determines which one is better. We show that AVR-Eval properly identifies good from broken or mismatched content. Second, we built AVR-Agent, a multi-agent system to generate JavaScript code from a bank of multimedia assets (audio, images, 3D models) and using AVR feedback. We show higher  AVR-Eval with AVR-Agent than one-shot prompt. However, while humans benefit from high-quality assets and audio-visual feedback, they do not significantly increase AVR-Eval for LLMs. This reveals a gap between humans and AI content creation."},"_bibtex":{"value":"@misc{\njolicoeur-martineau2026multiagent,\ntitle={Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings},\nauthor={Alexia Jolicoeur-Martineau},\nyear={2026},\nurl={https://openreview.net/forum?id=IrGJvFKuX2}\n}"},"title":{"value":"Multi-Agent Game Generation and Evaluation via Audio-Visual Recordings"},"pdf":{"value":"/pdf/e4942e574975ea3d37767b049af18ea5f117b1c0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"jolicoeurmartineau|multiagent_game_generation_and_evaluation_via_audiovisual_recordings"},"authorids":{"value":["~Alexia_Jolicoeur-Martineau1"]},"authors":{"value":["Alexia Jolicoeur-Martineau"]}},"version":2},{"content":{"summary":{"value":"This paper investigates whether large language models (LLMs) are able to solve graph algorithm problems in natural language. A benchmark NLGraph contains 29,370 problems, covering 8 graph reasoning tasks with varying complexity from simple tasks such as connectivity, cycle, and shortest path to more complex problems such as topological sort, maximum flow, bipartite graph matching, Hamilton path, and simulating graph neural networks. Various prompting techniques have been applied to evaluate the capabilities of LLMs for performing complex reasoning on graphs. Several valuable insights are provided based on the experimental results and analyses. Two prompting techniques Build-a-Graph and Algorithmic Prompting are further introduced for enhancing graph reasoning."},"presentation":{"value":"4 excellent"},"contribution":{"value":"4 excellent"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The proposed NLGraph benchmark is comprehensive, covering intuitively simple to more sophisticated tasks. Also, as the benchmark is synthetic, answers are unlikely to appear in the pretraining corpus of LLMs, making it a more robust benchmark for evaluating complex reasoning with LLMs. I think it would be a challenging and valuable testbed for evaluating the graph reasoning abilities of large language models.\n2. Extensive experiments and analyses are conducted on the proposed benchmark. Specifically, results of text-davinci-003 model with various prompting techniques (e.g. Chain-of-Thoughts, Least-to-Most, Self-Consistency) are reported.  Such results and analyses help the research community to better understand the capabilities/behavior of LLMs for graph reasoning or even complex reasoning. The four key phenomenons summarized in the paper are quite counter-intuitive and thought-provoking.\n3. The paper is clear and well-organized.\n"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I only have several minor concerns/questions about this paper as follows:\n1. Using a programming style of prompting techniques (e.g. PAL[1], PoT[2]) for solving these graph reasoning tasks is more intuitive. It could potentially address the problem of generating too many tokens of code-davinci-002.  \n2. Except for the results of text-davinici-003, only part of the results of code-davinci-002 are reported (i.e. cycle, shortest path, and hamilton path).  For GPT-3.5-turbo and GPT4, only 19 tasks across 8 tasks are provided. I am not sure the four key findings listed are universal for the LLMs. It would be interesting to know the overall performance of other LLMs on the benchmark. It would shed light on this problem.  \n3. It seems to me that including the Graph Neural Networks (GNNs) task is not well-motivated. Why do we want the LLM to perform graph convolution on a two-dimension node embedding? We already have an efficient way to calculate it even for node embeddings that have hundreds or thousands of dimensions. If the LLM can use tools (e.g. calculator, python interpreter), why do we want it to perform large number calculations itself? Similarly, LLMs can write the code of different GNNs to perform message propagation.\n\n\nReferences:\n\n[1]PAL: Program-aided Language Models. \n\n[2]Program of Thoughts Prompting. "},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"See above."},"rating":{"value":"8: Strong Accept: Technically strong paper, with novel ideas, excellent impact on at least one area, or high-to-excellent impact on multiple areas, with excellent evaluation, resources, and reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"The authors adequately addressed the limitations in the Appendix."}},"nonreaders":[],"tmdate":1702411271958,"tcdate":1688663609577,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission10162/Reviewer_E2y1"],"signatures":["NeurIPS.cc/2023/Conference/Submission10162/Reviewer_E2y1"],"forum":"UDqHhbqYJV","number":3,"license":"CC BY 4.0","cdate":1688663609577,"mdate":1702411271958,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission10162/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"UDqHhbqYJV","id":"alWbjdXDEs","forumContent":{"venue":{"value":"NeurIPS 2023 spotlight"},"keywords":{"value":["large language models","graph reasoning","structured reasoning"]},"supplementary_material":{"value":"/attachment/6e3b61522a39dda6fc5831a757f63263134d6280.zip"},"_bibtex":{"value":"@inproceedings{\nwang2023can,\ntitle={Can Language Models Solve Graph Problems in Natural Language?},\nauthor={Heng Wang and Shangbin Feng and Tianxing He and Zhaoxuan Tan and Xiaochuang Han and Yulia Tsvetkov},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=UDqHhbqYJV}\n}"},"title":{"value":"Can Language Models Solve Graph Problems in Natural Language?"},"paperhash":{"value":"wang|can_language_models_solve_graph_problems_in_natural_language"},"abstract":{"value":"Large language models (LLMs) are increasingly adopted for a variety of tasks with implicit graphical structures, such as planning in robotics, multi-hop question answering or knowledge probing, structured commonsense reasoning, and more. While LLMs have advanced the state-of-the-art on these tasks with structure implications, whether LLMs could explicitly process textual descriptions of graphs and structures, map them to grounded conceptual spaces, and perform structured operations remains underexplored. To this end, we propose NLGraph (Natural Language Graph), a comprehensive benchmark of graph-based problem solving designed in natural language. NLGraph contains 29,370 problems, covering eight graph reasoning tasks with varying complexity from simple tasks such as connectivity and shortest path up to complex problems such as maximum flow and simulating graph neural networks. We evaluate LLMs (GPT-3/4) with various prompting approaches on the NLGraph benchmark and find that 1) language models do demonstrate preliminary graph reasoning abilities, 2) the benefit of advanced prompting and in-context learning diminishes on more complex graph problems, while 3) LLMs are also (un)surprisingly brittle in the face of spurious correlations in graph and problem settings. We then propose Build-a-Graph Prompting and Algorithmic Prompting, two instruction-based approaches to enhance LLMs in solving natural language graph problems. Build-a-Graph and Algorithmic prompting improve the performance of LLMs on NLGraph by 3.07% to 16.85% across multiple tasks and settings, while how to solve the most complicated graph reasoning tasks in our setup with language models remains an open research question."},"pdf":{"value":"/pdf/756b1f7b1f822853cd68c1925e52eb13bdda2e78.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Heng_Wang10","~Shangbin_Feng1","~Tianxing_He1","~Zhaoxuan_Tan1","~Xiaochuang_Han1","~Yulia_Tsvetkov1"]},"authors":{"value":["Heng Wang","Shangbin Feng","Tianxing He","Zhaoxuan Tan","Xiaochuang Han","Yulia Tsvetkov"]}},"version":2},{"content":{"venue":{"value":"Complex Intell. Syst. 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s40747-025-01806-y.pdf"},"venueid":{"value":"dblp.org/journals/COMINTSYS/2025"},"paperhash":{"value":"huang|\\text_h^2\\text_can_heterogeneous_hypergraph_attention_network_with_counterfactual_learning_for_multimodal_sentiment_analysis"},"authorids":{"value":["","","~Qionghao_Huang1","~Xiaodi_Huang3","",""]},"html":{"value":"https://doi.org/10.1007/s40747-025-01806-y"},"_bibtex":{"value":"@article{DBLP:journals/comintsys/HuangLHHJC25,\n  author={Changqin Huang and Zhenheng Lin and Qionghao Huang and Xiaodi Huang and Fan Jiang and Jili Chen},\n  title={$$\\text {H}^2\\text {CAN}$$: heterogeneous hypergraph attention network with counterfactual learning for multimodal sentiment analysis},\n  year={2025},\n  cdate={1735689600000},\n  journal={Complex Intell. Syst.},\n  volume={11},\n  number={4},\n  url={https://doi.org/10.1007/s40747-025-01806-y}\n}\n"},"abstract":{"value":"Multimodal sentiment analysis (MSA) has garnered significant attention for its immense potential in human-computer interaction. While cross-modality attention mechanisms are widely used in MSA to capture inter-modality interactions, existing methods are limited to pairwise interactions between two modalities. Additionally, these methods can not utilize the causal relationship to guide attention learning, making them susceptible to bias information. To address these limitations, we introduce a novel method called Heterogeneous Hypergraph Attention Network with Counterfactual Learning \\((\\text {H}^2\\text {CAN}).\\) The method constructs a heterogeneous hypergraph based on sentiment expression characteristics and employs Heterogeneous Hypergraph Attention Networks (HHGAT) to capture interactions beyond pairwise constraints. Furthermore, it mitigates the effects of bias through a Counterfactual Intervention Task (CIT). Our model comprises two main branches: hypergraph fusion and counterfactual fusion. The former uses HHGAT to capture inter-modality interactions, while the latter constructs a counterfactual world using Gaussian distribution and additional weighting for the biased modality. The CIT leverages causal inference to maximize the prediction discrepancy between the two branches, guiding attention learning in the hypergraph fusion branch. We utilize unimodal labels to help the model adaptively identify the biased modality, thereby enhancing the handling of bias information. Experiments on three mainstream datasets demonstrate that \\(\\text {H}^2\\text {CAN}\\) sets a new benchmark."},"title":{"value":"$$\\text {H}^2\\text {CAN}$$: heterogeneous hypergraph attention network with counterfactual learning for multimodal sentiment analysis"},"authors":{"value":["Changqin Huang","Zhenheng Lin","Qionghao Huang","Xiaodi Huang","Fan Jiang","Jili Chen"]}},"tmdate":1784547982429,"pdate":1767139200000,"tcdate":1776728919852,"writers":["~"],"signatures":["~Xiaodi_Huang3"],"forum":"1FKytspnki","license":"CC BY-SA 4.0","number":868555,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1784547982429,"domain":"DBLP.org","id":"1FKytspnki","version":2},{"content":{"summary":{"value":"The paper introduces the Spectral-Inspired Neural Operator (SINO), a novel architecture designed to learn PDE dynamics from extremely limited data (as few as 2-5 trajectories) without prior knowledge of the governing equations. Inspired by classical spectral methods, SINO uses modules to learn spectral multipliers for derivatives and nonlinear interactions (with de-aliasing). SINO is evaluated on 2D and 3D PDE benchmarks, outperforming several data-driven & physics-encoded baselines, and demonstrating robust out-of-distribution generalization."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"See Weaknesses above"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- The OOD generalization and super-resolution experiments demonstrate that SINO is able to approximate the underlying operator using only a very limited number of trajectories.\n- The architecture is motivated by first principles, resulting in a learnable (physics-agnostic) version of classical spectral solvers and providing a good inductive bias for low-data regimes.\n- The paper provides ablation studies confirming the necessity of each key component."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"There are two main concerns:\n\n1. In almost all experiments, the parameter count of the baselines is significantly larger than the one of SINO. Together with the relatively small grid of hyperparameters used for each method, it seems that all these methods are overfitting to the limited amount of training data. For a fair comparison, each baseline should be evaluated with several configurations that lead to a comparable parameter count. Moreover, it is unclear how much samples from each trajectory are taken for each method, in particular, since Table 3 shows different time-step sizes for DOL & POL methods.\n2. The proposed model is not a surrogate model for accelerating the numerical solution of PDEs. It incurs a similar inference time as a solver since it still relies on a RK4 integrator with sufficiently small steps. In this sense, there is no point in comparing to data-driven surrogate models. The focus of the paper should be on learning/discovering the underlying physical operator from limited data and comparing to other baselines in that domain. While the paper propose *operator distillation* as a workaround, this only shows that the underlying operator can be sufficiently well approximated by SINO such that it can act as a synthetic data generator. \n\nOther concerns:\n\n- Apart from the RK scheme used for time-stepping, the design of SINO is very similar to variants of FNOs, in particular https://arxiv.org/abs/2403.12553 for nonlinear mixing in the channel dimension and https://arxiv.org/pdf/2111.13587, https://arxiv.org/pdf/2403.03542 for MLPs in the Fourier space.\n- In its current form, the method is dependent on regular grids and periodic boundary conditions, excluding a vast number of real-world physics and engineering problems.\n- It would be good to evaluate other metrics such as spectra.\n- As mentioned above, only a few hyperparameter settings are tested for each baseline. In particular, this only includes a single training strategy (pushforward trick) with two learning rates.\n- The universality result is relatively standard and it seems to also follow from https://arxiv.org/abs/2304.13221\n- It is unclear how the low-pass filter is adapted in the super-resolution results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915831461,"tcdate":1761545799592,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1605/Reviewer_ZG9x"],"signatures":["ICLR.cc/2026/Conference/Submission1605/Reviewer_ZG9x"],"forum":"1GU1tqF2Ev","number":2,"license":"CC BY 4.0","cdate":1761545799592,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1605/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915831461,"domain":"ICLR.cc/2026/Conference","replyto":"1GU1tqF2Ev","id":"DKIhrUdIw5","forumContent":{"TLDR":{"value":"We introduce SINO, a neural operator that learns PDE dynamics from limited trajectories without any explicit PDE knowledge, achieving accurate inference, strong out‑of‑distribution generalization, and discretization invariance."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["PDE modeling","neural operators","spectral methods","AI for Physics"]},"supplementary_material":{"value":"/attachment/68ac10699e8e412c92fb453ec4479640b0ea2a38.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Learning PDE dynamics from limited data with unknown physics is challenging. Existing neural PDE solvers either require large datasets or rely on known physics (e.g., PDE residuals or handcrafted stencils), leading to limited applicability. To address these challenges, we propose Spectral-Inspired Neural Operator (SINO), which can model complex systems from just 2-5 trajectories, without requiring explicit PDE terms. Specially, SINO automatically captures both local and global spatial derivatives from frequency indices, enabling a compact representation of the underlying differential operators in physics-agnostic regimes. To model nonlinear effects, it employs a $\\Pi$-block that performs multiplicative operations on spectral features, complemented by a low-pass filter to suppress aliasing. Extensive experiments on both 2D and 3D PDE benchmarks demonstrate that SINO achieves state-of-the-art performance, with improvements of 1–2 orders of magnitude in accuracy. Particularly, with only 5 training trajectories, SINO outperforms data-driven methods trained on 1000 trajectories and remains predictive on challenging out-of-distribution cases where other methods fail."},"_bibtex":{"value":"@misc{\nwan2026spectralinspired,\ntitle={Spectral-inspired Operator Learning with Limited Data and Unknown Physics},\nauthor={Han Wan and Rui Zhang and Hao Sun},\nyear={2026},\nurl={https://openreview.net/forum?id=1GU1tqF2Ev}\n}"},"title":{"value":"Spectral-inspired Operator Learning with Limited Data and Unknown Physics"},"pdf":{"value":"/pdf/cc3212713f4ca0266d71c8212ecfb0ad6876619a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wan|spectralinspired_operator_learning_with_limited_data_and_unknown_physics"},"authorids":{"value":["~Han_Wan1","~Rui_Zhang22","~Hao_Sun4"]},"authors":{"value":["Han Wan","Rui Zhang","Hao Sun"]}},"version":2},{"content":{"comment":{"value":"> **Concern #5: Ablation Study**\n\nWe appreciate the reviewer's suggestion. To provide a comprehensive evaluation of the disentanglement capability, we extended the ablation study (originally Table 3) to include the CorrCoef metric for all photographic effects under the 'w/o Decoupled CA' setting. The results are presented in the table below. It can be observed that our full model consistently achieves superior performance across every photographic parameter compared to the naive approach ('w/o Decouple CA'). This confirms that the proposed mechanism effectively resolves the interference issues for all control signals.\n\n| Method | Bokeh | Zoom | Exposure | Color |\n| :--- | :---: | :---: | :---: | :---: |\n| w/o Decoupled CA | 0.4201 | 0.3975 | 0.4920 | 0.5031 |\n| **Full** | **0.5504** | **0.4550** | **0.5117** | **0.5176** |\n\n> **Concern #6: Details of Synthetic Data**\n\nWe have provided a comprehensive description of the synthetic data generation process in Section A.3 of the Supplementary Material. This section details the physical principles and specific rendering mechanisms for each photographic effect, as well as the parameter normalization strategies employed to ensure user-friendly control.\n\n> **Question #1 Base model**\n\nBoth the text-based baselines (in Tables 1 and 2) and our proposed CineCtrl are fine-tuned from the same backbone: the ReCamMaster model. The Stitching Baseline is constructed by stitching the physical rendering algorithms for all photographic effects and ReCamMaster for novel camera views rendering.\n\n> **Question #2 Subtle changes in qualitative results**\n\nWe respectfully disagree with the reviewer's observation that the qualitative effects produced by our method are subtle. The visual comparisons in Fig.4 and Fig.8, as well as the Supplementary Video, clearly demonstrate that our method is capable of generating distinct and significant photographic effects, showcasing superior fine-grained control and video quality compared to other baselines. Furthermore, the user study results in Table 4 serve as compelling evidence, confirming that human evaluators perceive a significant advantage in the visual performance and effect accuracy of our method.\n\n> **Question #3 Synthetic data**\n\nOur synthetic dataset is constructed such that each training sample contains exclusively a single photographic effect (e.g., a sample features only bokeh or only exposure adjustment). We do not include samples with complex mixtures of multiple effects in the training set. This isolated construction strategy was intentionally adopted to ensure the high visual fidelity of the training samples, preventing the quality degradation and artifacts that typically arise from stitching multiple simulation steps.\n\n> **Question #4 Meaning of Line 416**\nThe statement in Line 416 clarifies the rationale for excluding the stitching baseline from the quantitative evaluation. Calculating the CorrCoef requires a pseudo-Ground Truth (pseudo-GT) for each photographic effect to measure its correlation with the model's generated output. Since real-world videos lack intrinsic ground truth for these edited effects, we employ physics-based rendering methods to generate these pseudo-GTs. However, the stitching baseline itself is constructed using these exact same physics-based simulation algorithms. Consequently, evaluating it using CorrCoef would be tautological—essentially comparing the simulation against itself—which renders the metric meaningless. Therefore, to ensure a valid and fair evaluation, we exclude the stitching baseline from the CorrCoef comparison."},"title":{"value":"Response to Reviewer# WY3o (2/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764226293809,"tcdate":1764226293809,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2558/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission2558/Authors"],"forum":"sjIDaRwX84","number":12,"license":"CC BY 4.0","cdate":1764226293809,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2558/-/Official_Comment"],"mdate":1764226293809,"domain":"ICLR.cc/2026/Conference","replyto":"d1qCSLGvv6","id":"k0JiEaoO0G","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["diffusion model","video generative model","photographic effect"]},"supplementary_material":{"value":"/attachment/04590d448bd41209b999268a9edcf75054bd173e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Cinematic storytelling is profoundly shaped by the artful manipulation of photographic elements such as depth of field and exposure. These effects are crucial in conveying mood and creating aesthetic appeal.  However, controlling these effects in generative video models remains highly challenging, as most existing methods are restricted to camera motion control. In this paper, we propose CineCtrl, the first video cinematic editing framework that provides fine control over professional camera parameters (e.g., bokeh, shutter speed). \nWe introduce a decoupled cross-attention mechanism to disentangle camera motion from photographic inputs, allowing fine-grained, independent control without compromising scene consistency. To overcome the shortage of training data, we develop a comprehensive data generation strategy that leverages simulated photographic effects with a dedicated real-world collection pipeline, enabling the construction of a large-scale dataset for robust model training. Extensive experiments demonstrate that our model generates high-fidelity videos with precisely controlled, user-specified photographic camera effects."},"_bibtex":{"value":"@misc{\nsun2026generative,\ntitle={Generative Photographic Control for Scene-Consistent Video Cinematic Editing},\nauthor={Huiqiang Sun and Liao Shen and Zhan Peng and Kun Wang and Size Wu and Yuhang Zang and Tianqi Liu and Zihao Huang and Xingyu Zeng and Zhiguo Cao and Wei Li and Chen Change Loy},\nyear={2026},\nurl={https://openreview.net/forum?id=sjIDaRwX84}\n}"},"title":{"value":"Generative Photographic Control for Scene-Consistent Video Cinematic Editing"},"pdf":{"value":"/pdf/32664c104c63ec547ef8a47aed8eaa73c563489d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"sun|generative_photographic_control_for_sceneconsistent_video_cinematic_editing"},"authorids":{"value":["~Huiqiang_Sun1","~Liao_Shen1","~Zhan_Peng1","~Kun_Wang8","~Size_Wu1","~Yuhang_Zang1","~Tianqi_Liu3","~Zihao_Huang2","~Xingyu_Zeng1","~Zhiguo_Cao1","~Wei_Li51","~Chen_Change_Loy2"]},"authors":{"value":["Huiqiang Sun","Liao Shen","Zhan Peng","Kun Wang","Size Wu","Yuhang Zang","Tianqi Liu","Zihao Huang","Xingyu Zeng","Zhiguo Cao","Wei Li","Chen Change Loy"]}},"version":2},{"content":{"venue":{"value":"CoRR 2022"},"pdf":{"value":"http://arxiv.org/pdf/2209.03984v2"},"venueid":{"value":"dblp.org/journals/CORR/2022"},"paperhash":{"value":"costabal|pinns_physicsinformed_neural_networks_on_complex_geometries"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Francisco_Sahli_Costabal:","~Simone_Pezzuto1","https://dblp.org/search/pid/api?q=author:Paris_Perdikaris:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2209.03984"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2209-03984,\n  publtype={informal},\n  author={Francisco Sahli Costabal and Simone Pezzuto and Paris Perdikaris},\n  title={Δ-PINNs: physics-informed neural networks on complex geometries},\n  year={2022},\n  cdate={1640995200000},\n  journal={CoRR},\n  volume={abs/2209.03984},\n  url={https://doi.org/10.48550/arXiv.2209.03984}\n}\n"},"abstract":{"value":"Physics-informed neural networks (PINNs) have demonstrated promise in solving forward and inverse problems involving partial differential equations. Despite recent progress on expanding the class of problems that can be tackled by PINNs, most of existing use-cases involve simple geometric domains. To date, there is no clear way to inform PINNs about the topology of the domain where the problem is being solved. In this work, we propose a novel positional encoding mechanism for PINNs based on the eigenfunctions of the Laplace-Beltrami operator. This technique allows to create an input space for the neural network that represents the geometry of a given object. We approximate the eigenfunctions as well as the operators involved in the partial differential equations with finite elements. We extensively test and compare the proposed methodology against traditional PINNs in complex shapes, such as a coil, a heat sink and a bunny, with different physics, such as the Eikonal equation and heat transfer. We also study the sensitivity of our method to the number of eigenfunctions used, as well as the discretization used for the eigenfunctions and the underlying operators. Our results show excellent agreement with the ground truth data in cases where traditional PINNs fail to produce a meaningful solution. We envision this new technique will expand the effectiveness of PINNs to more realistic applications."},"title":{"value":"Δ-PINNs: physics-informed neural networks on complex geometries"},"authors":{"value":["Francisco Sahli Costabal","Simone Pezzuto","Paris Perdikaris"]}},"tmdate":1756313173282,"pdate":1640995200000,"externalIds":["dblp:journals/corr/abs-2209-03984"],"tcdate":1756313144457,"writers":["~"],"signatures":["~Simone_Pezzuto1"],"forum":"1rPqqajC9L","license":"CC BY-SA 4.0","number":618636,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1756313173282,"domain":"DBLP.org","id":"1rPqqajC9L","version":2},{"content":{"summary":{"value":"This paper proposes a unified foundation model for Artificial General Intelligence that can handle both generative and discriminative tasks. The model incorporates a central force field from physics, which enables the harmonization of energy-based and score-based models. The paper provides a clear explanation of the connections between this method and prior arts, and presents extensive experimental results that demonstrate its effectiveness in both image generation and classification benchmarks."},"presentation":{"value":"1 poor"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"The paper proposes a novel approach to unifying generative and discriminative models using a central force field from physics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper does not provide a comprehensive theoretical analysis of the method. Despite the claimed inspiration from physics, the necessity of introducing a central force field or a compelling motivation is not evident.\n2. The paper does not provide a detailed explanation of the model architecture and hyperparameters used in the experiments.\n3. The performance of the model is not good.\n4. The authors do not discuss the limitations."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. Is the table incomplete lacking some results?\n2. What is the significance of incorporating a central force field into the proposed method?\n3. Since the results (FID) is not promising compared to SOTA, is there any feasible way to improve the performance?\n4. Could the authors provide the discussion of the limitation of the method?\n5. Could the author provide a more comprehensive and understandable figure to illustrate the method?\n6. The theoretical experiments should be conducted to explain the motivation.\n7. The paper does not provide a detailed explanation of the underlying physics principles and assumptions that are used to derive the central force field.\n8. The paper does not provide a detailed analysis of the computational requirements and resource constraints of the proposed method,"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636563787,"tcdate":1699084994914,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5509/Reviewer_mAWA"],"signatures":["ICLR.cc/2024/Conference/Submission5509/Reviewer_mAWA"],"forum":"gVWnZVmpLP","number":3,"license":"CC BY 4.0","cdate":1699084994914,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5509/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636563787,"domain":"ICLR.cc/2024/Conference","replyto":"gVWnZVmpLP","id":"S1oZ7hcmqx","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Central Force Field","Generative Model","Discriminative Model","Energy-Based Model","Score-Based Model"]},"primary_area":{"value":"general machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"In the pursuit of Artificial General Intelligence, a prevalent approach is to establish a comprehensive unified foundation model that addresses multiple tasks concurrently. However, creating such a model that unifies generative and discriminative models presents significant challenges.\nThis paper aims to realize this unified model aspiration by suggesting the incorporation of a central force field from physics. \nMore precisely, within the framework of this central force field, the potential functions governing the data distribution and the joint data-label distribution become intricately interwoven with a standard discriminative classifier, rendering them well-suited for handling discriminative tasks.\nMoreover, the central force field exhibits a captivating characteristic: objects located within this field experience an attractive force that propels them towards the center. This phenomenon of centripetal motion, orchestrated by the force field, has the remarkable capability to progressively revert diffused data to its original configuration, thereby facilitating the execution of generative tasks.\nOur proposed method adeptly bridges the realms of energy-based and score-based models. Extensive experimental validation attests to the effectiveness of our approach, showcasing not only its prowess in image generation benchmarks but also its promising competitiveness in image classification benchmarks."},"_bibtex":{"value":"@misc{\nwang2024central,\ntitle={Central Force Field: Unifying Generative and Discriminative Models While Harmonizing Energy-Based and Score-Based Models},\nauthor={Guangrun Wang and Chen Lin and Philip Torr},\nyear={2024},\nurl={https://openreview.net/forum?id=gVWnZVmpLP}\n}"},"title":{"value":"Central Force Field: Unifying Generative and Discriminative Models While Harmonizing Energy-Based and Score-Based Models"},"pdf":{"value":"/pdf/3d3f9e355d11cfc995de6e109876e5d42d5b16fe.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|central_force_field_unifying_generative_and_discriminative_models_while_harmonizing_energybased_and_scorebased_models"},"authorids":{"value":["~Guangrun_Wang1","~Chen_Lin2","~Philip_Torr1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Guangrun Wang","Chen Lin","Philip Torr"]}},"version":2},{"content":{"venue":{"value":"IJCAI 2024"},"pdf":{"value":"https://www.ijcai.org/proceedings/2024/0696.pdf"},"venueid":{"value":"dblp.org/conf/IJCAI/2024"},"paperhash":{"value":"jassim|grasp_a_novel_benchmark_for_evaluating_language_grounding_and_situated_physics_understanding_in_multimodal_language_models"},"authorids":{"value":["~Serwan_Jassim1","https://dblp.org/search/pid/api?q=author:Mario_Holubar:","https://dblp.org/search/pid/api?q=author:Annika_Richter:","~Cornelius_Wolff1","https://dblp.org/search/pid/api?q=author:Xenia_Ohmer:","https://dblp.org/search/pid/api?q=author:Elia_Bruni:"]},"html":{"value":"https://www.ijcai.org/proceedings/2024/696"},"_bibtex":{"value":"@inproceedings{DBLP:conf/ijcai/JassimHRWOB24,\n  author={Serwan Jassim and Mario Holubar and Annika Richter and Cornelius Wolff and Xenia Ohmer and Elia Bruni},\n  title={GRASP: A Novel Benchmark for Evaluating Language GRounding and Situated Physics Understanding in Multimodal Language Models},\n  year={2024},\n  cdate={1704067200000},\n  pages={6297-6305},\n  url={https://www.ijcai.org/proceedings/2024/696},\n  booktitle={IJCAI},\n  crossref={conf/ijcai/2024}\n}\n"},"abstract":{"value":"This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This evaluation is accomplished via a two-tier approach leveraging Unity simulations. The first level tests for language grounding by assessing a model's ability to relate simple textual descriptions with visual information. The second level evaluates the model's understanding of \"Intuitive Physics\" principles, such as object permanence and continuity. In addition to releasing the benchmark, we use it to evaluate several state-of-the-art multimodal LLMs. Our evaluation reveals significant shortcomings in the language grounding and intuitive physics capabilities of these models. Although they exhibit at least some grounding capabilities, particularly for colors and shapes, these capabilities depend heavily on the prompting strategy. At the same time, all models perform below or at the chance level of 50% in the Intuitive Physics tests, while human subjects are on average 80% correct. These identified limitations underline the importance of using benchmarks like GRASP to monitor the progress of future models in developing these competencies."},"title":{"value":"GRASP: A Novel Benchmark for Evaluating Language GRounding and Situated Physics Understanding in Multimodal Language Models"},"authors":{"value":["Serwan Jassim","Mario Holubar","Annika Richter","Cornelius Wolff","Xenia Ohmer","Elia Bruni"]}},"tmdate":1747739392259,"pdate":1704067200000,"tcdate":1747301942652,"writers":["~"],"signatures":["~Cornelius_Wolff1"],"forum":"mifAi1nxaN","license":"CC BY-SA 4.0","number":481376,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747739392259,"domain":"DBLP.org","id":"mifAi1nxaN","version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2311.09048v3"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"jassim|grasp_a_novel_benchmark_for_evaluating_language_grounding_and_situated_physics_understanding_in_multimodal_language_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Serwan_Jassim:","https://dblp.org/search/pid/api?q=author:Mario_Holubar:","https://dblp.org/search/pid/api?q=author:Annika_Richter:","~Cornelius_Wolff1","https://dblp.org/search/pid/api?q=author:Xenia_Ohmer:","https://dblp.org/search/pid/api?q=author:Elia_Bruni:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2311.09048"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2311-09048,\n  publtype={informal},\n  author={Serwan Jassim and Mario Holubar and Annika Richter and Cornelius Wolff and Xenia Ohmer and Elia Bruni},\n  title={GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2311.09048},\n  url={https://doi.org/10.48550/arXiv.2311.09048}\n}\n"},"abstract":{"value":"This paper presents GRASP, a novel benchmark to evaluate the language grounding and physical understanding capabilities of video-based multimodal large language models (LLMs). This evaluation is accomplished via a two-tier approach leveraging Unity simulations. The first level tests for language grounding by assessing a model's ability to relate simple textual descriptions with visual information. The second level evaluates the model's understanding of \"Intuitive Physics\" principles, such as object permanence and continuity. In addition to releasing the benchmark, we use it to evaluate several state-of-the-art multimodal LLMs. Our evaluation reveals significant shortcomings in the language grounding and intuitive physics capabilities of these models. Although they exhibit at least some grounding capabilities, particularly for colors and shapes, these capabilities depend heavily on the prompting strategy. At the same time, all models perform below or at the chance level of 50% in the Intuitive Physics tests, while human subjects are on average 80% correct. These identified limitations underline the importance of using benchmarks like GRASP to monitor the progress of future models in developing these competencies."},"title":{"value":"GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models"},"authors":{"value":["Serwan Jassim","Mario Holubar","Annika Richter","Cornelius Wolff","Xenia Ohmer","Elia Bruni"]}},"tmdate":1747301946352,"pdate":1672531200000,"tcdate":1747301943011,"writers":["~"],"signatures":["~Cornelius_Wolff1"],"forum":"ggjiO0jL87","license":"CC BY-SA 4.0","number":481379,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747301946352,"domain":"DBLP.org","id":"ggjiO0jL87","version":2},{"content":{"summary":{"value":"The paper proposes IFIN for physics-guided image reconstruction. IFIN embeds a pair of operators at every encoder–decoder stage: a FSO that maps current estimates into the measurement domain and an ISO that maps measurements back toward the image domain. A learnable, spatially varying PSF field conditions both operators across scales, supporting blind or mis-calibrated settings. Experiments on DiffuserCam and MultiWienerNet report consistent improvements over classical and recent learned/physics-guided baselines in PSNR/SSIM and LPIPS."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.The innovation mainly lies in system-level integration (multi-scale forward–inverse coupling with a learnable SV-PSF) rather than a conceptually new paradigm; the paper lacks a principled justification for why coupling both domains at all scales is necessary or theoretically superior to top-level or one-sided physics.\n\n2.The distinction from prior unfolded optimization or feature-domain deconvolution frameworks (e.g., UPDN, MWDN) is insufficiently clarified, making the novelty appear incremental without controlled counterexamples or failure analyses.\n\n3.Key design choices (where SI vs. SV operators are applied, how PSF encodings are injected, and the gating/residual paths) are scattered between main text and appendix, making the architectural logic difficult to reconstruct from the figures alone.\n\n4.The paper lacks a dedicated “Limitations” or “Failure Cases” section and does not discuss boundary conditions such as sampling mismatch, noise robustness, or nonlinear forward models, which would help frame applicability.\n\n5.The description of “physics-guided” components is mostly heuristic; it does not clearly connect to physical principles such as energy conservation, boundedness, or optical transfer constraints, which weakens the claimed physical interpretability.\n\n6.Some notation and figure captions are incomplete or inconsistent—operator modes (SI/SV), PSF-tile granularity k, and cross-scale data flow are not explicitly annotated, affecting readability and reproducibility.\n\n7.Computational trade-offs (accuracy–cost–latency) are unexplored; practitioners cannot judge when to switch between the SV and surrogate FSO modes or how scaling affects efficiency."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper presents a well-motivated architecture that tightly integrates forward and inverse physics across network stages to address limitations of one-sided models.\n2. It demonstrates consistent performance gains on multiple benchmarks, including real-world spatially varying data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The contribution reads more like a careful system integration than a genuinely new paradigm; the paper offers little formal grounding—there is no error-propagation or identifiability analysis, and the bias from using an averaged PSF under spatial variance is not quantified; ablations and robustness checks are thin (no systematic removal of modules or scale coupling, limited study of PSF tiling and losses, no variance/significance reporting or cross-device transfer); baseline fairness is uncertain because prior physics-guided methods are retrained under a unified loss without parallel results from their original settings or stronger recent baselines; presentation leaves gaps in where SI vs. SV operators are used and lacks clear captions/notation and a limitations section; and there is no clear commitment to release code, models, or data, which undercuts reproducibility."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926363272,"tcdate":1761581126802,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16203/Reviewer_2FZr"],"signatures":["ICLR.cc/2026/Conference/Submission16203/Reviewer_2FZr"],"forum":"q1TpQ6guwX","number":3,"license":"CC BY 4.0","cdate":1761581126802,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16203/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926363272,"domain":"ICLR.cc/2026/Conference","replyto":"q1TpQ6guwX","id":"R5dQXh9xlz","forumContent":{"TLDR":{"value":"We propose IFIN, the network couples forward physics and learned inverse at every layer with learnable calibration-free PSF, and shows state-of-the-art lensless imaging results under spatially varying blur and noise."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Computational Imaging","Lensless Imaging","Physics-guided Learning","Inverse Problem"]},"supplementary_material":{"value":"/attachment/9d05fbb0955cc21e2c4d9132124a9b7ae3b31195.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Inverse modeling plays a central role across computational optical imaging problems, including microscopy, imaging through scattering media, and lensless cameras, where the forward model often manifests as a severe blur. Discrepancies between the model and the actual imaging process further aggravate the ill-posed nature of the inverse problem. Physics-enabled methods that integrate analytical forward models with data-driven networks have been explored, but most incorporate physics only in a one-sided manner—either operating purely in the measurement space or only after inversion—thereby discarding complementary cues and reducing robustness to calibration errors.\nHere, we propose the Integrated Forward–Inverse Network (IFIN), a physics-guided deep neural network that interleaves differentiable forward operators with learnable inverse modules at every stage of the hierarchy. This design preserves physical consistency while shaping richer feature representations by jointly leveraging information from both measurement and image domains. A physics-guided kernel adaptation further compensates for inaccurate or unavailable PSF calibration, dynamically refining the kernel for blind deconvolution under system constraints.\nIFIN is especially effective when measurements are severely blurred by large point-spread functions, where conventional CNN-based inversion is limited by local receptive fields and underutilizes the measurement signal. On challenging lensless imaging benchmarks—including our newly introduced dataset, IFIN achieves state-of-the-art reconstruction quality and improved robustness under noise and model mismatch."},"_bibtex":{"value":"@misc{\nbae2026integrated,\ntitle={Integrated Forward{\\textendash}Inverse Network for Physics-Guided Image Reconstruction},\nauthor={Donggeon Bae and Jaewoo Jung and Yong Guk Kang and Kyung Chul Lee and Taeyoung Kim and Joonsik Park and Sangjun Byun and Jongho Kim and Nakkyu Baek and Hyeonyong Lee and Kyunghoon Jung and Seung Ah Lee},\nyear={2026},\nurl={https://openreview.net/forum?id=q1TpQ6guwX}\n}"},"title":{"value":"Integrated Forward–Inverse Network for Physics-Guided Image Reconstruction"},"pdf":{"value":"/pdf/0606a689890b9911fc407f24347161e12f8d2e59.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"bae|integrated_forwardinverse_network_for_physicsguided_image_reconstruction"},"authorids":{"value":["~Donggeon_Bae1","~Jaewoo_Jung1","~Yong_Guk_Kang1","~Kyung_Chul_Lee1","~Taeyoung_Kim11","~Joonsik_Park1","~Sangjun_Byun1","~Jongho_Kim3","~Nakkyu_Baek1","~Hyeonyong_Lee1","~Kyunghoon_Jung2","~Seung_Ah_Lee1"]},"authors":{"value":["Donggeon Bae","Jaewoo Jung","Yong Guk Kang","Kyung Chul Lee","Taeyoung Kim","Joonsik Park","Sangjun Byun","Jongho Kim","Nakkyu Baek","Hyeonyong Lee","Kyunghoon Jung","Seung Ah Lee"]}},"version":2},{"content":{"summary":{"value":"This paper proposes to use data-driven methods to replace the traditional solvers in the elastic simulation. The experiments show the effectiveness of the proposed method, leading to low errors and stable long-horizon stability. The experiments also show the proposed method is better for unseen conditions or inputs. Another contribution is that since the proposed method supervises the intermediate variables, it means the proposed method can use physics knowledge as contraints."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. For strengthening the experiments, do the authors think the baselines can augment with physics knowledge?\n2. What is the error for predicting just one step forward for each method?\n3. For baselines, have you tried non-autoregressive prediction, e.g., directly predicting all the following 100 steps."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper is easy to follow and well organized.\n2. This paper evaluates their method from multiple aspects: 1. Comparison with neural operator based baselines. 2.  Joint training versus separate training. 3. Comparison with physics simulator. 4. Inference by combining with traditional simulators. 5. Comparison with no direct physical constraint included."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The biggest weakness is that the proposed method is just substituting the immediate two steps that are done by traditional methods with data-driven methods. Thus, novelty is the biggest issue to me. Besides, the neural constitutive step is following the existing paper.\n2. Another contribution claimed by the authors is the separate training plus joint training. From what I understand, this training method was also proposed in the existing work to solve the “collapse” issue, which is also not new.\n3. For the experiments, the authors compare it with purely data-driven methods. However, their method uses physics. I don’t think the experiments are good enough for showing the advantages of their methods. For example, they should compare with the traditional methods and we can see the gap when physics is used for both methods.\n4. It will be good to see the advantages/disadvantages of the methods including baselines in terms of inference time.\n5. The method is only tested on the specific problem, which means the scope is limited."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917051548,"tcdate":1762185224077,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3825/Reviewer_7hFZ"],"signatures":["ICLR.cc/2026/Conference/Submission3825/Reviewer_7hFZ"],"forum":"qYVa1obqTZ","number":4,"license":"CC BY 4.0","cdate":1762185224077,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3825/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917051548,"domain":"ICLR.cc/2026/Conference","replyto":"qYVa1obqTZ","id":"5BjNXlFUkg","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"This paper presents Neural Modular Physics, a thorough modular framework for elastic simulation, enabling better generalization and stable long-horizon simulation."},"keywords":{"value":["Neural Simulation","Neural Modular Networks","Elastic Dynamics","learned simulation"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Learning-based methods have made significant progress in physics simulation, typically approximating dynamics with a monolithic end-to-end optimized neural network. Although these models offer an effective way to simulation, they may lose essential features compared to traditional numerical simulators, such as physical interpretability and reliability. Drawing inspiration from classical simulators that operate in a modular fashion, this paper presents Neural Modular Physics (NMP) for elastic simulation, which combines the approximation capacity of neural networks with the physical reliability of traditional simulators. Beyond the previous monolithic learning paradigm, NMP enables direct supervision of intermediate quantities and physical constraints by decomposing elastic dynamics into physically meaningful neural modules connected through intermediate physical quantities. With a specialized architecture and training strategy, our method transforms the numerical computation flow into a modular neural simulator, achieving improved physical consistency and generalizability. Experimentally, NMP demonstrates superior generalization to unseen initial conditions and resolutions, stable long-horizon simulation, better preservation of physical properties compared to other neural simulators, and greater feasibility in scenarios with unknown underlying dynamics than traditional simulators."},"_bibtex":{"value":"@misc{\nli2026neural,\ntitle={Neural Modular Physics for Elastic Simulation},\nauthor={Yifei Li and Haixu Wu and Zeyi Xu and Tuur Stuyck and Wojciech Matusik},\nyear={2026},\nurl={https://openreview.net/forum?id=qYVa1obqTZ}\n}"},"title":{"value":"Neural Modular Physics for Elastic Simulation"},"pdf":{"value":"/pdf/72ba4587211cfe2a7de32067a758b4ec539f8603.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|neural_modular_physics_for_elastic_simulation"},"authorids":{"value":["~Yifei_Li7","~Haixu_Wu1","~Zeyi_Xu2","~Tuur_Stuyck1","~Wojciech_Matusik2"]},"authors":{"value":["Yifei Li","Haixu Wu","Zeyi Xu","Tuur Stuyck","Wojciech Matusik"]}},"version":2},{"content":{"summary":{"value":"This paper introduces ChartMaster, a framework to improve MLLM performance on chart analysis. The framework consists of three main components: (1) ChartVerse, a new large-scale synthetic dataset; (2) Multi-Negative Direct Preference Optimization (MNDPO), a variant of DPO that uses hard negative samples; and (3) Reinforcement Learning with Dynamic Length Reward (DLR), a reward mechanism that encourages concise reasoning for simple queries and multi-step reasoning for complex ones. The authors show that this framework achieves state-of-the-art (SOTA) performance on six chart benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Justify the novelty of the work. Add further comparisons with existing similar literature and demonstrate how this work is a novel contribution."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The framework achieves state-of-the-art (SOTA) performance on six chart benchmarks.\n2. Their applied reasoning and perceptual optimization methods help the model to improve performance. \n3. Introduces a new large-scale synthetic dataset for MLLM training."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The major limitation of this paper is its limited novelty:\n\n(i) MNDPO is quite a straightforward extension of DPO to multiple negatives. \n\n(ii) DLR is also a simple heuristic that penalizes deviation from the shortest correct answer. \n\n(iii) ChartVerse contains synthetic data, which is also a common practice."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916360433,"tcdate":1762175970572,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2753/Reviewer_rj1g"],"signatures":["ICLR.cc/2026/Conference/Submission2753/Reviewer_rj1g"],"forum":"1UqnezioFf","number":4,"license":"CC BY 4.0","cdate":1762175970572,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2753/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916360433,"domain":"ICLR.cc/2026/Conference","replyto":"1UqnezioFf","id":"lyusrNuRq9","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"ChartMaster enhances MLLMs' chart analysis by improving their fine-grained perception and enabling efficient reasoning."},"keywords":{"value":["Chart Analysis","Multimodal Large Language Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multimodal Large Language Models (MLLMs) have demonstrated significant potential in understanding visual information, yet they often fall short in the complex domain of chart analysis. \nExisting models often struggle to accurately capture detailed visual elements and to perform efficient multi-step reasoning. \nTo address these challenges, we introduce ChartMaster, a holistic framework that systematically advances chart analysis by jointly optimizing data, perception, and reasoning.\nOur approach is built on three core innovations.\nFirst, we construct ChartVerse, a large-scale synthetic dataset with diverse chart types, rendering styles, and reasoning levels.\nBuilding on this foundation, we introduce a novel two-stage training paradigm: \n(i) Multi-Negative Direct Preference Optimization (MNDPO), which improves perceptual precision by training models to distinguish correct answers from carefully designed hard negative samples (i.e., plausible but incorrect alternatives); and (ii) Reinforcement Learning with Dynamic Length Reward (DLR), which adapts chain-of-thought reasoning to task complexity, encouraging concise solutions for simple queries and rigorous multi-step reasoning for complex ones. \nExtensive experiments across six benchmarks demonstrate that ChartMaster achieves state-of-the-art performance, surpassing prior chart-domain models and rivaling proprietary systems. These results highlight that coupling diverse data foundations with targeted perceptual and reasoning optimization provides an effective pathway toward robust chart understanding in MLLMs."},"_bibtex":{"value":"@misc{\nliu2026chartmaster,\ntitle={ChartMaster: Boosting {MLLM}s for Chart Analysis through Data, Perception, and Reasoning Optimization},\nauthor={Chaohu Liu and Desong Sui and Haoyu Cao and YongXiang Hua and Linli Xu},\nyear={2026},\nurl={https://openreview.net/forum?id=1UqnezioFf}\n}"},"title":{"value":"ChartMaster: Boosting MLLMs for Chart Analysis through Data, Perception, and Reasoning Optimization"},"pdf":{"value":"/pdf/9b5b8f6711e601dad37694430aa41f5ac4960e51.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"liu|chartmaster_boosting_mllms_for_chart_analysis_through_data_perception_and_reasoning_optimization"},"authorids":{"value":["~Chaohu_Liu1","~Desong_Sui1","~Haoyu_Cao1","~YongXiang_Hua1","~Linli_Xu1"]},"authors":{"value":["Chaohu Liu","Desong Sui","Haoyu Cao","YongXiang Hua","Linli Xu"]}},"version":2},{"content":{"comment":{"value":"Dear Reviewer, thanks for your comment. Understandably, your concerns are valid. We had this concern right from the beginning, and took many steps to ensure that a synthetic dataset is useful to real world applications as well. \n\n1. \"The synthetic data may not fully capture the complexity of real-world scenarios.\"\nWe ensured that the selection of objects, and the nature of interaction was as realistic as possible. Not only did we use textured real-world-like objects, but restricted object pairs to only those that are likely to be seen in the real world. Also the physics engine used ensures that the interaction is natural. \n\n2. \"The paper could benefit from a more extensive validation of the dataset's efficacy across a broader range of models and tasks.\"\nAgreed. We realize that there could be more experiments for the temporal part of STUPD. So we will include some results along those lines as well. \n\n3. \"How well do models trained on STUPD perform when applied to real-world data, considering the dataset is synthetic?\" \nWe demonstrated through the pretraining experiments (Table 3 and 4), that when models are first pretrained on STUPD, and then fine tuned on real-world datasets, the model performance is. consistently better. \n\n4. \"How does STUPD handle ambiguous or context-dependent prepositions where the spatial or temporal relationship might not be clear-cut?\"\nDifferent context dependent prepositions have distinct labels and videos corresponding to them. For example, the word \"against\" has two meanings, and different corresponding videos. Similarly, there are two different classes corresponding to \"into\". Within certain constraints, we cover all prepositional relations in our dataset. Maybe some less-frequently words are not covered, but their meaning is covered by one or the other of the classes in this dataset. There are 40 classes in total, and the full list can be seen in Appendix. \n\n5. \"What measures are in place to ensure that the synthetic data in STUPD is diverse and representative of real-world scenarios?\"\nWe started out by selecting a diverse range of objects from various supercategories (like vehicles, large grounded objects, small objects, larger objects, containers). Overall, the diversity of objects as well as the absolute number of objects is comparable to SpatialSense (a real world dataset). These synthetic objects have real-world-like textures. We also filter out object pairs that are unlikely to be seen in the real world (such as building on top of a car). \nBackgrounds are also selected so as to bring realness. Finally, a physics engine ensures that the interactions are real-world-like."},"title":{"value":"Viability of a synthetic dataset"}},"tmdate":1710476778363,"tcdate":1700538399195,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission7768/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission7768/Authors"],"forum":"eqz5aXtQv1","number":3,"license":"CC BY 4.0","cdate":1700538399195,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission7768/-/Official_Comment","ICLR.cc/2024/Conference/-/Edit"],"mdate":1710476778363,"domain":"ICLR.cc/2024/Conference","replyto":"O9xT61MYXL","id":"OqmLFLS7bF","forumContent":{"TLDR":{"value":"We propose a large scale synthetic pretraining dataset that can improve visual relation reasoning in real world settings"},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["spatial reasoning","visual relation detection"]},"supplementary_material":{"value":"/attachment/280cb63c566ded8e654fadd408337c74af4828c5.pdf"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Understanding relations between objects is crucial for understanding the semantics of a visual scene. It is also an essential step in order to bridge visual and language models. However, current state-of-the-art computer vision models still lack the ability to perform spatial reasoning well. Existing datasets mostly cover a relatively small number of spatial relations, all of which are static relations that do not intrinsically involve motion. In this paper, we propose the Spatial and Temporal Understanding of Prepositions Dataset (STUPD) – a large scale video dataset for understanding static and dynamic spatial relationships derived from prepositions of the English language. The dataset contains 150K visual depictions (videos and images), consisting of 30 distinct spatial prepositional senses, in the form of object interaction simulations generated synthetically using Unity3D. In addition to spatial relations, we also propose 50K visual depictions across 10 temporal relations, consisting of videos depicting event/time-point interactions. To our knowledge, no dataset exists that represents temporal relations through visual settings. In this dataset, we also provide 3D information about object interactions such as frame-wise coordinates, and descriptions of the objects used. The goal of this synthetic dataset is to help models perform better in visual relationship detection in real-world settings. We demonstrate an increase in the performance of various models over 2 real-world datasets (ImageNet VidVRD and Spatial Senses) when pretrained on the STUPD dataset, in comparison to other pretraining datasets."},"_bibtex":{"value":"@misc{\nagrawal2024stupd,\ntitle={{STUPD}: A Synthetic Dataset for Spatial and Temporal Relation Reasoning},\nauthor={Palaash Agrawal and Haidi Azaman and Cheston Tan},\nyear={2024},\nurl={https://openreview.net/forum?id=eqz5aXtQv1}\n}"},"title":{"value":"STUPD: A Synthetic Dataset for Spatial and Temporal Relation Reasoning"},"pdf":{"value":"/pdf/3c17d0412c20e944250bb91e2ccf788f4cd8179c.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"agrawal|stupd_a_synthetic_dataset_for_spatial_and_temporal_relation_reasoning"},"authorids":{"value":["~Palaash_Agrawal1","~Haidi_Azaman1","~Cheston_Tan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Palaash Agrawal","Haidi Azaman","Cheston Tan"]}},"version":2},{"content":{"summary":{"value":"This paper explores the use of MLFFs as general-purpose representation learners for local protein environments. Instead of relying on sequence-based or handcrafted descriptors, the authors repurpose latent embeddings from pretrained MLFFs (AIMNet, MACE, OrbNet, Egret) as compact, physics-grounded descriptors for atomic neighborhoods in proteins. The study benchmarks MLFF embeddings on diverse downstream tasks—secondary structure and amino acid classification, pKa prediction, and NMR chemical shift regression—and shows that these representations outperform or rival specialized baselines such as PropKa and pKa-ANI. Overall, the paper establishes MLFFs as reusable, physics-informed foundation models for structural biology."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Could the same MLFF-based representation transfer to nucleic acids or protein–ligand complexes, where local environments include non-canonical atoms and charges?\n\n2. MLFFs have multiple internal layers encoding different orders of interaction (0th, 1st, 2nd). Did the authors investigate which layer yields the most informative embeddings for downstream tasks?\n\n3. Since MLFF embeddings originate from networks with different scales and symmetries, how are they aligned or normalized across model families?\n\n4. The authors mention uncertainty-aware predictions. Is uncertainty derived from ensemble variance, likelihood width, or another calibration technique?\n\n5. Could MLFF embeddings serve as complementary features to pretrained protein language models, bridging physics-based and sequence-based representations?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The central insight, repurposing MLFF embeddings as general-purpose, transferable protein environment representations, is both novel and timely. Most prior MLFF applications focus on energy or force prediction for small molecules; extending them to protein representation learning is original and valuable. The paper effectively bridges quantum-chemistry-based potentials and protein machine learning.\n\n2. The experimental setup is comprehensive. The authors curate 165 k environments from 1048 proteins and evaluate four MLFF families on four biologically relevant tasks. Comparisons include classical and ML baselines (PropKa, pKa-ANI, UCBShift2-X). Results demonstrate meaningful improvements in both accuracy and interpretability. The inclusion of uncertainty quantification and physical consistency tests (e.g., ring-current effects) adds rigor.\n\n3. The methodology that extracting embeddings from pretrained MLFFs and mapping them to canonical residue-centered environments is sound and clearly motivated. Statistical reporting (mean absolute errors, standard deviations) is adequate, and all experiments appear well-controlled.\n\n4. The paper is well written and pedagogically organized. The introduction clearly motivates the challenge of representing local protein environments; figures (e.g., Fig. 1 and 2) effectively illustrate how embeddings are constructed and used. Terminology (canonical environment, focus residue, MLFF feature extraction) is consistent and accessible even to readers outside computational chemistry.\n\n5. This work could substantially influence both computational biology and machine-learning communities by providing a physics-consistent alternative to sequence-only protein language models. MLFF embeddings encode quantum-derived information unavailable in existing representations and show transferability to tasks requiring local chemical precision. The idea of using pretrained MLFFs as “foundation models for atoms” could be broadly significant."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The study restricts environments to 5 Å radius regions; while suitable for local chemistry, it omits long-range electrostatic or conformational effects. For tasks like folding or binding prediction, this locality may be insufficient. A discussion or experiment extending to multi-scale contexts would strengthen generality claims.\n\n2. MLFFs such as MACE and OrbNet require quantum-level pretraining on millions of molecules. Although embeddings are reused, the computational barrier to obtaining them limits accessibility compared to pretrained sequence models (e.g., ESM, ProtT5). The paper could better address scalability and efficiency trade-offs.\n\n3. Qualitative case studies (e.g., helix → strand unfolding) are insightful but anecdotal. Quantitative metrics (e.g., correlation between embedding distances and RMSD/chemical similarity) would solidify interpretability claims.\n\n4. While the authors compare to physics-based predictors, they do not benchmark against modern structure-aware geometric encoders (e.g., GVP-GNN, ProteinNeRF, FrameFold). Including such baselines would clarify whether MLFF embeddings provide advantages beyond standard geometric message-passing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942763567,"tcdate":1761564400479,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23683/Reviewer_fbDS"],"signatures":["ICLR.cc/2026/Conference/Submission23683/Reviewer_fbDS"],"forum":"9ZogcRkhoG","number":1,"license":"CC BY 4.0","cdate":1761564400479,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23683/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942763567,"domain":"ICLR.cc/2026/Conference","replyto":"9ZogcRkhoG","id":"HWRYLMQkQf","forumContent":{"TLDR":{"value":"We show that embeddings from machine learning force fields provide rich, transferable representations of local protein environments, enabling zero-shot generalization and state-of-the-art downstream performance."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Machine learning force fields","structural biology","NMR","representation learning"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"The local structure of a protein strongly impacts its function and interactions with other molecules. Representing local biomolecular environments remains a key challenge while applying machine learning approaches over protein structures. The structural and chemical variability of these environments makes them challenging to model, and performing representation learning on these objects remains largely under-explored.  In this work, we propose representations for local protein environments that leverage intermediate features from machine learning force fields (MLFFs). We extensively benchmark state-of-the-art MLFFs—comparing their performance across latent spaces and downstream tasks—and show that their embeddings capture local structural (e.g., secondary motifs) and chemical features (e.g., amino acid identity and protonation state), organizing protein environments into a structured manifold. We show that these representations enable zero-shot generalization and transfer across diverse downstream tasks. As a case study, we build a physics-informed, uncertainty-aware chemical shift predictor that achieves state-of-the-art accuracy in biomolecular NMR spectroscopy. Our results establish MLFFs as general-purpose, reusable representation learners for protein modeling, opening new directions in representation learning for structured physical systems."},"_bibtex":{"value":"@inproceedings{\nbojan2026representing,\ntitle={Representing local protein environments with machine learning force fields},\nauthor={Meital Bojan and Sanketh Vedula and Sai Advaith Maddipatla and Nadav Bojan and Anar Rzayev and Federico Napoli and Paul Schanda and Alexander Bronstein},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=9ZogcRkhoG}\n}"},"title":{"value":"Representing local protein environments with machine learning force fields"},"pdf":{"value":"/pdf/3da39fa684cf23630dfee0dd43015acfb6a38b2e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"bojan|representing_local_protein_environments_with_machine_learning_force_fields"},"authorids":{"value":["~Meital_Bojan1","~Sanketh_Vedula1","~Sai_Advaith_Maddipatla1","~Nadav_Bojan2","~Anar_Rzayev1","~Federico_Napoli1","~Paul_Schanda1","~Alexander_Bronstein1"]},"authors":{"value":["Meital Bojan","Sanketh Vedula","Sai Advaith Maddipatla","Nadav Bojan","Anar Rzayev","Federico Napoli","Paul Schanda","Alexander Bronstein"]}},"version":2},{"content":{"summary":{"value":"The paper trains a small reasoning model for materials discovery by using a larger model to generate reasoning traces and then perform physics-aware filtering (PaRS) to select “valid” traces for distillation. Experiments on QD-LED device property prediction show improved MAE and lower violation rates relative to simple rejection sampling baselines."},"soundness":{"value":2},"confidence":{"value":1},"questions":{"value":"see weakness.\n\nDue to my limited familiarity with materials discovery, I cannot fully assess the domain significance of the claimed contribution beyond the methodology itself."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper is clearly written and the problem of physics-admissible reasoning for materials tasks is practically relevant.\n\nThe evaluation protocol is relatively thorough within the chosen domain (teacher-side + student-side metrics).\n\nPhysics-based rejection constraints are reasonable and give some domain grounding compared with correctness-only filtering."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Core idea is conceptually straightforward and incremental\nThe pipeline is essentially:\n(big LRM → generate → filter → distill to small LRM).\nThis paradigm already appears widely in reasoning LLM alignment (e.g., rejection sampling + SFT/RFT/RL), and the physics-aware gating here amounts to domain-specific heuristics rather than a new training principle.\n\nTechnical novelty is limited\nThe PaRS mechanism is a direct instantiation of rejection sampling with thresholding; there is no new modeling component, no new learning objective, and no new algorithmic insight beyond domain-informed gating."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926896706,"tcdate":1761378219494,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16870/Reviewer_oa7n"],"signatures":["ICLR.cc/2026/Conference/Submission16870/Reviewer_oa7n"],"forum":"LxvL6iKzt6","number":1,"license":"CC BY 4.0","cdate":1761378219494,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16870/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926896706,"domain":"ICLR.cc/2026/Conference","replyto":"LxvL6iKzt6","id":"OZ7DswnnvF","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Materials design","LLM","Property prediction","Reasoning model"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Closing the experimental loop in materials discovery requires process-aware recipe to property predictors that are accurate, calibrated, and physically admissible. We approach this as a reasoning problem with large reasoning models (LRMs). To instill reasoning capability into language models, we curate reasoning traces from a teacher model to train a student model. However, most training pipelines select reasoning traces using binary correctness or learned preference signals that poorly reflect physical admissibility. We introduce Physics-aware Rejection Sampling (PaRS), a training-time trace selection scheme that favors traces consistent with fundamental physics and numerically close to targets, with lightweight halting to control compute. We instantiate our framework with a large student model fine-tuned on traces synthesized by a larger teacher model, and evaluate under matched token budgets against various rejection sampling baselines. Our method improves accuracy and calibration, reduces physics-violation rates, and lowers sampling cost relative to baselines. These results indicate that modest, domain-aware constraints combined with trace-level selection provide a practical path toward reliable, efficient LRMs for process-aware property prediction and closed-loop materials design."},"_bibtex":{"value":"@misc{\nhyun2025aligning,\ntitle={Aligning Reasoning {LLM}s for Materials Discovery with Physics-aware Rejection Sampling},\nauthor={Lee Hyun and Sohee Yoon and Jinwoo Park and Seongeon Park and Sue In Chae and Jooyeon Ahn and Yebin Jung and You Jung Chung and JINA KIM and Hogeun Chang and Sujin Park and Myeonginn Kang and Ho-Gyeong Kim and Myeonghun Jeong},\nyear={2025},\nurl={https://openreview.net/forum?id=LxvL6iKzt6}\n}"},"title":{"value":"Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling"},"pdf":{"value":"/pdf/d58f79b4c7f40637d8e6c2eeafaa8f327864a819.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hyun|aligning_reasoning_llms_for_materials_discovery_with_physicsaware_rejection_sampling"},"authorids":{"value":["~Lee_Hyun1","~Sohee_Yoon1","~Jinwoo_Park3","~Seongeon_Park1","~Sue_In_Chae1","~Jooyeon_Ahn1","~Yebin_Jung1","~You_Jung_Chung1","~JINA_KIM6","~Hogeun_Chang1","~Sujin_Park3","~Myeonginn_Kang1","~Ho-Gyeong_Kim1","~Myeonghun_Jeong1"]},"authors":{"value":["Lee Hyun","Sohee Yoon","Jinwoo Park","Seongeon Park","Sue In Chae","Jooyeon Ahn","Yebin Jung","You Jung Chung","JINA KIM","Hogeun Chang","Sujin Park","Myeonginn Kang","Ho-Gyeong Kim","Myeonghun Jeong"]}},"version":2},{"content":{"summary":{"value":"The paper Neural Modular Physics for Elastic Simulation , introduces a fully modular neural simulator for elastic dynamics. This represents a new direction relative to the monolithic, end-to-end neural simulation paradigms (NO, GNN) and hybrid simulators that replace partial components demonstrated through a complete modular neural architecture that mirrors traditional finite element method (FEM) computation flows. The approach enables direct supervision of intermediate physical quantities and interchangeable traditional-numerical and neural components, which is a novel result in the neural simulation literature. The contribution of this paper lies more in systematic integration and training methodology (the two-stage modular physics training) than in the individual components.\nGiven the surge of interest in physics-informed ML, differentiable simulation, and scientific machine learning. The paper addresses a key pain point of neural simulators—lack of interpretability and physical soundness—making it relevant to computational physics and graphics communities as well. The papers focus on elastic simulation is narrower than universal PDE solvers, but serves as an excellent proof-of-concept domain."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please address the following practical issues \n1. Discuss time to converge relative to traditional solvers on a real world engineering problem if possible\n2. discuss error analysis and boundary condition matching\n 3. provide more evidence on physics validation metrics: Energy conservation, stress-strain correlation, or physical invariants could further substantiate realism."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Concept: \nThe paper provides a fairly new approach to learning-based, physically interpretable simulation approach that modularizes elastic dynamics into neural subcomponents aligned with traditional numerical solvers. The idea of decomposing physics simulation into learnable yet interpretable modules, supervised via intermediate physical quantities, represents a meaningful conceptual and architectural advance.\nWhile the idea of combining neural networks with modular solvers has appeared before the innovation lies more in systematic integration and training methodology (the two-stage modular physics training) than in the individual components.\n2. Practicality/Application:  \nWhen combining neural networks with physics-based modular solvers, one of the hardest problems is ensuring boundary condition consistency between modules that are partly data-driven and partly physically constrained. In hybrid or modular networks, each subnetwork (e.g., constitutive law, integrator) may implicitly assume slightly different boundary behavior, leading to inconsistencies at their interfaces — this is known as a boundary-condition coupling error.  The NMP framework proposes to address this through architectural and Training level strategies.  The Neural Integration Module preserves the same interface as an implicit FEM integrator but replaces the internal update terms with a neural network. After the neural update, boundary condition enforcement and collision handling are explicitly re-applied.  \nHence, NMP preserves:\nDirichlet BCs \nNeumann BCs\nContact constraints\nThis design ensures neural predictions remain compatible with traditional BC enforcement — solving the “boundary leakage” problem found in prior hybrids\n3. Results: Published results show improved performance over top baselines as well as improved stability (reduced collapse) \n4. Generalizability: is demonstrated against unseen initial conditions and higher mesh resolutions. \n5. Evidence: The SPOT and BOB experiments are specifically designed to test BC handling\nSPOT: multiple fixed boundary points (head and tail)\nBOB: fixed vertices (head) and free elastic regions"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Practicality/Application:  solution process and validation\nComputational cost and scalability are not discussed (e.g., time vs. FEM or other neural simulators). This is a critical metric to be evaluated.\nLack of error analysis: No discussion on failure modes or interpretability visualization for intermediate variables.  This is also critical for real world applications on elastic body problems. \n2. Results: are limited to non real world elastici body  FEM problems  \n3. Generalizability: is limited to elastic solids only. Extensions to fluids or multi-physics are not shown and this gap is acknowledged by the authors.  \n4. Evidence and support: the following weaknesses exist:\nNo ablations on module architecture: How sensitive is performance to neural constitutive/integration model complexity?\nLimited physics validation metrics: Energy conservation, stress-strain correlation, or physical invariants could further substantiate realism.\n\nWriting clarity:\nline 121 -\"dynamics on top of analytical dynamics to account for unmodeled\" ... is an incomplete sentence\nline 135: \"In this paper, we notice the internal modularity\". ... please fix\nline 141: \"As aforementioned,\" --> as mentioned earlier ?\nline 160 : \"derives two disentangled modules\" --> two decoupled ?\nline 182: \"with vertice position\" --> vertex position ?\nline 329: Transolver (?). --> what is the \"? \" provide citation, Wu ?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358369884,"tcdate":1761766648026,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3825/Reviewer_b4yZ"],"signatures":["ICLR.cc/2026/Conference/Submission3825/Reviewer_b4yZ"],"forum":"qYVa1obqTZ","number":1,"license":"CC BY 4.0","cdate":1761766648026,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3825/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358369884,"domain":"ICLR.cc/2026/Conference","replyto":"qYVa1obqTZ","id":"Qe15gl9g84","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"This paper presents Neural Modular Physics, a thorough modular framework for elastic simulation, enabling better generalization and stable long-horizon simulation."},"keywords":{"value":["Neural Simulation","Neural Modular Networks","Elastic Dynamics","learned simulation"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Learning-based methods have made significant progress in physics simulation, typically approximating dynamics with a monolithic end-to-end optimized neural network. Although these models offer an effective way to simulation, they may lose essential features compared to traditional numerical simulators, such as physical interpretability and reliability. Drawing inspiration from classical simulators that operate in a modular fashion, this paper presents Neural Modular Physics (NMP) for elastic simulation, which combines the approximation capacity of neural networks with the physical reliability of traditional simulators. Beyond the previous monolithic learning paradigm, NMP enables direct supervision of intermediate quantities and physical constraints by decomposing elastic dynamics into physically meaningful neural modules connected through intermediate physical quantities. With a specialized architecture and training strategy, our method transforms the numerical computation flow into a modular neural simulator, achieving improved physical consistency and generalizability. Experimentally, NMP demonstrates superior generalization to unseen initial conditions and resolutions, stable long-horizon simulation, better preservation of physical properties compared to other neural simulators, and greater feasibility in scenarios with unknown underlying dynamics than traditional simulators."},"_bibtex":{"value":"@misc{\nli2026neural,\ntitle={Neural Modular Physics for Elastic Simulation},\nauthor={Yifei Li and Haixu Wu and Zeyi Xu and Tuur Stuyck and Wojciech Matusik},\nyear={2026},\nurl={https://openreview.net/forum?id=qYVa1obqTZ}\n}"},"title":{"value":"Neural Modular Physics for Elastic Simulation"},"pdf":{"value":"/pdf/72ba4587211cfe2a7de32067a758b4ec539f8603.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|neural_modular_physics_for_elastic_simulation"},"authorids":{"value":["~Yifei_Li7","~Haixu_Wu1","~Zeyi_Xu2","~Tuur_Stuyck1","~Wojciech_Matusik2"]},"authors":{"value":["Yifei Li","Haixu Wu","Zeyi Xu","Tuur Stuyck","Wojciech Matusik"]}},"version":2},{"content":{"comment":{"value":"> **W4.** Computational Cost Analysis is Missing.\n\n- Negligible compute cost of the synthetic data pipeline: Our model is based on MagicDriveDiT and fine-tuned for 8 hours on 8×H200 GPUs (2,000 iterations). Once fine-tuned, generating edited video data is lightweight: 7 scenes × 10 assets × 10 steps × 4 s ≈ **50 minutes** of GPU time. In contrast, training the downstream perception model on 28,130 real samples takes ≈**13 hours**, dominating the computational budget. Adding 420 synthetic samples adds negligible overhead. Overall, the synthetic pipeline’s compute cost is a small fraction of end-to-end training.\n- Synthetic data provides gains beyond additional epochs: Increasing training epochs mainly reduces optimization error, whereas synthetic data improves **distribution coverage**, especially for rare cases. Using only real data, the model achieves 36.1 → 42.2 → 43.1 mAP over 1–3 epochs. Adding **<2% synthetic samples** increases this to 40.7 → 43.6 → 44.5 mAP (+4.6, +1.4, +1.4), demonstrating that synthetic data **enhances performance beyond mere epoch increases and extends model generalization**.\n\n------\n\n> **Q2.** A discussion on the scene selection and potential bias.\n\n- We use an MLLM-assisted pipeline to select scenes. The model identifies whether a continuous safe driving trajectory exists in any direction (front, back, left, right) and outputs the initial asset insertion position (x, y, z in the ego-vehicle frame at the first frame). Detailed prompts are provided in the Appendix A.2.\n- Following the reviewer’s suggestion, we randomly selected two alternative sets of insertion scenes, forming three groups (A, B, C) including the original. Downstream task experiments show similar performance across all groups, indicating that any bias from scene selection is negligible.\n\n|       | 1x Epochs       |                 |                 | 2x Epochs       |                 |                 | 3x Epochs       |                 |                 |\n| ----- | --------------- | --------------- | --------------- | --------------- | --------------- | --------------- | --------------- | --------------- | --------------- |\n|       | Insert Scenes A | Insert Scenes B | Insert Scenes C | Insert Scenes A | Insert Scenes B | Insert Scenes C | Insert Scenes A | Insert Scenes B | Insert Scenes C |\n| mAP↑  | 40.7            | 40.8            | 40.7            | 43.6            | 43.5            | 43.6            | 44.5            | 44.3            | 44.5            |\n| mATE↓ | 64.2            | 64.1            | 64.2            | 61.6            | 61.5            | 61.4            | 59.8            | 59.8            | 59.9            |\n| mAOE↓ | 48.0            | 48.0            | 48.1            | 39.4            | 39.4            | 39.5            | 40.1            | 40.2            | 40.1            |\n| mAVE↓ | 27.1            | 26.9            | 27.1            | 27.4            | 27.5            | 27.3            | 27.2            | 27.3            | 27.2            |\n| NDS↑  | 52.0            | 52.1            | 52.1            | 54.3            | 54.2            | 54.3            | 55.0            | 55.0            | 55.1            |\n\n------\n\n> **Q3.** Does the generative model also learn to synthesize more complex physical interactions?\n\nWe thank the reviewer for this insightful question, which touches upon the current frontiers of generative modeling.\n\nOur approach prioritizes **semantic-level realism** by ensuring plausible appearance, correct shadows, and basic reflections, which is sufficient to enable robust feature learning and improve downstream performance. \n\nHowever, synthesizing complex physical interactions, such as dynamic splashes, dust clouds, or nuanced colored light reflections from traffic signals, falls **outside the scope** of our primary contribution. These effects require advanced physics-based rendering or highly sophisticated generative models (even the current SOTA diffusion-based models struggle with such fine-grained interactions). \n\nCrucially, **our method’s core contribution is identifying and synthesizing the valuable out-of-distribution scenarios for perception training**, not optimizing the underlying generative model to achieve perfect photorealism. \n\n-----\n\nFinally, thank you for your thoughts and valuable feedback! We hope we have addressed most of your concerns. If you have any questions, please don't hesitate to reach out to us."},"title":{"value":"Response to Reviewer yxwm 2/2"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763637108609,"tcdate":1763637108609,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission971/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission971/Authors"],"forum":"z3cFADf6zZ","number":10,"license":"CC BY 4.0","cdate":1763637108609,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission971/-/Official_Comment"],"mdate":1763637108609,"domain":"ICLR.cc/2026/Conference","replyto":"xTGLbqY6v3","id":"PD2QiJKtG3","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Autonomous Driving","Driving World Model","Perception Tasks","Synthetic Data"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Recent advancements in driving world models enable controllable generation of high-quality RGB videos or multimodal videos. \nExisting methods primarily focus on metrics related to generation quality and controllability. \nHowever, they often overlook the evaluation of downstream perception tasks, which are {\\bf really crucial} for the performance of autonomous driving. \nExisting methods usually leverage a training strategy that first pretrains on synthetic data and finetunes on real data, resulting in twice the epochs compared to the baseline (real data only). \nWhen we double the epochs in the baseline, the benefit of synthetic data becomes negligible.\nTo thoroughly demonstrate the benefit of synthetic data, we introduce Dream4Drive, a novel synthetic data generation framework designed for enhancing the downstream perception tasks.\nDream4Drive first decomposes the input video into several 3D-aware guidance maps and subsequently renders the 3D assets onto these guidance maps.\nFinally, the driving world model is fine-tuned to produce the edited, multi-view photorealistic videos, which can be used to train the downstream perception models.\nDream4Drive enables unprecedented flexibility in generating multi-view corner cases at scale, significantly boosting corner case perception in autonomous driving. \nTo facilitate future research, we also contribute a large-scale 3D asset dataset named DriveObj3D, covering the typical categories in driving scenarios and enabling diverse 3D-aware video editing.\nWe conduct comprehensive experiments to show that Dream4Drive can effectively boost the performance of downstream perception models under various training epochs. \nProject website: \\url{https://wm-research.github.io/Dream4Drive/}."},"_bibtex":{"value":"@inproceedings{\nzeng2026rethinking,\ntitle={Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks},\nauthor={Kai Zeng and Zhanqian Wu and Kaixin Xiong and Xiaobao Wei and Xiangyu Guo and Zhenxin Zhu and Kalok Ho and Lijun Zhou and Bohan Zeng and Ming Lu and Haiyang Sun and BING WANG and Guang Chen and Hangjun Ye and Wentao Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=z3cFADf6zZ}\n}"},"title":{"value":"Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks"},"pdf":{"value":"/pdf/71bda5c5ed57972d48727bc4da82c6856d114ee8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zeng|rethinking_driving_world_model_as_synthetic_data_generator_for_perception_tasks"},"authorids":{"value":["~Kai_Zeng7","~Zhanqian_Wu1","~Kaixin_Xiong1","~Xiaobao_Wei1","~Xiangyu_Guo5","~Zhenxin_Zhu1","~Kalok_Ho1","~Lijun_Zhou1","~Bohan_Zeng1","~Ming_Lu2","~Haiyang_Sun2","~BING_WANG23","~Guang_Chen2","~Hangjun_Ye1","~Wentao_Zhang1"]},"authors":{"value":["Kai Zeng","Zhanqian Wu","Kaixin Xiong","Xiaobao Wei","Xiangyu Guo","Zhenxin Zhu","Kalok Ho","Lijun Zhou","Bohan Zeng","Ming Lu","Haiyang Sun","BING WANG","Guang Chen","Hangjun Ye","Wentao Zhang"]}},"version":2},{"content":{"venue":{"value":"Web3D 2007"},"venueid":{"value":"dblp.org/conf/VRML/2007"},"paperhash":{"value":"coumans|collada_physics"},"authorids":{"value":["~Erwin_Coumans2","https://dblp.org/search/pid/api?q=author:Keith_Victor:"]},"html":{"value":"https://doi.org/10.1145/1229390.1229407"},"_bibtex":{"value":"@inproceedings{DBLP:conf/vrml/CoumansV07,\n  author={Erwin Coumans and Keith Victor},\n  title={COLLADA physics},\n  year={2007},\n  cdate={1167609600000},\n  pages={101-104},\n  url={https://doi.org/10.1145/1229390.1229407},\n  booktitle={Web3D},\n  crossref={conf/vrml/2007}\n}\n"},"abstract":{"value":"We will give an overview of the COLLADA 1.4 standard physics format [1][2][3] and its use in the 3D physics content pipeline. We describe design decisions, implementation, compatibility and interoperability aspects of adopters of this industry standard, as well as its relationship with other standards such as X3D, (ISO 19775) from the Web 3D Consortium. [4]"},"title":{"value":"COLLADA physics"},"authors":{"value":["Erwin Coumans","Keith Victor"]}},"tmdate":1762404822595,"pdate":1167609600000,"externalIds":["dblp:conf/vrml/CoumansV07"],"tcdate":1762404814740,"writers":["~"],"signatures":["~Erwin_Coumans2"],"forum":"uR1Z9XYFiD","license":"CC BY-SA 4.0","number":661854,"cdate":1167609600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762404822595,"domain":"DBLP.org","id":"uR1Z9XYFiD","version":2},{"content":{"summary":{"value":"The paper introduces a procedure to iteratively synthesize problem that a RLVR pipeline solves, and leads to boosted performance on various test set. The synthetic process consists of a few steps, such as conditioning on failed problems to synthesize new tasks, filtering and resolving those problem variations. The methods generally sound intuitive and improve the baseline significantly."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"=== *main takeaway* ===\n\nI think a main feeling after reading the paper is that it introduces yet another way to synthesize and augment dataset in RLVR process, and despite the performance improvements, I do not feel having learned too much insight from the paper, in additional to knowing yet another set of process to augment data for RLVR.\n\n=== *why certain filtering steps* ===\n\nI think a useful question worth answering is why we'd filter in one way vs another and it's useful to show why filtering the synthetic data according to the current criterion lead to performance vs. a more naive approach. I can see that we iterate a few times on the method while developing the paper and it's useful to understand and present initial failure cases and specific training comparison.\n\n=== *baseline* ===\n\nI think the paper lacks comparison to a few baselines. One is the prior work in the space on synthetic data, e.g., a few ideas related to synthesizing data from the initial training set has been proposed in [1], but maybe we should cite this paper + make proper comparisons, otherwise we do not see additional improvements over existing idea. In flagship results in table 2, such comparisons are warranted.\n\nAnother point is that plots like Fig 6, it seems that the x-axis measures the raw training samples experienced in the original dataset - however, the baseline leverages much more additional compute to arrive at a better performance, due to extra filtering etc. It will be useful to account for such additional compute when making such comparisons otherwise it is unfair to the baseline, one way is to compare against e.g. test time compute + baseline method to see if it's possible to make up for the performance gains, or increasing the number of parallel generation (group size) for baseline GRPO algorithm.\n\n[1] https://arxiv.org/abs/2505.03335\n\n=== *pass@1* ===\n\nWhile the paper motivates the writing with pass@1 metric, I don't see the connection between extra data augmentation with pass@1, after all we seek to augment data for which the model finds hard and try to patch such loopholes. It is not immediately clear to me why data augmentation addresses pass@1 issue at test time, it feels more like a generalization issue rather than an issue with metric (pass@1 vs. pass@32)"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper is interesting in that it provides another way to augment data for RLVR, leading to performance improvements. This affirms the knowledge that RLVR might benefit from synthetic data generation, and this helps push up even eval time performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I think overall the idea of synthetic data generation is not novel for RLVR, the idea of self-play and synthesizing new problems with filtering and verifications have been proposed and discussed in the literature. All in all, it feels that the idea is less novel and insights limited from the current paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917156940,"tcdate":1762600710563,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4059/Reviewer_yXhe"],"signatures":["ICLR.cc/2026/Conference/Submission4059/Reviewer_yXhe"],"forum":"Wjf3OMJxpn","number":4,"license":"CC BY 4.0","cdate":1762600710563,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4059/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917156940,"domain":"ICLR.cc/2026/Conference","replyto":"Wjf3OMJxpn","id":"oMK7VwOjmN","forumContent":{"TLDR":{"value":"We propose an online Self-play with Variational Problem Synthesis strategy for RLVR training that iteratively leverages model responses to synthesize variational problems for augmentation."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["LLM Reasoning; Reinforcement Learning; Self-envolving"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as a key paradigm for post-training Large Language Models (LLMs), particularly for complex reasoning tasks. However, vanilla RLVR training has been shown to improve Pass@1 performance at the expense of policy entropy, leading to reduced generation diversity and limiting the Pass@k performance, which typically represents the upper bound of LLM reasoning capability. In this paper, we systematically analyze the policy's generation diversity from the perspective of training problems and find that augmenting and updating training problems helps mitigate entropy collapse during training. Based on these observations, we propose an online Self-play with Variational problem Synthesis (SvS) strategy for RLVR training, which uses the policy's correct solutions to synthesize variational problems while ensuring their reference answers remain identical to the originals. This self-improving strategy effectively maintains policy entropy during training and substantially improves Pass@k compared with standard RLVR, sustaining prolonged improvements and achieving absolute gains of 18.3% and 22.8% in Pass@32 performance on the competition-level AIME24 and AIME25 benchmarks. Experiments on 12 reasoning benchmarks across varying model sizes from 3B to 32B consistently demonstrate the generalizability and robustness of SvS."},"_bibtex":{"value":"@inproceedings{\nliang2026beyond,\ntitle={Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains {RLVR}},\nauthor={Xiao Liang and Zhong-Zhi Li and Yeyun Gong and yelong shen and Ying Nian Wu and Zhijiang Guo and Weizhu Chen},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Wjf3OMJxpn}\n}"},"title":{"value":"Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVR"},"pdf":{"value":"/pdf/0d2ceab40f69d222a4fdfbb60fa04a9a233a8d2c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liang|beyond_pass_1_selfplay_with_variational_problem_synthesis_sustains_rlvr"},"authorids":{"value":["~Xiao_Liang9","~Zhong-Zhi_Li1","~Yeyun_Gong2","~yelong_shen1","~Ying_Nian_Wu1","~Zhijiang_Guo2","~Weizhu_Chen1"]},"authors":{"value":["Xiao Liang","Zhong-Zhi Li","Yeyun Gong","yelong shen","Ying Nian Wu","Zhijiang Guo","Weizhu Chen"]}},"version":2},{"content":{"review":{"value":"The paper is well motivated and addresses an important operational problem: CT scanner downtime due to X-ray tube failure in low-resource hospitals. The Moroccan healthcare context is clearly described, and the work is relevant to deployment-oriented AI in resource-constrained settings. The methodological contribution is interesting: the authors explicitly model data scarcity by calibrating a physics-based simulator from sparse hospital failure records and benchmarking several models under reduced training fractions.\n\nThe paper is generally clear and well structured. The methodology describes the degradation simulator, synthetic trajectory generation, RUL windowing, physics-informed monotonicity loss, and experimental comparison in sufficient detail. The results are also interpreted honestly: the authors do not overstate the benefit of the physics-informed model and show that simple LSTM models are more robust than complex architectures under scarcity.\n\nThe originality is moderate to good. Physics-informed RUL estimation and LSTM-based prognostics are not new, but applying them to CT X-ray tube maintenance in low-resource hospital settings is valuable and relatively underexplored.\n\n**Pros**\n- Important and practical problem for low-resource hospitals.\n- Strong contextual motivation using Moroccan public hospital data.\n- Clear focus on data scarcity and deployment constraints.\n- Good methodological structure: simulator calibration, synthetic trajectories, model benchmarking, and scarcity analysis.\n- Useful finding that simpler recurrent models are more robust than complex architectures under limited data.\n- Comparison includes SVR, Random Forest, LSTM, CNN-LSTM, Transformer, and PINN-LSTM.\n- Results are reported over multiple seeds with mean \\(\\pm\\) standard deviation.\n- Limitations are discussed honestly, especially the dependence on synthetic data.\n\n**Cons**\n- Main evaluation relies on synthetic trajectories rather than real CT sensor streams.\n- Simulator calibration is based on only 14 real tube failure events from two hospitals.\n- The simulator and physics-informed constraint share monotonic degradation assumptions, which may introduce some circularity.\n- No external validation on real hospital sensor data is provided.\n- The paper lacks important visual evidence such as predicted-vs-true RUL curves, degradation trajectories, error distributions, and uncertainty plots.\n- Only one physics-informed constraint and one value of \\(\\lambda = 0.3\\) are tested; sensitivity analysis would strengthen the conclusions.\n- The synthetic dataset contains only 100 tube trajectories, and failure modes appear dominated by filament failure.\n- The deployment pathway is discussed, but no prototype or real maintenance decision simulation is evaluated.\n\nOverall, this is a relevant and clearly written paper with a useful low-resource framing and an honest empirical message. However, its main conclusions should be interpreted as controlled simulation results rather than evidence of field-ready clinical deployment."},"confidence":{"value":4},"rating":{"value":6},"title":{"value":"The paper studies Remaining Useful Life (RUL) estimation for CT X-ray tubes under data scarcity in low-resource hospitals. The authors calibrate a physics-based degradation simulator using sparse real failure records from two Moroccan hospitals, generate synthetic RUL-labelled trajectories, and compare six models under decreasing training-data availability. The main finding is that simple recurrent models, especially LSTM-based approaches, are more robust under scarcity than more complex CNN-LSTM or Transformer architectures, while the proposed physics-informed monotonicity constraint provides limited and data-dependent benefits. The work is relevant, clearly motivated, and well aligned with predictive maintenance challenges in resource-constrained healthcare settings, but its clinical validity remains limited by the lack of real CT sensor-stream validation."}},"parentInvitations":"MICCAI.org/2026/Workshop/AFRICAI/-/Official_Review","nonreaders":[],"tmdate":1787850790678,"tcdate":1784755587765,"writers":["MICCAI.org/2026/Workshop/AFRICAI","MICCAI.org/2026/Workshop/AFRICAI/Submission10/Reviewer_m9oN"],"signatures":["MICCAI.org/2026/Workshop/AFRICAI/Submission10/Reviewer_m9oN"],"forum":"xPzjeyyEZz","number":2,"license":"CC BY 4.0","cdate":1784755587765,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/AFRICAI/Submission10/-/Official_Review","MICCAI.org/2026/Workshop/AFRICAI/-/Edit"],"mdate":1787850790678,"domain":"MICCAI.org/2026/Workshop/AFRICAI","replyto":"xPzjeyyEZz","id":"MHNGms6qPX","forumContent":{"TLDR":{"value":"Under clinical data scarcity, simple recurrent models beat complex architectures for CT X-ray tube RUL, while a physics-informed monotonicity constraint helps only when data is abundant."},"venue":{"value":"AFRICAI 2026 Oral"},"keywords":{"value":["Remaining Useful Life","CT X-ray Tubes","Data Scarcity"]},"venueid":{"value":"MICCAI.org/2026/Workshop/AFRICAI"},"paperhash":{"value":"oubnini|ct_xray_tube_rul_estimation_under_data_scarcity_in_lowresource_hospitals"},"authorids":{"value":["~ZAKARIA_OUBNINI1","wfmail@sina.com"]},"_bibtex":{"value":"@inproceedings{\noubnini2026ct,\ntitle={{CT} X-ray Tube {RUL} Estimation Under Data Scarcity in Low-Resource Hospitals},\nauthor={ZAKARIA OUBNINI and Wang Feng},\nbooktitle={First Workshop on Advancing African Medical AI through Global Integration},\nyear={2026},\nurl={https://openreview.net/forum?id=xPzjeyyEZz}\n}"},"title":{"value":"CT X-ray Tube RUL Estimation Under Data Scarcity in Low-Resource Hospitals"},"authors":{"value":["ZAKARIA OUBNINI","Wang Feng"]}},"version":2},{"content":{"summary":{"value":"The manuscript presents a unique approach to pre-training for offline reinforcement learning, utilizing synthetic data generated by a Markov Chain in lieu of traditional real-world language resources. The core premise is that this synthetic data can achieve comparable results to real-world data in downstream task performance, which is a significant assertion in the field of offline RL."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. Clarity of Presentation: The paper is well-structured, making it accessible even to those who may not be deeply versed in the domain. The significance of the research question is conveyed effectively, which facilitates a quick grasp of the paper's importance.\n\n2. Innovation in Data Construction: The methodology employed for the generation of synthetic data is both novel and straightforward, potentially offering a simpler alternative to more complex data generation strategies."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Methodological Justification: The rationale behind the adoption of a Markov Chain for synthetic data generation requires further elaboration. While the introduction suggests that understanding the underlying question is crucial for enhancing pre-training in deep reinforcement learning (DRL), the link between this understanding and the proposed method is not convincingly established.\n\n2. Need More Deep Analysis: The paper primarily demonstrates the efficacy of the proposed method without a robust analysis. It is advisable that the authors consider incorporating analysis akin to those found in the literature regarding Synthetic Data utilization in Transformer models (see Synthetic Pre-Training Tasks for Neural Machine Translation (ACL 2023) and its related works). This could potentially refine the proposed method and offer deeper insights through a more comprehensive analysis."},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. Could you provide a more detailed justification for the methodological choices, specifically the use of a Markov Chain for data synthesis?\n\n2. Are there illustrative examples of synthetic data that could be shared to better understand its characteristics and how it compares to real-world data?\n\n**After Rebuttal**\nThank you for the response. I have raised my score from 5 to 6 and my confidence is 2."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700819916999,"tcdate":1699277257448,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1152/Reviewer_dSfS"],"signatures":["ICLR.cc/2024/Conference/Submission1152/Reviewer_dSfS"],"forum":"PcxQgtHGj2","number":3,"license":"CC BY 4.0","cdate":1699277257448,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1152/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700819916999,"domain":"ICLR.cc/2024/Conference","replyto":"PcxQgtHGj2","id":"dB9dt84aHw","forumContent":{"venue":{"value":"ICLR 2024 poster"},"TLDR":{"value":"We show pre-training with synthetic Markov Chain and MDP data can significantly improve offline DRL performance."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Deep Reinforcement Learning","Offline Reinforcement Learning","Pretraining"]},"supplementary_material":{"value":"/attachment/25fde91c047ce8d710162318c68b791ee8f6d088.pdf"},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Recently, it has been shown that for offline deep reinforcement learning (DRL), pre-training Decision Transformer with a large language corpus can improve downstream performance (Reid et al., 2022). A natural question to ask is whether this performance gain can only be achieved with language pre-training, or can be achieved with simpler pre-training schemes which do not involve language. In this paper, we first show that language is not essential for improved performance, and indeed pre-training with synthetic IID data for a small number of updates can match the performance gains from pre-training with a large language corpus; moreover, pre-training with data generated by a one-step Markov chain can further improve the performance. Inspired by these experimental results, we then consider pre-training Conservative Q-Learning (CQL), a popular offline DRL algorithm, which is Q-learning-based and typically employs a Multi-Layer Perceptron (MLP) backbone. Surprisingly, pre-training with simple synthetic data for a small number of updates can also improve CQL, providing consistent performance improvement on D4RL Gym locomotion datasets. The results of this paper not only illustrate the importance of pre-training for offline DRL but also show that the pre-training data can be synthetic and generated with remarkably simple mechanisms."},"_bibtex":{"value":"@inproceedings{\nwang2024pretraining,\ntitle={Pre-training with Synthetic Data Helps Offline Reinforcement Learning},\nauthor={Zecheng Wang and Che Wang and Zixuan Dong and Keith W. Ross},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=PcxQgtHGj2}\n}"},"title":{"value":"Pre-training with Synthetic Data Helps Offline Reinforcement Learning"},"pdf":{"value":"/pdf/318129413e1d04dac4e9ef949bff927064665bee.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"wang|pretraining_with_synthetic_data_helps_offline_reinforcement_learning"},"authorids":{"value":["~Zecheng_Wang1","~Che_Wang1","~Zixuan_Dong1","~Keith_W._Ross1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zecheng Wang","Che Wang","Zixuan Dong","Keith W. Ross"]}},"version":2},{"content":{"summary":{"value":"This paper introduces I-PHYRE, a benchmark for intuitive physical reasoning capabilities in decision-making agents. It consists of four different block-removal games and benchmarks three planning strategies against them, implemented with both supervised and reinforcement learning. I-PHYRE's design centers on three principles: physical reasoning, multi-step planning, and in-situ intervention. \n\nThe four games are \"basic\" (teaching basic principles of physics), \"noisy\" (minor perturbations), \"compositional\" (combining various structures to require multi-step reasoning), and \"multi-ball\" (multiple dynamic events occurring concurrently, motivating carefully timed in-situ intervention). The planning strategies are \"planning-in-advance\" (generating an entire plan with timings based only on the initial state), \"planning-on-the-fly\" (generating the next action at each timestep given the observation - standard RL-type setup), and \"combined\" (generating the entire plan, then updating it after executing the first action). \n\nExperiments include results from model-free deep RL agents as well as some other baselines in supplementary on the games. Everything is compared to a human baseline. The paper specifically discusses how agents perform when generalizing (without further training) from the basic split to other splits, as well as how the three training strategies differ.\n\nFinally, the paper gives analysis of sources of difficulty in I-PHYRE, performance by offline algorithms, and limitations/future work."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"#### Quality\n- Overall, a solid paper. Intuitively, if an agent did well on I-PHYRE, I would believe that it had robust intuitive physics capabilities in certain domains, which is what a benchmark should convince me of.\n- Appendix G is very valuable. It does rest on the assumption that all failures are either of bad order or bad timing, which eliminates more basic failures, but since it is a simple simulator, and since Appendix G clearly shows that all the failures in Basic are timing-based (i.e. all the failed plans have correct order), I am convinced that the Basic split teaches nontrivial principles and the other splits \n\n#### Clarity\n- Very well-written! Clear language\n- The text organization is really useful, especially in the results section. The subsections make sense and each paragraph follows a claim-warrant structure. In terms of readability, papers often fall apart in the results section; this one does not. \n\n#### Originality\nI-PHYRE seems to have an original design. However, it's not the first interactive intuitive physics benchmark. The paper should compare to [1] and [2], though I do think it serves different, valuable purposes.\n\n#### Significance\nThe tasks in I-PHYRE are distinct from other related work and test compelling aspects of intuitive physics reasoning. So, I think if the paper can prove its claims, it is significant. \n\n[1] Physion: Evaluating Physical Prediction from Vision in Humans and Machines. Bear, D. M. et al. arxiv preprint: arXiv:2106.08261. 2021. \n[2] Jain, A. et al. \"Generalization to new actions in reinforcement learning.\" ICML 2020."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"#### Quality\n- The paper says that RL agents perform well on the noisy split, sort of justified by the fact that their noisy results \"correlate\" with their basic results. But that doesn't follow - the noisy results are clearly of lower magnitude even if they follow the same trends (across what factor of variation?), and we aren't given a correlation statistic to warrant this claim. The same applies to \"correlation diminishes in the compositional and multi-ball splits... inherent complexities of these tasks impact performance negatively.\" The correlation claim is similarly 1) nonobvious from looking at Fig 2, even if it does sort of look true, 2) not backed up by a number, 3) not clear why it matters - even if the performance on these two were correlated with perf on basic across whatever factor of variation, but lower in magnitude, I would accept that they are harder (and I of course do, just from looking at Fig 2). Then making a claim that this difficulty is due to the \"inherent complexities of these tasks\" is, while not unbelievable, strong. I'm not doubtful, just not convinced. \n- The discussion has a section titled \"why do current RL agents fail on I-PHYRE?\", which in my opinion is the most important aspect of a benchmark paper other than conceptual design and grounding in the environment. The paper claims three benefits: physics modeling being hard, multi-step interventions, and action timing - i.e. asserting that its design principles have effectively resulted in challenges for agents. However, this is all prose and little analysis of actual results - it would help to spell out for the reader which quantitative result comparison should lead to each conclusion. (Appendix G does a good job of this). \n\n#### Clarity\n- Figure 1 (repeated from a previous review I did of this paper, as the figure hasn't changed):\n  - Colors are hard to follow, maybe better to annotate split on box\n  - Since the boxes are much larger than the arrows, sort of look like two columns, and generally don't look like flowchart elements, it's confusing that the top-left box isn't the best one to start reading with. Bigger arrows and/or numbering would help.\n  - \"Wrong order leads to no overlap elimination timing for two balls\" - I don't understand this\n  - Compositional solution isn't that clear - annotation of key occurrences in each step might help\n- \"Combined\" planning strategy needs better explanation (a step-by-step could help) - it sounds like the whole plan is generated based on the initial state, then the first action is executed, then the plan is updated, and that's it. But I'm not totally sure if the entire plan is updated, if it's ever updated again later after subsequent actions, etc.\n- Figure 2 needs to be more organized and readable (in the previous version it was also hard to follow, but still more organized)\n  - Strategies should be put into separate subplots (with the same axes) or otherwise cued, especially since there are different numbers of each, making it hard to eyeball\n  - Baselines (especially the human comparison point) can be horizontal lines crossing the figure\n  - Not every algorithm gets every planning strategy, which doesn't seem to *only* be a function of the nature of the algorithm - so it's confusing to take all of this in just in the form of bars and text\n- Fig 3 would benefit from organization as well - e.g. group strategies by color and differentiate within them by line texture.\n\n#### Originality\n- The paper claims the three planning strategies as a contribution, but they aren't original - they are baselines just like the offline methods tested on I-PHYRE. I agree that the *results* are a contribution, but I would fold that into the third contribution bullet, or at least make it clear that the \"devised\" planning strategies themselves shouldn't be considered original. \n\nNits: \n- Different parts of the paper use different naming conventions for agents - e.g. \"SAC-I\" in one place, \"SAC Inadvance\" in another. Better to use the same thing throughout."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"- In a previous version of the paper, failure on I-PHYRE was ascribed to sparse action requirements and delayed reward. Now, those concepts are being pitched as being inherent to multi-step reasoning (delayed reward) and action timing (sparse action requirements), but it's not true that failing for those reasons means that if the agents were better at handling them, they would robustly learn to handle multi-step reasoning problems and in-situ intervention problems. Could you flesh this argument out more? \n- What exactly is the nature of the \"combined\" strategy, and could you say more on why it's effective?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635924214,"tcdate":1699124044115,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission9/Reviewer_BYPY"],"signatures":["ICLR.cc/2024/Conference/Submission9/Reviewer_BYPY"],"forum":"1bbPQShCT2","number":4,"license":"CC BY 4.0","cdate":1699124044115,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission9/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635924214,"domain":"ICLR.cc/2024/Conference","replyto":"1bbPQShCT2","id":"WJMxD7W7FR","forumContent":{"venue":{"value":"ICLR 2024 poster"},"TLDR":{"value":"Build human-level interactive agents that can control the physical dynamics by reasoning, planning, and intervention."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Intuitive physics","physical reasoning"]},"supplementary_material":{"value":"/attachment/bf4202ce3dfc8cb3a1fdecda0c659cf93a36497d.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Current evaluation protocols predominantly assess physical reasoning in stationary scenes, creating a gap in evaluating agents' abilities to interact with dynamic events. While contemporary methods allow agents to modify initial scene configurations and observe consequences, they lack the capability to interact with events in real time. To address this, we introduce I-PHYRE, a framework that challenges agents to simultaneously exhibit intuitive physical reasoning, multi-step planning, and in-situ intervention. Here, intuitive physical reasoning refers to a quick, approximate understanding of physics to address complex problems; multi-step denotes the need for extensive sequence planning in I-PHYRE, considering each intervention can significantly alter subsequent choices; and in-situ implies the necessity for timely object manipulation within a scene, where minor timing deviations can result in task failure. We formulate four game splits to scrutinize agents' learning and generalization of essential principles of interactive physical reasoning, fostering learning through interaction with representative scenarios. Our exploration involves three planning strategies and examines several supervised and reinforcement agents' zero-shot generalization proficiency on I-PHYRE. The outcomes highlight a notable gap between existing learning algorithms and human performance, emphasizing the imperative for more research in enhancing agents with interactive physical reasoning capabilities. The environment and baselines will be made publicly available."},"_bibtex":{"value":"@inproceedings{\nli2024iphyre,\ntitle={I-{PHYRE}: Interactive Physical Reasoning},\nauthor={Shiqian Li and Kewen Wu and Chi Zhang and Yixin Zhu},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=1bbPQShCT2}\n}"},"title":{"value":"I-PHYRE: Interactive Physical Reasoning"},"pdf":{"value":"/pdf/fad4695ed5caf629961f820cfffbd439e4662aa5.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"li|iphyre_interactive_physical_reasoning"},"authorids":{"value":["~Shiqian_Li1","~Kewen_Wu2","~Chi_Zhang12","~Yixin_Zhu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shiqian Li","Kewen Wu","Chi Zhang","Yixin Zhu"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS 2021"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-93409-5_40.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2021"},"paperhash":{"value":"jakovljevic|towards_building_a_digital_twin_of_complex_system_using_causal_modelling"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Luka_Jakovljevic:","https://dblp.org/search/pid/api?q=author:Dimitre_Kostadinov:","https://dblp.org/search/pid/api?q=author:Armen_Aghasaryan:","~Themis_Palpanas1"]},"html":{"value":"https://doi.org/10.1007/978-3-030-93409-5_40"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/JakovljevicKAP21,\n  author={Luka Jakovljevic and Dimitre Kostadinov and Armen Aghasaryan and Themis Palpanas},\n  title={Towards Building a Digital Twin of Complex System Using Causal Modelling},\n  year={2021},\n  cdate={1609459200000},\n  pages={475-486},\n  url={https://doi.org/10.1007/978-3-030-93409-5_40},\n  booktitle={COMPLEX NETWORKS},\n  crossref={conf/complexnetworks/2021-1}\n}\n"},"abstract":{"value":"Complex systems, such as communication networks, generate thousands of new data points about the system state every minute. Even if faults are rare events, they can easily propagate, which makes it challenging to distinguish root causes of errors from effects among the thousands of highly correlated alerts appearing simultaneously in high volumes of data. In this context, the need for automated Root Cause Analysis (RCA) tools emerges, along with the creation of a causal model of the real system, which can be regarded as a digital twin. The advantage of such model is twofold: (i) it assists in reasoning on the system state, given partial system observations; and (ii) it allows generating labelled synthetic data, in order to benchmark causal discovery techniques or create previously unseen faulty scenarios (counterfactual reasoning). The problem addressed in this paper is the creation of a causal model which can mimic the behavior of the real system by encoding the appearance, propagation and persistence of faults through time. The model extends Structural Causal Models (SCMs) with the use of logical noisy-OR gates to incorporate the time dimension and represent propagation behaviors. Finally, the soundness of the approach is experimentally verified by generating synthetic alert logs and discovering both the structure and parameters of the underlying causal model."},"title":{"value":"Towards Building a Digital Twin of Complex System Using Causal Modelling"},"authors":{"value":["Luka Jakovljevic","Dimitre Kostadinov","Armen Aghasaryan","Themis Palpanas"]}},"tmdate":1722418578474,"pdate":1609459200000,"tcdate":1722418526279,"writers":["~"],"signatures":["~Themis_Palpanas1"],"forum":"buhczYJh3X","license":"CC BY-SA 4.0","number":48576,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1722418578474,"domain":"DBLP.org","id":"buhczYJh3X","version":2},{"content":{"summary":{"value":"This paper proposes VIVIDCAM, a method for training text-to-video diffusion models to follow specific camera motions, particularly \"unconventional\" ones for which real-world training data is scarce. The core idea is to use low-poly, synthetic videos (rendered in Unity) as the training data.\n\nTo prevent the synthetic, \"low-poly\" appearance from bleeding into the final output, the authors propose a dual-adaptation training strategy. First, an \"appearance LoRA\" is trained on static synthetic videos to learn and isolate the synthetic visual style. This training is guided by a special \"virtual indicator\" tag in the prompt. Second, a camera motion module (either a LoRA for text-based control or fine-tuning a trajectory encoder) is trained on synthetic videos that do contain camera motion.\n\nAt inference time, the appearance LoRA is discarded, and the model is expected to produce realistic videos that follow the camera motions learned from the synthetic data. The method is evaluated on text-based and trajectory-based camera control, using automated metrics (FVD, TransErr) and human studies."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. The camera motion module (LoRA/encoder) is also trained with the <VIRTUAL> tag in the prompt. At inference, this tag is removed. How do you know the motion module isn't confused or degraded, as it was never trained on prompts without this tag? A key ablation is missing: what happens to motion accuracy if you keep the tag at inference (accepting the synthetic output)?\n\n2. Given that AnimateDiff's \"Domain Adapter\" + \"Motion Module\" training pipeline is functionally identical to this paper's, could the authors please state clearly what the technical novelty of this paper is, beyond just applying this pattern to a new synthetic dataset?\n\n3. Why was the decision made to degrade the perfect 3D trajectory data from Unity into coarse text prompts? Why not use the synthetic data for its primary strength and train a model that takes explicit 3D paths, target-of-interest coordinates, or specific orbit/dolly-zoom parameters as input?\n\n4. Table 1 lists \"Dolly zoom\" as a capability. A true Dolly zoom has a specific mathematical definition (simultaneously changing focal length and camera distance to keep the subject size constant). Can the text-based model actually reproduce this effect from a simple prompt, or is it just generating a standard \"zoom in\"?\n\n5. Could the authors provide a more in-depth analysis of why the appearance LoRA, trained only with a [VIRTUAL] tag, is sufficient to \"absorb\" all synthetic artifacts? The mechanism seems very similar to DreamBooth (specializing to a token) but is used for a different purpose (disentanglement), and this is not well-justified."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The problem itself is significant. Enabling controllable, complex, and artistic camera motions is a key challenge for creative video generation, and the lack of diverse, well-labeled real-world data is a real bottleneck.\n\nThe idea of using synthetic data is a logical approach to solving this data scarcity problem."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper suffers from a significant lack of novelty, a mismatch between its tools and goals, and an unconvincing evaluation.\n\n1. Critical Lack of Novelty: The central contribution, the \"dual adaptation\" or \"dual LoRA\" method for disentangling appearance from motion, is not new. This exact training pattern was established by AnimateDiff (Guo et al., 2023), which the authors cite as \"inspiration.\" AnimateDiff's \"Stage 1: Domain Adapter\" is functionally identical to this paper's \"Step 1: Appearance Adaptation.\" VIVIDCAM is a direct application of this existing 2-stage (adapter-then-motion) training strategy to a synthetic dataset. This is an incremental extension, not a novel framework, and represents a major overstatement of the paper's contribution.\n\n2. Mismatch of Tool and Goal: The paper uses a powerful 3D rendering engine (Unity) capable of generating perfect, precise, 6-DOF camera trajectories (i.e., extrinsics). However, a large part of the paper focuses on using this data to train a model on vague, ambiguous text prompts like \"pan left\" or \"push in.\" This is a baffling waste of the synthetic data's primary advantage. It's a step backward from existing works (including its own baselines like AC3D and CameraCtrl) that are focused on fine-grained, trajectory-based control.\n\n3. Unconvincing Rationale: The paper motivates the work by the need for \"unconventional\" motions. However, the motions listed in Table 1 (\"Simple\" and \"Composed\") are entirely conventional (push, tilt, truck). These are readily available in existing real-world datasets. While the \"Complex\" motions (e.g., \"seek object\") are more interesting, training them from ambiguous text prompts is far less compelling than training a model to follow an explicit 3D \"seeking\" path, which Unity could have easily generated.\n\n4. Insufficient Evaluation & Unconvincing Ablations: The primary evidence for the core disentanglement claim rests on the ablation in Figure 5. This figure not only fails to show temporal artifacts (by only showing static frames), but the visual difference between the full model and the \"w/o Style-aligned Prompts\" version is minimal. This does not convincingly demonstrate that the appearance LoRA and [VIRTUAL] tag are doing the heavy lifting the authors claim. Furthermore, the paper lacks direct side-by-side video comparisons to its baselines, making it impossible to truly judge the trade-off between motion accuracy and visual fidelity.\n\n5. Poor Qualitative Fidelity: When compared to the baselines (e.g., vanilla CogVideoX or AC3D), the visual performance of the proposed method appears to be a significant regression. The generated videos, while perhaps following the motion, are less realistic and have lower fidelity than the state-of-the-art baselines they are built upon. The method seems to trade visual quality for motion control, which is a poor trade-off."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920062483,"tcdate":1761887470039,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8073/Reviewer_5sQc"],"signatures":["ICLR.cc/2026/Conference/Submission8073/Reviewer_5sQc"],"forum":"qYhYrj91l2","number":4,"license":"CC BY 4.0","cdate":1761887470039,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8073/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920062483,"domain":"ICLR.cc/2026/Conference","replyto":"qYhYrj91l2","id":"pnSE8NuW3E","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"VividCam learns unconventional camera motions from surprisingly simple synthetic data, producing videos with motions that are often challenging for existing methods."},"keywords":{"value":["controllable video generation","camera control","diffusion models"]},"supplementary_material":{"value":"/attachment/d126123afd57b8783a0cceb17ddf5902c9bb1a99.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions, which is crucial in creating truly original and artistic videos. The challenge lies in the difficulty of finding sufficient training videos with the intended uncommon camera motions. To address this challenge, we propose VividCam, a training paradigm that enables diffusion models to learn complex camera motions from synthetic videos, releasing the reliance on collecting realistic training videos. VividCam incorporates multiple disentanglement strategies that isolates camera motion learning from synthetic appearance artifacts, ensuring more robust motion representation and mitigating domain shift. We demonstrate that our design synthesizes a wide range of precisely controlled and complex camera motions using surprisingly simple synthetic data. Notably, this synthetic data often consists of basic geometries within a low-poly 3D scene and can be efficiently rendered by engines like Unity. Our video results can be found in https://anonymoususers196.github.io/VividCamDemo/ ."},"_bibtex":{"value":"@misc{\nwu2026vividcam,\ntitle={VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos},\nauthor={Qiucheng Wu and Handong Zhao and Zhixin Shu and Jing Shi and Yang Zhang and Shiyu Chang},\nyear={2026},\nurl={https://openreview.net/forum?id=qYhYrj91l2}\n}"},"title":{"value":"VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos"},"pdf":{"value":"/pdf/ed7392347f737aeda8fad126bc8db60f8f4b6c86.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wu|vividcam_learning_unconventional_camera_motions_from_virtual_synthetic_videos"},"authorids":{"value":["~Qiucheng_Wu1","~Handong_Zhao3","~Zhixin_Shu1","~Jing_Shi1","~Yang_Zhang3","~Shiyu_Chang2"]},"authors":{"value":["Qiucheng Wu","Handong Zhao","Zhixin Shu","Jing Shi","Yang Zhang","Shiyu Chang"]}},"version":2},{"content":{"summary":{"value":"This paper presents RingLight-GS, a method that enhances 3D Gaussian Splatting (3DGS) by improving compactness and handling of complex lighting effects. Traditional 3DGS relies on large Spherical Harmonics (SH) coefficients to represent color, leading to high memory consumption. RingLight-GS addresses this by decomposing each Gaussian’s color into a view-independent base color and a view-dependent residual color. The core contribution is a Neural Tensor Ring (TR) Regression model that learns the view-dependent component. This TR model maps spatial coordinates and viewing directions to appearance features through a decomposed tensor representation that is both compact and expressive. As a result, RingLight-GS achieves approximately 2.5–3× storage reduction compared to standard 3DGS, while improving rendering quality, particularly for specular highlights and reflections."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"The Neural TR module plays a central role in modeling view-dependent effects. It remains unclear how its efficiency and compactness scale with scene complexity. In scenes with substantially more Gaussians or a larger spatial extent, would preserving rendering quality necessitate a proportional increase in TR rank or feature dimension, thereby diminishing the storage benefits?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Separating the base color from a tensor-factorized residual is an elegant and effective design. It reduces 3DGS’s memory overhead and improves its ability to represent complex, view-dependent effects like specular highlights that SHs handle poorly.\n\n2. The method achieves both compactness and improved performance: it reduces model size substantially (e.g., from 68 MB to 22 MB on NeRF-Synthetic) while also attaining higher PSNR, SSIM, and LPIPS scores across multiple datasets, outperforming existing compression and lighting-aware 3DGS variants."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The advanced Neural TR model incurs a computational cost. The paper reports substantially longer training times (e.g., ~29 minutes versus ~15 minutes for standard 3DGS on Shiny-Blender) and increased VRAM usage during training, making it less efficient to train and deploy than the original 3DGS.\n\nThe method introduces several new components (neural TR, render module, loss scaling factor γ) and associated hyperparameters (TR rank, feature dimension). The ablation studies show performance degrades if these are not set correctly, adding complexity and potentially limiting its ease of adoption.\n\nIn Figure 2, which aims to illustrate the overall scene color modeling framework, is too schematic and abstract. It uses generic blocks (e.g., \"TR Function Regression,\" \"Render Module\") without visually hinting at their unique internal mechanics (e.g., how the neural TR cores connect). This makes it difficult for a reader to grasp the innovative data flow and architecture without constantly cross-referencing the complex text in Section 3.2. A more detailed or annotated diagram would be greatly beneficial."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922216871,"tcdate":1761972747928,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11038/Reviewer_UJPX"],"signatures":["ICLR.cc/2026/Conference/Submission11038/Reviewer_UJPX"],"forum":"Lhsy6DYUtS","number":4,"license":"CC BY 4.0","cdate":1761972747928,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11038/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922216871,"domain":"ICLR.cc/2026/Conference","replyto":"Lhsy6DYUtS","id":"kZtkr5HN5H","forumContent":{"TLDR":{"value":"We present RingLight-GS, a compact view-dependent rendering framework that factorizes appearance into base color and directional residual, using Neural Tensor Ring decomposition to model high-frequency lighting with low storage."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Gaussian Splatting","Neural Radiance Field","Tensor Decomposition"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"3D Gaussian Splatting (3DGS) achieves impressive novel view synthesis in real-time by directly rendering Gaussian primitives. However, it incurs substantial storage demands and struggles to model high-frequency, view-dependent appearance effects under complex illumination. We introduce RingLight-GS, a compact framework that effectively models scene color in 3DGS, delivering high-quality rendering under complex lighting while greatly reducing storage costs. The scene color is separated into a view-independent base color and a view-dependent residual color by disentangling static albedo from dynamic lighting, with the base color learning similarity to the 3DGS opacity. Specifically, the residual color is derived from view-dependent appearance features via a neural tensor ring regression model, influenced by spatial positions and viewing directions. Extensive experiments on synthetic and real-world datasets demonstrate that RingLight-GS consistently outperforms both NeRF-based and 3DGS-based baselines. It delivers sharper highlights, better material consistency, and lower perceptual error with minimal memory overhead."},"_bibtex":{"value":"@misc{\nsheng2026ringlightgs,\ntitle={RingLight-{GS}: A compact and expressive framework for modeling scene color in 3{DGS}},\nauthor={Dian Sheng and Zhen Long and Kang Tan and Yingge Li and Xinyu Lin and Ce Zhu and Fan Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=Lhsy6DYUtS}\n}"},"title":{"value":"RingLight-GS: A compact and expressive framework for modeling scene color in 3DGS"},"pdf":{"value":"/pdf/fdd7a92a211a79906ac575badb421890a3eec05c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"sheng|ringlightgs_a_compact_and_expressive_framework_for_modeling_scene_color_in_3dgs"},"authorids":{"value":["~Dian_Sheng1","~Zhen_Long2","~Kang_Tan1","~Yingge_Li1","~Xinyu_Lin8","~Ce_Zhu1","~Fan_Zhang1"]},"authors":{"value":["Dian Sheng","Zhen Long","Kang Tan","Yingge Li","Xinyu Lin","Ce Zhu","Fan Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a framework to unify the influence-based reward into model training to guide the synthetic data generation.\nThe motivation stems from the observation that current synthetic data pipelines rely heavily on handcrafted rubrics and heuristic feedback loops, which are expensive, brittle, and poorly correlated with true downstream model performance. The authors introduce an influence-function–based estimator that quantifies each synthetic sample’s contribution to the fine-tuned model’s objective. This influence signal is then used to guide rubric optimization and synthetic data selection via a lightweight optimization loop, replacing manual rubric engineering with model-informed adaptation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Can the authors provide quantitative results about efficiency?\n2. Have the authors compared their influence-based signal with alternative feedback mechanisms?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"1. Motivation is clear. Synthetic data generation requires a lot rubric engineering and hard to scaling.\n2. Frameworks is good. Put influence functions into RL framework to make the whole procedure automatic. \n3. Good performance. The improvements are consistent and statistically meaningful."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Novelty and Conceptual Depth. While the paper presents a plausible framework, the underlying technical innovation appears incremental. The idea of integrating quantitative signals to guide synthetic data generation is intuitive and has been explored under different names (e.g., data valuation, reward-guided synthesis, active selection). The contribution primarily involves adopting influence-based rewards for rubric optimization, which builds directly on prior influence-function research rather than introducing a fundamentally new mechanism. As a result, the work feels more like a thoughtful application of existing tools than a conceptual breakthrough.\n2. Scalability and Efficiency Unclear. Traditional influence functions are notoriously computationally expensive. The paper claims to deploy an “optimizer-aware estimator” that scales efficiently, but does not provide sufficient empirical evidence or runtime comparisons to support this claim. Since the method is applied per-sample, a clear analysis of runtime overhead and system-level feasibility is crucial to evaluate its practicality for real-world synthetic data pipelines.\n3. Questionable Necessity of the Influence Signal. The paper argues that influence estimation provides a better feedback signal for synthetic data quality, but the empirical justification is not fully convincing. The improvement margins over simpler proxy metrics (e.g., downstream reward, embedding-similarity heuristics, or small proxy validation sets) are modest. Without ablation or sensitivity studies, it remains unclear whether the influence-based signal is truly superior or merely another heuristic choice. A comparison against alternative feedback mechanisms would strengthen the paper’s argument."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923695275,"tcdate":1761962517545,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12931/Reviewer_WCjb"],"signatures":["ICLR.cc/2026/Conference/Submission12931/Reviewer_WCjb"],"forum":"vFcm5sOitq","number":3,"license":"CC BY 4.0","cdate":1761962517545,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12931/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923695275,"domain":"ICLR.cc/2026/Conference","replyto":"vFcm5sOitq","id":"XyllAYhRcY","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["LLMs","data synthetic","instruction tuning"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large language models (LLMs) achieve strong downstream performance largely due to abundant supervised fine-tuning (SFT) data that imparts problem-solving capabilities. However, as applications expand, high-quality SFT data in knowledge-intensive verticals (e.g., humanities and social sciences, medicine, law, finance) is exceedingly scarce: expert curation is costly, privacy constraints are strict, and label consistency is hard to guarantee. Recent work turns to synthetic data, typically prompting a teacher model over domain documents and filtering with handcrafted rubrics. Yet, rubric design is expert-dependent and rarely transfers across domains; moreover, prevalent heuristic optimization follows a brittle loop (write rubric $\\rightarrow$ synthesize $\\rightarrow$ train $\\rightarrow$ inspect $\\rightarrow$ guess tweaks) that lacks reliable, quantitative feedback about a rubric's true contribution to downstream performance.\nWe argue for assessing synthetic data quality through its causal impact on the target model, using this feedback to guide data generation. Inspired by classic influence functions, we repurpose an optimizer-aware estimator that uses gradient information to quantify each synthetic sample's contribution to the objective of a given target model on specific tasks. Our analysis reveals a gap: although synthetic and real samples may be close in embedding space, their influence on learning can differ substantially. Building on this insight, we propose an optimization-based synthetic data framework that adapts rubrics with target-model feedback. Instead of manually engineering domain rubrics, we supply lightweight guiding text and delegate rubric generation to a rubric-specialized model conditioned on the task; crucially, rubric (and data) selection is supervised by estimated downstream impact rather than proxy formality. Empirically, the framework yields consistent gains across domains (HSS and health), target models (e.g., Qwen and Llama families), and data generators, demonstrating broad generalization and engineering portability without task-specific tuning."},"_bibtex":{"value":"@inproceedings{\nfan2026optimsyn,\ntitle={OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation},\nauthor={Zhiting Fan and Ruizhe Chen and Tianxiang Hu and Ru Peng and Zenan Huang and Haokai Xu and Yixin Chen and Jian Wu and Junbo Zhao and Zuozhu Liu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=vFcm5sOitq}\n}"},"title":{"value":"OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation"},"pdf":{"value":"/pdf/f96b5a93c80ede4022e02eb8d4f866ebc44ac43e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"fan|optimsyn_influenceguided_rubrics_optimization_for_synthetic_data_generation"},"authorids":{"value":["~Zhiting_Fan1","~Ruizhe_Chen1","~Tianxiang_Hu2","~Ru_Peng1","~Zenan_Huang1","~Haokai_Xu1","~Yixin_Chen27","~Jian_Wu6","~Junbo_Zhao1","~Zuozhu_Liu1"]},"authors":{"value":["Zhiting Fan","Ruizhe Chen","Tianxiang Hu","Ru Peng","Zenan Huang","Haokai Xu","Yixin Chen","Jian Wu","Junbo Zhao","Zuozhu Liu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes PHDME (Port-Hamiltonian Diffusion Model), a hybrid generative modeling framework that integrates Gaussian-process distributed Port-Hamiltonian systems (GP-dPHS) with diffusion models to learn physically consistent dynamics without explicit governing equations. A two-step process is put forth:\n\n- GP-dPHS training: limited observations are used to learn a probabilistic Hamiltonian representation of the dynamics, including uncertainty.\n\n- Diffusion model training: the learned physics prior (Hamiltonian structure and uncertainty) is embedded into the diffusion model’s loss function, enforcing energy-based consistency across diffusion steps.\n\nThen, conformal prediction is used post-hoc for uncertainty calibration. The method is evaluated on a 1-D nonlinear wave-like PDE (string vibration), showing faster generation and lower mean-square error than purely data-driven or partially physics-aware diffusion baselines."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- How does PHDME differ conceptually and empirically from prior physics-informed diffusion models (e.g., Bastek et al. 2024) and Latent SDE frameworks?\n\n- What is gained by coupling a GP-dPHS prior specifically, versus directly enforcing energy conservation or using a learned Hamiltonian neural network?\n\n- Why are methods like the Deep Markov Model, Neural ODE/SDE, or Neural Operator not included for comparison? These are natural comparators for dynamics learning without explicit equations.\n\n- Also, while classical Hamiltonian Neural Networks (HNNs) are deterministic, stochastic extensions do exist and could be used here as baselines to compare against.\n\n- How does the GP-dPHS scale with system dimensionality and number of spatial nodes?\n\n- Can the approach handle higher-dimensional PDEs or coupled multi-field systems?\n\n- How sensitive is the framework to noise or imperfect observation coverage?\n\n- Does the GP-uncertainty weighting truly improve robustness or simply regularize training?\n\n- Can the learned Hamiltonian be inspected or visualized to confirm that it corresponds to physically meaningful energy terms?\n\n- Have any field or experimental datasets (e.g., soft robotics, structural dynamics) been tested?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Combines physics-informed priors (via GP-based Port-Hamiltonian systems) with score-based diffusion modeling. This bridges structured physics modeling and generative uncertainty modeling.\n\n- Unlike PINNs or traditional physics-informed diffusion models, it infers latent physical structure directly from data.\n\n- The approach leverages limited observations to build a probabilistic physics prior before training the diffusion model.\n\n- Built-in uncertainty quantification: uses both GP posterior variance (during training) and conformal prediction (post-training) for calibrated uncertainty estimates.\n\n- Single-shot field generation avoids step-by-step numerical integration (claimed ~20× faster than GP-dPHS rollouts).\n\n- The two-stage process and the Hamiltonian residual penalty are well-defined and interpretable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Only a single PDE (the 1-D wave/string system) is tested; no real or high-dimensional datasets are used. Claims of generality (e.g., to soft robots, elasticity) are therefore speculative.\n\n- The “scarce data” scenario is simulated, but robustness to measurement noise or model misspecification is not demonstrated.\n\n- The reviewer would appreciate an elaboration on the contribution over existing GP-dPHS work (Beckers et al., 2022; Tan et al., 2024). The paper primarily extends these with a diffusion-based generator, but the gain over standard GP-dPHS rollout or other physics-aware surrogates is modest and not deeply analyzed.\n\n- Picking up from the above, it would help to explicitly quantify the benefit of each component (GP uncertainty weighting, Hamiltonian penalty, conformal calibration) (via an ablation study)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941874050,"tcdate":1761935442614,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21657/Reviewer_yDyr"],"signatures":["ICLR.cc/2026/Conference/Submission21657/Reviewer_yDyr"],"forum":"kTEOG9a2W3","number":2,"license":"CC BY 4.0","cdate":1761935442614,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21657/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941874050,"domain":"ICLR.cc/2026/Conference","replyto":"kTEOG9a2W3","id":"wnFQ7baoOF","forumContent":{"TLDR":{"value":"PHDM: diffusion guided by GP-learned Port-Hamiltonian energy gradients from sparse data. No exact fuction needed; structure is conserved and uncertainty is calibrated for spatiotemporal prediction."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed Learning","Diffusion Model","Port-Hamiltonian system","Uncertainty Quantification","Gaussian process"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Diffusion models are expressive priors for generating and predicting data from high-dimensional dynamical systems. Yet, purely data-driven approaches often lack reliability and trustworthiness, motivating growing interest in physics-informed machine learning (PIML). Most existing PIML methods, however, assume access to exact governing equations during training—an assumption that fails when the dynamics are unknown or too complex to model accurately. To address this gap, we introduce PHDME (Port-Hamiltonian Diffusion Model), a physics-informed diffusion framework that learns system dynamics without requiring exact equations. Our approach first trains a Gaussian process distributed Port-Hamiltonian system (GP-dPHS) on limited observations to capture an energy-based representation of the dynamics. The GP-dPHS is then used to generate a physically consistent and diverse dataset for diffusion training. To enforce physics-consistency, we embed the GP-dPHS structure directly into the diffusion training objective through a loss that penalizes deviations from the learned Hamiltonian dynamics, weighted by the GP’s predictive uncertainty. After training, we employ conformal prediction to provide distribution-free uncertainty quantification of the generated trajectories. In this way, PHDME is designed for regimes with scarce data and unknown equations, enabling data-efficient, physically valid trajectory generation with calibrated uncertainty estimates."},"_bibtex":{"value":"@misc{\ntan2026phdme,\ntitle={{PHDME}: Physics-Informed Diffusion Models without Explicit Governing Equations},\nauthor={Kaiyuan Tan and Kendra Lee Givens and Peilun Li and Thomas Beckers},\nyear={2026},\nurl={https://openreview.net/forum?id=kTEOG9a2W3}\n}"},"title":{"value":"PHDME: Physics-Informed Diffusion Models without Explicit Governing Equations"},"pdf":{"value":"/pdf/caa63d9bc69a07229f5b0d48dcdca801e0cc83ab.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"tan|phdme_physicsinformed_diffusion_models_without_explicit_governing_equations"},"authorids":{"value":["~Kaiyuan_Tan1","~Kendra_Lee_Givens1","~Peilun_Li3","~Thomas_Beckers1"]},"authors":{"value":["Kaiyuan Tan","Kendra Lee Givens","Peilun Li","Thomas Beckers"]}},"version":2},{"content":{"summary":{"value":"This paper introduces 3DGSim, an end-to-end differentiable framework that learns 3D physical simulation directly from multi-view RGB videos. 3DGSim combines a feed-forward inverse renderer based on MVSplat for reconstructing 3D Gaussian particles, a transformer-only dynamics model leveraging space-filling curves and temporal merging for efficient spatiotemporal reasoning, and a Gaussian splatting renderer for differentiable image supervision. Trained solely with image reconstruction loss, the model learns latent visuo-physical particle features that capture diverse dynamics without explicit physics priors. Experiments demonstrate that 3DGSim achieves high-fidelity, physically plausible long-horizon predictions, robust generalization to unseen scenarios such as ground removal or multi-object interactions."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Part of the questions are listed in the weakness part.\n- In Figure 12, what exactly do “input” and “reconstruction” views refer to? Is the “12 (4+5)” notation a typo or intentional (since 4+5 ≠ 12)?\n- In Figure 14, both “inference speed” and “cloth/rigid data FPS” are reported. What is the difference between inference FPS and dataset FPS? Is there any difference in inference FPS between the two dataset?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"-  The model generalizes to unseen scenarios such as ground removal and multi-object interactions, demonstrating robust physical understanding and scene editability.\n-  Extensive experiments across rigid, elastic, and cloth datasets show physically plausible long-horizon rollouts with strong quantitative and qualitative performance over 2D baselines like Cosmos."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper does not provide quantitative comparisons with existing 3D physical simulators. Although the authors mention that code and data for these baselines are unavailable, the absence of such comparisons weakens the rigor and positioning of the work.\n- Several implementation and presentation details are omitted or unclear, which may cause confusion for readers:\n  - The variable $p$ in Equation (2) is not explicitly defined.\n  - The process of extracting the feature $f$ is not clearly explained.\n  - The paper does not specify what dataset(s) were used for training or how many data samples are included.\n  - The meaning of “Groundtruth” in Figure 7 is unclear.\n  - It is not explained how the ground plane is represented or modeled in the method.\n- All experiments are on synthetic datasets. The model's performance on real-world videos remains to be verified.\n- It would be better if the author could add some ablation study of key design parts.\n- Missing related work: The proposed idea is also related to [1].\n\n[1] Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D Video. ICLR 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923853098,"tcdate":1761664824152,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13130/Reviewer_jav5"],"signatures":["ICLR.cc/2026/Conference/Submission13130/Reviewer_jav5"],"forum":"qwsCjNSHMz","number":1,"license":"CC BY 4.0","cdate":1761664824152,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13130/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923853098,"domain":"ICLR.cc/2026/Conference","replyto":"qwsCjNSHMz","id":"MEc9m9TtTa","forumContent":{"TLDR":{"value":"3DGSim learns 3D simulators from RGB videos by jointly training inverse rendering and dynamics forecasting."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["simulation learning","dynamics","forecasting","particle dynamics","learning from videos"]},"supplementary_material":{"value":"/attachment/6b2adca1cd3216ff4a3fa8411b518a0c3c17b7e3.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Realistic simulation is critical for applications ranging from robotics to animation. \nVideo generation models have emerged as a way to capture real-world physics from data, but they often face challenges in maintaining spatial consistency and object permanence, relying on memory mechanisms to compensate. \nAs a complementary direction, we present 3DGSim, a learned 3D simulator that directly learns physical interactions from multi-view RGB videos.\n3DGSim adopts MVSplat to learn a latent particle-based representation of 3D scenes, a Point Transformer for the particle dynamics, a Temporal Merging module for consistent temporal aggregation, and Gaussian Splatting to produce novel view renderings.\nBy jointly training inverse rendering and dynamics forecasting, 3DGSim embeds physical properties into point-wise latent features. This enables the model to capture diverse behaviors, from rigid and elastic to cloth-like dynamics and boundary conditions (e.g., fixed cloth corners), while producing realistic lighting effects. We show that 3DGSim can generate physically plausible results even in out of distribution cases, e.g. ground removal or multi-object interactions, despite being trained only on single-body collisions."},"_bibtex":{"value":"@misc{\nzhobro2026learning,\ntitle={Learning 3D-Gaussian Simulators from {RGB} Videos},\nauthor={Mikel Zhobro and Andreas Ren{\\'e} Geist and Georg Martius},\nyear={2026},\nurl={https://openreview.net/forum?id=qwsCjNSHMz}\n}"},"title":{"value":"Learning 3D-Gaussian Simulators from RGB Videos"},"pdf":{"value":"/pdf/7c4838c9e94225e4a8affa86c5379875eeb9057f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhobro|learning_3dgaussian_simulators_from_rgb_videos"},"authorids":{"value":["~Mikel_Zhobro1","~Andreas_René_Geist1","~Georg_Martius1"]},"authors":{"value":["Mikel Zhobro","Andreas René Geist","Georg Martius"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a benchmark for evaluating episodic memory capabilities in LLMs. The authors create a framework inspired by cognitive science to model episodic events with temporal and spatial contexts, entities, and detailed descriptions. They generate synthetic datasets and evaluate state-of-the-art LLMs across various recall and episodic reasoning tasks. The evaluation considers different memory strategies: in-context learning, RAG, and fine-tuning. The authors observe that even advanced models face challenges in handling episodic memory tasks, particularly when recalling sequences of related events or complex spatiotemporal relationships."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.How does the synthetic data generation process ensure realistic temporal and causal relationships between events?  \n2.Have you conducted rigorous validation studies comparing LLM judgments against human annotations or established metrics? What specific measures were taken to ensure consistency and reproducibility in the evaluation process?  \n3.How reliable is the process of scoring relevance \"against each ground truth item\"? Could you provide examples of how partial matches are handled?  \n4.In Table 3, the fine-tuned model performs well on single-event queries (F1=0.83) but poorly on multi-event queries (F1≤0.37). Could you elaborate on why naive fine-tuning fails to generalize beyond single-event memorization? What specific architectural or training modifications might address this limitation?  \n5.The gradient pattern in Figure 3 shows degrading performance from context to space to time cues. What specific aspects of temporal reasoning make it particularly challenging for current LLMs?  \n6.For the \"Latest state recall\" results in Table 4, what specific challenges prevent models from achieving higher accuracy in tracking entity states over time?  \n7.Have you tried other fine-tuning approaches beyond single-event memorization that might better capture the hierarchical and relational nature of episodic memory?  \n8.Have you tried other retrieval strategies beyond cosine similarity? How do you address the challenge of retrieving coherent information when relevant context is distributed across multiple chunks?  \n9.How might this benchmark contribute to developing novel training methodologies for episodic memory tasks in LLMs, beyond RAG, fine-tuning or parametric memory storage?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.The paper tries to address an important and timely challenge about the need for better episodic memory capabilities in LLMs;  \n2.The research takes a structured approach to modeling episodic memory by incorporating key concepts from cognitive science, focusing on key aspects of memory: temporal context, spatial grounding, and entity tracking;  \n3.The methodology demonstrates rigor through: (1) creating contamination-free synthetic benchmarks, (2) introducing multiple verification steps to ensure data quality, (3) providing flexibility in generating datasets of different sizes and complexities;  \n4.The study develops systematic ways to assess different aspects of memory (recall, chronological ordering, latest state)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The paper primarily utilizes LLM-generated synthetic data (Section 4.1), but does not adequately validate the quality and representativeness of the generated narratives. For example, while the authors claim to verify \"adherence to event meta-data,\" they do not provide quantitative metrics for assessing narrative coherence or natural language properties. The authors should establish clear validation metrics and demonstrate how their synthetic data captures the essential properties of real episodic memories.  \n2.The scope of the benchmark is unnecessarily limited. The current implementation: (1) only considers fictional narratives with human-like protagonists, (2) Uses oversimplified temporal representations, (3) Fails to address complex episodic memory scenarios involving interconnected events. The authors should expand the benchmark to include more diverse scenarios, complex temporal relationships, and interconnected event sequences that more accurately reflect real-world episodic memory challenges.  \n3.While Section 3.1 emphasizes the importance of entity state tracking, the experimental results in Table 3 do not adequately measure this capability. The evaluation focuses on simple recall rather than complex state changes. The paper claims to test \"understanding temporal sequences\" but does not properly evaluate how models handle causally related state changes. The authors should design specific test cases for complex state tracking, evaluate models' ability to handle causally related state changes, and include metrics for measuring state tracking accuracy.  \n4.The LLM-as-judge approach described in Section 4.3 lacks validation of inter-judge consistency across different evaluator LLMs and does not establish correlation with human judgments. This could be addressed by including human evaluation benchmarks and demonstrating consistent assessments across multiple judge models.  \n5.The RAG experiments in Section 5.1 use only basic paragraph-level chunking without exploring alternative strategies. The authors should investigate alternative chunking approaches, compare different retrieval mechanisms, and analyze how these choices impact episodic memory performance.  \n6.While Table 4 shows poor performance in chronological ordering tasks, the paper doesn't provide detailed error analysis or investigate specific failure patterns. The analysis in Section 5.2 focuses on aggregate metrics without examining individual failure cases. The authors should provide detailed case studies of failure modes, analyze patterns in chronological ordering errors, and investigate whether specific temporal relationships consistently challenge the models.  \n7.Although Section 5.2 mentions testing for hallucinations, the analysis is limited. The paper fails to examine when and why models confabulate, or how confabulation patterns vary across different model architectures and memory strategies. This could be improved by designing specific experiments to probe confabulation triggers and providing metrics for measuring confabulation severity."}},"nonreaders":[],"tmdate":1732770804565,"tcdate":1730127761307,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12017/Reviewer_uWQ8"],"signatures":["ICLR.cc/2025/Conference/Submission12017/Reviewer_uWQ8"],"forum":"6ycX677p2l","number":2,"license":"CC BY 4.0","cdate":1730127761307,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12017/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732770804565,"domain":"ICLR.cc/2025/Conference","replyto":"6ycX677p2l","id":"cgJFwcbq2p","forumContent":{"TLDR":{"value":"A new benchmark for testing episodic memory in AI language models, using synthetic stories to evaluate recall of specific events, their contexts, and chronological relationships."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Episodic Memory Modeling","Large Language Models","Synthetic Benchmark Generation","Cue-based Retrieval","Temporal-Spatial Reasoning","Long-context Understanding","Human-inspired AI"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Episodic memory -- the ability to recall specific events grounded in time and space -- is a cornerstone of human cognition, enabling not only coherent storytelling, but also planning and decision-making. Despite their remarkable capabilities, Large Language Models (LLMs) lack a robust mechanism for episodic memory: we argue that integrating episodic memory capabilities into LLM is essential for advancing AI towards human-like cognition, increasing their potential to reason consistently and ground their output in real-world episodic events, hence avoiding confabulations. To address this challenge, we introduce a comprehensive framework to model and evaluate LLM episodic memory capabilities. Drawing inspiration from cognitive science, we develop a structured approach to represent episodic events, encapsulating temporal and spatial contexts, involved entities, and detailed descriptions. We synthesize a unique episodic memory benchmark, free from contamination, and release open source code and datasets to assess LLM performance across various recall and episodic reasoning tasks. Our evaluation of state-of-the-art models, including GPT-4 and Claude variants, Llama 3.1, and o1-mini, reveals that even the most advanced LLMs struggle with episodic memory tasks, particularly when dealing with multiple related events or complex spatio-temporal relationships -- even in contexts as short as 10k-100k tokens."},"_bibtex":{"value":"@inproceedings{\nhuet2025episodic,\ntitle={Episodic Memories Generation and Evaluation Benchmark for Large Language Models},\nauthor={Alexis Huet and Zied Ben Houidi and Dario Rossi},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=6ycX677p2l}\n}"},"title":{"value":"Episodic Memories Generation and Evaluation Benchmark for Large Language Models"},"pdf":{"value":"/pdf/7a953bcf650b0ab0945dffa8bcfceac7ca25453f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"huet|episodic_memories_generation_and_evaluation_benchmark_for_large_language_models"},"authorids":{"value":["~Alexis_Huet1","~Zied_Ben_Houidi1","~Dario_Rossi1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Alexis Huet","Zied Ben Houidi","Dario Rossi"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a Physics-Informed Decentralized Federated Learning (PIDFL) framework that incorporates domain-specific knowledge, represented by differential equations, into decentralized federated learning (DFL)."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"see weakness"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- This paper is generally well-written and easy to follow\n- Incorporating domain knowledge through physics-informed constraints is novel and it seems to be enhancing model accuracy.\n- This paper provides theoretical convergence proof for the DFLA algorithm.\n- Experimental results indicate that PIDFL consistently outperforms traditional DFL approaches."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- It seems that the effectiveness of physics-informed constraints is contingent on the accuracy of the domain knowledge.\n- The sensitivity of the regularization parameter is not fully discussed.\n- The convergence is at $O(\\mu^2C^2)$, what does it say? is squared convergence good? more explanations, insights and remarks on this would be beneficial to this paper\n- Many other FL limitations are not discussed, e.g. it seems that this method requires frequent peer-to-peer communication, partial participation is not considered?"}},"nonreaders":[],"tmdate":1731428584408,"tcdate":1730697845419,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6147/Reviewer_f2zS"],"signatures":["ICLR.cc/2025/Conference/Submission6147/Reviewer_f2zS"],"forum":"ZXFJeR9Xm6","number":2,"license":"CC BY 4.0","cdate":1730697845419,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6147/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428584408,"domain":"ICLR.cc/2025/Conference","replyto":"ZXFJeR9Xm6","id":"dtxVKSGcrh","forumContent":{"TLDR":{"value":"New Decentralized Federated Learning Framework by Integration of Physics-Informed Neural Network"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Federated Learning","Physics-Informed Neural Network","Domain Knowledge","Decentralized"]},"supplementary_material":{"value":"/attachment/41baa9ef77c2dbd154cb2dd39b9445606e5cadba.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The integration of domain knowledge into the learning process of artificial intelligence (AI) has received significant attention in the last few years. Most of the approaches proposed so far have focused on centralized machine learning scenarios, with less emphasis on how domain knowledge can be effectively integrated in decentralized settings. In this paper, we address this gap by evaluating the effectiveness of domain knowledge integration in distributed settings, specifically in the context of Decentralized Federated Learning (DFL). We propose the Physics-Informed DFL (PIDFL) architecture by integrating domain knowledge expressed as differential equations. We introduce a serverless data aggregation algorithm for PIDFL, prove its convergence, and discuss its computational complexity. We performed comprehensive experiments across various datasets and demonstrated that  PIDFL significantly reduces average loss across diverse applications. This highlights the potential of PIDFL and offers a promising avenue for improving decentralized learning through domain knowledge integration."},"_bibtex":{"value":"@misc{\nalfano2024physicsinformed,\ntitle={Physics-Informed Decentralized Federated Learning},\nauthor={Gianvincenzo Alfano and Sergio Greco and Domenico Mandaglio and Francesco Parisi and Reza Shahbazian and Irina Trubitsyna},\nyear={2024},\nurl={https://openreview.net/forum?id=ZXFJeR9Xm6}\n}"},"title":{"value":"Physics-Informed Decentralized Federated Learning"},"pdf":{"value":"/pdf/e2d1aca8e693f7e237f0357a3de0af82c1aa0134.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"alfano|physicsinformed_decentralized_federated_learning"},"authorids":{"value":["~Gianvincenzo_Alfano1","~Sergio_Greco1","~Domenico_Mandaglio1","~Francesco_Parisi2","~Reza_Shahbazian1","~Irina_Trubitsyna2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Gianvincenzo Alfano","Sergio Greco","Domenico Mandaglio","Francesco Parisi","Reza Shahbazian","Irina Trubitsyna"]}},"version":2},{"content":{"summary":{"value":"The paper introduces \"PDC-Net,\" aimed at discerning the impact of mutations on protein-protein interaction (PPI) binding affinity changes, denoted as ΔΔ*G*. PDC-Net employs probability density clouds (PDCs) to model the magnitude and dynamics of protein movements during binding. The technique utilizes aligned networks for distributing equilibrium state representations in molecular systems. Two physics-inspired pretraining tasks are proposed, leveraging molecular dynamics simulations and a repository of static protein structures. Noteworthy contributions include:\n\n-   The novel representation of protein structures through PDCs, capturing equilibrium fluctuations and conformational changes at interfaces.\n-   Physics-derived pretraining tasks improving predictions of mutation effects on binding affinity.\n-   Comparative performance to empirical energy functions and other deep learning methods in predicting ΔΔ*G* upon mutations."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"**Originality**:  Conceptualizing atom positions in proteins as probability distributions rather than fixed coordinates is an intuitive idea. While the high-level PDC representation may be intuitive, the specific technical realization shows creativity. Propagating these distributions through the network is also novel and tailored to the PDC formulation.\n\n**Quality**: The proposed methods seem technically sound. The mathematical formulations for computing distribution parameters of geometric features are clearly derived (eq.2 can be improved). The network architecture modifications align well with the goal of propagating distributions. The experiments also appear rigorous.\n\n**Clarity**: The paper is in general well-written and easy to follow. \n\n**Significance**: The PDC representation and physics-inspired pretraining provide useful modeling advances. The superior performance demonstrates the significance of better capturing protein dynamics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Limited evaluations**: The evaluation is currently limited to only binding affinity prediction on a single dataset. Testing the methods on additional molecular modeling tasks and datasets would be beneficial to demonstrate broader utility. \n\n**Significance**: The gains over prior methods, while significant, are still arguably incremental.Discussing current limitations and potential future directions would enrich the discourse. Examining the model's performance in extreme cases (e.g., |ΔΔ*G*| > 5) could probably yield insights.\n\n**Clarity**: The intuition behind the physics-inspired pretraining (PIP) task could be explained more clearly. Some additional details on how optimizing the vectorized fluctuation target enables capturing thermodynamics would be helpful. The ablation studies could be elaborated on more to provide intuition about the contribution of each component. In particular, explaining the effect of the MRM in \"Per-Structure\" and \"Overall\" settings would make the gains more understandable. In general, some parts like the mathematical formulations are very technical and could use more plain language descriptions to make the concepts more accessible. The categorization of methods for comparison could be clarified further. Simply labeling methods like ESM-IF and MIF as \"unsupervised/semi-supervised\" is ambiguous, since all techniques are supervised trained on ∆∆G labels for the end task.\n\n**Spelling**: Minor typographical errors need correction, such as \"decpicts\" in the caption of Figure 2."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"**Extensions of PDC-Net**: Have you considered evaluating PDC-Net on additional molecular modeling tasks beyond binding affinity prediction? Testing generalizability to other problems like protein structure prediction could better validate its utility."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636035108,"tcdate":1699137798090,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1090/Reviewer_SdQS"],"signatures":["ICLR.cc/2024/Conference/Submission1090/Reviewer_SdQS"],"forum":"dFQL7nwksh","number":4,"license":"CC BY 4.0","cdate":1699137798090,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1090/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636035108,"domain":"ICLR.cc/2024/Conference","replyto":"dFQL7nwksh","id":"PCqc24szp2","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Protein-protein interaction","Thermodynamics","Geometric Deep Learning"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Understanding the ramifications of mutations at a protein level can have significant implications in various domains such as drug development, disease pathways, and the broader field of genomics. Despite the promise of data-driven and deep learning (DL) strategies, existing algorithms still face a significant challenge in integrating the dynamic changes of biomolecules to accurately predict protein-protein interaction binding affinity changes following mutations ($\\Delta \\Delta G$). Within this study, we introduce an inventive approach aimed at capturing the equilibrium fluctuations and discerning induced conformational changes at the interface, which is particularly important for forecasting mutational effects on binding. This novel technique harnesses probability density clouds (PDC) to describe the magnitude and intensity of their movement during and after the binding process and puts forth aligned networks to propagate distributions of the equilibrium of molecular systems. To fully unleash the potential of PDC-Net, we further present two physics-inspired pretraining tasks to employ the molecular dynamics (MD) simulation trajectories and the extensive collection of static crystal protein structures. Experiments demonstrate that our approach surpasses the performance of both empirical energy functions and alternative DL methods."},"_bibtex":{"value":"@misc{\nwu2024pdcnet,\ntitle={{PDC}-Net: Probability Density Cloud Representations of Proteins for Mutation Effect Prediction},\nauthor={Fang Wu and Shuting Jin and Shikun Feng and Le Song and Stan Z. Li},\nyear={2024},\nurl={https://openreview.net/forum?id=dFQL7nwksh}\n}"},"title":{"value":"PDC-Net: Probability Density Cloud Representations of Proteins for Mutation Effect Prediction"},"pdf":{"value":"/pdf/b477fe0a0e17483e5fd268a9493d1d0b845e4f2a.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|pdcnet_probability_density_cloud_representations_of_proteins_for_mutation_effect_prediction"},"authorids":{"value":["~Fang_Wu1","~Shuting_Jin1","~Shikun_Feng3","~Le_Song1","~Stan_Z._Li2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fang Wu","Shuting Jin","Shikun Feng","Le Song","Stan Z. Li"]}},"version":2},{"content":{"summary":{"value":"This paper investigates the performance of Graph Neural Networks (GNNs) in partially asynchronous inference settings, where nodes update in a staggered or asynchronous manner. The authors categorize GNNs into two types: \"explicitly-defined\" and \"implicitly-defined.\" They demonstrate that explicitly-defined GNNs are highly vulnerable to asynchronous updates. In contrast, implicitly-defined GNNs are shown to be robust. Additionally, the authors propose a novel explicitly-defined model, termed Energy GNN, which achieves notable improvements over existing methods on synthetic datasets."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- Could you provide a more intuitive explanation of implicitly defined GNNs, specifically highlighting the mechanisms contributing to their robustness under asynchronous updates?\n- Do you have any insight into why the performance on real-world datasets is not as great as on synthetic data?\n- Can you offer an explanation or hypothesis as to why explicitly defined GNNs did not perform as poorly as expected on the PPI dataset under asynchronous conditions?\n\nLine 369. There appears to be a typo in the index notation within the description of targets for the Sums dataset."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- Although I am not deeply familiar with the literature on asynchronous inference in GNNs, the proposed Energy GNN model introduces a novel and meaningful contribution to this field.\n- The paper is of high quality. The authors provide comprehensive mathematical proofs for the convergence of Energy GNNs. They also provide a detailed description of the experimental setup and results, which are well-organized and easy to follow.\n- The paper is generally well-structured, with a clear definition of the problem space and a concise summary of related GNN methods. The paper clearly defines the problem of asynchronous inference in explicitly-defined GNNs.\n- The proposed Energy GNN model demonstrates significant improvements on synthetic datasets, and competitive performance on real-world datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper lacks a clear and intuitive explanation of implicitly-defined GNNs, which is essential for understanding their robustness to asynchronous updates. While the authors offer detailed explanations for explicitly-defined GNNs, which are more straightforward, they do not provide the same depth of insight into implicitly-defined GNNs. This makes it difficult for readers unfamiliar with the topic to understand how implicitly-defined GNNs work and why they are resilient to asynchronous inference.\n\nAdditionally, although the proposed Energy GNN shows strong results on synthetic datasets, its performance on real-world datasets is rather competitive. On the PPI dataset, in particular, its performance is comparable to that of explicitly-defined GNNs, which were expected to fail under asynchrony. This discrepancy between synthetic and real dataset performance is not explained. A broader evaluation across various real-world datasets would increase the credibility of Energy GNNs as a robust solution."}},"nonreaders":[],"tmdate":1731428722196,"tcdate":1730706276913,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12109/Reviewer_GQeF"],"signatures":["ICLR.cc/2025/Conference/Submission12109/Reviewer_GQeF"],"forum":"WfxPVtYRlL","number":4,"license":"CC BY 4.0","cdate":1730706276913,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12109/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428722196,"domain":"ICLR.cc/2025/Conference","replyto":"WfxPVtYRlL","id":"ke2GVxOAeA","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["graph neural networks","multi-agent","asynchronous","decentralized"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Graph neural networks (GNNs) appear to be powerful tools to learn state representations for agents in distributed, decentralized multi-agent systems, but generate catastrophically incorrect predictions when nodes update asynchronously during inference.\n  This failure under asynchrony effectively excludes these architectures from many potential applications where synchrony is difficult or impossible to enforce, e.g., robotic swarms or sensor networks.\n  In this work we identify ''implicitly-defined'' GNNs as a class of architectures which is provably robust to asynchronous ''hogwild'' inference, adapting convergence guarantees from work in asynchronous and distributed optimization. \n  We then propose a novel implicitly-defined GNN architecture, which we call an energy GNN. \n  We show that this architecture outperforms other GNNs from this class on a variety of synthetic tasks inspired by multi-agent systems."},"_bibtex":{"value":"@inproceedings{\nsolodova2025graph,\ntitle={Graph Neural Networks Gone Hogwild},\nauthor={Olga Solodova and Nick Richardson and Deniz Oktay and Ryan P Adams},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=WfxPVtYRlL}\n}"},"title":{"value":"Graph Neural Networks Gone Hogwild"},"pdf":{"value":"/pdf/e8bd1d55ad9f030a6e0db669b9e72967f01edacb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"solodova|graph_neural_networks_gone_hogwild"},"authorids":{"value":["~Olga_Solodova1","~Nick_Richardson1","~Deniz_Oktay2","~Ryan_P_Adams1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Olga Solodova","Nick Richardson","Deniz Oktay","Ryan P Adams"]}},"version":2},{"content":{"venue":{"value":"Complex (1) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02466-5_23.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"zhou|personal_recommendation_in_userobject_networks"},"authorids":{"value":["~Tao_Zhou13"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02466-5_23"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/Zhou09,\n  author={Tao Zhou},\n  title={Personal Recommendation in User-Object Networks},\n  year={2009},\n  cdate={1230768000000},\n  pages={247-253},\n  url={https://doi.org/10.1007/978-3-642-02466-5_23},\n  booktitle={Complex (1)},\n  crossref={conf/complex/2009-1}\n}\n"},"abstract":{"value":"Thanks to the Internet and the World Wide Web, we live in a world of many possibilities we can choose from thousands of movies, millions of books, and billions of web pages. Far exceeding our personal processing capacity, this excessive freedom of choice calls for automated ways to find the relevant information. As a result, the field of information filtering is very active and rich with unanswered challenges. In this short paper, I will give a brief introduction on the design of recommender systems, which recommend objects to users based on the historical records of users’ activities. A diffusion-based recommendation algorithm, as well as two improved algorithms are investigated. Numerical results on a benchmark data set have demonstrated the advantages in algorithmic accuracy."},"title":{"value":"Personal Recommendation in User-Object Networks"},"authors":{"value":["Tao Zhou"]}},"tmdate":1767616579074,"pdate":1230768000000,"externalIds":["dblp:conf/complex/Zhou09"],"tcdate":1767616531956,"writers":["~"],"signatures":["~Tao_Zhou13"],"forum":"iCUTlli0UZ","license":"CC BY-SA 4.0","number":715692,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1767616579074,"domain":"DBLP.org","id":"iCUTlli0UZ","version":2},{"content":{"venue":{"value":"Complex (1) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02466-5_37.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"lai|mania_a_gene_network_reverse_algorithm_for_compounds_modeofaction_and_genes_interactions_inference"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Darong_Lai:","~Hongtao_Lu1","https://dblp.org/search/pid/api?q=author:Mario_Lauria:","https://dblp.org/search/pid/api?q=author:Diego_di_Bernardo:","https://dblp.org/search/pid/api?q=author:Christine_Nardini:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02466-5_37"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/LaiLLBN09,\n  author={Darong Lai and Hongtao Lu and Mario Lauria and Diego di Bernardo and Christine Nardini},\n  title={MANIA: A Gene Network Reverse Algorithm for Compounds Mode-of-Action and Genes Interactions Inference},\n  year={2009},\n  cdate={1230768000000},\n  pages={389-399},\n  url={https://doi.org/10.1007/978-3-642-02466-5_37},\n  booktitle={Complex (1)},\n  crossref={conf/complex/2009-1}\n}\n"},"abstract":{"value":"Understanding the complexity of the cellular machinery represents a grand challenge in molecular biology. To contribute to the deconvolution of this complexity, a novel inference algorithm based on linear ordinary differential equations is proposed, based on high-throughput gene expression data. The algorithm can infer (i) gene-gene interactions from steady state expression profiles AND (ii) mode-of-action of the components that can trigger changes in the system. Results demonstrate that the proposed algorithm can identify both information with high performances, thus overcoming the limitation of current algorithms that can infer reliably only one."},"title":{"value":"MANIA: A Gene Network Reverse Algorithm for Compounds Mode-of-Action and Genes Interactions Inference"},"authors":{"value":["Darong Lai","Hongtao Lu","Mario Lauria","Diego di Bernardo","Christine Nardini"]}},"tmdate":1749785911888,"pdate":1230768000000,"tcdate":1749785820776,"writers":["~"],"signatures":["~Hongtao_Lu1"],"forum":"Rj0ggwAwzj","license":"CC BY-SA 4.0","number":562132,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1749785911888,"domain":"DBLP.org","id":"Rj0ggwAwzj","version":2},{"content":{"venue":{"value":"Complex (1) 2009"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-02466-5_5.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEX/2009"},"paperhash":{"value":"ching|optimal_service_capacities_in_a_competitive_multipleserver_queueing_environment"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Wai-Ki_Ching:","https://dblp.org/search/pid/api?q=author:Sin-Man_Choi:","~Min_Huang3"]},"html":{"value":"https://doi.org/10.1007/978-3-642-02466-5_5"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complex/ChingCH09,\n  author={Wai-Ki Ching and Sin-Man Choi and Min Huang},\n  title={Optimal Service Capacities in a Competitive Multiple-Server Queueing Environment},\n  year={2009},\n  cdate={1230768000000},\n  pages={66-77},\n  url={https://doi.org/10.1007/978-3-642-02466-5_5},\n  booktitle={Complex (1)},\n  crossref={conf/complex/2009-1}\n}\n"},"abstract":{"value":"The study of economic behavior of service providers in a competition environment is an important and interesting research issue. A two-server queueing model has been proposed in Kalai et al. [11] for this purpose. Their model aims at studying the role and impact of service capacity in capturing larger market share so as to maximize the long-run expected profit. They formulate the problem as a two-person strategic game and analyze the equilibrium solutions. The main aim of this paper is to extend the results of the two-server queueing model in [11] to the case of multiple servers. We will only focus on the case when the queueing system is stable."},"title":{"value":"Optimal Service Capacities in a Competitive Multiple-Server Queueing Environment"},"authors":{"value":["Wai-Ki Ching","Sin-Man Choi","Min Huang"]}},"tmdate":1741229434556,"pdate":1230768000000,"tcdate":1741229261063,"writers":["~"],"signatures":["~Min_Huang3"],"forum":"a1i22JNETQ","license":"CC BY-SA 4.0","number":356558,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741229434556,"domain":"DBLP.org","id":"a1i22JNETQ","version":2},{"content":{"venue":{"value":"Eng. Appl. Artif. Intell. 2024"},"venueid":{"value":"dblp.org/journals/EAAI/2024"},"paperhash":{"value":"costabal|pinns_physicsinformed_neural_networks_on_complex_geometries"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Francisco_Sahli_Costabal:","~Simone_Pezzuto1","~Paris_Perdikaris1"]},"html":{"value":"https://doi.org/10.1016/j.engappai.2023.107324"},"_bibtex":{"value":"@article{DBLP:journals/eaai/CostabalPP24,\n  author={Francisco Sahli Costabal and Simone Pezzuto and Paris Perdikaris},\n  title={Δ-PINNs: Physics-informed neural networks on complex geometries},\n  year={2024},\n  month={January},\n  cdate={1704067200000},\n  journal={Eng. Appl. Artif. Intell.},\n  volume={127},\n  number={Part B},\n  pages={107324},\n  url={https://doi.org/10.1016/j.engappai.2023.107324}\n}\n"},"abstract":{"value":"Physics-informed neural networks (PINNs) have demonstrated promise in solving forward and inverse problems involving partial differential equations. Despite recent progress on expanding the class of problems that can be tackled by PINNs, most of existing use-cases involve simple geometric domains. To date, there is no clear way to inform PINNs about the topology of the domain where the problem is being solved. In this work, we propose a novel positional encoding mechanism for PINNs based on the eigenfunctions of the Laplace–Beltrami operator. This technique allows to create an input space for the neural network that represents the geometry of a given object. We approximate the eigenfunctions as well as the operators involved in the partial differential equations with finite elements. We extensively test and compare the proposed methodology against different types of PINNs in complex shapes, such as a coil, a heat sink and the Stanford bunny, with different physics, such as the Eikonal equation and heat transfer. We also study the sensitivity of our method to the number of eigenfunctions used, as well as the discretization used for the eigenfunctions and the underlying operators. Our results show excellent agreement with the ground truth data in cases where traditional PINNs fail to produce a meaningful solution. We envision this new technique will expand the effectiveness of PINNs to more realistic applications. Code available at: https://github.com/fsahli/Delta-PINNs."},"title":{"value":"Δ-PINNs: Physics-informed neural networks on complex geometries"},"authors":{"value":["Francisco Sahli Costabal","Simone Pezzuto","Paris Perdikaris"]}},"tmdate":1756313144698,"pdate":1704067200000,"tcdate":1727708537746,"writers":["~"],"signatures":["~Paris_Perdikaris1"],"forum":"h5HBduRHKs","license":"CC BY-SA 4.0","number":119606,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1756313144698,"domain":"DBLP.org","id":"h5HBduRHKs","version":2},{"content":{"venue":{"value":"IEEE Trans. Neural Networks Learn. Syst. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/5962385/10517792/10255379.pdf"},"venueid":{"value":"dblp.org/journals/TNN/2024"},"paperhash":{"value":"kapoor|physicsinformed_neural_networks_for_solving_forward_and_inverse_problems_in_complex_beam_systems"},"authorids":{"value":["~Taniya_Kapoor1","https://dblp.org/search/pid/api?q=author:Hongrui_Wang_0001:","https://dblp.org/search/pid/api?q=author:Alfredo_Núñez:","https://dblp.org/search/pid/api?q=author:Rolf_P._B._J._Dollevoet:"]},"html":{"value":"https://doi.org/10.1109/TNNLS.2023.3310585"},"_bibtex":{"value":"@article{DBLP:journals/tnn/KapoorWND24,\n  author={Taniya Kapoor and Hongrui Wang and Alfredo Núñez and Rolf P. B. J. Dollevoet},\n  title={Physics-Informed Neural Networks for Solving Forward and Inverse Problems in Complex Beam Systems},\n  year={2024},\n  month={May},\n  cdate={1714521600000},\n  journal={IEEE Trans. Neural Networks Learn. Syst.},\n  volume={35},\n  number={5},\n  pages={5981-5995},\n  url={https://doi.org/10.1109/TNNLS.2023.3310585}\n}\n"},"abstract":{"value":"This article proposes a new framework using physics-informed neural networks (PINNs) to simulate complex structural systems that consist of single and double beams based on Euler–Bernoulli and Timoshenko theories, where the double beams are connected with a Winkler foundation. In particular, forward and inverse problems for the Euler–Bernoulli and Timoshenko partial differential equations (PDEs) are solved using nondimensional equations with the physics-informed loss function. Higher order complex beam PDEs are efficiently solved for forward problems to compute the transverse displacements and cross-sectional rotations with less than $1e-3$ % error. Furthermore, inverse problems are robustly solved to determine the unknown dimensionless model parameters and applied force in the entire space–time domain, even in the case of noisy data. The results suggest that PINNs are a promising strategy for solving problems in engineering structures and machines involving beam systems."},"title":{"value":"Physics-Informed Neural Networks for Solving Forward and Inverse Problems in Complex Beam Systems"},"authors":{"value":["Taniya Kapoor","Hongrui Wang","Alfredo Núñez","Rolf P. B. J. Dollevoet"]}},"tmdate":1736500797129,"pdate":1704067200000,"tcdate":1736500789482,"writers":["~"],"signatures":["~Taniya_Kapoor1"],"forum":"7bKfuFvSDw","license":"CC BY-SA 4.0","number":261147,"cdate":1714521600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1736500797129,"domain":"DBLP.org","id":"7bKfuFvSDw","version":2},{"content":{"summary":{"value":"***Late review: Apologies my review is in one day late. I will aim to engage as quickly as needed to makeup.***\n\nThis very interesting paper considers a very important application, although using some well-established tools in machine learning and computer vision. The paper focuses on the difficult weak-lensing regime, where deflections are too weak to be interpretable reliably for a single galaxy. The idea is that attempting to understand how weak lensing alters (or shears) images of galaxies in dense fields, one could reconstruct a density map of the underlying matter. This is a difficult problem as the matter density is usually analyzed in 2D: the measured shear map is used to recover a projected 2D estimate. The paper aims to solve the inverse problem i.e. to obtain a 3D reconstruction of the dark matter field from 2D images. From a modeling perspective, this becomes challenging as galaxies are observed from a single view. Further, unlensed shapes of known visible sources are themselves not fully understood, introducing a large amount of uncertainty. The main idea of the paper is to use the underlying physics of gravitational lensing to recover a continuous 3D field, with a focus on also capturing non-Gaussian features, which are not easy with traditional methods, which use strong priors. The idea is to use coordinate-based neural fields augmented with the underlying physics (clearly described in figures 1 and 2), with the total loss function described in equation 7. The experiments, over simulated data, are able to show that the method has promise in both reconstructing the 3D matter field and also non-Gaussian features."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"-- Could the authors elaborate on the training protocols and how difficult it was to get stabilized training? From the writeup, it is not clear how easy it is to get the model to work."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"-- The paper attacks a problem of quite some importance. While the machine learning used is standard, its combination with the underlying physics allows the authors to obtain a promising solution, although it is only tested on simulations (standard for this area of physics at the moment). \n\n-- The paper is well-written, well-structured and easy to follow. While I am not an expert in this specific application, I have worked in some adjacent applications. I think the results support the underlying proposal and motivation well, and are creative in combining standard ML tools with physics-based models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-- More of a question: Could the authors elaborate in the paper why are strong Gaussian priors usually made? This will not be obvious to a standard ML audience. I see that it could make sense from the perspective of filtering. But I am not sure I understand why this should drop out from the physics models. I would expect the features to be highly non-Gaussian. I find this confusing whenever I venture to look at papers in this area."}},"nonreaders":[],"tmdate":1731469524461,"tcdate":1731469457487,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7909/Reviewer_p3hC"],"signatures":["ICLR.cc/2025/Conference/Submission7909/Reviewer_p3hC"],"forum":"Ax0i933gtp","number":4,"license":"CC BY 4.0","cdate":1731469457487,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7909/-/Official_Review"],"mdate":1731469524461,"domain":"ICLR.cc/2025/Conference","replyto":"Ax0i933gtp","id":"e0trP7PN38","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["computational imaging","signal processing","inverse problems","astrophysics","cosmology","neural fields","machine learning for physical sciences"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Weak gravitational lensing is the slight distortion of galaxy shapes caused primarily by the gravitational effects of dark matter in the universe. In our work, we seek to invert the weak lensing signal from 2D telescope images to reconstruct a 3D map of the universe's dark matter field. While inversion typically yields a 2D projection of the dark matter field, accurate 3D maps of the dark matter distribution are essential for localizing structures of interest and testing theories of our universe. However, 3D inversion poses significant challenges. First, unlike standard 3D reconstruction that relies on multiple viewpoints, in this case, images are only observed from a single viewpoint. This challenge can be partially addressed by observing how galaxy emitters throughout the volume are lensed. However, this leads to the second challenge: the shapes and exact locations of unlensed galaxies are unknown, and can only be estimated with a very large degree of uncertainty. This introduces an overwhelming amount of noise which nearly drowns out the lensing signal completely. Previous approaches tackle this by imposing strong assumptions about the structures in the volume. We instead propose a methodology using a gravitationally-constrained neural field to flexibly model the continuous matter distribution. We take an analysis-by-synthesis approach, optimizing the weights of the neural network through a fully differentiable physical forward model to reproduce the lensing signal present in image measurements. We showcase our method on simulations, including realistic simulated measurements of dark matter distributions that mimic data from upcoming telescope surveys. Our results show that our method can not only outperform previous methods, but importantly is also able to recover potentially surprising dark matter structures."},"_bibtex":{"value":"@inproceedings{\nzhao2025revealing,\ntitle={Revealing the 3D Cosmic Web through Gravitationally Constrained Neural Fields},\nauthor={Brandon Zhao and Aviad Levis and Liam Connor and Pratul P. Srinivasan and Katherine Bouman},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Ax0i933gtp}\n}"},"title":{"value":"Revealing the 3D Cosmic Web through Gravitationally Constrained Neural Fields"},"pdf":{"value":"/pdf/14dff818c86b7716f40ef9c613ca69f02c15287e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|revealing_the_3d_cosmic_web_through_gravitationally_constrained_neural_fields"},"authorids":{"value":["~Brandon_Zhao1","~Aviad_Levis1","~Liam_Connor1","~Pratul_P._Srinivasan1","~Katherine_Bouman1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Brandon Zhao","Aviad Levis","Liam Connor","Pratul P. Srinivasan","Katherine Bouman"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Neural Gaussian Force Field (NGFF) to address the challenge of predicting physical dynamics from raw visual data. It learns an implicit force field to drive the temporal evolution of 3D Gaussian Splatting, thereby modeling 4D dynamics. Experimental results demonstrate that NGFF achieves state-of-the-art performance in the 4D generation of collision motion in both simulation and real-world cases."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"My questions are mainly based on the above weaknesses:\n1. Do the predictions from NGFF perform better compared to the MPM simulation results? Is it possible to train the model with real captured data?\n2. The predicted simulation in the paper is limited to the collision of rigid or soft body objects. Can this method be extended to more complex physical motions?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The idea of using an implicit force field, rather than an explicit physics engine, to govern 3D GS evolution is highly innovative. It effectively tackles the critical bottlenecks of large computational overhead and insufficient robustness inherent in tightly coupling 3D GS with physical simulation.\n2. The paper provides extensive quantitative and qualitative results across diverse and challenging evaluations, which demonstrate NGFF's effectiveness in both 4D dynamic prediction and the more challenging real-world simulation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Performance: Since the model is primarily trained on data generated by the Material Point Method (MPM) simulation, its learned dynamics are inherently limited by the accuracy, fidelity, and approximations of the underlying MPM solver. This raises a concern that the performance of the trained NGFF model is limited to the quality of the synthetic ground truth.\n2. Generalization to Unseen Physics: While NGFF performs excellently on the provided datasets, its ability to generalize to unseen material properties (e.g., how a model trained only on rigid and soft bodies handles sand or fluids) or unseen complex constraints (e.g., complex joints or hinges) remains unclear. This is a crucial factor for real-world applicability.\n3. Missing Implementation Details: Modeling and training details are not clear enough, making this paper hard to follow. For example, what attributes of Gaussian kernels need to be supervised, and how is L' calculated in Eq.1? What is $s$ in Eq.6, and how to convert it to the state of Gaussian kernels?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928346225,"tcdate":1761740239899,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18635/Reviewer_sf96"],"signatures":["ICLR.cc/2026/Conference/Submission18635/Reviewer_sf96"],"forum":"KxvboPqav6","number":2,"license":"CC BY 4.0","cdate":1761740239899,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18635/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928346225,"domain":"ICLR.cc/2026/Conference","replyto":"KxvboPqav6","id":"cyHRqoqAi8","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Physical reasoning","video prediction"]},"supplementary_material":{"value":"/attachment/90ac022886fc42c632a139045ebfc3349691daec.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Predicting physical dynamics from raw visual data remains a major challenge in AI. While recent video generation models have achieved impressive visual quality, they still cannot consistently generate physically plausible videos due to a lack of modeling of physical laws. Recent approaches combining 3D Gaussian splatting and physics engines can produce physically plausible videos, but are hindered by high computational costs in both reconstruction and simulation, and often lack robustness in complex real-world scenarios. To address these issues, we introduce **Neural Gaussian Force Field (NGFF)**, an end-to-end neural framework that integrates 3D Gaussian perception with physics-based dynamic modeling to generate interactive, physically realistic 4D videos from multi-view RGB inputs, achieving two orders of magnitude faster than prior Gaussian simulators. To support training, we also present **GSCollision**, a 4D Gaussian dataset featuring diverse materials, multi-object interactions, and complex scenes, totaling over 640k rendered physical videos (∼4 TB). Evaluations on synthetic and real 3D scenarios show NGFF’s strong generalization and robustness in physical reasoning, advancing video prediction towards physics-grounded world models."},"_bibtex":{"value":"@inproceedings{\nli2026learning,\ntitle={Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields},\nauthor={Shiqian Li and Ruihong Shen and Junfeng Ni and Chang Pan and Chi Zhang and Yixin Zhu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KxvboPqav6}\n}"},"title":{"value":"Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields"},"pdf":{"value":"/pdf/d24aad8cb8c4f83bd8e677b5af43fdd97540be37.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|learning_physicsgrounded_4d_dynamics_with_neural_gaussian_force_fields"},"authorids":{"value":["~Shiqian_Li1","~Ruihong_Shen1","~Junfeng_Ni1","~Chang_Pan2","~Chi_Zhang12","~Yixin_Zhu1"]},"authors":{"value":["Shiqian Li","Ruihong Shen","Junfeng Ni","Chang Pan","Chi Zhang","Yixin Zhu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a novel method to incorporate information from a physics-based force field to improve sampling quality of a structural antibody diffusion model."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The authors focus on improvement to the binding affinity of generated antibodies, but the guiding force field is optimising for stability and overall energy rather than purely binding. Could the authors comment on how much the total energy of the protein/antibody improves when sampling with their guiding strategy, and whether this has other desirable properties, such as reducing the need for post-processing/relaxation?\n\nIn Algorithm 1, two hyperparameters are introduced that control the magnitude of the force applied during sampling. In the following section 4.2 and 4.3, a specific value is taken for both of these parameters, without a discussion of the tradeoff or impact of taking higher or lower coefficients. Could a discussion be added to expand on how table 1 would change for different parameter choices? There is a brief discussion in Appendix G, but it would be useful to provide a more qualitative understanding.\n\nIn Figure 4, it seems that the choice of when to apply the force field (at 70% in their experiments) plays an important role in changes to the output quality. Would starting the force field earlier in the sampling, or with a higher coefficient, substantially modify the results shown in this figure?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The idea of incorporating information of a physics-based force field into ML-based sampling seems very promising, and could lead to substantial improvements in practical applicability of these methods which often generate clashes or physically impossible structures."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The benchmarking is relatively limited, comparing only to DiffAb, a very similar model without the guided sampling, and RAbD, a physics-based method. It would be useful to include comparisons with other recent ML work, such as dyMEAN, HERN, RFdiffusion or IgDiff. In table 1, it would be helpful to include standard deviation for each metric, to understand the statistical significance of these results, especially as numbers are given to 2 or 3 decimal points.\nThough the incorporation of force field is very interesting, it is not completely novel and the experiments shown in this article are not completely convincing that this leads to significant and robust improvements in output quality."}},"nonreaders":[],"tmdate":1731427981668,"tcdate":1730674508822,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3919/Reviewer_UNs3"],"signatures":["ICLR.cc/2025/Conference/Submission3919/Reviewer_UNs3"],"forum":"9120xQKmcN","number":3,"license":"CC BY 4.0","cdate":1730674508822,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3919/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427981668,"domain":"ICLR.cc/2025/Conference","replyto":"9120xQKmcN","id":"Hs1Zeuy453","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion models","antibody design"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Antibodies, crucial for immune defense, primarily rely on complementarity-determining regions (CDRs) to bind and neutralize antigens, such as viruses. The design of these CDRs determines the antibody's affinity and specificity towards its target. Generative models, particularly denoising diffusion probabilistic models (DDPMs), have shown potential to advance the structure-based design of CDR regions. However, only a limited dataset of bound antibody-antigen structures is available, and generalization to out-of-distribution interfaces remains a challenge. Physics based force-fields, which approximate atomic interactions, offer a coarse but universal source of information to better mold designs to target interfaces. Integrating this foundational information into diffusion models is, therefore, highly desirable. Here, we propose a novel approach to enhance the sampling process of diffusion models by integrating force field energy-based feedback. Our model, DiffForce, employs forces to guide the diffusion sampling process, effectively blending the two distributions. Through extensive experiments, we demonstrate that our method guides the model to sample CDRs with lower energy, enhancing both the structure and sequence of the generated antibodies."},"_bibtex":{"value":"@misc{\nkulyt{\\.{e}}2024improving,\ntitle={Improving Antibody Design with Force-Guided Sampling in Diffusion Models},\nauthor={Paulina Kulyt{\\.{e}} and Francisco Vargas and Simon V Mathis and Yu Guang Wang and Jos{\\'e} Miguel Hern{\\'a}ndez-Lobato and Pietro Lio},\nyear={2024},\nurl={https://openreview.net/forum?id=9120xQKmcN}\n}"},"title":{"value":"Improving Antibody Design with Force-Guided Sampling in Diffusion Models"},"pdf":{"value":"/pdf/124980a44db47e9687c36abfdc153a6907cde169.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"kulyt|improving_antibody_design_with_forceguided_sampling_in_diffusion_models"},"authorids":{"value":["~Paulina_Kulytė1","~Francisco_Vargas1","~Simon_V_Mathis1","~Yu_Guang_Wang1","~José_Miguel_Hernández-Lobato1","~Pietro_Lio1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Paulina Kulytė","Francisco Vargas","Simon V Mathis","Yu Guang Wang","José Miguel Hernández-Lobato","Pietro Lio"]}},"version":2},{"content":{"summary":{"value":"The paper proposes DA-Transformer, a physics-informed spatiotemporal forecasting model for air quality (PM2.5) that embeds diffusion and advection operators directly inside Transformer encoder layers. In each encoder layer, Diffusion is implemented through a discrete Laplacian with a temperature-dependent coefficient, while advection module shifts the features along the direction and strength of local wind fields to simulate how pollutants are carried through the air.  The outputs of these two modules are then combined with learnable weights and added back into the Transformer’s residual connections. The model adopts learnable positional encodings, global spatiotemporal self-attention over flattened station–time tokens, and a simple non-autoregressive prediction head that maps the encoded window to multi-horizon outputs. Experiments on Beijing, UK (AURN), and a newly curated California dataset report consistent MAE/RMSE improvements over several baselines. Moreover, ablation studies suggest modest but positive contributions from the physics operators and positional encodings ."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Is the final model encoder-only with a non-autoregressive head, or encoder–decoder? Please reconcile Sections 3.1 vs. Implementation Details, and state which variant produced Table 1/Table 2. \n\n2. How robust is the station ordering by proximity under varying wind directions? Have you tried graph-based spatial derivatives (e.g., learned graph with edge directions aligned to wind)?\n\n3. Have you conducted any efficiency analysis comparing DA-Transformer against standard encoder–decoder baselines?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Integrating learnable diffusion and advection operators directly within Transformer encoder layers offers a distinct alternative to prior graph-based PDE solvers and neural ODE formulations. It avoids reliance on external solvers or predefined spatial graphs, and maintains the Transformer’s scalability and parallelism.\n\n2. Introduces a coherent architecture combining the physics block, global attention, learnable positional encodings, and a non-autoregressive prediction head. \n\n3. Experimental results across three geographical regions demonstrate consistent improvements over a diverse set of baselines, and ablation studies indicate that both the physics modules and positional encodings contribute to model performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Decoder Inconsistency** :  The paper claims a non-autoregressive head that “removes the decoder stack,” (Section 3.1), yet Implementation Details state the model “consists of 3 encoder layers and 2 decoder layers.” This causes a discrepancy which should be resolved.\n\n2. **Issues with Diffusion Formulation:**\n- The diffusion operator (Eq. 5) applies the discrete Laplacian simultaneously along the station index and the feature-channel index.\nWhile diffusion over the spatial axis is physically consistent with the PDE’s $D\\nabla^{2}u$ term, diffusing along the feature dimension has no physical analogue in pollutant transport. Feature channels in $H^{(\\ell)}$ represent heterogeneous or latent variables, not spatial coordinates. Therefore, smoothing across them mixes unrelated semantics. This turns the physics-motivated Laplacian into a numerical regularizer rather than a faithful discretization of spatial diffusion, undermining the claim of physical interpretability.\n\n- The implementation assumes periodic boundaries along both station and feature axes, effectively connecting the last station back to the first in a cyclic topology. Such a condition is unrealistic for irregularly distributed monitoring networks, where geographic continuity does not prevail. In practice, pollutant dispersion at boundary stations should follow conditions like Dirichlet (fixed-value) or be modeled through a graph-based Laplacian reflecting true spatial adjacency. The periodic assumption can distort spatial relationships and introduce non-physical transport between distant or unconnected locations.\n\n3. **Issues with the Advection Formulation:** In Section 3.4.2, Advection uses directional derivatives formed by  differences along station order and differences along the feature axis. \n\n- The first derivative $(\\delta_{\\text{stat}} H)$, assumes that stations can be arranged on a 1-D chain based on nearest-neighbor distance. This simplification ignores the true 2-D (latitude–longitude) geometry of the monitoring network. It risks misaligning wind direction with station indices. Furthermore, Spatial directional dependence (e.g., north–south vs. east–west transport) cannot be captured in a single 1-D stencil.\n\n- The second finite-difference term, $(\\delta_{\\text{feat}} H)_{b,s,i,c}$ treats the feature dimension $c$ as if it were a spatial coordinate, coupling it with wind components. However, $c$ indexes observed features (e.g. PM2.5, temperature or hidden embeddings), not spatial directions. Transporting information along feature channels with real wind velocities is physically meaningless: the wind field acts in geographic space, not across neural feature dimensions.\n\n4. **Evaluation Weaknesses:**\n\n- Baselines considered in the study are insufficient. Performance should be compared against Deep Learning based air quality prediction methods (PM25GNN), other transformer based (Airformer) and physics guided (Air-DualODE) models . \n- It's better to provide hyperparameter details of baseline models for reproducibility in the Appendix.\n- The data curation steps, splits, smoke episode handling, and code release details are insufficient in the California Dataset to be considered as a contribution.\n- Ablations are too narrow. They do not test alternative discretizations, boundary conditions, or station-ordering choices. Given the physical concerns above, these ablations are critical to validating the approach. Moreover, the study omits attention ablations, even though the architecture heavily relies on both global spatial-temporal self-attention and the physics module. \n\n5. Minor formatting weaknesses throughout the manuscript. (Ex: Some equations are not properly numbered) Please recheck and refine for clarity"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943482962,"tcdate":1761176515119,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission25577/Reviewer_4pVz"],"signatures":["ICLR.cc/2026/Conference/Submission25577/Reviewer_4pVz"],"forum":"PLO1gjCMk5","number":1,"license":"CC BY 4.0","cdate":1761176515119,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission25577/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943482962,"domain":"ICLR.cc/2026/Conference","replyto":"PLO1gjCMk5","id":"EhnFj4gdXl","forumContent":{"TLDR":{"value":"A physics-informed Transformer that learns temperature-conditioned diffusion and wind-driven advection to improve long-horizon PM2.5 forecasting across regions."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Spatiotemporal Forecasting","Air Quality","Physics-informed Learning","Transformers","Diffusion and Advection"]},"supplementary_material":{"value":"/attachment/127ec3d3ecc63d58a3d88e1354dd42e816827af4.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Air pollution is a major concern for public health and the environment globally, which highlights the need for effective monitoring and predictive modeling to mitigate its impact. Although data-driven models have shown promising results in air quality prediction, they still struggle to model the underlying physical mechanisms of pollutant dispersion, where diffusion governs small-scale spreading and advection drives large-scale directional transport. To address this limitation, we propose the Diffusion-Advection Transformer (DA-Transformer), a novel physics-informed architecture. Specifically, the model integrates the two key physical mechanisms by embedding diffusion and advection as differential equation-based components. These physics-informed modules are incorporated into a Transformer framework to enable the model to better capture pollutant transport dynamics, such as local diffusion-driven smoothing and wind-induced directional propagation in air quality data. Experiments on three real-world datasets demonstrate that DA-Transformer consistently outperforms baseline models in $\\mathrm{PM}_{2.5}$ concentration prediction and achieves substantial gains over its variants that exclude diffusion and advection in their model design."},"_bibtex":{"value":"@misc{\nanonymous2026diffusionadvection,\ntitle={Diffusion-Advection Transformer for Air Quality Prediction},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=PLO1gjCMk5}\n}"},"title":{"value":"Diffusion-Advection Transformer for Air Quality Prediction"},"pdf":{"value":"/pdf/e159a98d7df32c21638799cb787ba60fd05cb848.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"zhang|diffusionadvection_transformer_for_air_quality_prediction"},"authorids":{"value":["~Luyang_Zhang3","~Chunbo_Luo2","~Geyong_Min1"]},"authors":{"value":["Luyang Zhang","Chunbo Luo","Geyong Min"]}},"version":2},{"content":{"venue":{"value":"CogSci 2023"},"pdf":{"value":"https://escholarship.org/content/qt0rg3z8f6/qt0rg3z8f6.pdf?t=rxyaz2"},"venueid":{"value":"dblp.org/conf/COGSCI/2023"},"paperhash":{"value":"tatai|people_use_newtonian_physics_in_intuitive_sensorimotor_decisions_under_risk"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Fabian_Tatai:","~Dominik_Straub1","~Constantin_A._Rothkopf1"]},"html":{"value":"https://escholarship.org/uc/item/0rg3z8f6"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cogsci/TataiSR23,\n  author={Fabian Tatai and Dominik Straub and Constantin A. Rothkopf},\n  title={People use Newtonian physics in intuitive sensorimotor decisions under risk},\n  year={2023},\n  cdate={1672531200000},\n  url={https://escholarship.org/uc/item/0rg3z8f6},\n  booktitle={CogSci},\n  crossref={conf/cogsci/2023}\n}\n"},"abstract":{"value":"Author(s): Tatai, Fabian; Straub, Dominik; Rothkopf, Constantin | Abstract: Decisions under risk have been classically studied with tasks involving lotteries with explicit monetary rewards and uncertain gambles. More recently, sensorimotor decisions, specifically single movements to targets yielding rewards and losses, have been conceptualized as decisions under risk. While human choices between gambles have long been known not to maximize expected gains, sensorimotor decisions have been well described by statistical decision theory in many tasks. However, because many naturalistic scenarios of sensorimotor decisions are inescapably governed by the laws of physics, the question arises, how people act under such circumstances. Here, participants slid pucks to target areas, providing gains and losses in a virtual environment so that the uncertainty inherent in motor control interacts with the physical relationships governing objects' motion. Using model comparison with several generative models of participants' sliding actions, we find evidence that human motor decisions in scenarios with prospective economic outcomes take Newtonian physics into account."},"title":{"value":"People use Newtonian physics in intuitive sensorimotor decisions under risk"},"authors":{"value":["Fabian Tatai","Dominik Straub","Constantin A. Rothkopf"]}},"tmdate":1747225068793,"pdate":1672531200000,"tcdate":1741353512338,"writers":["~"],"signatures":["~Constantin_Rothkopf1"],"forum":"WD5Zvppv8m","license":"CC BY-SA 4.0","number":364881,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747225068793,"domain":"DBLP.org","id":"WD5Zvppv8m","version":2},{"content":{"summary":{"value":"This paper proposes a novel framework called PHYCO, which aims to learn implicit physical laws from GS’s monocular observation data. The method effectively addresses two major challenges existing in current implicit learning approaches: unstable geometric learning and lack of physical interpretability. Specifically, PHYCO introduces Edge-Aware Depth Consensus Anchors (EADCA) to stabilize geometric reconstruction and designs a Physics-Consistent Loss that integrates physical laws into the training of implicit functions. This enables robust, interpretable, and highly generalizable learning of complex physical processes."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"see above"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The authors demonstrate significant innovation and advantages in applying implicit learning to physical modeling:\n1.\tThe paper introduces EADCA to effectively tackle the issue of inaccurate or locally optimal geometric representations in implicit methods under noisy monocular supervision. This mechanism ensures high-quality geometric reconstruction, providing a solid foundation for subsequent physical parameter inversion.\n2.\tOne of the core contributions of this paper is the design of MHPV, which embeds known physical conservation laws (such as momentum conservation) as hard constraints into the training of implicit constitutive laws. This ensures that the learned constitutive functions are physically reasonable and reliable, greatly enhancing the model’s interpretability — something that purely data-driven methods can hardly achieve.\n3.\tThe PHYCO framework successfully combines the efficient rendering capability of GS with the expressive power of implicit functions, enabling the direct learning of complex, nonlinear, and non-elastic constitutive laws from monocular videos. This avoids dependence on traditional predefined explicit constitutive equations and significantly broadens the model’s generalization and modeling capability for various complex materials and physical phenomena."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"However, I do have several concerns about this work:\n1.\tThe core components of the framework rely on several pre-trained modules. Although the authors partially address lighting variations by changing illumination conditions in the dataset, if these modules perform poorly under certain conditions (for example, large-scale deformations), the overall performance of the framework could be greatly affected.\n2.\tDue to the high structural complexity and diverse optimization objectives, the framework may suffer from high debugging and training costs, leading to potential instability during optimization.\n3.\tIn MHPV, the authors select a set of classical constitutive models to validate material physical behaviors. However, if these selected material models do not adequately approximate or represent real material physics, the “physical rationality” constraints imposed by MHPV might become counterproductive rather than optimizing PHYCO. Essentially, it strongly assume that the chosen constitutive equations are trustworthy but lacks rigorous proof of their validity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923186703,"tcdate":1761879188707,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12246/Reviewer_TWS7"],"signatures":["ICLR.cc/2026/Conference/Submission12246/Reviewer_TWS7"],"forum":"XyHbp7Y2T5","number":3,"license":"CC BY 4.0","cdate":1761879188707,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12246/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923186703,"domain":"ICLR.cc/2026/Conference","replyto":"XyHbp7Y2T5","id":"uLILb4HJBN","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Gaussian splatting","physics-informed learning","implicit constitutive laws"]},"supplementary_material":{"value":"/attachment/831d2a297e1a2ec1a3e64d35f8a5d86acb52a56b.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"We present **PhyCo**, a framework for learning implicit constitutive laws from \\textbf{monocular dynamic observations} of Gaussian splatting. Existing implicit methods often suffer from local minima under noisy supervision and lack physical interpretability, while explicit approaches rely on predefined constitutive equations, limiting generalizability. To address these issues, our framework, **PhyCo**, introduces two key innovations. First, **initializing from a static multi-view scan, we propose *Edge-Aware Depth Consensus Anchors* to establish robust geometric constraints from subsequent monocular dynamic observations**, circumventing unreliable pixel-level supervision. Second, a *Multi-Hypothesis Physics Verifier* integrates classical constitutive models as differentiable hypotheses, providing strong physical priors to regularize the optimization while preserving the flexibility of implicit modeling. This unified approach ensures physical plausibility without sacrificing generality. Extensive experiments on synthetic, real-to-sim, and real-world datasets demonstrate that **PhyCo** significantly outperforms existing methods, achieving state-of-the-art performance in learning accurate and generalizable physical dynamics from monocular videos."},"_bibtex":{"value":"@misc{\nliu2026phyco,\ntitle={PhyCo: Physics-Consistent Learning of Implicit Constitutive Laws via Monocular Observations of 3D Gaussians},\nauthor={Xiaoyang Liu and Kai Han},\nyear={2026},\nurl={https://openreview.net/forum?id=XyHbp7Y2T5}\n}"},"title":{"value":"PhyCo: Physics-Consistent Learning of Implicit Constitutive Laws via Monocular Observations of 3D Gaussians"},"pdf":{"value":"/pdf/230c2a74b4fdefee465eb282bf5252b5dd6e30e3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"liu|phyco_physicsconsistent_learning_of_implicit_constitutive_laws_via_monocular_observations_of_3d_gaussians"},"authorids":{"value":["~Xiaoyang_Liu5","~Kai_Han1"]},"authors":{"value":["Xiaoyang Liu","Kai Han"]}},"version":2},{"content":{"summary":{"value":"1. **Feature Distribution Issue**: Models trained on synthetic data show discrete and clustered features, limiting generalization due to synthetic data distribution.\n\n2. **SynDR-IQA Framework**: Introduces DDCUp and DRCDown strategies to enhance diversity and balance density, reshaping data for better generalization.\n\n3. **Validation**: SynDR-IQA proves effective across cross-dataset scenarios and integrates with existing models without extra inference costs."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. **Lack of Comparisons with Advanced Methods**: : The paper does not include comparisons with state-of-the-art methods like Q-align[1], LIQE[2], CLIPIQA[3], and TOPIQ[4], which would provide a clearer benchmark of SynDR-IQA’s performance.\n\n2. **Complex Implementation**: The DDCUp and DRCDown strategies, while theoretically valuable, add complexity in implementation. Practical adoption may be challenging without detailed instructions or code, possibly limiting its use in broader BIQA applications.\n\n﻿3. **Synthetic Data Focus**: The framework mainly addresses synthetic-to-real transitions, which may limit its applicability and effectiveness on purely real-world distortions where data characteristics and distortion patterns differ substantially from synthetic ones.\n\n\n\n\n[1] Wu, Haoning, et al. \"Q-align: Teaching lmms for visual scoring via discrete text-defined levels.\" arXiv preprint arXiv:2312.17090 (2023).\n\n[2] Zhang, Weixia, et al. \"Blind image quality assessment via vision-language correspondence: A multitask learning perspective.\" Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023.\n\n[3] Wang, Jianyi, Kelvin CK Chan, and Chen Change Loy. \"Exploring clip for assessing the look and feel of images.\" Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 37. No. 2. 2023.\n\n[4] Chen, Chaofeng, et al. \"Topiq: A top-down approach from semantics to distortions for image quality assessment.\" IEEE Transactions on Image Processing (2024)."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. **Innovative Approach**: Proposes SynDR-IQA, a novel framework that reshapes synthetic data distribution to improve BIQA generalization.\n\n2. **Effective Strategies**: Introduces DDCUp and DRCDown, two strategies that enhance data diversity and reduce redundancy, addressing core issues in synthetic datasets.\n\n3. **Strong Experimental Validation**: Demonstrates effectiveness across multiple cross-dataset settings, showing that SynDR-IQA enhances performance without increasing inference costs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **lack of IQA result images**. Visual examples comparing predictions with ground truth would clarify SynDR-IQA's impact on quality assessment.\n2. **Synthetic Data Focus**: Primarily targets synthetic data limitations, which may limit applicability to real-world data.\n3. **Limited Analysis on Failure Cases**: The paper lacks discussion on SynDR-IQA's potential limitations or specific cases where it may underperform, such as certain distortion types or datasets. A balanced evaluation could clarify the framework's boundaries."}},"nonreaders":[],"tmdate":1732731768370,"tcdate":1730764647558,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13695/Reviewer_aBSv"],"signatures":["ICLR.cc/2025/Conference/Submission13695/Reviewer_aBSv"],"forum":"pwNIOcr8fU","number":4,"license":"CC BY 4.0","cdate":1730764647558,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13695/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732731768370,"domain":"ICLR.cc/2025/Conference","replyto":"pwNIOcr8fU","id":"EpJUAnCNuP","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Blind Image Quality Assessment; Data Distribution Reshaping; Synthetic Data"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Blind Image Quality Assessment (BIQA) has advanced significantly through deep learning, but the scarcity of large-scale labeled datasets remains a challenge. While synthetic data offers a promising solution, models trained on existing synthetic datasets often show limited generalization ability. In this work, we make a key observation that representations learned from synthetic datasets often exhibit a discrete and clustered pattern that hinders regression performance: features of high-quality images cluster around reference images, while those of low-quality images cluster based on distortion types. Our analysis reveals that this issue stems from the distribution of synthetic data rather than model architecture. Consequently, we introduce a novel framework SynDR-IQA, which reshapes synthetic data distribution to enhance BIQA generalization. Based on theoretical derivations of sample diversity and redundancy's impact on generalization error, SynDR-IQA employs two strategies: distribution-aware diverse content upsampling, which enhances visual diversity while preserving content distribution, and density-aware redundant cluster downsampling, which balances samples by reducing the density of densely clustered areas. Extensive experiments across three cross-dataset settings (synthetic-to-authentic, synthetic-to-algorithmic, and synthetic-to-synthetic) demonstrate the effectiveness of our method. Additionally, as a data-based approach, SynDR-IQA can be coupled with model-based methods without increasing inference costs. The source code will be publicly available."},"_bibtex":{"value":"@misc{\nli2025towards,\ntitle={Towards Syn-to-Real {IQA}: A Novel Perspective on Reshaping Synthetic Data Distributions},\nauthor={Aobo Li and Jinjian Wu and Yongxu Liu and Leida Li and Weisheng Dong},\nyear={2025},\nurl={https://openreview.net/forum?id=pwNIOcr8fU}\n}"},"title":{"value":"Towards Syn-to-Real IQA: A Novel Perspective on Reshaping Synthetic Data Distributions"},"pdf":{"value":"/pdf/eb697de053a898cb44da0f2b135be854402f559a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|towards_syntoreal_iqa_a_novel_perspective_on_reshaping_synthetic_data_distributions"},"authorids":{"value":["~Aobo_Li1","~Jinjian_Wu1","~Yongxu_Liu2","~Leida_Li3","~Weisheng_Dong1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Aobo Li","Jinjian Wu","Yongxu Liu","Leida Li","Weisheng Dong"]}},"version":2},{"content":{"venue":{"value":"ICONIP (2) 2015"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-26535-3_47.pdf"},"venueid":{"value":"dblp.org/conf/ICONIP/2015"},"paperhash":{"value":"popa|conjugate_gradient_algorithms_for_complexvalued_neural_networks"},"authorids":{"value":["~Călin-Adrian_Popa1"]},"html":{"value":"https://doi.org/10.1007/978-3-319-26535-3_47"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iconip/Popa15,\n  author={Calin-Adrian Popa},\n  title={Conjugate Gradient Algorithms for Complex-Valued Neural Networks},\n  year={2015},\n  cdate={1420070400000},\n  pages={412-422},\n  url={https://doi.org/10.1007/978-3-319-26535-3_47},\n  booktitle={ICONIP (2)},\n  crossref={conf/iconip/2015-2}\n}\n"},"abstract":{"value":"In this paper, conjugate gradient algorithms for complex-valued feedforward neural networks are proposed. Since these algorithms yielded better training results for the real-valued case, an extension to the complex-valued case is a natural option to enhance the performance of the complex backpropagation algorithm. The full deduction of the classical variants of the conjugate gradient algorithm is presented, and the resulting training methods are exemplified on synthetic and real-world applications. The experimental results show a significant improvement over the complex gradient descent algorithm."},"title":{"value":"Conjugate Gradient Algorithms for Complex-Valued Neural Networks"},"authors":{"value":["Calin-Adrian Popa"]}},"tmdate":1772378575149,"pdate":1451520000000,"externalIds":["dblp:conf/iconip/Popa15"],"tcdate":1772378543503,"writers":["~"],"signatures":["~Călin-Adrian_Popa1"],"forum":"hJ3G6ogOoB","license":"CC BY-SA 4.0","number":841182,"cdate":1420070400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772378575149,"domain":"DBLP.org","id":"hJ3G6ogOoB","version":2},{"content":{"TLDR":{"value":"Physics-constrained symbolic regression, Monte-Carlo tree search with symmetric invariant representations and graph neural networks."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Symbolic Regression","Physics-constrained","Graph Neural Network","Reinforcement Learning","Monte-Carlo Tree Search","Expression Tree","Automated Feature Engineering","Symbolic Graph"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"As data-driven scientific discovery increasingly demands explainable over ‘black-box’ machine learning (ML) methods, Symbolic Regression (SR) that derives analytical expressions can help identify key functional dependencies in complex systems. However, traditional SR methods often suffer from (a) inefficient exploration due to their inability to compress the search space of equivalent expressions, and (b) non-physical solutions that violate fundamental physics constraints. We here introduce a symmetric invariant representation of candidate analytical expressions using a Symbolic Graph (SG), on which the Symbolic Graph Neural Network (SGNN) encodes operators, symmetries,   constraints and constant fitting knowledge. We further develop reinforcement learning (RL) algorithms with Monte-Carlo Tree Search (MCTS) on our SGNN for SR. Such a physics-constrained graph symbolic regression (PCGSR) method effectively compresses the search space for efficient SR. Experiments on synthetic and real-world scientific datasets demonstrate the efficiency and accuracy of our PCGSR in discovering underlying expressions and adhering to physical laws, yielding physically meaningful solutions."},"_bibtex":{"value":"@misc{\nxiang2025physicsconstrained,\ntitle={Physics-constrained Graph Symbolic Regression},\nauthor={Ziyu Xiang and Kenna Ashen and Xiaofeng Qian and Xiaoning Qian},\nyear={2025},\nurl={https://openreview.net/forum?id=Ia17iAtr0P}\n}"},"title":{"value":"Physics-constrained Graph Symbolic Regression"},"pdf":{"value":"/pdf/f49a5930438a1b85984252300a705d7a4a0f259c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"xiang|physicsconstrained_graph_symbolic_regression"},"authorids":{"value":["~Ziyu_Xiang1","~Kenna_Ashen1","~Xiaofeng_Qian1","~Xiaoning_Qian2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ziyu Xiang","Kenna Ashen","Xiaofeng Qian","Xiaoning Qian"]}},"tmdate":1738735754520,"tcdate":1727372625886,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7682/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission7682/Authors"],"forum":"Ia17iAtr0P","license":"CC BY 4.0","number":7682,"cdate":1727372625886,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/Submission7682/-/Full_Submission","ICLR.cc/2025/Conference/Submission7682/-/Rebuttal_Revision","ICLR.cc/2025/Conference/-/Edit"],"mdate":1738735754520,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"Ia17iAtr0P","version":2},{"content":{"summary":{"value":"This paper introduces FOLIAGE, a geometry-centric latent world model designed to predict and control accretive surface growth — phenomena where surfaces add new material and evolve in morphology (e.g., leaves, tissues, 4D-printed materials).\n\nThe framework combines:\n\n- A multimodal perception stack that fuses images, point clouds, and meshes using correspondence-constrained fusion and vertex age embeddings to highlight regions where growth occurs.\n\n- An action-conditioned latent dynamics model that predicts geometric evolution based on material parameters (stretch, shear, bend coefficients) without directly simulating physics.\n\n- A physics-guided training branch using energy-gated message passing—privileged physical energy signals available only at training time to improve the representation’s physical awareness.\n\nThe system is trained and evaluated on a new dataset, SURF-GARDEN, and benchmark suite, SURF-BENCH, which include diverse growth simulations with ground-truth correspondences and counterfactual branching."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- Can the latent state learned in FOLIAGE transfer to unseen material classes or natural growth data without retraining?\n\n- How does the model behave if mesh topology changes drastically (e.g., splits or merges)?\n\n- Does the learned latent preserve interpretable geometric quantities such as strain or curvature distributions?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Novel problem formulation:\nAccretive surface growth is an underexplored domain bridging geometry, control, and materials. The paper convincingly motivates why traditional differentiable simulators and pixel-based world models fail in this regime.\n\n- Conceptually elegant design:\nFOLIAGE neatly decouples perception (multimodal geometry fusion) from dynamics (action-conditioned latent evolution), preserving counterfactual semantics while learning from physics signals only during training."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited real-world validation:\nEvaluation is purely synthetic. Although SURF-BENCH is rich, testing on real sensor data (e.g., growth of plants or soft materials) would solidify claims of generalization and practical relevance.\n\n- High system complexity:\nThe model combines multimodal encoders, correspondence graphs, hierarchical pooling, energy-gated message passing, and action-conditioned transformers—raising questions about interpretability and reproducibility despite code release promises.\n\n- Physical interpretability:\nWhile FOLIAGE learns from per-vertex energies, it doesn’t explicitly model or verify physical correctness (e.g., energy conservation, material law adherence). It functions as a learned surrogate rather than a physically grounded simulator."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925644689,"tcdate":1762799526367,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15361/Reviewer_tTmg"],"signatures":["ICLR.cc/2026/Conference/Submission15361/Reviewer_tTmg"],"forum":"eTXTOUrrhY","number":4,"license":"CC BY 4.0","cdate":1762799526367,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15361/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925644689,"domain":"ICLR.cc/2026/Conference","replyto":"eTXTOUrrhY","id":"BUxH5ZjDln","forumContent":{"TLDR":{"value":"We present FOLIAGE, a latent world model for unbounded accretive surface evolution"},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Surface Growth","World Model","Multimodal"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Accretive surfaces grow by adding material and changing rest metrics, producing emergent, complex, and changing morphologies. We introduce FOLIAGE, a geometry-centric latent world model that infers a deployable state from heterogeneous, partial sensors and predicts its action-conditioned evolution. The perception stack aligns images, point clouds, and meshes through correspondence-constrained fusion and age features, then pools into global and young-region summaries that emphasize where change will occur next. Dynamics input act only on the latent, taking material coefficients and a horizon code to produce counterfactual roll-outs without entangling perception with control. Training-time physics guides representation via a target encoder that receives per-vertex energies and energy-gated message passing, while the deployable path relies solely on observable inputs. On the SURF-GARDEN data platform and the SURF-BENCH suite, FOLIAGE improves mesh topology classification by ~3 pp, reduces dense-correspondence geodesic error by ~10\\%, lifts cross-modal retrieval by ~25\\% mAP@100, increases growth-stage recognition by ~8 pp, lowers 5-step Chamfer by ~20\\%, and cuts inverse-material error by ~40\\% relative to strong baselines. Stress tests show graceful degradation under sensor loss, stable long-horizon roll-outs, and gains from train-only physics without test-time privileges. Code and datasets used in this study will be made publicly available upon publication to facilitate reproducibility and further research."},"_bibtex":{"value":"@misc{\nanonymous2026foliage,\ntitle={{FOLIAGE}: a Latent World Model for Accretive Surface Growth},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=eTXTOUrrhY}\n}"},"title":{"value":"FOLIAGE: a Latent World Model for Accretive Surface Growth"},"pdf":{"value":"/pdf/10b1072240dbcdc709a9e9ac2c4a5deccfedb375.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"liu|foliage_a_latent_world_model_for_accretive_surface_growth"},"authorids":{"value":["~Xiaoyi_Liu7","~Hao_Tang6"]},"authors":{"value":["Xiaoyi Liu","Hao Tang"]}},"version":2},{"content":{"summary":{"value":"This paper provides the identifiability analysis of linear Ordinary Differential Equation (ODE) systems, particularly in scenarios where latent variables interact with the system. In detail, it investigates two specific cases. In the first scenario, latent confounders do not exhibit causal relationships, but their evolution follows specific functional forms, such as polynomial functions of time. The analysis is then extended to a second, more complex scenario, where hidden confounders have causal dependencies described by a Directed Acyclic Graph (DAG). The authors perform a series of simulations to substantiate their theoretical results."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"How practical are the proposed assumptions in the paper? Can the authors discuss their validity and design experiments to test their validity for real datasets?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"This paper makes a significant contribution by extending the understanding of identifiability in linear ODE systems to include cases with latent variables, thereby enhancing the reliability of causal inferences in more complex systems. The simulated experimental results provide strong support for the theoretical findings, making this a robust and valuable study in the field."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The symbols $x'$ and $A'$ in the paper are not clearly defined. This may lead to confusion for readers when understanding the derivation process and the results. It is recommended to clearly define and explain these symbols in the paper.\n2. The assumptions and theorems lack intuitive explanations. It would be better to provide some intuitive explanations or examples after each assumption and theorem to help readers better understand the essence of these theories and their roles in practical applications."},"limitations":{"value":"See the weaknesses above."}},"nonreaders":[],"tmdate":1730879207168,"tcdate":1720699871543,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission7973/Reviewer_LnTR"],"signatures":["NeurIPS.cc/2024/Conference/Submission7973/Reviewer_LnTR"],"forum":"8271eFxojN","number":1,"license":"CC BY 4.0","cdate":1720699871543,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission7973/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879207168,"domain":"NeurIPS.cc/2024/Conference","replyto":"8271eFxojN","id":"nmOajlusEi","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Linear ODEs","Identifiability analysis","Hidden confounders","Causality"]},"supplementary_material":{"value":"/attachment/2d0dfc4da566f56f537954f5b4c34c80fde6dcc7.zip"},"primary_area":{"value":"causal_inference"},"abstract":{"value":"The identifiability analysis of linear Ordinary Differential Equation (ODE) systems is a necessary prerequisite for making reliable causal inferences about these systems. While identifiability has been well studied in scenarios where the system is fully observable, the conditions for identifiability remain unexplored when latent variables interact with the system. This paper aims to address this gap by presenting a systematic analysis of identifiability in linear ODE systems incorporating hidden confounders. Specifically, we investigate two cases of such systems. In the first case, latent confounders exhibit no causal relationships, yet their evolution adheres to specific functional forms, such as polynomial functions of time $t$. Subsequently, we extend this analysis to encompass scenarios where hidden confounders exhibit causal dependencies, with the causal structure of latent variables described by a Directed Acyclic Graph (DAG). The second case represents a more intricate variation of the first case, prompting a more comprehensive identifiability analysis. Accordingly, we conduct detailed identifiability analyses of the second system under various observation conditions, including both continuous and discrete observations from single or multiple trajectories. To validate our theoretical results, we perform a series of simulations, which support and substantiate our findings."},"_bibtex":{"value":"@inproceedings{\nwang2024identifiability,\ntitle={Identifiability Analysis of Linear {ODE} Systems with Hidden Confounders},\nauthor={Yuanyuan Wang and Biwei Huang and Wei Huang and Xi Geng and Mingming Gong},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=8271eFxojN}\n}"},"title":{"value":"Identifiability Analysis of Linear ODE Systems with Hidden Confounders"},"pdf":{"value":"/pdf/3cbb99334696d36cb119a6f10d3a253978283f71.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"wang|identifiability_analysis_of_linear_ode_systems_with_hidden_confounders"},"authorids":{"value":["~Yuanyuan_Wang5","~Biwei_Huang1","~Wei_Huang8","~Xi_Geng1","~Mingming_Gong1"]},"authors":{"value":["Yuanyuan Wang","Biwei Huang","Wei Huang","Xi Geng","Mingming Gong"]}},"version":2},{"content":{"summary":{"value":"The paper considers the problem of supervised learning with a class imbalance and proposes a new algorithm for oversampling the minority class based on topological data analysis. The method, Simplicial SMOTE, forms a simplicial complex from data in the minority class, and then generates synthetic samples based on the convex combinations of points sampled from simplices of this simplicial complex. The authors additionally propose variants of other SMOTE algorithms based on simplicial complexes. The utility of the simplicial SMOTE algorithm is validated on synthetic and real datasets, and suggests the value of the approach for a wide variety of empirical applications."},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"soundness":{"value":"4 excellent"},"strengths":{"value":"The paper is extremely well presented and provides an original application of topological data analysis in a machine learning setting.\n\nThe quality of the work is generally high: The empirical results are presented over a wide set of synthetic and empirical datasets with class imbalances which provide a full picture of the proposed algorithm's value.\n\nThe proposed algorithm is presented in an extremely clear way and the figures help to highlight why this approach is different than prior methods. Overall, the presentation is excellent and the paper was enjoyable to read.\n\nThe work ultimately provides a valuable step in using topological data analysis (TDA) for machine learning: TDA methods are typically computationally intensive (as noted by the authors), and are quick to be dismissed in machine learning applications. However, this paper shows that TDA methods can still add value to empirical performance of learning algorithms and hence provides a foundation for a wide variety of future work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Some minor points:\n\n1) Although the paper is quite interesting from the lens of topological data analysis, it is presented as a simplicial extension of SMOTE and hence feels limited in terms of significance for a machine learning audience. \n\n2) The authors could be a bit more clear on why a simplicial complex may be better than a graph for creating synthetic points for oversampling-- it feels like there is some type of local decision-boundary type of argument which would make clear when this method should be valuable.\n\n3) There empirical results could be presented better. The tables are fine and are valuable because they give access to the raw data. However, they could benefit from the addition of confidence intervals. Alternatively, additional visualizations may more clearly summarize the value of the proposed algorithm for each dataset."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"N/A -- it may be helpful if the authors can address point number 2 above."},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636717043,"tcdate":1698575371607,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6425/Reviewer_zg3w"],"signatures":["ICLR.cc/2024/Conference/Submission6425/Reviewer_zg3w"],"forum":"zqXTZ3B7fU","number":1,"license":"CC BY 4.0","cdate":1698575371607,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6425/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636717043,"domain":"ICLR.cc/2024/Conference","replyto":"zqXTZ3B7fU","id":"Gj2UKsdn2E","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["data augmentation","oversampling","imbalanced learning problem"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"SMOTE is the established geometric approach to random oversampling to balance classes in the imbalanced classes learning problem, followed by many extensions. Its idea is to introduce synthetic data points of the minor class, with each new point being the convex combination of an existing data point and one of its k-nearest neighbors. This could be viewed as a sampling from the edges of a geometric neighborhood graph. Borrowing tools from the topological data analysis, we propose a generalization of the sampling approach, thus sampling from the simplices of the geometric neighborhood simplicial complex. That is, a new point is defined by the barycentric coordinates with respect to a simplex spanned by an arbitrary number of data points being sufficiently close, rather than a pair. We evaluate the generalized technique which we call Simplicial SMOTE on 23 benchmark datasets, and conclude that it outperforms the original SMOTE and its extensions. Moreover, we show how simplicial sampling can be integrated into several popular SMOTE extensions, with our simplicial generalization of Borderline SMOTE further improves the performance on benchmarks datasets."},"_bibtex":{"value":"@misc{\nkachan2024simplicial,\ntitle={Simplicial {SMOTE}: Oversampling Solution to the Imbalanced Learning Problem},\nauthor={Oleg Kachan and Andrey Savchenko and Gleb Gennadjevich Gusev},\nyear={2024},\nurl={https://openreview.net/forum?id=zqXTZ3B7fU}\n}"},"title":{"value":"Simplicial SMOTE: Oversampling Solution to the Imbalanced Learning Problem"},"pdf":{"value":"/pdf/e26beb03ed1e66ffa1b320a4fd052b5c5e7d2e2d.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"kachan|simplicial_smote_oversampling_solution_to_the_imbalanced_learning_problem"},"authorids":{"value":["~Oleg_Kachan1","~Andrey_Savchenko1","~Gleb_Gennadjevich_Gusev1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Oleg Kachan","Andrey Savchenko","Gleb Gennadjevich Gusev"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a bilevel optimization framework for estimating the physics of a scene. This is done with an LLm at the high level iteratively refining the physical formulas governing the particles in the scene (through an evolutionary algorith) and a low-level differentiable particle-based simulator refining the individual particle properties."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"- **Q.1:** More of a comment: Your contributions are weirdly written. I think contibution 1 is actually made up of of contibutions 2 and 3. So that should be one bullet point. And running experiments that validate your method is not a contribution, so contribution 4 should be removed.\n- **Q.2:** Another suggestion: since MPM simulation is so important to your method, maybe spend a short paragraph explaining it to everyone in the main body of the paper."},"rating":{"value":8},"details_of_ethics_concerns":{"value":"n/a"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":4},"strengths":{"value":"- **S.1:** Great idea. I think this is a really beautiful and elegant idea. Having the physics defined in the outer loop and the parameters tuned in the inner loop is great.\n- **S.2:** Clear Results. As far as I can tell, the results look pretty impressive and the paper compares its method to a variety of contemporary baselines, which is great.\n- **S.3:** Reproducibility. I appreciate that the authors released their source code, additional videos, and all LLM prompts."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **W.1:** The writing is incredibly dense. I've worked in parameter identification for simulator tuning and I've written my own physics engines and I was barely able to follow all of this. A bit simpler writing would greatly benefit this paper. For example, it took me a while to understand why you're not just using the gradient that you get from the differentiable simulator rollout to also tune the high-level physics."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358902813,"tcdate":1762233234763,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16648/Reviewer_HDyU"],"signatures":["ICLR.cc/2026/Conference/Submission16648/Reviewer_HDyU"],"forum":"eWoUcwEtLt","number":4,"license":"CC BY 4.0","cdate":1762233234763,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16648/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358902813,"domain":"ICLR.cc/2026/Conference","replyto":"eWoUcwEtLt","id":"ospPyuwakA","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["3D Gaussian Splatting","Physics-Driven 4D Interaction","Scientific Discovery","Intrinsic Dynamics"]},"supplementary_material":{"value":"/attachment/c4addf925e9c8e7366b5147b0971bec365061c81.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"The intrinsic dynamics of an object governs its physical behavior in the real world, playing a critical role in enabling physically plausible interactive simulation with 3D assets. Existing methods have attempted to infer the intrinsic dynamics of objects from visual observations, but generally face two major challenges: one line of work relies on manually defined constitutive priors, making it difficult to align with actual intrinsic dynamics; the other models intrinsic dynamics using neural networks, resulting in limited interpretability and poor generalization. To address these challenges, we propose VisionLaw, a bilevel optimization framework that infers interpretable expressions of intrinsic dynamics from visual observations. At the upper level, we introduce an LLMs-driven decoupled constitutive evolution strategy, where LLMs are prompted to act as physics experts to generate and revise constitutive laws, with a built-in decoupling mechanism that substantially reduces the search complexity of LLMs. At the lower level, we introduce a vision-guided constitutive evaluation mechanism, which utilizes visual simulation to evaluate the consistency between the generated constitutive law and the underlying intrinsic dynamics, thereby guiding the upper-level evolution. Experiments on both synthetic and real-world datasets demonstrate that VisionLaw can effectively infer interpretable intrinsic dynamics from visual observations. It significantly outperforms existing state-of-the-art methods and exhibits strong generalization for interactive simulation in novel scenarios."},"_bibtex":{"value":"@inproceedings{\nlin2026visionlaw,\ntitle={VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization},\nauthor={Jiajing Lin and Shu Jiang and Qingyuan Zeng and Zhenzhong Wang and Min Jiang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=eWoUcwEtLt}\n}"},"title":{"value":"VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization"},"pdf":{"value":"/pdf/b53b15cdf6ad888f73077e1f5a14c67ff3ff2e0d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lin|visionlaw_inferring_interpretable_intrinsic_dynamics_from_visual_observations_via_bilevel_optimization"},"authorids":{"value":["~Jiajing_Lin1","~Shu_Jiang3","~Qingyuan_Zeng2","~Zhenzhong_Wang2","~Min_Jiang1"]},"authors":{"value":["Jiajing Lin","Shu Jiang","Qingyuan Zeng","Zhenzhong Wang","Min Jiang"]}},"version":2},{"content":{"venue":{"value":"IEEE Transactions on Medical Imaging"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/42/11165775/10330700.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"zhang|toward_better_generalization_using_synthetic_data_a_domain_adaptation_framework_for_t2_mapping_via_multiple_overlappingecho_acquisition"},"html":{"value":"https://doi.org/10.1109/TMI.2023.3335212"},"abstract":{"value":"The generation of synthetic data using physics-based modeling provides a solution to limited or lacking real-world training samples in deep learning methods for rapid quantitative magnetic resonance imaging (qMRI). However, synthetic data distribution differs from real-world data, especially under complex imaging conditions, resulting in gaps between domains and limited generalization performance in real scenarios. Recently, a single-shot qMRI method, multiple overlapping-echo detachment imaging (MOLED), was proposed, quantifying tissue transverse relaxation time ( $\\text {T}_{{2}}$ ) in the order of milliseconds with the help of a trained network. Previous works leveraged a Bloch-based simulator to generate synthetic data for network training, which leaves the domain gap between synthetic and real-world scenarios and results in limited generalization. In this study, we proposed a  $\\text {T}_{{2}}$  mapping method via MOLED from the perspective of domain adaptation, which obtained accurate mapping performance without real-label training and reduced the cost of sequence research at the same time. Experiments demonstrate that our method outshined in the restoration of MR anatomical structures."},"title":{"value":"Toward Better Generalization Using Synthetic Data: A Domain Adaptation Framework for T2 Mapping via Multiple Overlapping-Echo Acquisition"},"authors":{"value":[{"fullname":"Chi Zhang","username":"~Chi_Zhang50"},{"fullname":"Qizhi Yang"},{"fullname":"Linyu Fan"},{"fullname":"Shaocong Yu"},{"fullname":"Liyan Sun"},{"fullname":"Congbo Cai"},{"fullname":"Xinghao Ding"}]}},"tmdate":1789090880062,"pdate":1756684800000,"externalIds":["doi:10.1109/tmi.2023.3335212"],"tcdate":1764574673714,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Chi_Zhang50"],"forum":"f3qW171UcX","license":"CC BY-SA 4.0","number":19187,"cdate":1764162164179,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789090880062,"domain":"OpenReview.net/Public_Article","id":"f3qW171UcX","version":2},{"content":{"summary":{"value":"This paper proposes an alternative computational unit, Generative Matching Units (GMUs), for feedforward supervised learning architectures. The generalization ability of GMU is compared with MLPs in comprehensive synthetic experiments and real-data experiments. It is shown in these experiments that GMU-MLPs generalize better than the MLP baselines in most cases. On vision datasets, when compared with the performance of the CNN baseline, GMU-CNNs demonstrate better generalization."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. In this paper, GMU is compared with MLP. How does the performance of GMU compare with other newly proposed unit structures, for example, those structures mentioned in the Introduction section?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper proposed an alternative computational unit called GMU and its advantage is verified through synthetic and real data experiments.    \n\n2. Synthetic data experiments with different types of distribution are conducted, and GMU variants show better generalization, especially to out-of-distribution test data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The structure of this paper could be clearer to help readers understand it better.  For example, the Introduction section includes the literature review together with the motivation and definition of GMU, and it is followed by an alternative definition of GMU in Section 1.1. Perhaps the clarity of the organization could be improved, and the summary of the organization of the whole paper could be added in the Introduction section.\n\n2. The experiments on real datasets (Section 6) could be explained in more detail. Additionally, the results in Table 6 are not referenced in the paper.\n\n3. There are some typographical errors in this paper, such as ‘mathcalN’ in line 425.\n\n4. The abstract contains many specific equations and notations. The authors could provide a more intuitive overview of this paper here without mentioning too many detailed notations and equations."}},"nonreaders":[],"tmdate":1731428973514,"tcdate":1730449654048,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9193/Reviewer_vWvJ"],"signatures":["ICLR.cc/2025/Conference/Submission9193/Reviewer_vWvJ"],"forum":"PJojB68YBu","number":2,"license":"CC BY 4.0","cdate":1730449654048,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9193/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428973514,"domain":"ICLR.cc/2025/Conference","replyto":"PJojB68YBu","id":"j1P41fyBMS","forumContent":{"TLDR":{"value":"We propose a new computational unit for feedforward supervised learning architectures, called generative matching units."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["supervised learning","classification","robustness"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We propose an alternative computational unit for feedforward supervised learning architectures, called Generative Matching Units (GMUs). To understand GMUs, we start with the standard perceptron unit and view it as an undirected symmetric measure of computation between the weights $W=[w_1,w_2,..w_d]$ and each input datapoint $X=[x_1,x_2,..,x_d]$. Perceptrons forward $W^TX+b$, which is usually followed by an activation function. In contrast, GMUs compute a directed asymmetric measure of computation that estimates the degree of functional dependency $f$ of the input elements $x_i$ of each datapoint to the weights $w_i$ in terms of latent generative variables $\\theta$, i.e,  $f(w_i,\\theta) \\rightarrow x_i$.  In order to estimate the functional dependency, GMUs measure the minimum error $\\sum (f(w_i,\\theta)-x_i)^2$ incurred in the generation process by optimizing $\\theta$ for each input datapoint. Subsequently, GMUs map the error into a functional dependency measure via an appropriate scalar function, and forward it to the next layer for further computation. In GMUs, the weights $[w_1,w_2,..,w_d]$ can therefore be interpreted as the $\\textit{generative weights}$. We first compare the generalization ability of GMUs and multi-layered-perceptrons (MLPs) via comprehensive synthetic experiments across a range of diverse settings. The most notable finding is that when the input is a sparse linear combination of latent generating variables, GMUs generalize significantly better than MLPs. Subsequently, we evaluate Resnet MLP networks where the first feedforward layer is replaced by GMUs (GMU-MLP) on 30 tabular datasets and find that in most cases, GMU-MLPs generalize better than the MLP baselines. We also compare GMU-MLP to a set of other benchmarks, including TabNet, XGBoost, etc. Lastly, we evaluate GMU-CNNs on three standard vision datasets and find that in all cases they generalize better than the corresponding CNN baselines. We also find that GMU-CNNs are significantly more robust to test-time corruptions."},"_bibtex":{"value":"@misc{\nghosh2024generative,\ntitle={Generative Matching Units for Supervised Learning},\nauthor={Rohan Ghosh and Mehul Motani},\nyear={2024},\nurl={https://openreview.net/forum?id=PJojB68YBu}\n}"},"title":{"value":"Generative Matching Units for Supervised Learning"},"pdf":{"value":"/pdf/013240e50ee7e6774b528709d49eac028abfe20f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"ghosh|generative_matching_units_for_supervised_learning"},"authorids":{"value":["~Rohan_Ghosh1","~Mehul_Motani1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rohan Ghosh","Mehul Motani"]}},"version":2},{"content":{"summary":{"value":"Fair4Free presents a novel approach to generating synthetic data that is fair and unbiased using data-free distillation. This method is particularly useful when access to training data is restricted due to privacy concerns. The process involves distilling knowledge from a pre-trained teacher model to a smaller student model without using real data, relying instead on noise input. The paper claims significant improvements over other models in terms of fairness, utility, and synthetic quality."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"see weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.  The data-free approach allows the generation of fair data even when access to original datasets is restricted, addressing significant privacy concerns.\n2.  By using a smaller student model, Fair4Free reduces computational resources and enables potential deployment on edge devices.\n3.  The model demonstrates superior performance across fairness, utility, and synthetic quality metrics compared to state-of-the-art models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Complexity and Overhead: While reducing computational costs via a smaller model, the initial setup of training and distilling between models can be complex and resource-intensive.\n2.  How does the performance of synthetic samples generated by Fair4Free compare to non-synthetic data in terms of real-world utility and fairness?\n3. How robust is Fair4Free to shifts in data distributions that might occur in practical scenarios?"}},"nonreaders":[],"tmdate":1731428297434,"tcdate":1730481272144,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6131/Reviewer_emFZ"],"signatures":["ICLR.cc/2025/Conference/Submission6131/Reviewer_emFZ"],"forum":"iRgzG5DKgA","number":2,"license":"CC BY 4.0","cdate":1730481272144,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6131/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428297434,"domain":"ICLR.cc/2025/Conference","replyto":"iRgzG5DKgA","id":"Ccsr1B04Yf","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["data fairness","fair generative models","knowledge distillation","latent space distillation","synthetic data","biased data"]},"supplementary_material":{"value":"/attachment/96403b6eb869aa033b81be0ee2f8d2ad8174de3a.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This work presents Fair4Free, a novel generative model to generate synthetic fair data using data-free distillation in the latent space. Fair4Free can work on the situation when the data is private or inaccessible.  In our approach, we first train a teacher model to create fair representation and then distil the knowledge to a student model (using a smaller architecture). The process of distilling the student model is data-free, i.e. the student model does not have access to the training dataset while distilling. After the distillation, we use the distilled model to generate fair synthetic samples. Our extensive experiments show that our synthetic samples outperform state-of-the-art models in all three criteria (fairness, utility and synthetic quality) with a performance increase of 5\\% for fairness, 8\\% for utility and 12\\% in synthetic quality for both tabular and image datasets."},"_bibtex":{"value":"@misc{\nsikder2025fairfree,\ntitle={Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data-Free Distillation},\nauthor={Md Fahim Sikder and Daniel de Leng and Fredrik Heintz},\nyear={2025},\nurl={https://openreview.net/forum?id=iRgzG5DKgA}\n}"},"title":{"value":"Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data-Free Distillation"},"pdf":{"value":"/pdf/bf09a39d8fc93c7601b5f410a98be727c91cb446.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"sikder|fair4free_generating_highfidelity_fair_synthetic_samples_using_datafree_distillation"},"authorids":{"value":["~Md_Fahim_Sikder1","~Daniel_de_Leng1","~Fredrik_Heintz1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Md Fahim Sikder","Daniel de Leng","Fredrik Heintz"]}},"version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2303.01055v2"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"kapoor|physicsinformed_neural_networks_for_solving_forward_and_inverse_problems_in_complex_beam_systems"},"authorids":{"value":["~Taniya_Kapoor1","https://dblp.org/search/pid/api?q=author:Hongrui_Wang_0001:","https://dblp.org/search/pid/api?q=author:Alfredo_Núñez:","https://dblp.org/search/pid/api?q=author:Rolf_P._B._J._Dollevoet:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2303.01055"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2303-01055,\n  publtype={informal},\n  author={Taniya Kapoor and Hongrui Wang and Alfredo Núñez and Rolf P. B. J. Dollevoet},\n  title={Physics-informed neural networks for solving forward and inverse problems in complex beam systems},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2303.01055},\n  url={https://doi.org/10.48550/arXiv.2303.01055}\n}\n"},"abstract":{"value":"This paper proposes a new framework using physics-informed neural networks (PINNs) to simulate complex structural systems that consist of single and double beams based on Euler-Bernoulli and Timoshenko theory, where the double beams are connected with a Winkler foundation. In particular, forward and inverse problems for the Euler-Bernoulli and Timoshenko partial differential equations (PDEs) are solved using nondimensional equations with the physics-informed loss function. Higher-order complex beam PDEs are efficiently solved for forward problems to compute the transverse displacements and cross-sectional rotations with less than 1e-3 percent error. Furthermore, inverse problems are robustly solved to determine the unknown dimensionless model parameters and applied force in the entire space-time domain, even in the case of noisy data. The results suggest that PINNs are a promising strategy for solving problems in engineering structures and machines involving beam systems."},"title":{"value":"Physics-informed neural networks for solving forward and inverse problems in complex beam systems"},"authors":{"value":["Taniya Kapoor","Hongrui Wang","Alfredo Núñez","Rolf P. B. J. Dollevoet"]}},"tmdate":1736500796422,"pdate":1672531200000,"tcdate":1736500789489,"writers":["~"],"signatures":["~Taniya_Kapoor1"],"forum":"OY8tZW5aBq","license":"CC BY-SA 4.0","number":261149,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1736500796422,"domain":"DBLP.org","id":"OY8tZW5aBq","version":2},{"content":{"summary":{"value":"This paper introduces a framework to progressively refine the resolution of physics solvers during neural network training. The authors demonstrate that training a neural network with a physics solver scheduled to increase in iterations $K$ over training can significantly reduce computational costs, especially by using fewer iterations in the early training phases (*progressive refinement* savings). Additionally, they observe that a neural network can be effectively trained even when the physics solver has not fully converged, eliminating the need for an extensive number of solver steps to achieve high accuracy (*incomplete convergence* savings). To automate this process and determine optimal parameters for these refinements, the authors propose an algorithm that monitors validation set metrics, incrementally increasing solver refinement when performance plateaus. The framework is validated across four use cases: a linear inverse solver, linear neural emulator learning, nonlinear neural emulator learning, and a neural-hybrid emulator."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* How sensitive are the parameters of PRDP? Did you experiment with many different parameters for each problem before achieving the results, or was it relatively straightforward?\n* Do you expect the method to achieve similar savings in training time for domains with complex boundary conditions and irregular meshing?\n* Do you think your method could work where the physics solver contains incomplete physics (e.g., the parameters are not perfectly calibrated, and some terms of the equations could be missing)?\n* How would the method be affected by noise in the observations?\n* How well do you think this could help reduce the training time of neural GCMs [1]?\n\n[1] Kochkov et al. Neural general circulation models for weather and climate. Nature, 2024."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The paper is well-written, with clear explanations of the intuitions and motivations behind the method.\n* The algorithm is straightforward and effectively delivers the intended results, as seen in the reduction of validation loss with the progressive refinement of the physics solver, particularly notable in Figure 4 for the 2D Heat and 2D Navier-Stokes cases.\n* The savings in training time and computational resources are substantial.\n* The appendix is thorough and well-organized, with especially valuable details on iterative linear solvers and detailed derivations for each problem."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Overall, the technical contribution of the paper is somewhat limited, with the main novelty being the proposed algorithm for iterative refinements.\n* All physics solvers employed rely on iterative linear solvers. It would have been interesting to see if the method also applies with other physics solvers.\n* With the exception of the final example (the neural-hybrid approach), the other examples appear to be simplified or illustrative cases without clear, concrete applications."}},"nonreaders":[],"tmdate":1731427710340,"tcdate":1730463566580,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3012/Reviewer_5MWA"],"signatures":["ICLR.cc/2025/Conference/Submission3012/Reviewer_5MWA"],"forum":"9Fh0z1JmPU","number":3,"license":"CC BY 4.0","cdate":1730463566580,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3012/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427710340,"domain":"ICLR.cc/2025/Conference","replyto":"9Fh0z1JmPU","id":"KdIZWHyvvY","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["differentiable physics","iterative PDE solvers","neural surrogate"]},"supplementary_material":{"value":"/attachment/21fd955829533ffcb5edab8123b77632a930e068.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The physics solvers employed for neural network training are primarily iterative, and hence, differentiating through them introduces a severe computational burden as iterations grow large. Inspired by works in bilevel optimization, we show that full accuracy of the network is achievable through physics significantly coarser than fully converged solvers. We propose *progressively refined differentiable physics* (PRDP), an approach that identifies the level of physics refinement sufficient for full training accuracy. By beginning with coarse physics, adaptively refining it during training, and stopping refinement at the level adequate for training, it enables significant compute savings without sacrificing network accuracy. Our focus is on differentiating iterative linear solvers for sparsely discretized differential operators, which are fundamental to scientific computing. PRDP is applicable to both unrolled and implicit differentiation. We validate its performance on a variety of learning scenarios involving differentiable physics solvers such as inverse problems, autoregressive neural emulators, and correction-based neural-hybrid solvers. In the challenging example of emulating the Navier-Stokes equations, we reduce training time by 62%."},"_bibtex":{"value":"@inproceedings{\nbhatia2025prdp,\ntitle={{PRDP}: Progressively Refined Differentiable Physics},\nauthor={Kanishk Bhatia and Felix Koehler and Nils Thuerey},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=9Fh0z1JmPU}\n}"},"title":{"value":"PRDP: Progressively Refined Differentiable Physics"},"pdf":{"value":"/pdf/82abf59987428b8451c36c339fb2fc08f5b0afdf.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"bhatia|prdp_progressively_refined_differentiable_physics"},"authorids":{"value":["~Kanishk_Bhatia1","~Felix_Koehler1","~Nils_Thuerey1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kanishk Bhatia","Felix Koehler","Nils Thuerey"]}},"version":2},{"content":{"summary":{"value":"The paper investigates whether classifiers trained on synthetic data can match the performance of those trained on real data, especially in visual tasks. Through comparative analysis, the authors find that although synthetic classifiers achieve similar overall accuracy to real-data-trained counterparts, they underperform in challenging scenarios such as fine-grained classification, extreme object scales, and brightness variations. The authors attribute these limitations to the inability of current generative models to fully capture the complexity and diversity of real-world data. To address these issues, they propose **RealTune**, a method that fine-tunes synthetic classifiers with a small amount of real data, significantly improving performance in these complex scenarios. The results demonstrate that combining synthetic and real data is essential to create more robust and efficient classifiers."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1.*How well does RealTune generalize across different visual datasets or domains?*\nRealTune was tested on ImageNet and its smaller subset, ImageNet-100, which are certainly widely used benchmarks. But I wonder—would the same level of improvement hold if we used a dataset with a lot more variation, like COCO, or even a domain-specific dataset like medical imagery? These types of datasets have different complexities, like intricate scene layouts or highly specialized objects, which might expose weaknesses that did not appear with ImageNet. It would be helpful to know if RealTune is a broadly applicable method for any kind of visual domain or if its effectiveness is more specific to the characteristics of ImageNet.\n\n2.*How does the quality of the fine-tuning data affect RealTune's success?*\nThe authors showed that RealTune works well with a small amount of real data, but it made me wonder about the specifics of that fine-tuning data. What if the real data used for fine-tuning is biased or lacks diversity? Would RealTune still perform as well? Real-world data collection often has biases, like unbalanced classes or limited examples of certain challenging conditions, and it would be useful to understand how that affects the final model. Does RealTune require a carefully balanced and curated fine-tuning set, or can it adapt well even if the real data is subpar? It would be great to see more on whether the composition of this fine-tuning data matters.\n\n3.*How does RealTune's computational efficiency compare with other methods in terms of energy use and scalability?*\nRealTune was described as efficient, running on a single GPU in a relatively short time. But I would like to know more about how it stacks up against other methods in terms of energy usage, especially considering the push for more environmentally friendly AI solutions. How does RealTune compare, for example, to larger pretraining strategies or other fine-tuning techniques when it comes to energy consumption or the practicality of scaling to larger models? Understanding these trade-offs would be really important for someone trying to decide between RealTune and other enhancement methods, especially in scenarios with limited computational resources. It would also help gauge if the method's efficiency benefits hold up when scaling to bigger models or if there are diminishing returns."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper provides a detailed comparison between classifiers trained on synthetic versus real data, especially in challenging scenarios like fine-grained classification and rare situations (e.g., extreme object scales, brightness variations, object occlusion). The authors use a range of quantitative metrics, such as employing YOLO models for object scale assessment and CIELAB color space for brightness evaluation, making their results systematic and reliable. This thorough analysis adds significant value to understanding the limitations of synthetic classifiers and potential improvements for future model development.\n2. Generating synthetic data is relatively easy compared to collecting real-world data, but the quality and diversity often fall short. RealTune effectively addresses these shortcomings using minimal real data (around 3% of ImageNet), significantly enhancing the classifier’s performance. The proposed method also reduces computational and energy costs, making it particularly useful in resource-constrained environments. From a review perspective, this solution is practical, cost-effective, and demonstrates a clear advancement in addressing synthetic data challenges.\n3. Beyond RealTune, the authors also explore the effectiveness of mixed data pretraining, showcasing the potential of combining real and synthetic data. This strategy not only improves classifier accuracy but also outperforms using either real or synthetic data alone. The experiments on mixed pretraining provide new insights into leveraging limited data resources effectively, which is crucial in scenarios where collecting large-scale real data is impractical. These contributions not only provide an empirical foundation for future research but also serve as valuable guidance for practical training strategies. The discussion of data-mixing approaches adds depth to the paper and offers an innovative direction for making optimal use of synthetic data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper exclusively focuses on vision tasks without exploring other modalities like text or audio, which limits the generalizability of the findings. Given the growing relevance of synthetic data in various domains beyond vision (such as natural language and speech), this is a significant limitation. The authors briefly mention extending the study to other modalities in future work (Section 6), but as it stands, the narrow focus weakens the overall contribution. Extending the experiments to non-visual modalities or providing some preliminary analysis could have increased the paper’s broader applicability.\n2. In the evaluation of rare scenarios, such as object occlusion, extreme object scales, and brightness variations, the authors use ImageNet-X (Figure 3). However, the analysis lacks an adequate diversity of datasets that could highlight other real-world challenges (e.g., dynamic backgrounds, motion blur, or domain shifts). ImageNet-X, though useful, does not fully encompass the variety of rare scenarios that might be seen in real-world applications. For a thorough evaluation, incorporating other benchmarks like ImageNet-C or ObjectNet could have provided a more comprehensive assessment of classifier robustness. This would also have helped to better validate the claims about synthetic data limitations.\n3. Some figures in the paper are difficult to interpret due to suboptimal presentation choices. For example, Figure 4, which aims to demonstrate class consistency and frequency for fine-grained class confusion, uses bar plots that make it challenging to interpret the differences across models at a glance. The class consistency rates presented are not clearly distinguished, leading to potential confusion. Additionally, the visualization of synthetic data distribution shows class imbalance, but the representation could have been more effective with a clearer breakdown across different class labels to provide insight into how imbalance specifically affects model performance. The authors could benefit from using more readable visualizations like heatmaps or swarm plots that better convey the underlying relationships.\n4. While RealTune demonstrates promising improvements, the experimental evaluation lacks adequate baseline comparisons against other established methods for enhancing synthetic classifiers. For example, the paper introduces SynTune as a counterpart to RealTune, but it would have been more informative to compare RealTune with other fine-tuning or data augmentation techniques that are popular in the field. Including a detailed comparison with transfer learning or data distillation methods could have added value to demonstrate RealTune's efficiency more convincingly. Moreover, while Figure 6 shows accuracy improvements, adding more baselines could help to better gauge the significance of the presented results.\n5. The mixed training strategy, explored in Section 4.3 and illustrated in Table 3, shows that combining real and synthetic data during pretraining can enhance the performance of classifiers. However, the analysis lacks depth in explaining why certain combinations outperform others and how different ratios of real to synthetic data impact the results. The choice of using only 7.7% of real data (for ImageNet-100) in the mixed dataset is arbitrary and not well justified, limiting the ability to generalize findings to other datasets or settings. Additionally, there are no detailed ablation studies exploring various ratios between real and synthetic data, which would have provided a better understanding of the trade-offs and the optimal way to mix these data types. This omission reduces the experimental rigor and leaves questions regarding the optimal strategy for practical scenarios."}},"nonreaders":[],"tmdate":1732286793139,"tcdate":1729023440629,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8604/Reviewer_rppm"],"signatures":["ICLR.cc/2025/Conference/Submission8604/Reviewer_rppm"],"forum":"oClr2P7V0T","number":1,"license":"CC BY 4.0","cdate":1729023440629,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8604/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732286793139,"domain":"ICLR.cc/2025/Conference","replyto":"oClr2P7V0T","id":"9vJbeIdisq","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["generative model","representation learning","synthetic data"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Foundation models have achieved significant advancements across various domains, yet their training demands vast amounts of real-world data, which is becoming increasingly scarce. To address this challenge, synthetic data has garnered substantial interest as an alternative for augmenting training datasets in fields such as computer vision and natural language processing. However, skepticism remains regarding whether synthetic classifiers can match the performance of those trained on real data. In this paper, we investigate this question by conducting a detailed analysis within the realm of visual tasks, comparing classifiers trained on synthetic versus real data using CLIP and ViT. Our results reveal that synthetic classifiers exhibit deficiencies in a range of challenging real-world scenarios, such as fine-grained classification, extreme object scales and extreme brightness despite achieving comparable overall accuracy to their real-data-trained counterparts. We find that the limitations of synthetic classifiers can be traced back to the limitations of current generative models in capturing the complexity and diversity of real-world data in these aspects. To mitigate these issues efficiently, we explore \\textbf{RealTune}, a simple method that enhances synthetic classifiers by finetuning them with a small amount of real data. Experimental evaluations demonstrate that RealTune significantly improves the performance of synthetic classifiers using only a limited real dataset (e.g., 40k images,  3% of ImageNet) with minimal training time (e.g., 1hour on a single NVIDIA RTX 3090 GPU). Our findings indicate that while synthetic data is a valuable resource, integrating real and synthetic data is essential to achieve robust and efficient classifiers. This work underscores the necessity of leveraging both data types to bridge the performance gap and enhance the overall effectiveness of foundation models."},"_bibtex":{"value":"@misc{\nzhang2025are,\ntitle={Are Synthetic Classifiers Really as Good as Real Classifiers?},\nauthor={Jizhe Zhang and Yifei Wang and Stefanie Jegelka and Yisen Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=oClr2P7V0T}\n}"},"title":{"value":"Are Synthetic Classifiers Really as Good as Real Classifiers?"},"pdf":{"value":"/pdf/5c7c6c96146c459ff24309065eadc679ba88b158.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|are_synthetic_classifiers_really_as_good_as_real_classifiers"},"authorids":{"value":["~Jizhe_Zhang1","~Yifei_Wang1","~Stefanie_Jegelka3","~Yisen_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jizhe Zhang","Yifei Wang","Stefanie Jegelka","Yisen Wang"]}},"version":2},{"content":{"summary":{"value":"This paper explores an unusual route of training a large-scale 3D reconstruction model using synthetic data. It demonstrates that high-quality reconstruction can be achieved solely with synthetic procedural data, bypassing the need for real, hand-crafted 3D models, which are challenging to collect. The paper, trains two reconstruction models, one with objaverse dataset and the other with synthetic dataset which the paper proposes (Zeroverse). Test results of these models on ABO and Google Scanned Dataset shows that competitative reconstruction quality can be achieved by just synthetic data. With this the paper showcases that, global semantics of an object are not crucial for reconstruction. Consequently, similar reconstruction quality can be attained using complex geometric synthetic data with rich textures, even if they lack global semantics."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to weakness."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- **Novelty** - The use of synthetic data for reconstruction task is novel. The synthetic data generated by the method in this paper can also be used for data augmentation in other tasks. \n- **Clarity** - The paper is well written with good attention to detail. \n- **Results** - The approach has been appropriately validated on different datasets (Google Scanned and ABO).\n- **Related Work** - Comprehensive releated works are covered in the paper."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper lacks enough technical contributions. It seems that the entire paper is about ways to create synthetic data using predefined primitives and applying augmentations over different combinations of these primitives. After this once the data is prepared, an off-the-shelf gaussian splatting based reconstruction model is trained for performing experiments. \n- From the results in Table 7 in supplementary, it seems that reconstruction quality significantly suffers (~2-3.5 PSNR) when the input views are very sparse (4 views) as compared to results with 8 input views as shown in Table 1. What is the reason behind this performance dip? Is it the object semantics? \n- As this approach requires comparatively denser views (at least 8), I am concerned about its application in learning 3D priors, as generally 3D prior models work with extremely sparse views.\n- The paper compares results by training models with same quantity of data for both real and synthetic datasets. It would be interesting to see if the reconstruction quality of the model trained with synthetic data improves or achieves comparable results when the dataset size is increased."},"limitations":{"value":"Adequate limitations are discussed by the authors in Appendix."}},"nonreaders":[],"tmdate":1730878621136,"tcdate":1721355923879,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission49/Reviewer_xM78"],"signatures":["NeurIPS.cc/2024/Conference/Submission49/Reviewer_xM78"],"forum":"MtRvzJBsBA","number":2,"license":"CC BY 4.0","cdate":1721355923879,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission49/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878621136,"domain":"NeurIPS.cc/2024/Conference","replyto":"MtRvzJBsBA","id":"mSF0xAg40w","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"We train a Large Reconstruction Model (LRM) on a synthesized 3D dataset and closely match the performance of the state-of-the-art LRM model trained on real 3D data."},"keywords":{"value":["3D Reconstruction","Transformer","Pre-training","Synthetic Data"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"We present LRM-Zero, a Large Reconstruction Model (LRM) trained entirely on synthesized 3D data, achieving high-quality sparse-view 3D reconstruction. The core of LRM-Zero is our procedural 3D dataset, Zeroverse, which is automatically synthesized from simple primitive shapes with random texturing and augmentations (e.g., height fields, boolean differences, and wireframes). Unlike previous 3D datasets (e.g., Objaverse) which are often captured or crafted by humans to approximate real 3D data, Zeroverse completely ignores realistic global semantics but is rich in complex geometric and texture details that are locally similar to or even more intricate than real objects. We demonstrate that our LRM-Zero, trained with our fully synthesized Zeroverse, can achieve high visual quality in the reconstruction of real-world objects, competitive with models trained on Objaverse. We also analyze several critical design choices of Zeroverse that contribute to LRM-Zero's capability and training stability. Our work demonstrates that 3D reconstruction, one of the core tasks in 3D vision, can potentially be addressed without the semantics of real-world objects. The Zeroverse's procedural synthesis code and interactive visualization are available at: https://desaixie.github.io/lrm-zero/."},"_bibtex":{"value":"@inproceedings{\nxie2024lrmzero,\ntitle={{LRM}-Zero: Training Large Reconstruction Models with Synthesized Data},\nauthor={Desai Xie and Sai Bi and Zhixin Shu and Kai Zhang and Zexiang Xu and Yi Zhou and Soren Pirk and Arie Kaufman and Xin Sun and Hao Tan},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=MtRvzJBsBA}\n}"},"title":{"value":"LRM-Zero: Training Large Reconstruction Models with Synthesized Data"},"pdf":{"value":"/pdf/e1be4c6318db6da389fa5e7cc8d7250e92650ba6.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"xie|lrmzero_training_large_reconstruction_models_with_synthesized_data"},"authorids":{"value":["~Desai_Xie1","~Sai_Bi1","~Zhixin_Shu1","~Kai_Zhang7","~Zexiang_Xu1","~Yi_Zhou1","~Soren_Pirk1","~Arie_Kaufman1","~Xin_Sun11","~Hao_Tan1"]},"authors":{"value":["Desai Xie","Sai Bi","Zhixin Shu","Kai Zhang","Zexiang Xu","Yi Zhou","Soren Pirk","Arie Kaufman","Xin Sun","Hao Tan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a physics-informed hybrid neural network for radial-phase retrieval from intensity-only measurements, addressing the challenging problem of outer-ring generalization. The authors combine three key components: (i) radial priors with smooth exponentiated splines and a monotone outer-radius booster, (ii) dual differentiable PDE branches (Kerr-NLSE for high-frequency synthesis and TIE for coarse structure), and (iii) strict radial projection with radius-dependent α-fusion. The method is trained on only 1-3 inner rings and tested on 4-9 outer rings. Experimental results demonstrate better reconstruction quality, better stability in peak positions and amplitude calibration, and improved generalization compared to baseline methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Do the authors have plans to validate the method on real experimental optical data? What specific challenges do they anticipate in transitioning from synthetic to real data, and how might the method be adapted to handle domain shift?\n\n2. *Can the authors provide quantitative results on how the method degrades under increasing levels of asymmetry (e.g., varying degrees of astigmatism)? At what point does the radial symmetry assumption break down?\n\n3. Given the sensitivity to parameter mismatch, what calibration accuracy is needed for practical deployment? Could the method be extended to estimate or adapt to unknown system parameters jointly?\n\n4. The local PL inequality (Theorem 2) requires specific assumptions. How often are these assumptions satisfied in practice, and what happens when they are violated?\n\n5. Beyond the empirical generalization results, are there theoretical bounds on how many rings the method can extrapolate to, given training on N rings?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1) The paper combines multiple physics-informed components: (i) radial priors with monotone outer-radius boosting, (ii) dual PDE branches (Kerr-NLSE and TIE), and (iii) strict radial projection with radius-dependent α-fusion. This approach is original in addressing the outer-ring generalization challenge in phase retrieval.\n\n2) The work demonstrates theoretical foundations, including Lipschitz stability analysis (Propositions 1-3), identifiability guarantees (Theorem 1), and local PL-type convergence (Theorem 2). The experimental design is thorough with comprehensive evaluations across multiple dimensions (E1-E7), including ablations, data efficiency studies, and robustness tests.\n\n3) The paper makes contributions to physics-informed neural networks for inverse problems. The demonstrated ability to generalize from 1-3 rings to 4-9 rings addresses a fundamental challenge in computational imaging. The work has broad applicability beyond phase retrieval to other optical inverse problems, MRI, ultrasound, and PDE-based reconstruction tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) The paper relies on synthetic data with controlled Fraunhofer diffraction patterns. While robustness tests include noise and parameter mismatch, the absence of real experimental optical data raises concerns about domain transfer and practical applicability in actual settings.\n\n2) The strict radial projection and entire framework assume effective radial symmetry. The paper acknowledges this limitation (Section 8) but does not provide an empirical evaluation of performance degradation under asymmetric conditions, such as astigmatism, ellipticity, or off-axis aberrations, which commonly occur in real optical systems.\n\n3) The robustness experiments (E4) reveal that parameter mismatch in wavelength and focal length is the dominant failure mode, with precision dropping from 0.75 to 0.43 and recovered rings from 5.1 to 2.8. This sensitivity suggests the method may require precise system calibration, limiting practical deployment.\n\n4) While the paper compares against U-Net and FNO, it does not evaluate against recent physics-informed phase retrieval methods or classical iterative approaches like Gerchberg-Saxton variants with regularization, making it difficult to assess the relative contribution of the specific architectural choices.\n\n5) Although latency is reported (27.6ms vs 21.3ms for U-Net), the paper does not discuss training time, convergence speed, or the practical cost of the dual-PDE forward pass during inference at scale. The 30% increase in latency may be significant for real-time applications."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925478362,"tcdate":1761505413103,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15167/Reviewer_5Dxa"],"signatures":["ICLR.cc/2026/Conference/Submission15167/Reviewer_5Dxa"],"forum":"jS3EKPSaAR","number":1,"license":"CC BY 4.0","cdate":1761505413103,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15167/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925478362,"domain":"ICLR.cc/2026/Conference","replyto":"jS3EKPSaAR","id":"QibZjes95e","forumContent":{"TLDR":{"value":"a pde based physic informed neural network that performs well in generalization"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["phase retrieval;PDE networks; outer-ring extrapolation; inverse problems"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Phase retrieval from intensity-only measurements is severely ill-posed due to global-gauge and rotational symmetries. We consider outer-ring generalization: training with supervision from only a few inner rings and testing the model’s ability to reconstruct a broader set of unseen outer rings. We introduce a physics-informed hybrid network that combines (i) radial priors encoded by a smooth exponentiated spline and a \\emph{monotone} outer-radius booster, (ii) two differentiable PDE branches---a Strang-split Kerr--NLSE pathway for high-frequency synthesis and a TIE-based low-pass pathway for coarse structure---and (iii) a strict radial projection enforcing output symmetry, together with a radius-dependent $\\alpha$-fusion. Across the tested configurations, when trained only on a few rings (1-3), our model reconstructs more rings(4-9) than conventional methods, and achieves better stability in peak\npositions and amplitude calibration under out-of-distribution settings. This provides some inspiration for enhancing the generalization of physics-informed neural networks when applied to optical inverse problems. Ablations isolate the contribution of the alpha fusion, PDE coupling, and monotone\nboosting. We will release pseudo-code to facilitate reproducibility."},"_bibtex":{"value":"@misc{\nyao2026physicsinformed,\ntitle={{PHYSICS}-{INFORMED} {RADIAL} {PHASE} {RETRIEVAL} {NEURAL} {NETWORK} {WITH} {HYBRID} {DEEP} {PRIORS} {AND} {DUAL} {PDE}},\nauthor={ZIYONG YAO},\nyear={2026},\nurl={https://openreview.net/forum?id=jS3EKPSaAR}\n}"},"title":{"value":"PHYSICS-INFORMED RADIAL PHASE RETRIEVAL NEURAL NETWORK WITH HYBRID DEEP PRIORS AND DUAL PDE"},"pdf":{"value":"/pdf/d0ee395178c85a98450cfb5c748c70d7faeb7393.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yao|physicsinformed_radial_phase_retrieval_neural_network_with_hybrid_deep_priors_and_dual_pde"},"authorids":{"value":["~ZIYONG_YAO1"]},"authors":{"value":["ZIYONG YAO"]}},"version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Physics simulation","video generation","diffusion models"]},"supplementary_material":{"value":"/attachment/d78b0ed68e9aec228a565b1b8a23614ab692d817.zip"},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for physics-grounded image-to-video generation with physical parameters and force control. At its core is a generative physics network that learns the distribution of physical dynamics across four materials (elastic, sand, plasticine, and rigid) via a diffusion model conditioned on physics parameters and applied forces. We represent physical dynamics as 3D point trajectories and train on a large-scale synthetic dataset of 550K animations generated by physics simulators. We enhance the diffusion model with a novel spatiotemporal attention block that emulates particle interactions and incorporates physics-based constraints during training to enforce physical plausibility. Experiments show that PhysCtrl generates realistic, physics-grounded motion trajectories which, when used to drive image-to-video models, yield high-fidelity, controllable videos that outperform existing methods in both visual quality and physical plausibility. Our code, model and data will be made publicly available upon publication."},"_bibtex":{"value":"@inproceedings{\nwang2025physctrl,\ntitle={PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation},\nauthor={Chen Wang and Chuhao Chen and Yiming Huang and Zhiyang Dou and Yuan Liu and Jiatao Gu and Lingjie Liu},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=AHEKhff4Oa}\n}"},"title":{"value":"PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation"},"pdf":{"value":"/pdf/f004af8c5e76c1b850e491839d37683ebf36cfac.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"wang|physctrl_generative_physics_for_controllable_and_physicsgrounded_video_generation"},"authorids":{"value":["~Chen_Wang13","~Chuhao_Chen1","~Yiming_Huang5","~Zhiyang_Dou1","~Yuan_Liu3","~Jiatao_Gu1","~Lingjie_Liu1"]},"authors":{"value":["Chen Wang","Chuhao Chen","Yiming Huang","Zhiyang Dou","Yuan Liu","Jiatao Gu","Lingjie Liu"]}},"tmdate":1783627457584,"pdate":1758216540463,"tcdate":1744519310695,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission2097/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission2097/Authors"],"forum":"AHEKhff4Oa","license":"CC BY 4.0","number":2097,"cdate":1744519310695,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission2097/-/Full_Submission","NeurIPS.cc/2025/Conference/Submission2097/-/Supplementary_Material","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission2097/-/Camera_Ready_Revision"],"mdate":1783627457584,"odate":1761704737831,"domain":"NeurIPS.cc/2025/Conference","id":"AHEKhff4Oa","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2509.20358v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"wang|physctrl_generative_physics_for_controllable_and_physicsgrounded_video_generation"},"authorids":{"value":["","","","~Zhiyang_Dou1","","",""]},"html":{"value":"https://doi.org/10.48550/arXiv.2509.20358"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2509-20358,\n  publtype={informal},\n  author={Chen Wang and Chuhao Chen and Yiming Huang and Zhiyang Dou and Yuan Liu and Jiatao Gu and Lingjie Liu},\n  title={PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation},\n  year={2025},\n  month={September},\n  cdate={1756684800000},\n  journal={CoRR},\n  volume={abs/2509.20358},\n  url={https://doi.org/10.48550/arXiv.2509.20358}\n}\n"},"abstract":{"value":"Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for physics-grounded image-to-video generation with physical parameters and force control. At its core is a generative physics network that learns the distribution of physical dynamics across four materials (elastic, sand, plasticine, and rigid) via a diffusion model conditioned on physics parameters and applied forces. We represent physical dynamics as 3D point trajectories and train on a large-scale synthetic dataset of 550K animations generated by physics simulators. We enhance the diffusion model with a novel spatiotemporal attention block that emulates particle interactions and incorporates physics-based constraints during training to enforce physical plausibility. Experiments show that PhysCtrl generates realistic, physics-grounded motion trajectories which, when used to drive image-to-video models, yield high-fidelity, controllable videos that outperform existing methods in both visual quality and physical plausibility. Project Page: https://cwchenwang.github.io/physctrl"},"title":{"value":"PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation"},"authors":{"value":["Chen Wang","Chuhao Chen","Yiming Huang","Zhiyang Dou","Yuan Liu","Jiatao Gu","Lingjie Liu"]}},"tmdate":1772313705320,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2509-20358"],"tcdate":1772313688226,"writers":["~"],"signatures":["~Zhiyang_Dou1"],"forum":"OSUasXC6Ds","license":"CC BY-SA 4.0","number":840304,"cdate":1756684800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772313705320,"domain":"DBLP.org","id":"OSUasXC6Ds","version":2},{"content":{"summary":{"value":"The paper tackles the problem of inconsistent dataset taxonomies and prevents using multi-dataset combining in training. The authors propose training a shared feature extractor with two competing heads, a gating head to determine dataset-specific classes and a classification head to map samples to shared taxonomy emerging from training. \n\nThe authors perform some evaluation using a synthetic benchmark. They demonstrate validation on a few small-scale classification tasks. The methods provide a comprehensive discussion to a few related work, e.g. DynAlign and LMSeg, but it is yet clear how they would differentiate in the proposed settings and evaluations. This is part the method can potentially be further improved."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"I am overall less familiar with the literature in this space, so I raise the questions whether there are any improvements the authors can make to push this study closer to real world scenarios and validate with more recent works. Please help me clarify if there are clear distinctions that the suggestions are not appropriate. I will also be open to other reviewers' suggestion on this."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"SLAMDUNKS introduces a novel approach that unifies label alignment and multi-dataset training within a single end-to-end model. This design allows automatic discovery of semantic relations across datasets without manual mappings or reliance on external language models.\n\nThe paper also creates a synthetic benchmark with ground-truth taxonomies. It allows the paper and future work to define ground truth relations explicitly and a taxonomy setting which be validated. Under this setting, the methods can outperform existig baselines Missing Link and AUT."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the small-scale experiments on synthetic dataset is clean and nice, it was not intuitive to me how the approach would extend to real-world tasks such as segmentation or detection. There are plenty existing classic datasets in this spaces. Making some evaluation with efforts constructing a similar setting will be impactful. \n\nIn addition, while the paper shows clear advantage over the two classical baselines, both baselines Missing Link and AUT are not frontier according to the discussions in the related work. It would be good if the authors can bringer the mentioned newer work such as DynAlign, LMSeg in the evaluation loop, which would be enhance the claim to outperform state-of-the-art."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943155016,"tcdate":1761977114724,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24664/Reviewer_beds"],"signatures":["ICLR.cc/2026/Conference/Submission24664/Reviewer_beds"],"forum":"9ttjYx3NJQ","number":3,"license":"CC BY 4.0","cdate":1761977114724,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24664/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943155016,"domain":"ICLR.cc/2026/Conference","replyto":"9ttjYx3NJQ","id":"h5vu3fQJQd","forumContent":{"TLDR":{"value":"This paper considers automated relation discovery between classes in different datasets as a standalone task and proposes an architecture that simultaneously discovers those relations and trains the model using this information."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["computer vision","multi-dataset training","image classification","aligning taxonomies"]},"supplementary_material":{"value":"/attachment/3b4e67361d6f04169458f292725cd2d187cab411.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multi-dataset training is a key strategy for improving the versatility and robustness of deep models, but its effectiveness is often hindered by unaligned and contradictory dataset taxonomies. These inconsistencies introduce training noise and prevent effective knowledge sharing. To address this, we propose SLAMDUNKS, a framework for simultaneous multi-dataset training and label alignment. Its core is a shared feature extractor trained with two competing heads: a gating head that determines which dataset-specific classes should be shared, and a classification head that maps samples to the emerging shared taxonomy. To rigorously evaluate alignment quality, we introduce a synthetic benchmark where ground-truth relations are modeled as bipartite graphs. Our method demonstrates remarkable precision, perfectly recovering the true taxonomy (a Graph Edit Distance of 0) for same-domain datasets. Across more challenging cross-domain pairs, SLAMDUNKS achieves an Average Precision of 0.8, outperforming the state-of-the-art by 0.1 to 0.2 and validating its superior alignment capabilities."},"_bibtex":{"value":"@misc{\nbevandic2025finding,\ntitle={Finding agreement in disagreement: Simultaneous label alignment and multi-dataset training with {SLAMDUNKS}},\nauthor={Petra Bevandi{\\'c} and Riza Velioglu and Barbara Hammer},\nyear={2025},\nurl={https://openreview.net/forum?id=9ttjYx3NJQ}\n}"},"title":{"value":"Finding agreement in disagreement: Simultaneous label alignment and multi-dataset training with SLAMDUNKS"},"pdf":{"value":"/pdf/5f207b6a9ef40c685106d15f2466f203ee1a3a04.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"bevandi|finding_agreement_in_disagreement_simultaneous_label_alignment_and_multidataset_training_with_slamdunks"},"authorids":{"value":["~Petra_Bevandić1","~Riza_Velioglu1","~Barbara_Hammer4"]},"authors":{"value":["Petra Bevandić","Riza Velioglu","Barbara Hammer"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a real-to-sim approach to learning locomotion and navigation from pixels. The paper scales up real2sim visual data for training autonomous agents with a mix of data sources including iphone captures, video generative model outputs as well as existing multi-view datasets like arkit. The paper utilizes of off-the-shelf tools like VGGT and NeRFstudio to create 3D Gaussian Splats of these scenes which act as drop in renders for physics sim such as Isaac-gym. Experiments demonstrate sim2real transfer of locomtion policies. Other experiments demonstrate navigation can also be achieved with a similar method."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See questions and comments in the weakness section. I am looking forward to author's responses in the rebuttal."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"In my opinion, following are the strengths of the paper:\n\n1. Scaling up photorealistic real2sim data for policy learning is a major advantage. While still in static scenes, it shows the utility of current zero-shot foundation models can be robustly used to scale up visual data to train policies. \n\n2. Zero-shot sim2real deployment is a nice result demonstrating the method is able to get good performance with democratized data collection such as with iphone cameras or existing captures. \n\n3. The paper is nicely written and the visuals/diagrams support and complement the text very well. \n\n4. A vectorized rendering support for existing physics sim is a great feature to have in existing physics simulators to scale up real2sim learning, albeit with an initial capturing overhead."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"In my opinion, below are the weaknesses of the method:\n\n1. Lack of comparisons to existing similar works in this space: I think the paper is lacking comparisons or discussions to existing closely related works [1,2,3]. A comparison interms of visual realism as well as PSNR as well as efficiency would be great for the community to understand which approach is the most useful in terms of ease of acquiring real2sim vs accuracy axes. \n\n2. The evaluation setting for the single-task multi-environment real2sim policy that the paper trains for their main results is unclear. Can the authors show results on a common benchmark, perhaps something similar to EmbodiedSplat [2] where it is clear what the training data distribution is and if there are any gains from training a large policy on all the real2sim collected data on unseen enviornments i.e. whether the approach is able to quickly finetune etc.?\n\n3. With a high throughput, did the authors also experiment with real-world RL?\n\n4. Are the physics parameter tuned to achieve sim2real transfer? Details of these are missing in the paper. \n\n[Minor]\n\n4. I am curious does the same results hold for manipulation as well where other factors can come into play i.e. occlusions, visual fidelity of the embodiment itself i.e. hands considering the current embodiment is a synthetic one. \n\n[1] Xie et al. Vid2Sim: Realistic and Interactive Simulation from Video for Urban Navigation, CVP 2025\n[2] Yu et al. Real2Render2Real Scaling Robotic Manipulation Data Without Dynamics Simulation or Robot Hardware, CORL 2025 Oral\n[3] Chablani et al. EmbodiedSplat: Personalized Real-to-Sim-to-Real Navigation with Gaussian Splats from a Mobile Device, ICCV 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942856253,"tcdate":1762632258110,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23922/Reviewer_rmNs"],"signatures":["ICLR.cc/2026/Conference/Submission23922/Reviewer_rmNs"],"forum":"w1xbvA3rBk","number":4,"license":"CC BY 4.0","cdate":1762632258110,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23922/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942856253,"domain":"ICLR.cc/2026/Conference","replyto":"w1xbvA3rBk","id":"hIYqSlm4FJ","forumContent":{"TLDR":{"value":"We integrate 3D Gaussian Splatting into fast vectorized simulators, achieving photorealism at over 1M FPS on consumer GPUs and reducing the sim-to-real gap across diverse tasks."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["simulation","photoreal","robotics","real2sim","sim2real"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present a photorealistic robot simulator that integrates 3D Gaussian Splatting as a drop-in renderer within vectorized physics simulators such as IsaacGym. This enables unprecedented speed—exceeding 100,000 steps per second on consumer GPUs—while maintaining high visual fidelity, which we showcase across diverse tasks. We additionally demonstrate its applicability in a sim-to-real robotics setting. Beyond depth-based sensing, our results highlight how rich visual semantics improve navigation and decision-making, such as avoiding undesirable regions. We further showcase the ease of incorporating thousands of environments from iPhone scans, large-scale scene datasets (e.g., GrandTour, ARKit), and outputs from generative video models like Veo, enabling rapid creation of realistic training worlds. This work bridges high-throughput simulation and high-fidelity perception, advancing scalable and generalizable robot learning, and allowing researchers to benchmark their visual locomotion algorithms. All code and data will be open-sourced for the community to build upon. Videos, code, and data are available on the project website: https://gauss-gym.com"},"_bibtex":{"value":"@misc{\nescontrela2026gaussgym,\ntitle={GaussGym: An open-source real-to-sim framework for learning locomotion from pixels},\nauthor={Alejandro Escontrela and Justin Kerr and Jonas Frey and Yan Duan and Carmelo Sferrazza and Pieter Abbeel},\nyear={2026},\nurl={https://openreview.net/forum?id=w1xbvA3rBk}\n}"},"title":{"value":"GaussGym: An open-source real-to-sim framework for learning locomotion from pixels"},"pdf":{"value":"/pdf/80be68216979bbb3bf9d63b30050632ddf0471af.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"escontrela|gaussgym_an_opensource_realtosim_framework_for_learning_locomotion_from_pixels"},"authorids":{"value":["~Alejandro_Escontrela1","~Justin_Kerr1","~Jonas_Frey1","~Yan_Duan1","~Carmelo_Sferrazza1","~Pieter_Abbeel2"]},"authors":{"value":["Alejandro Escontrela","Justin Kerr","Jonas Frey","Yan Duan","Carmelo Sferrazza","Pieter Abbeel"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2504.17077v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"seo|physicsguided_and_fabricationaware_inverse_design_of_photonic_devices_using_diffusion_models"},"authorids":{"value":["","","","~Jong_Chul_Ye1",""]},"html":{"value":"https://doi.org/10.48550/arXiv.2504.17077"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2504-17077,\n  publtype={informal},\n  author={Dongjin Seo and Soobin Um and Sangbin Lee and Jong Chul Ye and Haejun Chung},\n  title={Physics-guided and fabrication-aware inverse design of photonic devices using diffusion models},\n  year={2025},\n  month={April},\n  cdate={1743465600000},\n  journal={CoRR},\n  volume={abs/2504.17077},\n  url={https://doi.org/10.48550/arXiv.2504.17077}\n}\n"},"abstract":{"value":"Designing free-form photonic devices is fundamentally challenging due to the vast number of possible geometries and the complex requirements of fabrication constraints. Traditional inverse-design approaches--whether driven by human intuition, global optimization, or adjoint-based gradient methods--often involve intricate binarization and filtering steps, while recent deep learning strategies demand prohibitively large numbers of simulations (10^5 to 10^6). To overcome these limitations, we present AdjointDiffusion, a physics-guided framework that integrates adjoint sensitivity gradients into the sampling process of diffusion models. AdjointDiffusion begins by training a diffusion network on a synthetic, fabrication-aware dataset of binary masks. During inference, we compute the adjoint gradient of a candidate structure and inject this physics-based guidance at each denoising step, steering the generative process toward high figure-of-merit (FoM) solutions without additional post-processing. We demonstrate our method on two canonical photonic design problems--a bent waveguide and a CMOS image sensor color router--and show that our method consistently outperforms state-of-the-art nonlinear optimizers (such as MMA and SLSQP) in both efficiency and manufacturability, while using orders of magnitude fewer simulations (approximately 2 x 10^2) than pure deep learning approaches (approximately 10^5 to 10^6). By eliminating complex binarization schedules and minimizing simulation overhead, AdjointDiffusion offers a streamlined, simulation-efficient, and fabrication-aware pipeline for next-generation photonic device design. Our open-source implementation is available at https://github.com/dongjin-seo2020/AdjointDiffusion."},"title":{"value":"Physics-guided and fabrication-aware inverse design of photonic devices using diffusion models"},"authors":{"value":["Dongjin Seo","Soobin Um","Sangbin Lee","Jong Chul Ye","Haejun Chung"]}},"tmdate":1772729942269,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2504-17077"],"tcdate":1772729915286,"writers":["~"],"signatures":["~Jong_Chul_Ye1"],"forum":"1eQdHR8Zge","license":"CC BY-SA 4.0","number":848647,"cdate":1743465600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772729942269,"domain":"DBLP.org","id":"1eQdHR8Zge","version":2},{"content":{"summary":{"value":"This paper proposes GESPI, a general framework for leveraging synthetic data in statistical inference with guaranteed error-rate control. The method wraps around a base inference algorithm. It runs the base method three times: (1) standard, (2) a more relaxed guardrail version using only real data, and (3) pooled real+synthetic. It then aggregates the outputs to ensure the final decision does not exceed an error-rate bound, even if the synthetic data is misaligned. Experiments span risk-controlled protein structure prediction, model comparison, out-of-distribution detection, and interpretability."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Can you provide one clear decision-making example where the guardrail version changes the outcome in a scientifically significant way?\n2. How will the other baseline perform in those tasks?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. It addresses an important and timely problem: using synthetic data safely.\n2. The proposed framework is general, applies to various inference tasks, with finite-sample guarantees providing rigor and confidence in decision-making."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The core idea resembles Synthetic-Powered Predictive Inference (SPI) (Bashari et al. 2025) with a wrapper reformulation. The conceptual advance over existing selective-use-of-synthetic-data methods is modest.\n2. The applications largely show numerical improvements but do not demonstrate meaningful decision changes attributable to GESPI. For example, in AlphaFold experiments, why should abstaining on slightly fewer residues matter biologically?\n3. Lack of baseline: each application should have more baselines. For example, the hypothesis testing should have PPI and PPI++ as baseline. Similar for other applications.\n4. The synthetic data in the experiments is unusually high-quality (e.g., AlphaFold MSAs), biasing results toward success scenarios. In practice, the synthetic data may not be of such high quality"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359241406,"tcdate":1761942032905,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14457/Reviewer_MhsE"],"signatures":["ICLR.cc/2026/Conference/Submission14457/Reviewer_MhsE"],"forum":"Ww3cow5dO3","number":3,"license":"CC BY 4.0","cdate":1761942032905,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14457/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359241406,"domain":"ICLR.cc/2026/Conference","replyto":"Ww3cow5dO3","id":"2rUWwaR5mV","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Hypothesis Testing","Multiple Hypothesis Testing","Conformal Prediction","Uncertainty Quantification","Auxiliary Data"]},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"abstract":{"value":"The rapid proliferation of high-quality synthetic data---generated by advanced AI models or collected as auxiliary data from related tasks---presents both opportunities and challenges for statistical inference. This paper introduces a GEneral Synthetic-Powered Inference (GESPI) framework that wraps around any statistical inference procedure to safely enhance sample efficiency by combining synthetic and real data. Our framework leverages high-quality synthetic data to boost statistical power, yet adaptively defaults to the standard inference method using only real data when synthetic data is of low quality. \nThe error of our method remains below a user-specified bound without any distributional assumptions on the synthetic data, and decreases as the quality of the synthetic data improves.\nThis flexibility enables seamless integration with conformal prediction, risk control, hypothesis testing, and multiple testing procedures, all without modifying the base inference method.\nWe demonstrate the benefits of our method on challenging tasks with limited labeled data, including AlphaFold protein structure prediction, and comparing large reasoning models on complex math problems."},"_bibtex":{"value":"@misc{\nbashari2025statistical,\ntitle={Statistical Inference Leveraging Synthetic Data with Distribution-Free Guarantees},\nauthor={Meshi Bashari and Yonghoon Lee and Roy Maor Lotan and Edgar Dobriban and Yaniv Romano},\nyear={2025},\nurl={https://openreview.net/forum?id=Ww3cow5dO3}\n}"},"title":{"value":"Statistical Inference Leveraging Synthetic Data with Distribution-Free Guarantees"},"pdf":{"value":"/pdf/b5cf30a2565db935a92c995597bb1a6c462eb7c6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"bashari|statistical_inference_leveraging_synthetic_data_with_distributionfree_guarantees"},"authorids":{"value":["~Meshi_Bashari1","~Yonghoon_Lee2","~Roy_Maor_Lotan1","~Edgar_Dobriban2","~Yaniv_Romano1"]},"authors":{"value":["Meshi Bashari","Yonghoon Lee","Roy Maor Lotan","Edgar Dobriban","Yaniv Romano"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2409.14393v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"tessler|maskedmimic_unified_physicsbased_character_control_through_masked_motion_inpainting"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Chen_Tessler:","~Yunrong_Guo1","https://dblp.org/search/pid/api?q=author:Ofir_Nabati:","https://dblp.org/search/pid/api?q=author:Gal_Chechik:","https://dblp.org/search/pid/api?q=author:Xue_Bin_Peng:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2409.14393"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2409-14393,\n  publtype={informal},\n  author={Chen Tessler and Yunrong Guo and Ofir Nabati and Gal Chechik and Xue Bin Peng},\n  title={MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2409.14393},\n  url={https://doi.org/10.48550/arXiv.2409.14393}\n}\n"},"abstract":{"value":"Crafting a single, versatile physics-based controller that can breathe life into interactive characters across a wide spectrum of scenarios represents an exciting frontier in character animation. An ideal controller should support diverse control modalities, such as sparse target keyframes, text instructions, and scene information. While previous works have proposed physically simulated, scene-aware control models, these systems have predominantly focused on developing controllers that each specializes in a narrow set of tasks and control modalities. This work presents MaskedMimic, a novel approach that formulates physics-based character control as a general motion inpainting problem. Our key insight is to train a single unified model to synthesize motions from partial (masked) motion descriptions, such as masked keyframes, objects, text descriptions, or any combination thereof. This is achieved by leveraging motion tracking data and designing a scalable training method that can effectively utilize diverse motion descriptions to produce coherent animations. Through this process, our approach learns a physics-based controller that provides an intuitive control interface without requiring tedious reward engineering for all behaviors of interest. The resulting controller supports a wide range of control modalities and enables seamless transitions between disparate tasks. By unifying character control through motion inpainting, MaskedMimic creates versatile virtual characters. These characters can dynamically adapt to complex scenes and compose diverse motions on demand, enabling more interactive and immersive experiences."},"title":{"value":"MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting"},"authors":{"value":["Chen Tessler","Yunrong Guo","Ofir Nabati","Gal Chechik","Xue Bin Peng"]}},"tmdate":1741236681948,"pdate":1704067200000,"tcdate":1741236673469,"writers":["~"],"signatures":["~Kelly_Guo1"],"forum":"iStYpGD4Sj","license":"CC BY-SA 4.0","number":357196,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741236681948,"domain":"DBLP.org","id":"iStYpGD4Sj","version":2},{"content":{"venue":{"value":"ACM Trans. Graph. 2024"},"venueid":{"value":"dblp.org/journals/TOG/2024"},"paperhash":{"value":"tessler|maskedmimic_unified_physicsbased_character_control_through_masked_motion_inpainting"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Chen_Tessler:","~Yunrong_Guo1","https://dblp.org/search/pid/api?q=author:Ofir_Nabati:","https://dblp.org/search/pid/api?q=author:Gal_Chechik:","https://dblp.org/search/pid/api?q=author:Xue_Bin_Peng:"]},"html":{"value":"https://doi.org/10.1145/3687951"},"_bibtex":{"value":"@article{DBLP:journals/tog/TesslerGNCP24,\n  author={Chen Tessler and Yunrong Guo and Ofir Nabati and Gal Chechik and Xue Bin Peng},\n  title={MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting},\n  year={2024},\n  month={December},\n  cdate={1733011200000},\n  journal={ACM Trans. Graph.},\n  volume={43},\n  number={6},\n  pages={209:1-209:21},\n  url={https://doi.org/10.1145/3687951}\n}\n"},"abstract":{"value":"Crafting a single, versatile physics-based controller that can breathe life into interactive characters across a wide spectrum of scenarios represents an exciting frontier in character animation. An ideal controller should support diverse control modalities, such as sparse target keyframes, text instructions, and scene information. While previous works have proposed physically simulated, scene-aware control models, these systems have predominantly focused on developing controllers that each specializes in a narrow set of tasks and control modalities. This work presents MaskedMimic, a novel approach that formulates physics-based character control as a general motion inpainting problem. Our key insight is to train a single unified model to synthesize motions from partial (masked) motion descriptions, such as masked keyframes, objects, text descriptions, or any combination thereof. This is achieved by leveraging motion tracking data and designing a scalable training method that can effectively utilize diverse motion descriptions to produce coherent animations. Through this process, our approach learns a physics-based controller that provides an intuitive control interface without requiring tedious reward engineering for all behaviors of interest. The resulting controller supports a wide range of control modalities and enables seamless transitions between disparate tasks. By unifying character control through motion inpainting, MaskedMimic creates versatile virtual characters. These characters can dynamically adapt to complex scenes and compose diverse motions on demand, enabling more interactive and immersive experiences."},"title":{"value":"MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting"},"authors":{"value":["Chen Tessler","Yunrong Guo","Ofir Nabati","Gal Chechik","Xue Bin Peng"]}},"tmdate":1741236679825,"pdate":1704067200000,"tcdate":1741236673470,"writers":["~"],"signatures":["~Kelly_Guo1"],"forum":"axvb62RY8J","license":"CC BY-SA 4.0","number":357197,"cdate":1733011200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741236679825,"domain":"DBLP.org","id":"axvb62RY8J","version":2},{"content":{"summary":{"value":"This paper focuses on Personalized Federated Learning (pFL) scenarios characterized by both model heterogeneity (clients use diverse architectures like ConvNet, MLP, and ResNet-9 due to varying computational resources) and data heterogeneity (non-IID data via Dirichlet distribution or quantity-based label imbalance). The core framework, MH-pFedHNDD, leverages a server-maintained hypernetwork to generate client-specific model parameters. Key steps include: 1) Clients use their locally generated models (from the hypernetwork) and local data to distill synthetic datasets; 2) The server aggregates all clients’ synthetic data into a global synthetic dataset and distributes it back to clients; 3) Clients train their personalized models using both local data and the global synthetic data, then upload parameter updates to the server; 4) The server optimizes the hypernetwork parameters and clients’ customized embedding vectors to refine future personalized model generation. The method aims to balance generalization (via global synthetic data) and personalization (via local data) while accommodating model heterogeneity."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Regarding theoretical support for L_CC and Reg Loss: Have you conducted any theoretical analysis to formalize why these terms balance generalization and personalization? If not, could you provide empirical evidence to strengthen this link?\nOn privacy risks of synthetic data: Have you evaluated privacy leaks from synthetic data using standard FL attacks (e.g., membership inference, where an attacker infers if a sample was in a client’s local data)? How does the privacy risk of sharing synthetic data compare to sharing gradients—e.g.      , using metrics like membership advantage or privacy loss (ε) for DP?\nAbout resource overhead quantification: Could you provide quantitative results on (a) the time to distill synthetic data per client (vs. local training time), (b) the size of synthetic data uploads (vs. gradient uploads for baselines like FedAvg), and (c) memory usage for distillation on edge devices (e.g., mobile GPUs)? This would clarify the method’s practicality for resource-constrained clients.\nRegarding scenario suitability: Why does MH-pFedHNDD show limited gains on CIFAR-10 (10 classes) but significant gains on CIFAR-100/Tiny-ImageNet (100/200 classes)?  Is the method inherently better suited for fine-grained classification (many classes) or high data heterogeneity?\nFor hypernetwork parameter truncation: The paper uses truncation (θ_i = h(v_i;φ)_[1:K_i]) to support model heterogeneity, but why is this necessary?       Would a hypernetwork with architecture-specific heads (instead of truncation) achieve similar or better performance? Have you compared these two approaches to justify truncation’s value?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Minimal scenario assumptions: MH-pFedHNDD operates effectively under realistic constraints—supporting heterogeneous computing resources (clients use diverse architectures like MLP, LeNet, ResNet-9) and heterogeneous data (Dirichlet-distributed and quantity-based label-imbalanced non-IID settings). This makes it more applicable to real-world edge devices (e.g., phones, IoT devices) than methods limited to homogeneous models.\nInteresting integration of hypernetworks and distillation: The core insight—using hypernetworks for model heterogeneity while distilling synthetic data to share global knowledge—is novel. It fills a gap left by prior works: hypernetwork-based pFL methods (e.g., pFedHN) ignore data-driven signals, while federated distillation methods (e.g., FedDM) overlook model structure, and this integration addresses both heterogeneities.\nComprehensive experimental validation: Experiments cover critical dimensions to support claims: (a) Two non-IID settings (Dirichlet, quantity-based label imbalance); (b) Homogeneous/heterogeneous model architectures; (c) Generalization to held-out clients; (d) Sensitivity to IPC (images per class) of synthetic data; (e) Ablation of Contrastive Condensation Loss (L_CC) and Universum Negatives (showing explicit accuracy drops when removed); (f) Differential Privacy (DP) compatibility (only ~1% accuracy loss); (g) “Plug-and-play” gains (up to +13% accuracy) when integrating distillation into baselines; (h) Scalability to 200 clients. These results collectively confirm the method’s robustness.\nClear and accessible presentation: The paper follows a logical structure (Introduction → Related Works → Method → Experiments → Conclusion) that unfolds the framework coherently. Key components (hypernetwork backbone, synthetic data generation, local training with regularization) are explained in detail, and figures (e.g., workflow in Figure 1, feature distribution in Figure 2) effectively illustrate core mechanisms. Experiments are well-organized, making it easy to follow how results support claims.\nIncremental contributions to model-heterogeneous pFL: It is among the first to integrate hypernetworks with data distillation for pFL, addressing both model and data heterogeneity from a data-driven perspective. The design of L_CC and Reg Loss (with Universum Negatives) also provides practical tools to enhance synthetic data quality and local model generalization, filling gaps in existing heterogeneous pFL methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weak theoretical support: The paper relies solely on visualizations (e.g., Figure 2) to justify L_CC and Reg Loss, with no formal analysis of why these regularization terms balance generalization and personalization. There is no proof that L_CC-induced compact synthetic features reduce cross-client distribution shift, nor why Universum Negatives in Reg Loss enhance inter-class discriminability for personalized models. This lack of theory makes it hard to generalize the method to new datasets or architectures—undermining methodological soundness.\nInadequate privacy risk assessment: A core innovation of MH-pFedHNDD is sharing synthetic data (instead of gradients), but the paper provides no privacy analysis. There is no threat model (e.g., membership inference, attribute inference) or empirical testing to measure if synthetic data leaks more local information than gradient updates. Without this, practitioners cannot assess its suitability for privacy-sensitive domains (e.g., healthcare)—a critical oversight for federated systems.\nUnquantified resource overhead: Distilling synthetic data (3000 local iterations per client) and uploading synthetic datasets introduce computational and communication costs, but these are not measured. The paper does not compare: (a) Distillation time vs. local training time; (b) Synthetic data upload size vs. gradient upload size (a key FL efficiency metric); (c) GPU/CPU memory usage for distillation on edge devices. This omits practicality checks for resource-constrained clients, weakening the method’s real-world applicability.\nLack of scenario analysis: The method shows modest gains on simple datasets (e.g., CIFAR-10, 10 classes) but significant gains on complex datasets (e.g., CIFAR-100, 100 classes; Tiny-ImageNet, 200 classes). However, the paper does not explain why this discrepancy exists—e.g., whether synthetic data adds more value for fine-grained classification (many classes) or high data heterogeneity. No explicit discussion of ideal application scenarios limits its utility for practitioners.\nWeak module coupling: The three core modules—(a) hypernetwork parameter truncation (for model heterogeneity), (b) synthetic data aggregation/distribution, (c) hypernetwork co-training—operate nearly independently. Removing any one module (e.g., truncation, synthetic data) leaves a functional (though less effective) method, and there is no evidence of synergistic (“1+1>2”) benefits. For example, it does not show that hypernetwork-generated models produce better synthetic data than standalone models, nor that synthetic data improves hypernetwork parameter generation beyond local updates alone. This reduces the method’s coherence and weakens claims about its integrated design.\nScattered notation definitions: Key notations (e.g., client embedding vector v_i, model parameter count K_i, Reg Loss L_Reg) are scattered across Sections 3.3–3.5, requiring readers to cross-reference to understand their roles. This minor flaw hinders readability, even though the overall presentation is clear.\nLimited transformative innovation: Hypernetworks and data distillation are existing techniques, and their combination does not introduce a paradigm shift in pFL. There are no new principles for balancing heterogeneity and collaboration—only a practical integration of existing tools. This limits the contribution’s impact compared to works that introduce novel theoretical or methodological paradigms."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925496509,"tcdate":1761915971123,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15193/Reviewer_S4AY"],"signatures":["ICLR.cc/2026/Conference/Submission15193/Reviewer_S4AY"],"forum":"etHBN33ny5","number":3,"license":"CC BY 4.0","cdate":1761915971123,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15193/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925496509,"domain":"ICLR.cc/2026/Conference","replyto":"etHBN33ny5","id":"KPW2ln3ujf","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We are the first to propose a data-driven perspective  for model heterogeneity in pFL. Our MH-pFedHNDD is the first to integrate data distillation into the pFL hypernetwork as well as better balance personalization and generalization."},"keywords":{"value":["Hypernetwork","Personalized Federated Learning","Data Distillation"]},"supplementary_material":{"value":"/attachment/a4ade09ebb19aed96027b73ca9f0656986d29aa0.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Personalized federated learning (pFL) aims to provide each client with a customized model based on global knowledge. However, in highly heterogeneous scenarios, pFL often struggles to obtain effective global information and faces a trade-off between personalization and generalization, which can degrade overall generalization performance.\nTo address this issue, we propose a **M**odel-**H**eterogeneous **p**ersonalized **Fed**erated learning framework based on **H**yper**N**etworks with **D**ata **D**istillation, **MH-pFedHNDD**, which, for the first time, incorporates data distillation into a hypernetwork-based federated learning framework, introducing a data-driven perspective to tackle this problem.\nWe design two effective regularization terms:\n(1) Contrastive Condensation Loss, which encourages the latent embeddings of synthetic data to be more compact and closely aligned with the local data of clients used as anchors;\n(2) Reg Loss, which integrates the latent embeddings of all clients’ synthetic data as anchors to guide the optimization direction for generalization, thereby enhancing each client’s personalized optimization performance on its local data along with the use of universum negatives.\nBy leveraging synthetic data distilled with more robust global information, our method enhances local training on clients, is the first to alleviate the imbalance between commonality and personalization for hypernetworks, and improves the performance and generalization of the hypernetwork. \nExtensive experiments under various settings demonstrate the effectiveness of our MH-pFedHNDD in personalized federated learning. Our code is available at \\url{https://anonymous.4open.science/r/MH-pFedHNDD}."},"_bibtex":{"value":"@misc{\nzhang2025exploring,\ntitle={Exploring Hypernetwork to Enhance Model Heterogeneous Personalized Federated Learning with Data Distillation},\nauthor={Chen Zhang and Husheng Li and Xiang Liu and Linshan Jiang and Danxin Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=etHBN33ny5}\n}"},"title":{"value":"Exploring Hypernetwork to Enhance Model Heterogeneous Personalized Federated Learning with Data Distillation"},"pdf":{"value":"/pdf/567b156e565ec94fa34cece2acb80d0c3bf957af.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|exploring_hypernetwork_to_enhance_model_heterogeneous_personalized_federated_learning_with_data_distillation"},"authorids":{"value":["~Chen_Zhang27","~Husheng_Li4","~Xiang_Liu15","~Linshan_Jiang1","~Danxin_Wang1"]},"authors":{"value":["Chen Zhang","Husheng Li","Xiang Liu","Linshan Jiang","Danxin Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Morpheus, a new benchmark designed to evaluate the physical reasoning capabilities of video generative models (VGMs). The authors argue that current evaluation methods, which rely on subjective human judgment or simple trajectory matching, are insufficient for rigorously assessing physical plausibility. Morpheus consists of a dataset of real-world videos capturing nine core physical phenomena governed by Newtonian mechanics (e.g., falling objects, projectile motion, pendulums). The proposed methodology extracts object trajectories from both real and generated videos and assesses their physical plausibility using two novel, physics-informed metrics: a \"Dynamical Score,\" which measures adherence to the governing equations of motion via Physics-Informed Neural Networks (PINNs), and a \"Physical Invariance Score,\" which quantifies the conservation of physical quantities like energy and momentum. The authors evaluate several state-of-the-art VGMs and find that, despite generating aesthetically pleasing videos, they struggle to adhere to these fundamental physical principles, highlighting a significant gap in their world modeling capabilities."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.Could the authors elaborate on the decision to use Depth Anything V2 for checking depth consistency in videos? Were video-specific depth estimation models considered, and would they potentially offer more temporally consistent results that could benefit the analysis?\n2.The discard rates for some models are notably high (e.g., 47% for PyramidalFlow in single-frame mode). This suggests a high rate of catastrophic failures before any fine-grained physics analysis can even be performed. Should this high rate of \"unevaluable\" generations itself be considered a primary metric for physical reasoning, perhaps weighted more heavily in the final assessment of a model's capabilities?\n3.How robust is the evaluation pipeline to potential errors from the upstream object segmentation and tracking modules? Since the entire calculation of velocity and acceleration depends on the accuracy of the extracted centroids from SAM-2, even small tracking errors could be amplified into large errors in the final physics scores. Was any sensitivity analysis performed on this?\n4.The paper states that the benchmark is restricted to Newtonian physics and assumes negligible friction and air resistance. However, these forces are always present in real-world videos. How does the framework handle a generated video that plausibly models friction (e.g., an object slowing down correctly), which would violate the ideal conservation of energy assumption? Could this lead to physically realistic videos being unfairly penalized?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"The authors have included an ethics statement. The work is focused on benchmarking publicly available models using a newly collected dataset of inanimate objects in controlled experiments. The research does not involve sensitive data, human subjects, or any apparent negative societal impacts. I have no ethical concerns with this paper."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The main strength of this work is its shift from qualitative, subjective assessments to a quantitative, physics-informed evaluation framework. By grounding the metrics in fundamental conservation laws and equations of motion, the benchmark provides a more objective and rigorous way to measure the physical reasoning of VGMs. This is a significant step forward for the field. The creation of a new dataset based on controlled laboratory experiments is a valuable contribution. This controlled setting allows for the targeted evaluation of specific physical principles and ensures that the comparisons between models are fair and reproducible, which is often a challenge with in-the-wild video data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.Limited Scope of Physical Phenomena: The benchmark's reliance on trajectory extraction from rigid bodies inherently limits its scope. Many important real-world physical phenomena do not involve a clearly trackable rigid object, such as fluid dynamics (e.g., water flowing, smoke rising), thermodynamics (e.g., ice melting), non-rigid body dynamics, or events like explosions. While the focus on Newtonian mechanics is a reasonable starting point, the claim of \"benchmarking physical reasoning\" is very broad, whereas the method is confined to a relatively narrow fundamental subset of physics.\n2.Simplicity of a Majority of the Physical Scenarios: Most of the experiments focus on the motion of a single simple object (e.g., a falling ball, a sliding book). While these are excellent for isolating variables, they may be too \"trivial\" for the rapid pace of VGM development. State-of-the-art models may quickly learn to master these simple scenarios, potentially limiting the long-term utility of the benchmark. The benchmark could be strengthened by including more complex multi-object interactions or scenarios where secondary physical effects (e.g., air resistance, complex friction) are more prominent.\n3.Lack of Correlation with Human Judgment: The paper rightly critiques the subjectivity of human evaluation. However, it does not provide any analysis of how the Morpheus scores correlate with human perception of physical plausibility. A strong benchmark should ideally produce rankings that are consistent with human intuition. Without a correlation study, it is difficult to know if the proposed metrics might sometimes penalize generations that are perceptually plausible (e.g., due to unmodeled physics like friction) or reward generations that are mathematically sound but visually awkward. This comparison is crucial for validating that the metrics are capturing meaningful aspects of physical realism.\nConstraints on Evaluation Modality: The proposed method is best suited for image-to-video or video-to-video generation, where the initial state (position, and sometimes velocity) of the object is clearly defined by the conditioning frame(s). This makes it more challenging to fairly evaluate pure text-to-video models, where the model generates the initial state from scratch, introducing ambiguity that the evaluation framework is not designed to handle. This limits the benchmark's applicability to a subset of VGMs."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924363536,"tcdate":1761895301234,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13841/Reviewer_MqUN"],"signatures":["ICLR.cc/2026/Conference/Submission13841/Reviewer_MqUN"],"forum":"1E6pburMKc","number":2,"license":"CC BY 4.0","cdate":1761895301234,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13841/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924363536,"domain":"ICLR.cc/2026/Conference","replyto":"1E6pburMKc","id":"9hckciQ1SB","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video generation","physical realism"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in image and video generation raise hopes that these models possess world modeling capabilities—the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous driving, and scientific simulation. However, before treating these models as world models, we must ask: Do they adhere to physical laws?  Current evaluation methods rely on subjective judgments or trajectory matching, limiting their usage for physical reasoning estimation, where many generations could be physically plausible. \nThus, we introduce **Morpheus**, a new benchmark for evaluating video generation models on physical reasoning. It features 130 real-world videos capturing physical phenomena, guided by conservation laws. Using those as conditioning for video generation, we assess physical plausibility using physics-informed metrics evaluated with respect to infallible conservation laws known per physical setting, leveraging advances in physics-informed neural networks and vision-language foundation models.\n% Since artificial generations lack ground truth, we assess physical plausibility using physics-informed metrics evaluated with respect to infallible conservation laws known per physical setting, leveraging advances in physics-informed neural networks and vision-language foundation models. \nOur findings reveal that even with advanced prompting and video conditioning, current models struggle to encode physical principles despite generating aesthetically pleasing videos."},"_bibtex":{"value":"@misc{\nzhang2026benchmarking,\ntitle={Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments},\nauthor={Chenyu Zhang and Daniil Cherniavskii and Antonios Tragoudaras and Antonios Vozikis and Thijmen Nijdam and Derck W. E. Prinzhorn and Mark Bodracska and Nicu Sebe and Andrii Zadaianchuk and Stratis Gavves},\nyear={2026},\nurl={https://openreview.net/forum?id=1E6pburMKc}\n}"},"title":{"value":"Benchmarking Physical Reasoning of Video Generative Models with Real Physical Experiments"},"pdf":{"value":"/pdf/e73a054680d0548d6cd3f77e26615d63b30d45b0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|benchmarking_physical_reasoning_of_video_generative_models_with_real_physical_experiments"},"authorids":{"value":["~Chenyu_Zhang4","~Daniil_Cherniavskii1","~Antonios_Tragoudaras1","~Antonios_Vozikis1","~Thijmen_Nijdam1","~Derck_W._E._Prinzhorn1","~Mark_Bodracska1","~Nicu_Sebe1","~Andrii_Zadaianchuk1","~Stratis_Gavves1"]},"authors":{"value":["Chenyu Zhang","Daniil Cherniavskii","Antonios Tragoudaras","Antonios Vozikis","Thijmen Nijdam","Derck W. E. Prinzhorn","Mark Bodracska","Nicu Sebe","Andrii Zadaianchuk","Stratis Gavves"]}},"version":2},{"content":{"summary":{"value":"The paper introduces **STAN**, a reinforcement‑learning policy for satellite–debris collision avoidance under continuous low‑thrust. The key architectural idea is a Spatio‑Temporal Attention (ST‑Attention) layer that injects a physics‑motivated bias—computed from each object’s distance of closest approach (DCA) and time to closest approach (TCA)—into scaled dot‑product attention over a variable‑sized set of debris. STAN outputs both a thrust vector and a termination signal; training uses PPO across four scenario families (single, strict multistage, probability‑based multistage, complex). The manuscript claims large improvements in reward, fuel, and collision probability versus FC/CNN/LSTM baselines, with ablations for attention and termination."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Every item raised in the Weaknesses section can be viewed as a question for the authors.\nI may well be mistaken on several of these points, and I would sincerely appreciate clarification or correction wherever appropriate.\nIf the authors can address or resolve even part of these concerns—whether by showing that I misunderstood something or by providing additional detail—it would be very helpful."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"I think the problem choice is timely and practically important: multistage conjunction handling with **continuous** low-thrust and variable-length inputs is closer to real electric-propulsion operations than impulsive-burn abstractions. The physics-aware inductive bias—folding DCA/TCA into attention—offers an intuitive way to prioritize likely threats without heavy modeling; the mechanism is clearly written with $b_i=-\\sum_j\\gamma_j\\Phi_{ij}$ and $W=\\mathrm{softmax}(QK^\\top/\\sqrt{D}+\\lambda B)$. Conceptually, the **termination head** is a good interface for fuel economy in continuous-thrust settings. I also appreciated the scenario taxonomy and visuals, and that some PPO hyperparameters and environment details are spelled out. In terms of significance, a robust policy for variable-$N$ debris could matter for on-board autonomy. And in terms of clarity, the high-level pipeline and figures communicate the intent well, even if some equations/indices need tightening."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1) Termination gating appears inverted and non-differentiable.**  \nEquation (7) computes the final action as\n$ \\Delta v_t^{\\text{final}} = H(p_{\\text{done}}-0.5)\\,[\\alpha S(t)+(1-\\alpha)a] $,\nwhich *enables* thrust when the model believes the task is “done,” contradicting the prose (“stop thrusting when appropriate”) and discarding the continuous probability via a hard Heaviside. I strongly recommend flipping the logic (e.g., multiply by $1-p_{\\text{done}}$ or use a smooth sigmoid gate) and **re-running all experiments**.\n\n**2) ST-Attention bias broadcasting is ambiguous and may be ineffective.**  \nYou define a per-debris bias $b_i$ from $\\Phi_i=(d_i,t_i)$ and then state “each row is identical ($B_{ij}=b_j$)” before adding $\\lambda B$ to the logits. If $B_{i*}$ is row-constant, adding the same constant to a row cancels in the row-wise softmax; if it’s column-constant (per-key), the bias is global and not pairwise. Please pin down the tensor shapes and show an ablation that the bias actually changes attention maps and outcomes.\n\n**3) Internal inconsistencies weaken trust.**  \nThe manuscript alternates between a collision-probability threshold of $0.02$ vs $0.002$; the ablation-section narrative claims rewards “around $750$” in the 10-debris probabilistic setting while the table lists $\\approx 276$; scenario size is “up to $116$” debris in one place and $1178$ elsewhere. Some tables also show FC outperforming STAN on reward (and reporting implausibly tiny probabilities) despite the text claiming consistent dominance. Please take a look at these items.\n\n**4) Collision-probability ($P_c$) modeling is under-specified and numerically suspect.**  \nThe analytic form lacks the covariance definitions $\\sigma_R,\\sigma_S,\\sigma_W,\\sigma_{SW}$ and how they are estimated/propagated. Reported values span $10^{-34}$–$10^{-1}$, which looks implausible without a careful uncertainty model and numerically stable evaluation. Safety claims based on $P_c$ are hard to interpret or reproduce.\n\n**5) Baselines do not represent current best set/attention policies (and may be unfair).**  \nGiven STAN’s set-attention core, comparisons against FC/CNN/LSTM are not enough. Please include capacity-matched **vanilla Transformer/Set Transformer** and consider a simple **model-based** low-thrust planner, with clearly documented padding/masking for permutation robustness.\n\n**6) Physics realism gap.**  \nEverything uses a two-body model, yet “complex” scenarios reference real-world (e.g., CelesTrak) data. Over a 1.2-hour window some regimes may be fine, but ignoring J2/drag (especially in LEO) and omitting the TLE→state pipeline (e.g., SGP4) risks mischaracterizing DCA/TCA and $P_c$.\n\n**7) Methodology and statistics need tightening.**  \nResults appear single-seed; no mean ± std or clear train/test splits. The “strict multistage” scenario is policy-dependent (training continues until success, then the next collision is injected on the updated trajectory), which risks leakage between training and evaluation. Runtime scaling claims should be backed by matched FLOPs/wall-clock under identical widths/batching.\n\n**8) Reproducibility and clarity gaps.**  \nKey actor sizes, number of heads/embedding $D$, normalization/dropout, and values/schedules for $\\lambda,\\gamma,\\alpha$ are missing. Orbit deviation is defined on normalized elements but table captions/values sometimes read like kilometers."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920421883,"tcdate":1761156192419,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8566/Reviewer_gUnw"],"signatures":["ICLR.cc/2026/Conference/Submission8566/Reviewer_gUnw"],"forum":"aZs6DkGM2I","number":1,"license":"CC BY 4.0","cdate":1761156192419,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8566/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920421883,"domain":"ICLR.cc/2026/Conference","replyto":"aZs6DkGM2I","id":"LlXAOBioo2","forumContent":{"TLDR":{"value":"We propose STAN, a spatio-temporal attention network for real-time low-thrust avoidance of multistage space debris collisions using deep reinforcement learning."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Space Debris Collision Avoidance","Deep Reinforcement Learning","Spatio-Temporal Attention","Continuous Low-Thrust Control","Policy Network Design"]},"supplementary_material":{"value":"/attachment/1a3a7b0f9fbd5927ca9f72aeab75a707522e3ea9.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"The rapid expansion of space missions has led to an exponential increase in space debris, posing severe threats to spacecraft. Existing approaches struggle to handle multistage collision risks in cluttered orbital environments, and the use of continuous low-thrust propulsion further complicates avoidance planning. To address these challenges, we propose the Spatio-Temporal Attention Network (**STAN**), which employs novel Spatio-Temporal Attention (**ST-Attention**) layers in place of conventional attention mechanisms. STAN encodes satellite-debris pairs and integrates time and distance into attention weight computation, enabling the model to generate context-aware low-thrust maneuvers. The model is trained using deep reinforcement learning across four representative multistage collision scenarios, jointly optimizing collision probability, fuel consumption, and orbital deviation. Experimental results show that STAN outperforms baseline methods in safety performance, fuel efficiency, and orbit preservation."},"_bibtex":{"value":"@misc{\nyang2026stan,\ntitle={{STAN}: A Spatio-Temporal Attention Network for Space Debris Multistage Collision Avoidance},\nauthor={Liwen Yang and Xue Bai and Xiaoyi Wang and Ming Xu},\nyear={2026},\nurl={https://openreview.net/forum?id=aZs6DkGM2I}\n}"},"title":{"value":"STAN: A Spatio-Temporal Attention Network for Space Debris Multistage Collision Avoidance"},"pdf":{"value":"/pdf/706d4e60cd5e28f3485ea7f7a132a3dc192f4593.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yang|stan_a_spatiotemporal_attention_network_for_space_debris_multistage_collision_avoidance"},"authorids":{"value":["~Liwen_Yang1","~Xue_Bai3","~Xiaoyi_Wang10","~Ming_Xu14"]},"authors":{"value":["Liwen Yang","Xue Bai","Xiaoyi Wang","Ming Xu"]}},"version":2},{"content":{"summary":{"value":"This paper  investigates the impact of AI-generated images on the performance of online continual learning (CL) models. It introduces a novel method called Entropy Selection with Real-synthetic similarity Maximization (ESRM) to mitigate the negative effects of synthetic data contamination. ESRM leverages entropy-based sample selection and a contrastive learning approach to align the feature embeddings of real and synthetic data, thereby enhancing the robustness of CL models against the degradation caused by synthetic data."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please refer to the weaknesses above."},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The authors clearly articulate the problem, methodology, and results, enhancing the paper's accessibility and understanding for a broad readership.\n\n- The paper uniquely identifies the issue of synthetic data contamination in online continual learning, a significant challenge for the future of this field. The work has substantial implications for the ML community, offering a pioneering approach to maintaining the integrity of continual learning models in the presence of synthetic data.\n\n- This paper proposes ESRM, an innovative method that combines entropy selection and contrastive learning to mitigate the negative effects of synthetic data, demonstrating creativity in addressing this new problem.\n\n- The paper is underpinned by robust technical approaches, ensuring the quality and reliability of the proposed solution through comprehensive experimental validation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper evaluates the impact of synthetic data using a limited set of generative models. And the experiments are primarily conducted on image classification datasets.  Expanding this to include a broader range of tasks, particularly truly text-to-images, text-to-text, could strengthen the findings.\n\n- The method for generating synthetic datasets is straightforward, using simple prompts. Incorporating more complex and diverse prompts, potentially using large language models to simulate user queries, could better reflect real-world synthetic data contamination.\n\n- The reliance on entropy as a key metric for distinguishing real from synthetic data might not be universally applicable. Further exploration of this metric's effectiveness across different domains and its theoretical underpinnings could bolster the method's credibility.\n\n- While the paper provides insights into the method's effectiveness, there is limited discussion on its computational efficiency and scalability, especially when dealing with large-scale datasets or high contamination ratios. Addressing these aspects could be crucial for practical applications where computational resources and time are critical constraints."},"limitations":{"value":"The authors have discussed the limitations in the paper."}},"nonreaders":[],"tmdate":1730878981464,"tcdate":1719588963535,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission5166/Reviewer_h2ny"],"signatures":["NeurIPS.cc/2024/Conference/Submission5166/Reviewer_h2ny"],"forum":"Lc8gemv97Y","number":1,"license":"CC BY 4.0","cdate":1719588963535,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission5166/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878981464,"domain":"NeurIPS.cc/2024/Conference","replyto":"Lc8gemv97Y","id":"kRSlgzx7OE","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"Investigating the dataset contamination caused by synthetic data in Online Continual Learning, and proposing a method to alleviate the performance degradation with entropy selection and real-synthetic similarity maximization."},"keywords":{"value":["Online Continual Learning","Image Generation","Replay-based method","Entropy Selection"]},"supplementary_material":{"value":"/attachment/1f7335eb5371418c8bc715b2c98b548a425b3b5b.zip"},"primary_area":{"value":"online_learning"},"abstract":{"value":"Image generation has shown remarkable results in generating high-fidelity realistic images, in particular with the advancement of diffusion-based models. However, the prevalence of AI-generated images may have side effects for the machine learning community that are not clearly identified. Meanwhile, the success of deep learning in computer vision is driven by the massive dataset collected on the Internet. The extensive quantity of synthetic data being added to the Internet would become an obstacle for future researchers to collect \"clean\" datasets without AI-generated content. Prior research has shown that using datasets contaminated by synthetic images may result in performance degradation when used for training. In this paper, we investigate the potential impact of contaminated datasets on Online Continual Learning (CL) research. We experimentally show that contaminated datasets might hinder the training of existing online CL methods. Also, we propose Entropy Selection with Real-synthetic similarity Maximization (ESRM), a method to alleviate the performance deterioration caused by synthetic images when training online CL models. Experiments show that our method can significantly alleviate performance deterioration, especially when the contamination is severe. For reproducibility, the source code of our work is available at https://github.com/maorong-wang/ESRM."},"_bibtex":{"value":"@inproceedings{\nwang2024dealing,\ntitle={Dealing with Synthetic Data Contamination in Online Continual Learning},\nauthor={Maorong Wang and Nicolas Michel and Jiafeng Mao and Toshihiko Yamasaki},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=Lc8gemv97Y}\n}"},"title":{"value":"Dealing with Synthetic Data Contamination in Online Continual Learning"},"pdf":{"value":"/pdf/eeeaa1a535b3be8d4d63b515ebc76e0021839555.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"wang|dealing_with_synthetic_data_contamination_in_online_continual_learning"},"authorids":{"value":["~Maorong_Wang1","~Nicolas_Michel1","~Jiafeng_Mao1","~Toshihiko_Yamasaki1"]},"authors":{"value":["Maorong Wang","Nicolas Michel","Jiafeng Mao","Toshihiko Yamasaki"]}},"version":2},{"content":{"venue":{"value":"Vis. Comput. 2010"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s00371-009-0396-3.pdf"},"venueid":{"value":"dblp.org/journals/VC/2010"},"paperhash":{"value":"málková|an_intuitive_polygon_morphing"},"authorids":{"value":["","","","~Bedrich_Benes1"]},"html":{"value":"https://doi.org/10.1007/s00371-009-0396-3"},"_bibtex":{"value":"@article{DBLP:journals/vc/MalkovaPKB10,\n  author={Martina Málková and Jindrich Parus and Ivana Kolingerová and Bedrich Benes},\n  title={An intuitive polygon morphing},\n  year={2010},\n  cdate={1262304000000},\n  journal={Vis. Comput.},\n  volume={26},\n  number={3},\n  pages={205-215},\n  url={https://doi.org/10.1007/s00371-009-0396-3}\n}\n"},"abstract":{"value":"We present a new algorithm for morphing simple polygons that is inspired by growing forms in nature. While previous algorithms require user-assisted definition of complicated correspondences between the morphing objects, our algorithm defines the correspondence by overlapping the input polygons. Once the morphing of one object into another is defined, very little or no user interaction is necessary to achieve intuitive results. Our algorithm is suitable namely for growth-like morphing. We present the basic algorithm and its three variations. One of them is suitable mainly for convex polygons, the other two are for more complex polygons, such as curved or spiral polygonal forms."},"title":{"value":"An intuitive polygon morphing"},"authors":{"value":["Martina Málková","Jindrich Parus","Ivana Kolingerová","Bedrich Benes"]}},"tmdate":1772212998000,"pdate":1293753600000,"externalIds":["dblp:journals/vc/MalkovaPKB10"],"tcdate":1772212926152,"writers":["~"],"signatures":["~Bedrich_Benes1"],"forum":"AnfNgee7hH","license":"CC BY-SA 4.0","number":834085,"cdate":1262304000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772212998000,"domain":"DBLP.org","id":"AnfNgee7hH","version":2},{"content":{"summary":{"value":"The paper addresses the problem of catastrophic forgetting in neural networks during continual learning. The paper aims to find optimal task-selection protocols to minimize forgetting and maximize performance, moving beyond heuristic approaches. This is done through combining statistical physics with optimal control theory to derive closed-form formulas for training dynamics.\n\n**Contributions:**\n- Derive optimal task-selection protocols and learning rate schedules as functions of task similarity and problem parameters.\n- Proposes a \"pseudo-optimal\" replay strategy with distinct focus and revision phases to minimize forgetting."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- In many results, the pseudo-optimal method is comparable in performance to the naive interleaved. Could you better explain the advantages of using this pseudo-optimal method instead of a naive interleave?\n- The \"pseudo-optimal\" strategy provides practical insights, but the underlying mechanics of the optimal task-selection structure could be further explained to make the protocol more interpretable.\n- (see weaknesses for more suggestions)"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"- Figures are clear, well explained, and easy to understand.\n- The paper uses an original and refreshing approach with strong theoretical backing.\n- Demonstrates robust alignment between theoretical predictions and experimental results.\n- Complex mathematical formulations are well-supported."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Please try to use shorter sentences. The sentences tend to be overly long and complex.\n- The mathematics and equations are dense and could limit accessibility for readers unfamiliar with optimal control or statistical physics.\n- The methodology might benefit from additional real-world datasets or further exploration in complex architectures for broader generalizability. Additional datasets with higher complexity (e.g., CIFAR-10/100 or domain-specific continual learning datasets) would better demonstrate the generalizability of the proposed protocols.\n- Experiments primarily focus on two-task scenarios. Including more tasks would test the robustness and scalability of the approach, offering more insight into how it handles diverse, real-world continual learning settings.\n- Although the paper suggests a posteriori interpretations, actionable guidelines on why specific phases are structured as they are (e.g., conditions where certain strategies excel) would improve applicability and user understanding.\n- Theoretical findings rely on idealized assumptions (e.g., i.i.d. Gaussian inputs, simplified two-layer networks) which may not hold in complex, structured data environments. Expanding to other data models would strengthen claims of broader applicability.\n- The model does not address relative task difficulty or task order impact, which are critical in real-world scenarios. Future extensions could explore adaptive protocols that consider task-specific attributes dynamically."}},"nonreaders":[],"tmdate":1731427959953,"tcdate":1730652331750,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3852/Reviewer_pAPx"],"signatures":["ICLR.cc/2025/Conference/Submission3852/Reviewer_pAPx"],"forum":"rhhQjGj09A","number":2,"license":"CC BY 4.0","cdate":1730652331750,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3852/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427959953,"domain":"ICLR.cc/2025/Conference","replyto":"rhhQjGj09A","id":"g1zHFrdPTL","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"This theoretical work combines statistical physics and control theory to design optimal task-selection protocols mitigating catastrophic forgetting while preserving performance in continual learning, validated on synthetic and real-world data."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["machine learning theory","statistical physics","online learning","continual learning","optimal control theory"]},"supplementary_material":{"value":"/attachment/dc97de97da7cea6229777e72c9dad623a14bcda1.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Artificial neural networks often struggle with _catastrophic forgetting_ when learning multiple tasks sequentially, as training on new tasks degrades the performance on previously learned tasks. Recent theoretical work has addressed this issue by analysing learning curves in synthetic frameworks under predefined training protocols. However, these protocols relied on heuristics and lacked a solid theoretical foundation assessing their optimality. In this paper, we fill this gap by combining exact equations for training dynamics, derived using statistical physics techniques, with optimal control methods. We apply this approach to teacher-student models for continual learning and multi-task problems, obtaining a theory for task-selection protocols maximising performance while minimising forgetting. Our theoretical analysis offers non-trivial yet interpretable strategies for mitigating catastrophic forgetting, shedding light on how optimal learning protocols modulate established effects, such as the influence of task similarity on forgetting. Finally, we validate our theoretical findings with experiments on real-world data."},"_bibtex":{"value":"@inproceedings{\nmori2025optimal,\ntitle={Optimal Protocols for Continual Learning via Statistical Physics and Control Theory},\nauthor={Francesco Mori and Stefano Sarao Mannelli and Francesca Mignacco},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=rhhQjGj09A}\n}"},"title":{"value":"Optimal Protocols for Continual Learning via Statistical Physics and Control Theory"},"pdf":{"value":"/pdf/fc8906565869a5db2dff20baaae05782eff0226d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"mori|optimal_protocols_for_continual_learning_via_statistical_physics_and_control_theory"},"authorids":{"value":["~Francesco_Mori1","~Stefano_Sarao_Mannelli1","~Francesca_Mignacco1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Francesco Mori","Stefano Sarao Mannelli","Francesca Mignacco"]}},"version":2},{"content":{"summary":{"value":"This paper introduces the Physics-guided Parametric Augmentation Network (PANet), designed to improve real-world image dehazing. \n\nPANet combines physics-based modeling with data-driven techniques to generate diverse hazy images, aiming to bridge the gap between synthetic and real-world hazy datasets. \n\nBy mapping haze characteristics into a parametric space, PANet can resample parameters and generate new, physically realistic hazy images."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. For the colour shift issue, what is the reason?\n\n2. It seems for the sky, white object and road, the results are not promising, what is the reason?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- physics-guided + data-driven make sense. \n\n- The physics-guided and parametric approach to generating realistic hazy images also makes sense."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The progress of daytime dehazing or defogging has been significant over the past 10 years. These methods can handle many problems, particularly when the haze or fog is relatively thin. Non-uniform haze/fog is also not a significant issue, as many methods can handle it well. (If there is any disagreement, the paper should provide evidence of existing methods failing to deal with non-uniform haze.) The main challenge of dehazing arises when the haze/fog is significantly thick. Unfortunately, the proposed method does not address this thick haze/fog problem specifically, as evidenced by the results. Moreover, the proposed method has no specific mechanism or treatment in dealing with the thick haze/fog and its characteristics.\n\n- The qualitative experimental results do not show that the proposed method outperforms the existing methods. \nIn Fig. 1 and 6, when the fog/haze is thick, the method still suffers from it and suffers from colour shift.\n\n- The proposed method does not have any specific features that differentiate it from existing methods in terms of the haze/fog problem it aims to solve. The results presented in the paper could be achieved by existing methods, including non-deep learning methods, with comparable quality.\n\n- Missing citation and dataset in [1]\n[1] Structure Representation Network and Uncertainty Feedback Learning for Dense Non-Uniform Fog Removal"}},"nonreaders":[],"tmdate":1731427577922,"tcdate":1730740757280,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2323/Reviewer_wXTj"],"signatures":["ICLR.cc/2025/Conference/Submission2323/Reviewer_wXTj"],"forum":"YS5zdlSzvv","number":4,"license":"CC BY 4.0","cdate":1730740757280,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2323/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427577922,"domain":"ICLR.cc/2025/Conference","replyto":"YS5zdlSzvv","id":"wSITpdrD1N","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Image dehazing","Image rehazing","Data augmentation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Image dehazing faces significant challenges in real-world scenarios due to the large domain gap between synthetic and real-world hazy images, which often hinders dehazing performance. Collecting real-world datasets is particularly difficult, as hazy and clean image pairs must be captured under identical conditions. To address this, we propose a Physics-guided Parametric Augmentation Network (PANet) that generates realistic hazy and clean training pairs, enhancing dehazing performance in real-world applications. PANet consists of two components: a Haze-to-Parameter Mapper (HPM), which projects hazy images into a parametric space representing haze characteristics, and a Parameter-to-Haze Mapper (PHM), which converts resampled haze parameters back into hazy images. By resampling individual haze parameter maps at the pixel level in the parametric space, PANet generates diverse hazy images with physically explainable haze conditions that are not present in the training data. Our experimental results show that PANet effectively enriches existing hazy image benchmarks, significantly improving the performance of current dehazing models."},"_bibtex":{"value":"@misc{\ntsai2025unleashing,\ntitle={Unleashing the Power of Deep Dehazing Models: A Physics-guided Parametric Augmentation Net for Image Rehazing},\nauthor={Fu-Jen Tsai and Chih-Ling Chang and ZILING HUANG and Lin Gu and Chia-Wen Lin},\nyear={2025},\nurl={https://openreview.net/forum?id=YS5zdlSzvv}\n}"},"title":{"value":"Unleashing the Power of Deep Dehazing Models: A Physics-guided Parametric Augmentation Net for Image Rehazing"},"pdf":{"value":"/pdf/850a9c3316fcebe8c517f7e530bba83e391d1b3e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"tsai|unleashing_the_power_of_deep_dehazing_models_a_physicsguided_parametric_augmentation_net_for_image_rehazing"},"authorids":{"value":["~Fu-Jen_Tsai1","~Chih-Ling_Chang1","~ZILING_HUANG3","~Lin_Gu4","~Chia-Wen_Lin1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fu-Jen Tsai","Chih-Ling Chang","ZILING HUANG","Lin Gu","Chia-Wen Lin"]}},"version":2},{"content":{"summary":{"value":"This manuscript introduces a unified neuro-symbolic framework named Euclid-Omni, designed to solve Euclidean geometry problems. The framework combines LLMs and VLMs with a novel, versatile symbolic geometry solver called Euclidea. The Euclidea solver, which is the core of this framework, fuses logical deduction (a deductive database) with algebraic computation (an algebraic engine) to automatically generate human-readable reasoning steps, enabling it to handle both proving-style and calculation-style geometry problems. Furthermore, the manuscript constructs a comprehensive data generation pipeline. This pipeline can automatically synthesize symbolic problems, render corresponding diagrams, generate solution steps using the Euclidea solver, and finally translate this symbolic content into natural language . The resulting synthetic datasets can be used to train LLMs and VLMs."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Details see the Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The manuscript is well-structured and content-rich, constructing a complex neuro-symbolic framework that represents a significant workload. The method demonstrates high efficiency in solving Olympiad-level problems. Compared to the significant overhead of AlphaGeometry, Euclid-Omni's hybrid system achieves competitive performance using only a small number of training samples, reducing the training data by orders of magnitude."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In Section 3.1, the problem formalization relies on both \"metric relations\" and \"diagrammatic relations\". The paper states that diagrammatic relations (e.g., SameSide(a, b, c, d)) are \"implicit but can be extracted from the diagram\" . This extraction process is a critical, non-trivial step, yet its implementation is not detailed. How is this extraction performed?\n2. The algebraic system in Euclidea (Section 3.1) categorizes equations into four types. While the first three (linear and log-linear) are solved systematically via Gaussian elimination , the handling of the fourth \"complex\" category (e.g., trigonometric or higher-order polynomial relations) is vague. The paper states simplification and substitution work \"in many cases\". This raises a key question: What happens when a complex equation cannot be simplified by this method? Does the system abandon that reasoning branch, or are there fallback mechanisms?\n3. The claim of generating \"human-readable\" reasoning steps is made throughout the paper. However, the examples provided in Appendix C.1 consist of highly symbolic, step-by-step logical and algebraic derivations . While these proofs are more compact than those from PyEuclid (e.g., 13 steps vs. 32) , they are not \"human-readable\" in the sense of a prose proof and remain difficult for a non-expert to parse. How do the authors define and evaluate \"human-readability\"?\n4. In Section 4.2, the VLM is trained on a mixed dataset of 10K synthetic samples from Euclid-Omni and 10K samples from Geo170K . This introduces a significant confounding variable. The strong performance reported could be largely attributable to the existing Geo170K dataset, which the paper notes was added to \"better match the out-of-distribution diagrams present in existing benchmarks\". To properly assess the contribution of the proposed data generation pipeline, a crucial ablation study is missing. What is the performance of the VLM when trained only on the 10K (or 20K) samples generated by Euclid-Omni?\n5. The entire framework's success in training models relies on the quality of its synthetic data (Section 3.2). However, the paper does not provide a clear validation of this synthetic data's quality. While the generation pipeline is described, what measures or standards were used to assess the quality, diversity, and difficulty distribution of the resulting problems?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942203749,"tcdate":1761831959171,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22403/Reviewer_pzHX"],"signatures":["ICLR.cc/2026/Conference/Submission22403/Reviewer_pzHX"],"forum":"1GQv7jhmtV","number":2,"license":"CC BY 4.0","cdate":1761831959171,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22403/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942203749,"domain":"ICLR.cc/2026/Conference","replyto":"1GQv7jhmtV","id":"QxMUyBKDod","forumContent":{"TLDR":{"value":"We propose a unified neuro-symbolic framework that integrates our formal geometry system with LLMs and VLMs to tackle a broad range of geometric reasoning tasks."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Geometry Problem Solving","Neuro-Symbolic","LLM","VLM","Synthetic Data Generation"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Euclidean geometry presents a compelling testbed for AI reasoning capabilities, requiring seamless integration of diagram understanding, logical deduction, and algebraic computation. Existing systems have either been narrowly scoped or struggled with challenging problems. We introduce Euclid-Omni, a unified neuro-symbolic framework that combines a formal geometry system with Large (Vision)–Language Models (LLMs and VLMs) to address both calculation- and proving-style problems across formal and natural languages, up to Olympiad-level difficulty. At its core, we develop Euclidea, a versatile geometry symbolic solver that automatically generates human-readable reasoning steps through logical deduction and algebraic solving. On top of this, we implement a comprehensive data generation pipeline that synthesizes symbolic problems, renders diagrams, and translates problems into natural language, yielding large-scale, diverse datasets for training LLMs and VLMs in different reasoning settings.\nExperiments on multiple benchmarks demonstrate that Euclidea can tackle a broader range of problems than prior symbolic systems. \nOur trained VLMs achieve superior results on calculation tasks, while combining LLMs with Euclidea remains competitive with state-of-the-art systems on Olympiad-level theorem proving problems, despite using orders of magnitude less compute and data."},"_bibtex":{"value":"@misc{\nli2026euclidomni,\ntitle={Euclid-Omni: A Unified Neuro-Symbolic Framework for Geometry Problem Solving},\nauthor={Zhaoyu Li and Hangrui Bi and Youyuan Zhang and Wenjie Ma and Zenan Li and Xujie Si and Kaiyu Yang},\nyear={2026},\nurl={https://openreview.net/forum?id=1GQv7jhmtV}\n}"},"title":{"value":"Euclid-Omni: A Unified Neuro-Symbolic Framework for Geometry Problem Solving"},"pdf":{"value":"/pdf/d33d32b6465b42cebd4b18d4addd86cae0607ac5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|euclidomni_a_unified_neurosymbolic_framework_for_geometry_problem_solving"},"authorids":{"value":["~Zhaoyu_Li3","~Hangrui_Bi1","~Youyuan_Zhang1","~Wenjie_Ma1","~Zenan_Li3","~Xujie_Si1","~Kaiyu_Yang1"]},"authors":{"value":["Zhaoyu Li","Hangrui Bi","Youyuan Zhang","Wenjie Ma","Zenan Li","Xujie Si","Kaiyu Yang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a formulation for inverse problems, which the authors term the \"Ensemble Inverse Problem\" (EIP). The goal of EIP is to recover a \"truth\" distribution, $p(x)$, given a set of observations, $\\mathcal{Y}$, which are generated by feeding samples from $p(x)$ through a fixed but unknown forward model and adding noise to them to build a likelihood distribution $p(y|x)$. This problem is common in science (e.g., \"unfolding\" in particle physics) and imaging, especially when the prior $p(x)$ is unknown.\nThe authors distinguish two versions of the problem:\n1. EIP-I (Prior): Recover the prior $p(x)$ from the ensemble $\\mathcal{Y}$.\n2. EIP-II (Posterior): Recover the posterior $p(x|y)$ for a single observation $y$, using the entire ensemble $\\mathcal{Y}$ as context.\n\nThe core methodological contribution is what the authors call \"ensemble inverse generative models,\" that are trained to approximate $p(x|y, \\mathcal{Y})$. This is achieved by conditioning a generative model (e.g., DDPM or Flow Matching) on both the single measurement $y$ and a learned, permutation-invariant embedding of the entire observation set $\\mathcal{Y}$, $\\phi_w(\\mathcal{Y})$. To enable generalization, the model is trained on a collection of datasets from different priors, all passed through the same unknown forward operator. \nThe authors demonstrate through experiments on synthetic data, a 7D particle physics unfolding problem, and a synthetic image inversion task that their method achieves state-of-the-art results. They show it can successfully generalize to unseen priors that were not included in the training data, a task where naive conditional models fail."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. About the fact that the physical model seems to be required to be fixed across M: In most scientific applications I’ve seen, when the true physical model is unknown, what is available is typically a whole menu of physical simulators that approximate this underlying physical model to various degrees of fidelity (and various levels of compute cost). It’s those approximate simulators that allows building the mapping from x to y. But in that context, 1- it’s never guaranteed that the true mapping in within the set of simulations we can access, 2- the physical model would be changing between simulations and real experiment where we want to apply the trained model. This is related to point 1 in the weaknesses, can you clarify the assumptions you are working under, and if the physical model does need to be fixed, I’d like it to be explained more clearly how this is supposed to applied to a realistic real-world experimental setup.\n\n2. Similarly, it’s not clear from the paper if the proposed method would work is the different priors were over a heterogenous parameter space, meaning that $p_m(x)$ have different dimensionalities for $x$ across different $m$? This is another context where I could see this applied in a scientific context where one would like to do Bayesian model comparison/fit different model with varying number of parameters. Is that something that could be possible?  \n3.  I was surprised to see that GDDPM-v (with more information) performs worse than GDDPM (with less). This suggests potential overfitting or non-optimal training, which weakens the claim of SOTA performance. Adding information should not impair recovery if the model is well-trained, do you have an explanation for why this is not the case?\n4. Could you provide more details on how the experiments were set up? For example, for Fig 2, how was the prior obtained from observations for the conditional diffusion model and flow matching? Was it by getting the MAP or the mean of the posterior samples for multiple observations? Also, could you provide more details on how the conditioning was done for the other models for the MNIST experiment?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The formalization of the \"Ensemble Inverse Problem\" (EIP) is a significant contribution. This is essentially a new \"in-context learning\" or \"meta-learning\" framework for scientific inverse problems, which is potentially impactful. The proposed solution—conditioning on a learned, permutation-invariant embedding of the observation set $\\mathcal{Y}$—is an elegant and general way to solve the problem. Instead of using hard-coded statistics (like moments), it allows the model to learn the most relevant features of the observation set that define the prior, and does not require knowledge of the forward model. \n2. The experiments are designed to test the paper's central claim.\n    * In the synthetic 2D Gaussian and MNIST-mixture experiments, the model is explicitly trained on a disjoint set of priors and tested on the \"in-between\" region (e.g.,  interpolation between seen regions).\n    * The results in Figures 3 and 6 are unambiguous . The proposed method (EI-FM/EI-DDPM) successfully generalizes and solves the problem in the unseen regions, while the baseline (cFM/cDDPM) that lacks the ensemble information completely fails. This is a very powerful demonstration.\n3. The method achieves SOTA results on a real world HEP unfolding task, outperforming multiple baselines (including the similar GDDPM) on 4 unseen physics processes. This shows the approach scales and is practical for complex scientific data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Ambiguous Experimental Setups: A few points in the paper would need to be clarified, specifically about the experimental setups. \n    - For the HEP experiment:  My understanding was that the paper's premise is that the forward model and the noise model that allow building $p(y|x)$ need to be fixed. However, the data description mentions using \"various parton distribution functions and parton shower models\". This sounds like the underlying physics, and thus $p(y|x)$, is changing across the different datasets in M. This would really need clarification. \n\n    - For the MNIST digit mixture problem, it’s not clear what is the goal of the task. It is to recover one of the deblended digits (or both), or is it to recover the blended mixture at some given time (making it much more like an inpainting+denoising task)? From the text I understood it was the former, but Figure 6 seems to suggest the latter. This needs to be clarified in the text.   \n2. Limited baseline comparisons: The paper positions itself as a general solution for inference, but I was surprised to see that baseline comparisons were limited. While it’s true that most inference methods that have been developed with scientific applications in mind rely on knowledge of a forward model (except on naive conditioning on observation with diffusion and flow matching), with the paired dataset that is assumed to be available, it would be natural to train an emulator for the forward model, which could then be used in a standard inference framework like SBI. This is the natural baseline to compare the proposed method against (and, as I understand, a standard method for HEP applications as well). This would potentially be impactful since SBI methods are known to struggle with covariate shifts, so it’d be natural to show the strength on the proposed method against this specific baseline, and not including it is a clear lack in the paper in its current state (7dimension is very low-dimensional, definitely within the applicability of SBI methods). Moreover, once an emulator has been trained, it would be natural to compare against population-level inference frameworks to empirically adapt or learn to correct misspecified priors, such as 2402.07808, or more general expectation-maximization methods like those proposed in 2405.13712 or 2407.17667. Seeing how the proposed methods compare to these more recent proposed approaches is a natural question for a practitioner.  \n\n3. Limited Evaluation Metrics: The paper evaluates the quality of their inference using metrics like SWD, WD, MSE, and SSIM. These metrics compare bulk properties, means, or medians, and these can miss key differences between high-dimensional distributions.  Coverage tests like TARP (2302.03026) or accuracy scores like PQMass (2402.04355) would be needed to assess if the inferred posteriors/priors here accurate in terms of the entire distribution. This is especially a problem for the image-space application with MNIST, since just the MSE or the SSIM of the mean of the posterior is assessed. The mean, specifically in image space, could look nothing like any individual sample and this could affect the results specifically for structural integrity. It could also favor methods that are overconfident.  \n\n4. Unclear Scaling: The paper claims scalability, but the \"high-dim\" MNIST test is only $28 \\times 28$. A real study of how this method would scale in real high-dimensional settings (e.g., ImageNet-scale), is missing. \n\n5. Limited scientific applications: The performance of the method is demonstrated in a single scientific context (HEP), but the generals claims are that the method is useful for general scientific applications. Applications to other scientific contexts like those in the benchmark in 2503.11043 would really strengthen the paper. Other domains of applications where a forward model is generally not known are, for example, weather forecasting, climate modelling, and cosmology."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931318290,"tcdate":1762269798659,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19400/Reviewer_Ay8h"],"signatures":["ICLR.cc/2026/Conference/Submission19400/Reviewer_Ay8h"],"forum":"aVXXZAp41g","number":4,"license":"CC BY 4.0","cdate":1762269798659,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19400/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931318290,"domain":"ICLR.cc/2026/Conference","replyto":"aVXXZAp41g","id":"yQtqi84a3F","forumContent":{"TLDR":{"value":"We introduce the ensemble inverse problem and propose a posterior sampling method based on generative models to solve it."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Inverse problems","conditional generative models","posterior sampling","permutation invariant neural network"]},"supplementary_material":{"value":"/attachment/c25721b5053dc87abe458c387415015bf452f63c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We introduce a new multivariate statistical problem that we refer to as the Ensemble Inverse Problem (EIP). The aim of EIP is to invert for an ensemble that is distributed according to the pushforward of a prior under a forward process. In high energy physics (HEP), this is related to a widely known problem called unfolding, which aims to reconstruct the true physics distribution of quantities, such as momentum and angle, from measurements that are distorted by detector effects. In recent applications, the EIP also arises in inverse imaging with unknown priors. We propose non-iterative inference-time methods that construct posterior samplers based on a new class of conditional generative models, which we call  ensemble inverse generative models. For the posterior modeling, these models additionally use the ensemble information contained in the observation set on top of single measurements.  Unlike existing methods, our proposed methods avoid explicit and iterative use of the forward operator at inference time via training across several sets of truth-observation pairs that are consistent with the same forward operator, but originate from a wide range of priors. We demonstrate that this training procedure implicitly encodes the likelihood model. The use of ensemble information helps posterior inference and enables generalization to unseen priors. We benchmark the proposed method on several synthetic and real datasets in HEP and inverse imaging."},"_bibtex":{"value":"@misc{\nhuan2026the,\ntitle={The Ensemble Inverse Problem: Applications and Methods},\nauthor={Zhengyan Huan and Camila Pazos and Martin Klassen and Vincent Croft and Pierre-Hugues Beauchemin and Shuchin Aeron},\nyear={2026},\nurl={https://openreview.net/forum?id=aVXXZAp41g}\n}"},"title":{"value":"The Ensemble Inverse Problem: Applications and Methods"},"pdf":{"value":"/pdf/7efc21d8ad2fc1796d9ea8a75ce49b47e338d44b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"huan|the_ensemble_inverse_problem_applications_and_methods"},"authorids":{"value":["~Zhengyan_Huan2","~Camila_Pazos1","~Martin_Klassen1","~Vincent_Croft1","~Pierre-Hugues_Beauchemin1","~Shuchin_Aeron2"]},"authors":{"value":["Zhengyan Huan","Camila Pazos","Martin Klassen","Vincent Croft","Pierre-Hugues Beauchemin","Shuchin Aeron"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"GenCP introduces a generative paradigm for coupled physics simulation that not only offers excellent theoretical guarantee, but also provides concrete guidance for practical implementation."},"keywords":{"value":["Coupled Physics Simulation","Flow Matching","Operator Splitting","Joint Sampling"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Real-world physical systems are inherently complex, often involving the coupling of multiple physics, making their simulation both highly valuable and challenging. Many mainstream approaches face challenges when dealing with decoupled data. Besides, they also suffer from low efficiency and fidelity in strongly coupled spatio-temporal physical systems. Here we propose GenCP, a novel and elegant generative paradigm for coupled multiphysics simulation. By formulating coupled-physics modeling as a probability modeling problem, our key innovation is to integrate probability density evolution in generative modeling with iterative multiphysics coupling, thereby enabling training on data from decoupled simulation and inferring coupled physics during sampling. We also utilize operator-splitting theory in the space of probability evolution to establish error controllability guarantees for this “conditional-to-joint” sampling scheme. We evaluate our paradigm on a synthetic setting and three challenging multiphysics scenarios to demonstrate both principled insight and superior application performance of GenCP. Code is available at this repo: https://github.com/AI4Science-WestlakeU/GenCP."},"_bibtex":{"value":"@inproceedings{\ngao2026gencp,\ntitle={Gen{CP}: Towards Generative Modeling Paradigm of Coupled physics},\nauthor={Tianrun Gao and Haoren Zheng and Wenhao Deng and Haodong Feng and Tao Zhang and Ruiqi Feng and Qianyi Chen and Tailin Wu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=tn2VAi1KIO}\n}"},"title":{"value":"GenCP: Towards Generative Modeling Paradigm of Coupled physics"},"pdf":{"value":"/pdf/314cb0e8881cc1c85d631f529e0106933e8c46e1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gao|gencp_towards_generative_modeling_paradigm_of_coupled_physics"},"authorids":{"value":["~Tianrun_Gao3","~Haoren_Zheng2","~Wenhao_Deng2","~Haodong_Feng1","~Tao_Zhang56","~Ruiqi_Feng1","~Qianyi_Chen1","~Tailin_Wu1"]},"authors":{"value":["Tianrun Gao","Haoren Zheng","Wenhao Deng","Haodong Feng","Tao Zhang","Ruiqi Feng","Qianyi Chen","Tailin Wu"]}},"tmdate":1775876903400,"pdate":1769435684601,"tcdate":1757050144466,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2289/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission2289/Authors"],"forum":"tn2VAi1KIO","license":"CC BY 4.0","number":2289,"cdate":1757050144466,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission2289/-/Full_Submission","ICLR.cc/2026/Conference/Submission2289/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission2289/-/Camera_Ready_Revision"],"mdate":1775876903400,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"tn2VAi1KIO","version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Physical reasoning","video prediction"]},"supplementary_material":{"value":"/attachment/90ac022886fc42c632a139045ebfc3349691daec.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Predicting physical dynamics from raw visual data remains a major challenge in AI. While recent video generation models have achieved impressive visual quality, they still cannot consistently generate physically plausible videos due to a lack of modeling of physical laws. Recent approaches combining 3D Gaussian splatting and physics engines can produce physically plausible videos, but are hindered by high computational costs in both reconstruction and simulation, and often lack robustness in complex real-world scenarios. To address these issues, we introduce **Neural Gaussian Force Field (NGFF)**, an end-to-end neural framework that integrates 3D Gaussian perception with physics-based dynamic modeling to generate interactive, physically realistic 4D videos from multi-view RGB inputs, achieving two orders of magnitude faster than prior Gaussian simulators. To support training, we also present **GSCollision**, a 4D Gaussian dataset featuring diverse materials, multi-object interactions, and complex scenes, totaling over 640k rendered physical videos (∼4 TB). Evaluations on synthetic and real 3D scenarios show NGFF’s strong generalization and robustness in physical reasoning, advancing video prediction towards physics-grounded world models."},"_bibtex":{"value":"@inproceedings{\nli2026learning,\ntitle={Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields},\nauthor={Shiqian Li and Ruihong Shen and Junfeng Ni and Chang Pan and Chi Zhang and Yixin Zhu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KxvboPqav6}\n}"},"title":{"value":"Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields"},"pdf":{"value":"/pdf/d24aad8cb8c4f83bd8e677b5af43fdd97540be37.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|learning_physicsgrounded_4d_dynamics_with_neural_gaussian_force_fields"},"authorids":{"value":["~Shiqian_Li1","~Ruihong_Shen1","~Junfeng_Ni1","~Chang_Pan2","~Chi_Zhang12","~Yixin_Zhu1"]},"authors":{"value":["Shiqian Li","Ruihong Shen","Junfeng Ni","Chang Pan","Chi Zhang","Yixin Zhu"]}},"tmdate":1775877273485,"pdate":1769436402225,"tcdate":1758289642050,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18635/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission18635/Authors"],"forum":"KxvboPqav6","license":"CC BY 4.0","number":18635,"cdate":1758289642050,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission18635/-/Full_Submission","ICLR.cc/2026/Conference/Submission18635/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission18635/-/Camera_Ready_Revision"],"mdate":1775877273485,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"KxvboPqav6","version":2},{"content":{"summary":{"value":"The paper introduces the $r$-loopy Weisfeiler-Leman ($r$-$l$WL) test, an innovative hierarchy of graph isomorphism tests, and the corresponding GNN framework, $r$-$l$MPNN. This new approach extends the counting capabilities of previous algorithms, specifically allowing the counting of cycles up to length $r+2$ and homomorphisms of cactus graphs. Empirical validation demonstrates the expressiveness and performance of $r$-$l$MPNN on both synthetic and real-world datasets."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The paper is generally well-written and clear, although some sections could benefit from additional explanations or examples to aid understanding. I have a few questions:\n\n1. How does the computational complexity of $r$-$l$WL compared to existing higher-order WL variants like $3$-WL in practice, especially when dealing with large and dense graphs?\n\n2. What are the limitations of $r$-$l$WL in terms of scalability and memory usage, particularly when applied to real-world datasets with varying degrees of sparsity?\n\n3. In Table 1, why didn't your method perform well on the Extension (100) and CFI (100) datasets compared to 3-WL and PPGN?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"The paper's strengths are rooted in its originality, introducing a novel algorithm ($r$-$l$WL) and corresponding GNNs ($r$-$l$GIN) that significantly enhance the expressivity of graph neural networks. These contributions are supported by rigorous theoretical proofs and empirical validation. Specifically, $r$-$l$WL demonstrates the ability to count cycles up to length $r+2$ and homomorphisms of cactus graphs, substantiated with detailed mathematical proofs. The experiments use several synthetic datasets to validate the counting power and expressiveness of $r$-$l$MPNN effectively. Furthermore, the paper contextualizes its contributions within prior work, highlighting the limitations of existing methods and demonstrating how $r$-$l$WL and $r$-$l$MPNN address these gaps. Overall, the claims are well-supported by theoretical proofs and empirical results, indicating a clear improvement over existing methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Some of the mathematical proofs are complex and may be difficult for readers without a strong background in graph theory and GNNs. Providing additional intuitive explanations or examples could improve accessibility. While the empirical validation is strong, it could be expanded to include a broader range of real-world datasets to further demonstrate the robustness and generalizability of the approach."},"limitations":{"value":"The authors have addressed the limitations related to the complexity of higher-order GNNs and the scalability issues associated with $k$-WL. However, a more detailed discussion on the limitations of $r$-$l$WL in terms of computational overhead and potential impact on large-scale applications would be beneficial."}},"nonreaders":[],"tmdate":1730879883092,"tcdate":1719037526243,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission16891/Reviewer_1SHr"],"signatures":["NeurIPS.cc/2024/Conference/Submission16891/Reviewer_1SHr"],"forum":"9O2sVnEHor","number":1,"license":"CC BY 4.0","cdate":1719037526243,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission16891/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879883092,"domain":"NeurIPS.cc/2024/Conference","replyto":"9O2sVnEHor","id":"a6h9ZaTuaO","forumContent":{"TLDR":{"value":"We introduce GNNs that can count cycles and homomorphisms of cactus graphs, surpassing the limitations of existing GNNs while being scalable on real-world graphs."},"venue":{"value":"NeurIPS 2024 oral"},"keywords":{"value":["Graph Neural Networks","Weisfeiler-Leman (WL) Test","Homomorphism Counting","Theory and Expressivity in GNNs","Cactus Graphs"]},"primary_area":{"value":"graph_neural_networks"},"abstract":{"value":"We introduce $r$-loopy Weisfeiler-Leman ($r$-$\\ell$WL), a novel hierarchy of graph isomorphism tests and a corresponding GNN framework, $r$-$\\ell$MPNN, that can count cycles up to length $r{+}2$. Most notably, we show that $r$-$\\ell$WL can count homomorphisms of cactus graphs. This extends 1-WL, which can only count homomorphisms of trees and, in fact, is incomparable to $k$-WL for any fixed $k$. We empirically validate the expressive and counting power of $r$-$\\ell$MPNN on several synthetic datasets and demonstrate the scalability and strong performance on various real-world datasets, particularly on sparse graphs."},"_bibtex":{"value":"@inproceedings{\npaolino2024weisfeiler,\ntitle={Weisfeiler and Leman Go Loopy: A New Hierarchy for Graph Representational Learning},\nauthor={Raffaele Paolino and Sohir Maskey and Pascal Welke and Gitta Kutyniok},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=9O2sVnEHor}\n}"},"title":{"value":"Weisfeiler and Leman Go Loopy: A New Hierarchy for Graph Representational Learning"},"pdf":{"value":"/pdf/160b0368f27f6ae00575a4abc8d44870237c95f9.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"paolino|weisfeiler_and_leman_go_loopy_a_new_hierarchy_for_graph_representational_learning"},"authorids":{"value":["~Raffaele_Paolino1","~Sohir_Maskey1","~Pascal_Welke1","~Gitta_Kutyniok2"]},"authors":{"value":["Raffaele Paolino","Sohir Maskey","Pascal Welke","Gitta Kutyniok"]}},"version":2},{"content":{"summary":{"value":"The authors present a new architecture, specifically they integrate some clever ideas of Dirac structure in order to unify the port-Hamiltonian and Poisson formulations from geometric mechanics.  Both nice innovations for physics-constrained neural networks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The models considered all seem to be linear.  Is that correct?  If so, more standard system ID or linear methods should be considered instead of all this sophisticated ML/AI architectures.  \n\nCan this generalize to nonlinear models?  Or simply be applied to nonlinear models with success?  Although you did Fitzhugh-Nagumo and  Chua's model, these are only \"slightly\" nonlinear and system ID models can work pretty well with those.\n\nWhat have you not considered non-neural network methods like DMD... or system ID methods like mamba/S4 not been compared?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The integration of physics-based principles is always welcome in the dynamical systems world.  The authors have a very nice contribution to make here potentially as the integration of physics principles into neural networks is very important."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The models used seem to be all linear: which begs the question about simple linear model regressions such as dynamic mode decomposition, dynamic mode decomposition with control and time-delay embedded DMD which models missing/coupled physics.  These more baseline (non-NN) methods are simply not talked about or considered and I think they should be.\n\nThere are statements that are simply not true:  \"two key limitations remain in modeling dynamical systems, especially those described by ordinary differential equations (ODEs). The first limitation is the narrow focus on mechanical systems.\"  This suggests the authors don't know the field well and it is a concern.  People are modeling all kinds of dynamical systems with ML/AI architectures well beyond mechanical systems.\n\nFurther: \"The second limitation is that most methods treat the system as a single, monolithic entity.\"  This is also not true.  Many people are working on coupled systems where time-delay embeddings are often used to gather information for missing, coupled and unmeasured variables.  Again, it is odd that these statements exist in the paper which suggests the authors are not aware of the great body of work on model discovery and dynamical systems methods with ML/AI for systems which are not mechanical and which indeed have coupling."}},"nonreaders":[],"tmdate":1731428090709,"tcdate":1730636763057,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4280/Reviewer_h9yM"],"signatures":["ICLR.cc/2025/Conference/Submission4280/Reviewer_h9yM"],"forum":"U1DjXQeJRx","number":5,"license":"CC BY 4.0","cdate":1730636763057,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4280/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428090709,"domain":"ICLR.cc/2025/Conference","replyto":"U1DjXQeJRx","id":"h1Sy6LQKut","forumContent":{"TLDR":{"value":"Poisson-Dirac formulation with ports enables neural networks to model various dynamical systems across domains and identify their internal structures."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["neural ordinary differential equations","coupled system","Poisson system","Dirac structure"]},"supplementary_material":{"value":"/attachment/376c68399daaa997b7f3994fa1de6490412f27aa.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Deep learning has achieved great success in modeling dynamical systems, providing data-driven simulators to predict complex phenomena, even without known governing equations. However, existing models have two major limitations: their narrow focus on mechanical systems and their tendency to treat systems as monolithic. These limitations reduce their applicability to dynamical systems in other domains, such as electrical and hydraulic systems, and to coupled systems. To address these limitations, we propose Poisson-Dirac Neural Networks (PoDiNNs), a novel framework based on the Dirac structure that unifies the port-Hamiltonian and Poisson formulations from geometric mechanics. This framework enables a unified representation of various dynamical systems across multiple domains as well as their interactions and degeneracies arising from couplings. Our experiments demonstrate that PoDiNNs offer improved accuracy and interpretability in modeling unknown coupled dynamical systems from data."},"_bibtex":{"value":"@inproceedings{\nkhosrovian2025poissondirac,\ntitle={Poisson-Dirac Neural Networks for Modeling Coupled Dynamical Systems across Domains},\nauthor={Razmik Arman Khosrovian and Takaharu Yaguchi and Hiroaki Yoshimura and Takashi Matsubara},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=U1DjXQeJRx}\n}"},"title":{"value":"Poisson-Dirac Neural Networks for Modeling Coupled Dynamical Systems across Domains"},"pdf":{"value":"/pdf/cbe5869996e1da69833fe07385317929fb66b5f9.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"khosrovian|poissondirac_neural_networks_for_modeling_coupled_dynamical_systems_across_domains"},"authorids":{"value":["~Razmik_Arman_Khosrovian1","~Takaharu_Yaguchi1","~Hiroaki_Yoshimura1","~Takashi_Matsubara1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Razmik Arman Khosrovian","Takaharu Yaguchi","Hiroaki Yoshimura","Takashi Matsubara"]}},"version":2},{"content":{"venue":{"value":"Physics Letters B"},"pdf":{"value":"https://www.sciencedirect.com/science/article/pii/S0370269320304317/pdfft?md5=6b31c11595a0c54fbdcf2d3f6b753aca&pid=1-s2.0-S0370269320304317-main.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"khan|physicsinspired_deep_learning_to_characterize_the_signal_manifold_of_quasicircular_spinning_nonprecessing_binary_black_hole_mergers"},"html":{"value":"https://doi.org/10.1016/j.physletb.2020.135628"},"_bibtex":{"value":"@article{Khan_2020,\n\tdoi = {10.1016/j.physletb.2020.135628},\n\turl = {https://doi.org/10.1016%2Fj.physletb.2020.135628},\n\tyear = 2020,\n\tmonth = {sep},\n\tpublisher = {Elsevier {BV}},\n\tvolume = {808},\n\tpages = {135628},\n\tauthor = {Asad Khan and E.A. Huerta and Arnav Das},\n\ttitle = {Physics-inspired deep learning to characterize the signal manifold of quasi-circular, spinning, non-precessing binary black hole mergers},\n\tjournal = {Physics Letters B}\n}"},"abstract":{"value":"The spin distribution of binary black hole mergers contains key information concerning the formation channels of these objects, and the astrophysical environments where they form, evolve and coalesce. To quantify the suitability of deep learning to estimate the individual spins, effective spin and mass-ratio of quasi-circular, spinning, non-precessing binary black hole mergers, we introduce a modified version of WaveNet trained with a novel optimization scheme that incorporates general relativistic constraints of the spin properties of astrophysical black holes. The neural network model is trained, validated and tested with 1.5 million ℓ=|m|=2 waveforms generated within the regime of validity of NRHybSur3dq8, i.e., mass-ratios q≤8 and individual black hole spins |s|{1,2}z≤0.8. To reduce time-to-insight, we deployed a distributed training algorithm at the IBM Power9 Hardware-Accelerated Learning cluster at the National Center for Supercomputing Applications to reduce the training stage from 1 month, using a single V100 NVIDIA GPU, to 12.4 hours using 64 V100 NVIDIA GPUs. We have also fully trained this model using 1536 V100 GPUs (256 nodes) in the Summit supercomputer at Oak Ridge National Laboratory, achieving state-of-the-art accuracy within just 1.2 hours. Using this neural network model, we quantify how accurately we can infer the astrophysical parameters of black hole mergers in the absence of noise. We do this by computing the overlap between waveforms in the testing data set and the corresponding signals whose mass-ratio and individual spins are predicted by our neural network. We find that the convergence of high performance computing and physics-inspired optimization algorithms enable an accurate reconstruction of the mass-ratio and individual spins of binary black hole mergers across the parameter space under consideration. This is a significant step towards an informed utilization of physics-inspired deep learning models to reconstruct the spin distribution of binary black hole mergers in realistic detection scenarios."},"title":{"value":"Physics-inspired deep learning to characterize the signal manifold of quasi-circular, spinning, non-precessing binary black hole mergers"},"authors":{"value":[{"fullname":"Asad Khan","username":"~Asad_Khan3"},{"fullname":"E.A. Huerta"},{"fullname":"Arnav Das"}]}},"tmdate":1789091750440,"pdate":1598918400000,"externalIds":["doi:10.1016/j.physletb.2020.135628"],"tcdate":1761445306785,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Asad_Khan3"],"forum":"PsPF8o9dmS","license":"CC BY-SA 4.0","number":1993,"cdate":1595870776846,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789091750440,"domain":"OpenReview.net/Public_Article","id":"PsPF8o9dmS","version":2},{"content":{"venue":{"value":"Journal of Computational Physics"},"pdf":{"value":"/pdf/29c8584c566e11a58f2889be02ccdcfe323d58c9.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"liu|modular_operator_superposition_mos_a_physicsguided_machine_learning_framework_for_addressing_the_curse_of_dimensionality_and_multiscale_challenges_in_computational_fluid_dynamics"},"authorids":{"value":["~Kai_Liu43","~S._Balachandar1","~Haochen_Li14"]},"abstract":{"value":"We introduce Modular Operator Superposition (MOS), a physics-guided and AI-augmented framework for efficient, scalable, and generalizable flow field modeling in high-dimensional and multiscale fluid systems. Rather than globally resolving flow fields via mesh-based discretization, MOS decomposes the system into physically meaningful flow primitives, each represented by a reusable modular operator. These operators are trained offline using a parameterized physics-informed neural network (P-PINN) in a single pre-processing step, and later composed through a physics-guided superposition strategy to approximate the full system-level mapping. The core advantage of MOS lies in its modularization strategy. By learning only small-scale flow primitives offline, MOS reduces the training cost to a fixed, minimal investment independent of system-level complexity. In the online stage, MOS dynamically solves for primitive-level interactions for any specific configuration of a system, then reconstructs the global flow field through the superposition of modular outputs. This two-stage online process, comprising both solving and inference, enables scalable and generalizable predictions. As a result, MOS addresses the curse of dimensionality by reducing high-dimensional systems to tractable compositions of modular operators and overcomes multiscale challenges through a scale-adaptive operator assembly that flexibly resolves flow features with minimal overhead. We demonstrate MOS for static and dynamic arrays of up to $15{,}000$ cylinders in a channel cross-flow (corresponding to roughly $10^5$ input parameters). All of these configurations are solved using a shared single-cylinder cross-flow modular operator, trained offline in $30$ hours using a data-free, physics-informed machine learning strategy. In the online stage, MOS achieves end-to-end flow field prediction at $3$ to $5$ orders of magnitude speedup over conventional numerical solvers, while maintaining high fidelity ($R^2 > 0.85$ for all cases). Moreover, the MOS solution format requires $3$ to $5$ orders of magnitude lower memory usage than conventional numerical outputs. Once solved, the solution can be queried in real-time to infer flow variables at arbitrary spatial resolutions or scattered points, enabling flexible and efficient visualization across scales. Additional tests indicate that MOS remains robust to polydispersity and translational/rotational motion of the cylinders."},"title":{"value":"Modular operator superposition (MOS): A physics-guided machine learning framework for addressing the curse of dimensionality and multiscale challenges in computational fluid dynamics"},"authors":{"value":["Kai Liu","S. Balachandar","Haochen Li"]}},"tmdate":1778185792312,"pdate":1767240000000,"tcdate":1778185792312,"writers":["~Kai_Liu43","~S._Balachandar1","~Haochen_Li14"],"signatures":["~Kai_Liu43"],"forum":"l64DC2Nvnw","license":"CC BY 4.0","number":49694,"cdate":1778185792312,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1778185792312,"domain":"OpenReview.net/Archive","id":"l64DC2Nvnw","version":2},{"content":{"summary":{"value":"This paper introduces CryoCCD, a synthesis framework designed to generate realistic cryo-electron microscopy (cryo-EM) micrographs. It integrates biophysical modeling with a conditional cycle-consistent diffusion model, which captures the inherent heterogeneity and complex noise characteristics in cryo-EM data. The method employs a biophysical engine that simulates structural diversity and imaging physics, ensuring accurate representation of biological specimens. CryoCCD also incorporates cycle-consistent translation and mask-guided contrastive learning to improve the noise generation process, providing high-fidelity synthetic data. The results show that CryoCCD outperforms existing methods in multiple tasks, including particle picking, pose estimation, and generalization across diverse protein families. The framework demonstrates strong potential to reduce the need for extensive manual annotations and accelerate the development of cryo-EM tools."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"The direction of this paper is promising, and I appreciate the substantial engineering effort and technical details presented. However, the weaknesses outlined above are too significant for me to confidently recommend acceptance at this stage. I encourage the authors to address these issues to substantially strengthen the paper. Thanks!"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. **Originality**: This paper builds upon the foundational work of CryoGEM, but its key innovation lies in the introduction of the CycleDiffusion approach. By combining the power of diffusion models with cycle consistency, the authors present a more robust framework for generating realistic cryo-EM micrographs. The novel integration of diffusion loss and cycle loss enhances the method’s stability and performance, especially in preserving structural fidelity while modeling complex noise. Moreover, the paper expands on the training process by incorporating a larger, more diverse dataset, enabling the model to generalize effectively across mixed datasets, which is a significant advancement over previous generative models.\n\n2. **Quality**: The authors provide sufficient technical details, making the methodology both clear and reproducible. The experimental results are credible and demonstrate realistic simulations, showcasing the method's ability to generate synthetic cryo-EM micrographs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Clarity and Methodological Ambiguity**:\nSeveral parts of the paper lack clarity and precise definition, especially in Section 3 and Section 4. The role of the mesh extracted from isosurfaces is never clearly explained—although the paper mentions mesh simplification in the Multi-Scale Volume Modeling stage, it is unclear how or whether the mesh is used in the final projection process. Similarly, the mask generation process is never defined; Section 4 directly introduces masks as inputs without explaining how they are computed or derived. Furthermore, in Section 4, both the notation and logical flow are confusing. After defining two diffusion models, G_AB and G_BA, these are not used in Equations (3)–(6), and several variables such as 𝜖_𝜃 and mask appear without prior definition. The calculation of L_diff lacks a clear description of how G_AB and G_BA participate in the diffusion process. The same issue extends to Equation (9), where it is unclear whether the forward diffusion process involves a single or multiple sampling steps, and the variables x and y are never defined. This lack of formal clarity makes it difficult to follow or reproduce the method.\n\n2. **Conceptual Inconsistency in Motivation (Section 4.2)**:\nThe paper criticizes GAN-based methods like CryoGEM for not considering cycle consistency, which the authors argue limits their performance. However, the authors still incorporate both contrastive loss and GAN loss—the same losses used in CryoGEM. This creates a conceptual inconsistency: while they criticize CryoGEM for not using cycle consistency, they continue to use the same losses from CryoGEM and do not clearly justify the motivation for their inclusion in the new framework. This redundancy weakens the theoretical coherence of the paper and calls for a clearer explanation of why these losses are necessary in the proposed method.\n\n3. **Weak Baseline Comparison**:\nThe evaluation of CryoGEM as a baseline is highly confusing. The reported CryoGEM results are far worse than those in the original paper, even though the CryoGEM code, pretrained weights, and datasets are publicly available (e.g., EMPIAR-10028). For instance, in Table 2, the AUPRC for particle picking on PhageMS2 drops from 0.915 (in the original paper) to only 0.159, and in Figure 4, the Poisson noise simulations fail completely. This raises doubts about whether CryoGEM was correctly reproduced or appropriately configured. One possibility is that CryoGEM, which was designed as a per-scene training method, may have been trained on mixed datasets (15 different scenes), which is inconsistent with its intended use. The paper should provide a clear justification for why the baseline fails to converge and, for fairness, compare CryoCCD against CryoGEM’s official pretrained models or datasets.\n\n4. **Experimental Limitations and Generalization Concerns**: Although CryoCCD claims to generalize across heterogeneous data, the experimental results suggest otherwise. In Figure 4, the generated synthetic micrographs show substantial visual differences from real data, and the FID scores consistently exceed 100, indicating limited realism. Conceptually, it is also unclear how a diffusion model can learn to generate realistic micrographs without conditioning on microscope parameters such as defocus or dose rate. Since each biological structure in the dataset corresponds to only one specific noise distribution, training a single diffusion model to generalize across multiple noise conditions seems questionable. Without additional conditioning variables, the model lacks the information needed to decide what type of noise to produce, which may explain the poor quantitative metrics and unstable training dynamics. The paper would benefit from ablation studies or additional experiments clarifying how CryoCCD handles this inherent heterogeneity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917199118,"tcdate":1761570630502,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4150/Reviewer_3N8p"],"signatures":["ICLR.cc/2026/Conference/Submission4150/Reviewer_3N8p"],"forum":"SdGgEcsx2W","number":2,"license":"CC BY 4.0","cdate":1761570630502,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4150/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917199118,"domain":"ICLR.cc/2026/Conference","replyto":"SdGgEcsx2W","id":"Xo2TEx7bjs","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["cryo-EM; diffusion; synthesis; simulator"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Single-particle cryo-electron microscopy (cryo-EM) has become a cornerstone of structural biology, enabling near-atomic resolution analysis of macromolecules through advanced computational methods. However, the development of cryo-EM processing tools is constrained by the scarcity of high-quality annotated datasets. Synthetic data generation offers a promising alternative, but existing approaches lack thorough biophysical modeling of heterogeneity and fail to reproduce the complex noise observed in real imaging. To address these limitations, we present CryoCCD, a synthesis framework that unifies versatile biophysical modeling with the first conditional cycle-consistent diffusion model tailored for cryo-EM. The biophysical engine provides multi-functional generation capabilities to capture authentic biological organization, and the diffusion model is enhanced with cycle consistency and mask-guided contrastive learning to ensure realistic noise while preserving structural fidelity. Extensive experiments demonstrate that CryoCCD generates structurally faithful micrographs, enhances particle picking and pose estimation, as well as achieves superior performance over state-of-the-art baselines, while also generalizing effectively to held-out protein families."},"_bibtex":{"value":"@misc{\njiang2025cryoccd,\ntitle={Cryo{CCD}: Conditional Cycle-consistent Diffusion with Biophysical Modeling for Cryo-{EM} Synthesis},\nauthor={Runmin Jiang and Genpei Zhang and Yuntian Yang and Siqi Wu and Minhao Wu and Wanyue Feng and Yizhou Zhao and Xi Xiao and Xiao Wang and Tianyang Wang and Xingjian Li and Muyuan Chen and Min Xu},\nyear={2025},\nurl={https://openreview.net/forum?id=SdGgEcsx2W}\n}"},"title":{"value":"CryoCCD: Conditional Cycle-consistent Diffusion with Biophysical Modeling for Cryo-EM Synthesis"},"pdf":{"value":"/pdf/2a43d64a5f5ac6a8ba92b52ac416392d2821b906.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"jiang|cryoccd_conditional_cycleconsistent_diffusion_with_biophysical_modeling_for_cryoem_synthesis"},"authorids":{"value":["~Runmin_Jiang1","~Genpei_Zhang2","~Yuntian_Yang1","~Siqi_Wu6","~Minhao_Wu1","~Wanyue_Feng1","~Yizhou_Zhao2","~Xi_Xiao2","~Xiao_Wang21","~Tianyang_Wang1","~Xingjian_Li1","~Muyuan_Chen1","~Min_Xu4"]},"authors":{"value":["Runmin Jiang","Genpei Zhang","Yuntian Yang","Siqi Wu","Minhao Wu","Wanyue Feng","Yizhou Zhao","Xi Xiao","Xiao Wang","Tianyang Wang","Xingjian Li","Muyuan Chen","Min Xu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces the framework of understanding the LLM synthetic data generation process, by incorporating the reverse-bottleneck perspective. Using a reverse-bottleneck perspective, the authors model the generation of synthetic data and propose an information-theoretic metric called Generalization Gain via Mutual Information (GGMI) to assess the generalization performance of post-trained models. It also reports the information gain and generalization bound from previous literature."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"See above"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The motivation is easy to follow and the figure is well represented.\n\n2. The equation derived from previous literature, such as Lemma 3.1 and 4.4 is clear. \n\n3. The paper tries to inject the LLM synthetic generation perspective using previously derived framework, showing a potential direction to understanding the synthetic data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks the experimental support, which cannot convince the readers of their assumption and derivations.\n\n2. The paper is not organized in a good shape for guiding the readers for understanding their contributions. Instead, it enumerates the equations of generalization errors and gains, but lacking sufficient experimental results to support this.\n\n3. Some derivations may be challenging for readers unfamiliar with advanced information theory, which may requires providing backgrounds about this. Since the authors did not conduct real-world experiments, it is unsure whether this can generalize to real-world settings. The framework assumes ideal conditions for synthetic data and LLMs, which may not generalize well to complex real-world scenarios with noisy data or biases."}},"nonreaders":[],"tmdate":1733149221744,"tcdate":1730721858635,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5944/Reviewer_qaT5"],"signatures":["ICLR.cc/2025/Conference/Submission5944/Reviewer_qaT5"],"forum":"UxkznlcnHf","number":4,"license":"CC BY 4.0","cdate":1730721858635,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5944/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733149221744,"domain":"ICLR.cc/2025/Conference","replyto":"UxkznlcnHf","id":"MZsFEyZ9Zr","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"This paper explores the critical role of synthetic data in enhancing the post-training performance of large language models (LLMs) from a novel reverse-bottleneck perspective."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["large language models; synthetic data; information bottleneck"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Synthetic data has become a pivotal resource in post-training tasks for large language models (LLMs) due to the scarcity of high-quality, specific data. While various methods have been developed to generate synthetic data, there remains a discernible gap between the practical effects of synthetic data and our theoretical comprehension. To address this challenge, we commence by presenting a detailed modeling of the prevalent synthetic data generation process. Building upon this modeling, we demonstrate that the generalization capability of the post-trained model is critically determined by the information gain derived from the generative model, as analyzed from a novel reverse-bottleneck perspective. Moreover, we introduce the concept of Generalization Gain via Mutual Information (GGMI) and elucidate the relationship between generalization gain and information gain. This analysis serves as a theoretical foundation for synthetic data generation and further highlights its connection with the generalization capability of post-trained models, offering an understanding about the design of synthetic data generation techniques and the optimization of the post-training process. We open-source our code at https://github.com/ZyGan1999/Towards-a-Theoretical-Understanding-of-Synthetic-Data-in-LLM-Post-Training."},"_bibtex":{"value":"@inproceedings{\ngan2025towards,\ntitle={Towards a Theoretical Understanding of Synthetic Data in {LLM} Post-Training: A Reverse-Bottleneck Perspective},\nauthor={Zeyu Gan and Yong Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=UxkznlcnHf}\n}"},"title":{"value":"Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective"},"pdf":{"value":"/pdf/127d76775eb769452b3e1f3cffc5359d9e886a32.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"gan|towards_a_theoretical_understanding_of_synthetic_data_in_llm_posttraining_a_reversebottleneck_perspective"},"authorids":{"value":["~Zeyu_Gan1","~Yong_Liu7"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyu Gan","Yong Liu"]}},"version":2},{"content":{"comment":{"value":"### **Appendix B.1  Physics-Inspired Nature of PCMFA**\n\n\nTo formalize the intuition behind PCMFA, we frame the module as the fusion of two norm-stable branches—a geometric path operating in hyperbolic space and a quantum-inspired path operating in a complex Hilbert space. We first state a precise definition of a physics-inspired attention mechanism and then prove that PCMFA satisfies this definition.\n\n### **Definition 1 (Physics-inspired attention mechanism)**\nLet $\\mathcal{X}$ be an input feature space, $\\mathcal{M}$ a Riemannian\nmanifold of constant negative curvature (e.g., a Poincar\\'e ball or Lorentz\nhyperboloid), and $\\mathcal{H}$ a finite-dimensional complex Hilbert space\nwith inner product $\\langle\\cdot,\\cdot\\rangle$. An attention mechanism $T : \\mathcal{X} \\to [0,1]^e$ is called \n\\emph{physics-inspired} if there exist maps\n$E^{\\mathrm{geo}} : \\mathcal{X} \\to \\mathcal{M}$ and\n$\\Psi : \\mathcal{X} \\to \\mathcal{H}$ such that:\n\n**(G1) Geometric branch.**\n  $E^{\\mathrm{geo}}$ is realized via a Riemannian exponential-type map on\n  $\\mathcal{M}$ (e.g., Poincar\\'e ball or Lorentz hyperboloid), and the\n  corresponding attention scores depend only on geodesic distances or norms\n  in $\\mathcal{M}$.\n\n**(Q1) Quantum-inspired branch.**\n  $\\Psi(x)$ is a complex state whose channel-wise scores are proportional to\n  Born-rule amplitudes $|\\Psi(x)_k|^2$, i.e., squared magnitudes in\n  $\\mathcal{H}$, normalized to a probability vector, following the standard\n  probabilistic interpretation of quantum states.\n\n**(N1) Physical normalization and stability.**\n  The final attention weights $T(x)$ are obtained from these physically\n  derived scores by a probability-preserving normalization (e.g., Softmax or sigmoid) and induce a non-expansive multiplicative update on \n  features in every $\\ell_p$ norm.\n\n\n### **Theorem 1 (PCMFA is a physics-inspired attention mechanism)**\nFor each EHF layer $i$, the PCMFA map\n$$\nA_i = \\mathrm{PCMFA}\\bigl(x_i^{S}, x'_{i+1}\\bigr)\n$$\ndefined in **Eq.(1)** is a physics-inspired attention mechanism in the sense of\n**Definition (1)**.\n\n\n### ***Proof***. \nWe verify (G1)--(Q1)--(N1) for the PCMFA construction in **Sec. 3.1.2**.}\n\n**(G1) Geometric branch.**\nWithin PCMFA, the MHDGA block takes\n$\\psi_i = \\mathrm{GAP}(\\mathrm{DCT}(x'_i))$ (**Eq. 3**) and maps it to\nhyperbolic embeddings in two constant-negative-curvature models:\nthe Poincar\\'e ball and the Lorentz hyperboloid (**Eqs.4--8**).\nBoth mappings are exponential-type projections with learnable but bounded\ncurvature (see the curvature parameterization and clipping in **Sec. 3.1.2** and\n**Lemma 2** below), so\n$E^{\\mathrm{geo}}(\\psi_i)$ is a well-defined element of\n$\\mathcal{M}^P \\times \\mathcal{M}^L$.\nThe corresponding attentions $A^P_i, A^L_i$ depend only on hyperbolic radii\nand Lorentzian norms, and are fused into $A^D_i$ by an affine map plus\nsigmoid (**Eq. 3**), satisfying (G1).\n\n\n**(Q1) Quantum-inspired branch.**\nMQIA maps the same $\\psi_i$ into a complex vector\n$q_i \\in \\mathbb{C}^e$ via learned real and imaginary coefficients (**Eq. 9**),\nso $q_i$ lies in a Hilbert space $\\mathcal{H} \\cong \\mathbb{C}^e$.\nThe channel-wise quantities $|q_{i,k}|^2$ are Born-rule amplitudes, and\nSoftmax over a scaled version of $|q_i|^2$ together with a Lorentz-norm term, yields a probability vector $A^Q_i$ (**Eq. 9**), satisfying (Q1). \n\n\n**(N1) Normalization and stability.**\nMAFG fuses $A^D_i$ and $A^Q_i$ channel-wise using bounded coefficients\n$c_i$ (**Eq.10**).\n**Lemma 2** below implies that all components of\n$A^D_i$ lie in $(0,1)$ and $A^Q_i$ is a probability vector; under the bounded\ngate assumption, the fused gate $A_i$ satisfies $\\|A_i\\|_\\infty \\le 1$.\n**Theorem 3** then shows that the multiplicative update\n$z \\mapsto z \\odot A_i$ is non-expansive in every $\\ell_p$ norm.\n \n\nTogether, these properties establish that PCMFA has (i) a geometric branch on\nhyperbolic manifolds, (ii) a quantum-inspired branch in a complex Hilbert\nspace, and (iii) a normalized, non-expansive fusion of the two, so it is\nphysics-inspired in the sense of **Definition 1**."},"title":{"value":"Author Response (Part 2/6)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1770917693715,"tcdate":1764170870347,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20215/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission20215/Authors"],"forum":"mZJM8hXmVg","number":3,"license":"CC BY 4.0","cdate":1764170870347,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20215/-/Official_Comment","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770917693715,"domain":"ICLR.cc/2026/Conference","replyto":"5JXy3DkI6z","id":"duQKHgBGmC","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Medical Imaging","Multimodal Fusion","Attention","Deep Learning"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Multimodal fusion learning (MFL) paradigm (a framework to jointly learn from heterogeneous data sources) has shown great potential in various fields such as Medicine, Science, Engineering, etc. It is extremely desirable in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges. First, they struggle to capture complex cross-modal interactions effectively. Second, they are often designed and evaluated for narrow, fixed modality configurations (e.g., imaging-only, or specific pairs such as image and omics or image and clinical text), which limits evidence of their adaptability and generalizability to broader collections of heterogeneous medical modalities. Finally, they incur high computational costs, restricting their applicability in resource-constrained healthcare AI. To address these challenges, we propose a novel MFL framework – Efficient Hybrid-fusion Physics-inspired Attention Learning Network (EHPAL-Net) – a lightweight and scalable framework that integrates various modalities through novel Efficient Hybrid Fusion (EHF) layers. Each EHF layer captures rich modality-specific multi-scale spatial information, followed by a Physics-inspired Cross-modal Fusion Attention module to model fine-grained, structure-preserving cross-modal interactions, thereby learning robust complementary shared representations. Furthermore, EHF layers are sequentially learned for each modality, making them adaptable and generalizable. Extensive evaluations on 15 public datasets show that EHPAL-Net outperforms leading multimodal fusion methods, boosting performance by up to 3.97% and lowering computational costs by up to 87.8%, ensuring more effective and reliable predictions."},"_bibtex":{"value":"@misc{\ndhar2026advancing,\ntitle={Advancing Multimodal Fusion on Heterogeneous Data with Physics-inspired Attention},\nauthor={Joy Dhar and Chen Chen and Nayyar Zaidi and Maryam Haghighat and Puneet Goyal and Manish Kumar Pandey and Ferdous Sohel},\nyear={2026},\nurl={https://openreview.net/forum?id=mZJM8hXmVg}\n}"},"title":{"value":"Advancing Multimodal Fusion on Heterogeneous Data with Physics-inspired Attention"},"pdf":{"value":"/pdf/dcca73e1c79fba073f627966bc41011969e062a4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dhar|advancing_multimodal_fusion_on_heterogeneous_data_with_physicsinspired_attention"},"authorids":{"value":["~Joy_Dhar1","~Chen_Chen18","~Nayyar_Zaidi1","~Maryam_Haghighat1","~Puneet_Goyal1","~Manish_Kumar_Pandey1","~Ferdous_Sohel1"]},"authors":{"value":["Joy Dhar","Chen Chen","Nayyar Zaidi","Maryam Haghighat","Puneet Goyal","Manish Kumar Pandey","Ferdous Sohel"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a new neural network architecture to simulate the evolution of a given physical system that can be seen as a generalized weighted combination of linearizations around a set reference pairs of state and measurements $(x_{ref, i}, y_{ref, i})_{i=1}^{N}$. The paper proposes to use this method when solving inverse problems in physics through MAP (maximum a posteriori) estimations. It then shows that using this architecture they can achieve smaller RMSE w.r.t the state $x$ that was used to generate the observation $y$ with different regularizations (priors) for the problem of ocean acoustic tomography when compared to other forward models."},"soundness":{"value":"2 fair"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"* If $P_y$ is a matrix, then I do not see why the choice of separating $W$ and $P_y$ in equation 8, since $WP_Y$ is just another matrix of dimension $2431 \\times 800$, that could be directly parameterized.\n\n* It is known that self-normalisation of weights that correspond to inner products in big dimensions lead to weight collapse (after self normalisation one of the reference ($x_ref, y_ref$) can have a weight of almost one while the others 0). In this case, the inverse problem would be basically based upon $A_ref$ for this given pair. Is this something that was seen during the optimization for solving the inverse problem? If this is the case, the model would be basically selecting one of the reference solutions, and could be simplified.\n\n* How sensitive is the model to the reference solution that were chosen? How would it be impacted if another set was chosen? \n\n*The MLP model seems to be overfitting (lines 275 and 276). If i'm not mistaken, standard techniques to reduce overfitting have not been tried (dropout, batch normalisation, etc...). How would those impact the performance?\n\n\n"},"rating":{"value":"3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"2 fair"},"contribution":{"value":"1 poor"},"strengths":{"value":"The paper proposes a reasonable model that is based on reference set. The presentation of the inverse problem applications to the physics domain is good and the application to ocean acoustic tomography is original and not often found amongst the NeurIPS community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"One of the main weakness of the paper is the lack of support of the claim that the proposed architecture is a efficient tool for modelling general forward problems in physics. There are two main problems with the validation of the method:\n\n1. Lack of replicates for the validation metrics:\n - Neither table 1 or 2 contain confidence intervals. Those kind of metrics are expected to have aggregated results from several replications of initializations of the network before training as well as several initializations of the optimization problem for solving the inverse problem. Without it, it is quite hard to know if the RMSE differences are signficative in say MLP vs Petal in table 1 or PETAL vs WAN in table 2.\n\n2. Comparison with other baselines:\nI think the comparison with other possible forward models is lacking. Indeed I'd expect a comparison to a more expressive meta model than a MLP, by either training a more complex network that better exploits the 2d nature of the problem. Other methods that take advantage of some reference set, such as Gaussian processes could have been exploited as well.\nI'd like to know as well if the proposed method scales to other domains than Ocean Acoustic Tomography. This would be of great interest in reproducibility, since the dataset used is not available and therefore the community has no way of replicating the experiments.\n"},"limitations":{"value":"NA"}},"nonreaders":[],"tmdate":1702411149131,"tcdate":1688462496114,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission8113/Reviewer_1TNX"],"signatures":["NeurIPS.cc/2023/Conference/Submission8113/Reviewer_1TNX"],"forum":"CXrRMfs5eY","number":2,"license":"CC BY 4.0","cdate":1688462496114,"mdate":1702411149131,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission8113/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"CXrRMfs5eY","id":"VEuxfu25Vx","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Inverse Problems","Neural Adjoint","Hybrid Machine Learning","Physics"]},"supplementary_material":{"value":"/attachment/3d275badafcb3a26c11fa4e34a53f5e47288f0dc.zip"},"_bibtex":{"value":"@inproceedings{\njin2023petal,\ntitle={{PETAL}: Physics Emulation Through Averaged Linearizations for Solving Inverse Problems},\nauthor={Jihui Jin and Etienne Ollivier and Richard Touret and Matthew McKinley and Karim Sabra and Justin Romberg},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=CXrRMfs5eY}\n}"},"title":{"value":"PETAL: Physics Emulation Through Averaged Linearizations for Solving Inverse Problems"},"paperhash":{"value":"jin|petal_physics_emulation_through_averaged_linearizations_for_solving_inverse_problems"},"TLDR":{"value":"PETAL proposes an explicit approach to embed physics-based linearizations into surrogate modeling for solving inverse problems."},"abstract":{"value":"Inverse problems describe the task of recovering an underlying signal of interest given observables. Typically, the observables are related via some non-linear forward model applied to the underlying unknown signal. Inverting the non-linear forward model can be computationally expensive, as it often involves computing and inverting a linearization at a series of estimates. Rather than inverting the physics-based model, we instead train a surrogate forward model (emulator) and leverage modern auto-grad libraries to solve for the input within a classical optimization framework. Current methods to train emulators are done in a black box supervised machine learning fashion and fail to take advantage of any existing knowledge of the forward model. In this article, we propose a simple learned weighted average model that embeds linearizations of the forward model around various reference points into the model itself, explicitly incorporating known physics. Grounding the learned model with physics based linearizations improves the forward modeling accuracy and provides richer physics based gradient information during the inversion process leading to more accurate signal recovery. We demonstrate the efficacy on an ocean acoustic tomography (OAT) example that aims to recover ocean sound speed profile (SSP) variations from acoustic observations (e.g. eigenray arrival times) within simulation of ocean dynamics in the Gulf of Mexico."},"pdf":{"value":"/pdf/16abe1d60824e68317418bb5e8abc77b50f3ef26.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Jihui_Jin1","eollivier3@gatech.edu","rtouret3@gatech.edu","mmckinley31@gatech.edu","~Karim_Sabra1","~Justin_Romberg1"]},"authors":{"value":["Jihui Jin","Etienne Ollivier","Richard Touret","Matthew McKinley","Karim Sabra","Justin Romberg"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"TLDR":{"value":"PS3Simulator generates physics-parametrised  synthetic sonar data enabling structure-aware self-supervised learning for sim-to-real sonar classification without any real data collection."},"venue":{"value":"MaCVi Poster"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"keywords":{"value":["self-supervised learning","synthetic data generation","sim-to-real transfer","sonar image classification","vision transformers","masked image modeling","domain adaptation","underwater robotics","physics-based rendering"]},"venueid":{"value":"thecvf.com/CVPR/2026/Workshop/MaCVi"},"paperhash":{"value":"s|ps3simulator_physicsparametrised_synthetic_sonar_for_selfsupervised_simtoreal_transfer"},"authorids":{"value":["~Kamal_Basha_S1","~Athira_Nambiar2"]},"abstract":{"value":"Sonar imaging plays a critical role in underwater object detection and classification for maritime security, autonomous underwater vehicle operations, and environmental monitoring. However, real sonar datasets are scarce due to high collection costs, confidentiality restrictions, and expensive annotation, limiting the development of deep learning systems for underwater perception. Although synthetic data generation offers a scalable alternative, existing pipelines lack explicit physical parameters, and models trained on synthetic data suffer from a persistent sim-to-real domain gap. To address these limitations, three contributions are presented in this work. First, PS3Simulator, \nA Physics parametrised synthetic Side-Scan Sonar dataset is black presented by incorporating per-image physical parameters including seabed material, grazing angle, altitude, and object rotation, enabling controllable dataset generation at scale. Second, the first application of I-JEPA~\\cite{assran2023self} to the sonar domain proposes a structure-aware self-supervised learning framework based on masked latent prediction. Third, a cross-domain, cross-paradigm evaluation compares pretraining across ImageNet and PS3Simulator domains under a strict synthetic-train real-test protocol on KSLG and SCTD real-world sonar datasets. The proposed method achieves $70.9\\%$ accuracy, outperforming DINO~cite{caron2021emerging} on identical PS3Simulator data by $+12.1\\%$ ($\\pm4.5\\%$ vs $\\pm11.9\\%$ variance). At scale, I-JEPA ViT-H/14 attains $86.0\\%$. These results confirm that structure-aware SSL transfers more effectively than appearance-based methods for sim-to-real sonar transfer without requiring real sonar data."},"title":{"value":"PS3Simulator: Physics-Parametrised Synthetic Sonar for Self-Supervised Sim-to-Real Transfer"},"authors":{"value":["Kamal Basha S","Athira Nambiar"]}},"tmdate":1777291345092,"pdate":1777291343633,"tcdate":1772683564576,"writers":["thecvf.com/CVPR/2026/Workshop/MaCVi","thecvf.com/CVPR/2026/Workshop/MaCVi/Submission4/Authors"],"signatures":["thecvf.com/CVPR/2026/Workshop/MaCVi/Submission4/Authors"],"forum":"A9wFWcUhGX","license":"CC BY 4.0","number":4,"cdate":1772683564576,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Workshop/MaCVi/-/Submission","thecvf.com/CVPR/2026/Workshop/MaCVi/-/Submission_Change_Before_Reviewing","thecvf.com/CVPR/2026/Workshop/MaCVi/-/Submission_Change_Before_Bidding","thecvf.com/CVPR/2026/Workshop/MaCVi/-/Submission_Release"],"mdate":1777291345092,"odate":1777291343633,"domain":"thecvf.com/CVPR/2026/Workshop/MaCVi","id":"A9wFWcUhGX","version":2},{"content":{"summary":{"value":"This paper proposes a Predictor-Corrector mechanism to address the problem of error accumulation in multi-step time-series forecasting. The core idea is to treat any pre-trained time-series model as a Predictor and use its output forecasts to drive a Corrector model. This Corrector is implemented as a Neural Controlled Differential Equation (Neural CDE).\n\nThe mechanism works as follows:\n1. The Predictor generates a forecast trajectory.\n2. This forecast trajectory is used as the continuous control path for the Neural CDE.\n3. The Neural CDE is trained to predict the error (the residual between the Predictor's forecast and the ground truth).\n4. The final, corrected forecast is the sum of the original prediction and the CDE's predicted error.\n\nThe authors demonstrate that this framework can be applied to diverse Predictors, including continuous-time (NODE, ContiFormer) and discrete-time (DLinear) models, and works on both regular and irregular data. The method is validated on synthetic, physics simulation, and real-world LTSF datasets, showing consistent performance improvements over the Predictor-alone baselines. The paper also introduces two regularization strategies, \"Variable-Length Control Paths\" and \"Sparse Control Paths,\" to improve the Corrector's extrapolation capabilities."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Following Weakness #1: Could the authors provide an ablation study on the complexity of the decoder ($\\xi_{\\varphi}$)? Specifically, how does the Corrector's performance change if the 4-layer $FC(400)_4$ decoder is replaced with a simpler linear layer, or a 1-layer MLP? This would be crucial to disentangle the contribution of the CDE's dynamic modeling from the powerful function approximation capabilities of the decoder.\n2. Following Weakness #2: Given that other concurrent work (e.g., Jhin et al., 2024) also uses continuous-time models (NODEs) to address related problems like prediction delay, how do the authors position the conceptual novelty of their contribution? Is the novelty in the idea of using CDEs for time-series, or purely in the application as a post-hoc correction module?"},"rating":{"value":2},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The primary strength of this work is its agnosticism to the Predictor model. The ability to use the Corrector as a plug-and-play module to improve the performance of existing, pre-trained models—ranging from NODE to ContiFormer and even the non-continuous DLinear—is a significant practical advantage.\n2. The paper provides a comprehensive set of experiments across synthetic data (Lorenz, FHN) , high-dimensional physics simulations (MuJoCo) , and challenging real-world LTSF benchmarks (Weather, Exchange, etc.). The consistent improvement in MSE/MAE across all these settings (e.g., in Table 1, 2, and 4) makes a strong case for the method's effectiveness.\n3. The introduction of Variable-Length and Sparse Control Paths  is a clever and well-reasoned contribution. These regularizers directly address a known weakness of neural differential equations (extrapolation) and are shown to improve generalization while also accelerating training by reducing NFEs (as shown in Figure 4)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Over-reliance on the Decoder's Expressiveness: The entire framework hinges on the assumption that the CDE's latent state $z(t)$, controlled by the forecast path $X(s)$, will capture the necessary information to predict the error $e(t)$. However, the paper provides no structural or theoretical guarantee for this. The connection between $z(t)$ and $e(t)$ is established solely by a very powerful 4-layer MLP decoder ($FC(400)_4$). This design feels contrived. It is difficult to ascertain whether the CDE is truly learning the \"error dynamics\" or if it is merely acting as a complex feature extractor, with the heavy-lifting (i.e., fitting the complex $z$-to-$e$ mapping) being done by the high-capacity MLP decoder.\n\n2. Limited Methodological Novelty: The core idea of using continuous-time neural models (NODEs/CDEs) to address shortcomings in time-series forecasting is not entirely unique. For instance, the provided paper by Jhin et al. (2024, \"CONTIME\") also employs a continuous-time (NODE-based GRU) architecture to specifically solve the problem of prediction delays in time-series models. While this paper's application (post-hoc error magnitude correction) is different from CONTIME's (end-to-end delay correction), the foundational concept of using CDEs/NODEs to fix a known flaw in time-series prediction is very similar. This concurrent work suggests that this methodological tool is a \"next logical step\" in the field, which somewhat diminishes the paper's claim to high conceptual novelty."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359021389,"tcdate":1761840952871,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15025/Reviewer_EhtC"],"signatures":["ICLR.cc/2026/Conference/Submission15025/Reviewer_EhtC"],"forum":"biZBdFpOzu","number":2,"license":"CC BY 4.0","cdate":1761840952871,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15025/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359021389,"domain":"ICLR.cc/2026/Conference","replyto":"biZBdFpOzu","id":"w7qGzrbWKz","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["time series; dynamical systems; forecasting; predictor-corrector"]},"supplementary_material":{"value":"/attachment/00ba8ffc11922fd08a17bacf5e6018020fac0e07.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Learned time-series models, whether continuous- or discrete-time, are widely used to forecast the states of a dynamical system. Such models generate multi-step forecasts either directly, by predicting the full horizon at once, or iteratively, by feeding back their own predictions at each step. In both cases, the multi-step forecasts are prone to errors. To address this, we propose a Predictor-Corrector mechanism where the Predictor is any learned time-series model and the Corrector is a neural controlled differential equation. The Predictor forecasts, and the Corrector predicts the errors of the forecasts. Adding these errors to the forecasts improves forecast performance. The proposed Corrector works with irregularly sampled time series and continuous- and discrete-time Predictors. Additionally, we introduce two regularization strategies to improve the extrapolation performance of the Corrector with accelerated training. We evaluate our Corrector with diverse Predictors, e.g., neural ordinary differential equations, Contiformer, and DLinear, on synthetic, physics simulation, and real-world forecasting datasets. The experiments demonstrate that the Predictor-Corrector mechanism consistently improves the performance compared to Predictor alone."},"_bibtex":{"value":"@misc{\nshahid2026neural,\ntitle={Neural {CDE}s as Correctors for learned time series models},\nauthor={Muhammad Bilal Shahid and Prajwal Koirala and Cody Fleming},\nyear={2026},\nurl={https://openreview.net/forum?id=biZBdFpOzu}\n}"},"title":{"value":"Neural CDEs as Correctors for learned time series models"},"pdf":{"value":"/pdf/35c1008907c429a9f65e1fd76caa803968d4187c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"shahid|neural_cdes_as_correctors_for_learned_time_series_models"},"authorids":{"value":["~Muhammad_Bilal_Shahid1","~Prajwal_Koirala1","~Cody_Fleming2"]},"authors":{"value":["Muhammad Bilal Shahid","Prajwal Koirala","Cody Fleming"]}},"version":2},{"content":{"summary":{"value":"Current MLLMs can see but often don’t really understand the physical/social structure of the world. So they define that missing layer as “visual knowledge”, a zone between pixels and reasoning (gravity, affordances, etc). They build VKBench: 1,249 videos, 1,680 MCQs, 8 dimensions split into world-centric (physics, affordance, material, spatial) and human-centric (event anticipation, mental state, social relation, intention). They very deliberately filter out audio and language shortcuts so models can’t just guess from text. Then they show: even strong video MLLMs are worse than humans overall, and the gap is especially bad on world-centric stuff (intuitive physics, spatial). To show this isn’t hopeless, they build Video-VK+: basically Qwen2.5-VL-7B with a See–Think–Answer format + GRPO RL + a visual-knowledge reward that checks whether your description was good enough to answer."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Some tasks (event anticipation, social relation) clearly like longer context. Why did you fix at 32 instead of reporting a “long-context” track?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Proposes 8 tasks that map neatly to cognitive/vision literature.\n2. Complete anti-shortcut pipeline.\n3. The papers shows models are OK on human-centric but bad on world-centric.\n4. The benchmark is balanced and well-scoped."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. A lot of the difficulty comes from filtering existing datasets, not from filming new, adversarial, physics-centric videos, leading to any upstream video biases it originally uses.\n2. The visual-knowledge reward uses a frozen MLLM as verifier. That’s convenient, but it bakes the verifier’s biases right back into training.\n3. The paper shows correlations, but not what models actually get wrong?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915777518,"tcdate":1761987254741,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1472/Reviewer_LbT3"],"signatures":["ICLR.cc/2026/Conference/Submission1472/Reviewer_LbT3"],"forum":"P798W8Ag7L","number":3,"license":"CC BY 4.0","cdate":1761987254741,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1472/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915777518,"domain":"ICLR.cc/2026/Conference","replyto":"P798W8Ag7L","id":"g7iL4rLhwM","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multimodal Large Language Model","Video Large Language Model","Visual Knowledge","Benchmark","Datasets"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical and social principles. This capability, which we term visual knowledge, forms a bridge between perception and reasoning, yet remains an underexplored gap in current systems.\nTo systematically measure this capability, we present VKBench, a comprehensive video benchmark featuring 1,680 questions in 1,249 videos, covering eight core types of visual knowledge spanning both world-centric (e.g., intuitive physics) and human-centric (e.g., subjective intentions). Results show that leading models still fall short of human performance, with particularly notable gaps in world-centric visual knowledge.\nTo bridge this gap, we introduce VKQA, a new dataset, and Video-VK+, a baseline model that explicitly incorporates visual knowledge into MLLMs. Video-VK+ follows a structured See–Think–Answer format and adopts reinforcement learning with visual knowledge reward. This approach improves performance on VKBench by 3.7% and surpasses existing models on multiple video benchmarks.\nOur findings highlight visual knowledge as a key component for developing more robust and generalizable MLLMs that can not only see but also truly understand our world."},"_bibtex":{"value":"@misc{\njiang2025benchmarking,\ntitle={Benchmarking Visual Knowledge in Multimodal Large Language Models},\nauthor={Tianxiang Jiang and Sheng Xia and Yicheng Xu and Linquan Wu and Xiangyu Zeng and Limin Wang and Yu Qiao and Yi Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=P798W8Ag7L}\n}"},"title":{"value":"Benchmarking Visual Knowledge in Multimodal Large Language Models"},"pdf":{"value":"/pdf/186de1c1275cf653df49719895560101bf3f1938.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"jiang|benchmarking_visual_knowledge_in_multimodal_large_language_models"},"authorids":{"value":["~Tianxiang_Jiang1","~Sheng_Xia1","~Yicheng_Xu1","~Linquan_Wu1","~Xiangyu_Zeng4","~Limin_Wang1","~Yu_Qiao1","~Yi_Wang19"]},"authors":{"value":["Tianxiang Jiang","Sheng Xia","Yicheng Xu","Linquan Wu","Xiangyu Zeng","Limin Wang","Yu Qiao","Yi Wang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a new dataset for referrer image segmentation and a baseline on this dataset. The new dataset, namely RIS-CQ, built upon RefCOCO and Visual Genome contains complex text queries for each image. The complex text queries contain more specific information, generated from GPT3.5 and manually refined. Since other works performed poorly on the new dataset, they propose a dual-modality graph alignment model. The model makes use of two types of information: graph alignment of sense graph and text sequence, and feature alignment of image embedding and text embedding. The baseline mode outperforms previous work on the new dataset."},"soundness":{"value":"2 fair"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Please see Weaknesses."},"rating":{"value":"3: reject, not good enough"},"details_of_ethics_concerns":{"value":"(Minor) Since the text descriptions are generated from GPT3.5, it would be great to discuss the ethical problem that may be caused by generative AI."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"strengths":{"value":"- They provide a new dataset with complex text queries. The complex text queries are generated based on GPT 3.5 with inputs being the triplets of object relationships.\n- The baseline is intuitive and works well. They fuse multimodal information Graph alignment: scene graph and text sentence."},"flag_for_ethics_review":{"value":["Yes, Discrimination / bias / fairness concerns"]},"weaknesses":{"value":"1. Only a simple comparison with RefCOCOg is given. What are the advantages compared with RefCOCOg? Is there essential difference between query lengths of 8.43 and 13.18?\n2. The comparison with `PhraseCut' is missed, which is also a large-scale dataset with complex queries.\n3. The key in the proposed method is scene-graph-based cross-modal alignment. However, this way is widely used in cross-modal retrieval works.\n4. How to get mask annotations?\n5. VCTree is a method to generate bounding boxes. How to generate mask predictions?"}},"nonreaders":[],"tmdate":1699647575602,"tcdate":1699077406556,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission866/Reviewer_V6x1"],"signatures":["ICLR.cc/2024/Conference/Submission866/Reviewer_V6x1"],"forum":"SrzrUGWoRq","number":2,"license":"CC BY 4.0","cdate":1699077406556,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission866/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699647575602,"domain":"ICLR.cc/2024/Conference","replyto":"SrzrUGWoRq","id":"ERuBEJiPZa","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"We propose a novel benchmark dataset, RIS-CQ, which challenges the existing RIS with294 complex queries, and propose a novel SOTA method for RIS tasks called dual-modality alignment with graph learning (DUMOGA)."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Referring Image Segmentation; Complex Language Query; Dual-Modality Alignment"]},"supplementary_material":{"value":"/attachment/266777bea0c704815ea6141cc41449dbbd06ca19.pdf"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Referring Image Understanding (RIS) has been extensively studied over the past decade, leading to the development of advanced algorithms. However, there has been a lack of research investigating how existing algorithms should be benchmarked with complex language queries, which include more informative descriptions of surrounding objects and backgrounds (e.g., \"the black car.\" vs. \"the black car is parking on the road and beside the bus.\"). Given the significant improvement in the semantic understanding capability of large pre-trained models, it is crucial to take a step further in RIS by incorporating complex language that resembles real-world applications. To close this gap, building upon the existing RefCOCO and Visual Genome datasets, we propose a new RIS benchmark with complex queries, namely RIS-CQ. The RIS-CQ dataset is of high quality and large scale, which challenges the existing RIS with enriched, specific and informative queries, and enables a more realistic scenario of RIS research. Besides, we present a nichetargeting method to better task the RIS-CQ, called dual-modality graph alignment model (DuMoGa), which outperforms a series of RIS methods. To provide a valuable foundation for future advancements in the field of RIS with complex queries, we release the datasets, preprocessing and synthetic scripts, and the algorithm implementations."},"_bibtex":{"value":"@misc{\nji2024towards,\ntitle={Towards Complex-query Referring Image Segmentation: A Novel Benchmark},\nauthor={Wei Ji and Li Li and Hao Fei and Xiangyan Liu and Xun Yang and Juncheng Li and Roger Zimmermann},\nyear={2024},\nurl={https://openreview.net/forum?id=SrzrUGWoRq}\n}"},"title":{"value":"Towards Complex-query Referring Image Segmentation: A Novel Benchmark"},"pdf":{"value":"/pdf/7ab86767879af77091a4ddf50bf41d2452a7d87c.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"ji|towards_complexquery_referring_image_segmentation_a_novel_benchmark"},"authorids":{"value":["~Wei_Ji1","~Li_Li18","~Hao_Fei1","~Xiangyan_Liu1","~Xun_Yang1","~Juncheng_Li3","~Roger_Zimmermann1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wei Ji","Li Li","Hao Fei","Xiangyan Liu","Xun Yang","Juncheng Li","Roger Zimmermann"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Physically Inspired Neural Dynamics Symbolic Regression, a method for automatically discovering symbolic expressions that describe complex network dynamics. The authors present a two-part solution: a Physically Inspired Neural Dynamics component that augments and denoises trajectory data through interpolation, and a coordinated genetic search algorithm that derives symbolic expressions using references from the neural dynamics. The method addresses key limitations of existing approaches by handling noisy observations from multiple trajectories without requiring pre-defined function libraries or estimated time derivatives. The authors demonstrate PI-NDSR's effectiveness through evaluation on both synthetic datasets with various dynamics and real datasets on disease spreading, showing improvements in recovery probability and error rates compared to existing methods."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- What could be the potential issues when using the model for highly complex formulations?\n\n- For simulated datasets, would it be better to report both filtered and unfiltered MSE, or to use a separate metric for skeleton recovery?\n\n- It would be beneficial to expand the discussion on the limitations of the model."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The model introduced in this paper is groundbreaking, offering a method designed to overcome limitations of other approaches, such as high sensitivity to noise.\n\n- Comprehensive experiments are provided, comparing the proposed model with established methods across synthetic and real-world datasets.\n\n- The paper is generally well-written (except for the conclusion) with a clear and linear structure that makes it easy to follow the proposed methodology, experiments, and results. The explanations of the Physically Inspired Neural Dynamics Symbolic Regression framework and the challenges it addresses are presented in a logical sequence, enhancing reader comprehension."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The evaluation of the model is somewhat lacking, particularly in the choice of metrics for synthetic and real datasets, which is not fully explained. For example, filtering out models that don’t fully recover the skeleton in synthetic datasets may overlook valuable insights into how different methods handle structural errors. Often, a method’s ability to approximate the skeleton, even imperfectly, can still yield good performance on downstream tasks.\n\n- The conclusion is brief, as it mainly summarizes the process without emphasizing key findings. Improvements over existing methods are not highlighted, and it lacks an impact statement or discussion of broader implications. The future work section is vague; it’s unclear what is meant by “more complex systems” and “more real-world applications.” Although limitations are mentioned, it’s not evident why the method fails for highly complex formulations."}},"nonreaders":[],"tmdate":1731428022686,"tcdate":1730587330031,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4051/Reviewer_CABM"],"signatures":["ICLR.cc/2025/Conference/Submission4051/Reviewer_CABM"],"forum":"RdFpj6z4nE","number":3,"license":"CC BY 4.0","cdate":1730587330031,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4051/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428022686,"domain":"ICLR.cc/2025/Conference","replyto":"RdFpj6z4nE","id":"1MB0cGYHLV","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We propose PI-NDSR, a new neural network and genetic search algorithm for symbolic regression of complex network dynamics."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["network dynamics","symbolic regression","complex network"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex networks describe important structures in nature and society, composed of nodes and the edges that connect them. The evolution of these networks is typically described by dynamics, which are labor-intensive and require expert knowledge to derive. However, because the complex network involves noisy observations from multiple trajectories of nodes, existing symbolic regression methods are either not applicable or ineffective on its dynamics. In this paper, we propose Physically Inspired Neural Dynamics Symbolic Regression (PI-NDSR), a method based on neural networks and genetic programming to automatically learn the symbolic expression of dynamics. Our method consists of two key components: a Physically Inspired Neural Dynamics (PIND) to augment and denoise trajectories through observed trajectory interpolation; and a coordinated genetic search algorithm to derive symbolic expressions. This algorithm leverages references of node dynamics and edge dynamics from neural dynamics to avoid overfitted expressions in symbolic space. We evaluate our method on synthetic datasets generated by various dynamics and real datasets on disease spreading. The results demonstrate that PI-NDSR outperforms the existing method in terms of both recovery probability and error."},"_bibtex":{"value":"@misc{\nqiu2025neural,\ntitle={Neural Symbolic Regression of Complex Network Dynamics},\nauthor={Haiquan Qiu and Shuzhi Liu and Yong Li and Quanming Yao},\nyear={2025},\nurl={https://openreview.net/forum?id=RdFpj6z4nE}\n}"},"title":{"value":"Neural Symbolic Regression of Complex Network Dynamics"},"pdf":{"value":"/pdf/065a6edf214767f40ec312fa840284169100c3e0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"qiu|neural_symbolic_regression_of_complex_network_dynamics"},"authorids":{"value":["~Haiquan_Qiu1","~Shuzhi_Liu1","~Yong_Li7","~Quanming_Yao3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haiquan Qiu","Shuzhi Liu","Yong Li","Quanming Yao"]}},"version":2},{"content":{"summary":{"value":"The work proposes a complex-value deep neural network for computing vision tasks. First, the authors propose an inevitable real-to-complex transformation. Then, the work proposes an architecture comprising spectral convolution and a complex T2T module. \nThe authors evaluated their model on image classification, smooth object detection, and defocus blur detection."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Line 317: Citation link broken for T2T-ViTYuan et al. (2021)\n\n2. Line 269: Do you use complex-valued  “ normalization” as discussed in [1]\n\n3. How well does the model perform if we consider trivial real to complex conversion that considers the real numbers as complex numbers with $0$ imaginary part? This is also a very crucial ablation that the authors should perform.\n\n4. Does using spectral convolution make it challenging to capture local features as it performs global convolution?\n\n\n[1] DEEP COMPLEX NETWORKS"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The work proposes a novel invertible real-to-complex conversion for RGB images in complex-valued neural networks. The procedures are clearly stated using pseudocode and figures."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The authors seem to miss the seminal work on complex values networks and didn’t compare/discuss with the techniques discussed in [1]\n2. The work is directed at using complex-valued networks for real-valued images. The majority of the paper involves devising an invertible conversion from real to complex representation. However, the paper fails to demonstrate its utility. For example, for classification on ImageNets, the authors did not consider state-of-the-art models, such as Vit, Swin-v2, etc., which achieve above 90% accuracy. \n\n3. The paper does not discuss the motivation of the specific real to complex conversion. There are many invertible conversions between real and complex.\n\n[1] DEEP COMPLEX NETWORKS"}},"nonreaders":[],"tmdate":1731427401786,"tcdate":1730527140121,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1276/Reviewer_F89z"],"signatures":["ICLR.cc/2025/Conference/Submission1276/Reviewer_F89z"],"forum":"9hmDl8fFDs","number":3,"license":"CC BY 4.0","cdate":1730527140121,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1276/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427401786,"domain":"ICLR.cc/2025/Conference","replyto":"9hmDl8fFDs","id":"c03jsl6RYW","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"A robust complex-valued approach in Spatio-spectral domain for multiple tasks on both real and complex data."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Deep Complex Newtworks","Complex-valued color transformation"]},"supplementary_material":{"value":"/attachment/d3c1bebc77a7f4edd23a402e6580595b34f472e4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex-valued neural networks have attracted growing attention for their ability to handle complex-valued data with enhanced representational capacity. However, their potential in computer vision remains relatively untapped. \nIn this paper, we introduce Deep Complex Spatio-Spectral Network (DCSNet), a fully complex-valued token-based, end-to-end neural network designed for binary segmentation tasks. Additionally, our DCSNet encoder can be used for image classification in the complex domain. We also propose an invertible real-to-complex (R2C) transform, which generates two complex-valued input channels, complex intensity and complex hue, while producing complex-valued images with distinct real and imaginary components.\nDCSNet operates in both spatial and spectral domains by leveraging complex-valued inputs and complex Fourier transform.\nAs a result, the complex-valued representation is maintained throughout DCSNet, and we avoid the information loss typically associated with Real$\\leftrightarrow$Complex transformations. Extensive experiments show that DCSNet surpasses existing complex-valued methods across various tasks on both real and complex-valued data and achieves competitive performance compared to existing real-valued methods, establishing a robust framework for handling both data types effectively."},"_bibtex":{"value":"@misc{\nyadav2025deep,\ntitle={Deep Complex Spatio-Spectral Networks with Complex Visual Inputs},\nauthor={Saurabh Yadav and Koteswar Rao Jerripothula},\nyear={2025},\nurl={https://openreview.net/forum?id=9hmDl8fFDs}\n}"},"title":{"value":"Deep Complex Spatio-Spectral Networks with Complex Visual Inputs"},"pdf":{"value":"/pdf/78685d9476ab42033341d680f3e952693b0cbc64.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yadav|deep_complex_spatiospectral_networks_with_complex_visual_inputs"},"authorids":{"value":["~Saurabh_Yadav2","~Koteswar_Rao_Jerripothula3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Saurabh Yadav","Koteswar Rao Jerripothula"]}},"version":2},{"content":{"venue":{"value":"MICCAI (2) 2023"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-43895-0_42.pdf"},"venueid":{"value":"dblp.org/conf/MICCAI/2023"},"paperhash":{"value":"heo|physicsbased_decoding_improves_magnetic_resonance_fingerprinting"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Juyeon_Heo:","https://dblp.org/search/pid/api?q=author:Pingfan_Song:","~Weiyang_Liu1","https://dblp.org/search/pid/api?q=author:Adrian_Weller:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-43895-0_42"},"_bibtex":{"value":"@inproceedings{DBLP:conf/miccai/HeoSLW23,\n  author={Juyeon Heo and Pingfan Song and Weiyang Liu and Adrian Weller},\n  title={Physics-Based Decoding Improves Magnetic Resonance Fingerprinting},\n  year={2023},\n  cdate={1672531200000},\n  pages={446-456},\n  url={https://doi.org/10.1007/978-3-031-43895-0_42},\n  booktitle={MICCAI (2)},\n  crossref={conf/miccai/2023-2}\n}\n"},"abstract":{"value":"Magnetic Resonance Fingerprinting (MRF) is a promising approach for fast Quantitative Magnetic Resonance Imaging (QMRI). However, existing MRF methods suffer from slow imaging speeds and poor generalization performance on radio frequency pulse sequences generated in various scenarios. To address these issues, we propose a novel MRI physics-informed regularization for MRF. The proposed approach adopts a supervised encoder-decoder framework, where the encoder performs the main task, i.e. predicting the target tissue properties from input magnetic responses, and the decoder servers as a regularization via reconstructing the inputs from the estimated tissue properties using a Bloch-equation based MRF physics model. The physics-based decoder improves the generalization performance and uniform stability by a considerable margin in practical out-of-distribution settings. Extensive experiments verified the effectiveness of the proposed approach and achieved state-of-the-art performance on tissue property estimation."},"title":{"value":"Physics-Based Decoding Improves Magnetic Resonance Fingerprinting"},"authors":{"value":["Juyeon Heo","Pingfan Song","Weiyang Liu","Adrian Weller"]}},"tmdate":1747344778442,"pdate":1672531200000,"tcdate":1747344762713,"writers":["~"],"signatures":["~Weiyang_Liu1"],"forum":"TGrPGTUcFa","license":"CC BY-SA 4.0","number":498971,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747344778442,"domain":"DBLP.org","id":"TGrPGTUcFa","version":2},{"content":{"venue":{"value":"FIMH (2) 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-94562-5_4.pdf"},"venueid":{"value":"dblp.org/conf/FIMH/2025"},"paperhash":{"value":"magaña|\\vardelta_poissonn_learning_atrial_activation_map_from_the_ecg_with_physicsinformed_neural_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Efraín_Magaña:","https://dblp.org/search/pid/api?q=author:Elena_Zappon:","https://dblp.org/search/pid/api?q=author:Gernot_Plank:","~Simone_Pezzuto1","https://dblp.org/search/pid/api?q=author:Francisco_Sahli_Costabal:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-94562-5_4"},"_bibtex":{"value":"@inproceedings{DBLP:conf/fimh/MaganaZPPC25,\n  author={Efraín Magaña and Elena Zappon and Gernot Plank and Simone Pezzuto and Francisco Sahli Costabal},\n  title={$\\varDelta $-PoIssoNN: Learning Atrial Activation Map from the ECG with Physics-Informed Neural Networks},\n  year={2025},\n  cdate={1735689600000},\n  pages={31-40},\n  url={https://doi.org/10.1007/978-3-031-94562-5_4},\n  booktitle={FIMH (2)},\n  crossref={conf/fimh/2025-2}\n}\n"},"abstract":{"value":"Cardiac digital twins have shown promise to personalize treatments. However, there are multiple challenges to incorporate patient-specific information from non-invasive data. For instance, recovering the activation sequence in atria from the standard electrocardiogram (ECG) remains elusive. Recent studies have tackled this task on the ventricles, where the ECG signal is much stronger. This work presents a novel methodology to recover the atrial electrical activity with physics-informed neural networks. Instead of focusing on the activation times, we predict the direction of propagation of the electrical wave at each point with a neural network. Then, by solving a linear system for the Poisson equation, we recover the activation times that satisfy the anisotropic eikonal equation. The proposed methodology is compared with a methodology that predicts directly the electrical propagation and does not enforce the propagation model. We compare it to a traditional physics-informed neural network formulation, where the eikonal equation is only weakly imposed. We validate our methodology in a biatrial synthetic case using realistic lead fields for ECG calculation. We then learn the activation sequence from patient data, recovering a physiological activation pattern. We believe this is a first step toward digital twinning of the atria."},"title":{"value":"$\\varDelta $-PoIssoNN: Learning Atrial Activation Map from the ECG with Physics-Informed Neural Networks"},"authors":{"value":["Efraín Magaña","Elena Zappon","Gernot Plank","Simone Pezzuto","Francisco Sahli Costabal"]}},"tmdate":1756313183016,"pdate":1735689600000,"externalIds":["dblp:conf/fimh/MaganaZPPC25"],"tcdate":1756313143431,"writers":["~"],"signatures":["~Simone_Pezzuto1"],"forum":"9AznCjtUhT","license":"CC BY-SA 4.0","number":618609,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1756313183016,"domain":"DBLP.org","id":"9AznCjtUhT","version":2},{"content":{"summary":{"value":"The paper introduces **SAIR**, a million-scale synthetic structure–activity dataset ($\\approx$ 1.05M protein–ligand systems) built by curating bioactivity from ChEMBL and BindingDB, then generating 3D complexes using Boltz-1x co-folding. The curation removes entries with missing identifiers, validity flags, inequalities, and measurements outside a defined dynamic range; ligands are standardized (desalting, neutral-pH protonation, RDKit canonical SMILES). To avoid leakage from experimental structures, (UniProt, CCD) pairs present in PDB are excluded. The authors assess structural plausibility with PoseBusters, report overall low failure rates (only $\\approx$ 0.53% of complexes have all five poses failing), analyze confidence metrics (iPTM, interaction-PTM, complex iPLDDT) vs potency, and benchmark a small set of affinity predictors (Vina/Vinardo, OnionNet-2, AEV-PLIG)."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. **Affinity aggregation strategy over five structures.** For each protein–ligand pair, do you report model performance per-structure, or do you aggregate (mean/min/best) across the five predicted structures?\n2. **Pocket diversity in main text.** I saw that there are no information about pocket diversity in the main text despite its detailed explanation is provided in appendix. In my opinion, high pocket diversity for several proteins may introduce unintended bias in data. I'd like to know about the authors opinion of this issue. If this is problematic, giving diversity of protein pocket and let the users handling about that issue themselves can be a solution.\n3. **Cofactors and tiny ligands.** How are cofactors/ions handled (e.g., heavy-atom count near 1, as described in table 3)? Are such entries filtered or annotated, and how do they affect pocket identification and affinity correlations?\n4. **About the degradation of performance of OnionNet-2.** In Fig. 6, OnionNet-2 attains only Pearson r $\\approx$ 0.35 on SAIR, whereas its reported performance on the CASF-2016 scoring set is r $\\approx$ 0.864. Given that OnionNet-2 relies primarily on protein–ligand contact patterns, this large gap suggests potential issues with pocket localization/contacts in the co-folded structures."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- **New dataset for proper demand.** If fully released, SAIR could be a valuable pretraining/analysis data source for structure-based binding affinity prediction models. Since the major limitation of training PLI models is lack of paired data, this approach is suitable in terms of its purpose and time.\n- **Scale and transparency.** Clear end-to-end pipeline from bioactivity ingestion to synthetic structure generation; explicit de-duplication and PDB-leakage control. As the authors stated, the computational cost for making this data is very expensive and thus this dataset is invaluable.\n- **Interesting diagnostics.** Introduction/usage of interface-focused confidence metrics (interaction-PTM, iPTM) and assay-aware correlation analyses (biochemical > cell)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Despite the strengths of the proposed dataset, my major concern is rooted in the fact that creating a synthetic structure-affinity dataset should minimize uncertainty in both structure and affinity.\n1. **More rigorous validation of predicted structure.** All complexes are treated as monomers with canonical UniProt sequences, despite many targets (e.g., GPCRs) being oligomeric or using non-canonical constructs. Recent studies show that co-folding models can predict plausible-looking yet physically inconsistent pockets/poses (including persistent ligand placement after pocket disruption), yielding false positives in downstream analyses. [1,2] Thus, I suggest that the authors include an additional assessment of pocket validity. For instance, they could examine whether the predicted pocket residues are consistent with experimentally validated binding sites observed in highly similar structures associated with proteins with similar sequence, especially for the proteins from rare protein family.\n2. **Potential distillation bias.** Because Boltz-1x uses deterministic MSA subsampling and results improve most on a high-confidence subset, the dataset may over-represent 'easy' (well-constrained/high-MSA or training-set–similar) proteins and under-represent rare/low-MSA targets. I would welcome analyses quantifying how confidence filtering shifts family/MSA composition and whether conclusions hold under family-balanced or MSA-balanced evaluations. I think this may results in making false positive data for rare protein family.\n3. **Confidence wording vs effect size.** In the page 7's first and second paragraphs, calling $r_{s}\\approx 0.25$ \"strong\" overstates the effect; the paper's own narrative elsewhere acknowledges low absolute correlations comparable to interface metrics.\n\n[1] Masters, Matthew R., Amr H. Mahmoud, and Markus A. Lill. \"Investigating whether deep learning models for co-folding learn the physics of protein-ligand interactions.\" _Nature Communications_ 16.1 (2025): 8854.\n[2] Škrinjar, Peter, et al. \"Have protein-ligand co-folding methods moved beyond memorisation?.\" _BioRxiv_ (2025): 2025-02."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923248597,"tcdate":1761754798011,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12325/Reviewer_rhF2"],"signatures":["ICLR.cc/2026/Conference/Submission12325/Reviewer_rhF2"],"forum":"qgk2F6jxH4","number":2,"license":"CC BY 4.0","cdate":1761754798011,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12325/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923248597,"domain":"ICLR.cc/2026/Conference","replyto":"qgk2F6jxH4","id":"UbO8aCAl0o","forumContent":{"TLDR":{"value":"We introduce SAIR, the largest public dataset of protein-ligand 3D structures with activity data"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Protein","Ligand","Dataset","Affinity"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Accurate prediction of protein-ligand binding affinities remains a cornerstone problem in drug discovery. While binding affinity is inherently dictated by the 3D structure and dynamics of protein-ligand complexes, current deep learning approaches are limited by the lack of high-quality experimental structures with annotated binding affinities. To address this limitation, we introduce the Structurally Augmented IC50 Repository (SAIR), the largest publicly available dataset of protein-ligand 3D structures with associated activity data. The dataset comprises $5,244,285$ structures across $1,048,857$ unique protein-ligand systems, curated from the ChEMBL and BindingDB databases, which were then computationally folded using the Boltz-1x model. We provide a comprehensive characterization of the dataset, including distributional statistics of proteins and ligands, and evaluate the structural fidelity of the folded complexes using PoseBusters. Our analysis reveals that approximately $3 \\%$ of structures exhibit physical anomalies, predominantly related to internal energy violations. As an initial demonstration, we benchmark several binding affinity prediction methods, including empirical scoring functions (Vina, Vinardo), a 3D convolutional neural network (Onionnet-2), and a graph neural network (AEV-PLIG). While machine learning-based models consistently outperform traditional scoring function methods, neither exhibit a high correlation with ground truth affinities, highlighting the need for models specifically fine-tuned to synthetic structure distributions. This work provides a foundation for developing and evaluating next-generation structure and binding-affinity prediction models and offers insights into the structural and physical underpinnings of protein-ligand interactions. \nThe link to the data will be added upon publication, to preserve anonymity of the submission."},"_bibtex":{"value":"@inproceedings{\nlemos2026sair,\ntitle={{SAIR}: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset},\nauthor={Pablo Lemos and Zane Beckwith and Sasaank Bandi and Maarten Van Damme and Jordan Crivelli-Decker and Benjamin J. Shields and Thomas Merth and Punit K Jha and Nicola De Mitri and Tiffany Callahan and AJ Nish and Paul Abruzzo and Romelia Salomon-Ferrer and Martin Ganahl},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=qgk2F6jxH4}\n}"},"title":{"value":"SAIR: Enabling Deep Learning for Protein-Ligand Interactions with a Synthetic Structural Dataset"},"pdf":{"value":"/pdf/65927dac1971e4118c090555b328a21f175367de.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lemos|sair_enabling_deep_learning_for_proteinligand_interactions_with_a_synthetic_structural_dataset"},"authorids":{"value":["~Pablo_Lemos1","~Zane_Beckwith1","~Sasaank_Bandi1","~Maarten_Van_Damme1","~Jordan_Crivelli-Decker1","~Benjamin_J._Shields1","~Thomas_Merth1","~Punit_K_Jha1","~Nicola_De_Mitri1","~Tiffany_Callahan1","~AJ_Nish1","~Paul_Abruzzo1","~Romelia_Salomon-Ferrer1","~Martin_Ganahl1"]},"authors":{"value":["Pablo Lemos","Zane Beckwith","Sasaank Bandi","Maarten Van Damme","Jordan Crivelli-Decker","Benjamin J. Shields","Thomas Merth","Punit K Jha","Nicola De Mitri","Tiffany Callahan","AJ Nish","Paul Abruzzo","Romelia Salomon-Ferrer","Martin Ganahl"]}},"version":2},{"content":{"venue":{"value":"ICRA 2015"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/7128761/7138973/07139730.pdf"},"venueid":{"value":"dblp.org/conf/ICRA/2015"},"paperhash":{"value":"kim|physicsbased_hand_interaction_with_virtual_objects"},"authorids":{"value":["~Jun-Sik_Kim2","https://dblp.org/search/pid/api?q=author:Jung-Min_Park:"]},"html":{"value":"https://doi.org/10.1109/ICRA.2015.7139730"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icra/KimP15,\n  author={Jun-Sik Kim and Jung-Min Park},\n  title={Physics-based hand interaction with virtual objects},\n  year={2015},\n  cdate={1420070400000},\n  pages={3814-3819},\n  url={https://doi.org/10.1109/ICRA.2015.7139730},\n  booktitle={ICRA},\n  crossref={conf/icra/2015}\n}\n"},"abstract":{"value":"We propose an intuitive user interaction system with virtual objects by directly using user's hands through a physics simulation. This system aims to provide a more direct way to manipulate and interact with virtual objects and to induce more realistic reactions of the virtual objects in the interaction than the existing methods. The key contribution of this paper is to model the deformations of the user's hands in two ways: global deformation by the hand poses and the local deformation by contacts with objects. By the two deformation models, objects can be manipulated with better stability and robustness without any predefinition of interaction methods depending on the object shapes, by solely using a physics simulation. The experiment shows that the proposed method enables a user to intuitively manipulate virtual object accurately."},"title":{"value":"Physics-based hand interaction with virtual objects"},"authors":{"value":["Jun-Sik Kim","Jung-Min Park"]}},"tmdate":1731563701616,"pdate":1420070400000,"tcdate":1731563695216,"writers":["~"],"signatures":["~Jun-Sik_Kim2"],"forum":"TjTMi7ZmxR","license":"CC BY-SA 4.0","number":238223,"cdate":1420070400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731563701616,"domain":"DBLP.org","id":"TjTMi7ZmxR","version":2},{"content":{"venue":{"value":"MILCOM 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10773620/10773624/10773796.pdf"},"venueid":{"value":"dblp.org/conf/MILCOM/2024"},"paperhash":{"value":"wang|adapting_complex_event_detection_to_perceptual_domain_shifts"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Brian_Wang:","https://dblp.org/search/pid/api?q=author:Julian_de_Gortari_Briseno:","https://dblp.org/search/pid/api?q=author:Liying_Han:","https://dblp.org/search/pid/api?q=author:Henry_Phillips:","https://dblp.org/search/pid/api?q=author:Jeffrey_Craighead:","https://dblp.org/search/pid/api?q=author:Ben_Purman:","~Lance_M._Kaplan1","https://dblp.org/search/pid/api?q=author:Mani_Srivastava_0001:"]},"html":{"value":"https://doi.org/10.1109/MILCOM61039.2024.10773796"},"_bibtex":{"value":"@inproceedings{DBLP:conf/milcom/WangBHPCPK024,\n  author={Brian Wang and Julian de Gortari Briseno and Liying Han and Henry Phillips and Jeffrey Craighead and Ben Purman and Lance M. Kaplan and Mani Srivastava},\n  title={Adapting Complex Event Detection to Perceptual Domain Shifts},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-6},\n  url={https://doi.org/10.1109/MILCOM61039.2024.10773796},\n  booktitle={MILCOM},\n  crossref={conf/milcom/2024}\n}\n"},"abstract":{"value":"Human decision-making, as well as control of autonomous systems, have deployed deep learning models for detecting complex events from unstructured sensory data. However, the strong performance of these models is restricted to events with short intervals of time and space due to the limited context memory of their architectures. Thus, detecting events that transpire over long periods of time with multiple spatially distant sensor sources (known as complex events) remains challenging for these purely neural-based methods, particularly as environmental conditions and object appearances change. In recent years, neurosymbolic approaches have been proposed that use both neural-based perception and symbolic reasoning for capturing complex events. However, these approaches still face issues of adaptation to perceptual domain shift in complex events. We address these problems in the context of a prototype neurosymbolic system called DANCER, which performs Domain Adaptation and Neurosymbolic inference in Complex Event Reasoning. DANCER aims to provide domain adaptation in a post-deployment setting while minimizing runtime user burden for annotation. To enable training and evaluation of DANCER, we also provide a physics-based synthetic sensor data generator to create videos given complex scenario specifications. We evaluate DANCER on a dataset of generated synthetic data. We show that DANCER yields a 48% increase in accuracy of complex event detection using domain adaptation while significantly reducing the annotation time of our synthetic complex events by up to 2.7x, demonstrating DANCER’s ability to effectively detect complex events under perceptual domain shift."},"title":{"value":"Adapting Complex Event Detection to Perceptual Domain Shifts"},"authors":{"value":["Brian Wang","Julian de Gortari Briseno","Liying Han","Henry Phillips","Jeffrey Craighead","Ben Purman","Lance M. Kaplan","Mani Srivastava"]}},"tmdate":1747251960368,"pdate":1704067200000,"tcdate":1747251948038,"writers":["~"],"signatures":["~Lance_M._Kaplan1"],"forum":"JV0YPjxLmc","license":"CC BY-SA 4.0","number":465614,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747251960368,"domain":"DBLP.org","id":"JV0YPjxLmc","version":2},{"content":{"summary":{"value":"DiffuPhyGS presents an end-to-end text-to-video pipeline that generates 3D Gaussian objects with physics-driven motion. The system refines prompts using an LLM loop, stabilizes geometry with multi-view diffusion guidance and Densification-by-Adaptive-Splitting, and learns per-Gaussian material mixtures. Motion is driven through an MPM simulator using both implicit diffusion cues and an explicit velocity loss. On a small set of handcrafted prompts, the method reports higher metrics scores and is preferred in a small user study over PhysDreamer, OmniPhysGS, and PhysGaussian."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses"},"rating":{"value":2},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper tackles joint generation of text-aligned appearance and dynamics in a single 3D Gaussian framework, combining prompt processing, diffusion-based 3D synthesis, and differentiable physics.\n- Leveraging video diffusion priors with Score Distillation Sampling (SDS) to model motion and material is an interesting idea."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The evaluation is very limited, four template prompts plus qualitative figures, without standardized datasets or complex multi-object interactions. It’s hard to assess generalization beyond curated cases.\n- Baselines depend on geometry produced by DiffuPhyGS, and the main metrics blend appearance and motion. There are no physics-grounded metrics, so claims of physical fidelity aren’t well supported.\n- The 3D generation component lags behind recent text-to-3D/image-to-3D methods. SDS-based approaches tend to produce lower quality and run slowly for both geometry and appearance, which makes it hard to judge the pipeline’s full potential since results are limited by the underlying generator.\n- There isn’t a clear novelty claim on the 3D generation side. Using 2D and multi-view diffusion with SDS for 3D Gaussians has been explored before, and the LLM-CoT-IPR module mainly uses an existing LLM to refine prompts without introducing new techniques.\n- I would suggest the authors apply the motion and material learning techniques on more recent and higher-quality 3D generation methods such as Trellis[1].\n\n[1] Structured 3D Latents for Scalable and Versatile 3D Generation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917103218,"tcdate":1761996162986,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3940/Reviewer_hkzr"],"signatures":["ICLR.cc/2026/Conference/Submission3940/Reviewer_hkzr"],"forum":"mq43BAAos0","number":3,"license":"CC BY 4.0","cdate":1761996162986,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3940/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917103218,"domain":"ICLR.cc/2026/Conference","replyto":"mq43BAAos0","id":"ja9CjYNkQ3","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Text-to-Video","Gaussian Splatting","Diffusion Model","Dynamic 3D Generation","LLM"]},"supplementary_material":{"value":"/attachment/7d953bf9100e1b07d3a33399bdb531b407c0d505.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Generating realistic 3D object videos is crucial for virtual reality and digital content creation. However, existing 3D dynamics generation methods often struggle to achieve high-quality appearance and physics-aware motion, relying on manual inputs and pre-existing models. To address these challenges, we propose DiffuPhyGS, a novel framework that generates high-quality 3D objects with realistic and learnable physical motion directly from text prompts. Our approach features an LLM-Chain-of-Thought-based Iterative Prompt Refinement (LLM-CoT-IPR) method, which obtains prompt-aligned 2D and multi-view 3D diffusion priors to guide Gaussian Splatting (GS) to generate 3D objects. We further enhance 3D generation quality with a Densification-by-Adaptive-Splitting (DAS) mechanism. Next, we employ a material property decoder that utilizes a Mixture-of-Experts Material Constitutive Models (MoEMCMs) to predict the mixed material properties of the 3D object. We then apply the Material Point Method (MPM) to deform 3D Gaussian kernels, ensuring physics-grounded motion guided by implicit and explicit physical priors from the video diffusion model and a velocity loss function. Extensive experiments show DiffuPhyGS outperforms other methods in generating realistic physics-grounded motion across diverse materials."},"_bibtex":{"value":"@misc{\nwang2025diffuphygs,\ntitle={DiffuPhy{GS}: Text-to-Video Generation with 3D Gaussians and Learnable Physical Properties via Diffusion Priors},\nauthor={Wenqing Wang and Yun Fu},\nyear={2025},\nurl={https://openreview.net/forum?id=mq43BAAos0}\n}"},"title":{"value":"DiffuPhyGS: Text-to-Video Generation with 3D Gaussians and Learnable Physical Properties via Diffusion Priors"},"pdf":{"value":"/pdf/47a9347799011259548f308f32392a5726638e36.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|diffuphygs_texttovideo_generation_with_3d_gaussians_and_learnable_physical_properties_via_diffusion_priors"},"authorids":{"value":["~Wenqing_Wang4","~Yun_Fu1"]},"authors":{"value":["Wenqing Wang","Yun Fu"]}},"version":2},{"content":{"summary":{"value":"The authors propose a SO(3)-equivariant network operating on scalars, vectors and 2-tensors. They identify the corresponding equivariant linear layers and come up with a mixing strategy to mix different representations. Furthermore, SO(2)-equivariant linear layers are proposed to allow the scenario that SO(3) symmetry breaks into a subgroup SO(2) axial symmetry along a specific axis $\\hat{j}$.  Finally, they demonstrate how it is appllied to b-tagging in High Energy Physics (HEP), where the data is rotational symmetric w.r.t. jet axis $\\hat{j}$."},"soundness":{"value":"3 good"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"- How to understand Line 54 \"We show that this kind of equivariant neuron is generally only possible with the introduction of order-2 tensor representations\"? For me it seems everything should work well even if we only use scalars and vectors (like VectorPFN). "},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"strengths":{"value":"- The authors provide a good and simple implementation of SO(3) equivariant networks on scalars, vectors and 2-tensors. The weight matrix is carefully designed to preserve the symmetry. The way to mix different representations is also intuitive and easy to understand. \n- The discussion on axial symmetry is well motivated and easy to follow, and the analysis is solid and sound. "},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- [Novelty] The idea of tensor-product-based representations is not new [1]. The main method (except the SO(2) part) looks a simple variation and the technique involved is pretty standard. Add discussion and comparison with existing tensor-product-representation-based methods could make the work more solid.\n- [Evaluation] To show the effectiveness of mixing and SO(2) linear layers, I think it is better to put more intermidiate results (e.g., w/ and w/o SO(2) linear layers) in the main table. \n- [Minor issues]: Eq. (1, 2, 3) use Einstein summation without declaration, which may cause confusion to readers without physics background. Line 168 should be \"isotropic linear neuron of Eq. (2)\" instead of \"Eq. (1)\".\n\n[1] Finkelshtein, Ben, et al. \"A simple and universal rotation equivariant point-cloud network.\" _Topological, Algebraic and Geometric Learning Workshops 2022_. PMLR, 2022.\n"},"limitations":{"value":"Limitations are not included in the manuscript."}},"nonreaders":[],"tmdate":1702411514055,"tcdate":1688482454378,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission15315/Reviewer_ybFn"],"signatures":["NeurIPS.cc/2023/Conference/Submission15315/Reviewer_ybFn"],"forum":"Ai40Gvt2wj","number":3,"license":"CC BY 4.0","cdate":1688482454378,"mdate":1702411514055,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission15315/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"Ai40Gvt2wj","id":"MLMurWUNDE","forumContent":{"venue":{"value":"Submitted to NeurIPS 2023"},"keywords":{"value":["equivariance","so(3) symmetry","tensor data","physics"]},"_bibtex":{"value":"@misc{\nshimmin2023rethinking,\ntitle={Rethinking {SO}(3)-equivariance with Bilinear Tensor Networks},\nauthor={Chase Shimmin and Zhelun Li and Ema Catalina Smith},\nyear={2023},\nurl={https://openreview.net/forum?id=Ai40Gvt2wj}\n}"},"title":{"value":"Rethinking SO(3)-equivariance with Bilinear Tensor Networks"},"paperhash":{"value":"shimmin|rethinking_so3equivariance_with_bilinear_tensor_networks"},"TLDR":{"value":"We design a modular, equivariant neural architecture that generalizes affine layers to SO(3) representations and exploits expressive bilinear operations, to improve learning on scalar, vector, and tensor-valued data."},"abstract":{"value":"Many datasets in scientific and engineering applications are comprised of objects which have\nspecific geometric structure. A common example is data which inhabits a representation of the\ngroup SO(3) of 3D rotations: scalars, vectors, tensors, etc. One way for a neural network to\nexploit prior knowledge of this structure is to enforce SO(3)-equivariance throughout its layers, and\nseveral such architectures have been proposed. While general methods for handling arbitrary SO(3)\nrepresentations exist, they computationally intensive and complicated to implement. We show that\nby judicious symmetry breaking, we can efficiently increase the expressiveness of a network operating\nonly on vector and order-2 tensor representations of SO(2). We demonstrate the method on an\nimportant problem from High Energy Physics known as b-tagging, where particle jets originating\nfrom b-meson decays must be discriminated from an overwhelming QCD background. In this task,\nwe find that augmenting a standard architecture with our method results in a 2.3× improvement in\nrejection score."},"pdf":{"value":"/pdf/fea1c6b4e1bec8a75690eee435b79c521a758dcd.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference/Rejected_Submission"},"authorids":{"value":["~Chase_Shimmin1","~Zhelun_Li2","~Ema_Catalina_Smith1"]},"authors":{"value":["Chase Shimmin","Zhelun Li","Ema Catalina Smith"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a neuro-symbolic framework that integrates symbolic regression with trajectory-guided image-to-video models to achieve physics-grounded video generation. The method extracts motion trajectories from input videos using CoTracker, discovers governing equations through a novel retrieval-based symbolic regression approach (ReSR) initialized with a curated physics equation bank, and uses the discovered equations to forecast future trajectories that guide video generation models without fine-tuning. The framework is evaluated on classical mechanics systems including spring-mass oscillators, pendulums, and projectile motion, demonstrating improved equation recovery and physical alignment compared to baseline methods."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Your equation bank construction (Section 3.3) involves manually replacing time-dependent variables with \"t\" and time-independent variables with constants like \"10\". Can you provide evidence that this substitution process preserves the mathematical structure relevant for trajectory fitting? \n\n\n2. Your trajectory extraction selects the \"top-5 trajectories with the highest temporal variance\" from a 10×10 grid (Section 3.2), but Table 2 shows substantially degraded performance on real initial frames versus synthetic ones (e.g., Kling: FVD 1064 vs 641, TraEr 404 vs 325). Can you quantify how often your heuristic correctly identifies the object of interest in your real-world test cases?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. Novel neuro-symbolic integration for video generation. The paper presents an innovative approach combining interpretable symbolic equation discovery with data-driven video generation models, bridging two typically separate domains. The ReSR method with retrieval-based pre-training from a physics equation bank is a well-motivated contribution that significantly improves convergence speed (44.3 iterations vs 61.4 for PySR) and equation accuracy (TED 0.80 vs 0.47) as shown in Table 1.\n2. Comprehensive evaluation framework with meaningful metrics. The paper employs multiple complementary evaluation metrics including symbolic similarity (TED), trajectory error (MSE), convergence speed (ITB), and video quality measures (FVD, FID, smoothness, trajectory error), alongside human evaluation studies with substantial inter-annotator agreement (Fleiss' Kappa 0.73). The ablation studies systematically analyze the impact of initialization weight w and training/test split proportions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Equation bank construction relies heavily on manual curation. Section 3.3 describes constructing the equation bank by manually adapting 106 Feynman equations through time-variable substitution where \"time-dependent variables (e.g., velocity, acceleration, momentum) with the time variable t\" and \"variables that are independent of time (e.g., mass, density) are replaced with constant values (e.g., 10).\" \n\n2. The trajectory extraction pipeline uses heuristic selection. The paper extracts trajectories by sampling a uniform 10×10 grid and selecting \"top-k trajectories with the highest motion magnitude\" based on \"temporal variance\" (Section 3.2), motivated by the observation that \"target objects in physics-driven videos typically exhibit the most motion.\" However, this heuristic may fail when background elements move more than objects of interest (e.g., camera motion, wind effects on curtains), or when the physical system involves multiple interacting objects with different motion scales. \n\n3. Limited scope to idealized classical mechanics scenarios restricts practical applicability. The paper acknowledges in Section 4.1 that evaluation is limited to \"classical physics systems\" in \"controlled laboratory environment\" because it enables \"direct evaluation against ground-truth equations.\" While this is reasonable for proof-of-concept, the scenarios tested (spring-mass, pendulums, projectile motion) represent highly simplified physics with analytical solutions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918834301,"tcdate":1761940540856,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6442/Reviewer_Jxx2"],"signatures":["ICLR.cc/2026/Conference/Submission6442/Reviewer_Jxx2"],"forum":"dDvLeDjBOa","number":3,"license":"CC BY 4.0","cdate":1761940540856,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6442/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918834301,"domain":"ICLR.cc/2026/Conference","replyto":"dDvLeDjBOa","id":"rVXhQK4wBM","forumContent":{"TLDR":{"value":"We propose a framework combining physical equation learning and trajectory-guided video models to improve physical realism in video generation by forecasting dynamics with derived equations of motion."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["equation learning","video understanding","video forecasting"]},"supplementary_material":{"value":"/attachment/a4cc7cd8b23a7139de1d8f3f9a998ffb9704690b.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Recent advances in video generation models have achieved remarkable visual realism. However, these models typically lack accurate physical alignment, failing to replicate real-world dynamics in object motion. This limitation arises primarily from their reliance on learned statistical correlations rather than capturing mechanisms adhering to physical laws. To address this issue, we introduce a novel framework that integrates symbolic regression (SR) and trajectory-guided image-to-video (I2V) models for physics-grounded video forecasting. Our approach extracts motion trajectories from input videos, uses a retrieval-based pre-training mechanism to enhance symbolic regression, and discovers equations of motion to forecast physically accurate future trajectories. These trajectories then guide video generation without requiring fine-tuning of existing models. We evaluate our framework on scenarios from classical mechanics, including spring-mass, pendulums, and projectile motions. In these settings, our method successfully recovers ground-truth analytical equations and improves the physical alignment of generated videos compared to baseline methods. This work provides a first step toward integrating equation discovery with video generation."},"_bibtex":{"value":"@misc{\nfeng2025physicsgrounded,\ntitle={Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation},\nauthor={Tao Feng and Xianbing Zhao and Zhenhua Chen and Tien-Tsin Wong and Hamid Rezatofighi and Gholamreza Haffari and Lizhen Qu},\nyear={2025},\nurl={https://openreview.net/forum?id=dDvLeDjBOa}\n}"},"title":{"value":"Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation"},"pdf":{"value":"/pdf/dc64e8ee9c2500724ce0a1baaa49b141cc903847.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"feng|physicsgrounded_motion_forecasting_via_equation_discovery_for_trajectoryguided_imagetovideo_generation"},"authorids":{"value":["~Tao_Feng2","~Xianbing_Zhao1","~Zhenhua_Chen3","~Tien-Tsin_Wong2","~Hamid_Rezatofighi1","~Gholamreza_Haffari2","~Lizhen_Qu2"]},"authors":{"value":["Tao Feng","Xianbing Zhao","Zhenhua Chen","Tien-Tsin Wong","Hamid Rezatofighi","Gholamreza Haffari","Lizhen Qu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces BrowserAgent, an interactive web agent designed to solve complex, real-world problems by browsing the web. Unlike previous agents that rely on external tools to convert web pages into static text, BrowserAgent operates directly on raw web pages. It uses a set of human-inspired actions—such as scrolling, clicking, and typing—to navigate and interact with dynamic web environments via the Playwright framework.\n\nTo train the agent, the authors use a lightweight two-stage training pipeline:\n1. Supervised Fine-Tuning (SFT): This initial stage teaches the base model (Qwen-7B-Instruct) basic reasoning capabilities and the correct format for its thoughts and actions.\n2. Rejection Fine-Tuning (RFT): In this second stage, the SFT model generates multiple potential reasoning paths. The system then selects the best correct answer (specifically, the one with the most reasoning steps) to further fine-tune the model, encouraging deeper and more robust reasoning.\n\nA key feature of BrowserAgent is its explicit memory mechanism. This allows the agent to store key conclusions it finds during its browsing, helping it to maintain context and reason effectively over long-horizon, multi-step tasks."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Above"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"It proposes a framework where the agent interacts directly with raw web pages using fine-grained actions like click, scroll, and type, which is a departure from relying on static text summaries.\n\nThe paper introduces an efficient SFT and RFT training strategy that improves the model's reasoning abilities without complex reinforcement learning techniques.\n\nIt introduces a memory module that stores key findings, significantly enhancing the agent's performance on complex multi-hop reasoning tasks.\n\n BrowserAgent-7B achieves significant performance gains over strong baselines like Search-R1, notably showing around a 20% improvement on multi-hop QA tasks (such as HotpotQA, 2Wiki, and Bamboogle) while using significantly less training data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The agent operates on an \"accessibility tree\" parsed by Playwright, which is \"pure text-based\". This is a known limitation for all agents of this type. It means the agent is blind to visual layout, non-text elements (like complex JavaScript-rendered charts), and website designs that do not have a clean, descriptive accessibility tree.\n\n2. The agent uses a \"minimal yet expressive\" set of predefined actions (e.g., click, type, scroll). This fixed set may not be sufficient for more complex, dynamic web interactions like drag-and-drop, handling complex pop-up modals, or solving CAPTCHAs, which are not mentioned in the action list.\n\n3. Scalability of Training Data: While the authors developed a system to parallelize Playwright instances and improve data collection throughput, the process is still inherently complex. The training was conducted on only 5.3K trajectories, which is noted as being \"significantly less\" than baselines, but also points to the difficulty of creating this type of high-quality, browser-native data at scale."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917363680,"tcdate":1762342515534,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4437/Reviewer_CZep"],"signatures":["ICLR.cc/2026/Conference/Submission4437/Reviewer_CZep"],"forum":"dtExkh4l4z","number":3,"license":"CC BY 4.0","cdate":1762342515534,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4437/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917363680,"domain":"ICLR.cc/2026/Conference","replyto":"dtExkh4l4z","id":"cAZZsCuH7p","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Brosweragent","RFT","LLM"]},"supplementary_material":{"value":"/attachment/eaad41cf4b037266779f1634daba4eabe3a94eae.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Efficiently solving real-world problems with LLMs increasingly hinges on their ability to interact with dynamic web environments and autonomously acquire external information. While recent research like Search-R1 and WebDancer demonstrates strong performance in solving web tasks, they heavily rely on additional tools to convert the interactive web environment into static text content. This is in contrast to human browsing behaviors, which involve diverse interactions with the browser, such as scrolling, clicking, and typing. In this paper, we propose BrowserAgent, a more interactive agent that solves complex tasks through human-inspired browser actions. BrowserAgent operates directly on raw web pages via Playwright through a set of predefined browser actions. We adopt a two-stage training (Supervised Fine-Tuning (SFT) and Rejection Fine-Tuning (RFT)) to improve the model's generalization abilities. Despite using significantly less training data than Search-R1, BrowserAgent achieves more competitive results across different Open-QA tasks.  Additionally, we introduce an explicit memory mechanism to store key conclusions across steps, further enhancing the model's reasoning capabilities for long-horizon tasks. Notably, BrowserAgent-7B can achieve around 20\\% improvement over Search-R1 on multi-hop QA tasks like HotpotQA, 2Wiki, and Bamboogle. These results indicate that BrowserAgent can serve as a more advanced framework for more interactive and scalable web agents."},"_bibtex":{"value":"@misc{\nyu2025browseragent,\ntitle={BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions},\nauthor={Tao Yu and Zhengbo Zhang and Zhiheng Lyu and Junhao Gong and Hongzhu Yi and Xinming Wang and Yuxuan Zhou and Jiabing Yang and Ping Nie and Yan Huang and Wenhu Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=dtExkh4l4z}\n}"},"title":{"value":"BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions"},"pdf":{"value":"/pdf/b9d3673d5deb8b8d203884b5851ffd834873b471.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yu|browseragent_building_web_agents_with_humaninspired_web_browsing_actions"},"authorids":{"value":["~Tao_Yu15","~Zhengbo_Zhang2","~Zhiheng_Lyu2","~Junhao_Gong1","~Hongzhu_Yi3","~Xinming_Wang1","~Yuxuan_Zhou12","~Jiabing_Yang2","~Ping_Nie1","~Yan_Huang2","~Wenhu_Chen3"]},"authors":{"value":["Tao Yu","Zhengbo Zhang","Zhiheng Lyu","Junhao Gong","Hongzhu Yi","Xinming Wang","Yuxuan Zhou","Jiabing Yang","Ping Nie","Yan Huang","Wenhu Chen"]}},"version":2},{"content":{"summary":{"value":"This paper explores the effectiveness of generative data augmentation for image classification, focusing on models trained on internal datasets as well as those trained on large-scale external datasets. The authors conduct experiments on CIFAR-10 and ImageNet and find that while generative data augmentation can enhance classification performance, it generally requires significantly more synthetic data to achieve similar effects as real data augmentation. The study provides practical guidelines for quantifying the ratio between generative and real data needed to achieve comparable improvements. The results suggest that high-quality synthetic data can significantly improve performance when real data is limited, though it remains less efficient compared to real data."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Could the authors provide more theoretical insights or justification for the empirical equivalence between synthetic and real data augmentation? How generalizable is this relationship across different types of datasets and generative models?\n\n2. Have the authors considered combining internal and external generative models to leverage the benefits of both? If so, how would such a hybrid approach affect classification performance, particularly in cases with limited labeled data?\n\n3. Could the authors provide an analysis comparing the computational cost of generating synthetic data to the benefits it provides in terms of performance gains? This would help in assessing the practical feasibility of using generative models for data augmentation in real-world applications.\n\n4. Would the authors consider including other evaluation metrics, such as robustness to adversarial attacks or generalization to out-of-distribution samples, to provide a more holistic assessment of the classifier's performance when trained with synthetic data?\n\n5. Is there an optimal ratio of synthetic to real data that consistently yields the best performance, as observed in Figure 5? If so, how is this ratio affected by the quality of the synthetic data or the complexity of the underlying dataset?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper presents a novel investigation into the role of generative data augmentation for image classification, focusing on both internal (training set only) and external (pre-trained on large datasets) generative models. It uniquely addresses the equivalence between synthetic and real data augmentation, providing empirical quantification to match their effects.\n\n2. The paper is well-constructed, featuring rigorous experimentation on widely recognized datasets like CIFAR-10 and ImageNet. The thorough comparison of real vs. synthetic data, using multiple generative models, ensures that the findings are robust and generalizable.\n\n3. The findings offer practical guidelines for leveraging synthetic data when real data is scarce, providing valuable insights into the trade-offs between real and synthetic data. The results have broad implications for data augmentation in machine learning, especially in contexts where data collection is challenging."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper's empirical analysis primarily focuses on CIFAR-10 and ImageNet, which limits the generalizability of the findings to more complex or diverse datasets. It would be beneficial to include experiments on additional datasets with varying complexity to validate the proposed approach across different domains.\n\n2. The equivalence formulation between real and synthetic data is presented without sufficient theoretical backing. While the empirical results are insightful, the lack of a theoretical framework makes it difficult to generalize the equivalence relationship beyond the specific datasets and models used in the study.\n\n3. The paper does not deeply address the quality of synthetic data and how it impacts classification performance beyond a certain point. For instance, Figures 1(b) and 1(c) highlight the impact of synthetic data on performance, but a more detailed analysis of the characteristics of synthetic data (e.g., diversity, fidelity) is needed to explain the diminishing returns observed in these figures.\n\n4. The experiments on internal and external generative models are treated as largely separate investigations, without adequately exploring potential synergies. Figure 2 shows equivalence curves for internal and external settings separately, but it would be valuable to examine whether combining internal and external generative models could lead to better data augmentation strategies, especially in settings with limited labeled data.\n\n5. The discussion on the trade-offs between real and synthetic data augmentation could be expanded to consider the computational cost of generating synthetic data. Practical considerations, such as the time and resources required to generate large synthetic datasets compared to collecting real data, are not sufficiently addressed. Figure 5, which illustrates the diminishing returns of synthetic data augmentation, could benefit from a discussion on the associated computational costs to assess the feasibility of the proposed approach in real-world scenarios.\n\n6. The study mainly uses accuracy as the evaluation metric, which may not fully capture the nuances of model performance. For example, Figures 4 and 7 show accuracy improvements with different synthetic data ratios, but incorporating additional metrics, such as robustness to adversarial attacks or generalization to out-of-distribution samples, could provide a more comprehensive evaluation of the benefits and limitations of synthetic data augmentation."}},"nonreaders":[],"tmdate":1731475967827,"tcdate":1729024239272,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9246/Reviewer_TBMi"],"signatures":["ICLR.cc/2025/Conference/Submission9246/Reviewer_TBMi"],"forum":"MyAqAYCjP5","number":1,"license":"CC BY 4.0","cdate":1729024239272,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9246/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731475967827,"domain":"ICLR.cc/2025/Conference","replyto":"MyAqAYCjP5","id":"Ks1JKkAnm5","forumContent":{"TLDR":{"value":"This paper investigates the effectiveness of generative data augmentation in image classification, revealing that both internal and external data can significantly improve performance, with empirical guidelines established for using synthetic data."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["analysis by synthesis","image classification","data generation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this paper, we address a key question in machine learning: **How effectively can generative data augmentation enhance image classification?** We begin by examining the differences and similarities between real and synthetic data generated by advanced text-to-image models. Through comprehensive experiments, we provide systematic insights into leveraging synthetic data for improved classification performance. Our findings show that: 1). Generative data augmentation by models trained solely on the internal (available training) set can effectively improve classification performance, validating the long-held hypothesis that synthesis enhances analysis by enriching modeling capability.\n2). For generative data augmentation by models trained on both internal and external data (e.g. large-scale image-text pairs) separately, the size of equivalent synthetic dataset augmentation can be determined empirically. In addition to being aligned with a common intuition that real data augmentation is always preferred, our empirical formulation also provides a guideline for quantitatively estimating how much larger the size of generative dataset augmentation is, over the real data augmentation, to achieve comparable improvements. Our CIFAR-10 and ImageNet results also demonstrate its impact w.r.t. the size of the baseline training set and the quality of generative models."},"_bibtex":{"value":"@misc{\nwang2024mousterian,\ntitle={Mousterian: exploring the equivalence of generative and real data augmentation in classification},\nauthor={Haowen Wang and Guowei Zhang and Xiang Zhang and Zeyuan Chen and Haiyang Xu and Dou Hoon Kwark and Zhuowen Tu},\nyear={2024},\nurl={https://openreview.net/forum?id=MyAqAYCjP5}\n}"},"title":{"value":"Mousterian: exploring the equivalence of generative and real data augmentation in classification"},"pdf":{"value":"/pdf/a4d0aa10edeebfade8c0639ea74caddebddfeea0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|mousterian_exploring_the_equivalence_of_generative_and_real_data_augmentation_in_classification"},"authorids":{"value":["~Haowen_Wang4","~Guowei_Zhang1","~Xiang_Zhang13","~Zeyuan_Chen2","~Haiyang_Xu2","~Dou_Hoon_Kwark1","~Zhuowen_Tu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haowen Wang","Guowei Zhang","Xiang Zhang","Zeyuan Chen","Haiyang Xu","Dou Hoon Kwark","Zhuowen Tu"]}},"version":2},{"content":{"venue":{"value":"AI4Physics"},"pdf":{"value":"/pdf/6af8d51b12e7fc952f12c419d14a615bbae9db75.pdf","readers":["everyone"]},"keywords":{"value":["Neural Network Interpretability","Weight-Space Representations","Hypernetworks","Lattice Field Theory","Joint-Embedding Predictive Architecture","Normalizing Flows","Machine Learning","ICML"]},"venueid":{"value":"ICML.cc/2026/Workshop/AI4Physics"},"paperhash":{"value":"göbel|weightspace_physics_interpretable_hypernetworks_for_lattice_quantum_field_theories"},"authorids":{"value":["~Tobias_Göbel1","~Julian_R._Ebelt1","~Zier_Mensch1","~Mathis_Gerdes1","~Miranda_C._N._Cheng1"]},"abstract":{"value":"Lattice field theory is the workhorse of non-perturbative physics, used to simulate phenomena from the strong nuclear force to critical phenomena in materials. Its Boltzmann distributions are parametrized analytically by coupling constants, but these bare parameters are weak predictors of observables---extracting physics typically requires extensive simulation. While normalizing flows have emerged as effective samplers at fixed couplings, it remains difficult to interpret what these networks have learned. This raises a natural question: can the flow network parameters themselves be generated for new theories, and can their physics be read off directly from the network weights? We propose lattice field theory as a testbed for neural network interpretability: because the target physics is qualitatively well-understood and smoothly varying, it provides ideal synthetic data with known ground truth. To this end we introduce JEPAWG, a Joint-Embedding Predictive Architecture--based Weight Generator that maps couplings directly to flow weights via a learned latent space. On a scalar theory at lattice sizes $6^2$, $8^2$ and $10^2$, the JEPAWG latent space recovers the correct intrinsic dimension of the underlying manifold, locates the phase transition, and encodes a finite-size shift aligned with the 2D Ising exponent $\\nu \\approx 1$, allowing us to uncover physical structure by studying the network weights alone. As a generator, JEPAWG also interpolates and extrapolates to unseen couplings effectively and remains robust to weight-space discontinuities introduced by multi-seed training data, outperforming PCA, AE, and VAE baselines."},"_bibtex":{"value":"@inproceedings{\ngobel2026weightspace,\ntitle={Weight-Space Physics: Interpretable Hypernetworks for Lattice Quantum Field Theories},\nauthor={Tobias G{\\\"o}bel and Julian R. Ebelt and Zier Mensch and Mathis Gerdes and Miranda C. N. Cheng},\nbooktitle={ICML 2026 Workshop on AI for Physics},\nyear={2026},\nurl={https://openreview.net/forum?id=eHSiINhFoo}\n}"},"title":{"value":"Weight-Space Physics: Interpretable Hypernetworks for Lattice Quantum Field Theories"},"authors":{"value":["Tobias Göbel","Julian R. Ebelt","Zier Mensch","Mathis Gerdes","Miranda C. N. Cheng"]}},"tmdate":1784564009829,"pdate":1780110850499,"tcdate":1778239349594,"writers":["ICML.cc/2026/Workshop/AI4Physics","ICML.cc/2026/Workshop/AI4Physics/Submission155/Authors"],"signatures":["ICML.cc/2026/Workshop/AI4Physics/Submission155/Authors"],"forum":"eHSiINhFoo","license":"CC BY 4.0","number":155,"cdate":1778239349594,"readers":["everyone"],"invitations":["ICML.cc/2026/Workshop/AI4Physics/-/Submission","ICML.cc/2026/Workshop/AI4Physics/-/Post_Submission","ICML.cc/2026/Workshop/AI4Physics/-/Edit","ICML.cc/2026/Workshop/AI4Physics/Submission155/-/Camera_Ready_Revision"],"mdate":1784564009829,"odate":1782846810361,"domain":"ICML.cc/2026/Workshop/AI4Physics","id":"eHSiINhFoo","version":2},{"content":{"summary":{"value":"The paper introduces RESuM (Rare Event Surrogate Model), designed to optimize physics detector design, specifically for reducing background events in neutrinoless double-beta decay (NLDBD) detection. The authors frame the challenge as a Rare Event Design (RED) problem, where background events are so rare that traditional simulation methods become computationally prohibitive. RESuM employs a multi-fidelity approach, using a Conditional Neural Process (CNP) model to learn from limited data and a Gaussian Process (GP) model to efficiently explore the design space. The authors apply RESuM to the LEGEND experiment’s neutron moderator design and achieve an  efficient performance."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"See Weaknesses for questions."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The model targets a critical challenge in physic.\n2. The integration of CNP and MFGP is a novel approach for handling rare event problems.\n3. The framework’s formulation makes it generalizable to other simulation-heavy domains."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I do not find major concerns.\nMinor issues:\n1. The success of RESuM’s multi-fidelity modeling approach relies on access to both high-fidelity and low-fidelity simulation data. In situations where these data sources are unavailable or highly dissimilar, would the model's effectiveness be reduced?\n2. It seems to me that the RESuM does not embed physics constraints or validation check to the model, thus it might be possible that RESuM would suggest some invalid designs. In practice, do we need to first define the valid range of parameters for RESuM? And thus for different application contexts, we need to modify the model accordingly?\n3. Section 5 uses subsubsections without subsections.\n4. Can RESuM scale to applications with larger or more complex detector designs? How to estimate the relationship between the dimension of design space and model performance? Also, can RESuM incorporate expert knowledge on the design space to make it more efficient?"}},"nonreaders":[],"tmdate":1731428173322,"tcdate":1730485646562,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4920/Reviewer_TJgd"],"signatures":["ICLR.cc/2025/Conference/Submission4920/Reviewer_TJgd"],"forum":"lqTILjL6lP","number":1,"license":"CC BY 4.0","cdate":1730485646562,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4920/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428173322,"domain":"ICLR.cc/2025/Conference","replyto":"lqTILjL6lP","id":"5WQJqGBRJS","forumContent":{"TLDR":{"value":"We developed a Rare Event Surrogate Model (RESuM) powered by Conditional Neural Process to optimize particle physics detector design under high-variance design metrics."},"venue":{"value":"ICLR 2025 Spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["surrogate model","simulation","rare event search","AI4Sci","AI for physics","conditional neural process","Bayesian methods","emulator","Multi-Fidelity Gaussian Process"]},"supplementary_material":{"value":"/attachment/61607f968727ee55b905590bd08e4b8019ad0fb3.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The experimental discovery of neutrinoless double-beta decay (NLDBD) would answer one of the most important questions in physics: Why is there more matter than antimatter in our universe? To maximize the chances of discovery, NLDBD experiments must optimize their detector designs to minimize the probability of background events contaminating the detector. Given that this probability is inherently low, design optimization either requires extremely costly simulations to generate sufficient background counts or contending with significant variance. In this work, we formalize this dilemma as a Rare Event Design (RED) problem: identifying optimal design parameters when the design metric to be minimized is inherently small. We then designed the Rare Event Surrogate Model (RESuM) for physics detector design optimization under RED conditions. RESuM uses a pre-trained Conditional Neural Process (CNP) model to incorporate additional prior knowledge into a Multi-Fidelity Gaussian Process model. We applied RESuM to optimize neutron shielding designs for the LEGEND NLDBD experiment, identifying an optimal design that reduces the neutron background by $(66.5 \\pm 3.5)$% while using only 3.3% of the computational resources compared to traditional methods. Given the prevalence of RED problems in other fields of physical sciences, especially in rare-event searches, the RESuM algorithm has broad potential for accelerating simulation-intensive applications."},"_bibtex":{"value":"@inproceedings{\nschuetz2025resum,\ntitle={{RES}uM: A Rare Event Surrogate Model for  Physics Detector Design},\nauthor={Ann-Kathrin Schuetz and A.W.P. Poon and Aobo Li},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=lqTILjL6lP}\n}"},"title":{"value":"RESuM: A Rare Event Surrogate Model for  Physics Detector Design"},"pdf":{"value":"/pdf/4c5722e366a600768c810a33c24a205a720be28c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"schuetz|resum_a_rare_event_surrogate_model_for_physics_detector_design"},"authorids":{"value":["~Ann-Kathrin_Schuetz1","~A.W.P._Poon1","~Aobo_Li2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ann-Kathrin Schuetz","A.W.P. Poon","Aobo Li"]}},"version":2},{"content":{"summary":{"value":"This paper introduces $EFO_{k}-CQA$, a framework for generating datasets and evaluating models on complex query answering (CQA) tasks over knowledge graphs (KGs). Unlike previous datasets limited to set operations and tree-based structures, EFOk-CQA supports the broader family of Existential First-Order (EFO) queries, extending to multiple variables and complex graph structures. The authors propose an end-to-end pipeline to generate EFO queries, sample answers, train models, and evaluate performance. The dataset is shown to cover 741 query types, aiming to capture a wide combinatorial space of queries. Through extensive experiments, the paper highlights structural biases in existing datasets and evaluates six prominent CQA models, revealing the influence of query hardness and graph topology on model performance."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Given the vast combinatorial space, how scalable is the framework for generating queries with more complex conditions or additional free variables? Are there specific trade-offs in terms of time or computational resources that become prohibitive?\n2.  How often do the assumptions used to generate non-trivial query graphs (e.g., no decomposition) align with real-world query requirements? Would relaxing these assumptions reveal additional complexities or insights about query-answering models?\n3.  Can the authors provide an error analysis highlighting specific failure modes of the evaluated models? For example, are there certain query types (e.g., cyclic vs. acyclic) where models consistently underperform? This information could clarify which specific capabilities of current SOTA works require improvement."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The EFOk-CQA framework is a meaningful extension to existing CQA benchmarks, expanding beyond single-variable, set operation-based queries to include multiple-variable and complex query graphs. It also addresses a notable gap in CQA evaluation, offering a comprehensive benchmark that reflects real-world query complexity.\n2.  The paper demonstrates a high degree of rigor in constructing the EFOk-CQA dataset. The authors establish well-defined assumptions to exclude trivial cases, implement a pipeline for end-to-end evaluation, and provide empirical results with various CQA models across a comprehensive range of query structures and complexities.\n3. The paper is well-organized, with clear definitions and illustrations that help in understanding complex query structures. Concepts such as query graphs, grounding processes, and evaluation metrics are presented with detailed explanations and visual aids, facilitating comprehension.\n4. The EFOk-CQA dataset and framework have implications for advancing CQA research. By covering a broader range of query types, the benchmark could lead to the development of more generalized and capable CQA models. The analysis of model performance based on query topology and difficulty also provides useful insights for future CQA model improvements."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In section 3.1, while the framework supports complex queries, the combinatorial space of queries could grow exponentially as query parameters increase. This might make it challenging to scale up the dataset generation process efficiently or to apply the framework to larger, more complex knowledge graphs.\n2.  The use of joint metrics to evaluate models with multiple free variables is valuable, but it is noted in the paper that joint rankings could lack reliability due to inherent difficulty. This limitation could affect the robustness of evaluation for complex queries, particularly when interpreting the rankings for CQA tasks."}},"nonreaders":[],"tmdate":1731428721734,"tcdate":1730479326652,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6686/Reviewer_F5D3"],"signatures":["ICLR.cc/2025/Conference/Submission6686/Reviewer_F5D3"],"forum":"K7XiXLfFSP","number":2,"license":"CC BY 4.0","cdate":1730479326652,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6686/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428721734,"domain":"ICLR.cc/2025/Conference","replyto":"K7XiXLfFSP","id":"EsuV9GJpHm","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"This paper extends the task of complex query answering on knowledge graph to a brand new frontier."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["complex query answering","knowledge graph"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"To answer complex queries on knowledge graphs, logical reasoning over incomplete knowledge needs learning-based methods because they are capable of generalizing over unobserved knowledge. Therefore, an appropriate dataset is fundamental to both obtaining and evaluating such methods under this paradigm. In this paper, we propose a comprehensive framework for data generation, model training, and method evaluation that covers the combinatorial space of Existential First-order Queries with multiple variables ($EFO_{k}$). The combinatorial query space in our framework significantly extends those defined by set operations in the existing literature. Additionally, we construct a dataset, $EFO_{k}$-CQA, with 741 query types for empirical evaluation, and our benchmark results provide new insights into how query hardness affects the results. Furthermore, we demonstrate that the existing dataset construction process is systematically biased and hinders the appropriate development of query-answering methods, highlighting the importance of our work. Our code and data are provided in~\\url{https://anonymous.4open.science/r/EFOK-CQA/README.md}."},"_bibtex":{"value":"@misc{\nyin2025efokcqa,\ntitle={\\${EFO}\\_\\{k\\}\\$-{CQA}: Towards Knowledge Graph Complex Query Answering beyond Set Operation},\nauthor={Hang Yin and Zihao Wang and Weizhi Fei and Yangqiu Song},\nyear={2025},\nurl={https://openreview.net/forum?id=K7XiXLfFSP}\n}"},"title":{"value":"$EFO_{k}$-CQA: Towards Knowledge Graph Complex Query Answering beyond Set Operation"},"pdf":{"value":"/pdf/f130423487689cd42f045f9a55e985fbb29e1af2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yin|efo_kcqa_towards_knowledge_graph_complex_query_answering_beyond_set_operation"},"authorids":{"value":["~Hang_Yin3","~Zihao_Wang11","~Weizhi_Fei1","~Yangqiu_Song1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hang Yin","Zihao Wang","Weizhi Fei","Yangqiu Song"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a framework for designing sequential indirect experiments for estimating targeted scientific queries in complex, nonlinear environments with potential unobserved confounding when direct intervention is impractical or impossible. The authors formulate the problem as the sequential instrument design, using instrumental variables and minimax optimization to estimate the upper and lower bounds of targeted causal effects. The proposed method then tightens these bounds iteratively through adaptive strategies. Experiments on simulated data demonstrate the proposed method's efficacy compared to non-adaptive experimental design baselines."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Please see the questions in the Weaknesses part.\n\n2. It seems that the proposed method is closely related to causal Bayesian Optimization [1]. I wonder if the authors could discuss the connection in detail. \n\n3. Experiments demonstrate that the proposed method performs well for local causal effect queries. However, it is usually unknown whether the target causal effect is local or more global (i.e., long-range) in real-world applications. I wonder if the proposed method could estimate long-range causal effects accurately.\n\n4. What is the difference/connection between the proposed method and targeted indirect experiment design in an active learning setting?\n\n[1] Aglietti, V., Lu, X., Paleyes, A., & González, J. (2020, June). Causal bayesian optimization. In International Conference on Artificial Intelligence and Statistics (pp. 3155-3164). PMLR."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper focuses on targeted sequential indirect experimental design, which is an important problem in scientific discovery where direct intervention is often impractical or impossible.\n\n2. The proposed method is more flexible as it considers nonlinear, multi-variate, and confounded settings, and formulating the problem as a sequential underspecified instrumental variable estimation is intuitive and sound.\n\n3. The authors develop closed-form estimators for the bounds given targeted queries when the mechanism is in an RKHS.\n\n4. The proposed method outperforms non-adaptive baselines in synthetic experiments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the proposed method focuses on indirect experimental design, it still requires the causal structure between variables known, which is also one of the main challenges in scientific discovery. I am curious about the method's sensitivity to imperfections in the causal structure, assumptions regarding instrumental variables, and the presence of confounders. It would be great if the authors could discuss further about them.\n\n2. The authors only conduct synthetic experiments in a simple setting when the number of variables is small. I wonder if the authors could conduct more experiments in a more complex setting with a larger number of nodes, and also discuss the scalability of the proposed method (since kernel-based techniques are used for optimization as mentioned in the paper)."},"limitations":{"value":"The authors adequately addressed the limitations of their work."}},"nonreaders":[],"tmdate":1730879407146,"tcdate":1720775001539,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission10694/Reviewer_aNmJ"],"signatures":["NeurIPS.cc/2024/Conference/Submission10694/Reviewer_aNmJ"],"forum":"U3Rgdb4li9","number":2,"license":"CC BY 4.0","cdate":1720775001539,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission10694/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879407146,"domain":"NeurIPS.cc/2024/Conference","replyto":"U3Rgdb4li9","id":"zgcM4a3mVB","forumContent":{"TLDR":{"value":"We adaptively learn optimal indirect experiments to narrow the bounds on a functional of f in high-dimensional, non-linear settings with unobserved confounding."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["causality","experiment design","instrumental variables","indirect experiments"]},"supplementary_material":{"value":"/attachment/7800e7cc4583297806bf169b20565f622396a2ba.zip"},"primary_area":{"value":"causal_inference"},"abstract":{"value":"Scientific hypotheses typically concern specific aspects of complex, imperfectly understood or entirely unknown mechanisms, such as the effect of gene expression levels on phenotypes or how microbial communities influence environmental health. Such queries are inherently causal (rather than purely associational), but in many settings, experiments can not be conducted directly on the target variables of interest, but are indirect. Therefore, they perturb the target variable, but do not remove potential confounding factors. If, additionally, the resulting experimental measurements are high-dimensional and the studied mechanisms nonlinear, the query of interest is generally not identified. We develop an adaptive strategy to design indirect experiments that optimally inform a targeted query about the ground truth mechanism in terms of sequentially narrowing the gap between an upper and lower bound on the query. While the general formulation consists of a bi-level optimization procedure, we derive an efficiently estimable analytical kernel-based estimator of the bounds for the causal effect, a query of key interest, and demonstrate the efficacy of our approach in confounded, multivariate, nonlinear synthetic settings."},"_bibtex":{"value":"@inproceedings{\nailer2024targeted,\ntitle={Targeted Sequential Indirect Experiment Design},\nauthor={Elisabeth Ailer and Niclas Dern and Jason Hartford and Niki Kilbertus},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=U3Rgdb4li9}\n}"},"title":{"value":"Targeted Sequential Indirect Experiment Design"},"pdf":{"value":"/pdf/d6dbe3094fc9ef1881705acaf521213cc8dcb314.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"ailer|targeted_sequential_indirect_experiment_design"},"authorids":{"value":["~Elisabeth_Ailer1","~Niclas_Dern1","~Jason_Hartford1","~Niki_Kilbertus1"]},"authors":{"value":["Elisabeth Ailer","Niclas Dern","Jason Hartford","Niki Kilbertus"]}},"version":2},{"content":{"TLDR":{"value":"We define and analyze the concept of verification ceiling in synthetic code data generation pipelines and present major findings to construct better ways to verify synthetic code that are provably helpful in building stronger codegen models."},"venue":{"value":"ICLR 2026 Workshop VerifAI-2"},"keywords":{"value":["Synthetic code data","Verification via synthetic unit tests","Code Generation Models","LLM Code Evaluation"]},"abstract":{"value":"Large language models for code generation increasingly rely on synthetic data, where both problem solutions and verification tests are generated by models. While this enables scalable data creation, it introduces a previously unexplored bottleneck: the verification ceiling, in which the quality and diversity of training data are fundamentally constrained by the capabilities of synthetic verifiers. In this work, we systematically study how verification design and strategies influence model performance. We investigate (i) what we verify by analyzing the impact of test complexity and quantity: richer test suites improve code generation capabilities (on average +3 pass@1), while quantity alone yields diminishing returns, (ii) how we verify by exploring relaxed pass thresholds: rigid 100% pass criteria can be overly restrictive. By allowing for relaxed thresholds or incorporating LLM-based soft verification, we can recover valuable training data, leading to a 2-4 point improvement in pass@1 performance. However, this benefit is contingent upon the strength and diversity of the test cases used, and (iii) why verification remains necessary through controlled comparisons of formally correct versus incorrect solutions and human evaluation: retaining diverse correct solutions per problem yields consistent generalization gains. Our results show that Verification as currently practiced is too rigid, filtering out valuable diversity. But it cannot be discarded, only recalibrated. By combining calibrated verification with diverse, challenging problem-solution pairs, we outline a path to break the verification ceiling and unlock stronger code generation models."},"_bibtex":{"value":"@inproceedings{\ngureja2026verification,\ntitle={Verification Limits Code {LLM} Training},\nauthor={Srishti Gureja and Marzieh Fadaee and Sara Hooker and Matthias Gall{\\'e} and Jingyi He and Elena Tommasone},\nbooktitle={ICLR 2026 Workshop: VerifAI-2: The Second Workshop on AI Verification in the Wild},\nyear={2026},\nurl={https://openreview.net/forum?id=KP2nLjrkhD}\n}"},"title":{"value":"Verification Limits Code LLM Training"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"venueid":{"value":"ICLR.cc/2026/Workshop/VerifAI-2"},"paperhash":{"value":"gureja|verification_limits_code_llm_training"},"authorids":{"value":["~Srishti_Gureja1","~Marzieh_Fadaee2","~Sara_Hooker2","~Matthias_Gallé1","~Jingyi_He1","~Elena_Tommasone1"]},"Track":{"value":"long paper (up to 8 pages)"},"authors":{"value":["Srishti Gureja","Marzieh Fadaee","Sara Hooker","Matthias Gallé","Jingyi He","Elena Tommasone"]}},"tmdate":1773260208136,"pdate":1772424058019,"tcdate":1770220535309,"writers":["ICLR.cc/2026/Workshop/VerifAI-2","ICLR.cc/2026/Workshop/VerifAI-2/Submission22/Authors"],"signatures":["ICLR.cc/2026/Workshop/VerifAI-2/Submission22/Authors"],"forum":"KP2nLjrkhD","license":"CC BY 4.0","number":22,"cdate":1770220535309,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/VerifAI-2/-/Submission","ICLR.cc/2026/Workshop/VerifAI-2/-/Post_Submission","ICLR.cc/2026/Workshop/VerifAI-2/-/Edit","ICLR.cc/2026/Workshop/VerifAI-2/Submission22/-/Camera_Ready"],"mdate":1773260208136,"odate":1772718030290,"domain":"ICLR.cc/2026/Workshop/VerifAI-2","id":"KP2nLjrkhD","version":2},{"content":{"venue":{"value":"Journal of High Energy Physics"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"cerro|spectral_clustering_for_jet_physics"},"abstract":{"value":"We present a new approach to jet definition alternative to clustering methods, such as the anti-kT scheme, that exploit kinematic data directly. Instead the newmethod uses kinematic information to represent the particles in a multidimensional space, as in spectral clustering. After confirming its Infra-Red (IR) safety, we compare its performance in analysing gg → H125 GeV → H40 GeVH40 GeV → b¯bb¯b, gg → H500 GeV → H125 GeVH125 GeV → b¯bb¯b and gg, qq¯ → tt¯ → b¯bW+W− → b¯bjjlνl events from Monte Carlo (MC) samples, specifically, in reconstructing the relevant final states, to that of the anti-kT algorithm. Finally, we show that the results for spectral clustering are obtained without any change in the parameter settings of the algorithm, unlike the anti-kT case, which requires the cone size to be adjusted to the physics process under study."},"_bibtex":{"value":"@article{244bd96830044b60acdee9018d306a79,\n  title     = \"Spectral clustering for jet physics\",\n  abstract  = \"We present a new approach to jet definition alternative to clustering methods, such as the anti-kT scheme, that exploit kinematic data directly. Instead the newmethod uses kinematic information to represent the particles in a multidimensional space, as in spectral clustering. After confirming its Infra-Red (IR) safety, we compare its performance in analysing gg → H125 GeV → H40 GeVH40 GeV → b¯bb¯b, gg → H500 GeV → H125 GeVH125 GeV → b¯bb¯b and gg, qq¯ → tt¯ → b¯bW+W− → b¯bjjlνl events from Monte Carlo (MC) samples, specifically, in reconstructing the relevant final states, to that of the anti-kT algorithm. Finally, we show that the results for spectral clustering are obtained without any change in the parameter settings of the algorithm, unlike the anti-kT case, which requires the cone size to be adjusted to the physics process under study.\",\n  author    = \"Giorgio Cerro and Srinandan Dasmahapatra and Day-Hall, {Henry A.} and Billy Ford and Stefano Moretti and Shepherd-Themistocleous, {Claire H.}\",\n  year      = \"2022\",\n  month     = feb,\n  day       = \"21\",\n  doi       = \"10.1007/JHEP02(2022)165\",\n  language  = \"English\",\n  volume    = \"2022\",\n  journal   = \"Journal of High Energy Physics\",\n  issn      = \"1126-6708\",\n  publisher = \"Springer\",\n}"},"title":{"value":"Spectral clustering for jet physics"},"authors":{"value":[{"fullname":"Giorgio Cerro"},{"fullname":"Srinandan Dasmahapatra","username":"~Srinandan_Dasmahapatra1"},{"fullname":"Henry A. Day-Hall"},{"fullname":"Billy  Ford"},{"fullname":"Stefano Moretti"},{"fullname":"Claire H. Shepherd-Themistocleous"}]}},"tmdate":1789089738504,"pdate":1645401600000,"externalIds":["doi:10.1007/jhep02(2022)165"],"tcdate":1762330561811,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Srinandan_Dasmahapatra1"],"forum":"vlzhDSmupc","license":"CC BY-SA 4.0","number":6038,"cdate":1749667099686,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789089738504,"domain":"OpenReview.net/Public_Article","id":"vlzhDSmupc","version":2},{"content":{"summary":{"value":"The authors build 3D digital phantoms of the male pelvis by combining four MRI/CT patient anatomies with 62 ex-vivo prostate SOS maps, resulting in 248 model variants. They then simulate two-dimensional acoustic wave propagation using a finite-difference time-domain solver, producing paired input–output samples: ultrasound waveforms and their corresponding ground-truth SOS maps. By slicing and rotating these models, they generate about 280K 2D examples meant to reflect the geometric and acoustic variability encountered in trans-rectal and trans-abdominal USCT setups.\n\nThey also use the dataset to compare some baselines:\n1) classical beamforming,\n2) iterative physics-based USCT inversion,\nand 3,4) two learned models: InversionNet (cnn) and a ViTvariant.\n\nThe learned models achieve lower pixelwise errors and run orders of magnitude faster than physics-based inversion, though their outputs remain too smooth to resolve fine prostate boundaries or lesions. Unfortunately out-of-distribution tests, where one patient anatomy is held out, reveal major performance degradation. This i s evidence of poor generalization beyond the limited training anatomies."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1) The dataset is generated from four patient anatomies and 62 ex-vivo prostate samples. This needs to become more transparent in the paper -- from the very start.\n2) How were the four patient anatomies selected? Are they representative? Please provide some info about where they come from: healthy, diseased, varied in size, age, pathology?\n3) The paper slices 3D anatomies into many 2D samples. how do you ensure that (almost) identical slices or planes are not split across train/validation/test sets?\n4) The large difference between ID and OOD results suggests overfitting to specific anatomies. This worries me a lot. Maybe the authors can also report confidence intervals across anatomies and make the OOD evaluation a central table in the main paper?\n5) How is the performance affected by sensor or probe minor placement errors or SOS background shifts? A robustness study would help. \n6) It is not clear to me what will be finally released by the authors. waveforms, SOS maps, code, and 3D phantoms? What about the original MRI/CT data?\n7) in terms of related work, a more SOTA approach would be to use neural operator learning -- something like this https://arxiv.org/abs/2304.03297 . I suggest you discuss this in the conclusions?\n8-9) A couple of more technical questions: \nthe forward model only focuses on SOS. Out-of-plane scattering, attenuation, and density variations are ignored. Have you tried to validate the simulator against 3D k-Wave simulations? How big are the errors when ignoring such things? \nAnd how does ex-vivo SOS differ from in-vivo values at body temperature?"},"rating":{"value":8},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Addresses a clinically important and computationally demanding imaging problem\n2. Provides an open and standardized resource that can make it easier for other people to design and evaluate their own models\n3. The selected benchmarks are reasonable and cover both physics-based and learned approaches\n4. The paper is technically good and the simulation details are sufficient, I think, if someone wants to reproduce the results"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My main concern: the dataset is based on only four patient anatomies and this is mentioned only deep in the methods and appendix (I think it should be mentioned in the abstract or, worst-case, introduction). The anatomical diversity is therefore extremely limited!\n\n- All simulations are 2D and reconstruct only SOS, ignoring attenuation and density. \n- The machine-learning component is incremental (standard CNN and ViT baselines) without any more recent architectures or hybrid physics-learning integration. But this is not a major issue because I recognize that the main contribution of the paper is teh dataset.\n- Evaluation focuses on image-level metrics rather than clinically more meaningful goals such as gland boundary accuracy or lesion detection\n- The ex-vivo data used for SOS ground truth may differ from in-vivo conditions. This is not something the paper discusses or tries to quantify."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916249363,"tcdate":1762002142314,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2477/Reviewer_kNhg"],"signatures":["ICLR.cc/2026/Conference/Submission2477/Reviewer_kNhg"],"forum":"kcFEpBagea","number":4,"license":"CC BY 4.0","cdate":1762002142314,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2477/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916249363,"domain":"ICLR.cc/2026/Conference","replyto":"kcFEpBagea","id":"zRHZcTwzJZ","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Ultrasound Computed Tomography","Prostate Imaging","Benchmark Dataset","Medical Imaging"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Prostate cancer is one of the most common and lethal cancers among men, making its early detection critically important. Ultrasound computed tomography (USCT) has emerged as an accessible and cost-effective method that reconstructs quantitative tissue parameters, which can serve as potential biomarkers for malignancy. However, current prostate USCT faces considerable barriers: limited-angle acquisitions due to anatomical constraints, tissue heterogeneity, proximity to organs and bony pelvic structures, and lengthy processing times. The lack of large-scale, anatomically precise datasets significantly hampers the development of high-quality, efficient, and generalizable methods. To address this gap, we introduce OpenPros, the first large-scale benchmark dataset for limited-angle prostate USCT, designed to evaluate machine learning algorithms for inverse problems systematically. Our dataset includes over 280,000 paired samples of realistic 2D speed-of-sound (SOS) phantoms and corresponding ultrasound full-waveform data, generated from anatomically accurate 3D digital prostate models derived from 4 real clinical MRI/CT scans and 62 ex vivo prostate specimens with experimental ultrasound measurements, annotated by medical experts. Simulations are conducted under clinically realistic configurations using advanced finite-difference time-domain (FDTD) and Runge-Kutta acoustic wave solvers, both provided as open-source components. Through comprehensive benchmarking, we find that deep learning methods significantly outperform traditional physics-based algorithms in inference efficiency and reconstruction accuracy. However, our results also reveal that current machine learning methods fail to deliver clinically acceptable, high-resolution reconstructions, underscoring critical gaps in generalization, robustness, and uncertainty quantification. By publicly releasing OpenPros, we provide the community with a rigorous benchmark that not only enables fair method comparison but also motivates new advances in physics-informed learning, foundation models for scientific imaging, and uncertainty-aware reconstruction—bridging the gap between academic ML research and real-world clinical deployment. The dataset is publicly accessible at https://open-pros.github.io/."},"_bibtex":{"value":"@inproceedings{\nwang2026openpros,\ntitle={OpenPros: A Large-Scale Dataset for Limited View Prostate Ultrasound Computed Tomography},\nauthor={Hanchen Wang and Yixuan Wu and Yinan Feng and Peng Jin and Luoyuan Zhang and Shihang Feng and James Wiskin and Baris Turkbey and Peter Pinto and Bradford J Wood and Songting Luo and Yinpeng Chen and Emad Boctor and Youzuo Lin},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=kcFEpBagea}\n}"},"title":{"value":"OpenPros: A Large-Scale Dataset for Limited View Prostate Ultrasound Computed Tomography"},"pdf":{"value":"/pdf/dde14fc3358d4c372d349210f070c720f91ab713.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|openpros_a_largescale_dataset_for_limited_view_prostate_ultrasound_computed_tomography"},"authorids":{"value":["~Hanchen_Wang3","~Yixuan_Wu4","~Yinan_Feng1","~Peng_Jin6","~Luoyuan_Zhang2","~Shihang_Feng1","~James_Wiskin1","~Baris_Turkbey1","~Peter_Pinto1","~Bradford_J_Wood1","~Songting_Luo1","~Yinpeng_Chen1","~Emad_Boctor1","~Youzuo_Lin1"]},"authors":{"value":["Hanchen Wang","Yixuan Wu","Yinan Feng","Peng Jin","Luoyuan Zhang","Shihang Feng","James Wiskin","Baris Turkbey","Peter Pinto","Bradford J Wood","Songting Luo","Yinpeng Chen","Emad Boctor","Youzuo Lin"]}},"version":2},{"content":{"summary":{"value":"This paper aims to investigate why transformers possess the capability of in-context learning from a theoretical perspective. Previous works have shown simplified self-attention layer’s capability of learning linear functions in context. This work conducts further research based on softmax regression, as softmax-based algorithms are more complex and closer to algorithms used in actual LLMs. Through mathematical analysis and experimental validation, the authors conclude that the update acquired through gradient descent and in-context learning are similar when training simplified transformers for softmax regression tasks.\n\nI am not fully follow this paper and I wish to learn more intuitive understanding of this paper during rebuttal (from the authors and other reviewers). I may adjust my rating score after rebuttal period."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"None"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.Explaining the reasons why LLMs can learn from context at a theoretical level is of significant importance, and this paper has made a valuable exploration into this issue.\n\n2.There is a thorough and comprehensive mathematical analysis to demonstrate the similarity between the models learned by transformers and gradient-descent on softmax regression.\n\n3.Theoretical results and experimental results corroborate each other in this paper."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.This paper appears to build upon previous research that explored in-context learning capabilities of transformer using linear regression, but instead opts for a softmax regression approach, lacking significant innovation and contribution.\n\n2.Although this paper represents a further advancement over previous studies based on linear regression, it still employs a highly simplified Transformer model, which is insufficient to fully explain the principles of LLMs’ in-context learning ability.\n\n3.I suggest that the structure of the paper could be organized more clearly, allowing readers to grasp the overall framework first. The details of the two models compared in the experiments should be explained more thoroughly. (1) The section 2 and 3 are not so well-organized, where some formula lacks necessary comments and explanations. (2) The introduction section is not so intuitive to follow.  (3) There is almost no textual descriptions about the proposed method section (section 3). I believe most NeurIPS reviewers are hard to follow the section 3 without enough descriptions."},"limitations":{"value":"The limitations are fine with me."}},"nonreaders":[],"tmdate":1730879197838,"tcdate":1720845980896,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission7836/Reviewer_K6fK"],"signatures":["NeurIPS.cc/2024/Conference/Submission7836/Reviewer_K6fK"],"forum":"SFaEENfEyw","number":1,"license":"CC BY 4.0","cdate":1720845980896,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission7836/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879197838,"domain":"NeurIPS.cc/2024/Conference","replyto":"SFaEENfEyw","id":"17injgF0nR","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["in-context learning"]},"primary_area":{"value":"optimization_for_deep_networks"},"abstract":{"value":"Large language models (LLMs) are known for their exceptional performance in natural language processing, making them highly effective in many human life-related tasks. The attention mechanism in the Transformer architecture is a critical component of LLMs, as it allows the model to selectively focus on specific input parts. The softmax unit, which is a key part of the attention mechanism, normalizes the attention scores. Hence, the performance of LLMs in various NLP tasks depends significantly on the crucial role played by the attention mechanism with the softmax unit.\n\nIn-context learning is one of the celebrated abilities of recent LLMs. \nWithout further parameter updates, Transformers can learn to predict based on few in-context examples. \nHowever, the reason why Transformers becomes in-context learners is not well understood.\nRecently, in-context learning has been studied from a mathematical perspective with simplified linear self-attention without softmax unit. \nBased on a linear regression formulation $\\min_x\\| Ax  - b \\|_2$, existing works show linear Transformers' capability of learning linear functions in context. The capability of Transformers with softmax unit approaching full Transformers, however, remains unexplored.\n\nIn this work, we study the in-context learning based on a softmax regression formulation $\\min_{x} \\| \\langle \\exp(Ax), {\\bf 1}_n \\rangle^{-1} \\exp(Ax) - b \\|_2$. We show the upper bounds of the data transformations induced by a single self-attention layer with softmax unit and by gradient-descent on a $\\ell_2$ regression loss for softmax prediction function.\nOur theoretical results imply that when training self-attention-only Transformers for fundamental regression tasks, the models learned by gradient-descent and Transformers show great similarity."},"_bibtex":{"value":"@inproceedings{\nli2024the,\ntitle={The Closeness of In-Context Learning and Weight Shifting for Softmax Regression},\nauthor={Shuai Li and Zhao Song and Yu Xia and Tong Yu and Tianyi Zhou},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=SFaEENfEyw}\n}"},"title":{"value":"The Closeness of In-Context Learning and Weight Shifting for Softmax Regression"},"pdf":{"value":"/pdf/9b0fac1a2e82221eaedb68182d5435c9ddb47ba6.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"li|the_closeness_of_incontext_learning_and_weight_shifting_for_softmax_regression"},"authorids":{"value":["~Shuai_Li3","~Zhao_Song3","~Yu_Xia9","~Tong_Yu3","~Tianyi_Zhou4"]},"authors":{"value":["Shuai Li","Zhao Song","Yu Xia","Tong Yu","Tianyi Zhou"]}},"version":2},{"content":{"summary":{"value":"The paper proposes an approach for fine-tuning diffusion models for the design of antibodies. The core diffusion-based generative model comes from Luo et al. [36] and to my understanding there are no technical changes to it. The second component is direct preference based optimization, inspired by fine-tuning of large language models. This is the main technical contribution and builds entirely on the works by Rafailov et al [41] and Wallace et al [46]. The reward signal comes from binding free energy that is decomposed at the residue level. Different components include attraction and repulsion forces that can be linked to antibody function. To overcome possibly diverging gradients associated with different energy components, the authors propose to leverage gradient surgery from Yu et al [51] that essentially increases cosine similarity between gradients for different tasks/energy components.\n\nEmpirical evaluation is focused on qualitative aspects and whether the approach is able to discover better binders than initially given complexes, measured using binding free energy."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Could you also list how many mutations from the original CDR sequence are the improved binders (Table 1)?  \n\nCan you share the structural similarity scores between epitopes in the train and test folds?\n\nIf you limit the fraction of hydrophobic residues (PHR) to the value for the original complex and accept only designs with lower score, how does Table 1 look? Can you list relative to how many test fold complexes each method achieves better binding free energy (e.g., MEAN wins on N, AbDPO wins on 20, etc.)?\n\nHow does Table 1 look like if you apply ITA to the vanilla baselines?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"I think it is an interesting approach to merry molecular simulations and physics based energy calculations with diffusion and generative models. Especially given the small number of available crystal structures in the SAbDAB database.\n\nAdding energy-based signal via direct preference optimization is an interesting re-purposing of that method.\n\nEmpirical evaluation goes beyond amino-acid recovery rate and RMSD metrics. The “success” at generating better binders quantified via improvement in binding free energy relative to the initial complex is an interesting metric."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Table 1 indicates that the approach is able to design better binders, measured via binding free energy relative to the initial complex. However, this comes at the expense of increasing the number of hydrophobic residues which is typically associated with non-specific binding. This would in all likelihood be useless binders and from the perspective of function no better than baselines.\n\nIt would be interesting to see the results of a baseline that “fine-tunes” relative to ddG score directly. In my understanding, the results in Table 1 are vanilla baselines. None of them (e.g., MEAN, dyMEAN, HERN) has for instance been used in combination with iterative improvement algorithm and some physics based simulator. How would this compare to fine-tuning relative to the directed preference optimization?\n\nIt would be interesting to see the results relative to different physics-based simulators. I’m not sure how many different simulators were used to generate “fine-tuning” signal?\n\nWould it be possible to include additional metrics such as lDDT, TM, and some metrics characterizing the fit on angles?"},"limitations":{"value":"NA"}},"nonreaders":[],"tmdate":1730879678673,"tcdate":1720713398061,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission14012/Reviewer_X6SE"],"signatures":["NeurIPS.cc/2024/Conference/Submission14012/Reviewer_X6SE"],"forum":"GN2GXjPyN8","number":3,"license":"CC BY 4.0","cdate":1720713398061,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission14012/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879678673,"domain":"NeurIPS.cc/2024/Conference","replyto":"GN2GXjPyN8","id":"cvBb9vzhu2","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["direct preference optimization","diffusion model","antibody design"]},"primary_area":{"value":"machine_learning_for_other_sciences_and_fields"},"abstract":{"value":"Antibody design, a crucial task with significant implications across various disciplines such as therapeutics and biology, presents considerable challenges due to its intricate nature. In this paper, we tackle antigen-specific antibody sequence-structure co-design as an optimization problem towards specific preferences, considering both rationality and functionality. Leveraging a pre-trained conditional diffusion model that jointly models sequences and structures of antibodies with equivariant neural networks, we propose direct energy-based preference optimization to guide the generation of antibodies with both rational structures and considerable binding affinities to given antigens. Our method involves fine-tuning the pre-trained diffusion model using a residue-level decomposed energy preference. Additionally, we employ gradient surgery to address conflicts between various types of energy, such as attraction and repulsion. Experiments on RAbD benchmark show that our approach effectively optimizes the energy of generated antibodies and achieves state-of-the-art performance in designing high-quality antibodies with low total energy and high binding affinity simultaneously, demonstrating the superiority of our approach."},"_bibtex":{"value":"@inproceedings{\nzhou2024antigenspecific,\ntitle={Antigen-Specific Antibody Design via Direct Energy-based Preference Optimization},\nauthor={Xiangxin Zhou and Dongyu Xue and Ruizhe Chen and Zaixiang Zheng and Liang Wang and Quanquan Gu},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=GN2GXjPyN8}\n}"},"title":{"value":"Antigen-Specific Antibody Design via Direct Energy-based Preference Optimization"},"pdf":{"value":"/pdf/1707cccb06a5edc814908e30e85b89e886aed8f5.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"zhou|antigenspecific_antibody_design_via_direct_energybased_preference_optimization"},"authorids":{"value":["~Xiangxin_Zhou1","~Dongyu_Xue1","~Ruizhe_Chen3","~Zaixiang_Zheng2","~Liang_Wang3","~Quanquan_Gu1"]},"authors":{"value":["Xiangxin Zhou","Dongyu Xue","Ruizhe Chen","Zaixiang Zheng","Liang Wang","Quanquan Gu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a new foundation model TimeRCD for zero-shot time series anomaly detection (TSAD) that overcomes the limitations of reconstruction-based approaches. Traditional TSAD models often fail to detect subtle or contextual anomalies and frequently misclassify complex normal patterns due to an objective mismatch between reconstruction and anomaly detection. To address this, the authors propose a new pre-training paradigm called Relative Context Discrepancy (RCD), which trains the model to identify anomalies by comparing discrepancies between adjacent time windows rather than reconstructing inputs. Built on a standard Transformer architecture, TimeRCD captures relational dependencies across time and uses a large-scale, fully labeled synthetic corpus that provides diverse anomaly patterns for robust pre-training. Extensive experiments across several datasets demonstrate that TimeRCD achieves state-of-the-art performance in zero-shot settings and remains competitive with fully supervised baselines."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"S1. TimeRCD demonstrates impressive performance in strict zero-shot settings, outperforming or matching specialized and fully supervised models on diverse datasets.\n\nS2. The paper designs a large-scale, fully labeled synthetic corpus that captures diverse and complex anomaly types, including point, contextual, collective, and cross-variate anomalies.\n\nS3. The experiments demonstrate that performance improves predictably with larger synthetic datasets, indicating scalable pre-training benefits similar to those observed in NLP and vision foundation models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1. While the paper includes some ablation analysis (e.g., testing the effect of synthetic data), it lacks fine-grained ablations to isolate the contribution of individual components within the TimeRCD framework. For instance, the effect of auxiliary anomaly head and RCD window  is not examined in detail. Without such targeted ablations, it is difficult to clearly attribute which design choices most contribute to performance gains.\n\nW2. The paper does not provide sufficient efficiency or scalability analyses. Since TimeRCD relies on extremely long context windows (up to 13k timesteps), understanding its computational feasibility is crucial for real-world deployment. The absence of such experiments leaves open questions about whether the model can scale efficiently to resource-constrained or latency-sensitive environments.\n\nW3. Although the paper compares against several established baselines, it omits more recent or stronger full-shot models—such as modern deep contextual detectors (e.g., DCdetector[1] and TFMAE[2]). Including these would provide a clearer picture of where TimeRCD stands relative to the current state of the art in both zero-shot and fully supervised regimes.\n\n[1] DCdetector: Dual Attention Contrastive Representation Learning for Time Series Anomaly Detection\n\n[2] Temporal-Frequency Masked Autoencoders for Time Series Anomaly Detection"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920635088,"tcdate":1761558080989,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8873/Reviewer_YVt4"],"signatures":["ICLR.cc/2026/Conference/Submission8873/Reviewer_YVt4"],"forum":"Z4T26VztkU","number":1,"license":"CC BY 4.0","cdate":1761558080989,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8873/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920635088,"domain":"ICLR.cc/2026/Conference","replyto":"Z4T26VztkU","id":"V3eh2PY5wx","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Anomaly Detection for Time Series","Time-series Foundational Model","Zero-shot","Supervised Learning"]},"supplementary_material":{"value":"/attachment/a33143782a4140a903eefb5cf450856578004361.pdf"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Time series anomaly detection (TSAD) is a critical task, but developing models that generalize to unseen data in a zero-shot manner remains a major challenge. Prevailing foundation models for TSAD predominantly rely on reconstruction-based objectives, which suffer from a fundamental objective mismatch: they struggle to identify subtle anomalies while often misinterpreting complex normal patterns, leading to high rates of false negatives and positives. To overcome these limitations, we introduce TimeRCD, a novel foundation model for TSAD built upon a new pre-training paradigm: Relative Context Discrepancy (RCD). Instead of learning to reconstruct inputs, TimeRCD is explicitly trained to identify anomalies by detecting significant discrepancies between adjacent time windows. This relational approach, implemented with a standard Transformer architecture, enables the model to capture contextual shifts indicative of anomalies that reconstruction-based methods often miss. To facilitate this paradigm, we develop a large-scale, diverse synthetic corpus with token-level anomaly labels, providing the rich supervisory signal necessary for effective pre-training. Extensive experiments demonstrate that TimeRCD significantly outperforms existing general-purpose and anomaly-specific foundation models in zero-shot TSAD across diverse datasets. Our results validate the superiority of the RCD paradigm and establish a new, effective path toward building robust and generalizable foundation models for time series anomaly detection. The code is available in https://anonymous.4open.science/r/TimeRCD-5BE1/"},"_bibtex":{"value":"@misc{\nlan2026towards,\ntitle={Towards Foundation Models for Zero-Shot Time Series Anomaly Detection: Leveraging Synthetic Data and Relative Context Discrepancy},\nauthor={Tian Lan and Hao Duong Le and Jinbo Li and Wenjun He and Meng Wang and Chenghao Liu and Chen Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=Z4T26VztkU}\n}"},"title":{"value":"Towards Foundation Models for Zero-Shot Time Series Anomaly Detection: Leveraging Synthetic Data and Relative Context Discrepancy"},"pdf":{"value":"/pdf/20b101da3013b2a492a76e2c9726e2cc076b764a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"lan|towards_foundation_models_for_zeroshot_time_series_anomaly_detection_leveraging_synthetic_data_and_relative_context_discrepancy"},"authorids":{"value":["~Tian_Lan10","~Hao_Duong_Le1","~Jinbo_Li2","~Wenjun_He1","~Meng_Wang49","~Chenghao_Liu1","~Chen_Zhang2"]},"authors":{"value":["Tian Lan","Hao Duong Le","Jinbo Li","Wenjun He","Meng Wang","Chenghao Liu","Chen Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a hierarchical policy to help LLMs to tackle complex reasoning problems. This policy consists of (1) a high-level leader to explore solution direction and a low-level follower to generate a detailed solution, and (2) a tournament-based approach to select desired reasoning chains during exploration. All the modules are implemented by prompting large language models without additional model training. The results show that the hierarchical policy is able to achieve better accuracy on solving complex math question tasks compared to several SOTA approaches."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- Experiment results have interesting findings that: when the number of reasoning chain samples increases, the increase of recall of correct solution and accuracy of the final answer is not aligned. This reveals the potential of LLMs to solve complex questions and can provide insights to other future researchers.\n\n- Results show that the hierarchical policy outperforms other prior approaches in solving complex math questions.\n\n- The paper is well-written. The problem is well-motivated by grounding on the prior work and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The evaluation datasets are not comprehensive enough. The authors only evaluate the approaches on a single dataset (the MATH dataset), and the relative evaluation size is small. Other math datasets (e.g., GSM8K, PRM800K) or other domain datasets in MMLU (e.g., Physics, Chemistry, etc) should be evaluated to demonstrate the generalization.\n\n- It is uncertain whether the improvement is from the hierarchical policy or the \"self-evaluation\" process when choosing the better reasoning chains. Previous research (e.g., \"Language Models (Mostly) Know What They Know\") suggests that LLMs like GPT-4 possess the ability to assess the likelihood that their output is correct."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"- What's the motivation of the \"Grouped-Majority Recall\" metric? A more intuitive idea may be the percentage of questions whose ground truth answer exists in at least one of the answers.\n\n- In the \"tournament-based approach\", GPT-4 is used to select the better reasoning chain as the final solution. Because of the \"self-evaluation\" ability of LLMs, have you tried to use the majority vote strategy to obtain the final answer as an ablation experiment and compute the accuracy? In GPT-3.5 based approaches, is the \"tournament\" based on GPT-4 or GPT-3.5 (you mentioned that the GPT-4 is prompted to compare the current chains with (i+1)-th chain (Section 3) )?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635971723,"tcdate":1698771806389,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission453/Reviewer_Yers"],"signatures":["ICLR.cc/2024/Conference/Submission453/Reviewer_Yers"],"forum":"adQ2YC2IV7","number":2,"license":"CC BY 4.0","cdate":1698771806389,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission453/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635971723,"domain":"ICLR.cc/2024/Conference","replyto":"adQ2YC2IV7","id":"WlDYuJrR0G","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large Language Models","Reasoning","Hierarchical Policy"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Large Language Models (LLMs) have achieved tremendous progress, yet they still often struggle with challenging reasoning problems. Current approaches address this challenge by sampling or searching detailed and low-level reasoning chains. However, these methods are still limited in their exploration capabilities, making it challenging for correct solutions to stand out in the huge solution space. In this work, we unleash LLMs' creative potential for exploring multiple diverse problem solving strategies by framing an LLM as a hierarchical policy via in-context learning. This policy comprises of a visionary leader that proposes multiple diverse high-level problem-solving tactics as hints, accompanied by a follower that executes detailed problem-solving processes following each of the high-level instruction. The follower uses each of the leader's directives as a guide and samples multiple reasoning chains to tackle the problem, generating a solution group for each leader proposal. Additionally, we propose an effective and efficient tournament-based approach to select among these explored solution groups to reach the final answer. Our approach produces meaningful and inspiring hints, enhances problem-solving strategy exploration, and improves the final answer accuracy on challenging problems in the MATH dataset."},"_bibtex":{"value":"@misc{\nling2024unleashing,\ntitle={Unleashing the Creative Mind: Language Model As Hierarchical Policy For Improved Exploration on Challenging Problem Solving},\nauthor={Zhan Ling and Yunhao Fang and Xuanlin Li and Tongzhou Mu and Mingu Lee and Reza Pourreza and Roland Memisevic and Hao Su},\nyear={2024},\nurl={https://openreview.net/forum?id=adQ2YC2IV7}\n}"},"title":{"value":"Unleashing the Creative Mind: Language Model As Hierarchical Policy For Improved Exploration on Challenging Problem Solving"},"pdf":{"value":"/pdf/5645b6271ab72c5bbc719ddc190bb28ff36f820d.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"ling|unleashing_the_creative_mind_language_model_as_hierarchical_policy_for_improved_exploration_on_challenging_problem_solving"},"authorids":{"value":["~Zhan_Ling2","~Yunhao_Fang1","~Xuanlin_Li1","~Tongzhou_Mu1","~Mingu_Lee1","~Reza_Pourreza1","~Roland_Memisevic1","~Hao_Su1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhan Ling","Yunhao Fang","Xuanlin Li","Tongzhou Mu","Mingu Lee","Reza Pourreza","Roland Memisevic","Hao Su"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a thermodynamic interpretation of deep learning where gradient descent is viewed as an energy-entropy exchange process. The authors formalize training as free-energy minimization, where loss acts as energy, parameter entropy as disorder, and learning rate as effective temperature. Experiments on synthetic datasets with MLP/CNN/Transformer architectures show entropy declining during training, which the authors interpret as \"cooling\" toward equilibrium states."},"correctness":{"value":"2: Fair"},"soundness":{"value":"2: Fair"},"strengths":{"value":"**TL;DR**: The paper provides an interesting physical metaphor connecting optimization dynamics to thermodynamic principles and presents a unified view linking information geometry, flat minima research, and energy-based models.\n\n**Long form**:\n- (major) Provides a cohesive framework (EEF) that connects several existing concepts (information bottleneck, flat minima, energy-based models) under a single thermodynamic interpretation, potentially offering intuitive insights for practitioners\n- (minor) Clear presentation with helpful visualizations showing entropy and energy trajectories across different architectures\n- (minor) Proposes concrete, actionable schedules (Eq. 3) and phase-shift detection mechanisms based on the thermodynamic view"},"weaknesses":{"value":"**TL;DR**: The contributions are very unclear as the main components (entropy regularization, flat minima preference, thermodynamically-inspired schedules) are well-established. Critical experimental validation is compromised by using an undisclosed synthetic dataset.\n\n**Long form**:\n- (major) Lack of novelty: The paper reframes existing techniques without clear added value. Weight decay as Gaussian entropy regularization (Eq. 4) is standard; flat minima improving generalization is well-known [2-4]; SAM already encourages high-entropy flat regions.\n- (major) No new algorithmic contribution: The proposed EEF-SGD (Eq. 4) reduces to standard weight decay. The \"thermodynamic\" interpretation doesn't lead to new optimization methods that outperform existing approaches\n- (major) Prior work on thermodynamic learning: Learning rate schedules inspired by physics already exist (e.g., SGDR/cosine annealing from simulated annealing, cyclical learning rates). The paper doesn't clearly distinguish its contribution from these\n- (major) Non-reproducible experiments: Uses an undisclosed \"synthetic, learnable image dataset (no downloads)\" making results impossible to verify or reproduce. This violates basic scientific standards\n- (major) Weak empirical validation: Only toy synthetic experiments on small models. No comparison with baseline schedules, no ablations, no large-scale validation to demonstrate whether the thermodynamic view provides practical benefits\n- (major) Questionable entropy proxy: Assumes Gaussian parameter distributions for entropy calculation (acknowledged in §6), which may not hold, making the \"thermodynamic\" measurements potentially meaningless\n- (minor) The metaphor linking \"universe organizing energy\" to \"networks organizing uncertainty\" (§1) is vague and doesn't add scientific rigor\n- (minor) Missing comparisons: How does the proposed T_k schedule (Eq. 3) compare quantitatively to standard cosine/step decay schedules?\n- (minor) Phase transitions are mentioned but not rigorously defined or detected algorithmically\n- (minor) The effective temperature interpretation of learning rate, while intuitive, isn't validated (e.g., does it predict generalization better than existing theory?)\n- (minor) No discussion of when/why the thermodynamic analogy might break down or be misleading"},"confidence":{"value":4},"rating":{"value":2}},"parentInvitations":"NLDL.org/2026/Abstracts_Track/-/Official_Review","nonreaders":[],"tmdate":1762345638130,"tcdate":1761986645170,"writers":["NLDL.org/2026/Abstracts_Track","NLDL.org/2026/Abstracts_Track/Submission43/Reviewer_xLGN"],"signatures":["NLDL.org/2026/Abstracts_Track/Submission43/Reviewer_xLGN"],"forum":"6aU0wfyIoz","number":2,"license":"CC BY 4.0","cdate":1761986645170,"readers":["everyone"],"invitations":["NLDL.org/2026/Abstracts_Track/Submission43/-/Official_Review","NLDL.org/2026/Abstracts_Track/-/Edit"],"mdate":1762345638130,"domain":"NLDL.org/2026/Abstracts_Track","replyto":"6aU0wfyIoz","id":"fvej0WEDH9","forumContent":{},"version":2},{"content":{"summary":{"value":"I am by no means a physics expert. I tried my best to review this paper from the viewpoint of ICLR. \n\nThis paper presents JetBench, a benchmark for multi-parameter classification of relativistic heavy-ion collision events. Each event is represented as a 32×32 jet image, labeled with three physics parameters: (i) Energy loss module, (ii) Strong coupling constant, and (iii) Virtuality separation scale (Q0).\n\nThe dataset, built from JETSCAPE simulations, contains 7.2 million aggregated jet-event images. The authors benchmark several modern vision architectures, namely EfficientNetV2, ConvNeXt V2, ViT-CoMer, Swin V2, and Mamba, for joint prediction of these parameters.  A key methodological component is the Virtual Image Aggregation technique, which averages multiple sparse jet events. The paper also analyzes calibration and confusion patterns, showing that model errors follow smooth, physically consistent transitions. ViT-CoMer achieves the best overall performance."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I do not really have any specific questions. The authors could try to convince us about why this paper is relevant for the ICLR community."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"Comprehensive benchmarking: Systematic evaluation of multiple model families under unified settings.\n\nClear methodology and reproducibility: Dataset construction, pseudocode, and ablation details are well documented.\n\nCross-disciplinary value: Establishes a link between deep vision models and heavy-ion collision physics, potentially useful for domain researchers.\n\nHigh-quality presentation: The paper is well written, figures are clear, and experiments are carefully organized."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited ML novelty: The work benchmarks existing architectures without introducing new algorithms, losses, or theoretical insights.\n\nNo statistical robustness: Single-seed results are reported without standard deviations or significance testing.\n\nDomain specificity: The contribution’s relevance to the ICLR audience is narrow, as the study’s core advances lie in computational physics rather than machine learning.\n\nThe paper is a well-executed application and benchmarking study, but its contribution to ICLR topics is incremental. It will be valuable for the physics community as a dataset and benchmarking resource, yet it lacks methodological or conceptual innovation aligned with ICLR’s core scope."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924968130,"tcdate":1762022763520,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14583/Reviewer_5ZrH"],"signatures":["ICLR.cc/2026/Conference/Submission14583/Reviewer_5ZrH"],"forum":"Hcux3hyTMH","number":3,"license":"CC BY 4.0","cdate":1762022763520,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14583/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924968130,"domain":"ICLR.cc/2026/Conference","replyto":"Hcux3hyTMH","id":"PzRgma17Kl","forumContent":{"TLDR":{"value":"We study multi-task learning on simulated heavy-ion collisions, benchmarking CNNs, Transformers, and state-space models. Models reach ~100% (energy), ~95% ($\\alpha_s$), ~78% ($Q_0$). Loss weighting highlights inter-task trade-offs."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-Task Learning","Multi-Parameter Classification","Benchmark Datasets","Vision Transformers","State Space Models","Loss Weighting in Multi-Task Learning","Model Calibration and Generalization","Training Dynamics and Optimization","Scientific Machine Learning","Relativistic Heavy Ion Collisions","Quark–Gluon Plasma","Physics-Informed Deep Learning"]},"supplementary_material":{"value":"/attachment/79fc4ac9e88bc1d55c992b467e675ecce3afc966.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Relativistic heavy-ion collisions provide a window into quark--gluon plasma formation, but extracting parameters such as the energy loss mechanism, strong coupling $\\alpha_s$, and virtuality scale $Q_0$ has traditionally required costly Bayesian inference. We introduce \\textbf{JetBench}, a benchmark for multi-parameter classification of heavy-ion events using the ANONYMIZED dataset. Each event is encoded as a $32\\times32$ jet image with three targets: energy loss module, $\\alpha_s$, and $Q_0$. We systematically evaluate CNNs (EfficientNetV2, ConvNeXt V2), Transformers (ViT-CoMer, Swin V2), and state space models (Mamba) under unified training. Results show saturated performance on energy loss ($\\sim$100\\%), strong accuracy on $\\alpha_s$ ($\\sim$95\\%), and up to $78\\%$ on $Q_0$, with ViT-CoMer achieving the best joint accuracy (74.5\\%). Loss-weight ablations reveal trade-offs between tasks, with $Q_0$ emphasis improving recall at modest cost to $\\alpha_s$. Probability calibration confirms errors follow physics continuity (e.g., $\\alpha_s=0.2/0.3$, $Q_0=2.0/2.5$). These findings establish JetBench as a scalable complement to Bayesian approaches. Code and preprocessing scripts are available at ANONYMIZED url."},"_bibtex":{"value":"@misc{\nmehryar2026jetbench,\ntitle={JetBench: Benchmarking Vision Models for Jet Observables' Classification in Heavy-Ion Physics},\nauthor={Haydar Mehryar and Loren Schwiebert},\nyear={2026},\nurl={https://openreview.net/forum?id=Hcux3hyTMH}\n}"},"title":{"value":"JetBench: Benchmarking Vision Models for Jet Observables' Classification in Heavy-Ion Physics"},"pdf":{"value":"/pdf/82984570c1c0025e4fde3e4e9c48741f96608c8a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"mehryar|jetbench_benchmarking_vision_models_for_jet_observables_classification_in_heavyion_physics"},"authorids":{"value":["~Haydar_Mehryar1","~Loren_Schwiebert1"]},"authors":{"value":["Haydar Mehryar","Loren Schwiebert"]}},"version":2},{"content":{"summary":{"value":"In this work, author try to augment simulator with DNN to learn hidden non-linear dynamics not captured by first principle solver. \nprior work explicitly define a system of simulator + DNN learned residual. \nthis work propose an implicit combination of simulator and dnn modules, integrate a dnn into the numerical solver."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- could this framework be integrated into other simulation engines, are there any constraints on this?\n- since here one need to backpropagation through the DNN within the numerical solver, would it be a bottleneck for large system?\n- how does it work for complex simulation on meshes?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- introduce hybrid simulation engine that integrate DNN into numerical methods to enable reusable and data-efficient learned modules. \n- hybrid method with more accuracy grounding \n- speed up simulation by model complex part with DNN"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"demonstrate on relative simple physics-based model, not sure how this work for more complex system with non-linear/complex physics-based simulator."}},"nonreaders":[],"tmdate":1731427938500,"tcdate":1730881130980,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4821/Reviewer_6XvB"],"signatures":["ICLR.cc/2025/Conference/Submission4821/Reviewer_6XvB"],"forum":"sSWiZr8QU7","number":5,"license":"CC BY 4.0","cdate":1730881130980,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4821/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427938500,"domain":"ICLR.cc/2025/Conference","replyto":"sSWiZr8QU7","id":"Bse5ght03a","forumContent":{"TLDR":{"value":"We present a new simulation paradigm that directly integrates DNNs with numerical engines of physics-based solvers to enable simulation of a fully implicit gray box modeling."},"venue":{"value":"ICLR 2025 Conference Desk Rejected Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["gray box modeling","simulation","neural networks"]},"supplementary_material":{"value":"/attachment/d46e857012ac5443346dbdb14f6d311a7a70729f.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Simulation is vital for scientific and engineering disciplines, as it enables the prediction and design of physical systems. However, the computational challenges inherent to large-scale simulations often arise from complex device models featuring high degrees of nonlinearities or hidden physical behaviors not captured by first principles. Gray-box models that combine deep neural networks (DNNs) with physics-based models have been proposed to address the computational challenges in modeling complex physical systems. A well-crafted gray box model capitalizes on the interpretability and accuracy of a physical model while incorporating deep neural networks to capture hidden physical behaviors and mitigate computational load associated with highly nonlinear components. Previously, gray box models have been constructed by defining an explicit combination of physics-based and black-box models to represent the behavior of sub-systems; however this alone cannot represent the coupled interactions that define the behavior of the entire physical system. We, therefore, explore an implicit gray box model, where both DNNs (trained on measurement and simulated data) and physical equations share a common set of state-variables. While this approach captures coupled interactions at the boundary of data-driven and physics-based models, simulating the implicit gray box model remains an open-ended problem. In this work, we introduce a new hybrid simulation that directly integrates DNNs into the numerical solvers of simulation engines to fully simulate implicit gray box models of large physical systems. This is accomplished by backpropagating through the DNN to calculate specific Jacobian values during each iteration of the numerical method. The hybrid simulation of implicit gray-box models improves the accuracy and runtime compared to full physics-based simulation and enables reusable DNN models with lower data requirements for training. For demonstration, we explore the advantages of this approach as compared to physics-based, black box, and other gray box methods for simulating the steady-state and electromagnetic transient behavior of power systems."},"_bibtex":{"value":"@misc{\nagarwal2024a,\ntitle={A Hybrid Simulation of {DNN}-based Gray Box Models},\nauthor={Aayushya Agarwal and Yihan Ruan and Lawrence Pileggi},\nyear={2024},\nurl={https://openreview.net/forum?id=sSWiZr8QU7}\n}"},"title":{"value":"A Hybrid Simulation of DNN-based Gray Box Models"},"pdf":{"value":"/pdf/91be2685b8516df0dec43e00cb865f1ad6f4029a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"agarwal|a_hybrid_simulation_of_dnnbased_gray_box_models"},"authorids":{"value":["~Aayushya_Agarwal1","~Yihan_Ruan1","~Lawrence_Pileggi1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Aayushya Agarwal","Yihan Ruan","Lawrence Pileggi"]}},"version":2},{"content":{"summary":{"value":"This paper explores the robustness of Graph Neural Networks (GNNs) against adversarial attacks. Drawing inspiration from principles in physics, the authors propose a novel model called Hamiltonian Neural Flows for constructing GNN models. The effectiveness of the proposed method is evaluated on various benchmark datasets.\n\n"},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"The problem addressed in this paper is both significant and intriguing, and the proposed method is well-supported by principles from physics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I am a non-expert in physics-based methods, my evaluation is based on my understanding of GNN methods and educated guesses. I suggest AC to ignore my evaluation if there are expert reviewers in the field.\n\nHere are a couple of suggestions to improve the paper:\n\nIn my opinion, the comparison with existing baselines seems inadequate. The TDGIA paper was published in 2021, while MetaGIA was published in early 2022 or late 2021. It would be beneficial if the authors could include more recent graph attack papers and compare their proposed method against them.\n\nAdditionally, it would be interesting to see a comparison between the proposed method and the Graph Isomorphism Network (GIN)."},"confidence":{"value":"1: Your assessment is an educated guess. The submission is not in your area or the submission was difficult to understand. Math/other details were not carefully checked."},"questions":{"value":"see cons above"},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1702411119097,"tcdate":1688705560546,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission7559/Reviewer_Yk3z"],"signatures":["NeurIPS.cc/2023/Conference/Submission7559/Reviewer_Yk3z"],"forum":"xtADRDRsM2","number":3,"license":"CC BY 4.0","cdate":1688705560546,"mdate":1702411119097,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission7559/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"xtADRDRsM2","id":"7PRWNrRySc","forumContent":{"venue":{"value":"NeurIPS 2023 spotlight"},"keywords":{"value":["adversarial robustness","graph neural networks"]},"supplementary_material":{"value":"/attachment/ac853cb7c389364364f49d080ef61de51593a175.zip"},"_bibtex":{"value":"@inproceedings{\nzhao2023adversarial,\ntitle={Adversarial Robustness in Graph Neural Networks: A Hamiltonian Approach},\nauthor={Kai Zhao and Qiyu Kang and Yang Song and Rui She and Sijie Wang and Wee Peng Tay},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=xtADRDRsM2}\n}"},"title":{"value":"Adversarial Robustness in Graph Neural Networks: A Hamiltonian Approach"},"paperhash":{"value":"zhao|adversarial_robustness_in_graph_neural_networks_a_hamiltonian_approach"},"abstract":{"value":"Graph neural networks (GNNs) are vulnerable to adversarial perturbations, including those that affect both node features and graph topology. This paper investigates GNNs derived from diverse neural flows, concentrating on their connection to various stability notions such as BIBO stability, Lyapunov stability, structural stability, and conservative stability. We argue that Lyapunov stability, despite its common use, does not necessarily ensure adversarial robustness. Inspired by physics principles, we advocate for the use of conservative Hamiltonian neural flows to construct GNNs that are robust to adversarial attacks. The adversarial robustness of different neural flow GNNs is empirically compared on several benchmark datasets under a variety of adversarial attacks. Extensive numerical experiments demonstrate that GNNs leveraging conservative Hamiltonian flows with Lyapunov stability substantially improve robustness against adversarial perturbations. The implementation code of experiments  is available at \\url{https://github.com/zknus/NeurIPS-2023-HANG-Robustness}."},"pdf":{"value":"/pdf/7d9d2708759abbade7f882eec2fad4f890aec1ae.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Kai_Zhao7","~Qiyu_Kang2","~Yang_Song7","~Rui_She1","~Sijie_Wang1","~Wee_Peng_Tay1"]},"authors":{"value":["Kai Zhao","Qiyu Kang","Yang Song","Rui She","Sijie Wang","Wee Peng Tay"]}},"version":2},{"content":{"comment":{"value":"**W3:** It is questionable that whether the findings on the small Transfomers can be generalized to more complex and sizable architectures.\n\n**A for W3:** We evaluated the performance of GPT2-small (12 layers, 12 heads, 117M parameters) and GPT2-large (36 layers, 20 heads, 774M parameters) on the FTCT data, with results detailed in Appendix J [Lines 1502-1611] of the revised version. These models exhibit the same pattern: compositional reasoning ability emerges with increased shot numbers and relative knowledge ratios, suggesting our findings are generalizable to more complex and sizable architectures. However, the performance of these larger models is less stable compared to smaller ones, likely due to overfitting.\n    \n\n[1] Press, Ofir, et al. \"Measuring and narrowing the compositionality gap in language models.\" *arXiv preprint arXiv:2210.03350* (2022).\n\n[2] Zhou, Denny, et al. \"Least-to-most prompting enables complex reasoning in large language models.\" *arXiv preprint arXiv:2205.10625* (2022).\n\n[3] Khot, Tushar, et al. \"Decomposed prompting: A modular approach for solving complex tasks.\" *arXiv preprint arXiv:2210.02406* (2022).\n\n[4] Bubeck, Sébastien, et al. \"Sparks of artificial general intelligence: Early experiments with gpt-4.\" *arXiv preprint arXiv:2303.12712* (2023).\n\n[5] Chan, Stephanie, et al. \"Data distributional properties drive emergent in-context learning in transformers.\" *Advances in Neural Information Processing Systems* 35 (2022): 18878-18891.\n\n[6] Allen-Zhu, Zeyuan, and Yuanzhi Li. \"Physics of language models: Part 1, context-free grammar.\" *arXiv preprint arXiv:2305.13673* (2023).\n\n[7] Bietti, Alberto, et al. \"Birth of a transformer: A memory viewpoint.\" *Advances in Neural Information Processing Systems* 36 (2024).\n\n[8] Hupkes, Dieuwke, et al. \"Compositionality decomposed: How do neural networks generalise?.\" *Journal of Artificial Intelligence Research* 67 (2020): 757-795.\n\n[9] Arora, Sanjeev, and Anirudh Goyal. \"A theory for emergence of complex skills in language models.\" *arXiv preprint arXiv:2307.15936* (2023).\n\n[10] Yu, Dingli, et al. \"Skill-Mix: A flexible and expandable family of evaluations for AI models.\" *arXiv preprint arXiv:2310.17567*(2023).\n\n[11] Xu, Zhuoyan, Zhenmei Shi, and Yingyu Liang. \"Do large language models have compositional ability? an investigation into limitations and scalability.\" *ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models*. 2024.\n\n[12] Treutlein, Johannes, et al. \"Connecting the dots: Llms can infer and verbalize latent structure from disparate training data.\" *arXiv preprint arXiv:2406.14546* (2024).\n\n[13] Ramesh, Rahul, et al. \"How capable can a transformer become? a study on synthetic, interpretable tasks.\" *arXiv preprint arXiv:2311.12997* (2023).\n\n[14] Ye, Tian, et al. \"Physics of language models: Part 2.1, grade-school math and the hidden reasoning process.\" *arXiv preprint arXiv:2407.20311* (2024)."},"title":{"value":"Response for Reviewer EQsG (PART 2)"}},"tmdate":1732226197210,"tcdate":1732226197210,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2678/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission2678/Authors"],"forum":"1Xg4JPPxJ0","number":8,"license":"CC BY 4.0","cdate":1732226197210,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2678/-/Official_Comment"],"mdate":1732226197210,"domain":"ICLR.cc/2025/Conference","replyto":"bJy3XKyRX3","id":"SFB9nijFGP","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Transformer; Chain-of-Thought; In-Context-Learning; Compositional Generalization"]},"supplementary_material":{"value":"/attachment/b8c4e05383b92ddd3e548442492bdbb451408314.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Humans exhibit remarkable compositional reasoning by integrating knowledge from various sources. For example, if someone learns ( B = f(A) ) from one source and ( C = g(B) ) from another, they can deduce ( C=g(B)=g(f(A)) ) even without encountering ( ABC ) together, showcasing the generalization ability of human intelligence.  In this paper, we introduce a synthetic learning task, \"FTCT\" (Fragmented at Training, Chained at Testing), to validate the potential of  Transformers in replicating this skill and interpret its inner mechanism. During training, data consist of separated knowledge fragments from an overall causal graph. In testing, Transformers must combine these fragments to infer complete causal traces. Our findings demonstrate that few-shot Chain-of-Thought prompting enables Transformers to perform compositional reasoning on FTCT by revealing correct combinations of fragments, even if such combinations were absent in training data. Furthermore, the emergence of compositional reasoning ability is strongly correlated with model complexity and training-testing data similarity. We propose, both theoretically and empirically, that Transformers learn an underlying generalizable program from training, enabling effective compositional reasoning during testing."},"_bibtex":{"value":"@inproceedings{\nyin2025are,\ntitle={Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?},\nauthor={Yutong Yin and Zhaoran Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=1Xg4JPPxJ0}\n}"},"title":{"value":"Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?"},"pdf":{"value":"/pdf/2d58475d0979133db2c7fb27c2e5abc6899ae229.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"yin|are_transformers_able_to_reason_by_connecting_separated_knowledge_in_training_data"},"authorids":{"value":["~Yutong_Yin1","~Zhaoran_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yutong Yin","Zhaoran Wang"]}},"version":2},{"content":{"venue":{"value":"ICML 2024"},"pdf":{"value":"https://raw.githubusercontent.com/mlresearch/v235/main/assets/cho24b/cho24b.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"cho|parameterized_physicsinformed_neural_networks_for_parameterized_pdes"},"html":{"value":"https://proceedings.mlr.press/v235/cho24b.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icml/ChoJLL00P24,\n  author={Woojin Cho and Minju Jo and Haksoo Lim and Kookjin Lee and Dongeun Lee and Sanghyun Hong and Noseong Park},\n  title={Parameterized Physics-informed Neural Networks for Parameterized PDEs},\n  year={2024},\n  cdate={1704067200000},\n  pages={8510-8533},\n  url={https://proceedings.mlr.press/v235/cho24b.html},\n  booktitle={ICML},\n  crossref={conf/icml/2024}\n}\n"},"abstract":{"value":"Complex physical systems are often described by partial differential equations (PDEs) that depend on parameters such as the Raynolds number in fluid mechanics. In applications such as design optimization or uncertainty quantification, solutions of those PDEs need to be evaluated at numerous points in the parameter space. While physics-informed neural networks (PINNs) have emerged as a new strong competitor as a surrogate, their usage in this scenario remains underexplored due to the inherent need for repetitive and time-consuming training. In this paper, we address this problem by proposing a novel extension, parameterized physics-informed neural networks (P$^2$INNs). P$^2$INNs enable modeling the solutions of parameterized PDEs via explicitly encoding a latent representation of PDE parameters. With the extensive empirical evaluation, we demonstrate that P$^2$INNs outperform the baselines both in accuracy and parameter efficiency on benchmark 1D and 2D parameterized PDEs and are also effective in overcoming the known “failure modes”."},"title":{"value":"Parameterized Physics-informed Neural Networks for Parameterized PDEs"},"authors":{"value":[{"fullname":"Woojin Cho","username":"~Woojin_Cho1"},{"fullname":"Minju Jo","username":""},{"fullname":"Haksoo Lim","username":""},{"fullname":"Kookjin Lee","username":""},{"fullname":"Dongeun Lee","username":""},{"fullname":"Sanghyun Hong","username":""},{"fullname":"Noseong Park","username":""}]}},"tmdate":1784190894600,"pdate":1735603200000,"externalIds":["dblp:conf/icml/ChoJLL00P24"],"tcdate":1784190890716,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Woojin_Cho1"],"forum":"1QWaIrHKGg","license":"CC BY-SA 4.0","number":59171,"cdate":1704067200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784190894600,"domain":"OpenReview.net/Public_Article","id":"1QWaIrHKGg","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"https://arxiv.org/pdf/2408.09446v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"cho|parameterized_physicsinformed_neural_networks_for_parameterized_pdes"},"html":{"value":"https://doi.org/10.48550/arXiv.2408.09446"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2408-09446,\n  publtype={informal},\n  author={Woojin Cho and Minju Jo and Haksoo Lim and Kookjin Lee and Dongeun Lee and Sanghyun Hong and Noseong Park},\n  title={Parameterized Physics-informed Neural Networks for Parameterized PDEs},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2408.09446},\n  url={https://doi.org/10.48550/arXiv.2408.09446}\n}\n"},"abstract":{"value":"Complex physical systems are often described by partial differential equations (PDEs) that depend on parameters such as the Reynolds number in fluid mechanics. In applications such as design optimization or uncertainty quantification, solutions of those PDEs need to be evaluated at numerous points in the parameter space. While physics-informed neural networks (PINNs) have emerged as a new strong competitor as a surrogate, their usage in this scenario remains underexplored due to the inherent need for repetitive and time-consuming training. In this paper, we address this problem by proposing a novel extension, parameterized physics-informed neural networks (P$^2$INNs). P$^2$INNs enable modeling the solutions of parameterized PDEs via explicitly encoding a latent representation of PDE parameters. With the extensive empirical evaluation, we demonstrate that P$^2$INNs outperform the baselines both in accuracy and parameter efficiency on benchmark 1D and 2D parameterized PDEs and are also effective in overcoming the known \"failure modes\"."},"title":{"value":"Parameterized Physics-informed Neural Networks for Parameterized PDEs"},"authors":{"value":[{"fullname":"Woojin Cho","username":"~Woojin_Cho1"},{"fullname":"Minju Jo","username":""},{"fullname":"Haksoo Lim","username":""},{"fullname":"Kookjin Lee","username":""},{"fullname":"Dongeun Lee","username":""},{"fullname":"Sanghyun Hong","username":""},{"fullname":"Noseong Park","username":""}]}},"tmdate":1784190894380,"pdate":1735603200000,"externalIds":["dblp:journals/corr/abs-2408-09446"],"tcdate":1784190890633,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Woojin_Cho1"],"forum":"3roTMF4Gxx","license":"CC BY-SA 4.0","number":59170,"cdate":1704067200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784190894380,"domain":"OpenReview.net/Public_Article","id":"3roTMF4Gxx","version":2},{"content":{"venue":{"value":"ICML 2024 Oral"},"abstract":{"value":"Complex physical systems are often described by partial differential equations (PDEs) that depend on parameters such as the Raynolds number in fluid mechanics. In applications such as design optimization or uncertainty quantification, solutions of those PDEs need to be evaluated at numerous points in the parameter space. While physics-informed neural networks (PINNs) have emerged as a new strong competitor as a surrogate, their usage in this scenario remains underexplored due to the inherent need for repetitive and time-consuming training. In this paper, we address this problem by proposing a novel extension, parameterized physics-informed neural networks (P$^2$INNs). P$^2$INNs enable modeling the solutions of parameterized PDEs via explicitly encoding a latent representation of PDE parameters. With the extensive empirical evaluation, we demonstrate that P$^2$INNs outperform the baselines both in accuracy and parameter efficiency on benchmark 1D and 2D parameterized PDEs and are also effective in overcoming the known “failure modes”."},"_bibtex":{"value":"@inproceedings{\ncho2024parameterized,\ntitle={Parameterized Physics-informed Neural Networks for Parameterized {PDE}s},\nauthor={Woojin Cho and Minju Jo and Haksoo Lim and Kookjin Lee and Dongeun Lee and Sanghyun Hong and Noseong Park},\nbooktitle={Forty-first International Conference on Machine Learning},\nyear={2024},\nurl={https://openreview.net/forum?id=n3yYrtt9U7}\n}"},"title":{"value":"Parameterized Physics-informed Neural Networks for Parameterized PDEs"},"pdf":{"value":"/pdf/1e8948a36b6442ae1c500a2faab703c6283348f1.pdf"},"venueid":{"value":"ICML.cc/2024/Conference"},"paperhash":{"value":"cho|parameterized_physicsinformed_neural_networks_for_parameterized_pdes"},"authorids":{"value":["~Woojin_Cho1","~Minju_Jo1","~Haksoo_Lim1","~Kookjin_Lee1","~Dongeun_Lee1","~Sanghyun_Hong1","~Noseong_Park1"]},"authors":{"value":["Woojin Cho","Minju Jo","Haksoo Lim","Kookjin Lee","Dongeun Lee","Sanghyun Hong","Noseong Park"]}},"tmdate":1719388289372,"pdate":1714610380171,"tcdate":1706778033867,"writers":["ICML.cc/2024/Conference","ICML.cc/2024/Conference/Submission5087/Authors"],"signatures":["ICML.cc/2024/Conference/Submission5087/Authors"],"forum":"n3yYrtt9U7","license":"CC BY 4.0","number":5087,"cdate":1706778033867,"readers":["everyone"],"invitations":["ICML.cc/2024/Conference/-/Submission","ICML.cc/2024/Conference/-/Post_Submission","ICML.cc/2024/Conference/-/Edit","ICML.cc/2024/Conference/Submission5087/-/Camera_Ready_Revision"],"mdate":1719388289372,"odate":1717693061431,"domain":"ICML.cc/2024/Conference","id":"n3yYrtt9U7","version":2},{"content":{"summary":{"value":"This paper introduces GENIE, a hybrid scene representation that marries the editability of 3D Gaussian Splatting with the photorealistic rendering of NeRFs to enable scene editing and physics-driven manipulation. Each Gaussian consists of a learnable feature embedding that conditions a NeRF; nearby Gaussians are selected efficiently with Ray-Traced Gaussian Proximity Search (RT-GPS), and the system supports densification/pruning to balance quality and speed. This design allows direct primitive manipulation or mesh-driven deformations to immediately reflect in rendered views without fine-tuning. Experiments on NeRF-Synthetic show reconstruction on par with strong baselines and competitive with 3DGS while outperforming editable baselines like RIP-NeRF on most scenes; on Mip-NeRF 360 the method remains editable though PSNR trails the best static models. The authors provide code, configs, and reproduction details."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. The paper claims that the method is real-time (L42): but Table 4 shows that rendering is not real time. What does this claim mean?\n2. Can you incorporate metrics for editing performance as well?\n3. Why not use pure Gaussian Splatting editing? Given that Gaussian Splatting achieves better quality and faster rendering, what specific advantages justify the NeRF hybrid beyond avoiding \"gaps\" during super-resolution?\n4. Can the method handle view dependent effects, such as specular highlights, as a result of editing?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is well written and easy to follow\n2. code, configs, hardware/software details, and workflows provided.\n3. Novel hybrid architecture that meaningfully combines implicit and explicit representations, addressing a real limitation of pure NeRF methods (difficult editing)\n4. Competitive reconstruction on NeRF-Synthetic and better than editable baselines like RIP-NeRF on most scenes; rendering speed competitive given editability.\n5. Thorough ablations isolate the value of Splash Grid Encoding, RT-GPS, learnable means, and densification."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. On Mip-NeRF 360, the proposed method achieves 19.47 PSNR vs 25.25 for Gaussian Splatting on bicycle scene - a 5.78 dB drop. Similar gaps appear across all outdoor scenes (Table 2). Even on NeRF-Synthetic, it underperforms GS (34.67 vs 35.82) and LagHash. This quality sacrifice is substantial and undermines the hybrid approach's value proposition.\n2. While editing capabilities are a critical contribution of this work, there are no quantitative metrics for editing quality. Only qualitative comparison with PhysGaussian/GASP (Figure 7) are shown. Quantitative metrics, and potentially a user study would go a long way in emphasizing the editing capabilities of this method.\n3. The rendering speed (Table 4) is considerably slower than prior works.\n4. The system is extremely complex: there are many interacting components (RT-GPS with quantile Q, k neighbors, Splash Grid with multiple LoDs, pruning with confidence vectors, densification thresholds τ_s and τ_α). While ablation helps, the hyperparameter sensitivity and design choices aren't fully explored."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918586578,"tcdate":1762213735050,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6281/Reviewer_oxnh"],"signatures":["ICLR.cc/2026/Conference/Submission6281/Reviewer_oxnh"],"forum":"wyauZpoqRK","number":4,"license":"CC BY 4.0","cdate":1762213735050,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6281/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918586578,"domain":"ICLR.cc/2026/Conference","replyto":"wyauZpoqRK","id":"2ijLjCnNm6","forumContent":{"TLDR":{"value":"his work introduces GENIE, a hybrid model that fuses NeRF’s photorealism with 3D Gaussian Splatting’s editability, enabling interactive 3D scene editing via Gaussian-based conditioning and fast nearest Gaussian search."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Gaussian Splatting","NeRF","3D objects","physical simulations"]},"supplementary_material":{"value":"/attachment/b8f83c0e55c5d12a6c9fc74eec83e45ec06b0cd4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have recently transformed 3D scene representation and rendering. NeRF achieves high-fidelity novel view synthesis by learning volumetric representations through neural networks, but its implicit encoding makes editing and physical interaction challenging. In contrast, 3DGS represents scenes as explicit collections of Gaussian primitives, enabling real-time rendering, faster training, and more intuitive manipulation. This explicit structure has made 3DGS particularly well-suited for interactive editing and integration with physics-based simulation. In this paper, we introduce GENIE (Gaussian Encoding for Neural Radiance Fields Interactive Editing), a hybrid model that combines the photorealistic rendering quality of NeRF with the editable and structured representation of 3DGS. Instead of using spherical harmonics for appearance modeling, we assign each Gaussian a trainable feature embedding. These embeddings are used to condition a NeRF network based on the \n nearest Gaussians to each query point. To make this conditioning efficient, we introduce Ray-Traced Gaussian Proximity Search (RTGPS), a fast nearest Gaussian search based on a modified ray-tracing pipeline. We also integrate a multi-resolution hash grid to initialize and update Gaussian features. Together, these components enable real-time, locality-aware editing: as Gaussian primitives are repositioned or modified, their interpolated influence is immediately reflected in the rendered output. By combining the strengths of implicit and explicit representations, GENIE supports intuitive scene manipulation, dynamic interaction, and compatibility with physical simulation, bridging the gap between geometry-based editing and neural rendering."},"_bibtex":{"value":"@misc{\nzielinski2025genie,\ntitle={{GENIE}: Gaussian Encoding for Neural Radiance Field Interactive Editing},\nauthor={Miko{\\l}aj Zieli{\\'n}ski and Krzysztof Byrski and Tomasz Szczepanik and Przemys{\\l}aw Spurek},\nyear={2025},\nurl={https://openreview.net/forum?id=wyauZpoqRK}\n}"},"title":{"value":"GENIE: Gaussian Encoding for Neural Radiance Field Interactive Editing"},"pdf":{"value":"/pdf/9683e9431e57f996c107fc4e7f18a38fd6e208e7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zieliski|genie_gaussian_encoding_for_neural_radiance_field_interactive_editing"},"authorids":{"value":["~Mikołaj_Zieliński1","~Krzysztof_Byrski1","~Tomasz_Szczepanik1","~Przemysław_Spurek1"]},"authors":{"value":["Mikołaj Zieliński","Krzysztof Byrski","Tomasz Szczepanik","Przemysław Spurek"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a novel approach called CoT Inversion, designed to evaluate the faithfulness of generated Chain-of-Thought (CoT) reasoning processes. The method falls under the framework of the variational information bottleneck (VIB), in which separate encoder and decoder networks are trained. Beyond its interpretability perspective, CoT Inversion can also enhance CoT reasoning performance by aligning the inferred latent CoT representations with the generated CoT through a similarity-based optimization objective. While the underlying idea is intuitive, the practical implementation pipeline is relatively complex."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please see questions in the weaknesses section."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The core idea behind the proposed method is intuitive: the entire reasoning process should be abstracted and represented within a latent space, and this representation should be learned through the variational information bottleneck framework."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper is difficult to follow. For example:\n  \n  - Methodology: The entropy-weighted edit distance is not clearly formulated, and several notations are undefined (e.g., ( w ) in Equation 1). It is also unclear how ( p(\\tilde{c} \\mid x) ) is computed.\n  - Experiments: The implementation details of the CoT Inversion method used for comparison against the Raw CoT method are not well described.\n  - Overall pipeline: The general workflow of the proposed approach is hard to grasp, and Figure 2 is too abstract to effectively aid understanding.\n- The paper makes an intuitive but strong assumption—that the causal reasoning process in discrete space can be represented by a single latent vector. Such a claim requires comprehensive theoretical justification and empirical validation, which are lacking in the current work.\n- The reported performance improvements are marginal and inconsistent, as shown in Table 2 and Figure 3.\n- The evaluation metric is problematic. In Table 3, the authors report “accuracy based on the proportion of CoT tokens to total output tokens,” but do not include the standard accuracy metric for comparison. Furthermore, this custom metric should be formally defined."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915710925,"tcdate":1761827466701,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1216/Reviewer_vsTH"],"signatures":["ICLR.cc/2026/Conference/Submission1216/Reviewer_vsTH"],"forum":"bpmQnmnvp1","number":2,"license":"CC BY 4.0","cdate":1761827466701,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1216/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915710925,"domain":"ICLR.cc/2026/Conference","replyto":"bpmQnmnvp1","id":"LQmI0NIRr0","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["language model"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Chain-of-Thought (CoT) enables large language models (LLMs) to tackle complex reasoning tasks by generating intermediate steps. Although CoT provides opportunities for improved interpretability and facilitates the monitoring of AI safety, the consistency between generated CoT and the model's actual reasoning process is not guaranteed. Models can output seemingly reasonable CoT that fails to reflect the true computational trajectory leading to the final answer. IIn this work, we introduce a novel approach called CoT Inversion to evaluate CoT faithfulness. We cast the problem in a probabilistic framework, viewing the genuine reasoning chain as a latent variable that mediates between the input and the answer. Leveraging variational inference with a scoring function, we infer this hidden CoT by effectively reversing the model's answer generation process; under our chosen variational family, the optimization reduces to an instance of the Expectation Maximization (EM) algorithm.\nFurthermore, we propose an explicit alignment objective that promotes similarity between the inferred latent CoT and the model's directly generated CoT, considering the explicit CoT as an informative and possibly unfaithful signal. Our approach enables the quantitative assessment of agreement between articulated and inferred reasoning processes, offering a practical metric of CoT faithfulness and strengthening our ability to interpret and trust the reasoning of language models."},"_bibtex":{"value":"@misc{\nyu2025enabling,\ntitle={Enabling Reasoning Language Models to Reveal Their True Thoughts via CoT Inversion},\nauthor={Chunlin Yu and Ruiqi Xu and Chuliang Weng},\nyear={2025},\nurl={https://openreview.net/forum?id=bpmQnmnvp1}\n}"},"title":{"value":"Enabling Reasoning Language Models to Reveal Their True Thoughts via CoT Inversion"},"pdf":{"value":"/pdf/e43bd8dbb57de9cbf888d635f15263c73ec70251.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yu|enabling_reasoning_language_models_to_reveal_their_true_thoughts_via_cot_inversion"},"authorids":{"value":["~Chunlin_Yu2","~Ruiqi_Xu5","~Chuliang_Weng1"]},"authors":{"value":["Chunlin Yu","Ruiqi Xu","Chuliang Weng"]}},"version":2},{"content":{"venue":{"value":"ECML/PKDD (2) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-70344-7_24.pdf"},"venueid":{"value":"dblp.org/conf/PKDD/2024"},"paperhash":{"value":"gao|physicsinformed_spatiotemporal_model_for_human_mobility_prediction"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Quanyan_Gao:","~Chao_Li24","~Qinmin_Yang1"]},"html":{"value":"https://doi.org/10.1007/978-3-031-70344-7_24"},"_bibtex":{"value":"@inproceedings{DBLP:conf/pkdd/GaoLY24,\n  author={Quanyan Gao and Chao Li and Qinmin Yang},\n  title={Physics-Informed Spatio-Temporal Model for Human Mobility Prediction},\n  year={2024},\n  cdate={1704067200000},\n  pages={409-425},\n  url={https://doi.org/10.1007/978-3-031-70344-7_24},\n  booktitle={ECML/PKDD (2)},\n  crossref={conf/pkdd/2024-2}\n}\n"},"abstract":{"value":"Human mobility prediction is the fundamental problem in studying human social behaviors. However, current approaches overlook the dynamic physics processes inherent in the movement of human, especially in an ever-changing temporal environments. In this work, we propose a physics-informed spatial-temporal prediction network (Physics-ST) for human mobility prediction. Specifically, we employ the assumption that human mobility is driven by a potential energy filed and establish a physics-informed equation to describe the dynamics of human mobility in the city. Then, we propose a multi-head node attention graph convolutional network to express the process of potential energy transfer within human mobility between regions. To model the impact of temporal information on mobility, we introduce correction terms in the physics-informed equation and propose a history-future information reasoning module which can extrapolate the human mobility energy under the future scenarios. Experiments conducted on real-world datasets demonstrate that our model exhibits superior performance, particularly in its understanding of physics mechanism."},"title":{"value":"Physics-Informed Spatio-Temporal Model for Human Mobility Prediction"},"authors":{"value":["Quanyan Gao","Chao Li","Qinmin Yang"]}},"tmdate":1767888590132,"pdate":1704067200000,"tcdate":1738584592633,"writers":["~"],"signatures":["~Chao_Li24"],"forum":"YG0Km0K5dT","license":"CC BY-SA 4.0","number":290744,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1767888590132,"domain":"DBLP.org","id":"YG0Km0K5dT","version":2},{"content":{"summary":{"value":"The authors introduce a new interpretability framework grounded in statistical physics and Bayesian learning theory. The key idea is to view a neural network as a Bayesian statistical mechanical system, where the model’s parameters and interactions can be explored through small perturbations of the data distribution. Small shifts in the input data (like moving from natural text to programming code) produce first-order linear responses in specific components of the model. These responses (called susceptibilities) reveal how strongly and in what direction each component reacts to data changes, providing a principled way to quantify its sensitivity and functional role within the network."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"* as I mentioned in the weaknesses section, I think the paper would be stronger if it demonstrated how the proposed method could be applied to analyze more complex patterns or behaviors (such as bias or knowledge acquisition)."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"* The framework is theoretically rigorous, grounded in statistical physics and Bayesian learning theory, while also being empirically validated through concrete experiments.\n\n* The proposed approach is novel:  Susceptibility analysis connects the functional behavior of model components to shifts in the training distribution and shows that heads with similar response patterns cluster into interpretable groups."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The analysis (Sec 4) focuses on a very simple set of patterns (Word Start, Word Part, Word End, Induction Pattern, Right Delimiter). In addition, the framework is demonstrated only on a small toy language model (3M, 2 attention layers, without MLP). Although the authors note that they do not anticipate major obstacles in scaling the method to larger models, applying it to a larger model could have strengthened the work by demonstrating the ability to capture more complex behaviors.\n\n* This is only a suggestion, but I think the presentation could be made more accessible to readers who are less familiar with the theoretical background."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762932948768,"tcdate":1761781671488,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20053/Reviewer_B7pi"],"signatures":["ICLR.cc/2026/Conference/Submission20053/Reviewer_B7pi"],"forum":"J4GYMiE3JT","number":2,"license":"CC BY 4.0","cdate":1761781671488,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20053/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762932948768,"domain":"ICLR.cc/2026/Conference","replyto":"J4GYMiE3JT","id":"oEDHO9EWJU","forumContent":{"TLDR":{"value":"Introduces susceptibilities to study the internal structure of language models"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Interpretability","Statistical Physics","Singular Learning Theory"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of an observable localized on a chosen component of the network. The resulting susceptibility can be estimated efficiently with local SGLD samples and factorizes into signed, per-token contributions that serve as attribution scores. We combine these susceptibilities into a response matrix whose low-rank structure separates functional modules such as multigram and induction heads in a 3M-parameter transformer."},"_bibtex":{"value":"@inproceedings{\nbaker2026structural,\ntitle={Structural Inference: Interpreting Small Language Models with Susceptibilities},\nauthor={Garrett Baker and George Wang and Jesse Hoogland and Vinayak Pathak and Daniel Murfet},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=J4GYMiE3JT}\n}"},"title":{"value":"Structural Inference: Interpreting Small Language Models with Susceptibilities"},"pdf":{"value":"/pdf/ab5b372e641ebfacee71034a517279f9c94ed0c8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"baker|structural_inference_interpreting_small_language_models_with_susceptibilities"},"authorids":{"value":["~Garrett_Baker1","~George_Wang1","~Jesse_Hoogland1","~Vinayak_Pathak1","~Daniel_Murfet1"]},"authors":{"value":["Garrett Baker","George Wang","Jesse Hoogland","Vinayak Pathak","Daniel Murfet"]}},"version":2},{"content":{"venue":{"value":"EJNMMI Physics"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1186/s40658-024-00646-y.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"benetti|dosepatch_physicsinspired_cropping_layout_for_patchbased_monte_carlo_simulations_to_provide_fast_and_accurate_internal_dosimetry"},"html":{"value":"https://doi.org/10.1186/s40658-024-00646-y"},"abstract":{"value":"Dosimetry-based personalized therapy was shown to have clinical benefits e.g. in liver selective internal radiation therapy (SIRT). Yet, there is no consensus about its introduction into clinical practice, mainly as Monte Carlo simulations (gold standard for dosimetry) involve massive computation time. We addressed the problem of computation time and tested a patch-based approach for Monte Carlo simulations for internal dosimetry to improve parallelization. We introduce a physics-inspired cropping layout for patch-based MC dosimetry, and compare it to cropping layouts of the literature as well as dosimetry using organ-S-values, and dose kernels, taking whole-body Monte Carlo simulations as ground truth. This was evaluated in five patients receiving Yttrium-90 liver SIRT."},"title":{"value":"DosePatch: physics-inspired cropping layout for patch-based Monte Carlo simulations to provide fast and accurate internal dosimetry"},"authors":{"value":[{"fullname":"Francesca De Benetti","username":"https://orcid.org/orcid-search/search?searchQuery=Francesca%20De%20Benetti"},{"fullname":"Julia Brosch-Lenz","username":"https://orcid.org/orcid-search/search?searchQuery=Julia%20Brosch-Lenz"},{"fullname":"Jorge Mario Guerra González","username":"https://orcid.org/orcid-search/search?searchQuery=Jorge%20Mario%20Guerra%20Gonz%C3%A1lez"},{"fullname":"Carlos Uribe","username":"https://orcid.org/orcid-search/search?searchQuery=Carlos%20Uribe"},{"fullname":"Matthias Eiber","username":"https://orcid.org/orcid-search/search?searchQuery=Matthias%20Eiber"},{"fullname":"Nassir Navab","username":"https://orcid.org/orcid-search/search?searchQuery=Nassir%20Navab"},{"fullname":"Thomas Wendler","username":"~Thomas_Wendler1"}]}},"tmdate":1788300678261,"pdate":1719360000000,"externalIds":["doi:10.1186/s40658-024-00646-y"],"tcdate":1788300670032,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Thomas_Wendler1"],"forum":"vO94hpyv7K","license":"CC BY-SA 4.0","number":107333,"cdate":1719379543292,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1788300678261,"domain":"OpenReview.net/Public_Article","id":"vO94hpyv7K","version":2},{"content":{"summary":{"value":"The paper provides a theoretical characterization of the generalization error for optimal transport mappings parameterized by input-convex neural networks, focusing on a class of continuous Wasserstein-2 distance solvers. The analysis decomposes the overall error into two components, estimation error and approximation error, each of which is independently derived and subsequently combined to yield the final generalization bound. The theoretical findings are supported by experimental validation on a simple, low-dimensional synthetic dataset, which confirm the consistency of the results with the proposed theoretical framework."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"In Figure 3, the slope appears to increase as the dimensionality rises. Could the authors provide further insights into the underlying reasons for this behavior? Additionally, as the dimensionality continues to grow, do the authors expect the slope to remain below or equal to 0.5?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper makes an original theoretical contribution by deriving generalization error bounds for quadratic optimal transport (OT) solvers parameterized by neural networks, particularly within the class of input-convex models. This work provides valuable insights into the learnability and error behavior of neural OT mappings and offers practical guidance for their use in learning-based transport problems. The theoretical framework is rigorous, decomposing the overall generalization error into estimation and approximation components with clear mathematical justification.\n\nThe paper is clearly written, logically structured, and easy to follow, even when presenting technically complex material. Although the experiments are limited to simple synthetic datasets, they effectively validate the theoretical results. The significance of this work lies in establishing foundational learnability guarantees for neural OT solvers, which bridge the gap between theoretical understanding and practical implementation in modern OT-based learning frameworks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"A key limitation of the paper is its exclusive reliance on low-dimensional synthetic datasets for empirical validation. While the theoretical analysis is sound, the absence of experiments on higher-dimensional or real-world datasets limits the assessment of the framework’s practical relevance and robustness. Extending the experiments to more complex domains would provide stronger empirical support for the theoretical claims and demonstrate the scalability of the proposed bounds.\n\nAdditionally, the paper assumes \\beta-strong convexity, a condition that is often difficult to guarantee in neural network parameterizations. While the authors acknowledge this assumption as restrictive, a deeper discussion of its implications and potential violations would be valuable.\n\nSome missing references, e.g. [proof ref.] in multiple theorems and propositions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918245779,"tcdate":1762180604536,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5764/Reviewer_aRUC"],"signatures":["ICLR.cc/2026/Conference/Submission5764/Reviewer_aRUC"],"forum":"FJTdyG8jeJ","number":2,"license":"CC BY 4.0","cdate":1762180604536,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5764/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918245779,"domain":"ICLR.cc/2026/Conference","replyto":"FJTdyG8jeJ","id":"pNSVbbSWoL","forumContent":{"TLDR":{"value":"Statistical generalization bounds for (semi-dual) quadratic Neural Optimal Transport solvers"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["optimal transport","semi-dual optimal transport","statistical learning theory","approximation bounds"]},"supplementary_material":{"value":"/attachment/6c7a9f074771df7579f461a86a560c4b5ffd03db.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Neural network-based optimal transport (OT) is a recent and fruitful direction in the generative modeling community. It finds its applications in various fields such as domain translation, image super-resolution, computational biology and others. Among the existing OT approaches, of considerable interest are adversarial minimax solvers based on semi-dual formulations of OT problems. While promising, these methods lack theoretical investigation from a statistical learning perspective. Our work fills this gap by establishing upper bounds on the generalization error of an approximate OT map recovered by the minimax quadratic OT solver. Importantly, the bounds we derive depend solely on some standard statistical and mathematical properties of the considered functional classes (neural nets). While our analysis focuses on the quadratic OT, we believe that similar bounds could be derived for general OT case, paving the promising direction for future research. Our experimental illustrations are available online https://github.com/milenagazdieva/StatOT."},"_bibtex":{"value":"@inproceedings{\ntarasov2026a,\ntitle={A Statistical Learning Perspective on Semi-dual Adversarial Neural Optimal Transport Solvers},\nauthor={Roman Tarasov and Petr Mokrov and Milena Gazdieva and Evgeny Burnaev and Alexander Korotin},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=FJTdyG8jeJ}\n}"},"title":{"value":"A Statistical Learning Perspective on Semi-dual Adversarial Neural Optimal Transport Solvers"},"pdf":{"value":"/pdf/789b645db43c96f1010d5c6211d57b4dafe9227f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"tarasov|a_statistical_learning_perspective_on_semidual_adversarial_neural_optimal_transport_solvers"},"authorids":{"value":["~Roman_Tarasov1","~Petr_Mokrov1","~Milena_Gazdieva1","~Evgeny_Burnaev1","~Alexander_Korotin2"]},"authors":{"value":["Roman Tarasov","Petr Mokrov","Milena Gazdieva","Evgeny Burnaev","Alexander Korotin"]}},"version":2},{"content":{"venue":{"value":"NICE 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10548018/10548050/10548180.pdf"},"venueid":{"value":"dblp.org/conf/NICE/2024"},"paperhash":{"value":"theilman|spiking_physicsinformed_neural_networks_on_loihi_2"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Bradley_H._Theilman:","https://dblp.org/search/pid/api?q=author:Qian_Zhang:","https://dblp.org/search/pid/api?q=author:Adar_Kahana:","~Eric_C._Cyr1","https://dblp.org/search/pid/api?q=author:Nathaniel_Trask:","~James_B._Aimone1","https://dblp.org/search/pid/api?q=author:George_Em_Karniadakis:"]},"html":{"value":"https://doi.org/10.1109/NICE61972.2024.10548180"},"_bibtex":{"value":"@inproceedings{DBLP:conf/nice/TheilmanZKCTAK24,\n  author={Bradley H. Theilman and Qian Zhang and Adar Kahana and Eric C. Cyr and Nathaniel Trask and James B. Aimone and George Em Karniadakis},\n  title={Spiking Physics-Informed Neural Networks on Loihi 2},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-6},\n  url={https://doi.org/10.1109/NICE61972.2024.10548180},\n  booktitle={NICE},\n  crossref={conf/nice/2024}\n}\n"},"abstract":{"value":"Neuromorphic computing platforms hold the promise to dramatically reduce power requirements for calculations that are computationally intensive. One such application space is scientific machine learning (SciML). Techniques in this space use neural networks to approximate solutions of scientific problems. For instance, the popular physics-informed neural network (PINN) approximates the solution to a partial differential equation by using a trained feed-forward neural network, and injecting the knowledge of the physics through the loss function. Recent efforts have demonstrated how to convert a trained PINN to a spiking network architecture. In this work, we discuss our approach to quantization and implementation required to migrate these spiking PINNs to Intel’s Loihi 2 neuromorphic hardware. We explore the effect of quantization on the model accuracy, as well as the energy and throughput characteristics of the implementation. It is our intent that this serve as a starting point for additional SciML implementations on neuromorphic hardware."},"title":{"value":"Spiking Physics-Informed Neural Networks on Loihi 2"},"authors":{"value":["Bradley H. Theilman","Qian Zhang","Adar Kahana","Eric C. Cyr","Nathaniel Trask","James B. Aimone","George Em Karniadakis"]}},"tmdate":1747083573038,"pdate":1704067200000,"tcdate":1746720436466,"writers":["~"],"signatures":["~Eric_C_Cyr1"],"forum":"UUr0HfLXTH","license":"CC BY-SA 4.0","number":413605,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747083573038,"domain":"DBLP.org","id":"UUr0HfLXTH","version":2},{"content":{"venue":{"value":"NICE 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10548018/10548050/10548180.pdf"},"venueid":{"value":"dblp.org/conf/NICE/2024"},"paperhash":{"value":"theilman|spiking_physicsinformed_neural_networks_on_loihi_2"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Bradley_H._Theilman:","https://dblp.org/search/pid/api?q=author:Qian_Zhang:","https://dblp.org/search/pid/api?q=author:Adar_Kahana:","~Eric_C._Cyr1","https://dblp.org/search/pid/api?q=author:Nathaniel_Trask:","https://dblp.org/search/pid/api?q=author:James_B._Aimone:","https://dblp.org/search/pid/api?q=author:George_Em_Karniadakis:"]},"html":{"value":"https://doi.org/10.1109/NICE61972.2024.10548180"},"_bibtex":{"value":"@inproceedings{DBLP:conf/nice/TheilmanZKCTAK24,\n  author={Bradley H. Theilman and Qian Zhang and Adar Kahana and Eric C. Cyr and Nathaniel Trask and James B. Aimone and George Em Karniadakis},\n  title={Spiking Physics-Informed Neural Networks on Loihi 2},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-6},\n  url={https://doi.org/10.1109/NICE61972.2024.10548180},\n  booktitle={NICE},\n  crossref={conf/nice/2024}\n}\n"},"abstract":{"value":"Neuromorphic computing platforms hold the promise to dramatically reduce power requirements for calculations that are computationally intensive. One such application space is scientific machine learning (SciML). Techniques in this space use neural networks to approximate solutions of scientific problems. For instance, the popular physics-informed neural network (PINN) approximates the solution to a partial differential equation by using a trained feed-forward neural network, and injecting the knowledge of the physics through the loss function. Recent efforts have demonstrated how to convert a trained PINN to a spiking network architecture. In this work, we discuss our approach to quantization and implementation required to migrate these spiking PINNs to Intel’s Loihi 2 neuromorphic hardware. We explore the effect of quantization on the model accuracy, as well as the energy and throughput characteristics of the implementation. It is our intent that this serve as a starting point for additional SciML implementations on neuromorphic hardware."},"title":{"value":"Spiking Physics-Informed Neural Networks on Loihi 2"},"authors":{"value":["Bradley H. Theilman","Qian Zhang","Adar Kahana","Eric C. Cyr","Nathaniel Trask","James B. Aimone","George Em Karniadakis"]}},"tmdate":1746720440376,"pdate":1704067200000,"tcdate":1746720426582,"writers":["~"],"signatures":["~Eric_C_Cyr1"],"forum":"qhHw2fKfLQ","license":"CC BY-SA 4.0","number":413592,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1746720440376,"domain":"DBLP.org","id":"qhHw2fKfLQ","version":2},{"content":{"venue":{"value":"biorxiv"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"zhang|physdock_a_physicsguided_allatom_diffusion_model_for_proteinligand_complex_prediction"},"authorids":{"value":["~Kexin_Zhang7","~Yuanyuan_Ma3","~Jiale_Yu7","luoht2024@shanghaitech.edu.cn","~Jinyu_Lin1","~Yifan_Qin1","lixch2023@shanghaitech.edu.cn","~Qian_Jiang4","baifang@shanghaitech.edu.cn","~Jiayi_Dou1","~Jie_Zheng4","~Jingyi_Yu5","~Liping_Sun1"]},"html":{"value":"https://doi.org/10.1101/2025.04.28.650887"},"abstract":{"value":"Accurate prediction of protein-ligand complexes remains a central challenge in structural biology. Traditional methods are computationally inefficient and prone to local minima, whereas deep learning approaches struggle to capture structural flexibility and physical plausibility. We introduce PhysDock, a physics-guided diffusion model that uniquely integrates (i) all-atom diffusion to model ligand flexibility and protein precision-flexibility (i.e., subtle conformational adjustments); (ii) physical priors as diffusion conditioning, alongside two-phase physics guidance during the denoising diffusion to ensure physical plausibility. PhysDock demonstrates state-of-the-art performance in redocking benchmarks and excels in the more challenging cross-docking assessments. For practical utility, PhysDock (i) resolves cannabinoid receptor selectivity across diverse molecules, achieving accuracy comparable to experiments; (ii) distinguishes most drug candidates from weak binders in virtual screening of NTRK3 kinase, while uncovering novel candidates with structural insights. PhysDock serves as a versatile tool for protein-ligand complex prediction, with substantial potential to accelerate structure-based drug discovery."},"title":{"value":"PhysDock: A Physics-Guided All-Atom Diffusion Model for Protein-Ligand Complex Prediction"},"authors":{"value":["Kexin Zhang","Yuanyuan Ma","Jiale Yu","Huiting Luo","Jinyu Lin","Yifan Qin","Xiangcheng Li","Qian Jiang","Fang Bai","Jiayi Dou","Jie Zheng","Jingyi Yu","Liping Sun"]}},"tmdate":1778051523195,"pdate":1749916800000,"tcdate":1778051523195,"writers":["~Kexin_Zhang7","~Yuanyuan_Ma3","~Jiale_Yu7","luoht2024@shanghaitech.edu.cn","~Jinyu_Lin1","~Yifan_Qin1","lixch2023@shanghaitech.edu.cn","~Qian_Jiang4","baifang@shanghaitech.edu.cn","~Jiayi_Dou1","~Jie_Zheng4","~Jingyi_Yu5","~Liping_Sun1"],"signatures":["~Liping_Sun1"],"forum":"egt4o7T5mK","license":"arXiv.org perpetual, non-exclusive license","number":48872,"cdate":1778051523195,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1778051523195,"domain":"OpenReview.net/Archive","id":"egt4o7T5mK","version":2},{"content":{"summary":{"value":"This paper proposes a method to learn the implicit constitutive laws of deformable objects via monocular video observation. The key innovation is how the authors introduce regularizations through depth-based loss with the help of a monocular depth network and through a library of explicit physics rules."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Why is the color-space supervision not used? Does it hurt the result? \n\nWhat is the rationale behind the depth loss based on the rank correlation instead of metric-based losses?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"I like that this approach learns a neural implicit dynamics model with the help of existing explicit physics laws as guidance. The final learning outcome can be correlated back to the explicit models while not constrained by them. The way the explicit physics laws are used is in a spirit similar to expectation maximization. \n\nThe experimental result shows superior performance of the proposed model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I'm not sure how realistic the problem and experiment setup is. The deformation of the objects is very significant and uncommon in the real world. The dropping motion is doable but also limiting and not very common in real-world experience. I wonder what the practical application of this task are, given that it takes more than 1 hour to reason about one video of one object. \n\nThere is a lack of ablation study of the methodology. Only comparisons shown are without either of the two loss functions. This is a very coarse comparison. I think more detailed experiments are helpful especially given that there are very specific designs of the loss functions (e.g., global vs anchor-level supervision, rank-based loss formulation, procedure of the Multi-Hypothesis Physics Verifier, etc)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923187390,"tcdate":1761867601081,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12246/Reviewer_uUW6"],"signatures":["ICLR.cc/2026/Conference/Submission12246/Reviewer_uUW6"],"forum":"XyHbp7Y2T5","number":1,"license":"CC BY 4.0","cdate":1761867601081,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12246/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923187390,"domain":"ICLR.cc/2026/Conference","replyto":"XyHbp7Y2T5","id":"25vB9XtR9w","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Gaussian splatting","physics-informed learning","implicit constitutive laws"]},"supplementary_material":{"value":"/attachment/831d2a297e1a2ec1a3e64d35f8a5d86acb52a56b.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"We present **PhyCo**, a framework for learning implicit constitutive laws from \\textbf{monocular dynamic observations} of Gaussian splatting. Existing implicit methods often suffer from local minima under noisy supervision and lack physical interpretability, while explicit approaches rely on predefined constitutive equations, limiting generalizability. To address these issues, our framework, **PhyCo**, introduces two key innovations. First, **initializing from a static multi-view scan, we propose *Edge-Aware Depth Consensus Anchors* to establish robust geometric constraints from subsequent monocular dynamic observations**, circumventing unreliable pixel-level supervision. Second, a *Multi-Hypothesis Physics Verifier* integrates classical constitutive models as differentiable hypotheses, providing strong physical priors to regularize the optimization while preserving the flexibility of implicit modeling. This unified approach ensures physical plausibility without sacrificing generality. Extensive experiments on synthetic, real-to-sim, and real-world datasets demonstrate that **PhyCo** significantly outperforms existing methods, achieving state-of-the-art performance in learning accurate and generalizable physical dynamics from monocular videos."},"_bibtex":{"value":"@misc{\nliu2026phyco,\ntitle={PhyCo: Physics-Consistent Learning of Implicit Constitutive Laws via Monocular Observations of 3D Gaussians},\nauthor={Xiaoyang Liu and Kai Han},\nyear={2026},\nurl={https://openreview.net/forum?id=XyHbp7Y2T5}\n}"},"title":{"value":"PhyCo: Physics-Consistent Learning of Implicit Constitutive Laws via Monocular Observations of 3D Gaussians"},"pdf":{"value":"/pdf/230c2a74b4fdefee465eb282bf5252b5dd6e30e3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"liu|phyco_physicsconsistent_learning_of_implicit_constitutive_laws_via_monocular_observations_of_3d_gaussians"},"authorids":{"value":["~Xiaoyang_Liu5","~Kai_Han1"]},"authors":{"value":["Xiaoyang Liu","Kai Han"]}},"version":2},{"content":{"correctness":{"value":"Yes"},"summary_and_contributions":{"value":"The submission introduces M2Lingual, a fully synthetic multilingual, multi-turn instruction fine-tuning (IFT) dataset designed to improve large language models' performance across a diverse set of languages and tasks. \n\nKey contributions include:\n1. Creation of a synthetic dataset with 182K instruction-response pairs from 70 languages and 19 NLP tasks.\n2. Demonstration of the dataset's effectiveness in fine-tuning LLMs to achieve better performance on benchmarks like MT-Bench, XQuAD, MGSM, TyDiQA, MLQA, XNLI, and XLSUM.\n3. Highlighting the utility of the dataset for smaller LLMs (1.8B parameters) showing performance gains.\n4. Detailed analysis of the synthesis steps and their importance in the dataset's effectiveness."},"confidence":{"value":2},"documentation":{"value":"Yes"},"rating":{"value":7},"title":{"value":"A new fully synthetic multilingual and multi-turn instruction fine-tuning dataset for large language models"},"ethics":{"value":"No"},"clarity":{"value":"The paper is well-written, with clear explanations of the methodology, experiments, and results. Figures and tables are effectively used to illustrate key points."},"review":{"value":"The paper is well-structured and provides a comprehensive overview of the dataset creation process and its benefits. The methodology is sound, and the experiments are thorough, demonstrating improvements over existing datasets.\n\nPros:\n1. New method for generating synthetic multilingual, multi-turn data.\n2. Extensive evaluation on multiple benchmarks showing consistent improvements.\n3. Detailed ablation studies providing insights into the dataset's effectiveness.\n4. Contributions to improving performance in low-resource languages.\n\nCons:\n1. The paper discusses improvements in low-resource languages but does not delve deeply into specific challenges and solutions for these languages beyond dataset size.\n2. While the methodology is outlined, the paper does not provide enough detailed implementation steps to allow other researchers to easily reproduce the results.\n3. The comparative analysis with existing datasets, while extensive, could be more thorough by including additional qualitative analyses of the differences in dataset characteristics.\n4. The comparative analysis with existing datasets, while extensive, could be more thorough by including additional qualitative analyses of the differences in dataset characteristics."},"strengths":{"value":"1. The paper presents a ne approach to generating a synthetic multilingual, multi-turn instruction fine-tuning dataset.\n\n2. Demonstrates the dataset's effectiveness across multiple benchmarks, consistently outperforming existing datasets.\n\n3. Ensures equal representation of 70 languages, significantly benefiting low-resource languages."},"flag_for_ethics_review":{"value":"2: No, there are no or only very minor ethics concerns"},"relation_to_prior_work":{"value":"The paper discusses how it builds upon and differs from previous contributions. It highlights the limitations of existing multilingual IFT datasets and how M2Lingual addresses these issues."},"opportunities_for_improvement":{"value":"See review"},"additional_feedback":{"value":"NA"},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1731500806662,"tcdate":1721423471166,"writers":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission2194/Reviewer_JJxi"],"signatures":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission2194/Reviewer_JJxi"],"forum":"V6891G9dWu","number":1,"license":"CC BY 4.0","cdate":1721423471166,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission2194/-/Official_Review","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1731500806662,"domain":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","replyto":"V6891G9dWu","id":"OmbNsGjxav","forumContent":{"TLDR":{"value":"A multi-lingual, multi-turn, evolved instruction finetuning dataset that leads to state-of-the-art results on several multilingual evaluation benchmarks and multi-turn eval benchmarks."},"venue":{"value":"Submitted to NeurIPS 2024 Track Datasets and Benchmarks"},"keywords":{"value":["Multilingual LLM alignment","Multilingual IFT dataset","Multilingual multi-turn IFT dataset"]},"supplementary_material":{"value":"/attachment/b6271c6995e517741602a85cf55bd6b39ed553d9.zip"},"abstract":{"value":"Instruction finetuning (IFT) is critical for aligning Large Language Models (LLMs) to follow instructions. Numerous effective IFT datasets have been proposed in the recent past, but most focus on rich resourced languages such as English. In this work, we propose a diverse, task taxonomy guided, fully synthetic Multilingual, Multi-turn evoled instruction finetuning dataset, called M2Lingual, to better align LLMs on a diverse set of languages and tasks. M2Lingual contains a total of 182K IFT pairs that are built upon diverse seeds collected from Aya collection and Aya dataset covering 70 languages, 19 NLP tasks and general instruction-response pairs. LLMs finetuned with M2Lingual substantially outperform the majority of existing multilingual IFT datasets. Importantly, LLMs trained with M2Lingual consistently competitive results across wide variety of evaluation benchmarks compared to existing multilingual IFT datasets that enable LLMs performance in only one or a few subset of the benchmarks. Specifically, LLMs finetuned with M2Lingual achieve strong performance on multi-turn evaluation benchmarks such as MT-bench and across wide-variety of multilingual tasks such as XQuAD, MGSM, TyDiQA, MLQA, XNLI and XLSUM. We show efficacy of M2Lingual across LLMs with different sizes, especially smaller LLMs with 1.8B size which benefit massively from our dataset. Lastly, we present key analyses to highlight importance of each synthesis step of M2Lingual."},"_bibtex":{"value":"@misc{\nanonymous2024mlingual,\ntitle={M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models},\nauthor={Anonymous},\nyear={2024},\nurl={https://openreview.net/forum?id=V6891G9dWu}\n}"},"title":{"value":"M2Lingual: Enhancing Multilingual, Multi-Turn Instruction Alignment in Large Language Models"},"pdf":{"value":"/pdf/899235d13a3aa79ae00aff618da99c77290eaf22.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Rejected_Submission"},"paperhash":{"value":"maheshwary|m2lingual_enhancing_multilingual_multiturn_instruction_alignment_in_large_language_models"},"authorids":{"value":["~Rishabh_Maheshwary2","~Vikas_Yadav2","~Hoang_H_Nguyen1","~Khyati_Mahajan1","~Sathwik_Tejaswi_Madhusudhan2"]},"authors":{"value":["Rishabh Maheshwary","Vikas Yadav","Hoang H Nguyen","Khyati Mahajan","Sathwik Tejaswi Madhusudhan"]}},"version":2},{"content":{"summary":{"value":"Although T2V models have shown great progress in generating good media-level content, this paper challenges their capability to become the real world simulator. This paper first proposes a PhyGenBench, 160 T2V prompts composed of several physics categories, then proposes a hierarchical framework to evaluate semantic alignment and physics commonsense alignment. It shows that current models, even ones with large scales, struggle with physical commonsense."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"(1) What is the efficiency of using your auto-evaluator? Could you provide an estimation?\n\n(2) Could you provide some error analysis on the bad cases where PhyGenEval is opposite to the human eval? Maybe this can provide some insight on how to further improve the reward modeling of world simulators."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"(1) This paper handcrafts T2V prompts in a fine-grained way.\n(2) This paper provides a carefully designed pipeline to conduct the evaluation of physics commonsense."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) The paper lacks decent novelty in terms of benchmark and evaluator itself. The way it constructs the evaluator heavily relies on several generative models. For example, using GPT4o to do information extraction and create questions sometimes brings about hallucination. Also, using VLMs in different stages can also lead to hallucination. Since it is a complex pipeline composed of different stages, error propagation might happen.\n\n(2) The comparison with other baselines is unfair. The comparisons with other baselines are biased. Although they acknowledge that alternative auto-evaluators lack robustness, they do not demonstrate whether their own auto-evaluator performs effectively on prompts from different benchmarks as part of a generalization analysis. Typically, like concurrent work, these kinds of auto-evaluators are tailored to specific prompt distributions. Basically, the generalization of the reward modeling for world simulators should be enough for another paper.\n\n(3) The number of prompts engaged in this paper are limited, which might be a weak signal for evaluating the video generation models as a world simulator."}},"nonreaders":[],"tmdate":1732690100887,"tcdate":1730681800040,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2363/Reviewer_wXRw"],"signatures":["ICLR.cc/2025/Conference/Submission2363/Reviewer_wXRw"],"forum":"6rMHcLWxl4","number":5,"license":"CC BY 4.0","cdate":1730681800040,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2363/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732690100887,"domain":"ICLR.cc/2025/Conference","replyto":"6rMHcLWxl4","id":"d0XBO5Em8S","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["World Simulator","Physical Commonsense","Video Generation","Evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-to-video (T2V) models like Sora have made significant strides in visualizing complex prompts, which is increasingly viewed as a promising path towards constructing the universal world simulator. Cognitive psychologists believe that the foundation for achieving this goal is the ability to understand intuitive physics. However, the capacity of these models to accurately represent intuitive physics remains largely unexplored. To bridge this gap, we introduce PhyGenBench, a comprehensive \\textbf{Phy}sics \\textbf{Gen}eration \\textbf{Ben}chmark designed to evaluate physical commonsense correctness in T2V generation. PhyGenBench comprises 160 carefully crafted prompts across 27 distinct physical laws, spanning four fundamental domains, which could comprehensively assesses models' understanding of physical commonsense. Alongside PhyGenBench, we propose a novel evaluation framework called PhyGenEval. This framework employs a hierarchical evaluation structure utilizing appropriate advanced vision-language models and large language models to assess physical commonsense. Through PhyGenBench and PhyGenEval, we can conduct large-scale automated assessments of T2V models' understanding of physical commonsense, which align closely with human feedback. Our evaluation results and in-depth analysis demonstrate that current models struggle to generate videos that comply with physical commonsense. Moreover, simply scaling up models or employing prompt engineering techniques is insufficient to fully address the challenges presented by PhyGenBench (e.g., dynamic scenarios). We hope this study will inspire the community to prioritize the learning of physical commonsense in these models beyond entertainment applications. We will release the data and codes at https://github.com/PhyGenBench/PhyGenBench"},"_bibtex":{"value":"@misc{\nmeng2025towards,\ntitle={Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation},\nauthor={Fanqing Meng and Jiaqi Liao and Xinyu Tan and Wenqi Shao and Quanfeng Lu and Kaipeng Zhang and Yu Cheng and Dianqi Li and Yu Qiao and Ping Luo},\nyear={2025},\nurl={https://openreview.net/forum?id=6rMHcLWxl4}\n}"},"title":{"value":"Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation"},"pdf":{"value":"/pdf/1814f0c3473ab9a04aca4edcd8aab3e678055bdd.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"meng|towards_world_simulator_crafting_physical_commonsensebased_benchmark_for_video_generation"},"authorids":{"value":["~Fanqing_Meng1","~Jiaqi_Liao2","~Xinyu_Tan1","~Wenqi_Shao2","~Quanfeng_Lu1","~Kaipeng_Zhang1","~Yu_Cheng1","~Dianqi_Li1","~Yu_Qiao1","~Ping_Luo2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fanqing Meng","Jiaqi Liao","Xinyu Tan","Wenqi Shao","Quanfeng Lu","Kaipeng Zhang","Yu Cheng","Dianqi Li","Yu Qiao","Ping Luo"]}},"version":2},{"content":{"review":{"value":"This work presents a simulation of porous media using physics-informed ML models. The idea is not new and has already been done in the literature; see, for example,\n\n1. Faroughi, S. A., Soltanmohammadi, R., Datta, P., Mahjour, S. K., & Faroughi, S. (2023). Physics-informed neural networks with periodic activation functions for solute transport in heterogeneous porous media. Mathematics, 12(1), 63.\n2. Lehmann, François, Marwan Fahs, Ali Alhubail, and Hussein Hoteit. \"A mixed pressure-velocity formulation to model flow in heterogeneous porous media with physics-informed neural networks.\" Advances in Water Resources 181 (2023): 104564.\n3. Faroughi, Salah A., Ramin Soltanmohammadi, Pingki Datta, Seyed Kourosh Mahjour, and Shirko Faroughi. \"Physics-informed neural networks with periodic activation functions for solute transport in heterogeneous porous media.\" Mathematics 12, no. 1 (2023): 63."},"confidence":{"value":2},"rating":{"value":4},"title":{"value":"Physics-Informed Machine Learning for Fluid Flow Prediction in Porous Media"}},"nonreaders":[],"tmdate":1709481615017,"tcdate":1708465736817,"writers":["ICLR.cc/2024/Workshop/AI4DiffEqtnsInSci","ICLR.cc/2024/Workshop/AI4DiffEqtnsInSci/Submission82/Reviewer_EK2Y"],"signatures":["ICLR.cc/2024/Workshop/AI4DiffEqtnsInSci/Submission82/Reviewer_EK2Y"],"forum":"c3mS9jMXUX","number":2,"license":"CC BY 4.0","cdate":1708465736817,"readers":["everyone"],"invitations":["ICLR.cc/2024/Workshop/AI4DiffEqtnsInSci/Submission82/-/Official_Review","ICLR.cc/2024/Workshop/AI4DiffEqtnsInSci/-/Edit"],"mdate":1709481615017,"domain":"ICLR.cc/2024/Workshop/AI4DiffEqtnsInSci","replyto":"c3mS9jMXUX","id":"7WZXOCkiVx","forumContent":{"TLDR":{"value":"This work presents a physics-informed machine learning model for predicting grid-level flow fields in porous media, ensuring physical consistency and accuracy across diverse media variations."},"venue":{"value":"AI4DiffEqtnsInSci @ ICLR 2024 Poster"},"pdf":{"value":"/pdf/ff5a9923c14a87f2ee873e1db42b624403052869.pdf"},"keywords":{"value":["Physics-Informed Machine Learning","Fluid Flow","Porous Media"]},"venueid":{"value":"ICLR.cc/2024/Workshop/AI4DiffEqtnsInSci"},"paperhash":{"value":"takbiriborujeni|physicsinformed_machine_learning_for_fluid_flow_prediction_in_porous_media"},"authorids":{"value":["alitakb@amazon.com","~Mohammad_Kazemi2","sam.takbiri@gmail.com"]},"abstract":{"value":"The objective of this work is to predict grid-level flow fields in porous media as a priori to determining the permeability of porous media.\n\t\tA physics-informed ML model is developed by using the results from numerical fluid flow simulations of randomly distributed circular grains to represent the porous media. The deep U-Net and ResNet neural network architectures are combined to train a deep learning model that avoids vanishing gradient issues. \n\t\tThe model integrates continuity and momentum conservation equations into the loss function to ensure physical consistency. Additionally, we modify the padding function in convolutional layers to use circular paddings, mimicking periodic boundary conditions in LB simulations. By learning inter-grid communications, the ML model achieves precise flow predictions for new simulation sets with high accuracy. The robustness of the developed model is then tested for numerous variations of porous media that have not been used for developing the model."},"_bibtex":{"value":"@inproceedings{\ntakbiri-borujeni2024physicsinformed,\ntitle={Physics-Informed Machine Learning for Fluid Flow Prediction in Porous Media},\nauthor={Ali Takbiri-Borujeni and Mohammad Kazemi and Sam Takbiri},\nbooktitle={ICLR 2024 Workshop on AI4DifferentialEquations In Science},\nyear={2024},\nurl={https://openreview.net/forum?id=c3mS9jMXUX}\n}"},"title":{"value":"Physics-Informed Machine Learning for Fluid Flow Prediction in Porous Media"},"authors":{"value":["Ali Takbiri-Borujeni","Mohammad Kazemi","Sam Takbiri"]}},"version":2},{"content":{"summary":{"value":"In this work, the authors address problem of phase retrieval from intensity-only measurements and explores outer-ring generalization, where a model trained on inner rings is evaluated on unseen outer rings. The authors propose a hybrid physics-informed network that integrates several components: (i) radial priors using spline-based and monotone boosting mechanisms, (ii) two PDE-based branches (Kerr-NLSE and TIE), and (iii) a radial projection with a radius-dependent α-fusion. The method is claimed to outperform baselines on synthetic optical datasets and to provide better generalization to unseen spatial frequencies."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"Given how difficult to read is the paper, here are some suggestions for improvement:\n\n1. Begin with a clear and intuitive definition of the phase retrieval problem, including its mathematical formulation and physical motivation. Provide images if possible.\n2. Greatly simplify and clarify the exposition. Replace jargon-heavy descriptions with explanations of what each component contributes conceptually.\n3. Provide qualitative results (images) alongside quantitative metrics.\n4. Clarify dataset generation and all experimental settings.\n5. Remove trivial statements or overly technical statements (e.g., Proposition 1) or make them meaningful by linking them to design insights.\n6. Revise the writing tone to be more informative and guide the reader through your reasoning. At the moment it seems “technical for the sake of sounding technical.”"},"rating":{"value":0},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"1. Improving generalization of physics-informed neural networks in optical inverse problems is interesting and relevant.\n2. The attempt to connect PDE-based modeling and learned priors is conceptually valuable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper is not written well enough to be accepted and requires a lot more work to be in a readable format.\n\n1. The paper assumes a high level of prior knowledge about the phase retrieval problem, but never explicitly defines it. For most readers at a general machine learning venue (such as ICLR), the problem setup (what is measured, what is reconstructed, and why it is ill-posed) needs to be clearly introduced. In its current form, the paper is not understandable. For example: (a) the writing is overly dense and often boils down to unexplained jargon (e.g., “monotone curvature reduces ring,” “Strang-split Kerr-NLSE pathway”) with undefined acronyms (the authors mention SLM, PINNs and much more but never define it!), which significantly hinders readability. (b)  Several sections read like lists of technical keywords rather than structured explanations, e.g. the “Design rules” paragraph is essentially a bullet list of concepts without narrative or connection.\n2. Proposition 1 is trivial and adds little value. If it is meant to motivate the model’s architecture, the reasoning should be expanded or omitted.\n3. The paper presents multiple architectural and mathematical components (radial priors, PDE branches, $\\alpha$-fusion), but their interrelations and necessity are not clearly motivated. It is unclear why these components are combined or how they interact theoretically.\n4. Despite being an imaging paper, no reconstructed images are shown, which severely limits interpretability. Visual examples are crucial to assess the claimed improvements in ring generalization, amplitude stability, or spatial structure.\n5. Although the abstract mentions that pseudo-code will be released, the current version lacks sufficient experimental detail to reproduce the results. Precise dataset specifications, training protocols, and implementation details (e.g., optimizer, learning rate, loss definition) should be included. The dataset description is vague: “synthetic Fraunhofer (centered FFT, energy-normalized)” is not sufficient to reproduce the setup. Details such as the aperture geometry, sampling conditions, number of rings, and noise model must be specified.\n6. Sections 6.4–6.7 are particularly difficult to parse; results are presented without context or interpretation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925476627,"tcdate":1761945406758,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15167/Reviewer_23eb"],"signatures":["ICLR.cc/2026/Conference/Submission15167/Reviewer_23eb"],"forum":"jS3EKPSaAR","number":5,"license":"CC BY 4.0","cdate":1761945406758,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15167/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925476627,"domain":"ICLR.cc/2026/Conference","replyto":"jS3EKPSaAR","id":"yf8v52XZLl","forumContent":{"TLDR":{"value":"a pde based physic informed neural network that performs well in generalization"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["phase retrieval;PDE networks; outer-ring extrapolation; inverse problems"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Phase retrieval from intensity-only measurements is severely ill-posed due to global-gauge and rotational symmetries. We consider outer-ring generalization: training with supervision from only a few inner rings and testing the model’s ability to reconstruct a broader set of unseen outer rings. We introduce a physics-informed hybrid network that combines (i) radial priors encoded by a smooth exponentiated spline and a \\emph{monotone} outer-radius booster, (ii) two differentiable PDE branches---a Strang-split Kerr--NLSE pathway for high-frequency synthesis and a TIE-based low-pass pathway for coarse structure---and (iii) a strict radial projection enforcing output symmetry, together with a radius-dependent $\\alpha$-fusion. Across the tested configurations, when trained only on a few rings (1-3), our model reconstructs more rings(4-9) than conventional methods, and achieves better stability in peak\npositions and amplitude calibration under out-of-distribution settings. This provides some inspiration for enhancing the generalization of physics-informed neural networks when applied to optical inverse problems. Ablations isolate the contribution of the alpha fusion, PDE coupling, and monotone\nboosting. We will release pseudo-code to facilitate reproducibility."},"_bibtex":{"value":"@misc{\nyao2026physicsinformed,\ntitle={{PHYSICS}-{INFORMED} {RADIAL} {PHASE} {RETRIEVAL} {NEURAL} {NETWORK} {WITH} {HYBRID} {DEEP} {PRIORS} {AND} {DUAL} {PDE}},\nauthor={ZIYONG YAO},\nyear={2026},\nurl={https://openreview.net/forum?id=jS3EKPSaAR}\n}"},"title":{"value":"PHYSICS-INFORMED RADIAL PHASE RETRIEVAL NEURAL NETWORK WITH HYBRID DEEP PRIORS AND DUAL PDE"},"pdf":{"value":"/pdf/d0ee395178c85a98450cfb5c748c70d7faeb7393.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yao|physicsinformed_radial_phase_retrieval_neural_network_with_hybrid_deep_priors_and_dual_pde"},"authorids":{"value":["~ZIYONG_YAO1"]},"authors":{"value":["ZIYONG YAO"]}},"version":2},{"content":{"summary":{"value":"The authors consider the problem of compute-efficiency for continued pretraining on synthetic data augmentations. Their approach is to (1) form a knowledge graph over entities contained in documents, (2) apply entity and edge centrality metrics to obtain an importance score for each edge, and (3) to sample synthetic data proportional to these edge importance scores. They compare to random sampling (EntiGraph) on the QuALITY dataset of articles and books, and provide a simple mathematical model justifying their weighted sampling scheme given oracle importance weights."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"Questions\n* Can the authors improve the articulation of their main claim, which is improving the compute-efficiency of domain-specific continued pretraining? They should justify why studying compute-efficiency in a data-constrained setting is useful. As a devil's advocate, is it really so many FLOPs to just train on 100x or 1000x the tokens of a small proprietary corpus?\n* Alternatively, can the authors run larger-scale experiments which demonstrate CosyCPT reaches a higher asymptotic loss than EntiGraph? This would support the claim that CosyCPT improves data efficiency. \n* Can the author run CosyCPT and baselines on another domain (e.g., code or another specialized domain) to support their empirical claims?\n\nI would be willing to increase my score if the improved experiments support the data-efficiency claim or the paper provides a more focused pitch for compute-efficiency.\n\nClarifications\n* EntiGraph constructs entity graphs for each document. It is unclear reading the exposition of your method whether the entity graphs are on a per-document or per-corpus level.\n* The 5.4 Instruction Following experiment is just a guardrail to demonstrate that the synthetic CPT does not harm instruction following capabilities? Do you do any replay on instruction tuning data to get this to work?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper is generally well-written and clear.\n- The problem setting (improving synthetic data generation for continued pretraining) is timely and of interest to the ICLR community.\n- The method is intuitive: it is analogous to data curation, focusing pretraining compute on high importance relationships in the text. The task-- and domain-agnostic nature of the graph centrality metrics is also appealing (this could be better articulated by the authors)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The authors confusingly refer to their approach as a data-efficiency method. The usual definition of data efficiency in the pretraining and synthetic data literature is in improving performance given a fixed number of unique tokens in a seed corpus [1]. The present work is really an improvement in compute-efficiency, that is, the slope of the \"val loss versus log(synthetic tokens)\" curve is better. The paper's claims should be rewritten to reflect this; e.g., data-efficiency gains are claimed in Line 3, 70, 109, etc.\n- Given the above, the motivation for pursuing compute-efficiency of domain-specific continued pretraining is not clearly presented. For example, the authors could argue that many less well-resourced companies and organizations will want to do synthetic CPT on their proprietary datasets.\n- The empirical gains do not appear very significant over random sampling: an inconsistent 1-4% accuracy improvement, on a single dataset (QuaLITY). I would appreciate if the authors could run out larger-scale synthetic CPT runs until one of the approaches begins to asymptote, and ideally add another domain for continued pretraining. This is a particularly useful experiment because if CosyCPT does end up reaching a higher asymptote than EntiGraph, then this paper can actually make a data-efficiency as opposed to compute-efficiency claim (which is more salient in the setting of small proprietary datasets).\n\n[1] Muennighoff et al., 2023. Scaling Data-Constrained Language Models. In NeurIPS."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942130123,"tcdate":1761942847798,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22241/Reviewer_aeQX"],"signatures":["ICLR.cc/2026/Conference/Submission22241/Reviewer_aeQX"],"forum":"G3dW21Geb6","number":3,"license":"CC BY 4.0","cdate":1761942847798,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22241/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942130123,"domain":"ICLR.cc/2026/Conference","replyto":"G3dW21Geb6","id":"ynCqgo47qt","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["synthetic continued pretraining","knowledge acquisition","large language models","graph mining","sampling","data augmentation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Synthetic continued pretraining adapts LLMs to specific domains by fine-tuning them on synthetic data that augments real domain data. However, existing methods are often data-inefficient (requiring massive synthetic corpora to enumerate all relational facts) and fail to account for the relative importance of different entity relationships. In this paper, we propose coreness-aware synthetic continued pretraining (CosyCPT), a systematic pipeline that addresses both limitations. Our method (1) constructs a graph representation of entity relations in a document, (2) quantifies relation importance via coreness scores derived from the graph, and (3) leverages these scores to guide synthetic data sampling and augmentation for continued pretraining. We investigate four definitions of entity coreness and four formulations of relation coreness, verifying that multiple variants of coreness-aware sampling can outperform random sampling of augmented data for synthetic continued pretraining. We offer a mathematical analysis, proving that (1) given a learning budget, maximizing the expected accuracy on a query set about relational knowledge in a document collection is an NP-complete problem, (2) coreness-aware sampling is the optimal solution when each query examines one entity pair, and (3) coreness-aware sampling has a better upper bound for expected accuray than random sampling."},"_bibtex":{"value":"@misc{\nzhao2026cosycpt,\ntitle={Cosy{CPT}: Coreness-Aware Synthetic Continued Pretraining},\nauthor={Wenlong Zhao and Shuang Yang and Abhishek Bhandwaldar and Eshwar Prasad Sivaramakrishnan and Seungwook Han and Shivchander Sudalairaj and Aldo Pareja and Krishnateja Killamsetty and Elvira Rui Xiong and Hao Wang and Kai Xu and Akash Srivastava and Andrew McCallum},\nyear={2026},\nurl={https://openreview.net/forum?id=G3dW21Geb6}\n}"},"title":{"value":"CosyCPT: Coreness-Aware Synthetic Continued Pretraining"},"pdf":{"value":"/pdf/aa9f13c967adc91ebf5c3b7ebe4bc67ab8e87d54.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhao|cosycpt_corenessaware_synthetic_continued_pretraining"},"authorids":{"value":["~Wenlong_Zhao1","~Shuang_Yang14","~Abhishek_Bhandwaldar1","~Eshwar_Prasad_Sivaramakrishnan1","~Seungwook_Han1","~Shivchander_Sudalairaj1","~Aldo_Pareja1","~Krishnateja_Killamsetty1","~Elvira_Rui_Xiong2","~Hao_Wang22","~Kai_Xu4","~Akash_Srivastava1","~Andrew_McCallum1"]},"authors":{"value":["Wenlong Zhao","Shuang Yang","Abhishek Bhandwaldar","Eshwar Prasad Sivaramakrishnan","Seungwook Han","Shivchander Sudalairaj","Aldo Pareja","Krishnateja Killamsetty","Elvira Rui Xiong","Hao Wang","Kai Xu","Akash Srivastava","Andrew McCallum"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the limitations of automatic differentiation (AD) in PINNs for non-analytic PDEs. The authors propose a hybrid approach that combines numerical solvers with deep learning models to replace AD for gradient calculations. The proposed method enables the exact imposition of Dirichlet boundary conditions. The proposed approach is flexible and can be incorporated into any physics-informed model. Gradient computation is up to two orders of magnitude faster than automatic differentiation."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- How does the proposed method perform on 3D PDEs or larger-scale problems?\n- Can the method handle more complex boundary types (e.g., Neumann, Robin) just as effectively?\n- Handling of Nonlinear Operators: How well does the hybrid method generalize to nonlinear PDEs?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- **Addresses AD Limitations**. The proposed method addresses cases where AD fails, e.g., when the PDE coefficients do not have an analytic form, or when enriched input data is fed to the deep learning model. \n- **BC Imposition**. The proposed method allows strong constraints of Dirichlet BCs.\n- **Scalability and Efficiency**. The proposed approach shows significant speedups in gradient computation compared to AD-based PINNs. The numerical cost is independent of the complexity of the trained model. \n- **Flexiblility**. The proposed approach is flexible and can be incorporated into any physics-informed model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Generalization**. The proposed work could handle 1D and 2D PDEs with Dirichlet BCs. Yet, whether it could generalize to higher dimensions or more complex BCs remains unknown. \n- **Dependency on External Numerical Solvers**. The reliance on external numerical solvers makes the model more complex. \n- **Insufficient PINN Baselines**: The experiments do not thoroughly compare with SOTA neural operators (e.g., FNO, GNO), which are considered an important baseline for PINN-based neural routines.\n- **Presentation**. The presentation of this paper could be further improved. There are also typos (e.g., Eq. (19) is not properly aligned)."}},"nonreaders":[],"tmdate":1731428919217,"tcdate":1730078329900,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9786/Reviewer_tLLa"],"signatures":["ICLR.cc/2025/Conference/Submission9786/Reviewer_tLLa"],"forum":"R5FzCFR5yU","number":1,"license":"CC BY 4.0","cdate":1730078329900,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9786/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428919217,"domain":"ICLR.cc/2025/Conference","replyto":"R5FzCFR5yU","id":"2eQ8CO72tH","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Scientific Machine Learning","Physics-Informed Neural Networks","Automatic Differentiation","Partial Differential Equations"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This work demonstrates that automatic differentiation has strong limitations when employed to compute physical derivatives in a general physics-informed framework, therefore limiting the range of applications that these methods can address. A hybrid approach is proposed, combining deep learning and traditional numerical solvers such as the finite element method, to address the shortcomings of automatic differentiation. This novel approach enables the exact imposition of Dirichlet boundary conditions in a seamless manner, and more complex, non analytical problems can be solved. Finally, enriched inputs can be used by the model to help convergence. The proposed approach is flexible and can be incorporated into any physics-informed model. Our hybrid gradient computation proposal is also up to two orders of magnitude faster than automatic differentiation, as its numerical cost is independent of the complexity of the trained model. Several numerical applications are provided to illustrate the discussion."},"_bibtex":{"value":"@misc{\nchenaud2025hybrid,\ntitle={Hybrid Numerical {PINN}s: On the effectiveness of numerical differentiation for non-analytic problems},\nauthor={Marien Chenaud and Alves Jos{\\'e} and Frederic Magoules},\nyear={2025},\nurl={https://openreview.net/forum?id=R5FzCFR5yU}\n}"},"title":{"value":"Hybrid Numerical PINNs: On the effectiveness of numerical differentiation for non-analytic problems"},"pdf":{"value":"/pdf/c15501b2dd799f11c280ffdb1cedc6c590c25c11.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"chenaud|hybrid_numerical_pinns_on_the_effectiveness_of_numerical_differentiation_for_nonanalytic_problems"},"authorids":{"value":["~Marien_Chenaud1","~Alves_José1","~Frederic_Magoules1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Marien Chenaud","Alves José","Frederic Magoules"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Integrated Forward–Inverse Network (IFIN), an encoder–decoder architecture that embeds a Forward System Operator (FSO) and a learnable Inverse System Operator (ISO) at every feature scale, conditioned by a shared, learnable spatially varying PSF field. The goal is to keep raw measurement information useful  throughout the network while enforcing forward–inverse consistency, which the authors argue is crucial under severe, large-kernel blur in lensless imaging. The method initializes with a PSF-aware inverse, then runs two coupled streams, measurement to and from images,  through multi-scale IFIB blocks. Experiments on DiffuserCam, an SV-lensless dataset, and a MultiWienerNet benchmark show improvements over classical, data-driven, and physics-guided baselines."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"NA"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"* The problem motivation is clear. Conventional CNN/ViT backbones under-utilize information under large-kernel blur; one-sided physics pipelines discard complementary cues. The paper articulates this trade-off well.  \n\n* Robust empirical evaluation is provided. Diverse baselines (classical, data-driven, physics-guided) and three benchmarks dataset are used to report the results."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The score of the work is too narrow. The title suggests that the work could be applied to image reconstruction, while the paper and the contribution only targets the lensless imaging. It is not clear what is the use of the proposed method on the general family of inverse problems. \n\n* The contribution is very limited. While the authors describe the integration of forward and inverse operators at every scale as novel, this idea parallels well-established unrolled or feature-space deconvolution frameworks. In these methods, forward and adjoint physics are embedded at every iteration or scale. The “forward–inverse coupling at each layer” is not new; it’s conceptually equivalent to \"Unrolled primal–dual or ADMM networks (Adler & Öktem 2018 and  Zhang et al. 2020)\", which apply both forward and adjoint operators at every iteration, and \"Deep Wiener and Multi-Wiener Deconvolution Networks (DWDN/MWDN) (Dong et al and  Li et al.)\", which already apply learned deconvolutions at multiple scales in feature space and can incorporate a physical forward model. \n\n* The use of a learnable PSF field is a straightforward extension of prior PSF-grid and blind deconvolution frameworks. Encoding the PSF and conditioning the network hierarchically is an implementation detail, not a conceptual advance, given that hypernetwork conditioning on optical parameters is widely explored in blind or self-calibrating deconvolution literature.\n\n* The method is described as physics-integrated, yet both the forward and inverse modules reduce to convolutional operations without explicit physics enforcement or constraints. Thus, the “integration” is architectural rather than physical or theoretical in nature.\n\n* The contribution claim on “integration” improving performance, but there is no ablation showing the involvement of each component in the final performance. \n\n* The performance comparison of couple of baselines is missing ( FISTA-Net, MoDL, and newer deep unfolding methods such as \"Robust unrolled network for lensless imaging with enhanced resistance to model mismatch and noise\")."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926364613,"tcdate":1761245112311,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16203/Reviewer_BU3i"],"signatures":["ICLR.cc/2026/Conference/Submission16203/Reviewer_BU3i"],"forum":"q1TpQ6guwX","number":1,"license":"CC BY 4.0","cdate":1761245112311,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16203/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926364613,"domain":"ICLR.cc/2026/Conference","replyto":"q1TpQ6guwX","id":"a1p6CMpFvT","forumContent":{"TLDR":{"value":"We propose IFIN, the network couples forward physics and learned inverse at every layer with learnable calibration-free PSF, and shows state-of-the-art lensless imaging results under spatially varying blur and noise."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Computational Imaging","Lensless Imaging","Physics-guided Learning","Inverse Problem"]},"supplementary_material":{"value":"/attachment/9d05fbb0955cc21e2c4d9132124a9b7ae3b31195.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Inverse modeling plays a central role across computational optical imaging problems, including microscopy, imaging through scattering media, and lensless cameras, where the forward model often manifests as a severe blur. Discrepancies between the model and the actual imaging process further aggravate the ill-posed nature of the inverse problem. Physics-enabled methods that integrate analytical forward models with data-driven networks have been explored, but most incorporate physics only in a one-sided manner—either operating purely in the measurement space or only after inversion—thereby discarding complementary cues and reducing robustness to calibration errors.\nHere, we propose the Integrated Forward–Inverse Network (IFIN), a physics-guided deep neural network that interleaves differentiable forward operators with learnable inverse modules at every stage of the hierarchy. This design preserves physical consistency while shaping richer feature representations by jointly leveraging information from both measurement and image domains. A physics-guided kernel adaptation further compensates for inaccurate or unavailable PSF calibration, dynamically refining the kernel for blind deconvolution under system constraints.\nIFIN is especially effective when measurements are severely blurred by large point-spread functions, where conventional CNN-based inversion is limited by local receptive fields and underutilizes the measurement signal. On challenging lensless imaging benchmarks—including our newly introduced dataset, IFIN achieves state-of-the-art reconstruction quality and improved robustness under noise and model mismatch."},"_bibtex":{"value":"@misc{\nbae2026integrated,\ntitle={Integrated Forward{\\textendash}Inverse Network for Physics-Guided Image Reconstruction},\nauthor={Donggeon Bae and Jaewoo Jung and Yong Guk Kang and Kyung Chul Lee and Taeyoung Kim and Joonsik Park and Sangjun Byun and Jongho Kim and Nakkyu Baek and Hyeonyong Lee and Kyunghoon Jung and Seung Ah Lee},\nyear={2026},\nurl={https://openreview.net/forum?id=q1TpQ6guwX}\n}"},"title":{"value":"Integrated Forward–Inverse Network for Physics-Guided Image Reconstruction"},"pdf":{"value":"/pdf/0606a689890b9911fc407f24347161e12f8d2e59.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"bae|integrated_forwardinverse_network_for_physicsguided_image_reconstruction"},"authorids":{"value":["~Donggeon_Bae1","~Jaewoo_Jung1","~Yong_Guk_Kang1","~Kyung_Chul_Lee1","~Taeyoung_Kim11","~Joonsik_Park1","~Sangjun_Byun1","~Jongho_Kim3","~Nakkyu_Baek1","~Hyeonyong_Lee1","~Kyunghoon_Jung2","~Seung_Ah_Lee1"]},"authors":{"value":["Donggeon Bae","Jaewoo Jung","Yong Guk Kang","Kyung Chul Lee","Taeyoung Kim","Joonsik Park","Sangjun Byun","Jongho Kim","Nakkyu Baek","Hyeonyong Lee","Kyunghoon Jung","Seung Ah Lee"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2412.05133v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"kag|learning_hidden_physics_and_system_parameters_with_deep_operator_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Vijay_Kag:","https://dblp.org/search/pid/api?q=author:Dibakar_Roy_Sarkar:","https://dblp.org/search/pid/api?q=author:Birupaksha_Pal:","~Somdatta_Goswami1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2412.05133"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2412-05133,\n  publtype={informal},\n  author={Vijay Kag and Dibakar Roy Sarkar and Birupaksha Pal and Somdatta Goswami},\n  title={Learning Hidden Physics and System Parameters with Deep Operator Networks},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2412.05133},\n  url={https://doi.org/10.48550/arXiv.2412.05133}\n}\n"},"abstract":{"value":"Big data is transforming scientific progress by enabling the discovery of novel models, enhancing existing frameworks, and facilitating precise uncertainty quantification, while advancements in scientific machine learning complement this by providing powerful tools to solve inverse problems to identify the complex systems where traditional methods falter due to sparse or noisy data. We introduce two innovative neural operator frameworks tailored for discovering hidden physics and identifying unknown system parameters from sparse measurements. The first framework integrates a popular neural operator, DeepONet, and a physics-informed neural network to capture the relationship between sparse data and the underlying physics, enabling the accurate discovery of a family of governing equations. The second framework focuses on system parameter identification, leveraging a DeepONet pre-trained on sparse sensor measurements to initialize a physics-constrained inverse model. Both frameworks excel in handling limited data and preserving physical consistency. Benchmarking on the Burgers' equation and reaction-diffusion system demonstrates state-of-the-art performance, achieving average $L_2$ errors of $\\mathcal{O}(10^{-2})$ for hidden physics discovery and absolute errors of $\\mathcal{O}(10^{-3})$ for parameter identification. These results underscore the frameworks' robustness, efficiency, and potential for solving complex scientific problems with minimal observational data."},"title":{"value":"Learning Hidden Physics and System Parameters with Deep Operator Networks"},"authors":{"value":["Vijay Kag","Dibakar Roy Sarkar","Birupaksha Pal","Somdatta Goswami"]}},"tmdate":1747358506680,"pdate":1704067200000,"tcdate":1747358476078,"writers":["~"],"signatures":["~Somdatta_Goswami1"],"forum":"2e90oYVpRM","license":"CC BY-SA 4.0","number":501493,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747358506680,"domain":"DBLP.org","id":"2e90oYVpRM","version":2},{"content":{"summary":{"value":"This paper introduces Ctrl&Shift, a novel diffusion framework for high-quality, geometry-aware object manipulation in images and videos. The core challenge it addresses is the fusion of precise controllability from geometry-based methods with the realism and generalization of diffusion-based approaches.\n\nThe paper's main contribution is to cleverly decompose the complex manipulation task into two sub-tasks: 1) object removal and 2) reference-guided inpainting under explicit camera pose control. The authors design an 8-dimensional relative camera pose descriptor, f, which is injected as a control signal into the diffusion model. This enables end-to-end geometric control during inference without requiring explicit 3D modeling.\n\nTo achieve this, the paper proposes a multi-task, multi-stage training strategy, supplemented by a sophisticated data construction pipeline to generate paired, geometry-supervised, real-world data. Experimental results on the authors' new GeoEditBench and the ObjectMover-A benchmark demonstrate state-of-the-art performance, outperforming existing methods in fidelity, identity preservation, and geometric control accuracy (e.g., pose MAPE and IoU)."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"As weakness."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Conceptual Innovation: The paper's primary strength lies in its core idea. It represents a conceptual shift: instead of relying on expensive or unstable 3D reconstruction (like NeRF or Mesh) at inference time, it injects precise geometric control (a relative pose vector) as a condition into the 2D diffusion process. This is a very clever decoupling that elegantly combines the advantages of both domains.\n\nSystematic Framework Design: The Ctrl&Shift architecture is designed with systematic and sound principles. The multi-task training (main task, removal, inpainting) is clearly motivated and helps the model disentangle the functions of different control signals (identity, location, pose), which is strongly validated by the ablation study (Table 3). The multi-stage training strategy (Stage 1 for priors, Stage 2 for real-world backgrounds) is logical. It first learns the core geometric knowledge on controllable synthetic data before generalizing to complex real-world scenes, effectively balancing geometric understanding and realism.\n\n3.Strong Empirical Results and Evaluation:The method significantly outperforms existing SOTA approaches in both quantitative (Tables 1, 2) and qualitative (Fig. 4) comparisons.The evaluation metrics are comprehensive. They include not only traditional fidelity (PSNR) and identity preservation (DINO, CLIP) metrics but also creatively introduce geometric control accuracy metrics (Pose MAPE, Obj IoU), which are crucial for evaluating \"controllable\" generation tasks.The construction of the GeoEditBench benchmark is also a valuable contribution to the community, providing a standardized evaluation platform for this specific challenge."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I agree with the authors that this is excellent and inspiring work. To make the paper more complete and rigorous, I strongly recommend the authors add a 'Limitations and Future Work' discussion to the final version (e.g., in the conclusion or appendix).\n\nI would like the authors to specifically address the following points:\n\nFrom Technical Controllability to Practical Usability:\nThe authors should discuss the challenge of mapping intuitive user interactions (e.g., 2D mouse drags, rotational gestures) to the proposed 8D relative pose descriptor f. A powerful technology is of limited value without an intuitive interface. \n\nBeyond Geometry: The Challenge of Physical Realism:\nThe authors should acknowledge that the current framework focuses primarily on geometric consistency, while the modeling of physical interactions (especially lighting, shadows, and reflections) is limited. Generating physically correct new shadows and reflections based on the new location's lighting conditions is a major challenge for this method (and all manipulation methods) and a valuable direction for future research.\n\nGeneralization Boundaries Induced by the Data Pipeline:\nThe authors should discuss potential failure cases arising from the dependency on the Image2Mesh pipeline. The method will likely struggle to generalize to: Non-rigid objects (e.g., cloth, hair, pets) Transparent or highly reflective objects (e.g., glass, metal)\nTopologically complex objects (e.g., trees, smoke) Scenes with complex occlusion (e.g., moving an object behind another object in the scene; the current 'remove + inpaint' framework may not correctly handle this new depth relationship).\n\nAnd I expect the open-source of the benchmark."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918557094,"tcdate":1761740702283,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6225/Reviewer_19Wi"],"signatures":["ICLR.cc/2026/Conference/Submission6225/Reviewer_19Wi"],"forum":"T6T0JUgGFQ","number":2,"license":"CC BY 4.0","cdate":1761740702283,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6225/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918557094,"domain":"ICLR.cc/2026/Conference","replyto":"T6T0JUgGFQ","id":"3duE5oAbOF","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion Model","Image Editing","Video Editing"]},"supplementary_material":{"value":"/attachment/074c59f5cd746fbd772659c69a433bb39245d5d0.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Object-level manipulation—relocating or reorienting objects in images or videos while preserving scene realism—is central to film post-production, AR, and creative editing. Yet existing methods struggle to jointly achieve three core goals: background preservation, geometric consistency under viewpoint shifts, and user-controllable transformations. Geometry-based approaches offer precise control but require explicit 3D reconstruction and generalize poorly; diffusion-based methods generalize better but lack fine-grained geometric control. We present **Ctrl&Shift**, an end-to-end diffusion framework to achieve geometry-consistent object manipulation without explicit 3D representations. Our key insight is to decompose manipulation into two stages—object removal and reference-guided inpainting under explicit camera pose control—and encode both within a unified diffusion process. To enable precise, disentangled control, we design a multi-task, multi-stage training strategy that separates background, identity, and pose signals across tasks. To improve generalization, we introduce a scalable real-world dataset construction pipeline that generates paired image and video samples with estimated relative camera poses. Extensive experiments demonstrate that **Ctrl&Shift** achieves state-of-the-art results in fidelity, viewpoint consistency, and controllability. *To our knowledge, this is the first framework to unify fine-grained geometric control and real-world generalization for object manipulation—without relying on any explicit 3D modeling.*"},"_bibtex":{"value":"@inproceedings{\nruan2026ctrlshift,\ntitle={{CTRL}\\&{SHIFT}: High-quality Geometry-Aware Object Manipulation in Visual Generation},\nauthor={Penghui Ruan and Bojia Zi and Xianbiao Qi and Youze Huang and Rong Xiao and Pichao WANG and Jiannong Cao and Yuhui Shi},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=T6T0JUgGFQ}\n}"},"title":{"value":"CTRL&SHIFT: High-quality Geometry-Aware Object Manipulation in Visual Generation"},"pdf":{"value":"/pdf/7d47655f48d07ad3093338605338af3abff547a2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ruan|ctrlshift_highquality_geometryaware_object_manipulation_in_visual_generation"},"authorids":{"value":["~Penghui_Ruan1","~Bojia_Zi2","~Xianbiao_Qi2","~Youze_Huang1","~Rong_Xiao3","~Pichao_WANG3","~Jiannong_Cao1","~Yuhui_Shi2"]},"authors":{"value":["Penghui Ruan","Bojia Zi","Xianbiao Qi","Youze Huang","Rong Xiao","Pichao WANG","Jiannong Cao","Yuhui Shi"]}},"version":2},{"content":{"summary":{"value":"They suggested a zero-shot quantization method to retain the detailed texture feature distribution and introduced the mixup knowledge distillation module to diversify synthetic samples for finetuning"},"soundness":{"value":"2 fair"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"1. please provide more detail on why the texture feature distribution calibration is important when generating synthetic data.\n\n2. PTQ methods such as AdaRound, AdaQuant, and Brecq can employ synthetic data to quantize models. how about adapting these post-training quantization schemes for your methods?\n\n3. I expect to see more evaluation on various models in order to convince the superiority of your work."},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"strengths":{"value":"They identified the new feature required when generating synthetic data for quantization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"They should compare their work with MixMix [1], Genie[2] and KW[3].\n\nThe authors only empirically showed their superiority. i.e. lt lacks explanations of intuitive or mathematic. The author should give more reasons.\n\nThe image they generated showed a little bit of poor quality to argue that it has captured the texture feature distributions. please see the synthetic images in [1], [2], [3].\n\n[1] Li, Yuhang, et al. \"Mixmix: All you need for data-free compression are feature and data mixing.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021.\n\n[2] Jeon, Yongkweon, Chungman Lee, and Ho-young Kim. \"Genie: Show Me the Data for Quantization.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.\n\n[3] Haroush, Matan, et al. \"The knowledge within: Methods for data-free model compression.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020."},"limitations":{"value":"It lacks a literature survey."}},"nonreaders":[],"tmdate":1702410891895,"tcdate":1688650398321,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission3309/Reviewer_zKQV"],"signatures":["NeurIPS.cc/2023/Conference/Submission3309/Reviewer_zKQV"],"forum":"r8LYNleLf9","number":5,"license":"CC BY 4.0","cdate":1688650398321,"mdate":1702410891895,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission3309/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"r8LYNleLf9","id":"P3YqJUgzqS","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Zero-shot quantization","Texture feature calibration","Post-training quantization","low bit width","Neural network compression"]},"_bibtex":{"value":"@inproceedings{\nchen2023texq,\ntitle={TexQ: Zero-shot Network Quantization with Texture Feature Distribution Calibration},\nauthor={Xinrui Chen and Yizhi Wang and Renao Yan and Yiqing Liu and Tian Guan and Yonghong He},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=r8LYNleLf9}\n}"},"title":{"value":"TexQ: Zero-shot Network Quantization with Texture Feature Distribution Calibration"},"paperhash":{"value":"chen|texq_zeroshot_network_quantization_with_texture_feature_distribution_calibration"},"TLDR":{"value":"We proposed a novel zero-shot quantization method TexQ, which takes advantage of texture feature energy distribution calibration and mixup knowledge distillation, achieving state-of-the-art performance especially in low bit width."},"abstract":{"value":"Quantization is an effective way to compress neural networks. By reducing the bit width of the parameters, the processing efficiency of neural network models at edge devices can be notably improved. Most conventional quantization methods utilize real datasets to optimize quantization parameters and fine-tune. Due to the inevitable privacy and security issues of real samples, the existing real-data-driven methods are no longer applicable. Thus, a natural method is to introduce synthetic samples for zero-shot quantization (ZSQ). However, the conventional synthetic samples fail to retain the detailed texture feature distributions, which severely limits the knowledge transfer and performance of the quantized model. In this paper, a novel ZSQ method, TexQ is proposed to address this issue. We first synthesize a calibration image and extract its calibration center for each class with a texture feature energy distribution calibration method. Then, the calibration centers are used to guide the generator to synthesize samples. Finally, we introduce the mixup knowledge distillation module to diversify synthetic samples for fine-tuning. Extensive experiments on CIFAR10/100 and ImageNet show that TexQ is observed to perform state-of-the-art in ultra-low bit width quantization. For example, when ResNet-18 is quantized to 3-bit, TexQ achieves a 12.18% top-1 accuracy increase on ImageNet compared to state-of-the-art methods. Code at https://github.com/dangsingrue/TexQ."},"pdf":{"value":"/pdf/036e1e59e632aecf014bd56a4e35796ee2be8cdd.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Xinrui_Chen1","~Yizhi_Wang3","~Renao_Yan1","~Yiqing_Liu1","~Tian_Guan1","~Yonghong_He1"]},"authors":{"value":["Xinrui Chen","Yizhi Wang","Renao Yan","Yiqing Liu","Tian Guan","Yonghong He"]}},"version":2},{"content":{"summary":{"value":"This paper introduces **NewtonBench**, a benchmark for evaluating LLMs' scientific law discovery capabilities that addresses fundamental limitations in existing approaches. Current benchmarks suffer from a methodological trilemma (forcing trade-offs between scientific relevance, scalability, and memorization resistance) and oversimplify discovery as static function fitting rather than interactive exploration. NewtonBench resolves these issues through two key innovations: (1) **metaphysical shifts** that systematically mutate canonical physical laws to generate 324 scientifically grounded yet memorization-resistant tasks across 12 physics domains, and (2) **interactive model discovery** requiring agents to actively probe virtual environments to uncover hidden laws embedded in complex systems with confounding variables. Evaluation of 11 state-of-the-art LLMs reveals that while frontier models like GPT-5 and Gemini-2.5-pro demonstrate emerging capability (65-73% accuracy), their performance degrades precipitously with increasing complexity, exhibits extreme sensitivity to observational noise, and paradoxically suffers when given code assistance due to premature exploitation over exploration. The benchmark demonstrates that robust, generalizable scientific discovery in complex, interactive environments remains the core unsolved challenge for AI-driven science."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the Weaknesses."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This paper exhibits **high originality** through its novel benchmark design principles, **strong quality** in experimental rigor and scope, **good clarity** in presentation and visualization, and **substantial significance** for both AI research and automated science. The work makes meaningful contributions across all four dimensions, with particularly notable strengths in originality (metaphysical shifts resolution of trilemma) and clarity (outstanding visual communication and narrative structure). The significance is enhanced by the timeliness of addressing LLM scientific reasoning at a critical juncture in model capability development."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### LLM-as-Judge for Primary Metric Introduces Circularity\n\n**Problem**: Symbolic Accuracy, the main evaluation metric, relies on LLM verification of equation equivalence. Using LLMs to judge LLM-generated discoveries creates potential systematic biases.\n\n**Actionable Fixes**:\n\n1. Provide detailed error analysis: On what types of equations does LLM-as-judge fail? Are these random or systematic?\n2. Report inter-rater reliability among human experts\n3. Include traditional symbolic equivalence checking (even if imperfect) as a secondary validation\n4. Test whether different judge models (GPT vs. Gemini vs. Claude) produce consistent verdicts\n\n### Framing Claims \"Scientific Law Discovery\" but Evaluates Only Physics Equations\n\n**Problem**: The paper is titled \"Benchmarking Generalizable **Scientific Law Discovery**\" and makes broad claims about \"AI-driven science,\" \"automated science,\" and \"genuine scientific intelligence\" throughout. However, the evaluation exclusively covers **physics equations**—12 domains, all closed-form algebraic expressions. This creates a fundamental misalignment between claims and evidence. Scientific discovery encompasses far more than physics equations: chemistry involves molecular structures and reaction mechanisms, biology includes qualitative evolutionary principles and statistical patterns, social sciences rarely use closed-form equations, and even within physics the benchmark excludes quantum mechanics, field theories, differential equations, and computational models. The paper's significance claims about measuring \"scientific intelligence\" and guiding development of agents \"capable of genuine scientific discovery\" substantially overreach what physics equation discovery can demonstrate.\n\n**Actionable Fixes**: \n\n(1) **Reframe title and claims** to match actual scope—change to \"Benchmarking Physics Equation Discovery\" or \"Mathematical Law Discovery\"; (2) **Add explicit scope limitations** in Abstract/Introduction acknowledging this represents one facet of scientific discovery; (3) **Discuss generalization boundaries**—which findings likely transfer to non-physics domains (e.g., noise sensitivity) versus which are physics-specific (e.g., dimensional consistency); (4) **Provide expansion roadmap** for chemistry (reaction mechanisms), biology (population dynamics), or qualitative theory discovery to validate claimed generalizability. Honest scoping doesn't diminish the contribution—it makes claims defensible and clarifies what the benchmark actually measures: rigorous evaluation of mathematical law discovery from experimental data, a valuable but bounded capability for AI-assisted science."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919396110,"tcdate":1761735501895,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7266/Reviewer_cQE7"],"signatures":["ICLR.cc/2026/Conference/Submission7266/Reviewer_cQE7"],"forum":"Gk6umqW74m","number":3,"license":"CC BY 4.0","cdate":1761735501895,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7266/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919396110,"domain":"ICLR.cc/2026/Conference","replyto":"Gk6umqW74m","id":"5Vh62YuMGE","forumContent":{"TLDR":{"value":"We introduce NewtonBench, a benchmark for evaluating LLMs’ scientific law discovery via interactive, generalizable experiments."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["large language models","benchmark","virtual environment","generalization","agent","scientific law discovery"]},"supplementary_material":{"value":"/attachment/644261b3f64d1a719a04f6b1e6ebeac1a56e1865.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large language models (LLMs) are emerging as powerful tools for scientific law discovery, a foundational challenge in AI-driven science.\nHowever, existing benchmarks for this task suffer from a fundamental methodological trilemma, forcing a trade-off between scientific relevance, scalability, and resistance to memorization. Furthermore, they oversimplify discovery as static function fitting, failing to capture the authentic scientific process of uncovering embedded laws through the interactive exploration of complex model systems. To address these critical gaps, we introduce **NewtonBench**, a benchmark comprising 324 scientific law discovery tasks across 12 physics domains. Our design mitigates the evaluation trilemma by using counterfactual law shifts - systematic alterations of canonical laws - to generate a vast suite of problems that are scalable, scientifically relevant, and memorization-resistant.\nMoreover, we elevate the evaluation from static function fitting to interactive model discovery, requiring agents to experimentally probe simulated complex systems to uncover hidden principles. Our extensive evaluation of 11 state-of-the-art LLMs reveals a clear but fragile capability for discovery in frontier models: this ability degrades precipitously with increasing system complexity and exhibits extreme sensitivity to observational noise. Notably, we uncover a paradoxical effect of tool assistance: providing a code interpreter can hinder more capable models by inducing a premature shift from exploration to exploitation, causing them to satisfice on suboptimal solutions. These results demonstrate that robust, generalizable discovery in complex, interactive environments remains the core challenge for the future of automated science. By providing a scalable, robust, and scientifically authentic testbed, NewtonBench offers a crucial tool for measuring true progress and guiding the development of next-generation AI agents capable of genuine scientific discovery."},"_bibtex":{"value":"@inproceedings{\nzheng2026newtonbench,\ntitle={NewtonBench: Benchmarking Generalizable Scientific Law Discovery in {LLM} Agents},\nauthor={Tianshi Zheng and Kelvin Kiu Wai Tam and Newt Nguyen Kim Hue Nam and Baixuan Xu and Zhaowei Wang and Cheng Jiayang and Hong Ting Tsang and Weiqi Wang and Jiaxin Bai and Tianqing Fang and Yangqiu Song and Ginny Wong and Simon See},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Gk6umqW74m}\n}"},"title":{"value":"NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents"},"pdf":{"value":"/pdf/6375dd7eb877e02f27026fa1429638e0d73abf11.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zheng|newtonbench_benchmarking_generalizable_scientific_law_discovery_in_llm_agents"},"authorids":{"value":["~Tianshi_Zheng1","~Kelvin_Kiu_Wai_Tam1","~Newt_Nguyen_Kim_Hue_Nam1","~Baixuan_Xu1","~Zhaowei_Wang2","~Cheng_Jiayang1","~Hong_Ting_Tsang1","~Weiqi_Wang1","~Jiaxin_Bai1","~Tianqing_Fang1","~Yangqiu_Song1","~Ginny_Wong1","~Simon_See1"]},"authors":{"value":["Tianshi Zheng","Kelvin Kiu Wai Tam","Newt Nguyen Kim Hue Nam","Baixuan Xu","Zhaowei Wang","Cheng Jiayang","Hong Ting Tsang","Weiqi Wang","Jiaxin Bai","Tianqing Fang","Yangqiu Song","Ginny Wong","Simon See"]}},"version":2},{"content":{"TLDR":{"value":"IMBench evaluates intuitive manipulation on challenging physical tasks, revealing that current foundation models struggle to translate physical reasoning into executable actions."},"venue":{"value":"RSS SemRob 2026 Poster"},"pdf":{"value":"/pdf/096b1f8a3c8ca89851cd3cc28b1e181949f67b8d.pdf"},"keywords":{"value":["Benchmarks and datasets for robot learning","Robot manipulation"]},"venueid":{"value":"roboticsfoundation.org/RSS/2026/Workshop/SemRob"},"paperhash":{"value":"maurya|imbench_a_benchmark_for_intuitive_robotic_manipulation"},"authorids":{"value":["~Anurag_Maurya1","~Sukhvansh_Jain1","~Prajwal_Avhad1","~Gautham_Balachandran1","~Ziyi_Zhou12","~Atharva_Kshirsagar2","~SATYAM_SINGH2","~Bowen_Li7","~Rishabh_Mukund1","~Ritul_Singh1","~Jatin_Vira1","~Suvonil_Chatterjee1","~Devesh_K._Jha1"]},"abstract":{"value":"Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to\nthis capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBench, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision-language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBench as a step toward evaluating and enabling more integrated, adaptive physical intelligence."},"_bibtex":{"value":"@inproceedings{\nmaurya2026imbench,\ntitle={{IMB}ench: A Benchmark for Intuitive Robotic Manipulation},\nauthor={Anurag Maurya and Sukhvansh Jain and Prajwal Avhad and Gautham Balachandran and Ziyi Zhou and Atharva Kshirsagar and SATYAM SINGH and Bowen Li and Rishabh Mukund and Ritul Singh and Jatin Vira and Suvonil Chatterjee and Devesh K. Jha},\nbooktitle={3rd Workshop on Semantic Reasoning and Goal Understanding in Robotics (RSS 2026)},\nyear={2026},\nurl={https://openreview.net/forum?id=emePaiT4dY}\n}"},"title":{"value":"IMBench: A Benchmark for Intuitive Robotic Manipulation"},"authors":{"value":["Anurag Maurya","Sukhvansh Jain","Prajwal Avhad","Gautham Balachandran","Ziyi Zhou","Atharva Kshirsagar","SATYAM SINGH","Bowen Li","Rishabh Mukund","Ritul Singh","Jatin Vira","Suvonil Chatterjee","Devesh K. Jha"]}},"tmdate":1784334355279,"pdate":1782782064459,"tcdate":1781766135216,"writers":["roboticsfoundation.org/RSS/2026/Workshop/SemRob","roboticsfoundation.org/RSS/2026/Workshop/SemRob/Submission27/Authors"],"signatures":["roboticsfoundation.org/RSS/2026/Workshop/SemRob/Submission27/Authors"],"forum":"emePaiT4dY","license":"CC BY 4.0","number":27,"cdate":1781766135216,"readers":["everyone"],"invitations":["roboticsfoundation.org/RSS/2026/Workshop/SemRob/-/Submission","roboticsfoundation.org/RSS/2026/Workshop/SemRob/-/Post_Submission","roboticsfoundation.org/RSS/2026/Workshop/SemRob/-/Edit","roboticsfoundation.org/RSS/2026/Workshop/SemRob/Submission27/-/Camera_Ready"],"mdate":1784334355279,"odate":1784334355162,"domain":"roboticsfoundation.org/RSS/2026/Workshop/SemRob","id":"emePaiT4dY","version":2},{"content":{"summary":{"value":"The paper presents WavNAF, designed to improve room acoustics modeling by integrating wave propagation physics directly into the acoustic synthesis process. It overcomes the limitations of geometric methods, which cannot model intricate wave phenomena like diffraction, refraction, and complex reflections, by utilizing the Finite-Difference Time-Domain (FDTD) method to numerically solve the wave equation.\nThe key contributions include generating physically-informed wave propagation priors by deriving acoustic parameters (sound speed and density) from visual scene geometry encoded by NeRF. Crucially, the framework introduces a Neural Acoustic Scaling Module that efficiently addresses the high computational cost of full FDTD simulations by learning adaptive, time-dependent transformations to accurately estimate full-scale RIRs from compressed simulations. This physics-based approach leads to substantial improvements in acoustic quality across standard metrics compared to existing state-of-the-art methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Given that the FDTD CUDA kernel accounts for only 5.1% of the total inference time, and data preparation (acoustic parameter grid generation) accounts for 71.4%, what specific optimizations are being considered for the data preparation pipeline to make the transition to full 3D FDTD simulations feasible? Besides, if computational resources were unlimited, what is the potential performance gain anticipated by moving from 2D to 3D simulations? In particular, how would the inclusion of vertical propagation specifically impact the accuracy of metrics like EDT and C50?\n- Have the authors experimented with different initialization strategies for the MLPs that are specifically designed to emphasize correction for late reverberation, such as initializing $\\nu$ with a slight non-zero bias, and how did this affect convergence stability?\n- The few-shot learning study shows that WavNAF achieves comparable performance to a baseline (NeRAF) trained on 75% of data when WavNAF is trained on only 50% of data. Does this data efficiency hold true when testing on the RAF real-world dataset, where the noise and complexity may be greater than in the synthetic SoundSpaces data?"},"rating":{"value":8},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- WavNAF integrates physically-informed wave propagation priors derived directly from FDTD-simulated pressure maps, providing a strong inductive bias for neural acoustic field learning, even with simplified material parameters.\n- The Neural Acoustic Scaling Module addresses the high computational cost and stability constraints of FDTD by learning adaptive, time-dependent transformations to estimate accurate full-scale RIRs from temporally compressed simulations.\n- The physics-informed wave propagation priors enable better generalization from limited data\n- Using NeRF to extract visual scene geometry, converting volumetric density fields into 2D acoustic parameter grids (sound speed and density) allows physics-informed simulations without explicit material annotations\n- The paper is well-written, and I really like the detailed background section."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- WavNAF shows sensitivity to the alignment between the visual and acoustic coordinate systems. The model currently requires empirical tuning to ensure the necessary precise coordinate alignment between the geometry derived from NeRF and the simulation space\n- The integration of on-the-fly FDTD simulations significantly increases the computational complexity compared to purely neural or geometric baselines\n- To model energy dissipation and numerical stability, the framework uses a combination of an empirically defined sponge layer near the boundaries and a simple global damping factor applied uniformly across the domain to model natural air attenuation. While effective, this is a simplified model of complex, frequency-dependent air absorption effects."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926286886,"tcdate":1761798756596,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16111/Reviewer_3MYx"],"signatures":["ICLR.cc/2026/Conference/Submission16111/Reviewer_3MYx"],"forum":"FKZtlVQcFo","number":2,"license":"CC BY 4.0","cdate":1761798756596,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16111/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926286886,"domain":"ICLR.cc/2026/Conference","replyto":"FKZtlVQcFo","id":"7sRRUZSk8g","forumContent":{"TLDR":{"value":"We propose WavNAF, a neural acoustic synthesis framework that leverages physically-informed wave propagation priors to explicitly capture complex acoustic interactions."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["neural acoustic fields","spatial audio","audio scenes","implicit representations","applications"]},"supplementary_material":{"value":"/attachment/f9dcf7b8ebde03e9c554b33ac8c44ec49ad68b89.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Room acoustics modeling requires capturing intricate wave phenomena such as reflections, refractions, and diffractions beyond direct sound propagation. Recent neural acoustic synthesis methods have improved acoustic realism but typically focus only on straight sound paths and coarse reverberation, missing detailed interactions like diffraction or multi-order reflections. We propose WavNAF, a neural framework that leverages physically-informed wave propagation priors to explicitly capture complex acoustic interactions. We generate these priors by numerically solving the wave equation with the Finite-Difference Time-Domain (FDTD) method, which directly simulates wave-based acoustic behavior that geometric methods cannot capture. Specifically, we extract essential acoustic parameters for FDTD, such as wave speed and density, from visual scene geometry encoded by Neural Radiance Fields (NeRF). We then generate physically-informed pressure maps and encode them via a feature extractor to learn wave propagation priors that capture intricate acoustic phenomena. To address the inherent computational cost issue of FDTD, we introduce a novel Neural Acoustic Scaling Module, inspired by traditional acoustic scale model. This module adaptively recalibrates encoded pressure map features from temporally compressed simulations to efficiently estimate accurate full-scale Room Impulse Responses. Experimental results demonstrate that WavNAF achieves substantial improvements in acoustic quality across various evaluation metrics compared to existing state-of-the-art methods."},"_bibtex":{"value":"@misc{\nkim2026wavnaf,\ntitle={Wav{NAF}: Learning Wave Propagation Priors for Neural Acoustic Fields},\nauthor={Taeho Kim and Hyunjun Kim and MinKyu Lee and SuBeen Lee and Inkyu An and Jae-Pil Heo},\nyear={2026},\nurl={https://openreview.net/forum?id=FKZtlVQcFo}\n}"},"title":{"value":"WavNAF: Learning Wave Propagation Priors for Neural Acoustic Fields"},"pdf":{"value":"/pdf/6b0d30442a9d96c80f04f237b20eab0f64b9a437.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|wavnaf_learning_wave_propagation_priors_for_neural_acoustic_fields"},"authorids":{"value":["~Taeho_Kim1","~Hyunjun_Kim1","~MinKyu_Lee1","~SuBeen_Lee1","~Inkyu_An1","~Jae-Pil_Heo3"]},"authors":{"value":["Taeho Kim","Hyunjun Kim","MinKyu Lee","SuBeen Lee","Inkyu An","Jae-Pil Heo"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a physics framework for learning B-spline control points in systems with varying initial and boundary conditions. The approach allows for using analytical expressions as physics informed losses since B-spline function gradients can be expressed as analytical expressions. The proposed approach is used for estimating high dimensional surfaces described by families of ODEs/PDEs. The authors provide theoretical bounds and empirical evidence for their claims."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"1.\tHow does the proposed approach compare against [1], [2], [3]?\n\n[1] Wandel, Nils, et al. \"Spline-pinn: Approaching pdes without data using fast, physics-informed hermite-spline cnns.\" Proceedings of the AAAI conference on artificial intelligence. Vol. 36. No. 8. 2022.\n\n[2] Doległo, Kamil, et al. \"Deep neural networks for smooth approximation of physics with higher order and continuity B-spline base functions.\" arXiv preprint arXiv:2201.00904 (2022).\n\n[3] Zhu, Xuedong, et al. \"A Best-Fitting B-Spline Neural Network Approach to the Prediction of Advection–Diffusion Physical Fields with Absorption and Source Terms.\" Entropy 26.7 (2024): 577."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"Originality: This paper has good mathematical intuition and a novel way to estimate the solution of PDEs using B-splines. This approach makes it easier to use PINN loses using analytical expressions improving efficiency over autograd.  The use of control points allows for any Dirichlet boundary conditions and initial conditions to be directly assigned. \n\nQuality: The method is very clearly described, and the authors provide a good mathematical background. They clearly state the difference between their proposed approach and KAN based approaches. The paper provides theoretical bounds and empirical evidence. \n\nSignificance: This work has a lot of potential in the field of scientific ML, as the authors show significant speedup over Physics informed ML techniques, and a reasonable improvement in accuracy of predictions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the authors provide a good theoretical basis for B-spline based PINNs, the experimental section is quite small. They evaluate on 2 settings. However, it would be good to showcase the performance on other standard evaluation benchmarks (Navier-Stokes, Darcy Flow, Burgers, Kolmogorov flow etc.) \n\nThe use of splines to train PINNs is not new. [1] uses Hermite splines to approximate PDEs. [2], [3] proposes using neural networks to learn the coefficients of B-spline functions. It would be good to show the differences between the proposed approach and these works. \n\n[1] Wandel, Nils, et al. \"Spline-pinn: Approaching pdes without data using fast, physics-informed hermite-spline cnns.\" Proceedings of the AAAI conference on artificial intelligence. Vol. 36. No. 8. 2022.\n\n[2] Doległo, Kamil, et al. \"Deep neural networks for smooth approximation of physics with higher order and continuity B-spline base functions.\" arXiv preprint arXiv:2201.00904 (2022).\n\n[3] Zhu, Xuedong, et al. \"A Best-Fitting B-Spline Neural Network Approach to the Prediction of Advection–Diffusion Physical Fields with Absorption and Source Terms.\" Entropy 26.7 (2024): 577."}},"nonreaders":[],"tmdate":1731427986194,"tcdate":1730130913390,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12276/Reviewer_KSqP"],"signatures":["ICLR.cc/2025/Conference/Submission12276/Reviewer_KSqP"],"forum":"rCvdAVQpAe","number":1,"license":"CC BY 4.0","cdate":1730130913390,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12276/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427986194,"domain":"ICLR.cc/2025/Conference","replyto":"rCvdAVQpAe","id":"aAdjB9QNq8","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed machine learning","B-splines","Partial differential equations (PDEs)"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Physics-informed machine learning provides an approach to combing data and governing physics laws for solving complex partial differential equations (PDEs). However, efficiently solving PDEs with varying parameters and changing initial conditions and boundary conditions (ICBCs) remains an open challenge. We propose a hybrid framework that uses a neural network to learn B-spline control points to approximate solutions to PDEs with varying system and ICBC parameters. The proposed network can be trained efficiently as one can directly specify ICBCs without imposing losses, calculate physics-informed loss functions through analytical formulas, and requires only learning the weights of B-spline functions as opposed to both weights and basis as in traditional neural operator learning methods. We show theoretical guarantees that the proposed B-spline networks are universal approximators of arbitrary dimensional PDEs under certain conditions. We also demonstrate in experiments that the proposed B-spline network can solve problems with discontinuous ICBCs and outperforms existing methods, and is able to learn solutions of 3D heat equations with diverse initial conditions."},"_bibtex":{"value":"@misc{\nwang2025physicsinformed,\ntitle={Physics-Informed Deep B-Spline Networks},\nauthor={Zhuoyuan Wang and Raffaele Romagnoli and Jasmine Ratchford and Yorie Nakahira},\nyear={2025},\nurl={https://openreview.net/forum?id=rCvdAVQpAe}\n}"},"title":{"value":"Physics-Informed Deep B-Spline Networks"},"pdf":{"value":"/pdf/b4fca5b407d63b3e6e53576c83d34f3d3bdcf008.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|physicsinformed_deep_bspline_networks"},"authorids":{"value":["~Zhuoyuan_Wang1","~Raffaele_Romagnoli1","~Jasmine_Ratchford1","~Yorie_Nakahira2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhuoyuan Wang","Raffaele Romagnoli","Jasmine Ratchford","Yorie Nakahira"]}},"version":2},{"content":{"venue":{"value":"IECON 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10904979/10905066/10905551.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"meiramov|enhancing_human_pose_estimation_accuracy_using_synthetic_data"},"html":{"value":"https://doi.org/10.1109/IECON55916.2024.10905551"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iecon/MeiramovBVY24,\n  author={Rakhat Meiramov and Zarema Balgabekova and Huseyin Atakan Varol and Adnan Yazici},\n  title={Enhancing Human Pose Estimation Accuracy Using Synthetic Data},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-6},\n  url={https://doi.org/10.1109/IECON55916.2024.10905551},\n  booktitle={IECON},\n  crossref={conf/iecon/2024}\n}\n"},"abstract":{"value":"In industrial applications, Human Pose Estimation (HPE) is crucial for enhancing both automation and human-computer interaction. This study investigates the impact of synthetic data on HPE model efficacy, particularly examining the performance of the YOLOv8 algorithm. Using Nvidia Omniverse Isaac Sim, we created a synthetic dataset called ISAAC, designed for various complex scenarios. This tool was chosen for its ability to simulate highly realistic and intricate industrial contexts with advanced physics and AI capabilities. The inclusion of this synthetic dataset significantly enhances model accuracy, evidenced by up to a 19% increase in mean Average Precision (mAP) at an Intersection over Union (IoU) of 0.5, and a 12% improvement across the 0.5-0.95 IoU range compared to traditional datasets. These results highlight the substantial advantages of synthetic data in training more accurate and robust HPE models, advocating for the integration of innovative data solutions in the field of computer vision."},"title":{"value":"Enhancing Human Pose Estimation Accuracy Using Synthetic Data"},"authors":{"value":[{"fullname":"Rakhat Meiramov","username":"~Rakhat_Meiramov1"},{"fullname":"Zarema Balgabekova","username":""},{"fullname":"Huseyin Atakan Varol","username":"~Huseyin_Atakan_Varol2"},{"fullname":"Adnan Yazici","username":""}]}},"tmdate":1779712299054,"pdate":1735603200000,"externalIds":["dblp:conf/iecon/MeiramovBVY24"],"tcdate":1779709233841,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~HUSEYIN_ATAKAN_VAROL1"],"forum":"LmSkdsBDYm","license":"CC BY-SA 4.0","number":22885,"cdate":1704067200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1779712299054,"domain":"OpenReview.net/Public_Article","id":"LmSkdsBDYm","version":2},{"content":{"TLDR":{"value":"We introduce SINO, a neural operator that learns PDE dynamics from limited trajectories without any explicit PDE knowledge, achieving accurate inference, strong out‑of‑distribution generalization, and discretization invariance."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["PDE modeling","neural operators","spectral methods","AI for Physics"]},"supplementary_material":{"value":"/attachment/68ac10699e8e412c92fb453ec4479640b0ea2a38.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Learning PDE dynamics from limited data with unknown physics is challenging. Existing neural PDE solvers either require large datasets or rely on known physics (e.g., PDE residuals or handcrafted stencils), leading to limited applicability. To address these challenges, we propose Spectral-Inspired Neural Operator (SINO), which can model complex systems from just 2-5 trajectories, without requiring explicit PDE terms. Specially, SINO automatically captures both local and global spatial derivatives from frequency indices, enabling a compact representation of the underlying differential operators in physics-agnostic regimes. To model nonlinear effects, it employs a $\\Pi$-block that performs multiplicative operations on spectral features, complemented by a low-pass filter to suppress aliasing. Extensive experiments on both 2D and 3D PDE benchmarks demonstrate that SINO achieves state-of-the-art performance, with improvements of 1–2 orders of magnitude in accuracy. Particularly, with only 5 training trajectories, SINO outperforms data-driven methods trained on 1000 trajectories and remains predictive on challenging out-of-distribution cases where other methods fail."},"_bibtex":{"value":"@misc{\nwan2026spectralinspired,\ntitle={Spectral-inspired Operator Learning with Limited Data and Unknown Physics},\nauthor={Han Wan and Rui Zhang and Hao Sun},\nyear={2026},\nurl={https://openreview.net/forum?id=1GU1tqF2Ev}\n}"},"title":{"value":"Spectral-inspired Operator Learning with Limited Data and Unknown Physics"},"pdf":{"value":"/pdf/cc3212713f4ca0266d71c8212ecfb0ad6876619a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wan|spectralinspired_operator_learning_with_limited_data_and_unknown_physics"},"authorids":{"value":["~Han_Wan1","~Rui_Zhang22","~Hao_Sun4"]},"authors":{"value":["Han Wan","Rui Zhang","Hao Sun"]}},"tmdate":1770804810537,"tcdate":1756896770751,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1605/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission1605/Authors"],"forum":"1GU1tqF2Ev","license":"CC BY 4.0","number":1605,"cdate":1756896770751,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission1605/-/Full_Submission","ICLR.cc/2026/Conference/Submission1605/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770804810537,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"1GU1tqF2Ev","version":2},{"content":{"summary":{"value":"This paper studies a critical problem about data synthesis for improving visual grounding capabilities of vision and language models. It explores various strategies for generating synthetic image-text pairs and image-text-box triplets to enhance model training, comparing synthetic data with real and web-crawled data. The proposed SynGround pipeline demonstrates that synthetic data can effectively improve the localization capabilities of existing models. Notably, SynGround boosts pointing game accuracy for models like ALBEF and BLIP on benchmarks like RefCOCO+ and Flickr30k, showing the potential of synthetic data for scalable improvements in visual grounding tasks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. The selection of base vision and language models. Why not applied to more recent SOTAs. Is SynGround data benefiting recent SOTAs in visual grounding also?\n2. How does SynGround compare with exisiting visual grounding data collected from public source with bounding boxes synthesized as well?"},"rating":{"value":3},"details_of_ethics_concerns":{"value":"1. Discrimination / bias / fairness: Images with human-generated content may reflect biases, leading to fairness concerns in model training and predictions.\n2. Legal compliance: Images containing identifiable human features may raise GDPR and copyright concerns if used without consent or proper authorization.\n3. Responsible research: Releasing datasets with human-generated images requires careful handling to protect privacy and prevent potential misuse, especially if individuals are recognizable."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Visual grounding is an essential problem with current vision and language models. It's important to study an effective approach to build synthetic data to further scale up models' visual grounding capabilities. This paper is one of the approaches that study how to generate such data, and with comparisons of various approaches to generate such data.\n2. SynGround improves the pointing game accuracy of pretrained ALBEF and BLIP significantly."},"flag_for_ethics_review":{"value":["Yes, Discrimination / bias / fairness concerns","Yes, Legal compliance (e.g., GDPR, copyright, terms of use)","Yes, Responsible research practice (e.g., human subjects, data release)"]},"weaknesses":{"value":"1. Previous synthetic visual grounding dataset are missing, for example, GRIT data - \"a Ground-and-Refer Instruction-Tuning dataset with 1.1M samples.\nGRIT contains multiple levels of spatial knowledge, covering objects, relationships, region descriptions, and complex reasoning\" - proposed in Ferret is not compared with. It's not clear how proposed SynGround is differing from previous synthetic visual grounding data, and how it surpasses previous data generation approaches.\n2. The main tables lack important SOTA baselines, for example, Shrika and Ferret on RefCOCO+ and Flickr, which are a lot better than the model fine-tuned on SynGround on RefCOCO+, and similar on Flickr.\n3. In Table 1, also the proposed approach get average 0.36 marginal improvement, and also no better than directly fine-tuning on existing VG data, which get average 0.96 improvement."}},"nonreaders":[],"tmdate":1731427462584,"tcdate":1730780431059,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1676/Reviewer_pC1Y"],"signatures":["ICLR.cc/2025/Conference/Submission1676/Reviewer_pC1Y"],"forum":"EuoHhIqvRD","number":4,"license":"CC BY 4.0","cdate":1730780431059,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1676/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427462584,"domain":"ICLR.cc/2025/Conference","replyto":"EuoHhIqvRD","id":"rD7rnLJT5B","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Visual Grounding","Referring Expression Comprehension","Learning from Models","Synthetic Data"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper extensively investigates the effectiveness of synthetic training data to improve the capabilities of vision-and-language models for grounding textual descriptions to image regions. We explore various strategies to best generate image-text pairs and image-text-box triplets using a series of pretrained models under different settings and varying degrees of reliance on real data. Through comparative analyses with synthetic, real, and web-crawled data, we identify factors that contribute to performance differences, and propose SynGround, an effective pipeline for generating useful synthetic data for visual grounding. Our findings show that SynGround can improve the localization capabilities of off-the-shelf vision-and-language models and offers the potential for infinite data generation. Particularly, SynGround improves the pointing game accuracy of pretrained ALBEF and BLIP models by 4.81% and 17.11% absolute percentage points, respectively, across the RefCOCO+ and the Flickr30k benchmarks."},"_bibtex":{"value":"@misc{\nhe2024is,\ntitle={Is Synthetic Data Ready for Improving Visual Grounding?},\nauthor={Ruozhen He and Ziyan Yang and Paola Cascante-Bonilla and Alexander C. Berg and Vicente Ordonez},\nyear={2024},\nurl={https://openreview.net/forum?id=EuoHhIqvRD}\n}"},"title":{"value":"Is Synthetic Data Ready for Improving Visual Grounding?"},"pdf":{"value":"/pdf/0451ed16c34402e803305d2c3eed2bcf1e789c97.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"he|is_synthetic_data_ready_for_improving_visual_grounding"},"authorids":{"value":["~Ruozhen_He1","~Ziyan_Yang1","~Paola_Cascante-Bonilla1","~Alexander_C._Berg1","~Vicente_Ordonez2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruozhen He","Ziyan Yang","Paola Cascante-Bonilla","Alexander C. Berg","Vicente Ordonez"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a novel framework for training diffusion samplers by jointly optimizing both the generative and destruction processes in discrete time. Unlike prior approaches that fix the destruction process or assume continuous-time SDE dynamics, the authors propose learning both forward and reverse transitions, with decoupled, state-dependent variances parameterized via neural networks. This flexibility enables better adaptation to limited-step sampling regimes and complex energy landscapes."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"See the weakness part."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Novel joint training of generation and destruction processes in diffusion samplers, enabling improved convergence and sampling quality, especially in few-step regimes.\n\n2. Flexible design with state-dependent, decoupled variances for both processes—only possible in discrete-time formulation—leading to enhanced adaptability to complex energy landscapes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited visual results: The paper presents few qualitative or visual examples (only human faces in Fig. 4), making it difficult to fully assess sampling quality, especially in image-related tasks.\n\n2. No discussion of limitations: The paper lacks a section acknowledging potential limitations (e.g., scalability to more complex distributions, sensitivity to architecture choices), which raises concerns about generalizability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920785682,"tcdate":1762016395523,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9079/Reviewer_sTsU"],"signatures":["ICLR.cc/2026/Conference/Submission9079/Reviewer_sTsU"],"forum":"38Ey1FrSDt","number":3,"license":"CC BY 4.0","cdate":1762016395523,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9079/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920785682,"domain":"ICLR.cc/2026/Conference","replyto":"38Ey1FrSDt","id":"mMnIBX2mI3","forumContent":{"TLDR":{"value":"Training the destruction (noising) process with state-dependent means and variances in diffusion-based samplers of unnormalised densities improves few-step modelling on standard benchmarks and high-dim applications."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["diffusion samplers","diffusion models","sampling","amortized inference","GFlowNets","reinforcement learning"]},"supplementary_material":{"value":"/attachment/e76d57b2c7ed43ceef394524275f013e4d38be39.zip"},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"abstract":{"value":"This paper explores the challenges and benefits of a trainable destruction process in diffusion samplers -- diffusion-based generative models trained to sample an unnormalised density without access to data samples. Contrary to the majority of work that views diffusion samplers as approximations to an underlying continuous-time model, we view diffusion models as discrete-time policies trained to produce samples in very few generation steps. We propose to trade some of the elegance of the underlying theory for flexibility in the definition of the generative and destruction policies. In particular, we decouple the generation and destruction variances, enabling both transition kernels to be learnt as unconstrained Gaussian densities. We show that, when the number of steps is limited, training both generation and destruction processes results in faster convergence and improved sampling quality on various benchmarks. Through a robust ablation study, we investigate the design choices necessary to facilitate stable training. Finally, we show the scalability of our approach through experiments on GAN latent space sampling for conditional image generation."},"_bibtex":{"value":"@misc{\ngritsaev2026adaptive,\ntitle={Adaptive Destruction Processes for Diffusion Samplers},\nauthor={Timofei Gritsaev and Nikita Morozov and Kirill Tamogashev and Daniil Tiapkin and Sergey Samsonov and Alexey Naumov and Dmitry Vetrov and Nikolay Malkin},\nyear={2026},\nurl={https://openreview.net/forum?id=38Ey1FrSDt}\n}"},"title":{"value":"Adaptive Destruction Processes for Diffusion Samplers"},"pdf":{"value":"/pdf/fd492bb74232049ced6c7f0084501c8680dc665d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"gritsaev|adaptive_destruction_processes_for_diffusion_samplers"},"authorids":{"value":["~Timofei_Gritsaev1","~Nikita_Morozov1","~Kirill_Tamogashev1","~Daniil_Tiapkin1","~Sergey_Samsonov1","~Alexey_Naumov1","~Dmitry_Vetrov1","~Nikolay_Malkin1"]},"authors":{"value":["Timofei Gritsaev","Nikita Morozov","Kirill Tamogashev","Daniil Tiapkin","Sergey Samsonov","Alexey Naumov","Dmitry Vetrov","Nikolay Malkin"]}},"version":2},{"content":{"summary":{"value":"The paper focuses on tackling the challenge of generating robust and reliable SQL for a given natural language question. The authors propose SGU-SQL which is a novel framework leveraging the structural relationship between entities in user questions and the database tables for solving Text to SQL tasks using LLMs. There are three main challenges that the authors focus on:\n1. Ambiguous user intents which lead to LLM's making uninformed decisions about SQL generation\n2. Complex database schemas which may lead the LLM to confuse between different columns or tables\n3. Complexity of SQL - LLMs are trained mainly for text generation whereas SQL is complex like programming language where only well-defined syntax must be used.\n\nThe authors tackle these issues by creating this framework where entities in natural language questions and the database are represented using graphs. These entities from questions to database are linked using a structure-linking model forming connections. Then, finally a prompt based method decomposes the complex question into subquestions and tackle these subquestions separately while aggregating them later. This method has been experimented with GPT-4 and a comparison has been reported using several models including SOTA from the leaderboard. Overall, the framework is proved to be successful while also achieving SOTA results on"},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1. In a real world, the databases are changed very often and there may be challenges where the proposed framework might not work same as proposed. In such cases, what components of the framework can still be relevant and what components must be changed is particularly industrial practitioners care about. Atleast including it in future directions might suffice.\n2. How does this framework vary for dialects other than SQL? Not a super important question to clarify for the purpose of this conference, but having information about this can be helpful for practitioners."},"rating":{"value":8},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"1. The paper focuses on one of the critical issues with LLMs for SQL generation which requires the LLMs to tackle them in a unique way where reliability and robustness are a strong focus.\n2. The framework is novel, intuitive and well designed to bridge the gap between an LLM which is tuned effectively for text generation and the SQL generation task which requires a structured understanding where the user queries can be ambiguous and answers can be complex.\n3. The paper demonstrates strong results on the most popular SQL datasets SPIDER and BIRD which certifies the potential of the approach.\n4. Well articulated work and not complex to follow from beginning to end. Conducted experiments including ablation studies are designed with thorough understanding.\n5. The real world applicability and potential of this framework can be very high and effective when tackling large production grade databases in industrial applications."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The framework is novel and intuitive but because of SPIDER not being the best representation of production grade SQL, it would also beneficial to probably test them using SPIDER-2. BIRD is somewhat closer to a real-world SQL but SPIDER-2 has large-scale and diverse databases which evaluates the proposal on a much better dataset revealing how the proposal works on large scale datasets.\n2. There's not much information provided on the future direction of this work which might hurt the significance/potential of the work a little. I'm mainly looking for how can this work be extended as it can provides opportunities for other researchers trying to explore in the same direction.\n3. Specific discussion on the model used for structure linking seem insufficient. Like what training methodology and architecture choices.\n4. When we discuss about constructing graphs and linking nodes, it may be important to learn about the complexities at a high level."}},"nonreaders":[],"tmdate":1731428796591,"tcdate":1730704676292,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7028/Reviewer_vRvu"],"signatures":["ICLR.cc/2025/Conference/Submission7028/Reviewer_vRvu"],"forum":"ReKWjKvkJE","number":2,"license":"CC BY 4.0","cdate":1730704676292,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7028/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428796591,"domain":"ICLR.cc/2025/Conference","replyto":"ReKWjKvkJE","id":"Ob4F8TNWRA","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-SQL","large language model","structure learning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in large language models (LLMs) have shown promise in bridging the gap between natural language queries and database management systems, enabling users to interact with databases without the background of SQL. However, LLMs often struggle to fully exploit and comprehend the user intention and complex structures of databases. Decomposition-based methods have been proposed to enhance the performance of LLMs on complex tasks, but decomposing SQL generation into subtasks is non-trivial due to the declarative structure of SQL syntax and the intricate connections between query concepts and database elements. In this paper, we propose a novel $\\textbf{S}$tructure $\\textbf{GU}$ided text-to-$\\textbf{SQL}$ framework ($\\textbf{SGU-SQL}$) that incorporates syntax-based prompting to enhance the SQL generation capabilities of LLMs. Specifically, SGU-SQL establishes structure-aware links between user queries and database schema and recursively decomposes the complex generation task using syntax-based prompting to guide LLMs in incrementally constructing target SQLs. Extensive experiments on two benchmark datasets demonstrate that SGU-SQL consistently outperforms state-of-the-art text-to-SQL baselines. These results highlight the importance of incorporating structural syntax information for effective text-to-SQL generation and pave the way for more robust and reliable interfaces to databases in the era of artificial intelligence."},"_bibtex":{"value":"@misc{\nzhang2025structureguided,\ntitle={Structure-Guided Large Language Models for Text-to-{SQL} Generation},\nauthor={Qinggang Zhang and Hao Chen and Junnan Dong and Wentao Li and Feiran Huang and Xiao Huang},\nyear={2025},\nurl={https://openreview.net/forum?id=ReKWjKvkJE}\n}"},"title":{"value":"Structure-Guided Large Language Models for Text-to-SQL Generation"},"pdf":{"value":"/pdf/eb8009e7c7e32d38bb9c9118af71052bcf1ca5cf.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|structureguided_large_language_models_for_texttosql_generation"},"authorids":{"value":["~Qinggang_Zhang2","~Hao_Chen18","~Junnan_Dong1","~Wentao_Li3","~Feiran_Huang1","~Xiao_Huang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Qinggang Zhang","Hao Chen","Junnan Dong","Wentao Li","Feiran Huang","Xiao Huang"]}},"version":2},{"content":{"summary":{"value":"This work focuses on enhancing the CoT capabilities of VLMs to serve complex vision-language reasoning tasks. The proposed approach includes generating synthetic data, SFT for basic CoT capability, and DPO for further calibration of complex reasoning ability stages.\n\nThe method is overall reasonable. With results on 9 datasets, the authors demonstrate the effectiveness of the proposed method. The experiments are clear and the analysis is comprehensive. Overall, this work positively contributes to the applications that involve VLM reasoning."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Given the concerns raised in the Weakness section, more discussion on each point summarized in lines 88-92 would help to showcase the novelty and contributions of the work. \n\nTo be more clear, lines 88-92 claim the contribution is to (A) provide a GPT-generated dataset, (B) verify that this dataset can be used for SFT, and (C) use DPO after SFT through an approach similar to [3] is meaningful. This makes me feel that this work is more of a resource paper. I would like to learn more about the technical contributions.\n\nReference:\n\n[3] RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold: https://arxiv.org/pdf/2406.14532"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The overall design makes sense. As demonstrated by the benchmark results, this could be a possible choice for implementing vision-language applications.\n2. The release of synthetic data generated by GPT-4o contributes to the VLM finetuning.\n3. The failure analysis in this paper is insightful and potentially helpful for other VLM research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The novelty is unclear.\n- From the idea level, tuning for CoT ability is being actively explored in LLM-centric research (e.g., [1], [2], [3])\n- For the overall design, the proposed method has recently been widely used. Including,\n   - Step 1 directly leverages the commonly seen approach, i.e., knowledge distillation from a larger teacher model (GPT-4). (e.g., [4])\n   - Two-stage tuning (SFT-RL) has been shown to be effective in improving reasoning ability in many LLM works, either through \n       - Pertaining (e.g., Llama-3) \n       - SFT on synthetic CoT data, then use RL to refine the LLM's reasoning abilities (e.g., [5]) \n\nBesides, one minor problem is missing related work in language-based reasoning. \nSpecifically, \"How does LLM research tackle the author-targeted challenges in pure language tasks?\" This background discussion would be beneficial because this paper particularly studies the Vision-Language Model, which utilizes a vision encoder to project visual information onto language space; thus, it may share certain characteristics with its underlying LLM.\n\nReference:\n\n[1] Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models. https://arxiv.org/pdf/2312.06585\n\n[2] LogiCoT: Logical Chain-of-Thought Instruction Tuning. https://arxiv.org/pdf/2305.12147\n\n[3] RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold: https://arxiv.org/pdf/2406.14532\n\n[4] LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents. https://arxiv.org/pdf/2311.05437\n\n[5] Tackling Vision Language Tasks Through Learning Inner Monologues. https://arxiv.org/pdf/2308.09970"}},"nonreaders":[],"tmdate":1731428684331,"tcdate":1730453208893,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12273/Reviewer_TsHB"],"signatures":["ICLR.cc/2025/Conference/Submission12273/Reviewer_TsHB"],"forum":"XgYZT35N76","number":2,"license":"CC BY 4.0","cdate":1730453208893,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12273/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428684331,"domain":"ICLR.cc/2025/Conference","replyto":"XgYZT35N76","id":"sn9S45aY71","forumContent":{"TLDR":{"value":"We improve vision language model chain-of-thought reasoning"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Vision Language Model","Chain-of-thought Reasoning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Chain-of-thought (CoT) reasoning in vision language models (VLMs) is crucial for improving interpretability and trustworthiness. \nHowever, current training recipes lack robust CoT reasoning data, relying on datasets dominated by short annotations with minimal rationales. In this work, we first evaluate the CoT abilities of existing VLMs and show that training on short answers does not generalize well to reasoning tasks that require more detailed responses. To address this, we propose a two-fold approach. First, we distill rationales from GPT-4o model to enrich the training data and fine-tune VLMs, boosting their CoT performance. Second, we apply reinforcement learning to further calibrate reasoning quality by constructing positive (correct) and negative (incorrect) pairs of model-generated reasoning chains, based on the comparisons with annotated short answers. We then use the Direct Preference Optimization algorithm on this pairwise data to refine the model’s reasoning abilities. Our experiments demonstrate significant improvements in CoT reasoning on benchmark datasets and better generalization to direct answer prediction as well. This work emphasizes the importance of incorporating detailed rationales in training and leveraging reinforcement learning to strengthen the reasoning capabilities of VLMs."},"_bibtex":{"value":"@misc{\nzhang2024improve,\ntitle={Improve Vision Language Model Chain-of-thought Reasoning},\nauthor={Ruohong Zhang and Bowen Zhang and Yanghao Li and Haotian Zhang and Zhiqing Sun and Zhe Gan and Yinfei Yang and Ruoming Pang and Yiming Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=XgYZT35N76}\n}"},"title":{"value":"Improve Vision Language Model Chain-of-thought Reasoning"},"pdf":{"value":"/pdf/b72cc41c24b6688d70c14008b88cd77708cdb8dd.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|improve_vision_language_model_chainofthought_reasoning"},"authorids":{"value":["~Ruohong_Zhang1","~Bowen_Zhang2","~Yanghao_Li1","~Haotian_Zhang3","~Zhiqing_Sun1","~Zhe_Gan1","~Yinfei_Yang1","~Ruoming_Pang2","~Yiming_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruohong Zhang","Bowen Zhang","Yanghao Li","Haotian Zhang","Zhiqing Sun","Zhe Gan","Yinfei Yang","Ruoming Pang","Yiming Yang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes an ensemble (teacher-student) based physics-informed and data-driven model for traffic state estimation. Specifically, due to the well-known complexity of coupled modeling of physics information (in the form of PDE based conservation conditions) and data-driven (e.g., Mean-squared error) losses, the paper proposes a de-coupled method of modeling these two steps by first employing a physics-informed neural network to model the PDE based losses and a Bayesian Neural Network to model the data-driven losses. Further, a feed-forward neural network is employed to ensemble the predictions of the physics-informed and data-driven components by employing the latent representations from each of the two (I.e., physics-informed and data-driven) models. The ensemble model (infused with both data-driven and physics-driven knowledge) is trained in a data-driven manner and knowledge from the ensemble (teacher) model is distilled into each individual physics-driven, data-driven (student) models. In this manner, employing a combination of ensemble modeling, physics informed neural networks and knowledge distillation, the paper proposes a Physics-Informed Deep Learning (PIDL) solution for traffic state estimation."},"presentation":{"value":"3 good"},"contribution":{"value":"1 poor"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. Two interesting parts about the paper are (i) the alternating teacher / student training (ii) the uncertainty fusion employing the BNN covariance matrix.  However, not enough is discussed about either of these aspects in the main paper. \n  \n\n2. Overall the paper is cohesive and well-written.  The quantitative and qualitative results (although incomplete) demonstrate that the proposed framework yields good performance relative to baselines in addition to highlighting the importance of each component (ablation analysis)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **[Limited Novelty].** The proposed model is a combination of a physics informed neural network (PINN) employed with a relatively simple PDE, in addition to another standard Bayesian Neural Network (BNN) paradigm combined with knowledge distillation (KD). These three paradigms (PINN, BNN, KD) are all extremely well studied, well-understood and the current paper doesn’t propose any novel extensions of the actual paradigms or increase the characteristic understanding of any of the aforementioned paradigms. It is simply an exercise in the application (specifically, the combination) of these paradigms in a (somewhat) creative way to address the problem of traffic state estimation.  \n    - Further, the knowledge that the physics losses and data-driven loss don’t always play well together (mentioned in contribution 1) isn’t new and has been well researched in the context of physics-informed neural networks applied to PDEs [3, 4, 5]\t \n \n\n- **[Important Related Work Missing].** The paper completely misses mentioning operator learning paradigms which are more recent updates to traditional PINNs which learn families of PDEs as opposed to single instances of PDEs (albeit in a slightly different manner) however the reviewer believes that a brief description of operators like [1, 2] would contextualize the current work in the physics informed deep learning space.  \n \n\n- **[Incomplete Performance comparison].** A related paper Trafficflowgan [6] that employs physics information, uncertainty aware GAN for traffic state estimation has not been compared against. \n\n### References:\n\n1. Lu, Lu, Pengzhan Jin, and George Em Karniadakis. \"Deeponet: Learning nonlinear operators for identifying differential equations based on the universal approximation theorem of operators.\" arXiv preprint arXiv:1910.03193 (2019). \n \n\n2. Li, Zongyi, et al. \"Fourier neural operator for parametric partial differential equations.\" arXiv preprint arXiv:2010.08895 (2020). \n \n\n3. Krishnapriyan, Aditi, et al. \"Characterizing possible failure modes in physics-informed neural networks.\" Advances in Neural Information Processing Systems 34 (2021): 26548-26560. \n \n\n4. Wang, Sifan, Yujun Teng, and Paris Perdikaris. \"Understanding and mitigating gradient flow pathologies in physics-informed neural networks.\" SIAM Journal on Scientific Computing 43.5 (2021): A3055-A3081. \n \n\n5. Kim, Jungeun, et al. \"DPM: A novel training method for physics-informed neural networks in extrapolation.\" Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 35. No. 9. 2021. \n \n\n6. Mo, Zhaobin, et al. \"Trafficflowgan: Physics-informed flow based generative adversarial network for uncertainty quantification.\" Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Cham: Springer Nature Switzerland, 2022."},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. Why has TrafficflowGAN [6] not been compared despite being a related / physics-informed + uncertainty aware model for traffic state estimation?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636262436,"tcdate":1698617274630,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3151/Reviewer_gduu"],"signatures":["ICLR.cc/2024/Conference/Submission3151/Reviewer_gduu"],"forum":"GszBQ3ZTzk","number":2,"license":"CC BY 4.0","cdate":1698617274630,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3151/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636262436,"domain":"ICLR.cc/2024/Conference","replyto":"GszBQ3ZTzk","id":"DvBFJkJUH7","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed deep learning","Traffic state estimation","Knowledge distillation","Ensemble learning"]},"supplementary_material":{"value":"/attachment/a2e6b5b1436004e0907405012360cce043adef51.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Traditional physics-informed deep learning combines the data-driven methods with the model-based methods by incorporating physics loss as a constraint in total loss function in general, which aims to enforce the neural network to behave according to the physics property. However, this simple integration makes physical knowledge submerged in data information since data loss and physics loss could have large magnitude differences, conflicting directions of the gradients, and varying convergence rates so that the physics law may not work as expected and inhibits the model from working effectively furthermore, especially for traffic state estimation (TSE). To alleviate these issues, we propose a Physical knowledge combined Data information neural network with Ensemble Distillation framework (PDED) to first disentangle the data-driven model and physics-based model, and then reassemble them to take advantages of label information and physics property. Practically, we separately train data-driven model based on true labels and physics-based model according to physics laws. Then, we introduce the ensemble learning and knowledge distillation to assemble their representations of these two models for constructing a more competitive learnable online teacher model, which in turn distills knowledge to guide the update of them for learning richer knowledge to improve the performance of student models. Through extensive experiments on both synthetic dataset and real-world datasets, our model demonstrates better performance than the existing state-of-the-art methods."},"_bibtex":{"value":"@misc{\nfu2024pded,\ntitle={{PDED}: Revitalize physics laws submerged in data information for Traffic State Estimation},\nauthor={Yao Fu and Hong Zhao and Xiaoyu Cai and Ruiheng Yang and Weihao Jiang and Shiliang Pu},\nyear={2024},\nurl={https://openreview.net/forum?id=GszBQ3ZTzk}\n}"},"title":{"value":"PDED: Revitalize physics laws submerged in data information for Traffic State Estimation"},"pdf":{"value":"/pdf/dfcd70006ce024e9a689e1fc20180c31fce193e5.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"fu|pded_revitalize_physics_laws_submerged_in_data_information_for_traffic_state_estimation"},"authorids":{"value":["~Yao_Fu7","~Hong_Zhao5","~Xiaoyu_Cai1","~Ruiheng_Yang1","~Weihao_Jiang2","~Shiliang_Pu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yao Fu","Hong Zhao","Xiaoyu Cai","Ruiheng Yang","Weihao Jiang","Shiliang Pu"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Vis. Comput. Graph. 1996"},"pdf":{"value":"https://ieeexplore.ieee.org/iel1/2945/10582/00489389.pdf"},"venueid":{"value":"dblp.org/journals/TVCG/1996"},"paperhash":{"value":"qin|dnurbs_a_physicsbased_framework_for_geometric_design"},"authorids":{"value":["~Hong_Qin1","https://dblp.org/search/pid/api?q=author:Demetri_Terzopoulos:"]},"html":{"value":"https://doi.org/10.1109/2945.489389"},"_bibtex":{"value":"@article{DBLP:journals/tvcg/QinT96,\n  author={Hong Qin and Demetri Terzopoulos},\n  title={D-NURBS: A Physics-Based Framework for Geometric Design},\n  year={1996},\n  cdate={820454400000},\n  journal={IEEE Trans. Vis. Comput. Graph.},\n  volume={2},\n  number={1},\n  pages={85-96},\n  url={https://doi.org/10.1109/2945.489389}\n}\n"},"abstract":{"value":"Presents dynamic non-uniform rational B-splines (D-NURBS), a physics-based generalization of NURBS. NURBS have become a de facto standard in commercial modeling systems. Traditionally, however, NURBS have been viewed as purely geometric primitives, which require the designer to interactively adjust many degrees of freedom-control points and associated weights-to achieve the desired shapes. The conventional shape modification process can often be clumsy and laborious. D-NURBS are physics-based models that incorporate physical quantities into the NURBS geometric substrate. Their dynamic behavior, resulting from the numerical integration of a set of nonlinear differential equations, produces physically meaningful, and hence intuitive shape variation. Consequently, a modeler can interactively sculpt complex shapes to required specifications not only in the traditional indirect fashion, by adjusting control points and setting weights, but also through direct physical manipulation, by applying simulated forces and local and global shape constraints. We use Lagrangian mechanics to formulate the equations of motion for D-NURBS curves, tensor-product D-NURBS surfaces, swung D-NURBS surfaces and triangular D-NURBS surfaces. We apply finite element analysis to reduce these equations to efficient numerical algorithms computable at interactive rates on common graphics workstations. We implement a prototype modeling environment based on D-NURBS and demonstrate that D-NURBS can be effective tools in a wide range of computer-aided geometric design (CAGD) applications."},"title":{"value":"D-NURBS: A Physics-Based Framework for Geometric Design"},"authors":{"value":["Hong Qin","Demetri Terzopoulos"]}},"tmdate":1731492014485,"pdate":820454400000,"tcdate":1731488909105,"writers":["~"],"signatures":["~Hong_Qin1"],"forum":"3ZosB6E47o","license":"CC BY-SA 4.0","number":219638,"cdate":820454400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731492014485,"domain":"DBLP.org","id":"3ZosB6E47o","version":2},{"content":{"summary":{"value":"This paper introduces a hierarchical framework, PHYLOMAN, for Generative Behavior Control (GBC), combining language-driven planning, diffusion-based motion generation, and physics-based control, and constructs a large-scale hierarchical text-to-motion dataset."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer the Weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.This paper introduces a hierarchical framework, PHYLOMAN, for Generative Behavior Control (GBC), combining language-driven planning, diffusion-based motion generation, and physics-based control.\n2.The paper constructs a large-scale hierarchical text-to-motion dataset with three levels of structured annotations: BehaviorScript, PoseScript, and MotionScript."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.While the proposed PHYLOMAN framework is structurally coherent, its components—an LLM-based planner, a motion diffusion model, and a physics controller—are largely based on existing paradigms.\n2.Although the paper cites MotionAgent (Wu et al., 2024) as a representative language-to-motion framework, there is no direct experimental comparison and analysis.\n3.In the main experimental section, PHYLOMAN is not included in the key comparison Table 2, which presents quantitative results across baselines. The authors should include PHYLOMAN in Table 2 using the same configuration and evaluation metrics as other methods. \n4.The GBC-100K dataset, described as containing 123.7K motion sequences and 250 hours of video, introduces hierarchical annotations: BehaviorScript, PoseScript, and MotionScript. While this is valuable, several issues arise:\nThe reported W-MPJPE ≈ 222 mm (Table 7) remains quite large for a high-quality motion dataset. The evaluation only includes PA-MPJPE, W-MPJPE, and RTE,and while the authors acknowledge the presence of typical error types in their data, it is necessary to address additional aspects of physical consistency, such as foot sliding, body penetration, failure ratio, and temporal jitter.\nThere is no analysis of long-horizon temporal consistency, which is crucial for “ultra-long” behaviors.\n5. The paper repeatedly refers to the motion diffusion backbone as “parallel-in-time”, implying computational efficiency. However, there is no quantitative evidence (e.g., speedup, training cost, or memory footprint) to support this claim."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915817947,"tcdate":1761991515779,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1568/Reviewer_1oWQ"],"signatures":["ICLR.cc/2026/Conference/Submission1568/Reviewer_1oWQ"],"forum":"SjjXr9mnlk","number":3,"license":"CC BY 4.0","cdate":1761991515779,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1568/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915817947,"domain":"ICLR.cc/2026/Conference","replyto":"SjjXr9mnlk","id":"jokJsHq69W","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We introduce Generative Behavior Control (GBC) and the GBC-100K dataset to generate long-horizon, goal-directed, and coherent humanoid behaviors via LLM planning and physics-informed control."},"keywords":{"value":["Human motion generation","Long-horizon synthesis","Task planning","Motion planning","Behavior control"]},"supplementary_material":{"value":"/attachment/e7d49c4fb719cea22b3a4182d317118447f766a0.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Human motion generative modeling aims to synthesize complex motions from daily activities. However, current research is fragmented, focusing on either low-level, short-horizon motions or high-level, disembodied action planning, thereby neglecting the hierarchical and goal-oriented nature of human activities. This work shifts the research focus from motion generation to the more holistic task of humanoid behavior modeling. To formally address this, we first introduce Generative Behavior Control (GBC), a new task focused on generating long-term, physically plausible, and semantically coherent behaviors from high-level intentions. To tackle this task, we present a novel framework that aligns motion synthesis with hierarchical plans generated by large language models (LLMs), leveraging principles from task and motion planning. Concurrently, to overcome the limitations of existing benchmarks, we introduce the GBC-100K dataset, a large-scale corpus annotated with hierarchical semantic and motion plans. Experimental results demonstrate our framework, trained on GBC-100K, generates more diverse and purposeful human behaviors with up to 10$\\times$ longer horizons than existing methods. This work lays a foundation for future research in behavior-centric modeling, with all code and data to be made publicly available."},"_bibtex":{"value":"@misc{\nzhang2025from,\ntitle={From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control},\nauthor={Jusheng zhang and Jinzhou Tang and Sidi Liu and Jian Wang and Keze Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=SjjXr9mnlk}\n}"},"title":{"value":"From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control"},"pdf":{"value":"/pdf/7c1888332dfe48654f7506c3c66c081e25486693.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|from_motion_to_behavior_hierarchical_modeling_of_humanoid_generative_behavior_control"},"authorids":{"value":["~Jusheng_zhang6","~Jinzhou_Tang1","~Sidi_Liu1","~Jian_Wang10","~Keze_Wang1"]},"authors":{"value":["Jusheng zhang","Jinzhou Tang","Sidi Liu","Jian Wang","Keze Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces **LLM-SR**, a novel framework for discovering scientific equations using **Large Language Models (LLMs)**. The approach integrates LLMs' scientific knowledge and code generation capabilities with evolutionary search and optimization techniques to generate, evaluate, and refine mathematical equation hypotheses iteratively. The key steps involve generating equation skeletons, optimizing parameters using Python-based tools, and managing an experience buffer to iteratively improve the search process.\n\n**Main Contributions:**\n1. **Novel Framework**: LLM-SR combines LLMs' strengths with evolutionary algorithms to navigate complex equation discovery.\n2. **Integration of Scientific Priors**: The method leverages LLMs' embedded scientific knowledge for more efficient hypothesis generation.\n3. **Benchmark Creation**: The authors design new benchmark problems in physics, biology, and materials science to test the method's capabilities and avoid memorization of well-known equations.\n4. **Performance Advantage**: LLM-SR reportedly outperforms traditional symbolic regression methods, showing better accuracy and generalization, especially in out-of-domain tests.\n5. **Ablation Study**: Demonstrates the importance of various components, such as problem specification and iterative refinement, in the overall performance of LLM-SR."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1. **Clarification on the Necessity of Complex Mechanisms**:\n   - *Question*: Could the authors elaborate on the rationale behind incorporating complex mechanisms such as the sampling strategy and island model? How do these elements contribute to the overall performance, especially compared to simpler alternatives?\n   - *Suggestion*: Including comparative results or ablation studies that illustrate the impact of these mechanisms on the method’s performance would be helpful to better understand their necessity and specific contributions.\n\n2. **Verification of LLM Prior Knowledge Utilization**:\n   - *Question*: How do the authors verify that the LLM’s embedded scientific prior knowledge is effectively utilized during the equation generation process? Are there examples or analyses demonstrating cases where this prior knowledge played a significant role?\n   - *Suggestion*: We suggest adding controlled experiments that isolate and assess the contribution of prior knowledge. Comparing models with and without domain-specific pretraining might further strengthen this aspect.\n\n3. **Evaluation on Real-World Datasets with Noise and Outliers**:\n   - *Question*: Do the authors plan to evaluate LLM-SR on more challenging, real-world datasets that contain noise, outliers, or measurement errors? How do they foresee the method performing under such conditions?\n   - *Suggestion*: Expanding the experiments to include data with inherent noise and outliers, or discussing the method’s anticipated robustness in these cases, could help demonstrate its applicability to complex, real-world scenarios.\n\n4. **Impact of LLM Programming Performance on Results**:\n   - *Question*: Have the authors considered examining the extent to which the LLM’s programming capabilities, rather than scientific reasoning, might influence the results? How do they ensure that the outcomes reflect genuine equation discovery rather than sophisticated code generation?\n   - *Suggestion*: Providing an analysis or qualitative examples that differentiate between the model’s programming skills and its scientific reasoning in generating equations would be valuable to clarify the contributions and limitations of the approach.\n\n5. **Interpretability of Results**:\n   - *Question*: The authors state that the generated equations are more interpretable and scientifically meaningful. Could they provide examples or a detailed comparison showing how LLM-SR’s output stands out in terms of interpretability relative to baseline methods?\n   - *Suggestion*: Presenting case studies or examples where the equations discovered by LLM-SR provide clearer or more scientifically insightful interpretations would reinforce this claim.\n\n6. **Broader Comparison with Hybrid Methods**:\n   - *Question*: How does LLM-SR compare with recent hybrid approaches that incorporate domain knowledge or leverage deep learning for symbolic regression? Are there particular strengths or limitations of LLM-SR compared to these methods?\n   - *Suggestion*: Including comparisons with other state-of-the-art hybrid methods or discussing potential improvements and challenges in adapting LLM-SR could provide more context on its positioning and competitive advantages."},"rating":{"value":8},"details_of_ethics_concerns":{"value":"For **ethical considerations**, we recommend the authors provide the following details:\n\n- **Licenses and Usage Permission Details for Datasets**:\n   - *Suggestion*: We encourage the authors to report the type of licenses and usage permissions associated with the datasets used in the study. Clarifying whether each dataset is publicly available and whether its use complies with the original publisher’s license agreement would enhance the transparency and compliance of the paper.\n   - *Rationale*: This is particularly important in academic research to ensure that the use of datasets does not violate copyright, privacy, or other relevant legal considerations. Clear documentation of dataset usage can help prevent potential ethical issues and provide future researchers with well-defined guidelines."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. **Innovative Application of LLMs for Scientific Equation Discovery**:\n   - The authors introduce an innovative approach by leveraging the scientific knowledge and code generation capabilities of large language models (LLMs) for equation discovery. Representing equations as code and integrating data-driven feedback and optimization provides a novel perspective for scientific modeling.\n\n2. **Experience Buffer and Iterative Optimization**:\n   - The use of an experience buffer to maintain high-quality equation examples and facilitate iterative optimization is a unique feature. This mechanism helps the model learn and recall effective generation patterns, improving the quality and efficiency of equation exploration.\n\n3. **Comprehensive Multi-Domain Benchmarking**:\n   - The paper demonstrates the method’s applicability across multiple scientific domains, including physics, biology, and materials science. The newly designed benchmarks, which aim to prevent memorization of known equations, reflect the authors' commitment to scientific rigor.\n\n4. **Promising Experimental Results**:\n   - Despite certain challenges and areas for improvement, LLM-SR shows promising performance compared to existing symbolic regression methods, particularly in out-of-domain tests. This indicates potential for tackling complex scientific problems and suggests a valuable direction for future research."},"flag_for_ethics_review":{"value":["Yes, Legal compliance (e.g., GDPR, copyright, terms of use)"]},"weaknesses":{"value":"**Complexity and Necessity of the Method**:\nThe authors introduce complex mechanisms such as sampling strategies and an island model, showcasing innovative efforts. However, clearer justification is needed to explain the necessity of these elements and their specific contributions to enhancing the method’s performance. Simplifying and elucidating these mechanisms' actual benefits would improve the clarity and persuasiveness of the method.\n\n**Utilization and Verification of Prior Knowledge**:\nAlthough the authors assume that the LLM leverages its scientific prior knowledge to generate reasonable equations, the paper lacks detailed analysis and experimental validation of how this mechanism functions and its tangible effects. Further exploration and evidence would more convincingly support this claim and make the method’s core advantage clearer to readers.\n\n**Breadth and Realism of Experimental Design**:\nThe current experiments are based on data generated from known equations, providing a basic evaluation. However, data in real-world scientific research often contains noise and outliers, and the method’s performance under such conditions is a critical indicator of its practical applicability. Including datasets with measurement errors and outliers would help demonstrate the method’s robustness and generalizability in complex, realistic scenarios.\n\n**Impact of Programming Capabilities on Experimental Results**:\nThe paper emphasizes the LLM's ability to generate programming code but lacks analysis of whether these programming capabilities interfere with the experimental results. The boundary between the LLM’s programming skills and scientific reasoning remains unclear, which might mean that experimental outcomes reflect coding proficiency rather than pure scientific inference. Clarifying the role of these aspects in equation generation would aid in better understanding the method’s true strengths and limitations."}},"nonreaders":[],"tmdate":1732722877522,"tcdate":1730720385657,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2272/Reviewer_89Tg"],"signatures":["ICLR.cc/2025/Conference/Submission2272/Reviewer_89Tg"],"forum":"m2nmp8P5in","number":3,"license":"CC BY 4.0","cdate":1730720385657,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2272/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732722877522,"domain":"ICLR.cc/2025/Conference","replyto":"m2nmp8P5in","id":"mZrXklOXz9","forumContent":{"TLDR":{"value":"We introduce LLM-SR, an approach that harnesses Large Language Models (LLMs) to discover governing equations from data in an efficient, knowledge-guided manner."},"venue":{"value":"ICLR 2025 Oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Symbolic Regression","Equation Discovery","Large Language Models","Evolutionary Search"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Mathematical equations have been unreasonably effective in describing complex natural phenomena across various scientific disciplines. However, discovering such insightful equations from data presents significant challenges due to the necessity of navigating extremely large combinatorial hypothesis spaces. Current methods of equation discovery, commonly known as symbolic regression techniques, largely focus on extracting equations from data alone, often neglecting the domain-specific prior knowledge that scientists typically depend on. They also employ limited representations such as expression trees, constraining the search space and expressiveness of equations. To bridge this gap, we introduce LLM-SR, a novel approach that leverages the extensive scientific knowledge and robust code generation capabilities of Large Language Models (LLMs) to discover scientific equations from data. Specifically, LLM-SR treats equations as programs with mathematical operators and combines LLMs' scientific priors with evolutionary search over equation programs. The LLM iteratively proposes new equation skeleton hypotheses, drawing from its domain knowledge, which are then optimized against data to estimate parameters. We evaluate LLM-SR on four benchmark problems across diverse scientific domains (e.g., physics, biology), which we carefully designed to simulate the discovery process and prevent LLM recitation. Our results demonstrate that LLM-SR discovers physically accurate equations that significantly outperform state-of-the-art symbolic regression baselines, particularly in out-of-domain test settings. We also show that LLM-SR's incorporation of scientific priors enables more efficient equation space exploration than the baselines."},"_bibtex":{"value":"@inproceedings{\nshojaee2025llmsr,\ntitle={{LLM}-{SR}: Scientific Equation Discovery via Programming with Large Language Models},\nauthor={Parshin Shojaee and Kazem Meidani and Shashank Gupta and Amir Barati Farimani and Chandan K. Reddy},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=m2nmp8P5in}\n}"},"title":{"value":"LLM-SR: Scientific Equation Discovery via Programming with Large Language Models"},"pdf":{"value":"/pdf/adefa2fe60eb12f3bd4e8145e6e921fbd5d96d33.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"shojaee|llmsr_scientific_equation_discovery_via_programming_with_large_language_models"},"authorids":{"value":["~Parshin_Shojaee1","~Kazem_Meidani1","~Shashank_Gupta3","~Amir_Barati_Farimani2","~Chandan_K._Reddy1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Parshin Shojaee","Kazem Meidani","Shashank Gupta","Amir Barati Farimani","Chandan K. Reddy"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Efficient Hybrid-fusion Physics-inspired Attention Learning Network (EHPAL-Net), a lightweight and scalable multimodal fusion framework designed for heterogeneous biomedical data, such as imaging, multi-omics, and EHR. Its Efficient Hybrid Fusion (EHF) layer sequentially integrates modalities through: 1) Efficient Multimodal Residual Convolution (EMRC) for multi-scale spatial representations, 2) Physics-inspired Cross-Modal Fusion Attention (PCMFA) combining hyperbolic and quantum-inspired attention to model complex cross-modal interactions, and 3) Shared Information Refinement (SIR) for representational diversity. A large number of heterogeneous medical datasets are used to validate the proposed method."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Please see the weaknesses section above."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- The problems to be addressed, the performance, generalization, and efficiency of multimodal fusion learning, are significant.\n- The integration of the hyperbolic dual-geometry attention has not been seen in the field of multimodal learning, and it looks novel to me.\n- The reported reduction in the number of model parameters (98.3%) and FLOPs (97.6) is impressive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The presentation of the paper could be improved. As a researcher working on multimodal learning, I find it difficult to follow the logic of the paper, which looks to me quite diffuse and redundant. Key ideas such as \"physics-inspired attention\", \"dual-geometry interaction\",  and \"hierarchical structure preservation\" are mentioned repeatedly, but I could not find their precise definitions or theoretical justifications. The narrative often cycles through the same claims without further clarification and justification of the mechanisms enabling the desired effects.\n- The proposed framework claims to be \"physics-inspired\", but the connection to physical principles appears largely metaphorical rather than mechanistic. The framework design and choice of the submodules are not evidently grounded in physical modeling.\n- I do not understand the exact meaning of \"modality\" in this paper. In lines 450, for example, it seems to imply that HAM10000, SIPaKMeD, and PathMNIST are treated as three modalities, each with its distinct data sources and label sets. The source codes provided by the authors seem to confirm my guess. The usual setting of multimodal fusion means that one sample has data of different modalities, e.g., a patient has a CT scan image and a pathological image, where information from these modalities is combined to make a prediction for the sample concerned. I think the exact definition of \"modality\" in this paper should be explicitly defined, and what the modalities are in each dataset should be clearly detailed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933713645,"tcdate":1761989818741,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20215/Reviewer_7JKi"],"signatures":["ICLR.cc/2026/Conference/Submission20215/Reviewer_7JKi"],"forum":"mZJM8hXmVg","number":3,"license":"CC BY 4.0","cdate":1761989818741,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20215/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933713645,"domain":"ICLR.cc/2026/Conference","replyto":"mZJM8hXmVg","id":"5JXy3DkI6z","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Medical Imaging","Multimodal Fusion","Attention","Deep Learning"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Multimodal fusion learning (MFL) paradigm (a framework to jointly learn from heterogeneous data sources) has shown great potential in various fields such as Medicine, Science, Engineering, etc. It is extremely desirable in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges. First, they struggle to capture complex cross-modal interactions effectively. Second, they are often designed and evaluated for narrow, fixed modality configurations (e.g., imaging-only, or specific pairs such as image and omics or image and clinical text), which limits evidence of their adaptability and generalizability to broader collections of heterogeneous medical modalities. Finally, they incur high computational costs, restricting their applicability in resource-constrained healthcare AI. To address these challenges, we propose a novel MFL framework – Efficient Hybrid-fusion Physics-inspired Attention Learning Network (EHPAL-Net) – a lightweight and scalable framework that integrates various modalities through novel Efficient Hybrid Fusion (EHF) layers. Each EHF layer captures rich modality-specific multi-scale spatial information, followed by a Physics-inspired Cross-modal Fusion Attention module to model fine-grained, structure-preserving cross-modal interactions, thereby learning robust complementary shared representations. Furthermore, EHF layers are sequentially learned for each modality, making them adaptable and generalizable. Extensive evaluations on 15 public datasets show that EHPAL-Net outperforms leading multimodal fusion methods, boosting performance by up to 3.97% and lowering computational costs by up to 87.8%, ensuring more effective and reliable predictions."},"_bibtex":{"value":"@misc{\ndhar2026advancing,\ntitle={Advancing Multimodal Fusion on Heterogeneous Data with Physics-inspired Attention},\nauthor={Joy Dhar and Chen Chen and Nayyar Zaidi and Maryam Haghighat and Puneet Goyal and Manish Kumar Pandey and Ferdous Sohel},\nyear={2026},\nurl={https://openreview.net/forum?id=mZJM8hXmVg}\n}"},"title":{"value":"Advancing Multimodal Fusion on Heterogeneous Data with Physics-inspired Attention"},"pdf":{"value":"/pdf/dcca73e1c79fba073f627966bc41011969e062a4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dhar|advancing_multimodal_fusion_on_heterogeneous_data_with_physicsinspired_attention"},"authorids":{"value":["~Joy_Dhar1","~Chen_Chen18","~Nayyar_Zaidi1","~Maryam_Haghighat1","~Puneet_Goyal1","~Manish_Kumar_Pandey1","~Ferdous_Sohel1"]},"authors":{"value":["Joy Dhar","Chen Chen","Nayyar Zaidi","Maryam Haghighat","Puneet Goyal","Manish Kumar Pandey","Ferdous Sohel"]}},"version":2},{"content":{"summary":{"value":"The paper presents a computational framework (Poly2Graph) and large-scale dataset that map non-Hermitian Hamiltonians of one-dimensional crystals into spectral graphs, i.e., geometric multigraphs embedded in the complex energy plane. The framework provides efficient construction of spectral graphs from Hamiltonians, enabling systematic dataset generation across varying scales. The largest constructed dataset, HSG-12M, contains 11.6 million graphs spanning 1,401 distinct Hamiltonian (characteristic polynomial) classes. The authors also benchmark multiple graph neural network (GNN) architectures on a graph-level classification task: predicting the underlying Hamiltonian class from the corresponding spectral graph."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See Weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Non-Hermitian physics has emerged as an important area in condensed matter and atomic physics. This work establishes a physics-grounded benchmark for evaluating and developing machine learning models on spatial multigraphs derived from physical systems.\n- The dataset is impressive in scale ($\\approx$12 million graphs), and the automated spectral graph extraction pipeline is technically solid, with potential benefits for the broader physics community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper does not clearly explain whether the proposed task, predicting the Hamiltonian class from its spectral graph, corresponds to a physically meaningful scenario. Specially, my question is two-fold:\n- This formulation implicitly assumes situations where the Hamiltonian is unknown but the spectral graph itself can somehow be measured, possibly through experiments. A key question is whether the spectral graph, as defined in this work, represents a physically measurable quantity that can be constructed from experimental data.\n- Also, the authors may elaborate on why predicting the Hamiltonian class represents a meaningful task in physics. \n\nI would be glad to raise my score if the authors can adequately address these concerns."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359165916,"tcdate":1761948105744,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17371/Reviewer_JuyR"],"signatures":["ICLR.cc/2026/Conference/Submission17371/Reviewer_JuyR"],"forum":"YxuKCME576","number":2,"license":"CC BY 4.0","cdate":1761948105744,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17371/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359165916,"domain":"ICLR.cc/2026/Conference","replyto":"YxuKCME576","id":"MCePrSHuXF","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["graph-level learning","spatial network","multigraph","dataset generator","large-scale dataset","condensed matter physics","non-Hermitian physics","topological physics","AI4Science","Toeplitz matrix"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"AI is transforming scientific research by revealing new ways to understand complex physical systems, but its impact remains constrained by the lack of large, high-quality domain-specific datasets. A rich, largely untapped resource lies in non-Hermitian quantum physics, where the energy spectra of crystals form intricate geometries on the complex plane—termed as $\\textit{Hamiltonian spectral graphs}$. Despite their significance as fingerprints for electronic behavior, their systematic study has been intractable due to the reliance on manual extraction. To unlock this potential, we introduce $\\textbf{Poly2Graph}$ (https://github.com/sarinstein-yan/Poly2Graph): a high-performance, open-source pipeline that automates the mapping of 1-D crystal Hamiltonians to spectral graphs. Using this tool, we present $\\textbf{HSG-12M}$ (https://github.com/sarinstein-yan/HSG-12M): a dataset containing 11.6 million static and 5.1 million dynamic Hamiltonian spectral graphs across 1401 characteristic-polynomial classes, distilled from 177 TB of spectral potential data. Crucially, HSG-12M is the first large-scale dataset of $\\textit{spatial multigraphs}$—graphs embedded in a metric space where multiple geometrically distinct trajectories between two nodes are retained as separate edges. This simultaneously addresses a critical gap, as existing graph benchmarks overwhelmingly assume simple, non-spatial edges, discarding vital geometric information. Benchmarks with popular GNNs expose new challenges in learning spatial multi-edges at scale. Beyond its practical utility, we show that spectral graphs serve as universal topological fingerprints of polynomials, vectors, and matrices, forging a new algebra-to-graph link. HSG-12M lays the groundwork for data-driven scientific discovery in condensed matter physics, new opportunities in geometry-aware graph learning and beyond."},"_bibtex":{"value":"@inproceedings{\nyan2026hsgm,\ntitle={{HSG}-12M: A Large-Scale Benchmark of Spatial Multigraphs from the Energy Spectra of Non-Hermitian Crystals},\nauthor={Xianquan Yan and Hakan Akg{\\\"u}n and Kenji Kawaguchi and N. Duane Loh and Ching Hua Lee},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=YxuKCME576}\n}"},"title":{"value":"HSG-12M: A Large-Scale Benchmark of Spatial Multigraphs from the Energy Spectra of Non-Hermitian Crystals"},"pdf":{"value":"/pdf/a3cae5d1b873a4b0f8658cab3767a016c8992d1c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yan|hsg12m_a_largescale_benchmark_of_spatial_multigraphs_from_the_energy_spectra_of_nonhermitian_crystals"},"authorids":{"value":["~Xianquan_Yan1","~Hakan_Akgün1","~Kenji_Kawaguchi1","~N._Duane_Loh1","~Ching_Hua_Lee1"]},"authors":{"value":["Xianquan Yan","Hakan Akgün","Kenji Kawaguchi","N. Duane Loh","Ching Hua Lee"]}},"version":2},{"content":{"summary":{"value":"The paper proposes two models, PI-RGSM and PI-RGSM-K, for groundwater seepage prediction using physics-informed neural networks (PINNs). These models integrate physical constraints into neural networks to enhance prediction accuracy in groundwater flow, reducing dependency on observational data and adapting well to complex seepage conditions. PI-RGSM-K extends the base model by incorporating heterogeneous hydraulic conductivity fields, showing improved adaptability in complex conditions."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. \"However, the ”black box” nature of DNN exhibiting a lack of transparency in their decision-making processes and the significant dependence on extensive training data, limit their use in groundwater research.\" These are general drawbacks of DNN, are there any unique challenges to groundwater modeling?\n2. \"Although significant progress has been made in improving groundwater seepage models using PINNs, current models still heavily depend on the adequacy and quality of observed data and remain sensitive to outliers.\" Are there references supporting this? If these methods have known limitations, why not compare them experimentally?\n3. It's better to define $\\mu$ and $t$ explicitly in Section 2.1 for clarity. \n4. The study uses randomly generated hydraulic conductivities and source/sink terms to enrich the training data. Is random generation commonly accepted in groundwater modeling, and does it effectively capture real-world scenarios?\n5. Do $\\phi$ and $\\varphi$ represent the same meaning in this paper? Their mixed usage leads to confusion.\n6. $H_{EC}$ is not clearly defined, nor is its role in the model explained.\n7. Why use $RES$ in section 2 while use $LOSS$ in section 3?\n8. Section 4.1 states, \"The experiment did not use observations,\" which conflicts with Section 2.2(4), which introduces observed data constraints. In addition, Section 2.2 is confusing since it introduces multiple terms while it looks like most of them ($RES_{BC}, RES_{IC}, RES_{OC}$) are not used in the experiments.\n9. Is there a specific reason for choosing $K=-0.01x+0.8$?\n\nSome minor suggestions:\n1. Use proper LaTeX notation for quotes: \\`\\`example text'' for left and right quotes.\n2. Figures 2 and 3 appear nearly identical. Consider combining them into a single figure that highlights the distinctions between PI-RGSM and PI-RGSM-K for conciseness."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. The paper effectively applies physics-informed neural networks to model groundwater seepage, enhancing model interpretability and physical consistency.\n2. By integrating hard constraints, the models reduce reliance on labeled data, making them suitable for scenarios with limited observational data.\n3. The models achieve promising predictive performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper is poorly organized and lacks clarity, with essential content relegated to the appendix, leaving the main text insufficiently self-contained. See questions below for specific issues.\n2. Although applying PINNs to groundwater seepage is relatively novel within the specific application area, the paper contributes little in terms of new machine learning techniques.\n3. The model relies on manual tuning of hyperparameters (e.g., loss weights, threshold $H_{EC}$), but the paper lacks experiments analyzing the impact of these parameters.\n4. The paper fails to compare the proposed method with existing machine learning models for groundwater seepage."}},"nonreaders":[],"tmdate":1731427887692,"tcdate":1730662160871,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3582/Reviewer_CiS8"],"signatures":["ICLR.cc/2025/Conference/Submission3582/Reviewer_CiS8"],"forum":"GeMWhBIzrk","number":2,"license":"CC BY 4.0","cdate":1730662160871,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3582/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427887692,"domain":"ICLR.cc/2025/Conference","replyto":"GeMWhBIzrk","id":"YaACPtYIRF","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-Informed Neural Networks","Groundwater Prediction","Definite condition","Hard constraints","Self-Supervised"]},"supplementary_material":{"value":"/attachment/3c0ed59cd065eff86a3448d42bfb3942f93c64a3.pdf"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Neural networks, especially deep learning, have achieved revolutionary advances in several domains, including image and speech recognition, with excellent results. However, their reliance on labeled data, lack of interpretability, and inconsistency with physical principles limit their applicability in groundwater seepage prediction and other scientific disciplines. Physics-Informed Neural Networks (PINNs) significantly improve these issues by integrating physical knowledge with neural networks. This study focuses on modeling the groundwater flow field and proposes a physics-informed river-canal groundwater seepage model (PI-RGSM). This model enables self-supervised learning by incorporating hard constraints of boundary and initial conditions, utilizing hydrogeological parameters and boundary conditions as direct inputs, thus diminishing dependence on observable data. Compared to the baseline PINNs, the PI-RGSM adapts to and accurately predicts diverse seepage situations with just one training session, achieving a mean coefficient of determination of 0.978. To further enhance applicability in complex dynamic groundwater seepage situations, we propose PI-RGSM-K, which builds upon PI-RGSM. This model simulates heterogeneous groundwater seepage fields and improves performance in complex seepage environments through parameterized hydraulic conductivity field $K(x,y)$ and fine-adjusted model architecture, attaining a mean coefficient of determination of 0.982. The physics-informed neural network models proposed in this study demonstrate exceptional efficacy in precisely forecasting groundwater seepage behavior."},"_bibtex":{"value":"@misc{\nchen2024groundwater,\ntitle={Groundwater Seepage Modeling in a River-Canal System based on Physics-Informed Neural Networks},\nauthor={Chong Chen and Yifan Li and Zongyu Han and Yixiao Niu and Xiaoyu Zhu and Yaru Xue},\nyear={2024},\nurl={https://openreview.net/forum?id=GeMWhBIzrk}\n}"},"title":{"value":"Groundwater Seepage Modeling in a River-Canal System based on Physics-Informed Neural Networks"},"pdf":{"value":"/pdf/776bbab10735af86de0fa6be660bc22dba65a388.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|groundwater_seepage_modeling_in_a_rivercanal_system_based_on_physicsinformed_neural_networks"},"authorids":{"value":["~Chong_Chen8","~Yifan_Li18","~Zongyu_Han2","~Yixiao_Niu1","~Xiaoyu_Zhu3","~Yaru_Xue1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chong Chen","Yifan Li","Zongyu Han","Yixiao Niu","Xiaoyu Zhu","Yaru Xue"]}},"version":2},{"content":{"summary":{"value":"The authors use pre-trained DNN image models to generate an uncertainty prediction regarding whether an input image is synthetic or real.  The method involves perturbing the model weights and observing a change in the features extracted from an image.  The results suggest that synthetic data features are more affected by model weight perturbations than real data features.  This is demonstrated using three benchmark datasets and DINOv2."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"None"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"The paper is well written and easy to follow.  The methodology is intuitive, and is implementable beyond the studies provided by the authors.  Distinguishing between real and synthetic data is an open question in the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The method present in this paper exclusively relies on access to deep learning models which have not been trained on any generated synthetic data.  Unfortunately, the proliferation of synthetic images means that such models will become harder and harder to find as this area of research progresses.  The method may have a built-in \"expiration date\", in that, future state-of-the-art models will likely be tainted (either knowingly or unknowingly) with generated data.  Similarly, as synthetic data becomes closer to natural data, I would expect the synthetic features to also approximate natural features.  \n\nThe authors directly acknowledge there is no theoretical justification for this method, and list it under future work. I appreciate their honesty and clarity, and agree that the result is very interesting.  However, without a theoretical justification or a more extensive analysis of the differences in natural-synthetic feature representation driving this metric, this work is unlikely to become high impact."}},"nonreaders":[],"tmdate":1731428055394,"tcdate":1730925559354,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13682/Reviewer_ANn9"],"signatures":["ICLR.cc/2025/Conference/Submission13682/Reviewer_ANn9"],"forum":"pIVOSU7TFQ","number":4,"license":"CC BY 4.0","cdate":1730925559354,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13682/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428055394,"domain":"ICLR.cc/2025/Conference","replyto":"pIVOSU7TFQ","id":"SXz7iE8Pvw","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["AI-generated image detection"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this work, we propose a novel approach for detecting AI-generated images by leveraging predictive uncertainty to mitigate misuse and associated risks. The motivation arises from the fundamental assumption regarding the distributional discrepancy between natural and AI-generated images. **The feasibility of distinguishing natural images from AI-generated ones is grounded in the distribution discrepancy between them**. Predictive uncertainty offers an effective approach for capturing distribution shifts, thereby providing insights into detecting AI-generated images. Namely, as the distribution shift between training and testing data increases, model performance typically degrades, often accompanied by increased predictive uncertainty. Therefore, we propose to employ predictive uncertainty to reflect the discrepancies between AI-generated and natural images. In this context, the challenge lies in ensuring that the model has been trained over sufficient natural images to avoid the risk of determining the distribution of natural images as that of generated images. We propose to leverage large-scale pre-trained models to calculate the uncertainty as the score for detecting AI-generated images. Inspired by MC Dropout, we perturb pre-trained models and find that the uncertainty can be captured by perturbing the weights of pre-trained models. This leads to a simple yet effective method for detecting AI-generated images using large-scale vision models: images that induce high uncertainty are identified as AI-generated. Comprehensive experiments across multiple benchmarks demonstrate the effectiveness of our method."},"_bibtex":{"value":"@misc{\nnie2025detecting,\ntitle={Detecting Discrepancies Between Generated and Natural Images Using Uncertainty},\nauthor={Jun Nie and Yonggang Zhang and Tongliang Liu and Yiu-ming Cheung and Bo Han and Xinmei Tian},\nyear={2025},\nurl={https://openreview.net/forum?id=pIVOSU7TFQ}\n}"},"title":{"value":"Detecting Discrepancies Between Generated and Natural Images Using Uncertainty"},"pdf":{"value":"/pdf/0dbf4bfd6d001c48a67cfff48b1beee4e3b5e483.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"nie|detecting_discrepancies_between_generated_and_natural_images_using_uncertainty"},"authorids":{"value":["~Jun_Nie1","~Yonggang_Zhang1","~Tongliang_Liu1","~Yiu-ming_Cheung1","~Bo_Han1","~Xinmei_Tian1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jun Nie","Yonggang Zhang","Tongliang Liu","Yiu-ming Cheung","Bo Han","Xinmei Tian"]}},"version":2},{"content":{"summary":{"value":"The paper introduce Unreal Multi-Agent Playground (UMAP), a 3D simulation environment for multi-agent reinforcement learning, built on Unreal Engine. UMAP addresses existing limitations by offering a flexible platform for complex tasks and research involving diverse agents and teams. Alongside UMAP, they present the Hybrid Multi-Agent Playground (HMAP), compatible with various algorithms. The paper showcases UMAP's capabilities through tasks, evaluations of MARL algorithms, and a sim-to-real experiment."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. Though it is a general platform, what's the special strength of the new environment built on this platform for evaluating existing MARL algorithms?\n2. How much effort should a new developer paid for introducing his own environment/problem into this platform?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Addressing a Significant Gap: The paper identifies and addresses a critical limitation in MARL research—the lack of simulation environments that are both realistic and capable of modeling complex, large-scale multi-agent interactions. By providing a physics-based environment, UMAP brings simulations closer to real-world scenarios, which is essential for the development and evaluation of practical MARL algorithms.\n2. High Extensibility and Customization: UMAP's hierarchical, modular architecture allows users to easily customize tasks at various levels, from high-level configurations to low-level implementations. This flexibility enables researchers to create a diverse range of scenarios tailored to specific research questions or application domains.\n3.Support for Diverse and Complex Tasks: The inclusion of scenarios featuring heterogeneous agents, large-scale populations, multiple teams, and sparse rewards demonstrates UMAP's ability to model a wide spectrum of complex MARL problems. This diversity is valuable for testing the robustness and generalization capabilities of MARL algorithms.\n4.Open-Source Contribution: By releasing UMAP and HMAP as open-source projects, the authors contribute valuable tools to the MARL community. This openness encourages collaboration, reproducibility, and further development by other researchers."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Depth of Experimental Analysis: The experimental results primarily focus on win rates and rewards. A deeper analysis of algorithm behaviors would provide more comprehensive insights into the challenges posed by the new environment.\n2. Comparison with Existing Environments: Although the paper mentions limitations of current MARL environments, a more detailed comparison highlighting specific features and performance metrics would better contextualize UMAP's advantages. Including benchmarks or case studies demonstrating UMAP's superiority could strengthen the argument.\n3. Accessibility and Cost of Use: While UMAP is described as user-friendly, potential barriers such as the need for familiarity with Unreal Engine and the complexity of setting up the environment could limit adoption. Providing comprehensive tutorials, documentation, and support could mitigate this issue. In addition, Physics-based environments, especially those built on game engines like Unreal Engine, can have significant computational overheads. Discussion about the performance implications, resource requirements, and optimizations would be beneficial for users with limited computing resources."}},"nonreaders":[],"tmdate":1732468951297,"tcdate":1730726446853,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6491/Reviewer_1Wxb"],"signatures":["ICLR.cc/2025/Conference/Submission6491/Reviewer_1Wxb"],"forum":"uYzJvP8HGl","number":6,"license":"CC BY 4.0","cdate":1730726446853,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6491/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732468951297,"domain":"ICLR.cc/2025/Conference","replyto":"uYzJvP8HGl","id":"pUgmVB4aBi","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-agent reinforcement learning","simulation environment","reinforcement learning"]},"supplementary_material":{"value":"/attachment/aa9431c89d7f69954578c157fd242f682481229d.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Existing simulation environments in the field of multi-agent reinforcement learning (MARL) either lack authenticity or complexity. The data generated by these environments significantly deviate from the requirements of the real world, hindering the practical application of MARL. To address this issue, we propose Unreal Multi-Agent Playground (UMAP), a highly extensible, physics-based 3D simulation environment implemented on the Unreal Engine. UMAP is user-friendly in terms of deployment, modification, and visualization, and all its components are open-sourced. Based on UMAP, we design a series of MARL tasks featuring heterogeneous agents, large-scale agents, multiple teams, and sparse team rewards.\nWe also develop an experimental framework compatible with algorithms ranging from \nrule-based to MARL-based provided by third-party frameworks. In the experimental section, we utilize the designed tasks to test several state-of-the-art algorithms. Additionally, We also conduct a physical experiment to demonstrate UMAP's potential in sim-to-real applications, which is a significant advantage due to the high extensibility and authenticity of UMAP. We believe UMAP can play an important role in the MARL field by evaluating existing algorithms and helping them apply to real-world scenarios, thus advancing the field of MARL."},"_bibtex":{"value":"@misc{\nhu2025umap,\ntitle={{UMAP}: A Highly Extensible and Physics-Based Simulation Environment for Multi-agent Reinforcement Learning},\nauthor={Tianyi Hu and Qingxu Fu and Zhiqiang Pu and Yuan Wang and Tenghai Qiu},\nyear={2025},\nurl={https://openreview.net/forum?id=uYzJvP8HGl}\n}"},"title":{"value":"UMAP: A Highly Extensible and Physics-Based Simulation Environment for Multi-agent Reinforcement Learning"},"pdf":{"value":"/pdf/f6ed3a0d7479f348bce2f242836bc24eb24fd4d9.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"hu|umap_a_highly_extensible_and_physicsbased_simulation_environment_for_multiagent_reinforcement_learning"},"authorids":{"value":["~Tianyi_Hu1","~Qingxu_Fu1","~Zhiqiang_Pu1","~Yuan_Wang34","~Tenghai_Qiu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Tianyi Hu","Qingxu Fu","Zhiqiang Pu","Yuan Wang","Tenghai Qiu"]}},"version":2},{"content":{"summary":{"value":"This paper investigated how physics-based image-quality degradations affect the robustness of deep learning models for medical imaging. The authors construct synthetic corruption pipelines for chest radiographs (JSRT dataset) and dermoscopic images (ISIC dataset), grounded in X-ray and dermoscopy physics (acquisition geometry, anti-scatter grid, beam energy, collimation, detector performance, focal spot size, mAs; and blur, focal plane, resolution, color reproduction for dermoscopy).\nThey then:\nTrain segmentation models (U-Net, Swin-Unet, SAM) and classification models (ResNet-18, ConvNeXt-tiny, ViT-b16) on clean, traditionally augmented, and physics-augmented images.\nEvaluate robustness under various corruption types and severity levels, using Dice for segmentation and accuracy for classification (Figs. 2–5).\nIntroduce an auxiliary task where classifiers predict corruption severity and type, reporting class-wise accuracies on JSRT (Table 1).\nResults show that incorporating physics-based corruptions during training generally improves performance on corrupted test images compared to training on clean data, and often outperforms simple augmentations (blur/noise/brightness), especially for U-Net, SAM, and ResNet-18."},"justification_of_the_preliminary_rating":{"value":"This paper addresses an important question: how physics-informed image-quality degradations affect the robustness of medical imaging models. The corruption pipelines are carefully motivated from X-ray and dermoscopy physics and mapped to explicit image-space operators, and the experiments span multiple architectures, tasks, and modalities. The results generally support the claim that training with realistic, physics-based degradations can improve model robustness and that models can predict corruption type and severity with reasonable accuracy.\n\nHowever, the study is limited by the small datasets (JSRT, ISIC 2016), the lack of validation against truly low-quality clinical images, and relatively shallow robustness analyses that rely on average Dice/accuracy without statistical testing or clear performance-vs-severity curves. Gains over traditional augmentation are present but sometimes modest or inconsistent, particularly for classification and for transformer-based models."},"confidentiality_llm_acknowledgment":{"value":"Yes"},"strengths":{"value":"Well-motivated, physics-grounded corruptions. The X-ray corruptions are systematically derived from imaging physics.\n\nTraining segmentation and classification models on clean, traditionally augmented, and physics-augmented images makes sense.\n\nThe paper covers two important modalities and tasks: JSRT chest X-rays (organ segmentation + nodule classification) and ISIC 2016 dermoscopy.\n\nThe corruption severity/type prediction task is an interesting addition; ResNet-18 achieves strong accuracy across most corruption types and levels.\nEvaluating robustness across various corruption types and severity levels, using Dice for segmentation and accuracy for classification, is important and relevant."},"weaknesses":{"value":"Limited dataset scale and diversity.\nJSRT (247 X-rays) and ISIC 2016 (1,279 images) are small, older benchmarks. There is no validation on larger or multi-centre datasets, limiting claims about generalizability to contemporary clinical data.\nNo direct validation against real low-quality acquisitions.\nRobustness analysis is somewhat shallow. Results are mostly average Dice/accuracy bars; there is no systematic analysis of performance vs severity level with confidence intervals, no robustness metrics (e.g., relative drop vs clean, worst-case over severities), and no statistical tests on model comparisons.\nMissing experimental details and ablations."},"confidence":{"value":5},"detailed_comments":{"value":"Clarify corruption parameterization and realism.\nIt would help to provide a table summarizing parameter ranges for each severity level for all corruptions and briefly explain how they map to realistic acquisition parameter deviations (e.g., approximate kVp shifts, grid misalignment angles, detector MTF drops) \nFor both JSRT and ISIC, show curves of performance vs corruption severity for at least one corruption type and model, with error bars over multiple seeds. This would better support statements like “performance generally increases as corrupted data is added and robustness improves.”\nAdd at least simple paired tests (e.g., Wilcoxon) to determine whether improvements from physics-based augmentation over traditional augmentation are statistically significant across test images.\nExactly how are physics-based corruptions mixed with clean data: per mini-batch, fixed percentage per epoch, or separate epochs? Are multiple corruption types sampled per image, or one type per copy?\nConsider reporting per-class metrics (sensitivity/specificity, AUC). Currently, accuracy can hide poor malignant performance. A brief experiment with class-balancing loss or sampling would be very informative."},"questions_to_address_in_the_rebuttal":{"value":"Exactly how are physics-based corruptions mixed with clean data: per mini-batch, fixed percentage per epoch, or separate epochs? Are multiple corruption types sampled per image, or one type per copy?\n\nCan you provide more quantitative evidence (e.g., relative performance drops or robustness metrics) that physics-based augmentation gives consistent and statistically significant robustness gains over traditional augmentation across severities and seeds?\n\nFor ISIC classification, can you report malignant-class performance (e.g., sensitivity/AUC) under clean and corrupted conditions, with and without physics-based augmentation? This is central for clinical relevance.\n\nPlease clarify the mixing strategy of corrupted vs clean images and how the intensities/severities of traditional vs physics-based augmentations were chosen so that comparisons are fair."},"preliminary_rating":{"value":3}},"parentInvitations":"MIDL.io/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1771530041529,"tcdate":1768003370669,"writers":["MIDL.io/2026/Conference","MIDL.io/2026/Conference/Submission214/Reviewer_mL8z"],"signatures":["MIDL.io/2026/Conference/Submission214/Reviewer_mL8z"],"forum":"fPXKcRF9fb","number":3,"license":"CC BY 4.0","cdate":1768003370669,"readers":["everyone"],"invitations":["MIDL.io/2026/Conference/Submission214/-/Official_Review","MIDL.io/2026/Conference/-/Edit"],"mdate":1771530041529,"domain":"MIDL.io/2026/Conference","replyto":"fPXKcRF9fb","id":"GKYTNlYY46","forumContent":{},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2505.23656v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"zhang|videorepa_learning_physics_for_video_generation_through_relational_alignment_with_foundation_models"},"html":{"value":"https://doi.org/10.48550/arXiv.2505.23656"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2505-23656,\n  publtype={informal},\n  author={Xiangdong Zhang and Jiaqi Liao and Shaofeng Zhang and Fanqing Meng and Xiangpeng Wan and Junchi Yan and Yu Cheng},\n  title={VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models},\n  year={2025},\n  month={May},\n  cdate={1746057600000},\n  journal={CoRR},\n  volume={abs/2505.23656},\n  url={https://doi.org/10.48550/arXiv.2505.23656}\n}\n"},"abstract":{"value":"Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability to accurately understand physics. We found that while the representations within T2V models possess some capacity for physics understanding, they lag significantly behind those from recent video self-supervised learning methods. To this end, we propose a novel framework called VideoREPA, which distills physics understanding capability from video understanding foundation models into T2V models by aligning token-level relations. This closes the physics understanding gap and enable more physics-plausible generation. Specifically, we introduce the Token Relation Distillation (TRD) loss, leveraging spatio-temporal alignment to provide soft guidance suitable for finetuning powerful pre-trained T2V models, a critical departure from prior representation alignment (REPA) methods. To our knowledge, VideoREPA is the first REPA method designed for finetuning T2V models and specifically for injecting physical knowledge. Empirical evaluations show that VideoREPA substantially enhances the physics commonsense of baseline method, CogVideoX, achieving significant improvement on relevant benchmarks and demonstrating a strong capacity for generating videos consistent with intuitive physics. More video results are available at https://videorepa.github.io/."},"title":{"value":"VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models"},"authors":{"value":[{"fullname":"Xiangdong Zhang","username":""},{"fullname":"Jiaqi Liao","username":"~Jiaqi_Liao2"},{"fullname":"Shaofeng Zhang","username":""},{"fullname":"Fanqing Meng","username":""},{"fullname":"Xiangpeng Wan","username":""},{"fullname":"Junchi Yan","username":""},{"fullname":"Yu Cheng","username":""}]}},"tmdate":1784955577039,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2505-23656"],"tcdate":1784955571209,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Jiaqi_Liao2"],"forum":"JmlNFghXw4","license":"CC BY-SA 4.0","number":103519,"cdate":1746057600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784955577039,"domain":"OpenReview.net/Public_Article","id":"JmlNFghXw4","version":2},{"content":{"summary":{"value":"This paper examines whether large language models can acquire general reasoning abilities solely from synthetic data, without relying on real-world knowledge.\nUsing reinforcement learning on fully artificial datasets such as PhantomWiki and GSM, the authors show significant performance improvements on real-world multi-hop QA benchmarks.\nThe results suggest that reasoning skills, such as knowledge composition, can transfer across domains, highlighting synthetic data as a scalable alternative to human-annotated reasoning datasets."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"None"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The study demonstrates that large language models can acquire generalizable reasoning skills purely from synthetic, knowledge-free data. It provides empirical evidence that these synthetic reasoning abilities transfer to real-world multi-hop QA tasks, achieving substantial performance gains. The approach offers a scalable and cost-effective framework for improving reasoning through verifiable, automatically generated training data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The experiments are limited to multi-hop QA. Even though the training data come from a synthetic world, the fact that performance improves on other multi-hop QA benchmarks is not particularly surprising.\n- The applicability of the approach to grammatically or semantically complex real-world texts remains unknown."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941724842,"tcdate":1761862103349,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21366/Reviewer_3JJe"],"signatures":["ICLR.cc/2026/Conference/Submission21366/Reviewer_3JJe"],"forum":"38nYZ5QBui","number":3,"license":"CC BY 4.0","cdate":1761862103349,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21366/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941724842,"domain":"ICLR.cc/2026/Conference","replyto":"38nYZ5QBui","id":"nFCyP6L0vI","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"RL fine-tuning LLMs on synthetic data improves real-world multi-hop reasoning by teaching knowledge composition skills"},"keywords":{"value":["multi-hop reasoning","large language models","reinforcement learning","synthetic data"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Reinforcement Learning (RL) has been shown to significantly boost reasoning capabilities of large language models (LLMs) in math, coding, and multi-hop reasoning tasks.\nHowever, RL fine-tuning requires abundant high-quality verifiable data, often sourced from human annotations, generated from frontier LLMs, or scored by LLM-based verifiers.\nAll three have considerable limitations: human-annotated datasets are small and expensive to curate, LLM-generated data is hallucination-prone and costly, and LLM-based verifiers are inaccurate and slow.\nIn this work, we investigate a cheaper alternative: RL fine-tuning on _rule-generated synthetic data_ for multi-hop reasoning tasks.\nWe discover that LLMs fine-tuned on synthetic data perform significantly better on popular real-world question-answering benchmarks, despite the synthetic data containing only fictional knowledge.\nOn stratifying performance by question difficulty, we find that synthetic data teaches LLMs to _compose knowledge_---a fundamental and generalizable reasoning skill.\nOur work highlights rule-generated synthetic reasoning data as a free and scalable resource to improve LLM reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\nkabra2026learning,\ntitle={Learning from Synthetic Data Improves Multi-hop Reasoning},\nauthor={Anmol Kabra and Yilun Yin and Albert Gong and Kamil{\\.{e}} Stankevi{\\v{c}}i{\\={u}}t{\\.{e}} and Dongyoung Go and Johann Lee and Katie Z Luo and Carla P Gomes and Kilian Q Weinberger},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=38nYZ5QBui}\n}"},"title":{"value":"Learning from Synthetic Data Improves Multi-hop Reasoning"},"pdf":{"value":"/pdf/a95ce444c53d4630d8b208129731802b6a16a9a1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"kabra|learning_from_synthetic_data_improves_multihop_reasoning"},"authorids":{"value":["~Anmol_Kabra1","~Yilun_Yin1","~Albert_Gong1","~Kamilė_Stankevičiūtė1","~Dongyoung_Go1","~Johann_Lee1","~Katie_Z_Luo1","~Carla_P_Gomes1","~Kilian_Q_Weinberger1"]},"authors":{"value":["Anmol Kabra","Yilun Yin","Albert Gong","Kamilė Stankevičiūtė","Dongyoung Go","Johann Lee","Katie Z Luo","Carla P Gomes","Kilian Q Weinberger"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (2) 2023"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-53499-7_17.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"asgari|mosaic_benchmark_networks_modular_link_streams_for_testing_dynamic_community_detection_algorithms"},"html":{"value":"https://doi.org/10.1007/978-3-031-53499-7_17"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/AsgariCB23,\n  author={Yasaman Asgari and Rémy Cazabet and Pierre Borgnat},\n  title={Mosaic Benchmark Networks: Modular Link Streams for Testing Dynamic Community Detection Algorithms},\n  year={2023},\n  cdate={1672531200000},\n  pages={209-222},\n  url={https://doi.org/10.1007/978-3-031-53499-7_17},\n  booktitle={COMPLEX NETWORKS (2)},\n  crossref={conf/complexnetworks/2023-2}\n}\n"},"abstract":{"value":"Community structure is a critical feature of real networks, providing insights into nodes’ internal organization. Nowadays, with the availability of highly detailed temporal networks such as link streams, studying community structures becomes more complex due to increased data precision and time sensitivity. Despite numerous algorithms developed in the past decade for dynamic community discovery, assessing their performance on link streams remains a challenge. Synthetic benchmark graphs are a well-accepted approach for evaluating static community detection algorithms. Additionally, there have been some proposals for slowly evolving communities in low-resolution temporal networks like snapshots. Nevertheless, this approach is not yet suitable for link streams. To bridge this gap, we introduce a novel framework that generates synthetic modular link streams with predefined communities. Subsequently, we evaluate established dynamic community detection methods to uncover limitations that may not be evident in snapshots with slowly evolving communities. While no method emerges as a clear winner, we observe notable differences among them."},"title":{"value":"Mosaic Benchmark Networks: Modular Link Streams for Testing Dynamic Community Detection Algorithms"},"authors":{"value":[{"fullname":"Yasaman Asgari","username":""},{"fullname":"Rémy Cazabet","username":""},{"fullname":"Pierre Borgnat","username":"~Pierre_Borgnat1"}]}},"tmdate":1787910542176,"pdate":1703980800000,"externalIds":["dblp:conf/complexnetworks/AsgariCB23"],"tcdate":1787910537337,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Pierre_Borgnat1"],"forum":"NcO3ea13MV","license":"CC BY-SA 4.0","number":144332,"cdate":1672531200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1787910542176,"domain":"OpenReview.net/Public_Article","id":"NcO3ea13MV","version":2},{"content":{"summary":{"value":"The paper introduces a physics-informed deep learning framework for climate and weather prediction. Unlike purely data-driven models (e.g., ClimODE, ClimaX), PA-TFNP embeds rotation-equivariant tensor-field neural operators directly on the sphere (to respect Earth’s geometry), incorporates spherical-transform-based gradient operators and physically consistent boundary treatments, and adds diffusion and momentum correction terms derived from the atmospheric primitive equations for long-term stability. Across global, regional, and monthly prediction tasks on ERA5 data, PA-TFNP achieves up to 78.9% RMSE reduction over ClimODE with 10× fewer parameters."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How is the spatially varying diffusion coefficient α(x) learned? Is it through per grid point or a continuous modeling?\n\n2. How sensitive is performance to the blending schedule τ₀? Could it be dynamically adapted instead of fixed exponential decay?\n\n3. How does PA-TFNP perform on 0.25° or sub-hourly resolutions? Does rotation-equivariance introduce computational overhead?\n\n4. Have you evaluated transfer performance to other reanalysis datasets or regional NWP outputs?\n\n5. Can tensor features or learned diffusion fields be visualized to reveal interpretable atmospheric patterns (e.g., jet streams, vortices)? Can you show some examples?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The idea is reasonable. Combining rotation-equivariant tensor fields, spherical gradients, and physics-based diffusion is a good blend of geometry, physics, and learning. The novelty is very incremental, concatenating ideas from many previous works, but reasonable.\n\n2. Demonstrates consistent and substantial accuracy gains across multiple spatiotemporal scales (hourly → monthly, regional → global). Clear modular structure (TFNP → PA-TFNP), with derivations and ablations validating each design choice.\n\n3. Addresses the pressing need for physically consistent ML weather models, aligning well with current research trends (e.g., GraphCast, Aurora)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. No code revealed for reproducibility check. The implementation details and reproducibility remain limited. Important training hyperparameters, computational costs at large scale, and code availability are not discussed in depth. These omissions make it challenging for others to replicate or extend the proposed approach.\n\n2. While the model introduces diffusion and drag terms inspired by the atmospheric primitive equations, these additions are applied uniformly across all variables. This uniform treatment overlooks the fundamental physical differences between quantities such as temperature, wind, and geopotential height. Each obeys distinct physical laws and should ideally have customized formulations rather than a single shared diffusion process.\n\n3. The benefits of rotation-equivariance, though central to the model’s innovation, appear to be confined largely to global-scale forecasting, and have been used by many previous works already. The only physics-aware part is according to the advection structure design of the PDE, which is also not new. ClimODE did the same already. **This is the main reason for me recommending paper reject**. The authors should try rebuttal on this first before everything else.\n\n4. While comparisons to ClimODE and ClimaX are thorough, the paper omits quantitative benchmarks against more recent large-scale foundation models like Aurora or GraphCast, both of which represent strong baselines in this field. Including such comparisons would strengthen the empirical claims of superiority. ClimODE is well-known to not perform well enough compared to large FMs, and its core contribution is more on guiding people to use physics and AI inherently together by methods other than just tuning loss functions. Simply showing an advantage over ClimODE is not entirely surprising or impactful, unless novel techniques/paradigms are proposed.\n\n5. Another limitation lies in the evaluation metrics. The study relies primarily on RMSE, without assessing physically meaningful diagnostics such as conservation of mass, energy, or vorticity, nor does it provide uncertainty estimates. This omission makes it difficult to judge whether the model’s forecasts are physically consistent, beyond being numerically accurate."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926125293,"tcdate":1761004493262,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15906/Reviewer_bGma"],"signatures":["ICLR.cc/2026/Conference/Submission15906/Reviewer_bGma"],"forum":"beV5wMTRIq","number":1,"license":"CC BY 4.0","cdate":1761004493262,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15906/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926125293,"domain":"ICLR.cc/2026/Conference","replyto":"beV5wMTRIq","id":"f8YgLMQOQz","forumContent":{"TLDR":{"value":"We introduce PA-TFNP, a physics-aware tensor-field neural PDE framework for weather and climate forecasting."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Weather Prediction","Neural ODE","Physics-Informed Machine Learning","Tensor Field Neural","Partial Differential Equations"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Climate and weather prediction has traditionally relied on computationally demanding numerical simulations grounded in atmospheric physics, yet deep-learning approaches are emerging as transformative alternatives. Existing methods, however, are often purely data-driven and physics-agnostic, overlooking essential physical principles and struggling to generalize. To address these challenges, we present the Physics-Aware Tensor Field Neural PDE (PA-TFNP), a forecasting framework that embeds rotation-equivariant tensor-field neural operators directly on the sphere, couples them with a numerically rigorous gradient operator based on spherical transforms and physically consistent boundary treatment, and augments the learned dynamics with diffusion terms derived from the atmospheric primitive equations. These innovations enable our model to achieve superior performance through strict physical fidelity and efficient learning. The proposed PA-TFNP achieves state-of-the-art performance in global and regional weather prediction, outperforming ClimODE by 78.92% on global hourly data with a comparable number of parameters."},"_bibtex":{"value":"@misc{\ncho2025physicsaware,\ntitle={Physics-Aware Tensor Field Neural {PDE} for Climate and Weather Prediction},\nauthor={Namkyeong Cho and Sung Woong Cho and Youngjoon Hong and Hyung Ju Hwang and Jae Yong Lee and Hwijae Son},\nyear={2025},\nurl={https://openreview.net/forum?id=beV5wMTRIq}\n}"},"title":{"value":"Physics-Aware Tensor Field Neural PDE for Climate and Weather Prediction"},"pdf":{"value":"/pdf/8b2fbbffd8e48eb20194e8b54fad719f70c564d5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"cho|physicsaware_tensor_field_neural_pde_for_climate_and_weather_prediction"},"authorids":{"value":["~Namkyeong_Cho1","~Sung_Woong_Cho1","~Youngjoon_Hong1","~Hyung_Ju_Hwang1","~Jae_Yong_Lee2","~Hwijae_Son1"]},"authors":{"value":["Namkyeong Cho","Sung Woong Cho","Youngjoon Hong","Hyung Ju Hwang","Jae Yong Lee","Hwijae Son"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the challenge of automated policy compliance assessment, a task that requires complex reasoning aligned with human-defined rules. The authors note that acquiring gold-standard expert reasoning traces for this task is prohibitively expensive. They propose a method called POLICY REASONING TRACES (PRT), which uses a powerful \"expert\" LLM to synthetically generate step-by-step reasoning traces, given only a case and its final compliance verdict (e.g., COMPLIANT/NONCOMPLIANT). These synthetically generated PRTs are then used as a \"reasoning bridge\"—either as few-shot in-context examples or as data for supervised fine-tuning—to improve the performance of a \"learner\" model. The authors demonstrate that this method improves accuracy and policy clause citation on the HIPAA and GDPR policy datasets."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. PRT Quality Validation: Was any human evaluation conducted to measure the quality, faithfulness, and factual correctness of the generated PRTs? How do we know the \"expert\" model's reasoning for a verdict is the correct reasoning, and not just a plausible hallucination?\n\n2. Forced Justification vs. Organic Reasoning: How does this method compare to simply prompting the expert model with standard Chain-of-Thought (e.g., \"Let's think step-by-step\") and having it generate both the reasoning and the verdict? By providing the gold verdict, aren't you forcing the model to justify a conclusion, which encourages rationalization over genuine reasoning?\n\n3. ModelSpec Performance Drop: Could you elaborate on the performance degradation for OpenAI models on ModelSpec? Does this not imply that PRTs are only beneficial as a crutch for weaker models, and may actually be harmful (by introducing noise) when a model's existing reasoning capability is already strong?\n\n4. Generalist vs. Specialist PRTs: Why do you hypothesize that the Generalist model's PRTs outperformed the Specialist (legal) model's PRTs? This seems to contradict the premise of using domain-specific expertise."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"No"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Important Problem Domain: The paper tackles a practical and high-stakes problem. Automating policy compliance has clear real-world applications, and the challenge of aligning LLM reasoning with complex, human-authored policy documents is a significant research barrier.\n\n2. Intuitive Core Idea: The central idea of using a more capable model to generate pseudo-expert reasoning, thereby avoiding the bottleneck of human expert annotation, is intuitive and a common-sense approach to data augmentation.\n\n3. Strong Empirical Gains on Some Datasets: The method demonstrates clear and significant performance boosts over baseline models, particularly for open-weight models on the HIPAA dataset. The claim of achieving a new state-of-the-art on the GDPR dataset is also a notable empirical contribution."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Critical Unaddressed Risk: PRT Quality and Faithfulness: The entire method hinges on the assumption that the \"expert\" LLM generates high-quality, faithful reasoning traces. The paper provides no substantial validation of this. The expert model is given the gold verdict and asked to produce reasoning for it. This setup strongly encourages post-hoc rationalization, where the model may generate a plausible-sounding but incorrect reasoning path to justify the given answer. If the PRTs are flawed, unfaithful, or contain factual errors (e.g., wrong clause citations), the learner model is simply being trained on high-quality noise. This \"garbage in, garbage out\" risk is a fundamental methodological flaw that is not adequately addressed.\n\n2. Questionable Novelty: The technique of using a strong model to generate synthetic data (especially reasoning steps or chains-of-thought) to train a weaker model is a well-established paradigm in NLP (e.g., data distillation, synthetic CoT generation). The paper fails to clearly articulate what makes \"PRT\" a novel method beyond applying this existing technique to the specific domain of policy compliance.\n\n3. Confusing/Contradictory Results: The results on the ModelSpec dataset are concerning. The authors note that for models already optimized for this policy (OpenAI models), adding PRTs decreases performance. This finding, which is somewhat downplayed, severely undermines the method's generalizability. It suggests that these \"expert-generated\" PRTs may be of lower quality than the model's own internal reasoning, further supporting the concern in Weakness 1.\n\n4. Superficial Analysis of PRT Sources: The paper generates PRTs from a \"Generalist\" (DeepSeek-R1) and \"Specialist\" (SAULLM) model but then states that the Generalist PRTs were \"majority advantage\" and used for subsequent experiments. This is counter-intuitive (why would a specialized legal model be worse?) and this finding is not explored. This feels like a missed opportunity for analysis and weakens the paper's contribution."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921068478,"tcdate":1761799048408,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9485/Reviewer_TNXz"],"signatures":["ICLR.cc/2026/Conference/Submission9485/Reviewer_TNXz"],"forum":"QgEDWbZQ6V","number":1,"license":"CC BY 4.0","cdate":1761799048408,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9485/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921068478,"domain":"ICLR.cc/2026/Conference","replyto":"QgEDWbZQ6V","id":"IWCg8wr4O7","forumContent":{"TLDR":{"value":"We propose Policy Reasoning Traces, a form of pseudo-expert reasoning imitation derived from frontier reasoning models to improve any off-the-shelf model's policy compliance capabilities."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["policy compliance","reasoning trace","chain-of-thought","large language model"]},"supplementary_material":{"value":"/attachment/3337a65ce66bb222dd05b1cb510d2a004ffe13b4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Policy compliance assessment is a fundamental task of evaluating whether an input case strictly complies with a set of human-defined rules, more generally known as *policies*. In practice, human experts follow a systematic, step-by-step process to identify violations with respect to specific stipulations outlined in the policy.  However, such documentation of gold-standard, expert-level reasoning processes is costly to acquire. In this paper, we introduce Policy Reasoning Traces (PRT), a form of specialized generated reasoning chains that serve as a *reasoning bridge* to improve an LLM's policy compliance assessment capabilities. Our empirical evaluations demonstrate that the use of PRTs for both inference-time and training-time scenarios significantly enhances the performance of open-weight and commercial models, setting a new state-of-the-art for HIPAA and GDPR policies. Beyond accuracy gains, we also highlight how PRTs can improve an LLM's ability to accurately cite policy clauses, as well as influence compliance decisions through their high utilization from the raw chains-of-thought."},"_bibtex":{"value":"@misc{\nimperial2026scaling,\ntitle={Scaling Policy Compliance Assessment in Language Models using Policy Reasoning Traces},\nauthor={Joseph Marvin Imperial and Harish Tayyar Madabushi},\nyear={2026},\nurl={https://openreview.net/forum?id=QgEDWbZQ6V}\n}"},"title":{"value":"Scaling Policy Compliance Assessment in Language Models using Policy Reasoning Traces"},"pdf":{"value":"/pdf/106850ab9b225cef6318559aa6eea0e3d8f97a4e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"imperial|scaling_policy_compliance_assessment_in_language_models_using_policy_reasoning_traces"},"authorids":{"value":["~Joseph_Marvin_Imperial1","~Harish_Tayyar_Madabushi1"]},"authors":{"value":["Joseph Marvin Imperial","Harish Tayyar Madabushi"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a climate model emulator that first resolves the advection equation, before refining the trajectory with a Neural SDE. The Neural SDE allows for accurate forecasting as well as uncertainty quantification, while incorporating the advection equation improves the skill of the model. \nThe model is evaluated on two datasets: ClimateSet, a dataset containing TAS, PR, and forcings data from monthly-resolution climate models, and FV3GFS, which represents 7 levels of temperatures and radiative fluxes. It outperforms competing models in TAS RMSE. \n\nOverall, the paper shows an interesting model, with a physics-backbone, but I am not sure that the chosen physics backbone is suited for the task of climate modelling and the limited evaluation does not provide enough evidence that it is."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Was there any hyperparameter search that was done when running baselines on the FV3GFS dataset? Or did you reuse the parameters from ClimateSet?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":3},"strengths":{"value":"The idea of combining a physics-guided backbone with a stochastic residual corrector is a very compelling idea, that could be very useful in the field of climate modelling as uncertainty representation + physical consistency are two main goals. Previous research had proposed to use a physics-guided backbone with a deterministic residual corrector for weather modelling (Kochkov et al., 2023). \n\nResolving the advection process with FDM makes sense, and the parameterization of the Neural SDE + Uncertainty quantification is sound."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"First, the advection diffusion model describes the evolution of gases in the atmosphere over weather timescales (i.e. high temporal resolution), but not over monthly timescale, where the dynamics of sea-surface temperature are mostly driven by the ocean. The diffusion of forcings at a monthly temporal resolution are not at all described by the advection diffusion model, but is rather dependent on the boundary conditions. Similarly, the climate-velocity index or \"local migration speed\" was originally defined for tracking climate change i.e. estimating the velocity between two 30-years periods (Gaponenko et al., 2022). Computing it from monthly data does not make much sense as it will be largely dependent on the initial conditions (it might depend on the season, for example). \n    \nAlthough the idea of using a physics-backbone for climate model emulation is compelling, I don't think that the chosen backbone (advection equation) is suited for the task. Such a model (on monthly data) would probably be better suited for modelling ocean dynamics. \n\nThe authors empirically show that the RMSE is lower when using the physics-backbone. However, the evaluation is very limited and this is not enough for proving that the advection equation truly represents the climate dynamics.   \n\nFirst, the result in figure 3 does not show statistical significance. I recommend training and evaluating the models with multiple seeds + initial conditions to get error bars and statistical significance. This could also be used to evaluate CRPS for competing models (as is traditionnally done with climate ensembles). \n\nSecond, when using the advection equation, the model is informed of the \"initial velocity\", which might lead to better MSE in the next months prediction only, and could lead to the results shown in figure 3. Using metrics that reflect long-term vs short-term prediction is necessary here. \n\nThird, RMSE is not by itself enough to verify whether a climate emulator performs well. It favors smooth models (double-penalty of MSE) but more importantly, the goal of climate modelling is to capture climate trends, as the exact temperature on a specific month is not predictable. Thus, I would recommend reporting statistics like the mean + std. dev., or the power spectral density of the climate indices as done in [1], or the time-dependent area-weighted global mean and the area-weighted global mean bias and RMSE of time-mean fields as done in [2]. \n\nFourth, the models are not evaluated on precipitation, although there are only two output variables in climateset, temperature and precipitation. No performance on precipitation is reported, and the second dataset contains temperature at different levels. I understand that the goal is to emulate TAS, as a response to forcings, but emulating the interactions between variables is key for this task, and understanding the model's performance on precipitation, and whether or not the advection equation helps is very important. \n\nThe model claims to perform uncertainty quantification. CRPS is reported but not compared. Also, no maps of the learned uncertainty is shown. It is also not clear if the uncertainty is well propagated during the autoregressive process. As mentioned above, baselines could be trained with multiple seeds and initial conditions to compute CRPS. \n\nThe limited evaluation prevents from understanding the importance of the two model components, whether PACER actually outperforms baselines, and whether the choice of advection equation as a physics-backbone is supported by empirical evidence.  \n\nFinally, the presentation of the paper could be improved. The zero-shot emulation section is hard to grasp (\"We report the RMSE and CRPS for TAS (surface air temperature) pre-trained on different climate models\"; no \"pretrained climate model\" column in table 3). Figure 2 illustrates \"stronger fidelity\", but does not show any comparison, and labels are small and blurry. In section 4 and Figure 1, it is not clear what u or x are (x,y are coordinates in the left part, but x is then the variable of interest in the right part of Figure 1?)\n\n[1] Hickman et al., Causal Climate Emulation with Bayesian Filtering, 2025\n\n[2] Watt-Meyer et al., ACE: A fast, skillful learned global atmospheric model for climate prediction, 2023"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927023777,"tcdate":1761669001235,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17000/Reviewer_Dh43"],"signatures":["ICLR.cc/2026/Conference/Submission17000/Reviewer_Dh43"],"forum":"RhW8yXgIxY","number":2,"license":"CC BY 4.0","cdate":1761669001235,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17000/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927023777,"domain":"ICLR.cc/2026/Conference","replyto":"RhW8yXgIxY","id":"z5tu5SvUo8","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"A 10-year auto-regressive, advection informed, uncertainty aware climate emulator trained on NLL objective."},"keywords":{"value":["Neural SDE","Uncertainty","Physics Informed Network","Climate Emulator"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics based numerical climate models serve as critical tools for evaluating the effects of climate change and projecting future climate scenarios. However, the reliance on numerical simulations of physical equations renders them computationally intensive and inefficient. While deep learning methodologies have made significant progress in weather forecasting, they are still unstable for longer roll-out climate emulation task. Here, we propose PACER, a relatively lightweight 2.1M parameter Physics Informed Uncertainty Aware Climate EmulatoR. PACER is trained across is trained across varying spatial resolutions and physics based climate models, enabling faithful and stable emulation of temperature fields at multiple surface levels over a 10 year horizon. We propose an auto-regressive ODE–SDE framework for climate emulation that integrates the fundamental physical law of advection, while being trained under a negative log-likelihood objective to enable principled uncertainty quantification of stochastic variability. We show PACER's emulation performance across 20 climate models outperforming relevant baselines and advancing towards explicit physics infusion in ML emulator."},"_bibtex":{"value":"@misc{\nsaleem2025pacer,\ntitle={{PACER}: Physics Informed and Uncertainty Aware Climate Emulator},\nauthor={Hira Saleem and Flora D. Salim and Cormac Purcell},\nyear={2025},\nurl={https://openreview.net/forum?id=RhW8yXgIxY}\n}"},"title":{"value":"PACER: Physics Informed and Uncertainty Aware Climate Emulator"},"pdf":{"value":"/pdf/efe47dca6a1e7f0bb6347335b2bbfb98d2b9ec39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"saleem|pacer_physics_informed_and_uncertainty_aware_climate_emulator"},"authorids":{"value":["~Hira_Saleem1","~Flora_D._Salim1","~Cormac_Purcell1"]},"authors":{"value":["Hira Saleem","Flora D. Salim","Cormac Purcell"]}},"version":2},{"content":{"venue":{"value":"Journal of High Energy Physics"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"dasmahapatra|spectral_clustering_for_jet_physics"},"authorids":{"value":["~Srinandan_Dasmahapatra1"]},"html":{"value":"https://link.springer.com/article/10.1007/JHEP02(2022)165"},"abstract":{"value":"We present a new approach to jet definition alternative to clustering methods,\nsuch as the anti-kT scheme, that exploit kinematic data directly.\nInstead the new method uses kinematic information to represent the\nparticles in a multidimensional space, as in spectral clustering.\nAfter confirming its Infra-Red (IR) safety, we compare its performance in analysing\n\\(gg\\to H_{125\\,\\text{GeV}} \\rightarrow H_{40\\,\\text{GeV}} H_{40\\,\\text{GeV}} \\rightarrow b \\bar{b} b \\bar{b}\\),\n\\(gg\\to H_{500\\,\\text{GeV}} \\rightarrow H_{125\\,\\text{GeV}} H_{125\\,\\text{GeV}} \\rightarrow b \\bar{b} b \\bar{b}\\)\nand\n\\(gg,q\\bar q\\to t\\bar t\\to b\\bar b W^+W^-\\to b\\bar b jj \\ell\\nu_\\ell\\) events from \nMonte Carlo (MC) samples, specifically, in reconstructing the relevant final states, to that of the \\antikt{} algorithm.\n   Finally, we show that the  results for spectral clustering are obtained  without any change in the parameter settings of the algorithm,\n   unlike the \\antikt{} case, which requires the cone size to be adjusted to the physics process under study."},"title":{"value":"Spectral clustering for jet physics"},"authors":{"value":["Srinandan Dasmahapatra"]}},"tmdate":1724067394905,"pdate":1645443388637,"tcdate":1724067394905,"writers":["~Srinandan_Dasmahapatra1"],"signatures":["~Srinandan_Dasmahapatra1"],"forum":"xxpEwIX2kd","license":"CC BY-SA 4.0","number":27794,"cdate":1724067394905,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1724067394905,"domain":"OpenReview.net/Archive","id":"xxpEwIX2kd","version":2},{"content":{"summary":{"value":"This paper systematically reviews both physical models—based on numerical weather prediction and turbine dynamics—and data-driven approaches, including modern deep learning architectures such as CNNs, LSTMs, and Transformers, as well as their hybrid integrations. \nThe authors identify key challenges in WPF, including data quality and heterogeneity, limited spatial generalization across sites, model adaptability under dynamic weather conditions, and the need for multi-objective optimization balancing accuracy, stability, and interpretability. \nThe paper also presents a taxonomy of forecasting horizons and evaluation metrics, illustrating how modeling strategies differ by temporal scale. \nOverall, it contributes a structured synthesis and roadmap for advancing WPF research, advocating for physics-informed learning, uncertainty modeling, multi-modal data fusion, and federated intelligence as critical directions toward more reliable and scalable wind power prediction."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Q1.\nYou discuss hybrid models as a promising approach to bridge the gap between physical and data-driven methods, but how do you see the integration of physical principles (e.g., turbine aerodynamics or meteorological models) with neural network architectures (such as CNNs or LSTMs)?\n\nQ2.\nYou mention data heterogeneity and spatial generalization as key challenges in wind power forecasting. How do you think transfer learning or domain adaptation methods can be employed to address these challenges, particularly when models trained on one wind farm or region must be adapted for another with different conditions (e.g., wind patterns, turbine models, etc.)?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"S1.\nThe authors systematically integrate studies across diverse timescales, geographies, and modeling approaches, ensuring breadth and depth of coverage.\n\nS2.\nThis paper is clear and well-organized.\nIt follows a logical progression from background principles to methodological categories, challenges, and future directions.\n\nS3.\nBy consolidating fragmented research across physics, meteorology, and machine learning, the paper serves as a reference for new entrants to the field and a strategic roadmap for researchers pursuing hybrid or AI-enhanced WPF systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1.\nThe paper introduces hybrid modeling (e.g., combining physical and data-driven approaches) as a promising direction but provides limited theoretical discussion on how these models bridge the gap between the physical and data-driven worlds. \nFor instance, the paper mentions that hybrid models can achieve better accuracy but lacks a detailed analysis on how these models integrate the complementary strengths of physics-based methods (e.g., understanding wind dynamics) and data-driven models (e.g., identifying complex patterns).\n\nW2.\nWhile the paper outlines the state of the art, it lacks empirical validation or case studies to illustrate how different forecasting approaches perform in real-world scenarios.\n\nW3.\nThe paper identifies several key challenges in wind power forecasting, such as data heterogeneity and spatial generalization, but the discussion of potential solutions for these challenges is somewhat superficial. \nFor instance, while data quality issues are mentioned, the paper does not dive deeply into how the field could overcome these issues (e.g., through data augmentation, unsupervised pretraining, or cross-domain transfer learning).\n\nW4.\nWhile the paper proposes several exciting future research directions, such as physics-informed learning and federated learning, the discussion could be more focused on how these approaches can be practically implemented in the WPF domain. \nFor instance, the authors mention uncertainty quantification but do not elaborate on specific models or techniques that can be applied to this area."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922155153,"tcdate":1761887735231,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10959/Reviewer_p8Ua"],"signatures":["ICLR.cc/2026/Conference/Submission10959/Reviewer_p8Ua"],"forum":"7WG3XbHV3V","number":2,"license":"CC BY 4.0","cdate":1761887735231,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10959/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922155153,"domain":"ICLR.cc/2026/Conference","replyto":"7WG3XbHV3V","id":"3mOvg1lhU8","forumContent":{"TLDR":{"value":"We propose a novel physics-guided architecture that traces upstream advection in background wind fields to enhance local wind speed forecasting at meteorological stations."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Time series forecasting","wind speed forecasting","physics-guided machine learning","multi-modal learning","context-enriched learning"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Wind energy is inherently intermittent and fluctuating, and the uncertainty in wind speed poses significant challenges for power system stability. Wind speed forecasting, particularly at wind farm sites, is crucial for balancing power generation and scheduling backup energy sources. Most existing studies rely solely on time-series forecasting methods, while ignoring the physical nature of wind as a spatiotemporal phenomenon primarily driven by atmospheric momentum advection. In this paper, we propose a framework that leverages the surrounding wind field as context to capture the wind dynamics. By explicitly modeling wind advection, our method traces atmospheric momentum transport, enabling the forecast of future trends. We further introduce a multi-modal dataset consisting of wind speed observations from multiple geographically distributed sites and their associated background wind field data. Experimental results demonstrate that our method achieves superior performance over traditional time-series forecasting models and state-of-the-art methods that leverage wind field information."},"_bibtex":{"value":"@misc{\nzhao2025chasing,\ntitle={Chasing the Wind: Background Flow Tracing for Wind Speed Forecasting},\nauthor={Junteng Zhao and Xutao Li and Yikai Li and Kaiqi Zhao and Yunming Ye and Guangbo Deng and Chenxunlai and Huiqi You},\nyear={2025},\nurl={https://openreview.net/forum?id=7WG3XbHV3V}\n}"},"title":{"value":"Chasing the Wind: Background Flow Tracing for Wind Speed Forecasting"},"pdf":{"value":"/pdf/599b9f556f61eedf99ccef3261ac1058e364c38b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhao|chasing_the_wind_background_flow_tracing_for_wind_speed_forecasting"},"authorids":{"value":["~Junteng_Zhao1","~Xutao_Li2","~Yikai_Li6","~Kaiqi_Zhao1","~Yunming_Ye1","~Guangbo_Deng1","~Chenxunlai2","~Huiqi_You1"]},"authors":{"value":["Junteng Zhao","Xutao Li","Yikai Li","Kaiqi Zhao","Yunming Ye","Guangbo Deng","Chenxunlai","Huiqi You"]}},"version":2},{"content":{"summary":{"value":"The paper introduces CL-DPS, a new diffusion-model framework for blind nonlinear inverse problems, removing the need to know measurement operators at inference. The method uses a contrastive learning trained encoder to approximate the conditional likelihood on some synthetic dataset. At test-time, the encoder is used to approximate the likelihood $p(y |x_t)$ and guide the diffusion sampling without operator inversion. The experiments show CL-DPS outperforms existing methods on complex nonlinear tasks and is competitive on linear benchmarks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. How well does the method generalize to real-world, non-synthetic degradations?\n\n2. Could the method be applied to more complex degradations such as unstructured degradations like rain for example?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The method uses a synthetic dataset to train the encoder making it still a blind method which is practical.\n\n2. The method outperforms the compared methods on different tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The method requires some training, which increases the computational cost compared to other zero-shot methods that are training-free.\n\n2. I have concerns about the fairness of the evaluation. Basically you are training your contrastive encoder on some degradations and then testing on them while the other methods do not. How do you think that this evaluation is conclusive on the superiority of the method given that the other methods do not have this advantage. A comparison with methods that use a synthetic dataset for training is needed.\n\n3. To me, to prove that the method is really interesting, it should be shown that it works better than the existing training methods in the out-of-distribution case. Otherwise, I do not see a big advantage compared to simply learning the degradation operator using the synthetic dataset."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915439251,"tcdate":1761998056715,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission32/Reviewer_s2GW"],"signatures":["ICLR.cc/2026/Conference/Submission32/Reviewer_s2GW"],"forum":"KoLYNHJRBY","number":4,"license":"CC BY 4.0","cdate":1761998056715,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission32/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915439251,"domain":"ICLR.cc/2026/Conference","replyto":"KoLYNHJRBY","id":"HkC1YHROg5","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion Models","Blind Inverse Problems","Contrastive Learning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Diffusion models (DMs) have recently become powerful priors for solving inverse problems. However, most work focuses on non-blind settings with known measurement operators, and existing DM-based blind solvers largely assume linear measurements, which limits practical applicability where operators are frequently nonlinear. We introduce CL-DPS, a contrastively trained likelihood for diffusion posterior sampling that requires no knowledge of the operator parameters at inference. To the best of our knowledge, CL-DPS is the first DM-based framework capable of solving blind nonlinear inverse problems. Our key idea is to train an auxiliary encoder offline, using a MoCo-style contrastive objective over randomized measurement operators, to learn a surrogate for the conditional likelihood \\$p(\\boldsymbol{y} | \\boldsymbol{x}\\_t)\\$. During sampling, we inject the surrogate's gradient as a guidance term along the reverse diffusion trajectory, which enables posterior sampling without estimating or inverting the forward operator. We further employ overlapping patch-wise inference to preserve fine structure and a lightweight color-consistency head to stabilize color statistics. The guidance is sampler-agnostic and pairs well with modern solvers (e.g., DPM-Solver++ (2M)). Extensive experiments show that CL-DPS effectively handles challenging nonlinear cases, such as rotational and zoom deblurring, where prior DM-based methods fail, while remaining competitive on standard linear benchmarks. Code: \\url{https://anonymous.4open.science/r/CL-DPS-4F5D}."},"_bibtex":{"value":"@inproceedings{\nye2026cldps,\ntitle={{CL}-{DPS}: A Contrastive Learning Approach to Blind Nonlinear Inverse Problem Solving via Diffusion Posterior Sampling},\nauthor={Linfeng Ye and Shayan Mohajer Hamidi and Mert Pilanci and Konstantinos N. Plataniotis},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KoLYNHJRBY}\n}"},"title":{"value":"CL-DPS: A Contrastive Learning Approach to Blind Nonlinear Inverse Problem Solving via Diffusion Posterior Sampling"},"pdf":{"value":"/pdf/22376416f87f5cc68e86d2b6a6d532638336d6ee.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ye|cldps_a_contrastive_learning_approach_to_blind_nonlinear_inverse_problem_solving_via_diffusion_posterior_sampling"},"authorids":{"value":["~Linfeng_Ye1","~Shayan_Mohajer_Hamidi1","~Mert_Pilanci3","~Konstantinos_N._Plataniotis2"]},"authors":{"value":["Linfeng Ye","Shayan Mohajer Hamidi","Mert Pilanci","Konstantinos N. Plataniotis"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Poly2Graph, an open-source pipeline that converts one-dimensional non-Hermitian crystal Hamiltonians into Hamiltonian spectral graphs (HSGs): spatial multigraphs representing the geometry of complex energy spectra. The authors build HSG-12M, a dataset of 11.6 M static and 5.1 M dynamic multigraphs spanning 1,401 characteristic-polynomial classes. Each graph encodes spectral topology and geometry via node coordinates and multi-edge trajectories on the complex-energy plane. The dataset is positioned as the first large-scale benchmark for spatial multigraph learning, and as a bridge between algebraic physics data and graph machine learning. Benchmarks with popular GNNs (GCN, GAT, GINE, GraphSAGE, etc.) show that edge-aware models outperform edge-agnostic ones, while overall Top-1 accuracies remain moderate ( 30–60 %) and Top-10 accuracies high (95 %)."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Why is a graph representation preferable to treating the spectra as 2-D arrays or polynomial coefficients?\n2. Can you show any downstream task where learning on graphs leads to qualitatively different or improved results than learning on images?\n3. How robust are the extracted graphs to numerical perturbations or thresholding choices?\n4. Would CNNs or transformers on spectral images achieve comparable or better accuracy?\n5. Are there physical or real-data applications planned where HSG-12M would provide measurable benefit?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Impressive scale and engineering: Poly2Graph automates a previously manual physics workflow, generating > 10 M graphs and compressing 177 TB of raw spectra into 256 GB.\n2. First large-scale multigraph dataset\n3. Novel cross-disciplinary framing: Establishes a link between non-Hermitian band theory and graph representation learning, potentially inspiring new geometry-aware GNNs.\n4. Solid baseline benchmarking: Eight GNNs are compared with consistent training budgets. Results reveal real performance gaps between edge-aware and edge-agnostic models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited justification of scientific or ML value: The paper convincingly shows that the data can be generated, but not why learning from these graphs is necessary or insightful. The benchmark task, namely classifying Hamiltonian families, appears artificial, with no demonstrated physical or methodological payoff.\n2. Unclear advantage of the graph representation: The authors do not compare to simpler baselines such as CNNs on spectral images, MLPs on polynomial coefficients, or models trained directly on spectral arrays. It remains unproven that representing spectra as graphs yields better or different information.\n3. Representation loss and stability: The extraction from spectra to graph skeletons may discard quantitative details and can fragment edges (acknowledged in Appendix H). No analysis quantifies how much information or robustness is lost.\n4. Synthetic and self-contained: All data are algorithmically generated. No connection is made to experimental measurements or to existing real-world multigraph domains.\n5. Benchmark insight is shallow: Standard GNNs achieve modest accuracies, but this mainly reflects task complexity and training budget, not necessarily new modeling challenges.\n6. Incremental as a benchmark contribution: The work is an engineering milestone rather than a conceptual one. Its usefulness for advancing ML is not clear."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927280459,"tcdate":1762000026386,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17371/Reviewer_K5Rg"],"signatures":["ICLR.cc/2026/Conference/Submission17371/Reviewer_K5Rg"],"forum":"YxuKCME576","number":4,"license":"CC BY 4.0","cdate":1762000026386,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17371/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927280459,"domain":"ICLR.cc/2026/Conference","replyto":"YxuKCME576","id":"99zalDNDMc","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["graph-level learning","spatial network","multigraph","dataset generator","large-scale dataset","condensed matter physics","non-Hermitian physics","topological physics","AI4Science","Toeplitz matrix"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"AI is transforming scientific research by revealing new ways to understand complex physical systems, but its impact remains constrained by the lack of large, high-quality domain-specific datasets. A rich, largely untapped resource lies in non-Hermitian quantum physics, where the energy spectra of crystals form intricate geometries on the complex plane—termed as $\\textit{Hamiltonian spectral graphs}$. Despite their significance as fingerprints for electronic behavior, their systematic study has been intractable due to the reliance on manual extraction. To unlock this potential, we introduce $\\textbf{Poly2Graph}$ (https://github.com/sarinstein-yan/Poly2Graph): a high-performance, open-source pipeline that automates the mapping of 1-D crystal Hamiltonians to spectral graphs. Using this tool, we present $\\textbf{HSG-12M}$ (https://github.com/sarinstein-yan/HSG-12M): a dataset containing 11.6 million static and 5.1 million dynamic Hamiltonian spectral graphs across 1401 characteristic-polynomial classes, distilled from 177 TB of spectral potential data. Crucially, HSG-12M is the first large-scale dataset of $\\textit{spatial multigraphs}$—graphs embedded in a metric space where multiple geometrically distinct trajectories between two nodes are retained as separate edges. This simultaneously addresses a critical gap, as existing graph benchmarks overwhelmingly assume simple, non-spatial edges, discarding vital geometric information. Benchmarks with popular GNNs expose new challenges in learning spatial multi-edges at scale. Beyond its practical utility, we show that spectral graphs serve as universal topological fingerprints of polynomials, vectors, and matrices, forging a new algebra-to-graph link. HSG-12M lays the groundwork for data-driven scientific discovery in condensed matter physics, new opportunities in geometry-aware graph learning and beyond."},"_bibtex":{"value":"@inproceedings{\nyan2026hsgm,\ntitle={{HSG}-12M: A Large-Scale Benchmark of Spatial Multigraphs from the Energy Spectra of Non-Hermitian Crystals},\nauthor={Xianquan Yan and Hakan Akg{\\\"u}n and Kenji Kawaguchi and N. Duane Loh and Ching Hua Lee},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=YxuKCME576}\n}"},"title":{"value":"HSG-12M: A Large-Scale Benchmark of Spatial Multigraphs from the Energy Spectra of Non-Hermitian Crystals"},"pdf":{"value":"/pdf/a3cae5d1b873a4b0f8658cab3767a016c8992d1c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yan|hsg12m_a_largescale_benchmark_of_spatial_multigraphs_from_the_energy_spectra_of_nonhermitian_crystals"},"authorids":{"value":["~Xianquan_Yan1","~Hakan_Akgün1","~Kenji_Kawaguchi1","~N._Duane_Loh1","~Ching_Hua_Lee1"]},"authors":{"value":["Xianquan Yan","Hakan Akgün","Kenji Kawaguchi","N. Duane Loh","Ching Hua Lee"]}},"version":2},{"content":{"venue":{"value":"CogSci 2025"},"pdf":{"value":"https://escholarship.org/content/qt2fx7q9mg/qt2fx7q9mg.pdf"},"venueid":{"value":"dblp.org/conf/COGSCI/2025"},"paperhash":{"value":"wu|clicking_fast_and_slow_towards_intuitive_and_analytical_behaviors_modeling_for_recommender_systems"},"authorids":{"value":["~Youlin_Wu1","~Haoxi_Zhan1","https://dblp.org/search/pid/api?q=author:Yuanyuan_Sun_0002:","https://dblp.org/search/pid/api?q=author:Haohao_Zhu:","~Bo_Xu8","https://dblp.org/search/pid/api?q=author:Liang_Yang_0003:","https://dblp.org/search/pid/api?q=author:Hongfei_Lin:"]},"html":{"value":"https://escholarship.org/uc/item/2fx7q9mg"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cogsci/WuZ0Z00L25,\n  author={Youlin Wu and Haoxi Zhan and Yuanyuan Sun and Haohao Zhu and Bo Xu and Liang Yang and Hongfei Lin},\n  title={Clicking, Fast and Slow: Towards Intuitive and Analytical Behaviors Modeling for Recommender Systems},\n  year={2025},\n  cdate={1735689600000},\n  url={https://escholarship.org/uc/item/2fx7q9mg},\n  booktitle={CogSci},\n  crossref={conf/cogsci/2025}\n}\n"},"abstract":{"value":"Author(s): Wu, Youlin; Zhan, Haoxi; Sun, Yuanyuan; Zhu, Haohao; Xu, Bo; Yang, Liang; Lin, Hongfei | Abstract: Recommender systems personalize content delivery based on user's interaction history. However, not all clicks result from deliberate decisionsâ€”many arise from intuitive reactions. Inspired by the dual process theory, we argue that intuitive clicks are primarily driven by System 1, reacting to superficial cues, while analytical clicks involve deeper processing by System 2, considering the semantic meaning and long-term preference. However, existing models overlook these cognitive mechanisms. To address this, we propose DualRec, a novel recommendation method that models both intuitive and analytical behaviors. DualRec encodes items using language models, leveraging shallow layers for superficial understanding (System 1) and deep layers for semantic comprehension (System 2). It employs Transformer-based encoders with two attention mechanisms to capture intuitive \"fast\" and analytical \"slow\" click patterns. A learnable fusion layer balances these behaviors. Extensive experiments demonstrate that DualRec outperforms existing methods and highlights the importance of integrating both cognitive processes in recommendations."},"title":{"value":"Clicking, Fast and Slow: Towards Intuitive and Analytical Behaviors Modeling for Recommender Systems"},"authors":{"value":["Youlin Wu","Haoxi Zhan","Yuanyuan Sun","Haohao Zhu","Bo Xu","Liang Yang","Hongfei Lin"]}},"tmdate":1784616715410,"pdate":1767139200000,"externalIds":["dblp:conf/cogsci/WuZ0Z00L25"],"tcdate":1771040641556,"writers":["~"],"signatures":["~Bo_Xu8"],"forum":"joGJXtkTw8","license":"CC BY-SA 4.0","number":823186,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1784616715410,"domain":"DBLP.org","id":"joGJXtkTw8","version":2},{"content":{"summary":{"value":"The framework, FlexMotion, leverages a computationally efficient diffusion model in the latent space, eliminating the need for physics simulators and enabling fast training. It employs a multimodal pre-trained Transformer encoder-decoder that integrates various motion parameters like joint locations, contact forces, joint actuations, and muscle activations to ensure the physical plausibility of the generated motions. FlexMotion also introduces a plug-and-play module for spatial control over motion parameters, enhancing its applicability across different domains."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"1. The author should discuss more about how to obatin the physcis related inputs, for example torque.\n\n2. The author should give more examples to show their method significantly superpass other methods."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1.The use of a diffusion model in the latent space is a novel approach that significantly reduces computational costs compared to traditional methods that rely on physics engines.\n\n2.The integration of joint locations, contact forces, joint actuations, and muscle activations into a single framework is a comprehensive way to ensure physically plausible motion generation.\n\n3.The plug-and-play module for spatial control over a range of motion parameters adds versatility to the framework, making it suitable for various applications.\n\n4.The paper demonstrates superior performance in terms of realism, physical plausibility, and controllability over existing methods, as shown through evaluations on extended datasets.\n\n5.FlexMotion's lightweight design and efficient training process make it suitable for real-time applications, which is a significant advantage over computationally intensive methods"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This paper does not provide any videos to show their qualitative results, which are important to prove their contribution and progress in this research area.\n\n2. In this paper, the property of physics-aware motion is weak. Indeed, we need such properties on flat ground, but the physics-aware property also should work on some uneven terrains and human-human interactions.\n\n3. This paper ignores some baselines, for example, the TL-control in ECCV 2024.\n\n4. The physics-based results are still not as good as the simulation-based method, such as phydiff.\n\n5. Some citation issues, for example the ``Adding conditional control to text-to-image diffusion models'' has two different reference."}},"nonreaders":[],"tmdate":1731429585139,"tcdate":1730355552327,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13568/Reviewer_UeMU"],"signatures":["ICLR.cc/2025/Conference/Submission13568/Reviewer_UeMU"],"forum":"7652tHbbVE","number":2,"license":"CC BY 4.0","cdate":1730355552327,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13568/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429585139,"domain":"ICLR.cc/2025/Conference","replyto":"7652tHbbVE","id":"I50UQQWhu2","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["3D human motion generation","diffusion models","conditional generation","physics aware"]},"supplementary_material":{"value":"/attachment/ce8425b95043b839b3d2c1deb897fea342af8a4b.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Lightweight, controllable, and physically plausible human motion synthesis is crucial for animation, virtual reality, robotics, and human-computer interaction applications. Existing methods often compromise between computational efficiency, physical realism, or spatial controllability. We propose FlexMotion, a novel framework that leverages a computationally lightweight diffusion model operating in the latent space, eliminating the need for physics simulators and enabling fast and efficient training. FlexMotion employs a multimodal pre-trained Transformer encoder-decoder, integrating joint locations, contact forces, joint actuations and muscle activations to ensure the physical plausibility of the generated motions. FlexMotion also introduces a plug-and-play module, which adds spatial controllability over a range of motion parameters (e.g., joint locations, joint actuations, contact forces, and muscle activations). Our framework achieves realistic motion generation with improved efficiency and control, setting a new benchmark for human motion synthesis. We evaluate FlexMotion on extended datasets and demonstrate its superior performance in terms of realism, physical plausibility, and controllability."},"_bibtex":{"value":"@misc{\ntashakori2025flexmotion,\ntitle={FlexMotion: Lightweight, Physics-Aware, and Controllable Human Motion Generation},\nauthor={Arvin Tashakori and Arash Tashakori and Gongbo Yang and Z. Jane Wang and Peyman Servati},\nyear={2025},\nurl={https://openreview.net/forum?id=7652tHbbVE}\n}"},"title":{"value":"FlexMotion: Lightweight, Physics-Aware, and Controllable Human Motion Generation"},"pdf":{"value":"/pdf/55acf31294a70f74ab80cd15e304566fda5d445e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"tashakori|flexmotion_lightweight_physicsaware_and_controllable_human_motion_generation"},"authorids":{"value":["~Arvin_Tashakori1","~Arash_Tashakori1","~Gongbo_Yang1","~Z._Jane_Wang1","~Peyman_Servati1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Arvin Tashakori","Arash Tashakori","Gongbo Yang","Z. Jane Wang","Peyman Servati"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new measure of feature complexity based on an information-theoretic metric, the $\\nu$-information metric. Utilizing this complexity measure, the paper shows (1) visualization of features of different complexities, (2) simple features are propagated through the residual connections to reach the final layer, (3) simple features are learned earlier than complex features during training, and (4) simple features usually have a higher importance (weight) to the output score of the model."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.\tThe paper claims that simple features tend to propagate through the residual connection, while complex features tend to propagate through the main branch of the network. Then a natural question arises: how are simple and complex features propagate in CNNs without residual connections (e.g., AlexNet, VGG)?\n\n2.\tAbout the visualization of simple and complex features in Appendix B. The feature *fences* is among the most simple features, while the features *whiskers* and *dotted texture* are among the most complex features. However, I do not see an essential difference between the features *fences*, *whiskers*, and *dotted texture* from the sample images in Figure 7 and Figure 8. For example, *fences* and *dotted texture* both seem to consist of lines that cross over each other to form holes. Similarly, the *whiskers* feature is also composed of lines, although the lines do not cross over each other but appear parallel. Why is *fences* a simple feature, but *whiskers* and *dotted texture* complex features?\n\nAnother question is that previous studies show that CNNs are usually biased towards textures, i.e., they tend to learn textures as a shortcut solution, but why is the *dotted texture* here measured to be a complex feature?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1.\tThe paper is well presented and is quite easy to follow. I appreciate the summary of different experiments into the words “what” “where” and ”when.”\n\n2.\tThe paper provides abundant visualization to help readers grasp the intuitive behind simple features and complex features."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My major concerns for this paper are its coherence and significance.\n\n1.\tThe paper lacks coherence because multiple components (e.g., metrics, algorithms) introduced in this paper do not come from the same theoretical framework. For example, the notion of feature complexity in this paper (the $\\nu$-information metric) is drawn from information theory, while the method for extracting features (the Craft method) is based on non-negative matrix factorization and has little connection with information theory. The same issue applies to the feature visualization method for demonstrating features of different complexities, and CKA metric, and the importance metric $\\Gamma(z_i)$. These metrics/algorithms are borrowed from different theoretical frameworks and contexts, so it is in doubt whether they can be used together. I would appreciate it if the authors are able to re-organize all components of the paper under a coherent framework, e.g., the information theoretic perspective (if at all possible).\n\n2.\tMy second concern is about the paper’s significance, since many claims have been discussed in previous works and appears non-novel. The conclusion that simple features are usually color/edge detectors that are located in earlier layers is not surprising, nor do the conjecture that simple features tend to propagate through the residual connection. For the conclusion that simple features are learned earlier than complex features, there are many studies leading to this conclusion, from the perspective of spectral bias [cite1], game-theoretic interactions [cite2], or frequency [cite3]. Overall speaking, I do not learn new insights from the paper. Nevertheless, as mentioned in the 1st point in weaknesses, if the authors are able to re-organize the paper from the purely information-theoretic perspective, it will be a more intriguing work.\n\n3.\tUsing K-means clustering to aggregate features into meta-features may lead to incorrect clustering results. K-means clustering is known for its poor performance for data clusters that are not circular-shaped and are of discrepant sizes. I am not sure how features are distributed in this paper’s experiments, but more advanced clustering strategies such as the spectral clustering or hierarchical clustering can be applied to mitigate this issue.\n\n[cite1] Rahaman et al. On the Spectral Bias of Neural Networks. ICML, 2019.\n\n[cite2] Liu et al. Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different Complexities. NeurIPS, 2023.\n\n[cite3] Xu et al. Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks. ICLR, 2020."},"limitations":{"value":"NA"}},"nonreaders":[],"tmdate":1730880128325,"tcdate":1720853355105,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission20708/Reviewer_5SRh"],"signatures":["NeurIPS.cc/2024/Conference/Submission20708/Reviewer_5SRh"],"forum":"NhqZpst42I","number":5,"license":"CC BY 4.0","cdate":1720853355105,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission20708/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730880128325,"domain":"NeurIPS.cc/2024/Conference","replyto":"NhqZpst42I","id":"f7VLvWO1kQ","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Explainability"]},"primary_area":{"value":"interpretability_and_explainability"},"abstract":{"value":"Recent studies suggest that deep learning models' inductive bias towards favoring simpler features may be an origin of shortcut learning. Yet, there has been limited focus on understanding the complexities of the myriad features that models learn. In this work, we introduce a new metric for quantifying feature complexity, based on V-information and capturing whether a feature requires complex computational transformations to be extracted. Using this V-information metric, we analyze the complexities of 10,000 features—represented as directions in the penultimate layer—that were extracted from a standard ImageNet-trained vision model. Our study addresses four key questions:\n\nFirst, we ask what features look like as a function of complexity, and find a spectrum of simple-to-complex features present within the model. Second, we ask when features are learned during training. We find that simpler features dominate early in training, and more complex features emerge gradually. Third, we investigate where within the network simple and complex features \"flow,\" and find that simpler features tend to bypass the visual hierarchy via residual connections. Fourth, we explore the connection between features' complexity and their importance for driving the network's decision. We find that complex features tend to be less important. Surprisingly, important features become accessible at earlier layers during training, like a \"sedimentation process,\" allowing the model to build upon these foundational elements."},"_bibtex":{"value":"@inproceedings{\nfel2024understanding,\ntitle={Understanding Visual Feature Reliance through the Lens of Complexity},\nauthor={Thomas FEL and Louis B{\\'e}thune and Andrew Kyle Lampinen and Thomas Serre and Katherine Hermann},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=NhqZpst42I}\n}"},"title":{"value":"Understanding Visual Feature Reliance through the Lens of Complexity"},"pdf":{"value":"/pdf/62f199b75b56724285a417564dfc48bb27bb43a4.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"fel|understanding_visual_feature_reliance_through_the_lens_of_complexity"},"authorids":{"value":["~Thomas_FEL1","~Louis_Béthune1","~Andrew_Kyle_Lampinen1","~Thomas_Serre1","~Katherine_Hermann1"]},"authors":{"value":["Thomas FEL","Louis Béthune","Andrew Kyle Lampinen","Thomas Serre","Katherine Hermann"]}},"version":2},{"content":{"summary":{"value":"This work proposes a method for wind-driven object dynamics reconstruction. It leverages a combination of physics-based simulations and machine learning techniques to accurately model the behavior of objects subjected to wind forces. Given observational data, the method learns to infer the underlying physical parameters used in the Material Point Method (MPM) simulations, enabling realistic reconstructions of object dynamics in windy environments. A physics-aware optimization strategy is employed to improve the fidelity of the reconstructions."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What's the computational cost and runtime of the proposed method compared to baselines?\n2. How does the method initialize the physical parameters for optimization? A comparison between the initialization and the final learned parameters would be insightful."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The coupling of differentiable physics (LBM + MPM) with 3DGS for realistic wind–object interaction modeling is novel. \n2. Experiments on synthetic and real-world datasets demonstrate clear performance gains over state-of-the-art methods. \n3. The introduction of WD-Objects and the novel “wind retargeting” task broaden research potential."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This work can be viewed as exploration in generative simulation. However, the experimental results in the paper are still toy examples. The author should justify the practicality of the proposed method in real-world applications. What can this method be used for in practice? \n2. The novelty of the proposed method should be further emphasized. How does it compare to existing approaches in the literature? Based on PhysGaussian, the authors should discuss more about the differences and improvements over prior optimization-based methods for physical parameter estimation, such as PhysDreamer, Physics3D, DreamPhysics, OmniPhysGS, etc. \n3. In the supplementary material, the authors only provide backward reconstruction results for clean-background cases. It seems that the method can only handle simple scenarios, which limits its applicability. The authors should provide more analysis and discussion on the limitations of the proposed method."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916209360,"tcdate":1761714500478,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2366/Reviewer_QPR1"],"signatures":["ICLR.cc/2026/Conference/Submission2366/Reviewer_QPR1"],"forum":"vKVzihkbQo","number":2,"license":"CC BY 4.0","cdate":1761714500478,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2366/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916209360,"domain":"ICLR.cc/2026/Conference","replyto":"vKVzihkbQo","id":"IFXphT2ygs","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Physics-based Modeling","3D Dynamics","System Identification","Differentiable Physics"]},"supplementary_material":{"value":"/attachment/7ceccac6970d2edd40c720683c1eac6f76904ac4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio–temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics-informed differentiable framework that unifies wind–object interaction modeling, video-based reconstruction, and forward simulation. Specifically, we represent wind as a grid-based physical field and objects as particle systems derived from 3D Gaussian Splatting, with their interaction modeled by the Material Point Method (MPM). To recover wind-driven object dynamics, we introduce a reconstruction framework that jointly optimizes the spatio–temporal wind force field and object motion through differentiable rendering and simulation. To ensure physical validity, we incorporate the Lattice Boltzmann Method (LBM) as a physics-informed constraint, enforcing compliance with fluid dynamics laws. Beyond reconstruction, our method naturally supports forward simulation under novel wind conditions and enable new applications such as wind retargeting. We further introduce WD-Objects, a dataset of synthetic and real-world wind-driven scenes. Extensive experiments demonstrate that our method significantly outperforms prior dynamic scene modeling approaches in both reconstruction accuracy and simulation fidelity, opening a new avenue for video-based wind–object interaction modeling. The project page is available at: [https://zju3dv.github.io/DiffWind/](https://zju3dv.github.io/DiffWind/)."},"_bibtex":{"value":"@inproceedings{\nlei2026diffwind,\ntitle={DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics},\nauthor={Yuanhang Lei and Boming Zhao and Zesong Yang and Xingxuan Li and Tao Cheng and Haocheng Peng and Ru Zhang and Yang Yang and Siyuan Huang and Yujun Shen and Ruizhen Hu and Hujun Bao and Zhaopeng Cui},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=vKVzihkbQo}\n}"},"title":{"value":"DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics"},"pdf":{"value":"/pdf/3422bdeef670a30b81dac5199f82667d5ca1d1fa.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lei|diffwind_physicsinformed_differentiable_modeling_of_winddriven_object_dynamics"},"authorids":{"value":["~Yuanhang_Lei1","~Boming_Zhao2","~Zesong_Yang1","~Xingxuan_Li2","~Tao_Cheng4","~Haocheng_Peng1","~Ru_Zhang2","~Yang_Yang156","~Siyuan_Huang2","~Yujun_Shen1","~Ruizhen_Hu1","~Hujun_Bao1","~Zhaopeng_Cui1"]},"authors":{"value":["Yuanhang Lei","Boming Zhao","Zesong Yang","Xingxuan Li","Tao Cheng","Haocheng Peng","Ru Zhang","Yang Yang","Siyuan Huang","Yujun Shen","Ruizhen Hu","Hujun Bao","Zhaopeng Cui"]}},"version":2},{"content":{"summary":{"value":"The paper derives complexity-dependent bounds that formally characterize the impact of aligned prior domain knowledge on learning rates in physics-informed machine learning settings. This is approached by focusing on bounds of the excess risk in regularized ERM under dependent data generated from non-linear dynamical systems. Results under the combined assumptions of physics-informed regularization and non-IID data have yet to be shown in existing literature, and the derived bounds reveal new insights into the relationship between statistical learning theory and physics-informed ML."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"- Are there any promising avenues for achieving the stronger burn-in requirement while relaxing $s\\ge 2d_X$ toward the standard $s\\ge d_X/2$, or are there clear reasons to believe this is a necessary tradeoff? Can factors/assumptions be loosened in its place while maintaining the original bound, giving another route to the same result?\n- In the numerical experiment and Figure 2, does the reported slope fit account for burn-in (if even needed in this case)? It would be instructive to see how this shows up here, and/or when burn-in exceeds realistically attainable sample sizes in practice."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper is well-organized and has a presentation that's easy to follow from start to finish. I appreciated the clear groups in Section 3 and the visual distinction for all stated assumptions.\n- The authors address an important theoretical challenge in the physics-informed ML space, with a fundamental connection to all PIML architectures and practitioners looking to tackle scientific problems. Characterizing the impact of well-aligned domain knowledge takes a clear step toward helping researchers and practitioners alike when it comes to justifying additional time spent refining assumptions and calibrating inductive biases.\n- Relaxing the IID assumptions from prior work (e.g., Doumèche et al. 2024) appears to be an important adjustment that broadens the reach of the resulting bounds. In most PIML settings, one is working with heavily dependent sequential data, so this seems a prudent step toward establishing useful bounds that mirror the real world."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Assumption 5 is quite strong, limiting the applicability of the results in many real-world physics-informed modeling scenarios. For instance, unless additional constraints are applied, popular methods like physics-informed neural networks would presumably violate this assumption, and in these cases it's unclear how one should think about the applicability of the results. I understand details on approximations were stated as out of scope, but it would be nice to include some discussion/analysis of the likely ways practitioners may violate assumptions in practice and what elements of the original bound still apply (if any). \n- Similar to the above point, Assumption 4 also excludes common practical scenarios that leverage non-linear PDE priors. Discussion on what remains of the bound in these cases would be helpful for position the paper's analysis in a broader context.\n- This is somewhat tangential to the core theoretical aims of the paper, but it would be instructive to include a more holistic case study. In particular, the impact of using misaligned/incomplete prior physics knowledge and how this empirically threads between the \"with knowledge\" and \"without knowledge\" scenarios. The sample sizes at which various levels of knowledge alignment outperform others, if only for short ranges of $T$, would also be helpful to characterize insofar as they relate to the paper's analysis of sample size."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922856807,"tcdate":1761998319789,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11838/Reviewer_3SEU"],"signatures":["ICLR.cc/2026/Conference/Submission11838/Reviewer_3SEU"],"forum":"IvLVPbeoRx","number":3,"license":"CC BY 4.0","cdate":1761998319789,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11838/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922856807,"domain":"ICLR.cc/2026/Conference","replyto":"IvLVPbeoRx","id":"FZP46QPJNB","forumContent":{"TLDR":{"value":"We prove that adding correct prior domain knowledge to nonparametric learning with dependent data speeds up learning"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["learning with dependent data","physics-informed machine learning","convergence rates","complexity-dependent bounds"]},"supplementary_material":{"value":"/attachment/a00585da11ea98b0c0c19b8e1c5d1c3a2b58963e.pdf"},"primary_area":{"value":"learning theory"},"abstract":{"value":"A major challenge in physics-informed machine learning is to understand how the incorporation of prior domain knowledge affects learning rates when data are dependent. Focusing on empirical risk minimization with physics-informed regularization, we derive complexity-dependent bounds on the excess risk in probability and in expectation. We prove that, when the physical prior information is aligned, the learning rate improves from the (slow) Sobolev minimax rate to the (fast) optimal i.i.d. one without sample-size deflation due to data dependence."},"_bibtex":{"value":"@inproceedings{\nscampicchio2026physicsinformed,\ntitle={Physics-informed learning under mixing: How physical knowledge speeds up learning},\nauthor={Anna Scampicchio and Leonardo Felipe Toso and Rahel Rickenbach and James Anderson and Melanie Zeilinger},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=IvLVPbeoRx}\n}"},"title":{"value":"Physics-informed learning under mixing: How physical knowledge speeds up learning"},"pdf":{"value":"/pdf/9917146046b56820383565485dace6196f5b94c5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"scampicchio|physicsinformed_learning_under_mixing_how_physical_knowledge_speeds_up_learning"},"authorids":{"value":["~Anna_Scampicchio1","~Leonardo_Felipe_Toso1","~Rahel_Rickenbach1","~James_Anderson6","~Melanie_Zeilinger1"]},"authors":{"value":["Anna Scampicchio","Leonardo Felipe Toso","Rahel Rickenbach","James Anderson","Melanie Zeilinger"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on the scarcity of heterogeneous graph datasets and overcomes the need for such datasets in the domain of GNN explanations. In particular, it proposes a novel SynHING framework for generating synthetic HINs. This method leverages the real-world HINs as references and systematically generates the major motifs for explanations. Extensive experiments have demonstrated its generality and practicality."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. I would like to know if this method can assist with graph-level explanation tasks in heterogeneous graphs."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Originality: This paper introduces the first framework for generating synthetic heterogeneous graphs with ground-truth explanations.\n\n2. Clarity: The paper's writing in the methodology section is logically clear.  The reason behind each step of the method design is also explained in detail.\n\n3. Significance: The framework for generating diverse synthetic HINs that can be flexibly adjusted will provide a solid foundation for future research on heterogeneous GNN explanations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The size of many figures in this paper needs to be adjusted. For example, Figure 2 is too small, which is not conducive to reading.\n\n2.  The writing of this paper could be improved. For instance, there is a significant gap between the first and second paragraphs of the introduction.\n\n3.  The techniques used in this paper for generating synthetic HINs are relatively normal, focusing only on the degree of nodes in the graph to consider the relationship with real graphs. \n\n4.  This paper only validates the methods on node-level tasks, and it is unclear whether this method can be transferred to graph-level explanation task scenarios.\n\n5. The design of the paper's experimental section needs improvement. First, more tabular types of experiments should be used to conduct a more intuitive quantitative analysis. Second, important experimental results should be presented at the forefront."}},"nonreaders":[],"tmdate":1732589508930,"tcdate":1730190819267,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8484/Reviewer_kxr8"],"signatures":["ICLR.cc/2025/Conference/Submission8484/Reviewer_kxr8"],"forum":"ZbHIDgDFN0","number":1,"license":"CC BY 4.0","cdate":1730190819267,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8484/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732589508930,"domain":"ICLR.cc/2025/Conference","replyto":"ZbHIDgDFN0","id":"VUCWORBYFg","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic graph generation","heterogeneous information networks","graph neural networks","explainable artificial intelligence"]},"supplementary_material":{"value":"/attachment/1a596c472d696653ea9d00ed17f09eac6df0c5af.zip"},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Graph Neural Networks (GNNs) excel in modeling graph structures across diverse domains, such as community analysis and recommendation systems. As the need for GNN interpretability grows, there is an increasing demand for robust baselines and comprehensive graph datasets, especially within the realm of Heterogeneous Information Networks (HIN). To address this, we introduce SynHING, a framework for Synthetic Heterogeneous Information Network Generation designed to advance graph learning and explanation.\nAfter identifying key motifs in a target HIN, SynHING systematically employs a bottom-up generation process with intra-cluster and inter-cluster merge modules. This process, along with post-pruning techniques, ensures that the synthetic HIN accurately mirrors the structural and statistical properties of the original graph. The effectiveness of SynHING is validated using four datasets - IMDB, Recipe, ACM, and DBLP - spanning three distinct application categories, demonstrating both its generality and practicality.\nFurthermore, SynHING provides ground-truth motifs for evaluating GNN explainer models, establishing a new benchmark for explainable, synthetic HIN generation. This contributes significantly to advancing interpretable machine learning in complex network environments."},"_bibtex":{"value":"@misc{\nhong2025synhing,\ntitle={Syn{HING}: Synthetic Heterogeneous Information Network Generation for Graph Learning and Explanation},\nauthor={Ming-Yi Hong and Yi-Hsiang Huang and Shao-En Lin and You-Chen Teng and Chih-Yu Wang and Che Lin},\nyear={2025},\nurl={https://openreview.net/forum?id=ZbHIDgDFN0}\n}"},"title":{"value":"SynHING: Synthetic Heterogeneous Information Network Generation for Graph Learning and Explanation"},"pdf":{"value":"/pdf/6ef74ba4277eeea71e09dd3cf7b2e42402e1a696.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"hong|synhing_synthetic_heterogeneous_information_network_generation_for_graph_learning_and_explanation"},"authorids":{"value":["~Ming-Yi_Hong1","~Yi-Hsiang_Huang1","~Shao-En_Lin1","~You-Chen_Teng1","~Chih-Yu_Wang1","~Che_Lin1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ming-Yi Hong","Yi-Hsiang Huang","Shao-En Lin","You-Chen Teng","Chih-Yu Wang","Che Lin"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a physics-informed deep learning framework for forecasting nearshore harmful algal blooms (HABs) along the California coast. Using multi-sensor satellite data, ocean–atmosphere reanalysis fields, and geospatial predictors, the authors compare three models: a ConvLSTM, a Temporal Fusion Transformer (TFT), and a physics-informed ConvLSTM (PINN) that embeds a 2-D advection–diffusion equation as a soft constraint.\nThe model trained on an 18-year, 4-km dataset with blocked spatiotemporal cross-validation, the PINN achieves modest but consistent gains in spatial fidelity and generalization, particularly for moderate-to-high bloom regimes and short lead times (8–16 days). SHAP and partial-dependence analyses reveal interpretable physical drivers (e.g., Kd490, SST, river proximity), enhancing ecological insight. Overall, the work offers a scalable and interpretable approach for physics-guided coastal forecasting with direct implications for environmental management."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Q1: In Section 2.5, the author introduces the residual-loss weight (λ) and the diffusivity parameter (κ = 25 m² s⁻¹) as key components of the physics constraint. Could you clarify whether you tested different values or schedules for these parameters? If not, how are you sure that the reported improvements are robust to changes in λ and κ, rather than dependent on a specific setting?\n\nQ2: It is mentioned in the Abstract and Discussion that the framework is scalable and transferable to other upwelling systems. Could you elaborate on whether this transferability is theoretical or based on preliminary cross-region experiments? If not tested yet, what modifications would be needed to adapt the model to other coastal systems?\n\nQ2: In Section 3.4, the author evaluates spatial fidelity using metrics such as relative scale error, spectral energy ratio, and cross-shelf RMSE, concluding that the PINN improves spatial robustness. Could you clarify how you define “high spatial fidelity” and how fidelity differs from accuracy? What aspects of the physics constraint most contribute to this improvement?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"S1: The paper has Figures that are well-annotated, equations properly derived, and every methodological step, feature engineering, model architecture, evaluation, and explainability are presented in reproducible detail. The explicit disclosure of limited LLM use for grammar use only shows ethical transparency.\n\nS2: The paper shows prediction accuracy by representing physical interpretability through SHAP and partial-dependence analyses. Showed key environmental drivers such as Kd490, SST, and river influence improve scientific understanding and credibility, thus making an important advance in a field often dominated by black-box models.\n\nS3: The author's contribution is by bringing physics-informed neural networks (PINNs) into the domain of nearshore harmful algal bloom (HAB) forecasting, a field traditionally dominated by empirical statistical models or purely mechanistic simulations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1: The paper ends with a detailed discussion, but no explicit conclusion. A dedicated “Conclusion” section is missing, making it difficult to understand the what are final takeaway.\n\nW2: Generality beyond the California coast is suggested but not very clearly demonstrated.\n\nW3: The probabilistic skill is assessed only for TFT, but PINN uncertainty quantification is absent. The PINN evaluation in Section 2.6 only uses basic error metrics and briefly mentions physics residuals without defining how they’re measured or interpreted.\n\nW4: The paper does not have a Related Work section explaining how the related studies have been conducted and how they are different from the proposed approach. There have been many PINN approaches for different applications. They should be clearly compared."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916525585,"tcdate":1761693514650,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3047/Reviewer_CSUo"],"signatures":["ICLR.cc/2026/Conference/Submission3047/Reviewer_CSUo"],"forum":"JS6H6R60I6","number":1,"license":"CC BY 4.0","cdate":1761693514650,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3047/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916525585,"domain":"ICLR.cc/2026/Conference","replyto":"JS6H6R60I6","id":"VlAEn1Kwqc","forumContent":{"TLDR":{"value":"This study develops physics-informed deep learning models that combine data to more accurately forecast harmful algal blooms along California’s coast, improving spatial fidelity and ecological insight for management and public health."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["harmful algal blooms","spatiotemporal modeling","physics-informed neural networks","ocean forecasting","remote sensing"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Harmful algal blooms (HABs) are increasing in frequency, duration, and extent along the California coast, driven by climate variability, nutrient enrichment, and complex physical–biogeochemical interactions. Forecasting HAB development and spread remains a challenge, especially in Eastern Boundary Upwelling Systems where advection, stratification, and episodic river inputs strongly shape bloom dynamics. Existing approaches often trade physical realism for statistical flexibility, limiting generalization across bloom regimes. We present a physics-informed deep learning framework for nearshore chlorophyll-a forecasting in the California Current System, integrating multi-sensor satellite products, atmospheric and ocean reanalysis fields, and static geospatial predictors. Three architectures are evaluated: a convolutional long short-term memory network (ConvLSTM), a Temporal Fusion Transformer (TFT), and a physics-informed ConvLSTM (PINN) incorporating the two-dimensional advection–diffusion equation as a soft training constraint. A multi-year, 4 km-resolution dataset (2003–2021) is processed via a tailored feature engineering pipeline with quality-controlled gap-filling, rolling statistics, lagged predictors, and climatology-based anomalies. Models are assessed with strict spatiotemporal cross-validation, emphasizing spatial fidelity, bloom footprint representation, and predictor interpretability. Post-hoc explainability analyses identify key environmental drivers consistent with known upwelling–bloom linkages in the region. We present comparative skill assessments, spatial bias analyses, and predictor attribution results, highlighting the advantages and trade-offs of adding physical constraints to coastal HAB forecasting models. This work delivers a scalable and transferable methodology with direct implications for ecosystem management, fisheries, and public health."},"_bibtex":{"value":"@misc{\nmohanty2026physicsguided,\ntitle={Physics-Guided Neural Forecasts of Nearshore Harmful Algal Blooms in the California Current System},\nauthor={Yashnil Mohanty},\nyear={2026},\nurl={https://openreview.net/forum?id=JS6H6R60I6}\n}"},"title":{"value":"Physics-Guided Neural Forecasts of Nearshore Harmful Algal Blooms in the California Current System"},"pdf":{"value":"/pdf/c08cc651a6ddeba4fd8a92f3b258a53d17814ed0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"mohanty|physicsguided_neural_forecasts_of_nearshore_harmful_algal_blooms_in_the_california_current_system"},"authorids":{"value":["~Yashnil_Mohanty1"]},"authors":{"value":["Yashnil Mohanty"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a new Recurrent Attention-based Neural Cellular Automata (RA-NCA) in modeling complex systems involving many agents with local, often stochastic interactions.  RA-NCA's innovation lies in a recurrent cellular attention module that combines long short-term memory (LSTM) with cellular self-attention. By evaluating on three simulated multi-agent datasets, RA-NCA shows three good properties which are 1.) Robustness in the presence of stochastic interactions; 2.) Data Efficiency which requires less training data due to the permutation invariance inductive bias; and 3.) Scalability, as it can be trained on small systems and successfully applied to significantly larger systems without the need for re-training, and without a decrease in performance."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The paper proposes a new model that extends existing NCA methods to capture long-term local and stochastic interactions among agents.\n2. The \bmethod is technically sound and the writing is in general easy to follow.\n3. The experiment results compared with selected baselines show the superior of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. My major question is the comparison with existing baselines. I understand the paper targets developing a new NCA method to address long-term local interaction among agents and the stochastic property. However, for the specific task it is dealing with, there are much more methods to compare with in literature on multi-agent dynamical system modeling. Examples include discrete GNN-based methods [1][2], and continuous GNN-based methods [3][4] where the key idea is similar to capture the influence from neighbors and past timestamps to make predictions in the future. Also there are some reinforcement learning literature that can address the same task. I believe at least the authors should have a thorough discussion about these directions.\n\n\n[1] Alvaro Sanchez-Gonzalez et.al.  Learning to simulate complex physics with graph networks.\n\n\n[2] Peter W. Battaglia et.al. Interaction Networks for Learning about Objects,Relations and Physics.\n\n\n[3] Zijie Huang et.al. Learning continuous system dynamics from irregularly-sampled partial observations.\n\n\n[4] Chengxi Zang et.al. Neural Dynamics on Complex Networks."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1.  many-agent --> multi-agent?\n2. In Figure 4, can you also plot the visualization from baselines as comparisons?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637010325,"tcdate":1698827303179,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8157/Reviewer_rEW6"],"signatures":["ICLR.cc/2024/Conference/Submission8157/Reviewer_rEW6"],"forum":"YhpgUWE4Rt","number":4,"license":"CC BY 4.0","cdate":1698827303179,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8157/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637010325,"domain":"ICLR.cc/2024/Conference","replyto":"YhpgUWE4Rt","id":"vdGR5NiwFq","forumContent":{"TLDR":{"value":"A neural cellular automata model that effectively discovers the dynamic and stochastic local interaction with limited data"},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["neural cellular automata","complex system","self-organization","multi-agent system"]},"supplementary_material":{"value":"/attachment/848be84c1d57ebd088f32bf3b146d90b9a486964.zip"},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Many-agent systems, such as epidemic spread, rumor propagation through crowd, prey-predator model, and forest fire, exhibit complex global dynamics originated from local, relatively simple, and often stochastic interactions between agents. Despite significant advancements in predictive modeling through deep learning, such interactions among many agents have rarely explored as a specific domain for predictive modeling. We present Recurrent Attention-based Neural Cellular Automata (RA-NCA), to effectively discover the local stochastic interaction by associating the temporal information between neighboring agents in a permutation-invariant manner. RA-NCA exhibits the superior generalizability across various agent configurations (i.e., spatial distribution of agents), data efficiency and robustness in extremely data-limited scenarios even with the presence of stochastic interactions, and scalability through spatial dimension-independent prediction. We compare and evaluate RA-NCA with other NCA networks and scene prediction networks in the three synthetic multi-agent systems with thousands of agents, such as forest fire, host-pathogen, and stock market models."},"_bibtex":{"value":"@misc{\nkang2024recurrent,\ntitle={Recurrent Neural Cellular Automata with Self-Attention for Multi-agent System},\nauthor={Beomseok Kang and Harshit Kumar and Minah Lee and Biswadeep Chakraborty and Saibal Mukhopadhyay},\nyear={2024},\nurl={https://openreview.net/forum?id=YhpgUWE4Rt}\n}"},"title":{"value":"Recurrent Neural Cellular Automata with Self-Attention for Multi-agent System"},"pdf":{"value":"/pdf/6890ec75812596570a3c95a497c9118c77194d23.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"kang|recurrent_neural_cellular_automata_with_selfattention_for_multiagent_system"},"authorids":{"value":["~Beomseok_Kang1","hkumar64@gatech.edu","~Minah_Lee1","~Biswadeep_Chakraborty1","~Saibal_Mukhopadhyay2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Beomseok Kang","Harshit Kumar","Minah Lee","Biswadeep Chakraborty","Saibal Mukhopadhyay"]}},"version":2},{"content":{"summary":{"value":"The authors leverage the framework of score matching, which has become popular for training diffusion-based models for generative tasks, to reverse physical processes defined by forward stochastic differential equations (SDEs). Given a system state at time t=T, they propose to solve for the initial conditions at t=0 by iteratively applying a reverse-time diffusion process defined by a reverse physics simulator, a diffusion term, and the score of the data distribution. In the main contribution of the paper, the authors introduce both a 1-step and multi-step loss for training a network to learn the diffusion and score terms for solving the reverse SDE. They prove the equivalence of their proposed training objectives to vanilla denoising score matching (and a related variational objective in the multi-step case). Extensive experiments demonstrate the efficacy of their method compared to baselines, and ablation studies show the utility of the proposed multi-step objective.               "},"soundness":{"value":"4 excellent"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"- What is your reasoning for training your model with the 1-step loss between $x_m$ and $x_{m-1}$ instead of $x_m$ and $x_0$ (as in typical diffusion-based generative models)? As I understand it, the forward and reverse physical processes are not time-dependent, and arbitrarily large time steps can be taken by simply increasing $\\Delta t$. Intuitively, this method makes more sense to me than your proposed approach, and would also remove the need for multi-step training. \n- How did you deal with the memory expense during multi-step training? Did you find that you needed to use smaller data and network sizes to fit the gradients in memory, or were there tricks and optimizations you used to reduce memory use?\n- One concern is that your model will overfit to the physical parameters at training and perform poorly if there is a test-time distribution shift. How robust is your proposed approach to test-time shifts in the SDE parameters (namely, the coefficients of the simulator and diffusion terms)? I understand that there is limited time for responses, but a small experiment would be appreciated and may convince me to raise my score.  "},"rating":{"value":"7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"4 excellent"},"strengths":{"value":"### Originality\n\nTo my knowledge, this is the first work that considers the drift term in the typical forward SDE used in diffusion-based generative models to be a realistic physical process. Other works [1,2] have considered different diffusion processes besides additive Gaussian noise, but even these use artificial processes such as synthetic blur and pixel masking. I appreciate that the authors of the current paper identify the connection between physically-defined differential equations and those used in training score-based models, and propose a novel method for solving physics problems.\n\n### Quality\n\n- The authors perform extensive experiments grounded in real-world physical processes and show the effects of varying numerous design choices in each experiment\n- The proposed approach demonstrates superior performance compared to baselines across various settings\n- Ablation studies show the utility of the multi-step objective over the 1-step variant\n- The figures are well-made, particularly Figure 1 which clearly lays out the key ideas of the proposed approach  \n\n### Clarity\n\n- The introduction does a good job of laying out the current state of diffusion-based generative models, the connection to physics processes, and the contribution of the current work\n- Clear descriptions of the problem setup, model training and inference, and hyperparameters are given for each experiment\n\n### Significance\n\nPhysics-based inverse problems arise in many fields such as astronomy, geophysics, and wireless communication, so finding better solutions is a highly significant problem with broad interest. Furthermore, quantifying the uncertainty of these solutions is crucial to the downstream decision-making process. The authors of the current paper provide a principled approach to both solve and provide uncertainty estimates (using multiple samples from the SDE) for inverse problems in physics.                    \n\n### References:\n\n[1] G. Daras, M. Delbracio, H. Talebi, A. Dimakis, and P. Milanfar, “Soft Diffusion: Score Matching with General Corruptions,” Transactions on Machine Learning Research, 2023, [Online].\n\n[2] A. Bansal et al., Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise. 2022."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- I believe that the loss in Eq (2) is incorrect. The quantity within the squared L2 norm should be $x_m$ minus the quantity on the right-hand side of Eq (1). In its current form, Eq (2) does not match Eq (3) when Eq (3) is expanded with window S=2.  \n- While the thorough experimental details are appreciated, I believe some of that can be pushed to the appendix. More space should be dedicated to expanding on motivation and design choices. As it currently stands, section 2 reads as a constant flow of information with insufficient context for the proposed objectives.\n- I believe that the organization of the sections could be improved by introducing the typical score matching SDE subject matter first, identifying the differences between that formulation and the physics-based inverse problem formulation, then motivating your proposed approach within this context.\n- An obvious concern to me regarding the multi-step training is the huge memory expense arising from the recursive calls to $s_{\\theta}$. However, the authors did not address this point in the main paper.\n- The authors state that, for decreasing time step $\\Delta t$, the reverse physics simulator is equivalent to the negative of the forward simulator (line 113). However, it is not clarified whether this is a simplifying assumption, true in general, or true in the specific case of the reverse-time ODE. The authors expand on the specific choices of the reverse simulators in the experiments section, but the relationship between the forward and reverse simulators remains unclear to me. \n"},"limitations":{"value":"The authors mention limitations in their conclusion."}},"nonreaders":[],"tmdate":1702410810287,"tcdate":1688253144137,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1882/Reviewer_3QqF"],"signatures":["NeurIPS.cc/2023/Conference/Submission1882/Reviewer_3QqF"],"forum":"2BpoGPSDCR","number":1,"license":"CC BY 4.0","cdate":1688253144137,"mdate":1702410810287,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1882/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"2BpoGPSDCR","id":"dHgNUpRTct","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["inverse problems","diffusion models","learned corrections","score matching"]},"_bibtex":{"value":"@inproceedings{\nholzschuh2023solving,\ntitle={Solving Inverse Physics Problems with Score Matching},\nauthor={Benjamin Holzschuh and Simona Vegetti and Nils Thuerey},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=2BpoGPSDCR}\n}"},"title":{"value":"Solving Inverse Physics Problems with Score Matching"},"paperhash":{"value":"holzschuh|solving_inverse_physics_problems_with_score_matching"},"TLDR":{"value":"We propose a novel methodology for solving inverse problems that involve the temporal evolution of physical systems."},"abstract":{"value":"We propose to solve inverse problems involving the temporal evolution of physics systems by leveraging recent advances from diffusion models. \nOur method moves the system's current state backward in time step by step by combining an approximate inverse physics simulator and a learned correction function. \nA central insight of our work is that training the learned correction with a single-step loss is equivalent to a score matching objective, while recursively predicting longer parts of the trajectory during training relates to maximum likelihood training of a corresponding probability flow.\nWe highlight the advantages of our algorithm compared to standard denoising score matching and implicit score matching, as well as fully learned baselines for a wide range of inverse physics problems. The resulting inverse solver has excellent accuracy and temporal stability and, in contrast to other learned inverse solvers, allows for sampling the posterior of the solutions. Code and experiments are available at https://github.com/tum-pbs/SMDP."},"pdf":{"value":"/pdf/e14d868c93d4bb63690c02fd9c5fcb4f28eed65c.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Benjamin_Holzschuh1","~Simona_Vegetti1","~Nils_Thuerey1"]},"authors":{"value":["Benjamin Holzschuh","Simona Vegetti","Nils Thuerey"]}},"version":2},{"content":{"summary":{"value":"This paper presents a transformation to map RGB images into complex domain and an associate network comprising a loss function to handle such complex inputs.\n\nOverall, the contributions are significant and may be of interest to the community, but the paper organization could be improved and more focus should be placed on the transformation."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1) I am very curious about the invertibility of the proposed transform. Given the Algorithm 2 in Appendix A, would it be possible to have some experiments to prove its effectiveness? I think that this transformation is the real contribution of the paper, as it allows a direct mapping between greyscale and RGB images, which was lacking in complex and quaternion papers that often struggle to do so.\n2) Which is the computational load in terms of FLOPs, runtime memory, and time of the proposed model?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1) The proposed transformation is novel and may be an important contribution for the community working in complex and hypercomplex domains.\n2) Although not novel, using the fourier filter module is good to handle complex inputs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) While the method sounds, the results are not impressive and surely they are not statistically significant. As of my experience, complex, quaternion, and in general hypercomplex models clearly outperform real-valued counterparts when they are able to catch some underlying physical process intrinsic into data. Maybe, the tasks chosen by the authors do not highlight the effectiveness of their method. Maybe, the authors could stress more the parameters saving of using a complex model with respect to a real-valued one, which can help reducing the computational load while obtaining comparable results.\n2) The authors should have focused more on the transformation, which is a novel contribution, and better show its properties (see questions).\n3) Some key references to related works are missing, the authors should at least give credit to them, or better try to compare their method with them. Some of them follow, but I encourage the authors to better explore previous literature on complex, quaternion and hypercomplex networks.\n\n[1] C. Trabelsi, O. Bilaniuk, Dmitriy Serdyuk, Sandeep Subramanian, J. F. Santos, Soroush Mehri, Negar Rostamzadeh, Yoshua Bengio, C. Pal, \"Deep Complex Networks\", ICLR 2017.\n\n[2] E. Grassucci, A. Zhang, D. Comminiello, \"PHNNs: Lightweight Neural Networks via Parameterized Hypercomplex Convolutions\", IEEE Transactions on Neural Networks and Learning Systems, (Volume: 35, Issue: 6, June 2024).\n\n\nMinor comments:\n\nThe Saxon Genitive should be avoided in scientific writing, although I know that both ChatGPT and Grammarly insert it. I suggest the authors to remove all the Saxon genitives in the paper."}},"nonreaders":[],"tmdate":1733085709762,"tcdate":1730543839409,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1276/Reviewer_BUvX"],"signatures":["ICLR.cc/2025/Conference/Submission1276/Reviewer_BUvX"],"forum":"9hmDl8fFDs","number":4,"license":"CC BY 4.0","cdate":1730543839409,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1276/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733085709762,"domain":"ICLR.cc/2025/Conference","replyto":"9hmDl8fFDs","id":"tcCEo24jc3","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"A robust complex-valued approach in Spatio-spectral domain for multiple tasks on both real and complex data."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Deep Complex Newtworks","Complex-valued color transformation"]},"supplementary_material":{"value":"/attachment/d3c1bebc77a7f4edd23a402e6580595b34f472e4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex-valued neural networks have attracted growing attention for their ability to handle complex-valued data with enhanced representational capacity. However, their potential in computer vision remains relatively untapped. \nIn this paper, we introduce Deep Complex Spatio-Spectral Network (DCSNet), a fully complex-valued token-based, end-to-end neural network designed for binary segmentation tasks. Additionally, our DCSNet encoder can be used for image classification in the complex domain. We also propose an invertible real-to-complex (R2C) transform, which generates two complex-valued input channels, complex intensity and complex hue, while producing complex-valued images with distinct real and imaginary components.\nDCSNet operates in both spatial and spectral domains by leveraging complex-valued inputs and complex Fourier transform.\nAs a result, the complex-valued representation is maintained throughout DCSNet, and we avoid the information loss typically associated with Real$\\leftrightarrow$Complex transformations. Extensive experiments show that DCSNet surpasses existing complex-valued methods across various tasks on both real and complex-valued data and achieves competitive performance compared to existing real-valued methods, establishing a robust framework for handling both data types effectively."},"_bibtex":{"value":"@misc{\nyadav2025deep,\ntitle={Deep Complex Spatio-Spectral Networks with Complex Visual Inputs},\nauthor={Saurabh Yadav and Koteswar Rao Jerripothula},\nyear={2025},\nurl={https://openreview.net/forum?id=9hmDl8fFDs}\n}"},"title":{"value":"Deep Complex Spatio-Spectral Networks with Complex Visual Inputs"},"pdf":{"value":"/pdf/78685d9476ab42033341d680f3e952693b0cbc64.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yadav|deep_complex_spatiospectral_networks_with_complex_visual_inputs"},"authorids":{"value":["~Saurabh_Yadav2","~Koteswar_Rao_Jerripothula3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Saurabh Yadav","Koteswar Rao Jerripothula"]}},"version":2},{"content":{"review":{"value":"This paper presents a new method for creating efficient ML surrogates for physics-based simulations. The proposed approach trains faster, and can better capture both complex geometries and dynamics over state-of-the-art surrogates modeling techniques. \n\nThis approach would address a key painpoint in creating accurate surrogates for numerical simulations. The proposed dynamic lifting based on synthetic dynamics is a very interesting self-supervised approach to generate large volumes of training data without running full, computationally-expensive simulations. The authors demonstrate that this works well for a some industry benchmark datasets for aerodynamics, hydrodynamics and car crash. The paper is well written and the results are compelling to be accepted for ICLR. The work appears to be theoretically grounded and can apply to other domains. \n\nThe key drawbacks here are that (1) the benchmarks are still relatively simple (NASA-CRM, Driv-Aer-ML), so it might not scale to more complex PDE-based simulations, (2) the authors claim that this handles simulations with complex geometries, but the industrial examples are not as complex compared to other real-world use cases (e.g., earth system dynamics), and (3) it's not clear how well the synthetic dynamics generated using the relatively simple transport process during training would match  actual dynamics, and whether this would affect model performance, particularly when real-world conditions diverge significantly from the modeled paths used for training.  The reduction in data requirements are notable although the need to generate a good volume of simulation data will still remain an issue. The improvement in accuracy compared to the Transolver benchmarks is also relatively small (about 10-20% from Tables 7-10) but significant given the reduced data requirements. Still the primary contribution of incorporating both dynamics and geometry (using only geometry) is novel, and the paper is a good  proof of concept demonstration for this new approach."},"confidence":{"value":2},"rating":{"value":9},"title":{"value":"GeoPT neural simulator for surrogate modeling"}},"parentInvitations":"ICLR.cc/2026/Workshop/FM4Science/-/Official_Review","nonreaders":[],"tmdate":1772549370026,"tcdate":1771982703426,"writers":["ICLR.cc/2026/Workshop/FM4Science","ICLR.cc/2026/Workshop/FM4Science/Submission77/Reviewer_cyfG"],"signatures":["ICLR.cc/2026/Workshop/FM4Science/Submission77/Reviewer_cyfG"],"forum":"N9qIqvanBj","number":2,"license":"CC BY 4.0","cdate":1771982703426,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/FM4Science/Submission77/-/Official_Review","ICLR.cc/2026/Workshop/FM4Science/-/Edit"],"mdate":1772549370026,"domain":"ICLR.cc/2026/Workshop/FM4Science","replyto":"N9qIqvanBj","id":"z1st9VCxqF","forumContent":{"TLDR":{"value":"This paper presents GeoPT as a unified model pre-trained from large-scale geometries for general physics simulation."},"venue":{"value":"ICLR 2026 Workshop FM4Science Poster"},"pdf":{"value":"/pdf/3df9a842bf6bf180192840eca36de1f7eda780d3.pdf"},"keywords":{"value":["Neural Simulation","PDE Solving","Self-Supervised Pre-Training"]},"venueid":{"value":"ICLR.cc/2026/Workshop/FM4Science"},"paperhash":{"value":"wu|geopt_scaling_physics_simulation_via_lifted_geometric_pretraining"},"authorids":{"value":["~Haixu_Wu1","~Minghao_Guo1","~Zongyi_Li1","~Zhiyang_Dou1","~Mingsheng_Long5","~Kaiming_He2","~Wojciech_Matusik2"]},"abstract":{"value":"Neural simulators promise efficient surrogates for physics simulation, but scaling them is bottlenecked by the prohibitive cost of generating high-fidelity training data. Pre-training on abundant off-the-shelf geometries offers a natural alternative, yet faces a fundamental gap: supervision on static geometry alone ignores dynamics and can lead to negative transfer on physics tasks. We present GeoPT, a unified pre-trained model for general physics simulation based on lifted geometric pre-training. The core idea is to augment geometry with synthetic dynamics, enabling dynamics-aware self-supervision without physics labels. Pre-trained on over one million samples, GeoPT consistently improves industrial-fidelity benchmarks spanning fluid mechanics for cars, aircraft, and ships, and solid mechanics in crash simulation, reducing labeled data requirements by 20-60% and accelerating convergence by 2$\\times$. These results show that lifting with synthetic dynamics bridges the geometry-physics gap, unlocking a scalable path for neural simulation and potentially beyond. Code is available at https://github.com/Physics-Scaling/GeoPT."},"_bibtex":{"value":"@inproceedings{\nwu2026geopt,\ntitle={Geo{PT}: Scaling Physics Simulation via Lifted Geometric Pre-Training},\nauthor={Haixu Wu and Minghao Guo and Zongyi Li and Zhiyang Dou and Mingsheng Long and Kaiming He and Wojciech Matusik},\nbooktitle={ICLR 2026 Workshop on Foundation Models for Science: Real-World Impact and Science-First Design},\nyear={2026},\nurl={https://openreview.net/forum?id=N9qIqvanBj}\n}"},"title":{"value":"GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training"},"authors":{"value":["Haixu Wu","Minghao Guo","Zongyi Li","Zhiyang Dou","Mingsheng Long","Kaiming He","Wojciech Matusik"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a physics-constrained neural network approach for temperature profile prediction in furnaces. The authors introduce a regularization technique based on the Hottel Zone method and demonstrate its effectiveness across various neural network architectures (MLP, LSTM, KAN, xLSTM)."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How does the proposed regularization technique affect inference time? Is it feasible for real-time applications, particularly in industrial settings? Could you provide specific inference time benchmarks or comparisons to baseline models without the physics-based regularization?\n\n2. Could the physics-constrained regularization approach be adapted for other high-temperature processes or domains where similar physical constraints exist, such as chemical reactors or solar power systems? What modifications, if any, would be required? Could you discuss which aspects of your approach are furnace-specific versus which could be more easily adapted to other domains? \n\n3. The paper distinguishes this approach from typical PINNs. Could the authors elaborate on specific advantages of their approach over traditional PINNs in similar scenarios?\n\n4. Has the proposed model been tested in actual industrial furnace setups? If not, what challenges or adaptations would be expected in deploying it for real-time industrial applications?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"**Originality**  \n   This work presents an innovative application of the Hottel Zone method for regularizing neural networks, which sets it apart from traditional data-driven models. Using a physics-based constraint to enhance temperature prediction accuracy is a novel concept.\n\n**Clarity**  \n   The methodology, especially the integration of the Hottel Zone method, is well-explained, allowing readers to understand how the physics constraint aids in capturing the furnace's temperature dynamics. Experimental design and metrics are clearly presented."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Domain-Specific Limitation**  \n   The method is specifically tailored to furnace temperature prediction, limiting its general applicability to other domains. This specificity reduces its potential impact within a broader range of applications or datasets outside of high-temperature industrial processes.\n\n**Inference Time Analysis**  \n   Although training time with regularization is detailed, the paper lacks explicit performance analysis for inference time on industrial setups. Given the potential complexity of incorporating physical constraints, a clear discussion on inference time efficiency for real-time applications would strengthen the paper's practical value.\n\n**Comparative Baseline**  \n   The paper does not provide a comprehensive comparison with other physics-informed neural networks or hybrid models, which makes it difficult to evaluate the novelty of this approach relative to existing techniques in similar settings.\n\n**Implementation Complexity**  \n   The integration of Hottel Zone constraints may complicate the model deployment process in practical, real-time environments. The paper would benefit from discussing strategies to mitigate the computational overhead associated with this regularization approach."}},"nonreaders":[],"tmdate":1731428320322,"tcdate":1730416647468,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission14235/Reviewer_96th"],"signatures":["ICLR.cc/2025/Conference/Submission14235/Reviewer_96th"],"forum":"hz3NtNpDNv","number":3,"license":"CC BY 4.0","cdate":1730416647468,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission14235/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428320322,"domain":"ICLR.cc/2025/Conference","replyto":"hz3NtNpDNv","id":"RnpniNsPR6","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Hottel Zone method","Physics-Informed Neural Networks","Radiation Heat Transfer","Furnaces"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper investigates a novel approach to improve the temperature profile prediction of furnaces in foundation industries, crucial for sustainable manufacturing. While existing methods like the Hottel Zone model are accurate, they lack real-time inference capabilities. Deep learning methods excel in speed and prediction but require careful generalization for real-world applications. We propose a regularization technique that leverages the Hottel Zone method to make deep neural networks physics-aware, improving prediction accuracy for furnace temperature profiles. Our approach demonstrates effectiveness on various neural network architectures, including Multi-Layer Perceptrons (MLP), Long Short-Term Memory (LSTM) and Kolmogorov-Arnold Networks (KANs). We also discussion the data generation involved."},"_bibtex":{"value":"@misc{\ndutta2024hottel,\ntitle={Hottel Zone Physics-Constrained Networks for Furnaces},\nauthor={Ujjal Kr Dutta and Aldo Lipani and Chuan Wang and Yukun Hu},\nyear={2024},\nurl={https://openreview.net/forum?id=hz3NtNpDNv}\n}"},"title":{"value":"Hottel Zone Physics-Constrained Networks for Furnaces"},"pdf":{"value":"/pdf/e986498612e3968972dbcf517810262453493c21.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"dutta|hottel_zone_physicsconstrained_networks_for_furnaces"},"authorids":{"value":["~Ujjal_Kr_Dutta1","~Aldo_Lipani1","~Chuan_Wang5","~Yukun_Hu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ujjal Kr Dutta","Aldo Lipani","Chuan Wang","Yukun Hu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a way to leverage process systems, which is a key model that can be used to emulate a number of physics models. The authors claim that process models are in general complex and difficult to understand and can also lead to incorrect results. In this paper they propose PAPM (physics-aware proxy model) which has the claimed benefit of including physics priors to accomplish better performance on prediction tasks."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. Paper is mostly well written\n2. Experiments are clear"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While I appreciate the intuitive explanations, process systems are not defined adequately, and this really impedes assessment of the paper. The terms describing this main concept are vague (abstract, introduction and in section 3), and qualitative. Nevertheless, I hope authors can clarify this in the discussion phase (see questions).\n2. It is unclear what is required in training vs. at inference\n3. The experiments seem to be run for one setting (no monte-carlo simulations)\n4. The experiments only consider classical, highly-structured pdes, it is unclear how the proposed model can be used for real-world settings where the dynamics are unknown and may not follow the underlying assumption of (eq.1)"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"### Understanding Process Models: \n\nWhile the contributions seem important it is difficult to understand what process models are. Following are questions which can help authors identify what the reviewer is struggling with, hopefully to help update the paper for a wider audience.\n1. Why are the dynamics/equations of the process model unknown? Isn't it defined by the practitioners?\n2. In relation to 1, it seems that authors consider dynamics which take the form of eq.1, while the exact values that these quantities take are unknown? Is this true? \n3. How are process models different from the proposed model in relation to eq1 and Fig. 3?\n\n### Understanding PAPM:\n4. \\lambda is defined as \"coefficients\" in sec 4.1, but it is unclear how they related to eq 1.\n5. During training the quantities, t, \\lambda, \\Phi_0 etc. are available, but during inference, what all inputs are assumed to be available?\n6. What is the impact of missing quantities on training, can the model still learn?\n7. The structures in Fig 3 (b and c) are still blackboxes, how do these assist in understanding the system as opposed to a process model?\n\n\n### Minor/semantics/other comments:\n1. Why use TSSM for temporal-spatial modeling method (TSSM), TSMM or TSM is more appropriate?\n2. The acronyms DF, CF, IST, and EST can be defined just below eq(1) for clarity.\n3. Decomposing pde as spatial and temporal modules has been studied in PIML. It is important to discuss these similarities in the present work; see Seo 2021. \n\n\nSeo et al. 2021, Physics-aware Spatiotemporal Modules with Auxiliary Tasks for Meta-Learning, IJCAI 2021."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636209927,"tcdate":1698872492789,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2688/Reviewer_hyse"],"signatures":["ICLR.cc/2024/Conference/Submission2688/Reviewer_hyse"],"forum":"UHIKtKzTj7","number":3,"license":"CC BY 4.0","cdate":1698872492789,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2688/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636209927,"domain":"ICLR.cc/2024/Conference","replyto":"UHIKtKzTj7","id":"OIdrvJvmSJ","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Process systems modeling","Physics-informed machine learning","Temporal-spatial stepping method","Out-of-sample generalizability."]},"supplementary_material":{"value":"/attachment/bba5ee17f6691d592c3e248133053c5555cbae92.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Process systems, which play a fundamental role in various scientific and engineering fields, often rely on computational models to capture their complex temporal-spatial dynamics. However, due to limited insights into the intricate physical principles, these models can be imprecise or inapplicable, coupled with a significant computational demand exacerbating inefficiencies. To address these challenges, we propose a physics-aware proxy model (PAPM) to explicitly incorporate partial prior mechanistic knowledge, including conservation and constitutive relations. Additionally, to enhance the inductive biases about strict physical laws and broaden the applicability scope, we introduce a holistic temporal and spatial stepping method (TSSM) aligned with the distinct equation characteristics of different process systems, resulting in better out-of-sample generalization. We systematically compare state-of-the-art pure data-driven models and physics-aware models, spanning five two-dimensional non-trivial benchmarks in nine generalization tasks. Notably, PAPM achieves an average absolute performance improvement of 6.4%, while requiring fewer FLOPs, and only 1% of the parameters compared to the prior leading method, PPNN. Through such analysis, the structural design and specialized spatio-temporal modeling schemes (i.e., TSSM) of PAPM exhibit not only the most balanced trade-off between accuracy and computational efficiency among all methods evaluated, but also an impressive out-of-sample generalization."},"_bibtex":{"value":"@misc{\nliu2024papm,\ntitle={{PAPM}: A Physics-aware Proxy Model for Process Systems},\nauthor={Pengwei Liu and Zhongkai Hao and Xingyu Ren and Hangjie Yuan and Dong Ni},\nyear={2024},\nurl={https://openreview.net/forum?id=UHIKtKzTj7}\n}"},"title":{"value":"PAPM: A Physics-aware Proxy Model for Process Systems"},"pdf":{"value":"/pdf/b922d6af06adf7a8c43b4720ecc76728df0e52a9.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|papm_a_physicsaware_proxy_model_for_process_systems"},"authorids":{"value":["~Pengwei_Liu1","~Zhongkai_Hao1","~Xingyu_Ren2","~Hangjie_Yuan1","~Dong_Ni3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pengwei Liu","Zhongkai Hao","Xingyu Ren","Hangjie Yuan","Dong Ni"]}},"version":2},{"content":{"summary":{"value":"The authors propose Physics-Aware Tensor Field Neural PDE (PA-TFNP), which uses rotation-equivariant Neural Operators on the sphere. It uses a gradient operator based on spherical transforms and uses boundary conditions.\n\nI think the application of the work is very important but in its current state it is missing too many relevant references as well as SOTA DLWP models to compare to."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Please clarify the novelty of the proposed approach compare to Spherical FNO (SFNO).\n2. Please clarify how the approach differes from Method of Lines (MOL) as stated in Section 3 and Neural ODEs.\n3. Is the central difference difference method stable? How is the time-step selected? \n4. Are other mroe advanced numerical methods compared to, e.g., finite volume and finite elements? Also were upwinding finite difference schemes tried?\n5. Do the authors understand hwy ClimODE is performing better on some variables?\n6. Do the authors see limitations of the deterministic approaches and future with probabilistic methods?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Important application of NeuralPDEs for weather\n- Nice emphasis on the importance of physical guidance\n- Nice use of a hybrid model to model finite differences with ML\n- Prediction of several atmospheric variables including geopotential, air temperature and wind speeds.\n- Nice ablation study that shows the importance of rotational equivariance\n- Nice to highlight that physics inductive biases help the long-term predictions.\n- Nice use of ERA5"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The sentences in the abstract are too long. Please break up into simpler sentences\n- SOTA weather models are not compared to, e.g., FourCastNet (Pathak et al., \"Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators\", 2022), FourCastNet v3 Bonev et al., https://arxiv.org/pdf/2507.12144, and importantly Bonev et al., \"Spherical Fourier Neural Operators: Learning Stable Dynamics on the Sphere\", ICML, 2023 which uses a related spherical harmonics basis.\n- See Karlbauer et al., \"Comparing and contrasting deep learning weather prediction backbones on Navier-Stokes and atmospheric dynamics\" for a benchmarking study comparing SOTA DLWP models which should be compared to as well as cited. This work also shows that computing the solutions on a HealPIX mesh can also be beneficial Karlbauer et al., \"Advancing Parsimonious Deep Learning Weather Prediction Using the HEALPix Mesh\", Journal of Advances in Modeling Earth Systems, 2024.\n- Lack of references also in Physics-Informed ML is much broader than PINNs. See Krishnapriyan et al., \"Characterizing possible failure modes in physics-informed neural networks\", NeurIPS on 2022 on limitations of PINNs and several hard-constrained works, e.g., Negiar et al., \"Learning differentiable solvers for systems with hard constraints\", ICLR, 2023, Chalapathi et al., \"Scaling physics-informed hard constraints with mixture-of-experts\", ICLR 2024,  Hansen et al., \"Learning Physical Models that Can Respect Conservation Laws\", ICML 2024, Mouli et al., \"Using uncertainty quantification to characterize and improve out-of-domain learning for pdes\", ICML, 2024, Saad et al., \"Guiding continuous operator learning through Physics-based boundary constraints\", ICLR, 2023, Utkarsh, U., \"End-to-End Probabilistic Framework for Learning with Hard Constraints\", https://arxiv.org/pdf/2506.07003?, 2025.\n- Saad et al., \"Guiding continuous operator learning through Physics-based boundary constraints\" is important reference for enforcing Neumann BC in Neural Operators.\n- The numerical discretization in this work is taken for granted and a simple finite difference scheme is used, which may not be accurate enough for the Navier Stokes equation. In particular, the spatial discretization of PDEs is very important.\n-With inputting finite differences, the ML problem simplifies to solving ODEs instead of PDEs. \n- The results in Table 1 are mixed with ClimODE sometimes performing better."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926122359,"tcdate":1762409015671,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15906/Reviewer_9yqv"],"signatures":["ICLR.cc/2026/Conference/Submission15906/Reviewer_9yqv"],"forum":"beV5wMTRIq","number":4,"license":"CC BY 4.0","cdate":1762409015671,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15906/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926122359,"domain":"ICLR.cc/2026/Conference","replyto":"beV5wMTRIq","id":"LcY0dpGmKP","forumContent":{"TLDR":{"value":"We introduce PA-TFNP, a physics-aware tensor-field neural PDE framework for weather and climate forecasting."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Weather Prediction","Neural ODE","Physics-Informed Machine Learning","Tensor Field Neural","Partial Differential Equations"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Climate and weather prediction has traditionally relied on computationally demanding numerical simulations grounded in atmospheric physics, yet deep-learning approaches are emerging as transformative alternatives. Existing methods, however, are often purely data-driven and physics-agnostic, overlooking essential physical principles and struggling to generalize. To address these challenges, we present the Physics-Aware Tensor Field Neural PDE (PA-TFNP), a forecasting framework that embeds rotation-equivariant tensor-field neural operators directly on the sphere, couples them with a numerically rigorous gradient operator based on spherical transforms and physically consistent boundary treatment, and augments the learned dynamics with diffusion terms derived from the atmospheric primitive equations. These innovations enable our model to achieve superior performance through strict physical fidelity and efficient learning. The proposed PA-TFNP achieves state-of-the-art performance in global and regional weather prediction, outperforming ClimODE by 78.92% on global hourly data with a comparable number of parameters."},"_bibtex":{"value":"@misc{\ncho2025physicsaware,\ntitle={Physics-Aware Tensor Field Neural {PDE} for Climate and Weather Prediction},\nauthor={Namkyeong Cho and Sung Woong Cho and Youngjoon Hong and Hyung Ju Hwang and Jae Yong Lee and Hwijae Son},\nyear={2025},\nurl={https://openreview.net/forum?id=beV5wMTRIq}\n}"},"title":{"value":"Physics-Aware Tensor Field Neural PDE for Climate and Weather Prediction"},"pdf":{"value":"/pdf/8b2fbbffd8e48eb20194e8b54fad719f70c564d5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"cho|physicsaware_tensor_field_neural_pde_for_climate_and_weather_prediction"},"authorids":{"value":["~Namkyeong_Cho1","~Sung_Woong_Cho1","~Youngjoon_Hong1","~Hyung_Ju_Hwang1","~Jae_Yong_Lee2","~Hwijae_Son1"]},"authors":{"value":["Namkyeong Cho","Sung Woong Cho","Youngjoon Hong","Hyung Ju Hwang","Jae Yong Lee","Hwijae Son"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the problem of fusing information from various modalities for video retrieval. The authors propose CLAMR, a framework that applies a late-interaction mechanism, inspired by text-retrieval models like ColBERT, to a multimodal context encompassing video, audio, OCR, and metadata. The core idea is to jointly encode all modalities within a single vision-language model (VLM) to create contextualized representations, and then perform token-level similarity matching. To train the model to focus on relevant modalities, the paper introduces a modality-aware contrastive loss and a new large-scale synthetic dataset, MULTIVENT 2.0++, created using an LLM to generate modality-specific queries. The experiments show that CLAMR achieves strong performance, particularly on this new synthetic benchmark, outperforming several unimodal and multimodal baselines."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How can you ensure that the substantial performance gains on MULTIVENT 2.0++ are not primarily due to the model learning the specific artifacts of the Gemma-3-27b-it query generator? Have you considered a cross-validation experiment where you train on queries generated by one LLM and test on queries generated by a completely different LLM?\n\n\n2. Could you provide a clearer justification for the discrepancy between the training objective and the inference objective? If the goal is to teach the model modality selection, why not use a scoring function at inference that also reflects this selection process? Does this design choice not suggest that the model isn't actually performing explicit modality selection at test time?\n\n3. The paper's baselines for multimodal fusion include simple averaging and a \"hard\" router. A potentially stronger and more relevant baseline would be a \"soft-routing\" or attention-based mechanism that learns to dynamically weight the similarity scores from different modalities based on the query. How would you expect CLAMR to perform against such a baseline?\n\n4. Given the high indexing cost of late-interaction and the more moderate performance gains on MSR-VTT, what is the compelling practical argument for deploying CLAMR over a simpler, highly optimized dual-encoder model that is fine-tuned on the same comprehensive data?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The paper tackles a well-recognized and challenging problem in multimodal retrieval—how to effectively combine signals from diverse and potentially noisy sources without performance degradation.\n\n2. The application of late-interaction mechanisms from the text domain to a complex multimodal video scenario is a logical and interesting direction. It provides an alternative to more common early-fusion or simple late-fusion (score averaging) techniques.\n\n3. The authors make a notable effort to address the scarcity of suitable training data by generating a large-scale synthetic dataset with modality-targeted queries. This resource could potentially benefit future research in the area."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The core technical contribution can be viewed as an application of the existing ColBERT architecture to a multimodal setting using a standard VLM backbone. While the engineering is non-trivial, the conceptual novelty is somewhat incremental, as it primarily combines and adapts existing components rather than introducing a fundamentally new retrieval paradigm.\n\n2. The most impressive results (e.g., +25.6 nDCG@10) are reported on the authors' own synthetic dataset, MULTIVENT 2.0++. This raises significant concerns about evaluation validity. The model might be overfitting to the specific patterns, vocabulary, and artifacts of the LLM used for query generation, rather than learning a truly generalizable retrieval capability. The performance gains on the established, human-annotated MSR-VTT benchmark are far more modest, which may indicate that the practical impact on real-world queries is less significant than claimed.\n\n3. The paper proposes a modality-aware loss (LImw) for training, which encourages the model to identify the single most relevant modality. However, at inference, a different, holistic scoring function (LIcontext) that aggregates scores across all modalities is used. This disconnect between the training objective and the inference procedure weakens the central claim of \"dynamic modality selection.\" It's unclear whether the model is truly learning to select modalities or if the performance gain is simply a byproduct of a more complex training objective that acts as a form of regularization.\n\n4.The late-interaction approach introduces significant computational and memory overhead during the offline indexing phase compared to standard dual-encoder models. Given that the performance improvement on MSR-VTT is not as dramatic as on the synthetic dataset, the practical utility of CLAMR is questionable for applications where efficiency is a key concern."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926097510,"tcdate":1761567892450,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15879/Reviewer_f3Zs"],"signatures":["ICLR.cc/2026/Conference/Submission15879/Reviewer_f3Zs"],"forum":"AXFuBS3ujj","number":2,"license":"CC BY 4.0","cdate":1761567892450,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15879/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926097510,"domain":"ICLR.cc/2026/Conference","replyto":"AXFuBS3ujj","id":"tKxkfN6Vll","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["multimodal retrieval","text-video retrieval","RAG"]},"supplementary_material":{"value":"/attachment/30f3f46b897718fb535393e45c6585c6c2f48bcf.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Online video content is richly multimodal: a single video might blend vision, speech, ambient audio, and on-screen text. Conventional retrieval systems typically treat these modalities as independent retrieval sources, which can lead to noisy and subpar results.\nIn this work, we explore multimodal video content retrieval, where relevance can be scored from a single modality or jointly across multiple modalities. Consequently, an effective retriever must dynamically determine which modality (or set of modalities) best address a given query. We introduce CLaMR, a multimodal, late-interaction retriever that jointly indexes four modalities: video frames, transcribed speech, on-screen text, and metadata. CLaMR jointly encodes all modalities within a unified multimodal backbone for improved contextualization and is trained to enhance dynamic modality selection via two key innovations. First, to overcome the lack of suitable training data, we introduce MultiVent 2.0++, a large-scale synthetic dataset built on MultiVent 2.0 (a collection of event-centric videos in various languages paired with English queries) with modality-targeted queries to teach modality selection. Next, we propose a modality-aware contrastive loss that trains the model on both a standard contrastive objective and an objective for learning correct modality usage. On the test sets of MultiVent 2.0++ and MSRVTT, we observe that conventional aggregation strategies, such as averaging similarities for baseline retrievers, often degrade performance by introducing noise from irrelevant modalities. In contrast, CLaMR consistently outperforms existing retrievers: on MultiVent 2.0++, CLaMR improves nDCG@10 by 25.6 points over the best-performing single-modality retriever and by 35.4 points over the best-performing multi-modality retriever. We illustrate the downstream utility of CLaMR with experiments on long-video QA, where it improves performance by 3.50% over LanguageBind on Video-MME and 1.42% over dense frame sampling on LongVideoBench."},"_bibtex":{"value":"@misc{\nwan2026clamr,\ntitle={{CL}a{MR}: Contextualized Late-Interaction for Multimodal Content Retrieval},\nauthor={David Wan and Han Wang and Elias Stengel-Eskin and Jaemin Cho and Mohit Bansal},\nyear={2026},\nurl={https://openreview.net/forum?id=AXFuBS3ujj}\n}"},"title":{"value":"CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval"},"pdf":{"value":"/pdf/64dc1ed7cf0393bc3ae6e960de0123892d1a795c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wan|clamr_contextualized_lateinteraction_for_multimodal_content_retrieval"},"authorids":{"value":["~David_Wan1","~Han_Wang9","~Elias_Stengel-Eskin1","~Jaemin_Cho1","~Mohit_Bansal2"]},"authors":{"value":["David Wan","Han Wang","Elias Stengel-Eskin","Jaemin Cho","Mohit Bansal"]}},"version":2},{"content":{"summary":{"value":"The paper attempts to model the meteorological dynamics with physics guided deep leaning method.  Authors propose a physics guided deep learning framework, which can be combined with existing deep learning networks. Experiments are conducted on real world datasets to validate the model."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. First of all, weather forecasting is a very important problem and should get more attention from the ai community to defending the climate change.\n2. Generally, I like the idea of combining physics mechanism with deep learning to improve the performance and generalization ability of ai methods. The proposed method seems reasonable. \n3. Experiments are conducted on the ERA5 dataset, which is one of the most high-quality weather data. The results show that the proposed method can obtain performance gain for both downscaling and forecasting tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The writing of the paper could be further improved, especially the use of symbols. For example, what is (u, v, w) in figure 1, is Q the same as Q_pi? What is epsilon?\n2. One the mentioned advantage of physics guide is physical consistent. However, it is not clear how to measure physical consistent, do you mean the analysis in sec 4.3?\n3. There are some important recent works and background information missed in the related work part. Section 2.1 missed some recent progress in weather forecasting such as GraphCast, ClimaX, and FengWu. The relate work part also lacks discussions about PINNs.\n4. It is also not clear to me what is the core technique contribution of this paper when aligning it to the broad PINN family. \n5. Regarding the experiments, it is hard to compare this work with existing deep learning works since it is conducted in the special region. It would be good to present the results in the full ERA5 dataset or a smaller resolution (e.g., weatherbench) if the resource is limited. \n6. The assumption that deep learning model is accurate enough does not hold in this paper.\n7. It is not clear about the detailed train, validation, and test settings. There is also no code shared for reproduce ability checking."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"please check the weakness part."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636080624,"tcdate":1699086564586,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1523/Reviewer_HSXq"],"signatures":["ICLR.cc/2024/Conference/Submission1523/Reviewer_HSXq"],"forum":"QMkYEau02q","number":4,"license":"CC BY 4.0","cdate":1699086564586,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1523/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636080624,"domain":"ICLR.cc/2024/Conference","replyto":"QMkYEau02q","id":"W1vZ5Cq2iS","forumContent":{"TLDR":{"value":"We propose a physics-guided learning framework to capture the nonlinear dynamics of meteorology and align deep learning models with the corrected physical mechanism to improve generalization."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Weather Forecasting","Weather Downscaling","Spatiotemporal Modeling","Partial Differential Equations","Physics-Guided Learning"]},"supplementary_material":{"value":"/attachment/567c290bd8ad3d6c0e4029accff29999f21cf00c.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Weather forecasting is of paramount importance for a myriad of societal and scientific applications. Traditionally, numerical weather prediction (NWP) methods based on physical principles are computationally intensive and can struggle with the inherent complexity of atmospheric dynamics. Recently, deep learning techniques have shown promise in weather prediction, but the long-term generalization and physical consistency of pure data-driven approaches remain challenging. In this paper, we introduce a novel physics-guided approach for numerical weather prediction that combines the strengths of both physical mechanism and deep learning, namely PhyDL-NWP. Our method can capture the nonlinear dynamics of meteorology and align deep learning models with the underlying physical mechanism to improve generalization. Extensive experiments on real-world weather datasets show that our model can significantly improve the performance of deep learning methods in a wide range of tasks from forecasting to downscaling."},"_bibtex":{"value":"@misc{\nluo2024physicsguided,\ntitle={Physics-Guided Learning of Meteorological Dynamics for Weather Forecasting and Downscaling},\nauthor={Yingtao Luo and Shikai Fang and Binqing Wu and Qingsong Wen and Liang Sun},\nyear={2024},\nurl={https://openreview.net/forum?id=QMkYEau02q}\n}"},"title":{"value":"Physics-Guided Learning of Meteorological Dynamics for Weather Forecasting and Downscaling"},"pdf":{"value":"/pdf/a7920e01cfa49437681f9a44d58380cc7d960ddc.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"luo|physicsguided_learning_of_meteorological_dynamics_for_weather_forecasting_and_downscaling"},"authorids":{"value":["~Yingtao_Luo1","~Shikai_Fang2","~Binqing_Wu1","~Qingsong_Wen2","~Liang_Sun2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yingtao Luo","Shikai Fang","Binqing Wu","Qingsong Wen","Liang Sun"]}},"version":2},{"content":{"summary":{"value":"This paper proposes IR4Net, a physics-informed neural network for non-contact optical side-channel attacks on isolated screens. By leveraging wall-reflected speckle patterns, the method reconstructs screen content using a physically-regularized inversion module and a semantic completion module. Experiments on four\nscene categories, show strong performance under various perturbations such as low light, noise, and material changes."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Refer to the weakness section."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"No Ethics Concerns."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- The paper has thoroughly discussed the related work. \n- Physics-guided design: Embeds radiative transfer modeling into learning, improving stability and inversion accuracy.\n- Strong performance: Outperforms existing methods across PSNR, SSIM, RMSE, and LPIPS, under various conditions (e.g., low brightness, noise).\n- Thorough evaluation: Includes ablation, noise tests, luminance robustness, material variation, and temporal consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The evaluation is primarily conducted on synthetic datasets. It remains unclear how well the method generalizes to real-world environments with complex conditions such as multipath reflections or dynamic lighting. Could the authors provide results or discussion in such scenarios?\n\n2. Analysis on inference time, computational cost, or resource requirements should be added, which limits understanding of its practicality in real-world attacks. Please include runtime metrics or efficiency evaluations.\n\n3. The proposed method may be sensitive to wall reflectance properties, but there is no quantitative analysis of this dependency. How robust is the method under varying surface materials or reflectance levels?\n\n4. The operational feasibility of the attack is not fully discussed. What are the requirements in terms of camera resolution, alignment precision, or environmental constraints?\n\n5. No clear defense strategies are proposed. It would be valuable to include potential countermeasures or mitigation suggestions to guide practical security considerations.\n\n6. Visual quality needs improvement: Teaser, pipeline, and visual examples appear blurry and should be presented with higher clarity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927283587,"tcdate":1761993212517,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17374/Reviewer_J4Kx"],"signatures":["ICLR.cc/2026/Conference/Submission17374/Reviewer_J4Kx"],"forum":"evepIXBxL8","number":3,"license":"CC BY 4.0","cdate":1761993212517,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17374/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927283587,"domain":"ICLR.cc/2026/Conference","replyto":"evepIXBxL8","id":"a2HlxR87qL","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Computer Vision","Deep Learning","Side Channel Attack","Information Security","Information Theft"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Noncontact exfiltration of electronic screen content poses a security challenge, with side-channel incursions as the principal vector. We introduce an optical projection side-channel paradigm that confronts two core instabilities: (i) the near-singular Jacobian spectrum of projection mapping breaches Hadamard stability, rendering inversion hypersensitive to perturbations; (ii) irreversible compression in light transport obliterates global semantic cues, magnifying reconstruction ambiguity. Exploiting passive speckle patterns formed by diffuse reflection, our Irradiance Robust Radiometric Inversion Network (IR$^4$Net) fuses a Physically Regularized Irradiance Approximation (PRIrr‑Approximation), which embeds the radiative transfer equation in a learnable optimizer, with a contour-to-detail cross-scale reconstruction mechanism that arrests noise propagation. Moreover, an Irreversibility Constrained Semantic Reprojection (ICSR) module reinstates lost global structure through context-driven semantic mapping. Evaluated across four scene categories, IR$^4$Net achieves fidelity beyond competing neural approaches while retaining resilience to illumination perturbations."},"_bibtex":{"value":"@inproceedings{\nzheng2026physicallyguided,\ntitle={Physically-Guided Optical Inversion Enable Non-Contact Side-Channel Attack on Isolated Screens},\nauthor={Zhiwen Zheng and Yuheng Qiao and Xiaoshuai Zhang and Zhao Huang and Tao Zhang and Huiyu Zhou and Shaowei Jiang and Jin Liu and Wenwen Tang and Xingru Huang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=evepIXBxL8}\n}"},"title":{"value":"Physically-Guided Optical Inversion Enable Non-Contact Side-Channel Attack on Isolated Screens"},"pdf":{"value":"/pdf/99a232d6166f3dd0056102d8de06270e22ac5ac4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zheng|physicallyguided_optical_inversion_enable_noncontact_sidechannel_attack_on_isolated_screens"},"authorids":{"value":["~Zhiwen_Zheng1","~Yuheng_Qiao1","~Xiaoshuai_Zhang2","~Zhao_Huang2","~Tao_Zhang5","~Huiyu_Zhou3","~Shaowei_Jiang1","~Jin_Liu22","~Wenwen_Tang1","~Xingru_Huang1"]},"authors":{"value":["Zhiwen Zheng","Yuheng Qiao","Xiaoshuai Zhang","Zhao Huang","Tao Zhang","Huiyu Zhou","Shaowei Jiang","Jin Liu","Wenwen Tang","Xingru Huang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a parametric SDF for dynamic surface reconstruction. The parametric SDF is built basd on SDF and temporally varying signal modeling with basis functions. In implementation, this parametric function is modeled by a grid and the appearance is modeled by a 4D hash grid. With differentiable rendering, the parametric SDF and appearance are optimized. In experiments, the work curates a new synthetic benchmark, SynthoMotion-360. The experimental results on SynthoMotion-360, DiVa-360 and CMU Panoptic Studio show that the method achieves promoising results, with better fidelity and coherence."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Many works focus on dynamic scene reconstruction and they evaluated their methods on some human datasets. It is better to also evaluate the proposed methods on human dataset, like ZJU-Mocap and PeopleSnapshot.\n2. It seems that NeuS2 is a SOTA SDF-based method that can reconstruct dynamic scenes. It is better to compare with this method.\n3. Can you discuss training efficiency for different methods, including NeuS2.\n4. For the detail reconstruction, is it possible to increase the grid resolution to improve the reconstruciton?\n5. Can you show some reconstruciton examples under challenging lighting conditions (Line103)?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The method introduces a parametric SDF for dynamic surface reconstuction, which leverages the advantages of SDF and temporlly varying signal modeling with basis functions.\n2. The method distanges geometry and appearance to optimize the parametric SDF and material properties. \n3. The work curates a new synthetic benchmark to evaluate the dynamic surface reconstruction.\n4. The method achieves promoising results for dynamic surface reconstruction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The quantitative surface reconsutruction evaluation is only test on the synthetic benchmark. \n2. For more complex scenes, like the scenes on CMU Panoptic Studio dataset, the method fails to reconstruct high-fidelity surfaces.\n3. The method cannot recover complex geometris and details."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918845328,"tcdate":1761970845098,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6455/Reviewer_XrhY"],"signatures":["ICLR.cc/2026/Conference/Submission6455/Reviewer_XrhY"],"forum":"mhHlsuGNRW","number":4,"license":"CC BY 4.0","cdate":1761970845098,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6455/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918845328,"domain":"ICLR.cc/2026/Conference","replyto":"mhHlsuGNRW","id":"JWCY6wkncK","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["3D Reconstruction","Dynamic Reconstruction","Mesh Extraction"]},"supplementary_material":{"value":"/attachment/18ea179a1dcb0a24e3253d0114c2bafc07de7477.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Reconstructing high-fidelity, temporally coherent surfaces of dynamic scenes remains a critical challenge in computer vision. While recent methods excel at novel view synthesis, they often fail to recover accurate geometry, yielding noisy or temporally inconsistent meshes that are suboptimal for downstream applications such as simulation or editing. In this work, we introduce a new paradigm for dynamic surface reconstruction based on a parametric Signed Distance Function ({\\nameshort}). Our key insight is to generalize static SDF fields—where each spatial point stores a constant value—into time-dependent parametric curves, where each curve defines a temporally evolving SDF trajectory. Such a parametric SDF modeling provides a principled way to capture complex temporal variations, naturally enforcing smoothness and continuity in shape dynamics. At each timestamp, a static SDF field can be queried from {\\nameshort} and converted into an explicit surface mesh via differentiable iso-surfacing. By rendering these meshes with a physically based differentiable renderer, we optimize the underlying parametric curves end-to-end against 2D image observations. Our framework produces high-fidelity, temporally coherent surfaces and inherently disentangles geometry, material, and lighting from multi-view videos. It robustly reconstructs geometry under large-scale motions and resolves appearance ambiguities caused by challenging lighting and occlusions. Experiments on both synthetic and real-world scenes demonstrate that our method achieves state-of-the-art geometric accuracy and temporal consistency, delivering delicate meshes that surpass prior work."},"_bibtex":{"value":"@misc{\ngao2026parametric,\ntitle={Parametric {SDF} for Dynamic Surface Reconstruction},\nauthor={Chong Gao and Kai Ye and Qiyu Dai and Yiming Shao and Qiong Zeng and Ding Liang and Yan-Pei Cao and Guanbin Li and Wenzheng Chen},\nyear={2026},\nurl={https://openreview.net/forum?id=mhHlsuGNRW}\n}"},"title":{"value":"Parametric SDF for Dynamic Surface Reconstruction"},"pdf":{"value":"/pdf/a242314267bdfdc588119307a93a30be37b05c2a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"gao|parametric_sdf_for_dynamic_surface_reconstruction"},"authorids":{"value":["~Chong_Gao1","~Kai_Ye4","~Qiyu_Dai1","~Yiming_Shao2","~Qiong_Zeng1","~Ding_Liang1","~Yan-Pei_Cao1","~Guanbin_Li1","~Wenzheng_Chen1"]},"authors":{"value":["Chong Gao","Kai Ye","Qiyu Dai","Yiming Shao","Qiong Zeng","Ding Liang","Yan-Pei Cao","Guanbin Li","Wenzheng Chen"]}},"version":2},{"content":{"summary":{"value":"The purpose of this paper is to accurately restore lighting in indoor scenes using a limited number of observed views, enabling realistic lighting effects for applications such as virtual object insertion and lighting editing.\n\nTo achieve this, a neural-field-based parametric lighting representation is used. This model predicts the intensity and direction of lighting at each location in 3D space and enhances restoration accuracy by limiting possible light source locations through semantic embeddings. Additionally, differentiable path tracing is used to optimize alignment with observed images, and voxel-based sampling is introduced to reduce noise during path tracing and improve efficiency.\n\nThe dataset is an extended version of 3D-Front, incorporating Lambertian materials and texture-based lighting to create synthetic data under diverse indoor lighting conditions, which is used for training and evaluation."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Obtaining Complete Geometry and Material Information in Realistic Scenarios\n\nIn realistic scenarios, obtaining complete geometry and material information for a space requires precise Lidar scanning or highly dense photographic data. Could the authors propose a plausible scenario in which the experiments in this paper could be reproduced in real-world settings with only sparse views to reconstruct the scene’s geometry and material properties?\n\n2. Integration with Modern Radiance Field and Semantic Segmentation Techniques\n\nTo achieve a more realistic scene reconstruction, could the authors suggest a scenario in which modern radiance field techniques and compatible semantic segmentation methods (e.g., OpenNeRF) could be integrated with the methodology proposed in this paper?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Tackling the Challenge of Lighting Reconstruction in Complex Indoor Scenes \n\nIndoor scenes are particularly challenging for lighting reconstruction due to the variety of light sources, complex geometries, and multiple reflections involved. This paper’s approach appears to address these complexities by enabling accurate lighting reconstruction even with limited observed views, which could potentially broaden the scope for realistic lighting manipulation in intricate indoor environments.\n\n2. Thoughtful Integration of Multiple Techniques \n\nThe paper combines neural field-based lighting representation, semantic embedding, differentiable path tracing, and voxel-based sampling in what seems to be an efficient and well-balanced way. This integration could be a strength as it leverages the individual advantages of each technique, allowing for both high-quality reconstruction and computational efficiency in complex lighting conditions. The combination might help enhance the model’s accuracy and robustness, making it better suited to the intricate lighting variations characteristic of indoor scenes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Reliance on Synthetic Datasets\n\nThis paper primarily uses synthetic datasets such as 3D-Front to train and evaluate the model. While synthetic data is useful for demonstrating the feasibility of lighting reconstruction in complex indoor scenes, it may not fully capture the variability and complex material characteristics of real-world environments. Incorporating real-world datasets (e.g., ScanNet, Replica) could help validate the model's generalization performance in actual settings.\n\n2. Assumption of Complete Geometry Conflicts with Limited View Claims\n\nThis paper claims to enable lighting reconstruction with limited views, but it is based on the assumption that precise geometry of the entire scene is already available. In an approach intended to work with limited views, assuming pre-existing complete geometry may be logically inconsistent. The notion of limited views generally implies incomplete scene information; however, by assuming full geometry is already provided, the paper may face limitations in fully evaluating the effectiveness of reconstruction with truly limited views."}},"nonreaders":[],"tmdate":1731428386950,"tcdate":1730658217277,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6948/Reviewer_skt1"],"signatures":["ICLR.cc/2025/Conference/Submission6948/Reviewer_skt1"],"forum":"fy4rCv3s5i","number":2,"license":"CC BY 4.0","cdate":1730658217277,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6948/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428386950,"domain":"ICLR.cc/2025/Conference","replyto":"fy4rCv3s5i","id":"GiQngvwwwL","forumContent":{"TLDR":{"value":"We introduce a neural representation for emissive light sources and learn a prior over them using a new dataset to better constrain the sparse-view inverse rendering problem"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["lighting representation","prior learning","neural field","3D","computer graphics"]},"supplementary_material":{"value":"/attachment/96259d4186e4f110f6a216de6596bfe51fdbbd78.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We introduce Neural Lighting Priors, a learned surface emission model for indoor scenes. Given multi-view observations as well as the geometry of a scene, we decouple spatially varying lighting and material parameters. Existing inverse rendering methods typically use hand-crafted emission models or require a large number of views to better constrain the highly ambiguous appearance decomposition task. We aim to overcome these limitations by introducing an expressive learned parametric emission model and utilizing semantic information to sufficiently constrain the optimization, thus allowing us to infer light sources, even if they are not visible in the observations. We model the emitted radiance with a neural field parameterized by the emitting direction and a local latent code stored in a voxel grid. At test time, we fit the local latent codes to the scene using differentiable path tracing, optimizing the reconstruction loss. Our reconstruction allows us to insert virtual objects in a scene and gives us control over the emitters to change their emission color and intensity. Thanks to the learned 3D prior, our method requires fewer views than state-of-the-art relighting methods, gives more control, and also improves the relighting quality."},"_bibtex":{"value":"@misc{\nkocsis2025neural,\ntitle={Neural Lighting Priors for Indoor Scenes},\nauthor={Peter Kocsis and Vincent Sitzmann and Matthias Nie{\\ss}ner},\nyear={2025},\nurl={https://openreview.net/forum?id=fy4rCv3s5i}\n}"},"title":{"value":"Neural Lighting Priors for Indoor Scenes"},"pdf":{"value":"/pdf/d32c2a8963da7130d8d7ff52dec69ae88f2ec62c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"kocsis|neural_lighting_priors_for_indoor_scenes"},"authorids":{"value":["~Peter_Kocsis1","~Vincent_Sitzmann1","~Matthias_Nießner2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Peter Kocsis","Vincent Sitzmann","Matthias Nießner"]}},"version":2},{"content":{"summary":{"value":"The paper introduces PhysUniBench, a large-scale multimodal benchmark for undergraduate-level physics reasoning. It contains 3,304 problems across 8 sub-disciplines, each paired with a diagram image, with both open-ended and multiple-choice formats and five difficulty levels calibrated via model-in-the-loop rollouts. The authors evaluate a range of MLLMs. An ablation replacing images with auto-generated captions improves accuracy, suggesting current models’ weaknesses in physics-diagram understanding. The benchmark, construction pipeline, and evaluation scripts are released for reproducibility."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Where do the ground-truth answers come from? The authors mention that they evaluate against a “known gold solution” in the Appendix, then fall back to SymPy equivalence or an LLM judge (GPT-4/4o). It seems like the paper implies the gold solutions are from the source materials (textbooks/exams/exercise sets), but it doesn’t spell out a per-item list, or explain if there are specific cases where they had to involve humans for ground-truth answers\n- Are there any overlap tests that were done to see how similar physics-related benchmarks (e.g., UGPhysics, PhysicsArena, PHYX, PhysReason) are related to PhysUniBench? How did the authors ensure that there were no duplicate questions (within PhysUniBench itself and also compared to other benchmarks)?\n- Beyond just describing the accuracy results from the experiment tables, it would be nice to know what exactly caused the mistakes in each subcategory. For example, is it diagram misread? Using the wrong formula? Algebra slip? etc."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The benchmark targets undergrad physics with calibrated difficulty and multilingual EN/ZH support\nThe paper uses a model-in-the-loop curation to remove trivially solvable items and to stratify difficulty; this is a thoughtful twist on dataset construction\n- There is a clear dataset stats and coverage across different subfields, with balanced difficulty bins\n- The caption ablation is a smart diagnostic revealing a likely bottleneck in visual/diagram understanding for physics\n- The results tables disaggregate by sub-discipline and difficulty, highlighting where models fail"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Difficulty calibration (Qwen2.5-VL rollouts) and LLM judging with GPT-4o may inject model-specific biases; it’s unclear how sensitive results are to the choice of judge/rollout model or to prompt templates. A cross-judge analysis (or human spot-checks) would strengthen validity\n- Potential data contamination: It is not clear where exactly the datasets are from. The paper mentioned that problems came from textbooks/exams/competitions, but many may already appear online. The paper does not quantify near-duplicate overlap with existing pretraining/test corpora (e.g., UGPhysics, PhysicsArena, PhysReason), risking optimistic or inconsistent estimates\n- The “caption > image” result is interesting, but captions are generated by GPT-4o -- a strong prior that may inject solution-relevant structure (naming vectors/relations), not just a faithful description. In other words, the result doesn't necessarily prove that \"text > images\" for physics reasoning; it might just prove \"having GPT-4o pre-analyze the problem helps.\"\n- As the paper claims that having multilingual evaluation is one of their main contributions (what separates their benchmark from other similar benchmarks), the English/Chinese evaluation results should be included in the main sections of the paper"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915994272,"tcdate":1761947127148,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2029/Reviewer_87P4"],"signatures":["ICLR.cc/2026/Conference/Submission2029/Reviewer_87P4"],"forum":"TqgPgCvBnF","number":3,"license":"CC BY 4.0","cdate":1761947127148,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2029/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915994272,"domain":"ICLR.cc/2026/Conference","replyto":"TqgPgCvBnF","id":"JeJ2yfmiER","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics","reasoning","benchmark","large language model","multi-modal large language model"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Physics problem-solving is a challenging domain for large AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Existing evaluations fail to capture the full breadth and complexity of undergraduate physics, whereas this level provides a rigorous yet standardized testbed for pedagogically relevant assessment of multi-step physical reasoning. To this end, we present PhysUniBench, a large-scale multimodal benchmark designed to evaluate and improve the reasoning capabilities of multimodal large language models (MLLMs) specifically on undergraduate-level physics problems. PhysUniBench consists of 3,304 physics questions spanning 8 major sub-disciplines of physics, each accompanied by one visual diagrams. The benchmark includes both open-ended and multiple-choice questions, systematically curated and difficulty-rated through an iterative model-in-the-loop process. The benchmark's construction involved a rigorous multi-stage process, including multiple roll-outs, expert-level evaluation, automated filtering of easily solved problems, and a nuanced difficulty grading system with five levels. Through extensive experiments, we observe that current state-of-the-art models encounter substantial challenges in physics reasoning. For example, GPT-5 achieves only about 53.7% accuracy in the proposed PhysUniBench. These results highlight that current MLLMs struggle with advanced physics reasoning, especially on multi-step problems and those requiring precise diagram interpretation. By providing a broad and rigorous assessment tool, PhysUniBench aims to drive progress in AI for Science, encouraging the development of models with stronger physical reasoning, problem-solving skills, and multimodal understanding. The benchmark and evaluation scripts are available at https://anonymous.4open.science/r/PhysUniBenchmark-5784."},"_bibtex":{"value":"@misc{\nwang2026physunibench,\ntitle={PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level},\nauthor={Lintao Wang and Encheng Su and Jiaqi Liu and Pengze Li and Jiabei Xiao and Wenlong Zhang and Xi Chen and Yuan Meng and LEI BAI and Wanli Ouyang and SHIXIANG TANG and Aoran Wang and Xinzhu Ma},\nyear={2026},\nurl={https://openreview.net/forum?id=TqgPgCvBnF}\n}"},"title":{"value":"PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level"},"pdf":{"value":"/pdf/006ead177bfaf2dd134dbde4641e5de03d19d8e8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|physunibench_a_multimodal_physics_reasoning_benchmark_at_undergraduate_level"},"authorids":{"value":["~Lintao_Wang1","~Encheng_Su1","~Jiaqi_Liu7","~Pengze_Li3","~Jiabei_Xiao1","~Wenlong_Zhang3","~Xi_Chen20","~Yuan_Meng2","~LEI_BAI1","~Wanli_Ouyang1","~SHIXIANG_TANG1","~Aoran_Wang1","~Xinzhu_Ma1"]},"authors":{"value":["Lintao Wang","Encheng Su","Jiaqi Liu","Pengze Li","Jiabei Xiao","Wenlong Zhang","Xi Chen","Yuan Meng","LEI BAI","Wanli Ouyang","SHIXIANG TANG","Aoran Wang","Xinzhu Ma"]}},"version":2},{"content":{"summary":{"value":"This paper proposes two ways to incorporate physical priors into generative models: (1) they inject distributional priors by choosing the proper equivariant models and (2) they incorporate physical feasibility priors by decomposing nonlinear constraints into elementary cases. The key is to identify elementary cases where Jensen's gap can be omitted and decompose complicated constraints into elementary cases. In general, this paper is a good contribution to the AI4Science community, where many scientific priors are available and should be incorporated into learning systems."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"* Why does performance improve only incrementally after adding priors?\n* In Equation (4) and line 205, are there typos? \\nabla x -> \\nabla_x?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The motivation is clear - this paper is a clear contribution to the AI4Science (esp AI4Physics) community where data-driven models should be combined with scientific inductive biases. Although physics-informed learning has been mainstream in scientific machine learning, physics-informed learning in the context of generative modeling is relatively new. \n* This paper is both theoretically (supported by theorems) and practically sound (supported by experiments). The idea of decomposing a complex constraint into elementary cases is a good strategy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Empirical improvement is marginal (Table 1, 2, 3).\n* This framework is useful when (1) we don't understand the system fully but (2) we know some partial information (existence of energy etc). However, all the examples in the paper are synthetic, so we know both the underlying equation and energy but \"pretend\" that we don't know about the equation. I understand this is only a proof of concept, but a more realistic example would greatly strengthen the paper."}},"nonreaders":[],"tmdate":1731427301606,"tcdate":1729965878660,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission688/Reviewer_2sjz"],"signatures":["ICLR.cc/2025/Conference/Submission688/Reviewer_2sjz"],"forum":"eNjXcP6C0H","number":1,"license":"CC BY 4.0","cdate":1729965878660,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission688/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427301606,"domain":"ICLR.cc/2025/Conference","replyto":"eNjXcP6C0H","id":"61hGXB72jA","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion models","generative models","physical dynamics","priors"]},"supplementary_material":{"value":"/attachment/47ab44b471ce3705a1cad3019ba8f11534475c8c.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Generating physically feasible dynamics in a data-driven context is challenging, especially when adhering to physical priors expressed in specific equations or formulas. Existing methodologies often overlook the integration of ''physical priors'', resulting in violation of basic physical laws and suboptimal performance. In this paper, we introduce a novel framework that seamlessly incorporates physical priors into diffusion-based generative models to address this limitation. Our approach leverages two categories of priors: 1) distributional priors, such as roto-translational invariance, and 2) physical feasibility priors, including energy and momentum conservation laws and PDE constraints. By embedding these priors into the generative process, our method can efficiently generate physically realistic dynamics, encompassing trajectories and flows. Empirical evaluations demonstrate that our method produces high-quality dynamics across a diverse array of physical phenomena with remarkable robustness, underscoring its potential to advance data-driven studies in AI4Physics. Our contributions signify a substantial advancement in the field of generative modeling, offering a robust solution to generate accurate and physically consistent dynamics."},"_bibtex":{"value":"@inproceedings{\nzhou2025generating,\ntitle={Generating Physical Dynamics under Priors},\nauthor={Zihan Zhou and Xiaoxue Wang and Tianshu Yu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=eNjXcP6C0H}\n}"},"title":{"value":"Generating Physical Dynamics under Priors"},"pdf":{"value":"/pdf/672b228aaf31555595ab6bb814183dcd91d04134.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhou|generating_physical_dynamics_under_priors"},"authorids":{"value":["~Zihan_Zhou7","~Xiaoxue_Wang2","~Tianshu_Yu2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zihan Zhou","Xiaoxue Wang","Tianshu Yu"]}},"version":2},{"content":{"venue":{"value":"FPI-NEURIPS2025 Poster"},"pdf":{"value":"/pdf/843d87269260f0134a0d9a74003ab10a75efec1c.pdf"},"keywords":{"value":["Physics-Informed Machine Learning; Time Series; Diffusion Model; Transformer Diffusion; Physics-Injection Inference"]},"venueid":{"value":"NeurIPS.cc/2025/Workshop/FPI"},"abstract":{"value":"Scientific time series data, particularly in climate science, present unique challenges at the intersection of probabilistic inference and physical constraints. While large pre-trained Time Series Diffusion Transformers excel at capturing complex data distributions, they lack mechanisms to enforce physical consistency. To address this gap, we present a novel framework for adapting pre-trained generative models for scientific tasks. Our contribution is a model-agnostic physics-injection module that employs Langevin dynamics at inference time to steer predictions toward physically consistent solutions without costly retraining. We provide theoretical guarantees for convergence under physical constraints and empirically validate our method across multiple synthetic partial differential equations and climate systems, offering insights into the synergy between machine learning and physics-based sampling for scientific applications."},"_bibtex":{"value":"@inproceedings{\nanonymous2025energybased,\ntitle={Energy-Based Physics-Informed Diffusion Transformers Sampling for Time Series Forecasting},\nauthor={Anonymous},\nbooktitle={2nd edition of Frontiers in Probabilistic Inference: Learning meets Sampling},\nyear={2025},\nurl={https://openreview.net/forum?id=rdiElRhUU1}\n}"},"title":{"value":"Energy-Based Physics-Informed Diffusion Transformers Sampling for Time Series Forecasting"},"track":{"value":"Main Track"}},"tmdate":1764113831093,"pdate":1758595821071,"tcdate":1756795341276,"writers":["NeurIPS.cc/2025/Workshop/FPI","NeurIPS.cc/2025/Workshop/FPI/Submission98/Authors"],"signatures":["NeurIPS.cc/2025/Workshop/FPI/Submission98/Authors"],"forum":"rdiElRhUU1","license":"CC BY 4.0","number":98,"cdate":1756795341276,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Workshop/FPI/-/Submission","NeurIPS.cc/2025/Workshop/FPI/-/Post_Submission","NeurIPS.cc/2025/Workshop/FPI/-/Edit","NeurIPS.cc/2025/Workshop/FPI/Submission98/-/Camera-Ready"],"mdate":1764113831093,"odate":1764094029337,"domain":"NeurIPS.cc/2025/Workshop/FPI","id":"rdiElRhUU1","version":2},{"content":{"summary":{"value":"HASARD is a benchmark testing platform specifically designed for safe reinforcement learning, based on ViZDoom, providing a diverse range of 3D environments.\n\n1. The tasks on this platform require agents to pursue high rewards while considering safety strategies, moving beyond simple 2D navigation to incorporate complex elements such as spatial understanding.\n2. HASARD offers three difficulty levels and supports both soft and hard safety constraints, flexibly adapting to varying safety requirements.\n3. The platform integrates Sample-Factory, enabling high-speed simulation that allows agents to address real-world safety challenges while reducing computational costs.\n4. HASARD includes six environments based on ViZDoom and benchmarks various methods to demonstrate the limitations of existing technologies."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. The article does not provide an in-depth analysis of performance under different safety budgets. Is there a plan to supplement the experiments with varying safety thresholds to comprehensively demonstrate the trade-offs between reward and safety for each algorithm? This would be very helpful in understanding the adaptability of different methods under various safety requirements.\n2. Considering the limitations of ViZDoom in simulating real-world physics, have the authors explored other engines with superior physical simulation capabilities (e.g., Isaac Gym)?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"No ethics concerns."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The authors tested six baseline algorithms on HASARD and provided an analysis of the results.\n2. The tasks move beyond simple 2D navigation to incorporate complex elements such as spatial understanding"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The reviewer believes that if the distinction between soft and hard constraints is merely based on whether the threshold is $0$, then other benchmarks share this characteristic, making this claim somewhat unsubstantiated.\n2. Although multiple methods were tested in the current experiments, there is a lack of analysis on performance under different safety budgets. It is recommended to include experiments with varying safety thresholds to better understand the trade-off between safety and reward for each algorithm.\n3. HASARD is based on the ViZDoom game engine, which, while computationally inexpensive, lacks detailed simulation of real-world physics.\n4. The anonymous video link provided by the authors is inaccessible."}},"nonreaders":[],"tmdate":1732509002005,"tcdate":1730710315709,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10610/Reviewer_NMGZ"],"signatures":["ICLR.cc/2025/Conference/Submission10610/Reviewer_NMGZ"],"forum":"5BRFddsAai","number":3,"license":"CC BY 4.0","cdate":1730710315709,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10610/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732509002005,"domain":"ICLR.cc/2025/Conference","replyto":"5BRFddsAai","id":"9ixF7aYzIA","forumContent":{"TLDR":{"value":"A Safe RL benchmark for vision-based learning in complex navigable 3D environments."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["reinforcement learning","AI safety","safe RL","constrained RL","benchmark","vizdoom","3D","difficulty levels"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advancing safe autonomous systems through reinforcement learning (RL) requires robust benchmarks to evaluate performance, analyze methods, and assess agent competencies. Humans primarily rely on embodied visual perception to safely navigate and interact with their surroundings, making it a valuable capability for RL agents. However, existing vision-based 3D benchmarks only consider simple navigation tasks. To address this shortcoming, we introduce **HASARD**, a suite of diverse and complex tasks to **HA**rness **SA**fe **R**L with **D**oom, requiring strategic decision-making, comprehending spatial relationships, and predicting the short-term future. HASARD features three difficulty levels and two action spaces. An empirical evaluation of popular baseline methods demonstrates the benchmark's complexity, unique challenges, and reward-cost trade-offs. Visualizing agent navigation during training with top-down heatmaps provides insight into a method's learning process. Incrementally training across difficulty levels offers an implicit learning curriculum. HASARD is the first safe RL benchmark to exclusively target egocentric vision-based learning, offering a cost-effective and insightful way to explore the potential and boundaries of current and future safe RL methods. The environments and baseline implementations are open-sourced."},"_bibtex":{"value":"@inproceedings{\ntomilin2025hasard,\ntitle={{HASARD}: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents},\nauthor={Tristan Tomilin and Meng Fang and Mykola Pechenizkiy},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=5BRFddsAai}\n}"},"title":{"value":"HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents"},"pdf":{"value":"/pdf/fde61410f00ccf22d9510dc28b79cc37337c80d3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"tomilin|hasard_a_benchmark_for_visionbased_safe_reinforcement_learning_in_embodied_agents"},"authorids":{"value":["~Tristan_Tomilin1","~Meng_Fang1","~Mykola_Pechenizkiy1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Tristan Tomilin","Meng Fang","Mykola Pechenizkiy"]}},"version":2},{"content":{"venue":{"value":"Nat. 2021"},"pdf":{"value":"https://www.nature.com/articles/s41586-021-03382-w.pdf"},"venueid":{"value":"dblp.org/journals/NATURE/2021"},"paperhash":{"value":"hatfield|the_datadriven_future_of_highenergydensity_physics"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Peter_W._Hatfield:","https://dblp.org/search/pid/api?q=author:Jim_A._Gaffney:","https://dblp.org/search/pid/api?q=author:Gemma_J._Anderson:","https://dblp.org/search/pid/api?q=author:Suzanne_Ali:","https://dblp.org/search/pid/api?q=author:Luca_Antonelli:","https://dblp.org/search/pid/api?q=author:Suzan_Basegmez_du_Pree:","https://dblp.org/search/pid/api?q=author:Jonathan_Citrin:","https://dblp.org/search/pid/api?q=author:Marta_Fajardo:","https://dblp.org/search/pid/api?q=author:Patrick_F._Knapp:","https://dblp.org/search/pid/api?q=author:Brendan_Kettle:","https://dblp.org/search/pid/api?q=author:Bogdan_Kustowski:","https://dblp.org/search/pid/api?q=author:Michael_J._MacDonald:","https://dblp.org/search/pid/api?q=author:Derek_Mariscal:","https://dblp.org/search/pid/api?q=author:Madison_E._Martin:","https://dblp.org/search/pid/api?q=author:Taisuke_Nagayama:","https://dblp.org/search/pid/api?q=author:Charlotte_A._J._Palmer:","~Luc_Peterson1","https://dblp.org/search/pid/api?q=author:Steven_J._Rose:","https://dblp.org/search/pid/api?q=author:J._J._Ruby:","https://dblp.org/search/pid/api?q=author:Carl_Shneider:","https://dblp.org/search/pid/api?q=author:Matt_J._V._Streeter:","https://dblp.org/search/pid/api?q=author:Will_Trickey:","https://dblp.org/search/pid/api?q=author:Ben_Williams:"]},"html":{"value":"https://doi.org/10.1038/s41586-021-03382-w"},"_bibtex":{"value":"@article{DBLP:journals/nature/HatfieldGAAAPCF21,\n  author={Peter W. Hatfield and Jim A. Gaffney and Gemma J. Anderson and Suzanne Ali and Luca Antonelli and Suzan Basegmez du Pree and Jonathan Citrin and Marta Fajardo and Patrick F. Knapp and Brendan Kettle and Bogdan Kustowski and Michael J. MacDonald and Derek Mariscal and Madison E. Martin and Taisuke Nagayama and Charlotte A. J. Palmer and J. Luc Peterson and Steven J. Rose and J. J. Ruby and Carl Shneider and Matt J. V. Streeter and Will Trickey and Ben Williams},\n  title={The data-driven future of high-energy-density physics},\n  year={2021},\n  cdate={1609459200000},\n  journal={Nat.},\n  volume={593},\n  number={7859},\n  pages={351-361},\n  url={https://doi.org/10.1038/s41586-021-03382-w}\n}\n"},"abstract":{"value":"High-energy-density physics is the field of physics concerned with studying matter at extremely high temperatures and densities. Such conditions produce highly nonlinear plasmas, in which several phenomena that can normally be treated independently of one another become strongly coupled. The study of these plasmas is important for our understanding of astrophysics, nuclear fusion and fundamental physics—however, the nonlinearities and strong couplings present in these extreme physical systems makes them very difficult to understand theoretically or to optimize experimentally. Here we argue that machine learning models and data-driven methods are in the process of reshaping our exploration of these extreme systems that have hitherto proved far too nonlinear for human researchers. From a fundamental perspective, our understanding can be improved by the way in which machine learning models can rapidly discover complex interactions in large datasets. From a practical point of view, the newest generation of extreme physics facilities can perform experiments multiple times a second (as opposed to approximately daily), thus moving away from human-based control towards automatic control based on real-time interpretation of diagnostic data and updates of the physics model. To make the most of these emerging opportunities, we suggest proposals for the community in terms of research design, training, best practice and support for synthetic diagnostics and data analysis. This Perspective discusses how high-energy-density physics could tap the potential of AI-inspired algorithms for extracting relevant information and how data-driven automatic control routines may be used for optimizing high-repetition-rate&nbsp;experiments."},"title":{"value":"The data-driven future of high-energy-density physics"},"authors":{"value":["Peter W. Hatfield","Jim A. Gaffney","Gemma J. Anderson","Suzanne Ali","Luca Antonelli","Suzan Basegmez du Pree","Jonathan Citrin","Marta Fajardo","Patrick F. Knapp","Brendan Kettle","Bogdan Kustowski","Michael J. MacDonald","Derek Mariscal","Madison E. Martin","Taisuke Nagayama","Charlotte A. J. Palmer","J. Luc Peterson","Steven J. Rose","J. J. Ruby","Carl Shneider","Matt J. V. Streeter","Will Trickey","Ben Williams"]}},"tmdate":1716335663368,"pdate":1609459200000,"tcdate":1716335653494,"writers":["~"],"signatures":["~Luc_Peterson1"],"forum":"EZBXesLooV","license":"CC BY-SA 4.0","number":17740,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1716335663368,"domain":"DBLP.org","id":"EZBXesLooV","version":2},{"content":{"summary":{"value":"The paper analyzes algorithms for estimating local intrinsic dimension, focusing on ESS, NB, LIDL, and FLIPD. The experiments cover simple synthetic datasets with known dimension and realistic datasets with complex structure and unknown dimension. For the realistic case, the paper applies transformations that preserve manifold dimension or change it in a known way and then measures changes in reported dimension to evaluate robustness. The study investigates controlled factors that affect LID estimates, including manifold curvature, thickness, sample size, non uniform density, and proximity to boundaries."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"- What is the number of samples used in the Gaussian IDR experiment, and is it sufficient for reliable estimation in the tails? What are the number of samples used for the other experiments?\n\n- Can the authors define precisely distance to the neighbor on the same manifold versus distance to the nearest neighbor in the ambient space, and explain how these distances are instantiated in the Funnel and Spiral experiments?\n\n- In line 1471, it is claimed that \"Arrows (MS) has many V-shaped corners as artifact of translating and rotating\". Can this be further explained?\n\n- When two arrows overlap it is claimed that RGB values are added (Line 1471), but this effect is not visible in Figure 2. Can the authors clarify this?\n\n- Minor typos:\n\n  - Line 141: \"Sec. 1\" should be \"Sec. 2.\"\n\n  - Line 168: It is unclear whether the manifold is S^4 or embedded in a 6 dimensional space. Please clarify the intended statement.\n\n  - Line 242: The set description contains typos. Please correct the notation.\n\n  - Line 312: \"algorithm error\" should be \"algorithm's error.\"\n\n  - Line 317: \"Figure 8\" refers to another paper but links to the Figure 8 in the current paper.\n\n  - Line 412: \"too big our computational\" should be \"too big for our computational.\"\n\n  - Line 445: \"worse to\" should be \"worse than.\"\n\n  - Line 455: The word \"while\" should be removed.\n\n  - Figure references are inconsistent, alternating between \"Fig.\" and \"Figure\" throughout the paper."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The experimental design explicitly probes factors known to affect LID estimates, including curvature, thickness, sample size, non uniform density, and proximity to boundaries.\n\n- The experimental section is thorough.\n\n- The empirical design spans simple known dimension synthetic data and complex unknown dimension real data, which provides a broad stress test for the methods.\n\n- The paper systematically highlights shortcomings of existing LID methods and existing benchmarks through controlled experiments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The Gaussian (IDR) experiment does not report the number of samples used, and tail regions such as standard deviation beyond 2 standard deviations require high sample sizes to get enough samples in the region to have stable estimates, especially for model trained methods like FLIPD and LIDL. The number of training and test samples should be reported across other experiments as well.\n\n- The Funnel (IDR) experiment is placed under nearby manifolds, but as described it is a single component and seems more relevant to thin manifolds.\n\n- The Spiral (IDR) experiment under nearby manifolds appears to have extremely low point density on the outer spiral, which can cause any algorithm to fail.\n\n- Section 3.5 nearby manifolds uses terms like distance to the neighbor on the same manifold versus distance to the nearest neighbor in the ambient space without a precise definition, and the Funnel and Spiral experiments do not make these distances clear.\n\n- Figure 9 uses a log x axis and allocates a large fraction of the range to fewer than 100 samples, which is hard to interpret for FMNIST where the ambient dimension is 28*28.\n\n- Figure 23 (d) would be clearer with the x axis limited to a lower range because the current figure is unreadable.\n\n- The paper mentions using PyTorch interpolation for resizing in Upscaled ASE and possibly other experiments, and known aliasing artifacts in this method might affect LID [1].\n\n- Section 2 mentions audio but there are no audio experiments, and removing the audio discussion would improve focus.\n\n- Section 3.7 cites work showing ESS is invariant to sample size for artificial datasets, but comparable results for the other algorithms on synthetic data where ground truth is known are missing.\n\n[1] Parmar, Gaurav, Richard Zhang, and Jun-Yan Zhu. \"On aliased resizing and surprising subtleties in gan evaluation.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943340230,"tcdate":1761869087100,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission25138/Reviewer_Sef9"],"signatures":["ICLR.cc/2026/Conference/Submission25138/Reviewer_Sef9"],"forum":"ZEf03Uunvk","number":3,"license":"CC BY 4.0","cdate":1761869087100,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission25138/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943340230,"domain":"ICLR.cc/2026/Conference","replyto":"ZEf03Uunvk","id":"kmAGIt3Gij","forumContent":{"TLDR":{"value":"We show that LID estimation community needs new benchmarks for intrinsic dimension estimation and come to interesting conclusions on the performance of existing algorithms."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Local intrinsic dimension estimation","LIDL","FLIPD","Diffusion Models","Benhamark","Normalizing Flows","ESS","Normal Bundle","NB","LID"]},"supplementary_material":{"value":"/attachment/415de902ef6b8c0ea758232ce5bdfe6eda8a506e.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Neural Local Intrinsic Dimension (LID) estimators are typically bound to domain-specific architectures whose inductive biases can yield inconsistent estimates for the same underlying manifold. Existing evaluations either use overly simple synthetic data (with known LID) or real datasets (with unknown LID), obscuring true performance. We introduce a principled benchmarking framework that (i) maps the same manifold into multiple domain representations while preserving its structure, enabling like-for-like cross-architecture tests; (ii) designs harder variants of popular datasets that target key manifold properties; and (iii) applies controlled transformations with known LID shifts to stress-test methods even when absolute LID is unknown. Across this suite, including non-trivial synthetic datasets, we show that accuracy on simple manifolds does not transfer across domains and that state-of-the-art methods fail under targeted stressors, revealing clear failure modes and areas for improvement. Data and code are available: https://github.com/DominikFilipiak/LID-Benchmarks."},"_bibtex":{"value":"@inproceedings{\ntempczyk2026why,\ntitle={Why We Need New Benchmarks for Local Intrinsic Dimension Estimation},\nauthor={Piotr Tempczyk and Dominik Filipiak and {\\L}ukasz Garncarek and Ksawery Smoczy{\\'n}ski and Adam Kurpisz},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ZEf03Uunvk}\n}"},"title":{"value":"Why We Need New Benchmarks for Local Intrinsic Dimension Estimation"},"pdf":{"value":"/pdf/88bbda820d2f79b397171d3dca3bb65cc8637253.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"tempczyk|why_we_need_new_benchmarks_for_local_intrinsic_dimension_estimation"},"authorids":{"value":["~Piotr_Tempczyk1","~Dominik_Filipiak1","~Łukasz_Garncarek1","~Ksawery_Smoczyński1","~Adam_Kurpisz1"]},"authors":{"value":["Piotr Tempczyk","Dominik Filipiak","Łukasz Garncarek","Ksawery Smoczyński","Adam Kurpisz"]}},"version":2},{"content":{"summary":{"value":"This paper introduces FEABench for evaluating LLMs ability on solving physics, mathematics and engineering problems using finite element analysis software. The authors evaluate 3 LLMs on this benchmark and also develop a multi-turn LLM agent that can interact with the FEA software API to iteratively improve its solutions. Experimental results demonstrate that this benchmark is challenging."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Same as Weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"To the best of my knowledge, this is the first systematic work on large language models manipulating complex computer software.\nThe benchmark is challenging."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The FEABench Gold dataset is relatively too small, with only 15 problems. Is it comprehensive? Please compare with other relevant datasets to show whether your dataset is relatively too small, e,g, SWEBench[1] (Agent Dataset), OSWorld[2] (Software Agent Dataset).\n2. The work is not particularly outstanding in the context of LLM-Agent research and lacks sufficient references to related work in this area. Although the narrative is from a physics perspective, the overall approach is closely related to LLM Agent research.\n3. The experimental evaluation is not comprehensive enough. It only explores 3 closed-source models. Other models, including many open-source ones, e.g., Llama, were not tested.\n4. The study only considers one software, COMSOL Multiphysics, which may not fully capture the \"REAL WORLD PHYSICS REASONING ABILITY\" claimed in the title, e.g. ANSYS.\n5. The paper does not consider the graphical user interface of COMSOL Multiphysics, which is crucial for using this software, as is shown in Figure 5 in your paper.\n\n[1] SWE-bench: Can Language Models Resolve Real-World GitHub Issues?\n[2] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments"}},"nonreaders":[],"tmdate":1732819205423,"tcdate":1729248187160,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12171/Reviewer_oeq1"],"signatures":["ICLR.cc/2025/Conference/Submission12171/Reviewer_oeq1"],"forum":"hDkLpu1E64","number":1,"license":"CC BY 4.0","cdate":1729248187160,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12171/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732819205423,"domain":"ICLR.cc/2025/Conference","replyto":"hDkLpu1E64","id":"HctTgCMHDk","forumContent":{"TLDR":{"value":"How well can LLMs leverage FEA software to simulate and solve problems that require numerical analysis?"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["numerical analysis","finite element","benchmark","agents"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building precise simulations of the real world and invoking numerical solvers to answer quantitative problems is an essential requirement in engineering and science. We present FEABench, a benchmark to evaluate the ability of large language models (LLMs) and LLM agents to simulate and solve physics, mathematics and engineering problems using finite element analysis (FEA). We introduce a multipronged evaluation scheme to investigate the ability of LLMs to solve these problems by reasoning over natural language problem descriptions and operating COMSOL Multiphysics$^\\textregistered$, an FEA software, to compute the answers. In addition to testing state-of-the art-LLMs, we further design a language model agent equipped with the ability to interact with the software through its Application Programming Interface (API), examine its outputs and use tools to improve its solutions over multiple iterations. Our best performing strategy generates executable API calls 88\\% of the time. However, this benchmark still proves to be challenging enough that the LLMs and agents we tested were not able to completely and correctly solve any problem. LLMs that can successfully interact with and operate FEA software to solve problems such as those in our benchmark would significantly push the frontiers of their utility. Acquiring this capability would augment LLMs' reasoning skills with the precision of numerical solvers and advance the development of autonomous systems that can tackle complex problems in the real world."},"_bibtex":{"value":"@misc{\nmudur2025feabench,\ntitle={{FEAB}ench: Evaluating Language Models on Real World Physics Reasoning Ability},\nauthor={Nayantara Mudur and Hao Cui and Subhashini Venugopalan and Paul Raccuglia and Michael Brenner and Peter Christian Norgaard},\nyear={2025},\nurl={https://openreview.net/forum?id=hDkLpu1E64}\n}"},"title":{"value":"FEABench: Evaluating Language Models on Real World Physics Reasoning Ability"},"pdf":{"value":"/pdf/3e64111fb86b7cbb5ef6469de0f077b416722ed3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"mudur|feabench_evaluating_language_models_on_real_world_physics_reasoning_ability"},"authorids":{"value":["~Nayantara_Mudur1","~Hao_Cui3","~Subhashini_Venugopalan2","~Paul_Raccuglia1","~Michael_Brenner1","~Peter_Christian_Norgaard1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Nayantara Mudur","Hao Cui","Subhashini Venugopalan","Paul Raccuglia","Michael Brenner","Peter Christian Norgaard"]}},"version":2},{"content":{"summary":{"value":"This paper introduces ALEGP, a novel framework for Symbolic Regression (SR) that aims to address the challenges in Genetic Programming (GP), such as bloat, premature convergence, and inadequate simplification mechanisms. The core innovation of ALEGP is the strategic integration of Large Language Models (LLMs) with the evolutionary search process of GP, leveraging the LLMs' mathematical reasoning and simplification capabilities. Through experiments on 8 synthetic benchmarks and 5 real-world datasets, the authors demonstrate that ALEGP achieves significant improvements in solution accuracy compared to traditional GP baselines (Standard GP and gplearn)."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please refer to the Weaknesses section."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The idea of dynamically leveraging the powerful LLMs to simplify and refine mathematical expressions during the evolutionary process is highly novel. This paradigm, which combines symbolic search with neural model-based reasoning, opens a new avenue for solving complex optimization problems and holds significant research potential and value.\n2. The authors have designed a complex yet comprehensive synergistic mechanism to ensure effectiveness. This includes a multi-island architecture for maintaining population diversity, an adaptive intervention scheduler for efficient LLM calls, and a specificity strategy for integrating optimized solutions.\n3.  Through experiments on a variety of synthetic and real-world datasets,  the authors demonstrate that the ALEGP framework consistently outperforms traditional GP baselines (Standard GP and gplearn) in terms of solution accuracy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The use of deep learning models, particularly Transformers, for Symbolic Regression has become a major research direction, yielding many strong results (e.g., DSR, NeSymReS, E2E, SymbolicGPT). However, the Related Work section fails to adequately discuss the relationship and distinctions between ALEGP and these advanced methods. \n2. The paper proposes a system with multiple complex components but does not clearly highlight its single most crucial contribution. Beyond the high-level concept of \"using LLMs in GP,\" what is the primary innovation? Is it the adaptive scheduler, the multi-island co-evolutionary framework, or the LLM prompting strategy? The authors are encouraged to more clearly delineate and summarize the paper's contributions in a hierarchical manner to help readers grasp the key takeaways.\n3. The proposed method lacks necessary details regarding its LLM integration. The quality of an LLM's output is highly dependent on prompt design, hyperparameters like temperature, and the provided context. The paper should specify the exact prompt templates, key parameter settings, and few-shot examples used to guide the LLMs. Furthermore, the paper does not discuss how the reliability of the LLM output is ensured. For instance, LLMs can occasionally generate syntactically incorrect or mathematically nonsensical expressions. The authors should detail whether a validation mechanism exists to handle such illegal outputs and report on the frequency of these failure cases, which is crucial for assessing the method's robustness.\n4.  The results reveal an interesting phenomenon: different LLMs perform disparately on different types of datasets (e.g., GPT-4o-mini excels on synthetic data but its performance degrades on real-world data, while Gemini-2 shows the opposite trend). This is a finding worthy of in-depth discussion, yet the current paper offers no analysis. What are the underlying reasons? Is it due to the models' mathematical reasoning, generalization capabilities, or other factors? Additionally, the synthetic benchmark functions used are often classic problems whose forms and solutions may well be part of the LLMs' vast training data. The authors need to discuss the potential risk of data leakage for these synthetic benchmarks.\n5. The experimental section completely lacks performance comparisons SOTA symbolic regression methods, such as SBP-GP, GP-GOMEA, Operon. This omission makes it difficult for readers to accurately assess ALEGP's technical standing and practical advantages within the current SR landscape."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921637398,"tcdate":1761713424671,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10284/Reviewer_qAAw"],"signatures":["ICLR.cc/2026/Conference/Submission10284/Reviewer_qAAw"],"forum":"NtkWD4CQin","number":2,"license":"CC BY 4.0","cdate":1761713424671,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10284/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921637398,"domain":"ICLR.cc/2026/Conference","replyto":"NtkWD4CQin","id":"FyB2hzZHca","forumContent":{"TLDR":{"value":"ALEGP: An adaptive LLM-enhanced genetic programming approach that dynamically integrates large language models with multi-island evolutionary search to address bloating, premature convergence, and local optima in symbolic regression."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Symbolic Regression","Large Language Models (LLMs)","Genetic Programming","Adaptive Scheduling"]},"supplementary_material":{"value":"/attachment/93a5179f459e04d0c7455419eb253e275beec50e.zip"},"primary_area":{"value":"optimization"},"abstract":{"value":"Symbolic regression aims to discover mathematical expressions that capture underlying data relationships, but genetic programming (GP) approaches commonly encounter bloat, premature convergence, and inadequate expression simplification mechanisms. We propose ALEGP (Adaptive LLM-Enhanced Genetic Programming), a framework that strategically integrates large language models (LLMs) with evolutionary computation to address these interconnected challenges.\nALEGP incorporates three key components: (i) a multi-island evolutionary architecture employing specialized subpopulations with distinct optimization objectives to maintain population diversity, (ii) a context-aware intervention scheduler that triggers LLM assistance based on real-time evolutionary indicators including fitness stagnation, diversity loss, and expression bloat, and (iii) an island-specific integration protocol that reincorporates LLM-refined expressions while preserving beneficial evolutionary dynamics. This design enables targeted simplification of complex expressions, improved generalization performance, and reduced computational overhead through adaptive LLM utilization.\nExperiments on eight synthetic benchmark functions and five real-world regression datasets demonstrate that ALEGP achieves superior accuracy and interpretability while requiring 50–60\\% fewer LLM interventions than fixed-schedule strategies. Ablation studies validate the necessity of both adaptive scheduling and multi-island design for robust performance. These results establish ALEGP as an effective framework for resource-efficient symbolic regression, demonstrating principled integration of evolutionary algorithms with large language models. Code is provided as supplementary material."},"_bibtex":{"value":"@misc{\npalakonda2025how,\ntitle={How is Occam's Razor Realized in Symbolic Regression?: An Adaptive {LLM}-Enhanced Genetic Programming Approach for Efficient, Versatile, and Interpretable Representation Discovery through Simplification and Evolution},\nauthor={Vikas Palakonda and Jamshid Tursunboev and Samira Ghorbanpour and Il-Min Kim and Jae-Mo Kang and Sunghwan Moon},\nyear={2025},\nurl={https://openreview.net/forum?id=NtkWD4CQin}\n}"},"title":{"value":"How is Occam's Razor Realized in Symbolic Regression?: An Adaptive LLM-Enhanced Genetic Programming Approach for Efficient, Versatile, and Interpretable Representation Discovery through Simplification and Evolution"},"pdf":{"value":"/pdf/642d5e940755ae56bf51a78f4f0e0e034b57811a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"palakonda|how_is_occams_razor_realized_in_symbolic_regression_an_adaptive_llmenhanced_genetic_programming_approach_for_efficient_versatile_and_interpretable_representation_discovery_through_simplification_and_evolution"},"authorids":{"value":["~Vikas_Palakonda1","~Jamshid_Tursunboev1","~Samira_Ghorbanpour1","~Il-Min_Kim1","~Jae-Mo_Kang1","~Sunghwan_Moon1"]},"authors":{"value":["Vikas Palakonda","Jamshid Tursunboev","Samira Ghorbanpour","Il-Min Kim","Jae-Mo Kang","Sunghwan Moon"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2501.08428v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"karumuri|physicsinformed_latent_neural_operator_for_realtime_predictions_of_complex_physical_systems"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Sharmila_Karumuri:","https://dblp.org/search/pid/api?q=author:Lori_Graham-Brady:","~Somdatta_Goswami1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2501.08428"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2501-08428,\n  publtype={informal},\n  author={Sharmila Karumuri and Lori Graham-Brady and Somdatta Goswami},\n  title={Physics-Informed Latent Neural Operator for Real-time Predictions of Complex Physical Systems},\n  year={2025},\n  month={January},\n  cdate={1735689600000},\n  journal={CoRR},\n  volume={abs/2501.08428},\n  url={https://doi.org/10.48550/arXiv.2501.08428}\n}\n"},"abstract":{"value":"Deep operator network (DeepONet) has shown great promise as a surrogate model for systems governed by partial differential equations (PDEs), learning mappings between infinite-dimensional function spaces with high accuracy. However, achieving low generalization errors often requires highly overparameterized networks, posing significant challenges for large-scale, complex systems. To address these challenges, latent DeepONet was proposed, introducing a two-step approach: first, a reduced-order model is used to learn a low-dimensional latent space, followed by operator learning on this latent space. While effective, this method is inherently data-driven, relying on large datasets and making it difficult to incorporate governing physics into the framework. Additionally, the decoupled nature of these steps prevents end-to-end optimization and the ability to handle data scarcity. This work introduces PI-Latent-NO, a physics-informed latent operator learning framework that overcomes these limitations. Our architecture employs two coupled DeepONets in an end-to-end training scheme: the first, termed Latent-DeepONet, identifies and learns the low-dimensional latent space, while the second, Reconstruction-DeepONet, maps the latent representations back to the original physical space. By integrating governing physics directly into the training process, our approach requires significantly fewer data samples while achieving high accuracy. Furthermore, the framework is computationally and memory efficient, exhibiting nearly constant scaling behavior on a single GPU and demonstrating the potential for further efficiency gains with distributed training. We validate the proposed method on high-dimensional parametric PDEs, demonstrating its effectiveness as a proof of concept and its potential scalability for large-scale systems."},"title":{"value":"Physics-Informed Latent Neural Operator for Real-time Predictions of Complex Physical Systems"},"authors":{"value":["Sharmila Karumuri","Lori Graham-Brady","Somdatta Goswami"]}},"tmdate":1747358506964,"pdate":1735689600000,"tcdate":1747358476074,"writers":["~"],"signatures":["~Somdatta_Goswami1"],"forum":"gwcIfuW1zM","license":"CC BY-SA 4.0","number":501492,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747358506964,"domain":"DBLP.org","id":"gwcIfuW1zM","version":2},{"content":{"summary":{"value":"The authors contrast the learning and convergence properties of real-valued neural networks and complex-valued neural networks. Specifically, they study the problem of learning the function implemented by a single neuron using a 2-layer finite width network. Notably, they show that a complex valued neural network can learn functions expressed by a real-valued neuron as well as a complex-valued neuron. This result indicates the superior capacity of complex-valued neural networks as compared to their real-valued counterparts. Furthermore, they also study the convergence properties of learning a function expressed by a real-valued neuron with a complex-valued neuron and prove that its slower compared to learning with a real-valued neural network. Taken together, complex-valued neural networks have a better learning capacity at the cost of slower learning properties compared to real-valued neural networks. Overall, I think this paper makes a strong theoretical contribution to understanding the efficiency and capacity of complex-values neural networks in practice."},"soundness":{"value":"4 excellent"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. The caption of Fig. 2 is currently not very descriptive. Could you please update it such that the takeaway is clear to the reader without reading in detail the text?\n2. Is it possible to add a toy regression example where you can empirically demonstrate the different phases during the learning process?"},"rating":{"value":"7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"strengths":{"value":"1. The paper has strong theoretical foundations and presents arguments through rigorous proofs. \n2. Despite the involved mathematical machinery, the authors present a simplified view of the underlying learning dynamics and present the existence of phases in the learning process. \n3. Despite the assumption of the MSE loss, I feel the authors present a strong theoretical framework to analyze learning in complex-valued neural networks and contrast them with real-valued neural networks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The authors present strong theoretical results, and use specific assumptions to prove their results. However, it is unclear the extent to which these assumptions would hold in practice. \n2. In addition to their (very impressive) theoretical results, I feel the current version of the paper would benefit from some empirical validation. Having some toy experiments would also be sufficient to significantly improve the readability of the paper, while also increasing the reader pool. \n3. The manuscript is generally well-written, but there are certain typos or use of abbreviations/notations which make it confusing for the reader. E.g. it seems that there is a typo in Table 2, wherein the second row third column (complex-valued Neuron --> complex-valued neuron) should be $\\mathcal{O}(t^{-1})$. Also, abbreviations CVNN and RVNN or the notation $t$ is not introduced in the abstract."},"limitations":{"value":"N/A"}},"nonreaders":[],"tmdate":1702410865475,"tcdate":1688715991141,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission2826/Reviewer_cJ2H"],"signatures":["NeurIPS.cc/2023/Conference/Submission2826/Reviewer_cJ2H"],"forum":"qA0uHmaVKk","number":3,"license":"CC BY 4.0","cdate":1688715991141,"mdate":1702410865475,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission2826/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"qA0uHmaVKk","id":"T4MMzcQB6T","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Complex-valued Neural Networks; Learning Neurons; Real-valued Neural Networks; Convergence Rate"]},"supplementary_material":{"value":"/attachment/e0e8195000a1f4fde79c9bc8cf119b08fbfc8b4f.pdf"},"_bibtex":{"value":"@inproceedings{\nwu2023complexvalued,\ntitle={Complex-valued Neurons Can Learn More but Slower than Real-valued Neurons via Gradient Descent},\nauthor={Jin-Hui Wu and Shao-Qun Zhang and Yuan Jiang and Zhi-Hua Zhou},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=qA0uHmaVKk}\n}"},"title":{"value":"Complex-valued Neurons Can Learn More but Slower than Real-valued Neurons via Gradient Descent"},"paperhash":{"value":"wu|complexvalued_neurons_can_learn_more_but_slower_than_realvalued_neurons_via_gradient_descent"},"abstract":{"value":"Complex-valued neural networks potentially possess better representations and performance than real-valued counterparts when dealing with some complicated tasks such as acoustic analysis, radar image classification, etc. Despite empirical successes, it remains unknown theoretically when and to what extent complex-valued neural networks outperform real-valued ones. We take one step in this direction by comparing the learnability of real-valued neurons and complex-valued neurons via gradient descent. We show that a complex-valued neuron can efficiently learn functions expressed by any one real-valued neuron and any one complex-valued neuron with convergence rate $O(t^{-3})$ and $O(t^{-1})$ where $t$ is the iteration index of gradient descent, respectively, whereas a two-layer real-valued neural network with finite width cannot learn a single non-degenerate complex-valued neuron. We prove that a complex-valued neuron learns a real-valued neuron with rate $\\Omega (t^{-3})$, exponentially slower than the $O(\\mathrm{e}^{- c t})$ rate of learning one real-valued neuron using a real-valued neuron with a constant $c$. We further verify and extend these results via simulation experiments in more general settings."},"pdf":{"value":"/pdf/129de32bfd6b64ae584e3f42314707df638b5755.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Jin-Hui_Wu1","~Shao-Qun_Zhang1","~Yuan_Jiang1","~Zhi-Hua_Zhou2"]},"authors":{"value":["Jin-Hui Wu","Shao-Qun Zhang","Yuan Jiang","Zhi-Hua Zhou"]}},"version":2},{"content":{"summary":{"value":"The paper proposes the Lorentz Geometric Algebra Transformer (L-GATr) for high-energy physics tasks. This model extends the Geometric Algebra Transformer by incorporating relativistic considerations. Specifically, L-GATr supports partial and approximate symmetry for symmetry-breaking inputs and is applied to generative models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I think another interesting comparison would be to evaluate the proposed model against traditional sampling methods, especially if the goal is to design machine learning models that are more efficient than traditional methods (within an acceptable error tolerance). Could the paper include this comparison? I believe a positive result could attract attention from researchers in the high-energy physics community."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"- The motivation of the paper is strong, addressing an important application of equivariance models. While symmetries are prevalent in high-energy physics, there are few applications of equivariance models in this field.\n- The paper clearly distinguishes their work from existing research, emphasizing the significance of their contributions.\n- They connect the proposed model to the generative framework, making it potentially applicable to a broader range of tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The model's performance in Table 1 is not optimal.\n- The experiments on generative modeling lack comparisons with other equivariant models or SOTA models.\n- The scalability of the model remains limited, which reduces its suitability for high-energy physics applications."},"limitations":{"value":"The efficacy and scalability of the proposed framework may not yet suffice to replace traditional methods."}},"nonreaders":[],"tmdate":1730878888512,"tcdate":1720849408821,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission3969/Reviewer_srsp"],"signatures":["NeurIPS.cc/2024/Conference/Submission3969/Reviewer_srsp"],"forum":"X34GKv8sYT","number":3,"license":"CC BY 4.0","cdate":1720849408821,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission3969/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878888512,"domain":"NeurIPS.cc/2024/Conference","replyto":"X34GKv8sYT","id":"9LBwVu1bqd","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"A Lorentz-equivariant Transformer architecture plus Lorentz-equivariant flow matching for high-energy physics"},"keywords":{"value":["Geometric deep learning","equivariance","Lorentz symmetry","Transformer","flow matching","high-energy physics","particle physics"]},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Extracting scientific understanding from particle-physics experiments requires solving diverse learning problems with high precision and good data efficiency. We propose the Lorentz Geometric Algebra Transformer (L-GATr), a new multi-purpose architecture for high-energy physics. L-GATr represents high-energy data in a geometric algebra over four-dimensional space-time and is equivariant under Lorentz transformations, the symmetry group of relativistic kinematics. At the same time, the architecture is a Transformer, which makes it versatile and scalable to large systems. L-GATr is first demonstrated on regression and classification tasks from particle physics. We then construct the first Lorentz-equivariant generative model: a continuous normalizing flow based on an L-GATr network, trained with Riemannian flow matching. Across our experiments, L-GATr is on par with or outperforms strong domain-specific baselines."},"_bibtex":{"value":"@inproceedings{\nspinner2024lorentzequivariant,\ntitle={Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics},\nauthor={Jonas Spinner and Victor Breso Pla and Pim De Haan and Tilman Plehn and Jesse Thaler and Johann Brehmer},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=X34GKv8sYT}\n}"},"title":{"value":"Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics"},"pdf":{"value":"/pdf/8ddc6739ac0f4a9405f3368dba8b88248133a915.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"spinner|lorentzequivariant_geometric_algebra_transformers_for_highenergy_physics"},"authorids":{"value":["~Jonas_Spinner1","~Victor_Breso_Pla1","~Pim_De_Haan1","~Tilman_Plehn1","~Jesse_Thaler1","~Johann_Brehmer1"]},"authors":{"value":["Jonas Spinner","Victor Breso Pla","Pim De Haan","Tilman Plehn","Jesse Thaler","Johann Brehmer"]}},"version":2},{"content":{"venue":{"value":"ICML 2025 poster"},"keywords":{"value":["Surrogate Models","Uncertainty Quantification","Neural-PDE","Physics-Informed","Conformal Prediction","PDE Residuals","Nuclear Fusion"]},"_bibtex":{"value":"@inproceedings{\ngopakumar2025calibrated,\ntitle={Calibrated Physics-Informed Uncertainty Quantification},\nauthor={Vignesh Gopakumar and Ander Gray and Lorenzo Zanisi and Timothy Nunn and Daniel Giles and Matt Kusner and Stanislas Pamela and Marc Peter Deisenroth},\nbooktitle={Forty-second International Conference on Machine Learning},\nyear={2025},\nurl={https://openreview.net/forum?id=Z2uLBBck2X}\n}"},"title":{"value":"Calibrated Physics-Informed Uncertainty Quantification"},"paperhash":{"value":"gopakumar|calibrated_physicsinformed_uncertainty_quantification"},"TLDR":{"value":"Calibrated uncertainty quantification of neural-PDE solvers using physics residual errors as a non-conformity score for conformal prediction."},"primary_area":{"value":"applications->chemistry_physics_and_earth_sciences"},"abstract":{"value":"Simulating complex physical systems is crucial for understanding and predicting phenomena across diverse fields, such as fluid dynamics and heat transfer, as well as plasma physics and structural mechanics. Traditional approaches rely on solving partial differential equations (PDEs) using numerical methods, which are computationally expensive and often prohibitively slow for real-time applications or large-scale simulations. Neural PDEs have emerged as efficient alternatives to these costly numerical solvers, offering significant computational speed-ups. However, their lack of robust uncertainty quantification (UQ) limits deployment in critical applications. We introduce a model-agnostic, physics-informed conformal prediction (CP) framework that provides guaranteed uncertainty estimates without requiring labelled data. By utilising a physics-based approach, we can quantify and calibrate the model's inconsistencies with the physics rather than the uncertainty arising from the data. Our approach utilises convolutional layers as finite-difference stencils and leverages physics residual errors as nonconformity scores, enabling data-free UQ with marginal and joint coverage guarantees across prediction domains for a range of complex PDEs. We further validate the efficacy of our method on neural PDE models for plasma modelling and shot design in fusion reactors."},"link_to_code":{"value":"https://github.com/gitvicky/CP-PRE"},"pdf":{"value":"/pdf/636af320445522f17b899049f808ba441dc6106d.pdf"},"lay_summary":{"value":"This paper addresses a critical problem in using AI for scientific simulations: while neural networks can predict physical systems like weather or plasma behaviour 1000x faster than traditional methods, we don't know when to trust their predictions. We developed CP-PRE (Conformal Prediction with Physics Residual Error), a method that adds reliable \"confidence scores\" to AI predictions by checking how well they obey fundamental physics laws (like conservation of energy) rather than requiring expensive validation data. When tested on applications ranging from fluid dynamics to fusion reactor modelling, the method successfully identifies which predictions are trustworthy with statistical guarantees - essentially adding a physics-based \"confidence meter\" to AI models. This work presents a significant step in safely deploying AI in critical applications like nuclear fusion or aerospace engineering, where you need both speed and reliability, potentially accelerating scientific discovery while maintaining the safety standards these fields require."},"venueid":{"value":"ICML.cc/2025/Conference"},"authorids":{"value":["~Vignesh_Gopakumar1","~Ander_Gray1","~Lorenzo_Zanisi1","~Timothy_Nunn1","~Daniel_Giles1","~Matt_Kusner1","~Stanislas_Pamela1","~Marc_Peter_Deisenroth1"]},"authors":{"value":["Vignesh Gopakumar","Ander Gray","Lorenzo Zanisi","Timothy Nunn","Daniel Giles","Matt Kusner","Stanislas Pamela","Marc Peter Deisenroth"]}},"tmdate":1753293046784,"pdate":1746105548247,"tcdate":1737644809081,"writers":["ICML.cc/2025/Conference","ICML.cc/2025/Conference/Submission11215/Authors"],"signatures":["ICML.cc/2025/Conference/Submission11215/Authors"],"forum":"Z2uLBBck2X","license":"CC BY 4.0","number":11215,"cdate":1737644809081,"readers":["everyone"],"invitations":["ICML.cc/2025/Conference/-/Submission","ICML.cc/2025/Conference/-/Post_Submission","ICML.cc/2025/Conference/Submission11215/-/Full_Submission","ICML.cc/2025/Conference/-/Edit","ICML.cc/2025/Conference/Submission11215/-/Camera_Ready_Revision"],"mdate":1753293046784,"odate":1750231371113,"domain":"ICML.cc/2025/Conference","id":"Z2uLBBck2X","version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Video Generation","Video Diffusion Models"]},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability to accurately understand physics. We found that while the representations within T2V models possess some capacity for physics understanding, they lag significantly behind those from recent video self-supervised learning methods. To this end, we propose a novel framework called {VideoREPA}, which distills physics understanding capability from video understanding foundation models into T2V models by aligning token-level relations. This closes the physics understanding gap and enables more physics-plausible generation. Specifically, we introduce the {Token Relation Distillation (TRD) loss}, leveraging spatio-temporal alignment to provide soft guidance suitable for finetuning powerful pre-trained T2V models—a critical departure from prior representation alignment (REPA) methods. To our knowledge, VideoREPA is the first REPA method designed for finetuning T2V models and specifically for injecting physical knowledge. Empirical evaluations show that VideoREPA substantially enhances the physics commonsense of baseline method, CogVideoX, achieving significant improvement on relevant benchmarks and demonstrating a strong capacity for generating videos consistent with intuitive physics. Code and more video results are available at https://videorepa.github.io/."},"_bibtex":{"value":"@inproceedings{\nzhang2025videorepa,\ntitle={Video{REPA}: Learning Physics for Video Generation through Relational Alignment with Foundation Models},\nauthor={Xiangdong Zhang and Jiaqi Liao and Shaofeng Zhang and Fanqing Meng and Xiangpeng Wan and Junchi Yan and Yu Cheng},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=oHjLfABsK4}\n}"},"title":{"value":"VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models"},"pdf":{"value":"/pdf/4baec3b035c736937132f462800395cda167ca76.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"zhang|videorepa_learning_physics_for_video_generation_through_relational_alignment_with_foundation_models"},"authorids":{"value":["~Xiangdong_Zhang3","~Jiaqi_Liao2","~Shaofeng_Zhang1","~Fanqing_Meng1","~Xiangpeng_Wan1","~Junchi_Yan2","~Yu_Cheng1"]},"authors":{"value":["Xiangdong Zhang","Jiaqi Liao","Shaofeng Zhang","Fanqing Meng","Xiangpeng Wan","Junchi Yan","Yu Cheng"]}},"tmdate":1783627452920,"pdate":1758216563210,"tcdate":1744940517192,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission3019/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission3019/Authors"],"forum":"oHjLfABsK4","license":"CC BY 4.0","number":3019,"cdate":1744940517192,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission3019/-/Full_Submission","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission3019/-/Camera_Ready_Revision"],"mdate":1783627452920,"odate":1761704747973,"domain":"NeurIPS.cc/2025/Conference","id":"oHjLfABsK4","version":2},{"content":{"summary":{"value":"This paper addresses two core failure modes of Physics-Informed Neural Networks (PINNs) on complex PDEs — spectral bias and ill-conditioning — which often cause poor convergence. It identifies two major limitations of existing curriculum learning (CL) approaches for PINNs: (1) unreliable knowledge transfer between curriculum stages and (2) manual, ad-hoc curriculum design.\nTo overcome these issues, the authors propose Neural Operator-based Curriculum Learning (NOCL), a unified framework that leverages neural operators (mainly FNO) to perform functional-space knowledge transfer between stages, and uses the variance of the Neural Tangent Kernel (NTK) spectrum as a universal difficulty measure for automated curriculum construction. Additionally, PDE residual–based masking is used to filter high-quality points for PINN initialization. Experiments on heat, convection, and reaction–diffusion equations show that NOCL significantly improves convergence and generalization compared to baselines."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. **On the NTK variance metric:**  \n   Does its correlation with task difficulty hold under different network widths, activations, or sampling densities?  \n   Have you validated its consistency across multiple architectures or tasks?\n\n2. **Generalization to complex PDEs:**  \n   How does NOCL behave on highly nonlinear, multi-scale, or multi-physics PDEs (e.g., Navier–Stokes, Allen–Cahn, Burgers 2D/3D)?  \n   Are there observed failure cases or stability issues?\n\n3. **Efficiency and resource usage:**  \n   What is the total training time and GPU memory compared to the baselines?  \n   Does the operator training overhead offset the convergence benefits provided by the curriculum?\n\n4. **Operator bootstrapping bias:**  \n   Could early PINN errors accumulate within the operator during self-training?  \n   Beyond masking, have you tried uncertainty-weighted losses or consistency regularization between PDE residuals and operator outputs?\n\n5. **Stronger baselines and ablations:**  \n   Please include comparisons with PINO/DeepONet + physics constraints, and report how the mask ratio and FNO resolution jointly affect performance and stability."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. **Clear motivation and problem definition.**  \n   The paper systematically analyzes the two weaknesses of curriculum learning (CL) in PINNs — parameter inheritance instability and manual curriculum tuning — and provides empirical evidence supporting these claims.\n\n2. **Sound method design.**  \n   - Neural operators enable function-space transfer, avoiding parameter compatibility issues.  \n   - The PDE-residual-based mask filters unreliable operator outputs during initialization.  \n   - The NTK variance serves as a task-agnostic and theoretically interpretable difficulty metric, improving curriculum robustness.\n\n3. **Coherent algorithmic structure.**  \n   Algorithm 1 and Figure 1 clearly describe the closed-loop workflow (*train operator → operator-guided initialization → physics training → feedback*).  \n   This makes the framework reproducible and extensible.\n\n4. **Comprehensive experiments.**  \n   - **Heat equation:** NTK variance correlates with training error; NOCL yields consistent L2 improvements.  \n   - **Convection equation:** Better performance under large β and multi-stage α curricula, with masking ablation.  \n   - **Reaction–diffusion:** Strong results under both ν and ρ curricula, outperforming “causal” CL.\n\n5. **Rich implementation details.**  \n   The appendix includes hyperparameters, curriculum segmentation, and sampling strategy, which facilitate reproduction and adaptation to other PDE problems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited novelty boundary and comparison depth.**  \n   While the combination of operator + curriculum + mask + NTK variance is well-motivated, the paper lacks direct comparisons to *operator-enhanced PINNs* (e.g., PINO, DeepONet with physics constraints) or *NTK/spectral-based curriculum methods*.  \n   Claims of SOTA are only relative to weaker baselines, leaving uncertainty about the position of NOCL among stronger contemporaries.\n\n2. **Weak theoretical justification of NTK variance.**  \n   The metric’s robustness under varying network width, sampling density, or scaling is not systematically analyzed.  \n   Evidence remains empirical, without formal generalization guarantees or complexity estimates for large-scale PDE systems.\n\n3. **Potential bias in operator training.**  \n   The initial operator is trained on PINN predictions rather than high-fidelity ground truth, introducing a “bootstrapping” bias.  \n   Although masking mitigates this, quantifying or comparing against hybrid (small true data + bootstrap) strategies would strengthen the paper’s reliability.\n\n4. **Lack of efficiency evaluation.**  \n   The experiments only report relative L2 errors.  \n   There are no metrics on wall-clock time, GPU hours, convergence speed, or curriculum length–performance trade-offs, which are crucial for evaluating the practicality of NOCL.\n\n5. **Incomplete ablations.**  \n   - *CL w/o NO* only removes the operator component; more variants (e.g., using different difficulty metrics like curvature or loss spectral energy) should be tested.  \n   - The mask ratio α (0.3/0.5) is fixed; a sensitivity analysis or robustness curve is missing, which limits understanding of hyperparameter stability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926077189,"tcdate":1762850737602,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15858/Reviewer_yC3g"],"signatures":["ICLR.cc/2026/Conference/Submission15858/Reviewer_yC3g"],"forum":"2yWis3aAvj","number":4,"license":"CC BY 4.0","cdate":1762850737602,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15858/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926077189,"domain":"ICLR.cc/2026/Conference","replyto":"2yWis3aAvj","id":"RMTZHptDxi","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Curriculum Learning","PINNs","Nerual Operator","PDEs"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"In this paper, we tackle the critical failure modes of Physics-Informed Neural Networks (PINNs), such as spectral bias and ill-conditioning, which lead to poor convergence on complex PDEs. We identify two key shortcomings in existing curriculum learning methods for PINNs: unreliable knowledge transfer between stages and a reliance on manual, ad-hoc curriculum design. To overcome these limitations, we present Neural Operator-based Curriculum Learning (NOCL), a unified framework that leverages Neural Tangent Kernel (NTK) theory to automate curriculum generation and employs neural operators to enable robust, dynamic knowledge transfer across curriculum stages. By dynamically training the operator and filtering data for PINN initialization, our approach ensures scalable and effective learning across progressively difficult tasks. Experiments verify that our proposed NOCL achieves state-of-the-art performance, markedly improving convergence and generalization over existing methods."},"_bibtex":{"value":"@misc{\nmeng2025neural,\ntitle={Neural Operator-based Curriculum Learning for Physics-Informed Neural Networks},\nauthor={Lingshi MENG and Haosen Shi and Sinno Jialin Pan},\nyear={2025},\nurl={https://openreview.net/forum?id=2yWis3aAvj}\n}"},"title":{"value":"Neural Operator-based Curriculum Learning for Physics-Informed Neural Networks"},"pdf":{"value":"/pdf/33e57bacd2211759a9514a8c8ef2e2383a616b20.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"meng|neural_operatorbased_curriculum_learning_for_physicsinformed_neural_networks"},"authorids":{"value":["~Lingshi_MENG2","~Haosen_Shi1","~Sinno_Jialin_Pan1"]},"authors":{"value":["Lingshi MENG","Haosen Shi","Sinno Jialin Pan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes the concept of \"visual knowledge\", referring to the intuitive understanding of the physical world and social knowledge that MLLMs lack. To evaluate this ability systematically, the author constructed a VKBench video benchmark, covering 8 tasks such as intuitive physics and subjective intentions. The evaluation found that the leading model lags behind humans by 15% overall, especially in physical reasoning where the gap is significant. To address this issue, the paper introduces the VKQA dataset and Video VK+ baseline model. The model adopts an \"see think answer\" reasoning format and incorporates visual knowledge rewards through reinforcement learning, achieving a 3.7% improvement on VKBench and demonstrating excellent generalization ability on multiple video benchmarks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Refer to Weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"- 1. Carefully designed benchmark construction: The construction process of VKBench is very rigorous, not simply stacking existing data. The progressive filtering pipeline adopted (Minimize Audio Reliance ->Reduce Language Bias ->Enhance Wrong Candidates ->Human Verification) effectively removes language bias and audio dependencies, ensuring that the benchmark truly evaluates the model's ability to extract knowledge from visual signals, rather than its internal language model's prior knowledge.\n- 2. Large scale evaluation and analysis: The paper comprehensively evaluated 23 mainstream open-source and closed-source MLLMs, not only revealing the overall performance gap, but also conducting detailed analysis (such as differences between open-source and closed-source models, model scale effects, and the impact of reasoning version). Especially through correlation analysis, it was found that the world center and human center knowledge form two independent clusters, which verifies the rationality of their classification from the data.\n- 3. The paper shifts the research focus from traditional \"perception\" (object recognition) and \"complex reasoning\" to the intermediate layer of \"visual knowledge\", which is a relatively unexplored but crucial field. This evaluation and training data are meaningful to the MLLM community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- 1. The coverage of \"visual knowledge\" may still be incomplete: although the division of 8 dimensions is quite comprehensive, \"visual knowledge\" itself is an extremely rich concept. For example, the assessment of deeper cognitive abilities such as causality, counterfactual thinking, and visual metaphor may not have been fully covered in the proposed benchmark.\n- 2. VKBench only uses multiple-choice questions. This cannot fully evaluate whether the reasoning process behind it truly conforms to logic and is based on correct visual knowledge.\n- 3. During the RL training of the Video VK+model, an external frozen MLLM as the verifier is introduced, which is equivalent to handing over a key evaluation criterion to another potentially flawed model. Meanwhile, this significantly increases the complexity and computational cost of training.\n-4. The optimization goal of the Video VK+model is to make the visual description contain answer information as much as possible, thereby reducing the inference burden of LLM. However, it may introduce new risks: will this suppress the model from performing necessary and complex multi-step inference? Overemphasizing the \"self-contained\" visual knowledge may make models \"lazy\" for deep reasoning.\n- 5. The analysis of some experimental phenomena is relatively general, such as the performance of the InternVL3.5-38B-Think on VKBench, which is actually worse than InternVL3.5-38B due to \"excessive reasoning\". Video VK+performs poorly on Spatial Awareness tasks, even inferior to some baseline models. The paper simply attributes the task to requiring 'long-term visual memory'. These explanations are somewhat vague and deserve a deeper analysis."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915777836,"tcdate":1761883708820,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1472/Reviewer_gT1D"],"signatures":["ICLR.cc/2026/Conference/Submission1472/Reviewer_gT1D"],"forum":"P798W8Ag7L","number":2,"license":"CC BY 4.0","cdate":1761883708820,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1472/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915777836,"domain":"ICLR.cc/2026/Conference","replyto":"P798W8Ag7L","id":"IJdeSP28ut","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multimodal Large Language Model","Video Large Language Model","Visual Knowledge","Benchmark","Datasets"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical and social principles. This capability, which we term visual knowledge, forms a bridge between perception and reasoning, yet remains an underexplored gap in current systems.\nTo systematically measure this capability, we present VKBench, a comprehensive video benchmark featuring 1,680 questions in 1,249 videos, covering eight core types of visual knowledge spanning both world-centric (e.g., intuitive physics) and human-centric (e.g., subjective intentions). Results show that leading models still fall short of human performance, with particularly notable gaps in world-centric visual knowledge.\nTo bridge this gap, we introduce VKQA, a new dataset, and Video-VK+, a baseline model that explicitly incorporates visual knowledge into MLLMs. Video-VK+ follows a structured See–Think–Answer format and adopts reinforcement learning with visual knowledge reward. This approach improves performance on VKBench by 3.7% and surpasses existing models on multiple video benchmarks.\nOur findings highlight visual knowledge as a key component for developing more robust and generalizable MLLMs that can not only see but also truly understand our world."},"_bibtex":{"value":"@misc{\njiang2025benchmarking,\ntitle={Benchmarking Visual Knowledge in Multimodal Large Language Models},\nauthor={Tianxiang Jiang and Sheng Xia and Yicheng Xu and Linquan Wu and Xiangyu Zeng and Limin Wang and Yu Qiao and Yi Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=P798W8Ag7L}\n}"},"title":{"value":"Benchmarking Visual Knowledge in Multimodal Large Language Models"},"pdf":{"value":"/pdf/186de1c1275cf653df49719895560101bf3f1938.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"jiang|benchmarking_visual_knowledge_in_multimodal_large_language_models"},"authorids":{"value":["~Tianxiang_Jiang1","~Sheng_Xia1","~Yicheng_Xu1","~Linquan_Wu1","~Xiangyu_Zeng4","~Limin_Wang1","~Yu_Qiao1","~Yi_Wang19"]},"authors":{"value":["Tianxiang Jiang","Sheng Xia","Yicheng Xu","Linquan Wu","Xiangyu Zeng","Limin Wang","Yu Qiao","Yi Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces generalized stationary SGD with infinite memory, a theoretical framework for accelerating stochastic gradient descent on infinite-dimensional quadratic problems with power-law spectral conditions. The authors build upon classical results showing that deterministic Gradient Descent achieves a convergence rate of $O(t^{-\\zeta})$, and that Heavy Ball with a non-stationary Jacobi schedule can accelerate this rate to $O(t^{-2\\zeta})$. However, in stochastic mini-batch settings, no such results are known.\nThe contribution of the paper is a geometric representation of stationary (S)GD algorithms as contours in the complex plane. The authors show that contours with a corner of external angle $\\theta\\pi$ accelerate the plain GD rate from $O(t^{-\\zeta})$ to $O(t^{-\\theta\\zeta})$. In deterministic settings, increasing $\\theta$ up to $2$ gives rates arbitrarily close to $O(t^{-2\\zeta})$. In stochastic setting, increasing $\\theta$ also amplifies noise, leading to a trade-off, with optimal rate $\\theta_{\\max} = \\min(2, \\nu, \\tfrac{2}{\\zeta + 1/\\nu})$, where $\\nu$ and $\\zeta$ are the spectral exponents coming from the capacity and source conditions.\n\nTo make Corner SGD practical, the paper proposes finite-memory approximations of the infinite-memory algorithms using rational approximations of the power function $z^\\theta$. Empirical experiments on a synthetic problem and MNIST show accelerated convergence consistent with the theoretical predictions."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"1. Can the authors comment on how the framework might extend beyond quadratic case?\n\n2. How should one choose the memory size $M$ in practice?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The contour-based geometric representation of stationary (S)GD algorithms is new and interesting, connecting optimization dynamics with complex analysis.\n\n2. The paper unifies acceleration, stability, and noise amplification under a single geometric-spectral framework.\n\n3. Rigorous theoretical results.\n\n4. The finite-memory approximation gives a practical way to approximate the idealized algorithms."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper is technically rigorous and presents a new geometric perspective on SGD acceleration, but several assumptions (only quadratic case, SE, infinite-dimensional limit) narrow the practical relevance. It is unclear if the main conclusions hold beyond this idealized regime.\n\n2. Limited experimental validation (one synthetic experiment with MNIST). However, I acknowledge that the paper is of theoretical nature."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924242853,"tcdate":1761764065056,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13681/Reviewer_yRoH"],"signatures":["ICLR.cc/2026/Conference/Submission13681/Reviewer_yRoH"],"forum":"nOXCfIdhD9","number":3,"license":"CC BY 4.0","cdate":1761764065056,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13681/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924242853,"domain":"ICLR.cc/2026/Conference","replyto":"nOXCfIdhD9","id":"1iX8p40D1y","forumContent":{"TLDR":{"value":"Generalized (S)GD algorithms = contours in $\\mathbb C$. Contours having a corner with external angle $\\theta\\pi, 1<\\theta<2,$ accelerate loss convergence rates $t^{-\\xi}$ to $t^{-\\theta\\xi}$."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["mini-batch stochastic gradient descent","momentum","sampling noise","convergence rates","acceleration","power laws","phase diagram","contour integration","rational approximations","asymptotic methods","MNIST","frequency response function"]},"supplementary_material":{"value":"/attachment/c2514bd7d34c3f980bbaab207f885b2a8d237bda.zip"},"primary_area":{"value":"optimization"},"abstract":{"value":"We consider SGD-type optimization on infinite-dimensional quadratic problems with power law spectral conditions. It is well-known that on such problems deterministic GD has loss convergence rates $L_t=O(t^{-\\zeta})$, which can be improved to $L_t=O(t^{-2\\zeta})$ by using Heavy Ball with a non-stationary Jacobi-based schedule (and the latter rate is optimal among fixed schedules). However, in the mini-batch Stochastic GD setting, the sampling noise causes the Jacobi HB to diverge; accordingly no $O(t^{-2\\zeta})$ algorithm is known. In this paper we show that rates up to $O(t^{-2\\zeta})$ can be achieved by a generalized stationary SGD with infinite memory. We start by identifying  generalized (S)GD algorithms with contours in the complex plane. We then show that contours that have a corner with external angle $\\theta\\pi$ accelerate the plain GD rate $O(t^{-\\zeta})$ to $O(t^{-\\theta\\zeta})$. For deterministic GD, increasing $\\theta$ allows to achieve rates arbitrarily close to $O(t^{-2\\zeta})$. However, in Stochastic GD, increasing $\\theta$ also amplifies the sampling noise, so in general $\\theta$ needs to be optimized by balancing the acceleration and noise effects. We prove that the optimal rate is given by $\\theta_{\\max}=\\min(2,\\nu,\\tfrac{2}{\\zeta+1/\\nu})$, where $\\nu,\\zeta$ are the exponents appearing in the capacity and source spectral conditions. Furthermore, using fast rational approximations of the power functions, we show that ideal corner algorithms can be efficiently approximated by practical finite-memory algorithms."},"_bibtex":{"value":"@inproceedings{\nyarotsky2026corner,\ntitle={Corner Gradient Descent},\nauthor={Dmitry Yarotsky},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=nOXCfIdhD9}\n}"},"title":{"value":"Corner Gradient Descent"},"pdf":{"value":"/pdf/f22a5e7ef3ee671039e702d5f9ab13984be1e906.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yarotsky|corner_gradient_descent"},"authorids":{"value":["~Dmitry_Yarotsky1"]},"authors":{"value":["Dmitry Yarotsky"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewer for the insightful comments. Below, we clarify our metric design, discuss the release of an end-to-end evaluation pipeline, and address the relation to prior benchmarks.\n\n**Re: (1) Lack of a clear standalone pipeline or evaluation API**\n\nWe do provide a full standalone evaluation pipeline, and the complete codebase is included in the supplementary material. The released package contains a detailed `README.md` with step-by-step instructions for running evaluations.\n\nIn short, users can directly apply **`evaluate_videos.py`** to any preprocessed set of videos without needing to modify or generate additional metadata. The workflow is identical to other widely used evaluation pipelines—researchers simply supply their model’s generated videos, and the evaluator handles all scoring automatically. Please refer to the supplementary material for full instructions.\n\n---\n\n**Re: (2) Metrics such as physics consistency are described too abstractly**\n\nBoth **Physics Consistency** and **Semantic Adherence** are defined using *explicit, fine-grained, and sample-specific standards*, not high-level abstractions. Figure 5 in the main paper shows an example of this structure, and the full set of standards for all 1,050 prompts is provided in the supplementary file **`prompts-with-standard-and-index.json`**.\n\nAs noted in line 267 of the main paper, our terminology follows prior work:  \n*“Following (Bansal et al., 2024; Meng et al., 2024), we use Semantic Adherence (SA) and Physical Commonsense (PC) to evaluate video performance.”*  \nThese metrics are therefore consistent with existing evaluation literature.\n\nRegarding the binary scoring: when a model “rationalizes” an action—producing outputs that avoid or obscure physics violations rather than following the intended scenario—the video is **not** adhering to the prompt or the underlying physics. In these cases, the correct and reproducible behavior is to mark the sample as *false* for physics consistency. This provides a clear and unambiguous scoring rule.\n\nAll metric definitions and standards are therefore fully reproducible from the released materials and the accompanying JSON specification.\n\n---\n\n**Re: (3) Limited real-world coverage**\n\nOur benchmark is designed specifically for **evaluating text-to-video generation models**, and therefore does not involve real-world video datasets or robotic interaction data. Because current generative models only produce synthetic outputs, it is natural and appropriate for the benchmark to evaluate synthetic videos as well.  \n\nAdditionally, we would like to clarify that **MuJoCo** and **Unity** are *not* used anywhere in our paper. If the reviewer was referring to another benchmark or a misunderstanding, we would appreciate further clarification so we can address the concern more precisely."},"title":{"value":"Rebuttal 1/2"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764146065757,"tcdate":1764146065757,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8464/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission8464/Authors"],"forum":"rlZeILv3fm","number":5,"license":"CC BY 4.0","cdate":1764146065757,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8464/-/Official_Comment"],"mdate":1764146065757,"domain":"ICLR.cc/2026/Conference","replyto":"tgrWH9jead","id":"nCKBEAQoa2","forumContent":{"venue":{"value":"ICLR 2026 Oral"},"TLDR":{"value":"Large-scale, multidimensional video generation for physics"},"keywords":{"value":["Video Generation","Video Evaluation"]},"supplementary_material":{"value":"/attachment/bd8cdaa60661c40d3e703b3dd09ff6689653ba72.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents $PhyWorldBench$\n, a comprehensive benchmark designed to evaluate video generation models based on their adherence to the laws of physics. The benchmark covers multiple levels of physical phenomena, ranging from fundamental principles like object motion and energy conservation to more complex scenarios involving rigid body interactions and human or animal motion. Additionally, we introduce a novel \"Anti-Physics\" category, where prompts intentionally violate real-world physics, enabling the assessment of whether models can follow such instructions while maintaining logical consistency. Besides large-scale human evaluation, we also design a simple yet effective method that could utilize current MLLM to evaluate the physics realism in a zero-shot fashion. We evaluate 10 state-of-the-art text-to-video generation models, including five open-source and five proprietary models, with a detailed comparison and analysis. we identify pivotal challenges models face in adhering to real-world physics. Through systematic testing of their outputs across 1,050 curated prompts—spanning fundamental, composite, and anti-physics scenarios—we identify pivotal challenges these models face in adhering to real-world physics. We then rigorously examine their performance on diverse physical phenomena with varying prompt types, deriving targeted recommendations for crafting prompts that enhance fidelity to physical principles."},"_bibtex":{"value":"@inproceedings{\ngu2026phyworldbench,\ntitle={\\$PhyWorldBench\\$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models},\nauthor={Jing Gu and Xian Liu and Yu Zeng and Ashwin Nagarajan and Fangrui Zhu and Daniel Hong and Yue Fan and Qianqi Yan and Kaiwen Zhou and Ming-Yu Liu and Xin Eric Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=rlZeILv3fm}\n}"},"title":{"value":"$PhyWorldBench$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models"},"pdf":{"value":"/pdf/6138b5d05836a6d9ee27ae5fc6f3bbbe3667ff02.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gu|phyworldbench_a_comprehensive_evaluation_of_physical_realism_in_texttovideo_models"},"authorids":{"value":["~Jing_Gu2","~Xian_Liu1","~Yu_Zeng1","~Ashwin_Nagarajan1","~Fangrui_Zhu1","~Daniel_Hong1","~Yue_Fan3","~Qianqi_Yan1","~Kaiwen_Zhou3","~Ming-Yu_Liu1","~Xin_Eric_Wang2"]},"authors":{"value":["Jing Gu","Xian Liu","Yu Zeng","Ashwin Nagarajan","Fangrui Zhu","Daniel Hong","Yue Fan","Qianqi Yan","Kaiwen Zhou","Ming-Yu Liu","Xin Eric Wang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a method for effectively learning video representations from synthetic videos without the need for training on natural videos. It introduces simple generation steps, such as moving, transforming, and accelerating shapes, to create a series of synthetic video datasets. The authors explore the performance of models pre-trained on these synthetic datasets in downstream tasks, focusing primarily on action recognition using VideoMAE. Experiments on the UCF101 and HMDB51 datasets show that the performance of models pre-trained on synthetic data is comparable to those pre-trained on natural data. Additionally, results from UCF101-P demonstrate that models pre-trained on synthetic data exhibit similar robustness."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Could you provide experimental results from other datasets, such as Something-SomethingV2?\n2. What type of generator is used in this study? The paper mentions an on-the-fly generation strategy for training (Page 4, Line 215). Does this approach consume more computational resources than the original pre-training? Please provide relevant experimental data.\n3. In the paragraph on Page 6, Line 292, how were the experimental data obtained? Is there a missing table or figure? How was the conclusion reached? Additionally, how is the “97.2%” figure mentioned multiple times in the paper derived? How are the experimental data in Section 4.3 calculated?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The proposed method generates synthetic data through operations on shapes (e.g., accelerating, transforming), leading to models that perform comparably to those trained on natural data in downstream tasks. This approach is relatively simple and novel, distinguishing itself from previous work centered on human-based synthetic data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks completeness:\n   - a) It primarily focuses on synthetic data for action recognition (Page 6, Line 511). Still, it does not discuss or compare its experiments with similar works by [2] and [3]. Furthermore, the models and datasets used (VideoMAE and UCF101/HMDB51) are less robust compared to these studies.\n   - b) The UCF101 and HMDB51 datasets are relatively simple and do not adequately demonstrate the advantages of synthetic data for video pre-training. Experiments on more complex datasets, like Something-SomethingV2, are needed to support the claims.\n   - c) Section 6 mentions that this is merely a preliminary work and suggests additional tasks or models will be discussed in future work, leading to the conclusion that this paper does not constitute a complete study.\n\n2. The overall writing resembles a technical (experimental) report, particularly in Section 5, which is structured similarly to [1]. While the paper presents work on videos, it does not adequately expand on this area compared to the image work done by [1].\n\n3. The two synthetic datasets generated, “Accelerating Transforming StyleGAN Crops” and “Accelerating Transforming Textures,” yield good performance for the pre-trained models. However, these datasets are based on data generation techniques introduced by [1].\n\n4. There are writing issues:\n   - Page 3, Line 140: Incorrect citation for “Section 3.1.”\n   - Page 10, Line 537: The number “92.%” is incomplete.\n   - Page 8, Line 424: The origin of “28 datasets” is not explained.\n\n5. Section 5 analyzes how synthetic data benefits video pre-training only from the perspective of static images or individual frames (spatial information). Given that videos are characterized by temporal dynamics and motion cues, the paper lacks relevant experiments and analyses in these areas.\n\n6. The accuracy of the “Dynamic StyleGAN videos” setting in Table 3 is reported to be only 68.7%, but the paper does not provide an explanation or conclusion regarding this result.\n\n[1]\tManel Baradad, Jonas Wulff, Tongzhou Wang, Phillip Isola, and Antonio Torralba. Learning to see by looking at noise. In A. Beygelzimer, Y. Dauphin, P. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, 2021. URL https://openreview. net/forum?id=RQUl8gZnN7O.\n\n[2]\tYoWhan Kim, Samarth Mishra, SouYoung Jin, Rameswar Panda, Hilde Kuehne, Leonid Karlinsky, Venkatesh Saligrama, Kate Saenko, Aude Oliva, and Rogerio Feris. How transferable are video representations based on synthetic data? In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. URL https://openreview. net/forum?id=lRUCfzs5Hzg.\n\n[3]\tHoward Zhong, Samarth Mishra, Donghyun Kim, SouYoung Jin, Rameswar Panda, Hilde Kuehne, Leonid Karlinsky, Venkatesh Saligrama, Aude Oliva, Rogerio Feris. Learning Human Action Recognition Representations Without Real Humans. In Thirty-seventh Conference on Neural Information Processing Systems Track on Datasets and Benchmarks. URL https://openreview.net/forum?id=UBbm5embIB"}},"nonreaders":[],"tmdate":1731427352303,"tcdate":1730706526750,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission985/Reviewer_dyKo"],"signatures":["ICLR.cc/2025/Conference/Submission985/Reviewer_dyKo"],"forum":"xz3dmxfFva","number":3,"license":"CC BY 4.0","cdate":1730706526750,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission985/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427352303,"domain":"ICLR.cc/2025/Conference","replyto":"xz3dmxfFva","id":"5n1JbNy16U","forumContent":{"TLDR":{"value":"We learn robust video representations from synthetic videos and natural images"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video representation learning","learning from synthetic data"]},"supplementary_material":{"value":"/attachment/c06aa60b8770ad18fba72b26d7a3a025452e6a1d.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this paper, we show that useful video representations can be learned from synthetic videos and natural images, without incorporating natural videos in the training. We propose a progression of video datasets synthesized by simple generative processes, that model a growing set of natural video properties (e.g. motion, acceleration, and shape transformations). The downstream performance of video models pre-trained on these generated datasets gradually increases with the dataset progression. A VideoMAE model pre-trained on our synthetic videos closes 97.2\\% of the performance gap on UCF101 action classification between training from scratch and self-supervised pre-training from natural videos, and outperforms the pre-trained model on HMDB51. Introducing crops of static images to the pre-training stage results in similar performance to UCF101 pre-training and outperforms the UCF101 pre-trained model on 11 out of 14 out-of-distribution datasets of UCF101-P. Analyzing the low-level properties of the datasets, we identify correlations between frame diversity, frame similarity to natural data, and downstream performance. Our approach provides a more controllable and transparent alternative to video data curation processes for pre-training."},"_bibtex":{"value":"@misc{\nyu2024video,\ntitle={Video Representation Learning Without Natural Videos},\nauthor={Xueyang Yu and Xinlei Chen and Yossi Gandelsman},\nyear={2024},\nurl={https://openreview.net/forum?id=xz3dmxfFva}\n}"},"title":{"value":"Video Representation Learning Without Natural Videos"},"pdf":{"value":"/pdf/8a00313bef9b8b167c320215d1076d1e040eb836.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"yu|video_representation_learning_without_natural_videos"},"authorids":{"value":["~Xueyang_Yu1","~Xinlei_Chen1","~Yossi_Gandelsman2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xueyang Yu","Xinlei Chen","Yossi Gandelsman"]}},"version":2},{"content":{"summary":{"value":"This paper tackles causal effect estimation under hidden confounding using nonparametric instrumental variable regression. Existing spectral feature learning methods rely on features from the top singular subspaces of the treatment–instrument operator, which are outcome-agnostic and may fail when the true causal function is misaligned with these subspaces. To overcome this, the authors propose Augmented Spectral Feature Learning, which introduces outcome-awareness by constructing an augmented operator that integrates information from the outcome variable. A novel contrastive loss enables learning task-specific spectral features robust to spectral misalignment. The paper provides theoretical guarantees and empirical results demonstrating improved accuracy and robustness on challenging synthetic and semi-synthetic benchmarks."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. The paper assumes sub-Gaussian distributions (Assumption 4), which might be restrictive in practice. Could the authors clarify how sensitive their theoretical results are to this assumption, and whether the framework could extend to heavier-tailed (e.g., sub-exponential) settings?\n2. The paper is quite dense and may be difficult to follow for readers unfamiliar with spectral feature learning."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper identifies a fundamental limitation of spectral NPIV methods (outcome agnosticism leading to failure under spectral misalignment) and proposes a principled solution.\n2. The proposed method is supported by rigorous theoretical analysis.\n3. Experimental results demonstrate that the proposed approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The method enhances the operator using a rank-one augmentation, which simplifies theory but limits its ability to capture complex, multidimensional dependencies between $Y$ and the instrument–treatment relationship.\n\n2. All experiments are conducted on synthetic or semi-synthetic datasets (e.g., dSprites), leaving the method’s practical effectiveness in noisy, unstructured real-world environments untested.\n\n3. The paper evaluates the proposed method against only a few baselines (e.g., KIV, DFIV), missing more recent state-of-the-art approaches for nonparametric IV regression. For an ICLR submission, it would be beneficial to include comparisons with modern representation learning–based IV methods, to better contextualize the contribution and demonstrate broader relevance."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927939111,"tcdate":1761750538309,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18188/Reviewer_3F2d"],"signatures":["ICLR.cc/2026/Conference/Submission18188/Reviewer_3F2d"],"forum":"ZPam0jRjTC","number":1,"license":"CC BY 4.0","cdate":1761750538309,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18188/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927939111,"domain":"ICLR.cc/2026/Conference","replyto":"ZPam0jRjTC","id":"VYvoWFTDb3","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["instrumental variable regression","NPIV","nonparametric statistics","feature learning","causal inference","operator learning"]},"primary_area":{"value":"learning theory"},"abstract":{"value":"We address the problem of causal effect estimation in the presence of hidden confounders using nonparametric instrumental variable (IV) regression. An established approach is to use estimators based on learned \\emph{spectral features}, that is, features spanning the top singular subspaces of the operator linking treatments to instruments. While powerful, such features are agnostic to the outcome variable. Consequently, the method can fail when the true causal function is poorly represented by these dominant singular functions.\n\nTo mitigate, we introduce **Augmented Spectral Feature Learning**, a framework that makes the feature learning process **outcome-aware**. Our method learns features by minimizing a novel contrastive loss derived from an **augmented** operator that incorporates information from the outcome. By learning these task-specific features, our approach remains effective even under spectral misalignment. We provide a theoretical analysis of this framework and validate our approach on challenging benchmarks."},"_bibtex":{"value":"@misc{\nmeunier2026outcomeaware,\ntitle={Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression},\nauthor={Dimitri Meunier and Jakub Wornbard and Vladimir R Kostic and Antoine Moulin and Karim Lounici and Alek Fr{\\\"o}hlich and Massimiliano Pontil and Arthur Gretton},\nyear={2026},\nurl={https://openreview.net/forum?id=ZPam0jRjTC}\n}"},"title":{"value":"Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression"},"pdf":{"value":"/pdf/f3fc802ef938b77f0cb0d5838ed128f9bb459536.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"meunier|outcomeaware_spectral_feature_learning_for_instrumental_variable_regression"},"authorids":{"value":["~Dimitri_Meunier1","~Jakub_Wornbard1","~Vladimir_R_Kostic1","~Antoine_Moulin1","~Karim_Lounici1","~Alek_Fröhlich1","~Massimiliano_Pontil4","~Arthur_Gretton1"]},"authors":{"value":["Dimitri Meunier","Jakub Wornbard","Vladimir R Kostic","Antoine Moulin","Karim Lounici","Alek Fröhlich","Massimiliano Pontil","Arthur Gretton"]}},"version":2},{"content":{"summary":{"value":"This paper presents a new approach to integrate physics-based tumor growth constraint with multi-modal imaging data to predict tumor cell distribution and thus enhance tumor treatment planning for glioblastomas. The core of the methods includes a discrete physics residual and initial assumptions encoding initial tumor distribution and symmetrical pattern between hemispheres. Experiments were performed on a small real dataset, and the success was measured by the “recurrence coverage” defined as the percentage of tumor (based on follow-up MRI) within the model-predicted volume. Experimental results demonstrate improved recurrence coverage of the presented method compared to fully data-driven or fully physics-based methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Clarifications on the data split used to derive the reported results will be appreciated. \n\nThe statistical significance of the reported performance (Table 1) will be appreciated. \n\nClarifications on the results related to Fig 5 (see questions above) will be appreciated. \n\nHow sensitive are the presented methods to the different terms in the loss functions, and which are the most important terms?\n\nHow is asymmetry measured as an optimization objective?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Ability to integrate physics-based constraints and multi-modal imaging data is important and a difficult problem to address. The presented work is thus of important potential in bringing physics-informed learning into biomedical tasks.\n\nThe dataset and task considered are quite complex involving multi-modal images and pre-operative and follow-up comparisons. I applaud the authors for the efforts devoted in this type of challenging tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The experimental evaluation of the presented work was performed on a very small dataset with unclear training-test split. This raises some question about the general conclusion that can be drawn from the results obtained on limited samples.\n\nGiven the reported standard deviation (eg. Table 1), the obtained margin of improvement appears to be marginal without statistical significance. \n\nThe writing of the paper lacks clarify at places. For instance, the results presented in Figure 5 was unclear. What are the “greater”, “less”, and “equal” categories, and why a bigger gap between the “greater” and “less” categories indicate a favorable performance?\n\nThe objective of the presented model includes a large number of terms. The sensitivity of the model performance to the include/exclusion of these different terms and their hyperparameters deserves substantial analyses that are missing in the current paper."},"limitations":{"value":"The authors presented some discussion about the limitations of the current work. Adding discussion about limitations related to the limited sample size for the experiments would be appreciated."}},"nonreaders":[],"tmdate":1730879912866,"tcdate":1720773407169,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission17469/Reviewer_YNED"],"signatures":["NeurIPS.cc/2024/Conference/Submission17469/Reviewer_YNED"],"forum":"YfVMcbcDqo","number":3,"license":"CC BY 4.0","cdate":1720773407169,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission17469/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879912866,"domain":"NeurIPS.cc/2024/Conference","replyto":"YfVMcbcDqo","id":"2gY27ETx49","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"We propose a physics-regularized learning approach on dynamic discrete meshes to address the complex inverse problem of tumor localization."},"keywords":{"value":["Inverse Problems","System Identification","Physics-Informed","Biomechanical Modeling","Tumor Growth"]},"primary_area":{"value":"machine_learning_for_healthcare"},"flagged_for_ethics_review":{"value":true},"abstract":{"value":"Physical models in the form of partial differential equations serve as important priors for many under-constrained problems. One such application is tumor treatment planning, which relies on accurately estimating the spatial distribution of tumor cells within a patient’s anatomy. While medical imaging can detect the bulk of a tumor, it cannot capture the full extent of its spread, as low-concentration tumor cells often remain undetectable, particularly in glioblastoma, the most common primary brain tumor. Machine learning approaches struggle to estimate the complete tumor cell distribution due to a lack of appropriate training data. Consequently, most existing methods rely on physics-based simulations to generate anatomically and physiologically plausible estimations. However, these approaches face challenges with complex and unknown initial conditions and are constrained by overly rigid physical models. In this work, we introduce a novel method that integrates data-driven and physics-based cost functions, akin to Physics-Informed Neural Networks (PINNs). However, our approach parametrizes the solution directly on a dynamic discrete mesh, allowing for the effective modeling of complex biomechanical behaviors. Specifically, we propose a unique discretization scheme that quantifies how well the learned spatiotemporal distributions of tumor and brain tissues adhere to their respective growth and elasticity equations. This quantification acts as a regularization term, offering greater flexibility and improved integration of patient data compared to existing models. We demonstrate enhanced coverage of tumor recurrence areas using real-world data from a patient cohort, highlighting the potential of our method to improve model-driven treatment planning for glioblastoma in clinical practice."},"_bibtex":{"value":"@inproceedings{\nbalcerak2024physicsregularized,\ntitle={Physics-Regularized Multi-Modal Image Assimilation for Brain Tumor Localization},\nauthor={Michal Balcerak and Tamaz Amiranashvili and Andreas Wagner and Jonas Weidner and Petr Karnakov and Johannes C. Paetzold and Ivan Ezhov and Petros Koumoutsakos and Benedikt Wiestler and bjoern menze},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=YfVMcbcDqo}\n}"},"title":{"value":"Physics-Regularized Multi-Modal Image Assimilation for Brain Tumor Localization"},"pdf":{"value":"/pdf/d652352e6b706e6110a4a40f48dfc2eec542165a.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"balcerak|physicsregularized_multimodal_image_assimilation_for_brain_tumor_localization"},"authorids":{"value":["~Michal_Balcerak1","~Tamaz_Amiranashvili1","~Andreas_Wagner1","~Jonas_Weidner1","~Petr_Karnakov1","~Johannes_C._Paetzold1","~Ivan_Ezhov1","~Petros_Koumoutsakos1","~Benedikt_Wiestler1","~bjoern_menze1"]},"authors":{"value":["Michal Balcerak","Tamaz Amiranashvili","Andreas Wagner","Jonas Weidner","Petr Karnakov","Johannes C. Paetzold","Ivan Ezhov","Petros Koumoutsakos","Benedikt Wiestler","bjoern menze"]}},"version":2},{"content":{"summary":{"value":"This paper provides a theoretical analysis comparing single-modal and multi-modal contrastive learning. The authors develop a unified framework to analyze the optimization dynamics and generalization capabilities of both approaches. Key findings include:\n\n- Both single-modal and multi-modal contrastive learning can achieve small training error.\n- Multi-modal contrastive learning generalizes better to downstream tasks compared to single-modal learning.\n- The advantage of multi-modal learning comes from cooperation between modalities and higher quality data in the second modality.\n\nThe analysis is based on a two-stage optimization process and uses a signal-noise data generation model. Theoretical results are supported by synthetic experiments."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Line 205: “On the contrary, augmentation often maintains the same SNR as the original data, so single-modal learning hardly benefits from the augmentation and can only memorize the noise from the data.” This claim significantly contradicts empirical experiments such as SimCLR and MoCo. How would you justify this?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Develops a unified framework to analyze both single-modal and multi-modal contrastive learning, allowing for direct comparisons.\n\n- Provides detailed theoretical analysis, including convergence guarantees and generalization bounds.\n\n- Identifies key factors (signal-to-noise ratio, cooperation between modalities) that explain the superior performance of multi-modal learning.\n\n- Supports theoretical findings with synthetic experiments that align well with the analysis."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper could benefit from more intuitive explanations of the key theoretical results to make them more accessible to a broader audience. For example, can you provide intuition for why the cooperation between modalities leads to better signal learning in the multi-modal case? How might this insight be leveraged to improve single-modal contrastive learning?\n\n- The paper lacks discussion on how the insights could be applied to improve existing contrastive learning methods or guide the development of new approaches. Do your theoretical insights suggest any practical strategies for improving multi-modal contrastive learning, such as ways to select or preprocess data to increase the effective signal-to-noise ratio?\n\n- The theoretical setup and assumption of the linear data generation model are somewhat simple and restricted. Do you expect the main insights to hold for more complex data distributions?\n\n- The experimental validation is limited to synthetic data. Including experiments on real-world datasets, even if simplified, would strengthen the paper's impact."},"limitations":{"value":"See weaknesses."}},"nonreaders":[],"tmdate":1730878650243,"tcdate":1720842137078,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission555/Reviewer_bcfY"],"signatures":["NeurIPS.cc/2024/Conference/Submission555/Reviewer_bcfY"],"forum":"O2UwxfhY1P","number":3,"license":"CC BY 4.0","cdate":1720842137078,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission555/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878650243,"domain":"NeurIPS.cc/2024/Conference","replyto":"O2UwxfhY1P","id":"0Uan1Fqbmk","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["multi-modal contrastive learning","single-modal contrastive learning","optimization","learning theory"]},"supplementary_material":{"value":"/attachment/5351bd5a1322b1fac330ef5e0f4bda82e6a6c674.zip"},"primary_area":{"value":"other"},"abstract":{"value":"Multi-modal contrastive learning with language supervision has presented a paradigm shift in modern machine learning. By pre-training on a web-scale dataset, multi-modal contrastive learning can learn high-quality representations that exhibit impressive robustness and transferability. Despite its empirical success, the theoretical understanding is still in its infancy, especially regarding its comparison with single-modal contrastive learning. In this work, we introduce a feature learning theory framework that provides a theoretical foundation for understanding the differences between multi-modal and single-modal contrastive learning. Based on a data generation model consisting of signal and noise, our analysis is performed on a ReLU network trained with the InfoMax objective function. Through a trajectory-based optimization analysis and generalization characterization on downstream tasks, we identify the critical factor, which is the signal-to-noise ratio (SNR), that impacts the generalizability in downstream tasks of both multi-modal and single-modal contrastive learning. Through the cooperation between the two modalities, multi-modal learning can achieve better feature learning, leading to improvements in performance in downstream tasks compared to single-modal learning. Our analysis provides a unified framework that can characterize the optimization and generalization of both single-modal and multi-modal contrastive learning. Empirical experiments on both synthetic and real-world datasets further consolidate our theoretical findings."},"_bibtex":{"value":"@inproceedings{\nhuang2024on,\ntitle={On the Comparison between Multi-modal and Single-modal Contrastive Learning},\nauthor={Wei Huang and Andi Han and Yongqiang Chen and Yuan Cao and zhiqiang xu and Taiji Suzuki},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=O2UwxfhY1P}\n}"},"title":{"value":"On the Comparison between Multi-modal and Single-modal Contrastive Learning"},"pdf":{"value":"/pdf/c9933b975dba8940756388b4decd66b29013d72e.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"huang|on_the_comparison_between_multimodal_and_singlemodal_contrastive_learning"},"authorids":{"value":["~Wei_Huang6","~Andi_Han1","~Yongqiang_Chen1","~Yuan_Cao1","~zhiqiang_xu1","~Taiji_Suzuki1"]},"authors":{"value":["Wei Huang","Andi Han","Yongqiang Chen","Yuan Cao","zhiqiang xu","Taiji Suzuki"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on comparing the sample efficiency advantages of MARL and SARL approaches in LLM agentic systems. Based on the PAC framework, it formally defines the LLM sequential decision problem and presents a comparison of the efficiency of MARL and SARL under different subtask conditions, along with quantitative analysis and verification. The conclusions are validated using synthetic tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.\tThe theoretical bounds depend on hyperparameters such as $d,d_i,B,B_i,L_{\\rm step},T_{\\rm max}$. How do these quantities relate to the actual number of parameters in a high-dimensional LLM? Can you provide a theoretical and practical connection?\n\n2.\tCan you add supplementary experiments to bridge the gap between the paper's background and experiments, and provide additional experimental support for the relevant theory?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"LLM agentic systems have recently been frequently explored for complex task processing based on MARL. While many empirical algorithms have emerged for MARL-based approaches, theoretical guidance is currently lacking. \n\nThis paper presents a rigorous theoretical derivation and, in the LLM agentic setting, provides a systematic PAC comparison and testable thresholds for SARL vs. MARL. In particular, the introduced \"task alignment\" factor, $\\alpha$, quantifies the sample cost of MARL strategies in complex real-world situations, providing strong theoretical guidance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tExperimental Verification: First, I think there is a gap between the experimental verification and and motivation & background, which are based on a complex LLM agentic system. Section 4.4 only uses a lightweight synthetic linear task. The \"dependence\" in the synthetic experiment (i.e., the current output depends on the mean of the previous output) is completely different from the complex semantics, logic, and state dependencies in the LLM agentic system. Therefore, while this experiment mathematically verifies the theory, it does not prove that the theory can be transferred to the LLM problem it claims to solve. Second, the experiment lacks some key theoretical verifications. For example, in Section 4.3, the paper proposes a task misalignment and defines a \"misalignment factor\" α, deriving new sample bounds. However, the experimental section does not verify α and its PAC bounds at all. Furthermore, the paper itself mentions in its limitations that the experiment only tests the two extreme cases of \"complete independence\" and \"complete dependence,\" lacking exploration of intermediate scenarios and coverage of \"partial dependence.\"\n\n2.\tLimitations of Theoretical Assumptions: Based on strong assumptions, the paper models MARL as segmented turn-based, sequence-level single scalar rewards, and text concatenation. This is significantly different from real-world complex agent interactions, such as parallel collaboration and multi-round debate."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920228326,"tcdate":1761659561567,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8299/Reviewer_oXhZ"],"signatures":["ICLR.cc/2026/Conference/Submission8299/Reviewer_oXhZ"],"forum":"RuP0GDhjWb","number":2,"license":"CC BY 4.0","cdate":1761659561567,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8299/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920228326,"domain":"ICLR.cc/2026/Conference","replyto":"RuP0GDhjWb","id":"AlDQTV9qQ7","forumContent":{"TLDR":{"value":"This paper compares the efficiency of llm-based multi-agent system and single agent system"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["learning efficiency","llm-based agentic system"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Reinforcement Learning (RL) has emerged as a crucial method for training or fine-tuning large language models (LLMs), enabling adaptive, task-specific optimizations through interactive feedback. Multi-Agent Reinforcement Learning (MARL), in particular, offers a promising avenue by decomposing complex tasks into specialized subtasks managed by distinct interacting agents, potentially enhancing the ability and efficiency of LLM systems. However, theoretical insights regarding when and why MARL outperforms Single-Agent RL (SARL) remain limited, creating uncertainty in selecting the appropriate RL framework. In this paper, we address this critical gap by rigorously analyzing the comparative sample efficiency of MARL and SARL within the context of LLM.\n% \\cwu{training/fine-tuning? js: applicable to both}. \nLeveraging the Probably Approximately Correct (PAC) framework, we formally define SARL and MARL setups for LLMs, derive explicit sample complexity bounds, and systematically characterize how task decomposition and alignment influence learning efficiency. Our results demonstrate that MARL improves sample complexity when tasks naturally decompose into independent subtasks, whereas dependent subtasks diminish MARL's comparative advantage. Additionally, we introduce and analyze the concept of task alignment, quantifying the trade-offs when enforcing independent task decomposition despite potential misalignments. These theoretical insights clarify empirical inconsistencies and provide practical criteria for deploying MARL strategies effectively in complex LLM scenarios."},"_bibtex":{"value":"@misc{\nsu2026when,\ntitle={When Do Multi-Agent Systems Outperform? Analysing the Learning Efficiency of Agentic Systems},\nauthor={Junwei Su and Chuan Wu},\nyear={2026},\nurl={https://openreview.net/forum?id=RuP0GDhjWb}\n}"},"title":{"value":"When Do Multi-Agent Systems Outperform? Analysing the Learning Efficiency of Agentic Systems"},"pdf":{"value":"/pdf/6c55fbcdd7d7871487169f135fad8c8e7780bdf3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"su|when_do_multiagent_systems_outperform_analysing_the_learning_efficiency_of_agentic_systems"},"authorids":{"value":["~Junwei_Su1","~Chuan_Wu1"]},"authors":{"value":["Junwei Su","Chuan Wu"]}},"version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2312.02235v2"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"zhang|genem_physicsinformed_generative_cryoelectron_microscopy"},"authorids":{"value":["~Jiakai_Zhang3","https://dblp.org/search/pid/api?q=author:Qihe_Chen:","https://dblp.org/search/pid/api?q=author:Yan_Zeng:","","https://dblp.org/search/pid/api?q=author:Xuming_He_0001:","https://dblp.org/search/pid/api?q=author:Zhijie_Liu:","~Jingyi_Yu5"]},"html":{"value":"https://doi.org/10.48550/arXiv.2312.02235"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2312-02235,\n  publtype={informal},\n  author={Jiakai Zhang and Qihe Chen and Yan Zeng and Wenyuan Gao and Xuming He and Zhijie Liu and Jingyi Yu},\n  title={GenEM: Physics-Informed Generative Cryo-Electron Microscopy},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2312.02235},\n  url={https://doi.org/10.48550/arXiv.2312.02235}\n}\n"},"abstract":{"value":"In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of proteins, such as the SARS- COV-2 spike protein. To achieve high-resolution reconstruction, a comprehensive data processing pipeline has been adopted. However, its performance is still limited as it lacks high-quality annotated datasets for training. To address this, we introduce physics-informed generative cryo-electron microscopy (CryoGEM), which for the first time integrates physics-based cryo-EM simulation with a generative unpaired noise translation to generate physically correct synthetic cryo-EM datasets with realistic noises. Initially, CryoGEM simulates the cryo-EM imaging process based on a virtual specimen. To generate realistic noises, we leverage an unpaired noise translation via contrastive learning with a novel mask-guided sampling scheme. Extensive experiments show that CryoGEM is capable of generating authentic cryo-EM images. The generated dataset can used as training data for particle picking and pose estimation models, eventually improving the reconstruction resolution."},"title":{"value":"GenEM: Physics-Informed Generative Cryo-Electron Microscopy"},"authors":{"value":["Jiakai Zhang","Qihe Chen","Yan Zeng","Wenyuan Gao","Xuming He","Zhijie Liu","Jingyi Yu"]}},"tmdate":1731510119137,"pdate":1672531200000,"tcdate":1730810837425,"writers":["~"],"signatures":["~Jiakai_Zhang3"],"forum":"GAYTsuMLMe","license":"CC BY-SA 4.0","number":170809,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731510119137,"domain":"DBLP.org","id":"GAYTsuMLMe","version":2},{"content":{"summary":{"value":"The paper presents PhysUniBench, a large-scale benchmark for evaluating undergraduate-level physics reasoning in multimodal large language models (MLLMs). It includes 3,304 human-verified problems with diagrams, five difficulty levels, and both open-ended and multiple-choice formats. Experiments show that current MLLMs perform poorly, underscoring the challenges of multimodal physical reasoning."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. In Table 3, for Open-ended Questions, GPT-5 significantly outperforms all other MLLMs on RE and QM, while the rest of the models exhibit nearly identical results (2.1% in RE, 0% in QM except GPT-4o with 6.2%). This trend deviates sharply from the patterns observed in other evaluation dimensions, suggesting potential systematic issues or bottlenecks in data construction, evaluation protocols, or metric design for this sub-task.\n\nMoreover, the reported 2.1% accuracy is mathematically inconsistent with a test set of 80 samples, raising concerns about statistical validity:\n\n    - If 1 sample were correct, the implied dataset size would be 1 / 0.021 ≈ 47.6, which contradicts the stated size of 80;\n\n    - If 2 samples were correct, the implied size becomes 2 / 0.021 ≈ 95.2, which exceeds 80.\nWithout clarification on additional scoring mechanisms (*e.g.*, partial credit, weighted averaging, repeated sampling, or a different denominator), this value lacks a consistent explanation. The authors are encouraged to provide the exact accuracy computation method or the raw number of correctly predicted samples.\n\n2. I suggest including **difficulty × sub-disciplines evaluation statistics**. In my view, analyzing difficulty stratification within each sub-discipline is essential to better characterize the model’s capability boundaries and failure modes, and would substantially strengthen the empirical findings."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- First comprehensive multimodal benchmark for undergraduate-level physics reasoning.\n\n- Rigorous curation process with difficulty calibration and quality control.\n\n- Extensive evaluation across sub-disciplines provides clear diagnostic insights.\n\n- Well-written and potentially impactful for advancing AI-for-Science."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### 1. Lack of actionable guidance for model improvement \nWhile the benchmark offers valuable insights into the limitations of current MLLMs, the paper does not sufficiently explore how architectural or training modifications (*e.g.*, physics-informed modules, structured reasoning layers, or symbolic integration) could help enhance physical perception and reasoning.\n\n### 2. Limited analysis of reasoning failures\nThe error analysis mainly reports accuracy drops across sub-disciplines and difficulty levels, but lacks qualitative or process-level diagnostics showing why models fail (*e.g.*, confusion between symbolic manipulation vs. conceptual understanding). Such analysis could better guide future model design.\n\n### 3. Benchmark orientation without methodological contribution\nPhysUniBench is primarily an evaluation resource, and while comprehensive, it lacks a corresponding methodological or modeling innovation that demonstrates how benchmark insights could translate into improved multimodal reasoning systems."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359358130,"tcdate":1761998124334,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2029/Reviewer_C6kN"],"signatures":["ICLR.cc/2026/Conference/Submission2029/Reviewer_C6kN"],"forum":"TqgPgCvBnF","number":4,"license":"CC BY 4.0","cdate":1761998124334,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2029/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359358130,"domain":"ICLR.cc/2026/Conference","replyto":"TqgPgCvBnF","id":"N9acJRbRqF","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics","reasoning","benchmark","large language model","multi-modal large language model"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Physics problem-solving is a challenging domain for large AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Existing evaluations fail to capture the full breadth and complexity of undergraduate physics, whereas this level provides a rigorous yet standardized testbed for pedagogically relevant assessment of multi-step physical reasoning. To this end, we present PhysUniBench, a large-scale multimodal benchmark designed to evaluate and improve the reasoning capabilities of multimodal large language models (MLLMs) specifically on undergraduate-level physics problems. PhysUniBench consists of 3,304 physics questions spanning 8 major sub-disciplines of physics, each accompanied by one visual diagrams. The benchmark includes both open-ended and multiple-choice questions, systematically curated and difficulty-rated through an iterative model-in-the-loop process. The benchmark's construction involved a rigorous multi-stage process, including multiple roll-outs, expert-level evaluation, automated filtering of easily solved problems, and a nuanced difficulty grading system with five levels. Through extensive experiments, we observe that current state-of-the-art models encounter substantial challenges in physics reasoning. For example, GPT-5 achieves only about 53.7% accuracy in the proposed PhysUniBench. These results highlight that current MLLMs struggle with advanced physics reasoning, especially on multi-step problems and those requiring precise diagram interpretation. By providing a broad and rigorous assessment tool, PhysUniBench aims to drive progress in AI for Science, encouraging the development of models with stronger physical reasoning, problem-solving skills, and multimodal understanding. The benchmark and evaluation scripts are available at https://anonymous.4open.science/r/PhysUniBenchmark-5784."},"_bibtex":{"value":"@misc{\nwang2026physunibench,\ntitle={PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level},\nauthor={Lintao Wang and Encheng Su and Jiaqi Liu and Pengze Li and Jiabei Xiao and Wenlong Zhang and Xi Chen and Yuan Meng and LEI BAI and Wanli Ouyang and SHIXIANG TANG and Aoran Wang and Xinzhu Ma},\nyear={2026},\nurl={https://openreview.net/forum?id=TqgPgCvBnF}\n}"},"title":{"value":"PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level"},"pdf":{"value":"/pdf/006ead177bfaf2dd134dbde4641e5de03d19d8e8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|physunibench_a_multimodal_physics_reasoning_benchmark_at_undergraduate_level"},"authorids":{"value":["~Lintao_Wang1","~Encheng_Su1","~Jiaqi_Liu7","~Pengze_Li3","~Jiabei_Xiao1","~Wenlong_Zhang3","~Xi_Chen20","~Yuan_Meng2","~LEI_BAI1","~Wanli_Ouyang1","~SHIXIANG_TANG1","~Aoran_Wang1","~Xinzhu_Ma1"]},"authors":{"value":["Lintao Wang","Encheng Su","Jiaqi Liu","Pengze Li","Jiabei Xiao","Wenlong Zhang","Xi Chen","Yuan Meng","LEI BAI","Wanli Ouyang","SHIXIANG TANG","Aoran Wang","Xinzhu Ma"]}},"version":2},{"content":{"venue":{"value":"Medical Imaging: Image Processing 2005"},"pdf":{"value":"https://www.spiedigitallibrary.org/conference-proceedings-of-spie/5747/0000/Physics-based-deformable-organisms-for-medical-image-analysis/10.1117/12.594856.pdf"},"venueid":{"value":"dblp.org/conf/MIIP/2005"},"paperhash":{"value":"hamarneh|physicsbased_deformable_organisms_for_medical_image_analysis"},"authorids":{"value":["~Ghassan_Hamarneh1","https://dblp.org/search/pid/api?q=author:Chris_McIntosh:"]},"html":{"value":"https://doi.org/10.1117/12.594856"},"_bibtex":{"value":"@inproceedings{DBLP:conf/miip/HamarnehM05,\n  author={Ghassan Hamarneh and Chris McIntosh},\n  title={Physics-based deformable organisms for medical image analysis},\n  year={2005},\n  cdate={1104537600000},\n  url={https://doi.org/10.1117/12.594856},\n  booktitle={Medical Imaging: Image Processing},\n  crossref={conf/miip/2005}\n}\n"},"abstract":{"value":"Previously, \"Deformable organisms\" were introduced as a novel paradigm for medical image analysis that uses artificial life modelling concepts. Deformable organisms were designed to complement the classical bottom-up deformable models methodologies (geometrical and physical layers), with top-down intelligent deformation control mechanisms (behavioral and cognitive layers). However, a true physical layer was absent and in order to complete medical image segmentation tasks, deformable organisms relied on pure geometry-based shape deformations guided by sensory data, prior structural knowledge, and expert-generated schedules of behaviors. In this paper we introduce the use of physics-based shape deformations within the deformable organisms framework yielding additional robustness by allowing intuitive real-time user guidance and interaction when necessary. We present the results of applying our physics-based deformable organisms, with an underlying dynamic spring-mass mesh model, to segmenting and labelling the corpus callosum in 2D midsagittal magnetic resonance images."},"title":{"value":"Physics-based deformable organisms for medical image analysis"},"authors":{"value":["Ghassan Hamarneh","Chris McIntosh"]}},"tmdate":1727454658651,"pdate":1104537600000,"tcdate":1727453021235,"writers":["~"],"signatures":["~Ghassan_Hamarneh1"],"forum":"3rA0RORYLo","license":"CC BY-SA 4.0","number":86422,"cdate":1104537600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727454658651,"domain":"DBLP.org","id":"3rA0RORYLo","version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text instructions are often too abstract to capture physical nuances, while target images are frequently infeasible to specify for dynamic tasks. To address this, we introduce Goal Force, a novel framework that allows users to define goals via explicit force vectors and intermediate dynamics, mirroring how humans conceptualize physical tasks. We train a video generation model on a curated dataset of synthetic causal primitives—such as elastic collisions and falling dominos—teaching it to propagate forces through time and space. Despite being trained on simple physics data, our model exhibits remarkable zero-shot generalization to complex, real-world scenarios, including tool manipulation and multi-object causal chains. Our results suggest that by grounding video generation in fundamental physical interactions, models can emerge as implicit neural physics simulators, enabling precise, physics-aware planning without reliance on external engines."},"_bibtex":{"value":"@inproceedings{\ngillman2026goal,\ntitle={Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals},\nauthor={Nate Gillman and Yinghua Zhou and Zitian Tang and Evan Luo and Arjan Chakravarthy and Daksh Aggarwal and Michael Freeman and Chen Sun},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=tjKxJMjOwL}\n}"},"title":{"value":"Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Gillman_Goal_Force_Teaching_Video_Models_To_Accomplish_Physics-Conditioned_Goals_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"gillman|goal_force_teaching_video_models_to_accomplish_physicsconditioned_goals"},"authorids":{"value":["~Nate_Gillman1","~Yinghua_Zhou2","~Zitian_Tang1","~Evan_Luo2","~Arjan_Chakravarthy1","~Daksh_Aggarwal1","~Michael_Freeman1","~Chen_Sun1"]},"authors":{"value":["Nate Gillman","Yinghua Zhou","Zitian Tang","Evan Luo","Arjan Chakravarthy","Daksh Aggarwal","Michael Freeman","Chen Sun"]}},"tmdate":1789656588123,"pdate":1789656437897,"tcdate":1765223359001,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission44111/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission44111/Authors"],"forum":"tjKxJMjOwL","license":"CC BY 4.0","number":44111,"cdate":1765223359001,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission44111/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/Submission44111/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656588123,"odate":1789656437897,"domain":"thecvf.com/CVPR/2026/Conference","id":"tjKxJMjOwL","version":2},{"content":{"summary":{"value":"The paper develops theoretical results for the convergence rates of complex and real valued neural networks on complex-valued problems. It develops analytical models for the training regimes of complex valued NNs. "},"soundness":{"value":"2 fair"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"- In Section 4.3, could you clarify how the outputs of the RVNN are matched to the outputs of the complex valued target function?\n- In Section 5, does the trainable zReLU make the network effectively deeper? Does the convergence rate change if psi is not trainable?\n- Define the schema of $L_{rr}, L_{cr},$ etc., and remind the reader every time you use them. By the time I got to Lemma 4 and Lemma 5, I did not know which result was which.\n- $t4 was never defined.\n- Define rvnn and cvnn in the abstract\n- Add equation numbers.\n"},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"strengths":{"value":"The results would be useful for those attempting to apply complex-valued networks to problem settings where they are not normally used. An intriguing result is that complex valued neurons have slower convergence rates for the same problems as real valued neurons. The theorems are thorough, and the diagrams provide excellent illustration to the proofs. They also provide a proof that real valued networks cannot learn complex valued functions, which is not surprising, but nice to have proved.\n"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"A fundamental flaw of the paper is that it does not have empirical verification of the theoretical results. The analysis makes many assumptions about different training regimes, but it is not obvious if complex valued neural networks actually behave like this. The complex-valued models in the theorems should be implemented and the convergence losses should be plotted alongside the theoretical predictions. Without supporting empirical evidence, the assumptions of the proof seem too strong. Fortunately, the problem settings should be simple to code up in any package, and the results would be easy to include in the author's presentation, given Figures 1 and 3.\n\nPage 4, line 156: “whereas a complex-valued neuron may only activate a small part as controlled by the parameter psi” This line makes it sound like the authors conclude that the primary benefit is the extra parameter in zReLU, not the complex arithmetic. The use of activation functions with trainable parameters is also a major distinction between the CVNN and the RVNN. Maybe all of the identified differences can come from a network with real weights, but an activation function with a learnable parameter? zReLU can be encoded in real-valued weights, if you interpret two real weights and use a learnable ReLU. Two variations are required to ablate this confounder:\n1. complex weights with a zReLU with a fixed psi (Would this be easier to study theoretically than the existing proofs? It would also be easy to demonstrate empirically.)\n2. real weights with a pseudo-zReLU with a learnable psi (Possibly hard to prove, but easy to empirically study.)\nI think considering case (1) is necessary to support the claim made by the paper that complex weights are the key differentiator."},"limitations":{"value":"A limitations section is missing from the paper.\n"}},"nonreaders":[],"tmdate":1702410865407,"tcdate":1688754360303,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission2826/Reviewer_Jsv8"],"signatures":["NeurIPS.cc/2023/Conference/Submission2826/Reviewer_Jsv8"],"forum":"qA0uHmaVKk","number":4,"license":"CC BY 4.0","cdate":1688754360303,"mdate":1702410865407,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission2826/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"qA0uHmaVKk","id":"a3NnyqG4yu","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Complex-valued Neural Networks; Learning Neurons; Real-valued Neural Networks; Convergence Rate"]},"supplementary_material":{"value":"/attachment/e0e8195000a1f4fde79c9bc8cf119b08fbfc8b4f.pdf"},"_bibtex":{"value":"@inproceedings{\nwu2023complexvalued,\ntitle={Complex-valued Neurons Can Learn More but Slower than Real-valued Neurons via Gradient Descent},\nauthor={Jin-Hui Wu and Shao-Qun Zhang and Yuan Jiang and Zhi-Hua Zhou},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=qA0uHmaVKk}\n}"},"title":{"value":"Complex-valued Neurons Can Learn More but Slower than Real-valued Neurons via Gradient Descent"},"paperhash":{"value":"wu|complexvalued_neurons_can_learn_more_but_slower_than_realvalued_neurons_via_gradient_descent"},"abstract":{"value":"Complex-valued neural networks potentially possess better representations and performance than real-valued counterparts when dealing with some complicated tasks such as acoustic analysis, radar image classification, etc. Despite empirical successes, it remains unknown theoretically when and to what extent complex-valued neural networks outperform real-valued ones. We take one step in this direction by comparing the learnability of real-valued neurons and complex-valued neurons via gradient descent. We show that a complex-valued neuron can efficiently learn functions expressed by any one real-valued neuron and any one complex-valued neuron with convergence rate $O(t^{-3})$ and $O(t^{-1})$ where $t$ is the iteration index of gradient descent, respectively, whereas a two-layer real-valued neural network with finite width cannot learn a single non-degenerate complex-valued neuron. We prove that a complex-valued neuron learns a real-valued neuron with rate $\\Omega (t^{-3})$, exponentially slower than the $O(\\mathrm{e}^{- c t})$ rate of learning one real-valued neuron using a real-valued neuron with a constant $c$. We further verify and extend these results via simulation experiments in more general settings."},"pdf":{"value":"/pdf/129de32bfd6b64ae584e3f42314707df638b5755.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Jin-Hui_Wu1","~Shao-Qun_Zhang1","~Yuan_Jiang1","~Zhi-Hua_Zhou2"]},"authors":{"value":["Jin-Hui Wu","Shao-Qun Zhang","Yuan Jiang","Zhi-Hua Zhou"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2405.02041v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"schnell|stabilizing_backpropagation_through_time_to_learn_complex_physics"},"authorids":{"value":["~Patrick_Schnell1","~Nils_Thuerey1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2405.02041"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2405-02041,\n  publtype={informal},\n  author={Patrick Schnell and Nils Thuerey},\n  title={Stabilizing Backpropagation Through Time to Learn Complex Physics},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2405.02041},\n  url={https://doi.org/10.48550/arXiv.2405.02041}\n}\n"},"abstract":{"value":"Of all the vector fields surrounding the minima of recurrent learning setups, the gradient field with its exploding and vanishing updates appears a poor choice for optimization, offering little beyond efficient computability. We seek to improve this suboptimal practice in the context of physics simulations, where backpropagating feedback through many unrolled time steps is considered crucial to acquiring temporally coherent behavior. The alternative vector field we propose follows from two principles: physics simulators, unlike neural networks, have a balanced gradient flow, and certain modifications to the backpropagation pass leave the positions of the original minima unchanged. As any modification of backpropagation decouples forward and backward pass, the rotation-free character of the gradient field is lost. Therefore, we discuss the negative implications of using such a rotational vector field for optimization and how to counteract them. Our final procedure is easily implementable via a sequence of gradient stopping and component-wise comparison operations, which do not negatively affect scalability. Our experiments on three control problems show that especially as we increase the complexity of each task, the unbalanced updates from the gradient can no longer provide the precise control signals necessary while our method still solves the tasks. Our code can be found at https://github.com/tum-pbs/StableBPTT."},"title":{"value":"Stabilizing Backpropagation Through Time to Learn Complex Physics"},"authors":{"value":["Patrick Schnell","Nils Thuerey"]}},"tmdate":1747332578047,"pdate":1704067200000,"tcdate":1727700994160,"writers":["~"],"signatures":["~Nils_Thuerey1"],"forum":"wM54GG2RZ4","license":"CC BY-SA 4.0","number":110841,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747332578047,"domain":"DBLP.org","id":"wM54GG2RZ4","version":2},{"content":{"venue":{"value":"ICLR 2024 poster"},"TLDR":{"value":"We show how modifying backpropagation and counteracting rotation in the resulting vector fields can improve learning in physical simulations."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["backpropagation through time","unrolling","differentiable physics","differentiable simulators","optimization"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Of all the vector fields surrounding the minima of recurrent learning setups, the gradient field with its exploding and vanishing updates appears a poor choice for optimization, offering little beyond efficient computability. We seek to improve this suboptimal practice in the context of physics simulations, where backpropagating feedback through many unrolled time steps is considered crucial to acquiring temporally coherent behavior. The alternative vector field we propose follows from two principles: physics simulators, unlike neural networks, have a balanced gradient flow and certain modifications to the backpropagation pass leave the positions of the original minima unchanged. As any modification of backpropagation decouples forward and backward pass, the rotation-free character of the gradient field is lost. Therefore, we discuss the negative implications of using such a rotational vector field for optimization and how to counteract them. Our final procedure is easily implementable via a sequence of gradient stopping and component-wise comparison operations, which do not negatively affect scalability. Our experiments on three control problems show that especially as we increase the complexity of each task, the unbalanced updates from the gradient can no longer provide the precise control signals necessary while our method still solves the tasks. Our code can be found at https://github.com/tum-pbs/StableBPTT."},"_bibtex":{"value":"@inproceedings{\nschnell2024stabilizing,\ntitle={Stabilizing Backpropagation Through Time to Learn Complex Physics},\nauthor={Patrick Schnell and Nils Thuerey},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=bozbTTWcaw}\n}"},"title":{"value":"Stabilizing Backpropagation Through Time to Learn Complex Physics"},"pdf":{"value":"/pdf/dcb8c6f1ff02f187589c32e89ff56702a5aa53d4.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"schnell|stabilizing_backpropagation_through_time_to_learn_complex_physics"},"authorids":{"value":["~Patrick_Schnell1","~Nils_Thuerey1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Patrick Schnell","Nils Thuerey"]}},"tmdate":1709912515065,"pdate":1705410854375,"tcdate":1695225835653,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2701/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission2701/Authors"],"forum":"bozbTTWcaw","number":2701,"cdate":1695225835653,"mdate":1709912515065,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/-/Submission","ICLR.cc/2024/Conference/-/Post_Submission","ICLR.cc/2024/Conference/Submission2701/-/Revision","ICLR.cc/2024/Conference/Submission2701/-/Rebuttal_Revision","ICLR.cc/2024/Conference/-/Edit","ICLR.cc/2024/Conference/Submission2701/-/Camera_Ready_Revision"],"odate":1697213872796,"domain":"ICLR.cc/2024/Conference","id":"bozbTTWcaw","version":2},{"content":{"summary":{"value":"This paper presents HA3C, a deep reinforcement learning algorithm designed to tackle the critical issue of sample inefficiency in Markov Decision Processes  with complex transition dynamics. The central premise is that augmenting the current state with historical information can simplify the learning of these complex causal relationships. To achieve this without suffering from high-dimensional inputs, HA3C employs a Convolutional Neural Network to compress the state trajectory into a concise, low-dimensional representation. This learned historical context is then integrated into a powerful actor-critic framework, leading to significant improvements in both sample efficiency and final performance across a suite of challenging continuous control benchmarks."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. What mechanisms within HA3C are designed to prevent the model from overfitting to spurious correlations within historical trajectories, and can you formalize the conditions under which history might become detrimental to learning?\n2. Instead of a fixed k, have the authors considered dynamic architectures, such as using an attention mechanism, that would allow the agent to adaptively learn the relevant temporal dependencies for a given task or state?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper is well-organized. \n2. The paper's central contribution rests on its compelling and counter-intuitive argument for leveraging history in MDPs. This core premise is rigorously validated through comprehensive empirical evidence.\n3. Experimental results show promising results on MuJoCo and DMC tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. On several MuJoCo tasks, such as HalfCheetah, the performance improvement over TD7, while positive, may not be statistically significant when considering the reported standard deviations. The authors should provide a more detailed analysis of the statistical significance of their results.\n2. The paper's core premise that history simplifies future prediction is not critically examined for its limitations. The authors' own finding that long histories introduce noise (for k=24) suggests potential failure modes, yet there is no broader discussion of which environmental properties might render historical information detrimental.\n3. The theoretical justification for why the proposed method reduces sample complexity is underdeveloped. The paper provides an intuitive motivation but lacks a formal analysis (e.g., sample complexity bounds) to explain the mechanism, positioning the contribution primarily as a strong empirical finding rather than a new theoretical framework.\n4. The algorithm's performance is highly sensitive to the historical window size k, a crucial hyperparameter that requires manual, task-specific tuning. The absence of a principled method for selecting k limits the algorithm's practicality and makes it difficult to apply to new environments without extensive tuning."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922989557,"tcdate":1761836375012,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11998/Reviewer_L1Ki"],"signatures":["ICLR.cc/2026/Conference/Submission11998/Reviewer_L1Ki"],"forum":"ZHIAjSDxve","number":2,"license":"CC BY 4.0","cdate":1761836375012,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11998/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922989557,"domain":"ICLR.cc/2026/Conference","replyto":"ZHIAjSDxve","id":"8VmdNzYAEF","forumContent":{"TLDR":{"value":"This paper investigates if augmenting the states with their historical information can simplify the complex causal relationships in MDPs and thus improve the sample efficiency of DRL."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Markov Decision Process","Deep Reinforcement Learning","Historical Augmentation","Sample Efficiency"]},"supplementary_material":{"value":"/attachment/94cc3461652539025b9ae4359cf09b587f84404e.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"Under the Markov assumption of Markov Decision Processes (MDPs), an optimal stationary policy does not need to consider history and is no worse than any non-stationary or history-dependent policy. Therefore, existing Deep Reinforcement Learning (DRL) algorithms usually model sequential decision-making as an MDP and then try to optimize a stationary policy by single-step state transitions. However, such optimization is often faced with sample inefficiency when the causal relationships of state transitions are complex. To address the above problem, this paper investigates if augmenting the states with their historical information can simplify the complex causal relationships in MDPs and thus improve the sample efficiency for DRL. First, we demonstrate that a complex causal relationship of single-step state transitions may be inferred by a simple causal function of the historically augmented states. Then, we propose a convolutional neural network architecture to learn the representation of the current state and its historical trajectory. This representation learning compresses the high-dimensional historical trajectories into a low-dimensional space to extract the simple causal relationships from historical information and avoid the overfitting caused by high-dimensional data. Finally, we formulate Historical Augmentation Aided Actor-Critic (HA3C) algorithm by adding the learned representations to the actor-critic method. The experiment on standard MDP tasks demonstrates that HA3C outperforms current state-of-the-art methods in terms of both sample efficiency and performance."},"_bibtex":{"value":"@misc{\nhuang2025beyond,\ntitle={Beyond Markov Assumption: Improving Sample Efficiency in {MDP}s by Historical Augmentation},\nauthor={Tianyi Huang and Zhiling Cai and Xiaolong Qin and Xin Yuan},\nyear={2025},\nurl={https://openreview.net/forum?id=ZHIAjSDxve}\n}"},"title":{"value":"Beyond Markov Assumption: Improving Sample Efficiency in MDPs by Historical Augmentation"},"pdf":{"value":"/pdf/c0b756fde412a0d07aa85929f2e653d112f57e46.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|beyond_markov_assumption_improving_sample_efficiency_in_mdps_by_historical_augmentation"},"authorids":{"value":["~Tianyi_Huang1","~Zhiling_Cai1","~Xiaolong_Qin1","~Xin_Yuan4"]},"authors":{"value":["Tianyi Huang","Zhiling Cai","Xiaolong Qin","Xin Yuan"]}},"version":2},{"content":{"venue":{"value":"Communications Physics"},"pdf":{"value":"/pdf/549b88bd6151eb3a8165232b5588a9261b262482.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"liu|multiresolution_partial_differential_equations_preserved_learning_framework_for_spatiotemporal_dynamics"},"authorids":{"value":["~Xin-Yang_Liu1","~Min_Zhu1","~Lu_Lu1","~Hao_Sun4","~Jian-Xun_Wang1"]},"html":{"value":"https://www.nature.com/articles/s42005-024-01521-z"},"abstract":{"value":"Traditional data-driven deep learning models often struggle with high training costs, error accumulation, and poor generalizability in complex physical processes. Physics-informed deep learning (PiDL) addresses these challenges by incorporating physical principles into the model. Most PiDL approaches regularize training by embedding governing equations into the loss function, yet this depends heavily on extensive hyperparameter tuning to weigh each loss term. To this end, we propose to leverage physics prior knowledge by “baking” the discretized governing equations into the neural network architecture via the connection between the partial differential equations (PDE) operators and network structures, resulting in a PDE-preserved neural network (PPNN). This method, embedding discretized PDEs through convolutional residual networks in a multi-resolution setting, largely improves the generalizability and long-term prediction accuracy, outperforming conventional black-box models. The effectiveness and merit of the proposed methods have been demonstrated across various spatiotemporal dynamical systems governed by spatiotemporal PDEs, including reaction-diffusion, Burgers’, and Navier-Stokes equations."},"title":{"value":"Multi-resolution partial differential equations preserved learning framework for spatiotemporal dynamics"},"authors":{"value":["Xin-Yang Liu","Min Zhu","Lu Lu","Hao Sun","Jian-Xun Wang"]}},"tmdate":1790622426841,"pdate":1705165202800,"tcdate":1714416115018,"writers":["~Xin-Yang_Liu1","~Min_Zhu1","~Lu_Lu1","~Hao_Sun4","~Jian-Xun_Wang1"],"signatures":["~Min_Zhu1"],"forum":"kfRPFgwM3C","license":"CC BY 4.0","number":26073,"cdate":1714416115018,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1790622426841,"domain":"OpenReview.net/Archive","id":"kfRPFgwM3C","version":2},{"content":{"venue":{"value":"CoRR 2019"},"pdf":{"value":"http://arxiv.org/pdf/1904.09860v1"},"venueid":{"value":"dblp.org/journals/CORR/2019"},"paperhash":{"value":"li|learning_manipulation_under_physics_constraints_with_visual_perception"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Wenbin_Li_0003:","https://dblp.org/search/pid/api?q=author:Ales_Leonardis:","~Jeannette_Bohg1","~Mario_Fritz1"]},"html":{"value":"http://arxiv.org/abs/1904.09860"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-1904-09860,\n  publtype={informal},\n  author={Wenbin Li and Ales Leonardis and Jeannette Bohg and Mario Fritz},\n  title={Learning Manipulation under Physics Constraints with Visual Perception},\n  year={2019},\n  cdate={1546300800000},\n  journal={CoRR},\n  volume={abs/1904.09860},\n  url={http://arxiv.org/abs/1904.09860}\n}\n"},"abstract":{"value":"Understanding physical phenomena is a key competence that enables humans and animals to act and interact under uncertain perception in previously unseen environments containing novel objects and their configurations. In this work, we consider the problem of autonomous block stacking and explore solutions to learning manipulation under physics constraints with visual perception inherent to the task. Inspired by the intuitive physics in humans, we first present an end-to-end learning-based approach to predict stability directly from appearance, contrasting a more traditional model-based approach with explicit 3D representations and physical simulation. We study the model's behavior together with an accompanied human subject test. It is then integrated into a real-world robotic system to guide the placement of a single wood block into the scene without collapsing existing tower structure. To further automate the process of consecutive blocks stacking, we present an alternative approach where the model learns the physics constraint through the interaction with the environment, bypassing the dedicated physics learning as in the former part of this work. In particular, we are interested in the type of tasks that require the agent to reach a given goal state that may be different for every new trial. Thereby we propose a deep reinforcement learning framework that learns policies for stacking tasks which are parametrized by a target structure."},"title":{"value":"Learning Manipulation under Physics Constraints with Visual Perception"},"authors":{"value":["Wenbin Li","Ales Leonardis","Jeannette Bohg","Mario Fritz"]}},"tmdate":1727595950346,"pdate":1546300800000,"tcdate":1717930419736,"writers":["~"],"signatures":["~Jeannette_Bohg1"],"forum":"oucrJJwbDF","license":"CC BY-SA 4.0","number":25543,"cdate":1546300800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1727595950346,"domain":"DBLP.org","id":"oucrJJwbDF","version":2},{"content":{"summary":{"value":"This paper studies how LLM performance scales when training on mixtures of real and synthetic data. It posits a *three‑phase* scaling pattern—Rapid‑Learning, Plateau, and Tail‑Learning—with two breakpoints tied to a truncation of tail knowledge in the synthetic distribution (Figures 2–3). The theory is formalized via **Theorem 1**, a generalization bound that depends on (i) empirical losses on real/synthetic subsets, (ii) distribution discrepancies between train and test, (iii) an NTK‑based term at initialization, and (iv) the real‑data proportion ω and sample size |S| (p. 5). Building on the bound, the authors propose a *retraining‑free* data valuation score (v(S)) (Eq. (8), p. 5) that combines weighted empirical losses, MK‑MMD discrepancies, an NTK term ($ \\hat y^\\top \\Theta_0^{-1} \\hat y / |S|$ ), and a small‑sample correction ( \\max($\\omega,1-\\omega)/|S| $). Experiments on CIFAR‑100/100‑C (image classification), IMDb/FinGPT (sentiment), Natural‑Instructions/Magpie (instruction following), and NuminaMath‑CoT (reasoning) report: (a) visual evidence for the three‑phase pattern under a long‑tail setup (Figures 4–6, p. 7), (b) higher correlations to “ground‑truth” contributor value than DAVINZ, Deviation, LOGRA, TracIn, and TRAK (Figure 7, Table 1, p. 8), (c) lower runtime (Figure 8, p. 8), and (d) stability of relative scores under subsampling (Table 2, p. 9)."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"See weakness above."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Timely focus and clear problem statement.** The paper tackles a pressing question—how to reason about data value and scaling when synthetic data increasingly dominates LLM training. \n2. **Readable presentation of the scaling picture.** The results are demonstrated in a clean and nice ways."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Experimental Setting is inconsistent with the Theory**\n The core assumption is that the synthetic distribution is *head‑only up to a cutoff k* (p. 4), leading to the two breakpoints. However, the “real vs synthetic” choice in the *image* experiment uses CIFAR‑100 vs CIFAR‑100‑C (corruptions) (p. 6). Corruptions degrade images but do **not** emulate head‑only token sampling or nucleus‑truncation; they don’t reduce class‑tail support in the sense assumed by the theory. This gap makes it hard to view Figures 4–5 (p. 7) as a direct validation of the truncation‑driven breakpoints. A more faithful test would generate synthetic **from a model trained on the real set** (e.g., top‑p/temperature sampling) and control tail truncation explicitly, then mix real and generated data as in the theory.) \n2. **Three‑phase evidence needs stronger quantification.**\n The “Plateau” region in Figure 5 (p. 7) is not convincingly flat; visually, the loss keeps trending without a clear slope break. Please quantify both breakpoints (e.g., via change‑point detection or piecewise‑power‑law fits with confidence intervals), report slopes and uncertainty across seeds, and show head/tail decompositions with error bars. \n3. **Bound terms need intuition.**\n Theorem 1 introduces two important terms: the initialization‑NTK term ( $\\hat y^\\top \\Theta_0^{-1}\\hat y / |S|$ ) and the finite‑sample term ($2\\max(\\omega,1-\\omega)\\log(8/\\delta)/|S|$) (p. 5). The paper should add intuitive explanations for how each term emerges. \n4. **Computational complexity.** Computing ( \\Theta_0^{-1} ) is (O(n^3)) naively; at LLM scale this is large.  This may be the reason why text experiments are only performed on LLMs smaller than 7B, which largely limits the practical value of the method.\n5. **Assumption setting needs empirical grounding.**\n The model assumes ($p_i \\propto i^{-\\varepsilon}$) for knowledge frequency. The paper should provide empirical evidence (e.g., token/knowledge histograms) that downstream tasks indeed follow the specified forms or offer more insights on this assumption.\n6. **Fairness and statistical power of comparisons.**\n Gradient‑based baselines are restricted to 1% of training data (p. 6), but it’s unclear whether your own method’s NTK/MMD computations were likewise subsampled or ran on full data. Please match compute budgets, report absolute FLOPs/time/memory, and repeat the comparison under equalized budgets.\n\n*I lean toward rejecting this paper, as it still requires refinement. However, I hope the authors do not take this as a harsh judgment—the work has potential, and with further effort, it could evolve into a solid and interesting contribution.*"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923938103,"tcdate":1761988994570,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13262/Reviewer_fFvw"],"signatures":["ICLR.cc/2026/Conference/Submission13262/Reviewer_fFvw"],"forum":"3iXyRG2nzT","number":3,"license":"CC BY 4.0","cdate":1761988994570,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13262/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923938103,"domain":"ICLR.cc/2026/Conference","replyto":"3iXyRG2nzT","id":"LhcmUJm2dz","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Data Valuation","LLM","Scaling Dynamics"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"The rapid progress of large language models (LLMs) is fueled by the growing reliance on datasets that blend real and synthetic data. While synthetic data offers scalability and cost-efficiency, it often introduces systematic distributional discrepancies, particularly underrepresenting long-tail knowledge due to truncation effects from data generation mechanisms like top-$p$ sampling, temperature scaling, and finite sampling. These discrepancies pose fundamental challenges in characterizing and evaluating the utility of mixed real-synthetic datasets. In this paper, we identify a three-phase scaling behavior characterized by two breakpoints that reflect transitions in model behavior across learning head and tail knowledge. We further derive an LLM generalization bound designed for real and synthetic mixtures, revealing several key factors that govern their generalization performance. Building on our theoretical findings, we propose an effective yet efficient data valuation method that scales to large-scale datasets. Comprehensive experiments across four tasks, including image classification, sentiment classification, instruction following, and complex reasoning, demonstrate that our method surpasses state-of-the-art baselines in data valuation with significantly low computational cost."},"_bibtex":{"value":"@misc{\nwang2025data,\ntitle={Data Value in the Age of Scaling: Understanding {LLM} Scaling Dynamics Under Real{\\textendash}Synthetic Data Mixtures},\nauthor={Haohui Wang and Jingyuan Qi and Jianpeng Chen and Jun Wu and Lifu Huang and Lecheng Zheng and Kevin Choi and Balaji Veeramani and Edward Bowen and Alison Hu and Tyler Cody and Dawei Zhou},\nyear={2025},\nurl={https://openreview.net/forum?id=3iXyRG2nzT}\n}"},"title":{"value":"Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real–Synthetic Data Mixtures"},"pdf":{"value":"/pdf/5aa1e4438f7169afa72c9a2942a6b8153822c667.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|data_value_in_the_age_of_scaling_understanding_llm_scaling_dynamics_under_realsynthetic_data_mixtures"},"authorids":{"value":["~Haohui_Wang1","~Jingyuan_Qi1","~Jianpeng_Chen1","~Jun_Wu3","~Lifu_Huang1","~Lecheng_Zheng1","~Kevin_Choi1","~Balaji_Veeramani1","~Edward_Bowen1","~Alison_Hu1","~Tyler_Cody1","~Dawei_Zhou1"]},"authors":{"value":["Haohui Wang","Jingyuan Qi","Jianpeng Chen","Jun Wu","Lifu Huang","Lecheng Zheng","Kevin Choi","Balaji Veeramani","Edward Bowen","Alison Hu","Tyler Cody","Dawei Zhou"]}},"version":2},{"content":{"TLDR":{"value":"A graph-augmented chemical language model that closes finite polymer chains with architecture-matched boundary conditions, giving repetition-invariant, transferable property prediction and design across random, alternating, and block polymers."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["polymer property prediction","chemical language model","graph neural networks","physics-guided machine learning","molecular representation learning"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Copolymers are macromolecules assembled from repeating structural units, arranged in random, alternating, or block architectures. Predicting and designing copolymer properties from molecular structure is of broad interest across energy storage, construction, medicine, and aerospace, yet remains fundamentally challenging because (1) there is no effective parameterization for the structural units or for the complex architectures connecting them, and (2) experimental measurements for any particular application are typically scarce and unbalanced. To resolve these challenges, we propose PolyBC, a physics-guided, graph-augmented chemical language model. In particular, we introduce a graph neural network over structural units at two levels of granularity — repeat units and sub-monomer fragments — whose nodes are encoded by Poly4mer, a chemical language model that supplies transferable latent embeddings of homopolymer segments. At either resolution, the GNN reasons over the polymer topology and closes the finite chain with architecture-matched boundary conditions: periodic for alternating copolymers, Neumann for block copolymers, and mean-field for random copolymers. This construction imposes a theoretically guaranteed, consistent bulk physics: training on one copolymer type and boundary condition generalizes to the others, and predictions are invariant to how many repeats are drawn. On both synthetic and experimental benchmarks, our model is repetition-invariant, data-efficient, and transferable across copolymer types at both granularities, outperforming state-of-the-art GNN and LLM baselines on prediction and design tasks."},"_bibtex":{"value":"@inproceedings{\nanonymous2026polybc,\ntitle={Poly{BC}: A Graph-Augmented Chemical Language Model for Complex Copolymer Learning},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6kjn4CTJHP},\nnote={under review}\n}"},"title":{"value":"PolyBC: A Graph-Augmented Chemical Language Model for Complex Copolymer Learning"},"pdf":{"value":"/pdf/66bb55f8f9d7b9ae844c7b6653c9ef415b20990d.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791234139918,"tcdate":1789763915631,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission48593/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission48593/Authors"],"forum":"6kjn4CTJHP","license":"CC BY 4.0","number":48593,"cdate":1789763915631,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission48593/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791234139918,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"6kjn4CTJHP","version":2},{"content":{"summary":{"value":"This paper proposes MagniLearning, an adaptive framework for improving generalization in neural PDE solvers, which integrates data-driven and knowledge-based supervision in scientific machine learning. The method dynamically adjusts the importance of different data subsets, spatial regions, temporal segments, and knowledge terms through three mechanisms: Leave-One-Region-Out (LORO), Leave-One-Time-Out (LOTO), and Leave-One-Knowledge-Out (LOKO). These adaptive weights prioritize underrepresented or informative subsets, theoretically reducing generalization error. An adaptive scheduling strategy further shifts training emphasis from data fitting to physics-based regularization over time."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. How frequently is the adaptive graph recomputed during training and inference, and how does this affect computational cost and stability?\n\n2. Can the authors provide theoretical or empirical evidence that the adaptive graph construction leads to better operator generalization rather than simply overfitting to localized regions?\n\n3. How does the proposed temporal attention mechanism compare to existing time-stepping adaptive schemes, such as variable Δt integrators or latent ODE-based operator models?\n\n4. Would the adaptive mechanism remain effective when applied to 3D turbulent flows or other complex systems with significantly higher resolution requirements?\n\n5. How sensitive is the performance to the hyperparameters controlling graph update frequency, attention temperature, or neighbor sampling?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The authors provide rigorous theoretical analysis that attempt to bound generalization risk under the proposed adaptive weighting scheme.\n\n2. The framework explicitly integrates both label-based and knowledge-based supervision and proposes a dynamic schedule to shift emphasis from labels to knowledge during training, which is a useful design for problems where labeled data are limited but physics knowledge exists."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The experimental scope is too limited and narrow: evaluations are limited to four beam problems with relatively simple physics and low-dimensional PDEs, which weakens claims that MagniLearning generalizes broadly to complex or high-dimensional PDEs (e.g., Navier-Stokes turbulence, multiphysics, or 3D domains).\n\n2. The baseline set is too limited. Comparing only to PINN and DKNN omits many relevant, modern baselines (e.g., domain-decomposed PINNs, XPINNs, adaptive curriculum or point-weighting schemes, operator learning methods, and other recent PINN variants). This makes it difficult to judge how much of the observed improvement is specific to MagniLearning versus attributable to baseline weaknesses.\n\n3. The computational and practical feasibility of the core weighting mechanisms is insufficiently addressed. LORO/LOTO/LOKO as described require computing model variants with regions/times/knowledge held out (which can be costly); the manuscript lacks concrete approximations, influence-function alternatives, or runtime measurements to show how this scales in practice.\n\n4. The claimed theoretical guarantees depend on strong, opaque technical conditions (many asymptotic/overparameterization bounds with complex dependencies). The results are difficult to interpret in realistic regimes, and the paper does not provide guidance on whether practical network sizes/ϕ choices satisfy these assumptions.\n\n5. The presentation and exposition are often dense and stylistically choppy. Notation is heavy, several definitions are introduced with minimal intuition, and transitions between theory, algorithm, and experiments are abrupt. This makes it harder for readers to connect the theoretical claims to empirical practice.\n\n6. The ML contribution is insufficiently emphasized relative to the application contribution. Much of the novelty lies in the choice of adaptive weighting rather than a fundamentally new learning algorithm. The manuscript needs to better articulate the general ML insights that would interest the ICLR community.\n\n7. Key ablations are missing. The paper does not isolate the relative contribution of LORO vs LOTO vs LOKO, nor does it quantify sensitivity to weight-scaling hyperparameters (β, λ, τ, κ) or to the ε-net region radius ϕ, which limits understanding of robustness and reproducibility.\n\n8. Training and implementation details are sparse and seem lightweight for a strong empirical claim: the experiments use a very small MLP (two layers × 50 units) and a single optimizer setting (Adam, lr=1e-2, 2000 epochs) without reporting hyperparameter sweeps, variance across seeds, training time, or cost of adaptive weighting. This undercuts claims about scalability and reliability.\n\n9. The manuscript claims broad applicability and improved generalization but does not demonstrate transfer experiments (e.g., training on one geometry/parameter regime and testing on another) nor does it show performance on noisy or sparse label regimes in a way that convincingly separates the benefit of adaptive weighting from simple reweighting heuristics."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927927591,"tcdate":1762480929049,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18173/Reviewer_SMLC"],"signatures":["ICLR.cc/2026/Conference/Submission18173/Reviewer_SMLC"],"forum":"rE1ggCutRz","number":4,"license":"CC BY 4.0","cdate":1762480929049,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18173/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927927591,"domain":"ICLR.cc/2026/Conference","replyto":"rE1ggCutRz","id":"PxBxFBITl3","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-Informed Neural Networks (PINNs)","Neural PDE solvers","Adaptive weighting","Generalization","Convergence acceleration"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Integrating domain knowledge into neural networks has advanced the development of Physics-Informed Neural Networks (PINNs), enabling solutions to partial differential equations across diverse applications. To enhance generalization in neural PDE solvers, we propose MagniLearning, an adaptive weighting strategy that dynamically adjusts the importance of spatial regions, knowledge components, and temporal segments during training. Our approach evaluates the impact of omitting each region or time block on model performance and assigns higher weights to the most influential data. This adaptive scheme accelerates convergence in neural PDE solvers by emphasizing the most informative regions and time segments, while enhancing robustness to noise and underrepresented physics. We formalize the method using an effective risk function that incorporates region- and time-dependent weights, and we provide theoretical guarantees for controlling the generalization error. Numerical experiments demonstrate that MagniLearning significantly improves both stability and accuracy."},"_bibtex":{"value":"@misc{\nfogh2026adaptive,\ntitle={Adaptive Spatial-Temporal Generalization for Physics-Informed Neural {PDE} Solvers},\nauthor={Fatemeh Fogh and Yufei Tang and Mohsen Ahmadi and Xingquan Zhu},\nyear={2026},\nurl={https://openreview.net/forum?id=rE1ggCutRz}\n}"},"title":{"value":"Adaptive Spatial-Temporal Generalization for Physics-Informed Neural PDE Solvers"},"pdf":{"value":"/pdf/9a603fabc6c0a0cbe2ae9c7bb60d57f96632be75.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"fogh|adaptive_spatialtemporal_generalization_for_physicsinformed_neural_pde_solvers"},"authorids":{"value":["~Fatemeh_Fogh1","~Yufei_Tang1","~Mohsen_Ahmadi1","~Xingquan_Zhu1"]},"authors":{"value":["Fatemeh Fogh","Yufei Tang","Mohsen Ahmadi","Xingquan Zhu"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We present PRISM-Physics, a benchmark and a process-level evaluation framework that encodes physics solutions as DAGs and employs rule-based symbolic equivalence checking for reliable, fine-grained scoring."},"keywords":{"value":["Physics Reasoning","Process-Level Evaluation","Symbolic Equivalence","Scientific Problem Solving"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Benchmarks for competition-style reasoning have advanced evaluation in mathematics and programming, yet physics remains comparatively underexplored. Most existing physics benchmarks evaluate only final answers, which fail to capture reasoning processes, while recent stepwise methods rely on heuristic LLM-as-judge scoring or restrictive linear assumptions, limiting reliability and diagnostic validity.\nWe introduce PRISM-Physics, a process-level evaluation framework and benchmark for complex physics reasoning problems. Solutions are represented as directed acyclic graphs (DAGs) of formulas, explicitly encoding causal dependencies among intermediate steps to enable fine-grained, interpretable, and theoretically grounded scoring. \nWe prove the optimality of the DAG representation and the corresponding scoring policy. Combining with a fully rule-based method for symbolic formula equivalence matching that we developed, we ensure consistent validation across diverse formulations without heuristic judgments. Results show that our evaluation framework is more aligned with human experts' scoring. \nExperiments on state-of-the-art LLMs reveal persistent reasoning failures in physics, while step-level scoring offers both diagnostic insight and rich signals for later training. By combining structural rigor, theoretical guarantees, and symbolic validation, PRISM-Physics provides a principled foundation for advancing process-level evaluation and guiding the development of models with deeper scientific reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\nzhao2026prismphysics,\ntitle={{PRISM}-Physics: Causal {DAG}-Based Process Evaluation for Physics Reasoning},\nauthor={Wanjia Zhao and Qinwei Ma and Jingzhe Shi and Shirley Wu and Jiaqi Han and Yijia Xiao and Si-Yuan Chen and Xiao Luo and Ludwig Schmidt and James Zou},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=4PZMeopXzP}\n}"},"title":{"value":"PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning"},"pdf":{"value":"/pdf/95b751e95b88437e4484cd3de1a315d0a89884f4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhao|prismphysics_causal_dagbased_process_evaluation_for_physics_reasoning"},"authorids":{"value":["~Wanjia_Zhao1","~Qinwei_Ma1","~Jingzhe_Shi1","~Shirley_Wu1","~Jiaqi_Han2","~Yijia_Xiao1","~Si-Yuan_Chen1","~Xiao_Luo3","~Ludwig_Schmidt1","~James_Zou1"]},"authors":{"value":["Wanjia Zhao","Qinwei Ma","Jingzhe Shi","Shirley Wu","Jiaqi Han","Yijia Xiao","Si-Yuan Chen","Xiao Luo","Ludwig Schmidt","James Zou"]}},"tmdate":1775877057510,"pdate":1769436000548,"tcdate":1758147726568,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9900/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission9900/Authors"],"forum":"4PZMeopXzP","license":"CC BY 4.0","number":9900,"cdate":1758147726568,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission9900/-/Full_Submission","ICLR.cc/2026/Conference/Submission9900/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission9900/-/Camera_Ready_Revision"],"mdate":1775877057510,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"4PZMeopXzP","version":2},{"content":{"summary":{"value":"The paper introduces OmniWorld, a large-scale multi-domain and multi-modal dataset (RGB, depth, pose, text, optical flow, mask) designed for 4D world modeling. It unifies new synthetic data (OmniWorld-Game) and re-annotated public sets, supporting benchmarks for 3D geometry prediction and camera-controllable video generation. The dataset features higher resolution, richer modalities, and more complex dynamics than existing resources like Sintel or TartanAir. Its automated annotation pipeline and unified benchmark design expose clear weaknesses of current geometric and generative models, and fine-tuning on OmniWorld improves performance across multiple datasets.\n\nHowever, the paper lacks clarity on how dynamic objects and camera motions are modeled, leaving its 4D claim partially unsubstantiated. The work also relies heavily on synthetic data, with limited validation on real-world domains (human, robot, internet). Moreover, the quality and uncertainty of automatically generated annotations are not quantitatively analyzed.\n\nOverall, OmniWorld is a valuable and ambitious contribution that could serve as a strong foundation for 3D and 4D perception research. Clearer handling of dynamic scenes, stronger real-world evaluation, and annotation-quality validation would further improve the paper. My current recommendation is Borderline Accept."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. **Significant scale and modality improvement.**\nCompared with existing synthetic datasets, OmniWorld-Game offers clear advantages in resolution, frame count, and modality diversity. It also covers more dynamic and complex scenes. Table 1’s comparison against Sintel, TartanAir, and HyperSim highlights this superiority.\n\n2. **Cross-domain integration and automated annotation pipeline.**\nThe paper clearly details an end-to-end annotation pipeline, from four-domain collection, video slicing and filtering, to modality annotation, and provides domain composition statistics, which support reproducibility and future extension.\n\n3. **Challenging benchmark design.**\nCompared with prior benchmarks featuring short sequences and smooth camera motion, OmniWorld-Game introduces longer sequences, complex trajectories, and dynamic scenes. On this benchmark, most state-of-the-art GFM/video generation models exhibit weaknesses in temporal consistency and camera control (as shown in Table 3/Table 4 and qualitative examples).\n\n4. **Demonstrated training value.**\nModels such as DUSt3R, CUT3R, and Reloc3R, when fine-tuned on OmniWorld, show consistent performance gains on standard public benchmarks (Sintel, KITTI, NYUv2), indicating that the dataset provides practical benefits for improving 3D geometry and temporal coherence."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Unclear handling of dynamic 4D content.**\nGiven the central claim of 4D world modeling, the paper should more explicitly describe how dynamic objects and camera motions are represented during dataset construction, particularly in the synthetic (simulation) domains. It is unclear whether the dataset systematically varies object motion patterns or camera trajectories, or includes a taxonomy distinguishing static versus dynamic regions. In evaluation, the study focuses mainly on static 3D foundation models; incorporating 4D reconstruction or dynamic scene methods (e.g., MegaSAM) would make the benchmark more consistent with its 4D ambition and strengthen its empirical validity.\n\n2. **High reliance on synthetic data and limited real-world evaluation.**\nAlthough the authors highlight the multi-domain nature of OmniWorld, both training and evaluation are largely centered on OmniWorld-Game, the synthetic/game-based subset. Systematic experiments on the human, robot, and internet domains are limited, which makes it difficult to assess how well the dataset and benchmarks generalize to real-world conditions. A stronger cross-domain analysis would clarify the practical robustness of the proposed data.\n\n3. **Lack of annotation-quality assessment and uncertainty analysis.**\nMany modalities (depth, pose, optical flow, and foreground masks) rely on automatic or pseudo-labeling pipelines, yet the paper provides no quantitative evaluation of their accuracy or reliability. It would be valuable to include error quantification (e.g., comparisons with manual ground truth or high-confidence models), failure-case analysis, and sensitivity studies showing how annotation noise affects training and evaluation outcomes. These analyses would help validate the trustworthiness of the dataset’s supervision signals."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918037963,"tcdate":1761826766373,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5398/Reviewer_XxMY"],"signatures":["ICLR.cc/2026/Conference/Submission5398/Reviewer_XxMY"],"forum":"1y1YFKb9pp","number":2,"license":"CC BY 4.0","cdate":1761826766373,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5398/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918037963,"domain":"ICLR.cc/2026/Conference","replyto":"1y1YFKb9pp","id":"HkWNNE0atr","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multi-Domain","Multi-Modal","World Modeling"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"The field of 4D world modeling—aiming to jointly capture spatial geometry and temporal dynamics—has witnessed  remarkable progress in recent years, driven by advances in large-scale generative models and multimodal learning.  However, the development of truly general 4D world models remains fundamentally constrained by the availability  of high-quality data. Existing datasets and benchmarks often lack the dynamic complexity, multi-domain diversity,  and spatial-temporal annotations required to support key tasks such as 4D geometric reconstruction, future  prediction, and camera-controlled video generation. To address this gap, we introduce OmniWorld, a large-scale,  multi-domain, multi-modal dataset specifically designed for 4D world modeling. OmniWorld consists of a newly  collected OmniWorld-Game dataset and several curated public datasets spanning diverse domains. Compared with  existing synthetic datasets, OmniWorld-Game provides richer modality coverage, larger scale, and more realistic  dynamic interactions. Based on this dataset, we establish a challenging benchmark that exposes the limitations of  current state-of-the-art (SOTA) approaches in modeling complex 4D environments. Moreover, fine-tuning existing  SOTA methods on OmniWorld leads to significant performance gains across 4D reconstruction and video generation  tasks, strongly validating OmniWorld as a powerful resource for training and evaluation. We envision OmniWorld  as a catalyst for accelerating the development of general-purpose 4D world models, ultimately advancing machines’  holistic understanding of the physical world."},"_bibtex":{"value":"@inproceedings{\nzhou2026omniworld,\ntitle={OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling},\nauthor={Yang Zhou and Yifan Wang and Jianjun Zhou and Wenzheng Chang and Haoyu Guo and Zizun Li and Kaijing Ma and Xinyue Li and Yating Wang and Haoyi Zhu and Mingyu Liu and Dingning Liu and Jiange Yang and Zhoujie Fu and Junyi Chen and Chunhua Shen and Jiangmiao Pang and Kaipeng Zhang and Tong He},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=1y1YFKb9pp}\n}"},"title":{"value":"OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling"},"pdf":{"value":"/pdf/094da53cbae5b8e6ec559bd94c67ca8f5e0b7844.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhou|omniworld_a_multidomain_and_multimodal_dataset_for_4d_world_modeling"},"authorids":{"value":["~Yang_Zhou34","~Yifan_Wang42","~Jianjun_Zhou1","~Wenzheng_Chang1","~Haoyu_Guo1","~Zizun_Li1","~Kaijing_Ma2","~Xinyue_Li13","~Yating_Wang3","~Haoyi_Zhu1","~Mingyu_Liu1","~Dingning_Liu4","~Jiange_Yang1","~Zhoujie_Fu2","~Junyi_Chen4","~Chunhua_Shen2","~Jiangmiao_Pang1","~Kaipeng_Zhang1","~Tong_He2"]},"authors":{"value":["Yang Zhou","Yifan Wang","Jianjun Zhou","Wenzheng Chang","Haoyu Guo","Zizun Li","Kaijing Ma","Xinyue Li","Yating Wang","Haoyi Zhu","Mingyu Liu","Dingning Liu","Jiange Yang","Zhoujie Fu","Junyi Chen","Chunhua Shen","Jiangmiao Pang","Kaipeng Zhang","Tong He"]}},"version":2},{"content":{"summary":{"value":"This paper introduces COMPOL, a neural operator framework aimed at modeling coupled multi-physics PDE systems. The key idea is to extend the Fourier Neural Operator (FNO) with additional latent feature aggregation mechanisms based on recurrent (GRU) and attention modules. The authors claim that COMPOL is architecture-agnostic and capable of generalizing to other operator learning paradigms (e.g., DeepONet, GNO, transformer-based solvers). The paper includes theoretical analyses (universal approximation, convergence, stability) and experimental evaluations on several multi-physics datasets, reporting improved predictive accuracy and generalization over baseline neural operators."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Would a self-attention layer in a standard transformer not already perform the proposed “cross-process aggregation”? How does COMPOL differ fundamentally from multi-head attention across channels?  \n\n2.  If COMPOL is truly architecture-agnostic, why are all experiments conducted on FNO backbones? Can the authors show results with DeepONet or transformer-based operators?  \n\n3. What happens when the number of coupled processes increases significantly (e.g., >10)? Is the recurrent aggregation stable and scalable?  \n\n4. Are the reported performance gains still evident when parameter count and compute budget are matched with strong baselines (e.g., models scaled to ~100M parameters) under identical training conditions?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Addresses an important problem: scalable multi-physics operator learning  \n- Incorporates latent aggregation ideas that could, in principle, enhance cross-process information flow.  \n- Provides a reasonably thorough set of experiments and ablation studies.  \n- Includes theoretical justification for stability and approximation, even if standard."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Clarity and Writing Quality  \n    - The paper is difficult to follow, especially in Section 3.2–3.3. The presentation is redundant, vague, and hand-waving.  \n    - Key terms such as “process-specific evolution,” “inter-process interactions” and “global context” are not rigorously defined.   \n\n2. Lack of Conceptual Novelty\n     - Despite being described as a new framework, COMPOL is largely built upon FNO, with some modifications—essentially adding a latent GRU or attention module between  FNO layers.  \n     - The claimed unified or architecture-agnostic nature is not demonstrated; all experiments are conducted with FNO backbones only.  \n\n3.  Insufficient Benchmarking\n       - Despite the availability of public multi-physics benchmarks(*PDEArena*, *PDEBench*, *the Well dataset*), the authors rely solely on custom datasets.  \n       - Stronger recent baselines such as DPOT, MPP, PDE-Transformer, and CViT are not included.  \n       - The main baselines (FNO, UFNO, Transolver, DeepONet) are either outdated or adapted in ways that may not reflect their best performance.  \n\n4.  Theoretical Part is Superficial  \n       - The so-called “proofs” of universal approximation, convergence, and stability are directly adapted from existing neural operator theorems (Kovachki et al., 2023) and do not depend on the new aggregation mechanism.  Hence, the theoretical contribution is minimal and not novel."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918366882,"tcdate":1760895332903,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5942/Reviewer_RALb"],"signatures":["ICLR.cc/2026/Conference/Submission5942/Reviewer_RALb"],"forum":"k3DrCkpCok","number":1,"license":"CC BY 4.0","cdate":1760895332903,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5942/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918366882,"domain":"ICLR.cc/2026/Conference","replyto":"k3DrCkpCok","id":"dW4tTvAXei","forumContent":{"TLDR":{"value":"A versatile multi-physics operator learning framwork"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["operator learning","physical simulations","coupled and multi-physics simulations"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Multi-physics simulations play an essential role in accurately modeling complex interactions across diverse scientific and engineering domains. Although neural operators, especially the Fourier Neural Operator (FNO), have significantly improved computational efficiency, they often fail to effectively capture intricate correlations inherent in coupled physical processes. To address this limitation, we introduce COMPOL, a novel coupled multi-physics operator learning framework. COMPOL extends conventional operator architectures by incorporating sophisticated recurrent and attention-based aggregation mechanisms, effectively modeling interdependencies among interacting physical processes within latent feature spaces. Our approach is architecture-agnostic and seamlessly integrates into various neural operator frameworks that involve latent space transformations. Extensive experiments on diverse benchmarks—including biological reaction-diffusion systems, pattern-forming chemical reactions, multiphase geological flows, and thermo-hydro-mechanical processes — demonstrate that COMPOL consistently achieves superior predictive accuracy compared to state-of-the-art methods."},"_bibtex":{"value":"@misc{\nsun2026compol,\ntitle={{COMPOL}: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations},\nauthor={Yifei Sun and Tao Wang and Junqi Qu and Yushun Dong and Hewei Tang and Shibo Li},\nyear={2026},\nurl={https://openreview.net/forum?id=k3DrCkpCok}\n}"},"title":{"value":"COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations"},"pdf":{"value":"/pdf/786ba3fc88eeb538ba6f9fce819f955aa4e43702.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"sun|compol_a_unified_neural_operator_framework_for_scalable_multiphysics_simulations"},"authorids":{"value":["~Yifei_Sun11","~Tao_Wang65","~Junqi_Qu1","~Yushun_Dong1","~Hewei_Tang1","~Shibo_Li1"]},"authors":{"value":["Yifei Sun","Tao Wang","Junqi Qu","Yushun Dong","Hewei Tang","Shibo Li"]}},"version":2},{"content":{"summary":{"value":"The paper introduces PHYX, a benchmark for physics grounded reasoning in multimodal settings. It offers 3k unique visual physics problems with both multiple choice and open ended formats, spans six domains and six reasoning types, and provides three text variants to probe text reliance. The authors evaluate many recent LLMs and MLLMs and present error analyses and takeaways about visual grounding and modality fusion. The topic is timely and the dataset could be useful to the community."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- How was the de redundancy edit validated. Did independent annotators confirm that answers remain unchanged and that difficulty is comparable?\n- See the weaknesses."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Important problem focus. Physical reasoning that integrates perception, symbolic manipulation, and real world constraints is a valuable target for evaluation.\n- Breadth of coverage across six physics domains and six reasoning categories with both MC and OE formats.\n- Three input variants to study redundancy and text dependence.\n- Integration with common eval toolkits and release plan for one click evaluation.\n- Evaluation includes both MLLMs and text only LLMs through captions, which enables cross modality comparisons.\n\nMy recommendation: reject.\n1. Methodological weaknesses in evaluation and judging reduce confidence in the reported gaps and rankings.\n2. Incremental novelty relative to recent physics reasoning benchmarks, with several claims that appear overstated."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Novelty claim appears overstated. Several 2025 benchmarks already target physics reasoning with images, for example PhysReason, UGPhysics, SeePhys, PhysUniBench, and others. PHYX is larger and uses some forms of de redundancy, but the claim of first large scale benchmark is not well supported.\n- Human baseline is too small and too weak (table 2 is basically empty for the human baselines). Only 15 students answered 18 questions each, with no per question overlap, no variance estimates, and seemingly only in one setting. This cannot support strong claims about a persistent 10 point human to model gap.\n- Moreover, MC accuracy for GPT-5 reaches 90.9! in test which conflicts with the headline message that \"all models struggle\", yet this is not reconciled with the human comparison.\n- Dataset scale and counts are confusing. The paper alternates between 3k and 6k questions. Table 1 says total new questions 6k yet unique questions 3k, and the text claims 3k problems.\n- Claims of realistic images (a central assertion in the paper, as shown in Fig. 3) are inconsistent with numerous examples that resemble textbook-style drawings. The paper itself later clarifies that these are not photographs. Therefore, the claim that PHYX uses realistic scenes should be moderated or better substantiated. Statements such as “We observe that the images in our dataset are highly realistic, often depicting concrete physical scenarios rather than stylized or abstract illustrations” must be supported, especially when they form part of the paper’s main claims.\n- Evaluation details are under specified. CoT prompting, temperatures, seeds, and retries are not fixed. \n- Filtering out the shortest 10 percent of questions as a quality control step is a weak proxy for difficulty. This is definitely not a \"rigorous process (that) plays a crucial role in maintaining the quality and difficulty of PHYX\". In reality, short questions can be the hardest ones!\n- The paper is generally difficult to read, with long sentences containing many redundant words and unnecessary phrases that do not contribute to the content and serve only as fillers."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924040423,"tcdate":1761770091861,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13409/Reviewer_cgQS"],"signatures":["ICLR.cc/2026/Conference/Submission13409/Reviewer_cgQS"],"forum":"UK0j7E1f4i","number":2,"license":"CC BY 4.0","cdate":1761770091861,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13409/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924040423,"domain":"ICLR.cc/2026/Conference","replyto":"UK0j7E1f4i","id":"H3lUqBfEtY","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["LMM Evaluation","Physical Benchmark","Reasoning Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Existing benchmarks fail to capture a crucial aspect of intelligence: physical reasoning, the integrated ability to combine domain knowledge, symbolic reasoning, and understanding of real-world constraints. To address this gap, we introduce PhyX: the first large-scale benchmark designed to assess models' capacity for physics-grounded reasoning in visual scenarios. PhyX includes 3K meticulously curated multimodal questions spanning 6 reasoning types across 25 sub-domains and 6 core physics domains: thermodynamics, electromagnetism, mechanics, modern physics, optics, and wave & acoustics. In our comprehensive evaluation, even state-of-the-art models struggle significantly with physical reasoning. GPT-o4-mini, Gemini-2.5-Pro, and GPT-5 achieve only 45.8%, 62.4%, and 65.2% accuracy respectively—performance gaps exceeding 10% compared to human experts. Our analysis exposes critical limitations in current models: over-reliance on memorized disciplinary knowledge, excessive dependence on mathematical formulations, and surface-level visual pattern matching rather than genuine physical understanding. We provide in-depth analysis through fine-grained statistics, detailed case studies, and multiple evaluation paradigms to thoroughly examine physical reasoning capabilities. To ensure reproducibility, we implement a compatible evaluation protocol based on widely-used toolkits such as VLMEvalKit and lmms-eval, enabling one-click evaluation."},"_bibtex":{"value":"@misc{\nshen2025phyx,\ntitle={PhyX: Does Your Model Have the ''Wits'' for Physical Reasoning?},\nauthor={Hui Shen and Taiqiang Wu and Qi Han and Yunta Hsieh and Jizhou Wang and Yuyue Zhang and Yuxin Cheng and Zijian Hao and Yuansheng Ni and Xin Wang and Zhongwei Wan and Kai Zhang and Wendong XU and Jing Xiong and Ping Luo and Wenhu Chen and Chaofan Tao and Zhuoqing Mao and Ngai Wong},\nyear={2025},\nurl={https://openreview.net/forum?id=UK0j7E1f4i}\n}"},"title":{"value":"PhyX: Does Your Model Have the \"Wits\" for Physical Reasoning?"},"pdf":{"value":"/pdf/11663e3de0b9f40c254399c8cc6c92a219eba44f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"shen|phyx_does_your_model_have_the_wits_for_physical_reasoning"},"authorids":{"value":["~Hui_Shen2","~Taiqiang_Wu1","~Qi_Han9","~Yunta_Hsieh1","~Jizhou_Wang2","~Yuyue_Zhang3","~Yuxin_Cheng2","~Zijian_Hao3","~Yuansheng_Ni1","~Xin_Wang71","~Zhongwei_Wan1","~Kai_Zhang10","~Wendong_XU2","~Jing_Xiong4","~Ping_Luo2","~Wenhu_Chen3","~Chaofan_Tao1","~Zhuoqing_Mao1","~Ngai_Wong1"]},"authors":{"value":["Hui Shen","Taiqiang Wu","Qi Han","Yunta Hsieh","Jizhou Wang","Yuyue Zhang","Yuxin Cheng","Zijian Hao","Yuansheng Ni","Xin Wang","Zhongwei Wan","Kai Zhang","Wendong XU","Jing Xiong","Ping Luo","Wenhu Chen","Chaofan Tao","Zhuoqing Mao","Ngai Wong"]}},"version":2},{"content":{"summary":{"value":"This paper presents D-REX, a differentiable real-to-sim-to-real pipeline that couples Gaussian Splat Representations for photorealistic 3D reconstruction with a differentiable physics engine for object-mass identification and force-aware grasp policy learning. The system jointly optimizes physical parameters (mass) from robot interaction videos and learns manipulation policies conditioned on the inferred mass, closing the sim-to-real loop for dexterous grasping tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see my weaknesses section, thanks!"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper demonstrates a technically competent system that merges 3DGS and differentiable physics for vision-based grasping. Authors have performed real-world experiments validating some of their claims."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Pipeline composition rather than a learning contribution.**\nThe full system is essentially a sequential pipeline: (1) Gaussian Splatting for 3D reconstruction with VLMs, (2) System identification to calibrate physical parameters, and (3) a procedural grasping policy that uses hand-designed grasp position and orientation heuristics. There is no novel algorithmic contribution or learning formulation that connects these modules beyond standard differentiable chaining. \n\n**2. Hand-designed grasp prediction.**\nThe grasping procedure relies on manually defined rules. This is not significantly different from prior grasp pipelines that use geometry-based scoring or analytical quality metrics.\n\n**3. No clear advantage over existing methods.**\nThe paper does not demonstrate how D-REX materially improves over existing differentiable grasping frameworks that already combine differentiable rendering and physics. The quantitative differences appear modest and could stem from tuning rather than a new principle. Additionally, just identifying mass, without taking materials into consideration, seems incomplete for robotics purposes."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915495237,"tcdate":1762201169636,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission333/Reviewer_snrT"],"signatures":["ICLR.cc/2026/Conference/Submission333/Reviewer_snrT"],"forum":"13jshGCK9i","number":4,"license":"CC BY 4.0","cdate":1762201169636,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission333/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915495237,"domain":"ICLR.cc/2026/Conference","replyto":"13jshGCK9i","id":"3nSMUZEWeR","forumContent":{"TLDR":{"value":"Differentiable Real-to-Sim-to-Real Engine for Learning Robotic Grasping"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Real-to-Sim-to-Real; Differentiable Simulation; Learning Robotic Policies from Videos; System Identification;"]},"supplementary_material":{"value":"/attachment/3bd2106e00972ff1710d086acfd45b53b8a4513e.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Simulation provides a cost-effective and flexible platform for data generation and policy learning to develop robotic systems. However, bridging the gap between simulation and real-world dynamics remains a significant challenge, especially in physical parameter identification. In this work, we introduce a real-to-sim-to-real engine that leverages the Gaussian Splat representations to build a differentiable engine, enabling object mass identification from real-world visual observations and robot control signals, while enabling grasping policy learning simultaneously. Through optimizing the mass of the manipulated object, our method automatically builds high-fidelity and physically plausible digital twins. Additionally, we propose a novel approach to train force-aware grasping policies from limited data by transferring feasible human demonstrations into simulated robot demonstrations. Through comprehensive experiments, we demonstrate that our engine achieves accurate and robust performance in mass identification across various object geometries and mass values. Those optimized mass values facilitate force-aware policy learning, achieving superior and high performance in object grasping, effectively reducing the sim-to-real gap."},"_bibtex":{"value":"@inproceedings{\nlou2026drex,\ntitle={D-{REX}: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping},\nauthor={Haozhe Lou and Mingtong Zhang and Haoran Geng and Hanyang Zhou and Sicheng He and Zhiyuan Gao and Siheng Zhao and Jiageng Mao and Pieter Abbeel and Jitendra Malik and Daniel Seita and Yue Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=13jshGCK9i}\n}"},"title":{"value":"D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping"},"pdf":{"value":"/pdf/f9a760a8d8f85c781c16d61cffc72bf5351d1c22.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lou|drex_differentiable_realtosimtoreal_engine_for_learning_dexterous_grasping"},"authorids":{"value":["~Haozhe_Lou2","~Mingtong_Zhang1","~Haoran_Geng1","~Hanyang_Zhou1","~Sicheng_He1","~Zhiyuan_Gao1","~Siheng_Zhao1","~Jiageng_Mao1","~Pieter_Abbeel2","~Jitendra_Malik2","~Daniel_Seita1","~Yue_Wang2"]},"authors":{"value":["Haozhe Lou","Mingtong Zhang","Haoran Geng","Hanyang Zhou","Sicheng He","Zhiyuan Gao","Siheng Zhao","Jiageng Mao","Pieter Abbeel","Jitendra Malik","Daniel Seita","Yue Wang"]}},"version":2},{"content":{"summary":{"value":"This paper explores how LMs capture the factual knowledge: by constructing an internal KG or simply memorized QA pairs.\nTo do that, the authors use LLaMA to construct a synthetic data set and train GPT-2 with it. Then, the verification is done by prompting and probing."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The topic is good. This paper studies a very interesting research problem. Figuring it out can help people better understand how LMs encode factual knowledge, and how to better enhance/edit factual knowledge. \n2. The synthetic setting is novel. Researching the factual knowledge captured during pretraining is fairly intractable, since it's hard to annotate facts from the pretraining corpus and run analysis experiments on pretraining. This paper constructs a synthetic pretraining environment using LLaMA and reproduces the pretraining of small GPT-2."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My main concern is the experiment setting. The general design is okay, but some essential points are missing. I'll elaborate in detail below.\n  1. The template to construct the synthetic dataset is problematic. \n\n      a) First of all, it's not diverse. Even if using LLaMA, it works more like a simple verbalizer. We know that factual knowledge is sparse in the text (compared to linguistic knowledge): some facts exist across multiple sentences. But in the synthetic dataset, facts are verbalized in a quite limited way. For analysis, it's fine since LM is trained in a general way. However, this paper includes the pretraining part. I don't think pretraining with this dataset is reasonable. It's easy for LM to learn shortcuts.\n      \n      b) Second, the biography is too short and quite similar to each other. LMs may capture and extract facts relating to one person or topic. Such short and similar text might cause LMs to be confused when learning facts. Even for humans, it's hard to differentiate persons/topics with such short and similar introductions.\n\n2. Contributions 1 and 2 are not clear to me. \n\n      a) The synthetic data is generated randomly, and there could be conflicts between each other. Researching the generalization is a bit confusing (contribution 2). \n\n      b) While for contribution 1, it seems the authors just try to test if GPT-2 generalizes well to understand the question and extract facts for answering. I don't know if GPT-2 is pretrained from scratch or not (see question 2). If yes, I don't think such pretraining data can let GPT-2 understand language. If not, then the setting is not reasonable, and facts may have been learned in GPT-2 before training.\n\nIn summary, I think this paper explores an interesting topic. However, the above concerns are essential to prove that the findings are correct. If the authors can address my concerns successfully, I'll be glad to raise my score to strong accept."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"1. Why use LLaMA-30B, not 65B or 7B?\n2. For the GPT-2 small, do you \"pretrain\" it from scratch or pretrained GPT-2?\n3. Why use LoRA for the finetuning? If I understand correctly, the model is quite small, and it's totally unnecessary to use parameter-efficient tuning. It not only reduces the test accuracy but also makes the setting harder to follow. And thus much information in Figure 2 is redundant.\n4. The augmentations in section 4.2 are confusing. Why this kind of augmentation? Do you follow existing work or propose it at first? From my view, I don't think this can work as augmentation. It can just alleviate the problem of the problematic templates.\n5. Why call the dataset a semi-synthetic biography? Generated from LLM doesn't mean the fact is true. I think it's still synthetic data.\n6. What does the \"(Part A)\" mean in the title?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700669138303,"tcdate":1698286801087,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission738/Reviewer_HX8x"],"signatures":["ICLR.cc/2024/Conference/Submission738/Reviewer_HX8x"],"forum":"rVnxymbsOS","number":1,"license":"CC BY 4.0","cdate":1698286801087,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission738/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700669138303,"domain":"ICLR.cc/2024/Conference","replyto":"rVnxymbsOS","id":"TYFyCX1v4c","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"Even if LLMs losslessly memorize the pretraining data, it may not be finetuned to extract knowledge from it. Probing techniques suggest that data augmentation is necessary on the pretrain level."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Interpretability","Transformers","Language Models","Linear Probing","Inner Working","Factual Knowledge"]},"supplementary_material":{"value":"/attachment/88c4d31b4153352b06cbd8b852d6e51ab228ffa2.zip"},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Large language models can store extensive world knowledge, often *extractable* through question-answering (e.g., \"What is Abraham Lincoln's birthday?\"). However, it's unclear whether the model answers questions based on exposure to exact/similar questions during training, or if it genuinely extracts knowledge from the source (e.g., Wikipedia biographies). \n\nIn this paper, we conduct an in-depth study of this problem using a controlled set of semi-synthetic biography data. We uncover a relationship between the model's knowledge extraction ability and different *diversity measures* of the training data. We conduct (nearly) linear probing, revealing a strong correlation between this relationship and whether the model (nearly) linearly encodes the knowledge attributes at the hidden embedding of the entity names, or across the embeddings of other tokens in the training text."},"_bibtex":{"value":"@misc{\nallen-zhu2024knowledge,\ntitle={Knowledge Storage and Extraction in Language Models (Part A)},\nauthor={Zeyuan Allen-Zhu and Yuanzhi Li},\nyear={2024},\nurl={https://openreview.net/forum?id=rVnxymbsOS}\n}"},"title":{"value":"Knowledge Storage and Extraction in Language Models (Part A)"},"pdf":{"value":"/pdf/20c78e4d75e010aff02861817ba67ab13ab0b7e1.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"allenzhu|knowledge_storage_and_extraction_in_language_models_part_a"},"authorids":{"value":["~Zeyuan_Allen-Zhu1","~Yuanzhi_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyuan Allen-Zhu","Yuanzhi Li"]}},"version":2},{"content":{"summary":{"value":"This paper presents PhysWorld, a framework for constructing physics-consistent and efficient world models of deformable objects from short real-world videos.\nThe method first builds a digital twin in an MPM simulator, where a Vision-Language Model (Qwen3) selects the appropriate constitutive material models and a global-to-local optimization refines physical parameters (e.g., friction, density, Young’s modulus).\nThe calibrated simulator then generates diverse synthetic demonstrations via Various Motion Pattern Generation (VMP-Gen) and Part-aware Physical Property Perturbation (P³-Pert).\nA GNN-based world model is trained on these demonstrations and fine-tuned using real videos.\nExperiments show that PhysWorld achieves accurate predictions and 47× faster inference than PhysTwin while maintaining visual and physical consistency.\nThe paper also provides a qualitative demonstration of model-based planning using MPPI for rope manipulation."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Could you provide **quantitative control results** (e.g., success rate, trajectory error, or control frequency) to demonstrate how PhysWorld’s high inference speed translates to better control performance?  \n2. How physically realistic are the synthetic demonstrations generated by VMP-Gen and P³-Pert?  \n3. How stable is the **VLM-based material selection** under noisy or complex real-world videos?  \n4. How transferable are the fine-tuned physical parameters (Φ) to new objects or scenes?  \n5. How does inference speed scale with larger particle counts or multiple interacting objects?  \n6. Could the model be extended to **action-free settings** through action inference or inverse dynamics?  \n7. Have you considered comparing the GNN to a Transformer-based or MLP-based world model for efficiency and accuracy?  \n8. Would adding explicit physical consistency losses (e.g., for momentum or volume preservation) further improve realism?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. **Comprehensive and well-integrated framework**  \n   The paper thoughtfully combines VLM-based material selection, multi-stage parameter optimization, and physics-guided data augmentation into a coherent pipeline bridging simulation and learning.\n\n2. **Strong accuracy–efficiency trade-off**  \n   The proposed GNN-based world model achieves high predictive accuracy while running at **799 FPS**, demonstrating its potential for real-time applications that require fast yet physically consistent inference.\n\n3. **Visual prediction capability**  \n   Integration of 3D Gaussian Splatting and Linear Blend Skinning enables **action-conditioned video generation**, evaluated with PSNR, SSIM, and LPIPS metrics.\n\n4. **Generalization to unseen interactions**  \n   The model generalizes well to new manipulation sequences and shows physically plausible deformation behavior (Fig. 3, Table 2).\n\n5. **Detailed ablation studies**  \n   The effects of global-to-local optimization, VMP-Gen, and P³-Pert are clearly quantified (Tables 3–5), and the results support the design choices."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Lack of quantitative downstream (control) evaluation — a key remaining limitation**  \n   While the framework’s real-time capability is convincingly demonstrated (47× faster inference), the paper does not provide **quantitative results in downstream control or planning tasks**.  \n   The MPPI example (Fig. 4) is qualitative only. Demonstrating success rates, trajectory errors, or computation times in model-based control would substantially strengthen the practical significance.  \n   Given that real-time inference is the method’s major advantage, connecting it to tangible control improvements would enhance the paper’s overall impact.\n\n2. **Action-conditioned requirement**  \n   The world model relies on explicit control inputs (\\(a_t\\)) extracted from video, and learning from action-free observational data is not addressed. This limits applicability to passive video datasets.\n\n3. **Architectural diversity**  \n   The study focuses solely on GNNs. Comparisons with alternative architectures (e.g., Transformers or MLP-based dynamics models) could clarify whether the observed advantages are model-specific or framework-driven.\n\n4. **Realism of synthetic demonstrations**  \n   Although the proposed perturbation and trajectory generation methods improve diversity, there is no quantitative analysis of how closely these synthetic motions resemble real physical dynamics.\n\n5. **VLM-based model selection robustness**  \n   The Qwen3-based constitutive model selection is promising, but evaluation is limited to 22 clean scenarios. Its reliability under noise, occlusion, or mixed-material settings remains unclear.\n\n6. **Scalability and physical consistency metrics**  \n   Experiments are conducted on relatively small particle systems (≈100–150 nodes). Performance and stability on larger, multi-object, or more complex scenes are not analyzed.  \n   Additionally, metrics directly assessing physical-law consistency (e.g., conservation of momentum or energy) are not reported.\n\n7. **Reproducibility details**  \n   While optimization procedures are well described, training sensitivity (e.g., to initialization, hyperparameters) and convergence analyses are not provided, making it difficult to assess robustness.\n\n8. **Conceptual scope of “World Model”**  \n   The method models object-level deformable dynamics rather than a full scene-level world model. Clarifying this scope would help set appropriate expectations."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918498943,"tcdate":1761976323468,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6146/Reviewer_YYty"],"signatures":["ICLR.cc/2026/Conference/Submission6146/Reviewer_YYty"],"forum":"dggfgPzGFW","number":5,"license":"CC BY 4.0","cdate":1761976323468,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6146/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918498943,"domain":"ICLR.cc/2026/Conference","replyto":"dggfgPzGFW","id":"WkTfglyr68","forumContent":{"TLDR":{"value":"We propose PhysWorld, a novel framework that synthesizes physics-aware demonstrations to learn efficient world models."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["World Models","Deformable Objects","Demonstration Synthesis"]},"supplementary_material":{"value":"/attachment/675f5ce44444cb7da292f6568ada0a79712054d7.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Interactive world models that simulate object dynamics are crucial for robotics, VR, and AR. However, it remains a significant challenge to learn physics-consistent dynamics models from limited real-world video data, especially for deformable objects with spatially-varying physical properties. To overcome the challenge of data scarcity, we propose PhysWorld, a novel framework that utilizes a simulator to synthesize physically plausible and diverse demonstrations to learn efficient world models. Specifically, we first construct a physics-consistent digital twin within MPM simulator via constitutive model selection and global-to-local optimization of physical properties. Subsequently, we apply part-aware perturbations to the physical properties and generate various motion patterns for the digital twin, synthesizing extensive and diverse demonstrations. Finally, using these demonstrations, we train a lightweight GNN-based world model that is embedded with physical properties. The real video can be used to further refine the physical properties. PhysWorld achieves accurate and fast future predictions for various deformable objects, and also generalizes well to novel interactions. Experiments show that PhysWorld has competitive performance while enabling inference speeds 47 times faster than the recent state-of-the-art method, i.e., PhysTwin. The code and pre-trained models will be publicly available."},"_bibtex":{"value":"@misc{\nyang2026physworld,\ntitle={PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis},\nauthor={Yu Yang and Zhilu Zhang and Xiang Zhang and Yihan Zeng and Hui Li and Wangmeng Zuo},\nyear={2026},\nurl={https://openreview.net/forum?id=dggfgPzGFW}\n}"},"title":{"value":"PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis"},"pdf":{"value":"/pdf/7af8d3b513222a70e305717fcbc2b09bc6b2e97a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yang|physworld_from_real_videos_to_world_models_of_deformable_objects_via_physicsaware_demonstration_synthesis"},"authorids":{"value":["~Yu_Yang40","~Zhilu_Zhang2","~Xiang_Zhang45","~Yihan_Zeng1","~Hui_Li40","~Wangmeng_Zuo3"]},"authors":{"value":["Yu Yang","Zhilu Zhang","Xiang Zhang","Yihan Zeng","Hui Li","Wangmeng Zuo"]}},"version":2},{"content":{"comment":{"value":"### Q3: Scope & Evaluation on External Benchmarks\n\nWe thank the reviewer for raising two points here, one on scope and one on external benchmarks.\n\n- **On Scope:** Our focus on particle-based continuum dynamics is intentional. Fluids are a notoriously difficult ([1]Zheng et al., 2024) but foundational starting point ([2]Lin et al., 2020). Our transfer experiments (Fig 4) to **cloth, sand, and smoke** already prove the method is not just 'about fluids' but about this general principle.\n\n  [1]Zheng, Z., Yan, X., Chen, Z., Wang, J., Lim, Q.Z., Tenenbaum, J.B., & Gan, C. (2024). ContPhy: Continuum Physical Concept Learning and Reasoning from Videos. *ArXiv, abs/2402.06119*.\n\n  [2]Lin, X., Wang, Y., Olkin, J., & Held, D. (2020). SoftGym: Benchmarking Deep Reinforcement Learning for Deformable Object Manipulation. *Conference on Robot Learning*.\n\n- **On External Benchmarks:** The reviewer's request for evaluation on external benchmarks brings up a core premise of our paper. As stated in our introduction (lines 049-052), high-level reasoning benchmarks (like ContPhy) \"entangle multiple capabilities\" (vision, language, logic, *and* physics). MLLMs fail at these tasks, but it's impossible to know *why*. Our paper's main contribution is to **disentangle** this problem and be the first to \"focus on its very first step... intuitive physics understanding.\" Our NFS and TCV tasks are specifically designed to isolate and evaluate this fundamental capability, which is a critical and overlooked bottleneck.  We evaluate our method on the high-level **ContPhy** benchmark. The results are shown in the table below:\n\n| Type      | Settings | Zero-shot | Finetune | SDF-Ours  |\n| --------- | -------- | --------- | -------- | --------- |\n| **Fluid** | Prop.    | 48.00     | 60.00    | 67.06     |\n|           | Pred.    | 5.77      | 6.73     | 10.38     |\n| **Soft**  | Prop.    | 50.67     | 42.66    | 50.33     |\n|           | Pred.    | 21.59     | 40.90    | 42.05     |\n| **Cloth** | Prop.    | 48.00     | 44.00    | 46.67     |\n|           | Pred.    | 44.34     | 51.13    | 58.37     |\n\nAs the table shows, our SDF-Ours model clearly improves prediction performance (e.g., fluid *prediction* from **5.77%** to **10.38%**; cloth *prediction* from **44.34%** to **58.37%**), while the property accuracy improvement is less significant. This is because our proposed SDF does not solve the high-level task outright. This is reasonable and expected: our method is designed to enhance the *first step* (perception). Fully solving these complex reasoning tasks would require incorporating more explicit high-level reasoning modules, which is an important, but distinct, future research direction. Our work provides the essential perceptual foundation that these future reasoning models will require.\n\n\n\nIn summary, our new comparisons against other baseline (e.g., optical flow) and our bias analysis with DINOv2 confirm our original claims. We hope this new evidence makes our contribution clear and resolves the reviewer's concerns."},"title":{"value":"Rebuttal to Reviewer Vi2b (Part 2/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763629591546,"tcdate":1763629591546,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7741/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission7741/Authors"],"forum":"Ax02eR2c3d","number":9,"license":"CC BY 4.0","cdate":1763629591546,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7741/-/Official_Comment"],"mdate":1763629591546,"domain":"ICLR.cc/2026/Conference","replyto":"7bc0LOYnOe","id":"v7qGKrgKen","forumContent":{"TLDR":{"value":"This paper introduces two low-level tasks to test intuitive physics understanding and proposes Scene Dynamic Field, a method to integrate visual representation from physics simulators to MLLMs while showcasing generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multi-modal LLM","Intuitive Physics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. \nIn this work, we investigate the first step of physical reasoning, i.e., **intuitive physics understanding**, revealing substantial limitations in understanding the dynamics of continuum objects. \nTo isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. \nTo address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. \nSDF substantially improves performance, achieving up to $20.7\\%$ gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field."},"_bibtex":{"value":"@inproceedings{\nli2026beyond,\ntitle={Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models},\nauthor={Nanxi Li and Xiang Wang and Yuanjie Chen and Haode Zhang and Hong Li and Yong-Lu Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Ax02eR2c3d}\n}"},"title":{"value":"Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models"},"pdf":{"value":"/pdf/3b199b74bbdbddb8166b5e799fc00ddb71a413ae.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|beyond_static_vision_scene_dynamic_field_unlocks_intuitive_physics_understanding_in_multimodal_large_language_models"},"authorids":{"value":["~Nanxi_Li1","~Xiang_Wang19","~Yuanjie_Chen2","~Haode_Zhang2","~Hong_Li5","~Yong-Lu_Li1"]},"authors":{"value":["Nanxi Li","Xiang Wang","Yuanjie Chen","Haode Zhang","Hong Li","Yong-Lu Li"]}},"version":2},{"content":{"summary":{"value":"The paper presents MetaSpatial, a reinforcement learning framework designed to enhance 3D spatial reasoning in vision-language models (VLMs), aimed at generating physically consistent, realistic 3D scene layouts without post-processing. The core contribution is 3D Spatial Policy Optimization (3D-SPO), an RL algorithm that integrates:\n\n- Physics-aware advantage modulation at the object level (masking x,y,z coordinate tokens and adjusting based on collision/constraint ratios),\n\n- Trajectory-level reward aggregation via multi-turn layout refinement, and\n\n- A three-tier reward structure combining format, physics, and rendering-based assessments.\n\nEmpirical results show substantial improvement in scene plausibility, formatting accuracy, and physical feasibility over baselines such as LayoutGPT, I-Design, and standard supervised Qwen2.5-VL models. MetaSpatial achieves better GPT-4o-based perceptual scores and lower collision rates. The paper also explores an SFT+RL hybrid strategy, showing efficiency improvements via pseudo-labeling high-reward rollouts."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- How sensitive is performance to the weighting factors (λ₁, λ₂, λ₃) and decay factor (γ)? Is 3D-SPO robust under different hyperparameter settings?\n\n- Can MetaSpatial generalize to real-world room scans (e.g., ScanNet or Matterport3D) beyond synthetic I-Design data?\n\n- How much of the performance gain stems from the multi-turn training vs. the physics-aware masking?\n\n- Could GPT-4o evaluation bias the model toward aesthetic rather than physical correctness?"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"The paper discusses potential misuse (e.g., synthetic environments for propaganda). These risks are minor and appropriately acknowledged. The main ethical concern is over-reliance on closed models (GPT-4o) for reward computation, which limits transparency and reproducibility in an academic context."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Well-structured problem formulation. The authors correctly identify that supervised fine-tuning (SFT) fails for spatial tasks with non-unique solutions and continuous coordinate spaces, positioning RL as a more suitable approach.\n\n- Technically sound algorithmic design. The proposed 3D-SPO builds meaningfully on GRPO by adding object-level physics-aware modulation and trajectory-level reward shaping, addressing sparse or unstable feedback in 3D reasoning.\n\n- Comprehensive reward design. The hybrid reward integrates low-level correctness, physical realism, and high-level perceptual alignment (via GPT-4o). The staged reward weighting strategy is well motivated."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Incremental conceptual novelty. While 3D-SPO is well engineered, it is primarily an adaptation of known ideas (GRPO + physics-informed masking + cumulative reward shaping). The theoretical grounding of the advantage modulation remains heuristic; no convergence or stability analysis is presented.\n\n- Evaluation largely internal. The results rely heavily on the authors’ synthetic dataset and GPT-4o-based evaluation. No human evaluation or cross-dataset validation is performed. It’s unclear how well the model generalizes to unseen real-world 3D scenes.\n\n- Ambiguity in quantitative metrics. The GPT-4o perceptual score and “format correctness” are proprietary or internal metrics with no clear definition of variance or statistical significance. Reported deltas (e.g., +0.2–0.3) may not be statistically meaningful.\n\n- Heavy dependence on GPT-4o for reward and evaluation. The rendering-based reward is computationally expensive and opaque; it conflates aesthetic preference with physical correctness, limiting reproducibility and interpretability.\n\n- Computational cost and scalability. Training requires rendering in Blender and external VLM calls (GPT-4o). The cost and latency make large-scale or real-time deployment questionable.\n\n- Lack of theoretical insight. The paper is strong empirically but weak in analysis, no ablation isolates whether improvements come from multi-turn refinement, reward design, or the masking mechanism.\n\n- Limited generality. The work focuses on interior design scenes; extension to robotics or embodied reasoning tasks is claimed but not demonstrated."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924757822,"tcdate":1762028793832,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14336/Reviewer_vWwd"],"signatures":["ICLR.cc/2026/Conference/Submission14336/Reviewer_vWwd"],"forum":"EdQzLC0Zra","number":4,"license":"CC BY 4.0","cdate":1762028793832,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14336/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924757822,"domain":"ICLR.cc/2026/Conference","replyto":"EdQzLC0Zra","id":"jleXTUAzug","forumContent":{"TLDR":{"value":"MetaSpatial leverages reinforcement learning to enhance 3D spatial reasoning in vision-language models (VLMs), enabling more structured, realistic, and adaptive scene generation."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["spatial reasoning","vision language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present MetaSpatial, the first reinforcement learning (RL) framework for enhancing 3D spatial reasoning in vision-language models (VLMs), enabling real-time 3D scene layout generation without post-processing. MetaSpatial addresses two key challenges: (i) the need for extensive post-processing, as existing VLMs lack inherent 3D spatial reasoning to generate realistic layouts; and (ii) the inefficiency of supervised fine-tuning (SFT) for layout generation due to scarcity of perfect annotations. Our core contribution is the 3D Spatial Policy Optimization (3D-SPO) algorithm, which incorporates physics-aware modulation into advantage estimates at the object level and trajectory-level reward from a training-only multi-turn refinement pipeline. This design enhances temporal credit assignment and encourages spatially consistent policy learning. Empirical evaluations across models of varying scales demonstrate that MetaSpatial improves spatial coherence, physical plausibility, and formatting stability, leading to more realistic and functionally coherent object placements applicable to metaverse environments."},"_bibtex":{"value":"@inproceedings{\npan2026metaspatial,\ntitle={MetaSpatial: Reinforcing 3D Spatial Reasoning in {VLM}s for the Metaverse},\nauthor={Zhenyu Pan and Han Liu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EdQzLC0Zra}\n}"},"title":{"value":"MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse"},"pdf":{"value":"/pdf/e9be0940303eba3e3eacf72042d92ce2235b2e49.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"pan|metaspatial_reinforcing_3d_spatial_reasoning_in_vlms_for_the_metaverse"},"authorids":{"value":["~Zhenyu_Pan1","~Han_Liu4"]},"authors":{"value":["Zhenyu Pan","Han Liu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a large-scale synthetic spoken dialogue dataset, ShareChatX, covering a wide range of scenarios, including emotion, audio, and music. It also introduces a dialogue system, OmniChat, for diverse scenarios. Then, it conducts extensive experiments to compare the proposed method and existing ones, on both a previous dataset, DailyTalk, and its proposed method, and demonstrates that OmniChat has the SOTA performance. It also provides analysis on the data scale impact, optimal synthetic data ratio, expert feature selection strategies, and complex scenarios."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. After listening to the demo audios, I observed that the synthetic data sounds relatively plain and do not deliver a strong emotion, compared to real-world conversations. How well do you think the synthetic data can capture emotional nuances? Is there any potential solution to make the synthetic data more real?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The paper contributes a very large-scale spoken dialogue dataset that covers a wide range of scenarios, which largely contributes to studies of related topics.\n\n2. The paper conducts extensive experiments and provides detailed and thorough analysis from a lot of perspectives, offering lots of useful insights."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. When trained wo/ real data, OmniChat does not leads to a large improvement, and is worse than Qwen2-Audio in metrics like METEOR, and GPT-eval. \n\n2. The paper provides some insights in finding the optimal synthetic data ratio. However, it may still be hard to tune the ratio in real-world scenarios, as the ratio choice rules discussed in the paper may not generalize to other dataset choices and tasks."}},"nonreaders":[],"tmdate":1731427733114,"tcdate":1730973132008,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3090/Reviewer_Jc5t"],"signatures":["ICLR.cc/2025/Conference/Submission3090/Reviewer_Jc5t"],"forum":"cVgOIjcNoQ","number":4,"license":"CC BY 4.0","cdate":1730973132008,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3090/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427733114,"domain":"ICLR.cc/2025/Conference","replyto":"cVgOIjcNoQ","id":"zcpR42NZhu","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"Enhancing Spoken Dialogue Systems with Scalable Synthetic Data"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Spoken Dialogue System","Synthetic Data","Multi-modal Large Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these systems still struggle to handle the full complexity of real-world conversations, including audio events, musical contexts, and emotional expressions, mainly because current dialogue datasets are constrained in both scale and scenario diversity. In this paper, we propose leveraging synthetic data to enhance the dialogue models across diverse scenarios. We introduce **ShareChatX**, the first comprehensive, large-scale dataset for spoken dialogue that spans diverse scenarios. Based on this dataset, we introduce **OmniChat**, a multi-turn dialogue system with a heterogeneous feature fusion module, designed to optimize feature selection in different dialogue contexts. In addition, we explored critical aspects of training dialogue systems using synthetic data. Through comprehensive experimentation, we determined the ideal balance between synthetic and real data, achieving state-of-the-art results on the real-world dialogue dataset DailyTalk. We also highlight the crucial importance of synthetic data in tackling diverse, complex dialogue scenarios, especially those involving audio and music. For more details, please visit our demo page at \\url{https://sharechatx.github.io/}."},"_bibtex":{"value":"@misc{\ncheng2024omnichat,\ntitle={OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios},\nauthor={Xize Cheng and Dongjie Fu and Xiaoda Yang and Minghui Fang and Ruofan Hu and Jingyu Lu and Bai Jionghao and Zehan Wang and Shengpeng Ji and Rongjie Huang and Linjun Li and Yu Chen and Tao Jin and Zhou Zhao},\nyear={2024},\nurl={https://openreview.net/forum?id=cVgOIjcNoQ}\n}"},"title":{"value":"OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios"},"pdf":{"value":"/pdf/18710de70969375ea1129596e69146d27b8f844c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"cheng|omnichat_enhancing_spoken_dialogue_systems_with_scalable_synthetic_data_for_diverse_scenarios"},"authorids":{"value":["~Xize_Cheng1","~Dongjie_Fu1","~Xiaoda_Yang1","~Minghui_Fang1","~Ruofan_Hu2","~Jingyu_Lu1","~Bai_Jionghao2","~Zehan_Wang2","~Shengpeng_Ji1","~Rongjie_Huang1","~Linjun_Li2","~Yu_Chen29","~Tao_Jin2","~Zhou_Zhao3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xize Cheng","Dongjie Fu","Xiaoda Yang","Minghui Fang","Ruofan Hu","Jingyu Lu","Bai Jionghao","Zehan Wang","Shengpeng Ji","Rongjie Huang","Linjun Li","Yu Chen","Tao Jin","Zhou Zhao"]}},"version":2},{"content":{"venue":{"value":"KDD (2) 2025"},"venueid":{"value":"dblp.org/conf/KDD/2025"},"paperhash":{"value":"li|slotpi_physicsinformed_objectcentric_reasoning_models"},"authorids":{"value":["~Jian_Li24","https://dblp.org/search/pid/api?q=author:Han_Wan:","https://dblp.org/search/pid/api?q=author:Ning_Lin:","https://dblp.org/search/pid/api?q=author:Yu-Liang_Zhan:","https://dblp.org/search/pid/api?q=author:Ruizhi_Chengze:","https://dblp.org/search/pid/api?q=author:Haining_Wang:","https://dblp.org/search/pid/api?q=author:Yi_Zhang:","https://dblp.org/search/pid/api?q=author:Hongsheng_Liu_0002:","https://dblp.org/search/pid/api?q=author:Zidong_Wang_0010:","https://dblp.org/search/pid/api?q=author:Fan_Yu:","https://dblp.org/search/pid/api?q=author:Hao_Sun_0002:"]},"html":{"value":"https://doi.org/10.1145/3711896.3737131"},"_bibtex":{"value":"@inproceedings{DBLP:conf/kdd/LiWLZCWZLWY025,\n  author={Jian Li and Han Wan and Ning Lin and Yu-Liang Zhan and Ruizhi Chengze and Haining Wang and Yi Zhang and Hongsheng Liu and Zidong Wang and Fan Yu and Hao Sun},\n  title={SlotPi: Physics-informed Object-centric Reasoning Models},\n  year={2025},\n  cdate={1735689600000},\n  pages={1376-1387},\n  url={https://doi.org/10.1145/3711896.3737131},\n  booktitle={KDD (2)},\n  crossref={conf/kdd/2025-2}\n}\n"},"abstract":{"value":"Understanding and reasoning about dynamics governed by physical laws through visual observation, akin to human capabilities in the real world, poses significant challenges. Currently, object-centric dynamic simulation methods, which emulate human behavior, have achieved notable progress but overlook two critical aspects: 1) the integration of physical knowledge into models. Humans gain physical insights by observing the world and apply this knowledge to accurately reason about various dynamic scenarios; 2) the validation of model adaptability across diverse scenarios. Real-world dynamics, especially those involving fluids and objects, demand models that not only capture object interactions but also simulate fluid flow characteristics. To address these gaps, we introduce SlotPi, a slot-based physics-informed object-centric reasoning model. SlotPi integrates a physical module based on Hamiltonian principles with a spatio-temporal prediction module for dynamic forecasting. Our experiments highlight the model's strengths in tasks such as prediction and Visual Question Answering (VQA) on benchmark and fluid datasets. Furthermore, we have created a real-world dataset encompassing object interactions, fluid dynamics, and fluid-object interactions, on which we validated our model's capabilities. The model's robust performance across all datasets underscores its strong adaptability, laying a foundation for developing more advanced world models."},"title":{"value":"SlotPi: Physics-informed Object-centric Reasoning Models"},"authors":{"value":["Jian Li","Han Wan","Ning Lin","Yu-Liang Zhan","Ruizhi Chengze","Haining Wang","Yi Zhang","Hongsheng Liu","Zidong Wang","Fan Yu","Hao Sun"]}},"tmdate":1762354108388,"pdate":1735689600000,"externalIds":["dblp:conf/kdd/LiWLZCWZLWY025"],"tcdate":1762354104736,"writers":["~"],"signatures":["~Jian_Li24"],"forum":"Q1zbCFcfeM","license":"CC BY-SA 4.0","number":659642,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762354108388,"domain":"DBLP.org","id":"Q1zbCFcfeM","version":2},{"content":{"comment":{"value":"We thank the reviewer **Zkv9** for high-quality review and the positive assessment, in particular the comments on impact, soundness, and the relevance of the sparse-observation setting. We address the main concerns below and will reflect all changes in the revised manuscript.\n\n---\n\n**Weakness (1)**\n\n>“One experiment only; The experimental validation needs more examples. I would suggest a synthetic example and one more real-data example. This would also give the authors the chance to explain inputs and outputs more clearly.”\n\n**Response**\n\nWe fully agree that a broader empirical picture is valuable. During the rebuttal period we have run two additional sets of experiments:\n\nWe add results on a second UK river system (Kielder), with three representative flood events plus a held-out test segment, comparing APILaNet against the same strong sequence baselines (CrossFormer, PatchTST, TSMixer, PatchMixer, Mamba-S4, iTransformer, N-HiTS, N-Beats). A condensed version of the results is:\n\n| Data    | Model  | APILaNet (MSE↓ / NSE↑) | CrossFormer (MSE↓ / NSE↑) | PatchTST (MSE↓ / NSE↑) | TSMixer (MSE↓ / NSE↑) | PatchMixer (MSE↓ / NSE↑) | Mamba S4 (MSE↓ / NSE↑) | iTransformer (MSE↓ / NSE↑) | N-HiTS (MSE↓ / NSE↑) | N-Beats (MSE↓ / NSE↑) |\n|---------|--------|-------------------------|----------------------------|------------------------|------------------------|---------------------------|-------------------------|-----------------------------|-----------------------|------------------------|\n| Kielder | Event 1 | **0.008 / 0.957** | 0.015 / 0.920 | 0.140 / 0.269 | 0.016 / 0.918 | *0.013 / 0.933* | 0.091 / 0.527 | 0.031 / 0.837 | 0.137 / 0.286 | 0.123 / 0.361 |\n| Kielder | Event 2 | *0.027 / 0.877* | 0.029 / 0.869 | 0.087 / 0.610 | 0.017 / 0.700 | **0.015 / 0.934** | 0.081 / 0.637 | 0.047 / 0.788 | 0.068 / 0.692 | 0.059 / 0.735 |\n| Kielder | Event 3 | **0.013 / 0.691** | 0.015 / 0.634 | 0.040 / 0.280 | *0.019 / 0.668* | 0.021 / 0.629 | 0.021 / 0.621 | 0.023 / 0.618 | 0.040 / 0.284 | 0.042 / 0.260 |\n| Kielder | **Test** | **0.003 / 0.962** | 0.004 / 0.942 | 0.014 / 0.826 | *0.004 / 0.951* | 0.004 / 0.946 | 0.009 / 0.894 | 0.005 / 0.940 | 0.013 / 0.844 | 0.013 / 0.845 |\n\nAPILaNet is either the best model or extremely close to the best baseline on all splits, and attains the lowest error on the held-out test segment (MSE 0.003 vs 0.004, NSE 0.962 vs 0.951). We hope that this addresses the request for an additional real-data example.\n\n**PDE benchmarks (synthetic, cross-domain)**\n\nTo complement the hydrological case studies with a synthetic, physics-focused benchmark, we evaluate APILaNet on three standard 1D PDEs with known IC/BC: viscous Burgers (fluid), the wave equation (propagation), and Allen–Cahn (phase-field / thermodynamic). We generate reference solutions with a finite-difference solver and compare to several established PINN variants:\n\n| Model                        | Burgers MSE    | Wave MSE        | Allen–Cahn MSE |\n| ---------------------------- | -------------- | --------------- | -------------- |\n| Vanilla PINN   [1]          | 5.80 × 10⁻⁴    | 2.62 × 10⁻⁴     | 1.18 × 10⁰     |\n| PINN-w  [2]     | 2.91 × 10⁻³    | 2.89 × 10⁻³     | 1.04 × 10⁰     |\n| gPINN  [3]  | 1.29 × 10⁻⁴    | 1.62 × 10⁻⁴     | 1.32 × 10⁻¹    |\n| vPINN [4]| 1.45 × 10⁻³    | 8.91 × 10⁻⁴     | 1.04 × 10⁰     |\n| **APILaNet (ours)**          | **4.5 × 10⁻⁵** | **1.52 × 10⁻⁴** | **1.18 × 10⁻¹** |\n\nTraditional PINNs perform well in the fully-specified setting (known geometry, full IC/BC, interior collocation points). APILaNet matches or outperforms these methods at the sensor while operating in a strictly weaker information regime (no explicit geometry, no IC/BC traces, no interior collocation). This addresses the request for a synthetic example and clarifies that the framework is not tied to a single application domain.\n\n---\n\n[1] M. Raissi, P. Perdikaris, and G. E. Karniadakis, “Physics-informed Neural Networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,” Journal of Computational Physics, vol. 378, pp. 686–707, Feb. 2019. doi:10.1016/j.jcp.2018.10.045\n \n[2] Tim De Ryck, Siddhartha Mishra, and Roberto Molinaro. wpinns: Weak physics informed neu-\nral networks for approximating entropy solutions of hyperbolic conservation laws, 2022. URL\nhttps://arxiv.org/abs/2207.08483.\n \n[3] Jeremy Yu, Lu Lu, Xuhui Meng, and George Em Karniadakis. Gradient-enhanced physics-informed\nneural networks for forward and inverse pde problems. Computer Methods in Applied Mechanics\nand Engineering, 393:114823, April 2022. ISSN 0045-7825. doi: 10.1016/j.cma.2022.114823.\nURL http://dx.doi.org/10.1016/j.cma.2022.114823.\n \n[4] Ehsan Kharazmi, Zhongqiang Zhang, and George Em Karniadakis. Variational physics-informed\nneural networks for solving partial differential equations. arXiv preprint arXiv:1912.00873, 2019."},"title":{"value":"Response to Reviewer Zkv9 (Part 1/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763490577828,"tcdate":1763490548987,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20627/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission20627/Authors"],"forum":"VScnURO2g1","number":14,"license":"CC BY 4.0","cdate":1763490548987,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20627/-/Official_Comment"],"mdate":1763490577828,"domain":"ICLR.cc/2026/Conference","replyto":"8m4Hcto94R","id":"D4LN4fqONg","forumContent":{"TLDR":{"value":"A physics-informed latent network with adaptive weighting and a learned weak-form measure enables robust single-sensor forecasting under sparse sensing, outperforming SOTA models"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed learning","conservation laws","adaptive loss weighting","latent field","monotone neural mapping","time-series forecasting"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Forecasting conservation-governed dynamics is often constrained by sparse sensing: in practice, we may have only a single boundary sensor and noisy exogenous variables. In this work we design an Adaptive Physics-Informed Latent Network (APILaNet) that learns a latent field and enforces 1D-conservation of physics law in the weak form using a learned, normalized space--time measure. Normalization makes physics enforcement insensitive to quadrature resolution and concentrates it on transient violations. A monotone, Lipschitz measurement layer maps latent variables to observed targets, improving identifiability from a single sensor. An adaptive, bounded scheduler scales the physics and smoothness loss terms with meaningful representations, emphasizing conservation of physics laws during events while preserving training stability. Learning a space-time measure for weak-form enforcement, combined with a monotone mapping and adaptive scheduling, enables accurate, data-efficient single-sensor forecasting in physics-governed systems. We evaluate APILaNet through a synthetic and hydrological case study, APILaNet outperforms strong sequence baselines and reduces MSE during extreme events, while improving Nash--Sutcliffe efficiency. Code will be released upon acceptance."},"_bibtex":{"value":"@misc{\nkucia2026apilanet,\ntitle={{APIL}aNet: Adaptive Physics-Informed Latent Network for Single-Sensor Forecasting},\nauthor={Adrian Kucia and Edward Rollason and Wai Lok Woo},\nyear={2026},\nurl={https://openreview.net/forum?id=VScnURO2g1}\n}"},"title":{"value":"APILaNet: Adaptive Physics-Informed Latent Network for Single-Sensor Forecasting"},"pdf":{"value":"/pdf/8cb8df5af300f217d6fe04bca6fa0d4678477421.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kucia|apilanet_adaptive_physicsinformed_latent_network_for_singlesensor_forecasting"},"authorids":{"value":["~Adrian_Kucia1","~Edward_Rollason1","~Wai_Lok_Woo1"]},"authors":{"value":["Adrian Kucia","Edward Rollason","Wai Lok Woo"]}},"version":2},{"content":{"summary":{"value":"The paper expresses various learning setups through variational functionals. The authors follow a well-known formalism from physics (Lagrange functions/principle of least action). The examples include active learning, the Bellman optimality equation, and supervised learning."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1) Time evolution in physics is unique and so is the solution of the corresponding variational formulations. How is this uniqueness interpreted in a learning context? Is there a unique best learning algorithm? \n\n2) How can a formulation in terms of variational principles benefit machine learning research? In math/physics, it allows to incorporate symmetries more easily, using certain proof techniques, geometric invariance, ... Can you demonstrate any of these advantages?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1) The paper studies analogies to variational calculus in physics/math. This could allow to bring well-established techniques from other fields to machine learning research.\n\n2) The authors attempt to unify different learning paradigms, which could create a deeper understanding of learning mechanisms."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) It is difficult to read the paper.\n\n- Generally, it would help to introduce the symbols in a more mathematical way (discrete/continuous, dimension).\n- At the beginning (L209,L216), sample-efficiency and compute time are introduced as a characterization of efficient learning but the latter seems to play no role in the paper. Omitting needless concepts could make it easier to read.\n- The purpose of Section 2 and Insight 1 is not clear to me. It appears to be a side observation that is not relevant for the core topic of the paper. \n- The standard Lagrangian/Hamiltonian frameworks, as I know them, usually starts with a clear declaration of the variables and their conjugate variables. I would recommend this for each of the examples.\n- Insight 2: Planning usually refers to using a learned/known model of the environment for better action selection. \n\n2) The contribution to derive learning algorithms from Lagrangians is not as strong as claimed. \n\n- Section 3.1 appears to me to be specific to linear regression. \n- The formulation of the Bellman equation through a variational principle is (as the authors say) not novel. \n- It is claimed that the Adam optimizer is derived. However, it looks more like a derivation of maximum likelihood estimation in Section 3.3 and that Adam is one way of approximating this. I doubt that Adam with all its features, such as normalization of update sizes, can be derived from a simple Lagrangian like the one in Equation 26."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928105411,"tcdate":1761264336128,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18398/Reviewer_N1Zv"],"signatures":["ICLR.cc/2026/Conference/Submission18398/Reviewer_N1Zv"],"forum":"FUEzlNM4jx","number":1,"license":"CC BY 4.0","cdate":1761264336128,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18398/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928105411,"domain":"ICLR.cc/2026/Conference","replyto":"FUEzlNM4jx","id":"Vz34BiVzdc","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"A physics perspective in efficient learning"},"keywords":{"value":["physics; learning; reinforcement learning; generative models; learning theory"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"We study the problem of building an efficient learning system. Efficient learning processes information in the least time, i.e., building a system that reaches a desired error threshold with the least number of observations. Building upon least action principles from physics, we derive classic learning algorithms, Bellman's optimality equation in reinforcement learning, and the Adam optimizer in generative models from first principles, i.e., the Learning $\\textit{Lagrangian}$. We postulate that learning searches for stationary paths in the Lagrangian, and learning algorithms are derivable by seeking the stationary trajectories."},"_bibtex":{"value":"@misc{\nguo2026physics,\ntitle={Physics of Learning: A Lagrangian perspective to different learning paradigms},\nauthor={Siyuan Guo and Bernhard Sch{\\\"o}lkopf},\nyear={2026},\nurl={https://openreview.net/forum?id=FUEzlNM4jx}\n}"},"title":{"value":"Physics of Learning: A Lagrangian perspective to different learning paradigms"},"pdf":{"value":"/pdf/d62dd55055d859ddad558fe8884679035dad36b9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"guo|physics_of_learning_a_lagrangian_perspective_to_different_learning_paradigms"},"authorids":{"value":["~Siyuan_Guo1","~Bernhard_Schölkopf1"]},"authors":{"value":["Siyuan Guo","Bernhard Schölkopf"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PhysUniBench, a large-scale multimodal benchmark specifically designed to evaluate AI models on undergraduate-level physics problems. The benchmark contains 3,304 carefully curated physics questions spanning 8 major sub-disciplines (Classical Mechanics, Electromagnetism, Optics, Quantum Mechanics, Thermodynamics, Solid Physics, Relativity Physics, and Molecular/Atomic/Subatomic Physics). The benchmark has EN and CN language coverage and has MC as well as Open-Ended QA (OE) in terms of question format, stratified across 5 difficulty level and evaluated using a variety of frontier (M)LLM models."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"I've mostly put the questions inside the last section of weakness. Some additional ones\n\n1. For evaluated models: Even for MLLMs, it's mostly GPT models, also o3 and o4-mini are o-series not GPT-series, why isn't Claude/Gemini/Grok/other series as comprehensively tested as GPT/OpenAI models? It would probably be nice to test scaling and time evolution of many families of models, or different size models of the same open-sourced families? In general I think the evaluation and discussion section can be improved significantly;\n\n2. For the OE judge logs can you show some examples of model output against LLM judge verdict, so that we can see how exactly they perform in terms of judging the output? I fear rubric-based grading may not be very optimal given GPT-4o's limited capability.\n\n3. Maybe the authors can better present the novelty of this work other than \"multi-phase construction by advanced AI models\" because one could always just use stronger models to get a much more difficult benchmarks in that case (e.g. replace Qwen-VL with Gemini-2.5-Pro or GPT-5-High?) then the benchmark would be much less prone to saturation over time comparing to the current version?\n\n4. Also, can you better articulate how OE questions are judged? The current Figure 1 has some of them just \"Steps\" and then the answer is contained in the last step, but some others has a dedicated \"Answer\"? Are OE questions judged based on \"Answer\" entry and not on the intermediate process? If on the intermediate process maybe you could shed some light (e.g. by giving an example in the main text, to show this grading, which would imo strengthen this work a lot)\n\n5. As previously mentioned could you better organize the appendix with a Table of Contents? It takes a lot of time for reviewers to read thru >20 pages of appendix without a clear structure.\n\n6. Although this may not warrant an ethics review, how exactly have the authors complied with licensing/data protection laws in the process of curating this benchmark? This may warrant a clearer presentation in Section 3. Currently most of the descriptions are really vague and high-level only."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Multimodal:  It's nice that each question has exactly one diagram, (altho it also raises the question as to sometimes in the real use case of AI4Science the ability to interpret multiple images and cross-ref between them could be a valuable capability to probe)\n2. Difficulty Rating: The authors wrote that each problem is annotated with a fine-grained difficulty level ranging from 1 to 5, which then in turn shows a corresponding performance gradient in the evaluation phase (harder problems = lower scores) which is good for testing MLLMs on a difficulty gradient. (Note: I later realized the difficulty rating was also based on LLM performance, which may be considered circular reasoning and could be strengthened by e.g. human-AI cross-validation and other checks for construct validity)\n3. I also noticed that the stratified difficulty has roughly similar number of questions around 660, which is a nice property for statistical analysis in general when one wants to test the model performance across different difficulty level without having to worry about sample size as a confounder"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Disclaimer: I come from a mix of CS+Physics background and have at least graduate student level of physics knowledge, so I believe I am qualified to speak to the validity to an undergraduate level physics benchmark.\n\nAPS: American Physical Society (which is generally regarded as one of the most credible non-profit org on physics research)\n\n1. Trivial mistakes in the Physics discipline: I found some of the mistakes made in this paper to be quite trivial:\n\nFor example:\"Solid Physics\" is not the correct term in physics, but it should be \"Solid State Physics\" or \"Solid-State Physics\" \nI encourage everyone to simply do a Google Search of the word \"Solid Physics\" you will find everything is about \"Solid-State Physics\" because \"Solid Physics\" is not really a term we use in Physics.\n\n[Start quote from Wiki] Solid-state physics is the study of rigid matter, or solids, through methods such as solid-state chemistry, quantum mechanics, crystallography, electromagnetism, and metallurgy. It is the largest branch of condensed matter physics. [End quote]\n\n(This is not some proprietary knowledge, but a simple Wikipedia search would yield this explanation/definition as to what it means and what the correct term is)\n\n2. Therefore, I respectfully question if there were indeed \"expert curation\" going on by the expert of physics in constructing such a benchmark, how could these \"experts\" have not noticed such a basic mistake in the terminology? This gave me no choice but to question the validity of the claimed expert construction process, because it's hard to believe that if any of the annotation was actually done by experts, or even inspected/guided by experts, they would not notice this basic mistake.\n\n3. Very similarly to the point I raised above,\"Relativity Physics\" is also not really a term we often use in Physics, because the Theory of Relativity has both \"Special Relativity\"(which applies to all physical phenomena in the absence of gravity.) as well as \"General Relativity\" (which explains the law of gravitation and its relation to other laws of physics) where GR is often classified into Gravitation, Cosmology and Astrophysics (this term is stipulated by APS). So it's also quite confusing to me as to why the authors choose to specifically list \"Relativity Physics\" in parallel on the same level of e.g. \"Quantum Mechanics\" because in Physics they are definitely not classified to be on the same level. (You can e.g. check PhySH published by the APS)\n\n4. Apart from trivial mistakes in Physics, the coverage of this benchmark has a couple of flaws:\n\na. The domain distribution is quite uneven, where EM and CM constitutes nearly half and the rest are much smaller\n\nb. There are domains which are totally not covered like Astrophysics and Cosmology, Quantum Optics, Quantum Info etc. depending on the curriculum you choose to follow I can often see some of these classes being offered to undergrad in many university-level institutions across various countries, especially to physics major in their final year(s).\n\n5. Many of the details aren't super clear to me: For example, if we are trusting LLM with stratification and LLM-as-a-judge for OE questions, are there any validity check by experts to see if they actually do a good job? There should be a inter-annotator or human-AI cross-validation so as to see if there's any false positive/negatives and report the statistics accordingly\n\n6. Also, there lacks a human baseline by different level of human candidates, or at least undergraduates?\n\n7. For difficulty stratification: If you are relying on MLLM 16-rollout success rate to classify the questions, aren't you kind of circular reasoning here in terms of:\n\na. Use model accuracy to classify the questions from 0 to 5\n\nb. Then say ok this classification is good because models indeed perform correspondingly, but that's actually by design, right? since it is indeed the standard by which you partition the questions in the very first place, so this point in discussion seems like circular reasoning to me. Granted you could say the partition is done by a single model and the eval on many models, but the pretraining corpora of these models are likely overlapping a lot, so I don't think this classification makes much sense to me as to how it's designed.\n\n8. The related work coverage and comparison can be more comprehensive, some general benchmarks like ARB has a physics portion and more dedicated benchmarks like HiPhO and SeePhys and Multi-Physics should be more comprehensively compared in this paper\n\n9. Also, the evaluated models could be more comprehensive (only 2 text-based LLM, maybe for multimodal benchmark you don't really need that? Or if you want to compare LLM against MLLM you should have roughly similar number of models evaluated on each side?) and more thorough error analysis could significantly strengthen this work. \n\nTherefore, I believe this work has significant improvement to be done that could better contribute to the community at large, I encourage the authors to take some more time for improvement. And while we appreciate the author's good effort in constructing this benchmark, it is essential to remember for all of us that AI4Science benchmark (not just in physics, but any other domains as well) should have some input from domain experts in respective field to ensure their construct validity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915996715,"tcdate":1760510040282,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2029/Reviewer_ax6M"],"signatures":["ICLR.cc/2026/Conference/Submission2029/Reviewer_ax6M"],"forum":"TqgPgCvBnF","number":1,"license":"CC BY 4.0","cdate":1760510040282,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2029/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915996715,"domain":"ICLR.cc/2026/Conference","replyto":"TqgPgCvBnF","id":"p9G8qgvu1q","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics","reasoning","benchmark","large language model","multi-modal large language model"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Physics problem-solving is a challenging domain for large AI models, requiring integration of conceptual understanding, mathematical reasoning, and interpretation of physical diagrams. Existing evaluations fail to capture the full breadth and complexity of undergraduate physics, whereas this level provides a rigorous yet standardized testbed for pedagogically relevant assessment of multi-step physical reasoning. To this end, we present PhysUniBench, a large-scale multimodal benchmark designed to evaluate and improve the reasoning capabilities of multimodal large language models (MLLMs) specifically on undergraduate-level physics problems. PhysUniBench consists of 3,304 physics questions spanning 8 major sub-disciplines of physics, each accompanied by one visual diagrams. The benchmark includes both open-ended and multiple-choice questions, systematically curated and difficulty-rated through an iterative model-in-the-loop process. The benchmark's construction involved a rigorous multi-stage process, including multiple roll-outs, expert-level evaluation, automated filtering of easily solved problems, and a nuanced difficulty grading system with five levels. Through extensive experiments, we observe that current state-of-the-art models encounter substantial challenges in physics reasoning. For example, GPT-5 achieves only about 53.7% accuracy in the proposed PhysUniBench. These results highlight that current MLLMs struggle with advanced physics reasoning, especially on multi-step problems and those requiring precise diagram interpretation. By providing a broad and rigorous assessment tool, PhysUniBench aims to drive progress in AI for Science, encouraging the development of models with stronger physical reasoning, problem-solving skills, and multimodal understanding. The benchmark and evaluation scripts are available at https://anonymous.4open.science/r/PhysUniBenchmark-5784."},"_bibtex":{"value":"@misc{\nwang2026physunibench,\ntitle={PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level},\nauthor={Lintao Wang and Encheng Su and Jiaqi Liu and Pengze Li and Jiabei Xiao and Wenlong Zhang and Xi Chen and Yuan Meng and LEI BAI and Wanli Ouyang and SHIXIANG TANG and Aoran Wang and Xinzhu Ma},\nyear={2026},\nurl={https://openreview.net/forum?id=TqgPgCvBnF}\n}"},"title":{"value":"PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level"},"pdf":{"value":"/pdf/006ead177bfaf2dd134dbde4641e5de03d19d8e8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|physunibench_a_multimodal_physics_reasoning_benchmark_at_undergraduate_level"},"authorids":{"value":["~Lintao_Wang1","~Encheng_Su1","~Jiaqi_Liu7","~Pengze_Li3","~Jiabei_Xiao1","~Wenlong_Zhang3","~Xi_Chen20","~Yuan_Meng2","~LEI_BAI1","~Wanli_Ouyang1","~SHIXIANG_TANG1","~Aoran_Wang1","~Xinzhu_Ma1"]},"authors":{"value":["Lintao Wang","Encheng Su","Jiaqi Liu","Pengze Li","Jiabei Xiao","Wenlong Zhang","Xi Chen","Yuan Meng","LEI BAI","Wanli Ouyang","SHIXIANG TANG","Aoran Wang","Xinzhu Ma"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a method for training spatio-temporal Gaussian processes which incorporate physics constraints including satisfying the governing equations at a number of collocation points and satisfying curl / divergence free constraints. Their approach scales both linearly in time and space by leveraging a number of approximations."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- As I understand it, you can always choose to satisfy the differential equations at enough collocation points such that you overwhelm the data. For example, say I have very limited data so I choose the number of collocation points to be much greater than the number of data points. Clearly, if I use too many collocation points, my model will ignore the data and just follow the differential equations. How do you think about this perspective in the context of uncertainty quantification?\n- As a follow up to this previous point, what are the implications of assuming that $0_n^{(C)} = g(F_n) + \\epsilon_C$? Is it correct to think about this as implicitly assuming that the PDE is stochastic? \n- Are there situations where you believe the Gaussian process assumption could be limiting?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- While many of the individual components (such as incorporating curl / divergence free constraints and satisfying governing equations at a number of collocation points) have been done before, I believe the combination of ideas here is novel. \n- Showing how to achieve linear complexity in the space and time dimension is valuable.\n- Showing how your approach recovers some prior works as a special case from a more general perspective is also valuable.\n- The bevy of test cases demonstrating the efficacy of your approach on both synthetic and real-world problems convincingly demonstrate the advantages offered by your approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The problem statement is poorly motivated in my perspective. It would be helpful to be more specific about what you hope to achieve with hybrid modeling. For example, is your goal to improve the predictive accuracy of mechanistic models by incorporating data? What computational efficiency do you hope to improve? i.e. do you want to reduce the amount of data needed to train physics surrogates? Do you argue your approach is more efficient than solving the mechanistic model using traditional methods? etc. While I understand space is limited, being more specific about the problem statement will also help guide more targeted numerical studies in future works.\n- I think you need slightly more of a discussion on virtual observations (see questions below). When virtual observations are competing with actual data it would be helpful to discuss in detail how you think about $\\epsilon_C$ and choosing the number of collocation points.\n- Since you state that one of the goals of your approach is to be useful in uncertainty quantification, the numerical studies would have been strengthened by comparing more than just RMSE. For example, comparing the continuous ranked probability score (CRPS) would have given an indication of how well your approach is estimating uncertainty.\n- Minor:\n  - L94 space cases -> special cases?"},"limitations":{"value":"I think the authors have done a good job of identifying some potential limitations. I think the paper would greatly benefit from a more in depth discussion on using collocation points to enforce governing equation constraints and the limitations this perspective brings."}},"nonreaders":[],"tmdate":1730879559156,"tcdate":1720206372931,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission12520/Reviewer_BaEt"],"signatures":["NeurIPS.cc/2024/Conference/Submission12520/Reviewer_BaEt"],"forum":"tCf7S75xFa","number":1,"license":"CC BY 4.0","cdate":1720206372931,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission12520/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879559156,"domain":"NeurIPS.cc/2024/Conference","replyto":"tCf7S75xFa","id":"pb7aA6oYMH","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["gaussian processes","variational approximations","state space gaussian processes","physics informed gaussian processes"]},"supplementary_material":{"value":"/attachment/9461298fce1715a838e9758280749bfe546e0f0a.zip"},"primary_area":{"value":"probabilistic_methods"},"abstract":{"value":"Differential equations are important mechanistic models that are integral to many scientific and engineering applications. With the abundance of available data there has been a growing interest in data-driven physics-informed models. Gaussian processes (GPs) are particularly suited to this task as they can model complex, non-linear phenomena whilst incorporating prior knowledge and quantifying uncertainty. Current approaches have found some success but are limited as they either achieve poor computational scalings or focus only on the temporal setting. This work addresses these issues by introducing a variational spatio-temporal state-space GP that handles linear and non-linear physical constraints while achieving efficient linear-in-time computation costs. We demonstrate our methods in a range of synthetic and real-world settings and outperform the current state-of-the-art in both predictive and computational performance."},"_bibtex":{"value":"@inproceedings{\nhamelijnck2024physicsinformed,\ntitle={Physics-Informed Variational State-Space Gaussian Processes},\nauthor={Oliver Hamelijnck and Arno Solin and Theodoros Damoulas},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=tCf7S75xFa}\n}"},"title":{"value":"Physics-Informed Variational State-Space Gaussian Processes"},"pdf":{"value":"/pdf/4424cd43913ea9035300e356e962178e273d1e61.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"hamelijnck|physicsinformed_variational_statespace_gaussian_processes"},"authorids":{"value":["~Oliver_Hamelijnck1","~Arno_Solin1","~Theodoros_Damoulas1"]},"authors":{"value":["Oliver Hamelijnck","Arno Solin","Theodoros Damoulas"]}},"version":2},{"content":{"summary":{"value":"The paper titled **\"FEABench: Evaluating Language Models on Real World Physics Reasoning Ability\"** introduces a novel benchmark for assessing the capability of large language models (LLMs) and LLM-based agents to perform finite element analysis (FEA) tasks. FEA is essential in many engineering domains for solving complex physical problems using numerical solvers. The paper tests LLMs' ability to read problem descriptions, reason over them, and interact with FEA software (COMSOL Multiphysics). It presents a benchmark called FEABench Gold, consisting of 15 manually verified FEA problems, and evaluates the performance of various state-of-the-art LLMs, such as Claude-3.5, GPT-4, and Gemini-1.5, in solving these tasks. The best strategy involved LLM agents capable of generating executable API calls for COMSOL. However, even the best-performing models struggled to solve any of the benchmark problems completely and correctly."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- Exploring the impact of training data and fine-tuning strategies is crucial. Current LLMs may not have adequate coverage of FEA-related knowledge in their training data, and domain-specific fine-tuning or data augmentation could potentially yield significant improvements. Additionally, discussing the potential of hybrid approaches that combine LLMs with other AI techniques, such as symbolic reasoning systems or neural-symbolic methods, could provide valuable insights into addressing the observed limitations.\n\n- Could the authors clarify the main factors that contributed to the models’ failure to solve any benchmark problems completely? Was it primarily API interaction errors, poor physical reasoning, or something else?\n- Have any attempts been made to include human feedback in the agent loops (e.g., human-in-the-loop evaluation)? If so, how does that impact the success of the LLM agents?\n- Given the difficulty of the FEA tasks, what improvements to LLM architectures or agent designs do the authors suggest to improve their performance in future iterations?\n- Could the authors include specific examples from the FEABench dataset to enhance the clarity and understanding of the benchmark? Providing concrete problem examples, along with their corresponding model specifications and expected outputs, would offer valuable insight into the types of tasks the LLMs are expected to handle."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The paper presents a novel benchmark focused on a crucial real-world domain (FEA) that has been underexplored in LLM evaluation. By requiring end-to-end reasoning from natural language to executable code for complex physics simulations, FEABench pushes the boundaries of what is currently possible with LLMs.\n\n- The authors have taken care to ensure the problems are solvable, self-contained, and quantitatively verifiable. The evaluation framework is comprehensive, with multiple metrics capturing different aspects of solution quality.\n- The inclusion of concrete examples (e.g., problem descriptions, code snippets) aids understanding.\n\n- FEABench addresses an important gap in LLM evaluation by focusing on complex engineering problems that require both high-level reasoning and low-level code generation. Success on this benchmark could have major implications for automating engineering workflows. The multi-turn agent design also provides a valuable template for future work on LLM-based problem-solving systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Lack of Comparison with Human Performance**: Including a baseline comparison of human performance on these FEA tasks would significantly strengthen the evaluation. This comparison should involve both novice and expert COMSOL users to provide a spectrum of human capabilities.  Expert evaluation is particularly crucial to validate task difficulty,  establish performance benchmarks, assess solution quality and check benchmark integrity. \n\n- **Limited Success of LLMs**: Despite the interesting problem posed, none of the LLMs tested were able to solve the benchmark problems completely and correctly. A detailed examination of where and why LLMs fail could provide valuable insights, potentially revealing common patterns in errors across different models or problem types. It would be beneficial to discuss how current LLM architectures might be fundamentally limited for these tasks and whether specialized components for mathematical reasoning or code generation could improve performance. \n\n- **Lack of LLMs Candidates**: To provide a more comprehensive evaluation of current LLM capabilities, it would be beneficial to expand the range of models tested. This should include open-source model families such as LLaMA, Mistral, and Qwen. Testing multiple scales within each model family could reveal how model size affects performance on FEABench.  If possible, evaluating domain-specific models trained on scientific or engineering corpora would add further depth to the analysis. This broader evaluation would provide a more complete picture of the current state of LLM capabilities on FEA tasks and could uncover interesting differences in how various model architectures and training approaches handle these challenges.\n\n- **Complexity of Task**: The paper could explore more how different levels of task complexity (e.g., simpler geometry or fewer physics variables) influence LLM performance. Additionally, the role of visual information from GUI could be better integrated into the evaluation process.\n\n  \n\n- **Not So Good Presentation**: The paper's presentation and structure could be improved to enhance clarity and readability.\n\n  - As a benchmark paper, it fails to provide a clear, concise overview of the datasets' composition and size. The distinction between FEABench Gold and FEABench Large is not well explained, leaving readers confused about the exact number of problems in each dataset and the rationale behind this split. A clear statement of the dataset sizes (15 problems in FEABench Gold and 120 in FEABench Large) early in the paper, along with a brief explanation of why two separate datasets were created, would significantly improve understanding.\n\n  - The evaluation metrics section is another area that lacks clarity. Some metrics, such as the Code Similarity Score, are introduced without sufficient justification or explanation of their relevance. The paper acknowledges that \"two different code blocks could generate equivalent model subtrees,\" which raises questions about the utility of this metric in evaluating solution quality. The authors should either provide a stronger rationale for including this metric or consider removing it if it doesn't contribute meaningfully to the evaluation. Furthermore, the paper fails to effectively communicate the motivation behind introducing domain-specific metrics. While these metrics may be crucial for evaluating FEA solutions, the paper doesn't clearly articulate why existing general code evaluation metrics are insufficient and how the new metrics address these shortcomings. A more detailed explanation of how these domain-specific metrics capture important aspects of FEA problem-solving that general metrics miss would strengthen the paper's contribution."}},"nonreaders":[],"tmdate":1731428346137,"tcdate":1729413727994,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12171/Reviewer_Yw8s"],"signatures":["ICLR.cc/2025/Conference/Submission12171/Reviewer_Yw8s"],"forum":"hDkLpu1E64","number":2,"license":"CC BY 4.0","cdate":1729413727994,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12171/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428346137,"domain":"ICLR.cc/2025/Conference","replyto":"hDkLpu1E64","id":"AF022L6GU4","forumContent":{"TLDR":{"value":"How well can LLMs leverage FEA software to simulate and solve problems that require numerical analysis?"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["numerical analysis","finite element","benchmark","agents"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building precise simulations of the real world and invoking numerical solvers to answer quantitative problems is an essential requirement in engineering and science. We present FEABench, a benchmark to evaluate the ability of large language models (LLMs) and LLM agents to simulate and solve physics, mathematics and engineering problems using finite element analysis (FEA). We introduce a multipronged evaluation scheme to investigate the ability of LLMs to solve these problems by reasoning over natural language problem descriptions and operating COMSOL Multiphysics$^\\textregistered$, an FEA software, to compute the answers. In addition to testing state-of-the art-LLMs, we further design a language model agent equipped with the ability to interact with the software through its Application Programming Interface (API), examine its outputs and use tools to improve its solutions over multiple iterations. Our best performing strategy generates executable API calls 88\\% of the time. However, this benchmark still proves to be challenging enough that the LLMs and agents we tested were not able to completely and correctly solve any problem. LLMs that can successfully interact with and operate FEA software to solve problems such as those in our benchmark would significantly push the frontiers of their utility. Acquiring this capability would augment LLMs' reasoning skills with the precision of numerical solvers and advance the development of autonomous systems that can tackle complex problems in the real world."},"_bibtex":{"value":"@misc{\nmudur2025feabench,\ntitle={{FEAB}ench: Evaluating Language Models on Real World Physics Reasoning Ability},\nauthor={Nayantara Mudur and Hao Cui and Subhashini Venugopalan and Paul Raccuglia and Michael Brenner and Peter Christian Norgaard},\nyear={2025},\nurl={https://openreview.net/forum?id=hDkLpu1E64}\n}"},"title":{"value":"FEABench: Evaluating Language Models on Real World Physics Reasoning Ability"},"pdf":{"value":"/pdf/3e64111fb86b7cbb5ef6469de0f077b416722ed3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"mudur|feabench_evaluating_language_models_on_real_world_physics_reasoning_ability"},"authorids":{"value":["~Nayantara_Mudur1","~Hao_Cui3","~Subhashini_Venugopalan2","~Paul_Raccuglia1","~Michael_Brenner1","~Peter_Christian_Norgaard1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Nayantara Mudur","Hao Cui","Subhashini Venugopalan","Paul Raccuglia","Michael Brenner","Peter Christian Norgaard"]}},"version":2},{"content":{"venue":{"value":"Physics of Fluids"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"shaotong|solving_the_one_dimensional_vertical_suspended_sediment_mixing_equation_with_arbitrary_eddy_diffusivity_profiles_using_temporal_normalized_physicsinformed_neural_networks"},"html":{"value":"https://pubs.aip.org/aip/pof/article-abstract/36/1/017132/3105970/Solving-the-one-dimensional-vertical-suspended?redirectedFrom=fulltext"},"abstract":{"value":"Analytical solutions are practical tools in ocean engineering, but their derivation is often constrained by the complexities of the real world. This underscores the necessity for alternative approaches. In this study, the potential of Physics-Informed Neural Networks (PINN) for solving the one-dimensional vertical suspended sediment mixing (settling-diffusion) equation which involves simplified and arbitrary vertical Ds profiles is explored. A new approach of temporal Normalized Physics-Informed Neural Networks (T-NPINN), which normalizes the time component is proposed, and it achieves a remarkable accuracy (Mean Square Error of 10^-5 and Relative Error Loss of 10^-4⁠). T-NPINN also proves its ability to handle the challenges posed by long-duration spatiotemporal models, which is a formidable task for conventional PINN methods. In addition, the T-NPINN is free of the limitations of numerical methods, e.g., the susceptibility to inaccuracies stemming from the discretization and approximations intrinsic to their algorithms, particularly evident within intricate and dynamic oceanic environments. The demonstrated accuracy and versatility of T-NPINN make it a compelling complement to numerical techniques, effectively bridging the gap between analytical and numerical approaches and enriching the toolkit available for oceanic research and engineering."},"title":{"value":"Solving the one dimensional vertical suspended sediment mixing equation with arbitrary eddy diffusivity profiles using temporal normalized physics-informed neural networks"},"authors":{"value":[{"fullname":"Zhang Shaotong"},{"fullname":"Deng Jiaxin"},{"fullname":"Li Xi-An"},{"fullname":"Zhao Zixi"},{"fullname":"Wu Jinran","username":"~Wu_Jinran1"},{"fullname":"Li Weide"},{"fullname":"Wang You-Gan"},{"fullname":"Jeng Dong-Sheng"}]}},"tmdate":1789092631224,"pdate":1706054400000,"externalIds":["doi:10.1063/5.0179223"],"tcdate":1769155442122,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Jinran_Wu1"],"forum":"BRMshAZ4ln","license":"CC BY-SA 4.0","number":36584,"cdate":1733705566844,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789092631224,"domain":"OpenReview.net/Public_Article","id":"BRMshAZ4ln","version":2},{"content":{"summary":{"value":"The paper introduces the Ensemble Inverse Problem (EIP) framework, in which the goal is to infer an unknown variable from a measurement without explicit access to the forward model or prior distribution. The authors describe two settings, EIP-I (recover prior) and EIP-II (recover posterior). The paper focuses on EIP-II and proposes ensemble inverse generative models (EI-DDPM and EI-FM) that condition posterior sampling on both a single observation and an observation ensemble. To extract information from the observation set, the method uses a permutation-invariant set encoder, such as Deep Sets or the Set Transformer. Experiments include 2-D Gaussian toy distributions, high-energy physics unfolding, and an MNIST image inversion benchmark. The authors claim improved generalization to unseen priors and avoidance of iterative inference."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- Can you provide some concrete, real-world settings where the forward operator is unknown, but large paired datasets exist for training?\n- How does performance scale with dimensionality beyond MNIST? Are there plans to test the proposed algorithms on natural images or scientific imaging?\n- What specific features of the set encoder contribute to generalization to unseen priors beyond simple moment-based approaches?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- The paper introduces a clear formalization of EIP and distinguishes EIP-I and EIP-II, clarifying the problem space and highlighting a setting where the forward operator is unknown. \n- The method uses established conditional generative models (DDPM and flow matching) and extends them to incorporate set-based ensemble conditioning. \n- A permutation-invariant representation is included to encode the ensemble of observations, aligning with prior set-modeling work. \n- Experiments across multiple synthetic and real-world tasks demonstrate feasibility (2-D Gaussian, HEP unfolding, MNIST inversion)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The method is largely a conditional generative model with an added set encoder. The paper frames this as a new problem, but the algorithmic contribution is incremental, in my opinion.\n- The practical relevance of EIP is not strongly justified. Real-world cases where forward operators are truly unavailable yet training observations exist remain unclear. \n- Toy Gaussian and MNIST blur/noise settings do not convincingly demonstrate scalability or relevance for complex inversion tasks. \n- No comparison to stronger modern inference baselines (e.g., existing diffusion-based inverse samplers outside EIP literature).\n- Duplicating samples when $N’<N$ is a rather ad-hoc approach. \n- Claims of generalization to unseen priors are qualitative and lack theoretical justification.\n- No ablations isolating the benefit of set conditioning versus simply conditioning on summary statistics (moments)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931319596,"tcdate":1761891240717,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19400/Reviewer_2G23"],"signatures":["ICLR.cc/2026/Conference/Submission19400/Reviewer_2G23"],"forum":"aVXXZAp41g","number":1,"license":"CC BY 4.0","cdate":1761891240717,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19400/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931319596,"domain":"ICLR.cc/2026/Conference","replyto":"aVXXZAp41g","id":"xiZsmvNdXI","forumContent":{"TLDR":{"value":"We introduce the ensemble inverse problem and propose a posterior sampling method based on generative models to solve it."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Inverse problems","conditional generative models","posterior sampling","permutation invariant neural network"]},"supplementary_material":{"value":"/attachment/c25721b5053dc87abe458c387415015bf452f63c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We introduce a new multivariate statistical problem that we refer to as the Ensemble Inverse Problem (EIP). The aim of EIP is to invert for an ensemble that is distributed according to the pushforward of a prior under a forward process. In high energy physics (HEP), this is related to a widely known problem called unfolding, which aims to reconstruct the true physics distribution of quantities, such as momentum and angle, from measurements that are distorted by detector effects. In recent applications, the EIP also arises in inverse imaging with unknown priors. We propose non-iterative inference-time methods that construct posterior samplers based on a new class of conditional generative models, which we call  ensemble inverse generative models. For the posterior modeling, these models additionally use the ensemble information contained in the observation set on top of single measurements.  Unlike existing methods, our proposed methods avoid explicit and iterative use of the forward operator at inference time via training across several sets of truth-observation pairs that are consistent with the same forward operator, but originate from a wide range of priors. We demonstrate that this training procedure implicitly encodes the likelihood model. The use of ensemble information helps posterior inference and enables generalization to unseen priors. We benchmark the proposed method on several synthetic and real datasets in HEP and inverse imaging."},"_bibtex":{"value":"@misc{\nhuan2026the,\ntitle={The Ensemble Inverse Problem: Applications and Methods},\nauthor={Zhengyan Huan and Camila Pazos and Martin Klassen and Vincent Croft and Pierre-Hugues Beauchemin and Shuchin Aeron},\nyear={2026},\nurl={https://openreview.net/forum?id=aVXXZAp41g}\n}"},"title":{"value":"The Ensemble Inverse Problem: Applications and Methods"},"pdf":{"value":"/pdf/7efc21d8ad2fc1796d9ea8a75ce49b47e338d44b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"huan|the_ensemble_inverse_problem_applications_and_methods"},"authorids":{"value":["~Zhengyan_Huan2","~Camila_Pazos1","~Martin_Klassen1","~Vincent_Croft1","~Pierre-Hugues_Beauchemin1","~Shuchin_Aeron2"]},"authors":{"value":["Zhengyan Huan","Camila Pazos","Martin Klassen","Vincent Croft","Pierre-Hugues Beauchemin","Shuchin Aeron"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2510.11878v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"pechko|gsverse_meshbased_gaussian_splatting_for_physicsaware_interaction_in_virtual_reality"},"authorids":{"value":["","~Piotr_Borycki1","","","","",""]},"html":{"value":"https://doi.org/10.48550/arXiv.2510.11878"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2510-11878,\n  publtype={informal},\n  author={Anastasiya Pechko and Piotr Borycki and Joanna Waczynska and Daniel Barczyk and Agata Szymanska and Slawomir Konrad Tadeja and Przemyslaw Spurek},\n  title={GS-Verse: Mesh-based Gaussian Splatting for Physics-aware Interaction in Virtual Reality},\n  year={2025},\n  month={October},\n  cdate={1759276800000},\n  journal={CoRR},\n  volume={abs/2510.11878},\n  url={https://doi.org/10.48550/arXiv.2510.11878}\n}\n"},"abstract":{"value":"As the demand for immersive 3D content grows, the need for intuitive and efficient interaction methods becomes paramount. Current techniques for physically manipulating 3D content within Virtual Reality (VR) often face significant limitations, including reliance on engineering-intensive processes and simplified geometric representations, such as tetrahedral cages, which can compromise visual fidelity and physical accuracy. In this paper, we introduce GS-Verse (Gaussian Splatting for Virtual Environment Rendering and Scene Editing), a novel method designed to overcome these challenges by directly integrating an object's mesh with a Gaussian Splatting (GS) representation. Our approach enables more precise surface approximation, leading to highly realistic deformations and interactions. By leveraging existing 3D mesh assets, GS-Verse facilitates seamless content reuse and simplifies the development workflow. Moreover, our system is designed to be physics-engine-agnostic, granting developers robust deployment flexibility. This versatile architecture delivers a highly realistic, adaptable, and intuitive approach to interactive 3D manipulation. We rigorously validate our method against the current state-of-the-art technique that couples VR with GS in a comparative user study involving 18 participants. Specifically, we demonstrate that our approach is statistically significantly better for physics-aware stretching manipulation and is also more consistent in other physics-based manipulations like twisting and shaking. Further evaluation across various interactions and scenes confirms that our method consistently delivers high and reliable performance, showing its potential as a plausible alternative to existing methods."},"title":{"value":"GS-Verse: Mesh-based Gaussian Splatting for Physics-aware Interaction in Virtual Reality"},"authors":{"value":["Anastasiya Pechko","Piotr Borycki","Joanna Waczynska","Daniel Barczyk","Agata Szymanska","Slawomir Konrad Tadeja","Przemyslaw Spurek"]}},"tmdate":1772139123349,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2510-11878"],"tcdate":1772139115367,"writers":["~"],"signatures":["~Piotr_Borycki1"],"forum":"5EqZCrKHMT","license":"CC BY-SA 4.0","number":828965,"cdate":1759276800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772139123349,"domain":"DBLP.org","id":"5EqZCrKHMT","version":2},{"content":{"summary":{"value":"This paper proposes SyNC, a new training paradigm for CLIP-based few-shot learning that aims to balance the fidelity and diversity of synthetic data representations. The method introduces two complementary losses:  \n(1) an ETF-based Neural Collapse loss to align real and synthetic features toward theoretically optimal geometric prototypes, thereby improving fidelity, and  \n(2) a regional supervised contrastive loss that enhances diversity by pushing apart misclassified synthetic features.  \nThe framework integrates these components with LoRA fine-tuning on both real and generated samples. Experiments on ten benchmark datasets show competitive or superior performance, especially on fine-grained datasets such as FGVC-Aircraft and Stanford Cars."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to Weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper is conceptually clear and easy to understand, providing a well-motivated formulation of the fidelity–diversity trade-off in synthetic-data-based few-shot learning.  \n- It accurately identifies the core limitation of prior works—overemphasis on either fidelity or diversity—and proposes a theoretically grounded direction to address both simultaneously. \n- The integration of Neural Collapse theory into few-shot CLIP adaptation is novel and insightful, offering a principled geometric interpretation for feature alignment."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the paper is highly intuitive, the experimental evidence does not convincingly validate its hypotheses:  \n  1. The presentation of results in the tables can be **misleading**, as only the proposed method’s results are boldfaced even when other methods achieve identical result (see Tables 1, 3, and 4).  \n  2. In Table 1, SyNC’s results are not significantly better than ImagineFSL, which contradicts the claim in Lines 59–62 about \"enhance diversity often pushes synthetic images further from the real data distribution\"\n  3. The ablation in Table 3 shows that the two proposed losses individually contribute marginal improvements on several datasets (e.g., IN, CAL, AirC, Pets, SUN, FLO). Furthermore, the paper’s justification in Section 3.1—“why not initialize ETF in a way better aligned with the input feature distribution?”—may be questionable, as the employed Kabsch algorithm merely aligns two point sets via least-squares rotation, functionally similar to random initialization followed by L2 fitting. This could explain the limited performance gains; thus, an error analysis would be necessary to substantiate the claimed benefits.  \n  4. The method involves too many hyperparameters (λ, λ₁, λ₂, ρ, β, τ), and the search ranges are inconsistent. For instance, the recommended range for λ₁ is [0, 1], yet in FGVC-Aircraft it is set to 15, far beyond the provided interval, suggesting a lack of systematic tuning."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918194105,"tcdate":1761842380428,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5689/Reviewer_qjUR"],"signatures":["ICLR.cc/2026/Conference/Submission5689/Reviewer_qjUR"],"forum":"nsUVWQiboS","number":4,"license":"CC BY 4.0","cdate":1761842380428,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5689/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918194105,"domain":"ICLR.cc/2026/Conference","replyto":"nsUVWQiboS","id":"x4qDRefkD3","forumContent":{"TLDR":{"value":"We propose a novel training algorithm for few-shot learning with synthetic data, balancing the model representation fidelity and diversity."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["few-shot learning","synthetic data","neural collapse"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"In few-shot learning, augmenting real data with synthesized images from text-to-image diffusion models has emerged as a promising direction. Although numerous studies have been proposed to improve the performance of this training framework, they often fail to adequately address the critical trade-off between fidelity and diversity when training with synthetic data. In this work, we propose SyNC, a novel training paradigm that explicitly balances these characteristics in the feature space through two complementary mechanisms. First, we leverage an optimal geometric prototype structure built upon the Neural Collapse phenomenon to increase fidelity, guiding the representations of both real and synthetic data toward their corresponding equiangular tight frame (ETF) prototypes. Second, we introduce an innovative regional contrastive loss function specifically designed to enhance diversity by improving the distinction between misclassified synthetic data features, thereby encouraging more varied and robust representations. Extensive experimental results demonstrate the effectiveness of our proposed method, which outperforms state-of-the-art approaches on average across few-shot image classification benchmarks and shows significant improvements on fine-grained datasets. Further analysis demonstrates that our method achieves a more favorable balance between representation fidelity and diversity, revealing a correlation between these factors and overall model performance."},"_bibtex":{"value":"@misc{\npham2026sync,\ntitle={Sy{NC}: Balancing Fidelity and Diversity of Synthetic Data Representations in {CLIP}-based Few-Shot Learning via Neural Collapse},\nauthor={Thanh Duc Pham and Lan-Cuong Nguyen and Dung D. Le and Minh-Tan Pham},\nyear={2026},\nurl={https://openreview.net/forum?id=nsUVWQiboS}\n}"},"title":{"value":"SyNC: Balancing Fidelity and Diversity of Synthetic Data Representations in CLIP-based Few-Shot Learning via Neural Collapse"},"pdf":{"value":"/pdf/20148c8b5902203177588e8c7a70df2c167fa4f9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"pham|sync_balancing_fidelity_and_diversity_of_synthetic_data_representations_in_clipbased_fewshot_learning_via_neural_collapse"},"authorids":{"value":["~Thanh_Duc_Pham1","~Lan-Cuong_Nguyen1","~Dung_D._Le2","~Minh-Tan_Pham2"]},"authors":{"value":["Thanh Duc Pham","Lan-Cuong Nguyen","Dung D. Le","Minh-Tan Pham"]}},"version":2},{"content":{"summary":{"value":"This paper proposes ARB, a new dataset for evaluating LLM reasoning in expert domains such as mathematics, physics, chemistry, biology and law. Though there is technical contribution, this is still an important contribution to the community due to the lack of good benchmarks."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"This paper proposes a new dataset ARB covering a wide range of domains for reasoning.\nThis paper evaluates 3 common models (ChatGPT, GPT-4 and Claude) on ARB. The authors also provide a breakdown of the error cases in GPT-4, which provide insights for future directions.\nThis paper proposes model-based rubric evaluation. The authors rigorously verify this approach by comparing the grading of GPT-4 and humans, which shows a moderately high correlation between them. This may be used as an evaluation tool for this dataset in the future."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"It’s not intuitive what questions are covered by ARB. Can you put a few examples in the paper?\nIt’s not clear what the position of ARB is compared to existing benchmarks. The authors claim ARB is more difficult. However, there is no strong supporing evidence except for the weak performance of models on the physics and math portions of ARB"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Section 3.1 “aspirational” -> Is it semantically correct here?\nFigure 1. Can you use a higher resolution or pdf instead? The y-axis has a wrong unit."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636512373,"tcdate":1698958204923,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5168/Reviewer_c9vS"],"signatures":["ICLR.cc/2024/Conference/Submission5168/Reviewer_c9vS"],"forum":"gsZAtAdzkY","number":2,"license":"CC BY 4.0","cdate":1698958204923,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5168/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636512373,"domain":"ICLR.cc/2024/Conference","replyto":"gsZAtAdzkY","id":"PnCElHYYdB","forumContent":{"TLDR":{"value":"Benchmark featuring questions in math, physics, biology, chemistry and law. We evaluate on recent LLMs, classify errors, propose a rubric-based auto-evaulation method"},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["benchmark","academic","mathematics","physics","LLM"]},"supplementary_material":{"value":"/attachment/97a01d29b3709dd8c6072185431a568ec2613997.pdf"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Large Language Models (LLMs) have demonstrated remarkable performance on various quantitative reasoning and knowledge benchmarks. However, many of these benchmarks are losing utility as LLMs get increasingly high scores, despite not yet reaching expert performance in these domains. We introduce ARB, a novel benchmark composed of advanced reasoning problems in multiple fields. ARB presents a more challenging test than prior benchmarks, featuring problems in mathematics, physics, biology, chemistry, and law. As a subset of ARB, we introduce a challenging set of math and physics problems which require advanced symbolic reasoning and domain knowledge. We evaluate recent models such as GPT-4 and Claude on ARB and demonstrate that current models score well below 50% on more demanding tasks. In order to improve both automatic and assisted evaluation capabilities, we introduce a rubric-based evaluation approach, allowing GPT-4 to score its own intermediate reasoning steps. Further, we conduct a human evaluation of the symbolic subset of ARB, finding promising agreement between annotators and GPT-4 rubric evaluation score."},"_bibtex":{"value":"@misc{\nsawada2024arb,\ntitle={{ARB}: Advanced Reasoning Benchmark for Large Language Models},\nauthor={Tomohiro Sawada and Daniel Paleka and Alexander Havrilla and Pranav Tadepalli and Paula Vidas and Alexander Perikles Kranias and John J Nay and Kshitij Gupta and Aran Komatsuzaki},\nyear={2024},\nurl={https://openreview.net/forum?id=gsZAtAdzkY}\n}"},"title":{"value":"ARB: Advanced Reasoning Benchmark for Large Language Models"},"pdf":{"value":"/pdf/243ad5c60e680cca1b67025762dad9cfb483cf58.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"sawada|arb_advanced_reasoning_benchmark_for_large_language_models"},"authorids":{"value":["~Tomohiro_Sawada1","~Daniel_Paleka1","~Alexander_Havrilla2","~Pranav_Tadepalli1","~Paula_Vidas1","~Alexander_Perikles_Kranias1","~John_J_Nay1","~Kshitij_Gupta1","~Aran_Komatsuzaki1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Tomohiro Sawada","Daniel Paleka","Alexander Havrilla","Pranav Tadepalli","Paula Vidas","Alexander Perikles Kranias","John J Nay","Kshitij Gupta","Aran Komatsuzaki"]}},"version":2},{"content":{"summary":{"value":"The paper presents a novel computational pipeline aimed at aligning the communication space of Multi-Agent Reinforcement Learning (MARL) agents with an embedding space of human natural language. The authors propose grounding agent communications on synthetic data generated by embodied Large Language Models (LLMs) in interactive teamwork scenarios."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weakness"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The use of synthetic data generated by LLMs to align MARL agents' communication is a creative application of existing technologies in a new way, demonstrating originality in methodology.\n- While the paper does not present new theoretical results, it does provide a solid empirical foundation for its approach, which is well-supported by experiments.\n- The proposed computational pipeline appears technically sound, with a clear explanation of how it aligns with human language and the rationale behind the design choices.\n- The authors provide clear explanations of complex concepts, such as the alignment of communication spaces and the grounding process, making the paper accessible to a broader audience."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper does not provide a theoretical framework or proofs to support the empirical findings. Developing a theoretical basis could strengthen the claims and provide deeper insights into why the approach works.\n- The experiments are conducted in controlled environments. To strengthen the claims, testing the approach in more diverse and complex scenarios could provide evidence of broader applicability.\n- The paper relies heavily on synthetic data generated by LLMs. There might be concerns about the representativeness of this data for real-world scenarios.\n- Some details regarding the implementation of the MARL agents and the interaction with LLMs could be better elaborated."},"limitations":{"value":"See weakness"}},"nonreaders":[],"tmdate":1730879500144,"tcdate":1720973485899,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission11830/Reviewer_hS12"],"signatures":["NeurIPS.cc/2024/Conference/Submission11830/Reviewer_hS12"],"forum":"DUHX779C5q","number":3,"license":"CC BY 4.0","cdate":1720973485899,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission11830/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879500144,"domain":"NeurIPS.cc/2024/Conference","replyto":"DUHX779C5q","id":"eLDITyRMoZ","forumContent":{"TLDR":{"value":"We propose a novel computational pipeline to ground MARL communication in human language using embodied LLM agents, enabling interpretable and generalizable communication in ad-hoc multi-agent teamwork."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Multi-Agent Reinforcement Learning","Emergent Communication","Ad-hoc Teamwork","Large Language Models"]},"supplementary_material":{"value":"/attachment/649adc931719a6db5c7fa89e2221b511fd2cb1da.zip"},"primary_area":{"value":"human-AI_interaction"},"abstract":{"value":"Multi-Agent Reinforcement Learning (MARL) methods have shown promise in enabling agents to learn a shared communication protocol from scratch and accomplish challenging team tasks. However, the learned language is usually not interpretable to humans or other agents not co-trained together, limiting its applicability in ad-hoc teamwork scenarios. In this work, we propose a novel computational pipeline that aligns the communication space between MARL agents with an embedding space of human natural language by grounding agent communications on synthetic data generated by embodied Large Language Models (LLMs) in interactive teamwork scenarios. Our results demonstrate that introducing language grounding not only maintains task performance but also accelerates the emergence of communication. Furthermore, the learned communication protocols exhibit zero-shot generalization capabilities in ad-hoc teamwork scenarios with unseen teammates and novel task states. This work presents a significant step toward enabling effective communication and collaboration between artificial agents and humans in real-world teamwork settings."},"_bibtex":{"value":"@inproceedings{\nli2024language,\ntitle={Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication},\nauthor={Huao Li and Hossein Nourkhiz Mahjoub and Behdad Chalaki and Vaishnav Tadiparthi and Kwonjoon Lee and Ehsan Moradi Pari and Charles Michael Lewis and Katia P. Sycara},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=DUHX779C5q}\n}"},"title":{"value":"Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication"},"pdf":{"value":"/pdf/6664b584d62b376a0ad3f0a353ad58f8ef896b2e.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"li|language_grounded_multiagent_reinforcement_learning_with_humaninterpretable_communication"},"authorids":{"value":["~Huao_Li1","~Hossein_Nourkhiz_Mahjoub1","~Behdad_Chalaki1","~Vaishnav_Tadiparthi1","~Kwonjoon_Lee1","~Ehsan_Moradi_Pari1","~Charles_Michael_Lewis1","~Katia_P._Sycara1"]},"authors":{"value":["Huao Li","Hossein Nourkhiz Mahjoub","Behdad Chalaki","Vaishnav Tadiparthi","Kwonjoon Lee","Ehsan Moradi Pari","Charles Michael Lewis","Katia P. Sycara"]}},"version":2},{"content":{"summary":{"value":"The proposed Depth Pro model employs a ViT architecture for zero-shot metric monocular depth estimation, targeting applications such as novel view synthesis. By employing patch-based, multi-scale processing, Depth Pro achieves high-resolution depth predictions while preserving real-time performance and edge sharpness. This paper introduces a two-stage training approach that integrates synthetic and real-world datasets, enhancing depth boundary accuracy. Furthermore, the model includes boundary evaluation metrics and a focal length estimation component, improving its robustness in handling images lacking metadata."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Structured Training Curriculum: The curriculum is specifically tailored to handle the unique strengths and weaknesses of synthetic and real-world datasets, resulting in a model that generalizes effectively while producing fine boundary details in depth maps.\n\n2. Enhanced Boundary Evaluation Metrics: New metrics for evaluating depth boundaries address a gap in existing benchmarks by focusing on boundary precision, which is critical for applications like view synthesis that demand fine details."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Dependence on pretrained ViT with Limited Architectural Innovation: Depth Pro benefits from pretrained ViT backbones, yet its architecture primarily builds on existing elements rather than introducing fundamentally new mechanisms for depth estimation, which limits its architectural novelty.\n\n2. Heavy Reliance on Synthetic Data for Boundary-Sensitive Training: The model’s second-stage training emphasizes synthetic datasets for boundary sharpness. This reliance on synthetic data could impact generalization to real-world environments, particularly in complex or unstructured settings where boundaries are less distinct."}},"nonreaders":[],"tmdate":1731427410114,"tcdate":1730538173568,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1325/Reviewer_a8Yt"],"signatures":["ICLR.cc/2025/Conference/Submission1325/Reviewer_a8Yt"],"forum":"aueXfY0Clv","number":3,"license":"CC BY 4.0","cdate":1730538173568,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1325/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427410114,"domain":"ICLR.cc/2025/Conference","replyto":"aueXfY0Clv","id":"dRIi6888Um","forumContent":{"TLDR":{"value":"A foundation model for zero-shot metric monocular depth estimation that synthesizes high-resolution depth maps with unparalleled sharpness and high-frequency detail"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["depth estimation","computer vision"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present a foundation model for zero-shot metric monocular depth estimation. Our model, Depth Pro, synthesizes high-resolution depth maps with unparalleled sharpness and high-frequency details. The predictions are metric, with absolute scale, without relying on the availability of metadata such as camera intrinsics. And the model is fast, producing a 2.25-megapixel depth map in 0.3 seconds on a standard GPU. These characteristics are enabled by a number of technical contributions, including an efficient multi-scale vision transformer for dense prediction, a training protocol that combines real and synthetic datasets to achieve high metric accuracy alongside fine boundary tracing, dedicated evaluation metrics for boundary accuracy in estimated depth maps, and state-of-the-art focal length estimation from a single image. Extensive experiments analyze specific design choices and demonstrate that Depth Pro outperforms prior work along multiple dimensions. We release code & weights at https://github.com/apple/ml-depth-pro"},"_bibtex":{"value":"@inproceedings{\nbochkovskiy2025depth,\ntitle={Depth Pro: Sharp Monocular Metric Depth in Less Than a Second},\nauthor={Alexey Bochkovskiy and Ama{\\\"e}l Delaunoy and Hugo Germain and Marcel Santos and Yichao Zhou and Stephan Richter and Vladlen Koltun},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=aueXfY0Clv}\n}"},"title":{"value":"Depth Pro: Sharp Monocular Metric Depth in Less Than a Second"},"pdf":{"value":"/pdf/2313c69543c424d5dd1279e35cdb1ab0e8fd22dc.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"bochkovskiy|depth_pro_sharp_monocular_metric_depth_in_less_than_a_second"},"authorids":{"value":["~Alexey_Bochkovskiy1","~Amaël_Delaunoy1","~Hugo_Germain2","~Marcel_Santos1","~Yichao_Zhou1","~Stephan_R._Richter1","~Vladlen_Koltun1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Alexey Bochkovskiy","Amaël Delaunoy","Hugo Germain","Marcel Santos","Yichao Zhou","Stephan Richter","Vladlen Koltun"]}},"version":2},{"content":{"summary":{"value":"This paper introduces **DiffWind**, a physics-informed generative framework that reconstructs and simulates **hidden wind fields** from **multi-view videos** of wind-driven objects.  \nThe core idea is to jointly optimize the **latent wind field** and the **object motion** under physical constraints.  \nDiffWind represents wind using an **Eulerian grid** and objects using **Lagrangian particles**, coupled via the **Material Point Method (MPM)** for differentiable wind–object interaction.  \nThe method further enforces **fluid dynamics consistency** through a loss derived from the **Lattice Boltzmann Method (LBM)**, ensuring physically plausible flow.  \n\nThe authors also construct a new dataset (**WD-Objects**) containing both synthetic and real scenes of deformable objects driven by wind.  \nExperiments show that DiffWind achieves high-quality 3D reconstruction, realistic forward simulation, and plausible “wind relocation” (transferring estimated wind fields to new objects or scenes).  \n\nOverall, this is a technically strong paper with detailed derivations — good job by the authors."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"**Q1.** How robust is DiffWind to **incomplete or single-view** input? Could it generalize if only a subset of views is available?  \n\n**Q2.** I am curious about the **training data setting**. The quantitative metrics in Table 1 are extremely high — for instance, a PSNR of 52.5 dB suggests a very dense camera setup. How many camera views were used during training?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"- **Novel problem formulation:** One of the first attempts to jointly reconstruct *hidden wind fields* and *object dynamics* from visual observations, bridging 3D reconstruction, differentiable physics, and generative modeling.  \n- **Strong technical depth:** The coupling of the Eulerian grid (wind) and Lagrangian particles (objects) via MPM is elegant and physically grounded. The inclusion of an LBM-based regularization further enforces physical realism.  \n- **Thorough experiments:** Evaluated on both **synthetic and real** WD-Objects datasets with multiple categories (cloth, flags, plants, etc.), demonstrating reconstruction quality (PSNR, SSIM, LPIPS) and perceptual realism via user studies.  \n- **Demonstrated versatility:** The model supports *forward simulation* and *wind relocation* tasks, showing potential for cross-scene generalization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**W1. Dependence on ideal inputs:** The method requires multi-view, calibrated video and accurate segmentation, which limits its applicability to real-world, in-the-wild scenarios.  \n\n**W2. Computational cost:** The framework integrates MPM, LBM, and differentiable rendering, but training time and memory requirements are not reported. Practical efficiency and scalability remain unclear."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916209560,"tcdate":1761656111093,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2366/Reviewer_8z3H"],"signatures":["ICLR.cc/2026/Conference/Submission2366/Reviewer_8z3H"],"forum":"vKVzihkbQo","number":1,"license":"CC BY 4.0","cdate":1761656111093,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2366/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916209560,"domain":"ICLR.cc/2026/Conference","replyto":"vKVzihkbQo","id":"KdS5oLAo1m","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Physics-based Modeling","3D Dynamics","System Identification","Differentiable Physics"]},"supplementary_material":{"value":"/attachment/7ceccac6970d2edd40c720683c1eac6f76904ac4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Modeling wind-driven object dynamics from video observations is highly challenging due to the invisibility and spatio–temporal variability of wind, as well as the complex deformations of objects. We present DiffWind, a physics-informed differentiable framework that unifies wind–object interaction modeling, video-based reconstruction, and forward simulation. Specifically, we represent wind as a grid-based physical field and objects as particle systems derived from 3D Gaussian Splatting, with their interaction modeled by the Material Point Method (MPM). To recover wind-driven object dynamics, we introduce a reconstruction framework that jointly optimizes the spatio–temporal wind force field and object motion through differentiable rendering and simulation. To ensure physical validity, we incorporate the Lattice Boltzmann Method (LBM) as a physics-informed constraint, enforcing compliance with fluid dynamics laws. Beyond reconstruction, our method naturally supports forward simulation under novel wind conditions and enable new applications such as wind retargeting. We further introduce WD-Objects, a dataset of synthetic and real-world wind-driven scenes. Extensive experiments demonstrate that our method significantly outperforms prior dynamic scene modeling approaches in both reconstruction accuracy and simulation fidelity, opening a new avenue for video-based wind–object interaction modeling. The project page is available at: [https://zju3dv.github.io/DiffWind/](https://zju3dv.github.io/DiffWind/)."},"_bibtex":{"value":"@inproceedings{\nlei2026diffwind,\ntitle={DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics},\nauthor={Yuanhang Lei and Boming Zhao and Zesong Yang and Xingxuan Li and Tao Cheng and Haocheng Peng and Ru Zhang and Yang Yang and Siyuan Huang and Yujun Shen and Ruizhen Hu and Hujun Bao and Zhaopeng Cui},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=vKVzihkbQo}\n}"},"title":{"value":"DiffWind: Physics-Informed Differentiable Modeling of Wind-Driven Object Dynamics"},"pdf":{"value":"/pdf/3422bdeef670a30b81dac5199f82667d5ca1d1fa.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lei|diffwind_physicsinformed_differentiable_modeling_of_winddriven_object_dynamics"},"authorids":{"value":["~Yuanhang_Lei1","~Boming_Zhao2","~Zesong_Yang1","~Xingxuan_Li2","~Tao_Cheng4","~Haocheng_Peng1","~Ru_Zhang2","~Yang_Yang156","~Siyuan_Huang2","~Yujun_Shen1","~Ruizhen_Hu1","~Hujun_Bao1","~Zhaopeng_Cui1"]},"authors":{"value":["Yuanhang Lei","Boming Zhao","Zesong Yang","Xingxuan Li","Tao Cheng","Haocheng Peng","Ru Zhang","Yang Yang","Siyuan Huang","Yujun Shen","Ruizhen Hu","Hujun Bao","Zhaopeng Cui"]}},"version":2},{"content":{"summary":{"value":"This paper investigates an intriguing problem regarding how machine learning models evolve during dynamic retraining using model-annotated samples, incorporating strategic human responses. The authors discover that it becomes increasingly likely for individuals to receive positive decisions as the model undergoes retraining, although the proportion of individuals with positive labels may decrease over time. To stabilize the dynamics, the authors propose a refined retraining process. They also examine how these retraining processes can impact algorithmic fairness and find that enforcing common fairness constraints in every retraining round may not benefit the disadvantaged groups in the long term. Experiments conducted on both synthetic and real-world data validate the findings of this study."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.\tThe paper thoroughly analyzes how humans, acting as strategic agents, adapt their behavior in response to ML systems and how this behavior impacts the retraining process of ML systems. By formalizing these interactions and analyzing their long-term dynamics, the paper provides a theoretical foundation for understanding and predicting these complex interactions. This in-depth analysis helps uncover potential systemic issues and offers concrete theoretical support for improving models.\n2.\tThe paper not only identifies potential risks of retraining ML models with model-annotated data but also proposes an improved retraining method using a probabilistic sampler to enhance the quality of model-annotated samples. This method aims to stabilize the dynamics of acceptance and qualification rates, reducing classifier bias. The proposed solution is innovative and practical, helping to mitigate negative impacts in real-world applications.\n3.\tThe paper combines theoretical analysis with experiments on semi-synthetic and real data to validate the findings. The experimental results, which show consistent dynamics with theoretical predictions, enhance the credibility and applicability of the research. This approach ensures the reliability of the research findings by providing both theoretical and empirical support."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tWhile the conclusions of the paper help understand the impact of human strategic behavior on ML systems, many of these conclusions are somewhat intuitive and straightforward. For instance, the increase in acceptance rates and the potential decrease in qualification rates over time are logically reasonable but do not provide particularly new insights. Delving deeper into underlying mechanisms or revealing more complex interactions could make the research more innovative and impactful.\n2.\tThe scale of the datasets used in experiments is relatively small, which may not fully capture the complexity of system dynamics in large-scale data environments.\n3.\tThe paper primarily focuses on linear models and specific distributions. The research conclusions might not fully apply to non-linear models and complex data distributions."},"limitations":{"value":"N.A"}},"nonreaders":[],"tmdate":1730878897997,"tcdate":1720720087104,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission4087/Reviewer_ynjw"],"signatures":["NeurIPS.cc/2024/Conference/Submission4087/Reviewer_ynjw"],"forum":"2UJLv3KPGO","number":1,"license":"CC BY 4.0","cdate":1720720087104,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission4087/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878897997,"domain":"NeurIPS.cc/2024/Conference","replyto":"2UJLv3KPGO","id":"OACb7m0ufh","forumContent":{"TLDR":{"value":"This paper studies the dynamics of welfare and fairness where strategic agents interact with an ML system retrained over time with model-annotated and human-annotated samples."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Strategic Classification","Long-term Fairness"]},"supplementary_material":{"value":"/attachment/47e5b67e423e5386c741bd55b9081353cf4a2dfa.zip"},"primary_area":{"value":"machine_learning_for_social_sciences"},"abstract":{"value":"As machine learning (ML) models are increasingly used in social domains to make consequential decisions about humans, they often have the power to reshape data distributions. Humans, as strategic agents, continuously adapt their behaviors in response to the learning system. As populations change dynamically, ML systems may need frequent updates to ensure high performance. However, acquiring high-quality *human-annotated* samples can be highly challenging and even infeasible in social domains. A common practice to address this issue is using the model itself to annotate unlabeled data samples. This paper investigates the long-term impacts when ML models are retrained with *model-annotated* samples when they incorporate human strategic responses. We first formalize the interactions between strategic agents and the model and then analyze how they evolve under such dynamic interactions. We find that agents are increasingly likely to receive positive decisions as the model gets retrained, whereas the proportion of agents with positive labels may decrease over time. We thus propose a *refined retraining process* to stabilize the dynamics. Last, we examine how algorithmic fairness can be affected by these retraining processes and find that enforcing common fairness constraints at every round may not benefit the disadvantaged group in the long run. Experiments on (semi-)synthetic and real data validate the theoretical findings."},"_bibtex":{"value":"@inproceedings{\nxie2024automating,\ntitle={Automating Data Annotation under Strategic Human Agents: Risks and Potential Solutions},\nauthor={Tian Xie and Xueru Zhang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=2UJLv3KPGO}\n}"},"title":{"value":"Automating Data Annotation under Strategic Human Agents: Risks and Potential Solutions"},"pdf":{"value":"/pdf/44841391152cc9c0c9a8aa7e562ff8698b75fdd0.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"xie|automating_data_annotation_under_strategic_human_agents_risks_and_potential_solutions"},"authorids":{"value":["~Tian_Xie4","~Xueru_Zhang2"]},"authors":{"value":["Tian Xie","Xueru Zhang"]}},"version":2},{"content":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Pseudo Physics","Data-Driven Physics Discovery","PDEs","Neural Operator","AI for science","Scientific Machine Learning"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in operator learning are transforming the landscape of computational physics and engineering, especially alongside the rapidly evolving field of physics-informed machine learning. The convergence of these areas offers\nexciting opportunities for innovative research and applications. However, merging\nthese two realms often demands deep expertise and explicit knowledge of physical systems, which may be challenging or even impractical in relatively complex applications. To address this limitation, we propose a novel framework: Pseudo\nPhysics-Informed Neural Operator (PPI-NO). In this framework, we construct a\nsurrogate physics system for the target system using partial differential equations\n(PDEs) derived from simple, rudimentary physics knowledge, such as basic differential operators. We then couple the surrogate system with the neural operator model, utilizing an alternating update and learning process to iteratively enhance\nthe model’s predictive power. While the physics derived via PPI-NO may not mirror the ground-truth underlying physical laws — hence the term “pseudo physics” — this approach significantly enhances the accuracy of current operator learning\nmodels, particularly in data scarce scenarios. Through extensive evaluations across\nfive benchmark operator learning tasks and an application in fatigue modeling,\nPPI-NO consistently outperforms competing methods by a significant margin. The\nsuccess of PPI-NO may introduce a new paradigm in physics-informed machine\nlearning, one that requires minimal physics knowledge and opens the door to\nbroader applications in data-driven physics learning and simulations."},"_bibtex":{"value":"@misc{\nchen2025pseudo,\ntitle={Pseudo Physics-Informed Neural Operators},\nauthor={Keyan Chen and Yile Li and Da Long and WEI W. XING and Jacob Hochhalter and Shandian Zhe},\nyear={2025},\nurl={https://openreview.net/forum?id=CrmUKllBKs}\n}"},"title":{"value":"Pseudo Physics-Informed Neural Operators"},"pdf":{"value":"/pdf/864b77c21caf8f31310746c5d9b464fe5feadfa1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|pseudo_physicsinformed_neural_operators"},"authorids":{"value":["~Keyan_Chen3","~Yile_Li1","~Da_Long1","~WEI_W._XING1","~Jacob_Hochhalter1","~Shandian_Zhe1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Keyan Chen","Yile Li","Da Long","WEI W. XING","Jacob Hochhalter","Shandian Zhe"]}},"tmdate":1738735703541,"tcdate":1727291893022,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4999/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission4999/Authors"],"forum":"CrmUKllBKs","license":"CC BY 4.0","number":4999,"cdate":1727291893022,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/Submission4999/-/Full_Submission","ICLR.cc/2025/Conference/-/Edit"],"mdate":1738735703541,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"CrmUKllBKs","version":2},{"content":{"summary":{"value":"PhysERL-Inv is a hybrid physics-encoded inverse modeling method to predict Arctic snow depth. The authors encode a hydrostatic balance equation into a sequence model, use supervised contrastive representation learning to shape latent representations, and invert the physics-encoded mapping to estimate hidden parameters that improve snow-depth prediction. They report substantial error reductions versus several neural baselines (reporting ≈20% improvement overall)."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Could you add ablations for (a) physics encoding strength (e.g., ablating terms in the hydrostatic relation), (b) contrastive loss weight, and (c) latent-dimension size."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Clear applied motivation. Snow depth over Arctic sea ice is an important but sparsely observed variable — the problem and practical need are well framed.\n\n- Nice hybrid approach. Combining an explicit physics relation (hydrostatic balance) with representation learning and an inverse mapping is conceptually sensible for data-sparse geoscience tasks. The architecture and training pipeline (encoder–decoder LSTM + self-attention, contrastive loss + MSE) are described and reasonable.\n\n- Solid empirical improvements & ablations. The paper compares to a set of neural baselines (LSTM, BiLSTM, NeuralODE, ResNet50) and shows quantitative wins (PhysERL-Inv MSE 0.3568 vs others; the table shows 17–32% relative changes on MSE/RMSE) and an ablation showing supervised contrastive learning (SCL) helps in low-data regimes. Those results are presented clearly."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited methodological novelty and analysis. The paper mainly integrates existing components — LSTM-based sequence modeling, supervised contrastive learning, and physics encoding via hydrostatic balance — without introducing a new learning algorithm or architectural mechanism. While the integration is well-motivated, the paper lacks deeper analysis or theoretical insight into why the combination works or how the proposed “surjective inversion” differs fundamentally from standard inverse modeling or PINN-style methods.\n\n- Baseline coverage is incomplete. The study compares primarily against conventional neural baselines (LSTM, BiLSTM, NeuralODE, ResNet50), but omits stronger physics-informed or operator-learning approaches (e.g., PINNs, DeepONet, FNO). Including or at least discussing these would be necessary to convincingly demonstrate advantages at the level expected for top ML venues.\n\n- Limited scope and generalization evaluation. Experiments are restricted to one Arctic dataset and temporal range, with no tests on cross-region or out-of-distribution generalization. Broader validation across different spatial or temporal domains would help establish robustness and strengthen claims of general applicability.\n\n- Formatting issue. The submission uses the 2025 ICLR template instead of the required 2026 version, which should be corrected for compliance."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931631496,"tcdate":1761956307915,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19786/Reviewer_Vw7G"],"signatures":["ICLR.cc/2026/Conference/Submission19786/Reviewer_Vw7G"],"forum":"49JcR9oeoo","number":3,"license":"CC BY 4.0","cdate":1761956307915,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19786/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931631496,"domain":"ICLR.cc/2026/Conference","replyto":"49JcR9oeoo","id":"0Qo1mQ9PBN","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Inverse modeling","Supervised representation learning","Surjective mapping","Arctic snow depth","Knowledge guidance"]},"supplementary_material":{"value":"/attachment/bf47f34f6a8ea7e7f0ff7790c09f265da7a14919.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"The accurate estimation of Arctic snow depth ($h_s$) remains a critical time-varying inverse problem due to the extreme scarcity and noise inherent in associated sea ice parameters. Existing process-based and data-driven models are either highly sensitive to sparse data or lack the physical interpretability required for climate-critical applications. To address this gap, we introduce PhysE-Inv, a novel framework that integrates a sophisticated sequential architecture, an LSTM Encoder-Decoder with Multi-head Attention and physics-guided contrastive learning, with physics-guided inference.Our core innovation lies in a surjective, physics-constrained inversion methodology. This methodology first leverages the hydrostatic balance forward model as a target-formulation proxy, enabling effective learning in the absence of direct $h_s$ ground truth; second, it uses reconstruction physics regularization over a latent space to dynamically discover hidden physical parameters from noisy, incomplete time-series input. Evaluated against state-of-the-art baselines, PhysE-Inv significantly improves prediction performance, reducing error by 20\\% while demonstrating superior physical consistency and resilience to data sparsity compared to empirical methods. This approach pioneers a path for noise-tolerant, interpretable inverse modeling, with wide applicability in geospatial and cryospheric domains."},"_bibtex":{"value":"@misc{\nsampath2026physerlinv,\ntitle={Phys{ERL}-Inv: A Physics-Encoded Inverse Modeling Approach for Arctic Snow Depth Prediction},\nauthor={Akila Sampath and Vandana Janeja and Jianwu Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=49JcR9oeoo}\n}"},"title":{"value":"PhysERL-Inv: A Physics-Encoded Inverse Modeling Approach for Arctic Snow Depth Prediction"},"pdf":{"value":"/pdf/7871ac9fc35149e3f88a3a37b1d92d3464bce152.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sampath|physerlinv_a_physicsencoded_inverse_modeling_approach_for_arctic_snow_depth_prediction"},"authorids":{"value":["~Akila_Sampath1","~Vandana_Janeja1","~Jianwu_Wang1"]},"authors":{"value":["Akila Sampath","Vandana Janeja","Jianwu Wang"]}},"version":2},{"content":{"summary":{"value":"The paper systematically studies how in-context learning (ICL) emerges in multimodal transformers using a controlled synthetic setup built from Gaussian mixtures. By varying data complexity factors and introducing quantitative diagnostics—PHStrength, IndStrength, TLA, and CLA—the authors trace the formation of induction-style attention circuits. The key findings are: (1) rotary and other relative positional encodings weaken ICL formation; (2) scaling increases the data complexity threshold for unimodal ICL, promoting memorization; (3) multimodal ICL is asymmetric, with the primary modality bootstrapping learning for the secondary; and (4) pretrained encoder quality is crucial for strong multimodal ICL. Results are validated on large models like Qwen2.5-VL and IDEFICS. The work provides a clear, mechanistic view of how architecture, scaling, and representation quality interact to produce ICL behavior in multimodal transformers."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"How might these findings inform training recipes for large production MLLMs (positional encoding choice, pretraining mix, encoder pretraining)? The paper gives implications — could you make them more prescriptive?\nHow sensitive are the RoPE-vs-absolute results to context length and dataset complexity? Is there a regime where RoPE still dominates (e.g., much longer contexts)?\nIn unimodal scaling experiments, if you scale data proportionally with model size, does the ICL threshold still increase? Please report model×data scaling curves."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Careful, controlled experimental design — synthetic GMM data + control over K, ε, B, α gives clear causal evidence for how data statistics drive ICL in both uni- and multimodal regimes. This leads to mechanistic progress measurements — PHStrength, IndStrength, TLA, CLA are well-motivated, quantitatively predictive, and allow the authors to track circuit formation over training. \n\nClear novel architectural insight: RoPE harms induction circuits — the paper demonstrates that RoPE (and ALiBi) consistently reduce ICL accuracy vs absolute PEs and produce more diffuse attention that weakens previous-token / induction heads. T\n\nImportant multimodal asymmetry finding — showing that a decoder pretrained on a high-diversity primary modality can bootstrap ICL such that the secondary modality needs far less diversity/burstiness is an intuitive and practically useful result for dataset and architecture design."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Synthetic → real generalization limited — while synthetic control is powerful, results hinge on idealized GMMs; the real-data validation is limited (Qwen2.5-VL analysis and a small Omniglot probe). Broader real-world tests are needed to confirm generality.\n\nPositional encoding recommendation could be risky in practice — the paper shows RoPE/ALiBi hurt ICL in these tasks, but RoPE brings other benefits (length generalization, training stability). The manuscript does not fully quantify tradeoffs (e.g., effect on other tasks, or hybrid encodings), which limits actionable guidance. \n\nScaling analysis might conflate capacity vs data regime — the unimodal result (“larger models need more complex data to show ICL”) is interesting, but the experiments use fixed data budgets. It remains unclear whether larger models trained with proportionally more data would still favor in-weight memorization. The compute/data scaling frontier isn’t fully explored."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916530768,"tcdate":1762789884218,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3055/Reviewer_qRdn"],"signatures":["ICLR.cc/2026/Conference/Submission3055/Reviewer_qRdn"],"forum":"V1vOjoWcvc","number":4,"license":"CC BY 4.0","cdate":1762789884218,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3055/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916530768,"domain":"ICLR.cc/2026/Conference","replyto":"V1vOjoWcvc","id":"BM74HSflfl","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["in-context learning","multimodal learning"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Multimodal large language models (MLLMs) often exhibit in-context learning (ICL) abilities, yet the conditions under which multimodal ICL emerges, and the mechanisms underlying it, remain poorly understood. In particular, how training data statistics and architectural choices jointly shape this capability is still an open question. To address this, we reverse-engineer multimodal ICL by training small transformer models on controlled synthetic classification tasks with varying data statistics and architectural choices.\nWe begin by revisiting core principles of unimodal ICL in modern transformers. While several prior findings replicate, our experiments yield two notable observations. First, Rotary Position Embeddings (RoPE), a standard component in contemporary LLMs, can delay the onset of ICL circuits. Second, larger models require stronger statistical cues in the training data for strong ICL to appear.\nExtending our analysis to the multimodal setting reveals a fundamental learning asymmetry. Once a primary modality has learned a core ICL circuit from statistically diverse data, a secondary modality can reach comparable ICL performance with far less data complexity. In contrast to the unimodal regime, we further find that model scaling consistently improves multimodal ICL.\nTo understand why these patterns emerge, we turn to mechanistic analysis. Using progress measures that track circuit formation during training, we show that ICL accuracy is tightly correlated with the strength of an induction-style circuit that copies labels from in-context exemplars that match the query. Both unimodal and multimodal ICL rely on this induction mechanism, while multimodal training primarily refines and extends it across modalities.\nTogether, these results provide a mechanism-level account of ICL in modern multimodal transformers, offer explanations for several empirical phenomena observed in MLLMs, and introduce a controlled testbed for future work on multimodal ICL."},"_bibtex":{"value":"@misc{\nhuang2026towards,\ntitle={Towards understanding multimodal in-context learning},\nauthor={Yiran Huang and Karsten Roth and Quentin Bouniot and Wenjia Xu and Zeynep Akata},\nyear={2026},\nurl={https://openreview.net/forum?id=V1vOjoWcvc}\n}"},"title":{"value":"Towards understanding multimodal in-context learning"},"pdf":{"value":"/pdf/8e584c012b4a1bacd045c9dc10131075230ec4c0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"huang|towards_understanding_multimodal_incontext_learning"},"authorids":{"value":["~Yiran_Huang2","~Karsten_Roth1","~Quentin_Bouniot1","~Wenjia_Xu1","~Zeynep_Akata1"]},"authors":{"value":["Yiran Huang","Karsten Roth","Quentin Bouniot","Wenjia Xu","Zeynep Akata"]}},"version":2},{"content":{"venue":{"value":"New Journal of Physics"},"pdf":{"value":"https://iopscience.iop.org/article/10.1088/1367-2630/ab60f8/pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"owen|corrigendum_number_of_hidden_states_needed_to_physically_implement_a_given_conditional_distribution_2019_new_journal_of_physics_21_013022"},"html":{"value":"https://doi.org/10.1088/1367-2630/ab60f8"},"abstract":{"value":"Corrigendum: Number of hidden states needed to physically implement a given conditional distribution (2019 New Journal of Physics 21 013022), Owen, Jeremy A, Kolchinsky, Artemy, Wolpert, David H"},"title":{"value":"Corrigendum: Number of hidden states needed to physically implement a given conditional distribution (2019 New Journal of Physics 21 013022)"},"authors":{"value":[{"fullname":"Jeremy A Owen","username":"https://orcid.org/orcid-search/search?searchQuery=Jeremy%20A%20Owen"},{"fullname":"Artemy Kolchinsky","username":"~Artemy_Kolchinsky1"},{"fullname":"David H Wolpert","username":"https://orcid.org/orcid-search/search?searchQuery=David%20H%20Wolpert"}]}},"tmdate":1789731685410,"pdate":1577836800000,"externalIds":["doi:10.1088/1367-2630/ab60f8"],"tcdate":1789731677970,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Artemy_Kolchinsky1"],"forum":"Ihpqvkhe3l","license":"CC BY-SA 4.0","number":110880,"cdate":1784410158153,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1789731685410,"domain":"OpenReview.net/Public_Article","id":"Ihpqvkhe3l","version":2},{"content":{"summary":{"value":"The paper deals with radial phase retrieval from intensity-only measurements, a highly ill-posed inverse problem, under the conditions of an outer ring generalization. A physics-informed hybrid network is proposed that (i) embeds radial priors via a smooth exponential spline and a monotonic outer radius booster, (ii) couples two differentiable PDE branches, and (iii) enforces a strict radial projection with a radius-dependent $\\alpha$-fusion and output symmetry. Experimental results using synthetic data show that the method can recover more outer rings and reduce hallucinations compared to baselines."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Since the phase retrieval is highly ill-posed problem, how is phase ambiguity handled, and what happens if the actual field deviates slightly from perfect radial symmetry?\n- What are the stability regions for NLSE (nonlinearity/dispersion) and TIE step size?\n- Can the hybrid model signal outer rings with low reliability to avoid hallucinations?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- The designed network adapts the inductive bias with the ring-structured phase retrieval problem, rather than relying on generic CNN priors. This is a thoughtful specialization that many “physic-informed” works propose but do not actually implement.\n-  The task “training on inner rings, testing on invisible outer rings” as an explicit setting for distribution shift is an original evaluation method for radial phase retrieval and could serve as a basis for subsequent benchmarks. Accurate recontruction of the outer ring is crucial for downstream optical tasks. The method that reduces hallucinations, improves ring fidelity, and remains stable at the same time is of great value to laboratories and systems that cannot afford dense multi-surface measurements. \n- The paper clearly separates the roles of (i) radial projection, (ii) outer radius boosting, (iii) NLSE vs. TIE paths, and (iv) $\\alpha(\\rho)$ fusion. This makes the method easier to understand and implement."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The experiment is conducted using synthetic data with limited distribution shifts. For the proposed method, which is to be generalized to invisible outer rings and be physics-informed, the empirical scope of application is too narrow.\n- For the comparison, the baselines are U-Net/FNO variants. Please compare the proposed method with physics-guided unrolled methods and plug-and-play / regularization by denoising (RED) algorithms [1, 2, 3]\n- Stability, identifiability, and convergence sketches are based on assumptions such as Lipschitzness, radial Hankel linearization, and noise models, which may not apply under realistic optical conditions. The theory presented is promising, but has not yet been sufficiently investigated for practical application.\n- The current \"6. Experiments\" and \"7.Overall Experimental Analysis\" sections are difficult to interpret. They offers only a limited interpretation of what the individual key figures reflect in terms of specific error modes, and contains figures/tables whose captions lack essential experimental details.\n\n[1] Ulyanov, Dmitry, Andrea Vedaldi, and Victor Lempitsky. \"Deep image prior.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.\n\n[2] Metzler, Christopher, et al. \"prDeep: Robust phase retrieval with a flexible deep network.\" International Conference on Machine Learning. PMLR, 2018.\n\n[3] Mardani, Morteza, et al. \"A Variational Perspective on Solving Inverse Problems with Diffusion Models.\" The Twelfth International Conference on Learning Representations."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925477028,"tcdate":1761941898047,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15167/Reviewer_o2RK"],"signatures":["ICLR.cc/2026/Conference/Submission15167/Reviewer_o2RK"],"forum":"jS3EKPSaAR","number":4,"license":"CC BY 4.0","cdate":1761941898047,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15167/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925477028,"domain":"ICLR.cc/2026/Conference","replyto":"jS3EKPSaAR","id":"3PxKsLzzDh","forumContent":{"TLDR":{"value":"a pde based physic informed neural network that performs well in generalization"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["phase retrieval;PDE networks; outer-ring extrapolation; inverse problems"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Phase retrieval from intensity-only measurements is severely ill-posed due to global-gauge and rotational symmetries. We consider outer-ring generalization: training with supervision from only a few inner rings and testing the model’s ability to reconstruct a broader set of unseen outer rings. We introduce a physics-informed hybrid network that combines (i) radial priors encoded by a smooth exponentiated spline and a \\emph{monotone} outer-radius booster, (ii) two differentiable PDE branches---a Strang-split Kerr--NLSE pathway for high-frequency synthesis and a TIE-based low-pass pathway for coarse structure---and (iii) a strict radial projection enforcing output symmetry, together with a radius-dependent $\\alpha$-fusion. Across the tested configurations, when trained only on a few rings (1-3), our model reconstructs more rings(4-9) than conventional methods, and achieves better stability in peak\npositions and amplitude calibration under out-of-distribution settings. This provides some inspiration for enhancing the generalization of physics-informed neural networks when applied to optical inverse problems. Ablations isolate the contribution of the alpha fusion, PDE coupling, and monotone\nboosting. We will release pseudo-code to facilitate reproducibility."},"_bibtex":{"value":"@misc{\nyao2026physicsinformed,\ntitle={{PHYSICS}-{INFORMED} {RADIAL} {PHASE} {RETRIEVAL} {NEURAL} {NETWORK} {WITH} {HYBRID} {DEEP} {PRIORS} {AND} {DUAL} {PDE}},\nauthor={ZIYONG YAO},\nyear={2026},\nurl={https://openreview.net/forum?id=jS3EKPSaAR}\n}"},"title":{"value":"PHYSICS-INFORMED RADIAL PHASE RETRIEVAL NEURAL NETWORK WITH HYBRID DEEP PRIORS AND DUAL PDE"},"pdf":{"value":"/pdf/d0ee395178c85a98450cfb5c748c70d7faeb7393.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yao|physicsinformed_radial_phase_retrieval_neural_network_with_hybrid_deep_priors_and_dual_pde"},"authorids":{"value":["~ZIYONG_YAO1"]},"authors":{"value":["ZIYONG YAO"]}},"version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2604.01313v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"xia|jetprism_diagnosing_convergence_for_generative_simulation_and_inverse_problems_in_nuclear_physics"},"html":{"value":"https://doi.org/10.48550/arXiv.2604.01313"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2604-01313,\n  publtype={informal},\n  author={Zeyu Xia and Tyler Kim and Trevor Reed and Judy Fox and Geoffrey Fox and Adam Szczepaniak},\n  title={JetPrism: diagnosing convergence for generative simulation and inverse problems in nuclear physics},\n  year={2026},\n  month={April},\n  cdate={1775001600000},\n  journal={CoRR},\n  volume={abs/2604.01313},\n  url={https://doi.org/10.48550/arXiv.2604.01313}\n}\n"},"abstract":{"value":"High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation. While Conditional Flow Matching (CFM) offers a robust acceleration approach, we demonstrate its standard training loss is fundamentally misleading. Specifically, utilizing a Jefferson Lab Nuclear Physics (NP) kinematic dataset ($γp \\to ρ^0 p \\to π^+π^- p$), we expose that CFM loss plateaus prematurely, obscuring ongoing physical refinement. To verify this disconnect is a dataset-agnostic pathology, we introduce ScatterPrism, an efficient generative surrogate evaluated against both the NP data and synthetic stress tests modeling challenging 1D distribution topologies. Coupling these benchmarks, we establish that physics-informed metrics continue improving long after standard loss converges. Consequently, we propose a multi-metric diagnostic protocol to ensure true kinematic fidelity without data memorization. Driven by NP challenges relevant to the forthcoming Electron-Ion Collider (EIC), this unified machinery has strong potential to extend to High-Energy Physics (HEP) applications, such as jet modeling. Furthermore, the framework holds promise for broader domains requiring rigorous generative reliability, including medical imaging, astrophysics, and quantitative finance."},"title":{"value":"JetPrism: diagnosing convergence for generative simulation and inverse problems in nuclear physics"},"authors":{"value":[{"fullname":"Zeyu Xia","username":""},{"fullname":"Tyler Kim","username":""},{"fullname":"Trevor Reed","username":""},{"fullname":"Judy Fox","username":""},{"fullname":"Geoffrey Fox","username":""},{"fullname":"Adam Szczepaniak","username":""}]}},"tmdate":1790235702412,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2604-01313"],"tcdate":1782290556783,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Zeyu_Xia1"],"forum":"aWUSNSJl4j","license":"CC BY-SA 4.0","number":47495,"cdate":1775001600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Author_Removal","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1790235702412,"domain":"OpenReview.net/Public_Article","id":"aWUSNSJl4j","version":2},{"content":{"comment":{"value":"We have now performed the extra experiments you suggested with a synthetic dataset. To briefly recap: you proposed to train the model on a synthetic dataset, constructed with the concepts (plus noise) of our choice, to understand how the model leverages the concepts. In our earlier experiments, this was hard to achieve, because the best and second-best values of the hyperparameter alpha were not close in value, and therefore not intuitive (almost all results for $\\alpha$ < 1 in Table 5 from Appendix F are within the same range by standard deviation, so the best and second-best settings do not carry that much weight. We included a new figure 10 in Appendix F that clearly illustrates this).\n\nThe results from the new experiments are presented in Appendix I. We generate a time series dataset as the sum of different sine functions, and then train an Autoformer model with a bottleneck on the attention heads of the second layer. We vary the value of hyperparameter $\\alpha$, and define each concept in the bottleneck as one of the underlying functions (for which we have the ground truth by construction). \n\nAs expected, we find that the similarity between the bottleneck components and the concepts increases with increasing $\\alpha$ (this is visible as the emergence of a yellow diagonal in layer 2 in Figure 23). At $\\alpha=0$, there is no concept bottleneck and the similarity to the predefined concepts is minimal. At $\\alpha=1.0$, the model is only optimized for similarity to the concepts, and the prediction performance is terrible. Interestingly, at $\\alpha=0.8$, we hit a sweet spot where similarity to the predefined concepts is high and the prediction performance is also at its maximum.\n\nWe believe these additional experiments help in understanding how the model leverages interpretable concepts. We would like to thank you again for the suggestion, and are curious to hear whether it is indeed exactly what you had in mind."},"title":{"value":"Synthetic dataset"}},"tmdate":1732738853372,"tcdate":1732738853372,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11701/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission11701/Authors"],"forum":"A0mk2Wi68Y","number":10,"license":"CC BY 4.0","cdate":1732738853372,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11701/-/Official_Comment"],"mdate":1732738853372,"domain":"ICLR.cc/2025/Conference","replyto":"KSvLNxUFds","id":"nc3MjQyUkB","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Interpretability","Concept Bottleneck Model","Centered Kernel Alignment","Autoformer","Time Series Transformer"]},"primary_area":{"value":"interpretability and explainable AI"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"There has been a recent push of research on Transformer-based models for long-term time series forecasting, even though they are inherently difficult to interpret and explain. While there is a large body of work on interpretability methods for various domains and architectures, the interpretability of Transformer-based forecasting models remains largely unexplored. To address this gap, we develop a framework based on Concept Bottleneck Models to enforce interpretability of time series Transformers. We modify the training objective to encourage a model to develop representations similar to predefined interpretable concepts. In our experiments, we enforce similarity using Centered Kernel Alignment, and the predefined concepts include time features and an interpretable, autoregressive surrogate model (AR). We apply the framework to the Autoformer model, and present an in-depth analysis for a variety of benchmark tasks. We find that the model performance remains mostly unaffected, while the model shows much improved interpretability. Additionally, interpretable concepts become local, which makes the trained model easily intervenable. As a proof of concept, we demonstrate a successful intervention in the scenario of a time shift in the data, which eliminates the need to retrain."},"_bibtex":{"value":"@misc{\nsprang2025enforcing,\ntitle={Enforcing Interpretability in Time Series Transformers: A Concept Bottleneck Framework},\nauthor={Angela van Sprang and Erman Acar and Willem Zuidema},\nyear={2025},\nurl={https://openreview.net/forum?id=A0mk2Wi68Y}\n}"},"title":{"value":"Enforcing Interpretability in Time Series Transformers: A Concept Bottleneck Framework"},"pdf":{"value":"/pdf/d779f2bb8968ce1255ab8c93764d1f228028c393.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"sprang|enforcing_interpretability_in_time_series_transformers_a_concept_bottleneck_framework"},"authorids":{"value":["~Angela_van_Sprang1","~Erman_Acar1","~Willem_Zuidema1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Angela van Sprang","Erman Acar","Willem Zuidema"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Weierstrass Positional Encoding (WePE), a mathematically grounded 2D positional encoding for Vision Transformers.\nInstead of using traditional sinusoidal or rotary encodings, the authors use the Weierstrass elliptic function, a doubly periodic complex function to map image coordinates onto the complex plane.\nThis design aims to preserve true 2D spatial continuity and provide better distance decay and relative-position properties.\nExperiments on CIFAR-100, ImageNet, and VTAB show small but consistent accuracy gains over existing encodings."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"What is the computational cost of WePE compared to standard sine/cosine positional encodings?\n\nDid the authors compare WePE to high-order Fourier or complex-valued positional encodings to isolate its specific benefit?\n\nIs the improvement primarily from the double periodicity, or could similar results be achieved with a simpler 2D periodic basis?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Elegant mathematical formulation: the use of a doubly periodic complex function is theoretically appealing and fits the 2D geometry of images.\n\nContinuous and resolution-independent: WePE naturally handles different image sizes without interpolation.\n\nEmpirical improvements: consistent accuracy gains and smoother attention maps in visualization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Lack of broader context: while the mathematical formulation is elegant, the paper could better clarify how WePE conceptually differs from other periodic or complex-valued encodings (e.g., Fourier- or rotary-based).\n\nLimited analysis: the paper mainly reports performance improvements but provides little analysis or visualization explaining why the proposed encoding helps attention behavior.\n\nGeneralization scope unclear: it remains uncertain whether WePE benefits generalize to larger-scale or non-vision tasks, since experiments are focused on a few image benchmarks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931387582,"tcdate":1761975654831,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19480/Reviewer_b6eQ"],"signatures":["ICLR.cc/2026/Conference/Submission19480/Reviewer_b6eQ"],"forum":"lVmnN8g6lW","number":2,"license":"CC BY 4.0","cdate":1761975654831,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19480/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931387582,"domain":"ICLR.cc/2026/Conference","replyto":"lVmnN8g6lW","id":"VGMMXfD4VU","forumContent":{"TLDR":{"value":"We propose WePE for ViTs. By exploiting the doubly periodicity of the Weierstrass elliptic function, it avoids flattening 2D images into 1D sequences and preserves the natural geometric inductive bias disrupted by conventional encodings."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Weierstrass Elliptic Function","Positional Encoding","Vision Transformers","Double periodicity"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Vision Transformers (ViTs) have demonstrated remarkable success in computer vision tasks.  However, their reliance on learnable one-dimensional positional encoding  disrupts the inherent two-dimensional spatial structure of images due to patch flattening. Existing positional encoding approaches lack geometric constraints and fail to preserve a monotonic correspondence between Euclidean spatial distances and sequential index distances, thereby limiting the model's capacity to leverage spatial proximity priors effectively. Recognizing that periodicity is particularly beneficial for positional encoding, we propose Weierstrass elliptic Positional Encoding (WePE), a mathematically principled approach that encodes two-dimensional coordinates in the complex domain. This method maps the normalized two-dimensional patch coordinates onto the complex plane and constructs a compact four-dimensional positional feature based on the Weierstrass elliptic function $\\wp(z)$ and its derivative. The doubly periodic property of $\\wp(z)$ enables a principled encoding of 2D positional information, while their intrinsic lattice structure aligns naturally with the geometric regularities of patch grids in images. Their nonlinear geometric characteristics enable faithful modeling of spatial distance relationships,  while the associated algebraic addition formula allows relative positional information between arbitrary patch pairs to be derived directly from their absolute encodings. WePE is a plug-and-play, resolution-agnostic positional module that integrates seamlessly with existing ViTs. Extensive experiments demonstrate that WePE delivers consistent performance gains in most scenarios, while its implementation with precomputed lookup tables ensures that these improvements incur no noticeable computational or memory overhead. In addition, several analyses and ablation studies bring further confirmation to the effectiveness of our method."},"_bibtex":{"value":"@misc{\nanonymous2026weierstrass,\ntitle={Weierstrass Positional Encoding for Vision Transformers},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=lVmnN8g6lW}\n}"},"title":{"value":"Weierstrass Positional Encoding for Vision Transformers"},"pdf":{"value":"/pdf/9499698f8b41d25eb5f7c9a9056a754d6d5f82be.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"xin|weierstrass_positional_encoding_for_vision_transformers"},"authorids":{"value":["~Zhihang_Xin1","~Rui_Wang14","~Xitong_Hu1","~Xiaojun_Wu2"]},"authors":{"value":["Zhihang Xin","Rui Wang","Xitong Hu","Xiaojun Wu"]}},"version":2},{"content":{"summary":{"value":"The paper \"InstaSHAP: Interpretable Additive Models Explain Shapley Values Instantly\" introduces InstaSHAP, an efficient method for calculating Shapley values to explain machine learning model predictions. InstaSHAP leverages a novel connection between Shapley values and Generalized Additive Models (GAMs), allowing Shapley values to be computed in a single forward pass. This variational approach addresses computational bottlenecks in traditional SHAP methods and improves accuracy in scenarios with complex feature interactions.\n\nKey contributions include a unified framework that connects GAMs and Shapley explanations, the development of InstaSHAP for fast Shapley computation, and extensive experiments showing its practical advantages over other methods like FastSHAP. This work makes Shapley-based explanations more accessible for real-time applications in high-stakes areas such as healthcare and finance which advancing the field of interpretable machine learning."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Questions:   \nValidation: Would you include experiments with more varied data types (e.g., time series, medical)? Have you considered comparing with Integrated Gradients or LIME?\n\nScalability and Performance: How does InstaSHAP scale on large datasets compared to FastSHAP? Can you share latency and runtime benchmarks?\n\nHandling Complexity: How does InstaSHAP manage complex, high-order dependencies? Have you tested it in scenarios where GAMs typically underperform?\n\nClarity Enhancements: Would you add intuitive explanations alongside complex proofs? Could you improve figure captions for clarity? Any real-world cases showing InstaSHAP's advantages? Are there plans for easy implementation or integration with SHAP tools?\n\nSuggestions\n\nExpand Empirical Analysis: Include experiments on more varied datasets to showcase broader applicability and compare InstaSHAP with methods like Integrated Gradients or LIME for a comprehensive evaluation.\n\nProvide Scalability Insights: Add benchmarks on training times and real-time performance metrics to support the paper’s scalability and real-time claims.\n\nClarify Theoretical Scope: Discuss how InstaSHAP handles high-order, non-linear interactions and elaborate on its limitations with complex models.\n\nEnhance Clarity: Add intuitive explanations for theoretical sections and improve figure annotations to make the paper more accessible.\n\nHighlight Practical Use: Present real-world case studies demonstrating InstaSHAP's benefits and share plans for making it user-friendly (e.g., open-source tools)."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Originality: The paper uniquely links Shapley values and Generalized Additive Models (GAMs), presenting an innovative approach to model interpretability. Introduction of \"InstaSHAP\" enables real-time Shapley value computation through GAM-based training which stands out as an original contribution.\n\nQuality: Strong, rigorous proofs link SHAP and GAMs can provide a detailed understanding of representation limits and power. Validated on synthetic, tabular, and image datasets showcase clear improvements over existing methods like FastSHAP and FaithSHAP.\n\nClarity: In this paper, complex concepts are made accessible with formal mathematical definitions and visual aids. Comprehensive context on Shapley values and GAMs ensures readers understand the motivation and methods.\n\nSignificance: The paper addresses gaps in traditional SHAP methods, especially for correlated and high-dimensional data, with real-time applicability. InstaSHAP’s instant computation is highly relevant for fields needing quick and accurate model insights, like healthcare and finance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Empirical Analysis: More diverse datasets (e.g., time series, medical) would better demonstrate robustness. A broader comparison with methods like Integrated Gradients or LIME is needed.\n\nScalability: The paper lacks details on computational efficiency for large datasets. Quantitative analysis of runtime and latency is needed to validate real-time claims.\n\nTheoretical Depth:  More discussion on handling complex non-linear dependencies would clarify limitations. The method's performance in non-GAM-friendly tasks needs clearer boundaries.\n\nClarity: Intuitive explanations alongside proofs would improve accessibility. More detailed captions and annotations are needed for clarity."}},"nonreaders":[],"tmdate":1732585310982,"tcdate":1730600854252,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1649/Reviewer_kSd9"],"signatures":["ICLR.cc/2025/Conference/Submission1649/Reviewer_kSd9"],"forum":"ky7vVlBQBY","number":2,"license":"CC BY 4.0","cdate":1730600854252,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1649/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732585310982,"domain":"ICLR.cc/2025/Conference","replyto":"ky7vVlBQBY","id":"sCvNcP63NY","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["additive models","GAM","SHAP","Shapley"]},"primary_area":{"value":"interpretability and explainable AI"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In recent years, the Shapley value and SHAP explanations have emerged as one\nof the most dominant paradigms for providing post-hoc explanations of blackbox models. Despite their well-founded theoretical properties, many recent works\nhave focused on the limitations in both their computational efficiency and their\nrepresentation power. The underlying connection with additive models, however,\nis left critically under-emphasized in the current literature. In this work, we find\nthat a variational perspective linking GAM models and SHAP explanations is able\nto provide deep insights into nearly all recent developments. In light of this connection, we borrow in the other direction to develop a new method to train interpretable GAM models which are automatically purified to compute the Shapley\nvalue in a single forward pass. Finally, we provide theoretical results showing the\nlimited representation power of GAM models is the same Achilles’ heel existing\nin SHAP and discuss the implications for SHAP’s modern usage in CV and NLP."},"_bibtex":{"value":"@inproceedings{\nenouen2025instashap,\ntitle={Insta{SHAP}: Interpretable Additive Models Explain Shapley Values Instantly},\nauthor={James Enouen and Yan Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=ky7vVlBQBY}\n}"},"title":{"value":"InstaSHAP: Interpretable Additive Models Explain Shapley Values Instantly"},"pdf":{"value":"/pdf/0fc58d283a75eea4ad672a45b5c738bf36ea4066.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"enouen|instashap_interpretable_additive_models_explain_shapley_values_instantly"},"authorids":{"value":["~James_Enouen1","~Yan_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["James Enouen","Yan Liu"]}},"version":2},{"content":{"summary":{"value":"This work introduces a foundation model tailored towards nuclear and particle physics. The authors develop a large scale open dataset with 10 million events and benchmark on three downstream tasks. The authors develop a self-supervised training methodology for their foundation models, and present benchmarks on downstream tasks in addition to scaling behaviors for their foundation models. They demonstrate improved performance of their model over baselines across all tasks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"* Why choose mamba over other linear transformers (deltanet, GLA..) and state space models? \n* How long are the sequence lengths for the input? State space models struggle with extremely long context. \n* Does the hierarchical raster scan have a strong impact on the amount of compute?\n* What is the full model architecture of the foundation model?\n* Can the code and models be made public?\n* Do you actually claim that self-supervised pre-training + fine-tuning outperforms a dedicated model trained from scratch on the downstream task in Fig. 5 (b)? Or are the architectures different? If you extend the axis, would Adaptor only eventually converge to the FM performance, as is seen in other FM papers in particle physics?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The authors present a well motivated problem that they address with a multi-step comprehensive framework \n* The self-supervised learning objective is physics motivated, and intuitively works with the detector setup\n* Very thorough benchmarking and ablation studies show improved performance for all pieces of the foundation model\n* Performance on downstream task show greatly increased performance over baselines\n* Scaling curves show increased performance for increased flops/parameter count, indicating scalability of the model\n* Shows greatly increased performance against state of the art models\n* Shows that the self-supervised learning technique improves performance on downstream tasks"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Does not justify choice of mamba over linear transformer models, or other state space model alternatives\n* Unclear how the hierarchical raster scan would impact the efficiency\n* No benchmarks against efficient transformer baselines such as HEP-T (locality sensitive hashing transformer for high energy physics applications, https://arxiv.org/abs/2402.12535), which performs very well on trackML datasets\n* For PID, also missing comparisons to dedicated physics ML algorithms, like MLPF (https://arxiv.org/abs/2503.00131) or HGPF (https://arxiv.org/abs/2410.23236)\n* Does not define the model architecture/construction (number of layers, parameters per layer etc)\n* There are no code or pretrained weights available\n* Unclear of pretraining + fine-tuning outperforms dedicated models trained from scratch; Most other literature on FMs in particle physics do not claim this and instead focus on data efficiency (i.e. better performance when training on smaller datasets)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919822885,"tcdate":1761973212640,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7782/Reviewer_2w6N"],"signatures":["ICLR.cc/2026/Conference/Submission7782/Reviewer_2w6N"],"forum":"qaI3cLFsiX","number":2,"license":"CC BY 4.0","cdate":1761973212640,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7782/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919822885,"domain":"ICLR.cc/2026/Conference","replyto":"qaI3cLFsiX","id":"QFwhIazxzv","forumContent":{"TLDR":{"value":"We present a scalable foundation model for sparse nuclear and particle physics detector data, capable of adaptation to diverse downstream tasks after pretraining."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Foundation Model","State Space Model","Neural Scaling","Particle Tracking","Nuclear Physics","Particle Physics"]},"supplementary_material":{"value":"/attachment/87460ebbfd31e692941657bb7ccc084990769618.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Large language models have revolutionized artificial intelligence by enabling large, generalizable models trained through self-supervision. This paradigm has inspired the development of scientific foundation models (FMs). However, applying this capability to experimental particle physics is challenging due to the sparse, spatially distributed nature of detector data, which differs dramatically from natural language. This work addresses if an FM for particle physics can scale and generalize across diverse tasks. We introduce a new dataset with more than 11 million particle collision events and a suite of downstream tasks and labeled data for evaluation. We propose a novel self-supervised training method for detector data and demonstrate its neural scalability with models that feature up to 188 million parameters. With frozen weights and task-specific adapters, this FM consistently outperforms baseline models across all downstream tasks. The performance also exhibits robust data-efficient adaptation. Further analysis reveals that the representations extracted by the FM are task-agnostic but can be specialized via a single linear mapping for different downstream tasks."},"_bibtex":{"value":"@inproceedings{\npark2026fmnpp,\ntitle={{FM}4{NPP}: A Scaling Foundation Model for Nuclear and Particle Physics},\nauthor={David Keetae Park and Shuhang Li and Yi Huang and Xihaier Luo and Haiwang Yu and Yeonju Go and Christopher Pinkenburg and Yuewei Lin and Shinjae Yoo and Joseph D. Osborn and Jin Huang and Yihui Ren},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=qaI3cLFsiX}\n}"},"title":{"value":"FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics"},"pdf":{"value":"/pdf/a8447a796e51fc50328ede271c7cf61e3c313155.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"park|fm4npp_a_scaling_foundation_model_for_nuclear_and_particle_physics"},"authorids":{"value":["~David_Keetae_Park2","~Shuhang_Li2","~Yi_Huang1","~Xihaier_Luo1","~Haiwang_Yu1","~Yeonju_Go1","~Christopher_Pinkenburg1","~Yuewei_Lin1","~Shinjae_Yoo1","~Joseph_D._Osborn1","~Jin_Huang9","~Yihui_Ren1"]},"authors":{"value":["David Keetae Park","Shuhang Li","Yi Huang","Xihaier Luo","Haiwang Yu","Yeonju Go","Christopher Pinkenburg","Yuewei Lin","Shinjae Yoo","Joseph D. Osborn","Jin Huang","Yihui Ren"]}},"version":2},{"content":{"venue":{"value":"ICIC (2) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-981-97-5581-3_26.pdf"},"venueid":{"value":"dblp.org/conf/ICIC/2024"},"paperhash":{"value":"song|physicsinformed_neural_networks_with_generalized_residualbased_adaptive_sampling"},"authorids":{"value":["~Xiaotian_Song1","https://dblp.org/search/pid/api?q=author:Shuchao_Deng:","~Jiahao_Fan4","~Yanan_Sun4"]},"html":{"value":"https://doi.org/10.1007/978-981-97-5581-3_26"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icic/SongDFS24,\n  author={Xiaotian Song and Shuchao Deng and Jiahao Fan and Yanan Sun},\n  title={Physics-Informed Neural Networks with Generalized Residual-Based Adaptive Sampling},\n  year={2024},\n  cdate={1704067200000},\n  pages={320-332},\n  url={https://doi.org/10.1007/978-981-97-5581-3_26},\n  booktitle={ICIC (2)},\n  crossref={conf/icic/2024-2}\n}\n"},"abstract":{"value":"Physics-informed neural networks (PINNs) are the powerful tools in solving partial differential equations (PDEs). In general, the performance of PINNs heavily relies on the sampling distribution of residual points. However, existing sampling methods still suffer from the following problems: 1) poor performance due to ignoring the locations with small PDE residuals; and 2) limited generalizability, i.e., the need to manually tune hyperparameters for every specific PDE. To address these issues, we propose a Generalized Residual-based Adaptive Sampling (G-RAS) method for PINNs. G-RAS incorporates a novel probability density function, which can concern locations with small PDE residuals. In addition, the hyperparameter setting is much less than others, i.e., various PDEs only need to set hyperparameter ranges rather than tuning for each one. Experiments on six widely used benchmarks demonstrate that G-RAS can improve prediction accuracy and convergence speed compared to 10 SOTA methods. The supplementary materials and source code are available at https://github.com/songxt3/G-RAS."},"title":{"value":"Physics-Informed Neural Networks with Generalized Residual-Based Adaptive Sampling"},"authors":{"value":["Xiaotian Song","Shuchao Deng","Jiahao Fan","Yanan Sun"]}},"tmdate":1745298121064,"pdate":1704067200000,"tcdate":1727709989011,"writers":["~"],"signatures":["~Jiahao_Fan4"],"forum":"ClxraKF4Ki","license":"CC BY-SA 4.0","number":120058,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1745298121064,"domain":"DBLP.org","id":"ClxraKF4Ki","version":2},{"content":{"summary":{"value":"This paper introduces a method to solve the k-Hyperplane Clustering problem in the 2-norm ($k-HC_2$) to global optimality using spatial branch-and-bound (SBB). The approach strengthens the classical mixed-integer quadratically constrained quadratic programming (MI-QCQP) formulation by incorporating constraints from polyhedral norm variants $(p = 1, \\infty)$. The authors show that including ∞-norm constraints allows the SBB method to obtain a nonzero lower bound in O(nk) nodes, compared to $\\Omega(2^{k(n−1)})$ nodes for the baseline. Experiments on synthetic datasets report speedups of 8–41 times, improving the number of instances solved to global optimality."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How does the method scale with the number of data points m, given that the assignment problem may become the bottleneck?\n2. Can the authors provide wall-clock time profiles and total node counts (not just medians) to better characterize tree size and convergence?\n3. How robust are the speedups under solver-default branching rules and parameters?\n4. Why does the multi-norm formulation perform worse than the individual $\\infty$-norm formulation in some cases?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Theoretical analysis proves that adding polyhedral norm constraints reduces the number of SBB nodes required to obtain a nonzero lower bound from exponential to polynomial.\n2. The method is evaluated on two synthetic testbeds (Low-dim and High-dim) with statistical tests (Wilcoxon signed-rank) confirming significant speedups.\n3. Clear MI-QCQP and MILP formulations are provided for the polyhedral norm constraints.\n4. The multi-norm relaxation framework is general and may be applied to other problems with nonconvex norm constraints."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Experiments are limited to small-scale instances (up to m = 30, n = 5, k = 5) and use synthetically generated data with a specific noise model.\n2. The theoretical advantage of $\\infty$-norm constraints assumes branching occurs first on the binary variables of the polyhedral norm formulation; performance under default branching strategies is not tested.\n3. The analysis focuses only on the number of nodes to the first nonzero bound, not the total tree size or gap closure.\n4. No comparison with specialized heuristics or state-of-the-art approximate methods is provided.\n5. Validation is limited to synthetic data; real-world applicability is not assessed.\n6. The method requires solving complex MI-QCQPs, which may not scale to very large instances.\n7. There is also no analysis of the trade-off between solution quality and computation time for practical use.\n8. The paper does not provide guidance on selecting which polyhedral norm constraints to include for best performance."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926851174,"tcdate":1762043190571,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16824/Reviewer_oow2"],"signatures":["ICLR.cc/2026/Conference/Submission16824/Reviewer_oow2"],"forum":"VJAqqtVXfD","number":3,"license":"CC BY 4.0","cdate":1762043190571,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16824/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926851174,"domain":"ICLR.cc/2026/Conference","replyto":"VJAqqtVXfD","id":"isTqlFO6oH","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["optimization; spatial branch-and-bound; clustering"]},"primary_area":{"value":"optimization"},"abstract":{"value":"We propose a method to solve $k$-HC$_2$—the $k$-Hyperplane Clustering problem that asks to find $k$ hyperplanes that minimize the sum of squared $2$-norm (Euclidean) distances between each point and its closest hyperplane—to global optimality via spatial branch-and-bound (SBB) techniques. Our method strengthens a mixed-integer quadratically constrained quadratic programming formulation for $k$-HC$_2$ with constraints that arise when formulating the problem in $p$-norms with $p \\ge 2$.\n\nIn particular, we show that, for every (suitably scaled) $p \\in \\mathbb{N} \\cup \\{\\infty\\}$, one obtains a variant of $k$-HC$_2$ whose optimal solutions yield lower bounds within a multiplicative approximation factor. We focus on the case of polyhedral norms where $p = 1, \\infty$ (which are disjunctive-programming representable), and prove that strengthening the original formulation by including, on top of its $2$-norm constraints, the constraints of one of the polyhedral norms leads to an SBB method where nonzero lower bounds are obtained in a number of nodes that is linear in $n$ and $k$ (rather than exponential).\n\nExperimentally, our method leads to very large speedups, reducing median solve times by up to $41\\times$ while increasing the total number of solved instances by up to $63\\%$, drastically improving the problem's solvability to global optimality."},"_bibtex":{"value":"@inproceedings{\nconiglio2026solving,\ntitle={Solving the 2-norm k-hyperplane clustering problem via multi-norm formulations},\nauthor={Stefano Coniglio},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=VJAqqtVXfD}\n}"},"title":{"value":"Solving the 2-norm k-hyperplane clustering problem via multi-norm formulations"},"pdf":{"value":"/pdf/7de066e87aa2e5f658657d48dd4a7a315ff9b426.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"coniglio|solving_the_2norm_khyperplane_clustering_problem_via_multinorm_formulations"},"authorids":{"value":["~Stefano_Coniglio2"]},"authors":{"value":["Stefano Coniglio"]}},"version":2},{"content":{"summary":{"value":"This paper investigates the relationship between initialization bias and trainability of DNNs, reconciling two theoretical perspectives:\n- Mean-Field (MF) theory — which studies gradient propagation and identifies the “edge of chaos” (EOC) as the optimal initialization boundary for stable training;\n- Initial Guessing Bias (IGB) framework — which describes how untrained networks can exhibit predictive bias (favoring one class) even before training.\n\nThe authors establish a formal equivalence between the MF and IGB frameworks, showing that the trainability boundary in MF corresponds to a biased initialization state in IGB. Contrary to the intuitive belief that the most trainable initializations should be neutral, they demonstrate theoretically and empirically that the optimal initialization is systematically biased rather than unbiased — a state they call transient deep prejudice."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"There can be some terminological confusion: It would be helpful to clarify that “bias,” “prejudice,” and “neutrality” refer to statistical asymmetry / symmetry in prediction space. Otherwise, readers may think about fairness or ethical bias."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Theoretical contribution: The paper establishes a clean mathematical connection between two previously distinct frameworks (MF and IGB), enriching both perspectives.\n2. Insight: It introduces a novel idea: bias at initialization can improve trainability; The new “prejudice-neutrality” phase view offers an intuitive explanation for initialization effects, linking bias to the dynamical stability of gradient flow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The theory is derived in the infinite-width limit and validated on small- to mid-scale settings. Its applicability to practical, large-scale deep networks (e.g., transformers with normalization and attention) is not demonstrated. Additionally, empirical evaluation focuses on synthetic and small vision datasets. It is of readers' interest to learn results on simple language tasks. \n2. Ambiguous practical relevance: While the “transient bias” insight is conceptually interesting, there is no clear recipe for practitioners (e.g., how to initialize weights to achieve the right level of prejudice). It would be helpful to add some executable takeaways.\n3. It would be helpful to introduce and compare with alternative trainability-enhancing initializations (e.g., orthogonal, LSUV, scaled ReLU, or NTK-based initializations)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933929674,"tcdate":1761964566424,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20499/Reviewer_h1rA"],"signatures":["ICLR.cc/2026/Conference/Submission20499/Reviewer_h1rA"],"forum":"75ce9hnKne","number":2,"license":"CC BY 4.0","cdate":1761964566424,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20499/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933929674,"domain":"ICLR.cc/2026/Conference","replyto":"75ce9hnKne","id":"jFJHcHLkRD","forumContent":{"TLDR":{"value":"We prove theoretically that the optimal trainable state is necessarily biased in a wide range of models."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Trainability","initial guessing bias","mean field regime","phase diagrams"]},"supplementary_material":{"value":"/attachment/75626e09e978ea4e6b4d10d504e83e848e58bb6b.zip"},"primary_area":{"value":"learning theory"},"abstract":{"value":"The statistical properties of deep neural networks (DNNs) at initialization play an important role to comprehend their trainability and the intrinsic architectural biases they possess before data exposure. Well established mean-field (MF) theories have uncovered that the distribution of parameters of randomly initialized networks strongly influences the behavior of the gradients, dictating whether they explode or vanish. Recent work has showed that untrained DNNs also manifest an initial-guessing bias (IGB), in which large regions of the input space are assigned to a single class. In this work, we provide a theoretical proof that links IGB to previous MF theories for a vast class of DNNs, showing that efficient learning is tightly connected to a network’s prejudice towards a specific class. This connection leads to a counterintuitive conclusion: the initialization that optimizes trainability is systematically biased rather than neutral."},"_bibtex":{"value":"@inproceedings{\nbassi2026when,\ntitle={When Bias Meets Trainability: Connecting Theories of Initialization},\nauthor={Alberto Bassi and Marco Baity-Jesi and Aurelien Lucchi and Carlo Albert and Emanuele Francazi},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=75ce9hnKne}\n}"},"title":{"value":"When Bias Meets Trainability: Connecting Theories of Initialization"},"pdf":{"value":"/pdf/07db85b6886df17dcc149f9d1185a5eaf7c6bc20.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"bassi|when_bias_meets_trainability_connecting_theories_of_initialization"},"authorids":{"value":["~Alberto_Bassi1","~Marco_Baity-Jesi1","~Aurelien_Lucchi1","~Carlo_Albert1","~Emanuele_Francazi1"]},"authors":{"value":["Alberto Bassi","Marco Baity-Jesi","Aurelien Lucchi","Carlo Albert","Emanuele Francazi"]}},"version":2},{"content":{"summary":{"value":"This paper implements a Lorentz-equivariant transformer and applies it to several problems in particle physics. The main contribution of the paper is not the definition of a conceptually different model. At least it doesn't claim to define a model significantly different from previously proposed models. But it provides a software library (under BSD-3-Clause-Clear license) and a nice set of numerical experiments inspired by physics problems. This is a nice contribution to the community."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please address the questions in the weaknesses section.\n\nHaving these clearly answered would improve the paper's clarity, reproducibility, and usability. \n\nI'd be willing to raise my score to 7-8 if all these aspects were clearly explained with concrete explanations and equations, since I believe this can turn into a great tool to the AI 4 physics community."},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper combines interesting ideas (imposing symmetries arising from physics using invariant theory and clifford algebras) with state-of-the-art machine learning models (transformers, Riemannian flow maching generative modeling). It comes with a software package and multiple experiments where several methods are compared. I see this as a tool that may be used by several members of the AI4science and AI4physics communities."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper is not self-contained. In order to be a valuable tool for our community it would benefit from a more comprehensive explanation of several of its technical aspects. \n\n- The implemented Lorentz equivariant model (what are the inputs?, are they a set of particles? how are they matched to queries, values and keys?)\n- The flow-matching approach. Again, what are the inputs, outputs? How's the model defined? What does it produce? The closest thing to an explanation is in page 24, in the context of describing the experiment, but still I don't find this explanation sufficient to understand the model conceptually.\n- What are the scientific goals of the experiments defined in the paper? (In the context of fundamental research in high energy physics)."},"limitations":{"value":"Not applicable."}},"nonreaders":[],"tmdate":1730878888789,"tcdate":1720735400748,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission3969/Reviewer_PeE1"],"signatures":["NeurIPS.cc/2024/Conference/Submission3969/Reviewer_PeE1"],"forum":"X34GKv8sYT","number":1,"license":"CC BY 4.0","cdate":1720735400748,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission3969/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878888789,"domain":"NeurIPS.cc/2024/Conference","replyto":"X34GKv8sYT","id":"DmEi30yI1L","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"A Lorentz-equivariant Transformer architecture plus Lorentz-equivariant flow matching for high-energy physics"},"keywords":{"value":["Geometric deep learning","equivariance","Lorentz symmetry","Transformer","flow matching","high-energy physics","particle physics"]},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Extracting scientific understanding from particle-physics experiments requires solving diverse learning problems with high precision and good data efficiency. We propose the Lorentz Geometric Algebra Transformer (L-GATr), a new multi-purpose architecture for high-energy physics. L-GATr represents high-energy data in a geometric algebra over four-dimensional space-time and is equivariant under Lorentz transformations, the symmetry group of relativistic kinematics. At the same time, the architecture is a Transformer, which makes it versatile and scalable to large systems. L-GATr is first demonstrated on regression and classification tasks from particle physics. We then construct the first Lorentz-equivariant generative model: a continuous normalizing flow based on an L-GATr network, trained with Riemannian flow matching. Across our experiments, L-GATr is on par with or outperforms strong domain-specific baselines."},"_bibtex":{"value":"@inproceedings{\nspinner2024lorentzequivariant,\ntitle={Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics},\nauthor={Jonas Spinner and Victor Breso Pla and Pim De Haan and Tilman Plehn and Jesse Thaler and Johann Brehmer},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=X34GKv8sYT}\n}"},"title":{"value":"Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics"},"pdf":{"value":"/pdf/8ddc6739ac0f4a9405f3368dba8b88248133a915.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"spinner|lorentzequivariant_geometric_algebra_transformers_for_highenergy_physics"},"authorids":{"value":["~Jonas_Spinner1","~Victor_Breso_Pla1","~Pim_De_Haan1","~Tilman_Plehn1","~Jesse_Thaler1","~Johann_Brehmer1"]},"authors":{"value":["Jonas Spinner","Victor Breso Pla","Pim De Haan","Tilman Plehn","Jesse Thaler","Johann Brehmer"]}},"version":2},{"content":{"summary":{"value":"The paper provides a new method for recovering the 3D cosmic web from galaxy weak lensing signal.\nThe proposed approach uses a coordinate-based neural field method with positional encoding. (Equation 6).\nThe reconstruction results are demonstrated in a series of experiments."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"* I would like to know some more details about the number of CPUs/GPUs used for training, memory requirements, and the training and inference time of the proposed method.\n\n* I'm also curious about the next steps. After the 3D cosmic web estimation step is done, what are the next science questions that can be answered with the new results?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"* The paper proposes a new method for an important physics problem.\n* The paper is well-written; I enjoyed reading it.\n* The physics background is exceptionally well-presented, and even non-experts can understand the main scientific ideas.\n* The proposed method works better than the leading baseline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I think this is a great paper, but in my opinion, ICLR might not be the best venue for this.\nThe new machine learning contributions are quite limited: neural field models and positional encoding have been around for some time, and I didn't find significantly new machine learning contributions in the paper.\n\nNonetheless, the proposed algorithm works well, and the physics application, estimating the 3D cosmic web, is a very important problem.\nSince the most significant contributions of the paper are in astrophysics, I think an astrophysics or cosmology journal would be a better venue for this paper."}},"nonreaders":[],"tmdate":1733109621902,"tcdate":1730700882731,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7909/Reviewer_x8ZD"],"signatures":["ICLR.cc/2025/Conference/Submission7909/Reviewer_x8ZD"],"forum":"Ax0i933gtp","number":3,"license":"CC BY 4.0","cdate":1730700882731,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7909/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733109621902,"domain":"ICLR.cc/2025/Conference","replyto":"Ax0i933gtp","id":"6wv5IlMdDF","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["computational imaging","signal processing","inverse problems","astrophysics","cosmology","neural fields","machine learning for physical sciences"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Weak gravitational lensing is the slight distortion of galaxy shapes caused primarily by the gravitational effects of dark matter in the universe. In our work, we seek to invert the weak lensing signal from 2D telescope images to reconstruct a 3D map of the universe's dark matter field. While inversion typically yields a 2D projection of the dark matter field, accurate 3D maps of the dark matter distribution are essential for localizing structures of interest and testing theories of our universe. However, 3D inversion poses significant challenges. First, unlike standard 3D reconstruction that relies on multiple viewpoints, in this case, images are only observed from a single viewpoint. This challenge can be partially addressed by observing how galaxy emitters throughout the volume are lensed. However, this leads to the second challenge: the shapes and exact locations of unlensed galaxies are unknown, and can only be estimated with a very large degree of uncertainty. This introduces an overwhelming amount of noise which nearly drowns out the lensing signal completely. Previous approaches tackle this by imposing strong assumptions about the structures in the volume. We instead propose a methodology using a gravitationally-constrained neural field to flexibly model the continuous matter distribution. We take an analysis-by-synthesis approach, optimizing the weights of the neural network through a fully differentiable physical forward model to reproduce the lensing signal present in image measurements. We showcase our method on simulations, including realistic simulated measurements of dark matter distributions that mimic data from upcoming telescope surveys. Our results show that our method can not only outperform previous methods, but importantly is also able to recover potentially surprising dark matter structures."},"_bibtex":{"value":"@inproceedings{\nzhao2025revealing,\ntitle={Revealing the 3D Cosmic Web through Gravitationally Constrained Neural Fields},\nauthor={Brandon Zhao and Aviad Levis and Liam Connor and Pratul P. Srinivasan and Katherine Bouman},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Ax0i933gtp}\n}"},"title":{"value":"Revealing the 3D Cosmic Web through Gravitationally Constrained Neural Fields"},"pdf":{"value":"/pdf/14dff818c86b7716f40ef9c613ca69f02c15287e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|revealing_the_3d_cosmic_web_through_gravitationally_constrained_neural_fields"},"authorids":{"value":["~Brandon_Zhao1","~Aviad_Levis1","~Liam_Connor1","~Pratul_P._Srinivasan1","~Katherine_Bouman1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Brandon Zhao","Aviad Levis","Liam Connor","Pratul P. Srinivasan","Katherine Bouman"]}},"version":2},{"content":{"venue":{"value":"American Journal of Physics 92, 655–662 (2024)"},"pdf":{"value":"/pdf/d6d4ec6a5b7caa933e1837799d7eebb0d43bfe0d.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"shah|data_science_education_in_undergraduate_physics_lessons_learned_from_a_community_of_practice"},"authorids":{"value":["~Karan_Shah1","butlerju@mountunion.edu","avknaub@gmail.com","anil@umd.edu","~William_Ratcliff_II1","msoltani@bu.edu"]},"abstract":{"value":"It is becoming increasingly important that physics educators equip their students with the skills to work with data effectively. However, many educators may lack the necessary training and expertise in data science to teach these skills. To address this gap, we created the Data Science Education Community of Practice (DSECOP), bringing together graduate students and physics educators from different institutions and backgrounds to share best practices and lessons learned from integrating data science into undergraduate physics education. In this article, we present insight and experiences from this community of practice, highlighting key strategies and challenges in incorporating data science into the introductory physics curriculum. Our goal is to provide guidance and inspiration to educators who seek to integrate data science into their teaching, helping to prepare the next generation of physicists for a data-driven world."},"title":{"value":"Data science education in undergraduate physics: Lessons learned from a community of practice"},"authors":{"value":["Karan Shah","Julie Butler","Alexis V. Knaub","Anıl Zenginoğlu","William Ratcliff","Mohammad Soltanieh-ha"]}},"tmdate":1777428024995,"pdate":1725174000000,"tcdate":1727817039891,"writers":["~Karan_Shah1","butlerju@mountunion.edu","avknaub@gmail.com","anil@umd.edu","william.ratcliff@nist.gov","msoltani@bu.edu"],"signatures":["~Karan_Shah1"],"forum":"84e0rn4qQa","license":"CC BY 4.0","number":30011,"cdate":1727817039891,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload","OpenReview.net/Archive/-/Edit"],"mdate":1777428024995,"domain":"OpenReview.net/Archive","id":"84e0rn4qQa","version":2},{"content":{"summary":{"value":"In this paper, a hybrid model (a model that combines machine learning with physics) is demonstrated for nowcasting and medium-range weather forecasting. WeatherGFT uses machine-learned weightings that combine two successful methods for weather forecasting (machine learning and the traditional numerical, PDE-based method). The combination of both methods allows for significantly increased temporal resolution compared to purely data-driven methods (6-hourly vs 15 minutes). This approach also provides a new framework for combining physics-based methods with data-driven methods, which allows for the ability to differentiate through the hybrid model."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"In section 4.3, I don’t understand how FourCastNet, ClimODE, or Keisler was used for nowcasting. They all have a time step of 6 hours, how is it possible to get 30-minute forecasts using frame interpolation methods? Are you using interpolation between ERA5 (e.g. the initial conditions) and a 6-hour forecast from these models? \n\nIn section 4.3, how are the forecasts including WeatherGFT initialized? Do they all use ERA5? \n\nI would like to see the effects of the PDE Kernel on other metrics. Does this inclusion of a physics kernel improve spectral bias or allow stability? What about the conservation of energy or momentum? \n\nWhat is the computational cost during inference of having a PDE kernel compared to just the transformer by itself?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Combining machine learning with physics-based approaches (e.g. hybrid modeling) is an exciting and highly researched topic for weather and climate. As mentioned in the paper, purely data-driven models are black-box systems even if they do perform well. Existing data-driven models also are too temporal coarse for some operational weather forecasting applications. This paper does a good job address both of these existing problems in the field."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The novelty of the paper is not as strong as claimed in the paper. Other hybrid models that combine machine learning with physic-based approaches exist (see Arcomano et. al. 2022 and 2023, Clark et. al. 2022, and others). The Arcomano et. al. 2022/2023 even includes a machine-learned weighting for the combination of the ML-based model and the physics-based model. \n\nThe time-embedding and multi-lead times in one model have also been demonstrated before, see Stormer (https://arxiv.org/abs/2312.03876) and MetNet (https://arxiv.org/abs/2306.06079)."},"limitations":{"value":"Overall the limitations of the paper are well laid out, however, some claims are not supported by the paper. \n\nClaims about being the first to move away from fixed lead times for data-driven weather forecasting are not true. Stormer (https://arxiv.org/abs/2312.03876) and MetNet (https://arxiv.org/abs/2306.06079) have demonstrated this previously. \n\nClaim “In addition, the prediction error of our model at the lead time of 6-hour is significantly smaller than that of the physical dynamic model ECMWF-IFS”. For Z500 this is not supported by Figure 4. \n\nFor Figure 5 the prediction of the subtropical high (it should be a subtropical ridge) isn’t convincing.  That seems to be a function of the contouring in Matplotlib as the difference plots show similar magnitude of errors. \n\nOverall, if some of these claims are toned down and the addition of the appropriate citations the manuscript will be improved."}},"nonreaders":[],"tmdate":1730878928998,"tcdate":1720807294345,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission4475/Reviewer_oEsH"],"signatures":["NeurIPS.cc/2024/Conference/Submission4475/Reviewer_oEsH"],"forum":"ioAlzcELTf","number":3,"license":"CC BY 4.0","cdate":1720807294345,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission4475/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878928998,"domain":"NeurIPS.cc/2024/Conference","replyto":"ioAlzcELTf","id":"OiIKCd3bcj","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["weather forecast","physics-AI hybrid model","partial differential equation","machine learning"]},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Data-driven artificial intelligence (AI) models have made significant advancements in weather forecasting, particularly in medium-range and nowcasting. However, most data-driven weather forecasting models are black-box systems that focus on learning data mapping rather than fine-grained physical evolution in the time dimension. Consequently, the limitations in the temporal scale of datasets prevent these models from forecasting at finer time scales. This paper proposes a physics-AI hybrid model (i.e., WeatherGFT) which generalizes weather forecasts to finer-grained temporal scales beyond training dataset. Specifically, we employ a carefully designed PDE kernel to simulate physical evolution on a small time scale (e.g., 300 seconds) and use a parallel neural networks with a learnable router for bias correction. Furthermore, we introduce a lead time-aware training framework to promote the generalization of the model at different lead times. The weight analysis of physics-AI modules indicates that physics conducts major evolution while AI performs corrections adaptively. Extensive experiments show that WeatherGFT trained on an hourly dataset, effectively generalizes forecasts across multiple time scales, including 30-minute, which is even smaller than the dataset's temporal resolution."},"_bibtex":{"value":"@inproceedings{\nxu2024generalizing,\ntitle={Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-{AI} Hybrid Modeling},\nauthor={Wanghan Xu and Fenghua Ling and Wenlong Zhang and Tao Han and Hao Chen and Wanli Ouyang and LEI BAI},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=ioAlzcELTf}\n}"},"title":{"value":"Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid Modeling"},"pdf":{"value":"/pdf/b8b7c1d96b59d20bcb7dfbdb207746fc2568d879.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"xu|generalizing_weather_forecast_to_finegrained_temporal_scales_via_physicsai_hybrid_modeling"},"authorids":{"value":["~Wanghan_Xu1","~Fenghua_Ling1","~Wenlong_Zhang3","~Tao_Han4","~Hao_Chen14","~Wanli_Ouyang1","~LEI_BAI1"]},"authors":{"value":["Wanghan Xu","Fenghua Ling","Wenlong Zhang","Tao Han","Hao Chen","Wanli Ouyang","LEI BAI"]}},"version":2},{"content":{"venue":{"value":"BMC Bioinform. 2024"},"pdf":{"value":"https://bmcbioinformatics.biomedcentral.com/counter/pdf/10.1186/s12859-024-05782-x"},"venueid":{"value":"dblp.org/journals/BMCBI/2024"},"paperhash":{"value":"giannantoni|biology_system_description_language_bisdl_a_modeling_language_for_the_design_of_multicellular_synthetic_biological_systems"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Leonardo_Giannantoni:","https://dblp.org/search/pid/api?q=author:Roberta_Bardini:","~Alessandro_Savino1","~Stefano_Di_Carlo1"]},"html":{"value":"https://doi.org/10.1186/s12859-024-05782-x"},"_bibtex":{"value":"@article{DBLP:journals/bmcbi/GiannantoniBSC24,\n  author={Leonardo Giannantoni and Roberta Bardini and Alessandro Savino and Stefano Di Carlo},\n  title={Biology System Description Language (BiSDL): a modeling language for the design of multicellular synthetic biological systems},\n  year={2024},\n  month={December},\n  cdate={1733011200000},\n  journal={BMC Bioinform.},\n  volume={25},\n  number={1},\n  pages={166},\n  url={https://doi.org/10.1186/s12859-024-05782-x}\n}\n"},"abstract":{"value":"The Biology System Description Language (BiSDL) is an accessible, easy-to-use computational language for multicellular synthetic biology. It allows synthetic biologists to represent spatiality and multi-level cellular dynamics inherent to multicellular designs, filling a gap in the state of the art. Developed for designing and simulating spatial, multicellular synthetic biological systems, BiSDL integrates high-level conceptual design with detailed low-level modeling, fostering collaboration in the Design-Build-Test-Learn cycle. BiSDL descriptions directly compile into Nets-Within-Nets (NWNs) models, offering a unique approach to spatial and hierarchical modeling in biological systems. BiSDL’s effectiveness is showcased through three case studies on complex multicellular systems: a bacterial consortium, a synthetic morphogen system and a conjugative plasmid transfer process. These studies highlight the BiSDL proficiency in representing spatial interactions and multi-level cellular dynamics. The language facilitates the compilation of conceptual designs into detailed, simulatable models, leveraging the NWNs formalism. This enables intuitive modeling of complex biological systems, making advanced computational tools more accessible to a broader range of researchers. BiSDL represents a significant step forward in computational languages for synthetic biology, providing a sophisticated yet user-friendly tool for designing and simulating complex biological systems with an emphasis on spatiality and cellular dynamics. Its introduction has the potential to transform research and development in synthetic biology, allowing for deeper insights and novel applications in understanding and manipulating multicellular systems."},"title":{"value":"Biology System Description Language (BiSDL): a modeling language for the design of multicellular synthetic biological systems"},"authors":{"value":["Leonardo Giannantoni","Roberta Bardini","Alessandro Savino","Stefano Di Carlo"]}},"tmdate":1740648441890,"pdate":1704067200000,"tcdate":1716358022435,"writers":["~"],"signatures":["~Stefano_Di_Carlo1"],"forum":"2tZqSlGxjK","license":"CC BY-SA 4.0","number":18037,"cdate":1733011200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1740648441890,"domain":"DBLP.org","id":"2tZqSlGxjK","version":2},{"content":{"summary":{"value":"The paper proposes LEAP, a learnable positional encoding for graphs based on the local Euler Characteristic Transform ($l$-ECT), which integrates both geometric and topological information and supports end-to-end training. The authors validate LEAP on synthetic tasks and multiple real-world graph datasets, demonstrating its ability to capture structural information even when node features are non-informative."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"The ablation study (Table 2 ) shows that 1-hop neighborhoods perform better than 2-hop or a 1,2-hop combination. This seems counter-intuitive to the MPNN notion that larger receptive fields are better. Why is 1-hop sufficient?\n\nHow does LEAP perform on very large graphs?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"(1) LEAP is the first work to integrate the local ECT into GNNs in an end-to-end trainable manner, combining geometric and topological insights.\n\n(2) The authors evaluate their method on synthetic data, small molecules, image graphs, and large-scale quantum chemistry datasets, showing consistent improvements. The synthetic task clearly demonstrates LEAP’s ability to classify graphs based solely on structural information, highlighting limitations of standard MPNNs.\n\n(3) The paper includes thorough ablation studies on various aspects (Projection strategies, Locality, PE dimension, and DECT hyperparameters) of LEAP, strengthening its conclusions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) As the authors acknowledge in the limitations, LEAP is not a purely structural PE. It requires node features $x(v)$ to compute the ECT. While the synthetic experiment shows LEAP works even with uninformative features, it is unclear how, or if, LEAP would be applied to graphs without any node features (e.g., graphs represented only by an adjacency matrix).\n\n(2) LEAP introduces more hyperparameters (e.g., number of directions, smoothing, discretization steps) compared to LaPE or RWPE.\n\n(3) Although not analyzed in depth, computing ECT for m-hop subgraphs for every node could be expensive for large graphs."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924239144,"tcdate":1761978223488,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13677/Reviewer_6i1u"],"signatures":["ICLR.cc/2026/Conference/Submission13677/Reviewer_6i1u"],"forum":"8XFPhByERc","number":3,"license":"CC BY 4.0","cdate":1761978223488,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13677/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924239144,"domain":"ICLR.cc/2026/Conference","replyto":"8XFPhByERc","id":"GyEVtrGknM","forumContent":{"TLDR":{"value":"Local ECT-Based Learnable Positional Encodings for Graphs"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Topology","Euler Characteristic Transform","Graph Neural Networks","Topological Data Analysis","TDA","Topological Deep Learning"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Graph neural networks (GNNs) largely rely on the message-passing paradigm, where nodes iteratively aggregate information from their neighbors. Yet, standard message passing neural networks (MPNNs) face well-documented theoretical and practical limitations. Graph positional encoding (PE) has emerged as a promising direction to address these limitations. The Euler Characteristic Transform (ECT) is an efficiently computable geometric–topological invariant that characterizes shapes and graphs. In this work, we combine the differentiable approximation of the ECT (DECT) and its local variant ($\\ell$-ECT) to propose LEAP, a new end-to-end trainable local structural PE for graphs. We evaluate our approach on multiple real-world datasets as well as on a synthetic task designed to test its ability to extract topological features. Our results underline the potential of LEAP-based encodings as a powerful component for graph representation learning pipelines."},"_bibtex":{"value":"@inproceedings{\namboage2026leap,\ntitle={{LEAP}: Local {ECT}-Based Learnable Positional Encodings for Graphs},\nauthor={Juan P Garcia Amboage and Ernst R{\\\"o}ell and Patrick Schnider and Bastian Rieck},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=8XFPhByERc}\n}"},"title":{"value":"LEAP: Local ECT-Based Learnable Positional Encodings for Graphs"},"pdf":{"value":"/pdf/c0d196a299fdc4e5569f26049b53d4770cf05761.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"amboage|leap_local_ectbased_learnable_positional_encodings_for_graphs"},"authorids":{"value":["~Juan_P_Garcia_Amboage1","~Ernst_Röell1","~Patrick_Schnider1","~Bastian_Rieck1"]},"authors":{"value":["Juan P Garcia Amboage","Ernst Röell","Patrick Schnider","Bastian Rieck"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2507.22099v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"li|runtime_failure_hunting_for_physics_engine_based_software_systems_how_far_can_we_go"},"authorids":{"value":["~Shuqing_Li2","https://dblp.org/search/pid/api?q=author:Qiang_Chen:","~Xiaoxue_Ren1","https://dblp.org/search/pid/api?q=author:Michael_R._Lyu:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2507.22099"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2507-22099,\n  publtype={informal},\n  author={Shuqing Li and Qiang Chen and Xiaoxue Ren and Michael R. Lyu},\n  title={Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?},\n  year={2025},\n  month={July},\n  cdate={1751328000000},\n  journal={CoRR},\n  volume={abs/2507.22099},\n  url={https://doi.org/10.48550/arXiv.2507.22099}\n}\n"},"abstract":{"value":"Physics Engines (PEs) are fundamental software frameworks that simulate physical interactions in applications ranging from entertainment to safety-critical systems. Despite their importance, PEs suffer from physics failures, deviations from expected physical behaviors that can compromise software reliability, degrade user experience, and potentially cause critical failures in autonomous vehicles or medical robotics. Current testing approaches for PE-based software are inadequate, typically requiring white-box access and focusing on crash detection rather than semantically complex physics failures. This paper presents the first large-scale empirical study characterizing physics failures in PE-based software. We investigate three research questions addressing the manifestations of physics failures, the effectiveness of detection techniques, and developer perceptions of current detection practices. Our contributions include: (1) a taxonomy of physics failure manifestations; (2) a comprehensive evaluation of detection methods including deep learning, prompt-based techniques, and large multimodal models; and (3) actionable insights from developer experiences for improving detection approaches. To support future research, we release PhysiXFails, code, and other materials at https://sites.google.com/view/physics-failure-detection."},"title":{"value":"Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?"},"authors":{"value":["Shuqing Li","Qiang Chen","Xiaoxue Ren","Michael R. Lyu"]}},"tmdate":1781799005140,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2507-22099"],"tcdate":1767791750944,"writers":["~"],"signatures":["~Xiaoxue_Ren1"],"forum":"ghOTzo3JFj","license":"CC BY-SA 4.0","number":728194,"cdate":1751328000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1781799005140,"domain":"DBLP.org","id":"ghOTzo3JFj","version":2},{"content":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Image Synthesis","Contrastive Learning","Cryo-EM"]},"primary_area":{"value":"generative_models"},"abstract":{"value":"In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of proteins, such as the SARS-COV-2 spike protein. To achieve high-resolution reconstruction, a comprehensive data processing pipeline has been adopted. However, its performance is still limited as it lacks high-quality annotated datasets for training. To address this, we introduce physics-informed generative cryo-electron microscopy (CryoGEM), which for the first time integrates physics-based cryo-EM simulation with a generative unpaired noise translation to generate physically correct synthetic cryo-EM datasets with realistic noises. Initially, CryoGEM simulates the cryo-EM imaging process based on a virtual specimen. To generate realistic noises, we leverage an unpaired noise translation via contrastive learning with a novel mask-guided sampling scheme. Extensive experiments show that CryoGEM is capable of generating authentic cryo-EM images. The generated dataset can be used as training data for particle picking and pose estimation models, eventually improving the reconstruction resolution."},"_bibtex":{"value":"@inproceedings{\nzhang2024cryogem,\ntitle={Cryo{GEM}: Physics-Informed Generative Cryo-Electron Microscopy},\nauthor={Jiakai Zhang and Qihe Chen and Yan Zeng and Wenyuan Gao and Xuming He and Zhijie Liu and Jingyi Yu},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=edOZifvwMi}\n}"},"title":{"value":"CryoGEM: Physics-Informed Generative Cryo-Electron Microscopy"},"pdf":{"value":"/pdf/28c9b2115b82c3166b55e72a5bbf65dc9b1611fc.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"zhang|cryogem_physicsinformed_generative_cryoelectron_microscopy"},"authorids":{"value":["~Jiakai_Zhang3","~Qihe_Chen2","~Yan_Zeng3","~Wenyuan_Gao1","~Xuming_He3","~Zhijie_Liu3","~Jingyi_Yu5"]},"authors":{"value":["Jiakai Zhang","Qihe Chen","Yan Zeng","Wenyuan Gao","Xuming He","Zhijie Liu","Jingyi Yu"]}},"tmdate":1736150876687,"pdate":1727287817205,"tcdate":1715596755867,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission6516/Authors"],"signatures":["NeurIPS.cc/2024/Conference/Submission6516/Authors"],"forum":"edOZifvwMi","license":"CC BY-NC 4.0","number":6516,"cdate":1715596755867,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/-/Submission","NeurIPS.cc/2024/Conference/-/Edit","NeurIPS.cc/2024/Conference/-/Post_Submission","NeurIPS.cc/2024/Conference/Submission6516/-/Revision","NeurIPS.cc/2024/Conference/Submission6516/-/Camera_Ready_Revision"],"mdate":1736150876687,"odate":1730873894161,"domain":"NeurIPS.cc/2024/Conference","id":"edOZifvwMi","version":2},{"content":{"venue":{"value":"SIGGRAPH 2025"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"xu|parc_physicsbased_augmentation_with_reinforcement_learning_for_character_controllers"},"authorids":{"value":["~Michael_Xu5","~Yi_Shi1","~KangKang_Yin2","~Xue_Bin_Peng1"]},"html":{"value":"https://arxiv.org/abs/2505.04002"},"abstract":{"value":"Humans excel in navigating diverse, complex environments with agile motor skills, exemplified by parkour practitioners performing dynamic maneuvers, such as climbing up walls and jumping across gaps. Reproducing these agile movements with simulated characters remains challenging, in part due to the scarcity of motion capture data for agile terrain traversal behaviors and the high cost of acquiring such data. In this work, we introduce PARC (Physics-based Augmentation with Reinforcement Learning for Character Controllers), a framework that leverages machine learning and physics-based simulation to iteratively augment motion datasets and expand the capabilities of terrain traversal controllers. PARC begins by training a motion generator on a small dataset consisting of core terrain traversal skills. The motion generator is then used to produce synthetic data for traversing new terrains. However, these generated motions often exhibit artifacts, such as incorrect contacts or discontinuities. To correct these artifacts, we train a physics-based tracking controller to imitate the motions in simulation. The corrected motions are then added to the dataset, which is used to continue training the motion generator in the next iteration. PARC's iterative process jointly expands the capabilities of the motion generator and tracker, creating agile and versatile models for interacting with complex environments. PARC provides an effective approach to develop controllers for agile terrain traversal, which bridges the gap between the scarcity of motion data and the need for versatile character controllers."},"title":{"value":"PARC: Physics-based Augmentation with Reinforcement Learning for Character Controllers"},"authors":{"value":["Michael Xu","Yi Shi","KangKang Yin","Xue Bin Peng"]}},"tmdate":1776907725793,"pdate":1754035200000,"tcdate":1763363884778,"writers":["mxa23@sfu.ca","~Yi_Shi1","~KangKang_Yin2","~Xue_Bin_Peng1"],"signatures":["~Yi_Shi1"],"forum":"epzNvDksKX","license":"CC BY 4.0","number":41455,"cdate":1763363884778,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload","OpenReview.net/Archive/-/Edit"],"mdate":1776907725793,"domain":"OpenReview.net/Archive","id":"epzNvDksKX","version":2},{"content":{"summary":{"value":"This paper introduces  a novel Complex-Aware RNA Design framework that addresses a critical limitation in existing methods by explicitly incorporating the protein-RNA interaction context into the inverse folding process. The core technical contributions include the Complex-Aware Transformer (CAFormer) and a biophysics-based distance-aware filtering mechanism. The paper further proposes a robust, scalable high-affinity design pipeline integrating sequence generation, affinity prediction, and structural validation, offering a computational alternative to resource-intensive experimental screening. Benchmarking against current baselines on challenging datasets (PRI30K, PRA201) demonstrates its superior performance and effectiveness."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Could the authors elaborate on how the ensemble scheme specifically leverages the \"respective advantages\" of the three lightweight models to ensure high screening accuracy? Furthermore, what is the rationale for not integrating a single, proven, high-accuracy predictor (such as Boltz2) directly into the pipeline, and have you benchmarked the resulting screening performance difference?\n\n2. Given that interface geometric complementarity is crucial for protein-RNA recognition, the current use of ESM-2 only encodes protein sequence information. Have the authors considered integrating models (such as SAProt) that explicitly utilize both protein sequence and structural information to potentially enrich the binding mode representation and further enhance the design fidelity?\n\nAll other questions reiterate the items listed under Weaknesses. I am very willing to increase my overall score if the authors successfully address these concerns."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is the first to systematically and successfully use the complete \"protein-RNA complex\" as the core design context, rather than the isolated RNA structure. This shift aligns far better with biological reality.\n\n2. The integrated high-affinity design framework seamlessly connects sequence generation, affinity evaluation, and structural validation. This comprehensive approach provides a powerful and computationally scalable solution for identifying functional sequences.\n\n3. The method demonstrates performance improvements on both the standard benchmark (PRI30K) and, critically, on the more challenging blind test set (PRA201), proving the method's effectiveness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The distance-aware filtering mechanism, while conceptually well-grounded in biophysics, empirically yields a slightly lower overall recovery rate than the simpler greedy distance selection (as shown in Table 3). The authors could address this trade-off by providing a deeper mechanistic explanation for this counter-intuitive result and clearly justifying the necessity of the complex biophysics-based approach when a simpler strategy performs better empirically.\n\n2. The affinity predictor, a cornerstone of the high-affinity screening pipeline, is currently only validated through internal cross-validation. A comprehensive external benchmark against contemporary state-of-the-art affinity prediction models is crucial to objectively establish its competitiveness\n\n3. The significant performance drop observed in Fold 3 compared to other cross-validation folds strongly suggests model sensitivity to data partitioning or distribution differences. This instability is concerning as it could introduce large, undesirable variance when the framework is applied in high-throughput, real-world screening contexts."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915907808,"tcdate":1761854602075,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1844/Reviewer_3Qg2"],"signatures":["ICLR.cc/2026/Conference/Submission1844/Reviewer_3Qg2"],"forum":"uZBKSOJftu","number":3,"license":"CC BY 4.0","cdate":1761854602075,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1844/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915907808,"domain":"ICLR.cc/2026/Conference","replyto":"uZBKSOJftu","id":"bXNTTghz4w","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"In this paper, we propose the Complex-Aware tertiary structure-based RNA Design model, CARD, that integrates complex-level information to enhance tertiary structure-based RNA sequence design."},"keywords":{"value":["tertiary structure-based RNA design","complex-aware attention","distance-aware filtering"]},"supplementary_material":{"value":"/attachment/60c6994afb367bf56ef2a7dfcbb14cb58ecea449.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Tertiary structure-based RNA design plays a crucial role in synthetic biology and therapeutics. While existing methods have explored structure-to-sequence mappings, they focus solely on RNA structures and overlook the role of complex-level information, which is crucial for effective RNA design. To address this limitation, we propose the Complex-Aware tertiary structure-based RNA Design model, CARD, that integrates complex-level information to enhance tertiary structure-based RNA sequence design. To be specific, our method incorporates protein features extracted by protein language model (e.g., ESM-2), enabling the design model to generate more accurate and complex relevant sequences. Considering the biological complexity of protein-RNA interactions, we introduce a distance-aware filtering for local features from protein representation. Furthermore, we design a high-affinity design framework that combines our CARD with an affinity evaluation model. In this framework, candidate RNA sequences are generated and rigorously screened based on affinity and structural alignment to produce high-affinity RNA sequences. Extensive experiments demonstrate the effectiveness of our method with an improvement of 7.3% compared with base model without our complex-aware feature integration. A concrete case study for 2LBS further validates the superiority of our CARD."},"_bibtex":{"value":"@misc{\nzhang2026beyond,\ntitle={Beyond {RNA} Structure Alone: Complex-Aware Fusion for Tertiary Structure-based {RNA} Design},\nauthor={Zixun Zhang and Jiayou Zheng and Yuzhe Zhou and Sheng Wang and Shuguang Cui and Zhen Li},\nyear={2026},\nurl={https://openreview.net/forum?id=uZBKSOJftu}\n}"},"title":{"value":"Beyond RNA Structure Alone: Complex-Aware Fusion for Tertiary Structure-based RNA Design"},"pdf":{"value":"/pdf/39bc3f837fd36d3b7b48e12eda09306411956af8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|beyond_rna_structure_alone_complexaware_fusion_for_tertiary_structurebased_rna_design"},"authorids":{"value":["~Zixun_Zhang1","~Jiayou_Zheng1","~Yuzhe_Zhou1","~Sheng_Wang9","~Shuguang_Cui1","~Zhen_Li6"]},"authors":{"value":["Zixun Zhang","Jiayou Zheng","Yuzhe Zhou","Sheng Wang","Shuguang Cui","Zhen Li"]}},"version":2},{"content":{"venue":{"value":"CAD/Graphics 2011"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/6062616/6062747/06062816.pdf"},"venueid":{"value":"dblp.org/conf/CADGRAPHICS/2011"},"paperhash":{"value":"lin|application_of_penbased_planar_haptic_interface_in_physics_education"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Liping_Lin:","https://dblp.org/search/pid/api?q=author:Yongtian_Wang:","~Yue_Liu1","https://dblp.org/search/pid/api?q=author:Makoto_Sato:"]},"html":{"value":"https://doi.org/10.1109/CAD/Graphics.2011.53"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cadgraphics/LinWLS11,\n  author={Liping Lin and Yongtian Wang and Yue Liu and Makoto Sato},\n  title={Application of Pen-Based Planar Haptic Interface in Physics Education},\n  year={2011},\n  cdate={1293840000000},\n  pages={375-378},\n  url={https://doi.org/10.1109/CAD/Graphics.2011.53},\n  booktitle={CAD/Graphics},\n  crossref={conf/cadgraphics/2011}\n}\n"},"abstract":{"value":"This paper proposes a pen-based interaction system with 2D co-located haptic and visual feedback for physics education. The system combines a 3DOF pen-based planar haptic device and simulated 2D physical world based on physical engine Box 2D and 2D game engine HGE. The inputs of haptic device are translation and rotation of pen tip and the outputs are translational and rotational force displayed at pen tip. In addition, visual display and haptic display are coincident and well integrated. With this system, user can create and design simulated physical world by drawing, selecting, moving or rotating objects on screen and setting their physical properties in a natural and intuitive way. This system provides an entertaining and cartoony tool for designing and carrying out physics experiments, and will greatly promote students' learning interest and creativity."},"title":{"value":"Application of Pen-Based Planar Haptic Interface in Physics Education"},"authors":{"value":["Liping Lin","Yongtian Wang","Yue Liu","Makoto Sato"]}},"tmdate":1730984984812,"pdate":1293840000000,"tcdate":1730984913108,"writers":["~"],"signatures":["~Yue_Liu1"],"forum":"Hz3EZh5ytc","license":"CC BY-SA 4.0","number":173877,"cdate":1293840000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1730984984812,"domain":"DBLP.org","id":"Hz3EZh5ytc","version":2},{"content":{"summary":{"value":"This paper investigates the role of synthetic data for LM pretraining.\n\n1. The main observation/finding is:\n\n> The mixture of synthetic data and real data is worse than real data only, which is evaluated with perplexity. (trained with Dolma + x% of Cosmopedia, and eval with wikitext-103, Redpajama, RefineWeb, C4-en)\n\n2. A new data synthesis method, TOEDIT, which resamples tokens with lower perplexity using a language model. The authors claim TOEDIT outperforms real data in continual pretraining and SFT scenarios."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"N/A"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"The analysis of Cosmopedia synthetic data is comprehensive and provides valuable insights for future research\n\nThe theoretical framework is interesting and well-developed\n\nThe paper addresses an important question in the field about the utility of synthetic data in pretraining"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Major Concerns**\n\n1. Significant Literature Gap and Contradictory Findings \n- The paper overlooks a crucial published work [1] that directly contradicts its main findings\n- [1] demonstrates that mixing synthetic and real data improves perplexity across various domains in the Pile and enhances downstream task performance\n- This fundamental contradiction needs to be addressed and explained\n\nFigure2 in [1] shows that by mixing the synthetic data with the real C4 data, LM can get lower perplexity on various domain corpus of the Pile. [1] also shows the introduce of synthetic data gave better results on downstream tasks. However, this submission claims the mixture of data is worse than real data only, by evaluating the perplexity. (I suspect the evaluation is incorrect which lead to this contrast observation. My 2nd concern below will detail more)\n\n2. Flawed Experimental Design for Main Claims \n\nThe pretraining real data is Dolma, which consist of wikipedia, C4. Hence, the more real data you use, the lower perplexity on Wikitext-103 and C4-en you should expect. So it is domain mismatching problem, instead of saying synthetic data is not useful. \nOn the other hand, the synthetic pretraining data Cosmpopedia was build from LLM's rephrase of web-crawled data, which means Cosmopedia's data distribution differs significantly from web-crawled data, making the comparison with RedPajama and RefineWeb potentially unfair.  I would suggest to use proper out of domain text for perplexity evaluation, and also introduce the downstream task evaluation.  I would suggest to follow the evaluation framework from [1] for better comparison\n\n3. TOEDIT Methodology Concerns\n- The non-autoregressive token replacement approach may compromise text coherence and grammatical correctness\n- No discussion of how grammatical consistency is maintained during token resampling\n\n4. Weak Empirical Results (Baseline performance is concerningly low)\n- MMLU: baseline (26.56) vs. proposed (23.63) vs. random (25)\n- WinoGrande and Hellaswag results near random guessing. Improvements are marginal (1.3, -0.08) and not statistically significant\n\n\n**Minor Concern**\n\n1.  The conclusion drawn from \"75% words under 0.6 probability\" (Lines 265-267) ignores fundamental linguistic entropy\n- English typically has 10-11 bits per word entropy\n- The observed probability distribution may be natural rather than indicating filtering potential\n\n\n[1][Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling](https://aclanthology.org/2024.acl-long.757) (Maini et al., ACL 2024)"}},"nonreaders":[],"tmdate":1731428152144,"tcdate":1730847197725,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11413/Reviewer_wNZi"],"signatures":["ICLR.cc/2025/Conference/Submission11413/Reviewer_wNZi"],"forum":"mVCcWCjeEz","number":4,"license":"CC BY 4.0","cdate":1730847197725,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11413/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428152144,"domain":"ICLR.cc/2025/Conference","replyto":"mVCcWCjeEz","id":"Zs0HXe2p79","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic data","model collapse"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We explore model collapse caused by synthetic data, where AI models trained on such data experience a gradual decline in performance. \nOur initial analysis examines language model pretraining on mixed human and synthetic data, highlighting performance degradation. Further statistical analysis reveals distributional shifts and an over-concentration of n-gram features caused by synthetic data. Inspired by these insights, we propose token-level editing on human data, to obtain semi-synthetic data instead of fully using model outputs. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conducted extensive experiments on pretraining, continual pretraining, and supervised fine-tuning of language models. The results validate our theoretical proof that token-level editing improves data quality and enhances model performance."},"_bibtex":{"value":"@misc{\nzhu2025toedit,\ntitle={ToEdit: How to Synthesize Text Data to Avoid Model Collapse?},\nauthor={Xuekai Zhu and Daixuan Cheng and Hengli Li and Kaiyan Zhang and Ermo Hua and Xingtai Lv and Ning Ding and Zhouhan Lin and Zilong Zheng and Bowen Zhou},\nyear={2025},\nurl={https://openreview.net/forum?id=mVCcWCjeEz}\n}"},"title":{"value":"ToEdit: How to Synthesize Text Data to Avoid Model Collapse?"},"pdf":{"value":"/pdf/8e156410d7c4f8ff0481b777f9bf4b79fb3eb266.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhu|toedit_how_to_synthesize_text_data_to_avoid_model_collapse"},"authorids":{"value":["~Xuekai_Zhu1","~Daixuan_Cheng1","~Hengli_Li1","~Kaiyan_Zhang1","~Ermo_Hua1","~Xingtai_Lv1","~Ning_Ding5","~Zhouhan_Lin1","~Zilong_Zheng1","~Bowen_Zhou8"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xuekai Zhu","Daixuan Cheng","Hengli Li","Kaiyan Zhang","Ermo Hua","Xingtai Lv","Ning Ding","Zhouhan Lin","Zilong Zheng","Bowen Zhou"]}},"version":2},{"content":{"summary":{"value":"The paper establishes connections between value functions used in optimal control and properties of path signatures, which are a powerful representation of paths with useful algebraic properties. Furthermore, they introduce a novel control framework called signature control, which efficiently generalizes the Bellman equation to the space of trajectories. This framework can naturally handle varying/adaptive time steps, propagate higher-level information more efficiently than value function updates, and is robust to dynamical system misspecification over long rollouts. They also devise a model predictive control method for path tracking within the signature control framework. This method generalizes integral control and is suitable for problems with unknown disturbances. The proposed algorithms are tested in simulation with differentiable physics models, including control and robotics tasks."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"The strengths of the paper are as follows:\n\n1) The paper establishes connections between value functions used in optimal control and properties of path signatures, providing a novel approach to trajectory following.\n\n2) The proposed signature control framework efficiently generalizes the Bellman equation to the space of trajectories, offering advantages such as handling varying/adaptive time steps and propagating higher-level information more efficiently than value function updates . The framework is robust to dynamical system misspecification over long rollouts, making it suitable for real-world applications . The paper presents a model predictive control method for path tracking within the signature control framework, which generalizes integral control and can handle problems with unknown disturbances. \n\n3) The algorithms proposed in the paper are tested in simulation with differentiable physics models, including control and robotics tasks, demonstrating their practical applicability. \n\n4) The paper offers a general approach that lifts the classical Bellman-based dynamic programming to the space of paths, providing more flexibility in devising algorithms without the need for problem-specific modifications"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) The paper lacks a detailed discussion on the limitations and potential challenges of the signature control framework and the proposed algorithms. It would be beneficial to address any potential drawbacks or scenarios where the framework may not perform optimally. \n\n2) The evaluation of the proposed algorithms is limited to simulation experiments with differentiable physics models. While this provides initial validation, it would be valuable to include real-world experiments or comparisons with existing control methods to demonstrate the effectiveness and practicality of the signature control framework. \n\n3) The paper does not provide a comprehensive analysis of the computational complexity or scalability of the signature control framework. It would be helpful to discuss the computational requirements and potential limitations when applying the framework to more complex and larger-scale problems."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"The paper could benefit from a more thorough discussion on the generalizability of the signature control framework to a wider range of control and robotics tasks. This would provide insights into the applicability of the framework beyond the specific tasks considered in the simulations."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636706401,"tcdate":1698813559820,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6382/Reviewer_ZzLg"],"signatures":["ICLR.cc/2024/Conference/Submission6382/Reviewer_ZzLg"],"forum":"8lLaS1ekDA","number":1,"license":"CC BY 4.0","cdate":1698813559820,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6382/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636706401,"domain":"ICLR.cc/2024/Conference","replyto":"8lLaS1ekDA","id":"mH0q6M0DgT","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Decision making","Path signature","Bellman equation","Integral control","Model predictive control","Robotics"]},"supplementary_material":{"value":"/attachment/31e087a732bf849de7221b5f837a4c104c4b0288.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Path signatures have been proposed as a powerful representation of paths that efﬁciently captures the path’s analytic and geometric characteristics, having useful algebraic properties including fast concatenation of paths through tensor products. Signatures have recently been widely adopted in machine learning problems for time series analysis. In this work we establish connections between value functions typically used in optimal control and intriguing properties of path signatures. These connections motivate our novel control framework with signature transforms that efﬁciently generalizes the Bellman equation to the space of trajectories. We analyze the properties and advantages of the framework, termed signature control. In particular, we demonstrate that (i) it can naturally deal with varying/adaptive time steps; (ii) it propagates higher-level information more efﬁciently than value function updates; (iii) it is robust to dynamical system misspeciﬁcation over long rollouts. As a speciﬁc case of our framework, we devise a model predictive control method for path tracking. This method generalizes integral control, being suitable for problems with unknown disturbances. The proposed algorithms are tested in simulation, with differentiable physics models including typical control and robotics tasks such as point-mass, curve following for an ant model, and a robotic manipulator."},"_bibtex":{"value":"@misc{\nohnishi2024signatures,\ntitle={Signatures Meet Dynamic Programming: Generalizing Bellman Equations for Trajectory Following},\nauthor={Motoya Ohnishi and Iretiayo Akinola and Jie Xu and Ajay Mandlekar and Fabio Ramos},\nyear={2024},\nurl={https://openreview.net/forum?id=8lLaS1ekDA}\n}"},"title":{"value":"Signatures Meet Dynamic Programming: Generalizing Bellman Equations for Trajectory Following"},"pdf":{"value":"/pdf/1735cec408b3b094754c527c1e255ea55d813fce.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"ohnishi|signatures_meet_dynamic_programming_generalizing_bellman_equations_for_trajectory_following"},"authorids":{"value":["~Motoya_Ohnishi1","~Iretiayo_Akinola1","~Jie_Xu7","~Ajay_Mandlekar1","~Fabio_Ramos1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Motoya Ohnishi","Iretiayo Akinola","Jie Xu","Ajay Mandlekar","Fabio Ramos"]}},"version":2},{"content":{"venue":{"value":"ASONAM 2021"},"venueid":{"value":"dblp.org/conf/ASUNAM/2021"},"paperhash":{"value":"polychronopoulou|distinguishability_of_graphs_a_case_for_quantuminspired_measures"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Athanasia_Polychronopoulou:","https://dblp.org/search/pid/api?q=author:Jumanah_Alshehri:","~Zoran_Obradovic1"]},"html":{"value":"https://doi.org/10.1145/3487351.3488330"},"_bibtex":{"value":"@inproceedings{DBLP:conf/asunam/Polychronopoulou21,\n  author={Athanasia Polychronopoulou and Jumanah Alshehri and Zoran Obradovic},\n  title={Distinguishability of graphs: a case for quantum-inspired measures},\n  year={2021},\n  cdate={1609459200000},\n  pages={48-55},\n  url={https://doi.org/10.1145/3487351.3488330},\n  booktitle={ASONAM},\n  crossref={conf/asunam/2021}\n}\n"},"abstract":{"value":"The question of graph similarity or graph distinguishability arises often in natural systems and their analysis over graphical networks. In many domains, graph similarity is used for graph classification, outlier detection or the identification of distinguished interaction patterns. Several methods have been proposed on how to address this topic, but graph comparison still presents many challenges. Recently, information physics has emerged as a promising theoretical foundation for complex networks. In many applications, it has been demonstrated that natural complex systems exhibit features that can be described and interpreted by measures typically applied in quantum mechanical systems. Therefore, a natural starting point for the identification of network similarity measures is information physics and a series of measures of distance for quantum states. In this work, we report experiments on synthetic and real-world data sets, and compare quantum-inspired measures to a series of state-of-the-art and well-established methods of graph distinguishability. We show that quantum-inspired methods satisfy the mathematical and intuitive requirements for graph similarities, while offering high interpretability."},"title":{"value":"Distinguishability of graphs: a case for quantum-inspired measures"},"authors":{"value":["Athanasia Polychronopoulou","Jumanah Alshehri","Zoran Obradovic"]}},"tmdate":1738866305624,"pdate":1609459200000,"tcdate":1738866235833,"writers":["~"],"signatures":["~Zoran_Obradovic1"],"forum":"kg2jeMvZ71","license":"CC BY-SA 4.0","number":309855,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1738866305624,"domain":"DBLP.org","id":"kg2jeMvZ71","version":2},{"content":{"summary":{"value":"This paper introduces a new technique for doing synthetic task augmentation and curriculum generation for RL. The proposed augmentation method \"h1\" simply composes multiple short-horizon problems by using the output of previous problems as variables for future ones. The empirical evaluation is focused on a Qwen2.5 3B model trained on a composable version of GSM8K and evaluated across other math domains."},"soundness":{"value":1},"confidence":{"value":3},"questions":{"value":"I would appreciate it if the authors could address the areas of improvement and questions raised in the  Weaknesses section of my review, and clarify any potential misconceptions for the discussion phase."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Synthetic data augmentation for long-horizon reasoning is an under-explored, but very relevant research direction, as we want to scale RL to increasingly complex domains.\n- Overall, the paper is well written and reads smoothly. I found the provided content and method description clear and intuitive.\n- I found the cost analysis result in Section 6 insightful, providing interesting results on the tradeoff between cost and performance for curriculum learning, something I have not seen in prior work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My **main concerns** about this work are that I found a large divide between the main claims made, the methodology, and the actual empirical results:\n1) The paper makes several resounding claims, e.g., that their method \"provides scalable synthetic long-horizon data, with explicit control over the horizon length and complexity without the need for new annotations.\" (lines 78-80). Yet even in the GSM8K settings, the proposed method requires annotating where variables of future claims can be augmented with previous answers, which is something extraneous to traditional RL datasets. Moreover, it is unclear to me how h1 would be applicable to domains beyond ones with numeric answers (e.g., game playing, coding, or scientific discovery, which the authors themselves mentioned in lines 38-39).\n2) The current main results focus on a 3B Qwen 2.5 model on the authors' own synthetic version of GSM8K for multi-hop problems with limited output context. If I am interpreting Table 1 correctly, gains from h1 on this task appear only on the synthetic examples, with performance seemingly lower than regular compute matched non-augmented training L1 on the original set of problems (column 1 of Table 1). While the performance improvements on these longer synthetic problems are interesting, I do not think this result alone provides evidence for generalized \"substantial performance gains on multi-step reasoning tasks\" (lines 475-476), which I think would require testing the h1 model on other explicitly multi-hop problems across different domains after training.\n3) When evaluating on a broader set of tasks in Table 2, the only baseline reported is the author's own RLVR implementation without *compute matching*. I am not sure why the other \"compute matched\" baselines from Table 1 are not also reported in Table 2.  Also, for the 7b Qwen experiments, I could not find the results on AIME@16. Given the 7B model has been widely used by prior reasoning works this would allow for comparison with a wider range of baselines. I would suggest expanding the range of evaluation tasks to include Minerva Math, OlympiadBench, and AMC as well, matching the evaluation setting of Dr.GRPO, and including this method as an independently trained baseline. Since Dr.GRPO itself is the method used by the authors for optimization, I think adding it as an independent baseline would allow readers to put into better perspective the actual impact of the methodology. At the moment, I do not think that claims like \"while improvements obtained from RLVR on standard data is bounded by the base model’s capabilities, our method performs significantly better\" (lines 76-78)or \"Furthermore, our results show that the model learns genuinely new reasoning capabilities, rather than just refining existing ones.\" (lines 478-479) are justified, given that all absolute performance score reported with the 3B model for the reasoning tasks (e.g., AIME, MATH) are quite low and well below prior reasoning work.\n\n\n**Other concerns**:\n1) In Algorithm 1, it is described that at each training stage, the model's previous solution $y_h$ is used to construct future prompts, prepended in the dataset, for later curriculum iterations (Algorithm line 8). Thus, it is not clear to me how the \"uniform mix\" and \"only long\" baselines' data is constructed. Is this using ground-truth solutions perhaps? I could not find this information in the main text. Additionally, it is not entirely clear to me how compute was \"equated\" for these baselines. Since the main algorithm was run for 200 iterations at each stage (up to 5 stages), for how many iterations were the other baselines run for?\n\n2) Qwen 2.5 3B instruct has not been post-trained with RL (unlike instruct models from the Qwen 3 family). Models from the Qwen 2.5 family have been shown to improve significantly with random rewards, and have often been tied to data leakage in the base model (e.g., [1]). To this end, I do not think that the sentence \"Improving an Instruct model with RL is generally considered more difficult and gains signify performance improvements beyond just instruction tuning, which cannot be directly inferred for improvements on base models\" (ln 250 253) is an accurate depiction of current findings.\n\n\n3) A key component emerging in reasoning models is their ability to self-verify, correct, and re-attempt problems at any point of their reasoning chain [2]. The analysis in Appendix C, initiated in Section 3 seems to ignore this (e.g., the Equation in line 185, assuming intermediate findings are fixed) and makes other assumptions not applicable to the \"single-turn\" evaluation tasks chosen, that I do not think justify some of the strong definitive statements in the paper (e.g., from the abstract \"Theoretically, we show that curriculum-based RL with outcome rewards achieves an exponential improvement in sample complexity over full-horizon training\"). I would recommend toning down these claims (e.g., achieves -> \"could achieve\") and explicitly mentioning that these are made under several simplifying assumptions.\n\n4) Looking at Table 4 in Appendix E, it seems the optimization was run with a quite low learning rate and gradient clipping magnitude. Were these choices due to a hyperparameter sweep and/or some instability observations? I could also not find this information in the paper.\n\n5) In Appendix E, the authors also mention that they evaluated their models every 50 steps for each training horizon and selected the checkpoint with the highest performance on a separate validation set. I found this quite unusual, as reasoning RL methods train on less than a single epoch anyway, meaning validation and training performance should match, and report results with their final checkpoint. I think it would be important to clarify exactly what validation data is being used and whether it is from a different distribution than the training set. Additionally, I think validation curves and the collected performances every 50 steps for both the h1 model and all baselines should be added to the text, as this information is currently missing.\n\n[1] Shao, Rulin, et al. \"Spurious rewards: Rethinking training signals in rlvr.\" arXiv preprint arXiv:2506.10947 (2025).\n\n[2] Guo, Daya, et al. \"Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.\" arXiv preprint arXiv:2501.12948 (2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942097572,"tcdate":1761579059580,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22165/Reviewer_WYvB"],"signatures":["ICLR.cc/2026/Conference/Submission22165/Reviewer_WYvB"],"forum":"SP7W06tphZ","number":1,"license":"CC BY 4.0","cdate":1761579059580,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22165/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942097572,"domain":"ICLR.cc/2026/Conference","replyto":"SP7W06tphZ","id":"NOLnremYNc","forumContent":{"TLDR":{"value":"We develop a method to improve the long-horizon reasoning capabilities of LLMs by scaling RL using only short-horizon data"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["long-horizon training","reasoning","LLMs","post-training","reinforcement learning"]},"supplementary_material":{"value":"/attachment/934b7e79570867afddbd40ca091a33e50d4dbdd4.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or costly step-level supervision, neither of which is scalable. In this work, we introduce a scalable method to bootstrap long-horizon reasoning capabilities using only existing, abundant short-horizon data. Our approach synthetically composes simple problems into complex, multi-step dependency chains of arbitrary length. We then train models on this data using outcome-only rewards under a curriculum that automatically increases in complexity, allowing RL training to be scaled much further without saturating. Empirically, our method generalizes remarkably well: curriculum training on composed 6th-grade math problems (GSM8K) boosts accuracy on unseen, Olympiad-level benchmarks (AIME) by up to 2.65x. Importantly, our long-horizon improvements are significantly higher than baselines even at high pass@k, showing that models can learn entirely new reasoning paths under RL. Theoretically, we show that curriculum-based RL with outcome rewards achieves an exponential improvement in sample complexity over full-horizon training, comparable to the gains from dense supervision, while providing strong training signal without additional annotations. h1 therefore introduces an efficient path towards scaling RL for longer horizons using existing data."},"_bibtex":{"value":"@misc{\nivanova2026h,\ntitle={h1: Bootstrapping Models to Reason over Longer Horizons via Reinforcement Learning},\nauthor={Alesia Ivanova and Sumeet Ramesh Motwani and Ziyang Cai and Philip Torr and Riashat Islam and Shital Shah and Christian Schroeder de Witt and Charles London},\nyear={2026},\nurl={https://openreview.net/forum?id=SP7W06tphZ}\n}"},"title":{"value":"h1: Bootstrapping Models to Reason over Longer Horizons via Reinforcement Learning"},"pdf":{"value":"/pdf/2583e51b8dcd4cb593d9e6ae4759bf80daedb354.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"ivanova|h1_bootstrapping_models_to_reason_over_longer_horizons_via_reinforcement_learning"},"authorids":{"value":["~Alesia_Ivanova2","~Sumeet_Ramesh_Motwani1","~Ziyang_Cai1","~Philip_Torr1","~Riashat_Islam1","~Shital_Shah1","~Christian_Schroeder_de_Witt1","~Charles_London1"]},"authors":{"value":["Alesia Ivanova","Sumeet Ramesh Motwani","Ziyang Cai","Philip Torr","Riashat Islam","Shital Shah","Christian Schroeder de Witt","Charles London"]}},"version":2},{"content":{"summary":{"value":"This paper presents a hybrid modeling method, which combines machine learning and physics, for a mechanical problem. This framework is explainable. Using 147 pasta buckling experiments, the authors trained an XGBoost model combining raw geometric features with a physics-derived term. The empirical results have shown that the model achieved good accuracy and surpassed the classical formula."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How does this method scale to larger datasets or higher-dimensional data? \n\n- Why did the authors not benchmark against extended buckling models (e.g., Timoshenko beam theory or imperfection-sensitive formulations)?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- This paper investigates an interesting problem that uses domain knowledge and a machine learning method. \n\n- This paper is easy to follow. \n\n- This work is a generalizable framework for physics-ML hybrid modeling of complex materials and may potentially influence future research in material science."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- This paper is of some value to specific domains. Its contributions are more on the domain science side rather than the machine learning side. I feel like this paper might be better suited for a materials science or applied mechanics venue. \n\n- The dataset used to evaluate the model's performance only includes 147 samples. It is difficult to see if this method is generalizable to different tasks or domains. Also, the experiment part is over-simplistic. I expect to see more baseline comparison, ablation study, different machine learning architectures, etc."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941912934,"tcdate":1761881656570,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21738/Reviewer_fS7v"],"signatures":["ICLR.cc/2026/Conference/Submission21738/Reviewer_fS7v"],"forum":"uMcI7jAe8Q","number":2,"license":"CC BY 4.0","cdate":1761881656570,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21738/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941912934,"domain":"ICLR.cc/2026/Conference","replyto":"uMcI7jAe8Q","id":"HYx2xG00F0","forumContent":{"TLDR":{"value":"Physics-informed, interpretable boosting learns residual corrections to Euler buckling; SHAP reveals a boundary-condition effect."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["explainable AI (XAI)","machine learning","SHAP","structural mechanics","Extreme Gradient Boosting (XGBoost)","buckling","materials science","physics-informed ML"]},"supplementary_material":{"value":"/attachment/f320d4bec85a9ceba3dc3a8b1badc0415f6e219e.zip"},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Predicting structural failure is a fundamental objective in materials science and mechanical engineering. Euler’s classical formula, the standard for predicting the buckling instability of slender columns for over 250 years, assumes idealized material properties that can lead to unreliable predictions and potentially catastrophic failures in critical infrastructure. This study proposes a solution by introducing a novel framework that synergizes machine learning and modern explainability techniques to model complex physical systems. We used pasta as a model non-ideal material for a comprehensive experimental analysis and a dataset from 147 controlled buckling experiments on four distinct pasta gauges. We then developed a physics-informed XGBoost model, incorporating both raw geometric measurements and a composite feature derived from Euler’s formula ($G = d^4/L^2$), and subsequently evaluated the model’s performance using a 5-fold cross-validation scheme. The model demonstrated an outstanding predictive power, achieving an average coefficient of determination (R²) of 0.97 and a Root Mean Squared Error (RMSE) of 0.14 N. We also examined the model’s internal decision-making process by employing SHAP (SHapley Additive exPlanations). The analysis confirmed the primary importance of the theoretically-derived feature but also revealed that the model learned to use raw geometric data as crucial correction factors. This study presents a powerful proof of concept for using interpretable machine learning to achieve not only predictive accuracy but also gain deeper physical insights into complex, non-ideal systems. The framework presented here has broad implications for advancing our understanding and design capabilities in materials science, engineering, and advanced manufacturing."},"_bibtex":{"value":"@misc{\nanonymous2026beyond,\ntitle={Beyond Euler: An Explainable Machine Learning Framework for Predicting and Interpreting Buckling Instabilities in Non-Ideal Materials},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=uMcI7jAe8Q}\n}"},"title":{"value":"Beyond Euler: An Explainable Machine Learning Framework for Predicting and Interpreting Buckling Instabilities in Non-Ideal Materials"},"pdf":{"value":"/pdf/3e488abd3c53880c4aebfb0ccbe9681c1e652d2e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"raichura|beyond_euler_an_explainable_machine_learning_framework_for_predicting_and_interpreting_buckling_instabilities_in_nonideal_materials"},"authorids":{"value":["~Pranil_Raichura1"]},"authors":{"value":["Pranil Raichura"]}},"version":2},{"content":{"venue":{"value":"Neural Networks"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"shi|towards_complex_dynamic_physics_system_simulation_with_graph_neural_ordinary_equations"},"html":{"value":"https://research.nottingham.edu.cn/en/publications/d4165d6e-94f8-491d-a5be-8f1d0dcbcf95"},"abstract":{"value":"The great learning ability of deep learning facilitates us to comprehend the real physical world, making learning to simulate complicated particle systems a promising endeavour both in academia and industry. However, the complex laws of the physical world pose significant challenges to the learning based simulations, such as the varying spatial dependencies between interacting particles and varying temporal dependencies between particle system states in different time stamps, which dominate particles’ interacting behavior and the physical systems’ evolution patterns. Existing learning based methods fail to fully account for the complexities, making them unable to yield satisfactory simulations. To better comprehend the complex physical laws, we propose a novel model – Graph Networks with Spatial–Temporal neural Ordinary Differential Equations (GNSTODE) – that characterizes the varying spatial and temporal dependencies in particle systems using a united end-to-end framework. Through training with real-world particle–particle interaction observations, GNSTODE can simulate any possible particle systems with high precisions. We empirically evaluate GNSTODE's simulation performance on two real-world particle systems, Gravity and Coulomb, with varying levels of spatial and temporal dependencies. The results show that GNSTODE yields better simulations than state-of-the-art methods, showing that GNSTODE can serve as an effective tool for particle simulation in real-world applications. Our code is made available at https://github.com/Guangsi-Shi/AI-for-physics-GNSTODE."},"_bibtex":{"value":"@article{d4165d6e94f8491da5be8f1d0dcbcf95,\n  title     = \"Towards complex dynamic physics system simulation with graph neural ordinary equations\",\n  abstract  = \"The great learning ability of deep learning facilitates us to comprehend the real physical world, making learning to simulate complicated particle systems a promising endeavour both in academia and industry. However, the complex laws of the physical world pose significant challenges to the learning based simulations, such as the varying spatial dependencies between interacting particles and varying temporal dependencies between particle system states in different time stamps, which dominate particles{\\textquoteright} interacting behavior and the physical systems{\\textquoteright} evolution patterns. Existing learning based methods fail to fully account for the complexities, making them unable to yield satisfactory simulations. To better comprehend the complex physical laws, we propose a novel model – Graph Networks with Spatial–Temporal neural Ordinary Differential Equations (GNSTODE) – that characterizes the varying spatial and temporal dependencies in particle systems using a united end-to-end framework. Through training with real-world particle–particle interaction observations, GNSTODE can simulate any possible particle systems with high precisions. We empirically evaluate GNSTODE's simulation performance on two real-world particle systems, Gravity and Coulomb, with varying levels of spatial and temporal dependencies. The results show that GNSTODE yields better simulations than state-of-the-art methods, showing that GNSTODE can serve as an effective tool for particle simulation in real-world applications. Our code is made available at https://github.com/Guangsi-Shi/AI-for-physics-GNSTODE.\",\n  keywords  = \"AI for physics science, Graph neural networks, Learning-based simulator, Neural Ordinary Differential Equations\",\n  author    = \"Guangsi Shi and Daokun Zhang and Ming Jin and Shirui Pan and Yu, \\{Philip S.\\}\",\n  note      = \"Publisher Copyright: {\\textcopyright} 2024 The Authors\",\n  year      = \"2024\",\n  month     = aug,\n  doi       = \"10.1016/j.neunet.2024.106341\",\n  language  = \"English\",\n  volume    = \"176\",\n  journal   = \"Neural Networks\",\n  issn      = \"0893-6080\",\n  publisher = \"Elsevier Ltd.\",\n}"},"title":{"value":"Towards complex dynamic physics system simulation with graph neural ordinary equations"},"authors":{"value":[{"fullname":"Guangsi Shi"},{"fullname":"Daokun Zhang","username":"~Daokun_Zhang1"},{"fullname":"Ming Jin"},{"fullname":"Shirui Pan"},{"fullname":"Philip S. Yu"}]}},"tmdate":1789091506405,"pdate":1722470400000,"externalIds":["doi:10.1016/j.neunet.2024.106341"],"tcdate":1763003938288,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Daokun_Zhang1"],"forum":"U1nKSPeOR9","license":"CC BY-SA 4.0","number":15745,"cdate":1760161837474,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789091506405,"domain":"OpenReview.net/Public_Article","id":"U1nKSPeOR9","version":2},{"content":{"venue":{"value":"Neural Networks 2024"},"pdf":{"value":"https://www.sciencedirect.com/science/article/pii/S089360802400265X/pdfft?md5=b94458764ca6d130ffa10d3718f6cbaf&pid=1-s2.0-S089360802400265X-main.pdf"},"venueid":{"value":"dblp.org/journals/NN/2024"},"paperhash":{"value":"shi|towards_complex_dynamic_physics_system_simulation_with_graph_neural_ordinary_equations"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Guangsi_Shi:","~Daokun_Zhang1","~Ming_Jin3","~Shirui_Pan1","~Philip_S._Yu1"]},"html":{"value":"https://doi.org/10.1016/j.neunet.2024.106341"},"_bibtex":{"value":"@article{DBLP:journals/nn/ShiZJPY24,\n  author={Guangsi Shi and Daokun Zhang and Ming Jin and Shirui Pan and Philip S. Yu},\n  title={Towards complex dynamic physics system simulation with graph neural ordinary equations},\n  year={2024},\n  cdate={1704067200000},\n  journal={Neural Networks},\n  volume={176},\n  pages={106341},\n  url={https://doi.org/10.1016/j.neunet.2024.106341}\n}\n"},"abstract":{"value":"The great learning ability of deep learning facilitates us to comprehend the real physical world, making learning to simulate complicated particle systems a promising endeavour both in academia and industry. However, the complex laws of the physical world pose significant challenges to the learning based simulations, such as the varying spatial dependencies between interacting particles and varying temporal dependencies between particle system states in different time stamps, which dominate particles’ interacting behavior and the physical systems’ evolution patterns. Existing learning based methods fail to fully account for the complexities, making them unable to yield satisfactory simulations. To better comprehend the complex physical laws, we propose a novel model – Graph Networks with Spatial–Temporal neural Ordinary Differential Equations (GNSTODE) – that characterizes the varying spatial and temporal dependencies in particle systems using a united end-to-end framework. Through training with real-world particle–particle interaction observations, GNSTODE can simulate any possible particle systems with high precisions. We empirically evaluate GNSTODE’s simulation performance on two real-world particle systems, Gravity and Coulomb, with varying levels of spatial and temporal dependencies. The results show that GNSTODE yields better simulations than state-of-the-art methods, showing that GNSTODE can serve as an effective tool for particle simulation in real-world applications. Our code is made available at https://github.com/Guangsi-Shi/AI-for-physics-GNSTODE."},"title":{"value":"Towards complex dynamic physics system simulation with graph neural ordinary equations"},"authors":{"value":["Guangsi Shi","Daokun Zhang","Ming Jin","Shirui Pan","Philip S. Yu"]}},"tmdate":1768496166490,"pdate":1704067200000,"tcdate":1723118923846,"writers":["~"],"signatures":["~Ming_Jin3"],"forum":"lijxn5ZCHp","license":"CC BY-SA 4.0","number":61858,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768496166490,"domain":"DBLP.org","id":"lijxn5ZCHp","version":2},{"content":{"summary":{"value":"The paper introduces ICE-CODER, a multi-agent framework for automatic code generation, specifically targeting complex (e.g., competitive programming) problems. The core idea is to improve code reliability by more closely simulating a human software engineering process. The system integrates existing black-box test generation with two key additions: (1) white-box, coverage-guided test generation to find edge cases, and (2) an LLM-based \"error diagnosis\" step to deliberate on and resolve conflicts between code outputs and test case outputs. The authors evaluate ICE-CODER on the LiveCodeBench-Hard dataset, demonstrating significant improvements over a baseline and claiming state-of-the-art performance by solving 72 out of 90 problems."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Could you please provide precise details on the exact modifications made to the problem descriptions to allow AgentCoder, MapCoder, etc., to run? This seems like a critical confounder. How can we be sure that these modifications did not make the task easier for ICE-Coder or harder for the other methods? A fair comparison would require all methods to run on the identical, unmodified problem set.\n2. Why was the evaluation limited only to LiveCodeBench-Hard? The authors' hypothesis is that this method is good for complex problems, but this does not excuse the omission of standard benchmarks.\n3. Could the authors more clearly articulate what the primary conceptual contribution is, beyond the successful integration and engineering of these known techniques?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper identifies a weakness in current code generation agents: their inability to handle subtle edge cases in complex problems. The motivation to move beyond simple black-box testing is strong and well-founded.\n2. The core idea of integrating white-box, coverage-guided testing is a logical and intuitive next step for improving test suite quality. Simulating a \"deliberation\" or \"diagnosis\" step to handle faulty tests is also a practical solution to a known problem (LLMs generating incorrect test expectations)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The primary weakness of this paper is its lack of a core, novel contribution. The work appears to be a complex engineering pipeline that skillfully integrates several existing techniques. The paper itself cites prior work on multi-agent code generation (e.g., AgentCoder), LLMs for coverage-guided testing (e.g., Pan et al., 2025; Pizzorno & Berger, 2025), and using execution traces for debugging (e.g., Zhong et al., 2024). The main contribution is this specific combination, which feels more like an incremental \"system-building\" paper than a fundamental research advance.\n2. The experimental support is not strong enough for the claims being made.\nThe entire evaluation rests on a single, 90-problem dataset (LiveCodeBench-Hard). While this is a relevant dataset, it's a very narrow benchmark. It is unknown if this complex, multi-stage pipeline generalizes to other coding benchmarks (e.g., HumanEval, MBPP) or if it might even regress on simpler problems due to its overhead and complexity.\n3. The writing is somewhat convoluted. The paper claims to combine \"simulating a software engineering environment\" (multi-agent) with \"simulating an individual developer's mind\" (single-agent), which makes the core metaphor of the framework confusing. It is difficult for the reader to disentangle which parts are re-implementations of prior art and what the single, clean takeaway contribution is."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919392326,"tcdate":1761974149358,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7259/Reviewer_rdR2"],"signatures":["ICLR.cc/2026/Conference/Submission7259/Reviewer_rdR2"],"forum":"EDgdbdjr4c","number":2,"license":"CC BY 4.0","cdate":1761974149358,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7259/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919392326,"domain":"ICLR.cc/2026/Conference","replyto":"EDgdbdjr4c","id":"IwyWgUv4Nh","forumContent":{"TLDR":{"value":"We layer multi-agent LLM coding with white-box test generation (inspired by coverage-guided testing and code reviews), standard black-box tests, and LLM deliberation on outputs, raising LiveCodeBench‑Hard solves from 55/90 to 72/90."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["code generation","test generation","multi-agent planning"]},"supplementary_material":{"value":"/attachment/d629aee3875ff434e8db854a7c6e0fe5bd4f6242.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"LLM-based coding agents are programs that utilise LLMs to automate code generation tasks. Typically, they incorporate code execution capabilities which, together with automated test generation and/or debugging methods, enhance the reliability of the generated code. However, the effectiveness of these approaches remains limited in complex problems (such as competitive programming problems) where bugs surface only in convoluted edge cases. This work builds upon multi-agent code generation techniques which emulate software engineering environments. In particular, to address obscure edge cases, we take inspiration from code coverage tools and code reviews to generate white-box tests, on top of existing black-box test generation approaches. Test case outputs are validated through a process of deliberation using the LLM. By increasing the quantity and quality of the test cases, we obtain more reliable generated code. We evaluated ICE-Coder on LiveCodeBench-Hard. Out of the 90 problems, it solves 72, compared to the baseline of 55."},"_bibtex":{"value":"@misc{\nkeat2026icecoder,\ntitle={{ICE}-Coder: Integrating White-box and Black-box Testing in Execution-guided Multi-agent Code Generation},\nauthor={Jed Koh Jin Keat and LI YUSU and Tianyi Zhang and BINGZHENG GAN and Yangkai Ding},\nyear={2026},\nurl={https://openreview.net/forum?id=EDgdbdjr4c}\n}"},"title":{"value":"ICE-Coder: Integrating White-box and Black-box Testing in Execution-guided Multi-agent Code Generation"},"pdf":{"value":"/pdf/89544eb522a6f63cdb081ad51bf19404d2ae6c93.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"keat|icecoder_integrating_whitebox_and_blackbox_testing_in_executionguided_multiagent_code_generation"},"authorids":{"value":["~Jed_Koh_Jin_Keat1","~LI_YUSU1","~Tianyi_Zhang15","~BINGZHENG_GAN1","~Yangkai_Ding1"]},"authors":{"value":["Jed Koh Jin Keat","LI YUSU","Tianyi Zhang","BINGZHENG GAN","Yangkai Ding"]}},"version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Physics-informed neural networks","Navier-Stokes equations","complex boundary conditions"]},"primary_area":{"value":"machine_learning_for_sciences"},"abstract":{"value":"Physics-informed neural networks (PINN) have achieved notable success in solving partial differential equations (PDE), yet solving the Navier-Stokes equations (NSE) with complex boundary conditions remains a challenging task. In this paper, we introduce a novel Hybrid Boundary PINN (HB-PINN) method that combines a pretrained network for efficient initialization with a boundary-constrained mechanism. The HB-PINN method features a primary network focused on inner domain points and a distance metric network that enhances predictions at the boundaries, ensuring accurate solutions for both boundary and interior regions. Comprehensive experiments have been conducted on the NSE under complex boundary conditions, including the 2D cylinder wake flow and the 2D blocked cavity flow with a segmented inlet. The proposed method achieves state-of-the-art (SOTA) performance on these benchmark scenarios, demonstrating significantly improved accuracy over existing PINN-based approaches."},"_bibtex":{"value":"@inproceedings{\nzhou2025hybrid,\ntitle={Hybrid Boundary Physics-Informed Neural Networks for Solving Navier-Stokes Equations with Complex Boundary},\nauthor={Chuyu Zhou and Tianyu Li and Chenxi Lan and Rongyu Du and Guoguo Xin and Wei Li and Guoqing Wang and Xun Liu and Hangzhou Yang},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=KQoVgPOM1S}\n}"},"title":{"value":"Hybrid Boundary Physics-Informed Neural Networks for Solving Navier-Stokes Equations with Complex Boundary"},"pdf":{"value":"/pdf/d22e73303ed04c477d3bd44620263d2b145970db.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"zhou|hybrid_boundary_physicsinformed_neural_networks_for_solving_navierstokes_equations_with_complex_boundary"},"authorids":{"value":["~Chuyu_Zhou1","~Tianyu_Li12","~Chenxi_Lan1","~Rongyu_Du1","~Guoguo_Xin1","~Wei_Li102","~Guoqing_Wang2","~Xun_Liu7","~Hangzhou_Yang1"]},"authors":{"value":["Chuyu Zhou","Tianyu Li","Chenxi Lan","Rongyu Du","Guoguo Xin","Wei Li","Guoqing Wang","Xun Liu","Hangzhou Yang"]}},"tmdate":1783627354158,"pdate":1758217100452,"tcdate":1746885979926,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission16527/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission16527/Authors"],"forum":"KQoVgPOM1S","license":"CC BY 4.0","number":16527,"cdate":1746885979926,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission16527/-/Full_Submission","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission16527/-/Camera_Ready_Revision"],"mdate":1783627354158,"odate":1761704933799,"domain":"NeurIPS.cc/2025/Conference","id":"KQoVgPOM1S","version":2},{"content":{"summary":{"value":"Authors propose an approach to learn physics based model of real world objects from multi-view video. Authors argue that 3D representation of the scene is important to reason about the physics and use per-frame Nerf to extract the object point cloud from every frame of the multi-view video. They build a graph of the object points and use spatial message passing to learn the dynamics.\nAuthors use Nvidia FleX to generate simulated data for testing and compare their method with NeRF-dy (Li et al. CoRL'21)."},"soundness":{"value":"3 good"},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"I'm not an expert in this area. I like the ideas presented here as they are simple and make intuitive sense. However I'm concerned about the novel contributions of the work. Please see pts. 1,2 above. I'd be happy to update my score after the rebuttal."},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"strengths":{"value":"+ The proposed method is simple and intuitive.\n+ The idea of leveraging Nerf to build a 3D model of the world and using it to model physics is interesting.\n+ The method compares favourably to baseline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"[Technical]\n1. Sec 3: Proposed method consists of two steps, extracting the scene point cloud using Nerf [41] and computing dynamics using message passing [7,36,37].\nIsn’t this similar to [35] where they predict 3D particle locations using NN and then use message passing + some other learning to predict the dynamics? If Nerf produces better points why can’t we replace the MLP based points from [35] with Nerf and use rest of the method? They also don’t make assumptions about physical properties of objects (also see pt. 2).\nIs the key idea here that [35] requires GT 3D supervision whereas the proposed method does not? But doesn't having multi-view images with known cameras (as used in the method) sort of provide this supervision?\n\n2. L154: what is captured in point attribute a_i^v? How is it obtained? Is it known beforehand? Adding dimensionality to variables would also help. \nUnlike proposed method, prior work [35] does not require prior knowledge about the physical properties of the object or the scene. Isn’t the proposed method restrictive? Please correct me if I misunderstood this.\n\n3. L141: How is object segmentation obtained? Is it provided as input? All inputs and outputs must be clearly explained.\nA good idea would be to mention these in the intro to method section. Similarly all assumptions about the problem setting such as in L168-171 should also be clarified early on.\n\n[Minor]\nL106-109: Authors write they learn 3D-IntPhysics without any 3D supervision which is not entirely true. They use multi-view images with known camera for Nerf based representation which is also not entirely \"not 3D\".\nAuthors also say that prior work uses GT trajectories which are hard to obtain in real setting, but getting multi-view images with GT camera parameters is also not trivial."},"limitations":{"value":"Authors briefly mention one of the limitations along with the conclusions but a more in-depth analysis would be appreciated. This could go in the supp. mat. and authors can at least leave a pointer in the main paper."}},"nonreaders":[],"tmdate":1702411095920,"tcdate":1689054413628,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission7124/Reviewer_493i"],"signatures":["NeurIPS.cc/2023/Conference/Submission7124/Reviewer_493i"],"forum":"Fp5uC6YHwe","number":4,"license":"CC BY 4.0","cdate":1689054413628,"mdate":1702411095920,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission7124/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"Fp5uC6YHwe","id":"mBUQehetMF","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Intuitive Physics","Computer Vision"]},"supplementary_material":{"value":"/attachment/d183adaec59dcb150929a16c875749851b24a450.pdf"},"_bibtex":{"value":"@inproceedings{\nxue2023dintphys,\ntitle={3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes},\nauthor={Haotian Xue and Antonio Torralba and Joshua B. Tenenbaum and Daniel LK Yamins and Yunzhu Li and Hsiao-Yu Tung},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=Fp5uC6YHwe}\n}"},"title":{"value":"3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes"},"paperhash":{"value":"xue|3dintphys_towards_more_generalized_3dgrounded_visual_intuitive_physics_under_challenging_scenes"},"abstract":{"value":"Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate the scene to achieve desired outcomes without relying on extensive trial and error. In this paper, we present a framework capable of learning 3D-grounded visual intuitive physics models from videos of complex scenes with fluids. Our method is composed of a conditional Neural Radiance Field (NeRF)-style visual frontend and a 3D point-based dynamics prediction backend, using which we can impose strong relational and structural inductive bias to capture the structure of the underlying environment. Unlike existing intuitive point-based dynamics works that rely on the supervision of dense point trajectory from simulators, we relax the requirements and only assume access to multi-view RGB images and (imperfect) instance masks acquired using color prior. This enables the proposed model to handle scenarios where accurate point estimation and tracking are hard or impossible. We generate datasets including three challenging scenarios involving fluid, granular materials, and rigid objects in the simulation. The datasets do not include any dense particle information so most previous 3D-based intuitive physics pipelines can barely deal with that. We show our model can make long-horizon future predictions by learning from raw images and significantly outperforms models that do not employ an explicit 3D representation space. We also show that once trained, our model can achieve strong generalization in complex scenarios under extrapolate settings."},"pdf":{"value":"/pdf/739e599bd87f8fd4fcb09b78dd8827c81af83052.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Haotian_Xue1","~Antonio_Torralba1","~Joshua_B._Tenenbaum1","~Daniel_LK_Yamins1","~Yunzhu_Li1","~Hsiao-Yu_Tung1"]},"authors":{"value":["Haotian Xue","Antonio Torralba","Joshua B. Tenenbaum","Daniel LK Yamins","Yunzhu Li","Hsiao-Yu Tung"]}},"version":2},{"content":{"summary":{"value":"This submission introduces PhysicsMinions, a coevolutionary multimodal multi-agent framework designed to elevate AI performance on challenging Olympiad-level physics problems. The system is structured into three specialized studios: a Visual Studio for diagram interpretation, a Logic Studio for solution generation and refinement, and a Review Studio for dual-stage verification. The agents interact through iterative feedback loops, combining structured visual extraction, explicit stepwise reasoning, and systematic error detection/correction. The framework is evaluated on the HiPhO benchmark (across 7 recent physics Olympiads), using four state-of-the-art multimodal language models. Results indicate consistent improvements for both open and closed-source models, including reaching gold medalist-level performance previously unattainable for open-source systems."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. It is necessary to clarify whether the potential research significance of your proposed method lies primarily in advancing the general development of large language models (LLMs), in contributing specifically to the \"LLM for Science\" direction, or in facilitating agent-based application development. At present, what I observe is essentially a finely crafted problem construction for a specific scenario. An adequate response should explain how your study can provide broader insights — for example, how it could inspire rapid design strategies when facing novel and challenging problems in future large models.\n\n2. Please provide quantitative details regarding the computational overhead (in terms of both time and token usage) incurred by your proposed method. In typical use cases, how many times greater is the overhead compared to the Single Model approach illustrated in Figure 1? How does the time consumption compare under these conditions?\n\n3. For *PhysicsMinions*, how is the decision made to completely reinitialize (i.e., discard and regenerate) a solution instead of continuing iterative refinement? Is this decision based on a fixed threshold, or is it determined adaptively?\n\nThe review scores may vary depending on the quality of your responses."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"I observed extensive and direct use of the “Minions” character design and name, which are copyrighted and trademarked by Universal Pictures/Illumination Entertainment. I hope the author will carefully consider whether this may pose unnecessary copyright risks."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"1. The manuscript is notably clear, demonstrating substantial effort in both visualization and written presentation.\n\n2. The study employs rigorous evaluation metrics (gold-medal thresholds and Pass@k) across seven recent and distinct Physics Olympiad competitions, comparing open-source and closed-source models to establish consistent performance benefits. Tables 1 and 2 convincingly show that *PhysicsMinions* elevates multiple models beyond the gold-medal line and achieves a historic milestone for open-source benchmarks.\n\n3. The ablation studies systematically examine the impact of the Visual Studio, the Review Studio (including physics-specific and general verification), and the key hyperparameter—Consecutive Verification (CV). The results support the claim that each architectural component is necessary to attain the highest performance tier.\n\n4. Section 5.2 candidly acknowledges current shortcomings in precise visual data extraction, pinpointing specific tools and failure cases within the system or available toolkits (see Fig. 7). This openness establishes a solid and honest foundation for future methodological improvements.\n\n5. The design rationale—encompassing visual, logical, and review components—is well-motivated, with clearly defined and differentiated roles for diagram parsing, symbolic reasoning, and correctness verification. The iterative co-evolution process and explicit feedback loops are logically coherent and empirically well-supported."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### The Generalizability of the Paper is Limited\n\nDespite claiming generalizability, the system is only evaluated using physics Olympiad problems derived from a single benchmark (HiPhO). No evidence is presented for transfer to related domains (such as mathematics Olympiads, chemistry Olympiads, or general science Q&A), nor are conceptual or empirical justifications provided for broader applicability beyond physics. I **would like to see evaluations targeting large-model agents on real scientific research problems** (even if sometimes limited to the physics domain, that would still be a significant breakthrough).\n\nMoreover, the experiments are heavily designed around the structure of the HiPhO Olympiad (including structured JSON extraction for diagrams that conform to this benchmark’s conventions), creating a risk of benchmark-specific optimization. The paper could be improved by incorporating cross-benchmark generalization tests or, at the very least, providing deeper error analyses on out-of-distribution tasks/problems.\n\n---\n\n### Lack of Fairness in Experimental Comparison\n\nIn fairness comparisons, the paper uses a unified evaluation benchmark and scoring standard, ensuring result comparability. However, the experiments do not disclose hardware configurations for each method during runtime, total reasoning time, overall GPU usage, or token costs—only a single example’s relative consumption ratio is provided in the hyperparameter analysis. Given that the proposed method is a multi-stage, multi-round generation-and-verification system, **its inference overhead is clearly higher than that of single models and other baselines**. It is recommended that the authors supplement the final version with resource consumption comparisons for each method to allow readers to fully assess the real-world application cost. **If the strong results are not achieved under equal token or time consumption, they are meaningless and unacceptable.** If resources can be used without limits, then comparing with a single agent or single LLM is pointless—frequent agent calls will be hard to inspire either product development or academic research. (I acknowledge that many LLM-Agent papers have this issue, but it is still something that must be carefully considered.)\n\n---\n\n### Many Large-Model Prompts May Have Been Generated by LLMs\n\nIn the appendix, I noticed numerous prompts that, when tested on GPTZero, are highly likely to have been generated by an LLM. Extremely long prompts can cause the model’s attention to collapse, and I am not certain that Markdown is the best choice for prompt formatting. I cannot definitively confirm the prompts were LLM-generated, but the design philosophy and generation process still need to be described in detail in the paper. After all, if each task requires such an extensive prompt design process, usability will be very poor."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915744418,"tcdate":1762012797423,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1350/Reviewer_YFh1"],"signatures":["ICLR.cc/2026/Conference/Submission1350/Reviewer_YFh1"],"forum":"kipQYpoZf1","number":3,"license":"CC BY 4.0","cdate":1762012797423,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1350/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915744418,"domain":"ICLR.cc/2026/Conference","replyto":"kipQYpoZf1","id":"wdmFZOi77A","forumContent":{"TLDR":{"value":"PhysicsMinions is a coevolutionary multimodal multi-agent system that achieves gold-medal performance in latest physics Olympiads, significantly outperforming single-model baselines and advancing open-source models to human-expert levels."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics Olympiad","multi-agent system","coevolutionary framework","multimodal model"]},"supplementary_material":{"value":"/attachment/33ae425d5b96c718e9f850d263c500e27f92b666.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics is central to understanding and shaping the real world, and the ability to solve physics problems is a key indicator of real-world physical intelligence. Physics Olympiads, renowned as the crown of competitive physics, provide a rigorous testbed requiring complex reasoning and deep multimodal understanding, yet they remain largely underexplored in AI research. Existing approaches are predominantly single-model based, and open-source MLLMs rarely reach gold-medal-level performance. To address this gap, we propose PhysicsMinions, a coevolutionary multi-agent system for Physics Olympiad. Its architecture features three synergistic studios: a Visual Studio to interpret diagrams, a Logic Studio to formulate solutions, and a Review Studio to perform dual-stage verification. The system coevolves through an iterative refinement loop where feedback from the Review Studio continuously guides the Logic Studio, enabling the system to self-correct and converge towards the ground truth. Evaluated on the HiPhO benchmark spanning 7 latest physics Olympiads, PhysicsMinions delivers three major breakthroughs: (i) Strong generalization: it consistently improves both open-source and closed-source models of different sizes, delivering clear benefits over their single-model baselines; (ii) Historic breakthroughs: it elevates open-source models from only 1–2 to 6 gold medals across 7 Olympiads, achieving the first-ever open-source gold medal in the latest International Physics Olympiad (IPhO) under the average-score metric; and (iii) Scaling to human expert: it further advances the open-source Pass@32 score to 26.8/30 points on the latest IPhO, ranking 4$^\\text{th}$ of 406 contestants and far surpassing the top single-model score of 22.7 (ranked 22$^\\text{nd}$). Generally, PhysicsMinions offers a generalizable framework for Olympiad-level problem solving, with the potential to extend across disciplines."},"_bibtex":{"value":"@misc{\nyu2026physicsminions,\ntitle={PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System},\nauthor={Fangchen Yu and Junchi Yao and Ziyi Wang and Haiyuan Wan and Youling Huang and Bo Zhang and Shuyue Hu and Dongzhan Zhou and Ning Ding and Ganqu Cui and LEI BAI and Wanli Ouyang and Peng Ye},\nyear={2026},\nurl={https://openreview.net/forum?id=kipQYpoZf1}\n}"},"title":{"value":"PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System"},"pdf":{"value":"/pdf/9b0cafdc54384f883a43e35363ab5ceebd7ea9e5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yu|physicsminions_winning_gold_medals_in_the_latest_physics_olympiads_with_a_coevolutionary_multimodal_multiagent_system"},"authorids":{"value":["~Fangchen_Yu1","~Junchi_Yao1","~Ziyi_Wang33","~Haiyuan_Wan1","~Youling_Huang2","~Bo_Zhang17","~Shuyue_Hu1","~Dongzhan_Zhou1","~Ning_Ding5","~Ganqu_Cui1","~LEI_BAI1","~Wanli_Ouyang1","~Peng_Ye4"]},"authors":{"value":["Fangchen Yu","Junchi Yao","Ziyi Wang","Haiyuan Wan","Youling Huang","Bo Zhang","Shuyue Hu","Dongzhan Zhou","Ning Ding","Ganqu Cui","LEI BAI","Wanli Ouyang","Peng Ye"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a modular surrogate for E3SM-ELM BGC spin-up: heterogeneous encoders (LSTM/CNN/FC), transformer fusion, physics soft/hard constraints (e.g., Softplus for non-negativity; NPP = GPP − AR penalty). Claims >60× acceleration by inferring slow pools from 20 years of data, then restarting ELM for 100 years to reach equilibrium. Outperforms MLP/CNN/FNO/PINN on R²/RMSE, generalizes from 1° → 0.5° with few-shot fine-tuning, and can produce restartable states (baselines often crash)."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Are the physics constraints hard-enforced during decoding, or only penalized via loss? What is the magnitude of constraint violation at inference time?\nHow is the trade-off parameter between physics and data loss tuned? Is it static or adaptive (e.g., gradient-balancing)?\nHow does the anomaly detector quantify distribution shift? What metric (e.g., Mahalanobis, reconstruction error) is used, and how is the threshold calibrated?\nHow transferable is PHASE to other PDE-based simulators (e.g., ocean, atmosphere)? Would the same modular hierarchy apply?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Practical, high-impact workflow contribution: generating restartable states that actually run forward in ELM is a meaningful bar beyond offline metrics.\n\nTackles heterogeneous inputs/targets systematically; the architecture diagram and data pipeline are helpful.\n\nSolid ablations (CNN/LSTM/Transformer/physics-loss) and cross-resolution study.\n\nThe domain-knowledge example (adding P in the tropics) is a nice, concrete illustration."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Physics integration is relatively light-touch.\nThe current mechanism is primarily soft constraints + architectural non-negativity. There’s no explicit conservation enforcement, differentiable constraints, or coupling to a solver. For a paper emphasizing “physics-integrated,” this is closer to regularized MTL than to, say, constrained operator learning. Please position claims accordingly and/or add stronger physics integration (e.g., constrained decoding, projected gradients, or differentiable diagnostic modules).\n\nRestart validation needs hard numbers.\nThe headline is the 60× speedup via restartability. Please report quantitative drift metrics after the 100-year continuation: e.g., tendencies of target pools, global/biome-wise mass balance errors, and failure rate across sites. Right now, the restart result is demonstrated but not statistically characterized.\n\nUncertainty & trust.\nThe use-case (initializing a coupled Earth system model) begs for uncertainty estimates (epistemic/aleatoric) and calibration checks. The OOD “anomaly” mechanism is mentioned, but thresholds, scoring, and deployment policy aren’t described. Even simple ensembles or MC-dropout with reliability diagrams would help.\n\nData access & reproducibility.\nThe dataset is built from internal ELM runs; without a releasable subset or scripts, end-to-end replication is hard. Consider releasing (i) the fusion/encoder templates, (ii) a toy public analog (e.g., open land model components), and (iii) restart evaluation scripts.\n\nComparators & fairness.\nFNO and a state-evolution PINN are relevant, but this task (restartable steady states) is unusual. Could you include a hybrid surrogate with constrained outputs (e.g., learned decoder followed by a small balance-enforcing reconciliation), or a solver-in-the-loop variant as a stronger baseline?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925405561,"tcdate":1761376390152,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15081/Reviewer_ky3X"],"signatures":["ICLR.cc/2026/Conference/Submission15081/Reviewer_ky3X"],"forum":"RBWVq21M2K","number":2,"license":"CC BY 4.0","cdate":1761376390152,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15081/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925405561,"domain":"ICLR.cc/2026/Conference","replyto":"RBWVq21M2K","id":"cOU8ymRTST","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Physics-informed deep learning","Scientific machine learning","Surrogate modeling","Earth system modeling"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Large‐scale numerical simulations underpin modern scientific discovery but remain constrained by prohibitive computational costs. AI surrogates offer acceleration, yet adoption in mission‑critical settings is limited by concerns over physical plausibility, trustworthiness, and the fusion of heterogeneous data. We introduce PHASE, a modular deep‑learning framework for physics‑integrated, heterogeneity‑aware surrogates in scientific simulations. PHASE combines data‑type–aware encoders for heterogeneous inputs with multi‑level physics‑based constraints that promote consistency from local dynamics to global system behavior. We validate PHASE on the biogeochemical (BGC) spin‑up workflow of the U.S. Department of Energy’s Energy Exascale Earth System Model (E3SM) Land Model (ELM), presenting—to our knowledge—the first scientifically validated AI‑accelerated solution for this task. Using only the first 20 simulation years, PHASE infers a near‑equilibrium state that otherwise requires more than 1,200 years of integration, yielding an effective reduction in required integration length by at least 60×. The framework is enabled by a pipeline for fusing heterogeneous scientific data and demonstrates strong generalization to higher spatial resolutions with minimal fine‑tuning. These results indicate that PHASE captures governing physical regularities rather than surface correlations, enabling practical, physically consistent acceleration of land‑surface modeling and other complex scientific workflows."},"_bibtex":{"value":"@misc{\ngao2025phase,\ntitle={{PHASE}: Physics\\nobreakdash-Integrated, Heterogeneity\\nobreakdash-Aware Surrogates for Scientific Simulations},\nauthor={Dawei Gao and Dali Wang and Zhuowei Gu and Qinglei Cao and Xiao Wang and Peter E. Thornton and Daniel Ricciuto and Yunhe Feng},\nyear={2025},\nurl={https://openreview.net/forum?id=RBWVq21M2K}\n}"},"title":{"value":"PHASE: Physics‑Integrated, Heterogeneity‑Aware Surrogates for Scientific Simulations"},"pdf":{"value":"/pdf/c3e3c1180de4599080869b4aa91ec4851b1c9a81.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"gao|phase_physicsintegrated_heterogeneityaware_surrogates_for_scientific_simulations"},"authorids":{"value":["~Dawei_Gao2","~Dali_Wang1","~Zhuowei_Gu1","~Qinglei_Cao2","~Xiao_Wang21","~Peter_E._Thornton1","~Daniel_Ricciuto1","~Yunhe_Feng2"]},"authors":{"value":["Dawei Gao","Dali Wang","Zhuowei Gu","Qinglei Cao","Xiao Wang","Peter E. Thornton","Daniel Ricciuto","Yunhe Feng"]}},"version":2},{"content":{"summary":{"value":"The thesis of this paper is that simple machine learning tasks should be addressed with simple, interpretable or explainable models, rather than with more complex (but also less transparent) deep learning approaches, even at the cost of a slightly lower performance. The paper makes the case of an image classification task, where simple color features and decision trees produce transparent and intuitive solutions that are only marginally worse than state-of-the-art deep architectures."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"+ I totally sympathize with the claim of the paper, since I agree that simpler methods should be preferred whenever possible, i.e., when there is no significant performance gap with respect to more complex (e.g., deep learning) models and where explainability is a crucial need.\n\n+ The paper is very well written, clear, and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The thesis of the paper is certainly not novel. The paper itself cites the work by Cynthia Rudin, but other works have recently proved the advantages of classical approaches like decision trees with respect to more complex deep learning approaches: e.g., see Grinsztajn et al., \"Why do tree-based models still outperform deep learning on typical tabular data?\", NeurIPS 2022 and references therein.\n\n- The paper only presents a case study on a single (image processing problem) which allows a solution based on very simple features (colors extracted from RGB). Therefore, it is vert hard to assess to what extent the methodology could be generalized, in fact the chosen features are in this case highly intuitiveand easy to understand for humans, and the explanation representation based on the RGB space depends on such features. It is not clear which other problems and tasks could present a similar scenario.\n\n- The paper looks in this sense more suited for a computer-vision workshop or conference, since the thesis is not novel, and the methodology is tailored to the case study. I am not sure how many real-world problems in computer vision fall in this category of tasks that can be easily solved by a decision tree (or a similar, simple technique). A study on this would be extremely interesting."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"* The proposed methodology exploits a pre-processing stage that transforms images from RGB to YUV and back to RGB after a luminance adjustment. It is not clear from the paper, but I think that also the deep learning approaches tested in the paper should work on the same images, processed with the same technique, to make a fair comparison.\n\n* On pag. 7, the paper states that \"for every non-leaf node, the DT learns a threshold value for one of its given features, thus producing two children\" -> This is true for the task considered in the paper, but not in general, where also discrete variables with multiple outcomes could be used as features (thus requiring no thresholds while producing multiple children).\n\n---\n\n- Pag. 2, \"for whom strategies\" -> \"for which strategies\"\n- Pag. 2, \"that are more intuitive\" -> \"that are most intuitive\"\n- Pag. 4, \"while the human performance\" -> \"while human performance\""},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636850399,"tcdate":1698444159195,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission7172/Reviewer_qoq3"],"signatures":["ICLR.cc/2024/Conference/Submission7172/Reviewer_qoq3"],"forum":"zqVvdn0NQM","number":2,"license":"CC BY 4.0","cdate":1698444159195,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission7172/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636850399,"domain":"ICLR.cc/2024/Conference","replyto":"zqVvdn0NQM","id":"Z3JSjNtFtg","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Explainable AI","Computer Vision","Natural Language Processing","Transformer","Decision Tree"]},"supplementary_material":{"value":"/attachment/be835a9a5984b50e7c2bed2ff4e78d68fc594d6f.pdf"},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The ability of deep learning-based approaches to extract features autonomously from raw data while outperforming traditional methods has led to several breakthroughs in artificial intelligence. However, it is well-known that deep learning models suffer from an intrinsic opacity, making it difficult to explain why they produce specific predictions. This is problematic not only because it hinders debugging but, most importantly, because it negatively affects the perceived trustworthiness of the systems. What is often overlooked is that many relatively simple tasks can be solved efficiently and effectively with data processing strategies paired with traditional models that are inherently more transparent. This work highlights the frequently neglected perspective of using knowledge-based and explainability-driven problem-solving in ML. To support our guidelines, we propose a simple strategy for solving the task of classifying the ripeness of banana crates. This is done by planning explainability and model design together. We showcase how the task can be solved using opaque deep learning models and more transparent strategies. Notably, there is a minimal loss of accuracy but a significant gain in explainability, which is truthful to the model’s inner workings. Additionally, we perform a user study to evaluate the perception of explainability by end users and discuss our findings."},"_bibtex":{"value":"@misc{\nrizzo2024stop,\ntitle={Stop overkilling simple tasks with black-box models, use more transparent models instead},\nauthor={Matteo Rizzo and Matteo Marcuzzo and Alessandro Zangari and Andrea Gasparetto and Andrea Albarelli},\nyear={2024},\nurl={https://openreview.net/forum?id=zqVvdn0NQM}\n}"},"title":{"value":"Stop overkilling simple tasks with black-box models, use more transparent models instead"},"pdf":{"value":"/pdf/72d8cd208324bd8d10e94e2b1e3c981928132b11.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"rizzo|stop_overkilling_simple_tasks_with_blackbox_models_use_more_transparent_models_instead"},"authorids":{"value":["~Matteo_Rizzo1","matteo.marcuzzo@unive.it","alessandro.zangari@unive.it","~Andrea_Gasparetto1","~Andrea_Albarelli1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Matteo Rizzo","Matteo Marcuzzo","Alessandro Zangari","Andrea Gasparetto","Andrea Albarelli"]}},"version":2},{"content":{"summary":{"value":"This paper explores the use of synthetic data in post-training tasks for large language models. It presents a theoretical framework to analyze the effects of synthetic data on model generalization, focusing on the reverse-bottleneck effect. The authors introduce key metrics like Generalization Gain via Mutual Information (GGMI) and propose a detailed modeling of the synthetic data generation process. They also provide upper bounds on generalization error based on information theory."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How would synthetic data that has not been processed by an LLM impact the proposed theory? For instance, data generated through specific rules, such as template-based concatenation or converting game-playing sequences into text.\n2. Does this theoretical analysis still hold for non-text, multimodal data? Would the same principles apply, or are there limitations in extending the theory to such data types?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. The paper provides a detailed theoretical analysis of the introduction of synthetic data, clearly explaining the information gain that synthetic data can bring.\n2. The authors present thorough and well-supported mathematical proofs to substantiate their claims, offering a solid foundation for their arguments.\n3. The reverse-bottleneck theory mathematically captures the essence of synthetic data's impact, offering valuable insights and guidance for the development of synthetic data methodologies."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper primarily focuses on theoretical analysis with minimal experimental validation. It is recommended to include empirical experiments to support the proposed theory.\n2. The use of GMM models and non-real numerical data raises questions about whether the findings can effectively reflect real-world LLM outcomes.\n3. The theory mainly emphasizes the performance benefits of introducing synthetic data, but lacks analysis in other areas. For example, it could explore potential drawbacks of over-reliance on synthetic data in LLMs, as highlighted in studies like https://www.nature.com/articles/s41586-024-07566-y."}},"nonreaders":[],"tmdate":1732529455479,"tcdate":1729500651118,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5944/Reviewer_7KR4"],"signatures":["ICLR.cc/2025/Conference/Submission5944/Reviewer_7KR4"],"forum":"UxkznlcnHf","number":1,"license":"CC BY 4.0","cdate":1729500651118,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5944/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732529455479,"domain":"ICLR.cc/2025/Conference","replyto":"UxkznlcnHf","id":"xXEk4Uyt8z","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"This paper explores the critical role of synthetic data in enhancing the post-training performance of large language models (LLMs) from a novel reverse-bottleneck perspective."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["large language models; synthetic data; information bottleneck"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Synthetic data has become a pivotal resource in post-training tasks for large language models (LLMs) due to the scarcity of high-quality, specific data. While various methods have been developed to generate synthetic data, there remains a discernible gap between the practical effects of synthetic data and our theoretical comprehension. To address this challenge, we commence by presenting a detailed modeling of the prevalent synthetic data generation process. Building upon this modeling, we demonstrate that the generalization capability of the post-trained model is critically determined by the information gain derived from the generative model, as analyzed from a novel reverse-bottleneck perspective. Moreover, we introduce the concept of Generalization Gain via Mutual Information (GGMI) and elucidate the relationship between generalization gain and information gain. This analysis serves as a theoretical foundation for synthetic data generation and further highlights its connection with the generalization capability of post-trained models, offering an understanding about the design of synthetic data generation techniques and the optimization of the post-training process. We open-source our code at https://github.com/ZyGan1999/Towards-a-Theoretical-Understanding-of-Synthetic-Data-in-LLM-Post-Training."},"_bibtex":{"value":"@inproceedings{\ngan2025towards,\ntitle={Towards a Theoretical Understanding of Synthetic Data in {LLM} Post-Training: A Reverse-Bottleneck Perspective},\nauthor={Zeyu Gan and Yong Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=UxkznlcnHf}\n}"},"title":{"value":"Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective"},"pdf":{"value":"/pdf/127d76775eb769452b3e1f3cffc5359d9e886a32.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"gan|towards_a_theoretical_understanding_of_synthetic_data_in_llm_posttraining_a_reversebottleneck_perspective"},"authorids":{"value":["~Zeyu_Gan1","~Yong_Liu7"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyu Gan","Yong Liu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes SP-VLA, a unified acceleration framework for Vision-Language-Action (VLA) models.\nIt combines:(1) Model Scheduling – dynamically switching between a full-scale VLA model and a lightweight ridge-regression-based generator depending on whether the action is “deliberative” or “intuitive,” to reduce temporal redundancy; and(2) Spatio-Semantic Token Pruning – pruning vision tokens using accumulated attention scores and Canny-edge information to remove spatial redundancy. Experiments on LIBERO and SimplerEnv demonstrate 1.5–2.4× inference acceleration with nearly lossless task performance, and improved inference frequency and latency."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Could you explain how the main hyperparameters (e.g., speed thresholds, pruning ratios, buffer size) were determined in practice?\nWere they tuned by trial-and-error, grid search, or guided by theoretical intuition or prior experiments? Are these parameter settings consistent across tasks and environments, or do they need adjustment for each scenario? \n\n2. Have you attempted to deploy SP-VLA on a physical robotic platform, or at least measured whether the acceleration results reported in simulation differ from those on real hardware?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. **Novel and conceptually clear idea**\n\nThe paper is among the first to jointly address temporal and spatial redundancies in VLA models.\nThe analogy to human dual-process control (deliberative vs. intuitive actions) is original and intuitively compelling, offering a behavioral perspective on model efficiency.\n\n2. **Comprehensive empirical validation**\n\nResults are reported across multiple VLA backbones (OpenVLA, CogACT) and environments (LIBERO, SimplerEnv) with consistent performance gains.\nAblation studies, sensitivity analyses, and visualizations are well-structured and strengthen empirical credibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Over-reliance on handcrafted heuristics**\n\nBoth the scheduling and pruning modules rely on manually designed heuristics—velocity thresholds, ridge regression fitting, Canny edges, and fixed attention thresholds—rather than adaptive or learnable components. This constrains scalability and contrasts with current trends toward learned compression/scheduling in multimodal LLM research.\n\n2. **Generalization and robustness not sufficiently validated**\n\nThe method’s stability under varying architectures, task types, or sensor conditions is not explored. Key hyperparameters (speed thresholds, pruning ratios, deliberative/intuitive ratio) might require re-tuning, reducing transferability. A cross-task or cross-model evaluation would greatly enhance the paper’s impact."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359436509,"tcdate":1761295052361,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission944/Reviewer_2cVD"],"signatures":["ICLR.cc/2026/Conference/Submission944/Reviewer_2cVD"],"forum":"RwdGIIjPlC","number":1,"license":"CC BY 4.0","cdate":1761295052361,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission944/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359436509,"domain":"ICLR.cc/2026/Conference","replyto":"RwdGIIjPlC","id":"6G8qpQbmRV","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Vision Language Action Model","Model Lightweighting","Acceleration","Embodied intelligence"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Vision-Language-Action (VLA) models have attracted increasing attention for their strong control capabilities. However, their high computational cost and low execution frequency hinder their suitability for real-time tasks such as robotic manipulation and autonomous navigation. Existing VLA acceleration methods primarily focus on structural optimization, overlooking the fact that these models operate in sequential decision-making environments. As a result, temporal redundancy in sequential action generation and spatial redundancy in visual input remain unaddressed. To this end, we propose SP-VLA, a unified framework that accelerates VLA models by jointly scheduling models and pruning tokens. Specifically, we design an action-aware model scheduling mechanism that reduces temporal redundancy by dynamically switching between VLA model and a lightweight generator.  Inspired by the human motion pattern of focusing on key decision points while relying on intuition for other actions, we categorize VLA actions into deliberative and intuitive, assigning the former to the VLA model and the latter to the lightweight generator, enabling frequency-adaptive execution through collaborative model scheduling. To address spatial redundancy, we further develop a spatio-semantic dual-aware token pruning method.  Tokens are classified into spatial and semantic types and pruned based on their dual-aware importance to accelerate VLA inference. These two mechanisms work jointly to guide the VLA in focusing on critical actions and salient visual information, achieving effective acceleration while maintaining high accuracy. Extensive experiments show that our method achieves 1.5$\\times$ lossless acceleration in LIBERO and 2.4$\\times$ in SimplerEnv, with up to 6\\% average performance gain. Inference frequency and latency improve by 2.2$\\times$ in SimplerEnv and 1.4$\\times$ in LIBERO."},"_bibtex":{"value":"@inproceedings{\nli2026spvla,\ntitle={{SP}-{VLA}: A Joint Model Scheduling and Token Pruning Approach for {VLA} Model Acceleration},\nauthor={Ye Li and Yuan Meng and Zewen Sun and Kangye Ji and Chen Tang and Jiajun Fan and Xinzhu Ma and Shu-Tao Xia and Zhi Wang and Wenwu Zhu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=RwdGIIjPlC}\n}"},"title":{"value":"SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration"},"pdf":{"value":"/pdf/0749a62cd579de7a7bcf030bcb8d83bae1a61f97.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|spvla_a_joint_model_scheduling_and_token_pruning_approach_for_vla_model_acceleration"},"authorids":{"value":["~Ye_Li2","~Yuan_Meng2","~Zewen_Sun3","~Kangye_Ji1","~Chen_Tang3","~Jiajun_Fan1","~Xinzhu_Ma1","~Shu-Tao_Xia1","~Zhi_Wang5","~Wenwu_Zhu1"]},"authors":{"value":["Ye Li","Yuan Meng","Zewen Sun","Kangye Ji","Chen Tang","Jiajun Fan","Xinzhu Ma","Shu-Tao Xia","Zhi Wang","Wenwu Zhu"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the problem of generating synthetic text data that preserves the privacy of the original data and is useful for downstream tasks. The paper proposes to use a pre-trained large language model and fine-tune it with differential privacy on a sensitive dataset, using different parameter-efficient methods such as prompt tuning and LoRa tuning. The paper shows that the synthetic data generated by this approach is of high quality and can achieve comparable or better performance than directly training a downstream classifier with differential privacy on the real data.  The paper also demonstrates that the synthetic data can be used for other tasks, such as hyperparameter tuning of the downstream models."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"- The authors conduct extensive evaluation and offer valuable empirical insights into DP synthetic text generation, such as highlighting the importance of prefix-LM that assigns zero weights to the prefix during training, random initialization for prompt tensors on DP prompt tuning, and the superior performance of LoRA compared to prompt tuning. \n- Additionally, the paper provides analysis of the synthetic data, such as the effects of synthetic data size, the rank correction of synthetic data for hyperparameter tuning, etc. These empirical findings provide a compelling and informative read.\n- The authors identify a critical issue regarding the overlap between pretrain data and finetuning data in some of the previous studies, underscoring the necessity to mitigate potential pitfalls in future research endeavors."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Novelty:\n- The novelty of the study may be limited, given that DP-SGD is a standard technique for DP synthetic text generation (Yue et al. 2022), and parameter-efficient fine-tuning has already been explored in DP LLM (Yu et al., 2021), albeit not directly applied to synthetic data generation.\n\n\nComparison to Yue et al. (2022): \n- The discussion and comparison with Yue et al. (2022) might be confusing to the readers. It would be helpful if the authors could clarify the difference between 'conditioning on some features' and 'augmenting the fine-tuning process.' mentioned in the Section 2 related work.  In Yue et al. (2022), labels are used in the prompt for conditional generation, with labels considered as non-private. Therefore,  there seems to be no “augmentation” during the finetuning process. This approach seems to be the same as the proposed method in section 4.1 of this paper, where the authors also use label names in the prefix as condition generation. \n- “obtaining good fidelity non-private synthetic data is possible, contrary to the results reported in (Yue et al., 2022) and Putta et al. (2023)” this statement may be confusing. Actually, Table 2 in Yue et al. (2022) shows a similar conclusion: synthetic data can outperform real data in terms of downstream model utility when both datasets are of the same size.\n\nDataset Choice:\n- While the authors acknowledge the presence of IMDB in Pile and perform deduplication accordingly, it might be more beneficial to directly use a dataset from an unseen domain, such as a medical dataset.\n\nDe-deplication:\n- “we used the suffix arrays to find common sequences of 50 or more tokens which appear in The Pile” A justification for the choice of the hyperparameter value of 50 in the use of suffix arrays would be appreciated.\n- Could the authors elucidate why this de-duplication is “stronger” than simply removing the datasets from the Pile?\n\nDownstream Tasks:\n- All downstream tasks in the study are binary sentiment classification tasks, which may appear monotonous and simplistic. The utility of synthetic data for more complex downstream tasks, such as multi-class classification or classification tasks beyond sentiment, as considered in previous studies, remains unclear. \n\n\nPresentation and Interpretation of Results:\n- The interpretation of results under different metrics in Table 3 could be more clearly explained, particularly for readers who may not be familiar with RBO, Spearman, and Kendall metrics. For example,  how good or bad is the value 0.56 under RBO 25% metric for Bert trained on real data with $\\epsilon=\\infty$?\n- “ For n-gram statistics, we determine the frequency of unigrams, bigrams, and sample lengths in characters for both the original and synthetic datasets. “ While the authors use these statistics to calculate the rank correction, some qualitative visualizations, such as plots comparing the bigrams/length distributions of original and synthetic data, would be a valuable addition to better illustrate the similarity between the two datasets.\n\n\nTypos: \n- In section 6,  there is a missing space after the comma: “monitoring,and sharing,” –> “ monitoring, and sharing,”"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Please see my questions in Weaknesses."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636002850,"tcdate":1698823572690,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission756/Reviewer_TdTq"],"signatures":["ICLR.cc/2024/Conference/Submission756/Reviewer_TdTq"],"forum":"TOE6N8dp4w","number":3,"license":"CC BY 4.0","cdate":1698823572690,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission756/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636002850,"domain":"ICLR.cc/2024/Conference","replyto":"TOE6N8dp4w","id":"c32PiKdSqj","forumContent":{"TLDR":{"value":"We propose a technique to generate a synthetic text dataset which is differentially private w.r.t. original text dataset. We show state of the art results in terms of quality of synthetic data, as measured by performance on downstream task."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Privacy","synthetic data","large language models"]},"supplementary_material":{"value":"/attachment/a4a173e8166c51b018d80e017233d59f7ffb349f.pdf"},"primary_area":{"value":"societal considerations including fairness, safety, privacy"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Differentially private training algorithms like DP-SGD protect sensitive training data by ensuring that trained models do not reveal private information. An alternative approach, which this paper studies, is to use a sensitive dataset to generate synthetic data that is differentially private with respect to the original data, and then non-privately training a model on the synthetic data.  Doing so has several advantages: synthetic data can be reused for other tasks (including for hyper parameter tuning), retained indefinitely, and shared with third parties without sacrificing privacy. \n\nHowever, generating private synthetic data is much harder than training a private model. To improve performance on text data, recent work has utilized public data by starting with a pre-trained generative language model and privately fine-tuning it on sensitive data. This model can be used to sample a DP synthetic dataset. While this strategy seems straightforward, executing it has proven problematic. Previous approaches either show significant performance loss, or have, as we show, critical design flaws.\n \nIn this paper we demonstrate that a proper training objective along with tuning fewer parameters results in excellent DP synthetic data quality. Our approach is competitive with direct DP-training of downstream classifiers in terms of performance on downstream tasks. Further, we demonstrate that our DP synthetic data is not only useful for downstream classifier training, but also to tune those same models."},"_bibtex":{"value":"@misc{\nkurakin2024harnessing,\ntitle={Harnessing large-language models to generate private synthetic text},\nauthor={Alexey Kurakin and Natalia Ponomareva and Umar Syed and Liam MacDermed and Andreas Terzis},\nyear={2024},\nurl={https://openreview.net/forum?id=TOE6N8dp4w}\n}"},"title":{"value":"Harnessing large-language models to generate private synthetic text"},"pdf":{"value":"/pdf/a01eed944f7131b60f836a4dd29be2772f6a6965.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"kurakin|harnessing_largelanguage_models_to_generate_private_synthetic_text"},"authorids":{"value":["~Alexey_Kurakin1","~Natalia_Ponomareva1","~Umar_Syed1","~Liam_MacDermed1","~Andreas_Terzis1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Alexey Kurakin","Natalia Ponomareva","Umar Syed","Liam MacDermed","Andreas Terzis"]}},"version":2},{"content":{"review":{"value":"This paper introduces FEABench, a novel benchmark for evaluating the ability of large language models (LLMs) to solve real-world physics and engineering problems using finite element analysis (FEA) software. The authors present a multi-faceted evaluation scheme to assess LLMs' capabilities in interacting with COMSOL Multiphysics software through its API to solve a range of physics and engineering problems.\n\nPros:\n1. Novel and important contribution: The paper addresses a significant gap in evaluating LLMs on real-world engineering tasks, specifically in the domain of finite element analysis.\n2. Dataset creation: FEABench includes 13 quantitatively verifiable problems across various physics domains, providing a comprehensive test of LLM capabilities.\n3. Multi-faceted evaluation: The authors develop a range of metrics to assess different aspects of LLM performance, including code executability, model tree similarity, and physics-specific metrics.\n4. Innovative approach: The authors develop an LLM agent equipped with the ability to interact with the FEA software API and iterate on solutions.\n5. Experimentation: The paper evaluates multiple state-of-the-art LLMs and compares different prompting strategies.\nClear presentation: The methodology, metrics, and results are generally well-explained and illustrated with helpful figures and tables.\n\nCons:\n1. Limited scope: The benchmark is currently restricted to problems solvable with COMSOL Multiphysics. Including other FEA software could enhance generalizability.\n2. Relatively small dataset: With only 13 problems, the benchmark may not fully capture the diversity of real-world FEA tasks.\n3. Lack of human baseline: The paper would benefit from comparing LLM performance to that of human engineers on the same tasks.\n4. Potential bias in problem selection: The criteria for selecting problems could be more thoroughly justified to ensure a representative sample of FEA tasks.\n5. Limited analysis of failure modes: While the paper presents performance metrics, a deeper analysis of why LLMs fail on certain tasks could provide valuable insights.\n6. Reproducibility concerns: The reliance on proprietary software (COMSOL) may limit the reproducibility of the benchmark by other researchers.\n7. Questionable real-world relevance: The paper doesn't adequately demonstrate how the ability to generate COMSOL API calls translates to actual problem-solving capabilities in engineering contexts.\n\nGeneral Comments:\n- The paper makes a sizable contribution to the field of AI evaluation, specifically in assessing LLMs' capabilities in complex engineering tasks. The FEABench provides a  tool for measuring progress in applying LLMs to real-world physics and engineering problems. \n- However, the benchmark's current limitations in scope and scale somewhat constrain its immediate impact. Expanding the number and diversity of problems, including problems from other FEA software packages, and providing a human performance baseline would significantly strengthen the work.\n- The paper does highlight the current limitations of LLMs in solving complex physics problems, but it fails to provide substantive insights or clear paths forward. In its current form, the work serves more as a proof-of-concept for evaluating LLMs on FEA tasks rather than a robust, widely applicable benchmark.\n- Despite these limitations, the paper represents an important step forward in evaluating LLMs on practical engineering tasks. It opens up new avenues for research in applying AI to complex scientific and engineering problems, potentially leading to significant advancements in how these tasks are approached in the future."},"confidence":{"value":4},"rating":{"value":6},"title":{"value":"Evaluation of Language Models for Physics Problem-Solving Using FEA Software"}},"nonreaders":[],"tmdate":1728525562955,"tcdate":1727589431913,"writers":["NeurIPS.cc/2024/Workshop/MATH-AI","NeurIPS.cc/2024/Workshop/MATH-AI/Submission40/Reviewer_89Ha"],"signatures":["NeurIPS.cc/2024/Workshop/MATH-AI/Submission40/Reviewer_89Ha"],"forum":"2z4U9reLm9","number":1,"license":"CC BY 4.0","cdate":1727589431913,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/MATH-AI/Submission40/-/Official_Review","NeurIPS.cc/2024/Workshop/MATH-AI/-/Edit"],"mdate":1728525562955,"domain":"NeurIPS.cc/2024/Workshop/MATH-AI","replyto":"2z4U9reLm9","id":"eigUUGyc50","forumContent":{"TLDR":{"value":"How well can LMs leverage FEA software to simulate and solve problems that require numerical analysis?"},"venue":{"value":"MATH-AI 24"},"pdf":{"value":"/pdf/25efbc7cd98dd6b0b91459d8a6b1135a2e6e7893.pdf"},"keywords":{"value":["numerical analysis","finite element","benchmark","agents"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/MATH-AI"},"paperhash":{"value":"mudur|feabench_evaluating_language_models_on_real_world_physics_reasoning_ability"},"concurrent_submissions":{"value":"Neurips Open World Agents Workshop '24 (submitted), ICLR '25 (Likely to be submitted)"},"authorids":{"value":["~Nayantara_Mudur1","~Hao_Cui3","~Subhashini_Venugopalan2","~Paul_Raccuglia1","~Michael_Brenner1","~Peter_Christian_Norgaard1"]},"abstract":{"value":"Building precise simulations of the real world and using numerical methods to solve quantitative problems is an essential task in engineering and physics. We present FEABench, a benchmark to evaluate the ability of language models (LLMs) and LLM agents to simulate and solve physics, mathematics and engineering problems using finite element analysis (FEA) software. We introduce a multipronged evaluation scheme to investigate the ability of LLMs to solve these problems using COMSOL Multiphysics$^\\textregistered$. We further design an LLM agent equipped with the ability to interact with the software through its Application Programming Interface (API), examine its outputs and use tools to improve its solution over several iterations. Our best performing strategy generates executable API calls 88\\% of the time. However, this benchmark still proves to be challenging enough that the LLMs and agents we tested were not able to completely and correctly solve any problem. LLMs that can successfully interact with and operate FEA software to solve problems such as those in our benchmark would push the frontiers of automation in engineering. Acquiring this capability would augment LLMs' reasoning skills with the precision of numerical solvers and advance the development of autonomous systems that can tackle complex problems in the real world."},"_bibtex":{"value":"@inproceedings{\nmudur2024feabench,\ntitle={{FEAB}ench: Evaluating Language Models on Real World Physics Reasoning Ability},\nauthor={Nayantara Mudur and Subhashini Venugopalan and Hao Cui and Paul Raccuglia and Michael Brenner and Peter Christian Norgaard},\nbooktitle={The 4th Workshop on Mathematical Reasoning and AI at NeurIPS'24},\nyear={2024},\nurl={https://openreview.net/forum?id=2z4U9reLm9}\n}"},"title":{"value":"FEABench: Evaluating Language Models on Real World Physics Reasoning Ability"},"authors":{"value":["Nayantara Mudur","Hao Cui","Subhashini Venugopalan","Paul Raccuglia","Michael Brenner","Peter Christian Norgaard"]}},"version":2},{"content":{"summary":{"value":"PIBNet is a novel machine learning method designed to accelerate the computationally intensive Boundary Element Method (BEM) used for solving scattering problems governed by PDEs like the Helmholtz equation. Its core innovation is replacing the slow iterative numerical solution for the boundary trace. It achieves this by utilizing a Physics-Informed Neural Network (PINN) approach where the loss function is specifically defined to minimize the residual of the Boundary Integral Equation (BIE), thus embedding the physics directly into the model's training. This approach yields significant inference speedups, demonstrating factors up to $200\\times$ faster than traditional BEM while maintaining accurate results."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"q1: The speedup of PIBNet relies on replacing the iterative GMRES solve. However, the BIE loss function still requires repeated evaluation of the Boundary Integral Operators, please clarify the computational cost of this BIE residual evaluation. Is this calculated analytically or numerically, and how does its complexity scale with the number of boundary elements ?\n\nq2: The BEM for the Helmholtz equation is known to suffer from the non-uniqueness issue at interior resonant frequencies (the \"fictitious frequencies\" problem). Does PIBNet, by training on the standard BIE formulation, inherit this instability?\n\nq3:  The paper demonstrates generalization across shapes and wavenumbers within the training distribution. Could the authors provide a more detailed analysis on the model's interpolation performance?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Technically robust framework validated across three key PDEs (Helmholtz/Laplace), demonstrating high-fidelity prediction and a substantial inference time speedup of up to $200\\times$.\n2. It replaces the GMRES solver bottleneck in BEM by minimizing the Boundary Integral Equation (BIE) residual directly, creating a novel, physics-constrained acceleration method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Generalization to Out-of-Distribution Inputs, different shapes not present in the training distribution (e.g., highly concave, asymmetric, or multiply connected domains).\n2. Comparison to State-of-the-Art Accelerated BEM: The comparison is made against the standard, non-accelerated BEM (solving the dense BIE matrix). The authors should include a quantitative comparison against the gold standard for large-scale BEM: the Fast Multipole Method (FMM)-accelerated BEM. \n3. Overall Cost Justification (Training Data Expense): The paper rightly focuses on the inference speedup, but this gain is predicated on the initial investment of generating a large dataset by performing thousands of expensive, full BEM solves (the ground truth). The authors should quantify the total computational cost (data generation + network training)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922425287,"tcdate":1761961224601,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11272/Reviewer_hhkB"],"signatures":["ICLR.cc/2026/Conference/Submission11272/Reviewer_hhkB"],"forum":"MiYStkjrfo","number":3,"license":"CC BY 4.0","cdate":1761961224601,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11272/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922425287,"domain":"ICLR.cc/2026/Conference","replyto":"MiYStkjrfo","id":"QV6Hx3hrOr","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Graph neural network","Multiple scattering","Boundary Element Method"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"The boundary element method (BEM) provides an efficient numerical framework for solving multiple scattering problems in unbounded homogeneous domains, since it reduces the discretization to the domain boundaries, thereby condensing the computational complexity. \nThe procedure first consists in determining the solution trace on the boundaries of the domain by solving a boundary integral equation, after which the volumetric solution can be recovered at low computational cost with a boundary integral representation.\nAs the first step of the BEM represents the main computational bottleneck, we introduce PIBNet, a learning-based approach designed to approximate the solution trace. The method leverages a physics-inspired graph-based strategy to model obstacles and their long-range interactions efficiently.\nThen, we introduce a novel multiscale graph neural network architecture for simulating the multiple scattering.\nTo train and evaluate our network, we present a benchmark consisting of several datasets of different types of multiple scattering problems. \nThe results indicate that our approach not only surpasses existing state-of-the-art learning-based methods on the considered tasks but also exhibits superior generalization to settings with an increased number of obstacles.\nCode available upon acceptance."},"_bibtex":{"value":"@misc{\nmarsal2026pibnet,\ntitle={{PIBN}et: a Physics-Inspired Boundary Network for Multiple Scattering Simulations},\nauthor={R{\\'e}mi Marsal and St{\\'e}phanie Chaillat},\nyear={2026},\nurl={https://openreview.net/forum?id=MiYStkjrfo}\n}"},"title":{"value":"PIBNet: a Physics-Inspired Boundary Network for Multiple Scattering Simulations"},"pdf":{"value":"/pdf/43cc0a3cc5762ac6c500b0051e99faa93a6b98d9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"marsal|pibnet_a_physicsinspired_boundary_network_for_multiple_scattering_simulations"},"authorids":{"value":["~Rémi_Marsal1","~Stéphanie_Chaillat1"]},"authors":{"value":["Rémi Marsal","Stéphanie Chaillat"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Geosci. Remote. Sens. 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/36/10006360/10102460.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"asiyabi|complexvalued_endtoend_deep_network_with_coherency_preservation_for_complexvalued_sar_data_reconstruction_and_classification"},"html":{"value":"https://doi.org/10.1109/TGRS.2023.3267185"},"_bibtex":{"value":"@article{DBLP:journals/tgrs/AsiyabiDAN23,\n  author={Reza Mohammadi Asiyabi and Mihai Datcu and Andrei Anghel and Holger Nies},\n  title={Complex-Valued End-to-End Deep Network With Coherency Preservation for Complex-Valued SAR Data Reconstruction and Classification},\n  year={2023},\n  cdate={1672531200000},\n  journal={IEEE Trans. Geosci. Remote. Sens.},\n  volume={61},\n  pages={1-17},\n  url={https://doi.org/10.1109/TGRS.2023.3267185}\n}\n"},"abstract":{"value":"Deep learning models have achieved remarkable success in many different fields and attracted many interests. Several researchers attempted to apply deep learning models to synthetic aperture radar (SAR) data processing, but it did not have the same breakthrough as the other fields, including optical remote sensing. SAR data are in complex domain by nature and processing them with real-valued (RV) networks neglects the phase component which conveys important and distinctive information. A complex-valued (CV) end-to-end deep network is developed in this study for the reconstruction and classification of CV-SAR data. Azimuth subaperture decomposition is utilized to incorporate physics-aware attributes of the CV-SAR into the deep model. Moreover, the correlation coefficient amplitude (coherence) of the CV-SAR images depends on the SAR system characteristics and physical properties of the target. This coherency should be considered and preserved in the processing chain of the CV-SAR data. The coherency preservation of the CV deep networks for CV-SAR images, which is mostly neglected in the literature, is evaluated in this study. Furthermore, a large-scale CV-SAR annotated dataset for the evaluation of the CV deep networks is lacking. A semantically annotated CV-SAR dataset from Sentinel-1 single look complex stripmap mode data [S1SLC_CVDL (complex-valued deep learning) dataset] is developed and introduced in this study. The experimental analysis demonstrated the better performance of the developed CV deep network for CV-SAR data classification and reconstruction in comparison with the equivalent RV model and more complicated RV architectures, as well as its coherency preservation and physics-aware capability."},"title":{"value":"Complex-Valued End-to-End Deep Network With Coherency Preservation for Complex-Valued SAR Data Reconstruction and Classification"},"authors":{"value":[{"fullname":"Reza Mohammadi Asiyabi","username":"~Reza_Mohammadi_Asiyabi1"},{"fullname":"Mihai Datcu","username":""},{"fullname":"Andrei Anghel","username":""},{"fullname":"Holger Nies","username":""}]}},"tmdate":1777646197193,"pdate":1703980800000,"externalIds":["dblp:journals/tgrs/AsiyabiDAN23"],"tcdate":1777646189747,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Reza_M._Asiyabi1"],"forum":"VLTd1BJWyX","license":"CC BY-SA 4.0","number":1748,"cdate":1672531200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1777646197193,"domain":"OpenReview.net/Public_Article","id":"VLTd1BJWyX","version":2},{"content":{"venue":{"value":"IJCNN 2020"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/9200848/9206590/09207122.pdf"},"venueid":{"value":"dblp.org/conf/IJCNN/2020"},"paperhash":{"value":"sunaga|similar_landform_discovery_complex_absolutevalue_max_pooling_in_complexvalued_convolutional_neural_networks_in_interferometric_synthetic_aperture_radar"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yuki_Sunaga:","~Ryo_Natsuaki1","~Akira_Hirose1"]},"html":{"value":"https://doi.org/10.1109/IJCNN48605.2020.9207122"},"_bibtex":{"value":"@inproceedings{DBLP:conf/ijcnn/SunagaNH20,\n  author={Yuki Sunaga and Ryo Natsuaki and Akira Hirose},\n  title={Similar land-form discovery: Complex absolute-value max pooling in complex-valued convolutional neural networks in interferometric synthetic aperture radar},\n  year={2020},\n  cdate={1577836800000},\n  pages={1-7},\n  url={https://doi.org/10.1109/IJCNN48605.2020.9207122},\n  booktitle={IJCNN},\n  crossref={conf/ijcnn/2020}\n}\n"},"abstract":{"value":"In a complex-valued convolutional neural network, its elementary unit consists of a complex-valued convolution layer and a complex pooling layer. The pooling layer has a variety in its dynamics. In this paper, we propose complex absolute-value max pooling to extract complex-amplitude feature patterns meaningful for discovery and/or adaptive classification of land form in interferometric synthetic aperture radar (InSAR). Experimental examination into amplitude and phase values in convolutional kernels reveals that useful land-shape features emerge through self-organization in high-magnitude kernels, which suggests that the proposed dynamics is successful in extracting important features."},"title":{"value":"Similar land-form discovery: Complex absolute-value max pooling in complex-valued convolutional neural networks in interferometric synthetic aperture radar"},"authors":{"value":["Yuki Sunaga","Ryo Natsuaki","Akira Hirose"]}},"tmdate":1749686748612,"pdate":1577836800000,"tcdate":1748057596677,"writers":["~"],"signatures":["~Akira_Hirose1"],"forum":"3pyUvmaB80","license":"CC BY-SA 4.0","number":550258,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1749686748612,"domain":"DBLP.org","id":"3pyUvmaB80","version":2},{"content":{"venue":{"value":"Crossref"},"pdf":{"value":"https://www.preprints.org/frontend/manuscript/46788ae3877c7feb6a00e9865b18f5cd/download_pub"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"ghosh|a_taxonomic_survey_of_physicsinformed_machine_learning"},"html":{"value":"https://doi.org/10.20944/preprints202305.0839.v1"},"abstract":{"value":"Physics-informed machine learning (PIML) refers to the emerging area of extracting physically relevant solutions to complex multiscale modeling problems and has enjoyed significant interest from the research community. This paper discusses the recent critical advancements in the PIML domain. Novel methods and applications of domain decomposition in physics-informed neural networks (PINN) in particular are highlighted. Additionally, we explore recent Works toward utilizing neural operator learning to intuit relationships in physics systems traditionally modeled by sets of complex governing equations and solved with expensive differentiation techniques. Finally, expansive applications of traditional physics-informed machine learning and potential limitations are discussed. In addition to summarizing recent work, we propose a novel taxonomic structure to catalog physics-informed machine learning based on how the physics-information is derived and injected into the machine learning process. The taxonomy assumes the explicit objectives of facilitating interdisciplinary collaboration in methodology, thereby promoting a wider characterization of what types of physics-problems are served by the physics-informed learning machines, and assisting in identifying apt targets for future work. To summarize, the major twofold goal of this work is to summarize recent advancements and introduce a taxonomic catalog for applications of physics-informed machine learning."},"title":{"value":"A Taxonomic Survey of Physics-Informed Machine Learning"},"authors":{"value":[{"fullname":"Preetam Ghosh"},{"fullname":"Joseph Pateras"},{"fullname":"Pratip Rana","username":"~Pratip_Rana1"}]}},"tmdate":1789092188601,"pdate":1683763200000,"externalIds":["doi:10.20944/preprints202305.0839.v1"],"tcdate":1772857826191,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Pratip_Rana1"],"forum":"IuFbkVxfmR","license":"CC BY-SA 4.0","number":53706,"cdate":1683855487450,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789092188601,"domain":"OpenReview.net/Public_Article","id":"IuFbkVxfmR","version":2},{"content":{"venue":{"value":"Applied Sciences"},"pdf":{"value":"https://www.mdpi.com/2076-3417/13/12/6892/pdf?version=1686110960"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"pateras|a_taxonomic_survey_of_physicsinformed_machine_learning"},"html":{"value":"https://doi.org/10.3390/app13126892"},"abstract":{"value":"Physics-informed machine learning (PIML) refers to the emerging area of extracting physically relevant solutions to complex multiscale modeling problems lacking sufficient quantity and veracity of data with learning models informed by physically relevant prior information. This work discusses the recent critical advancements in the PIML domain. Novel methods and applications of domain decomposition in physics-informed neural networks (PINNs) in particular are highlighted. Additionally, we explore recent works toward utilizing neural operator learning to intuit relationships in physics systems traditionally modeled by sets of complex governing equations and solved with expensive differentiation techniques. Finally, expansive applications of traditional physics-informed machine learning and potential limitations are discussed. In addition to summarizing recent work, we propose a novel taxonomic structure to catalog physics-informed machine learning based on how the physics-information is derived and injected into the machine learning process. The taxonomy assumes the explicit objectives of facilitating interdisciplinary collaboration in methodology, thereby promoting a wider characterization of what types of physics problems are served by the physics-informed learning machines and assisting in identifying suitable targets for future work. To summarize, the major twofold goal of this work is to summarize recent advancements and introduce a taxonomic catalog for applications of physics-informed machine learning."},"title":{"value":"A Taxonomic Survey of Physics-Informed Machine Learning"},"authors":{"value":[{"fullname":"Joseph Pateras"},{"fullname":"Pratip Rana","username":"~Pratip_Rana1"},{"fullname":"Preetam Ghosh"}]}},"tmdate":1789092114434,"pdate":1686096000000,"externalIds":["doi:10.3390/app13126892"],"tcdate":1772857475114,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Pratip_Rana1"],"forum":"KCfa936hHF","license":"CC BY-SA 4.0","number":53695,"cdate":1686133902070,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789092114434,"domain":"OpenReview.net/Public_Article","id":"KCfa936hHF","version":2},{"content":{"summary":{"value":"This paper presents a method to predict hard X-ray (HXR) energies emitted by the hot electrons in ICF implosions in nuclear physics. The method relies on engineering fusion-specific prompts for LLMs, along with \"signal-digesting channels\", to predict these time series data. The predictions themselves are then calibrated using a \"confidence scanner\" to try to characterize the confidence in predictions across the dataset."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The integration of LLMs with reservoir computing for scientific applications appears novel both in the nuclear physics domain and across scientific computing more generally. Additionally, the design of the signal-digesting channels is a novel addition.\n- Experimental results show that LPI-LLM achieves superior forecasting accuracy relative to the other data-driven approaches presented.\n- The introduction of the LPI4AI benchmark dataset provides a valuable resource for the scientific machine learning community.\n- The ablation studies provided in table 2C are useful."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The introduction leans heavily on specialized scientific language, potentially making it less accessible for the broader NeurIPS community unfamiliar with ICF or plasma physics.\n- The paper does not incorporate or mention any non-ML models that might be commonly used for these prediction tasks in plasma physics. This would provide a meaningful performance benchmark from within the field.\n- The confidence scanner’s accuracy is only demonstrated through select examples. However, there is no quantifiable metric provided. Could the authors perform any expected calibration test on these results? This is particularly valuable in light of the fact that the confidence scanner is not shown to produce a meaningful likelihood analytically."}},"nonreaders":[],"tmdate":1731428035960,"tcdate":1730758892509,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4098/Reviewer_yDn5"],"signatures":["ICLR.cc/2025/Conference/Submission4098/Reviewer_yDn5"],"forum":"JQrBYfD2gg","number":3,"license":"CC BY 4.0","cdate":1730758892509,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4098/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428035960,"domain":"ICLR.cc/2025/Conference","replyto":"JQrBYfD2gg","id":"g8zkv16A6q","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"LPI-LLM combines LLMs with reservoir computing to accurately predict laser-plasma instabilities in fusion experiments, outperforming existing methods and offering a cost-effective alternative to traditional simulations."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["AI for Science","Inertial Confinement Fusion"]},"supplementary_material":{"value":"/attachment/9c3452a6421e52d6d22f5980ea6badbd71b4ea68.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Controlled fusion energy is deemed pivotal for the advancement of human civilization. In this study, we introduce $\\textbf{LPI-LLM}$, a novel integration of Large Language Models (LLMs) with classical reservoir computing paradigms tailored to address a critical challenge, Laser-Plasma Instabilities ($\\texttt{LPI}$), in Inertial Confinement Fusion ($\\texttt{ICF}$). Our approach offers several key contributions: Firstly, we propose the $\\textit{LLM-anchored Reservoir}$, augmented with a $\\textit{Fusion-specific Prompt}$, enabling accurate forecasting of $\\texttt{LPI}$-generated-hot electron dynamics during implosion. Secondly, we develop $\\textit{Signal-Digesting Channels}$ to temporally and spatially describe the driver laser intensity across time, capturing the unique characteristics of $\\texttt{ICF}$ inputs. Lastly, we design the $\\textit{Confidence Scanner}$ to quantify the confidence level in forecasting, providing valuable insights for domain experts to design the $\\texttt{ICF}$ process. Extensive experiments demonstrate the superior performance of our method, achieving 1.90 CAE, 0.14 $\\texttt{top-1}$ MAE, and 0.11 $\\texttt{top-5}$ MAE in predicting Hard X-ray ($\\texttt{HXR}$) energies emitted by the hot electrons in $\\texttt{ICF}$ implosions, which presents state-of-the-art comparisons against concurrent best systems.  Additionally, we present $\\textbf{LPI4AI}$, the first $\\texttt{LPI}$ benchmark based on physical experiments, aimed at fostering novel ideas in $\\texttt{LPI}$ research and enhancing the utility of LLMs in scientific exploration. Overall, our work strives to forge an innovative synergy between AI and $\\texttt{ICF}$ for advancing fusion energy."},"_bibtex":{"value":"@misc{\nchen2025inertial,\ntitle={Inertial Confinement Fusion Forecasting via Large Language Models},\nauthor={Mingkai Chen and Taowen Wang and shihui cao and James Chenhao Liang and Chuan Liu and Chunshu Wu and Qifan Wang and Ying Nian Wu and Michael Huang and Chuang Ren and Ang Li and Tong Geng and Dongfang Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=JQrBYfD2gg}\n}"},"title":{"value":"Inertial Confinement Fusion Forecasting via Large Language Models"},"pdf":{"value":"/pdf/ef4deeaf56c930720600e6b4c708e39b829ed4e1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|inertial_confinement_fusion_forecasting_via_large_language_models"},"authorids":{"value":["~Mingkai_Chen1","~Taowen_Wang1","~shihui_cao1","~James_Chenhao_Liang1","~Chuan_Liu3","~Chunshu_Wu1","~Qifan_Wang2","~Ying_Nian_Wu1","~Michael_Huang1","~Chuang_Ren1","~Ang_Li11","~Tong_Geng1","~Dongfang_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mingkai Chen","Taowen Wang","shihui cao","James Chenhao Liang","Chuan Liu","Chunshu Wu","Qifan Wang","Ying Nian Wu","Michael Huang","Chuang Ren","Ang Li","Tong Geng","Dongfang Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes GENIE, a hybrid model that combines the high-quality rendering capabilities of NeRF with the editability of 3D Gaussian Splatting. The paper introduces the RT-GPS algorithm for fast nearest neighbor search and Splash Grid Encoding for multi-resolution feature encoding. The method supports real-time local editing and integration with physics-based simulation."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. What is the speedup ratio of RT-GPS compared to brute-force nearest neighbor search? Can you provide detailed time complexity analysis and experimental data?\n\n2. The ablation study in Table 3 shows that removing Splash Grid Encoding causes OOM in some scenes (Ship, Mic). Does this indicate a design flaw in this module?\n\n3. How do you ensure view consistency of rendering results after physics simulation? Are there artifacts from multiple viewpoints?\n\n4. Is the method applicable to dynamic scenes or time-varying illumination?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The overall structure is clear, with Figures 1 and 4 providing excellent overviews of the methodology.\n\n2. Rich visualization of physics simulation results (Figures 2, 3, and 6)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The notation is sometimes inconsistent. For example, the definition of G_GENIE in Section 4 is overly complex, which compromises readability.\n\n2. Some key concepts lack clear explanation. For instance, the use of \"triangle soup\" (Section 7) is not friendly to readers without a graphics background.\n\n3. The baseline selection criteria for Tables 1 and 2 are not explained. Why are there no RIP-NeRF results on the Mip-NeRF 360 dataset? \n\n4. The number of comparison methods is insufficient.\n\n5. Figure 8 has low visual quality with text labels that are too small.\n\n6. The experiments are insufficient in scope.\n\n7. Table 2 shows that GENIE's PSNR is 5-7 dB lower than Mip-NeRF on Mip-NeRF 360, with even larger gaps in SSIM and LPIPS metrics. While the authors attribute this to Gaussian density, there lacks systematic analysis. The authors need to strengthen their justification.\n\n8. There is no comparison of training time."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918586959,"tcdate":1761999547072,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6281/Reviewer_g8yM"],"signatures":["ICLR.cc/2026/Conference/Submission6281/Reviewer_g8yM"],"forum":"wyauZpoqRK","number":3,"license":"CC BY 4.0","cdate":1761999547072,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6281/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918586959,"domain":"ICLR.cc/2026/Conference","replyto":"wyauZpoqRK","id":"PhAKhMoJ6s","forumContent":{"TLDR":{"value":"his work introduces GENIE, a hybrid model that fuses NeRF’s photorealism with 3D Gaussian Splatting’s editability, enabling interactive 3D scene editing via Gaussian-based conditioning and fast nearest Gaussian search."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Gaussian Splatting","NeRF","3D objects","physical simulations"]},"supplementary_material":{"value":"/attachment/b8f83c0e55c5d12a6c9fc74eec83e45ec06b0cd4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have recently transformed 3D scene representation and rendering. NeRF achieves high-fidelity novel view synthesis by learning volumetric representations through neural networks, but its implicit encoding makes editing and physical interaction challenging. In contrast, 3DGS represents scenes as explicit collections of Gaussian primitives, enabling real-time rendering, faster training, and more intuitive manipulation. This explicit structure has made 3DGS particularly well-suited for interactive editing and integration with physics-based simulation. In this paper, we introduce GENIE (Gaussian Encoding for Neural Radiance Fields Interactive Editing), a hybrid model that combines the photorealistic rendering quality of NeRF with the editable and structured representation of 3DGS. Instead of using spherical harmonics for appearance modeling, we assign each Gaussian a trainable feature embedding. These embeddings are used to condition a NeRF network based on the \n nearest Gaussians to each query point. To make this conditioning efficient, we introduce Ray-Traced Gaussian Proximity Search (RTGPS), a fast nearest Gaussian search based on a modified ray-tracing pipeline. We also integrate a multi-resolution hash grid to initialize and update Gaussian features. Together, these components enable real-time, locality-aware editing: as Gaussian primitives are repositioned or modified, their interpolated influence is immediately reflected in the rendered output. By combining the strengths of implicit and explicit representations, GENIE supports intuitive scene manipulation, dynamic interaction, and compatibility with physical simulation, bridging the gap between geometry-based editing and neural rendering."},"_bibtex":{"value":"@misc{\nzielinski2025genie,\ntitle={{GENIE}: Gaussian Encoding for Neural Radiance Field Interactive Editing},\nauthor={Miko{\\l}aj Zieli{\\'n}ski and Krzysztof Byrski and Tomasz Szczepanik and Przemys{\\l}aw Spurek},\nyear={2025},\nurl={https://openreview.net/forum?id=wyauZpoqRK}\n}"},"title":{"value":"GENIE: Gaussian Encoding for Neural Radiance Field Interactive Editing"},"pdf":{"value":"/pdf/9683e9431e57f996c107fc4e7f18a38fd6e208e7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zieliski|genie_gaussian_encoding_for_neural_radiance_field_interactive_editing"},"authorids":{"value":["~Mikołaj_Zieliński1","~Krzysztof_Byrski1","~Tomasz_Szczepanik1","~Przemysław_Spurek1"]},"authors":{"value":["Mikołaj Zieliński","Krzysztof Byrski","Tomasz Szczepanik","Przemysław Spurek"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2503.20822v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"zhao|synthetic_video_enhances_physical_fidelity_in_video_synthesis"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Qi_Zhao:","https://dblp.org/search/pid/api?q=author:Xingyu_Ni:","https://dblp.org/search/pid/api?q=author:Ziyu_Wang:","~Feng_Cheng2","https://dblp.org/search/pid/api?q=author:Ziyan_Yang:","https://dblp.org/search/pid/api?q=author:Lu_Jiang:","https://dblp.org/search/pid/api?q=author:Bohan_Wang:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2503.20822"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2503-20822,\n  publtype={informal},\n  author={Qi Zhao and Xingyu Ni and Ziyu Wang and Feng Cheng and Ziyan Yang and Lu Jiang and Bohan Wang},\n  title={Synthetic Video Enhances Physical Fidelity in Video Synthesis},\n  year={2025},\n  month={March},\n  cdate={1740787200000},\n  journal={CoRR},\n  volume={abs/2503.20822},\n  url={https://doi.org/10.48550/arXiv.2503.20822}\n}\n"},"abstract":{"value":"We investigate how to enhance the physical fidelity of video generation models by leveraging synthetic videos derived from computer graphics pipelines. These rendered videos respect real-world physics, such as maintaining 3D consistency, and serve as a valuable resource that can potentially improve video generation models. To harness this potential, we propose a solution that curates and integrates synthetic data while introducing a method to transfer its physical realism to the model, significantly reducing unwanted artifacts. Through experiments on three representative tasks emphasizing physical consistency, we demonstrate its efficacy in enhancing physical fidelity. While our model still lacks a deep understanding of physics, our work offers one of the first empirical demonstrations that synthetic video enhances physical fidelity in video synthesis. Website: https://kevinz8866.github.io/simulation/"},"title":{"value":"Synthetic Video Enhances Physical Fidelity in Video Synthesis"},"authors":{"value":["Qi Zhao","Xingyu Ni","Ziyu Wang","Feng Cheng","Ziyan Yang","Lu Jiang","Bohan Wang"]}},"tmdate":1747688151703,"pdate":1735689600000,"tcdate":1747688148494,"writers":["~"],"signatures":["~Feng_Cheng2"],"forum":"Y7x8t2IRjL","license":"CC BY-SA 4.0","number":529650,"cdate":1740787200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747688151703,"domain":"DBLP.org","id":"Y7x8t2IRjL","version":2},{"content":{"venue":{"value":"Crossref"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-19833-5_24.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"duan|pip_physical_interaction_prediction_via_mental_simulation_with_span_selection"},"html":{"value":"https://doi.org/10.1007/978-3-031-19833-5_24"},"abstract":{"value":"Accurate prediction of physical interaction outcomes is a crucial component of human intelligence and is important for safe and efficient deployments of robots in the real world. While there are existing vision-based intuitive physics models that learn to predict physical interaction outcomes, they mostly focus on generating short sequences of future frames based on physical properties (e.g. mass, friction and velocity) extracted from visual inputs or a latent space. However, there is a lack of intuitive physics models that are tested on long physical interaction sequences with multiple interactions among different objects. We hypothesize that selective temporal attention during approximate mental simulations helps humans in physical interaction outcome prediction. With these motivations, we propose a novel scheme: Physical Interaction Prediction via Mental Simulation with Span Selection (PIP). It utilizes a deep generative model to model approximate mental simulations by generating future frames of physical interactions before employing selective temporal attention in the form of span selection for predicting physical interaction outcomes. To the best of our knowledge, attention has not been used with deep learning to tackle intuitive physics. For model evaluation, we further propose the large-scale SPACE+ dataset of synthetic videos with long sequences of three prime physical interactions in a 3D environment. Our experiments show that PIP outperforms human, baseline, and related intuitive physics models that utilize mental simulation. Furthermore, PIP’s span selection module effectively identifies the frames indicating key physical interactions among objects, allowing for added interpretability, and does not require labor-intensive frame annotations. PIP is available on https://sites.google.com/view/piphysics."},"title":{"value":"PIP: Physical Interaction Prediction via Mental Simulation with Span Selection"},"authors":{"value":[{"fullname":"Jiafei Duan","username":"https://orcid.org/orcid-search/search?searchQuery=Jiafei%20Duan"},{"fullname":"Samson Yu","username":"https://orcid.org/orcid-search/search?searchQuery=Samson%20Yu"},{"fullname":"Soujanya Poria","username":"~Soujanya_Poria1"},{"fullname":"Bihan Wen","username":"https://orcid.org/orcid-search/search?searchQuery=Bihan%20Wen"},{"fullname":"Cheston Tan","username":"https://orcid.org/orcid-search/search?searchQuery=Cheston%20Tan"}]}},"tmdate":1788084676968,"pdate":1640995200000,"externalIds":["doi:10.1007/978-3-031-19833-5_24"],"tcdate":1788084667347,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Soujanya_Poria1"],"forum":"8dhc40XZtC","license":"CC BY-SA 4.0","number":105751,"cdate":1667527725993,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1788084676968,"domain":"OpenReview.net/Public_Article","id":"8dhc40XZtC","version":2},{"content":{"summary":{"value":"This paper focuses on the Referring Expression Segmentation (RES) task. To address the issues of limited evaluation capability in existing benchmarks and suboptimal performance of models in real-world scenarios, this paper proposes the WildRES benchmark dataset and the SynRES synthetic data generation and augmentation pipeline, and validates the effectiveness of the proposed approach."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. WildRES only covers three types of cross-domain scenarios: dense crowds (CrowdHuman), autonomous driving (Cityscapes), and robotics (ARMBench) (Section 3.2). What is the main rationale for the authors choosing these three types of scenarios? Does it verify whether the domain gaps in these three scenarios can represent the real cross-domain requirements of RES?\n\n2. The paper only focuses on segmentation accuracy (gIoU/cIoU) and does not test the training efficiency of the model enhanced by SynRES—whether adding synthetic data for training affects the training speed—and should conduct a more thorough performance comparison with existing RES models.\n\n3. SynRES's text augmentation uses hypernym replacement (e.g., woman→person, Section 4.3), with a default replacement probability of p = 0.7 (Section 5.1). However, the authors need to explain whether hypernym replacement might cause semantic ambiguity when the class noun in the expression is strongly associated with an attribute (e.g., pregnant woman → person, where the attribute \"pregnant\" is tightly bound to \"female\"). It is necessary to clarify the adaptive differences under different replacement probabilities, and if replacement leads to semantic ambiguity, whether it could introduce new segmentation errors."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. WildRES addresses the evaluation of complex real-world scenarios in the RES field and can serve as a new standard for measuring models' complex reasoning abilities.\n\n2. Specifically, SynRES addresses the shortage of complex RES data at a low cost by automatically generating high-quality, densely annotated paired data. This pipeline can enhance model performance in complex queries and cross-domain scenarios, while remaining compatible with classic benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The mentioned issues include the lack of evaluation in complex scenarios, poor generalization to complex scenarios, and overly brief descriptions. However, these issues have been discussed in many existing RES or RIS benchmarks, such as MMR[1], LLM-Seg40K[2], ReasonSeg[3], MUSE[4], etc. Moreover, these benchmarks are more complex than the one proposed by the authors. The authors need to clearly explain the differences and advantages between the WildRES benchmark dataset and these existing benchmarks.\n\n2. The experiment only selects three open-source models (LISA, GSVA, and GLaMM) and fails to cover other typical RES architectures (e.g., lightweight Transformer-based models or cross-modal fusion models), which fails to fully demonstrate the generalizability of SynRES across different architectures. Additionally, the authors do not compare its performance with that of other existing models (such as [5][6][7]) on the RefCOCO series datasets. Furthermore, WildRES-DS only covers three types of scenarios, and this benchmark fails to truly achieve generalization to out-of-domain data.\n\n3. The SynRES synthetic data pipeline proposed by the authors is a combination of two types of data augmentation. Essentially, it merges Mosaic augmentation with text rewriting augmentation, both of which are already widely used in the RES field. The authors need to clarify the novelty of this pipeline.\n\n4. The authors obtain better test results by using more training data, rendering this comparison unfair. It is necessary to compare the speed and efficiency of this method with other methods—specifically in terms of training time and FLOPs—to determine if there are any differences.\n\n[1] MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation[C]//The Thirteenth International Conference on Learning Representations.\n\n[2] Llm-seg: Bridging image segmentation and large language model reasoning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 1765-1774.\n\n[3] Lisa: Reasoning segmentation via large language model[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 9579-9589.\n\n[4] Pixellm: Pixel reasoning with large multimodal model[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 26374-26383.\n\n[5]  Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks[J]. Advances in Neural Information Processing Systems, 2024, 37: 69925-69975.\n\n[6] Vltp: Vision-language guided token pruning for task-oriented segmentation[C]//2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2025: 9353-9363.\n\n[7]  F-lmm: Grounding frozen large multimodal models[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 24710-24721."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924823855,"tcdate":1761833540751,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14417/Reviewer_qvxs"],"signatures":["ICLR.cc/2026/Conference/Submission14417/Reviewer_qvxs"],"forum":"cgr5OAXe3q","number":3,"license":"CC BY 4.0","cdate":1761833540751,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14417/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924823855,"domain":"ICLR.cc/2026/Conference","replyto":"cgr5OAXe3q","id":"AQ6yFfnAqG","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Referring Expression Segmentation","Synthetic data","Multimodal augmentation"]},"supplementary_material":{"value":"/attachment/c086636c0c2aeea809f4586753e9c5607fc0af65.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Despite the advances in Referring Expression Segmentation (RES) benchmarks, their evaluation protocols remain constrained, primarily focusing on either single targets with short queries (containing minimal attributes) or multiple targets from distinctly different queries on a single domain. This limitation significantly hinders the assessment of more complex reasoning capabilities in RES models.\n    We introduce  WildRES, a novel benchmark that incorporates long queries with diverse attributes and non-distinctive queries for multiple targets. This benchmark spans diverse application domains, thus enabling more rigorous evaluation of complex reasoning capabilities in real-world settings. \n    Our analysis reveals that existing RES models demonstrate substantial performance deterioration when evaluated on WIldRES. To address this challenge, we introduce SynRES, an automated pipeline generating densely paired compositional synthetic training data through three innovations: (1) a dense caption-driven synthesis for attribute-rich image-mask-expression triplets, (2) reliable semantic alignment mechanisms rectifying caption-pseudo mask inconsistencies via Image-Text Aligned Grouping, and (3) domain-aware augmentations incorporating mosaic composition and superclass replacement to emphasize generalization ability and distinguishing attributes over object categories.\n    Experimental results demonstrate that models trained with SynRES achieve consistent improvements on not only our complex WildRES benchmark but also classic RES benchmarks (e.g. RefCOCO/+/g).\nCode is available at https://anonymous.4open.science/r/SynRES-Review-4B1F.\nDataset will be available upon acceptance."},"_bibtex":{"value":"@misc{\nkim2026towards,\ntitle={Towards Robust Referring Expression Segmentation for Complex Reasoning in the Wild},\nauthor={Dong-Hee Kim and Hyunjee Song and Donghyun Kim},\nyear={2026},\nurl={https://openreview.net/forum?id=cgr5OAXe3q}\n}"},"title":{"value":"Towards Robust Referring Expression Segmentation for Complex Reasoning in the Wild"},"pdf":{"value":"/pdf/a05b72f8da22c8e836233b9c978c4462034241ec.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|towards_robust_referring_expression_segmentation_for_complex_reasoning_in_the_wild"},"authorids":{"value":["~Dong-Hee_Kim1","~Hyunjee_Song1","~Donghyun_Kim2"]},"authors":{"value":["Dong-Hee Kim","Hyunjee Song","Donghyun Kim"]}},"version":2},{"content":{"summary":{"value":"This paper explores diffusion-based models for unnormalized sampling and introduces the Diffusion-PINN Sampler (DPS). Using physics-informed neural networks (PINNs), DPS directly approximates the log-density of SDE marginals, enabling more precise modeling of complex distributions. The authors provide theoretical convergence analysis and validate the method on several synthetic datasets."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- $\\ell_reg$ seems ad-hoc. Why regress onto the boundary condition of the score instead of (11b), which seems more aligned with PINN loss?\n- Given that x0 comes from LMC already, is the method computationally cheaper compared to standard MCMC methods?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper is generally easier to follow. Theoretical results in Sec 5 could be of interested for readers from other domains."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The proposed method is computationally expansive compared to other diffusion-based models to train—due to the Laplacian in PINN loss—and to sample from—due to the evaluation of gradient at every time step.\n- Insufficient comparison to modern baselines — in addition to  MC methods like HMC or SMC, there’re many other diffusion/SDE based methods, such as [1,2], just to name a few.\n- Experiments were only conducted on rather simple, synthetic, target. I’ll be more convinced to see some higher-dim experiments and/or real-world dataset. \n- Parametrize NN with log mu (12) seems like a strong inductive bias and can be infeasible for many practical applications (e.g., sample Boltzmann distribution for conformation generation) when querying energy functions are expansive. Can the authors provide ablation study without such parametrization? \n- Related works Section is rather short and should be extended.\n\n[1] Particle Denoising Diffusion Sampler (ICML 2024)\n[2] Improved sampling via learned diffusions (ICLR 2024)"}},"nonreaders":[],"tmdate":1732562202662,"tcdate":1730690911256,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5956/Reviewer_YacX"],"signatures":["ICLR.cc/2025/Conference/Submission5956/Reviewer_YacX"],"forum":"vxBvr5ZpIu","number":4,"license":"CC BY 4.0","cdate":1730690911256,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5956/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732562202662,"domain":"ICLR.cc/2025/Conference","replyto":"vxBvr5ZpIu","id":"v6yqppGYvC","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["posterior sampling","multi-modal sampling","mixing proportion identification","diffusion model","physics-informed neural network"]},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent success of diffusion models has inspired a surge of interest in developing sampling techniques using reverse diffusion processes. However, accurately estimating the drift term in the reverse stochastic differential equation (SDE) solely from the unnormalized target density poses significant challenges, hindering existing methods from achieving state-of-the-art performance. In this paper, we introduce the Diffusion-PINN Sampler (DPS), a novel diffusion-based sampling algorithm that estimates the drift term by solving the governing partial differential equation of the log-density of the underlying SDE marginals via physics-informed neural networks (PINN). We prove that the error of log-density approximation can be controlled by the PINN residual loss, enabling us to establish convergence guarantees of DPS. Experiments on a variety of sampling tasks demonstrate the effectiveness of our approach, particularly in accurately identifying mixing proportions when the target contains isolated components."},"_bibtex":{"value":"@misc{\nshi2025diffusionpinn,\ntitle={Diffusion-{PINN} Sampler},\nauthor={Zhekun Shi and Longlin Yu and Tianyu Xie and Cheng Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=vxBvr5ZpIu}\n}"},"title":{"value":"Diffusion-PINN Sampler"},"pdf":{"value":"/pdf/4bb3d35f78b16f58cfe9af9e1160df131f1b2e49.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"shi|diffusionpinn_sampler"},"authorids":{"value":["~Zhekun_Shi1","~Longlin_Yu1","~Tianyu_Xie1","~Cheng_Zhang3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhekun Shi","Longlin Yu","Tianyu Xie","Cheng Zhang"]}},"version":2},{"content":{"venue":{"value":"arXiv e-prints"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"xia|scatterprism_convergence_for_generative_simulation_and_inverse_problems_in_particle_and_nuclear_physics"},"abstract":{"value":"High-fidelity simulations and complex inverse problems, such as detector modeling and unfolding, are computationally intensive bottlenecks across subatomic physics, yet essential for accurate physical interpretation. While Conditional Flow Matching (CFM) offers a robust acceleration approach, we demonstrate its standard training loss is fundamentally misleading. Specifically, utilizing a Jefferson Lab Nuclear Physics (NP) kinematic dataset ($γp \\to ρ^0 p \\to π^+π^- p$), we expose that CFM loss plateaus prematurely, obscuring ongoing physical refinement. To verify this disconnect is a dataset-agnostic pathology, we introduce ScatterPrism, an efficient generative surrogate evaluated against both the NP data and synthetic stress tests modeling challenging 1D distribution topologies. Coupling these benchmarks, we establish that physics-informed metrics continue improving long after standard loss converges. Consequently, we propose a multi-metric diagnostic protocol to ensure true kinematic fidelity without data memorization. Driven by NP challenges relevant to the forthcoming Electron-Ion Collider (EIC), this unified machinery has strong potential to extend to High-Energy Physics (HEP) applications, such as jet modeling. Furthermore, the framework holds promise for broader domains requiring rigorous generative reliability, including medical imaging, astrophysics, and quantitative finance...."},"title":{"value":"ScatterPrism: convergence for generative simulation and inverse problems in particle and nuclear physics"},"authors":{"value":[{"fullname":"Zeyu Xia","username":""},{"fullname":"Tyler Kim","username":"https://orcid.org/orcid-search/search?searchQuery=Tyler%20Kim"},{"fullname":"Trevor Reed","username":"https://orcid.org/orcid-search/search?searchQuery=Trevor%20Reed"},{"fullname":"Judy Fox","username":"https://orcid.org/orcid-search/search?searchQuery=Judy%20Fox"},{"fullname":"Geoffrey Fox","username":"https://orcid.org/orcid-search/search?searchQuery=Geoffrey%20Fox"},{"fullname":"Adam Szczepaniak","username":"https://orcid.org/orcid-search/search?searchQuery=Adam%20Szczepaniak"}]}},"tmdate":1790234948386,"pdate":1775001600000,"externalIds":["doi:10.48550/arxiv.2604.01313"],"tcdate":1782290902696,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Zeyu_Xia1"],"forum":"rGt58xfHeb","license":"CC BY-SA 4.0","number":78352,"cdate":1775364499316,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim","OpenReview.net/Public_Article/-/Author_Removal"],"mdate":1790234948386,"domain":"OpenReview.net/Public_Article","id":"rGt58xfHeb","version":2},{"content":{"venue":{"value":"CVPR 2025"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2025/papers/Wu_PBR-NeRF_Inverse_Rendering_with_Physics-Based_Neural_Fields_CVPR_2025_paper.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2025"},"paperhash":{"value":"wu|pbrnerf_inverse_rendering_with_physicsbased_neural_fields"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Sean_Wu:","https://dblp.org/search/pid/api?q=author:Shamik_Basu:","~Tim_Broedermann1","https://dblp.org/search/pid/api?q=author:Luc_Van_Gool:","https://dblp.org/search/pid/api?q=author:Christos_Sakaridis:"]},"html":{"value":"https://openaccess.thecvf.com/content/CVPR2025/html/Wu_PBR-NeRF_Inverse_Rendering_with_Physics-Based_Neural_Fields_CVPR_2025_paper.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/WuBBGS25,\n  author={Sean Wu and Shamik Basu and Tim Broedermann and Luc Van Gool and Christos Sakaridis},\n  title={PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields},\n  year={2025},\n  cdate={1735689600000},\n  pages={10974-10984},\n  url={https://openaccess.thecvf.com/content/CVPR2025/html/Wu_PBR-NeRF_Inverse_Rendering_with_Physics-Based_Neural_Fields_CVPR_2025_paper.html},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2025}\n}\n"},"abstract":{"value":"We tackle the ill-posed inverse rendering problem in 3D reconstruction with a Neural Radiance Field (NeRF) approach informed by Physics-Based Rendering (PBR) theory, named PBR-NeRF. Our method addresses a key limitation in most NeRF and 3D Gaussian Splatting approaches: they estimate view-dependent appearance without modeling scene materials and illumination. To address this limitation, we present an inverse rendering (IR) model capable of jointly estimating scene geometry, materials, and illumination. Our model builds upon recent NeRF-based IR approaches, but crucially introduces two novel physics-based priors that better constrain the IR estimation. Our priors are rigorously formulated as intuitive loss terms and achieve state-of-the-art material estimation without compromising novel view synthesis quality. Our method is easily adaptable to other inverse rendering and 3D reconstruction frameworks that require material estimation. We demonstrate the importance of extending current neural rendering approaches to fully model scene properties beyond geometry and view-dependent appearance. Code is publicly available at: https://github.com/s3anwu/pbrnerf."},"title":{"value":"PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields"},"authors":{"value":["Sean Wu","Shamik Basu","Tim Broedermann","Luc Van Gool","Christos Sakaridis"]}},"tmdate":1762517923706,"pdate":1735689600000,"externalIds":["dblp:conf/cvpr/WuBBGS25"],"tcdate":1762505617592,"writers":["~"],"signatures":["~Tim_Broedermann1"],"forum":"I6D6nDisTL","license":"CC BY-SA 4.0","number":671565,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762517923706,"domain":"DBLP.org","id":"I6D6nDisTL","version":2},{"content":{"TLDR":{"value":"CVPR 2025 paper to appear that solves the ill-posed inverse rendering problem with a NeRF model optimized via physics-based priors and jointly estimating materials, illumination, and geometry"},"venue":{"value":"Greeks in AI 2025 Poster"},"pdf":{"value":"/pdf/5713e2c6dc55f30289fc9d75af202f3f9dc71d74.pdf"},"keywords":{"value":["neural rendering","inverse rendering","NeRF","3D reconstruction","physics-based rendering","BRDF","materials"]},"venueid":{"value":"greeksin.ai/Greeks_in_AI/2025/Symposium"},"paperhash":{"value":"wu|pbrnerf_inverse_rendering_with_physicsbased_neural_fields"},"authorids":{"value":["~Sean_Wu2","~Shamik_Basu1","~Tim_Broedermann1","~Luc_Van_Gool1","~Christos_Sakaridis1"]},"abstract":{"value":"We tackle the ill-posed inverse rendering problem in 3D reconstruction with a Neural Radiance Field (NeRF) approach informed by Physics-Based Rendering (PBR) theory, named PBR-NeRF. Our method addresses a key limitation in most NeRF and 3D Gaussian Splatting approaches: they estimate view-dependent appearance without modeling scene materials and illumination. To address this limitation, we present an inverse rendering (IR) model capable of jointly estimating scene geometry, materials, and illumination. Our model builds upon recent NeRF-based IR approaches, but crucially introduces two novel physics-based priors that better constrain the IR estimation. Our priors are rigorously formulated as intuitive loss terms and achieve state-of-the-art material estimation without compromising novel view synthesis quality. Our method is easily adaptable to other inverse rendering and 3D reconstruction frameworks that require material estimation. We demonstrate the importance of extending current neural rendering approaches to fully model scene properties beyond geometry and view-dependent appearance. Code is publicly available at: https://github.com/s3anwu/pbrnerf."},"_bibtex":{"value":"@inproceedings{\nwu2025pbrnerf,\ntitle={{PBR}-Ne{RF}: Inverse Rendering with Physics-Based Neural Fields},\nauthor={Sean Wu and Shamik Basu and Tim Broedermann and Luc Van Gool and Christos Sakaridis},\nbooktitle={Greeks in AI Symposium 2025},\nyear={2025},\nurl={https://openreview.net/forum?id=Eb6sSukTaW}\n}"},"title":{"value":"PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields"},"authors":{"value":["Sean Wu","Shamik Basu","Tim Broedermann","Luc Van Gool","Christos Sakaridis"]}},"tmdate":1750690429166,"pdate":1750690429137,"tcdate":1748340051363,"writers":["greeksin.ai/Greeks_in_AI/2025/Symposium","greeksin.ai/Greeks_in_AI/2025/Symposium/Submission68/Authors"],"signatures":["greeksin.ai/Greeks_in_AI/2025/Symposium/Submission68/Authors"],"forum":"Eb6sSukTaW","license":"CC BY 4.0","number":68,"cdate":1748340051363,"readers":["everyone"],"invitations":["greeksin.ai/Greeks_in_AI/2025/Symposium/-/Submission","greeksin.ai/Greeks_in_AI/2025/Symposium/Submission68/-/Full_Submission","greeksin.ai/Greeks_in_AI/2025/Symposium/-/Post_Submission","greeksin.ai/Greeks_in_AI/2025/Symposium/-/Edit"],"mdate":1750690429166,"odate":1750690429137,"domain":"greeksin.ai/Greeks_in_AI/2025/Symposium","id":"Eb6sSukTaW","version":2},{"content":{"summary":{"value":"This paper presents LOCA (Logical Chain Augmentation), a framework for automatically cleaning scientific corpora by addressing logical incompleteness in reasoning chains. The method employs an augment-and-review loop that (1) completes missing logical steps through chain completion, (2) decomposes each step into orthogonal principle and derivation components, and (3) uses specialized review agents to iteratively refine solutions. Applied to physics QA datasets (PHYBench, PHYSICS, ABench-Physics), LOCA reportedly reduces error rates from ~20% to below 2% while maintaining substantial dataset size."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What are computational costs? Report: (a) average LLM calls per question, (b) total tokens processed, (c) wall-clock time, (d) cost per 1000 questions, (e) comparison with baselines.\n\n2. What are hyperparameter sensitivities? Show how error rate and accepted set size vary across N_{corr}^{(max}} ∈ {1,2,3,4,5} and N_{wrg}^{(max)} ∈ {1,3,5,7,10}.\n\n3. How does the system determine when a step is \"non-atomic\"? How are principles identified from axiom space P? How is semantic equivalence determined for external consistency checks?\n\n4. Can you confirm template compliance? The bottom margins appear non-standard. Verify the manuscript uses the official ICLR 2026 LaTeX template without modifications."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses a genuine problem—high error rates in scientific reasoning datasets—with a structured approach. The principle-derivation decomposition provides interpretable intermediate representations that could facilitate both automated review and human verification.\n\n2. The evaluation compares against diverse baselines spanning reasoning methods (CoT, ToT, GoT), review methods (Review-SC), and iterative refinement (Self-Reflection) across multiple LLMs. The detailed example in Appendix A.2 illustrates how the method processes a physics problem through multiple iterations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Ground truth is created by experts using LOCA's structured outputs to identify errors, then LOCA is evaluated against these same errors. This makes the <2% error rate claim unverifiable.\n\n2. Only 300 questions evaluated across three datasets, far too limited for a method claiming to enable \"large-scale, high-quality scientific corpora.\" No computational cost analysis provided despite requiring up to 8 LLM calls per question.\n\n3. Please provide Cohen's kappa across multiple independent experts for error identification to validate ground truth reliability.\n\n4. Bottom margins appear excessively wide throughout, resulting in significantly less content per page than standard ICLR submissions. Authors must verify compliance with official ICLR 2026 template."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925122381,"tcdate":1761856168985,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14769/Reviewer_ro15"],"signatures":["ICLR.cc/2026/Conference/Submission14769/Reviewer_ro15"],"forum":"kdFjucrq7B","number":2,"license":"CC BY 4.0","cdate":1761856168985,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14769/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925122381,"domain":"ICLR.cc/2026/Conference","replyto":"kdFjucrq7B","id":"aR190QpKo2","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["scientific corpus cleaning","logical chain","AI for science","LLMs"]},"supplementary_material":{"value":"/attachment/9bd9e56c7096711b6b2d9ada26ff81d426ddc305.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"While Large Language Models (LLMs) excel in general domains, their reliability often falls short in scientific problem-solving. The advancement of scientific AI depends on large-scale, high-quality corpora. However, existing scientific question-answering (QA) datasets suffer from high error rates, frequently resulting from logical leaps and implicit reasoning within the answers. To address this issue, we introduce LOCA (Logical Chain Augmentation), a novel framework for automatically cleaning scientific corpora, implemented through an augment-and-review loop. At its core, LOCA enhances raw answers by completing missing logical steps and explicitly separating the underlying scientific principle from its subsequent derivation. By applying LOCA to challenging scientific corpora, we demonstrate that it can automatically filter noisy datasets, typically reducing the error rate from as high as 20\\% to below 2\\%. LOCA provides a scalable and effective methodology for creating high-quality scientific corpora, paving the way for more reliable training and evaluation of scientific AI."},"_bibtex":{"value":"@misc{\nfang2026loca,\ntitle={{LOCA}: Logical Chain Augmentation for Scientific Corpus Cleaning},\nauthor={Youle Fang and Dong-Shan Jian and Xiang Li and Ce Meng and Lingshi Meng and Chen-Xu Yan and Zhizhang Bian and Yan-Qing Ma},\nyear={2026},\nurl={https://openreview.net/forum?id=kdFjucrq7B}\n}"},"title":{"value":"LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning"},"pdf":{"value":"/pdf/6adf0ad32849f9012988103f077cd7de25b8d589.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"fang|loca_logical_chain_augmentation_for_scientific_corpus_cleaning"},"authorids":{"value":["~Youle_Fang1","~Dong-Shan_Jian1","~Xiang_Li161","~Ce_Meng1","~Lingshi_Meng3","~Chen-Xu_Yan1","~Zhizhang_Bian1","~Yan-Qing_Ma1"]},"authors":{"value":["Youle Fang","Dong-Shan Jian","Xiang Li","Ce Meng","Lingshi Meng","Chen-Xu Yan","Zhizhang Bian","Yan-Qing Ma"]}},"version":2},{"content":{"summary":{"value":"This paper introduces FlowSymm, a solver that augments a graph attention backbone with symmetry-preserving, physics-aware corrections for partially observed flow graphs. This method for estimating missing flows in networks (e.g., traffic, power, bike-sharing) where only a subset of edges have sensors.\n\nIts main contributions are:\n1、It reinterprets admissible divergence-free adjustments as elements of an Abelian group action. \n2、It combines this basis with edge-wise GATv2 embeddings. An attention mechanism scores each basis vector in a context-aware manner, enabling the model to inject corrective flows where the local structure demands, while preserving flow balance.\n3、It imposes soft physical consistency through a feature-conditioned Tikhonov refinement and trains all weights end-to-end via reverse-mode implicit differentiation within a bilevel optimization framework.\n4、Extensive experiments on three real-world benchmarks show that FlowSymm consistently improves upon nine baselines, reducing RMSE by up to ten percent and yielding performance gains over the previous state-of-the-art."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Questions：The entire method, starting with the \"balanced anchor\"  , assumes the net nodal injection vector c is known. In many real-world scenarios (e.g., unmonitored power grid feeders, unobserved traffic sources/sinks), c is also partially unknown or highly uncertain. How sensitive is FlowSymm's performance to errors in c?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"Nothing"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This is a very well-written and clearly presented paper. The core idea—leveraging a group-action basis for divergence-free corrections in network flow completion. The methodology is explained with remarkable clarity, and the experimental results are thorough and convincing. The core innovation is the reformulation of the flow completion problem through the lens of group theory. While GNNs typically encode spatial symmetries or permutation symmetries, this work defines a new, problem-specific \"algebraic symmetry\" derived directly from the graph's incidence matrix and the sensor mask. \n\nIn addition, the mathematical derivation is sound, from the initial balanced anchor using the Moore-Penrose pseudoinverse to the construction of the projector P_A and the implicit differentiation for the bilevel problem. The method is built on a firm algebraic foundation.\n\nFinally, the experimental design is robust. The use of three distinct, real-world domains (Traffic, Power, Bike) demonstrates generalizability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1、The introduction does an excellent job of covering related work in physics-informed ML and equivariant GNNs. However, the transition to the specific \"algebraic symmetry\" of flow conservation could be slightly more explicit for a reader unfamiliar with this background.\n2、The derivation of the projector P_A is technically correct but may be dense for some readers. The step of subtracting the projector onto the observed edges within the balanced space is crucial but could be briefly motivated with an intuitive phrase."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924958560,"tcdate":1761998291714,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14572/Reviewer_wnzU"],"signatures":["ICLR.cc/2026/Conference/Submission14572/Reviewer_wnzU"],"forum":"vu1IEpdUQh","number":4,"license":"CC BY 4.0","cdate":1761998291714,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14572/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924958560,"domain":"ICLR.cc/2026/Conference","replyto":"vu1IEpdUQh","id":"vofqEahAJu","forumContent":{"TLDR":{"value":"FlowSymm is an end-to-end graph neural model that completes missing flows by enforcing a divergence-free group-action prior, scoring corrections with attention, and refining with feature-conditioned Tikhonov regularization"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["graphs","networks","flow graphs","graph attention networks","group action","bilevel-optimization","physics-aware graph neural networks"]},"supplementary_material":{"value":"/attachment/64543588602e2348957ff8821e15ddc0f5206c70.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Recovering missing flows on the edges of a network, while exactly respecting local conservation laws, is a fundamental inverse problem that arises in many systems such as transportation, energy, and mobility. We introduce FlowSymm, a novel architecture that combines (i) a group-action on divergence-free flows, (ii) a graph-attention encoder to learn feature-conditioned weights over these symmetry-preserving actions, and (iii) a lightweight Tikhonov refinement solved via implicit bilevel optimization. The method first anchors the given observation on a minimum-norm divergence-free completion. We then compute an orthonormal basis for all admissible group actions that leave the observed flows invariant and parameterize the valid solution subspace, which shows an Abelian group structure under vector addition. A stack of GATv2 layers then encodes the graph and its edge features into per-edge embeddings, which are pooled over the missing edges and produce per-basis attention weights. This attention-guided process selects a set of physics-aware group actions that preserve the observed flows. Finally, a scalar Tikhonov penalty refines the missing entries via a convex least-squares solver, with gradients propagated implicitly through Cholesky factorization. Across three real-world flow benchmarks (traffic, power, bike), FlowSymm substantially outperforms state-of-the-art baselines in RMSE, MAE and correlation metrics."},"_bibtex":{"value":"@inproceedings{\ndemirci2026flowsymm,\ntitle={FlowSymm: Physics{\\textendash}Aware, Symmetry{\\textendash}Preserving Graph Attention for Network Flow Completion},\nauthor={Ege Demirci and Francesco Bullo and Ananthram Swami and Ambuj Singh},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=vu1IEpdUQh}\n}"},"title":{"value":"FlowSymm: Physics–Aware, Symmetry–Preserving Graph Attention for Network Flow Completion"},"pdf":{"value":"/pdf/5f45aa1b88ec06fd6198f68eb0b3063fdd10f8ce.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"demirci|flowsymm_physicsaware_symmetrypreserving_graph_attention_for_network_flow_completion"},"authorids":{"value":["~Ege_Demirci1","~Francesco_Bullo1","~Ananthram_Swami1","~Ambuj_Singh1"]},"authors":{"value":["Ege Demirci","Francesco Bullo","Ananthram Swami","Ambuj Singh"]}},"version":2},{"content":{"summary":{"value":"### Summary\n1. A Mamba architecture based retriever is presented which is trained as a discriminator model\n2. The retrievers are much smaller (130M and 1.3B) and outperform much larger open source (and non-finetuned) embedding based models\n3. Efficacy of all retrievers is tested by plugging the retrieved chunks (or sentences) into a RAG pipeline\n\n\n### Contributions\n1. The paper showcases how a Mamba architecture model could be finetuned to create an effective retriever for long context documents\n2. Methodology and ablation for creating synthetic data is well studied and supported via ablation\n3. Synthetic dataset would further research for long context document understanding (and retrieval)"},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. (Lines 233 to 239) Why is the retriever setup different for Mamba and embedding based retrievers? Mamba fetches Top 50 sentences whereas embedding based retrievers fetch 5 chunks. How do embedding models perform when you have identical setup?\n2. Why was the link based synthetic data not used to finetune an embedding model and compared as baseline? Perhaps the gains seen by Mamba retriever are due to synthetic data. A quick check would be to see how embedding retrievers perform (in terms of recall) on a held out set of the synthetic link based data. Maybe this number could be improved via finetuning and the finetuned embedding model tested as baseline\n3. (Lines 198-200) What decontamination pipeline was used to ensure that there is no overlap between synthetic data documents and documents used in the 41 benchmarks. Did you look for exact match of document content? Are the domains completely different?\n4. (Lines 230-232) How did you evaluate the quality of an answer using GPT-4o? Did you give some instructions in prompt on how to judge the quality of an answer? Did the model return a binary score (such as relevant or non-relevant)\n5. How was BM25 setup? Did it also retriever chunks? What happens when it retrieves sentences (same as Mamba)\n6. (Lines 204-207) Which hyper-parameters were optimized? How did you select the best checkpoint? How did you decide when to stop training?\n7. (Minor comment): Figures 6-10 in the Appendix seem to have the same exact prompt. Was this intentional or done by mistake?\n8. (Minor comment): Lines 89-90 say upto 256 k. Did the authors mean more than 256k?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This work presents a lightweight (small parameter) based retriever for long context documents\n2. Complete training code and dataset is available on Github for reproducibility\n3. New dataset would further research in the direction\n4. Ablations presented help understand the impact of synthetic data generation process (link based) as well as effect of context size on RAG system performance"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The baseline retriever (open source embedding model) is weak as it has not seen any synthetic data. It is unclear whether gains seen by Mamba retriever is due to synthetic data or the model's ability to handle long context\n2. The paper claims that Mamba is more efficient in handling long document context. No supportive results are given in terms of time taken. It's unclear how efficient the model is when compared to a retriever\n3. Embedding based retriever has a different retrieval granularity (chunks) as compared to Mamba (Top 50 sentences)\n4. The paper claims there is lack of synthetic data for long context documents, but ends up creating a dataset where maximum context length is 10k. Is my understanding correct?\n5. (Minor issue) No qualitative examples are presented in the paper which highlight how Mamba retrieves more relevant context as compared to Embedding based models. Perhaps some examples can be added to the Appendix, if there"}},"nonreaders":[],"tmdate":1733223875640,"tcdate":1729573886370,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13829/Reviewer_PSm7"],"signatures":["ICLR.cc/2025/Conference/Submission13829/Reviewer_PSm7"],"forum":"NJUzUq2OIi","number":2,"license":"CC BY 4.0","cdate":1729573886370,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13829/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733223875640,"domain":"ICLR.cc/2025/Conference","replyto":"NJUzUq2OIi","id":"vILBWMtRQN","forumContent":{"TLDR":{"value":"We propose an efficient retrieval model using Mamba architecture for full-context understanding of long documents."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Long-context retrieval","Efficient language models","Mamba architecture","Synthetic data generation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Long document question answering is challenging due to the quadratic computational cost of transformer-based LLMs. In resource-constrained environments, Retrieval-Augmented Generation (RAG) uses document chunking to maintain linear computational costs but often loses sight of the global context. We introduce the Mamba retriever 130M and Mamba retriever 1.3B retrievers, capable of processing entire long documents in linear time and integrating earlier context to retrieve relevant sentences for answering questions. Mamba retrievers outperform state-of-the-art embedding models across 41 long-document Q\\&A benchmarks while maintaining speed and computational efficiency. Their performance is comparable to GPT-4o on long documents over 256k tokens while using significantly fewer tokens. Mamba retrievers are trained on synthetic data generated from our novel link-based method, which enhances the retrievers' ability to leverage long-range document connections. We further demonstrate the effectiveness of our link-based method over baseline synthetic data methods. All code, datasets, and model checkpoints are available at https://github.com/MambaRetriever/MambaRetriever"},"_bibtex":{"value":"@misc{\ncao2025efficient,\ntitle={Efficient Full-Context Retrieval for Long Documents},\nauthor={Weili Cao and Jianyou Wang and Youze Zheng and Longtian Bao and Qirui Zheng and Taylor Berg-Kirkpatrick and Ramamohan Paturi and Leon Bergen},\nyear={2025},\nurl={https://openreview.net/forum?id=NJUzUq2OIi}\n}"},"title":{"value":"Efficient Full-Context Retrieval for Long Documents"},"pdf":{"value":"/pdf/0b8010402c2d211d3ab574c24916f6283ffd4b0a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"cao|efficient_fullcontext_retrieval_for_long_documents"},"authorids":{"value":["~Weili_Cao1","~Jianyou_Wang1","~Youze_Zheng1","~Longtian_Bao1","~Qirui_Zheng1","~Taylor_Berg-Kirkpatrick1","~Ramamohan_Paturi1","~Leon_Bergen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Weili Cao","Jianyou Wang","Youze Zheng","Longtian Bao","Qirui Zheng","Taylor Berg-Kirkpatrick","Ramamohan Paturi","Leon Bergen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Chain-of-Time, a method for improving physical simulation in Image Generation Models by generating intermediate frames step-by-step, inspired by mental simulation in humans and chain-of-thought reasoning in LLMs. The authors test on four domains (2D Motion, 2D Gravity, Fluids, Bouncing) and show improvements over direct prediction in most cases."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What fundamental properties of fluid simulation make Chain-of-Time perform worse than direct prediction?\n\n2. Can you provide results for simulations beyond 0.8 seconds to assess error accumulation?\n\n3. Have you tested on any established video prediction or physical reasoning benchmarks?\n\n4. Could you test additional VLM+IGM combinations to verify generalizability?\n\n5. Why do sample sizes differ across domains?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Moderately Novel approach: The connection between human mental simulation theory and in-context reasoning for IGMs is creative and well-motivated.\n\nInterpretability: The method provides interpretable intermediate steps that reveal the model's physical reasoning process.\n\nClear methodology: The paper clearly describes the de-rendering, simulation, and rendering components."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1: Insufficient analysis of fluid domain's fundamental challenges The authors observe performance degradation in the fluids domain but fail to analyze why fluids are fundamentally different from the other tested scenarios. The paper should discuss whether Chain-of-Time is inherently unsuitable for fluid domain due to some specific properties of it like continuous deformation or partial transparency, or whether the issue is specific to their experimental setup. The superficial explanation of \"flow rate estimation error\" doesn't address why the step-by-step approach that helps with projectile motion actually harms fluid simulation.\n\nW2: Limited temporal horizon evaluation All experiments are constrained to 0.8 seconds of simulation. For a method claiming to improve physical simulation, testing longer time horizons (e.g., 2-5 seconds) would better demonstrate scalability and compound error effects. The mental simulation literature the authors cite often involves longer-term predictions.\n\nW3: Insufficient experimental scope\n\nModel diversity: Only GPT-4o is tested. While the authors mention DALL-E 3's limitations, they should test other recent VLM+IGM combinations (e.g., Gemini Pro Vision, LLaVA variants with diffusion models, or Claude with image generation capabilities).\nDomain limitations: The four domains use overly simplistic objects and backgrounds (solid white, uniform blue). The authors should evaluate on some of the established video prediction datasets (Moving MNIST, KTH Actions, BAIR Robot Pushing) or physical reasoning benchmarks (IntPhys, CATER, Physion) to test generalization beyond synthetic scenes, whichever the authors think is suitable for validating their method.\n\nW4: Incomplete related work The related work section misses important connections:\n\nVideo prediction literature (e.g., stochastic video generation, physics-informed neural networks)\nWorld models that perform similar step-by-step physical prediction\n\nW5: Sample sizes vary across domains (N=5 to N=20) without justification."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942554059,"tcdate":1761867195123,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23193/Reviewer_7Pnk"],"signatures":["ICLR.cc/2026/Conference/Submission23193/Reviewer_7Pnk"],"forum":"f6ugrBWs3K","number":1,"license":"CC BY 4.0","cdate":1761867195123,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23193/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942554059,"domain":"ICLR.cc/2026/Conference","replyto":"f6ugrBWs3K","id":"0CVKYMciB5","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-modal Language Models","Spatial and Temporal Perception","Image Generation","Physical Reasoning"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"We propose a novel method to improve the physical simulation ability of vision-language models. This Chain-of-Time simulation is motivated by in-context reasoning in machine learning, and mental simulation in humans. The method involves generating a series of intermediate images during a simulation. Chain of Time is used at inference time and requires no additional fine-tuning for performance benefits. We apply the Chain-of-Time method to synthetic and real-world domains, including 2-D graphics simulations and natural 3-D videos. These domains test a variety of particular physical properties, including velocity, acceleration, fluid dynamics, and conservation of momentum. We found that using Chain-of-Time simulation substantially improves the performance of state-of-the-art Image Generation Model. Beyond examining performance, we also analyze the specific states of the world simulated by an image model at each time step, which sheds light on the dynamics underlying these simulations. This analysis reveals insights that are hidden from traditional evaluations of physical reasoning, including cases where an Image Generation Model is able to simulate physical properties that unfold over time, such as velocity, gravity, and collisions domain well. Our analysis also highlights particular cases where the Image Generation Model struggles to infer particular physical parameters from input images, despite being capable of simulating relevant physical processes."},"_bibtex":{"value":"@misc{\nwang2026chain,\ntitle={Chain of Time: In-Context Physical Simulation with Image Generation Models},\nauthor={YingQiao Wang and Eric Bigelow and Boyi Li and Tomer Ullman},\nyear={2026},\nurl={https://openreview.net/forum?id=f6ugrBWs3K}\n}"},"title":{"value":"Chain of Time: In-Context Physical Simulation with Image Generation Models"},"pdf":{"value":"/pdf/5d0ef4077bf57397482b6d12a29a564f7f8e8459.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|chain_of_time_incontext_physical_simulation_with_image_generation_models"},"authorids":{"value":["~YingQiao_Wang1","~Eric_Bigelow1","~Boyi_Li1","~Tomer_Ullman1"]},"authors":{"value":["YingQiao Wang","Eric Bigelow","Boyi Li","Tomer Ullman"]}},"version":2},{"content":{"venue":{"value":"Research in the Mathematical Sciences"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s40687-022-00329-z.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"xu|learning_generative_neural_networks_with_physics_knowledge"},"html":{"value":"https://doi.org/10.1007/s40687-022-00329-z"},"abstract":{"value":"Deep generative neural networks have enabled modeling complex distributions, but incorporating physics knowledge into the neural networks is still challenging and is at the core of current physics-based machine learning research. To this end, we propose a physics generative neural network (PhysGNN), a new class of generative neural networks for learning unknown distributions in a physical system described by partial differential equations (PDE). PhysGNN couples PDE systems with generative neural networks. It is a fully differentiable model that allows back-propagation of gradients through both numerical PDE solvers and generative neural networks, and is trained by minimizing the discrete Wasserstein distance between generated and observed probability distributions of the PDE outputs using the stochastic gradient descent method. Moreover, PhysGNN does not require adversarial training like standard generative neural networks, which offers better stability than adversarial training. We show that PhysGNN can learn complex distributions in stochastic inverse problems, where conventional methods such as maximum likelihood estimation and momentum matching methods may be inapplicable when little knowledge is known about the form of unknown distributions or the physical model is too complex. Our method allows physics-based generative neural network training for learning complex distributions in the context of differential equations."},"title":{"value":"Learning generative neural networks with physics knowledge"},"authors":{"value":[{"fullname":"Kailai Xu"},{"fullname":"Weiqiang Zhu"},{"fullname":"Eric Darve","username":"~Eric_Darve1"}]}},"tmdate":1789091012028,"pdate":1654041600000,"externalIds":["doi:10.1007/s40687-022-00329-z"],"tcdate":1767720574759,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Eric_Darve1"],"forum":"ciZaI0H5dx","license":"CC BY-SA 4.0","number":23083,"cdate":1652812178620,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789091012028,"domain":"OpenReview.net/Public_Article","id":"ciZaI0H5dx","version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS 2017"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-72150-7_37.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2017"},"paperhash":{"value":"ning|detection_of_fournode_motif_in_complex_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zhaolong_Ning:","https://dblp.org/search/pid/api?q=author:Lei_Liu:","~Shuo_Yu1","~Feng_Xia1"]},"html":{"value":"https://doi.org/10.1007/978-3-319-72150-7_37"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/NingLYX17,\n  author={Zhaolong Ning and Lei Liu and Shuo Yu and Feng Xia},\n  title={Detection of Four-Node Motif in Complex Networks},\n  year={2017},\n  cdate={1483228800000},\n  pages={453-462},\n  url={https://doi.org/10.1007/978-3-319-72150-7_37},\n  booktitle={COMPLEX NETWORKS},\n  crossref={conf/complexnetworks/2017}\n}\n"},"abstract":{"value":"Complex network analysis has gained research interests in a wide range of fields. Network motif, which is one of the most popular network properties, is a statistically significant network subgraph. In this paper, we propose a fast methodology, called Four-node Motif Detection Algorithm (FMDA), to extract four-node motifs in complex networks. Specifically, we employ a two-way spectral clustering method to cut big networks into small sub-graphs, and then identify motifs by recognition algorithm to reduce the computational complexity. After that, we use three isomorphic four-node motifs to analyze network structure by American Physical Society (APS) data set."},"title":{"value":"Detection of Four-Node Motif in Complex Networks"},"authors":{"value":["Zhaolong Ning","Lei Liu","Shuo Yu","Feng Xia"]}},"tmdate":1769133820989,"pdate":1483228800000,"externalIds":["dblp:conf/complexnetworks/NingLYX17"],"tcdate":1768477156348,"writers":["~"],"signatures":["~Feng_Xia1"],"forum":"lMPr7axk4Y","license":"CC BY-SA 4.0","number":762371,"cdate":1483228800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769133820989,"domain":"DBLP.org","id":"lMPr7axk4Y","version":2},{"content":{"summary":{"value":"Decentralized FL enables clients to own different local models and separately optimize local data. How can every client's local model learn generalizable representation is unknown. To address this question, This paper proposes a Decentralized FL technique by introducing Synthetic Anchors, as DESA. Authors leverage the synthetic anchors to implement 1) anchor loss that matches the distribution of the client's latent embedding with an anchor and 2) KD loss that enables clients learning from others. The proposed method doesn't presume access to real public or a global data generator."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The studied problem is novel and well motivated.\n2. Distilling local synthetic anchor is interesting.\n3. There are theoretical analysis of the proposed methods, in which the new generalization bound is better.\n4. Figure 3 is interesting, jointly considering worst local accuracy and global accuracy. Experiment results show significant improvements of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The local synthetic anchor dataset iss shared. Thus, the privacy of the synthesized anchor should be considerred. Although the DP is used to protect synthetic anchor. But could this defend against recovering the raw data?\n2. It would be better to conduct a more ablation study to decouple the effect of the sythetic anchor and the KD loss."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"See weaknesses."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635981435,"tcdate":1698847465472,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission545/Reviewer_2BZe"],"signatures":["ICLR.cc/2024/Conference/Submission545/Reviewer_2BZe"],"forum":"PcBJ4pA6bF","number":4,"license":"CC BY 4.0","cdate":1698847465472,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission545/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635981435,"domain":"ICLR.cc/2024/Conference","replyto":"PcBJ4pA6bF","id":"nAD3hcHU1C","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Federated Learning","Data Heterogeneity","Model Heterogeneity"]},"supplementary_material":{"value":"/attachment/9f59eb9a0983d84551769b886e45eac851ffafb7.zip"},"primary_area":{"value":"general machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Conventional Federated Learning (FL) involves collaborative training of a global model by multiple client local models. In this emerging paradigm, the central server assumes a critical role in aggregating local models and maintaining the global model. However, it encounters various challenges, including scalability, management, and inefficiencies arising from idle client devices. \nRecently, studies on serverless decentralized FL have shown advantages in overcoming these challenges, enabling clients to own different local models and separately optimize local data. Despite the promising advancements in decentralized FL, it is crucial to thoroughly investigate the implications of data and model heterogeneity, which pose unique challenges that must be overcome. Therefore, the research question to be answered in this study is: How can every client's local model learn generalizable representation?\nTo address this question, we propose a novel Decentralized FL technique by introducing Synthetic Anchors, dubbed as DeSA. Inspired by the theory of domain adaptation and Knowledge distillation (KD), we leverage the synthetic anchors to design two effective regularization terms for local training: 1) anchor loss that matches the distribution of the client's latent embedding with an anchor and 2) KD loss that enables clients learning from others. \nIn contrast to previous KD-based heterogeneous FL methods, we don’t presume access to real public or a global data generator. \nDeSA enables each client's model to become robust to distribution shift across different client-domains. Through extensive experiments on diverse client data distributions, we showcase the effectiveness of \\ours{} in enhancing both inter and intra-domain accuracy of each client."},"_bibtex":{"value":"@misc{\nhuang2024overcoming,\ntitle={Overcoming Data and Model heterogeneities in Decentralized Federated Learning via Synthetic Anchors},\nauthor={Chun-Yin Huang and Kartik Srinivas and Xin Zhang and Xiaoxiao Li},\nyear={2024},\nurl={https://openreview.net/forum?id=PcBJ4pA6bF}\n}"},"title":{"value":"Overcoming Data and Model heterogeneities in Decentralized Federated Learning via Synthetic Anchors"},"pdf":{"value":"/pdf/aa9370a559289b8e6c23498cbecfb3d977d6d6e3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"huang|overcoming_data_and_model_heterogeneities_in_decentralized_federated_learning_via_synthetic_anchors"},"authorids":{"value":["~Chun-Yin_Huang1","~Kartik_Srinivas1","~Xin_Zhang16","~Xiaoxiao_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chun-Yin Huang","Kartik Srinivas","Xin Zhang","Xiaoxiao Li"]}},"version":2},{"content":{"venue":{"value":"ICAT-EGVE (Posters & Demos) 2023"},"pdf":{"value":"http://localhost:4000/bitstreams/9ce09c67-cf06-43e1-aad0-e68743f4f83c/download"},"venueid":{"value":"dblp.org/conf/ICAT/2023"},"paperhash":{"value":"li|tangiblemrcreate_intuitive_authoring_of_mixed_reality_content"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xuyu_Li:","~John_Dingliana1"]},"html":{"value":"https://doi.org/10.2312/egve.20231339"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icat/LiD23,\n  author={Xuyu Li and John Dingliana},\n  title={TangibleMRCreate: Intuitive Authoring of Mixed Reality Content},\n  year={2023},\n  cdate={1672531200000},\n  pages={25-26},\n  url={https://doi.org/10.2312/egve.20231339},\n  booktitle={ICAT-EGVE (Posters & Demos)},\n  crossref={conf/egve/2023p}\n}\n"},"abstract":{"value":"In this poster, we present work-in-progress exploring the intuitive authoring of Mixed Reality (MR) content using physical object manipulation, tangible 3D user interfaces, and the creation of MR within MR itself. To assess this approach, we implemented an MR-based system enabling novice users to tangibly create MR scenes within MR. The system, prototyped on the Microsoft HoloLens 2, allows users to create digital scenes by manipulating a physical box and a tangible 3D user interface. A pilot study was conducted to provide a preliminary evaluation and to inform future system development and deeper study design."},"title":{"value":"TangibleMRCreate: Intuitive Authoring of Mixed Reality Content"},"authors":{"value":["Xuyu Li","John Dingliana"]}},"tmdate":1747846785121,"pdate":1672531200000,"tcdate":1747846747596,"writers":["~"],"signatures":["~John_Dingliana1"],"forum":"zxw4UzrbhG","license":"CC BY-SA 4.0","number":547683,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747846785121,"domain":"DBLP.org","id":"zxw4UzrbhG","version":2},{"content":{"summary":{"value":"The paper proposes PhyCAGE, a novel framework for generating physically constrained compositional 3D assets from a single 2D image.\nThe system integrates multi-view image generation, 3D Gaussian Splatting (3DGS) reconstruction, and a Physical Simulation-Enhanced Score Distillation Sampling (PSE-SDS) process. Unlike prior SDS-based or compositional 3D methods, PhyCAGE introduces a differentiable physical simulation (via MPM) that uses the SDS loss gradient as the initial velocity, enabling physics-guided correction of inter-object penetrations. Experiments on ComboVerse and other benchmarks demonstrate that PhyCAGE improves physical plausibility, visual quality, and component disentanglement."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How sensitive is PSE-SDS to the accuracy of Grounded-SAM segmentation or the inpainting model?\n2. What physical parameters (mass, stiffness) are assumed in the MPM simulation? Could they be learned automatically?\n3. How would the method behave in dynamic or non-rigid interactions (e.g., cloth over soft body)?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Novel integration of MPM-based physics into SDS optimization.\n2. Physically plausible compositional reconstruction, effectively addressing object penetration and stability.\n3. Quantitative improvements in PSNR, CLIP score, and penetration rate compared to strong baselines (Part123, ComboVerse).\n4. Extensive ablation and runtime analysis, demonstrating consistent improvements."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited novelty in theoretical contribution.\nThe PSE-SDS approach combines existing elements (SDS, MPM) rather than introducing a fundamentally new learning paradigm. The innovation is mainly in integration and practical application.\n\n2. Heuristic parameter dependence.\nThe method uses several empirically chosen constants (e.g., λ₁–λ₃, 500 simulation steps, timestep decay) without sensitivity analysis. Stability under different simulation conditions is not examined, which may limit reproducibility.\n\n3. Scalability and computational cost.\nThe runtime table shows PhyCAGE is slower than Part123, approaching ComboVerse’s cost (≈734s). Given its “single-image” input goal, the computational burden may hinder deployment.\n\n4. Assumption of accurate segmentation and inpainting.\nThe method relies heavily on Grounded-SAM and SyncDreamer outputs. Errors in segmentation or multi-view synthesis propagate to the physics simulation, potentially breaking physical consistency.\n\n5. Scope limitations.\nThe current pipeline handles only two-object interactions and assumes static scenes. Extending to multi-object or dynamic scenes would require further methodological development."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925345315,"tcdate":1761315122303,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15014/Reviewer_P5me"],"signatures":["ICLR.cc/2026/Conference/Submission15014/Reviewer_P5me"],"forum":"D8ypZxzIvO","number":1,"license":"CC BY 4.0","cdate":1761315122303,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15014/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925345315,"domain":"ICLR.cc/2026/Conference","replyto":"D8ypZxzIvO","id":"H8CKoCk6fh","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"PhyCAGE"},"keywords":{"value":["3D Generation","Image-to-3D","Physical Simulation"]},"supplementary_material":{"value":"/attachment/a0d62576046faeeec6dcf663efba55da998587c2.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We present PhyCAGE, the first approach for physically constrained compositional 3D asset generation from a single image. Given an input image, we first generate consistent multi-view images for components of the assets. These images are then fitted with 3D Gaussian Splatting representations. To ensure that the Gaussians representing objects are physically compatible with each other, we introduce a Physical Simulation-Enhanced Score Distillation Sampling (PSE-SDS) technique to further optimize the positions of the Gaussians. It is achieved by setting the gradient of the SDS loss as the initial velocity of the physical simulation, allowing the simulator to act as a physics-guided optimizer that progressively corrects the Gaussians' positions to a physically compatible state. Experimental results demonstrate that the proposed method can generate physically plausible compositional 3D assets given a single image."},"_bibtex":{"value":"@misc{\nyan2026phycage,\ntitle={Phy{CAGE}: Physically Constrained Compositional 3D Asset Generation from a Single Image},\nauthor={Han Yan and Mingrui Zhang and YANG LI and Chao Ma and Pan Ji},\nyear={2026},\nurl={https://openreview.net/forum?id=D8ypZxzIvO}\n}"},"title":{"value":"PhyCAGE: Physically Constrained Compositional 3D Asset Generation from a Single Image"},"pdf":{"value":"/pdf/5a08996c8b9147d733cf65ff8c9fcd37e8c4504c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yan|phycage_physically_constrained_compositional_3d_asset_generation_from_a_single_image"},"authorids":{"value":["~Han_Yan1","~Mingrui_Zhang4","~YANG_LI49","~Chao_Ma3","~Pan_Ji3"]},"authors":{"value":["Han Yan","Mingrui Zhang","YANG LI","Chao Ma","Pan Ji"]}},"version":2},{"content":{"summary":{"value":"Metal implants create severe artifacts that disrupt standard accelerated MRI acquisition and reconstruction pipelines. MASC introduces a unified framework that uses reinforcement learning to jointly optimize k-space sampling patterns and a downstream metal artifact reduction network. To enable supervised learning for this task, a physics-based simulation pipeline generates paired k-space data from CT volumes, providing ground-truth \"clean\" references for metal-corrupted samples. Experiments demonstrate that this co-adaptive approach allows the sampling policy to target k-space lines that maximize the artifact reduction network's performance, outperforming conventional sampling strategies."},"justification_of_final_rating":{"value":"The rebuttal effectively addressed key concerns regarding generalization and fair comparison. The added FastMRI cross-dataset evaluation and DQN baseline comparisons substantially strengthen validation, while clarifications on architecture and losses improve clarity. Despite persisting single-coil limitations and lack of physical phantom validation, the joint optimization framework demonstrates clear methodological value for metal-aware MRI acquisition."},"confidence":{"value":3},"final_rating":{"value":4},"justification_of_the_preliminary_rating":{"value":"MASC proposes a novel and methodologically sound approach to a complex problem that is often treated in isolation. Creating a physics-based paired dataset to enable reinforcement learning in this domain is a valuable contribution. However, the reliance on single-coil simulations and the lack of real-world validation data limit the immediate clinical applicability and prevent a higher score. Despite these limitations, the joint optimization strategy is promising and relevant to the MIDL community."},"confidentiality_llm_acknowledgment":{"value":"Yes"},"strengths":{"value":"- Tackling the intersection of active acquisition and metal artifact reduction addresses a specific but highly clinically relevant gap in the current literature.\n- Jointly optimizing the PPO agent with the reconstruction network allows the sampling policy to learn non-intuitive patterns specifically designed to mitigate metal-induced spectral corruption.\n- Constructing a paired dataset via rigorous physics-based simulation on AutoPET data offers a clever and necessary solution to the impossibility of obtaining real-world ground truth for metal artifacts.\n- Ablation studies clearly isolate the benefits of the end-to-end training strategy, proving that the policy effectively adapts to the specific capabilities of the MAR network.\n- Visualizations of the acquisition trajectory show the agent learning to prioritize high-information frequency bands, aligning well with MR physics intuition."},"weaknesses":{"value":"- Validating exclusively on simulated data creates a significant domain gap, as real-world metal artifacts involve complex B0 inhomogeneities and pile-up effects that may not be fully captured by the simulator.\n- Limiting the study to single-coil simulations reduces immediate clinical impact, given that parallel imaging is the standard for accelerated MRI in modern scanners.\n- Focusing solely on hip implants restricts the evaluation of the policy's robustness, leaving it unclear how the method handles more complex geometries like spinal fixation or dental hardware.\n- Comparing primarily against static sampling baselines (Random, Center-out) overlooks stronger recent deep learning-based reconstruction baselines (e.g., VarNet or unrolled networks) that might handle undersampling artifacts better even without active selection."},"detailed_comments":{"value":"- Discussing the inference latency of the PPO agent is crucial, as active acquisition requires real-time decision-making during the TR intervals of the scanner.\n- Expanding the discussion on how the specific simulation parameters (e.g., readout bandwidth) influence the difficulty of the task would provide better context for the results.\n- Clarifying whether the reward function's dependence on SSIM and NMSE drives the agent to prioritize high-frequency edges or low-frequency contrast would help interpret the learned masks.\n- Providing more details on the \"invalid actions masked\" step in the policy network would clarify how the agent handles the budget constraints during the rollout."},"questions_to_address_in_the_rebuttal":{"value":"- How does the framework perform when tested on implant materials or geometries significantly different from the training set (e.g., Titanium vs. CoCr)?\n- Is there a straightforward pathway to extend this method to multi-coil parallel imaging, and does the state space explosion become a bottleneck?\n- Have you attempted to apply the trained policy to any real phantom data (even without ground truth) to qualitatively assess if the simulation-learned features transfer to real k-space?"},"preliminary_rating":{"value":4}},"parentInvitations":"MIDL.io/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1771530043374,"tcdate":1767504025926,"writers":["MIDL.io/2026/Conference","MIDL.io/2026/Conference/Submission225/Reviewer_rFC2"],"signatures":["MIDL.io/2026/Conference/Submission225/Reviewer_rFC2"],"forum":"oq2ZSUiaaB","number":1,"license":"CC BY 4.0","cdate":1767504025926,"readers":["everyone"],"invitations":["MIDL.io/2026/Conference/Submission225/-/Official_Review","MIDL.io/2026/Conference/-/Edit"],"mdate":1771530043374,"domain":"MIDL.io/2026/Conference","replyto":"oq2ZSUiaaB","id":"Md6HlaoyO7","forumContent":{"TLDR":{"value":"This work presents the first unified approach to jointly address metal artifact reduction and accelerated MRI acquisition through co-adaptive reinforcement learning."},"venue":{"value":"MIDL 2026 Poster"},"midl_latex_submission_checklist":{"value":["The paper compiles correctly using the pdflatex compiler.","Created a single midl26_NNN.zip file with midl26_NNN.tex, midl26_NNN.bib, all necessary figures and files.","The LaTeX file includes the correct header commands before \\title.","Hyperref is preloaded; do not modify it.","The times package is not used.","All co-authors are listed with correct firstname, lastname, and any suffix/prefix.","Math in the title and abstract uses valid LaTeX.","References are provided through the .bib file only.","Tables and figures stay within the page margins.","The zip archive contains all necessary figures and no unused files.","Special formatting from rebuttal has been removed.","All special characters use LaTeX commands.","Appendices and supplementary material are included in the same PDF after references.","The main paper does not exceed 12 pages."]},"keywords":{"value":["MRI reconstruction","k-space sampling","metal artifact reduction","reinforcement learning"]},"reproducibility":{"value":"https://github.com/hrlblab/masc"},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Metal implants in MRI cause severe artifacts that degrade image quality and hinder clinical diagnosis. Traditional approaches address metal artifact reduction (MAR) and accelerated MRI acquisition as separate problems. We propose MASC, a unified reinforcement learning framework that jointly optimizes metal-aware k-space sampling and artifact correction for accelerated MRI. To enable supervised training, we construct a paired MRI dataset using physics-based simulation, generating k-space data and reconstructions for phantoms with and without metal implants. This paired dataset provides simulated 3D MRI scans with and without metal implants, where each metal-corrupted sample has an exactly matched clean reference, enabling direct supervision for both artifact reduction and acquisition policy learning. We formulate active MRI acquisition as a sequential decision-making problem, where an artifact-aware Proximal Policy Optimization (PPO) agent learns to select k-space phase-encoding lines under a limited acquisition budget. The agent operates on undersampled reconstructions processed through a U-Net-based MAR network, learning patterns that maximize reconstruction quality. We further propose an end-to-end training scheme where the acquisition policy learns to select k-space lines that best support artifact removal while the MAR network simultaneously adapts to the resulting undersampling patterns. Experiments demonstrate that MASC's learned policies outperform conventional sampling strategies, and end-to-end training improves performance compared to using a frozen pre-trained MAR network, validating the benefit of joint optimization. Cross-dataset experiments on FastMRI with physics-based artifact simulation further confirm generalization to realistic clinical MRI data. The code and models of MASC have been made publicly available: https://github.com/hrlblab/masc."},"_bibtex":{"value":"@inproceedings{\nlu2026masc,\ntitle={{MASC}: Metal-Aware Sampling and Correction via Reinforcement Learning for Accelerated {MRI}},\nauthor={Zhengyi Lu and Ming Lu and Chongyu Qu and Junchao Zhu and Junlin Guo and Marilyn Lionts and Yanfan Zhu and Yuechen Yang and Tianyuan Yao and Jayasai Rajagopal and Bennett Allan Landman and Xiao Wang and Xinqiang Yan and Yuankai Huo},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=oq2ZSUiaaB}\n}"},"title":{"value":"MASC: Metal-Aware Sampling and Correction via Reinforcement Learning for Accelerated MRI"},"latex_code":{"value":"/attachment/6e3cb240d68d361f65d7ee6594913d5299bf3cbc.zip"},"secondary_subject_area":{"value":"Safe and Trustworthy Learning-assisted Solutions for Medical Imaging"},"pdf":{"value":"/pdf/3ad3f101cf6a309c2086d04e75828aafcae7b654.pdf"},"copyright_form":{"value":"/attachment/cb5a5800e4d1982d2c981d6a1836f32a46972824.pdf"},"visa":{"value":"Yes"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Conference"},"paperhash":{"value":"lu|masc_metalaware_sampling_and_correction_via_reinforcement_learning_for_accelerated_mri"},"primary_subject_area":{"value":"Image Acquisition and Reconstruction"},"authorids":{"value":["~Zhengyi_Lu2","~Ming_Lu7","~Chongyu_Qu1","~Junchao_Zhu3","~Junlin_Guo1","~Marilyn_Lionts1","~Yanfan_Zhu1","~Yuechen_Yang1","~Tianyuan_Yao1","~Jayasai_Rajagopal1","~Bennett_A._Landman1","~Xiao_Wang21","~Xinqiang_Yan1","~Yuankai_Huo2"]},"registration":{"value":"Yes"},"authors":{"value":["Zhengyi Lu","Ming Lu","Chongyu Qu","Junchao Zhu","Junlin Guo","Marilyn Lionts","Yanfan Zhu","Yuechen Yang","Tianyuan Yao","Jayasai Rajagopal","Bennett Allan Landman","Xiao Wang","Xinqiang Yan","Yuankai Huo"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"This paper proposes One-Shot World Model (OSWM), a transformer-based world model trained entirely on synthetic data. Unlike traditional world models trained on environment observations, OSWM is pretrained entirely on synthetic data generated from a prior distribution composed of randomly initialised neural networks. OSWM can then be used as a world model for a given environment by leveraging in-context learning. Experiments demonstrate OSWM's ability to train RL agents that solve simple environments like GridWorld, CartPole, and a custom control task. However, the model struggles with more complex environments."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"- To study the effectiveness of synthetic pretraining, I think it would be useful to have a baseline world model that is trained from scratch using the in context samples provided to OSWM.\n- I am a bit unclear on how the momentum prior is used in the prior training. Section 3.2 mentions the NN prior and the momentum prior are concatenated. But Fig 1 shows that the entire state vector $x_t = [s_t^{1:d_s}, a_t^{1:d_a}$ so $x_t$ is just the NN prior, so where is the momentum prior concatenated? I do not see $v_t$ or $p_t$ in Fig 1."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- Training world models entirely on synthetic data generated from randomly initialized neural networks is a novel and intriguing idea. \n- Promising results on simple environments: The authors demonstrate successful agent training on simple environments purely from synthetic priors. It suggests the potential of this approach for rapid adaptation to new tasks. Leveraging in-context learning allows for quick adaptation to unseen environments without extensive retraining.\n- The paper provides a detailed analysis of the impact of different prior components, specifically analyzing the distribution of states produced by the NN prior and the momentum prior. They also study the impact of different context sampling strategies. \n- The paper provides a promising first step for further research in pretraining of RL world models with synthetic data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited applicability to complex environments: The current model struggles with harder environments highlighting the need for further development."}},"nonreaders":[],"tmdate":1731427585137,"tcdate":1730904161641,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2373/Reviewer_v5Tr"],"signatures":["ICLR.cc/2025/Conference/Submission2373/Reviewer_v5Tr"],"forum":"VjeT8VFhHo","number":4,"license":"CC BY 4.0","cdate":1730904161641,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2373/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427585137,"domain":"ICLR.cc/2025/Conference","replyto":"VjeT8VFhHo","id":"ZEn8Ii234k","forumContent":{"TLDR":{"value":"We propose a transformer world model trained on synthetic data sampled from randomly initialized, untrained neural networks. With minimal context, OSWM adapts and trains agents to solve simple environments like GridWorld and CartPole gym."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["World Models","Synthetic Pretraining","Reinforcement Learning","In-Context Learning"]},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"A World Model is a compressed spatial and temporal representation of a real world environment that allows one to train an agent or execute planning methods. However, world models are typically trained on observations from the real world environment, and they usually do not enable learning policies for other real environments. We propose One-Shot World Model (OSWM), a transformer world model that is learned in an in-context learning fashion from purely synthetic data sampled from a prior distribution. Our prior is composed of multiple randomly initialized neural networks, where each network models the dynamics of each state and reward dimension of a desired target environment. We adopt the supervised learning procedure of Prior-Fitted Networks by masking next-state and reward at random context positions and query OSWM to make probabilistic predictions based on the remaining transition context. During inference time, OSWM is able to quickly adapt to the dynamics of a simple grid world, as well as the CartPole gym and a custom control environment by providing 1k transition steps as context and is then able to successfully train environment-solving agent policies. However, transferring to more complex environments remains a challenge, currently. Despite these limitations, we see this work as an important stepping-stone in the pursuit of learning world models purely from synthetic data."},"_bibtex":{"value":"@misc{\nferreira2024oneshot,\ntitle={One-shot World Models Using a Transformer Trained on a Synthetic Prior},\nauthor={Fabio Ferreira and Moreno Schlageter and Raghu Rajan and Andr{\\'e} Biedenkapp and Frank Hutter},\nyear={2024},\nurl={https://openreview.net/forum?id=VjeT8VFhHo}\n}"},"title":{"value":"One-shot World Models Using a Transformer Trained on a Synthetic Prior"},"pdf":{"value":"/pdf/4e438753bbfe406dced55e1c081055bf1e1eef6a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"ferreira|oneshot_world_models_using_a_transformer_trained_on_a_synthetic_prior"},"authorids":{"value":["~Fabio_Ferreira1","~Moreno_Schlageter1","~Raghu_Rajan1","~André_Biedenkapp1","~Frank_Hutter1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fabio Ferreira","Moreno Schlageter","Raghu Rajan","André Biedenkapp","Frank Hutter"]}},"version":2},{"content":{"summary":{"value":"This paper proposes OptiCo, another new physics-guided CNN that embeds an optical phase (OP) kernel with complex convolutions for lithography mask optimization and simulation, showing lower MSE/EPE and some OOD gains on LithoBench with a TV prior for manufacturability."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Same as weakness."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"- Computational lithography is impactful for future computation and accelerating accurate simulation and optimization matters.\n\n- Encodes diffraction via an OP kernel shows empirical gains with reported improvements on benchmark tables."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Overall, I do not find this paper is interesting as the architectural change is domain-specific and incremental in the long line of physics-informed models and there isn’t a broadly useful for ML community. For a top ML venue, the ML contribution is weak the work seems better aligned with EDA venues where domain impact is the primary metric. If considering its contribution to AI for science and engineering, the real-world experiments would be required to justify its impact. \nMoreover, the evaluation cannot rule out cherry-picking with unclear splits/model selection, single-seed reporting, limited robustness and results may be benchmark and simulator-locked.\n\n- Limited ML novelty & generality. OP-kernel + complex convs is an incremental tweak as in previous lines of computational lithography papers they already have considered similar ideas. It is also task-specific and lacks a general learning insight likely to interest the broader ML community.\n- The experiment has unclear train/val/test protocol with no precise split definitions,  ambiguous use of validation for early stopping/hyper-params and appropriate usage of test-set is not guaranteed.\n- They almost use single-seed results; there are no mean±std or significance tests which cannot rule out cherry-picking risk.\n- OOD protocol is under-specified. Construction and difficulty balance of OOD sets are unclear. The stability across different split draws not shown.\n- The method is quite empirical. There is no theoretical insight on why the proposed architectural changes can be effective for lithographic modeling. This problem is non-negligible especially when demonstrating only on benchmarks. \n- The method is tied with specific simulator and process. No sensitivity to source/resist/NA changes or alternate simulators; gains could reflect over-tuning to one stack.\n- The ablations need to be cleaner. Right now it’s hard to tell what’s driving the gains. Please keep the training budget and seeds the same across variants, and pick checkpoints by validation, not by best test score. Also separate the effects such as show what changes when you switch to complex convs, when you change kernel size, and when you add TV.\n- The authors need to make baseline comparisons fairer. Tell us the exact hyperparams and training time you used for each baseline, and try to match hardware and budgets. If one model gets more epochs or a faster GPU, efficiency and accuracy claims won’t be comparable."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921744918,"tcdate":1761989060105,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10445/Reviewer_E2pw"],"signatures":["ICLR.cc/2026/Conference/Submission10445/Reviewer_E2pw"],"forum":"jhV81HJ06h","number":4,"license":"CC BY 4.0","cdate":1761989060105,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10445/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921744918,"domain":"ICLR.cc/2026/Conference","replyto":"jhV81HJ06h","id":"OytPEC53GY","forumContent":{"TLDR":{"value":"An optical physics-driven neural network for semiconductor mask optimization"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["AI4Science","Semiconductor","Mask Optimization","Optical Diffraction"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"In recent years, the increasing demand for smaller and more powerful semiconductors highlighted the critical role of lithography—a key stage in semiconductor manufacturing responsible for precise mask design and wafer patterning. To meet these demands, the semiconductor industry has increasingly adopted computational lithography, employing machine learning and deep learning techniques to accelerate advancements in lithographic technology. Despite the various research efforts and successes in computational lithography, there remains a lack of explicit incorporation of physical principles. This gap limits the ability of existing methods to fully capture the complex physical phenomena inherent in lithography behaviors. To bridge this gap, we propose OptiCo, a novel convolutional neural network that seamlessly integrates optical diffraction principles into its architecture. At its core, OptiCo employs an optical phase kernel to model phase variations resulting from light propagation, effectively capturing the physical interactions among light, masks, and wafers. We evaluate OptiCo on semiconductor lithography benchmarks, demonstrating its superior performance in mask optimization tasks, with its remarkable generalization capabilities in OOD datasets."},"_bibtex":{"value":"@misc{\nson2025optical,\ntitle={Optical Diffraction-based Convolution for Semiconductor Mask Optimization},\nauthor={Young-Han Son and Dong-Hee Shin and Deok-Joong Lee and Hyun Jung Lee and Tae-Eui Kam},\nyear={2025},\nurl={https://openreview.net/forum?id=jhV81HJ06h}\n}"},"title":{"value":"Optical Diffraction-based Convolution for Semiconductor Mask Optimization"},"pdf":{"value":"/pdf/8ae75180b8c0407aa71cf1be7cee8ebff3eff534.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"son|optical_diffractionbased_convolution_for_semiconductor_mask_optimization"},"authorids":{"value":["~Young-Han_Son1","~Dong-Hee_Shin1","~Deok-Joong_Lee1","~Hyun_Jung_Lee1","~Tae-Eui_Kam1"]},"authors":{"value":["Young-Han Son","Dong-Hee Shin","Deok-Joong Lee","Hyun Jung Lee","Tae-Eui Kam"]}},"version":2},{"content":{"TLDR":{"value":"We present Humanity's Sixth Sense (HSS), a benchmark for intuitive visual reasoning in multimodal large language models."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["intuitive visual reasoning; multimodal"]},"supplementary_material":{"value":"/attachment/1eee4c32859c96111c14145d3a79d604feca20af.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Humans perceive far more in a scene than what is explicitly depicted: a single glance captures past causes and future trajectories; a quick peek determines if a vehicle can fit between two parked cars; a few seconds of video reveals who holds authority in a room; and a fleeting clip highlights subtle abstract patterns like unwritten rules or hidden labels. This capacity reflects a form of humanity's *sixth sense*: an intuitive reasoning mechanism that recovers implicit information beyond raw sensory perception. Crucially, this rapid, zero-shot visual intuition underpins everyday navigation and social interaction, making it a vital capability for Multimodal Large Language Models (MLLMs) deployed alongside people. Existing visual benchmarks, however, target either deliberate expert-level analysis in academic and mathematical domains or low-level perception, leaving the intuitive reasoning that people perform largely untested. To bridge this gap, we introduce **Humanity's Sixth Sense (HSS)**, a benchmark for intuitive visual reasoning. HSS spans diverse image and video inputs, organizes items under a structured taxonomy, and pairs each with human-written prompts probing the implicit temporal, spatial, social, and abstract structure that people infer at a glance. State-of-the-art models fall well short of human performance on these tasks: participants reach 93.1\\% accuracy, while the strongest model, GPT-6-astra, reaches only 53.6\\% even at maximum reasoning effort. Despite excelling in many complex tasks that require advanced perception and knowledge, current models still struggle significantly on these visual tasks tasks intuitive for humans. We further explore agentic frameworks that apply dynamic visual manipulation to HSS, which narrows but does not close the gap. HSS establishes intuitive visual reasoning as a measurable axis of progress and directs attention to a capability that scaling on current benchmarks has so far left behind."},"_bibtex":{"value":"@inproceedings{\nanonymous2026humanitys,\ntitle={Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=1LQg9N8lly},\nnote={under review}\n}"},"title":{"value":"Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models"},"pdf":{"value":"/pdf/b6cad62de4364c3627f5333bc3a98313e6ff6f40.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791234196897,"tcdate":1789766375241,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission49023/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission49023/Authors"],"forum":"1LQg9N8lly","license":"CC BY 4.0","number":49023,"cdate":1789766375241,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission49023/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791234196897,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"1LQg9N8lly","version":2},{"content":{"summary":{"value":"This paper introduces a synthetic data generation strategy for creating high quality SFT and DPO datasets. The proposed strategy uses a multi-agent simulator that explicitly incorporates real-world user requirements into the data synthesis process. This simulator generates realistic and diverse scenarios using 1000 agents and a structured communication mechanism, where each agent is built based on a real-world human profile. These scenarios are then combined with specific user requirement categories (such as coding, dialogue, safety, etc.) to generate high quality instruction data that captures a wide range of real-world human needs.\n\nExperiments are conducted by training Llama3-8B model using synthetic datasets generated by the proposed approach and comparing it with Llama3-8B models trained using various alternative datasets. The results clearly demonstrate the effectiveness of the synthetic datasets generated using the proposed approach."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see weaknesses section for questions regarding the proposed method and experimental results. \n\nThe biggest concern I have is lack of necessary details about the proposed approach. In the current state, the paper feels more like a technical report rather than a scientific paper that others from the community could reproduce and build upon."},"rating":{"value":3},"details_of_ethics_concerns":{"value":"In the rebuttal phase, authors have indicated that the approach presented in the paper heavily relies on user data crawled from Twitter. Authors mentioned that they used LLMs to remove personal information from this data. This was done by simply prompting as LLM using the following prompt:\n***Given the user profile, please identify and remove any personal information such as names, ages, locations, or other identifiers from the following text. {User profile including their tweets}***\n\nIt is unclear how effective this processed is and how much of private user information is being used by the proposed approach."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The overall idea of creating a multi-agent simulator based on real-world human profiles is interesting. Grounding data synthesis in real-world human behavior can result in more realistic data. Specifically, the proposed approach goes beyond recent PersonaHub and leverages interactions between agents to create complex and diverse scenarios.\n\nExperimental results show that the synthetic datasets generated by the proposed approach are more effective than various existing real and synthetic SFT/DPO datasets."},"flag_for_ethics_review":{"value":["Yes, Privacy, security and safety"]},"weaknesses":{"value":"Presentation: The main problem with this paper is lack of specifics to fully understand and replicate the approach. The papers describes the components of the proposed multi-agent simulator at a high level without providing concrete details.\n\nThe proposed approach uses 1000 agents created based on real-world human profiles. Many things are unclear from the paper:\n-- What kind of user profiles are used in the simulator? What is the distribution? Paper does not provide details about these profiles.\n-- Which web sources are used to gather data for these user profiles?\n-- What kind of person-specific data is obtained from the web to create a user profile?\n-- How is this web data processed by the LLM to create a user profile? What are the LLM prompts used?\n-- Line 200 says \"Each profile includes a unique anonymized name, description, and a record of past actions, all processed to protect privacy.\" What does an action mean here? What kind of actions of real humans are obtained from internet?\n\nWhat are the LLM prompts used to create goals and action plans for each agent? Does each agent have multiple life goals?\n\nLine 154-155: \"For MATRIX-Gen-SFT, the instructions are generated with both simplicity and diversity. For MATRIX-Gen-DPO, the instructions are complex and specialized.\"\nThere are no details provided regarding the difference between \"simple SFT instructions\" and \"Complex/specialized DPO instructions\" and how the simplicity/complexity is varied. What types of instructions are considered simple and what are considered complex/specialized?\n\nIt was difficult to understand how the agents and modulators are operating/communicating. I couldn't follow what actions are being communicated and how the modulators actually operate. The paper neither provides concrete specifics nor examples describing the scenario generation mechanism.\n\nExperimental evaluation:\nAccording to Table 9 in Appendix, using specialized instructions (type 2) for SFT and simple instructions (Type 1) for DPO gives the best results (see the last row of Table. 9). However, the authors derive opposite conclusions and perform SFT with simple instructions and DPO with specialized instructions (as specified in lines 154-155). \n\nIt is unclear why authors chose different models as starting points for different domains when evaluating domain-specific Matrix-Gen SFT datasets (coding, safety and multi-turn dialogue).\n\nThe target responses for the synthesized instructions are generated using Llama3-8B-Instruct. Training with these responses can be interpreted as knowledge distillation. Despite distilling from Llama3-8B-Instruct, the Matrix-Tuned-Model outperforms Llama3-8B-Instruct (Table. 2 & 3). There is no discussion regarding this."}},"nonreaders":[],"tmdate":1732561274604,"tcdate":1731285220394,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6823/Reviewer_PZzj"],"signatures":["ICLR.cc/2025/Conference/Submission6823/Reviewer_PZzj"],"forum":"o83aL1nZJd","number":3,"license":"CC BY 4.0","cdate":1731285220394,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6823/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732561274604,"domain":"ICLR.cc/2025/Conference","replyto":"o83aL1nZJd","id":"kesCBciIIV","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["large language models","llm alignment","multi-agent simulation","llm society"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Post-training is essential for enabling large language models (LLMs) to follow human instructions. \nInspired by the recent success of using LLMs to simulate human society, we leverage multi-agent simulation to automatically generate diverse text-based scenarios, capturing a wide range of real-world human needs. \nWe introduce MATRIX, a multi-agent simulator that creates realistic and scalable scenarios. \nLeveraging these outputs, we introduce a novel scenario-driven instruction generator MATRIX-Gen for controllable and highly realistic data synthesis. Extensive experiments demonstrate that our framework effectively generates both general and domain-specific data. Notably, on AlpacaEval 2 and Arena-Hard benchmarks, Llama-3-8B-Base, post-trained on datasets synthesized by MATRIX-Gen with just 20K instruction-response pairs, outperforms Meta's Llama-3-8B-Instruct model, which was trained on over 10M pairs."},"_bibtex":{"value":"@misc{\ntang2025synthesizing,\ntitle={Synthesizing Post-Training Data for {LLM}s through Multi-Agent Simulation},\nauthor={Shuo Tang and Xianghe Pang and Zexi Liu and Bohan Tang and Rui Ye and Xiaowen Dong and Yanfeng Wang and Siheng Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=o83aL1nZJd}\n}"},"title":{"value":"Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation"},"pdf":{"value":"/pdf/1f867b1848afdd01de31ae1ee8c887aed765de7c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"tang|synthesizing_posttraining_data_for_llms_through_multiagent_simulation"},"authorids":{"value":["~Shuo_Tang2","~Xianghe_Pang1","~Zexi_Liu1","~Bohan_Tang1","~Rui_Ye1","~Xiaowen_Dong1","~Yanfeng_Wang1","~Siheng_Chen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shuo Tang","Xianghe Pang","Zexi Liu","Bohan Tang","Rui Ye","Xiaowen Dong","Yanfeng Wang","Siheng Chen"]}},"version":2},{"content":{"summary":{"value":"The paper introduces PivotMesh, a framework for generating compact and detailed 3D meshes at scale. The approach leverages a transformer-based auto-encoder to encode meshes into discrete tokens, which are then decoded hierarchically from face to vertex level. A key contribution is the use of pivot vertices as a coarse representation to guide the generation of complete mesh tokens, allowing for complex topology modeling. The method is evaluated on diverse datasets, including ShapeNet, Objaverse, and Objaverse-xl, demonstrating its versatility and effectiveness in generating high-quality 3D meshes."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"The paper is well-written and presents a contribution to the field of 3D mesh generation. I am impressed with the quality of the generated results, but there are some minor issues: the comparison with other similar baselines and more diverse generated shapes. With the above suggestions addressed, the paper would be a strong candidate for publication. I recommend accepting this paper with minor revisions.\nDetailed questions please refer to weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. PivotMesh addresses a significant challenge in 3D mesh generation by proposing a scalable framework that can handle large-scale datasets with simplified triangles. The use of pivot vertices as a coarse representation for guiding mesh generation is innovative and effectively handles complex topologies.\n2. The paper provides a thorough evaluation of PivotMesh across various datasets and applications, including mesh generation, variation, and refinement. The generated meshes are of high quality, with sharp details and complex geometries, outperforming existing methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Diversity of Generated Meshes: While the paper shows diverse mesh generation, a more systematic analysis or quantification of diversity could strengthen the results. The evaluations on more diverse shapes with complex topologies are encouraged to be conducted, such as some shapes with lots of holes, the thin structures (ficus in nerf-synthetic data), and so on.\n2. More direct comparisons with the current state-of-the-art methods, especially in terms of computational efficiency and mesh quality, could be beneficial. There is more relative work, such as meshanything(2), EdgeRunner, the comparison is very essential to demonstrate the superiority of the proposed method.\n3. The limitations and failure cases should be discussed comprehensively. And the paper acknowledges that the controlling ability of PivotMesh is not sufficient, and sometimes undesired geometries are produced. Enhancing the control mechanisms could be a valuable addition."}},"nonreaders":[],"tmdate":1731428443867,"tcdate":1730632166111,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5566/Reviewer_2uyx"],"signatures":["ICLR.cc/2025/Conference/Submission5566/Reviewer_2uyx"],"forum":"WAC8LmlKYf","number":3,"license":"CC BY 4.0","cdate":1730632166111,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5566/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428443867,"domain":"ICLR.cc/2025/Conference","replyto":"WAC8LmlKYf","id":"yQ7aCT1vGg","forumContent":{"TLDR":{"value":"We extend the native mesh generation to large-scale dataset with the proposed generic and scalable mesh generation framework, namely PivotMesh."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["mesh generation","auto-regressive generation","3D generation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Generating compact and sharply detailed 3D meshes poses a significant challenge for current 3D generative models. Different from extracting dense meshes from neural representation, some recent works try to model the native mesh distribution (i.e., a set of triangles), which generates more compact results as humans crafted. However, due to the complexity and variety of mesh topology, most of these methods are typically limited to generating meshes with simple geometry. In this paper, we introduce a generic and scalable mesh generation framework PivotMesh, which makes an initial attempt to extend the native mesh generation to large-scale datasets. We employ a transformer-based autoencoder to encode meshes into discrete tokens and decode them from face level to vertex level hierarchically. Subsequently, to model the complex typology, our model first learns to generate pivot vertices as coarse mesh representation and then generate the complete mesh tokens with the same auto-regressive Transformer. This reduces the difficulty compared with directly modeling the mesh distribution and further improves the model controllability. PivotMesh demonstrates its versatility by effectively learning from both small datasets like Shapenet, and large-scale datasets like Objaverse and Objaverse-xl. Extensive experiments indicate that PivotMesh can generate compact and sharp 3D meshes across various categories, highlighting its great potential for native mesh modeling."},"_bibtex":{"value":"@inproceedings{\nweng2025pivotmesh,\ntitle={PivotMesh: Generic 3D Mesh Generation via Pivot Vertices Guidance},\nauthor={Haohan Weng and Yikai Wang and Tong Zhang and C. L. Philip Chen and Jun Zhu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=WAC8LmlKYf}\n}"},"title":{"value":"PivotMesh: Generic 3D Mesh Generation via Pivot Vertices Guidance"},"pdf":{"value":"/pdf/090cb1e8b940da72c791c5f178c91fa0181a89aa.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"weng|pivotmesh_generic_3d_mesh_generation_via_pivot_vertices_guidance"},"authorids":{"value":["~Haohan_Weng1","~Yikai_Wang2","~Tong_Zhang14","~C._L._Philip_Chen1","~Jun_Zhu2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haohan Weng","Yikai Wang","Tong Zhang","C. L. Philip Chen","Jun Zhu"]}},"version":2},{"content":{"summary":{"value":"The paper extends recent works which leverages layouts to generate scenes corresponding to complex text-prompts. This work first shows that for complex prompts, existing layout to image generation methods have certain failure modes and proposes some practical modifications which are augmented with existing layout to scene generation methods.  First, the authors propose a scene blueprint to represent complex text-prompts; Secondly the authors design an iterative refinement process which improves the alignment of generated images with the complex text-prompts."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"- The research question is a very practical problem — usually most of the text-to-image generation models are not good at coherent images corresponding to complex prompts, so providing a solution for it is important.\n- The method consists of various components (a lot of these components are existing though) and is conceptually intuitive!\n- The framework obtains strong results on human-study for fidelity of images generated for long prompts."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Cons / Questions\n\n- While the writing is satisfactory, it can still be improved! The authors should provide more information in the paper on how Eq. (5) is used to guide the sampling process.\n- Can the authors provide more intuition on the interpolation step? \n- Given that there are stronger open-source diffusion models (e.g., DeepFloyd) — the authors should provide some context on how long prompts work in those cases, as they use a stronger text-encoder like T5. \n- While the authors comment that the size of the tokens (77 in CLIP) is one of the potential reasons on why SD cannot generate compositional prompts — I believe this is only partially true. Even for non-complex compositional prompts, SD is not able to generate coherent images. Can the authors comment in general on some potential reasons why these text-to-image models are not able to generate images corresponding to simple compositional or complex prompts? I think both are related somehow and it will be beneficial to provide some context regarding it."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"See Cons/Questions;\nOverall, the paper is practical, but the various components though intuitive are not technically novel. While I do agree that not everything in a paper needs to be novel — the authors should provide solid justifications on the design of each component.  \n\nI am happy to revisit my scores after the rebuttal!"},"rating":{"value":"5: marginally below the acceptance threshold"},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636474793,"tcdate":1699300622095,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4903/Reviewer_XXwY"],"signatures":["ICLR.cc/2024/Conference/Submission4903/Reviewer_XXwY"],"forum":"mNYF0IHbRy","number":4,"license":"CC BY 4.0","cdate":1699300622095,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4903/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636474793,"domain":"ICLR.cc/2024/Conference","replyto":"mNYF0IHbRy","id":"GSQu5jx5YE","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion","LLM"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in generating images from short, single-object descriptions, these models often struggle to faithfully capture all the nuanced details within longer and more elaborate textual inputs. In response, we present a novel approach leveraging Large Language Models (LLMs) to extract critical components from text prompts, including bounding box coordinates for foreground objects, detailed textual descriptions for individual objects, and a succinct background context. These components form the foundation of our layout-to-image generation model, which operates in two phases. The initial Global Scene Generation utilizes object layouts and background context to create an initial scene but often falls short in faithfully representing object characteristics as specified in the prompts. To address this limitation, we introduce an Iterative Refinement Scheme that iteratively evaluates and refines box-level content to align them with their textual descriptions, recomposing objects as needed to ensure consistency. Our evaluation on complex prompts featuring multiple objects demonstrates a substantial improvement in recall compared to baseline diffusion models. This is further validated by a user study, underscoring the efficacy of our approach in generating coherent and detailed scenes from intricate textual inputs. Our iterative framework offers a promising solution for enhancing text-to-image generation models' fidelity with lengthy, multifaceted descriptions, opening new possibilities for accurate and diverse image synthesis from textual inputs."},"_bibtex":{"value":"@inproceedings{\ngani2024llm,\ntitle={{LLM} Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts},\nauthor={Hanan Gani and Shariq Farooq Bhat and Muzammal Naseer and Salman Khan and Peter Wonka},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=mNYF0IHbRy}\n}"},"title":{"value":"LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts"},"pdf":{"value":"/pdf/6b9441d77313c54bc391372756461921c6c8b41e.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"gani|llm_blueprint_enabling_texttoimage_generation_with_complex_and_detailed_prompts"},"authorids":{"value":["~Hanan_Gani1","~Shariq_Farooq_Bhat1","~Muzammal_Naseer1","~Salman_Khan4","~Peter_Wonka1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hanan Gani","Shariq Farooq Bhat","Muzammal Naseer","Salman Khan","Peter Wonka"]}},"version":2},{"content":{"summary":{"value":"This paper studies the theoretical foundations of softmax attention by introducing a principled single-location regression (SLR) model. The authors derive analytical and asymptotic results using statistical physics tools (notably order parameters and replica analysis) to show that softmax attention achieves Bayes-optimal performance in high-dimensional limits, while linear attention and other alternatives (e.g., kernelized or element-wise nonlinearities) fundamentally fall short. The paper further provides a finite-sample characterization of test risk, confirming the statistical and computational advantages of softmax. Overall, it offers a clean theoretical framework that bridges information retrieval toy models and high-dimensional generalization analysis."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Can the proposed Single-Location Regression framework be extended to multi-location or multi-head attention, where multiple tokens jointly determine the output? Would the Bayes-optimality of softmax still hold in those cases?\n\nThe analysis assumes Gaussian and independent token embeddings. How sensitive are the results to these assumptions? Would correlated or structured embeddings affect the theoretical conclusions?\n\n\nHave the authors tested whether the predicted statistical gap between softmax and linear attention appears in small-scale Transformer experiments or synthetic retrieval tasks (e.g., Needle-in-a-Haystack)?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Clear and original theoretical contribution:\nThe paper introduces a new analytical framework — the Single-Location Regression (SLR) model — to study the statistical behavior of attention mechanisms. This formulation unifies previous “needle-in-a-haystack”-type setups under a mathematically tractable regime, enabling a principled comparison between softmax and linear attention.\n\nMethodological depth:\nThe work combines sequence multi-index models with replica-based high-dimensional analysis, bringing together tools from statistical physics and modern learning theory. This cross-disciplinary approach is technically sophisticated and extends recent progress in the theoretical understanding of attention networks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Idealized task setting:\nThe SLR model assumes the label depends on a single token’s linear transformation, which, while analytically convenient, is far from the multi-head, multi-layer structure of real Transformers. Hence, the practical relevance is limited.\n\nLimited empirical grounding:\nThe validation is restricted to synthetic setups, without experiments on realistic retrieval or sequence tasks. It remains unclear whether the predicted statistical advantage of softmax can be observed in real neural models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931278305,"tcdate":1761966912062,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19337/Reviewer_hz39"],"signatures":["ICLR.cc/2026/Conference/Submission19337/Reviewer_hz39"],"forum":"R44n1gZNjQ","number":2,"license":"CC BY 4.0","cdate":1761966912062,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19337/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931278305,"domain":"ICLR.cc/2026/Conference","replyto":"R44n1gZNjQ","id":"elY0l3pGcm","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We demonstrate advantages of softmax-based attention over alternatives in a high-dimensional information retrieval regression task."},"keywords":{"value":["attention","softmax attention","deep learning theory","sparse information retrieval","high-dimensional limit","high-dimensional statistics"]},"supplementary_material":{"value":"/attachment/44b82548cba3a05ec9eec92f65ad15b138b9ae2f.zip"},"primary_area":{"value":"learning theory"},"abstract":{"value":"Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly understood, and many theoretical works have focused on the easier-to-analyze linearized attention. In this work, we address this gap through a principled study of the single-location regression task, where the output depends on a linear transformation of a single input token at a random location. Building on ideas from statistical physics, we develop an analysis of attention-based predictors in the high-dimensional limit, where generalization performance is captured by a small set of order parameters. At the population level, we show that softmax achieves the Bayes risk, whereas linear attention fundamentally falls short. We then examine other activation functions to identify which properties are necessary for optimal performance. Finally, we analyze the finite-sample regime: we provide an asymptotic characterization of the test error and show that, while softmax is no longer Bayes-optimal, it consistently outperforms linear attention. We discuss the connection with optimization by gradient-based algorithms."},"_bibtex":{"value":"@inproceedings{\nduranthon2026statistical,\ntitle={Statistical Advantage of Softmax Attention: Insights from Single-Location Regression},\nauthor={O Duranthon and Pierre Marion and Claire Boyer and Bruno Loureiro and Lenka Zdeborov{\\'a}},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=R44n1gZNjQ}\n}"},"title":{"value":"Statistical Advantage of Softmax Attention: Insights from Single-Location Regression"},"pdf":{"value":"/pdf/bd399bd26b2e0078a342ef251b083271386861ad.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"duranthon|statistical_advantage_of_softmax_attention_insights_from_singlelocation_regression"},"authorids":{"value":["~O_Duranthon1","~Pierre_Marion1","~Claire_Boyer1","~Bruno_Loureiro1","~Lenka_Zdeborová1"]},"authors":{"value":["O Duranthon","Pierre Marion","Claire Boyer","Bruno Loureiro","Lenka Zdeborová"]}},"version":2},{"content":{"venue":{"value":"Knowl. Based Syst. 2024"},"venueid":{"value":"dblp.org/journals/KBS/2024"},"paperhash":{"value":"seol|a_novel_physicsaware_graph_network_using_highorder_numerical_methods_in_weather_forecasting_model"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yunchang_Seol:","https://dblp.org/search/pid/api?q=author:Suho_Kim:","https://dblp.org/search/pid/api?q=author:Minwoo_Jung:","~Youngjoon_Hong1"]},"html":{"value":"https://doi.org/10.1016/j.knosys.2024.112158"},"_bibtex":{"value":"@article{DBLP:journals/kbs/SeolKJH24,\n  author={Yunchang Seol and Suho Kim and Minwoo Jung and Youngjoon Hong},\n  title={A novel physics-aware graph network using high-order numerical methods in weather forecasting model},\n  year={2024},\n  cdate={1704067200000},\n  journal={Knowl. Based Syst.},\n  volume={300},\n  pages={112158},\n  url={https://doi.org/10.1016/j.knosys.2024.112158}\n}\n"},"abstract":{"value":"In this paper, we present a novel architecture for accurate weather forecasting within the framework of Physics-aware Graph Networks (PaGN), which aims to improve the prediction of weather time series by integrating both data and physical equations defined on sparsely distributed spatial domains. This approach employs geometric deep learning by accounting for both supervised loss and physics loss. Solely in supervised learning, data-driven training demands a large amount of data, whereas integrating physics information can reduce this requirement, resulting in faster training. The physics awareness also provides a stronger inductive bias and better generalization, directing the machine learning process by clarifying the complexity of weather data. Thus, a more accurate discretization of physics equations can enhance the physics awareness of the complex weather dynamics. Meanwhile, weather data is essentially gathered at observatories that are not located densely, since covering the Earth with numerous observatories is prohibitively expensive. Given its reliance on graph structure, the proposed PaGN is well-suited for handling sparsely distributed data particularly defined on arbitrary geometric domain. One of the main contributions of this study is the use of high-order numerical methods approximating physics equations defined on graph neural networks for achieving accuracy improvements. The fractional graph Laplacian operator is further integrated with the physics equations to account for the non-local characteristics of complex weather data. For computational efficiency, the accuracy improved PaGN architecture is smoothly implemented in the embedding space. A series of numerical tests on weather datasets is performed to verify the improved accuracy, robustness, and applicability of the proposed methodology. We first carry out the numerical study on the synthetic datasets obtained from the physics equations with an extra forcing term. We then extensively investigate the accuracies of weather forecasting for various configurations in terms of the geophysical regions and the accuracy-improved PaGN approaches for diffusion and wave equations. Moreover, we emphasize that the fractional graph Laplacian operator incorporated into our PaGN model can further improve the accuracy for weather forecasting."},"title":{"value":"A novel physics-aware graph network using high-order numerical methods in weather forecasting model"},"authors":{"value":["Yunchang Seol","Suho Kim","Minwoo Jung","Youngjoon Hong"]}},"tmdate":1727489410392,"pdate":1704067200000,"tcdate":1727489356605,"writers":["~"],"signatures":["~Youngjoon_Hong1"],"forum":"s25D3FJZYm","license":"CC BY-SA 4.0","number":97308,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727489410392,"domain":"DBLP.org","id":"s25D3FJZYm","version":2},{"content":{"summary":{"value":"This paper proposes ModelBench, a benchmark evaluating whether AI systems can extract executable, physics-based models from scientific papers. Each task provides a paper excerpt and experimental data; models generate runnable Python code implementing a physically meaningful model, fit parameters, and report metrics (MSE, R²). The benchmark includes 20 expert-curated photonics tasks with gold models, hierarchical weighted rubrics, and judging protocols. GPT-5 and Claude Opus 4.1 achieved 39% ± 18% and 28% ± 13% rubric satisfaction, respectively."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- What was the selection criteria for the 20 papers, and how much expert time was required per task?\n- What is the inter-judge agreement rate for rubric scoring?\n- The paper claims to move beyond code generation to “scientific modeling.” How is this distinction operationalized and measured?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- This work explores the challenge of reconstructing physics-based models from literature beyond function-level code generation in prior literatures.\n- Clear pipeline with gold models, rubrics, and reproducible scoring protocols."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Only 20 photonic-circuit tasks with evaluation restricted to two models (GPT-5, Claude Opus 4.1). This scale is insufficient for drawing generalizable conclusions about LLM capabilities.\n- The discussion of related work is limited, focusing primarily on PaperBench and ModelBench without sufficiently situating the benchmark in the broader context of scientific modeling, code generation, and physics-informed learning.\n- Minimal detail on rubric validation, inter-rater reliability, or quality assurance. Judge variance and limited human calibration raise concerns about score reproducibility.\n- The work feels closer to a technical report than a rigorous benchmark study. The limited scale and validation make it unsuitable for publication at a major venue."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920985654,"tcdate":1761939456355,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9367/Reviewer_wzC2"],"signatures":["ICLR.cc/2026/Conference/Submission9367/Reviewer_wzC2"],"forum":"Bw9LCBz9KW","number":2,"license":"CC BY 4.0","cdate":1761939456355,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9367/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920985654,"domain":"ICLR.cc/2026/Conference","replyto":"Bw9LCBz9KW","id":"Je16WU2k9l","forumContent":{"TLDR":{"value":"ModelBench is a benchmark for testing whether AI systems can read physics papers and produce executable, physics-based models."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Scientific AI benchmarks","Physics","LLM-as-judge","Rubric-based evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We introduce **ModelBench**, a benchmark for evaluating whether AI systems can extract\nexecutable physics-based models from scientific literature. ModelBench couples\n(i) gold-standard reference models,\n(ii) a hierarchical, weighted binary rubric covering physics correctness, code quality, and reproduction quality, and\n(iii) a judge protocol that produces pass/fail scores at rubric leaves.\nUnlike code-generation benchmarks that test function-level correctness, ModelBench\ntargets the end-to-end task of reconstructing physically grounded models from incomplete and underspecified scientific descriptions.\nWe release the benchmark specification, rubric generator and judge prompts,\nand an initial set of 20 gold models within the field of photonic integrated circuits, alongside scripts for fully reproducible evaluation.\nCandidate systems are required to produce a Python implementation of the model,\na plot of the fitted results, and evaluate MSE and $R^2$ metric of the fit.\nUsing general-purpose LLMs as neutral baselines, we report aggregate scores and case studies that reveal common failure modes\n(e.g., constraint violations, phenomenological overfitting) and show how rubric structure aids in diagnostic evaluation.\nWe discuss limitations (judge variance, dataset breadth, implicit-knowledge gaps) and outline a roadmap to expand domains,\ntighten constraint checking, and support multiple valid solutions. ModelBench provides a transparent platform\nfor tracking scientific modeling capabilities in AI under physical and empirical constraints."},"_bibtex":{"value":"@misc{\nschoolkate2026modelbench,\ntitle={ModelBench: A Benchmark for Extracting Executable, Physics-Based Models from Scientific Literature},\nauthor={Pim Schoolkate and Patrick Huembeli and Frank Sch{\\\"a}fer and Krystian Nowakowski and Carlos Arribalzaga Jov{\\'e} and Frank Koppens and Dirk Englund and Jacob M. Taylor},\nyear={2026},\nurl={https://openreview.net/forum?id=Bw9LCBz9KW}\n}"},"title":{"value":"ModelBench: A Benchmark for Extracting Executable, Physics-Based Models from Scientific Literature"},"pdf":{"value":"/pdf/76d73ba2e8b90fc035df3773654a2a2f446bc09b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"schoolkate|modelbench_a_benchmark_for_extracting_executable_physicsbased_models_from_scientific_literature"},"authorids":{"value":["~Pim_Schoolkate1","~Patrick_Huembeli1","~Frank_Schäfer1","~Krystian_Nowakowski1","~Carlos_Arribalzaga_Jové1","~Frank_Koppens1","~Dirk_Englund1","~Jacob_M._Taylor1"]},"authors":{"value":["Pim Schoolkate","Patrick Huembeli","Frank Schäfer","Krystian Nowakowski","Carlos Arribalzaga Jové","Frank Koppens","Dirk Englund","Jacob M. Taylor"]}},"version":2},{"content":{"correctness":{"value":"Yes, the synthetic data generation scheme is quite thorough and the LLMs benefit from this knowledge."},"summary_and_contributions":{"value":"The contributions of the work are as follows:\n\n- Introduces a new benchmark for evaluating computational thinking in LLMs, based in elementary visual programming. \n- Compares the performance of various generative models like GPT4o and Llama on the task. \n- Introduces a novel synthetic data generation technique using symbolic information, to boost performance with fine-tuning. \n- Analyzes how integrating symbolic information improves fine-tuning performance."},"confidence":{"value":3},"documentation":{"value":"Yes, there is sufficient detail on the synthetic data generation methodology."},"rating":{"value":6},"title":{"value":"Introduces new synthetic data generation technique for visual programming. Benchmarks performance of LLMs on small dataset and lacks error analysis."},"ethics":{"value":"No ethics concerns."},"clarity":{"value":"Yes, the paper is well written and clear."},"review":{"value":"The paper is clear and well-motivated in the task of evaluating computational thinking through visual programming. It contributes a unique synthetic data generation technique that is extensive and covers various types of cases. The symbolic information in the synthetic data improves the performance of the Llama model with finetuning. Some limitations of the work are that the evaluation dataset seems too small to make substantial claims and there is not much error analysis into where LLMs lack in the reasoning process."},"strengths":{"value":"- Contributes a synthetic data generation technique, in which tasks are intuitive and extensive.\n- Showcase results for improved finetuning performance using the above synthetic data.\n- Perform many experiments involving multiple open-source, closed-source, and finetuned LLMs, providing a benchmark for comparing visual programming capabilities of future models."},"flag_for_ethics_review":{"value":"2: No, there are no or only very minor ethics concerns"},"relation_to_prior_work":{"value":"Yes, they cover related works in benchmarking programming capabilities and reasoning abilities of LLMs, also discussing how prior works in visual programming do not evaluate vision performance or try finetuning."},"opportunities_for_improvement":{"value":"- What is the total number of samples in the benchmark? If it is just 20 HoC, 21 ACE, and 24 CT, then it is a very small dataset for evaluation and to make substantial claims.\n\n- It is unclear why the authors tried finetuning directly and did not experiment with in-context learning or few-shot prompting. I believe even the synthetic data could be provided as part of in-context learning. Providing those results would make the evaluation more complete and strengthen the argument for fine-tuning. \n\n- The paper can include more error analysis for where the LLMs fail eg. do they partially reason about the tracing correctly or get it wrong right off the bat? Why is Llama family performance on HoC just 0?\n\n- Whilst the synthetic data generation technique is novel, I am not sure of how extendible it is to other datasets and tasks for visual reasoning. Will be great if the authors can discuss this aspect if they believe it is extendible or include it as limitations."},"additional_feedback":{"value":"Question: Is Chain of Thought just zero-shot chain of thought or some other type?"},"limitations":{"value":"Yes, the authors have adequately covered the limitations. I believe the size of the evaluation dataset should also be considered as a limitation as it seems rather small (if there are just 20 HoC, 21 ACE, and 24 CT)."}},"nonreaders":[],"tmdate":1731500707873,"tcdate":1721881468700,"writers":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission1067/Reviewer_pLhh"],"signatures":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission1067/Reviewer_pLhh"],"forum":"q2WT19Ciad","number":3,"license":"CC BY 4.0","cdate":1721881468700,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission1067/-/Official_Review","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1731500707873,"domain":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","replyto":"q2WT19Ciad","id":"SHd5QJfUKx","forumContent":{"TLDR":{"value":"We curate a novel benchmark to assess computational thinking of state-of-the-art generative models and fine-tune models using a novel synthetic data generation methodology."},"venue":{"value":"NeurIPS 2024 Track Datasets and Benchmarks Poster"},"keywords":{"value":["generative models","computational thinking","problem-solving skills","elementary visual programming"]},"supplementary_material":{"value":"/attachment/e5f05ce9684882b0625ab1804c29a24326b8c019.pdf"},"abstract":{"value":"Generative models have demonstrated human-level proficiency in various benchmarks across domains like programming, natural sciences, and general knowledge. Despite these promising results on competitive benchmarks, they still struggle with seemingly simple problem-solving tasks typically carried out by elementary-level students. How do state-of-the-art models perform on standardized programming-related tests designed to assess computational thinking and problem-solving skills at schools? In this paper, we curate a novel benchmark involving computational thinking tests grounded in elementary visual programming domains. Our initial results show that state-of-the-art models like GPT-4o and Llama3 barely match the performance of an average school student. To further boost the performance of these models, we fine-tune them using a novel synthetic data generation methodology. The key idea is to develop a comprehensive dataset using symbolic methods that capture different skill levels, ranging from recognition of visual elements to multi-choice quizzes to synthesis-style tasks. We showcase how various aspects of symbolic information in synthetic data help improve fine-tuned models' performance. We will release the full implementation and datasets to facilitate further research on enhancing computational thinking in generative models."},"_bibtex":{"value":"@inproceedings{\np{\\u{a}}durean2024benchmarking,\ntitle={Benchmarking Generative Models on Computational Thinking Tests in Elementary Visual Programming},\nauthor={Victor-Alexandru P{\\u{a}}durean and Adish Singla},\nbooktitle={The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track},\nyear={2024},\nurl={https://openreview.net/forum?id=q2WT19Ciad}\n}"},"title":{"value":"Benchmarking Generative Models on Computational Thinking Tests in Elementary Visual Programming"},"pdf":{"value":"/pdf/a4f08d12aab06748278f56c0f1ec5f53ef6b556c.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track"},"paperhash":{"value":"pdurean|benchmarking_generative_models_on_computational_thinking_tests_in_elementary_visual_programming"},"authorids":{"value":["~Victor-Alexandru_Pădurean1","~Adish_Singla2"]},"authors":{"value":["Victor-Alexandru Pădurean","Adish Singla"]}},"version":2},{"content":{"venue":{"value":"The European Conference on Computer Vision (ECCV). Workshop on AI for Climate and Conservation (AICC) 2026"},"pdf":{"value":"/pdf/bd73ea18c6c057fa9f46a84ca7a45616b44739c7.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"singh|climate_physics_dynamic_matching"},"authorids":{"value":["~Gurjeet_Sangra_Singh1","~Frantzeska_Lavda1","~Alexandros_Kalousis1"]},"html":{"value":"https://arxiv.org/pdf/2602.17477"},"abstract":{"value":"Deep generative models such as flow matching and diffusion models have shown potential for learning complex dynamical systems, but typically act as black boxes that neglect underlying physical structure, while physics-based models governed by partial differential equations are often incomplete due to missing source terms, or uncertain parametrisations. We present Climate Physics Dynamic Matching (ClimPhyDM), a variational simulation-free dynamics informed framework for weather forecasting that combines an advection-type physics prior with data-driven components in a variational framework. % to capture the stochasticity and multi-modality of unresolved atmospheric dynamics. On the ERA5 benchmark at hourly (42-hour) and monthly (5-month) resolutions, ClimPhyDM outperforms ClimODE, and GB-DM, keeping the lower error at extended horizon, indicating improved temporal stability and resistance to error accumulation, while its simulation-free paradigm also enables training on a single modest 12 GB consumer GPU."},"title":{"value":"Climate Physics Dynamic Matching"},"authors":{"value":["Gurjeet Sangra Singh","Frantzeska Lavda","Alexandros Kalousis"]}},"tmdate":1790584900683,"pdate":1787781600000,"tcdate":1790584900683,"writers":["~Gurjeet_Sangra_Singh1","~Frantzeska_Lavda1","~Alexandros_Kalousis1"],"signatures":["~Gurjeet_Sangra_Singh1"],"forum":"KlQ11ZIlGj","license":"arXiv.org perpetual, non-exclusive license","number":55011,"cdate":1790584900683,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1790584900683,"domain":"OpenReview.net/Archive","id":"KlQ11ZIlGj","version":2},{"content":{"venue":{"value":"CoRR 2021"},"pdf":{"value":"http://arxiv.org/pdf/2108.10470v2"},"venueid":{"value":"dblp.org/journals/CORR/2021"},"paperhash":{"value":"makoviychuk|isaac_gym_high_performance_gpubased_physics_simulation_for_robot_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Viktor_Makoviychuk:","https://dblp.org/search/pid/api?q=author:Lukasz_Wawrzyniak:","~Yunrong_Guo1","https://dblp.org/search/pid/api?q=author:Michelle_Lu:","https://dblp.org/search/pid/api?q=author:Kier_Storey:","https://dblp.org/search/pid/api?q=author:Miles_Macklin:","https://dblp.org/search/pid/api?q=author:David_Hoeller:","https://dblp.org/search/pid/api?q=author:Nikita_Rudin:","https://dblp.org/search/pid/api?q=author:Arthur_Allshire:","~Ankur_Handa1","https://dblp.org/search/pid/api?q=author:Gavriel_State:"]},"html":{"value":"https://arxiv.org/abs/2108.10470"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2108-10470,\n  publtype={informal},\n  author={Viktor Makoviychuk and Lukasz Wawrzyniak and Yunrong Guo and Michelle Lu and Kier Storey and Miles Macklin and David Hoeller and Nikita Rudin and Arthur Allshire and Ankur Handa and Gavriel State},\n  title={Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning},\n  year={2021},\n  cdate={1609459200000},\n  journal={CoRR},\n  volume={abs/2108.10470},\n  url={https://arxiv.org/abs/2108.10470}\n}\n"},"abstract":{"value":"Isaac Gym offers a high performance learning platform to train policies for wide variety of robotics tasks directly on GPU. Both physics simulation and the neural network policy training reside on GPU and communicate by directly passing data from physics buffers to PyTorch tensors without ever going through any CPU bottlenecks. This leads to blazing fast training times for complex robotics tasks on a single GPU with 2-3 orders of magnitude improvements compared to conventional RL training that uses a CPU based simulator and GPU for neural networks. We host the results and videos at \\url{https://sites.google.com/view/isaacgym-nvidia} and isaac gym can be downloaded at \\url{https://developer.nvidia.com/isaac-gym}."},"title":{"value":"Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning"},"authors":{"value":["Viktor Makoviychuk","Lukasz Wawrzyniak","Yunrong Guo","Michelle Lu","Kier Storey","Miles Macklin","David Hoeller","Nikita Rudin","Arthur Allshire","Ankur Handa","Gavriel State"]}},"tmdate":1741236673656,"pdate":1609459200000,"tcdate":1731483778170,"writers":["~"],"signatures":["~Ankur_Handa1"],"forum":"C7pBDxtDUJ","license":"CC BY-SA 4.0","number":213702,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1741236673656,"domain":"DBLP.org","id":"C7pBDxtDUJ","version":2},{"content":{"venue":{"value":"Complex & Intelligent Systems"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s40747-025-01913-w.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"wang|hide_and_seek_in_transaction_networks_a_multiagent_framework_for_simulating_and_detecting_money_laundering_activities"},"html":{"value":"https://doi.org/10.1007/s40747-025-01913-w"},"abstract":{"value":"Detecting money laundering within financial networks presents a complex challenge due to the elusive behavior patterns of laundering agents, often resulting in data gaps. In this research, we propose a ‘Multiverse Simulation’ framework using a multi-agent system to generate synthetic datasets for anti-money laundering (AML) training and detection. This framework creates diverse virtual worlds, each with unique parameters to represent varying levels of illicit activity, thus mimicking the dynamics of money laundering and legitimate transactions. Our framework comprises two main types of agents: (1) the Detector, trained to identify laundering signs, and (2) Transaction agents, divided into those involved in laundering and those in legal transactions. These agents interact in a synthetic environment governed by rules that simulate real-world financial behaviors, enabling the generation of complex, realistic data. In the hide-and-seek multiverse simulation, the Detector learns to distinguish between licit and illicit transactions, a process refined by the evolving strategies of transaction agents to avoid detection. This adversarial setup fosters the co-evolution of laundering techniques and detection methods, enhancing system robustness. We demonstrate the efficacy of this approach by pre-training on synthetic cross-bank data, then evaluating with real-world data from the Elliptic dataset. Our results show that transfer learning significantly improves AML system performance, effectively bridging the gap between synthetic and authentic transaction patterns. The ‘Multiverse Simulation’ offers a scalable, dynamic approach to better understand and mitigate the gap between simulation and reality, contributing to more resilient and intelligent AML solutions."},"title":{"value":"Hide and seek in transaction networks: a multi-agent framework for simulating and detecting money laundering activities"},"authors":{"value":[{"fullname":"Qianyu Wang","username":"~Qianyu_Wang2"},{"fullname":"Wei-Tek Tsai","username":"https://orcid.org/orcid-search/search?searchQuery=Wei-Tek%20Tsai"},{"fullname":"Tianyu Shi","username":"https://orcid.org/orcid-search/search?searchQuery=Tianyu%20Shi"},{"fullname":"Wang Tang","username":"https://orcid.org/orcid-search/search?searchQuery=Wang%20Tang"},{"fullname":"Bowen Du","username":"https://orcid.org/orcid-search/search?searchQuery=Bowen%20Du"}]}},"tmdate":1784555451433,"pdate":1748736000000,"externalIds":["doi:10.1007/s40747-025-01913-w"],"tcdate":1784555447434,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Qianyu_Wang3"],"forum":"0Pd1BIZ7qa","license":"CC BY-SA 4.0","number":84238,"cdate":1746513264006,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784555451433,"domain":"OpenReview.net/Public_Article","id":"0Pd1BIZ7qa","version":2},{"content":{"summary":{"value":"This paper tackles the problem of generating physically plausible videos using text-to-video diffusion models The proposed method, DiffPhy, fine-tunes a pretrained video diffusion model with guidance from LLMs and Multimodal LLMs. LLM first infers physical context from the text prompt, and an MLLM evaluates the generated video against these physical rules. The MLLM’s feedback is converted into a continuous differentiable score that supervises the diffusion model, encouraging alignment with physical laws. A failure-aware refinement module further injects attention based on detected physical inconsistencies. Experiments on physics-oriented benchmarks show improved physical plausibility and semantic alignment compared to existing models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See the weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper proposes a reasonable integration of textual reasoning and multimodal verification to encourage physically consistent video generation, extending prior work on language-guided diffusion with an additional physics-aware supervision signal. The proposed continuous score estimator provides a practical way to incorporate non-differentiable LVLM feedback into gradient-based fine-tuning. The failure-aware refinement further introduces an interesting idea of using detected failure cases as textual cues to guide attention."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1) Incomplete Related Work & Baseline Coverage**  \nWhile the “Video Physics Reasoning” subsection is conceptually relevant, the cited works mainly focus on traditional physics simulation or early neural reasoning models, rather than recent diffusion-based approaches that explicitly address physical consistency in video generation. Incorporating more recent studies such as **PhyT2V [1]** and **Yang et al. [2]** would strengthen the connection between DiffPhy and the current landscape of physics-aware video diffusion. Also, the baseline comparison could be expanded to include these contemporary methods.\n\n**2)  Insufficient Comparison with Alignment-Based Methods**  \nThe paper does not adequately position its approach within the growing literature on alignment-based fine-tuning frameworks that leverage LVLM feedback (or more broadly, reward-model signals) for diffusion model adaptation (e.g., VADER [3]). While the proposed continuous score estimator provides a differentiable way to incorporate feedback, it remains unclear how this approach compares to or improves upon existing reward-based and preference-alignment methods in terms of training stability, supervision efficiency, or performance. A more systematic discussion or experimental comparison would help clarify the method’s advantages and situate it more clearly within this line of work.\n\n**3) Uncertain Effectiveness of Failure-Aware Refinement**  \nThe proposed failure-aware refinement introduces an interesting idea of using MLLM-identified failure cases as textual feedback to guide attention. However, this mechanism ultimately relies on the possibility that conditioning on such failure cues helps the model focus on physically inconsistent regions, rather than providing any guarantee of actual correction. Since the feedback is binary and lacks spatial-temporal grounding, it is unclear whether this refinement truly fixes the underlying physical inconsistencies or simply injects additional noise. Moreover, the paper does not include an ablation isolating this component during training, making it even more difficult to assess its actual impact. It would also be helpful to include inference-time ablations comparing this refinement against simpler alternatives, such as seed resampling, to demonstrate that the proposed strategy offers tangible advantages beyond random variation.\n\n**4) Citation Format**  \nCitation references are not consistently enclosed in parentheses throughout the paper.\n\n[1] *PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation*  \n[2] *Towards Physically Plausible Video Generation via VLM Planning*  \n[3] *Video Diffusion Alignment via Reward Gradients*"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918558642,"tcdate":1761858258103,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6229/Reviewer_pvJs"],"signatures":["ICLR.cc/2026/Conference/Submission6229/Reviewer_pvJs"],"forum":"lPKsPBstHg","number":3,"license":"CC BY 4.0","cdate":1761858258103,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6229/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918558642,"domain":"ICLR.cc/2026/Conference","replyto":"lPKsPBstHg","id":"tTq6QI3GLt","forumContent":{"TLDR":{"value":"We propose DiffPhy, a generic framework that enables physically-correct and semantically coherent video generation by fine-tuning a pre-trained video diffusion model."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Diffusion","Video Generation","Physical Commensense"]},"supplementary_material":{"value":"/attachment/dd73e8deab01feb18bdae06f4f6ecc4dca24d57c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions, interactions, and dynamics introduce great difficulties when learning physics from data. In this work, we propose DiffPhy, a generic framework that enables physically-correct and photo-realistic video generation by fine-tuning a pre-trained video diffusion model. Our method leverages large language models (LLMs) to infer rich physical context from the text prompt. To incorporate this context into the video diffusion model, we use a multimodal large language model (MLLM) to verify intermediate latent variables against the inferred physical rules, guiding the model’s gradient updates accordingly. MLLM’s textual output is transformed into continuous signals. We then formulate a set of training objectives that jointly ensure physical accuracy and semantic alignment with the input text.  Additionally, failure facts of physical phenomena are corrected via attention injection. We also establish a high-quality physical video dataset containing diverse phyiscal actions and events to facilitate effective finetuning. Extensive experiments on public benchmarks demonstrate that DiffPhy is able to produce state-of-the-art results across diverse physics-related scenarios. Code and data will be made available post-review."},"_bibtex":{"value":"@misc{\nzhang2026think,\ntitle={Think Before You Diffuse: Infusing Physical Rules into Video Diffusion},\nauthor={Ke Zhang and Cihan Xiao and Jiacong Xu and Yiqun Mei and Vishal M. Patel},\nyear={2026},\nurl={https://openreview.net/forum?id=lPKsPBstHg}\n}"},"title":{"value":"Think Before You Diffuse: Infusing Physical Rules into Video Diffusion"},"pdf":{"value":"/pdf/abed66725f92f4403a0e6585ebdac7432376d20d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|think_before_you_diffuse_infusing_physical_rules_into_video_diffusion"},"authorids":{"value":["~Ke_Zhang17","~Cihan_Xiao1","~Jiacong_Xu1","~Yiqun_Mei1","~Vishal_M._Patel1"]},"authors":{"value":["Ke Zhang","Cihan Xiao","Jiacong Xu","Yiqun Mei","Vishal M. Patel"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a masked pre-training technique for Graph Neural Networks (GNNs) aimed at solving physics simulations, particularly in computational fluid dynamics. The proposed approach uses an asymmetric encoder-decoder architecture in conjunction with gated multi-layer perceptrons. Extensive experiments are conducted, including independent training on multiple datasets, transfer learning, and multi-dataset pretraining. The results demonstrate that the MeshMask model delivers competitive performance compared to established baselines."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Given that your input mesh includes masked data, how do you effectively learn long-range interactions on the masked graph during the repeated message passing?\n2. In Sec 4.3,  given that the input and output physics quantities differ across the datasets, the potential for a significant gap in the encoded latent space raises concerns about the legality and effectiveness of the transfer learning approach. How do you ensure that the latent spaces of the two datasets are compatible? Secondly, the length of input and output quantities varies across datasets; how do you exactly fine-tune the model when faced with these differing lengths?\n3. What is the model's performance on out-of-distribution (OOD) mesh resolutions?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper introduces a novel approach by implementing masked pre-training for Graph Neural Networks (GNNs) in physics simulations. It is well-structured and easy to follow, with rich qualitative results and visualizations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The introduction of gated MLP and encoder-decoder architecture potentially increases the model's computational complexity. There is a lack of detailed discussion on the computational demands (e.g. training/inference time,  training/inference RAMs) compared with baselines, which is critical for evaluating the feasibility of deploying MeshMask in real-time applications or on large-scale datasets.\n2. The compared baselines are limited, authors should consider more GNN-based simulation models such as [1-4]. Furthermore, both GCN and U-Net baselines have only been tested on a few datasets.\n3. The novelty of replacing standard MLPs with gated MLPs in the processor is somewhat limited. The authors should consider a wider variety of network architectures such as attention-based models [1,2,4] for comparison to make their findings more persuasive.\n\n[1] Eagle: Large-scale learning of turbulent fluid dynamics with mesh transformers.\n\n[2] Learning flexible body collision dynamics with hierarchical contact mesh transformer.\n\n[3] Efficient Learning of Mesh-Based Physical Simulation with BSMS-GNN.\n\n[4] Transformer with Implicit Edges for Particle-based Physics Simulation."}},"nonreaders":[],"tmdate":1732512938393,"tcdate":1730276977554,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3755/Reviewer_i9XS"],"signatures":["ICLR.cc/2025/Conference/Submission3755/Reviewer_i9XS"],"forum":"bFHR8hNk4I","number":2,"license":"CC BY 4.0","cdate":1730276977554,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3755/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732512938393,"domain":"ICLR.cc/2025/Conference","replyto":"bFHR8hNk4I","id":"gCaphkFuSe","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"We introduce a novel masked pre-training technique for graph neural networks (GNNs) applied to computational fluid dynamics (CFD) problems"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["graph networks","simulation","mesh","physics"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We introduce a novel masked pre-training technique for graph neural networks (GNNs) applied to computational fluid dynamics (CFD) problems. By randomly masking up to 40\\% of input mesh nodes during pre-training, we force the model to learn robust representations of complex fluid dynamics. We pair this masking strategy with an asymmetric encoder-decoder architecture and gated multi-layer perceptrons to further enhance performance. The proposed method achieves state-of-the-art results on seven CFD datasets, including a new challenging dataset of 3D intracranial aneurysm simulations with over 250,000 nodes per mesh. Moreover, it significantly improves model performance and training efficiency across such diverse range of fluid simulation tasks. We demonstrate improvements of up to 60\\% in long-term prediction accuracy compared to previous best models, while maintaining similar computational costs. Notably, our approach enables effective pre-training on multiple datasets simultaneously, significantly reducing the time and data required to achieve high performance on new tasks.\nThrough extensive ablation studies, we provide insights into the optimal masking ratio, architectural choices, and training strategies."},"_bibtex":{"value":"@inproceedings{\ngarnier2025meshmask,\ntitle={MeshMask: Physics-Based Simulations with Masked Graph Neural Networks},\nauthor={Paul Garnier and Vincent Lannelongue and Jonathan Viquerat and Elie Hachem},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=bFHR8hNk4I}\n}"},"title":{"value":"MeshMask: Physics-Based Simulations with Masked Graph Neural Networks"},"pdf":{"value":"/pdf/50c39da508640afcf6bf4735a820fe773bd016f2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"garnier|meshmask_physicsbased_simulations_with_masked_graph_neural_networks"},"authorids":{"value":["~Paul_Garnier1","~Vincent_Lannelongue1","~Jonathan_Viquerat1","~Elie_Hachem1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Paul Garnier","Vincent Lannelongue","Jonathan Viquerat","Elie Hachem"]}},"version":2},{"content":{"summary":{"value":"The paper introduces GVFi, a framework for modeling the motion physics of complex dynamic 3D scenes using multi-view RGB videos without requiring additional annotations such as object shapes, types, or masks.\nBuilding on Deformable3DGS, GVFi incorporates constraints based on the laws of classical mechanics to guide motion predictions, ensuring that the Gaussian deformation estimated by the MLP aligns more closely with physical principles. By assuming that motion adheres to the laws of classical mechanics and explicitly learning the associated motion parameters, GVFi is capable of performing effective extrapolation rendering, allowing it to predict frames beyond the observed time span. Experimental results show that GVFi significantly outperforms existing methods, particularly excelling in future frame extrapolation tasks."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. Based on the methodology, there seem to be three possible approaches for interpolation rendering: (1) directly using $f_{defo}$ to predict the deformation at the given time $t$, (2) progressively calculating the Gaussian deformation at the given time $t$ from time 0 using the motion parameters predicted by $f_{trd}$, or (3) following the steps described in lines L261-L269. Which approach was used in the experiments? Are the results consistent across these three methods?\n2. For extrapolation rendering according to lines L261-L269, it seems feasible to use either the second or third approach from question 1. Which method was actually used by the authors? If the third approach was used, how does it perform over longer extrapolation periods? Could the authors provide visual results for extrapolations that extend beyond the time span covered in the dataset?\n3. The choice of baseline methods for comparison appears limited. For a comprehensive evaluation, it would be beneficial to compare against state-of-the-art methods in dynamic scene reconstruction, such as 4D-GS[2] and more recent work like E-D3DGS [3], which both have architectures similar to Deformable3DGS but differ in their motion representation. Could the authors verify if the proposed Translation Rotation Dynamics System can be integrated into these methods and whether it would yield similar performance gains?\n4. The authors claim that their framework is a general approach for modeling motion physics in complex dynamic 3D scenes. However, the datasets used, with only 60 frames in total, limit the complexity and extent of motion. Could the authors validate this claim by testing on more challenging synthetic and real-world datasets, such as the ParticleNeRF and PanopticSports datasets, to provide a more comprehensive evaluation of the framework’s effectiveness on complex scenes?\n5. In the ablation study, the authors provide a rationale for their choice of $\\delta t$, which is somewhat reasonable. However, this conclusion is based on results from only one dataset, which may not be sufficient, as each dataset could exhibit different motion characteristics. Could the authors clarify how to select an appropriate $\\delta t$ in practice across diverse datasets?\n6. The experimental details are insufficient, particularly regarding training time, required resources, storage size, and rendering speed. Could the authors provide more comprehensive information on these aspects?\n7. Please ensure that all abbreviations and technical terms are clearly defined, with full explanations and necessary citations. In the related work section, it would be helpful to explicitly clarify the differences from relevant works wherever possible.\n\n[2] Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.\n\n[3] Jeongmin Bae, Seoha Kim, Youngsik Yun, Hahyun Lee, Gun Bang, and Youngjung Uh. Per- gaussian embedding-based deformation for deformable 3d gaussian splatting. In Proceedings of the European Conference on Computer Vision (ECCV), 2024."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Modeling the motion of Gaussians through a Translation Rotation Dynamics System grounded in classical mechanics, resulting in a concise and conceptually elegant framework with solid mathematical and physical foundations.\n2. Introducing an effective method to train the motion parameters of the Translation Rotation Dynamics System, enabling the accurate estimation of translation and rotation dynamics for each particle in the scene.\n4. By explicitly learning motion parameters under classical mechanics, enabling effective extrapolation to unobserved frames and presenting potential for generation tasks that require plausible future frames in dynamic 3D scenes.\n3. The proposed approach is validated on two tasks, demonstrating superior performance compared to previous methods, highlighting its effectiveness in modeling motion dynamics in 3D scenes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The contributions of this work are somewhat incremental, as most of the methodological design heavily overlaps with the baseline method, Deformable3DGS [1]. The key difference lies in the incorporation of dynamical principles, primarily to enable extrapolation capabilities rather than introducing fundamentally novel approaches.\n2. The proposed motion modeling framework is overly restrictive, relying on an strong assumption of no external forces, disregarding energy transfer processes, and lacking the ability to handle non-rigid or nonlinear motion. These limitations significantly reduce the model’s applicability to real-world physics.\n3. Due to its reliance on idealized assumptions and limited scope, the model struggles to handle complex, real-world motion dynamics where varied forces, interactions, and non-rigid behaviors are prevalent, limiting its utility for practical applications in diverse environments.\n\n[1] Ziyi Yang, Xinyu Gao, Wen Zhou, Shaohui Jiao, Yuqing Zhang, and Xiaogang Jin. Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction. CVPR, 2024."}},"nonreaders":[],"tmdate":1732684266006,"tcdate":1730544569098,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1874/Reviewer_3rJ7"],"signatures":["ICLR.cc/2025/Conference/Submission1874/Reviewer_3rJ7"],"forum":"0Zot73kfLB","number":2,"license":"CC BY 4.0","cdate":1730544569098,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1874/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732684266006,"domain":"ICLR.cc/2025/Conference","replyto":"0Zot73kfLB","id":"rx84bWvvo3","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Dynamic Reconstruction","Physics","Motion Extrapolation"]},"supplementary_material":{"value":"/attachment/ba93674e0fa0a49347edf567ec9996ff62c59526.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this paper, we aim to model 3D scene geometry, appearance, and physical information just from dynamic multi-view videos in the absence of any human labels. By leveraging physics-informed losses as soft constraints or integrating simple physics models into neural networks, existing works often fail to learn complex motion physics, or doing so requires additional labels such as object types or masks. In this paper, we propose a new framework named **GVFi** to model the motion physics of complex dynamic 3D scenes. The key novelty of our approach is that, by formulating each 3D point as a rigid particle with size and orientation in space, we choose to directly learn a translation rotation dynamics system for each particle, explicitly estimating a complete set of physical parameters to govern the particle's motion over time. Extensive experiments on three existing dynamic datasets and two newly created challenging synthetic and real-world datasets demonstrate the extraordinary performance of our method over baselines in the task of future frame extrapolation. A nice property of our framework is that multiple objects or parts can be easily segmented just by clustering the learned physical parameters. Our datasets and code will be released at https://github.com/"},"_bibtex":{"value":"@misc{\nli2025gvfi,\ntitle={{GVF}i: Learning 3D Gaussian Velocity Fields from Dynamic Videos},\nauthor={Jinxi Li and Ziyang Song and Bo Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=0Zot73kfLB}\n}"},"title":{"value":"GVFi: Learning 3D Gaussian Velocity Fields from Dynamic Videos"},"pdf":{"value":"/pdf/b21a9a361227cef9dd98f4ed14b46af082aedb45.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|gvfi_learning_3d_gaussian_velocity_fields_from_dynamic_videos"},"authorids":{"value":["~Jinxi_Li2","~Ziyang_Song1","~Bo_Yang7"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jinxi Li","Ziyang Song","Bo Yang"]}},"version":2},{"content":{"summary":{"value":"MARS presents a dual-system multi-agent RL framework that unifies intuitive (System 1) and deliberate (System 2) reasoning within an LLM, jointly optimized via GRPO to improve deep research and reasoning performance across complex tasks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please refer to the weaknesses section for main questions."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper is clearly written and well-structured\n\n2. Proposes a dual-system multi-agent RL framework that explicitly models human-like System 1/System 2 reasoning, an interesting conceptual extension of existing multi-agent paradigms.\n\n3. Demonstrates measurable gains on challenging reasoning benchmarks"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed dual-system framework essentially resembles a standard RL-based tool-use pipeline augmented with a learnable summarizer that condenses the environment’s returned content before feeding it back. While the integration is well-engineered, the conceptual difference from existing RL tool-use or summarization-based reasoning systems is limited.\n\n2. Because the entire trajectory shares a single scalar reward, it is unclear how meaningful credit is assigned to System 1’s summarization behavior. Without step-level or component-wise feedback, System 1 receives only a weak and noisy learning signal, making it difficult to understand how it learns to produce more useful summaries. The paper could be strengthened by introducing more fine-grained supervision or ablation analyses that clarify how System 1’s updates contribute to overall improvement.\n\n3. In Table 1, several entries marked as best (bold) and second-best (underlined) appear to be incorrect. This is misleading to readers. \n\n4. The ablation study mainly analyzes the impact of removing different external tools (Google Search, Scholar, Python), but this aspect is peripheral to the paper’s main contribution. Since the core claim of MARS lies in the joint optimization and coordination between System 1 and System 2, the paper would benefit much more from ablations that directly test this interaction—for example, mixing trained and untrained versions of System 1/2, or disabling their shared optimization to assess whether the two systems truly co-adapt. \n\n5. Even so, I find the results in the Ablation Study on Tools for HLE rather confusing. For example, in Chem, the setup with all three tools performs the worst, while both without Scholar and Scholar-only achieve the best results — which makes it unclear whether Scholar is actually helpful or not; Also both without Search and Search-only achieve the Second. Similarly, in CS/AI, the best setup is without Search, yet Search-only also performs noticeably higher than most others; and in Engineering, Python-only gives the highest score, but without Python ranks second. Overall, the patterns seem quite inconsistent or even random. Given how close these numbers are, I wonder whether you ran multiple inference trials and averaged the results. The apparent randomness in this table makes it hard to trust the conclusions on HLE.\n\n6. Could you clarify whether the comparison between MARS and the baselines is fully fair in terms of tool usage? Specifically, do all methods have equal access to the same tools (Python, Search, and Scholar)? The results suggest that the presence or absence of certain tools has a large impact on performance, and in most subjects, removing a specific tool even makes MARS perform worse than most baselines. This raises concerns about whether the comparison setup is fully fair. It would be important to provide more details on the tool configurations for all baselines and ensure that all methods are evaluated under comparable conditions. Moreover, additional ablations are needed to justify that the reported gains on HLE truly come from the proposed dual-system RL framework, rather than differences in tool availability or usage."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926380776,"tcdate":1762072588247,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16220/Reviewer_f7bk"],"signatures":["ICLR.cc/2026/Conference/Submission16220/Reviewer_f7bk"],"forum":"abxVxyXNhW","number":3,"license":"CC BY 4.0","cdate":1762072588247,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16220/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926380776,"domain":"ICLR.cc/2026/Conference","replyto":"abxVxyXNhW","id":"OJMXH0s21N","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Deep Research"]},"supplementary_material":{"value":"/attachment/bef47cc9afe741af1ce80b8088ff968e7777c110.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Large Reasoning Models (LRMs) often exhibit a tendency for overanalysis in simple tasks, where the models excessively utilize System 2-type, deliberate reasoning, leading to inefficient token generation. \nFurthermore, these models face challenges in adapting their reasoning capabilities to rapidly changing environments due to the static nature of their pre-training data. \nTo address these issues, advancing Large Language Models (LLMs) for complex reasoning tasks requires innovative approaches that bridge intuitive and deliberate cognitive processes, akin to human cognition's dual-system dynamic. \nThis paper introduces a Multi-Agent System for Deep ReSearch (MARS) enabling seamless integration of System 1’s fast, intuitive thinking with System 2’s deliberate reasoning within LLMs. \nMARS strategically integrates multiple external tools—such as Google Search, Google Scholar, and Python Interpreter—to access up-to-date information and execute complex computations, while creating a specialized division of labor where System 1 efficiently processes and summarizes high-volume external information, providing distilled insights that expand System 2's reasoning context without overwhelming its capacity.\nFurthermore, we propose a multi-agent reinforcement learning framework extending Group Relative Policy Optimization to simultaneously optimize both systems with multi-turn tool interactions, bin-packing optimization, and sample balancing strategies that enhance collaborative efficiency.\nExtensive experiments demonstrate MARS achieves substantial improvements of 3.86\\% on the challenging Humanity's Last Exam (HLE) benchmark and an average gain of 8.9\\% across 7 knowledge-intensive tasks, validating the effectiveness of our dual-system paradigm for complex reasoning in dynamic information environments."},"_bibtex":{"value":"@misc{\nchen2026mars,\ntitle={{MARS}: Optimizing Dual-System Deep Research via Multi-Agent Reinforcement Learning},\nauthor={Guoxin Chen and Zile Qiao and Wenqing Wang and Donglei Yu and Xuanzhong Chen and Hao Sun and Minpeng Liao and Kai Fan and Yong Jiang and Pengjun Xie and Xin Zhao and Ruihua Song and Fei Huang},\nyear={2026},\nurl={https://openreview.net/forum?id=abxVxyXNhW}\n}"},"title":{"value":"MARS: Optimizing Dual-System Deep Research via Multi-Agent Reinforcement Learning"},"pdf":{"value":"/pdf/a61a6ccbaf829b85d64fd8586c6977e3169af9c6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"chen|mars_optimizing_dualsystem_deep_research_via_multiagent_reinforcement_learning"},"authorids":{"value":["~Guoxin_Chen1","~Zile_Qiao1","~Wenqing_Wang3","~Donglei_Yu1","~Xuanzhong_Chen1","~Hao_Sun9","~Minpeng_Liao1","~Kai_Fan1","~Yong_Jiang1","~Pengjun_Xie2","~Xin_Zhao10","~Ruihua_Song1","~Fei_Huang2"]},"authors":{"value":["Guoxin Chen","Zile Qiao","Wenqing Wang","Donglei Yu","Xuanzhong Chen","Hao Sun","Minpeng Liao","Kai Fan","Yong Jiang","Pengjun Xie","Xin Zhao","Ruihua Song","Fei Huang"]}},"version":2},{"content":{"summary":{"value":"The paper studies when synthetic data from diffusion models falls short when the target data distribution has a large distribution shift from the model's pretraining data distribution. Such distributional shift is specifically significant for satellite images which are rare and causes more significant noise misalignment in diffusion model's training and inference stages. Two methods are proposed to mitigate the misalignment and enhance the quality of synthetic satellite images."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"Please see above. My main concern is about empirical results validating the connection between noise misalignment and data distributional shift, as well as the originality of the proposed methods. I'd be open to adjust my score given the author's response."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is generally clearly written.\n\n2. The paper aims to analyze and establish the correlation between pretraining/target data distribution shift and the effectiveness of using corresponding synthetic data in training downstream models. Empirical results support this intuition.\n\n3. The proposed methods mitigate the noise misalignment problem. Qualitatively, the satellite images generated has higher quality (more faithful color) and yields better performances when used to train downstream models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Experiments in Section 3 does support the claim that larger distribution shift in training/target data distribution correlates with the effectiveness of using synthetic data in training downstream models. However, there seems to lack experimental supports for why noise misalignment may be more severe for data with larger distributional shift. Specifically, is there empirical results to validate the hypothesis mentioned in line 230-240?\n\n2. It seems to me that the offset noise method is a method that was used in prior work for slightly different purposes. It is thus not clear to me the technical contribution of this work. Could the authors clarify how the proposed methods different from existing techniques?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918308020,"tcdate":1762154423514,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5860/Reviewer_U8FA"],"signatures":["ICLR.cc/2026/Conference/Submission5860/Reviewer_U8FA"],"forum":"Xw0SudV0UB","number":4,"license":"CC BY 4.0","cdate":1762154423514,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5860/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918308020,"domain":"ICLR.cc/2026/Conference","replyto":"Xw0SudV0UB","id":"pSCydED057","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Latent Diffusion","Synthetic data","Text-to-image generation","Satellite Imagery"]},"primary_area":{"value":"generative models"},"abstract":{"value":"While satellite data is essential for applying computer vision to many real-world tasks, it remains expensive to acquire. Although other computer vision tasks have alleviated data procurement costs by augmenting training datasets with synthetic images from text-to-image models, such augmentation remains underdeveloped in the remote sensing domain. In this work, we propose an alternative approach for generating synthetic training data tailored to satellite imagery. To better understand the underlying problem, we begin by analyzing the impact of the target data distribution in comparison to the distributions used to train the text-to-image generation model. We find that data rarity is strongly correlated with the effectiveness of synthetic training data produced by Stable Diffusion fine-tuned on few-shot examples, suggesting that rarity can serve as a low-cost proxy for pre-evaluating the effectiveness of synthetic data generation. Notably, our analysis shows that Stable Diffusion struggles to produce useful training images for rare, out-of-distribution data. Building on this insight, we propose two modifications to the generation process tailored to satellite images: offset noise and leak-aligned noise. Both are designed to adjust the initial noise distribution and correct low-frequency characteristics. Our approaches enable improved training performance for classifiers trained on synthetic data, demonstrated on three satellite benchmarks."},"_bibtex":{"value":"@misc{\nhua2026aligning,\ntitle={Aligning Signal Leakage Matters for Synthetic Data Generation of Satellite Imagery},\nauthor={Ying Hua and Jessica Bader and Jae Myung Kim and Zeynep Akata},\nyear={2026},\nurl={https://openreview.net/forum?id=Xw0SudV0UB}\n}"},"title":{"value":"Aligning Signal Leakage Matters for Synthetic Data Generation of Satellite Imagery"},"pdf":{"value":"/pdf/d53e3fa4f61bc2b3d62a42efdbbc7c9ef7b81fc4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hua|aligning_signal_leakage_matters_for_synthetic_data_generation_of_satellite_imagery"},"authorids":{"value":["~Ying_Hua1","~Jessica_Bader1","~Jae_Myung_Kim1","~Zeynep_Akata1"]},"authors":{"value":["Ying Hua","Jessica Bader","Jae Myung Kim","Zeynep Akata"]}},"version":2},{"content":{"summary":{"value":"The paper presented a comprehensive empirical study of knowledge distillation in the post-training stage of LLMs. The authors conducted experiments on the Tulu 3 dataset and found that KD consistently outperforms SFT in low-data regimes. Moreover, in specific scenarios where high-quality datasets are scarce, the experiments demonstrate that KD provides a clear advantage for smaller student models. Finally, to address the shortage of human-labeled data, the authors propose a two-stage KD paradigm that first pre-trains the student using synthetic teacher-labeled data, followed by refinement with human-annotated data."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. Could the authors provide more discussion on the dataset scale at which the proposed two-stage distillation method is most effective?\n\n2. What is the appropriate ratio between human-labeled data and synthetic data that yields practical improvements in distillation performance?\n\n3. During synthetic data generation, did the authors consider data diversity or use any specific metrics to evaluate the quality of the generated samples?\n\n4. Could the authors include an ablation study to assess the actual contribution of each stage in the distillation process? Additionally, would reversing the order—training first on human data and then on synthetic data—affect the results?\n\n5. Can the authors compare their method with a broader range of existing distillation approaches to strengthen the experimental validation?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The authors conducted extensive and thorough experiments to investigate the behavior of knowledge distillation in the post-training stage of LLMs, examining the effects of teacher model strength, student training paradigm (SFT vs. KD), dataset size, and the use of synthetic data. The experiments are comprehensive and rigorously designed.\n\n2. The paper is well-structured and logically progressive. The authors present their analyses and reflections on the experimental phenomena with clarity and sound reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper primarily provides empirical conclusions based on observed experimental phenomena, but lacks deeper theoretical analysis and discussion to support these findings.\n\n2. Although the authors propose a two-stage distillation method, it appears to have limited novelty — essentially pretraining on synthetic data followed by distillation on high-quality human-labeled data. This approach seems tailored for scenarios with limited human annotations, but the paper would benefit from a more detailed discussion on the applicable data scale for this method. In addition, it would be valuable to analyze the optimal ratio between human-labeled and synthetic data that leads to practical improvements in distillation. The authors could also elaborate on whether they considered the diversity or quality metrics of the synthetic data during generation. Furthermore, including an ablation study to evaluate the contribution of each stage would strengthen the experimental section. It would also be interesting to explore whether reversing the order—training first on human data and then on synthetic data—would affect the results.\n\n3. The comparison methods in the experimental section are relatively limited. It would enhance the paper’s completeness to include comparisons with more recent or diverse distillation approaches."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942478623,"tcdate":1762440272465,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23022/Reviewer_JQ7A"],"signatures":["ICLR.cc/2026/Conference/Submission23022/Reviewer_JQ7A"],"forum":"DvrZRcoT6p","number":4,"license":"CC BY 4.0","cdate":1762440272465,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23022/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942478623,"domain":"ICLR.cc/2026/Conference","replyto":"DvrZRcoT6p","id":"BWlY1FCfRp","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Knowledge Distillation","Large Language Model","Post-Training"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Large language models (LLMs) achieve strong performance across many tasks, but their high computational cost limits deployment in resource-constrained environments. Knowledge Distillation (KD) offers a practical solution by transferring knowledge from a teacher model of a larger size to a smaller student model. While prior work has mainly examined task-specific or small-scale settings, the post-training stage for building general instruction-following models has received limited attention.\nIn this paper, we conduct a systematic study of KD in post-training using the large-scale Tulu 3 dataset. We find that KD outperforms supervised fine-tuning (SFT) in low-data regimes, but its advantage diminishes as more training data is added.\nDistilling from a stronger instruction-tuned teacher restores substantial gains even with abundant data, \nindicating that KD remains effective when the teacher provides knowledge that the student cannot easily acquire from the training data alone.\nWe further study domain-specific, low-resource scenarios and propose a two-stage KD strategy that leverages synthetic teacher-labeled data followed by refinement on human annotations. This method consistently improves student performance, providing practical guidance for building compact models in data-scarce environments."},"_bibtex":{"value":"@misc{\nliu2025understanding,\ntitle={Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails},\nauthor={Xin Liu and Simin Ma and Shujian Liu and Song Wang and Sathish Reddy Indurthi and Haoyun Deng and Lu Wang and Kaiqiang Song},\nyear={2025},\nurl={https://openreview.net/forum?id=DvrZRcoT6p}\n}"},"title":{"value":"Understanding Knowledge Distillation in Post-Training: When It Helps and When It Fails"},"pdf":{"value":"/pdf/d2f8e67e62278e99341bd7cf0ec72d9ca8e22fdd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|understanding_knowledge_distillation_in_posttraining_when_it_helps_and_when_it_fails"},"authorids":{"value":["~Xin_Liu18","~Simin_Ma1","~Shujian_Liu1","~Song_Wang10","~Sathish_Reddy_Indurthi2","~Haoyun_Deng1","~Lu_Wang9","~Kaiqiang_Song2"]},"authors":{"value":["Xin Liu","Simin Ma","Shujian Liu","Song Wang","Sathish Reddy Indurthi","Haoyun Deng","Lu Wang","Kaiqiang Song"]}},"version":2},{"content":{"summary":{"value":"This paper investigates the negative impact of synthetic data contamination on existing online continual learning methods. An entropy selection with the real-synthetic similarity maximization method is proposed to alleviate the performance deterioration."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Detailed analysis of synthetic data contamination and its influence on continual learning.\n2. This paper is technically clear and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The creation of simulated data is the cornerstone of this research, but there is a lack of detailed explanation on how to generate these data in the main paper.\n2. For Observation 4 \"With the limited diversity of synthetic data\", how about mixing the synthetic data from different generation models to increase the diversity, since the synthetic images on the internet also form different models? While the observation remains unchanged in this case?\n3. In order to simulate real data collection more realistically, in addition to synthetic data, new open-domain real data should also be incorporated."},"limitations":{"value":"The limitations are discussed while the potential negative societal impact is not discussed. However, for this work, I think it is not necessary to discuss this."}},"nonreaders":[],"tmdate":1730878981231,"tcdate":1721036253185,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission5166/Reviewer_ttuB"],"signatures":["NeurIPS.cc/2024/Conference/Submission5166/Reviewer_ttuB"],"forum":"Lc8gemv97Y","number":3,"license":"CC BY 4.0","cdate":1721036253185,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission5166/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878981231,"domain":"NeurIPS.cc/2024/Conference","replyto":"Lc8gemv97Y","id":"n7dlsw0Chs","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"Investigating the dataset contamination caused by synthetic data in Online Continual Learning, and proposing a method to alleviate the performance degradation with entropy selection and real-synthetic similarity maximization."},"keywords":{"value":["Online Continual Learning","Image Generation","Replay-based method","Entropy Selection"]},"supplementary_material":{"value":"/attachment/1f7335eb5371418c8bc715b2c98b548a425b3b5b.zip"},"primary_area":{"value":"online_learning"},"abstract":{"value":"Image generation has shown remarkable results in generating high-fidelity realistic images, in particular with the advancement of diffusion-based models. However, the prevalence of AI-generated images may have side effects for the machine learning community that are not clearly identified. Meanwhile, the success of deep learning in computer vision is driven by the massive dataset collected on the Internet. The extensive quantity of synthetic data being added to the Internet would become an obstacle for future researchers to collect \"clean\" datasets without AI-generated content. Prior research has shown that using datasets contaminated by synthetic images may result in performance degradation when used for training. In this paper, we investigate the potential impact of contaminated datasets on Online Continual Learning (CL) research. We experimentally show that contaminated datasets might hinder the training of existing online CL methods. Also, we propose Entropy Selection with Real-synthetic similarity Maximization (ESRM), a method to alleviate the performance deterioration caused by synthetic images when training online CL models. Experiments show that our method can significantly alleviate performance deterioration, especially when the contamination is severe. For reproducibility, the source code of our work is available at https://github.com/maorong-wang/ESRM."},"_bibtex":{"value":"@inproceedings{\nwang2024dealing,\ntitle={Dealing with Synthetic Data Contamination in Online Continual Learning},\nauthor={Maorong Wang and Nicolas Michel and Jiafeng Mao and Toshihiko Yamasaki},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=Lc8gemv97Y}\n}"},"title":{"value":"Dealing with Synthetic Data Contamination in Online Continual Learning"},"pdf":{"value":"/pdf/eeeaa1a535b3be8d4d63b515ebc76e0021839555.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"wang|dealing_with_synthetic_data_contamination_in_online_continual_learning"},"authorids":{"value":["~Maorong_Wang1","~Nicolas_Michel1","~Jiafeng_Mao1","~Toshihiko_Yamasaki1"]},"authors":{"value":["Maorong Wang","Nicolas Michel","Jiafeng Mao","Toshihiko Yamasaki"]}},"version":2},{"content":{"summary":{"value":"This manuscript aims to solve the problem of super-resolution of physical fields. Specifically, they draw on the framework in [1]. The difference lies in 1) weighting the denoising loss over pixels using the wavelet transform and 2) performing physical loss gradient descent over clean predictions.\n\n[1] A physics-informed diffusion model for high-fidelity flow field reconstruction"},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"see Weaknesses"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Authors present four turbulence datasets with different characteristics.\n\n2. I agree with one of the author's points in the introductory section. They note that some of the current work focuses on low-resolution data that is downsampled. This is not consistent with reality, and in fact, to apply these super-resolution models, such data should come from the low-resolution simulation of a PDE solver, which typically has lower fidelity than downsampled ones.\n\n2. The amount of experiments is rich enough, and all the graphs are clear."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. I have some questions about the authors' contributions to modeling. First is the training phase. The authors use the wavelet transform to determine which pixels need to be weighted. As can be seen from the visualization in Figure 10, the weighting seems to be added only at high frequencies. This raises the question of whether a complex wavelet transform needs to be introduced. This is because we can simply call the Laplace operator to achieve this effect. Furthermore, from the quantitative results in Table 1, it seems that this measure does not contribute much to the results.\n\n2. In addition, another contribution of the authors is to perform gradient descent on clean samples at each step to reduce the residuals on the equations. The authors ignore many related diffusion model guided generation literature [1,2,3] that the authors do not cite and discuss here.\n\n3. The authors claim that their high-fidelity data is actually simulated on a grid of 256. It is important to realize that for such a large Reynolds number (1000+), this resolution is usually insufficient. shu et al. [4]'s data was simulated on a grid of 2048.\n\n[1] DIFFUSION POSTERIOR SAMPLING FOR GENERAL NOISY INVERSE PROBLEMS\n\n[2] Denoising Diffusion Models for Plug-and-Play Image Restoration\n\n[3] DiffusionPDE: Generative PDE-Solving Under Partial Observation\n\n[4] A physics-informed diffusion model for high-fidelity flow field reconstruction"}},"nonreaders":[],"tmdate":1731428087137,"tcdate":1730643537288,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4262/Reviewer_ovnW"],"signatures":["ICLR.cc/2025/Conference/Submission4262/Reviewer_ovnW"],"forum":"EaiU4F5pwn","number":3,"license":"CC BY 4.0","cdate":1730643537288,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4262/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428087137,"domain":"ICLR.cc/2025/Conference","replyto":"EaiU4F5pwn","id":"m0mYRsVUfa","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed Neural Networks","Computational Fluid Dynamics"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Machine learning (ML) models are increasingly explored in fluid dynamics as a promising way to generate high-fidelity computational fluid dynamics data more efficiently. A common strategy is to use low-fidelity data as computational-efficient inputs, and employ ML techniques to reconstruct high-fidelity flow fields. However, existing work typically assumes that low-fidelity data is artificially downsampled from high-fidelity sources, which limits model performance. In real-world applications, low-fidelity data is generated directly by numerical solvers with a lower initial state resolution, resulting in large deviations from high-fidelity data. To address this gap, we propose PG-Diff, a novel diffusion model for reconstructing high-fidelity flow fields, where both low- and high-fidelity data are generated from numerical solvers. Our experiments reveal that state-of-the-art models struggle to recover fine-grained high-fidelity details when using solver-generated low-fidelity inputs, due to distribution shift. To overcome this challenge, we introduce an \\textit{Importance Weight} strategy during training as self-guidance and a training-free \\textit{Residual Correction} method during inference as physical inductive bias, guiding the diffusion model toward higher-quality reconstructions. Experiments on four 2D turbulent flow datasets demonstrate the effectiveness of our proposed method."},"_bibtex":{"value":"@misc{\nli2025physicsinformed,\ntitle={Physics-Informed Self-Guided Diffusion Model for High-Fidelity Simulations},\nauthor={Ruoyan Li and Zijie Huang and Yizhou Sun and Wei Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=EaiU4F5pwn}\n}"},"title":{"value":"Physics-Informed Self-Guided Diffusion Model for High-Fidelity Simulations"},"pdf":{"value":"/pdf/41a8baae3bbf4450e9eab5f3ce90c57260b3836a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|physicsinformed_selfguided_diffusion_model_for_highfidelity_simulations"},"authorids":{"value":["~Ruoyan_Li1","~Zijie_Huang1","~Yizhou_Sun1","~Wei_Wang13"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruoyan Li","Zijie Huang","Yizhou Sun","Wei Wang"]}},"version":2},{"content":{"summary":{"value":"This work establishes a formal gap between real and complex parameterizations of stable, diagonal SSMs. While complex parameters can trivially express any real SSM, the converse is not true and real SSMs need an arbitrarily large number of parameters to approximate complex SSMs in at least two important cases."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"See above."},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. Clear writing and presentation\n2. Theoretical results support a clear practical suggestion: use complex parametrizations for your SSM if you don't have input selectivity."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Experimental results on non-synthetic datasets would be great. Especially considering the theorems regarding practical learnability and the exponentiality of real parametrizations for random impulse responses with high probability."},"limitations":{"value":"The paper has a thorough discussion of the limitations of the analysis."}},"nonreaders":[],"tmdate":1730879946018,"tcdate":1720383198211,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission18021/Reviewer_SDv4"],"signatures":["NeurIPS.cc/2024/Conference/Submission18021/Reviewer_SDv4"],"forum":"h15RyEj151","number":3,"license":"CC BY 4.0","cdate":1720383198211,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission18021/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879946018,"domain":"NeurIPS.cc/2024/Conference","replyto":"h15RyEj151","id":"9ZMphy0Gzc","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"We establish formal gaps between real and complex SSMs in terms of expressiveness and practical learnability"},"keywords":{"value":["Neural Networks","Theory","Structured State Space Models","Mamba","S4","Complex parametrization"]},"primary_area":{"value":"learning_theory"},"abstract":{"value":"Structured state space models (SSMs), the core engine behind prominent neural networks such as S4 and Mamba, are linear dynamical systems adhering to a specified structure, most notably diagonal. In contrast to typical neural network modules, whose parameterizations are real, SSMs often use complex parameterizations. Theoretically explaining the benefits of complex parameterizations for SSMs is an open problem. The current paper takes a step towards its resolution, by establishing formal gaps between real and complex diagonal SSMs. Firstly, we prove that while a moderate dimension suffices in order for a complex SSM to express all mappings of a real SSM, a much higher dimension is needed for a real SSM to express mappings of a complex SSM. Secondly, we prove that even if the dimension of a real SSM is high enough to express a given mapping, typically, doing so requires the parameters of the real SSM to hold exponentially large values, which cannot be learned in practice. In contrast, a complex SSM can express any given mapping with moderate parameter values. Experiments corroborate our theory, and suggest a potential extension of the theory that accounts for selectivity, a new architectural feature yielding state of the art performance."},"_bibtex":{"value":"@inproceedings{\nran-milo2024provable,\ntitle={Provable Benefits of Complex Parameterizations for Structured State Space Models},\nauthor={Yuval Ran-Milo and Eden Lumbroso and Edo Cohen-Karlik and Raja Giryes and Amir Globerson and Nadav Cohen},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=h15RyEj151}\n}"},"title":{"value":"Provable Benefits of Complex Parameterizations for Structured State Space Models"},"pdf":{"value":"/pdf/36066871b1ea5ba0d5c2126254596aa0019d538c.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"ranmilo|provable_benefits_of_complex_parameterizations_for_structured_state_space_models"},"authorids":{"value":["~Yuval_Ran-Milo1","~Eden_Lumbroso1","~Edo_Cohen-Karlik1","~Raja_Giryes1","~Amir_Globerson1","~Nadav_Cohen1"]},"authors":{"value":["Yuval Ran-Milo","Eden Lumbroso","Edo Cohen-Karlik","Raja Giryes","Amir Globerson","Nadav Cohen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a benchmark that pairs real measurements with matched numerical simulations across five scenarios (cylinder wake, controlled cylinder, FSI, foil, and combustion; governing equations span Navier-Stokes, coupled FSI, and reactive Navier-Stokes with species transport). The benchmark defines three training regimes (train on simulation, train on real, pretrain on simulation then finetune on real), includes a mix of pixel and physics flavored metrics, and evaluates a set of neural PDE baselines including a pretrained foundation model. The headline empirical messages are: there is a nontrivial gap between simulation and laboratory data; pretraining on simulation generally helps downstream on real; and the codebase makes it straightforward to add models or datasets.\n\nI think this is timely and potentially useful. If we are serious about sim2real for scientific ML, we need carefully curated real data and a shared protocol. However, the current paper mixes benchmarking with a sim2real narrative in ways that are a bit loose, the experimental details are too thin for others to trust or extend the datasets, and some of the metrics and literature framing are not well aligned with fluid mechanics and combustion practice.\n\nAm open to potentially increasing my score if the authors narrow the claims, remove the robotics sim2real framing argument and instead point to real sim2real problems in fluid dynamics, add significantly more context to place this in the existing fluid dynamics literature, and substantially improve the experimental documentation and physics grounded evaluation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"[Several questions discussed above]"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The benchmark collects paired real and simulated trajectories for several nontrivial systems instead of yet another synthetic-only PDE suite. This is likely to be useful for the community.\n- The split of training regimes (simulation only, real only, pretrain on simulation then finetune on real) is useful, and the pretraining result is consistent with what many of us have seen in practice.\n- The code appears modular enough to add a new dataset or baseline without painful surgery, and using a single file format lowers friction for adoption.\n- Including both data oriented metrics and physics oriented diagnostics is better than reporting only RMSE. The autoregressive evaluation option is also a good idea.\n- The combustion scenario is ambitious and, if documented properly, could become a valuable stress test beyond the usual laminar toy problems.\n- The baseline measurements for their benchmark are very extensive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Major concerns:\n\n- The documentation about the experiment is unacceptably thin in its current form. If the experimental data was created or modified from another source, you need to cite it. If someone else created the dataset for you, you need them to write documentation for it. The current documentation on experimental data generation (which, of course, can be included in the appendix) is simply unacceptable for publication, especially for a paper which is supposed to be about this very dataset.\n- The paper motivates the sim2real gap by citing mostly work from robotics, which I found very strange, almost as if the authors are guessing there is a sim2real gap in fluids, without actually surveying the literature. Robotics has a much different sim2real gap than turbulence research does. In fluids, the sources of discrepancy, data acquisition, and noise models can be quite different. Please ground the narrative in fluids and combustion references. If you are addressing the sim2real gap in fluids, you must speak from the context of the fluids community, and discuss the ways in which the fluids community has quantified this gap.\n- Please state clearly whether you will release raw data (e.g., the PIV frames), calibration files, and the full processing scripts, not just the final HDF5 arrays, so that it can be checked by others. If raw data cannot be released, say so and justify it. Benchmarks live or die by their data hygiene, and biases in a benchmark can leak into biases in the community's preferred models.\n- Several figures quantify differences using frequency or Fourier errors over image-like arrays. This is admittedly a start, but it is not a physics grounded measure of mismatch between experiment and simulation, and certainly not something people in the fluids community would actually use as a robust measure of discrepancies between simulation and real data. In fact, for several real-world experimental problems, there is a difference here that is simply due to the nature of the real-world experiment, yet the simulation and experiment can actually have no discrepancy. This is because you would only care about some summary statistic, and not care about some wave mode that you know does not affect your statistic. Yet, your metric would completely miss this. I recommend reviewing and citing the fluids literature and how people measure discrepancies between simulations and real data. Note that this is a problem people have studied for literally decades in the fluids community. These additions would make your claims about a \"gap\" more convincing.\n- Simulation contains modalities that are not observed in the lab, and the current strategy randomly masks channels and adds noise. That is a start, but it does not reflect the actual sensor physics. Please consider sensor specific degradations (camera noise models, optical blur, saturation, PIV algorithmic artifacts) and state explicitly which channels are used for training and which are hidden. It would also help to define tasks that force parity, for example training all models only on the modalities that the lab provides.\n- The paper claims to be the first benchmark that integrates real-world measurements with paired numerical simulations across complex physical systems. Within fluids, this is far from true; as just one demonstration, \"ERCOFTAC\" has hosted combined experimental and numerical reference cases since 1995 across a wide range of flows. Please narrow the novelty claim to something that accurately represents existing datasets and benchmarks, and perhaps cite existing databases.\n\n\nAdditional comments and suggestions\n\n- The update ratio metric is interesting, but it conflates pretraining data scale with optimization effects.\n- Please make the train, validation, and test splits explicit at the parameter level so that generalization across Reynolds number, control frequency, mass ratio, or equivalence ratio is clear. I consider those to be most interesting axes. And if you already do this, show it more prominently.\n- The autoregressive evaluation stops very early. If you want to make claims about stability, show longer horizons and add probe based diagnostics, not only field RMSE.\n- The baselines are modern ML models, which is fine for a benchmark, but the story would be stronger if you add one or two domain baselines for each scenario (for example a simple reduced order model, or even a physics based filter) so readers have a calibration point.\n- Throughout the paper there are small terminology issues. I would prefer \"numerical error\" over \"computational error\" (computational error sounds like a code bug). Be precise about whether errors arise from discretization, closure modeling, boundary conditions, or measurement.\n- Not a criticism, but I think it would be better to rename the benchmark. \"RealBench\" seems far too broad and will clash with many domains. Something like \"RealPDEBench\" or \"RealFlowBench\" might be more appropriate."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916168832,"tcdate":1762241594100,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2270/Reviewer_6AF9"],"signatures":["ICLR.cc/2026/Conference/Submission2270/Reviewer_6AF9"],"forum":"y3oHMcoItR","number":4,"license":"CC BY 4.0","cdate":1762241594100,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2270/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916168832,"domain":"ICLR.cc/2026/Conference","replyto":"y3oHMcoItR","id":"F57t6Rab9l","forumContent":{"venue":{"value":"ICLR 2026 Oral"},"TLDR":{"value":"We propose the first benchmark for complex physical systems with paired real-world data and simulated data, and explore how to bridge simulated and real-world data."},"keywords":{"value":["complex physical system","PDE","benchmark","real-world data","prediction"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Predicting the evolution of complex physical systems remains a central problem in science and engineering. Despite rapid progress in scientific Machine Learning (ML) models, a critical bottleneck is the lack of expensive real-world data, resulting in most current models being trained and validated on simulated data. Beyond limiting the development and evaluation of scientific ML, this gap also hinders research into essential tasks such as sim-to-real transfer. We introduce RealPDEBench, the first benchmark for scientific ML that integrates real-world measurements with paired numerical simulations. RealPDEBench consists of five datasets, three tasks, nine metrics, and ten baselines. We first present five real-world measured datasets with paired simulated datasets across different complex physical systems. We further define three tasks, which allow comparisons between real-world and simulated data, and facilitate the development of methods to bridge the two. Moreover, we design nine evaluation metrics, spanning data-oriented and physics-oriented metrics, and finally benchmark ten representative baselines, including state-of-the-art models, pretrained PDE foundation models, and a traditional method. Experiments reveal significant discrepancies between simulated and real-world data, while showing that pretraining with simulated data consistently improves both accuracy and convergence. In this work, we hope to provide insights from real-world data, advancing scientific ML toward bridging the sim-to-real gap and real-world deployment. Our benchmark, datasets, and instructions are available at https://realpdebench.github.io/."},"_bibtex":{"value":"@inproceedings{\nhu2026realpdebench,\ntitle={Real{PDEB}ench: A Benchmark for Complex Physical Systems with Real-World Data},\nauthor={Peiyan Hu and Haodong Feng and Hongyuan Liu and Tongtong Yan and Wenhao Deng and Tianrun Gao and Rong Zheng and Haoren Zheng and Chenglei Yu and Chuanrui Wang and Kaiwen Li and Zhi-Ming Ma and Dezhi Zhou and Xingcai Lu and Dixia Fan and Tailin Wu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=y3oHMcoItR}\n}"},"title":{"value":"RealPDEBench: A Benchmark for Complex Physical Systems with Real-World Data"},"pdf":{"value":"/pdf/87a2cc33b38c708af3705ef8d5789740b5502dbf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"hu|realpdebench_a_benchmark_for_complex_physical_systems_with_realworld_data"},"authorids":{"value":["~Peiyan_Hu1","~Haodong_Feng1","~Hongyuan_Liu3","~Tongtong_Yan1","~Wenhao_Deng2","~Tianrun_Gao3","~Rong_Zheng3","~Haoren_Zheng2","~Chenglei_Yu1","~Chuanrui_Wang2","~Kaiwen_Li9","~Zhi-Ming_Ma1","~Dezhi_Zhou1","~Xingcai_Lu1","~Dixia_Fan1","~Tailin_Wu1"]},"authors":{"value":["Peiyan Hu","Haodong Feng","Hongyuan Liu","Tongtong Yan","Wenhao Deng","Tianrun Gao","Rong Zheng","Haoren Zheng","Chenglei Yu","Chuanrui Wang","Kaiwen Li","Zhi-Ming Ma","Dezhi Zhou","Xingcai Lu","Dixia Fan","Tailin Wu"]}},"version":2},{"content":{"comment":{"value":"We agree that broadly benchmarking physics-grounded text-to-3D motion remains important, but this task currently lacks standardized datasets. To ensure a fair comparison, we therefore follow the evaluation protocol in the existing works (e.g., PhysGaussian, PhysDreamer, OmniPhysGS), using representative template prompts and both quantitative and qualitative metrics. \n\nWhile our framework can naturally support complex multi-object interactions, in this work we focus on single-object scenarios to study the core challenge of learning mixed-material, physics-grounded dynamics from text without confounding factors due to scene complexity. However, we included more diverse video generation examples in the revised **supplementary materials** and the Appendix to better illustrate generalization."},"title":{"value":"Q1: Evaluation is limited, with no standard datasets or complex multi-object tests."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764579017543,"tcdate":1764567473540,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3940/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission3940/Authors"],"forum":"mq43BAAos0","number":11,"license":"CC BY 4.0","cdate":1764567473540,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3940/-/Official_Comment"],"mdate":1764579017543,"domain":"ICLR.cc/2026/Conference","replyto":"ja9CjYNkQ3","id":"C6A0oYzfzL","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Text-to-Video","Gaussian Splatting","Diffusion Model","Dynamic 3D Generation","LLM"]},"supplementary_material":{"value":"/attachment/7d953bf9100e1b07d3a33399bdb531b407c0d505.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Generating realistic 3D object videos is crucial for virtual reality and digital content creation. However, existing 3D dynamics generation methods often struggle to achieve high-quality appearance and physics-aware motion, relying on manual inputs and pre-existing models. To address these challenges, we propose DiffuPhyGS, a novel framework that generates high-quality 3D objects with realistic and learnable physical motion directly from text prompts. Our approach features an LLM-Chain-of-Thought-based Iterative Prompt Refinement (LLM-CoT-IPR) method, which obtains prompt-aligned 2D and multi-view 3D diffusion priors to guide Gaussian Splatting (GS) to generate 3D objects. We further enhance 3D generation quality with a Densification-by-Adaptive-Splitting (DAS) mechanism. Next, we employ a material property decoder that utilizes a Mixture-of-Experts Material Constitutive Models (MoEMCMs) to predict the mixed material properties of the 3D object. We then apply the Material Point Method (MPM) to deform 3D Gaussian kernels, ensuring physics-grounded motion guided by implicit and explicit physical priors from the video diffusion model and a velocity loss function. Extensive experiments show DiffuPhyGS outperforms other methods in generating realistic physics-grounded motion across diverse materials."},"_bibtex":{"value":"@misc{\nwang2025diffuphygs,\ntitle={DiffuPhy{GS}: Text-to-Video Generation with 3D Gaussians and Learnable Physical Properties via Diffusion Priors},\nauthor={Wenqing Wang and Yun Fu},\nyear={2025},\nurl={https://openreview.net/forum?id=mq43BAAos0}\n}"},"title":{"value":"DiffuPhyGS: Text-to-Video Generation with 3D Gaussians and Learnable Physical Properties via Diffusion Priors"},"pdf":{"value":"/pdf/47a9347799011259548f308f32392a5726638e36.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|diffuphygs_texttovideo_generation_with_3d_gaussians_and_learnable_physical_properties_via_diffusion_priors"},"authorids":{"value":["~Wenqing_Wang4","~Yun_Fu1"]},"authors":{"value":["Wenqing Wang","Yun Fu"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on efficiency of code generation: existing approaches either skip planning (which may lead to poor performance on complex tasks) or adopt a \"Plan-before-Trial\" (PbT) paradigm that wastes computational resources on simple tasks by enforcing planning regardless of task difficulty. \nTo this end, the authors propose PaT (Planning-after-Trial), a two-stage framework that prioritizes efficiency without sacrificing performance. \nIn the first stage, PaT leverages lightweight, low-cost small models to directly generate code for a given task; it then validates the generated code using test cases. If the code passes validation, PaT terminates early to save costs. \nIf validation fails (indicating a complex task), PaT invokes a more capable but expensive large model to perform task decomposition (planning), breaking the complex task into manageable sub-tasks that are solved iteratively.\nTo further optimize efficiency, PaT adopts a strategy: small models handle code generation, while large models only contribute to planning. \nThe authors evaluate PaT across multiple models and code benchmarks , showing that it matches the accuracy of larger models (e.g., Qwen3-32B) with smaller models (e.g., Qwen3-4B) while reducing costs, and outperforms prior methods like FunCoder on complex tasks at ~60% of the computational cost. The paper’s core contributions are the PaT framework, the role division of small/large models for cost-efficiency, and empirical validation of its effectiveness on diverse code generation tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"please refer to weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper proposes the use of heterogeneous models and first trial pipeline to resolve the overhead and performance problem. \nSpecifically, a smaller model is first employed for reasoning—if it passes the test cases, the process terminates; otherwise, a more powerful model and code generation workflow are invoked to solve complex problems. The proposed approach follows a straightforward rationale, and the authors have conducted extensive experiments across multiple baselines to validate its effectiveness. They claim that their method outperforms the previous SOTA approach in both performance and efficiency.\n\n2. The paper introduces a heterogeneous model framework to tackle complex problems: a more powerful large model is utilized for overall code generation planning, while a smaller model handles relatively deterministic code generation tasks. This collaborative approach aims to reduce the overall computational overhead in the code generation process."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The approach proposed in this paper is straightforward, with a key aspect being how to verify the correctness of the initial reasoning performed by the small model, which requires accurate test cases.\nHowever, the paper states that the authors used both the built-in test cases provided with the problems and additional test cases generated by the model based on problem requirements to jointly validate the small model's outputs. This raises a concern: does using the provided test cases to verify the small model's results risk potential data leakage?\n\n2. According to the experimental results, the proposed method demonstrates significantly and consistently superior performance compared to the previous SOTA method. While the paper claims to improve reasoning efficiency, the experiments show enhancements in both efficiency and performance over prior methods. Could the authors analyze the specific sources of these performance gains? Were the experimental settings fully aligned with those of the previous SOTA method.\n\n3. Could the authors provide additional statistical insights—such as the proportion of problems successfully solved by the small model alone, and a comparison of token overhead between the proposed method and pre-decomposition approaches like Funcoder for complex problems—to offer a more intuitive understanding of the results?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927574537,"tcdate":1761729144479,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17741/Reviewer_KVew"],"signatures":["ICLR.cc/2026/Conference/Submission17741/Reviewer_KVew"],"forum":"767aZTpsIl","number":3,"license":"CC BY 4.0","cdate":1761729144479,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17741/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927574537,"domain":"ICLR.cc/2026/Conference","replyto":"767aZTpsIl","id":"5EWlGjtsDz","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We propose Planning-after-Trial (PaT), a policy that attempts a direct solution and invokes an expensive planner only upon failure, significantly improving the cost-performance trade-off for code generation."},"keywords":{"value":["Code Generation","Divide-and-conquer"]},"supplementary_material":{"value":"/attachment/361b100aded8f48fbd2f60b3e2b4fb3eb0883692.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Large language models (LLMs) have demonstrated increasingly sophisticated capabilities for code generation. To extend the problem-solving reach of cost-efficient models to complex problems, strategic planning via problem decomposition has emerged as a key paradigm. However, most existing pipelines adopt a rigid Planning-before-Trial (PbT) policy, which inefficiently allocates test-time compute by incurring planning overhead even on directly solvable problems. We propose an adaptive Planning-after-Trial (PaT) policy that uses the outcome of a direct attempt as a feedback signal, invoking a planner only upon verification failure. This adaptive policy naturally enables a heterogeneous model configuration: a cost-efficient model handles generation attempts, while a powerful model is reserved for targeted planning interventions. Empirically, across multiple benchmarks and model families, our approach significantly advances the cost-accuracy Pareto frontier by judiciously avoiding indiscriminate planning on simple problems and concentrating test-time compute precisely where it is needed most."},"_bibtex":{"value":"@misc{\nyoon2025pat,\ntitle={PaT: Planning-after-Trial for Efficient Code Generation},\nauthor={Youngsik Yoon and Sungjae Lee and Seockbean Song and Siwei Wang and Wei Chen and Jungseul Ok},\nyear={2025},\nurl={https://openreview.net/forum?id=767aZTpsIl}\n}"},"title":{"value":"PaT: Planning-after-Trial for Efficient Code Generation"},"pdf":{"value":"/pdf/29c033f604fe3d1ccea84f32d67685fe642877b1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yoon|pat_planningaftertrial_for_efficient_code_generation"},"authorids":{"value":["~Youngsik_Yoon1","~Sungjae_Lee4","~Seockbean_Song1","~Siwei_Wang2","~Wei_Chen10","~Jungseul_Ok2"]},"authors":{"value":["Youngsik Yoon","Sungjae Lee","Seockbean Song","Siwei Wang","Wei Chen","Jungseul Ok"]}},"version":2},{"content":{"summary":{"value":"The paper studies how generative language models (LLMs) may interact recursively through training data that include other models’ outputs.\n\n1. It introduces a formal framework with two parameters — the synthetic data ratio ($\\alpha$) and the initial data weight ($\\beta$) — to describe cross-model data mixing.\n\n2. Theoretical analysis (Sec. 3) under a generalized linear model shows convergence and bias–variance behavior depending on $\\alpha$ and $\\beta$.\n\n3. Empirical results (Sec. 4; Fig. 3–5) using OPT-350M and LLaMA-1B demonstrate that moderate mixing ($\\alpha=\\beta=0.5$) improves both models’ performance but leads to representation homogenization.\n\nThe paper also concludes with a discussion of the implications for model diversity and long-term ecosystem dynamics (Sec. 5)."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"You may refe to the content in the “weakness” section. If you can address the doubts raised there effectively, I will consider giving a higher score. \n\nI hope valuable work will not be overlooked.\n\n\nFor example, in addition to theoretical explanations based on existing assumptions, it would be great if you could highlight some unique insights proposed in this paper.\nPerhaps the paper already includes them, but I did not notice."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"1. Clear and meaningful problem setting：The paper addresses a timely and practically significant question — how recursive data interactions among generative models affect learning stability and diversity. The motivation and background are well-articulated (Sec. 1–2), making the research goal both relevant and understandable.\n\n2. Comprehensive and interpretable theoretical framework：The proposed formalism based on the parameters $\\alpha$ and $\\beta$ (Sec. 3) systematically captures cross-model data mixing. The accompanying bias–variance and convergence analysis provides solid conceptual grounding for the empirical findings (Fig. 2–3). Even without verifying every derivation, the overall reasoning is coherent and accessible.\n\n3. Exceptional clarity and readability：The writing is well-structured and accessible to readers beyond the immediate subfield. Explanations, figures, and notation are consistently clear, enabling a broad audience to grasp the motivation, methodology, and conclusions (Sec. 1–5)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited novelty in the modeling of cross-model interaction：The description of data-mediated interactions between models (Sec. 3) is clear and well-structured but largely descriptive. While it helps readers understand the setup, this section mainly formalizes an intuitive process rather than introducing a new mechanism or theoretical insight. As a result, the contribution of this part feels limited in terms of originality.\n\n2. Gap between theoretical modeling and practical relevance：Most of the paper focuses on theoretical modeling and proofs (Sec. 4–5). Although the derivations appear sound, the connection to real-world large-scale training scenarios remains weak. The introduction of parameters $\\alpha$ and $\\beta$ is conceptually useful, yet in practice, their exact values or ratios are difficult to estimate or control during continuous training. The conclusions drawn from the linear or generalized linear setting may not easily transfer to nonlinear or high-dimensional models.In essence, while the problem definition is good and $\\alpha$–$\\beta$ reasoning is meaningful, it is unclear how the theory can concretely guide actual large-model training.\n\n3. Experimental validation is narrow and idealized：The experiments (Sec. 4–5) mainly serve to verify the theory, but they do not provide further insights into realistic settings. Only two medium-sized models (OPT-350M and LLaMA-1B) and two datasets (SciQ, GSM8K) are used, with highly controlled data composition. The synthetic data are assumed to represent model outputs cleanly, without considering realistic mixtures of human and synthetic text (I know in limitation part). Scaling experiments or additional ablations (e.g., varying model size, task diversity, or realistic data proportions) would make the findings more convincing.\n\nOverall, the experimental content is rather insufficient. The question itself is meaningful, but it does not provide much insight in terms of conclusions. However, considering that this might be a theoretical paper, it is difficult to for me to  assess the practical value of such a theory. Therefore, I would lower the confidence to mitigate the possible impact of this uncertainty."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928445682,"tcdate":1761988224524,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18737/Reviewer_rQht"],"signatures":["ICLR.cc/2026/Conference/Submission18737/Reviewer_rQht"],"forum":"JEU4PBaX85","number":4,"license":"CC BY 4.0","cdate":1761988224524,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18737/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928445682,"domain":"ICLR.cc/2026/Conference","replyto":"JEU4PBaX85","id":"MQ3x28akj3","forumContent":{"TLDR":{"value":"When generative AI models train on each others' generated outputs, they benefit by learning from other models' distributions but can become homogeneous."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["model collapse","generative AI"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"The internet serves as a common source of training data for generative AI (genAI) models but is increasingly populated with AI-generated content. This duality raises the possibility that future genAI models may be trained on other models'  generated outputs. Prior work has studied consequences of models training on their own generated outputs, but limited work has considered what happens if models ingest content produced by other models. Given society's increasing dependence on genAI tools, understanding such data-mediated model interactions is critical. This work provides empirical evidence for how data-mediated interactions might unfold in practice, develops a theoretical model for this interactive training process, and experimentally validates the theory. We find that data-mediated interactions can benefit models by exposing them to novel concepts perhaps missed in original training data, but also can homogenize their performance on shared tasks."},"_bibtex":{"value":"@inproceedings{\nvu2026what,\ntitle={What happens when generative {AI} models train recursively on each others' outputs?},\nauthor={Hung Anh Vu and Galen Reeves and Emily Wenger},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=JEU4PBaX85}\n}"},"title":{"value":"What happens when generative AI models train recursively on each others' outputs?"},"pdf":{"value":"/pdf/beb33cd6f02c26be5725ab1ea9fb8581d4b2366b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"vu|what_happens_when_generative_ai_models_train_recursively_on_each_others_outputs"},"authorids":{"value":["~Hung_Anh_Vu1","~Galen_Reeves1","~Emily_Wenger1"]},"authors":{"value":["Hung Anh Vu","Galen Reeves","Emily Wenger"]}},"version":2},{"content":{"venue":{"value":"ACL (Findings) 2025"},"pdf":{"value":"https://aclanthology.org/2025.findings-acl.660.pdf"},"venueid":{"value":"dblp.org/conf/ACL/2025"},"paperhash":{"value":"kuang|express_what_you_see_can_multimodal_llms_decode_visual_ciphers_with_intuitive_semiosis_comprehension"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jiayi_Kuang:","https://dblp.org/search/pid/api?q=author:Yinghui_Li:","https://dblp.org/search/pid/api?q=author:Chen_Wang_0049:","https://dblp.org/search/pid/api?q=author:Haohao_Luo:","~Ying_Shen1","https://dblp.org/search/pid/api?q=author:Wenhao_Jiang:"]},"html":{"value":"https://aclanthology.org/2025.findings-acl.660/"},"_bibtex":{"value":"@inproceedings{DBLP:conf/acl/KuangL0L0J25,\n  author={Jiayi Kuang and Yinghui Li and Chen Wang and Haohao Luo and Ying Shen and Wenhao Jiang},\n  title={Express What You See: Can Multimodal LLMs Decode Visual Ciphers with Intuitive Semiosis Comprehension?},\n  year={2025},\n  cdate={1735689600000},\n  pages={12743-12774},\n  url={https://aclanthology.org/2025.findings-acl.660/},\n  booktitle={ACL (Findings)},\n  crossref={conf/acl/2025f}\n}\n"},"abstract":{"value":"Bridging the gap between visual and language remains a pivotal challenge for the multimodal community. Traditional VQA benchmarks encounter a modality gap and over-reliance on language priors, whereas human cognition excels at intuitive semiosis, associating abstract visual symbols to linguistic semantics. Inspired by this neurocognitive mechanism, we focus on emojis, the visual cipher conveying abstract textual semantics. Specifically, we propose a novel task of generating abstract linguistics from emoji sequence images, where such reasoning underpins critical applications in cryptography, thus challenging MLLMs’ reasoning of decoding complex semantics of visual ciphers. We introduce eWe-bench (Express What you SeE), assessing MLLMs’ capability of intuitive semiosis like humans. Our data construction framework ensures high visual sensitivity and data quality, which can be extended to future data enhancement. Evaluation results on advanced MLLMs highlight critical deficiencies in visual intuitive symbolic reasoning. We believe our interesting insights for advancing visual semiosis in MLLMs will pave the way for cryptographic analysis and high-level intuitive cognition intelligence of MLLMs."},"title":{"value":"Express What You See: Can Multimodal LLMs Decode Visual Ciphers with Intuitive Semiosis Comprehension?"},"authors":{"value":["Jiayi Kuang","Yinghui Li","Chen Wang","Haohao Luo","Ying Shen","Wenhao Jiang"]}},"tmdate":1768972616045,"pdate":1767139200000,"externalIds":["dblp:conf/acl/KuangL0L0J25"],"tcdate":1768972613366,"writers":["~"],"signatures":["~Ying_Shen1"],"forum":"ECtHiqwEfr","license":"CC BY-SA 4.0","number":778063,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768972616045,"domain":"DBLP.org","id":"ECtHiqwEfr","version":2},{"content":{"summary":{"value":"This paper conducts a pioneering theoretical study on the expressive power of shallow polynomial neural networks (PNNs) over the finite field $\\mathbb{F}_q$. The authors define expressive power as the cardinality of their \"neural manifold\" and successfully provide exact formulas for this cardinality for several key classes of architectures (e.g., $r=2, k=1$ and $m=1,2$). In particular, the authors leverage the algebraic properties of the finite field's characteristic $p$ (when $p | r$) to simplify the problem. A core contribution is the revelation, through a key example ($d=(2,2,2), r=2$), of the significant difference in expressive power between finite fields and the complex numbers, showing that an architecture that \"fills\" the space over $\\mathbb{C}$ may only occupy a small fraction of the ambient space over $\\mathbb{F}_q$."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. could the authors further elaborate on the finding in Ex 3.6 that $\\mathcal{P}_{(2,2,2),r=2}$ only \"fills\" about 1/2 of the ambient space over $\\mathbb{F}_p$? Does this imply that the learning capacity of this quantized architecture is severely limited, or could this be seen as a form of \"implicit regularization\" that is beneficial for generalization?\n\n2. could the authors provide a simple example (even speculatively) to illustrate the connection with the Weil Conjectures? For instance, Lemma 4.2 concludes that $|\\overline{\\mathcal{P}_{(n,1,k),r}}| = |\\mathbb{P}^{n-1}||\\mathbb{P}^{k-1}|$. Does this imply that the Betti numbers of its complex manifold are related to those of $\\mathbb{P}^{n-1} \\times \\mathbb{P}^{k-1}$? \n\n3. the authors mention that for the analysis of $m>2$, the \"combinatorial complexity grows quickly.\" Could you specify where this complexity manifests? Is it related to the computation of subspaces of a specific rank (e.g., points on a Grassmannian), or are there more complex combinatorial dependencies?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"The main strength of this paper lies in its novelty. It provides a rigorous, algebraic-geometric framework for analysis for the important problem of \"quantifying the expressive power of neural networks\" (i.e., calculating the cardinality of the neural manifold over $\\mathbb{F}_q$). The paper not only proposes a framework but also delivers very solid and concrete mathematical results, providing exact cardinality counting formulas for many classes of architectures (as shown in Table 1). The authors skillfully apply advanced mathematical tools (such as projective geometry and finite field algebra) to solve complex counting problems in an elegant manner, for example, leveraging the property $(a+b)^p = a^p + b^p$ to simplify the case where $p|r$. A valuable contribution of this paper is its profound insight: it strongly demonstrates that the expressive power over finite fields can be fundamentally different from that over the complex numbers (Zariski closure), which serves as a cautionary note for theoretical research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although theoretical strength, this submission has some notable limitations. The most significant is the weak connection to practice, maybe I am not aware of any direct application at this point. \n\nThe paper is motivated by \"weight quantization\" but does not discuss the specific impact of its mathematical findings (e.g., the neural manifold occupying only half the space in Ex 3.6) on the practical training or generalization ability of quantized networks, which leaves the practical significance of the paper somewhat ambiguous. Furthermore, all analysis is strictly limited to single-hidden-layer (shallow) networks, and it is currently unclear to what extent these conclusions or analytical methods can be generalized to deep networks. At the same time, the authors repeatedly mention a connection to the \"Weil Conjectures\" (used to infer topological properties of the complex manifold). While this is an exciting motivation, the paper only completes the counting and does not actually demonstrate this connection, making the argument for the motivation incomplete."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942437067,"tcdate":1761997533561,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22921/Reviewer_FWBX"],"signatures":["ICLR.cc/2026/Conference/Submission22921/Reviewer_FWBX"],"forum":"tfGuvCp50e","number":4,"license":"CC BY 4.0","cdate":1761997533561,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22921/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942437067,"domain":"ICLR.cc/2026/Conference","replyto":"tfGuvCp50e","id":"ANT3fHDfzI","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Polynomial neural networks","finite fields","expressivity","neuromanifold"]},"supplementary_material":{"value":"/attachment/1ad9e6e8e4fc097bbfab70f15d239ed30369d0a0.zip"},"primary_area":{"value":"learning theory"},"abstract":{"value":"We study the expressivity of shallow polynomial neural networks (PNNs) with monomial activation functions over finite fields. For a given architecture, we define a neuromanifold as the image of the map from all possible network weights into the product of polynomial rings. We quantify the expressivity by the cardinality of the neuromanifold, and derive a natural lower and upper bound. This leads to counting rational points over finite fields, a problem closely linked to the Weil conjectures. Finally, we present an architecture that exhibits a striking difference in the neuromanifolds when considered over a characteristic zero versus a finite‐characteristic field, illustrating the critical role of field characteristic on the notion of expressivity."},"_bibtex":{"value":"@misc{\nzubkov2026expressivity,\ntitle={Expressivity of Shallow Neural Networks Over Finite Fields},\nauthor={Maksym Zubkov and Carol Wu and Shiwei Yang and Param Mody and Yifei Chen},\nyear={2026},\nurl={https://openreview.net/forum?id=tfGuvCp50e}\n}"},"title":{"value":"Expressivity of Shallow Neural Networks Over Finite Fields"},"pdf":{"value":"/pdf/11d526eec1b4c5d4dbd77ca0fe99090e01d37410.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zubkov|expressivity_of_shallow_neural_networks_over_finite_fields"},"authorids":{"value":["~Maksym_Zubkov1","~Carol_Wu2","~Shiwei_Yang1","~Param_Mody1","~Yifei_Chen3"]},"authors":{"value":["Maksym Zubkov","Carol Wu","Shiwei Yang","Param Mody","Yifei Chen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes chain of time simulation, or generating intermediate images during a simulation, to evaluate the capabilities of image generation models on predicting simulated and natural evaluations. They find that this significantly improves performance over the baseline, and suggests that image generation models are able to simulate some properties over time."},"soundness":{"value":1},"confidence":{"value":2},"questions":{"value":"How would you differentiate your method from the visual chain of thought line of work (which also produce intermediate images)? The related work mentions the final goal being that of producing an image, but there is existing literature applying CoT to image generation as well.\n\nCould you provide in the appendix additional images from the evaluation and across time steps? Image generation models often make other mistakes that are not related to physics, which may affect evaluation."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- Studies an unique topic - physics understanding in image generation models\n- Motivates the study from an interdisciplinary perspective"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- limited evaluation. image generation models span a variety of designs, and evaluating only one is insufficient.\n- given that gpt's image generation is a closed model with little public detail, it may be difficult to act on these findings to improve image generation models\n- limited context from related work\n- experimental results are unclear. for example, the paper mentions that figure 6 shows that IGM is able to simulate the projectile's motion because it is close to ground truth, but the pattern does not seem to behave in a way consistent with a physics equation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942553016,"tcdate":1762000357867,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23193/Reviewer_qjNx"],"signatures":["ICLR.cc/2026/Conference/Submission23193/Reviewer_qjNx"],"forum":"f6ugrBWs3K","number":4,"license":"CC BY 4.0","cdate":1762000357867,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23193/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942553016,"domain":"ICLR.cc/2026/Conference","replyto":"f6ugrBWs3K","id":"N2FBvXQlpe","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-modal Language Models","Spatial and Temporal Perception","Image Generation","Physical Reasoning"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"We propose a novel method to improve the physical simulation ability of vision-language models. This Chain-of-Time simulation is motivated by in-context reasoning in machine learning, and mental simulation in humans. The method involves generating a series of intermediate images during a simulation. Chain of Time is used at inference time and requires no additional fine-tuning for performance benefits. We apply the Chain-of-Time method to synthetic and real-world domains, including 2-D graphics simulations and natural 3-D videos. These domains test a variety of particular physical properties, including velocity, acceleration, fluid dynamics, and conservation of momentum. We found that using Chain-of-Time simulation substantially improves the performance of state-of-the-art Image Generation Model. Beyond examining performance, we also analyze the specific states of the world simulated by an image model at each time step, which sheds light on the dynamics underlying these simulations. This analysis reveals insights that are hidden from traditional evaluations of physical reasoning, including cases where an Image Generation Model is able to simulate physical properties that unfold over time, such as velocity, gravity, and collisions domain well. Our analysis also highlights particular cases where the Image Generation Model struggles to infer particular physical parameters from input images, despite being capable of simulating relevant physical processes."},"_bibtex":{"value":"@misc{\nwang2026chain,\ntitle={Chain of Time: In-Context Physical Simulation with Image Generation Models},\nauthor={YingQiao Wang and Eric Bigelow and Boyi Li and Tomer Ullman},\nyear={2026},\nurl={https://openreview.net/forum?id=f6ugrBWs3K}\n}"},"title":{"value":"Chain of Time: In-Context Physical Simulation with Image Generation Models"},"pdf":{"value":"/pdf/5d0ef4077bf57397482b6d12a29a564f7f8e8459.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|chain_of_time_incontext_physical_simulation_with_image_generation_models"},"authorids":{"value":["~YingQiao_Wang1","~Eric_Bigelow1","~Boyi_Li1","~Tomer_Ullman1"]},"authors":{"value":["YingQiao Wang","Eric Bigelow","Boyi Li","Tomer Ullman"]}},"version":2},{"content":{"venue":{"value":"CoRR 2018"},"pdf":{"value":"http://arxiv.org/pdf/1810.05095v2"},"venueid":{"value":"dblp.org/journals/CORR/2018"},"paperhash":{"value":"cimini|the_statistical_physics_of_realworld_networks"},"authorids":{"value":["~Giulio_Cimini1","https://dblp.org/search/pid/api?q=author:Tiziano_Squartini:","https://dblp.org/search/pid/api?q=author:Fabio_Saracco:","https://dblp.org/search/pid/api?q=author:Diego_Garlaschelli:","https://dblp.org/search/pid/api?q=author:Andrea_Gabrielli:","https://dblp.org/search/pid/api?q=author:Guido_Caldarelli:"]},"html":{"value":"http://arxiv.org/abs/1810.05095"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-1810-05095,\n  publtype={informal},\n  author={Giulio Cimini and Tiziano Squartini and Fabio Saracco and Diego Garlaschelli and Andrea Gabrielli and Guido Caldarelli},\n  title={The Statistical Physics of Real-World Networks},\n  year={2018},\n  cdate={1514764800000},\n  journal={CoRR},\n  volume={abs/1810.05095},\n  url={http://arxiv.org/abs/1810.05095}\n}\n"},"abstract":{"value":"In the last 15 years, statistical physics has been a very successful framework to model complex networks. On the theoretical side, this approach has brought novel insights into a variety of physical phenomena, such as self-organisation, scale invariance, emergence of mixed distributions and ensemble non-equivalence, that display unconventional features on heterogeneous networks. At the same time, thanks to their deep connection with information theory, statistical physics and the principle of maximum entropy have led to the definition of null models for networks reproducing some features of real-world systems, but otherwise as random as possible. We review here the statistical physics approach and the various null models for complex networks, focusing in particular on the analytic frameworks reproducing the local network features. We then show how these models have been used to detect statistically significant and predictive structural patterns in real-world networks, as well as to reconstruct the network structure in case of incomplete information. We further survey the statistical physics models that reproduce more complex, semi-local network features using Markov chain Monte Carlo sampling, as well as the models of generalised network structures such as multiplex networks, interacting networks and simplicial complexes."},"title":{"value":"The Statistical Physics of Real-World Networks"},"authors":{"value":["Giulio Cimini","Tiziano Squartini","Fabio Saracco","Diego Garlaschelli","Andrea Gabrielli","Guido Caldarelli"]}},"tmdate":1747118600577,"pdate":1514764800000,"tcdate":1747118566796,"writers":["~"],"signatures":["~Giulio_Cimini1"],"forum":"cf8T7Gxq0s","license":"CC BY-SA 4.0","number":439123,"cdate":1514764800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747118600577,"domain":"DBLP.org","id":"cf8T7Gxq0s","version":2},{"content":{"summary":{"value":"The proposed AtomWorld is a benchmark designed to evaluate spatial reasoning capabilities of LLMs in the domain of crystalline materials. The authors assume that LLMs possess the capability to understand and manipulate 3D atomic structures for crystal design. They designed a series of cognitive tasks in materials discovery and benchmarked several state-of-the-art LLMs, including Gemini 2.5 Pro, GPT-o3, GPT-o4-mini, Deepseek Chat, Llama-3 70B, and Qwen-3. All benchmark tasks uses CIF format as textual input of the LLMs along with a natural language action prompt, requiring certain actions to edit the CIF in order to output a valid modified CIF structure. The benchmark results suggest high success rates on simple operations like adding or changing single atoms, while failures on rotations and atom swapping. The work try to disentangle different failure modes, but the analysis suggests LLMs are tackling the tasks as pattern-matching rather than physics-informed reasoning."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. All demonstrated examples use relatively simple crystals with at most three elements. Have you benchmarked LLMs on materials with higher compositional complexity or complex crystal systems, e.g. from the Materials Project database? \n2. The CIF format contains many redundant information and can produce lengthy token sequences. Have you evaluated the LLMs' performance on token length of the CIFs? How does the token counts distribute for each task? How does performance correlate with the length of the input CIF file? Does the context window limitations affect performance of the LLMs?\n3. Specifically when atom coordinates involved, have you tested whether generating a structure where the target atom appears early or late in the CIF sequence lead to different results? \n4. For StructProp, have you tested whether models can correctly identify which property is being referenced?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The work aims at building a comprehensive benchmark that assesses whether LLMs can handle mechanical operations and higher-level cognitive tasks. It is working at building the foundation for future material discovery with LLM capabilities. \n2. Multiple LLMs across different model families are benchmarked in this work. And it reported an interesting observation on how all LLMs consistently fail at complex spatial reasoning tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The benchmarking data are oversimplified. This work evaluates LLMs exclusively on simple prototype crystals with at most three elements, which raises questions about whether the findings generalize to multi-component systems or more complex lattice structures. \n2. In addition, the benchmark rely on extremely small sample sizes often tens or hundreds. E.g. CIF-Repair was tested on 22 samples, CIF-Gen on 20 samples, and only 10 samples ran DFT in StructProp. Therefore the conclusions drawn from these tests about model capabilities or reasoning quality are not reliable.\n3. The StructProp benchmark design is intuitive but fundamentally flawed, as it assumes LLMs understand the property concepts being tested without validating this assumption. Interpretation of the StructProp results can only be regarded as bias from memorized patterns."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942282662,"tcdate":1762129249664,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22569/Reviewer_iYfW"],"signatures":["ICLR.cc/2026/Conference/Submission22569/Reviewer_iYfW"],"forum":"zvzrDQTk1n","number":4,"license":"CC BY 4.0","cdate":1762129249664,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22569/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942282662,"domain":"ICLR.cc/2026/Conference","replyto":"zvzrDQTk1n","id":"3syyfPKI04","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large Language Models","Spatial Reasoning","Benchmark Evaluation","Materials Science","Crystalline Materials","Geometric Operations"]},"supplementary_material":{"value":"/attachment/b108f855d60df5c6b0583c10f6288472a84ed8e4.pdf"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large Language Models (LLMs) excel at textual reasoning and are beginning to\ndevelop spatial understanding, prompting the question of whether these abilities\ncan be combined for complex, domain-specific tasks. This question is essential in\nfields like materials science, where deep understanding of 3D atomic structures\nis fundamental. While initial studies have successfully applied LLMs to tasks\ninvolving pure crystal generation or coordinate understandings, a standardized\nbenchmark to systematically evaluate their core reasoning abilities across diverse\natomic structures has been notably absent. To address this gap, we introduce\nthe AtomWorld benchmark to evaluate LLMs on tasks based in Crystallographic\nInformation Files (CIFs), a standard structure representation format. These tasks,\nincluding structural editing, CIF perception, and property-guided modeling, reveal\na critical limitation: current models, despite establishing promising baselines,\nconsistently fail in structural understanding and spatial reasoning. Our experiments\nshow that these models make frequent errors on structure modification tasks, and\neven in the basic CIF format understandings, potentially leading to cumulative\nerrors in subsequent analysis and materials insights. By defining these standardized\ntasks, AtomWorld lays the ground for advancing LLMs toward robust atomic-scale\nmodeling, crucial for accelerating materials research and automating scientific\nworkflows."},"_bibtex":{"value":"@misc{\nlv2026atomworld,\ntitle={AtomWorld: Benchmarking Spatial Reasoning in Large Language Models on Crystalline Materials},\nauthor={Taoyuze Lv and Alexander Chen and Fengyu Xie and Jeffrey Meng and Dongzhan Zhou and Bram Hoex and Zhicheng Zhong and Tong Xie},\nyear={2026},\nurl={https://openreview.net/forum?id=zvzrDQTk1n}\n}"},"title":{"value":"AtomWorld: Benchmarking Spatial Reasoning in Large Language Models on Crystalline Materials"},"pdf":{"value":"/pdf/a75f40dff50522a13c5ff1351f90d235b7b7f78d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"lv|atomworld_benchmarking_spatial_reasoning_in_large_language_models_on_crystalline_materials"},"authorids":{"value":["~Taoyuze_Lv1","~Alexander_Chen3","~Fengyu_Xie1","~Jeffrey_Meng1","~Dongzhan_Zhou1","~Bram_Hoex1","~Zhicheng_Zhong2","~Tong_Xie2"]},"authors":{"value":["Taoyuze Lv","Alexander Chen","Fengyu Xie","Jeffrey Meng","Dongzhan Zhou","Bram Hoex","Zhicheng Zhong","Tong Xie"]}},"version":2},{"content":{"summary":{"value":"This paper presents a new approach to generating synthetic face datasets for training face recognition (FR) models. The paper proposes a triplet based ID preservation loss to finetune the diffusion model. While the paper makes a strong case for its method, several areas require further clarification and improvement."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"Addressed in weakness section."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"Strength:\nTriplet Identity Training Objective: This is the central contribution. The authors introduce a novel triplet loss that considers both positive (images of the target identity) and negative (images of other identities) examples during training. The approach attempts to avoid overfitting to the training data while still ensuring high-quality synthetic images and adequate inter-identity separability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weakness:\nWeak Face recognition Baseline: The FR model trained on synthetic dataset is under-performing previous methods. Arc2Face or DCFace creates synthetic dataset for training FR models and the FR performance can be much higher than Table 3. The low performance needs an explanation.\nSpeed: To training large scale dataset, the model needs to be able to generating images efficiently. However, this method requires 1. finetuning and 2. subsequently generating which slows down the sampling speed. Speed analysis to show how much portion of time is spent on finetuning is spend would clarify this."}},"nonreaders":[],"tmdate":1731429045187,"tcdate":1730783758376,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13546/Reviewer_CvtJ"],"signatures":["ICLR.cc/2025/Conference/Submission13546/Reviewer_CvtJ"],"forum":"NWvsm2VxAM","number":4,"license":"CC BY 4.0","cdate":1730783758376,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13546/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429045187,"domain":"ICLR.cc/2025/Conference","replyto":"NWvsm2VxAM","id":"YqhFbkYtqZ","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Image synthesis","Diffusion models","Face recognition data"]},"supplementary_material":{"value":"/attachment/ced7b745e8a0ea638139919e948ed62e6249223d.pdf"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The recent retraction of large-scale biometric datasets, prompted by strict privacy  regulations, presents a critical challenge for future biometric research. This is evident with the face recognition task, for which large-scale datasets were often gathered through web-scraping without the consent of subjects. A potential solution entails the creation of synthetic data, suitable for training recognition models, with deep generative models. Existing generative approaches rely on conditioning and fine-tuning of powerful pretrained diffusion models to achieve the synthesis of realistic images of a desired identity. Yet, these methods often do not consider the identity of subjects during training, leading to poor consistency between generated and intended identities. In contrast, methods that employ identity-based training objectives tend to overfit on various aspects of the identity, and in turn, lower the diversity of images that can be generated. To address these issues, we present the ID-Booth fine-tuning framework, which utilizes a novel triplet identity training objective and enables identity-consistent image generation while retaining the synthesis capabilities of pretrained models. Experiments across two latent diffusion models with varying prompt complexity reveal that our method facilitates better intra-identity consistency and inter-identity separability while achieving higher image diversity. In turn, the produced data enables the training of better-performing recognition models than even real-world  datasets of a similar scale gathered with suitable consent. The source code for the ID-Booth framework is available at omitted_for_review."},"_bibtex":{"value":"@misc{\ntoma{\\v{s}}evi{\\'c}2024idbooth,\ntitle={{ID}-Booth: Identity-consistent image generation with diffusion models},\nauthor={Darian Toma{\\v{s}}evi{\\'c} and Fadi Boutros and Naser Damer and Vitomir Struc and Peter Peer},\nyear={2024},\nurl={https://openreview.net/forum?id=NWvsm2VxAM}\n}"},"title":{"value":"ID-Booth: Identity-consistent image generation with diffusion models"},"pdf":{"value":"/pdf/a8f446878873370d703c5633e7fd9c6efe4bd72c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"tomaevi|idbooth_identityconsistent_image_generation_with_diffusion_models"},"authorids":{"value":["~Darian_Tomašević1","~Fadi_Boutros1","~Naser_Damer1","~Vitomir_Struc1","~Peter_Peer1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Darian Tomašević","Fadi Boutros","Naser Damer","Vitomir Struc","Peter Peer"]}},"version":2},{"content":{"summary":{"value":"The paper present a method to simulate realistic surgical scene. The core method uses a physics-based Maxwell model to restore the complex deformations of tissues. The proposed system aims to improve the surgical training, planning, and robotic surgery systems by offering accurate surgery scene simulation."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. One major confusing point is: this paper states to simulate the dynamic surgery from monocular videos; however, why are the authors use stereo surgery videos, such as EndoNerf, to evaluate the proposed method's performance?\n\n2. Why put the quantitative evaluations with EndoNerf in the Appendix? The core evaluation results should be placed in the main paper."},"rating":{"value":3},"details_of_ethics_concerns":{"value":"Does not provide the IRB approval statement as the paper involves human subjects in user study."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The user study part is good and ensures the usability of the proposed system."},"flag_for_ethics_review":{"value":["Yes, Responsible research practice (e.g., human subjects, data release)"]},"weaknesses":{"value":"1. The methodology presented in this paper similar to the paper\"SimEndoGS: Efficient Data-driven Scene Simulation using Robotic Surgery Videos via Physics-embedded 3D Gaussians\" with no substantial novel improvements. The evaluation in Figure 3 does not suggest advancement in this work when comparing with the SimEndoGS.\n\n2. Lacks comprehensive quantitative evaluation of the proposed method itself, e.g. did not compare with SimEndoGS.\n\n3. Lacks significant testing."}},"nonreaders":[],"tmdate":1731427658534,"tcdate":1730052593059,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2800/Reviewer_Mqpg"],"signatures":["ICLR.cc/2025/Conference/Submission2800/Reviewer_Mqpg"],"forum":"faSfhqDpZP","number":1,"license":"CC BY 4.0","cdate":1730052593059,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2800/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427658534,"domain":"ICLR.cc/2025/Conference","replyto":"faSfhqDpZP","id":"H7TZyqaCwP","forumContent":{"TLDR":{"value":"We introduce an automatic system that reconstructs geomery consistent simulation scenes from monocular videos and perform simulation with Visco-Elastic physics model."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Surgical Simulation","Video-based Reconstruction","Robotic Surgery"]},"supplementary_material":{"value":"/attachment/c829d7810476857cd705d2a30706d4b5e4932c03.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper tackles the challenge of automatically constructing realistic surgical simulation systems from readily available surgical videos. Recent efforts have successfully integrated physically grounded dynamics within 3D Gaussians to perform high-fidelity simulations in well-reconstructed static simulation environments. However, they struggle with the geometry inconsistency of simulation environments and unrealistic physical deformations of soft tissues when it comes to dynamic and complex surgical processes. In this paper, we propose SurgiSim, a novel automatic simulation system to overcome these limitations. To build a surgical simulation environment, we maintain a canonical 3D scene composed of 3D Gaussians coupled with a deformation field to accurately model monocular dynamic surgical scenes. This process involves a multi-stage optimization with trajectory and anisotropic regularization, enhancing the geometry consistency of the canonical scene which serves the simulation environment. To improve the realism of physical simulations, we implement a Visco-Elastic deformation model based on the Maxwell model, effectively restoring the complex deformations of tissues. Additionally, we estimate the physical properties of tissues by minimizing the discrepancies between the input video and simulation results guided by predicted tissue motion, ensuring realistic simulation outcomes. Experiments across diverse surgical scenarios demonstrate SurgiSim's ability to perform realistic physical interactions of soft tissues among surgical procedures, showing its enormous potential for enhancing surgical training, planning, and robotic surgery systems."},"_bibtex":{"value":"@misc{\nwang2024realistic,\ntitle={Realistic Surgical Simulation from Monocular Videos},\nauthor={Kailing Wang and Chen Yang and Keyang Zhao and Xiaokang Yang and Wei Shen},\nyear={2024},\nurl={https://openreview.net/forum?id=faSfhqDpZP}\n}"},"title":{"value":"Realistic Surgical Simulation from Monocular Videos"},"pdf":{"value":"/pdf/30f733e502cf73bacbe3d5ed11cdad3a7996b0e0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|realistic_surgical_simulation_from_monocular_videos"},"authorids":{"value":["~Kailing_Wang1","~Chen_Yang16","~Keyang_Zhao1","~Xiaokang_Yang1","~Wei_Shen2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kailing Wang","Chen Yang","Keyang Zhao","Xiaokang Yang","Wei Shen"]}},"version":2},{"content":{"summary":{"value":"The paper critiques existing ICL theoretical analyses for relying on unrealistic i.i.d. (independent and identically distributed) assumptions and lacking an explanation of ICL emergence. To address these issues, the authors propose an auto-regressive next-token prediction (AR-NTP) framework, which mirrors real-world language learning by considering token dependencies within prompts. The paper introduces a pre-training and ICL framework, analyzing generalization bounds using PAC-Bayesian techniques and exploring how ICL arises through generalization of sequences and topics. The findings are validated through experiments on synthetic and real-world datasets, concluding that ICL capabilities emerge from the strong generalization of LLMs trained on diverse sequences and topics."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Why particularly choose PAC-Bayesian framework for ICL? There are numerous works study ICL using statistical frameworks like [1]\nWhy not based on uniform convergence or algorithmic stability, could help situate the results within a broader landscape of generalization theory?\n\n[1] Transformers as statisticians: Provable in-context learning with in-context algorithm selection"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This paper presents an original contribution to the understanding of in-context learning (ICL) by introducing the auto-regressive next-token prediction (AR-NTP) framework. The paper seek to addresses two major limitations in the current literature: the i.i.d. assumption in prompt tokens and the lack of an explanation for how ICL emerges from pre-trained large language models (LLMs). By shifting the focus to token dependencies in AR-NTP and offering a novel pre-training and ICL framework, the authors provide a fresh perspective on ICL emergence through generalization. The originality stems from the use of PAC-Bayesian generalization techniques, and the proposed framework opens up new avenues for exploring token dependencies and generalization beyond previous work that relied on supervised function learning. The authors also provide theoretical results that are supported by empirical experiments on both synthetic and real-world datasets, enhancing the quality and rigor of their contributions.\n\nthe paper is well-structured and flows well, guiding the reader through complex concepts such as generalization bounds, topic-dependent priors, and PAC-Bayesian techniques without overwhelming them with technical jargon. The theoretical framework is detailed yet accessible, and the results are presented in a way that is easy to follow, particularly with well-placed diagrams and examples. The significance of the work lies in its ability to explain ICL as an emergent property of pre-trained LLMs, which has strong implications for both academia and industry. The results contribute to understanding why larger models exhibit better ICL performance and provide concrete recommendations for improving model generalization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the paper makes significant strides in addressing limitations in existing ICL literature, there are several areas where the work could be strengthened. One key weakness lies in the assumptions behind the AR-NTP paradigm. Although the shift from i.i.d. settings to token-dependent sequences is crucial, the paper could have provided a more thorough discussion on the potential challenges this poses for practical implementation. For example, the dependency between tokens may introduce computational inefficiencies in scaling to larger datasets or more complex sequences, which could affect the generalization results. A more detailed exploration of how to handle these challenges efficiently during training would improve the practical applicability of the proposed framework. \n\nAltough the paper is mostly theoretical, there is no enough experiments in the main paper. Especially the claim that topic prior in pre-training would affect the ICL does not get justification. Appendix contains some synthetic topic results using GPT-2 but that is a bit too hand-wavy. \nThe authors validate their theoretical claims with experiments on synthetic and real-world datasets, which is commendable, but the diversity of the real-world datasets could be expanded. For example, the selection of datasets appears limited to common language tasks like sentiment analysis and topic classification. Including a wider variety of tasks such as training on Zipfian data but test on distributions like MMLU and GSM8k would better demonstrate the generalization and applicability of the proposed AR-NTP paradigm across different domains. Moreover, while the results show that increasing the number of pre-training topics, sequences per topic, and sequence length improves model performance, the experiments would benefit from more ablation studies. \n\nLastly, while the paper provides theoretical insights into how ICL emerges from pre-training, the connection between theory and practice could be more tightly integrated. The authors propose that ICL arises from excellent generalization of sequences and topics, but the specific impact of pre-training data diversity, model size, and optimization iterations could be made clearer. The paper could explore more explicitly how different pre-training regimes—such as training on more complex or diverse datasets [1] —might affect ICL performance in practice. This would provide practitioners with clearer guidelines on how to optimize their pre-training procedures for better ICL outcomes. In particular, the paper could explore limitations of their generalization bounds when models are applied in low-resource environments, where generalization may suffer in no enough ICL examples are given or the pre-training FLOPs are constrained.\n\n[1] Data Distributional Properties Drive Emergent In-Context Learning in Transformers Stephanie C.Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang, Aaditya Singh, Pierre H. Richemond, Jay McClelland, Felix Hill"}},"nonreaders":[],"tmdate":1732374890218,"tcdate":1729997035218,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7103/Reviewer_x2ai"],"signatures":["ICLR.cc/2025/Conference/Submission7103/Reviewer_x2ai"],"forum":"gK1rl98VRp","number":2,"license":"CC BY 4.0","cdate":1729997035218,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7103/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732374890218,"domain":"ICLR.cc/2025/Conference","replyto":"gK1rl98VRp","id":"hWsMX1mfm0","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["In-context learning","Auto-regressive next-token prediction","Generalization performance","PAC-Bayesian"]},"primary_area":{"value":"learning theory"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large language models (LLMs) have demonstrated remarkable in-context learning (ICL) abilities. However, existing theoretical analysis of ICL primarily exhibits two limitations: \\textbf{(a) Limited \\textit{i.i.d.} Setting.} Most studies focus on supervised function learning tasks where prompts are constructed with \\textit{i.i.d.} input-label pairs. This \\textit{i.i.d.} assumption diverges significantly from real language learning scenarios where prompt tokens are interdependent. \\textbf{(b) Lack of Emergence Explanation.} Most literature answers \\textbf{\\textit{what}} ICL does from an implicit optimization perspective but falls short in elucidating \\textbf{\\textit{how}} ICL emerges and the impact of pre-training phase on ICL. In our paper, to extend (a), we adopt a more practical paradigm, \\textbf{\\textit{auto-regressive next-token prediction (AR-NTP)}}, which closely aligns with the actual training of language models. Specifically, within AR-NTP, we emphasize prompt token-dependency, which involves predicting each subsequent token based on the preceding sequence. To address (b), we formalize a systematic pre-training and ICL framework, highlighting the layer-wise structure of sequences and topics, alongside a two-level expectation. In conclusion, we present data-dependent, topic-dependent and optimization-dependent PAC-Bayesian generalization bounds for pre-trained LLMs, investigating that \\textbf{\\textit{ICL emerges from the generalization of sequences and topics}}. Our theory is supported by experiments on numerical linear dynamic systems, synthetic GINC and real-world language datasets."},"_bibtex":{"value":"@inproceedings{\ngong2025towards,\ntitle={Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization},\nauthor={Zixuan Gong and Xiaolin Hu and Huayi Tang and Yong Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=gK1rl98VRp}\n}"},"title":{"value":"Towards Auto-Regressive Next-Token Prediction: In-context Learning Emerges from Generalization"},"pdf":{"value":"/pdf/51093c4c47c281e78f492bae809d9bad6f9af6a8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"gong|towards_autoregressive_nexttoken_prediction_incontext_learning_emerges_from_generalization"},"authorids":{"value":["~Zixuan_Gong1","~Xiaolin_Hu6","~Huayi_Tang1","~Yong_Liu7"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zixuan Gong","Xiaolin Hu","Huayi Tang","Yong Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes an LLM-based approach for detecting bugs with data flow analysis. The proposed approach LLMDFA consists of three parts: source/sinks extraction, dataflow summarization, and path feasibility validation. LLMDFA allows LLMs to interact with different program analysis tools. The authors conduct experiments on bug detection in both synthetic and real-world datasets."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Most baselines are non-deep-learning approaches. I wonder if there are any deep learning-based approaches previously proposed for similar bug detection tasks?\n\n- In my understanding, it is possible to write a single parser to extract the sources/sinks for all programs. What's the advantage in using LLM to generate different parsers for different programs? Also, I'd like to see more discussion on the impact of ASTs in source/sink extraction.\n\n- I’m a little confused about how the results of path feasibility validation can lead to the detection of bugs. Can the authors explain in more detail?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"+The authors proposed an interesting approach for solving an important question in program languages. \n\n+LLMDFA decomposes bug detection into three steps, which are proven useful for complex data flow reasoning.\n\n+Experiments are conducted on both synthetic and real-world datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-The results on synthetic datasets are very high, maybe the dataset is too easy?\n\n-All components in LLMDFA require multiple rounds of try-and-refine. It is better to discuss the cost of multiple-round generation."},"limitations":{"value":"The limitations have been well-discussed."}},"nonreaders":[],"tmdate":1730879167309,"tcdate":1718226047167,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission7502/Reviewer_UBFi"],"signatures":["NeurIPS.cc/2024/Conference/Submission7502/Reviewer_UBFi"],"forum":"QZ2d8E8Whu","number":1,"license":"CC BY 4.0","cdate":1718226047167,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission7502/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879167309,"domain":"NeurIPS.cc/2024/Conference","replyto":"QZ2d8E8Whu","id":"dp7AasgeOS","forumContent":{"TLDR":{"value":"Propose an LLM-powered dataflow analysis, showing the potential of LLMs in resolving complex code reasoning problems."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["LLM for code","code reasoning","dataflow analysis"]},"primary_area":{"value":"machine_learning_for_other_sciences_and_fields"},"abstract":{"value":"Dataflow analysis is a fundamental code analysis technique that identifies dependencies between program values. Traditional approaches typically necessitate successful compilation and expert customization, hindering their applicability and usability for analyzing uncompilable programs with evolving analysis needs in real-world scenarios. This paper presents LLMDFA, an LLM-powered compilation-free and customizable dataflow analysis framework. To address hallucinations for reliable results, we decompose the problem into several subtasks and introduce a series of novel strategies. Specifically, we leverage LLMs to synthesize code that outsources delicate reasoning to external expert tools, such as using a parsing library to extract program values of interest and invoking an automated theorem prover to validate path feasibility. Additionally, we adopt a few-shot chain-of-thought prompting to summarize dataflow facts in individual functions, aligning the LLMs with the program semantics of small code snippets to mitigate hallucinations. We evaluate LLMDFA on synthetic programs to detect three representative types of bugs and on real-world Android applications for customized bug detection. On average, LLMDFA achieves 87.10% precision and 80.77% recall, surpassing existing techniques with F1 score improvements of up to 0.35. We have open-sourced LLMDFA at https://github.com/chengpeng-wang/LLMDFA."},"_bibtex":{"value":"@inproceedings{\nwang2024llmdfa,\ntitle={{LLMDFA}: Analyzing Dataflow in Code with Large Language Models},\nauthor={Chengpeng Wang and Wuqi Zhang and Zian Su and Xiangzhe Xu and Xiaoheng Xie and Xiangyu Zhang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=QZ2d8E8Whu}\n}"},"title":{"value":"LLMDFA: Analyzing Dataflow in Code with Large Language Models"},"pdf":{"value":"/pdf/2cfd2ed22ae0deee46b1fb6367fa370ba0fca4e6.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"wang|llmdfa_analyzing_dataflow_in_code_with_large_language_models"},"authorids":{"value":["~Chengpeng_Wang2","~Wuqi_Zhang2","~Zian_Su1","~Xiangzhe_Xu1","~Xiaoheng_Xie2","~Xiangyu_Zhang3"]},"authors":{"value":["Chengpeng Wang","Wuqi Zhang","Zian Su","Xiangzhe Xu","Xiaoheng Xie","Xiangyu Zhang"]}},"version":2},{"content":{"venue":{"value":"Journal of Computational Physics"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"liu|binary_structured_physicsinformed_neural_networks_for_solving_equations_with_rapidly_changing_solutions"},"authorids":{"value":["~Yanzhi_Liu1","~Ruifan_Wu1","~Ying_Jiang1"]},"abstract":{"value":"Physics-informed neural networks (PINNs), rooted in deep learning, have emerged as a promising approach for solving partial differential equations (PDEs). By embedding the physical information described by PDEs into feedforward neural networks, PINNs are trained as surrogate models to ap\u0002proximate solutions without the need for label data. Nevertheless, even though PINNs have shown remarkable performance, they can face difficulties, especially when dealing with equations featur\u0002ing rapidly changing solutions. These difficulties encompass slow convergence, susceptibility to becoming trapped in local minima, and reduced solution accuracy. To address these issues, we pro\u0002pose a binary structured physics-informed neural network (BsPINN) framework, which employs binary structured neural network (BsNN) as the neural network component. By leveraging a binary structure that reduces inter-neuron connections compared to fully connected neural networks, BsPINNs excel in capturing the local features of solutions more effectively and efficiently. These features are particularly crucial for learning the rapidly changing in the nature of solutions. In a series of numerical experiments solving the Euler equation, Burgers equation, Helmholtz equation, and high-dimension Poisson equation, BsPINNs exhibit superior convergence speed and heightened accuracy compared to PINNs. Additionally, we compare BsPINNs with XPINNs, FBPINNs and FourierPINNs, finding that BsPINNs achieve the highest or comparable results in many experiments. Furthermore, our experiments reveal that integrating BsPINNs with Fourier feature mappings results in the most accurate predicted solutions when solving the Burgers equation and the Helmholtz equation on a 3D Möbius knot. This demonstrates the potential of BsPINNs as a foundational model. From these experiments, we discover that BsPINNs resolve the issues caused by increased hidden layers in PINNs resulting in over-fitting, and prevent the decline in accuracy due to non-smoothness of PDEs solutions."},"title":{"value":"Binary structured physics-informed neural networks for solving equations with rapidly changing solutions"},"authors":{"value":["Yanzhi Liu","Ruifan Wu","Ying Jiang"]}},"tmdate":1782268851517,"pdate":1722441600000,"tcdate":1782268513980,"writers":["~Yanzhi_Liu1","~Ruifan_Wu1","~Ying_Jiang1"],"signatures":["~Ying_Jiang1"],"forum":"rOjvmFAhlV","license":"CC BY 4.0","number":51137,"cdate":1782268513980,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1782268851517,"domain":"OpenReview.net/Archive","id":"rOjvmFAhlV","version":2},{"content":{"summary":{"value":"This paper targets the challenge of pose-guided portrait animation under complex human motions (e.g., flips, stunts) and proposes a concise DiT paradigm: it composes in latent space to inject both pose video and reference image, and uses a first-frame mask to suppress identity leakage. The core modification, SLF-RoPE, selectively enhances the low-frequency channels of RoPE along the spatial dimension, while learnable motion/space scaling factors improve global structure and identity stability under high-speed, nonlinear motion. Based on MotionX, the authors construct the Open-HyperMotionX dataset (automatically mining complex motion segments via wavelet energy, with OCR debiasing and caption cleaning) and release HyperMotionX Bench with 100 high-quality pose annotations (using XPose, removing unreliable hand keypoints in complex frames). Trained on the Wan2.1 backbone, the method achieves leading structural consistency (PCK), competitive pixel/perceptual/temporal metrics, and VBench-I2V gains in background/overall consistency and motion smoothness. Ablations show SLF-RoPE effectively reduces artifacts under extreme poses, and inference runs on a single 3090. Limitations include limited theoretical novelty, insufficient analysis of sensitivity to α/γ and the dynamics of the proposed scaling strategy, coarse-grained ablations, heavy reliance on XPose and hand keypoint quality, and incomplete evaluation for multi-person scenes, strong camera motion, and long sequences. The authors plan to open-source code, data, and evaluation, indicating solid engineering practicality and potential impact."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Method is simple yet targeted: latent-space composition in a DiT and a first-frame mask suppress identity leakage; SLF-RoPE selectively boosts spatial low-frequency channels and introduces learnable motion/space scaling, significantly improving global structure and identity stability under high-speed, non-linear motion.\n2. Significant data and evaluation contributions: constructed Open-HyperMotionX (automatic mining of complex motions, OCR debiasing, and caption cleaning) and released the high-quality HyperMotionX Bench (100 sequences with pose annotations), providing practical training and evaluation resources for complex motion scenarios.\n3. High engineering practicality: inference runs on a single 3090, with minimal modifications and easy integration into backbones like Wan2.1; plans to open-source code, data, and evaluation, facilitating community reproduction and extension."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited methodological novelty: The core modification (SLF-RoPE) scales the low-frequency band of RoPE’s frequencies. The idea is intuitive and engineering-oriented, but lacks theoretical depth and a justification of generality.\n\n2. I would like to see more visual results for this task, especially hand details (which are one of the core aspects). Please provide more examples focusing on hand motion.\n\n3. The sensitivity of the low-frequency ratio α and the scaling factor γ, as well as the learning dynamics of the motion scale and space scale under different motion intensities, remain unclear. How “motion intensity” is estimated and the consistency between training and inference phases are not elaborated.\n\n4. The paper only provides three comparisons (with/without SLF-RoPE and with/without dataset training). More fine-grained ablations are desirable (e.g., scaling H only or W only, fixing γ without dynamic modulation, using different α partitioning strategies, etc.).\n\n5. There is heavy reliance on XPose for pose annotations; hand keypoints are removed in complex segments. While this is reasonable, it limits the evaluation and optimization of fine-grained hand motion generation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915718501,"tcdate":1761496650604,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1255/Reviewer_rE7V"],"signatures":["ICLR.cc/2026/Conference/Submission1255/Reviewer_rE7V"],"forum":"pzRDviUmbH","number":1,"license":"CC BY 4.0","cdate":1761496650604,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1255/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915718501,"domain":"ICLR.cc/2026/Conference","replyto":"pzRDviUmbH","id":"wPziDksnb1","forumContent":{"TLDR":{"value":"DiT-based Pose-Guided Human  Image Animation"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Generation","Human Image Animation"]},"supplementary_material":{"value":"/attachment/914017d4ac5575fe6b734ac0142694d716dcd19c.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent animation sequences in regular motions and static scenes, there are still obvious limitations when facing complex human body motions that contain highly dynamic, non-standard motions, and the lack of a high-quality benchmark for evaluation of complex human motion animations. To address this challenge, we propose a simple yet powerful DiT-based video generation baseline and design spatial low-frequency enhanced RoPE, a novel module that selectively enhances low-frequency spatial feature modeling by introducing learnable frequency scaling. Furthermore, we introduce the Open-HyperMotionX Dataset and HyperMotionX Bench, which provide high-quality human pose annotations and curated video clips for evaluating and improving pose-guided human image animation models under complex human motion conditions. Our method significantly improves structural stability and appearance consistency in highly dynamic human motion sequences.  Extensive experiments demonstrate the effectiveness of our dataset and proposed approach in advancing the generation quality of complex human motion image animations. The codes and dataset will be made publicly available."},"_bibtex":{"value":"@misc{\nxu2026hypermotion,\ntitle={HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions},\nauthor={Shuolin Xu and Siming Zheng and Ziyi Wang and HC Yu and Jinwei Chen and Huaqi Zhang and Daquan Zhou and Bo Li and Peng-Tao Jiang},\nyear={2026},\nurl={https://openreview.net/forum?id=pzRDviUmbH}\n}"},"title":{"value":"HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions"},"pdf":{"value":"/pdf/605138d2cf34ffc4989039184adfaeb72620903f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"xu|hypermotion_ditbased_poseguided_human_image_animation_of_complex_motions"},"authorids":{"value":["~Shuolin_Xu1","~Siming_Zheng3","~Ziyi_Wang27","~HC_Yu1","~Jinwei_Chen3","~Huaqi_Zhang1","~Daquan_Zhou1","~Bo_Li20","~Peng-Tao_Jiang1"]},"authors":{"value":["Shuolin Xu","Siming Zheng","Ziyi Wang","HC Yu","Jinwei Chen","Huaqi Zhang","Daquan Zhou","Bo Li","Peng-Tao Jiang"]}},"version":2},{"content":{"venue":{"value":"Nature Communications"},"pdf":{"value":"https://www.nature.com/articles/s41467-026-71673-9.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"malik|hybrid_physicsmachine_learning_models_for_quantitative_electron_diffraction_refinements"},"html":{"value":"https://doi.org/10.1038/s41467-026-71673-9"},"abstract":{"value":"High accuracy electron microscopy simulations required for quantitative crystal structure refinements face a fundamental challenge: while physical interactions are well-described theoretically, real-world experimental effects are challenging to model analytically. To address this gap, we present a hybrid physics-machine learning framework that integrates differentiable physical simulations with neural networks. By leveraging automatic differentiation throughout the simulation pipeline, our method enables gradient-based joint optimization of physical parameters and neural network components representing experimental variables, offering superior scalability compared to traditional second-order methods. We demonstrate this framework through application to three-dimensional electron diffraction (3D-ED) structure refinement, where our approach learns complex thickness distributions directly from diffraction data rather than relying on simplified geometric models. This method achieves state-of-the-art refinement performance across synthetic and experimental datasets, recovering atomic positions, thermal displacements, and thickness profiles with high fidelity. The modular architecture proposed can naturally be extended to accommodate additional physical phenomena and extended to other electron microscopy techniques. This establishes differentiable hybrid modeling as a powerful paradigm for quantitative electron microscopy, where experimental complexities have historically limited analysis. A hybrid physics-machine learning framework enables scalable dynamical refinement of 3D-ED data by combining differentiable diffraction simulations with neural networks to jointly refine crystal structures and complex experimental parameters."},"title":{"value":"Hybrid physics-machine learning models for quantitative electron diffraction refinements"},"authors":{"value":[{"fullname":"Shreshth A. Malik","username":"~Shreshth_A_Malik1"},{"fullname":"Tiarnan A. S. Doherty","username":"https://orcid.org/orcid-search/search?searchQuery=Tiarnan%20A.%20S.%20Doherty"},{"fullname":"Benjamin Colmey","username":"https://orcid.org/orcid-search/search?searchQuery=Benjamin%20Colmey"},{"fullname":"Stephen J. Roberts","username":"https://orcid.org/orcid-search/search?searchQuery=Stephen%20J.%20Roberts"},{"fullname":"Yarin Gal","username":"https://orcid.org/orcid-search/search?searchQuery=Yarin%20Gal"},{"fullname":"Paul A. Midgley","username":"https://orcid.org/orcid-search/search?searchQuery=Paul%20A.%20Midgley"}]}},"tmdate":1789665925164,"pdate":1775865600000,"externalIds":["doi:10.1038/s41467-026-71673-9"],"tcdate":1789665916393,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Shreshth_A_Malik1"],"forum":"boLuRbiOnY","license":"CC BY-SA 4.0","number":110507,"cdate":1775929394298,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1789665925164,"domain":"OpenReview.net/Public_Article","id":"boLuRbiOnY","version":2},{"content":{"summary":{"value":"This paper revisits the framework of Conditional Whitney Forms (CWFs). The authors argue that the original CWF formulation leads to trivial conservation laws and propose a flux-regularized reformulation to enable true physics recovery. They propose a solution by reformulating the learning problem to include a flux reconstruction term as a regularizer. The theoretical claims are supported by two propositions and experimental validation on a series of advection-diffusion problems."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- The propositions assume an unconstrained $\\(\\hat{f}\\)$. However, in practice, \\(\\hat{f}\\) is output by a neural network conditioned on $\\(\\hat{u}\\)$. Can you provide empirical or theoretical evidence that the *coupled learning process* of $\\(W(z;\\theta)\\)$ and $\\(\\mathcal{NN}(\\hat{u}; \\phi, z_i)\\)$ is inherently biased towards the \"trivial\" solutions you construct, rather than just being capable of representing them?\n\n- By relaxing the hard physics constraint to a soft penalty, the method no longer provides exact conservation. How do you reconcile this with the stated goal of \"structure preservation\"? What are the quantitative trade-offs between the flux reconstruction accuracy and the exact satisfaction of the conservation law?\n\n- How does the method compare to other physics-informed or structure-preserving approaches (e.g., conservative PINNs, finite volume networks) in terms of accuracy, stability, and computational cost? A comparison with pure data-driven models should also be included.\n\n- The method is intrinsically dependent on a pre-defined, fixed mesh for the fine-grained partition. How would you apply this approach to a synthetic/real-world dataset where inputs come from irregular geometries or inconsistent discretizations? Does this not severely limit the applicability of the method as a general neural operator?\n\n- Have the author tested the method on more complex systems (e.g., 3D flows, systems with shocks, or multiphysics problems)? If not, what are the anticipated challenges?\n\n- Given that the flux regularization is a simple L2-loss, what is the novel *architectural* or *theoretical* insight here beyond a standard multi-task learning setup? Why is the CWF framework necessary to implement this idea?\n\n- The paper included one sensitivity study of the choice of $\\lambda$, and is there a principled way to select it for a new problem? \n\n- Explain why the unregularized CWF achieves near-perfect distribution reconstruction while failing completely at flux reconstruction. The significant discrepancy in the magnitude of errors between the two results raises concerns about their reliability.\n\n- Present the main architecture diagram of the model. Provide information on computational efficiency (e.g., the running time and memory) and the detailed model configuration summary table."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Timely Topic: The work addresses a relevant and emerging topic at the intersection of Finite Element Exterior Calculus (FEEC) and operator learning, focusing on the crucial aspect of physical structure preservation.\n\n- Clear Identification of a Problem: The paper successfully draws attention to a potential pitfall in the CWF framework (the possibility of fulfilling conservation laws in a physically meaningless way).\n\n- Practical Solution: The proposed flux regularization is a simple and intuitive first step to address the identified problem."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Major Technical Weakness**\n\n- Limited Conceptual Contribution: The core observation, that the original CWF constraint is underdetermined and allows trivial solutions, is more of a theoretical oversight in the original work than a novel insight. The proposed fix (adding a flux reconstruction loss) is a straightforward application of multi-task learning and does not introduce a new methodological or theoretical framework.\n\n- Superficial Theoretical Contribution: The core theoretical contribution (Propositions 1 and 2), while mathematically correct, is conceptually shallow. It demonstrates that an *unconstrained* flux field $\\hat{f}$ can be computed in a post-hoc manner to satisfy a discrete conservation law for any given $u$ and $s$. However, in the original CWF formulation by Kinch et al., the flux is \\emph{not} unconstrained; rather, it is parametrized by a neural network $\\mathcal{NN}(\\hat{u}; \\phi, z_i)$, whose structure and training are tightly coupled with the learning of the partition of unity. The paper fails to prove that this *coupled, learned* system will inevitably converge to the trivial solutions described in the propositions. It only shows that such solutions *exist*, which is a significantly weaker claim. As a result, the presented ``theoretical insight'' is better characterized as an \\emph{observation} regarding underdetermined linear systems, rather than a rigorous analysis of the CWF learning dynamics. Although two propositions are included on the underdetermined nature of the conservation constraint, these results are elementary from a linear algebra perspective and do not offer new theoretical understanding of the approximation properties or behavior of CWFs.\n\n**Major Conceptual Weakness**\n\n- Solution Lacks Physical Intuition and Rigor: The proposed solution, adding an L2-loss on the flux, reduces a hard physical constraint (the conservation law) to a soft, data-fitting penalty. This undermines the primary motivation of using CWF, which is to provide **guaranteed structure preservation**. The method now relies on the balance of a hyperparameter $\\(\\lambda\\)$ to *approximately* satisfy physics, a common paradigm in Physics-Informed Neural Networks (PINNs) which the authors implicitly critique. A more principled approach would be to impose structure on the flux network $\\(\\mathcal{NN}\\)$ itself (e.g., enforcing symmetry or other physical properties) to restrict the solution space without sacrificing the hard constraint.\n\n- Insufficient Analysis of Trade-offs: The paper does not adequately address the trade-off between structure preservation and model expressivity. For instance, the flux regularization term may force the model to learn overly smooth or simplistic flux representations, especially in highly nonlinear regimes, a concern only briefly mentioned in Appendix C.\n\n\n**Major Technical Flaw**\n\n - Inconsistent Error Metrics Suggest Fundamental Issues: The reported error metrics in Table 1 reveal a fundamental inconsistency that undermines the technical validity of the results. The distribution errors (e-7 to e-10) and flux errors (e-1 to e+0) differ by 7-10 orders of magnitude, which is mathematically implausible given the physical relationship between these quantities ($f$ contains $\\nabla u$). Until this discrepancy is rigorously explained and validated, the central claim, that the original CWF formulation leads to trivial conservation, rests on unreliable evidence, as the observed effect could be an artifact of the evaluation method or numerical instability rather than a profound physical insight. The authors must demonstrate that their error metrics on their $u$ and $f$ predictions are mathematically consistent and that the massive flux error is not a mere consequence of a poorly conditioned numerical scheme.\n\n- Lack of Meaningful Baselines: The comparison is limited to the regularized vs. unregularized versions of their own model. The work lacks comparisons against strong and relevant baselines, such as a pure data-driven model (e.g., a standard Transformer predicting both $\\(u\\)$ and $\\(f\\)$) or other physics-constrained architectures (e.g., conservative PINNs, finite volume-inspired methods). Without this, it is impossible to gauge the actual benefit of the complex CWF machinery over simpler approaches.\n\n- Toy Problems and Linearity: The experiments are conducted predominantly on linear or weakly nonlinear advection-diffusion problems. The claim of enabling \"physics recovery\" is not sufficiently tested against highly nonlinear or chaotic systems, where the identified problem and proposed solution would face a much greater challenge. There is no evidence that the method scales to 3D, complex geometries, or more challenging PDE systems (e.g., Navier-Stokes, wave equations). This severely limits the claimed generality of the approach.\n\n- Grid Dependency: The method is fundamentally tied to the fixed, underlying \"fine-grained partition of unity.\" The paper does not address or even acknowledge the limitation this imposes. It is unclear how the approach would generalize to synthetic datasets and real-world scenarios where training and test data come from different meshes or geometries, a key selling point of neural operators.\n\n\n**Minor issue**\n\n- Inadequate Discussion of Related Work: The discussion of structure-preserving methods is superficial. It does not engage deeply with the trade-offs between different approaches. For instance, how does the performance and guarantee of this method compare to methods that directly discretize and solve the PDEs in the loss function (e.g., FEM-based variational losses)? The claim that CWFs are the *only* framework that allows this transformation with \"minimal architectural intervention\" is strong but unsupported by a thorough comparison."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941717941,"tcdate":1761881507169,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21357/Reviewer_W241"],"signatures":["ICLR.cc/2026/Conference/Submission21357/Reviewer_W241"],"forum":"UVgNkQTScp","number":3,"license":"CC BY 4.0","cdate":1761881507169,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21357/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941717941,"domain":"ICLR.cc/2026/Conference","replyto":"UVgNkQTScp","id":"BollxfKPpm","forumContent":{"TLDR":{"value":"We show that existing conditional Whitney form formulations lead to trivial structure preservation, propose incorporating additive structure to achieve meaningful physics recovery and validate our approach experimentally."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["ai4science","physics-informed machine learning","finite elements","operator learning","interpretable scientific discovery"]},"supplementary_material":{"value":"/attachment/f3a488b6854318552fb614324c5cfe32e93e6f30.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Conditional Whitney forms have recently emerged as a promising framework at the intersection of scientific machine learning and finite element analysis. They offer a solid theoretical foundation for enforcing conservation laws in complex machine learning settings. However, their use so far has been restricted to learning tasks where structural constraints can be satisfied with simple, yet inaccurate, physics representations. In this work, we analyze why existing formulations reduce to typical unconstrained reformulations, circumventing physics recovery, and highlight the necessity of incorporating additive structure pertaining to the governing physics of the system. Based on the theoretical insights we first attain, we proceed to the reformulation of the learning problem to enable data-driven physics recovery and employ conditional Whitney forms to turn a Transformer-based architecture into a structure-preserving reduced-order model. We demonstrate the validity of our theoretical insights and the effectiveness of the subsequent proposed reformulation in a range of advection-diffusion systems of increasing difficulty. Our contributions can be viewed as a step towards understanding the capacity of conditional Whitney forms to build reliable structure-preserving models by harnessing the modeling power of state-of-the-art machine learning architectures in physical sciences."},"_bibtex":{"value":"@misc{\nkallinikidis2025revisiting,\ntitle={Revisiting Conditional Whitney Forms: From Structure Preservation to Physics Recovery},\nauthor={Pavlos Kallinikidis and Paris Perdikaris and George J. Pappas},\nyear={2025},\nurl={https://openreview.net/forum?id=UVgNkQTScp}\n}"},"title":{"value":"Revisiting Conditional Whitney Forms: From Structure Preservation to Physics Recovery"},"pdf":{"value":"/pdf/c00d44e75b31a888585ba4b5329f7effe91b703f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kallinikidis|revisiting_conditional_whitney_forms_from_structure_preservation_to_physics_recovery"},"authorids":{"value":["~Pavlos_Kallinikidis1","~Paris_Perdikaris1","~George_J._Pappas1"]},"authors":{"value":["Pavlos Kallinikidis","Paris Perdikaris","George J. Pappas"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PINNsFormer, a novel transformer-based framework for Physics-Informed Neural Networks (PINNs) to approximate solutions to partial differential equations (PDEs). PINNsFormer addresses the limitation of conventional PINNs in neglecting temporal dependencies within PDEs. Comprehensive experiments show PINNsFormer outperforms PINNs and variants in addressing failure modes and high-dimensional PDEs."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"PINNsFormer addresses a key limitation of PINNs by explicitly learning temporal dependencies, crucial for real-world physics systems. This significantly improves PINNs' generalization ability.\n\nThe proposed pseudo sequence representation and transformer architecture are clever approaches to adapt PINNs for sequential models.\n\nAblation studies provide insights into design choices and integration of existing PINNs schemes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the Wavelet activation function is theoretically justified to approximate arbitrary solutions, its advantages over other activations like ReLU, sigmoid, etc. require further empirical analysis and validation on practical problems. Conducting detailed empirical studies to evaluate Wavelet against various state-of-the-art activations under different settings can provide better insights into its benefits and limitations. This is important to fully understand its behavior and assess its effectiveness.\n\nThe paper only considers isotropic problems which have constant properties in all directions. However, most real-world physics systems exhibit anisotropic and nonlinear characteristics. Extending PINNsFormer to handle anisotropic problems modeled by direction-dependent PDEs, as well as nonlinear problems involving variable coefficients, would significantly broaden its applicability and demonstrate the approach's versatility. \n\nNo quantitative analysis was performed to evaluate important design choices like the pseudo sequence length and number of levels in coarsening. Without such ablation studies, it is difficult to justify critical hyperparameters and understand their impact on the model's performance as well as computational efficiency. These quantitative studies would provide further insights to validate the architectural design of PINNsFormer.\n\nAlthough various benchmark problems were tested, stronger validation would involve demonstrating the approach's effectiveness in entirely new physical domains beyond the existing test cases. Without such generalization to unseen problem classes, the claims regarding PINNsFormer's broad applicability remain partially unsubstantiated.\n\nWhile efficient on smaller problems, the inherent quadratic complexity of self-attention may pose scalability challenges for extremely large spatiotemporal datasets. Developing techniques to alleviate this computational limitation would enhance the method's practicality when dealing with massive real-world physics simulations."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"see weakness above"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636031165,"tcdate":1698634749636,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1048/Reviewer_mmgn"],"signatures":["ICLR.cc/2024/Conference/Submission1048/Reviewer_mmgn"],"forum":"DO2WFXU1Be","number":2,"license":"CC BY 4.0","cdate":1698634749636,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1048/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636031165,"domain":"ICLR.cc/2024/Conference","replyto":"DO2WFXU1Be","id":"2cwfvioX5W","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-Informed Neural Networks","Transformer","Self-Attention"]},"supplementary_material":{"value":"/attachment/e706d30a428f44ca8015da65b09bbd608a747a77.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Physics-Informed Neural Networks (PINNs) have emerged as a promising deep learning framework for approximating numerical solutions to partial differential equations (PDEs). However, conventional PINNs, relying on multilayer perceptrons (MLP), neglect the crucial temporal dependencies inherent in practical physics systems and thus fail to propagate the initial condition constraints globally and accurately capture the true solutions under various scenarios. In this paper, we introduce a novel Transformer-based framework, termed PINNsFormer, designed to address this limitation. PINNsFormer can accurately approximate PDE solutions by utilizing multi-head attention mechanisms to capture temporal dependencies. PINNsFormer transforms point-wise inputs into pseudo sequences and replaces point-wise PINNs loss with a sequential loss. Additionally, it incorporates a novel activation function, \\texttt{Wavelet}, which anticipates Fourier decomposition through deep neural networks. Empirical results demonstrate that PINNsFormer achieves superior generalization ability and accuracy across various scenarios, including PINNs failure modes and high-dimensional PDEs. Moreover, PINNsFormer offers flexibility in integrating existing learning schemes for PINNs, further enhancing its performance."},"_bibtex":{"value":"@inproceedings{\nzhao2024pinnsformer,\ntitle={{PINN}sFormer: A Transformer-Based Framework For Physics-Informed Neural Networks},\nauthor={Zhiyuan Zhao and Xueying Ding and B. Aditya Prakash},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=DO2WFXU1Be}\n}"},"title":{"value":"PINNsFormer: A Transformer-Based Framework For Physics-Informed Neural Networks"},"pdf":{"value":"/pdf/f06f4c86b32adde5a980de318a16e6874852f4d9.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"zhao|pinnsformer_a_transformerbased_framework_for_physicsinformed_neural_networks"},"authorids":{"value":["~Zhiyuan_Zhao1","~Xueying_Ding1","~B._Aditya_Prakash2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhiyuan Zhao","Xueying Ding","B. Aditya Prakash"]}},"version":2},{"content":{"summary":{"value":"This paper presents a new benchmark, SysMoBench, for assessing the ability of AI systems to generate valid formal models of complex real-world systems. The main contributions of the paper are threefold:\n1. The authors properly define key metrics to assess the quality of the AI generated system models (based on syntax correctness, runtime correctness, conformance to the system implementation and invariant correctness)\n2. The authors also clearly lay out an effective fully automated approach to compute those metrics \n3. An empirical evaluation that shows the limitation of state-of-the-art LLMs in generating system models for real-world complex systems."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"The authors write in section 3.2.4: \n>\"In principle, if a system model fully conforms to code, violations of these invariants would indicate bugs in system code; in practice, few\nAI-generated models achieved fine-grained conformance.\"\n\nLet's assume that, with the help of this new benchmark, state-of-the-art LLMs improve significantly at complex system modeling tasks. What approach would you recommend to determine whether a violation of an invariant stems from a rare LLM modeling error or from an actual bug in the system code?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. **Relevance**\nThe paper addresses a critical challenge: evaluating the capability of AI systems to move beyond basic code generation and comprehension, toward a deeper understanding and modeling of complex real-world systems.\n\n2. **Significance and Novelty**\nTo the best of my knowledge, the proposed approach is novel in its problem formulation, the techniques used to fully automate the quality assessment of AI-generated models, and its focus on real-world complex systems. The introduction of this new benchmark holds strong potential for significant impact, as it provides a vital resource that could accelerate progress in AI-driven modeling of complex systems.\n\n3. **Soundness and Experimental Evaluation**\nThe experimental results validate the effectiveness and robustness of the key LLM-assisted components within the automated evaluation pipeline. They also highlight that current state-of-the-art LLMs still face challenges in accurately modeling complex real-world systems, which underscores the importance of this benchmark in advancing research on AI modeling of intricate software systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I could not find any major problem with the paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917745698,"tcdate":1761933501170,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4900/Reviewer_f1Je"],"signatures":["ICLR.cc/2026/Conference/Submission4900/Reviewer_f1Je"],"forum":"SAeaTz8YoM","number":4,"license":"CC BY 4.0","cdate":1761933501170,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4900/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917745698,"domain":"ICLR.cc/2026/Conference","replyto":"SAeaTz8YoM","id":"VDva2Fl2QL","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"A benchmark for evaluating AI's ability to formally model real-world systems."},"keywords":{"value":["Specification","Benchmark","Distributed System","Concurrent System","Agentic AI","Large Language Model"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Formal models are essential to specifying large, complex computer systems and verifying their correctness, but are notoriously expensive to write and maintain. Recent advances in generative AI show promise in generating certain forms of specifications. However, existing work mostly targets small programs, not complete systems. It is unclear whether AI can deal with realistic system artifacts, as this requires abstracting their complex behavioral properties into formal models. We present SysMoBench, a benchmark that evaluates AI's ability to formally model large, complex systems. We focus on concurrent and distributed systems, which are keystones of today's critical infrastructure, encompassing operating systems and cloud infrastructure. We focus on TLA+, the de facto specification language for concurrent and distributed systems, though SysMoBench has been extended to other languages. We address the primary challenge of evaluating AI-generated models by automating metrics like syntactic and runtime correctness, conformance to system code, and invariant correctness. SysMoBench currently includes eleven diverse system artifacts: the Raft implementation of Etcd and Redis, ZooKeeper's leader election, the Spinlock, Mutex, and Ringbuffer in Asterinas OS, etc., with more being added. SysMoBench enables us to understand the capabilities and limitations of today's LLMs and agents, providing a firm footing for tools in this area and opening up promising new research directions."},"_bibtex":{"value":"@inproceedings{\ncheng2026sysmobench,\ntitle={SysMoBench: Evaluating {AI} on Formally Specifying Complex Real-World Systems},\nauthor={Qian Cheng and Ruize Tang and Emilie Ma and Finn Hackett and Peiyang He and Yiming Su and Ivan Beschastnikh and Yu Huang and Xiaoxing Ma and Tianyin Xu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=SAeaTz8YoM}\n}"},"title":{"value":"SysMoBench: Evaluating AI on Formally Specifying Complex Real-World Systems"},"pdf":{"value":"/pdf/d0eed08e290649943f7c04b29587678285b323e4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"cheng|sysmobench_evaluating_ai_on_formally_specifying_complex_realworld_systems"},"authorids":{"value":["~Qian_Cheng8","~Ruize_Tang1","~Emilie_Ma1","~Finn_Hackett1","~Peiyang_He2","~Yiming_Su3","~Ivan_Beschastnikh1","~Yu_Huang28","~Xiaoxing_Ma1","~Tianyin_Xu1"]},"authors":{"value":["Qian Cheng","Ruize Tang","Emilie Ma","Finn Hackett","Peiyang He","Yiming Su","Ivan Beschastnikh","Yu Huang","Xiaoxing Ma","Tianyin Xu"]}},"version":2},{"content":{"venue":{"value":"INDOCRYPT 2004"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-540-30556-9_2.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"sahai|secure_protocols_for_complex_tasks_in_complex_environments"},"html":{"value":"https://doi.org/10.1007/978-3-540-30556-9_2"},"_bibtex":{"value":"@inproceedings{DBLP:conf/indocrypt/Sahai04,\n  author={Amit Sahai},\n  title={Secure Protocols for Complex Tasks in Complex Environments},\n  year={2004},\n  cdate={1072915200000},\n  pages={14-16},\n  url={https://doi.org/10.1007/978-3-540-30556-9_2},\n  booktitle={INDOCRYPT},\n  crossref={conf/indocrypt/2004}\n}\n"},"abstract":{"value":"Over the last two decades, there has been tremendous success in placing cryptography on a sound theoretical foundation, and building an amazingly successful theory out of it. The key elements in this Modern Cryptographic Theory are the definitions capturing the intuitive, yet elusive notions of security in various cryptographic settings. The definitions of the early 80’s proved to be extremely successful in this regard. But with time, as the theory started addressing more and more complex concerns, further notions of security had to be introduced. One of the most important concerns theory ventured into is of complex environments where different parties are communicating with each other concurrently in many different protocols. A series of efforts in extending security definitions led to the paradigm of Universally Composable (UC) Security [1], which along with modeling a general complex network of parties and providing definitions of security in that framework, provided powerful tools for building protocols satisfying such definitions."},"title":{"value":"Secure Protocols for Complex Tasks in Complex Environments"},"authors":{"value":[{"fullname":"Amit Sahai","username":"~Amit_Sahai1"}]}},"tmdate":1779949810456,"pdate":1104451200000,"externalIds":["dblp:conf/indocrypt/Sahai04"],"tcdate":1779949760991,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Amit_Sahai1"],"forum":"BboatJLYqs","license":"CC BY-SA 4.0","number":35457,"cdate":1072915200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1779949810456,"domain":"OpenReview.net/Public_Article","id":"BboatJLYqs","version":2},{"content":{"summary":{"value":"This paper addresses the storage and analysis bottleneck for high-fidelity 5D gyrokinetic plasma turbulence simulations, which generate terabytes of data. The authors claim that standard lossy compression methods fail to preserve essential physical characteristics, particularly transient turbulence dynamics.\nThe core contributions are:\n1. A novel physics-informed loss function: This loss function is specifically designed for gyrokinetics and incorporates terms to preserve physical integrals heat flux, electrostatic potential, turbulence spectra, and monoticity.\n2. A Proposed Evaluation Framework: The paper proposes and uses a set of metrics to evaluate both spatial/steady-state quantities (quantitatively) and transient turbulence dynamics (the latter qualitatively).\n3. State-of-the-Art Compression: The PINC-VQ-VAE model achieves an extreme compression rate of 70,000x while maintaining significantly better physics fidelity than all baselines."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. On Baselines: Your related work review is thorough, but the experimental baselines are primarily traditional methods. Could you comment on why other learned compression methods (like VAPOR or others) were not included in the comparison?\n2. On \"Unified Evaluation Pipeline\": In practice, you introduced a new curated set of metrics. Can this really be called a unified evaluation pipeline?\n3. On Reproducibility: You state the 500GB dataset is too large to share. A much more practical solution for reproducibility would be to release the compressed test set (i.e., the latent codes), which would be negligibly small (MBs). This would allow anyone to reproduce your entire analysis pipeline (all tables and figures) without needing to re-run the GKW simulations. Would you be willing to add this to your supplementary materials?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Significance: The paper tackles a real, and high-value problem in the scientific ML, where data can be incredibly high-dimensional and sparse.\n2. Clever Loss Function: While the concept of physics-informed losses is well-established (e.g., PINNs), the specific formulation is a significant strength. The authors move beyond standard PDE residuals. The inclusion of losses on derived, non-local turbulence spectra—which can be difficult to compute—and the isotonic loss to enforce a physically-correct spectral shape is a non-trivial and highly effective application of this idea to the compression domain.\n3. Rigorous Experiments: The evaluation is thorough and backed by an impressive 500GB dataset (although difficult to share and reproduce)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Baselines: The related work section (Sec 2) mentions other relevant deep learning methods for scientific data (e.g., VAPOR, Anirudh et al., Cranganore et al.). However, the quantitative comparison in Section 4 is limited to traditional methods (ZFP, Wavelet, PCA, JPEG2000) and the authors' own non-PINC ablations. This makes it difficult to assess how PINC compares to other state-of-the-art learned compressors in this domain.\n2. Overstated \"Unified Evaluation Pipeline\" Contribution: The paper claims to contribute a \"unified evaluation pipeline\". This is strong language for what is, in practice, a curated set of metrics. While valuable, it is not a new automated framework. Furthermore, the authors admit in their limitations (Line 521) that the most novel part of this \"pipeline\"—the evaluation of transient dynamics—remains purely \"qualitative\"  and is not a quantitative metric.\n3. Reproducibility: In your statement, you note the dataset is too large to distribute. A far more effective solution for reproducibility would be to provide the compressed test set."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358855471,"tcdate":1761959497047,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18255/Reviewer_QH9G"],"signatures":["ICLR.cc/2026/Conference/Submission18255/Reviewer_QH9G"],"forum":"fixbsplpdw","number":2,"license":"CC BY 4.0","cdate":1761959497047,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18255/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358855471,"domain":"ICLR.cc/2026/Conference","replyto":"fixbsplpdw","id":"z3bu3u1pGz","forumContent":{"TLDR":{"value":"Neural compression methods enable extreme compression of plasma turbulence simulation data while maintaining low reconstruction error and preserving key physical characteristics."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics-inspired","turbulence","plasma","neural compression","autoencoders","neural fields"]},"supplementary_material":{"value":"/attachment/0814372529d7d0e61c5f895f3e8f4b70b3af92df.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"High-fidelity scientific simulations are now producing unprecedented amounts of data, creating a storage and analysis bottleneck. A single simulation can generate tremendous data volumes, often forcing researchers to discard valuable information. A prime example of this is plasma turbulence described by the Gyrokinetic equations: nonlinear, multiscale, and 5D in phase space. They represent one of the most computationally demanding frontiers of modern science, with runs taking weeks and resulting in tens of terabytes of data dumps.\nThe increasing storage demands underscore the importance of compression, however, compressed snapshots might not preserve essential physical characteristics after reconstruction. To assess whether such characteristics are captured, we propose a spatiotemporal evaluation pipeline which accounts for structural phenomena and multi-scale transient fluctuations. Indeed, we find that various compression techniques lack preservation of temporal turbulence characteristics. Therefore, we explore Physics-Informed Neural Compression (PINC), which incorporates physics-informed losses tailored to gyrokinetics and enables extreme compressions of over 100000x. This direction provides a viable and scalable solution to the prohibitive storage demands of gyrokinetics, enabling post-hoc analyses that were previously infeasible."},"_bibtex":{"value":"@misc{\ngalletti2026physicspreserving,\ntitle={Physics-Preserving Compression of High-Dimensional Plasma Turbulence Simulations},\nauthor={Gianluca Galletti and Gerald Gutenbrunner and Fabian Paischer and Sandeep Suresh Cranganore and William Hornsby and Naomi Carey and Lorenzo Zanisi and Stanislas Pamela and Johannes Brandstetter},\nyear={2026},\nurl={https://openreview.net/forum?id=fixbsplpdw}\n}"},"title":{"value":"Physics-Preserving Compression of High-Dimensional Plasma Turbulence Simulations"},"pdf":{"value":"/pdf/88f0fd80bf4066acbbabbaac93b2beb84178ea60.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"galletti|physicspreserving_compression_of_highdimensional_plasma_turbulence_simulations"},"authorids":{"value":["~Gianluca_Galletti1","~Gerald_Gutenbrunner1","~Fabian_Paischer1","~Sandeep_Suresh_Cranganore1","~William_Hornsby1","~Naomi_Carey1","~Lorenzo_Zanisi1","~Stanislas_Pamela1","~Johannes_Brandstetter1"]},"authors":{"value":["Gianluca Galletti","Gerald Gutenbrunner","Fabian Paischer","Sandeep Suresh Cranganore","William Hornsby","Naomi Carey","Lorenzo Zanisi","Stanislas Pamela","Johannes Brandstetter"]}},"version":2},{"content":{"venue":{"value":"SOCS 2019"},"pdf":{"value":"https://ojs.aaai.org/index.php/SOCS/article/download/18496/18287"},"venueid":{"value":"dblp.org/conf/SOCS/2019"},"paperhash":{"value":"alwala|intuitive_reliable_plans_with_contingencies_planning_with_safety_nets_for_landmarkbased_routing"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Kalyan_Vasudev_Alwala:","https://dblp.org/search/pid/api?q=author:Margarita_Safonova:","~Oren_Salzman2","https://dblp.org/search/pid/api?q=author:Maxim_Likhachev:"]},"html":{"value":"https://doi.org/10.1609/socs.v10i1.18496"},"_bibtex":{"value":"@inproceedings{DBLP:conf/socs/AlwalaSSL19,\n  author={Kalyan Vasudev Alwala and Margarita Safonova and Oren Salzman and Maxim Likhachev},\n  title={Intuitive, Reliable Plans with Contingencies: Planning with Safety Nets for Landmark-Based Routing},\n  year={2019},\n  cdate={1546300800000},\n  pages={2-9},\n  url={https://doi.org/10.1609/socs.v10i1.18496},\n  booktitle={SOCS},\n  crossref={conf/socs/2019}\n}\n"},"abstract":{"value":"We are interested in the problem of providing intuitive instructions for human agents to enable reliable navigation in unknown environments. Since the advent of GPS and digital maps, a common approach is to visually provide a planned path on a digital map defined in terms of actions to take at specific junctions. However, this approach relies on the agent to constantly and accurately localize itself. Furthermore, it comes in stark contrast to the way humans provide instructions—by leveraging known landmarks in the environment to both augment the description of the planned path as well as to allow to detect when the agent deviated from the planned path. Hence, there is need for assurable means of localization, an intuitive way of compactly conveying directions to agents and a systematic approach to account for human errors. To this end, our key insight is to employ known landmarks in the environment to overcome these challenges. We formally model this intuitive way to use landmarks for conveying instructions and for creating contingency plans. We present experiments demonstrating the efficacy of our approach both on synthetic environments as well as on realworld maps, computed using a smart-phone iOS application that we developed."},"title":{"value":"Intuitive, Reliable Plans with Contingencies: Planning with Safety Nets for Landmark-Based Routing"},"authors":{"value":["Kalyan Vasudev Alwala","Margarita Safonova","Oren Salzman","Maxim Likhachev"]}},"tmdate":1718683436839,"pdate":1546300800000,"tcdate":1718683329467,"writers":["~"],"signatures":["~Oren_salzman3"],"forum":"lY9WUucaSb","license":"CC BY-SA 4.0","number":33072,"cdate":1546300800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1718683436839,"domain":"DBLP.org","id":"lY9WUucaSb","version":2},{"content":{"summary":{"value":"This paper introduces Active Reading, a novel framework designed to improve the factual reliability of large language models. The problem is that LLMs often struggle to learn and recall facts, especially from the long tail of their training data. The proposed method has a model \"study\" a given corpus by first generating a diverse set of learning strategies specific to a document and then applying those strategies to create varied synthetic training data. This process is inspired by human learning techniques. The authors demonstrate that this method substantially improves factual recall on expert domain benchmarks, showing a 313% relative gain on a subset of SimpleQA and a 160% relative gain on FinanceBench compared to standard finetuning. They also scale this approach to create WikiExpert 8B, a model trained on 1 trillion synthetic tokens that surpasses the factual accuracy of much larger models like DeepSeekV2 and Llama 3.1 405B."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See Weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The Active Reading method is intuitive, scalable, and presents a clever way to generate highly diverse synthetic data by leveraging the model's own capabilities.\n2. The empirical results are extremely strong, particularly the performance of the 8B model on Simple WikiQA which nearly matches the gold context baseline.\n3. The release of WikiExpert 8B is a significant contribution, as it achieves state of the art factual recall for its size class and provides a powerful, compact model for fact intensive tasks.\n4. The scaling analysis in Section 4.2 is very insightful. The discovery that mixing in general pretraining data and using a higher learning rate is necessary to prevent degradation when scaling the knowledge corpus is a valuable finding for the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the method excels at information extraction, its performance on the full FinanceBench benchmark is notably weaker than the synthetic QA baseline. This suggests the generated strategies may not adequately cover complex reasoning, a point the paper acknowledges but does not fully resolve.\n2. The finding in Table 3 that data generated by a 70B model leads to worse performance for an 8B model than its own self generated data is highly counterintuitive. This result is not deeply investigated and raises more questions than it answers.\n3. The evaluation is heavily focused on Wikipedia based corpora (SimpleQA, NQ, and the main WikiExpert model). While FinanceBench provides one alternative, demonstrating this method's effectiveness in another distinct domain like medicine or law would strengthen the claims of generality."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762932936316,"tcdate":1762261992780,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20039/Reviewer_YUiv"],"signatures":["ICLR.cc/2026/Conference/Submission20039/Reviewer_YUiv"],"forum":"mRi2cJDtIS","number":3,"license":"CC BY 4.0","cdate":1762261992780,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20039/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762932936316,"domain":"ICLR.cc/2026/Conference","replyto":"mRi2cJDtIS","id":"14Yj77zseH","forumContent":{"TLDR":{"value":"We let the model generate self-learning strategies and train on them at scale to learn tail facts more consistently."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["factuality","tail knowledge","synthetic data","synthetic continued pretraining"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"LLMs are known to store vast amounts of knowledge in their parametric memory.\nHowever, learning and recalling facts from this memory is known to be unreliable, depending largely on the prevalence of particular facts in the training data and other factors which are poorly understood.\nPractitioners are lacking tools which will allow them to ensure that the models learn a given body of knowledge reliably and consistently.\nTo this end, we propose Active Reading: a framework where we train models to study a given set of material with self-generated learning strategies.\nFirst, we demonstrate models trained with Active Reading on expert domains absorb significantly more knowledge than vanilla finetuning and other data augmentations.\nWe train expert 8B models that achieve 66% on a Wikipedia-grounded subset of SimpleQA (+313% relative over vanilla finetuning) and 26% on FinanceBench (+160% relative over vanilla finetuning) by applying Active Reading to the source documents for each benchmark.\nFinally, we show that Active Reading can be utilized at pre-training scale to build more factual models.\nAs a demonstration of this, we release WikiExpert-8B, a Wikipedia-expert model trained on 1 trillion generated tokens, which outcompetes models with hundreds of billions of parameters on factual QA."},"_bibtex":{"value":"@inproceedings{\nlin2026learning,\ntitle={Learning Facts at Scale with Active Reading},\nauthor={Jessy Lin and Vincent-Pierre Berges and Xilun Chen and Wen-tau Yih and Gargi Ghosh and Barlas Oguz},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=mRi2cJDtIS}\n}"},"title":{"value":"Learning Facts at Scale with Active Reading"},"pdf":{"value":"/pdf/3b77539549dfce2c6bed6188d90e0b18cda42101.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lin|learning_facts_at_scale_with_active_reading"},"authorids":{"value":["~Jessy_Lin1","~Vincent-Pierre_Berges1","~Xilun_Chen1","~Wen-tau_Yih1","~Gargi_Ghosh3","~Barlas_Oguz1"]},"authors":{"value":["Jessy Lin","Vincent-Pierre Berges","Xilun Chen","Wen-tau Yih","Gargi Ghosh","Barlas Oguz"]}},"version":2},{"content":{"summary":{"value":"This paper shows neural collapse for multivariate regression is usually harmful to generalization, in contrast with neural collapse for classification being beneficial. Metrics for neural regression collapse are defined, with strong correlation between test error and derived NRC1/ID metrics for robotic control tasks. Intuitive explanations are given for why neural collapse harms generalization for regression tasks."},"significance":{"value":3},"workshop_fit":{"value":3},"strengths":{"value":"An intuitive description is given for why neural collapse is detrimental to having neural networks perform multivariate regression, an incredibly common task, with the metrics used being applicable even to complex domains such as robotic control. The presentation was clear and easy to follow. Defined metrics for intrinsic dimension and NRC1 are qualitatively consistent, giving strong evidence that the underlying phenomenon is being properly described. The question answered by this paper appears to be fundamental to the field of deep learning, at least for regression settings."},"confidence":{"value":2},"suggestions":{"value":"The only main recommendation I have is to explore more small-scale, limited problems where it's easier to draw direct causal relations between NRC1/ID and generalization. It's possible, though seemingly unlikely, that the results are purely correlational and dependent on hidden variables with a separate, causational explanation being difficult to spot. Without precise control over an incredibly simple toy problem, it's impossible to rule out that possibility.\n\nThis was likely not included in the main body for space constraints, but some discussion of the difference between NC in classification and regression should be included. Are the NRC1 and ID metrics useful when looking at NC in classifying networks? Why or why not? \n\nAdditionally one small point, in Figure 4, it seems the training curves were included in the (d, 2) and (e) plots, along with there being two different (d) labels."}},"parentInvitations":"ICLR.cc/2026/Workshop/Sci4DL/-/Official_Review","nonreaders":[],"tmdate":1772449477411,"tcdate":1771962750248,"writers":["ICLR.cc/2026/Workshop/Sci4DL","ICLR.cc/2026/Workshop/Sci4DL/Submission21/Reviewer_o1Js"],"signatures":["ICLR.cc/2026/Workshop/Sci4DL/Submission21/Reviewer_o1Js"],"forum":"6JIPhvyRzO","number":2,"license":"CC BY 4.0","cdate":1771962750248,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/Sci4DL/Submission21/-/Official_Review","ICLR.cc/2026/Workshop/Sci4DL/-/Edit"],"mdate":1772449477411,"domain":"ICLR.cc/2026/Workshop/Sci4DL","replyto":"6JIPhvyRzO","id":"MCATVE5J17","forumContent":{"TLDR":{"value":"Deep regression models risk collapsing their penultimate feature manifold to a lower intrinsic dimension than the target manifold, harming generalization."},"venue":{"value":"Sci4DL 2026"},"keywords":{"value":["multivariate regression","neural collapse","intrinsic dimension","deep learning","generalization"]},"abstract":{"value":"Neural multivariate regression underpins a wide range of domains such as control, robotics, and finance, yet the geometry of its learned representations remains poorly characterized. While neural collapse has been shown to benefit generalization in classification, we find that analogous collapse in regression consistently degrades performance. To explain this contrast, we analyze models through the lens of intrinsic dimension. Across control tasks and synthetic datasets, we estimate the intrinsic dimension of last-layer features ($ID_H$) and compare it with that of the regression targets ($ID_Y$). Collapsed models exhibit $ID_H < ID_Y$, leading to over-compression and poor generalization, whereas non-collapsed models typically maintain $ID_H > ID_Y$. For the non-collapsed models, performance with respect to $ID_H$ depends on the data quantity and noise levels. From these observations, we identify two regimes—over-compressed and under-compressed—that determine when expanding or reducing feature dimensionality improves performance. Our results provide new geometric insights into neural regression and suggest practical strategies for enhancing generalization."},"_bibtex":{"value":"@inproceedings{\nandriopoulos2026geometric,\ntitle={Geometric Properties of Neural Multivariate Regression: An Empirical Study},\nauthor={George Andriopoulos and Zixuan Dong and Bimarsha Adhikari and Keith W. Ross},\nbooktitle={Workshop on Scientific Methods for Understanding Deep Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=6JIPhvyRzO}\n}"},"title":{"value":"Geometric Properties of Neural Multivariate Regression: An Empirical Study"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"style_files":{"value":"I have used the style files."},"pdf":{"value":"/pdf/2d98c86dc111aa59830e89878ec0f5f99ebac582.pdf"},"venueid":{"value":"ICLR.cc/2026/Workshop/Sci4DL"},"paperhash":{"value":"andriopoulos|geometric_properties_of_neural_multivariate_regression_an_empirical_study"},"authorids":{"value":["~George_Andriopoulos1","~Zixuan_Dong1","~Bimarsha_Adhikari1","~Keith_W._Ross1"]},"challenge":{"value":"This submission is an entry to the science of DL improvement challenge."},"authors":{"value":["George Andriopoulos","Zixuan Dong","Bimarsha Adhikari","Keith W. Ross"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a post-generation curation framework for selecting high-quality synthetic data to improve visual recognition models. The authors argue that both fidelity (semantic similarity to real data) and diversity (novel variations) are crucial for maximizing the utility of synthetic datasets. They propose partitioning real data into homogeneous (HOMO) and heterogeneous (HETERO) subsets and scoring synthetic samples by their fidelity and diversity relative to each subset. The method, which is generator-agnostic and training-free, consistently enhances in-domain and out-of-domain accuracy across various datasets and models. Experiments demonstrate that carefully balancing fidelity and diversity yields better generalization and robustness than existing data selection strategies."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please address the comments in weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses a timely and practically important problem (i.e., how to effectively select synthetic data rather than merely generating it) and provides a simple yet principled solution.\n\n2. The paper is well-written, well-organized, and provides clear visualizations that make the proposed method and its effects intuitively understandable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed framework relies heavily on the quality and representation power of the feature extractor (e.g., CLIP, SigLIP). As a result, the selection outcome may be biased or unstable when different embedding spaces are used, limiting the generality of the method across various encoders or modalities.\n\n2. While the HOMO–HETERO partition is conceptually intuitive, it lacks a formal theoretical justification or analysis explaining why this separation leads to optimal or stable performance. Without such grounding, the effectiveness of the split may vary significantly across datasets with different intrinsic structures.\n\n3. Although the authors claim that the proposed selection method is efficient and scalable, these aspects are discussed only qualitatively. The paper would benefit from a quantitative evaluation of computational cost, such as runtime or memory usage as a function of dataset size, to substantiate the scalability claim.\n\n4. The proposed method appears to incorporate a core-set–like selection mechanism into the fidelity–diversity framework, aiming to choose a representative yet diverse subset of synthetic samples. While this idea is conceptually related to existing core-set selection principles, the paper does not clearly differentiate its approach or justify how it fundamentally extends beyond standard core-set algorithms in either theoretical formulation or empirical advantage."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918333115,"tcdate":1761978045071,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5896/Reviewer_Z8Sg"],"signatures":["ICLR.cc/2026/Conference/Submission5896/Reviewer_Z8Sg"],"forum":"6r0VuH8gGT","number":4,"license":"CC BY 4.0","cdate":1761978045071,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5896/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918333115,"domain":"ICLR.cc/2026/Conference","replyto":"6r0VuH8gGT","id":"Kgiai5CqYO","forumContent":{"TLDR":{"value":"Proposing scoring method, we select synthetic data for model training. Results show that balancing fidelity and diversity is key to unlock the potential of generative data in visual recognition."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Synthetic data selection","Generative model","Image classification"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"As generative models (GMs) advance, producing high-quality, photorealistic synthetic images is now feasible and increasingly common. Beyond use in entertainment, such synthetic data offers a promising solution to data scarcity in AI research. With the growing research interest in using synthetic data, a critical question arises: How can we maximize performance gains for downstream tasks using synthetic data? Fine-tuning or prompt engineering are two common strategies to fully exploit the potential of GMs to produce task-specific synthetic data. However, unlike these tedious optimizations, is it possible to exploit the data potential by selecting an optimal synthetic subset from a given pool? Motivated by this question, we propose an efficient data selection strategy to improve the utility of existing synthetic datasets without adjusting the GMs' output. Experiments on several benchmarks demonstrate that balancing the trade-off between fidelity and diversity in synthetic data benefits model performance and robustness. In summary, this paper presents a practical and scalable approach to harnessing synthetic data, particularly valuable in scenarios where customizing the outputs of generative models is infeasible."},"_bibtex":{"value":"@misc{\nliu2026balancing,\ntitle={Balancing Fidelity and Diversity: Synthetic data could stand on the shoulder of the real in visual recognition},\nauthor={Disheng Liu and Tuo Liang and Yu Yin},\nyear={2026},\nurl={https://openreview.net/forum?id=6r0VuH8gGT}\n}"},"title":{"value":"Balancing Fidelity and Diversity: Synthetic data could stand on the shoulder of the real in visual recognition"},"pdf":{"value":"/pdf/a5cb0c655d11b998a850b078982088222ddee9f1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|balancing_fidelity_and_diversity_synthetic_data_could_stand_on_the_shoulder_of_the_real_in_visual_recognition"},"authorids":{"value":["~Disheng_Liu2","~Tuo_Liang2","~Yu_Yin2"]},"authors":{"value":["Disheng Liu","Tuo Liang","Yu Yin"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a new framework, called SynMeter, to assess (tabular) synthetic data generators. They focus on three dimensions: \n\n- Fidelity: \n\t- Authors argue the need for a faithful and universal metric\n\t- They propose a Wasserstein distance-based metric to evaluate complex, high-dimensional tabular data distributions\n- Privacy\n\t- Authors argue syntactic privacy scores to not be adequate\n\t- Authors argue that existing MIAs are ineffective, as they are not well understood and no MIA is effective against all synthesizers. \n\t- They propose a new metric called membership disclosure. \n- Utility\n\t- Authors state that the traditionally used ML efficacy is not adequate, as they argue that there is no consensus on which evaluator should be used. \n\t\t○ They propose two new metrics: ML affinity and query error. \n\nThe paper then includes a holistic tuning objective as a combination of all metrics, to be used for hyperparameter selection. \n\nFinally, the paper includes comprehensive experiments evaluating all metric across datasets and generators."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"(also see weaknesses)\n\n- Could you come up with an experimental setup and results which would compellingly show why the Wasserstein-based fidelity metric is strictly better than other deployed methods?\n- Why are MIAs against tabular data synthesis not well understood? \n- There indeed does not exist one MIA effective across all synthesizers, but this does not seem like a justification why MIAs are not useful? The ineffectiveness of the MIA might also just reflect little privacy leakage?  \n- How does the MDS metric resolve your previously raised concerns regarding MIAs? To my understanding, you are in fact proposing a new MIA, but not evaluating it as such. \n- Could you implement the MIA developed by Houssiau et al, and explain why the MDS metric is superior to compute the MIA performance for records identified by Meeus et al? \n- Could authors clarify why the query error should be part of the utility and not part of the fidelity evaluation?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Comprehensively evaluating synthetic data generators is an important problem, and the paper provides a systematic, multi-dimensional evaluation framework to do so. \n2. The paper includes considers many datasets and synthetic data generators, and comparing them across metrics is valuable for the research domain as a whole.\n3. Proposes a way to pick hyperparameters across a multiple dimensions. \n4. Authors make the framework publicly available as a tool for people generating synthetic data"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While I understand the need for a holistic and widely agreed upon evaluation framework for synthetic data  generators, as a reader, I am not convinced that the metrics proposed by the authors are novel, or particularly better than previously proposed ones. I elaborate on each of the dimensions: \n\n**1. Fidelity.**\n\nWhile I find the notion of using Wasserstein distance to compute fidelity interesting, I remain to be convinced why this would be better than existing methods. \n- Could you come up with an experimental setup and results which would compellingly show why the Wasserstein-based fidelity metric is strictly better than other deployed methods?\n\n**2. Privacy.**\n\n I agree with the authors on the shortcomings of syntactic metrics, and like the example given for DCR. However:\n- I do not follow the arguments made for why MIAs are not sufficient. \n\t- Why are MIAs against tabular data synthesis not well understood? \n\t- There indeed does not exist one MIA effective across all synthesizers, but this does not seem like a justification why MIAs are not useful? The ineffectiveness of the MIA might also just reflect limited privacy leakage?  \n- I do not understand what the difference is between the MDS metric and an MIA. If I understand it correctly, you are building a shadow model setup to then compute an MIA scoring function (which you then not evaluate as an MIA). You then pick the record for which you get the best distinction for this scoring function.  To me this basically comes down to compute MIA performance for all records, and use the highest MIA performance as the privacy metric.  \n\t- How does this resolve your previously raised concerns regarding MIAs? \n\t- Moreover, with this, it is not clear whether this is the state-of-the-art MIA. \n- Finally, in this entire discussion, I believe authors fail to mention (and implement) important related work. Houssiau et al [1] propose a new MIA which beats the one proposed by Stadler et al, and Meeus et al [2] propose a principled way to identify most at-risk records. \n\n**3. Utility.**\n\nI agree with the authors that there is no consensus in the literature on which metric should be used to evaluate the utility of the synthetic data. My thoughts:\n\t- While the exact formulation of the MLA score is, at least to my knowledge, new, I believe its novelty to be very limited. For instance, Stadler et al (in Sec. 6.3) measure utility as a decrease in ML accuracy of a model trained on real compared to a model trained on synthetic data. The only difference with the MLA metric would be the averaging across ML models and the normalization. \n\t- Similarly, the query error seems very similar to the k-way marginals fidelity approach, which has also been studied in for instance Annamalai et al. [3] \n\t- Could authors clarify why the query error should be part of the utility and not part of the fidelity evaluation? \n\n**References**\n\n[1] Houssiau, F., Jordon, J., Cohen, S. N., Daniel, O., Elliott, A., Geddes, J., ... & Szpruch, L. (2022). Tapas: a toolbox for adversarial privacy auditing of synthetic data. arXiv preprint arXiv:2211.06550.\n\n[2] Meeus, M., Guepin, F., Creţu, A. M., & de Montjoye, Y. A. (2023, September). Achilles’ heels: vulnerable record identification in synthetic data publishing. In European Symposium on Research in Computer Security (pp. 380-399). Cham: Springer Nature Switzerland.\n\n[3] Annamalai, M. S. M. S., Gadotti, A., & Rocher, L. (2024). A linear reconstruction approach for attribute inference attacks against synthetic data."}},"nonreaders":[],"tmdate":1731427560601,"tcdate":1730399074773,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2222/Reviewer_SYTt"],"signatures":["ICLR.cc/2025/Conference/Submission2222/Reviewer_SYTt"],"forum":"3ANoEa7roV","number":2,"license":"CC BY 4.0","cdate":1730399074773,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2222/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427560601,"domain":"ICLR.cc/2025/Conference","replyto":"3ANoEa7roV","id":"Y9RIBmwIG7","forumContent":{"TLDR":{"value":"Proposed a set of new evaluation metrics and framework for tabular data synthesis from fidelity, privacy and utility"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Tabular Data Synthesis","Privacy","Evaluation Metric","Generative Models"]},"supplementary_material":{"value":"/attachment/7886ba29a1240cb14ed56ca5c3c71ecf87e8ce5d.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Data synthesis has been advocated as an important approach for utilizing data while protecting data privacy. In recent years, a plethora of tabular data synthesis algorithms (i.e., synthesizers) have been proposed. A comprehensive understanding of these synthesizers' strengths and weaknesses remains elusive due to the absence of principled evaluation metrics and head-to-head comparisons between state-of-the-art deep generative approaches and statistical methods. In this paper, we examine and critique existing evaluation metrics, and introduce a set of new metrics in terms of fidelity, privacy, and utility to address their limitations. Based on the proposed metrics, we also devise a unified objective for tuning, which can consistently improve the quality of synthetic data for all methods. We conducted extensive evaluations of 8 different types of synthesizers on 12 real-world datasets and identified some interesting findings, which offer new directions for privacy-preserving data synthesis."},"_bibtex":{"value":"@misc{\ndu2025systematic,\ntitle={Systematic Assessment of Tabular Data Synthesis},\nauthor={Yuntao Du and Ninghui Li},\nyear={2025},\nurl={https://openreview.net/forum?id=3ANoEa7roV}\n}"},"title":{"value":"Systematic Assessment of Tabular Data Synthesis"},"pdf":{"value":"/pdf/7d92d92de7eb050fc9b7d6381a802eb3b35def21.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"du|systematic_assessment_of_tabular_data_synthesis"},"authorids":{"value":["~Yuntao_Du3","~Ninghui_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuntao Du","Ninghui Li"]}},"version":2},{"content":{"summary":{"value":"This paper presents F2M-Reg, an unsupervised RGB-D registration framework that addresses the frame-to-frame registration task by dealing with multi-view inconsistencies with bootstrapping with a synthetic dataset."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"See weakness."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. F2M-Reg stands out by shifting from a frame-to-frame to a frame-to-model approach for RGB-D registration, which is an extension of existing approaches in the context of unsupervised 3D vision tasks. \n2. This use of a neural implicit field as a global scene model to capture broader scene-level information is a possible direction to handle complex conditions, such as low overlap and lighting changes, where traditional methods often fall short.\n 3. The introduction of a synthetic bootstrapping dataset, Sim-RGBD, bridges the gap between synthetic and real-world performance in unsupervised settings, which is a notable improvement in unsupervised model initialization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although F2M-Reg is compared with several baselines, it is unclear where the improvement comes from. According to Table 4, the results without bootstrapping are not exciting enough.\n2. In order to evaluate the effectiveness of the neural implicit field-guided mechanism, this paper needs additional experiments and comparisons with SOTA approaches without bootstrapping."}},"nonreaders":[],"tmdate":1731428884014,"tcdate":1730641671708,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7404/Reviewer_Sxy2"],"signatures":["ICLR.cc/2025/Conference/Submission7404/Reviewer_Sxy2"],"forum":"5G9PrHERql","number":3,"license":"CC BY 4.0","cdate":1730641671708,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7404/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428884014,"domain":"ICLR.cc/2025/Conference","replyto":"5G9PrHERql","id":"vCEqMxzJFw","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["RGB-D registation","unsupervised learning","frame-to-model optimization"]},"supplementary_material":{"value":"/attachment/7852bd05bcc793516055934855f4742cfe6a255b.pdf"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper focuses on training a robust RGB-D registration model without ground-truth pose supervision.\nExisting methods usually adopt a pairwise training strategy based on differentiable rendering, which enforces the photometric and the geometric consistency between the two registered frames as supervision. However, this frame-to-frame framework suffers from poor multi-view consistency due to factors such as lighting changes, geometry occlusion and reflective materials. In this paper, we present F2M-Reg, a novel frame-to-model optimization framework for unsupervised RGB-D registration. Instead of frame-to-frame consistency, we leverage the neural implicit field as a global model of the scene and use the consistency between the input and the rerendered frames for pose optimization. This design can significantly improve the robustness in scenarios with poor multi-view consistency and provides better learning signal for the registration model. Furthermore, to facilitate the neural field optimization, we create a synthetic dataset, Sim-RGBD, through a photo-realistic simulator to warm up the registration model. By first training the registration model on Sim-RGBD and later unsupervisedly fine-tuning on real data, our framework enables distilling the capability of feature extraction and registration from simulation to reality. Our method outperforms the state-of-the-art counterparts on two popular indoor RGB-D datasets, ScanNet and 3DMatch. Code and models will be released for paper reproduction."},"_bibtex":{"value":"@misc{\nyu2025fmreg,\ntitle={F2M-Reg: Unsupervised {RGB}-D Registration with Frame-to-Model Optimization},\nauthor={Zhinan Yu and Zheng Qin and Yijie Tang and Yongjun Wang and Renjiao Yi and Chenyang Zhu and Kai Xu},\nyear={2025},\nurl={https://openreview.net/forum?id=5G9PrHERql}\n}"},"title":{"value":"F2M-Reg: Unsupervised RGB-D Registration with Frame-to-Model Optimization"},"pdf":{"value":"/pdf/8c3e72cade9da1a6b2a4ede2bbeba98769aba04c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yu|f2mreg_unsupervised_rgbd_registration_with_frametomodel_optimization"},"authorids":{"value":["~Zhinan_Yu1","~Zheng_Qin2","~Yijie_Tang1","~Yongjun_Wang1","~Renjiao_Yi2","~Chenyang_Zhu1","~Kai_Xu5"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhinan Yu","Zheng Qin","Yijie Tang","Yongjun Wang","Renjiao Yi","Chenyang Zhu","Kai Xu"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (2) 2020"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-65351-4_29.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2020"},"paperhash":{"value":"papagiannis|deep_reinforcement_learning_for_control_of_probabilistic_boolean_networks"},"authorids":{"value":["~Georgios_Papagiannis1","~Sotiris_Moschoyiannis1"]},"html":{"value":"https://doi.org/10.1007/978-3-030-65351-4_29"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/PapagiannisM20,\n  author={Georgios Papagiannis and Sotiris Moschoyiannis},\n  title={Deep Reinforcement Learning for Control of Probabilistic Boolean Networks},\n  year={2020},\n  cdate={1577836800000},\n  pages={361-371},\n  url={https://doi.org/10.1007/978-3-030-65351-4_29},\n  booktitle={COMPLEX NETWORKS (2)},\n  crossref={conf/complexnetworks/2020-2}\n}\n"},"abstract":{"value":"Probabilistic Boolean Networks (PBNs) were introduced as a computational model for the study of complex dynamical systems, such as Gene Regulatory Networks (GRNs). Controllability in this context is the process of making strategic interventions to the state of a network in order to drive it towards some other state that exhibits favourable biological properties. In this paper we study the ability of a Double Deep Q-Network with Prioritized Experience Replay in learning control strategies within a finite number of time steps that drive a PBN towards a target state, typically an attractor. The control method is model-free and does not require knowledge of the network’s underlying dynamics, making it suitable for applications where inference of such dynamics is intractable. We present extensive experiment results on two synthetic PBNs and the PBN model constructed directly from gene-expression data of a study on metastatic-melanoma."},"title":{"value":"Deep Reinforcement Learning for Control of Probabilistic Boolean Networks"},"authors":{"value":["Georgios Papagiannis","Sotiris Moschoyiannis"]}},"tmdate":1758286823729,"pdate":1577836800000,"tcdate":1718015710489,"writers":["~"],"signatures":["~George_Papagiannis1"],"forum":"MrDjrgk1mD","license":"CC BY-SA 4.0","number":25756,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1758286823729,"domain":"DBLP.org","id":"MrDjrgk1mD","version":2},{"content":{"summary":{"value":"This paper introduces STARGAZER, a benchmark environment for evaluating AI agents on the iterative, multi-step scientific workflow of exoplanet discovery from radial velocity (RV) time series. The benchmark comprises 120 tasks (100 synthetic across three difficulty tiers, 20 from real archival data), with synthetic difficulty controlled by six physics-based factors (planet multiplicity, SNR, resonances, period coverage, observation count, correlated noise). Agents operate in a ReAct-style loop with a PythonREPL and a submit interface, receiving per-criterion feedback after each submission. Evaluation uses a conjunction of four criteria: two statistical (residual RMS, delta-BIC model selection) and two physical (parameter match score via Hungarian matching, planet count). Eight frontier models are evaluated, revealing three key findings: (1) agents reliably achieve good statistical fits but fail to recover correct physical parameters, especially on Hard tasks; (2) token usage does not predict performance  failed agents burn 10x more tokens in repetitive loops; and (3) self-generated skills improve efficiency on Easy tasks but do not transfer physical reasoning to Hard tasks. Notably, all 8 frontier models score 0% on the 20 real-data tasks, despite these being solved by human astronomers."},"ethics_flag":{"value":1},"reasons_to_reject":{"value":"1. **Classical baselines outperform all LLM agents on Easy tasks** (95% vs. best 80%), and match LLM performance on Medium tasks (~35%). The paper does not clearly establish at what difficulty level  if any  LLM agents provide an advantage over traditional methods. This undercuts the motivation for using LLMs as scientific agents in this domain.\n \n2. **No agent is tested with specialized RV tools** (RadVel, juliet) or domain-expert prompting strategies beyond vanilla skills injection. The benchmark currently measures general coding ability rather than the ceiling of what an LLM-based scientific agent could achieve with proper tooling. At least one such experiment is needed.\n \n3. **Three runs with high stochasticity**: the observed variance (Pass@3 of 95% vs. mean of 40-80%) suggests three runs may be insufficient for reliable mean estimates. No confidence intervals or error bars are reported on the main results.\n \n4. **Limited failure analysis**: the paper demonstrates that Hard tasks are very difficult but provides limited structured insight into *why* a systematic error taxonomy across Hard-tier episodes would be more actionable than two case studies.\n \n5. **The difficulty rubric is not validated against actual agent performance**  weights are set from domain knowledge rather than empirically verified to predict LLM agent difficulty."},"review":{"value":"## Strengths\n \n**S1. The statistical-physical dissociation is the paper's most important insight.**\n \nThe finding that agents achieve >90% on statistical criteria (delta-BIC, RMS) while collapsing to <6% on physical criteria (Match Score) on Hard tasks is genuinely revealing. This is not just a benchmark result it exposes a fundamental limitation of current agents: they optimize within a fixed hypothesis (curve fitting) rather than searching over physically plausible configurations (model selection). The conjunction gate that requires all four criteria to pass simultaneously is a well-designed evaluation choice that prevents trivially gaming the benchmark. This insight generalizes beyond exoplanet discovery to any scientific domain where goodness-of-fit does not imply correct physical interpretation.\n \n**S2. The environment design is exceptionally well-crafted.**\n \nThe seed-based reproducibility (every synthetic task is fully determined by a single random seed) enables infinite scalability new held-out suites can be generated on demand, preventing benchmark saturation. The six physics-based difficulty factors are grounded in established RV theory (Cumming 2004, Anglada-Escudé et al. 2009) rather than arbitrary complexity metrics. The iterative feedback loop (submit, receive per-criterion diagnostics, revise) faithfully mirrors how human astronomers actually work. The solvability filtering removing tasks that are physically non-identifiable under the realized observation window shows careful attention to construct validity.\n \n**S3. The case studies (Section 4.4) are outstanding pedagogical contributions.**\n \nThe success case (GPT-5.2, seed 96: textbook peel-and-search, 68K tokens, 1 submission, pass) contrasted with the failure case (GPT-5-mini, seed 196: alias convergence, 730K tokens, 10 submissions, all fail) is the most illuminating part of the paper. The step-by-step traces make the behavioral divide concrete: successful agents recognize when a fit converges to a search boundary and escalate model complexity, while failed agents perseverate on the same wrong answer. The 10.7x token ratio between failure and success elegantly demonstrates that more compute does not compensate for inability to revise hypotheses.\n \n**S4. The real-data subset provides a compelling ceiling test.**\n \nThe 0% pass rate across all 8 frontier models on 20 tasks that human astronomers have solved is a stark capability gap. The anonymization protocol (stripping star names, instrument names, literature references) is well-designed to prevent data contamination. The independent RadVel verification of ground-truth parameters adds credibility. The GJ 876 inclusion where the Keplerian superposition assumption itself breaks down, requiring N-body modeling is a particularly clever test of whether agents can recognize when their model family is misspecified.\n \n**S5. The match-score threshold analysis (Figure 3, Appendix B.6) is rigorous.**\n \nThe bimodal distribution of match scores (most submissions score either >0.9 or <0.5, with only 14% in the boundary region) and the stability of model rankings under ±10% threshold variation provide strong evidence that the S_match ≥ 0.8 threshold is not an arbitrary choice driving the results. This kind of sensitivity analysis is often missing from benchmark papers and its inclusion here reflects careful methodology.\n \n**S6. The skills analysis reveals an important negative result.**\n \nThe finding that self-generated skills boost Easy-tier pass rates (up to +28.3 pp) primarily through efficiency  compressing the workflow so previously budget-exceeding episodes reach the submission stage  rather than improving physical reasoning is valuable. The observation that skills can actually degrade Hard-tier RMS for weaker models (Table 2, Appendix B.4) is an important cautionary finding: procedural templates from simple tasks can interfere with the exploratory strategies needed for complex problems.\n \n \n## Weaknesses\n \n**W1. No agent with domain-specific prompting or tool access beyond vanilla ReAct.**\n \nAll 8 evaluated agents use the same vanilla ReAct loop with a PythonREPL and submit interface. No agent is given access to specialized RV analysis tools (RadVel, juliet, pyaneti) that human astronomers routinely use, nor are any agents prompted with domain-specific strategies beyond the optional skills.md. This means the benchmark currently measures \"general coding agent + physics knowledge from pretraining\" rather than \"best achievable agent performance on RV analysis.\" A single experiment with an agent given RadVel as a callable tool (rather than requiring the agent to implement fitting from scratch in Python) would help establish how much of the performance gap is due to reasoning limitations vs. implementation friction. The classical pipeline baseline partially addresses this, but it is a deterministic program, not an LLM agent with access to the same tools.\n \n**W2. The classical baselines outperform all LLM agents on Easy tasks.**\n \nBoth the Classical Pipeline and Nested Sampling baselines achieve 95% on Easy tasks higher than any LLM agent (best: GPT-5.3-codex at 80%). This is buried in the results but is a significant finding: on well-structured single-planet problems, traditional deterministic methods remain superior. The paper should discuss this more prominently. It raises the question: at what difficulty level do LLM agents actually provide an advantage over classical methods? On Medium tasks, the best LLM agents (Gemini-3.1-Pro: 35%) are comparable to classical methods (35%), suggesting LLM agents have not yet demonstrated a clear advantage over traditional approaches at any difficulty level in this benchmark.\n \n**W3. The difficulty scoring rubric weights are set a priori without validation.**\n \nThe six difficulty factor weights are described as \"set a priori from domain knowledge\" (line 122) and \"calibrated through pilot experiments\" (line 610-616). However, the paper does not validate whether these weights actually predict difficulty for LLM agents specifically. Figure 5 shows post-hoc correlations between factors and success, which is informative, but different weighting could produce a different difficulty stratification. An analysis of whether the difficulty score monotonically predicts agent pass rates at the task level (not just tier level) would strengthen the claim of construct validity.\n \n**W4. Three runs may not be sufficient for reliable estimates, given the high stochasticity observed.**\n \nThe paper notes \"substantial stochasticity\" in Pass@3 results four models reach 95% Pass@3 on Easy despite mean pass rates of 40-80% (Section 4.2). This level of variance suggests that three independent runs may not be enough for reliable mean estimates. Confidence intervals or standard errors on the main results (Table 1) would help readers assess the statistical reliability of the reported differences between models. For example, is the difference between GPT-5.3-codex (80.0%) and GPT-5-mini (76.7%) on Easy tasks meaningful or within noise?\n \n**W5. The token budget analysis could be much deeper.**\n \nThe paper makes the important claim that \"token usage does not predict performance\" and that failed agents fall into \"recursive failure loops.\" However, the evidence is largely anecdotal (the two case studies). A systematic analysis across all episodes would strengthen this: what is the correlation between token usage and pass rate within each tier? What fraction of Hard-tier tokens are spent on repeated submissions vs. novel exploration? Is there a token threshold beyond which additional computation never helps? The data clearly exists in the logs — extracting these patterns would turn an anecdote into a finding.\n \n**W6. The paper does not adequately discuss what would make Hard tasks solvable.**\n \nThe Hard-tier pass rate is <6% for all models, and the real-data pass rate is 0%. While this demonstrates the benchmark's difficulty, the paper offers limited insight into what capabilities would be needed to solve these tasks. Is the bottleneck model selection (deciding how many planets)? Parameter estimation (getting the orbital elements right)? Recognizing aliases and resonances? The per-criterion breakdown in Table 7 helps (Planet Count is often correct while Match Score fails), but a more structured error taxonomy — categorizing failure modes across all Hard-tier episodes — would provide actionable guidance for model developers.\n \n**W7. The match-score metric (Equation 3) has a design choice that could be scrutinized further.**\n \nThe count penalty term -0.25|n_truth - n_guess| in Equation 3 is additive with the exponential distance term. This means a submission with the correct number of planets but slightly wrong parameters could score lower than one with a wrong planet count but very accurate parameters for the planets it does detect. The paper acknowledges this partially in Appendix B.1 (the mean-over-matched vs. mean-over-truth formulation), but the interaction between S_match and ok_count as separate criteria deserves more discussion. In particular, the current formulation means the count penalty appears in both ok_match and ok_count, creating partial redundancy.\n \n**W8. Minor issues.**\n \n- Line 40: \"RV methods remains\" should be \"RV methods remain.\"\n- Line 109: \"centers\" is missing after \"explicitly on\" — \"it explicitly centers on open-weight local deployment\" (wait, this is from the other paper — ignore this).\n- The paper does not discuss computational cost for the evaluation: how many GPU-hours / API dollars were spent running 8 models × 120 tasks × 3 runs? This is relevant for reproducibility.\n- The skills are extracted from Opus 4.6 trajectories, which is not among the 8 evaluated models. While this avoids data leakage, it also means the skills may not be optimally formatted for the models that use them. Was this choice validated?\n \n \n## Minor Issues\n \n- Figure 1 (right panel) has a cartoon-style \"I burned nearly 1M tokens and got a wrong answer\" callout that, while memorable, may feel informal for a venue like COLM. Consider whether the humor serves the paper or undercuts its seriousness.\n \n- The paper would benefit from a \"Limitations\" section explicitly discussing the gap between what the benchmark tests (Keplerian model fitting) and the full scope of RV analysis (correlated noise modeling, activity indicators, Bayesian evidence, N-body dynamics).\n \n- Table 4 (real-data tasks) is excellent and should be highlighted more prominently  it is buried in the appendix but is one of the benchmark's strongest features."},"confidence":{"value":3},"reasons_to_accept":{"value":"1. **The statistical-physical dissociation is a genuinely important finding** that generalizes beyond exoplanet discovery: current agents can optimize loss functions but struggle with the model selection and physical interpretation that define real scientific reasoning. This insight should influence how the community designs evaluation metrics for scientific agents.\n \n2. **Exceptional environment design**: seed-based reproducibility, physics-grounded difficulty scaling, solvability filtering, iterative feedback, and a conjunction evaluation gate that prevents trivial gaming. The benchmark is infinitely scalable by design, addressing the saturation problem that plagues static benchmarks.\n \n3. **The 0% real-data pass rate** across all 8 frontier models on tasks provably solved by human astronomers — establishes a concrete, falsifiable capability gap between current AI agents and human researchers.\n \n4. **The case studies are outstanding**: the success/failure contrast (68K tokens → pass vs. 730K tokens → fail) makes abstract claims about agent reasoning concrete and actionable.\n \n5. **Thorough supplementary analysis**: match-score threshold sensitivity (Appendix B.6), per-criterion breakdowns (Table 7), difficulty factor correlations (Figure 5), and per-criterion skills analysis (Table 8) demonstrate a level of rigor unusual in benchmark papers."},"rating":{"value":7},"questions_to_authors":{"value":"1. **LLM agents vs. classical methods**: At what difficulty level do LLM agents genuinely outperform traditional methods? The Classical Pipeline matches or exceeds all LLM agents through Medium tier. Have you considered a hybrid baseline where an LLM agent orchestrates calls to RadVel or similar tools? This would help disentangle reasoning limitations from implementation friction.\n \n2. **Error taxonomy**: Could you categorize Hard-tier failures into a structured taxonomy  e.g., alias convergence (wrong periods), model under-specification (too few planets), model over-specification (too many planets), parameter inaccuracy (right periods but wrong K/e), format errors? Even a rough breakdown across all Hard-tier episodes would be more informative than the two case studies alone.\n \n3. **Statistical reliability**: Could you provide standard errors or confidence intervals on the main results in Table 1? Given the high stochasticity noted in Section 4.2, it would be helpful to know whether the differences between, say, GPT-5.3-codex (80.0%) and GPT-5-mini (76.7%) on Easy tasks are statistically meaningful.\n \n4. **Token-performance correlation**: You claim token usage does not predict performance. Could you show this systematically — e.g., a scatter plot of tokens consumed vs. match score across all episodes, colored by outcome? This would elevate the claim from anecdotal (two case studies) to empirical.\n \n5. **Difficulty validation**: Does the integer difficulty score d monotonically predict agent pass rates at the individual-task level (not just averaged across tiers)? A plot of pass rate vs. d across all 100 synthetic tasks would validate the rubric's construct validity.\n \n6. **Real-data near-misses**: Table 6 shows that several agents recover correct periods but overestimate semi-amplitudes K on real data. Is there a systematic explanation for the K overestimation? Could this stem from unmodeled stellar activity or correlated noise that is present in real data but absent (or differently parameterized) in synthetic tasks?\n \n7. **Evaluation cost**: What was the total computational cost (API calls, GPU-hours) of the full evaluation? This is important for labs considering reproducing or extending the benchmark."},"title":{"value":"Physics-Grounded Agentic Benchmark for Exoplanet Discovery excellent Environment Design and Insightful Findings, with Room for Stronger Baselines and Deeper Failure Analysis"}},"parentInvitations":"colmweb.org/COLM/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1787537428751,"tcdate":1776543585336,"writers":["colmweb.org/COLM/2026/Conference","colmweb.org/COLM/2026/Conference/Submission2935/Reviewer_xq8S"],"signatures":["colmweb.org/COLM/2026/Conference/Submission2935/Reviewer_xq8S"],"forum":"QLJaIenMoY","number":1,"license":"CC BY 4.0","cdate":1776543585336,"readers":["everyone"],"invitations":["colmweb.org/COLM/2026/Conference/Submission2935/-/Official_Review","colmweb.org/COLM/2026/Conference/-/Edit"],"mdate":1787537428751,"domain":"colmweb.org/COLM/2026/Conference","replyto":"QLJaIenMoY","id":"t9JB2kKBST","forumContent":{"TLDR":{"value":"STARGAZER is a dynamic environment for evaluating AI agents, uncovering a fundamental gap between numerical optimization and adherence to physical constraints."},"venue":{"value":"COLM 2026"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the COLM Code of Ethics on https://colmweb.org/CoE.html"},"LLM_usage":{"value":"In accordance with the Policy on the use of Large Language Models in the Call for Papers https://colmweb.org/cfp.html, I certify that my submission discloses any substantive use of LLMs in the research and paper-writing process."},"keywords":{"value":["AI for science","LLM evaluation","physical reasoning","AI agents"]},"supplementary_material":{"value":"/attachment/cb0d77737ba238990e4e3207b8939e03e4deddbd.zip"},"abstract":{"value":"The rise of autonomous AI agents suggests that dynamic benchmark envi-\nronments with built-in feedback on scientifically grounded tasks are needed\nto evaluate the capabilities of these agents in research work. We introduce\nSTARGAZER, a scalable environment for evaluating AI agents on dynamic,\niterative physics-grounded model-fitting tasks using inference on radial-\nvelocity (RV) time series data. STARGAZER comprises 120 tasks across three\ndifficulty tiers, including 20 real archival cases, covering diverse scenarios\nranging from high-SNR single-planet systems to complex multi-planetary\nconfigurations requiring involved low-SNR analysis. Our evaluation of\neight frontier agents reveals a gap between numerical optimization and\nadherence to physical constraints: although agents often achieve a good\nstatistical fit, they frequently fail to recover correct physical system pa-\nrameters, a limitation that persists even when agents are equipped with\ndomain-expert skills. Furthermore, increasing test-time compute yields\nonly marginal gains, with excessive token usage often reflecting recursive\nfailure loops rather than meaningful exploration. STARGAZER presents an\nopportunity to train, evaluate, scaffold, and scale strategies on a model-\nfitting problem of practical research relevance today. Our simulation-driven\nmethodology may also provide a useful design pattern for other scientific\nmodel-fitting domains."},"_bibtex":{"value":"@inproceedings{\nliu2026stargazer,\ntitle={Stargazer: A Scalable Model-fitting Benchmark Environment for {AI} Agents under Astrophysical Constraints},\nauthor={Xinge Liu and Terry Jingchen Zhang and Bernhard Sch{\\\"o}lkopf and Zhijing Jin and Kristen Menou},\nbooktitle={Third Conference on Language Modeling},\nyear={2026},\nurl={https://openreview.net/forum?id=QLJaIenMoY}\n}"},"title":{"value":"Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints"},"pdf":{"value":"/pdf/fba20bad468f9cf56ad40ba1f46c9d47743d452b.pdf"},"venueid":{"value":"colmweb.org/COLM/2026/Conference"},"paperhash":{"value":"liu|stargazer_a_scalable_modelfitting_benchmark_environment_for_ai_agents_under_astrophysical_constraints"},"authorids":{"value":["~Xinge_Liu1","~Terry_Jingchen_Zhang1","~Bernhard_Schölkopf1","~Zhijing_Jin1","~Kristen_Menou1"]},"author_guide":{"value":"I certify that this submission complies with the submission instructions as described on https://colmweb.org/AuthorGuide.html"},"authors":{"value":["Xinge Liu","Terry Jingchen Zhang","Bernhard Schölkopf","Zhijing Jin","Kristen Menou"]}},"version":2},{"content":{"summary":{"value":"I love the paper for its outperforming MGN; it has a solid baseline by DeepMind! It introduces PhyMPGN, a Physics-encoded Message Passing Graph Network that effectively models spatiotemporal PDE systems on irregular meshes using limited data. By integrating physics through a learnable Laplace-Beltrami operator and a novel boundary condition padding strategy, the approach ensures physically accurate predictions. PhyMPGN significantly outperforms existing methods, achieving over 50% performance gains and demonstrating strong generalization across various PDEs and conditions. The model's efficiency and robustness make it a valuable advancement for scientific simulations where data is sparse or complex geometries are involved."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"I have several questions to help me understand the method thoroughly. Firstly, how will the model behave if we delete (or do not use) the Mesh Laplace? Will it still be better than MGN? Secondly, why second-order RK for MOL? What about RK4 or forward Euler? Do these different choices matter? Thirdly, I understand graph Laplace preserves some diffusion physics, which is excellent and universal enough. Still, for examples of inviscid first order PDE such as wave/ advection equation, or inviscid Burger equation,  Euler equations, will add graph Laplacian make the traveling feature blur?"},"rating":{"value":10},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"Outperform a solid baseline; writing is good, presentation is good. The method is clear and new. The strength of this paper lies in its innovative integration of physics-based knowledge into graph neural networks, which enables accurate and efficient modeling of complex spatiotemporal PDE systems on irregular meshes with limited data. By employing a learnable Laplace block and a novel boundary condition padding strategy, PhyMPGN ensures solutions remain physically consistent and precise, overcoming limitations of traditional and purely data-driven methods. Additionally, the model's demonstrated ability to generalize well to different PDEs, geometries, and conditions, along with its significant performance gains over existing techniques, highlights its robustness and versatility for real-world scientific applications"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"There is some need for some clarification when it is for convection-dominant problem."}},"nonreaders":[],"tmdate":1732858045856,"tcdate":1730317907270,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5640/Reviewer_m5AJ"],"signatures":["ICLR.cc/2025/Conference/Submission5640/Reviewer_m5AJ"],"forum":"fU8H4lzkIm","number":2,"license":"CC BY 4.0","cdate":1730317907270,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5640/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732858045856,"domain":"ICLR.cc/2025/Conference","replyto":"fU8H4lzkIm","id":"5JVHmg36ri","forumContent":{"TLDR":{"value":"Presented a Physics-encoded Message Passing Graph Network for simulation of spatiotemporal PDE systems."},"venue":{"value":"ICLR 2025 Spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-encoded; Spatiotemporal PDEs; Graph Network; Deep Learning;"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Solving partial differential equations (PDEs) serves as a cornerstone for modeling complex dynamical systems. Recent progresses have demonstrated grand benefits of data-driven neural-based models for predicting spatiotemporal dynamics (e.g., tremendous speedup gain compared with classical numerical methods). However, most existing neural models rely on rich training data, have limited extrapolation and generalization abilities, and suffer to produce precise or reliable physical prediction under intricate conditions (e.g., irregular mesh or geometry, complex boundary conditions, diverse PDE parameters, etc.). To this end, we propose a new graph learning approach, namely, Physics-encoded Message Passing Graph Network (PhyMPGN), to model spatiotemporal PDE systems on irregular meshes given small training datasets. Specifically, we incorporate a GNN into a numerical integrator to approximate the temporal marching of spatiotemporal dynamics for a given PDE system. Considering that many physical phenomena are governed by diffusion processes, we further design a learnable Laplace block, which encodes the discrete Laplace-Beltrami operator, to aid and guide the GNN learning in a physically feasible solution space. A boundary condition padding strategy is also designed to improve the model convergence and accuracy. Extensive experiments demonstrate that PhyMPGN is capable of accurately predicting various types of spatiotemporal dynamics on coarse unstructured meshes, consistently achieves the state-of-the-art results, and outperforms other baselines with considerable gains."},"_bibtex":{"value":"@inproceedings{\nzeng2025phympgn,\ntitle={Phy{MPGN}: Physics-encoded Message Passing Graph Network for spatiotemporal {PDE} systems},\nauthor={Bocheng Zeng and Qi Wang and Mengtao Yan and Yang Liu and Ruizhi Chengze and Yi Zhang and Hongsheng Liu and Zidong Wang and Hao Sun},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=fU8H4lzkIm}\n}"},"title":{"value":"PhyMPGN: Physics-encoded Message Passing Graph Network for spatiotemporal PDE systems"},"pdf":{"value":"/pdf/3a945ae4beaef8cfd607d0b5ef1679eb3699e8d8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zeng|phympgn_physicsencoded_message_passing_graph_network_for_spatiotemporal_pde_systems"},"authorids":{"value":["~Bocheng_Zeng1","~Qi_Wang30","~Mengtao_Yan1","~Yang_Liu52","~Ruizhi_Chengze1","~Yi_Zhang92","~Hongsheng_Liu1","~Zidong_Wang1","~Hao_Sun4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Bocheng Zeng","Qi Wang","Mengtao Yan","Yang Liu","Ruizhi Chengze","Yi Zhang","Hongsheng Liu","Zidong Wang","Hao Sun"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LynX, a framework for equipping pretrained Visual Language Models (VLMs) with visual grounding capabilities without forgetting their existing image and language understanding skills. It proposes a Dual Mixture of Experts (MoE) architecture that allows the model to specialize in both image understanding and visual grounding simultaneously. Additionally, it introduces SCouT (Synthetic Chain-of-Thought with Grounding), a high-quality synthetic dataset with step-by-step grounded chain-of-thought annotations to facilitate effective training on grounding tasks. The framework also includes a step-by-step training methodology that breaks down complex tasks into intermediate steps with individual loss functions."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The paper presents a straightforward solution to incorporate grounding ability into VLMs, but the innovation is not substantial enough to provide new insights. A paper that is to be accepted should offer the community fresh perspectives.\n\nIn Table 5, MoE-LLaVA-phi2 is the base model. It appears that Lynx has fewer active parameters, which may be a mistake. Comparing the MoE model with a non-MoE model is not fair in Table 5 and Table 4."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"(1)  LynX uses a dual MoE module with one frozen MoE pretrained on image understanding and another trainable MoE for new grounding capabilities.\n     (2) The SCouT dataset provides rich supervision with step-by-step multimodal reasoning, an improvement upon previous synthetic datasets.\n     (3) LynX outperforms larger models such as Shikra-7B on complex tasks like phrase grounding, despite having fewer parameters (1.67B vs 7B)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) The novelty of this paper is limited in its content. In summary, the key innovations of this paper can be regarded as CoT[1] combined with adapter[2, 3], which does not provide any new insights. Additionally, the step-by-step loss function is a common auto-regressive (next token prediction) loss. The BERT classification module is not an elegant solution for the method.\n(2) The paper freezes the original network and only trains the added gated adapter. There is no doubt that the performance of the image understanding from the original base model cannot be affected. However, LynX, trained with additional data (REC, SCouT), shows limited performance improvement (83.9 -> 85.4) and is left behind by Shikra-7B.\nReference:\n[1] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V. and Zhou, D., 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35, pp.24824-24837.\n[2] Zhang, L., Rao, A. and Agrawala, M., 2023. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 3836-3847).\n[3] Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X. and Xu, J., 2023. Cogvlm: Visual expert for pretrained language models. arXiv preprint arXiv:2311.03079."}},"nonreaders":[],"tmdate":1731429250990,"tcdate":1731198666956,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10883/Reviewer_4zcz"],"signatures":["ICLR.cc/2025/Conference/Submission10883/Reviewer_4zcz"],"forum":"HjoYVtSkT8","number":4,"license":"CC BY 4.0","cdate":1731198666956,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10883/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429250990,"domain":"ICLR.cc/2025/Conference","replyto":"HjoYVtSkT8","id":"DJ1JkqdI4Y","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language models","visual grounding","mixture of experts","synthetic datasets"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Spatial awareness is key to enable embodied multimodal AI systems. Yet, without vast amounts of spatial supervision, current Visual Language Models (VLMs) struggle at this task. In this paper, we introduce LynX, a framework that equips pretrained VLMs with visual grounding ability without forgetting their existing image and language understanding skills.\nTo this end, we propose a Dual Mixture of Experts module that modifies only the decoder layer of the language model, using one frozen Mixture of Experts (MoE) pre-trained on image and language understanding and another learnable MoE for new grounding capabilities. This allows the VLM to retain previously learned knowledge and skills, while acquiring what is missing.\nTo train the model effectively, we generate a high-quality synthetic dataset we call SCouT, which mimics human reasoning in visual grounding. This dataset provides rich supervision signals, describing a step-by-step multimodal reasoning process, thereby simplifying the task of visual grounding. We evaluate LynX on several object detection and visual grounding datasets, demonstrating strong performance in object detection, zero-shot localization and grounded reasoning while maintaining its original image and language understanding capabilities on seven standard benchmark datasets."},"_bibtex":{"value":"@misc{\nbhowmik2024learning,\ntitle={Learning to Ground {VLM}s without Forgetting},\nauthor={Aritra Bhowmik and Mohammad Mahdi Derakhshani and Dennis Koelma and Martin R. Oswald and Yuki M Asano and Cees G. M. Snoek},\nyear={2024},\nurl={https://openreview.net/forum?id=HjoYVtSkT8}\n}"},"title":{"value":"Learning to Ground VLMs without Forgetting"},"pdf":{"value":"/pdf/804112808e91e3fbc20897a88e0fc87ec5ced894.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"bhowmik|learning_to_ground_vlms_without_forgetting"},"authorids":{"value":["~Aritra_Bhowmik1","~Mohammad_Mahdi_Derakhshani2","~Dennis_Koelma1","~Martin_R._Oswald1","~Yuki_M_Asano1","~Cees_G._M._Snoek1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Aritra Bhowmik","Mohammad Mahdi Derakhshani","Dennis Koelma","Martin R. Oswald","Yuki M Asano","Cees G. M. Snoek"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewer for their comments. To our understanding, the reviewer was primarily concerned with whether our findings suggest a strong physics prior or poor inductive reasoning, and how our experiments are connected to child cognition experiments.\n\n**Summary of the rebuttal**\n\n1. **Strong physics prior vs inductive physical reasoning (IPR)**: LMMs perform poorly in the SB scenario (Fig. 4, page 8), although it does not contradict any physics prior. This indicates the observed failure is not due to any strong physics prior.\n2. **Connection to cognition experiments**: [R4, R5, R6] showed that infants and children can adapt their internal model when provided with contradicting physical environments, in addition to the violation of expectation.\n3. **Simple dataset design** allowed us to establish the absence of IPR in LMMs without confounding factors. If an LMM struggles in IPR on a simple scene, it can only perform worse in a more complex scene. **New experiments on a more complex version** of InPhyRe (App. E.8, page 41) confirm this.\n4. **Explaining poor IPR**: We have **added new preliminary findings** (App. F, page 45) using linear probes and attention maps:\n    1. The hidden states of a corresponding regular-irregular scenario pair carry enough information to distinguish between the scenarios,\n    2. A model fine-tuned on InPhyRe adapted its hidden states to scenario changes, and\n    3. Image tokens get less attention by an order of magnitude compared to text tokens.\n\n> W1. My biggest concern is the positioning of this paper. The authors select examples that violate physical principles in order to avoid being covered by the LLM's pretrained parametric knowledge. However, I think the model's weak performance in such an environment only indicates that it is dominated by a strong physics prior, not necessarily that its inductive reasoning ability is weak. To the best of my knowledge, experiments related to the violation-of-expectation paradigm in psychology [1] have shown that human infants, when encountering phenomena that appear to violate their expectations, also tend to assume their internal world model is correct, rather than overturning their established mental simulation model based on a few sampled trajectories.\n\n> Q1. The authors argue that \"inductive physical reasoning is a hallmark of intelligence that humans develop at a very young age.\" Previous experiments on inductive and physical reasoning in infants were mostly conducted in real-world environments and did not involve scenarios that violate real-world physics. Could the authors provide further human studies to support this claim?\n\nThis is an excellent perspective. We provide two counter-arguments against the possibility of a strong physics prior causing our observations — (1) LMMs’ failure in the SB scenario, and (2) connection to cognitive science.\n\n1. **LMMs fail in SB despite no contradicting physics**. The SB scenario only challenges the visual prior “large = heavy” that LMMs may have. If LMMs had a strong physics prior, they should have succeeded in SB. Instead, they ignore the context (”the small object is heavier”) and fail in SB. A tabular form of the demos’ relation with physics prior is given below:\n\n    | Demos’ relation with physics prior | Performance | Experiment |\n    | --- | --- | --- |\n    | Align | Good | Sec. 4.3, page 7 |\n    | Contradict | Poor | Sec. 4.4, page 7 (all scenarios except SB) |\n    | Does not contradict | Poor | Sec. 4.4, page 7 (SB scenario) |\n\n    This confirms that LMMs lack inductive ability.\n\n2. **Cognitive experiments show that infants/children adapt to changing environments** after the initial surprise due to violation-of-expectation (VoE):\n    1. [R2]: Infants/toddlers adapted when the friction in opening/closing the drawer was temporally altered. See force adaptation and prior works discussed in the first paragraph of [R2].\n    2. [R3]: Infants adjusted their probabilistic understanding from physical constraints conditioned on color (like Red-LMC and Red-Pass). See the discussion on page 19 of [R3].\n    3. [R4]: Infants “used the habituation event to calibrate their predictions about the test events” when the object sizes in the experiment were changed. See abstract of [R4].\n    \n    These works and physical reasoning benchmarks that used VoE are also discussed in App. C (page 22)."},"title":{"value":"Response to reviewer q4tW -  part 1/2"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763729674531,"tcdate":1763729674531,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14468/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission14468/Authors"],"forum":"IIrPoZ28dN","number":6,"license":"CC BY 4.0","cdate":1763729674531,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14468/-/Official_Comment"],"mdate":1763729674531,"domain":"ICLR.cc/2026/Conference","replyto":"IvGTPyrkDR","id":"oyhV7SdSY0","forumContent":{"TLDR":{"value":"We evaluate whether large multimodal models can infer and apply unseen physical laws from demonstration samples."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["large multimodal models","physical reasoning","inductive reasoning"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large multimodal models (LMMs) encode universal physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potential collision event from visual input. However, since parametric knowledge includes only the physical laws seen during training, it is insufficient for reasoning when the inference scenario follows physical laws unseen during training. In contrast, humans can adapt their physical reasoning to unseen physical environments with only a few visual examples. This inductive physical reasoning ability is indispensable for LMMs if they are to replace human agents in safety-critical applications. Despite its importance, existing visual benchmarks evaluate only the parametric knowledge in LMMs, and not inductive physical reasoning. To this end, we propose InPhyRe, the first visual question answering benchmark to measure inductive physical reasoning in LMMs. InPhyRe evaluates LMMs on their ability to predict the outcome of collision events in algorithmically generated synthetic videos. By inspecting over 13 open-source and proprietary LMMs, InPhyRe informs us that (1) LMMs struggle to apply their limited parametric knowledge about universal physical laws to reasoning, (2) inductive physical reasoning in LMMs is weak when inference scenarios obey physical laws unseen during training, and (3) inductive physical reasoning in LMMs suffers from language bias and largely ignores the visual inputs, questioning the trustworthiness of LMMs regarding visual inputs."},"_bibtex":{"value":"@misc{\nsreekumar2026inphyre,\ntitle={InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning},\nauthor={Gautam Sreekumar and Vishnu Boddeti},\nyear={2026},\nurl={https://openreview.net/forum?id=IIrPoZ28dN}\n}"},"title":{"value":"InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning"},"pdf":{"value":"/pdf/b87a0e49daf5f5f05caff235d97e823b17b5dbfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sreekumar|inphyre_discovers_large_multimodal_models_struggle_in_inductive_physical_reasoning"},"authorids":{"value":["~Gautam_Sreekumar1","~Vishnu_Boddeti1"]},"authors":{"value":["Gautam Sreekumar","Vishnu Boddeti"]}},"version":2},{"content":{"TLDR":{"value":"We present the General Physics Transformer (GPhyT), trained on 1.8 TB of diverse simulation data, that demonstrates foundation model capabilities such as zero-shot generalization to unseen physics are achievable for physics."},"venue":{"value":"AI4Physics"},"pdf":{"value":"/pdf/5b5f53d8abd4813f99c4ecd54dc1a84b85df1f26.pdf","readers":["everyone"]},"keywords":{"value":["Physics Foundation Model","Multi-physics Learning","In-context Learning","Zero-shot Generalization","Scientific Machine Learning","Physics-Aware Machine Learning","Spatiotemporal Transformers"]},"venueid":{"value":"ICML.cc/2026/Workshop/AI4Physics"},"paperhash":{"value":"wiesner|towards_a_physics_foundation_model"},"authorids":{"value":["~Florian_Wiesner1","~Zoë_J._Gray1","~Matthias_Wessling1","~Stephen_Baek1"]},"abstract":{"value":"Foundation models have revolutionized natural language processing through a \"train once, deploy anywhere\" paradigm, where a single pre-trained model adapts to countless downstream tasks without retraining. Access to a Physics Foundation Model (PFM) would be transformative---democratizing access to high-fidelity simulations, accelerating scientific discovery, and eliminating the need for specialized solver development. Yet current physics-aware machine learning approaches remain fundamentally limited to single, narrow domains and require retraining for each new system. We present the General Physics Transformer (GPhyT), trained on 1.8 TB of diverse simulation data, that demonstrates foundation model capabilities are achievable for physics. Our key insight is that transformers can learn to infer governing dynamics from context, enabling a single model to simulate fluid-solid interactions, shock waves, thermal convection, and multi-phase dynamics without being told the underlying equations. GPhyT achieves three critical breakthroughs: (1) superior performance across multiple physics domains, outperforming SOTA multi-physics architectures by more than 7x, (2) plausible zero-shot generalization to entirely unseen physical systems through in-context learning, and (3) more stable long-term predictions through long-horizon rollouts. By establishing that a single model can learn generalizable physical principles from data alone, this work opens the path toward a universal PFM that could transform computational science and engineering."},"_bibtex":{"value":"@inproceedings{\nwiesner2026towards,\ntitle={Towards a Physics Foundation Model},\nauthor={Florian Wiesner and Zo{\\\"e} J. Gray and Matthias Wessling and Stephen Baek},\nbooktitle={ICML 2026 Workshop on AI for Physics},\nyear={2026},\nurl={https://openreview.net/forum?id=3OMwjNsx90}\n}"},"title":{"value":"Towards a Physics Foundation Model"},"authors":{"value":["Florian Wiesner","Zoë J. Gray","Matthias Wessling","Stephen Baek"]}},"tmdate":1784564011749,"pdate":1780110848354,"tcdate":1778178170239,"writers":["ICML.cc/2026/Workshop/AI4Physics","ICML.cc/2026/Workshop/AI4Physics/Submission98/Authors"],"signatures":["ICML.cc/2026/Workshop/AI4Physics/Submission98/Authors"],"forum":"3OMwjNsx90","license":"CC BY 4.0","number":98,"cdate":1778178170239,"readers":["everyone"],"invitations":["ICML.cc/2026/Workshop/AI4Physics/-/Submission","ICML.cc/2026/Workshop/AI4Physics/-/Post_Submission","ICML.cc/2026/Workshop/AI4Physics/-/Edit","ICML.cc/2026/Workshop/AI4Physics/Submission98/-/Camera_Ready_Revision"],"mdate":1784564011749,"odate":1782846807644,"domain":"ICML.cc/2026/Workshop/AI4Physics","id":"3OMwjNsx90","version":2},{"content":{"venue":{"value":"CIS (2) 2009"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/5375729/5375751/05375990.pdf"},"venueid":{"value":"dblp.org/conf/CIS/2009"},"paperhash":{"value":"mo|performance_analysis_of_the_artificial_physics_optimization_algorithm_with_simple_neighborhood_topologies"},"authorids":{"value":["","~Jianchao_Zeng1"]},"html":{"value":"https://doi.org/10.1109/CIS.2009.195"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cis/MoZ09,\n  author={Simin Mo and Jianchao Zeng},\n  title={Performance Analysis of the Artificial Physics Optimization Algorithm with Simple Neighborhood Topologies},\n  year={2009},\n  cdate={1230768000000},\n  pages={155-160},\n  url={https://doi.org/10.1109/CIS.2009.195},\n  booktitle={CIS (2)},\n  crossref={conf/cis/2009-2}\n}\n"},"abstract":{"value":"As a novel population-based optimization technique, the artificial physics optimization (APO) algorithm inspired by physics is presented recently. Although it is characteristic of rapid convergence speed, it also suffers from worse diversity and premature convergence. Accordingly, drawing lessons from those strategies for interactions among individuals in other algorithms, this paper presents the local artificial physics optimization (LAPO) algorithm both to apply it under some simply topologies and to get an insight into the effect of structures. For this end, the performances of LAPO algorithm under particular topologies are investigated by the gravitation constant G adjusting. Simulation results show that LAPO algorithm is valid under some neighborhood structures and that the gravitation constant G has a great influence on performance of LAPO algorithm with different topologies. Also, the results from simulation indicate that the presented LAPO algorithm is superior to APO algorithm so long as parameter G is selected properly under particular topologies."},"title":{"value":"Performance Analysis of the Artificial Physics Optimization Algorithm with Simple Neighborhood Topologies"},"authors":{"value":["Simin Mo","Jianchao Zeng"]}},"tmdate":1772244731386,"pdate":1262217600000,"externalIds":["dblp:conf/cis/MoZ09"],"tcdate":1772244570904,"writers":["~"],"signatures":["~Jianchao_Zeng1"],"forum":"AMgU6Flhh6","license":"CC BY-SA 4.0","number":837090,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772244731386,"domain":"DBLP.org","id":"AMgU6Flhh6","version":2},{"content":{"summary":{"value":"This paper proposes DiffPhy, a framework to improve the physical realism of video diffusion models.The method first uses a Large Language Model (LLM) to \"think\" about the input prompt, reasoning about physical laws, entities, and phenomena to create a set of physical rules and an \"enhanced prompt\". Experiments show DiffPhy achieves SOTA results on physics-based benchmarks, VideoPhy2 and PhyGenBench."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Overall, this paper presents a novel and well-executed framework for a challenging and important problem. The idea of using an MLLM as a \"physics verifier\" inside the diffusion training loop is a significant methodological contribution, and the strong empirical results validate its effectiveness. While I have concerns regarding the training scalability and the method's reliance on the quality of the MLLM supervisor, the novelty and strong results marginally outweigh these weaknesses. I am leaning towards acceptance."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper tackles a well-known and critical limitation of modern VDMs: their failure to adhere to basic physical laws, which breaks realism.\n- The \"Think before you diffuse\" paradigm is an intuitive and strong conceptual contribution. The use of an LLM for high-level physical reasoning to generate supervisory signals for a VDM is a novel training strategy.\n- The creation and release of the HQ-Phy dataset (8,000 curated, real-world videos labeled with physical phenomena) is a valuable contribution to the community, as existing datasets were often synthetic or for evaluation only."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The proposed training paradigm appears to be extremely computationally expensive. More computation and time cost analysis is needed here.\n- The entire method's effectiveness is contingent on the quality of the MLLM used for supervision. The paper itself notes in its limitations section that MLLMs \"struggle to interpret videos\" and their outputs can be \"shallow or generic.\" This raises a significant concern: if the MLLM is an imperfect evaluator (which is also suggested by the discrepancy between model-based and human-based scores), is DiffPhy simply learning to overfit to the specific biases of the VideoCon-Physics evaluator? How can we be sure it's learning generalizable physical rules rather than just optimizing for this specific MLLM's scoring function?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918558279,"tcdate":1761975082879,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6229/Reviewer_myHY"],"signatures":["ICLR.cc/2026/Conference/Submission6229/Reviewer_myHY"],"forum":"lPKsPBstHg","number":4,"license":"CC BY 4.0","cdate":1761975082879,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6229/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918558279,"domain":"ICLR.cc/2026/Conference","replyto":"lPKsPBstHg","id":"0lKAUr0rnl","forumContent":{"TLDR":{"value":"We propose DiffPhy, a generic framework that enables physically-correct and semantically coherent video generation by fine-tuning a pre-trained video diffusion model."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Diffusion","Video Generation","Physical Commensense"]},"supplementary_material":{"value":"/attachment/dd73e8deab01feb18bdae06f4f6ecc4dca24d57c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions, interactions, and dynamics introduce great difficulties when learning physics from data. In this work, we propose DiffPhy, a generic framework that enables physically-correct and photo-realistic video generation by fine-tuning a pre-trained video diffusion model. Our method leverages large language models (LLMs) to infer rich physical context from the text prompt. To incorporate this context into the video diffusion model, we use a multimodal large language model (MLLM) to verify intermediate latent variables against the inferred physical rules, guiding the model’s gradient updates accordingly. MLLM’s textual output is transformed into continuous signals. We then formulate a set of training objectives that jointly ensure physical accuracy and semantic alignment with the input text.  Additionally, failure facts of physical phenomena are corrected via attention injection. We also establish a high-quality physical video dataset containing diverse phyiscal actions and events to facilitate effective finetuning. Extensive experiments on public benchmarks demonstrate that DiffPhy is able to produce state-of-the-art results across diverse physics-related scenarios. Code and data will be made available post-review."},"_bibtex":{"value":"@misc{\nzhang2026think,\ntitle={Think Before You Diffuse: Infusing Physical Rules into Video Diffusion},\nauthor={Ke Zhang and Cihan Xiao and Jiacong Xu and Yiqun Mei and Vishal M. Patel},\nyear={2026},\nurl={https://openreview.net/forum?id=lPKsPBstHg}\n}"},"title":{"value":"Think Before You Diffuse: Infusing Physical Rules into Video Diffusion"},"pdf":{"value":"/pdf/abed66725f92f4403a0e6585ebdac7432376d20d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|think_before_you_diffuse_infusing_physical_rules_into_video_diffusion"},"authorids":{"value":["~Ke_Zhang17","~Cihan_Xiao1","~Jiacong_Xu1","~Yiqun_Mei1","~Vishal_M._Patel1"]},"authors":{"value":["Ke Zhang","Cihan Xiao","Jiacong Xu","Yiqun Mei","Vishal M. Patel"]}},"version":2},{"content":{"summary":{"value":"The paper proposes using graph neural networks (GNNs) to jointly learn interaction rules and heterogeneous structure in complex dynamical systems from data alone. Extensive experiments on simulated systems including particle interactions, wave propagation, reaction-diffusion, and signaling networks, showing its good performance."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The authors mention the method can infer the underlying governing equations, (line 17) but I do not see any analysis in the experiment part. It would be interesting to see how can we extract formula from a learned GNN."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is in general easy to follow and with clear writing flow. The problem is well-motivated, by using GNN to learn system dynamics over time and in the meanwhile, uncover the underlying latent properties in an interpretable way that facilitates further analysis.\n\n2. The evaluation of dynamical systems in the experiment sections are extensive, though adding some baselines for comparison would be better."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is no related work section. Some works are discussed in the introduction part, but there are many existing neural simulators that use GNN to rollout trajectories of multi-agent dynamical systems [1,2,3,4]. Discussion about existing work and comparison in the experiment section are helpful to provide a comprehensive analysis.\n\n2. As mentioned above, for rollout MSE across different datasets, it is suggested to compare against representative baselines. Also the run time comparison can be included across compared methods. \n\n\n\n\n\n[1]  Learning Continuous System Dynamics from Irregularly-Sampled Partial Observations.\n\n[2] Interaction Networks for Learning about Objects, Relations and Physics.\n\n[3] Learning to simulate complex physics with graph networks.\n\n[4] HOPE: High-order Graph ODE For Modeling Interacting Dynamics"}},"nonreaders":[],"tmdate":1731429580616,"tcdate":1731395986276,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12051/Reviewer_tcQf"],"signatures":["ICLR.cc/2025/Conference/Submission12051/Reviewer_tcQf"],"forum":"7FQDHv9fD4","number":4,"license":"CC BY 4.0","cdate":1731395986276,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12051/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429580616,"domain":"ICLR.cc/2025/Conference","replyto":"7FQDHv9fD4","id":"TCGxhEd22P","forumContent":{"TLDR":{"value":"We use graph neural networks to decompose heterogeneous dynamic systems and reveal the underlying interactions."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["graph neural networks","gnn","dynamic system","latent parameter discovery"]},"supplementary_material":{"value":"/attachment/a2a217bed4c64b6bc28fc615e84b2da38574623d.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Natural physical, chemical, and biological dynamical systems are often complex, with heterogeneous components interacting in diverse ways. We show how simple graph neural networks can be designed to jointly learn the interaction rules and the latent heterogeneity from observable dynamics. The learned latent heterogeneity and dynamics can be used to virtually decompose the complex system which is necessary to infer and parameterize the underlying governing equations. We tested the approach with simulation experiments of interacting moving particles, vector fields, and signaling networks. While our current aim is to better understand and validate the approach with simulated data, we anticipate it to become a generally applicable tool to uncover the governing rules underlying complex dynamics observed in nature."},"_bibtex":{"value":"@misc{\nallier2025decomposing,\ntitle={Decomposing heterogeneous dynamical systems with graph neural networks},\nauthor={Cedric Allier and Magdalena C. Schneider and Michael Innerberger and Larissa Heinrich and John A. Bogovic and Stephan Saalfeld},\nyear={2025},\nurl={https://openreview.net/forum?id=7FQDHv9fD4}\n}"},"title":{"value":"Decomposing heterogeneous dynamical systems with graph neural networks"},"pdf":{"value":"/pdf/7fe76b9b33733b68e6890d030f74d79cc7b0c68f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"allier|decomposing_heterogeneous_dynamical_systems_with_graph_neural_networks"},"authorids":{"value":["~Cedric_Allier1","~Magdalena_C._Schneider1","~Michael_Innerberger1","~Larissa_Heinrich1","~John_A._Bogovic1","~Stephan_Saalfeld1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Cedric Allier","Magdalena C. Schneider","Michael Innerberger","Larissa Heinrich","John A. Bogovic","Stephan Saalfeld"]}},"version":2},{"content":{"summary":{"value":"This paper introduces DITTO, a framework designed to scale instruction-based video editing models. The core contribution is a novel, large-scale, high-quality synthetic instruction-video-edit dataset (DITTO-Data), created through an innovative pipeline that leverages powerful pre-trained Large Language Models (LLMs) and diffusion models (DMs) to automatically generate diverse, complex editing instructions and the corresponding edited videos. Based on this dataset, the authors train DITTO-Model, a video editing model which demonstrates strong capabilities in instruction following, temporal consistency, and maintaining content fidelity. The experiments show DITTO-Model achieving state-of-the-art results on several benchmarks, particularly excelling in complex, style-based, and semantic edits, validated by both quantitative metrics and comprehensive human evaluation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Generalization to Real Edits: While the human study uses the synthetic dataset, how does DITTO-Model perform when asked to execute complex instructions on out-of-distribution real-world videos that might contain more unusual or messy degradation patterns not fully captured by the synthetic base videos?\n2. Ablation on LLM Prompting: Could the authors provide more detail, perhaps in the appendix, on the meta-prompts used to guide the LLM to generate the diverse and complex editing instructions? This \"prompt engineering\" is critical to the dataset's quality."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. High-Quality, Scalable Data Generation: The synthetic data pipeline is the major strength, addressing the prohibitive cost and complexity of manual video editing data collection. The use of LLMs for instruction diversity is particularly effective.\n2. State-of-the-Art Performance: DITTO-Model achieves superior results across multiple metrics, notably in human evaluation on Instruction Following and Temporal Consistency, which are crucial aspects of video editing.\n3. Instruction Complexity: The generated dataset and resulting model are shown to handle a wide range of instruction complexities, including appearance transformation, style transfer, and semantic manipulation, moving beyond simple object insertion/removal."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Black-Box Data Quality: While the paper describes the Quality Control module, the extent to which the synthetic data truly captures the complexity and subtle detail of real-world human-labeled edits is hard to quantify. Further analysis on the \"failure modes\" of the synthetic pipeline and the resulting data distribution bias would be beneficial.\n2. Model Architecture Novelty: The DITTO-Model architecture itself is largely an assembly of existing, robust components (latent diffusion model, motion modules). The novelty lies more in the data and training strategy than the architectural innovations."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915806348,"tcdate":1761402055489,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1545/Reviewer_AcEZ"],"signatures":["ICLR.cc/2026/Conference/Submission1545/Reviewer_AcEZ"],"forum":"qUJZX8LwMp","number":2,"license":"CC BY 4.0","cdate":1761402055489,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1545/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915806348,"domain":"ICLR.cc/2026/Conference","replyto":"qUJZX8LwMp","id":"AGKf1qP76S","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"To solve the data scarcity problem, we introduce a scalable pipeline Ditto for generating high-quality video editing data, which is used to train a new state-of-the-art instruction-based video editing model Editto."},"keywords":{"value":["Instruction-based video editing","diffusion models","synthetic ddataset"]},"supplementary_material":{"value":"/attachment/8daec567506b4e900bb3e7df424ee18ae21ccece.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel data generation pipeline that fuses the creative diversity of a leading image editor with an in-context video generator, overcoming the limited scope of existing models. To make this process viable, our framework resolves the prohibitive cost-quality trade-off by employing an efficient, distilled model architecture augmented by a temporal enhancer, which simultaneously reduces computational overhead and improves temporal coherence.  Finally, to achieve full scalability, this entire pipeline is driven by an intelligent agent that crafts diverse instructions and rigorously filters the output, ensuring quality control at scale. Using this framework, we invested over 12,000 GPU-days to build Ditto-1M, a new dataset of one million high-fidelity video editing examples. We trained our model, Editto, on Ditto-1M with a curriculum learning strategy.  The results demonstrate superior instruction-following ability and establish a new state-of-the-art in instruction-based video editing. We will release our dataset, models, and code to accelerate research in this field."},"_bibtex":{"value":"@misc{\nbai2025ditto,\ntitle={Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset},\nauthor={Qingyan Bai and Qiuyu Wang and Hao Ouyang and Hanlin Wang and Wen Wang and Ka Leong Cheng and Shuailei Ma and Yanhong Zeng and Yue Yu and Zichen Liu and Yinghao Xu and Yujun Shen and Qifeng Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=qUJZX8LwMp}\n}"},"title":{"value":"Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset"},"pdf":{"value":"/pdf/afa484d3376a91bed3da272265c4ecdd9e32ea7c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"bai|ditto_scaling_instructionbased_video_editing_with_a_highquality_synthetic_dataset"},"authorids":{"value":["~Qingyan_Bai1","~Qiuyu_Wang1","~Hao_Ouyang2","~Hanlin_Wang2","~Wen_Wang7","~Ka_Leong_Cheng2","~Shuailei_Ma1","~Yanhong_Zeng1","~Yue_Yu23","~Zichen_Liu7","~Yinghao_Xu1","~Yujun_Shen1","~Qifeng_Chen1"]},"authors":{"value":["Qingyan Bai","Qiuyu Wang","Hao Ouyang","Hanlin Wang","Wen Wang","Ka Leong Cheng","Shuailei Ma","Yanhong Zeng","Yue Yu","Zichen Liu","Yinghao Xu","Yujun Shen","Qifeng Chen"]}},"version":2},{"content":{"venue":{"value":"NeurIPS Datasets and Benchmarks 2021"},"pdf":{"value":"https://datasets-benchmarks-proceedings.neurips.cc/paper_files/paper/2021/file/28dd2c7955ce926456240b2ff0100bde-Paper-round2.pdf"},"venueid":{"value":"dblp.org/conf/NIPS/2021"},"paperhash":{"value":"makoviychuk|isaac_gym_high_performance_gpu_based_physics_simulation_for_robot_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Viktor_Makoviychuk:","https://dblp.org/search/pid/api?q=author:Lukasz_Wawrzyniak:","~Yunrong_Guo1","https://dblp.org/search/pid/api?q=author:Michelle_Lu:","https://dblp.org/search/pid/api?q=author:Kier_Storey:","https://dblp.org/search/pid/api?q=author:Miles_Macklin:","https://dblp.org/search/pid/api?q=author:David_Hoeller:","https://dblp.org/search/pid/api?q=author:Nikita_Rudin:","https://dblp.org/search/pid/api?q=author:Arthur_Allshire:","~Ankur_Handa1","https://dblp.org/search/pid/api?q=author:Gavriel_State:"]},"html":{"value":"https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/28dd2c7955ce926456240b2ff0100bde-Abstract-round2.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/nips/MakoviychukWGLS21,\n  author={Viktor Makoviychuk and Lukasz Wawrzyniak and Yunrong Guo and Michelle Lu and Kier Storey and Miles Macklin and David Hoeller and Nikita Rudin and Arthur Allshire and Ankur Handa and Gavriel State},\n  title={Isaac Gym: High Performance GPU Based Physics Simulation For Robot Learning},\n  year={2021},\n  cdate={1609459200000},\n  url={https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/28dd2c7955ce926456240b2ff0100bde-Abstract-round2.html},\n  booktitle={NeurIPS Datasets and Benchmarks},\n  crossref={conf/nips/2021db}\n}\n"},"abstract":{"value":"Isaac Gym offers a high-performance learning platform to train policies for a wide variety of robotics tasks entirely on GPU. Both physics simulation and neural network policy training reside on GPU and communicate by directly passing data from physics buffers to PyTorch tensors without ever going through CPU bottlenecks. This leads to blazing fast training times for complex robotics tasks on a single GPU with 2-3 orders of magnitude improvements compared to conventional RL training that uses a CPU-based simulator and GPUs for neural networks. We host the results and videos at https://sites.google.com/view/isaacgym-nvidia and Isaac Gym can be downloaded at https://developer.nvidia.com/isaac-gym. The benchmark and environments are available at https://github.com/NVIDIA-Omniverse/IsaacGymEnvs."},"title":{"value":"Isaac Gym: High Performance GPU Based Physics Simulation For Robot Learning"},"authors":{"value":["Viktor Makoviychuk","Lukasz Wawrzyniak","Yunrong Guo","Michelle Lu","Kier Storey","Miles Macklin","David Hoeller","Nikita Rudin","Arthur Allshire","Ankur Handa","Gavriel State"]}},"tmdate":1741236672614,"pdate":1609459200000,"tcdate":1731483778165,"writers":["~"],"signatures":["~Ankur_Handa1"],"forum":"RTwgPEORxp","license":"CC BY-SA 4.0","number":213700,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1741236672614,"domain":"DBLP.org","id":"RTwgPEORxp","version":2},{"content":{"summary":{"value":"This paper presents CMPhysBench, a new benchmark to test LLMs on graduate-level Condensed Matter Physics. It's made of 520+ hard calculation problems from textbooks. They also created a new scoring metric called SEED that gives partial credit for complex math answers (like equations or tuples) using tree-based analysis. Their tests on 18 LLMs show that even the best models, like Grok 4, perform poorly (36 SEED score, 28.9% accuracy), showing a big gap in this specific domain."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"For the error analysis, how do you know GPT-4o's categorizations are correct? Did you have any humans double-check its work? Does your evaluation look at the reasoning steps, or just the final answer in the box? It seems possible for a model to get the right answer by luck or by making mistakes that cancel out."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The paper's main strength is tackling a new, hard domain: graduate-level condensed matter physics. Most benchmarks are easier, so this is a needed step up. The SEED metric is also a big plus; it's a smart way to give partial credit on complex math answers instead of just right/wrong. This metric seems useful for other science benchmarks too. The testing of 18 models is thorough, and the error analysis in Figure 6 gives a good breakdown of why models fail, with \"Concept and Model Misuse\" being the biggest problem."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The main weakness I see is in the error analysis. The authors used GPT-4o to categorize all the model mistakes. While this is fast, it's not clear how accurate GPT-4o is at this task. It would be better if they had human experts check a sample of these to confirm the error breakdown. Also, the SEED score focuses on the final boxed answer. The prompt asks for step-by-step solutions, but it's not clear if the steps themselves are evaluated. A model could get the right answer with the wrong steps."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942126945,"tcdate":1761977881250,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22231/Reviewer_nemX"],"signatures":["ICLR.cc/2026/Conference/Submission22231/Reviewer_nemX"],"forum":"3d0FRYx0D0","number":3,"license":"CC BY 4.0","cdate":1761977881250,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22231/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942126945,"domain":"ICLR.cc/2026/Conference","replyto":"3d0FRYx0D0","id":"Cj56oo5WDr","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["LLM Benchmark","Condensed Matter Physics","LLM Evaluation","AI for Physics"]},"supplementary_material":{"value":"/attachment/7454ba38bf90b23ca876f4c77f54f001a24045f8.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We introduce CMPhysBench, designed to assess the proficiency of Large Language Models (LLMs) in Condensed Matter Physics, as a novel Benchmark. CMPhysBench is composed of more than 520 graduate-level meticulously curated questions covering both representative subfields and foundational theoretical frameworks of condensed matter physics, such as magnetism, superconductivity, strongly correlated systems, etc. To ensure a deep understanding of the problem-solving process,we focus exclusively on calculation problems, requiring LLMs to independently generate comprehensive solutions. Meanwhile, leveraging tree-based representations of expressions, we introduce the Scalable Expression Edit Distance (SEED) score, which provides fine-grained (non-binary) partial credit and yields a more accurate assessment of similarity between prediction and ground-truth. Our results show that even the best models, Grok-4, reach only 36 average SEED score and 29% accuracy on CMPhysBench, underscoring a significant capability gap, especially for this practical and frontier domain relative to traditional physics."},"_bibtex":{"value":"@inproceedings{\nwang2026cmphysbench,\ntitle={{CMP}hysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics},\nauthor={Weida Wang and Dongchen Huang and Jiatong LI and Tengchao Yang and Ziyang Zheng and Chuyi Peng and Di Zhang and Dong Han and Benteng Chen and Binzhao Luo and Zhiyu Liu and kunling liu and Zhiyuan Gao and Shiqigeng and Wei Ma and Jiaming Su and Xin Li and Shuchen Pu and Yuhan Shui and Qianjia Cheng and Zhihao Dou and Dongfei Cui and Changyong He and Jin Zeng and Zeke Xie and Mao Su and Dongzhan Zhou and Yuqiang Li and Wanli Ouyang and Yunqi Cai and Xi Dai and Shufei Zhang and LEI BAI and Jinguang Cheng and Zhong Fang and Hongming Weng},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=3d0FRYx0D0}\n}"},"title":{"value":"CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics"},"pdf":{"value":"/pdf/d44604df4f81ab08a67bb9ae040f05d3aaaccacc.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|cmphysbench_a_benchmark_for_evaluating_large_language_models_in_condensed_matter_physics"},"authorids":{"value":["~Weida_Wang3","~Dongchen_Huang1","~Jiatong_LI4","~Tengchao_Yang1","~Ziyang_Zheng3","~Chuyi_Peng1","~Di_Zhang8","~Dong_Han5","~Benteng_Chen1","~Binzhao_Luo1","~Zhiyu_Liu3","~kunling_liu1","~Zhiyuan_Gao4","~Shiqigeng1","~Wei_Ma9","~Jiaming_Su1","~Xin_Li105","~Shuchen_Pu1","~Yuhan_Shui1","~Qianjia_Cheng1","~Zhihao_Dou2","~Dongfei_Cui1","~Changyong_He1","~Jin_Zeng1","~Zeke_Xie1","~Mao_Su1","~Dongzhan_Zhou1","~Yuqiang_Li1","~Wanli_Ouyang1","~Yunqi_Cai2","~Xi_Dai1","~Shufei_Zhang1","~LEI_BAI1","~Jinguang_Cheng2","~Zhong_Fang2","~Hongming_Weng1"]},"authors":{"value":["Weida Wang","Dongchen Huang","Jiatong LI","Tengchao Yang","Ziyang Zheng","Chuyi Peng","Di Zhang","Dong Han","Benteng Chen","Binzhao Luo","Zhiyu Liu","kunling liu","Zhiyuan Gao","Shiqigeng","Wei Ma","Jiaming Su","Xin Li","Shuchen Pu","Yuhan Shui","Qianjia Cheng","Zhihao Dou","Dongfei Cui","Changyong He","Jin Zeng","Zeke Xie","Mao Su","Dongzhan Zhou","Yuqiang Li","Wanli Ouyang","Yunqi Cai","Xi Dai","Shufei Zhang","LEI BAI","Jinguang Cheng","Zhong Fang","Hongming Weng"]}},"version":2},{"content":{"summary":{"value":"The paper presents Graphical-TS, an interactive framework designed for causal discovery in multivariate time series (MTS) data. It addresses the challenges researchers face in generating and analyzing MTS data, particularly the redundancy in data generation efforts and the need for effective integration of expert knowledge. The authors argue that domain experts should define causal relationships, while algorithm developers focus on refining their algorithms.\n\nKey contributions of the paper include:\n\nUser Interface Development: The framework features a comprehensive user interface that allows users to create, edit, and visualize causal graphs, facilitating an iterative process of causal discovery and model refinement.\n\nIntegration with Causal Discovery Algorithms: Graphical-TS integrates state-of-the-art causal discovery algorithms, enabling users to generate initial causal graphs from data and refine them based on domain knowledge.\n\nSynthetic Data Generation: The interface supports the generation of synthetic MTS data that reflects specified causal relationships, which is crucial for hypothesis testing and model validation.\n\nCollaborative Environment: The framework promotes collaboration among researchers by allowing multiple users to work on the same causal graph, complete with version control features.\n\nEnhanced Causal Modeling: By enabling the simulation of complex spatiotemporal dependencies, Graphical-TS advances the methodology for evaluating time series forecasting and imputation algorithms, thereby improving the precision of algorithmic evaluations.\n\nOverall, Graphical-TS serves as a valuable tool for integrating human expertise with algorithmic processes, enhancing the representation of causal dynamics in various research fields, including healthcare and finance."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Empirical Validation of Graphical-TS\n\nQuestion: Can you provide specific examples of real-world datasets where Graphical-TS has been applied? What were the outcomes of these applications?\n\nSuggestion: Including empirical results from case studies would strengthen the paper. Consider adding a section that details the application of Graphical-TS in a specific domain, such as healthcare or finance, along with quantitative metrics that demonstrate its effectiveness compared to existing methods.\n\nLimitations of Synthetic Data\n\nQuestion: What measures have been taken to ensure that the synthetic data generated by Graphical-TS accurately reflects real-world complexities? How do you address potential biases in this data?\n\nSuggestion: A discussion on the limitations of synthetic data generation and how it may impact the validity of the causal models would be valuable. Consider exploring hybrid approaches that combine real-world data with synthetic data to enhance robustness.\n\nTheoretical Foundations\n\nQuestion: Can you elaborate on the theoretical principles that guided the design of Graphical-TS? Why were certain algorithms or methodologies chosen over others?\n\nSuggestion: Providing a more detailed theoretical justification for the framework's design choices would enhance its credibility. Including references to relevant literature that supports these choices could also strengthen the argument.\n\nScalability and High-Dimensional Data\n\nQuestion: How does Graphical-TS handle scalability issues, particularly when dealing with high-dimensional datasets? Are there any limitations in this regard?\n\nSuggestion: Addressing potential scalability challenges and discussing how the framework can be adapted to handle high-dimensional data would be important. Consider including performance benchmarks or examples that illustrate how the system performs under varying data conditions.\n\nIntegration of Expert Knowledge\n\nQuestion: How is expert knowledge integrated into the causal discovery process? What mechanisms are in place to ensure that this knowledge is accurately represented in the graphical models?\n\nSuggestion: A clearer explanation of how expert input is incorporated and refined within the framework would be beneficial. Discussing the iterative process of integrating human insights with algorithmic learning could provide a more comprehensive understanding of the system's functionality.\n\nFuture Directions and Improvements\n\nQuestion: What are the planned future enhancements for Graphical-TS? Are there specific features or functionalities that you aim to develop based on user feedback?\n\nSuggestion: Outlining future directions for the framework, including potential improvements or new features, would provide insight into the authors' vision for the system. This could also encourage collaboration and engagement from the research community."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Originality: The paper presents Graphical-TS, an innovative framework that integrates expert knowledge with causal discovery algorithms for multivariate time series (MTS) analysis. This originality is evident in several dimensions:\nNew Problem Formulation: The authors address the challenge of effectively incorporating domain expertise into causal modeling, which has been a limitation in traditional causal discovery methods. By emphasizing the role of human experts in defining causal relationships, the paper redefines the approach to causal discovery in time series data.\n\n\nCreative Combination of Ideas: The framework combines interactive user interfaces with advanced causal discovery algorithms, allowing for an iterative refinement process. This creative integration enhances the usability and applicability of causal discovery methods, making them more accessible to researchers and practitioners.\n\n\nApplication to New Domains: The focus on generating synthetic data that reflects real-world causal relationships opens up new avenues for research in various fields, including healthcare and finance, where accurate causal modeling is crucial.\n\n\nQuality: The quality of the paper is commendable, characterized by:\nRobust Methodology: The authors provide a comprehensive description of the framework's functionalities, including graph creation, editing, and synthetic data generation. The integration with state-of-the-art causal discovery algorithms is well-articulated, demonstrating a solid understanding of the field.\n\n\nEmpirical Validation: While the paper outlines the framework's capabilities, further empirical results and case studies would enhance the quality of the claims made. However, the theoretical foundation and proposed methodologies are sound and well-supported.\n\n\nClarity: The paper is generally well-structured and clear, with strengths in:\nWriting Style: The writing is coherent and accessible, making complex concepts understandable for a broad audience. The use of figures and examples effectively illustrates key points, aiding in reader comprehension.\n\n\nContextualization: The authors successfully situate their work within the existing literature, highlighting the limitations of current methodologies and the need for expert integration. This contextualization enhances the clarity of the paper's contributions.\n\n\nSignificance: The significance of the paper is substantial, as it addresses pressing issues in causal discovery and MTS analysis:\nImpact on Research Community: By providing a framework that enhances the representation of causal dynamics and facilitates rigorous evaluations of forecasting algorithms, the paper contributes valuable tools for researchers in various domains. The potential for improved accuracy in causal modeling has far-reaching implications for fields that rely on data-driven decision-making.\n\n\nBroader Applications: The ability to generate synthetic datasets that reflect real-world complexities is particularly significant in privacy-sensitive areas, allowing researchers to test hypotheses without compromising sensitive information."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the paper presents a compelling framework in Graphical-TS, there are several areas where it could be improved to enhance its contributions and effectiveness. Below are specific weaknesses along with constructive and actionable insights for improvement:\n\n1. Empirical Validation and Case Studies\nWeakness: The paper lacks comprehensive empirical validation of the Graphical-TS framework. While it outlines the functionalities and theoretical underpinnings, there are limited real-world case studies or experiments demonstrating its effectiveness in practice.\n\nActionable Insight:\n\nIncorporate Case Studies: The authors should include detailed case studies that showcase the application of Graphical-TS in real-world scenarios, particularly in fields like healthcare or finance. This could involve analyzing existing datasets to illustrate how the framework improves causal discovery compared to traditional methods.\n\nBenchmarking Against Existing Methods: Conduct systematic comparisons with established causal discovery methods (e.g., PCMCI, GES) using standardized datasets. Presenting quantitative metrics (e.g., precision, recall, F1-score) would provide a clearer picture of the framework's performance and its advantages.\n\n2. User Interface and Usability\n\nWeakness: While the paper mentions an intuitive user interface, it does not provide sufficient detail on how users interact with the system or the specific features that facilitate collaboration between domain experts and data scientists.\n\nActionable Insight:\n\nDetailed User Interface Description: Include screenshots or diagrams of the user interface to illustrate how users can manipulate causal relationships and input expert knowledge. A walkthrough of the user experience would help readers understand the practical implications of the framework.\nUser Feedback and Iteration: Consider conducting user studies or surveys with domain experts to gather feedback on the interface and usability. This could inform iterative improvements and ensure that the system meets the needs of its intended users.\n\n3. Limitations of Synthetic Data Generation\n\nWeakness: The paper discusses the generation of synthetic data but does not adequately address the limitations and potential biases that may arise from this approach. Relying solely on synthetic data could lead to overfitting or misrepresentation of real-world dynamics.\n\nActionable Insight:\n\nDiscuss Limitations: The authors should explicitly discuss the limitations of synthetic data generation, including potential biases and the risk of not capturing the full complexity of real-world systems. This could involve a section dedicated to the challenges and considerations when using synthetic data.\nHybrid Approaches: Explore the possibility of integrating real-world data with synthetic data to create a more robust dataset. This could involve using real data to inform the parameters of the synthetic data generation process, ensuring that the generated data reflects realistic scenarios.\n\n4. Theoretical Foundations and Justifications\n\nWeakness: The theoretical foundations of the framework could be more robustly articulated. While the paper references existing methodologies, it does not sufficiently justify the choices made in the design of Graphical-TS.\n\nActionable Insight:\n\nTheoretical Justification: Provide a more in-depth discussion of the theoretical principles underlying the framework, including why certain algorithms or approaches were chosen over others. This could involve citing relevant literature that supports these choices and discussing their implications for the framework's performance.\n\nAddressing Potential Critiques: Anticipate and address potential critiques of the framework, such as concerns about scalability or the handling of high-dimensional data. Discuss how Graphical-TS can be adapted or improved to address these challenges."}},"nonreaders":[],"tmdate":1731428147449,"tcdate":1730687607455,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13763/Reviewer_uQ8i"],"signatures":["ICLR.cc/2025/Conference/Submission13763/Reviewer_uQ8i"],"forum":"meY36sGyyv","number":2,"license":"CC BY 4.0","cdate":1730687607455,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13763/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428147449,"domain":"ICLR.cc/2025/Conference","replyto":"meY36sGyyv","id":"ojBXjYQaCH","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["time series","causal discovery","benchmarking","interface"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present \\texttt{Graphical-TS}, an interactive simulation framework for multivariate time series (MTS) incorporating spatiotemporal causal graphical models. The system offers extensive customizability, enabling users to define and modify causal dynamics with uncertainty in spatiotemporal relationships and functional mappings. \\texttt{Graphical-TS} integrates expert knowledge, supports MTS simulation, and allows for the input of real-world MTS data, facilitating a dynamic interplay between data-driven learning and domain expertise. The system iteratively enhances causal relationships and simulated data by simulating MTS data based on specified causal graphs, performing causal discovery from real or simulated MTS, and enabling the integration and refinement of expert knowledge with learned causality. This approach progressively improves the quality of causal models and the data they generate, supporting tasks such as time series forecasting, imputation, prediction, and robustness testing via scenario-driven distribution shifts. We compared state-of-the-art causal discovery methods on datasets generated by \\texttt{Graphical-TS}. The empirical results demonstrate the platform’s consistent performance compared to existing methods while offering versatility under distinct scenarios. This enables users to explore datasets more thoroughly and drive improvements in causal discovery research. With an intuitive user interface that connects domain experts and algorithm developers, \\texttt{Graphical-TS} empowers users to manipulate causal relationships, embedding domain knowledge into machine learning workflows. Originally developed to study physiological dynamics in patients, the system has broad applicability across various fields, offering a versatile platform for generating MTS datasets with known dynamics, validating causal discovery algorithms, and advancing research in time series analysis."},"_bibtex":{"value":"@misc{\nli2025graphicalts,\ntitle={Graphical-{TS}: An Interactive {AI} Pipeline for Multivariate Time Series with Ground-truth Graphical Modeling},\nauthor={Haixin Li and Yanke Li and Diego Paez-Granados},\nyear={2025},\nurl={https://openreview.net/forum?id=meY36sGyyv}\n}"},"title":{"value":"Graphical-TS: An Interactive AI Pipeline for Multivariate Time Series with Ground-truth Graphical Modeling"},"pdf":{"value":"/pdf/cf9bd358dcf783b866108cc4dcdedc4ba928c82d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|graphicalts_an_interactive_ai_pipeline_for_multivariate_time_series_with_groundtruth_graphical_modeling"},"authorids":{"value":["~Haixin_Li1","~Yanke_Li2","~Diego_Paez-Granados1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haixin Li","Yanke Li","Diego Paez-Granados"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a novel approach for recovering the density and velocity fields of inviscid fluids from sparse multiview videos. The model has two main contributions:\n- First, it incorporates physics-based losses to enforce the inference of a physically plausible velocity field that is divergence-free and drives the transport of density. This helps to deal with the visual ambiguities of fluid velocity. \n- Second, the model provides a hybrid neural velocity representation, which consists of a base neural velocity field capturing most irrotational energy and a vortex particle-based velocity modeling residual turbulent velocity. This representation enables the recovery of vortical flow details."},"soundness":{"value":"3 good"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"1. In the paper, a grid-based representation is used for density and velocity. When implementing $L_{density}$ and $L_{project}$, how exactly are these calculations performed? Are the loss functions computed for all grid positions, or is interpolation used to compute the losses at sampled points in space?\n\n2. What is the difference in performance between decomposing the velocity field into a base neural velocity field and a vortex particle-based velocity, versus solely using high-frequency position embedding?\n\n3. In the case of using physics-based loss functions in HyFluid, how does the predicted velocity field of HyFluid compare to the ground truth (GT) velocity field quantitatively? Can HyFluid be compared to other methods in terms of velocity field quantitatively?\n\n4. The paper does not provide experimental results and discussions regarding the constraint of constant radiance. What is the capability of HyFluid with real-world scenes that exhibit spatially-varying radiance? \n\n5. Can you provide more quantitative/ qualitative results of other baselines?"},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"strengths":{"value":"Originality: This paper primarily tackles the visual ambiguities in inverse rendering techniques for fluid data, focusing on resolving visual ambiguities. The integration of physics-based losses into the volume rendering framework is a rational approach. The results in Figure 8 illustrate the effectiveness of the newly proposed learning constraints in significantly improving the accuracy of the reconstruction. Moreover, the introduction of a hybrid representation of velocity fields and particle-based vortex flow showcases originality in the methodology.\n\nSignificance: The proposed model makes a significant contribution to the visual understanding of fluids, particularly smoke, fog, and gas.\n"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"For methodology: \n\n1. One contribution of this paper is to incorporate new forms of physical constraints in the framework of inverse rendering. However, the general idea is not entirely novel as previous work by Chu et al. (2022) has also presented a similar (albeit different) approach, which somewhat weakens the technical novelty of this paper. For example, using the density transport equation from incompressible fluid or NS equation as a loss function has been employed in other related papers. Please refer to the work from Baieri et al. (2023) and Li et al. (2023). \n\n- [Baieri et al., 2023] Fluid Dynamics Network: Topology-Agnostic 4D Reconstruction via Fluid Dynamics Priors. Arxiv, 2023. \n- [Li et al., 2023] PAC-NeRF: Physics Augmented Continuum Neural Radiance Fields for Geometry-Agnostic System Identification. ICLR, 2023. \n\n2. The constraint imposed on emitting radiance in HyFluid does present limitations in its applicability to real-world scenarios. The requirement of a constant emitting radiance may hinder accurate recovery in situations where spatially-varying lighting exists in the scene.\n\nLacking references:\n\n3. Some closely related work seems to be missing, such as NeuroFluid (Guan et al., 2022) and PAC-NeRF (Li et al., 2023), both of which also focus on visual physical inference through inverse rendering. It is important to acknowledge these works as they contribute to the existing body of literature in this field and provide valuable insights and techniques for comparison and benchmarking purposes.\n\n- [Guan et al., 2022] NeuroFluid: Fluid Dynamics Grounding with Particle-Driven Neural Radiance Fields. ICML, 2022. \n\nFor experiments: \n\n4. Is the proposed method limited to handling only inviscid fluids? It would be beneficial to evaluate the proposed method in a broader range of fluid scenarios, including different types of flows (e.g., laminar, turbulent) and varying fluid properties (e.g., viscosity, density). This would demonstrate the generalization ability of HyFluid under diverse conditions.\n\n5. Considering that the constraint of constant radiance may not be practical in complex real-world scenes, I highly recommend that the authors compare the reconstruction results obtained with and without the constraint. \n\n6. The proposed model is only compared with two existing models, which is not sufficient. To make the results more convincing, the authors could incorporate more advanced neural rendering techniques designed specifically for dynamic scenes. Additionally, given that the dataset comprises synthetic fluid simulation data, it would be beneficial for the authors to provide quantitative results regarding the reconstruction of the velocity field in comparison to the ground truth on these simulated scenes. "},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1702410782402,"tcdate":1688363291934,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1335/Reviewer_MRuw"],"signatures":["NeurIPS.cc/2023/Conference/Submission1335/Reviewer_MRuw"],"forum":"kRdaTkaBwC","number":1,"license":"CC BY 4.0","cdate":1688363291934,"mdate":1702410782402,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1335/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"kRdaTkaBwC","id":"imj5lHpqm8","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["neural scene representations","fluid dynamics","flow reconstruction","physics-based learning"]},"_bibtex":{"value":"@inproceedings{\nyu2023inferring,\ntitle={Inferring Hybrid Neural Fluid Fields from Videos},\nauthor={Hong-Xing Yu and Yang Zheng and Yuan Gao and Yitong Deng and Bo Zhu and Jiajun Wu},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=kRdaTkaBwC}\n}"},"title":{"value":"Inferring Hybrid Neural Fluid Fields from Videos"},"paperhash":{"value":"yu|inferring_hybrid_neural_fluid_fields_from_videos"},"TLDR":{"value":"Recovering fluid density and velocity from a few multiview videos. Website: https://kovenyu.com/HyFluid/"},"abstract":{"value":"We study recovering fluid density and velocity from sparse multiview videos. Existing neural dynamic reconstruction methods predominantly rely on optical flows; therefore, they cannot accurately estimate the density and uncover the underlying velocity due to the inherent visual ambiguities of fluid velocity, as fluids are often shapeless and lack stable visual features. The challenge is further pronounced by the turbulent nature of fluid flows, which calls for properly designed fluid velocity representations. To address these challenges, we propose hybrid neural fluid fields (HyFluid), a neural approach to jointly infer fluid density and velocity fields. Specifically, to deal with visual ambiguities of fluid velocity, we introduce a set of physics-based losses that enforce inferring a physically plausible velocity field, which is divergence-free and drives the transport of density. To deal with the turbulent nature of fluid velocity, we design a hybrid neural velocity representation that includes a base neural velocity field that captures most irrotational energy and a vortex particle-based velocity that models residual turbulent velocity. We show that our method enables recovering vortical flow details. Our approach opens up possibilities for various learning and reconstruction applications centered around 3D incompressible flow, including fluid re-simulation and editing, future prediction, and neural dynamic scene composition. Project website: https://kovenyu.com/HyFluid/"},"pdf":{"value":"/pdf/26199c18482ffee744acfac9b91bd7fd0dd8f740.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Hong-Xing_Yu1","~Yang_Zheng2","~Yuan_Gao11","~Yitong_Deng1","~Bo_Zhu2","~Jiajun_Wu1"]},"authors":{"value":["Hong-Xing Yu","Yang Zheng","Yuan Gao","Yitong Deng","Bo Zhu","Jiajun Wu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Guaranteed Inner-and-Outer Meshes (GIOM), a new algorithm for computing bounding meshes for neural implicit surfaces (especially neural signed distance functions (SDFs)). They treat SDF queries as neural network verification problems, enabling the use of CROWN, a linear bound propagation method, to compute affine upper and lower bounds on the SDF output over voxel regions. As a result, the method can obtain an explicit mesh representation that (1) guarantees correctness, (2) adapts to surface complexity, and (3) integrates directly with traditional mesh-based rendering and physics engines.\n\nGIOM is demonstrated on three tasks: neural rendering, collision detection, and constructive solid geometry (CSG). It shows up to 3× faster rendering, 5× faster physics simulation, and 10× lower reconstruction error in CSG compared to baselines such as Adaptive Shells, Spelunking the Deep, and CSG-nSDF."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Do the MLPs used in the paper include positional encodings (e.g. fourier or sinusoidal embedding)? If not, supplementing these features would be important for evaluating real-case implicit surfaces.\n\n- Could the authors report the actual time and memory consumption for the construction of the meshes before downstream applications?\n\n- How does GIOM compare in terms of time and memory with classical bounding-volume hierarchies such as octrees or BVHs, which also subdivide space adaptively?\n\n- Can the authors supplement and discuss some failure cases?"},"rating":{"value":8},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- The paper introduces a novel perspective by framing geometric queries on neural implicit surfaces as neural network verification problems, which in turn leads to the subsequent proposed GIOM algorithm.\n- The paper is technically sound. It offers a clear mathematical derivation of how affine upper and lower bounds (from CROWN-style linear bound propagation) can be used to construct certified inner and outer meshes.\n- The empirical results are strong. It demonstrates three applications in neural rendering, physics-based collision detection, and constructive solid geometry (CSG), and achieves significant performance gains in all cases.\n- Finally, the paper is also well-written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While the paper reports timing metrics such as FPS for rendering and milliseconds per frame for physics simulation, it omits the time required to construct the inner and outer meshes. When the underlying network of the implicit surfaces is MLPs, which is very common, each voxel’s bound computation entails a full backpropagation through the entire network. This step can become computationally expensive if the network is deep or if many voxels require refinement, and may result in higher costs in total than direct forward queries.\n\n- The idea of representing an object as multiple surfaces is not rare in literature and can be further discussed in the related works. For example, in [1][2][3].\n\n- If I am understanding correctly, CROWN’s linear relaxation could still potentially result in loose bounds (maybe for highly nonlinear or deep architectures), and there are only theoretical guarantees for the correctness of the bounds but not the tightness. It would be valuable to include an ablation showing how the error varies with subdivision depth. \n\n- In addition, no explicit failure cases are shown where GIOM’s bounds might degrade or where the linear approximation becomes visually noticeable. A discussion of such scenarios would improve completeness without undermining the paper’s overall contributions.\n\n[1] Esposito et. al. Volumetric Surfaces: Representing Fuzzy Geometries with Layered Meshes\n\n[2] Wang et. al. A Simple Approach to Differentiable Rendering of SDFs\n\n[3] Seyb etl. al. From microfacets to participating media: A unified theory of light transport with stochastic geometry"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919097108,"tcdate":1761515024956,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6833/Reviewer_pNRh"],"signatures":["ICLR.cc/2026/Conference/Submission6833/Reviewer_pNRh"],"forum":"RI3SmeizL2","number":1,"license":"CC BY 4.0","cdate":1761515024956,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6833/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919097108,"domain":"ICLR.cc/2026/Conference","replyto":"RI3SmeizL2","id":"AmCl8aT4bS","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Neural Verification","Bounding Volume","Rendering","Simulation"]},"supplementary_material":{"value":"/attachment/e6a8faaf98289f6c8de3061095ec4dcaf94683e8.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Geometric queries on neural implicit surfaces, such as ray tracing and collision detection, present a significant challenge since they require explicit spatial reasoning over neural networks. This work addresses this challenge by connecting these geometric queries to neural network verification problems. Inspired by the state-of-the-art neural verification tools, we propose a new framework utilizing linear bound propagation-based verifiers to solve these queries in real time, enabling applications such as real-time rendering and physics simulation with soundness guarantees. Instead of naively running neural network verifiers on-the-fly, we first classify a 3D input domain into multiple regions of interest, which can then assist in subsequent verifications. We achieve this objective by constructing explicit bounding volumes and then leveraging linear bounds generated by SOTA neural network verifiers to guide the generation of \\emph{sound piecewise linear bounding meshes}. In this paper, we propose Guaranteed Inner-and-Outer Meshes (GIOM), which can serve as bounding volumes and merge seamlessly with existing explicit geometry processors to accelerate queries on neural implicits. As tight and \\emph{sound} bounding meshes, GIOM enables accelerated neural SDF queries without sacrificing quality. With GIOM, we develop accelerated neural implicit ray casting, collision detection, and constructive solid geometry methods (CSG), achieving up to a 300\\% speedup in real-time rendering, a 500\\% speedup in physics simulation, and an optimization-free neural CSG procedure. \nExperiments show that GIOM significantly outperforms existing methods in the speed-quality trade‑off."},"_bibtex":{"value":"@misc{\ngao2026guaranteed,\ntitle={Guaranteed Bounding Meshes Extraction from Neural Implicit Surfaces via Neural Network Verification},\nauthor={Ruize Gao and Jorge Chavez and Xiangru Zhong and Shenlong Wang and Huan Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=RI3SmeizL2}\n}"},"title":{"value":"Guaranteed Bounding Meshes Extraction from Neural Implicit Surfaces via Neural Network Verification"},"pdf":{"value":"/pdf/e6983f406fbaef2bc10195e579a9d94718a7852d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"gao|guaranteed_bounding_meshes_extraction_from_neural_implicit_surfaces_via_neural_network_verification"},"authorids":{"value":["~Ruize_Gao3","~Jorge_Chavez1","~Xiangru_Zhong1","~Shenlong_Wang1","~Huan_Zhang1"]},"authors":{"value":["Ruize Gao","Jorge Chavez","Xiangru Zhong","Shenlong Wang","Huan Zhang"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"venue":{"value":"ASAB 2026 Oral"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"pdf":{"value":"/pdf/7bee7056bc60a0e9e05257c05a06999bfbaef860.pdf"},"keywords":{"value":["Robotics","Reactive Control","Active Perception"]},"venueid":{"value":"IEEE.org/ICRA/2026/Workshop/ASAB"},"paperhash":{"value":"cai|physicsinformed_force_prediction_with_reactive_surface_following_for_granular_material_manipulation_on_complex_hidden_surfaces"},"authorids":{"value":["~Zijian_Cai5","~Henrik_Andreasson2","~Yasemin_Bekiroglu1","~Johannes_A._Stork1"]},"abstract":{"value":"Granular material manipulation in containers with complex geometries is challenging due to hidden rigid surfaces and uncertain interaction forces. This paper proposes a framework that combines physics-informed force prediction with reactive control. A Gaussian Process model predicts the expected interaction force between the tool and the granular material, which is used as a baseline for an adaptive controller. Experiments show that the proposed method captures the forces trend and enables stable surface following, providing a basis for contact identification and probabilistic surface mapping."},"title":{"value":"Physics-Informed Force Prediction with Reactive Surface Following for Granular Material Manipulation on Complex Hidden Surfaces"},"authors":{"value":["Zijian Cai","Henrik Andreasson","Yasemin Bekiroglu","Johannes A. Stork"]}},"tmdate":1778893141933,"pdate":1778893141101,"tcdate":1775133621643,"writers":["IEEE.org/ICRA/2026/Workshop/ASAB","IEEE.org/ICRA/2026/Workshop/ASAB/Submission9/Authors"],"signatures":["IEEE.org/ICRA/2026/Workshop/ASAB/Submission9/Authors"],"forum":"jtMCaD9X04","license":"CC BY 4.0","number":9,"cdate":1775133621643,"readers":["everyone"],"invitations":["IEEE.org/ICRA/2026/Workshop/ASAB/-/Submission","IEEE.org/ICRA/2026/Workshop/ASAB/-/Submission_Change_Before_Reviewing","IEEE.org/ICRA/2026/Workshop/ASAB/-/Submission_Release"],"mdate":1778893141933,"odate":1778893141101,"domain":"IEEE.org/ICRA/2026/Workshop/ASAB","id":"jtMCaD9X04","version":2},{"content":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["gaussian processes","variational approximations","state space gaussian processes","physics informed gaussian processes"]},"supplementary_material":{"value":"/attachment/9461298fce1715a838e9758280749bfe546e0f0a.zip"},"primary_area":{"value":"probabilistic_methods"},"abstract":{"value":"Differential equations are important mechanistic models that are integral to many scientific and engineering applications. With the abundance of available data there has been a growing interest in data-driven physics-informed models. Gaussian processes (GPs) are particularly suited to this task as they can model complex, non-linear phenomena whilst incorporating prior knowledge and quantifying uncertainty. Current approaches have found some success but are limited as they either achieve poor computational scalings or focus only on the temporal setting. This work addresses these issues by introducing a variational spatio-temporal state-space GP that handles linear and non-linear physical constraints while achieving efficient linear-in-time computation costs. We demonstrate our methods in a range of synthetic and real-world settings and outperform the current state-of-the-art in both predictive and computational performance."},"_bibtex":{"value":"@inproceedings{\nhamelijnck2024physicsinformed,\ntitle={Physics-Informed Variational State-Space Gaussian Processes},\nauthor={Oliver Hamelijnck and Arno Solin and Theodoros Damoulas},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=tCf7S75xFa}\n}"},"title":{"value":"Physics-Informed Variational State-Space Gaussian Processes"},"pdf":{"value":"/pdf/4424cd43913ea9035300e356e962178e273d1e61.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"hamelijnck|physicsinformed_variational_statespace_gaussian_processes"},"authorids":{"value":["~Oliver_Hamelijnck1","~Arno_Solin1","~Theodoros_Damoulas1"]},"authors":{"value":["Oliver Hamelijnck","Arno Solin","Theodoros Damoulas"]}},"tmdate":1736854711252,"pdate":1727288009011,"tcdate":1715723409601,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission12520/Authors"],"signatures":["NeurIPS.cc/2024/Conference/Submission12520/Authors"],"forum":"tCf7S75xFa","license":"CC BY 4.0","number":12520,"cdate":1715723409601,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/-/Submission","NeurIPS.cc/2024/Conference/-/Post_Submission","NeurIPS.cc/2024/Conference/Submission12520/-/Revision","NeurIPS.cc/2024/Conference/-/Edit","NeurIPS.cc/2024/Conference/Submission12520/-/Camera_Ready_Revision"],"mdate":1736854711252,"odate":1730873949739,"domain":"NeurIPS.cc/2024/Conference","id":"tCf7S75xFa","version":2},{"content":{"TLDR":{"value":"A photorealistic physics benchmark with ground-truth trajectories, segmentation masks, and depth maps for evaluating and training video world models."},"venue":{"value":"ICLR 2026 Workshop World Models"},"pdf":{"value":"/pdf/ac0e41ec782602eb5f77fd03d0beefb5675b0523.pdf"},"keywords":{"value":["video generation","world models","physics simulation","benchmark","rigid-body","dynamics","video prediction","synthetic data","fine-tuning"]},"venueid":{"value":"ICLR.cc/2026/Workshop/World_Models"},"paperhash":{"value":"jain|rigidbench_evaluating_rigidbody_physics_in_video_generation_models"},"authorids":{"value":["~Swarnim_Jain1","~Shangzhe_Wu2"]},"abstract":{"value":"Video generation models are increasingly deployed as world model backbones for physical AI, yet their ability to predict rigid-body dynamics remains unreliable. Existing benchmarks either lack precise ground-truth annotations (relying on VLM judgment) or use synthetic primitives that create domain gaps from natural video. We introduce RigidBench, a benchmark combining Blender physics simulation with photorealistic scenes to provide exact 3D trajectories, segmentation masks, and depth maps across ten physics tasks. Evaluating seven leading models, we find that trajectory accuracy and perceptual quality are poorly correlated: models that best predict object motion often score worst on perceptual metrics. This decoupling demonstrates that standard video quality metrics cannot assess physics understanding, motivating the need for benchmarks with precise physics annotations. We also show that fine-tuning on RigidBench data improves physics prediction, suggesting a path toward more physically grounded world models."},"_bibtex":{"value":"@inproceedings{\njain2026rigidbench,\ntitle={RigidBench: Evaluating Rigid-Body Physics in Video Generation Models},\nauthor={Swarnim Jain and Shangzhe Wu},\nbooktitle={ICLR 2026 the 2nd Workshop on World Models: Understanding, Modelling and Scaling},\nyear={2026},\nurl={https://openreview.net/forum?id=HEjWQkFrzR}\n}"},"title":{"value":"RigidBench: Evaluating Rigid-Body Physics in Video Generation Models"},"authors":{"value":["Swarnim Jain","Shangzhe Wu"]}},"tmdate":1776292157537,"pdate":1772419703668,"tcdate":1770292488788,"writers":["ICLR.cc/2026/Workshop/World_Models","ICLR.cc/2026/Workshop/World_Models/Submission112/Authors"],"signatures":["ICLR.cc/2026/Workshop/World_Models/Submission112/Authors"],"forum":"HEjWQkFrzR","license":"CC BY 4.0","number":112,"cdate":1770292488788,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/World_Models/-/Submission","ICLR.cc/2026/Workshop/World_Models/-/Post_Submission","ICLR.cc/2026/Workshop/World_Models/-/Edit","ICLR.cc/2026/Workshop/World_Models/Submission112/-/Post-Decision_Revision"],"mdate":1776292157537,"odate":1772673287121,"domain":"ICLR.cc/2026/Workshop/World_Models","id":"HEjWQkFrzR","version":2},{"content":{"summary":{"value":"This paper presents VLMFP, a framework that uses two Vision-Language Models (VLMs) to automatically convert visual planning problems into formal PDDL domain and problem files. SimVLM is fine-tuned for visual understanding and action simulation. GenVLM generates the PDDL files and refines them based on SimVLM's feedback. The method works by iteratively generating PDDL files, checking their consistency against SimVLM's simulations, and updating them to correct errors. Evaluated on grid-world domains, VLMFP successfully generates valid plans for unseen problem instances and visual appearances, removing the need for predefined domain files or constant access to the environment."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1) The framework is demonstrated on 2D grid-worlds. What are the most significant technical hurdles you foresee in scaling this approach to more complex, realistic 3D environments, such as a simulated kitchen or a robot navigation task?\n\n2) The current method relies on discrete positions (e.g., pos-1-1). How could the approach be adapted to handle continuous state spaces, which are common in real-world robotics and control problems?\n\n3) The current setup assumes a full, top-down view of the entire state. How would VLMFP need to be modified to handle partially observable environments where the agent's view is limited?\n\n4) SimVLM required a massive, fine-tuned dataset (430k samples) for the grid-world domains. Is this level of domain-specific data a fundamental requirement, or do you see a path towards a more general-purpose SimVLM that could be applied to novel domains with minimal fine-tuning?\n\n5) The grid-worlds have a fixed set of object types with clear visual representations. How would the object recognition and spatial reasoning capabilities of SimVLM need to improve to handle novel, real-world objects with diverse and often ambiguous appearances?\n\n6) The iterative refinement process involving multiple VLM calls and PDDL simulations seems computationally expensive. What is the latency of the full VLMFP pipeline, and is it feasible for any real-time decision-making scenarios?\n\n7) Could the \"simulation consistency checking\" step be made more sample-efficient? Are there smarter strategies for generating the action sequences used for comparison, rather than random sampling, to identify logical flaws faster?\n\n8) The results show a sharp performance drop on the completely unseen \"freezing\" rule (Rule 5). Does this indicate a fundamental limitation in the system's ability to reason about truly novel action dynamics, as opposed to parametric variations of known rules?\n\n9) How robust is the iterative refinement process to persistent or cascading errors? If SimVLM itself makes a systematic error in simulation, could this lead GenVLM to converge on an incorrect but self-consistent PDDL domain?\n\n10) In your view, what single advancement—whether in model architecture, training data, or the core algorithm—would be the most critical for bridging the gap from these compelling grid-world results to a useful real-world application?"},"rating":{"value":8},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"It autonomously generates both PDDL domain and problem files from visual input, eliminating the need for human-predefined domains—a key bottleneck in prior work.\n\nThe separation into a specialized simulator (SimVLM) and a general-purpose generator (GenVLM) effectively leverages the strengths of different models.\n\nThe feedback loop that compares PDDL execution with VLM simulation is a reasonable mechanism for catching and correcting errors in the generated artefacts.\n\nThe framework demonstrates robust performance across unseen problem instances, visual appearances, and even some modified rules within the tested grid-world environments only.\n\nIt provides a preliminary basis for bridging the gap between high-dimensional visual perception and precise, symbolic reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The evaluation is confined to synthetic 2D grid-worlds. It remains unproven whether the approach can scale to realistic 3D environments, cluttered scenes, or tasks requiring complex physics reasoning.\n\nThe method likely relies on the discrete, structured nature of grid cells and positional predicates (e.g., pos-1-1). Translating continuous, real-world spaces into such a formalism is a major unsolved challenge.\n\nSimVLM was fine-tuned on a massive, domain-specific dataset (430k samples). Curating a similar dataset for every new real-world domain (e.g., kitchen manipulation, navigation) would be prohibitively expensive.\n\nThe iterative process of generating, executing, and comparing action sequences for refinement is computationally intensive and may not be feasible for real-time applications.\n\nWhile it handled some rule variations, performance dropped significantly on complex, novel rules (e.g., the \"freezing\" mechanic), indicating that reasoning about entirely new dynamics is still a limitation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941562056,"tcdate":1761232156798,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21170/Reviewer_rVGA"],"signatures":["ICLR.cc/2026/Conference/Submission21170/Reviewer_rVGA"],"forum":"7tlLpQpGlx","number":1,"license":"CC BY 4.0","cdate":1761232156798,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21170/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941562056,"domain":"ICLR.cc/2026/Conference","replyto":"7tlLpQpGlx","id":"PIfbfrDkdG","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We propose VLMFP, a Dual-VLM-guided framework that can autonomously generate both PDDL problem and domain files for formal visual planning."},"keywords":{"value":["Vision Language Models","Planning","PDDL","LLM Tool Use"]},"supplementary_material":{"value":"/attachment/53ca9786ae00bc100d35522b602f2707ae0ebeae.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Vision Language Models (VLMs) show strong potential for visual planning but struggle with precise spatial and long-horizon reasoning, while Planning Domain Definition Language (PDDL) planners excel at formal long-horizon planning but cannot interpret visual inputs. Recent works combine these complementary advantages by translating visual problems into PDDL. However, while  VLMs can generate PDDL problem files satisfactorily, accurately generating PDDL domain files, which encode planning rules, remains challenging and typically requires human expertise or environment interaction.\nWe propose VLMFP, a Dual-VLM-guided framework that autonomously generates both PDDL problem and domain files for formal visual planning. VLMFP combines a SimVLM that simulates action consequences with a GenVLM that generates and iteratively refines PDDL files by aligning symbolic execution with simulated outcomes, enabling multiple levels of generalization across unseen instances, visual appearances, and game rules.\nWe evaluate VLMFP on 6 grid-world domains and demonstrate its generalization capability. On average, SimVLM achieves 87.3\\% and 86.0\\% scenario understanding and action simulation for seen and unseen appearances, respectively. With the guidance of SimVLM, VLMFP attains 70.0\\%, 54.1\\% planning success on unseen instances in seen and unseen appearances, respectively. We further demonstrate that VLMFP scales to complex long-horizon 3D planning tasks, including multi-robot collaboration and assembly scenarios with partial observability and diverse visual variations. Project page: https://sites.google.com/view/vlmfp."},"_bibtex":{"value":"@inproceedings{\nhao2026simulation,\ntitle={Simulation to Rules: A Dual-{VLM} Framework for Formal Visual Planning},\nauthor={Yilun Hao and Yongchao Chen and Chuchu Fan and Yang Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=7tlLpQpGlx}\n}"},"title":{"value":"Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning"},"pdf":{"value":"/pdf/dfb4c968a6b6462df35ea16363eeefe98b0182a0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"hao|simulation_to_rules_a_dualvlm_framework_for_formal_visual_planning"},"authorids":{"value":["~Yilun_Hao1","~Yongchao_Chen1","~Chuchu_Fan2","~Yang_Zhang3"]},"authors":{"value":["Yilun Hao","Yongchao Chen","Chuchu Fan","Yang Zhang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2022"},"pdf":{"value":"http://arxiv.org/pdf/2206.04873v1"},"venueid":{"value":"dblp.org/journals/CORR/2022"},"paperhash":{"value":"chen|imitation_learning_via_differentiable_physics"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Siwei_Chen_0003:","~Xiao_Ma2","~Zhongwen_Xu1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2206.04873"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2206-04873,\n  publtype={informal},\n  author={Siwei Chen and Xiao Ma and Zhongwen Xu},\n  title={Imitation Learning via Differentiable Physics},\n  year={2022},\n  cdate={1640995200000},\n  journal={CoRR},\n  volume={abs/2206.04873},\n  url={https://doi.org/10.48550/arXiv.2206.04873}\n}\n"},"abstract":{"value":"Existing imitation learning (IL) methods such as inverse reinforcement learning (IRL) usually have a double-loop training process, alternating between learning a reward function and a policy and tend to suffer long training time and high variance. In this work, we identify the benefits of differentiable physics simulators and propose a new IL method, i.e., Imitation Learning via Differentiable Physics (ILD), which gets rid of the double-loop design and achieves significant improvements in final performance, convergence speed, and stability. The proposed ILD incorporates the differentiable physics simulator as a physics prior into its computational graph for policy learning. It unrolls the dynamics by sampling actions from a parameterized policy, simply minimizing the distance between the expert trajectory and the agent trajectory, and back-propagating the gradient into the policy via temporal physics operators. With the physics prior, ILD policies can not only be transferable to unseen environment specifications but also yield higher final performance on a variety of tasks. In addition, ILD naturally forms a single-loop structure, which significantly improves the stability and training speed. To simplify the complex optimization landscape induced by temporal physics operations, ILD dynamically selects the learning objectives for each state during optimization. In our experiments, we show that ILD outperforms state-of-the-art methods in a variety of continuous control tasks with Brax, requiring only one expert demonstration. In addition, ILD can be applied to challenging deformable object manipulation tasks and can be generalized to unseen configurations."},"title":{"value":"Imitation Learning via Differentiable Physics"},"authors":{"value":["Siwei Chen","Xiao Ma","Zhongwen Xu"]}},"tmdate":1747299150214,"pdate":1640995200000,"tcdate":1743245844771,"writers":["~"],"signatures":["~Xiao_Ma2"],"forum":"s89onTU2bP","license":"CC BY-SA 4.0","number":373982,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747299150214,"domain":"DBLP.org","id":"s89onTU2bP","version":2},{"content":{"comment":{"value":"Thank you for your response. We reiterate that our intent was *neither* to claim that the building blocks for interacting with the Multiphysics API interface is the *only* possible realization of these blocks *nor* to introduce an agentic algorithm that is the best possible algorithm for this task, but instead to design components to faciliate LLM-API interaction for a baseline agent on this benchmark. \n## On LLM-Agent Research\nWith regard to your concern on connections with existing agentic work, we believe we have discussed generalist agentic frameworks and algorithms in our draft [Quotes 1, 2]. These include some of the papers at the link you provide.  \n\n## On using other software\n\"*As you mentioned, the problems in your benchmark are solvable using COMSOL alone.*\" \nIt appears there might be a slight misunderstanding here. To reiterate Point 2 of our overall reply, the FEABench Gold problems do **not** preclude using other software. Since the FEABench Gold problems are physics / numerical problems, the Model Specifications field describes a physics / mathematical problem (example on Page 15) and the answer is **not** intended to be exclusively derivable via COMSOL Multiphysics®. \n\n## On real-world physics reasoning ability\nOur benchmark seeks to assess how well LLMs can represent complex objects and model interactions with physical phenomena to solve problems that require numerical (FE) analysis, and cannot typically be solved analytically alone -- a scenario encountered in diverse scientific and engineering contexts in the real world.\n\nWe chose COMSOL Multiphysics® as a medium for this exploration to maximize coverage across physical phenomena, and use a single tool to solve problems involving electromagnetism, structural mechanics and heat transfer among others. Solving the problems in our benchmark requires making correct and consistent physics / engineering decisions and correctly generating code to steer the API to encode these choices. While we did not seek to exhaustively explore the heterogeneous landscape of other simulation tools and packages used across myriad subfields in physics and engineering, we respectfully contend that many of the **abilities** required to excel at using other tools (such as Ansys® or physics-subfield specific packages) are similar to those required in this context: making correct and mutually consistent physics decisions and learning to encode those decisions by calling tools or writing code to steer domain-specific software. Nonetheless, we have **updated our title**.  \n*On combining FEA with analytical methods:* We included an element of this direction in our agentic setup. The VerifierLLM component of our Evaluator sets an analytical estimate for the target description at the start of the agent experiment. Whenever the executability of the LLM's solution crosses 0.90, this LLM additionally compares the numerically-computed answer with its a priori analytical estimate, when providing its feedback."},"title":{"value":"Reply (1/2)"}},"tmdate":1732586037460,"tcdate":1732586037460,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12171/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission12171/Authors"],"forum":"hDkLpu1E64","number":11,"license":"CC BY 4.0","cdate":1732586037460,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12171/-/Official_Comment"],"mdate":1732586037460,"domain":"ICLR.cc/2025/Conference","replyto":"2LMmoSYvDT","id":"M0FhalDNVk","forumContent":{"TLDR":{"value":"How well can LLMs leverage FEA software to simulate and solve problems that require numerical analysis?"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["numerical analysis","finite element","benchmark","agents"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building precise simulations of the real world and invoking numerical solvers to answer quantitative problems is an essential requirement in engineering and science. We present FEABench, a benchmark to evaluate the ability of large language models (LLMs) and LLM agents to simulate and solve physics, mathematics and engineering problems using finite element analysis (FEA). We introduce a multipronged evaluation scheme to investigate the ability of LLMs to solve these problems by reasoning over natural language problem descriptions and operating COMSOL Multiphysics$^\\textregistered$, an FEA software, to compute the answers. In addition to testing state-of-the art-LLMs, we further design a language model agent equipped with the ability to interact with the software through its Application Programming Interface (API), examine its outputs and use tools to improve its solutions over multiple iterations. Our best performing strategy generates executable API calls 88\\% of the time. However, this benchmark still proves to be challenging enough that the LLMs and agents we tested were not able to completely and correctly solve any problem. LLMs that can successfully interact with and operate FEA software to solve problems such as those in our benchmark would significantly push the frontiers of their utility. Acquiring this capability would augment LLMs' reasoning skills with the precision of numerical solvers and advance the development of autonomous systems that can tackle complex problems in the real world."},"_bibtex":{"value":"@misc{\nmudur2025feabench,\ntitle={{FEAB}ench: Evaluating Language Models on Real World Physics Reasoning Ability},\nauthor={Nayantara Mudur and Hao Cui and Subhashini Venugopalan and Paul Raccuglia and Michael Brenner and Peter Christian Norgaard},\nyear={2025},\nurl={https://openreview.net/forum?id=hDkLpu1E64}\n}"},"title":{"value":"FEABench: Evaluating Language Models on Real World Physics Reasoning Ability"},"pdf":{"value":"/pdf/3e64111fb86b7cbb5ef6469de0f077b416722ed3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"mudur|feabench_evaluating_language_models_on_real_world_physics_reasoning_ability"},"authorids":{"value":["~Nayantara_Mudur1","~Hao_Cui3","~Subhashini_Venugopalan2","~Paul_Raccuglia1","~Michael_Brenner1","~Peter_Christian_Norgaard1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Nayantara Mudur","Hao Cui","Subhashini Venugopalan","Paul Raccuglia","Michael Brenner","Peter Christian Norgaard"]}},"version":2},{"content":{"summary":{"value":"Here is a concise review based on my analysis.\nThe paper proposes AdaMeshNet, a GNN framework for fluid dynamics that aims to mitigate over-squashing by introducing an \"adaptive rewiring\" mechanism. This mechanism delays the addition of new edges to specific message-passing layers based on a \"rewiring delay score.\"\n\nWhile the paper addresses a relevant problem and shows slight empirical improvements, I must recommend **Reject**.\n\nThe work suffers from three critical flaws:\n1.  The core physics-informed rewiring strategy (identifying bottlenecks with Ollivier-Ricci Curvature and connecting them to nodes with maximal velocity difference) is not novel and is adopted directly from prior art, PIORF. The only novelty is an algorithmic tweak to *delay* this connection.\n2.  The entire justification for this delay is that instantaneous interactions are \"physically wrong\". This is fundamentally incorrect for the paper's primary benchmark, CylinderFlow which models *incompressible* flow. In this regime, information (pressure) propagates *instantaneously* by definition.\n3.  The core mechanism, the \"rewiring delay score\" (Eq. 8), is dimensionally incoherent. It compares a unitless GNN layer index $l$ to a score $s_{delay}$ that, by its formulation (unitless distance divided by velocity), must have physical units ($Time/Length$). This comparison is physically and mathematically meaningless.\n\nThe paper is an incremental work, while it's intuition is physically unjustified, and mathematically flawed."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.  How do the authors justify the core premise of \"gradual propagation\" when the CylinderFlow dataset models incompressible flow, where pressure propagation is instantaneous?\n2.  Can the authors provide a dimensional analysis for Equation 8? How can the unitless layer index $l$ be mathematically compared to $s_{delay}$, which (based on its components) has units of $Time/Length$?\n3.  The calculation of $W_1$ for ORC on triangular meshes is non-trivial and requires solving a linear program. How is this step practically implemented? Is a numerical optimizer used, and if so, why is this extremely expensive step placed *inside* the main training loop (Algorithm 1) rather than as a one-time preprocessing step? Can you list the method you used? Can you also list the timing?"},"rating":{"value":2},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1.  The paper tackles the important and challenging problem of over-squashing in mesh-based GNNs for physics simulations.\n2.  The empirical results in Tables 1 and 2 consistently show that AdaMeshNet achieves a lower RMSE than the baselines, including PIORF, across both datasets.\n3.  The ablation study effectively demonstrates that both the distance and velocity components of the heuristic score (Eq. 8) are necessary to achieve the reported performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  The paper presents its physics-informed selection strategy as novel. However, the entire logic—(1) identify topological bottlenecks using node-level ORC, and select a rewiring target $v_{i^*}$ that maximizes the velocity difference $||v_i - v_j||$—is identical to the strategy proposed in the PIORF paper[1]; The *only* novel contribution of this work is the calculation of $s_{delay}$ to decide *when* to add this pre-selected edge, which is an incremental modification, not a \"novel framework\".\n\n2. The paper's central premise is that modeling the \"gradual propagation of physical interactions\" is more realistic than assuming \"instantaneous interactions\". This justification is in direct conflict with the physics of the paper's main benchmark. The CylinderFlow dataset models incompressible flow. In an incompressible fluid, the speed of sound is infinite, and thus information (propagated via the pressure field) is, by definition, \"sensed instantaneously at all other points in the fluid\". The \"flaw\" this paper claims to fix is, in fact, the *correct* physical assumption for this system. The authors appear to be conflating the *advection* of mass (which takes time) with the *propagation* of information (which is instantaneous in this model).\n\n3.  The paper's core mechanism, Eq. 8, does not make sense dimensionally speaking, it's using some dimensionless quantity to divide realistic velocity. Also, the usage of geodesic minimal (edge-count) distance does not make sense too, as for FVM (finite volume method) simulation, the velocity is never defined on the edge.\n\n4.  The paper provides a standard definition for ORC, which involves computing the $W_1$ distance. This is known to be computationally expensive, requiring the solution of a linear programming problem for each edge. The only exception is for bi-partite graph or graph with girth >=5 [2]. But here we have triangular mesh (girth=3). PIORF performs this expensive calculation *once* as a preprocessing step. In contrast, AdaMeshNet's Algorithm 1 places the *entire* preprocessing step—including the ORC calculation and the $O(N^2)$ search for optimal partners—*inside* the main training loop, to be run *every epoch*. This suggests a prohibitive increase in computational cost, however the paper does not report implementation details or timing.\n\n[1] https://openreview.net/forum?id=qkBBHixPow\n[2] https://sites.stat.columbia.edu/sumitm/Ricci.pdf"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942678905,"tcdate":1761992765194,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23488/Reviewer_RkBo"],"signatures":["ICLR.cc/2026/Conference/Submission23488/Reviewer_RkBo"],"forum":"Tq2n0lFSOd","number":3,"license":"CC BY 4.0","cdate":1761992765194,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23488/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942678905,"domain":"ICLR.cc/2026/Conference","replyto":"Tq2n0lFSOd","id":"QBmyIxWnD2","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Mesh Simulation","Physics Simulation","Fluid Dynamics"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Mesh-based simulation using Graph Neural Networks (GNNs) has been recognized as a promising approach for modeling fluid dynamics. However, the mesh refinement techniques which allocate finer resolution to regions with steep gradients can induce the over-squashing problem in mesh-based GNNs, which prevents the capture of long-range physical interactions. Conventional graph rewiring methods attempt to alleviate this issue by adding new edges, but they typically complete all rewiring operations before applying them to the GNN. These approaches are physically unrealistic, as they assume instantaneous interactions between distant nodes and disregard the distance information between particles. To address these limitations, we propose a novel framework, called Adaptive Graph Rewiring in Mesh-Based Graph Neural Networks (AdaMeshNet), that introduces an adaptive rewiring process into the message-passing procedure to model the gradual propagation of physical interactions. Our method computes a rewiring delay score for bottleneck nodes in the mesh graph, based on the shortest-path distance and the velocity difference. Using this score, it dynamically selects the message-passing layer at which new edges are rewired, which can lead to adaptive rewiring in a mesh graph. Extensive experiments on mesh-based fluid simulations demonstrate that AdaMeshNet outperforms conventional rewiring methods, effectively modeling the sequential nature of physical interactions and enabling more accurate predictions."},"_bibtex":{"value":"@misc{\nseo2025adaptive,\ntitle={Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based {GNN}s for Fluid Dynamics Simulations},\nauthor={Sangwoo Seo and Hyunsung Kim and Jiwan Kim and Chanyoung Park},\nyear={2025},\nurl={https://openreview.net/forum?id=Tq2n0lFSOd}\n}"},"title":{"value":"Adaptive Graph Rewiring to Mitigate Over-Squashing in Mesh-Based GNNs for Fluid Dynamics Simulations"},"pdf":{"value":"/pdf/3449481a2cf58a46f69eeafbce6305c45a9cd9c3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"seo|adaptive_graph_rewiring_to_mitigate_oversquashing_in_meshbased_gnns_for_fluid_dynamics_simulations"},"authorids":{"value":["~Sangwoo_Seo1","~Hyunsung_Kim2","~Jiwan_Kim2","~Chanyoung_Park1"]},"authors":{"value":["Sangwoo Seo","Hyunsung Kim","Jiwan Kim","Chanyoung Park"]}},"version":2},{"content":{"summary":{"value":"In this paper a framework for efficiently scaling neural operators is introduced under the name of Universal Physics Transformers (UPTs). It is a novel paradigm that scales neural operators across diverse spatio-temporal problems without relying on specific grid or particle-based structures. Leveraging transformer architectures, UPTs propagate dynamics within a compressed latent space, enabling efficient and flexible simulations."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1/  The paper briefly mentions feature modulation for model conditioning but lacks detailed discussion on handling complex boundary conditions. Providing specific examples or additional experiments that illustrate UPTs' effectiveness in managing complex boundary conditions would strengthen the paper’s claims. Addressing this can clarify the practical applicability of UPTs in real-world scenarios where boundary conditions play a crucial role.\n\nHow does the UPT framework handle complex boundary conditions compared to traditional particle and grid-based methods?\n\n2/  While UPTs are presented as efficient and scalable, the paper does not delve into the specifics of their numerical stability and accuracy compared to traditional methods. Discussing strategies to maintain stability and accuracy over extended simulations and providing relevant experimental evidence would address potential concerns about the reliability of UPTs in long-term applications.\n\nWhat measures are in place to ensure the numerical stability and accuracy of UPTs over long simulation periods?\n\n3/ The paper demonstrates scalability and generalization within certain limits but does not extensively explore these aspects in more demanding scenarios. Providing additional experimental results or theoretical insights on UPTs’ performance in high-resolution and large-scale environments, as well as their ability to generalize to entirely new conditions, would significantly bolster the paper’s contribution and practical relevance.\n\n What are the scalability and generalization capabilities of UPTs when applied to extremely high-resolution grids, large particle systems, and unseen physical scenarios?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1/ Originality:\n\nThe paper presents an original contribution to the field of neural operators by introducing Universal Physics Transformers (UPTs). This novel paradigm removes the traditional reliance on grid or particle-based latent structures, enabling greater flexibility and scalability across various simulation types. The innovative use of transformer architectures to propagate dynamics within a compressed latent space is a quite original combination of existing ideas applied to a new domain.\n\n2/ Quality:\n\nThe quality of the overall contribution is good. The authors address a critical challenge in scaling neural operators for complex simulations and provide a robust framework with practical applications. The UPT framework is well-formulated, and the encoding and decoding schemes are designed to ensure efficient simulation rollouts. The experiments are comprehensive, covering multiple types of simulations, and the results demonstrate the superiority of UPTs in terms of performance and scalability. \nMoreover, the paper's methodology is sound, and the technical claims are well-supported by evidence from the experiments.\n\n3/ Clarity:\n\nThe paper is well-written and clearly presented. The authors provide a thorough background. \n\n4/ Significance:\n\nThe significance of the paper lies in its potential impact on the broader scientific and engineering communities. By providing a unified and scalable framework for neural operators, UPTs can be applied to a wide range of spatio-temporal problems, including those in fluid dynamics. The ability to efficiently handle large-scale simulations can lead to significant advancements in these fields, offering valuable insights and solutions to complex physical phenomena. \nMoreover, the experimental design is robust, with appropriate datasets and baselines used for comparison. The methodology is clearly described, allowing for reproducibility of the results. The use of inverse encoding and decoding techniques to enable latent rollouts is particularly noteworthy, as it demonstrates a deep understanding of the underlying principles and challenges in neural operator learning.\n\n5/ Latente space and scalability:\n\nThe Universal Physics Transformers (UPTs) framework is efficient since it allows to get compact yet expressive latent space representation, capturing the essential dynamics of physical systems while significantly reducing memory usage and computational overhead. This compactness, combined with efficient encoding and decoding schemes, ensures that the transformation to and from the latent space is both efficient and accurate. By unifying the encoding of various grids and particles, UPTs simplify the modeling process and improve generalization across different spatio-temporal problems.\n\nThe empirical validation provided in the paper, through steady-state and transient flow simulations, demonstrates that UPTs achieve lower mean-squared error (MSE) and faster computation times compared to other models, highlighting the effectiveness of the latent space representation.\n\nIn terms of scalability, UPTs are designed to handle large-scale simulations effectively. The fixed-size latent space, regardless of input size, allows UPTs to manage large inputs without a proportional increase in computational cost. This scalability is evident in the framework's ability to maintain performance and efficiency with large meshes and numerous particles. The efficient latent space rollouts further enhance scalability by enabling quick updates and predictions, making UPTs suitable for time-dependent simulations. Additionally, the transformer architecture provides a robust foundation for scalability, ensuring that UPTs can handle complex spatio-temporal simulations efficiently. \n\n\n6/ Lagrangian dynamic modeling:\n\nThe Universal Physics Transformers (UPTs) framework presents several notable strengths when viewed from the perspective of Lagrangian dynamic modeling. One of the primary advantages is the framework's ability to model particle-based simulations effectively without relying on traditional particle-structures. In Lagrangian simulations, particles move with the local deformation of the continuum, and modeling these dynamics accurately requires handling a large number of particles and their interactions.\n\nUPTs use a unified latent space to encode particle information, enabling the framework to handle varying numbers of particles flexibly. This flexibility is particularly beneficial for Lagrangian methods, such as Smoothed Particle Hydrodynamics (SPH), where particle counts can vary significantly based on the simulation's complexity. By compressing this information into a fixed-size latent space, UPTs manage to reduce computational overhead while maintaining the ability to capture intricate particle interactions.\n\nThe efficient encoding and decoding processes in UPTs ensure that particle dynamics are propagated accurately and swiftly within the latent space. This efficiency is crucial for large-scale Lagrangian simulations, where the computational cost can be prohibitive with traditional methods. The ability to perform latent rollouts enables UPTs to predict future states quickly, making the framework suitable for real-time or near-real-time applications in Lagrangian dynamic modeling.\n\nThe paper provides strong empirical evidence of UPTs' effectiveness in Lagrangian dynamic modeling through experiments on datasets such as the Taylor-Green vortex in three dimensions (TGV3D). The results demonstrate that UPTs can effectively learn the underlying field dynamics and predict particle velocities with lower error compared to traditional Graph Neural Networks (GNNs) and other baselines. This empirical validation underscores the practical applicability of UPTs in complex Lagrangian simulations\n\nUPTs exhibit strong generalization capabilities, which are critical for Lagrangian dynamic modeling. The ability to query the latent representation at any point in space-time allows UPTs to adapt to different particle distributions and simulation conditions without extensive retraining. This robustness ensures that UPTs can handle a wide range of Lagrangian problems, from small-scale particle systems to large-scale simulations involving a representative number of particles.\n\nThe use of transformers to manage the latent space representation in UPTs is a significant technical innovation. Transformers are known for their efficiency in handling large datasets and sequences, and their application in UPTs leverages this strength to manage the complexities of Lagrangian dynamics. This choice of architecture allows UPTs to model particle interactions accurately while maintaining computational efficiency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1/ Clarity and detail in methodology:\n\nThe methodology, particularly the detailed implementation of the encoding and decoding schemes, could be more clearly articulated. Some parts of the algorithms are difficult to follow. Providing additional diagrams, detailed explanations, or pseudo-code would help clarify these complex processes. Ensuring that each step of the process is well-explained and easily understandable would make the paper more accessible and replicable.\n\n2/ Scalability with extremely large datasets:\n\nWhile the paper demonstrates UPTs' scalability, there is limited discussion on how the framework performs with extremely large datasets or in distributed computing environments. Given the increasing size of datasets in fields like high-resolution climate modeling and genomics, a discussion or preliminary results on the scalability of UPTs in such contexts would be valuable. This could include potential challenges, proposed solutions, and the impact on computational resources.\n\n3/ Insufficient analysis of generalization capabilities:\n\nWhile UPTs are shown to generalize across different simulation types, the paper lacks a deep analysis of this capability. Neural operator networks are often praised for their ability to generalize across different boundary conditions and initial states. A more thorough examination of UPTs' generalization performance, especially in unseen or out-of-distribution scenarios, would strengthen the claims. Detailed experiments and discussions on how UPTs perform under varied conditions would be valuable.\n\n4/ Potential overfitting concerns:\n\nGiven the complexity and high capacity of transformer-based models, there is a risk of overfitting, especially on smaller datasets. The paper briefly mentions overfitting issues in some experiments but does not provide a detailed strategy for mitigating this. More discussion on regularization techniques, data augmentation strategies, or how UPTs handle overfitting would be beneficial. Insights into how the model can be generalized better and made more robust against overfitting are crucial for practical applications.\n\n5/ Analysis from particle and grid based methods:\n\nThe Universal Physics Transformers (UPTs) framework offers a novel approach, but there are specific areas where it could be improved, particularly when compared to traditional particle and grid-based methods.\n\nA/ Lack of detailed comparison with established methods:\n\nThe paper does not provide a thorough comparison with well-established particle-based methods (such as Smoothed Particle Hydrodynamics, SPH) and grid-based methods (such as Finite Volume Methods, FVM). Including detailed performance metrics, such as accuracy, computational cost, and scalability, would help highlight UPTs' advantages and areas needing improvement. Specifically, comparative studies on benchmark problems typically addressed by SPH (except figure 12) and FVM would provide more context.\n\nB/ Handling of boundary conditions:\n\nParticle and grid-based methods have well-established techniques for handling complex boundary conditions, which are often crucial in physical simulations. The paper does not sufficiently address how UPTs manage complex boundary conditions compared to these traditional methods. A more detailed discussion or experiments showcasing UPTs' effectiveness in dealing with various boundary conditions would strengthen the paper's claims.\n\nC/ Adaptation to high-resolution grids and large particle systems:\n\nWhile UPTs are shown to be scalable, the paper lacks detailed insights into their performance on extremely high-resolution grids or very large particle systems. Traditional methods often excel in these areas due to their specialized structures and optimizations. Providing more evidence on how UPTs handle such scenarios, including any potential bottlenecks and solutions, would be beneficial.\n\nD/ Numerical stability and accuracy:\n\nGrid-based methods, such as FVM, are known for their numerical stability and accuracy, especially in simulating fluid dynamics. The paper does not provide an in-depth analysis of UPTs' numerical stability and accuracy compared to these methods. Detailed experiments and discussions on how UPTs ensure stability and accuracy over long simulation times would enhance the paper’s credibility.\n\nE/ Computational efficiency in complex geometries:\n\nParticle and grid-based methods have specific strategies to efficiently handle complex geometries, such as adaptive meshing in grid-based methods or kernel adjustments in particle methods. The paper does not clearly explain how UPTs manage complex geometries and whether they can maintain computational efficiency in such scenarios. More detailed experiments or case studies involving complex geometrical domains would be informative.\n\nF/ Interoperability with existing simulation tools:\n\nTraditional methods are often integrated into comprehensive simulation tools (e.g., OpenFOAM for grid-based methods or LAMMPS for particle-based methods). The paper does not discuss how UPTs can be integrated or used alongside these existing tools, which is crucial for practical adoption. Providing insights into interoperability and how UPTs can complement or enhance traditional methods within established simulation frameworks would be useful.\n\nG/  Handling of multiscale phenomena:\n\nParticle and grid-based methods have developed sophisticated techniques to handle multiscale phenomena, such as adaptive mesh refinement (AMR) in grid-based methods. The paper lacks a discussion on how UPTs address multiscale phenomena, which are common in many physical simulations. Including experiments or theoretical discussions on UPTs' capabilities in multiscale modeling would be beneficial.\n\n\n6/ Model conditioning standpoint:\n\nModel conditioning is crucial for ensuring that neural networks adapt accurately to varying inputs and simulation conditions. The Universal Physics Transformers (UPTs) framework has room for improvement in this aspect. Here are specific areas where the paper could enhance its discussion and implementation of model conditioning.\n\nA/ Insufficient detail on conditioning mechanisms:\n\nThe paper briefly mentions the use of feature modulation (e.g., DiT modulation) for conditioning UPTs to various inputs, such as the current timestep and boundary conditions. However, it lacks a detailed explanation of how these conditioning mechanisms are implemented and their impact on model performance. Providing more comprehensive descriptions and theoretical justifications for the chosen conditioning methods would help readers understand their effectiveness and potential limitations.\n\nB/ Limited analysis of conditioning performance:\n\nWhile the paper demonstrates that UPTs can handle different flow regimes and domains, it does not provide a thorough analysis of how well the conditioning mechanisms work across a broader range of scenarios. Including detailed experiments that specifically evaluate the performance of UPTs under different conditioning inputs, such as varying boundary conditions, initial states, and external forces, would strengthen the paper's claims.\n\nC/ Scalability of conditioning methods:\n\nThe scalability of the conditioning mechanisms is not thoroughly discussed. As simulations grow in complexity, the effectiveness of conditioning methods can degrade if not properly scaled. The paper should address how the conditioning mechanisms scale with increasing input dimensions and simulation complexities, providing insights into any potential bottlenecks and how they are mitigated.\n\nD/ Generalization to unseen conditions:\n\nOne of the strengths of neural operator networks is their ability to generalize to unseen conditions. The paper does not sufficiently explore how well UPTs generalize to entirely new boundary conditions or physical scenarios that were not present in the training data. Including experiments that test the generalization capabilities of UPTs to unseen conditions would provide valuable insights into the robustness of the conditioning methods.\n\nE/  Comparison with other conditioning approaches:\n\nThe paper does not compare its conditioning mechanisms with other advanced conditioning techniques used in neural operator networks or related fields. A comparative analysis of different conditioning approaches, such as those used in Fourier Neural Operators and DeepONets, would highlight the strengths and weaknesses of the methods employed in UPTs.\n\nF/  Impact of conditioning on training stability:\n\nConditioning mechanisms can significantly impact the training stability of neural networks. The paper does not discuss how the chosen conditioning methods affect the stability of UPT training, especially in the presence of complex and noisy data. Analyzing and addressing potential stability issues arising from conditioning would be beneficial.\n\nG/ Real-world application scenarios:\n\nWhile the paper presents conditioning in the context of fluid dynamics simulations, it does not discuss its applicability to other real-world scenarios that require complex conditioning. Exploring and providing evidence of UPTs' effectiveness in diverse applications, such as climate modeling or structural analysis, where conditioning to various external factors is crucial, would enhance the practical relevance of the framework."},"limitations":{"value":"Yes."}},"nonreaders":[],"tmdate":1730879082773,"tcdate":1721732870653,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission6487/Reviewer_5AUz"],"signatures":["NeurIPS.cc/2024/Conference/Submission6487/Reviewer_5AUz"],"forum":"oUXiNX5KRm","number":5,"license":"CC BY 4.0","cdate":1721732870653,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission6487/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879082773,"domain":"NeurIPS.cc/2024/Conference","replyto":"oUXiNX5KRm","id":"1nGuGhHL5W","forumContent":{"TLDR":{"value":"We introduce Universal Physics Transformers, an efficiently scalable neural operator framework to model a wide range of spatio-temporal problems – for Lagrangian and Eulerian discretization schemes."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["neural operator","computational fluid dynamics","Lagrangian simulations","transformers","latent space modeling"]},"supplementary_material":{"value":"/attachment/ffbee57cf57b024cc7564d74926efc865fcb2da6.zip"},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Neural operators, serving as physics surrogate models, have recently gained increased interest. With ever increasing problem complexity, the natural question arises: what is an efficient way to scale neural operators to larger and more complex simulations - most importantly by taking into account different types of simulation datasets. This is of special interest since, akin to their numerical counterparts, different techniques are used across applications, even if the underlying dynamics of the systems are similar. Whereas the flexibility of transformers has enabled unified architectures across domains, neural operators mostly follow a problem specific design, where GNNs are commonly used for Lagrangian simulations and grid-based models predominate Eulerian simulations. \n\nWe introduce Universal Physics Transformers (UPTs), an efficient and unified learning paradigm for a wide range of spatio-temporal problems. UPTs operate without grid- or particle-based latent structures, enabling flexibility and scalability across meshes and particles. UPTs efficiently propagate dynamics in the latent space, emphasized by inverse encoding and decoding techniques. Finally, UPTs allow for queries of the latent space representation at any point in space-time. We demonstrate diverse applicability and efficacy of UPTs in mesh-based fluid simulations, and steady-state Reynolds averaged Navier-Stokes simulations, and Lagrangian-based dynamics."},"_bibtex":{"value":"@inproceedings{\nalkin2024universal,\ntitle={Universal Physics Transformers: A Framework For Efficiently Scaling Neural Operators},\nauthor={Benedikt Alkin and Andreas F{\\\"u}rst and Simon Lucas Schmid and Lukas Gruber and Markus Holzleitner and Johannes Brandstetter},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=oUXiNX5KRm}\n}"},"title":{"value":"Universal Physics Transformers: A Framework For Efficiently Scaling Neural Operators"},"pdf":{"value":"/pdf/4ffed9700454b3329a91e1a3f4304eef2c75e090.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"alkin|universal_physics_transformers_a_framework_for_efficiently_scaling_neural_operators"},"authorids":{"value":["~Benedikt_Alkin1","~Andreas_Fürst1","~Simon_Lucas_Schmid1","~Lukas_Gruber2","~Markus_Holzleitner1","~Johannes_Brandstetter1"]},"authors":{"value":["Benedikt Alkin","Andreas Fürst","Simon Lucas Schmid","Lukas Gruber","Markus Holzleitner","Johannes Brandstetter"]}},"version":2},{"content":{"summary":{"value":"The paper introduces APILaNet, a neural framework for forecasting physical systems when only one sensor is available. It learns a hidden space–time field where conservation laws are enforced in a weak, measure-weighted form, making the model robust to sparse spatial data. A dual LSTM captures slow and fast flow components, a monotone neural mapping ensures physically consistent observations, and an adaptive scheduler adjusts physics constraints based on signal difficulty. Applied to hydrological forecasting, APILaNet consistently outperforms leading deep sequence models in accuracy and stability, especially during extreme events."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"How sensitive is APILaNet’s performance to the specific choice of latent spatial discretization and the learned weighting measure?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper’s strength lies in tackling the single-sensor forecasting problem, which is a genuinely tricky and underexplored setup. The idea of using a weak-form physics constraint with a learned space–time weighting is a thoughtful technical twist: it avoids the brittle collocation sampling that usually plagues PINNs. The monotone mapping between discharge and stage is a sensible touch: it mirrors how water levels in rivers actually rise with flow instead of letting the network invent unrealistic relationships. The adaptive physics scheduling, while somewhat heuristic, shows awareness of practical training instability in physics-informed models and attempts to handle it dynamically. These pieces together make the framework conceptually interesting for sparse, physically constrained domains."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The overall novelty appears incremental, and the architecture feels somewhat overengineered relative to its contribution. The combination of multiple components (dual LSTMs, adaptive schedulers, weak-form latent mesh, and monotone mapping) adds complexity without a clear demonstration of which elements are essential or theoretically justified. The problem setup: \"forecasting from a single downstream sensor with exogenous rainfall\", is well-motivated but rather narrow, which may limit its broader relevance to general physics-informed or sequence modeling audiences. The theoretical component, particularly the “learned measure” weak form, seems to build on established weak PINN concepts with limited conceptual advancement. Empirical results show consistent but modest improvements, and could be strengthened by more robust statistical analysis and tests beyond hydrology. The paper would benefit from simplifying the model to highlight the key idea more clearly and from expanding the scope or validation to illustrate broader applicability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762934026277,"tcdate":1762807755030,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20627/Reviewer_vnKz"],"signatures":["ICLR.cc/2026/Conference/Submission20627/Reviewer_vnKz"],"forum":"VScnURO2g1","number":4,"license":"CC BY 4.0","cdate":1762807755030,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20627/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762934026277,"domain":"ICLR.cc/2026/Conference","replyto":"VScnURO2g1","id":"KDdst2g8VX","forumContent":{"TLDR":{"value":"A physics-informed latent network with adaptive weighting and a learned weak-form measure enables robust single-sensor forecasting under sparse sensing, outperforming SOTA models"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed learning","conservation laws","adaptive loss weighting","latent field","monotone neural mapping","time-series forecasting"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Forecasting conservation-governed dynamics is often constrained by sparse sensing: in practice, we may have only a single boundary sensor and noisy exogenous variables. In this work we design an Adaptive Physics-Informed Latent Network (APILaNet) that learns a latent field and enforces 1D-conservation of physics law in the weak form using a learned, normalized space--time measure. Normalization makes physics enforcement insensitive to quadrature resolution and concentrates it on transient violations. A monotone, Lipschitz measurement layer maps latent variables to observed targets, improving identifiability from a single sensor. An adaptive, bounded scheduler scales the physics and smoothness loss terms with meaningful representations, emphasizing conservation of physics laws during events while preserving training stability. Learning a space-time measure for weak-form enforcement, combined with a monotone mapping and adaptive scheduling, enables accurate, data-efficient single-sensor forecasting in physics-governed systems. We evaluate APILaNet through a synthetic and hydrological case study, APILaNet outperforms strong sequence baselines and reduces MSE during extreme events, while improving Nash--Sutcliffe efficiency. Code will be released upon acceptance."},"_bibtex":{"value":"@misc{\nkucia2026apilanet,\ntitle={{APIL}aNet: Adaptive Physics-Informed Latent Network for Single-Sensor Forecasting},\nauthor={Adrian Kucia and Edward Rollason and Wai Lok Woo},\nyear={2026},\nurl={https://openreview.net/forum?id=VScnURO2g1}\n}"},"title":{"value":"APILaNet: Adaptive Physics-Informed Latent Network for Single-Sensor Forecasting"},"pdf":{"value":"/pdf/8cb8df5af300f217d6fe04bca6fa0d4678477421.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kucia|apilanet_adaptive_physicsinformed_latent_network_for_singlesensor_forecasting"},"authorids":{"value":["~Adrian_Kucia1","~Edward_Rollason1","~Wai_Lok_Woo1"]},"authors":{"value":["Adrian Kucia","Edward Rollason","Wai Lok Woo"]}},"version":2},{"content":{"venue":{"value":"CoRR 2020"},"pdf":{"value":"http://arxiv.org/pdf/2012.03108v1"},"venueid":{"value":"dblp.org/journals/CORR/2020"},"paperhash":{"value":"mohandoss|generating_synthetic_multispectral_satellite_imagery_from_sentinel2"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Tharun_Mohandoss:","https://dblp.org/search/pid/api?q=author:Aditya_Kulkarni:","https://dblp.org/search/pid/api?q=author:Daniel_Northrup:","https://dblp.org/search/pid/api?q=author:Ernest_Mwebaze:","~Hamed_Alemohammad1"]},"html":{"value":"https://arxiv.org/abs/2012.03108"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2012-03108,\n  publtype={informal},\n  author={Tharun Mohandoss and Aditya Kulkarni and Daniel Northrup and Ernest Mwebaze and Hamed Alemohammad},\n  title={Generating Synthetic Multispectral Satellite Imagery from Sentinel-2},\n  year={2020},\n  cdate={1577836800000},\n  journal={CoRR},\n  volume={abs/2012.03108},\n  url={https://arxiv.org/abs/2012.03108}\n}\n"},"abstract":{"value":"Multi-spectral satellite imagery provides valuable data at global scale for many environmental and socio-economic applications. Building supervised machine learning models based on these imagery, however, may require ground reference labels which are not available at global scale. Here, we propose a generative model to produce multi-resolution multi-spectral imagery based on Sentinel-2 data. The resulting synthetic images are indistinguishable from real ones by humans. This technique paves the road for future work to generate labeled synthetic imagery that can be used for data augmentation in data scarce regions and applications."},"title":{"value":"Generating Synthetic Multispectral Satellite Imagery from Sentinel-2"},"authors":{"value":["Tharun Mohandoss","Aditya Kulkarni","Daniel Northrup","Ernest Mwebaze","Hamed Alemohammad"]}},"tmdate":1747277294041,"pdate":1577836800000,"tcdate":1747277285682,"writers":["~"],"signatures":["~Hamed_Alemohammad1"],"forum":"VvxBWt5XnR","license":"CC BY-SA 4.0","number":471114,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747277294041,"domain":"DBLP.org","id":"VvxBWt5XnR","version":2},{"content":{"summary":{"value":"This work explores the use of debate as a scalable oversight method for language models. The proposed method trains models to debate using self-play and assess whether this approach enhances the accuracy of AI-based evaluators in judging complex tasks, such as reading comprehension. The key contribution of the paper is demonstrating that models trained to debate in adversarial settings lead to more accurate judgments by AI evaluators compared to non-adversarial consultancy models."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. See weaknessnes.\n\n2. Figures 1, 2, and 5 are not referenced in the text. It would be helpful to cite them in the relevant sections to enhance clarity and support the descriptions."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. This paper introduces debate as a method for scalable oversight by leveraging self-play for training language models, which shows promise for improving evaluator accuracy in complex tasks.\n\n2. This paper discusses and compares multiple baselines (e.g., single, ensembled, and double consultancy)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper's experiments are limited to reading comprehension tasks, which raises concerns about the generalizability of the findings to other complex reasoning domains. Could the effectiveness of debate training extend to tasks beyond reading comprehension?\n\n2. The debate protocol is limited to a **two-turn** setting, which restricts the depth of argumentation and refutation between debaters. This simplified structure may not fully capture the complexities of real-world debates, where extended back-and-forth exchanges are often necessary to expose subtle flaws in reasoning. As a result, the findings may overestimate the effectiveness of debate in fostering accurate judgments, since more intricate discussions could reveal different dynamics or weaknesses in the model's performance​.\n\n3. While the experiments provide some insights, they are not entirely convincing. The study only employs GPT-4-Turbo as the judge and Llama3-8B-Instruct as the debate and consultancy models. Evaluating the proposed method across a broader range of model architectures and sizes would strengthen the assessment of its generalizability.\n\n4. Although debate is presented as a mechanism for scalable oversight, the paper finds little evidence that explicit refutation materially affects the judge’s decision-making. This undermines one of the key proposed benefits of debate as an oversight mechanism​."}},"nonreaders":[],"tmdate":1731428380976,"tcdate":1729586655433,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7884/Reviewer_3vo5"],"signatures":["ICLR.cc/2025/Conference/Submission7884/Reviewer_3vo5"],"forum":"gAEEjGv5Oa","number":1,"license":"CC BY 4.0","cdate":1729586655433,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7884/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428380976,"domain":"ICLR.cc/2025/Conference","replyto":"gAEEjGv5Oa","id":"JropzMue5I","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We test the effectiveness of debate as a method of scalable oversight, finding that language model evaluators can answer questions more accurately when exposed to models trained to debate with self-play."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["AI Safety","Scalable Oversight","Reinforcement Learning","Debate"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We test the robustness of debate as a method of scalable oversight by training models to debate with data generated via self-play. In a long-context reading comprehension task, we find that language model based evaluators answer questions more accurately when judging models optimized to win debates. By contrast, we find no such relationship for consultancy models trained to persuade a judge without an opposing debater present. In quantitative and qualitative comparisons between our debate models and novel consultancy baselines, we find evidence that debate training encourages stronger and more informative arguments, showing promise that it can help provide high-quality supervision for tasks that are difficult to directly evaluate."},"_bibtex":{"value":"@misc{\narnesen2025training,\ntitle={Training Language Models to Win Debates with Self-Play Improves Judge Accuracy},\nauthor={Samuel Arnesen and David Rein and Julian Michael},\nyear={2025},\nurl={https://openreview.net/forum?id=gAEEjGv5Oa}\n}"},"title":{"value":"Training Language Models to Win Debates with Self-Play Improves Judge Accuracy"},"pdf":{"value":"/pdf/0aaa9c0f990491a143568c769e82d0ca4509f99b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"arnesen|training_language_models_to_win_debates_with_selfplay_improves_judge_accuracy"},"authorids":{"value":["~Samuel_Arnesen1","~David_Rein1","~Julian_Michael1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Samuel Arnesen","David Rein","Julian Michael"]}},"version":2},{"content":{"venue":{"value":"IEEE Transactions on Instrumentation and Measurement"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/19/10012124/09975275.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"zhang|boosting_personalized_musculoskeletal_modeling_with_physicsinformed_knowledge_transfer"},"html":{"value":"https://doi.org/10.1109/TIM.2022.3227604"},"abstract":{"value":"Data-driven methods have become increasingly more prominent for musculoskeletal modeling due to their conceptually intuitive simple and fast implementation. However, the performance of a pretrained data-driven model using the data from specific subject(s) may be seriously degraded when validated using the data from a new subject, hindering the utility of the personalized musculoskeletal model in clinical applications. This article develops an active physics-informed deep transfer learning framework to enhance the dynamic tracking capability of the musculoskeletal model on the unseen data. The salient advantages of the proposed framework are twofold. First, for the generic model, physics-based domain knowledge is embedded into the loss function of the data-driven model as soft constraints to penalize/regularize the data-driven model. Second, for the personalized model, the parameters relating to the feature extraction will be directly inherited from the generic model, and only the parameters relating to the subject-specific inference will be fine-tuned by jointly minimizing the conventional data prediction loss and the modified physics-based loss. In this article, we use the synchronous muscle forces and joint kinematics prediction from surface electromyogram (sEMG) as the exemplar to illustrate the proposed framework. Moreover, convolutional neural network (CNN) is employed as the deep neural network to implement the proposed framework, and the physics law between muscle forces and joint kinematics is utilized as the soft constraints. Results of comprehensive experiments on a self-collected dataset from eight healthy subjects indicate the effectiveness and great generalization of the proposed framework."},"title":{"value":"Boosting Personalized Musculoskeletal Modeling With Physics-Informed Knowledge Transfer"},"authors":{"value":[{"fullname":"Jie Zhang"},{"fullname":"Yihui Zhao"},{"fullname":"Tianzhe Bao"},{"fullname":"Zhenhong Li"},{"fullname":"Kun Qian"},{"fullname":"Alejandro F. Frangi","username":"~Alejandro_F._Frangi1"},{"fullname":"Sheng Quan Xie"},{"fullname":"Zhi-Qiang Zhang"}]}},"tmdate":1789090559514,"pdate":1672531200000,"externalIds":["doi:10.1109/tim.2022.3227604"],"tcdate":1769105300213,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Alejandro_F._Frangi1"],"forum":"ksL7Xnoo6K","license":"CC BY-SA 4.0","number":35697,"cdate":1670526791974,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789090559514,"domain":"OpenReview.net/Public_Article","id":"ksL7Xnoo6K","version":2},{"content":{"summary":{"value":"This paper introduces a frequency-based, physics-informed approach to enhance motion quality and physical plausibility in video diffusion models without degrading visual fidelity or text alignment.\nRather than operating purely in pixel or latent space, the method injects a low-pass frequency-domain constraint that regularizes temporal dynamics and enforces physically consistent motion patterns.\nThe approach is applied as an auxiliary frequency-domain loss during training, compatible with existing architectures such as Open-Sora, MVDIT, and Hunyuan , by training a LoRA. It requires no modification to the diffusion backbone and adds only moderate computational cost.\nMain contributions are:\n-  A physics-guided frequency-domain regularization for video diffusion training that improves temporal motion realism.\n-  Efficient low-pass truncation scheme reducing computational cost. Differentiable frequency-domain least-squares loss integrated seamlessly into standard diffusion training loops.\n-  Extensive empirical validation on multiple video diffusion models."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"-  Does this method support multiple object motions? (e.g. if a single apple is cut in half and the two halves split away.)\n-  Does this method support the color, illumination, or texture change of an object within a video? If so, how does this loss reduce color flickering like the train in Appendix Fig. 4 ?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":4},"strengths":{"value":"The paper introduces a novel frequency-domain regularization for video diffusion models that leverages spectral signatures of translation, rotation, and scaling to guide learning without altering model architecture. The strengths are:\n-  The idea of combining classical ideas from Fourier analysis and the SIM(2) motion group with modern video diffusion models, demonstrating a creative synthesis of physics-based priors and deep generative modeling.\n-  The authors provide a thorough derivation connecting basic physical motions (translation, rotation, scaling) to spectral signatures, with attention to windowing, interpolation errors, and numerical stability. The breakdown of translational, rotational, and scaling motion losses, along with adaptive weighting, is logically organized and explained with intuitive interpretations.\n-  Results are reported on multiple video diffusion backbones and evaluated on diverse metrics. The experiments are comprehensive (no LoRA, +LoRA with other losses, +LoRA with proposed loss), and quantitative gains are consistent."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-  Although the theory is solid, as a paper in the video generation field, its presentation lacks some intuitive visualizations, such as visual demonstrations of spectral changes, and the qualitative evaluation is relatively limited;\n-  In the Abstract, “regularizer” is written as “regular- izer,” which looks like a copy-paste error;\n-  On the first page, in the “four groups” listing, why only (i) is bolded;\n-  As an important demonstration, the supplementary video is of poor asthetic quality and needs improvement.\n\nOverall, the paper has no significant issues in theory or experiments, but there are some minor presentation problems."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917207213,"tcdate":1762160819234,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4165/Reviewer_1ggH"],"signatures":["ICLR.cc/2026/Conference/Submission4165/Reviewer_1ggH"],"forum":"jhan3NJ5x1","number":3,"license":"CC BY 4.0","cdate":1762160819234,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4165/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917207213,"domain":"ICLR.cc/2026/Conference","replyto":"jhan3NJ5x1","id":"ega5RtREcI","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video generation","Diffusion model"]},"supplementary_material":{"value":"/attachment/61e5db89c3e0dcf46f52dfd534527c73711f043e.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Current video diffusion models generate visually compelling content but often violate \nbasic laws of physics, producing subtle artifacts like rubber-sheet deformations and \ninconsistent object motion. We introduce a frequency-domain physics prior that improves \nmotion plausibility without modifying model architectures. Our method decomposes common \nrigid motions (translation, rotation, scaling) into lightweight spectral losses, \nrequiring only 2.7% of frequency coefficients while preserving 97%+ of spectral energy. \nApplied to Open-Sora, MVDIT, and Hunyuan, our approach improves both motion accuracy and action recognition by ~11\\% on average on OpenVID-1M (relative), while maintaining visual quality. User studies show 74--83% preference for our physics-enhanced videos. It also reduces warping error by 22--37% (depending on the backbone) and improves temporal consistency scores. These results indicate that simple, global spectral cues are an effective drop-in regularizer for physically plausible motion in video diffusion."},"_bibtex":{"value":"@misc{\nanonymous2026physicsguided,\ntitle={Physics-Guided Motion Loss for Video Generation Model},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=jhan3NJ5x1}\n}"},"title":{"value":"Physics-Guided Motion Loss for Video Generation Model"},"pdf":{"value":"/pdf/5a7c6e0e967cec1df42675ec0f9003da99cdfd85.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"xue|physicsguided_motion_loss_for_video_generation_model"},"authorids":{"value":["~Bowen_Xue1","~Giuseppe_Claudio_Guarnera1","~Shuang_Zhao1","~Zahra_Montazeri1"]},"authors":{"value":["Bowen Xue","Giuseppe Claudio Guarnera","Shuang Zhao","Zahra Montazeri"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a pipeline for generating synthetic instruction tuning data. The method consists of the following steps: 1. data filtering is applied to seed coding data to select high quality examples; 2. base LLM is used to generate a set of coding concept and category based on the seed data; 3. base LLM is used to generate coding instruction, response and test; 4. generated examples are selected based on the code execution result."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Have you tried this framework using stronger LLM to generate synthetic data?\n2. Can you get even better performance by running several rounds of data generation with improved base model?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. the paper focuses on using base model to generate synthetic data to self-improve, which is an interesting and useful angle for synthetic data generation\n2. the method is evaluated on several different coding LLM benchmarks which shows the effectiveness of the method\n3. there are also ablation experiments verifying the contribution of specific design choices in the framework."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While using base model to self-improve is an interesting and useful direction, synthetic data generation could be improved by using a stronger LLM than the base model. It is not clear from the paper whether the proposed framework would be effective compared to previous methods if we use a stronger LLM to synthesize the data. The synthetic data generation could also be potentially improved by having multiple rounds of data generation process."},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1730879126833,"tcdate":1721052483419,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission7042/Reviewer_NEpV"],"signatures":["NeurIPS.cc/2024/Conference/Submission7042/Reviewer_NEpV"],"forum":"xXRnUU7xTL","number":4,"license":"CC BY 4.0","cdate":1721052483419,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission7042/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879126833,"domain":"NeurIPS.cc/2024/Conference","replyto":"xXRnUU7xTL","id":"j9oS0cI09q","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"Self-alignment for code LLMs without human annotations"},"keywords":{"value":["Large language models","Code generation","Instruction tuning","Self-Alignment"]},"supplementary_material":{"value":"/attachment/bcec3bbed475d368a97007019676fbf3d48cb111.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Instruction tuning is a supervised fine-tuning approach that significantly improves the ability of large language models (LLMs) to follow human instructions. For programming tasks, most models are finetuned with costly human-annotated instruction-response pairs or those generated by large, proprietary LLMs, which may not be permitted. We propose SelfCodeAlign, the first fully transparent and permissive pipeline for self-aligning code LLMs without extensive human annotations or distillation. SelfCodeAlign employs the same base model for inference throughout the data generation process. It first extracts diverse coding concepts from high-quality seed snippets to generate new tasks. It then samples multiple responses per task, pairs each with test cases, and validates them in a sandbox environment. Finally, passing examples are selected for instruction tuning. In our primary experiments, we use SelfCodeAlign with CodeQwen1.5-7B to generate a dataset of 74k instruction-response pairs. Finetuning on this dataset leads to a model that achieves a 67.1 pass@1 on HumanEval+, surpassing CodeLlama-70B-Instruct despite being ten times smaller. Across all benchmarks, this finetuned model consistently outperforms the original version trained with OctoPack, the previous state-of-the-art method for instruction tuning without human annotations or distillation. Additionally, we show that SelfCodeAlign is effective across LLMs of various sizes, from 3B to 33B, and that the base models can benefit more from alignment with their own data distribution. We further validate each component’s effectiveness in our pipeline, showing that SelfCodeAlign outperforms both direct distillation from GPT-4o and leading GPT-3.5-based distillation methods, such as OSS-Instruct and Evol-Instruct. SelfCodeAlign has also led to the creation of StarCoder2-Instruct, the first fully transparent, permissively licensed, and self-aligned code LLM that achieves state-of-the-art coding performance. Overall, SelfCodeAlign shows for the first time that a strong instruction-tuned code LLM can result from self-alignment rather than distillation."},"_bibtex":{"value":"@inproceedings{\nwei2024selfcodealign,\ntitle={SelfCodeAlign: Self-Alignment for Code Generation},\nauthor={Yuxiang Wei and Federico Cassano and Jiawei Liu and Yifeng Ding and Naman Jain and Zachary Mueller and Harm de Vries and Leandro Von Werra and Arjun Guha and LINGMING ZHANG},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=xXRnUU7xTL}\n}"},"title":{"value":"SelfCodeAlign: Self-Alignment for Code Generation"},"pdf":{"value":"/pdf/f4cd100d3f9f85fe8c929ea517dc4cbd24143e72.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"wei|selfcodealign_selfalignment_for_code_generation"},"authorids":{"value":["~Yuxiang_Wei2","~Federico_Cassano1","~Jiawei_Liu11","~Yifeng_Ding2","~Naman_Jain2","~Zachary_Mueller1","~Harm_de_Vries1","~Leandro_Von_Werra1","~Arjun_Guha3","~LINGMING_ZHANG2"]},"authors":{"value":["Yuxiang Wei","Federico Cassano","Jiawei Liu","Yifeng Ding","Naman Jain","Zachary Mueller","Harm de Vries","Leandro Von Werra","Arjun Guha","LINGMING ZHANG"]}},"version":2},{"content":{"TLDR":{"value":"Compatibility training avoids misspecification bias under uncertain physics."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Scientific machine learning","Physics-informed neural networks","Neural operators","Model uncertainty","Model misspecification","Uncertainty quantification"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics-informed machine learning usually penalizes violations of a chosen governing equation.  However, when that equation is misspecified, minimizing the nominal residual can instead increase the error in the predicted solution.  To make physics-informed machine learning effective under such misspecification, we propose compatibility training, which uses the smallest residual over a plausible family of equations.  At each training step, compatibility training selects the equation in this family that best fits the current prediction and penalizes only deviations from that equation.  Our theory establishes a mismatch-dependent bias floor for nominal training and characterizes the physics constraints that remain after the uncertain parameters are adjusted. We derive an exact local curvature identity that shows when the retained physics constraints and observations together identify the true solution.  This curvature identity connects solution recovery to the structure and extent of the uncertainty family.  Across multiple synthetic PINN benchmarks, PDEBench trajectory and operator tasks, and Darcy operator-learning experiments, compatibility training improves prediction under model mismatch, including in settings with sparse or noisy observations and field-valued uncertainty.  In field-valued Darcy flow, compatibility training also reduces the true residual by more than an order of magnitude."},"_bibtex":{"value":"@inproceedings{\nanonymous2026unbiased,\ntitle={Unbiased Physics Losses Under Model Uncertainty},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=vv6ySNHTI8},\nnote={under review}\n}"},"title":{"value":"Unbiased Physics Losses Under Model Uncertainty"},"pdf":{"value":"/pdf/accdc2c353d2e1686660d28f656923bd9acc2c9b.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791225621526,"tcdate":1789008552717,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission12736/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission12736/Authors"],"forum":"vv6ySNHTI8","license":"CC BY 4.0","number":12736,"cdate":1789008552717,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791225621526,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"vv6ySNHTI8","version":2},{"content":{"venue":{"value":"CoRR 2020"},"pdf":{"value":"http://arxiv.org/pdf/2006.04976v2"},"venueid":{"value":"dblp.org/journals/CORR/2020"},"paperhash":{"value":"wang|physics_regularized_gaussian_processes"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zheng_Wang_0042:","https://dblp.org/search/pid/api?q=author:Wei_W._Xing:","https://dblp.org/search/pid/api?q=author:Robert_Michael_Kirby:","~Shandian_Zhe1"]},"html":{"value":"https://arxiv.org/abs/2006.04976"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2006-04976,\n  publtype={informal},\n  author={Zheng Wang and Wei W. Xing and Robert Michael Kirby and Shandian Zhe},\n  title={Physics Regularized Gaussian Processes},\n  year={2020},\n  cdate={1577836800000},\n  journal={CoRR},\n  volume={abs/2006.04976},\n  url={https://arxiv.org/abs/2006.04976}\n}\n"},"abstract":{"value":"Deep kernel learning is a promising combination of deep neural networks and nonparametric function learning. However, as a data driven approach, the performance of deep kernel learning can still be restricted by scarce or insufficient data, especially in extrapolation tasks. To address these limitations, we propose Physics Informed Deep Kernel Learning (PI-DKL) that exploits physics knowledge represented by differential equations with latent sources. Specifically, we use the posterior function sample of the Gaussian process as the surrogate for the solution of the differential equation, and construct a generative component to integrate the equation in a principled Bayesian hybrid framework. For efficient and effective inference, we marginalize out the latent variables in the joint probability and derive a collapsed model evidence lower bound (ELBO), based on which we develop a stochastic model estimation algorithm. Our ELBO can be viewed as a nice, interpretable posterior regularization objective. On synthetic datasets and real-world applications, we show the advantage of our approach in both prediction accuracy and uncertainty quantification."},"title":{"value":"Physics Regularized Gaussian Processes"},"authors":{"value":["Zheng Wang","Wei W. Xing","Robert Michael Kirby","Shandian Zhe"]}},"tmdate":1747367845154,"pdate":1577836800000,"tcdate":1747367834464,"writers":["~"],"signatures":["~Shandian_Zhe1"],"forum":"Oea1HslYR8","license":"CC BY-SA 4.0","number":503920,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747367845154,"domain":"DBLP.org","id":"Oea1HslYR8","version":2},{"content":{"comment":{"value":"We appreciate the reviewer's suggestion to intuitively demonstrate the effectiveness of incorporating user history into item tokenization. Below, we provide two concrete examples:\n\n> Intuitive demonstration: Example 1\n\nIn our initial submission, we included a case study in Section 3.5 (Figure 4), which illustrates a real-world example observed in our \"Game\" dataset.\n\n**TL;DR**: The same game, StarCraft II, is tokenized into different Semantic IDs based on user history: one sequence reflects a user interested in story-driven games, while the other reflects a user interested in real-time strategy (RTS) games.\n\n**Details**: We observed two users with different preferences interacting with the same target item, StarCraft II:\n* **Story-driven context (user A)**: This user's history features narrative-heavy titles such as Tomb Raider and The Last of Us. Their interaction with StarCraft II is likely motivated by its extensive campaign mode.\n* **RTS context (user B)**: This user's history consists of classic strategy games like Warcraft III and Command & Conquer. Their interest aligns with the game's core RTS mechanics.\n\nAs shown in Figure 4, Pctx tokenizes StarCraft II into different Semantic IDs for these two users, resulting in:\n* StarCraft II (for user A): `[53, 395, 576, 770]`\n* StarCraft II (for user B): `[53, 412, 576, 770]`\n\nThis example shows that incorporating user history allows the tokenizer to capture the specific aspect of an item relevant to the user's context.\n\n> Intuitive demonstration: Example 2\n\nWe provide an additional example that directly echoes the reviewer's request regarding \"extremely similar items\". We examine a multi-faceted item: the \"Nintendo Switch - Mario Kart 8 Deluxe Bundle\" (which contains both a console and a game), compared against a pure game, \"Mario Kart Live\".\n\n**TL;DR**: Non-personalized tokenizers (*e.g.*, TIGER) yield nearly identical IDs for the bundle and the single game. In contrast, our Pctx assigns different semantic IDs for the bundle depending on user intent: it looks like a \"game\" when the user seeks content, but looks like \"accessories\" when the user is buying a new console.\n\n**Details**:\n* **Item A (The bundle)**: Nintendo Switch - Mario Kart 8 Deluxe bundle.\n* **Item B (The game)**: Mario Kart Live.\n\n1. **Static tokenization** (*e.g.*, TIGER): A static tokenizer maps both items to highly similar IDs sharing the first two tokens.:\n    * Item A (bundle): `[167, 466, 646, 771]`\n    * Item B (game): `[167, 466, 586, 770]`\n2. **Our personalized tokenization** (Pctx): Our model adaptively tokenizes item A (the bundle) based on the user's context:\n    * Gaming context: When the user history indicates a preference for racing games, the bundle is tokenized as `[191, 334, 744, 770]`. This shares the same prefix (191, 334) with item B (`[191, 334, 760, 770]`).\n    * Hardware/accessory context: When the user history implies a need for a new setup (*e.g.*, buying a new console), the bundle is tokenized as `[152, 334, 688, 770]`. Notably, this prefix (152, 334) aligns with hardware accessories like \"Nintendo Labo\" (`[152, 334, 519, 770]`), which are often purchased alongside a new Switch.\n\nThis demonstrates that Pctx effectively interprets the bundle, making its ID distinguishable based on whether the user views it as \"game\" or \"hardware\"."},"title":{"value":"Intuitive Demonstrations"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763982371954,"tcdate":1763982371954,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15948/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission15948/Authors"],"forum":"ahpO7S1Ppi","number":10,"license":"CC BY 4.0","cdate":1763982371954,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15948/-/Official_Comment"],"mdate":1763982371954,"domain":"ICLR.cc/2026/Conference","replyto":"0ynWRQLon9","id":"HvXcV6g38D","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Generative Recommendation","Personalization","Tokenization"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Generative recommendation (GR) models tokenize each action into a few discrete tokens (called semantic IDs) and autoregressively generate the next tokens as predictions, showing advantages such as memory efficiency, scalability, and the potential to unify retrieval and ranking.\nDespite these benefits, existing tokenization methods are static and non-personalized. They typically derive semantic IDs solely from item features, assuming a universal item similarity that overlooks user-specific perspectives. However, under the autoregressive paradigm, semantic IDs with the same prefixes always receive similar probabilities, so a single fixed mapping implicitly enforces a universal item similarity standard across all users. In practice, the same item may be interpreted differently depending on user intentions and preferences. To address this issue, we propose a personalized context-aware tokenizer that incorporates a user's historical interactions when generating semantic IDs. This design allows the same item to be tokenized into different semantic IDs under different user contexts, enabling GR models to capture multiple interpretive standards and produce more personalized predictions. Experiments on three public datasets demonstrate up to 8.9% improvement in NDCG@10 over non-personalized action tokenization baselines. Our code is available at https://anonymous.4open.science/r/Pctx-code-4246."},"_bibtex":{"value":"@misc{\nzhong2026pctx,\ntitle={Pctx: Tokenizing Personalized Context for Generative Recommendation},\nauthor={Qiyong Zhong and Jiajie Su and Yunshan Ma and Julian McAuley and Yupeng Hou},\nyear={2026},\nurl={https://openreview.net/forum?id=ahpO7S1Ppi}\n}"},"title":{"value":"Pctx: Tokenizing Personalized Context for Generative Recommendation"},"pdf":{"value":"/pdf/ad45a0e0b43c7fa044ff865a614abfd6604ed99d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhong|pctx_tokenizing_personalized_context_for_generative_recommendation"},"authorids":{"value":["~Qiyong_Zhong1","~Jiajie_Su1","~Yunshan_Ma1","~Julian_McAuley1","~Yupeng_Hou1"]},"authors":{"value":["Qiyong Zhong","Jiajie Su","Yunshan Ma","Julian McAuley","Yupeng Hou"]}},"version":2},{"content":{"venue":{"value":"Nature Machine Intelligence"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"buschoff|visual_cognition_in_multimodal_large_language_models"},"authorids":{"value":["~Luca_M._Schulze_Buschoff1","~Elif_Akata1","~Matthias_Bethge1","~Eric_Schulz1"]},"html":{"value":"https://www.nature.com/articles/s42256-024-00963-y.pdf"},"abstract":{"value":"A chief goal of artificial intelligence is to build machines that think like people. Yet it has been argued that deep neural network architectures fail to accomplish this. Researchers have asserted these models' limitations in the domains of causal reasoning, intuitive physics and intuitive psychology. Yet recent advancements, namely the rise of large language models, particularly those designed for visual processing, have rekindled interest in the potential to emulate human-like cognitive abilities. This paper evaluates the current state of vision-based large language models in the domains of intuitive physics, causal reasoning and intuitive psychology. Through a series of controlled experiments, we investigate the extent to which these modern models grasp complex physical interactions, causal relationships and intuitive understanding of others' preferences. Our findings reveal that, while some of these models demonstrate a notable proficiency in processing and interpreting visual data, they still fall short of human capabilities in these areas. Our results emphasize the need for integrating more robust mechanisms for understanding causality, physical dynamics and social cognition into modern-day, vision-based language models, and point out the importance of cognitively inspired benchmarks."},"title":{"value":"Visual cognition in multimodal large language models"},"authors":{"value":["Luca M. Schulze Buschoff","Elif Akata","Matthias Bethge","Eric Schulz"]}},"tmdate":1778151136820,"pdate":1732489200000,"tcdate":1778151136820,"writers":["~Luca_M._Schulze_Buschoff1","~Elif_Akata1","~Matthias_Bethge1","~Eric_Schulz1"],"signatures":["~Elif_Akata1"],"forum":"BZzalztUAL","license":"CC BY 4.0","number":49407,"cdate":1778151136820,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1778151136820,"domain":"OpenReview.net/Archive","id":"BZzalztUAL","version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2311.16093v3"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"buschoff|have_we_built_machines_that_think_like_people"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Luca_M._Schulze_Buschoff:","https://dblp.org/search/pid/api?q=author:Elif_Akata:","~Matthias_Bethge1","https://dblp.org/search/pid/api?q=author:Eric_Schulz:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2311.16093"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2311-16093,\n  publtype={informal},\n  author={Luca M. Schulze Buschoff and Elif Akata and Matthias Bethge and Eric Schulz},\n  title={Have we built machines that think like people?},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2311.16093},\n  url={https://doi.org/10.48550/arXiv.2311.16093}\n}\n"},"abstract":{"value":"A chief goal of artificial intelligence is to build machines that think like people. Yet it has been argued that deep neural network architectures fail to accomplish this. Researchers have asserted these models' limitations in the domains of causal reasoning, intuitive physics, and intuitive psychology. Yet recent advancements, namely the rise of large language models, particularly those designed for visual processing, have rekindled interest in the potential to emulate human-like cognitive abilities. This paper evaluates the current state of vision-based large language models in the domains of intuitive physics, causal reasoning, and intuitive psychology. Through a series of controlled experiments, we investigate the extent to which these modern models grasp complex physical interactions, causal relationships, and intuitive understanding of others' preferences. Our findings reveal that, while some of these models demonstrate a notable proficiency in processing and interpreting visual data, they still fall short of human capabilities in these areas. Our results emphasize the need for integrating more robust mechanisms for understanding causality, physical dynamics, and social cognition into modern-day, vision-based language models, and point out the importance of cognitively-inspired benchmarks."},"title":{"value":"Have we built machines that think like people?"},"authors":{"value":["Luca M. Schulze Buschoff","Elif Akata","Matthias Bethge","Eric Schulz"]}},"tmdate":1731493425568,"pdate":1672531200000,"tcdate":1731490359124,"writers":["~"],"signatures":["~Matthias_Bethge1"],"forum":"GUGUqDmmLE","license":"CC BY-SA 4.0","number":221222,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731493425568,"domain":"DBLP.org","id":"GUGUqDmmLE","version":2},{"content":{"summary":{"value":"This paper proposes a data-efficient learning framework for PDE dynamics forecasting by jointly learning from both the original PDEs and their simplified basic forms. Extensive experiments on a wide range of 1D/2D/3D PDE problems demonstrates the effectiveness of the proposed framework."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"See weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The proposed method is well-motivated. The authors provide a critical observation by evaluating existing SciML foundation models. They find a strong correlation between a model's performance on the original PDE and its performance on the fundamental components of that PDE (e.g., pure diffusion for a reaction-diffusion system). However, the absolute error on these basic terms remains high, indicating that even powerful models lack a robust understanding of the foundational physics, which motivates the need for explicit training on these concepts.\n- Methodological Innovation:​​ The paper proposes a simple yet effective multiphysics training framework. It first derive a \"basic form\" from the original PDE by retaining terms governing essential dynamics and removing terms that cause computational stiffness or high cost. The model is trained on a composite dataset from simulations of both the original PDE and the basic form."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Heuristic Nature of Decomposition:​​ The process for selecting terms for the \"basic form,\" while physically intuitive, remains heuristic. A more formalized principle or an ablation study discussing the impact of alternative decompositions more prominently would strengthen the methodology. \n- Inadequate Mechanistic Explanation for the Efficacy of Basic Form Data​: A significant weakness of the paper lies in its insufficient exploration of the underlying mechanisms by which the \"basic form\" data aids the learning of the original PDE. The attribution of performance gains solely to the incorporation of \"fundamental physics knowledge\" is a high-level concept that lacks granularity. A more rigorous analysis is required to dissect how the basic form data contributes. A possible explanation is that data from the basic form may provide more diverse initial conditions. Can the data from the basic form be replaced with an equivalent amount of original PDE data? Although this would incur greater simulation costs, it would help clarify the specific ways in which data from the basic form aids the model in learning the original PDE."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358546519,"tcdate":1761571490102,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14708/Reviewer_dWHm"],"signatures":["ICLR.cc/2026/Conference/Submission14708/Reviewer_dWHm"],"forum":"mJiPqOzc3O","number":1,"license":"CC BY 4.0","cdate":1761571490102,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14708/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358546519,"domain":"ICLR.cc/2026/Conference","replyto":"mJiPqOzc3O","id":"Rcm76I4eQ9","forumContent":{"TLDR":{"value":"We propose to incorporate fundamental physics knowledge into learning neural operators to enhance its data efficiency, long-term consistency, and OOD generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Neural Operator","PDE","Fundamental Physics Knowledge"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Recent advances in scientific machine learning (SciML) have enabled neural operators (NOs) to serve as powerful surrogates for modeling the dynamic evolution of physical systems governed by partial differential equations (PDEs). While existing approaches focus primarily on learning simulations from the target PDE, they often overlook more fundamental physical principles underlying these equations. Inspired by how numerical solvers are compatible with simulations of different settings of PDEs, we propose a multiphysics training framework that jointly learns from both the original PDEs and their simplified basic forms. Our framework enhances data efficiency, reduces predictive errors, and improves out-of-distribution (OOD) generalization, particularly in scenarios involving shifts of physical parameters and synthetic-to-real transfer. Our method is architecture-agnostic and demonstrates consistent improvements in normalized root mean square error (nRMSE) across a wide range of 1D/2D/3D PDE problems. Through extensive experiments, we show that explicit incorporation of fundamental physics knowledge significantly strengthens the generalization ability of neural operators.\nWe will release models and codes at https://sites.google.com/view/sciml-fundemental-pde."},"_bibtex":{"value":"@inproceedings{\nma2026learning,\ntitle={Learning Data-Efficient and Generalizable Neural Operators via Fundamental Physics Knowledge},\nauthor={Siying Ma and Mehrdad Momeni Zadeh and Mauricio Soroco and Wuyang Chen and Jiguo Cao and Vijay Ganesh},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=mJiPqOzc3O}\n}"},"title":{"value":"Learning Data-Efficient and Generalizable Neural Operators via Fundamental Physics Knowledge"},"pdf":{"value":"/pdf/27f1b69d2b552cb0e5d0a96e3231fc4148675ae2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ma|learning_dataefficient_and_generalizable_neural_operators_via_fundamental_physics_knowledge"},"authorids":{"value":["~Siying_Ma1","~Mehrdad_Momeni_Zadeh1","~Mauricio_Soroco1","~Wuyang_Chen1","~Jiguo_Cao1","~Vijay_Ganesh1"]},"authors":{"value":["Siying Ma","Mehrdad Momeni Zadeh","Mauricio Soroco","Wuyang Chen","Jiguo Cao","Vijay Ganesh"]}},"version":2},{"content":{"summary":{"value":"This paper explores problems with privacy metrics used to evaluate synthetic data (generated without differential privacy). The paper demonstrates the undesirable properties of these metrics and shows the viability of a reconstruction attack. The paper suggests a need to move away from ad hoc approaches to privacy . Extensive experiments are conducted demonstrating the viability of the attack on real synthetic data under a reasonable threat model."},"presentation":{"value":"3 good"},"contribution":{"value":"4 excellent"},"soundness":{"value":"3 good"},"strengths":{"value":"The paper is well motivated, the false promise of synthetic data generation (without DP) is a very important message. The overview of existing metrics for evaluating synthetic data alongside their weaknesses is a very clearly presented and a valuable contribution in its own right.\n\n The experiments are thorough and extensive detailed is given in the appendix. \n\nOverall the conclusion of the paper that effective attacks on synthetic data exist seems well supported by the experimental results, assuming the presentation can be improved."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"A thorough description of the attack is lacking in section 4 which significantly weakens the comprehensiveness of the paper in my view. I understand why the full Algorithm was relegated to the appendix but a paragraph explain the logic of the attack and where the information is being extracted would significantly improve the paper. Given the vulnerability of models trained with DP to this attack, the information is presumably coming from the metrics but an intuitive explanation of how that occurs is missing. \n\nFigure 2 should be presented in a way where it is possible to distinguish between the different training algorithms and datasets. I would suggest enlarging the figure despite the constrain on space since it is a key result for the paper. I find it more useful that FIgure 4, which is given much more space. \n\nI do not entirely understand what the takeaway is from Figure 3. The figure is referred to only in passing and never explained. More generally \nthe description of the results would benefit from more thorough explanation. \n\nAs far as I can tell, 'outliers' is never clearly defined but it is the denominator for all success rates reported and so it is important to explain what this set of data points is and how much of the training data they make up. \n\nStating that the attack works on models trained with DP in the abstract/intro is true in a literal sense but I feel somewhat disingenous since as noted later in the paper, the information is come from the statistics released without DP. Any model regarded as formally satisfying DP would never be able to look directly at the private data in this way. \n\nOverall, I believe the paper has a strong attack but these results could be better presented to really hammer the point to ensure the key message of this paper reaches the synthetic data community. Ideally there could be a figure that made it impossible to ignore the weaknesses of the synthetic data generation but the current figures are overly information loaded causing the key takeaway to be lost."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Can you report the privacy metrics with DP or would your recommendation be to not report distance metrics? \n\nHow do you define an outlier? As far as I can tell, 'outliers' is never clearly defined but it is the denominator for all success rates reported and so it is important to explain what this set of data points is and how much of the training data they make up."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1701465570470,"tcdate":1698709702216,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6385/Reviewer_XpV4"],"signatures":["ICLR.cc/2024/Conference/Submission6385/Reviewer_XpV4"],"forum":"g16vmAtJ8x","number":3,"license":"CC BY 4.0","cdate":1698709702216,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6385/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1701465570470,"domain":"ICLR.cc/2024/Conference","replyto":"g16vmAtJ8x","id":"alj7RyRBX2","forumContent":{"TLDR":{"value":"We demonstrate the inadequacy of commonly used similarity-based privacy metrics to guarantee privacy in synthetic data though analysis, counter-examples, and a novel reconstruction attack."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic data","privacy metrics","reconstruction attacks","differential privacy","generative models"]},"primary_area":{"value":"societal considerations including fairness, safety, privacy"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Training generative models to produce synthetic data is meant to provide a privacy-friendly approach to data release.\nHowever, we get robust guarantees only when models are trained to satisfy Differential Privacy (DP).\nAlas, this is not the standard in industry as many companies use ad-hoc strategies to empirically evaluate privacy based on the statistical {\\em similarity} between synthetic and real data.\n\nIn this paper, we review the privacy metrics offered by leading companies in this space and shed light on a few critical flaws in reasoning about privacy entirely via empirical evaluations.\nWe analyze the undesirable properties of the metrics and filters they use and demonstrate their unreliability and inconsistency through counter-examples.\nWe then present a reconstruction attack, \\emph{ReconSyn}, which successfully recovers (i.e., leaks all the attributes of) at least 78\\% of the low-density train records (or outliers) with only black-box access to a single fitted generative model and the privacy metrics.\nFinally, we show that applying DP or using generators with low utility does not successfully mitigate \\emph{ReconSyn} as the privacy leakage still comes from access to the metrics.\nOverall, our work serves as a warning to practitioners not to deviate from established privacy-preserving mechanisms."},"_bibtex":{"value":"@misc{\nganev2024on,\ntitle={On the Inadequacy of Similarity-based Privacy Metrics: Reconstruction Attacks against ``Truly Anonymous Synthetic Data''},\nauthor={Georgi Ganev and Emiliano De Cristofaro},\nyear={2024},\nurl={https://openreview.net/forum?id=g16vmAtJ8x}\n}"},"title":{"value":"On the Inadequacy of Similarity-based Privacy Metrics: Reconstruction Attacks against ``Truly Anonymous Synthetic Data''"},"pdf":{"value":"/pdf/f19b498393f0fdb577bfdaf510d64158672a15ba.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"ganev|on_the_inadequacy_of_similaritybased_privacy_metrics_reconstruction_attacks_against_``truly_anonymous_synthetic_data"},"authorids":{"value":["~Georgi_Ganev1","~Emiliano_De_Cristofaro1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Georgi Ganev","Emiliano De Cristofaro"]}},"version":2},{"content":{"summary":{"value":"This paper develop a generic framework for defining functions that are equiv-\nariant under the action of classical Lie groups acting diagonally on tensors.\nThe groups considered include the orthogonal group O(d), the indefinite orthogonal group O(s,k− s), and the symplectic group Sp(d) and other general group actions. The main theoretical contribution is a characterization of\nO(d)-equivariant polynomial functions mapping multiple tensor inputs to tensor outputs. The authors prove that any such functions can be expressed as linear\ncombination of tensor products of the inputs with O(d) isotropic tensors. An\nimportant result is Theorem 1, which provides an explicit parameterization\nof O(d)-equivariant polynomial functions. The authors also present a practical\ncorollary (Corollary 1) for the case where the inputs are vectors and the out-\nputs are vector spaces, showing that the equivariant functions can be written\nas linear combinations of basis elements formed by permutations of the input\nvectors. Furthermore, the author also present a generalization form of general\ntensors in Theorem 2 and Corollary 2.\n\nAs a proof of concept, They consider the challenge of recovering a planted\nsparse vector from a set of vectors forming an orthonormal basis of a sub-\nspace. By designing a machine learning model that learns an equivariant 2-\ntensor (which is the covariance matrix) from data, they use the top eigenvector\nof this tensor as an estimator for the sparse vector. The numerical experiments\ndemonstrate that the learned algorithms outperform state-of-the-art methods.\nThe models adapt effectively to various noise structures and data sampling\nmethods.\n\nIn physics, Lorentz group O(1,3) is a special case of the equivariant group.\nParticularly in general relativity, tensors are used to describe physical quantities such as the curvature of spacetime, energy-momentum distributions. The\nauthors’ work generally present one of the proof for how to comprehend this\nstructure in the regime of mathematical tensor products. they also provide a\nframework for building physics informed machine learning models that inherently respect Lorentz symmetry and applying for advanced physics studies."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"• The statement of Theorem 1 introduces a complex expression involving\nlinear combinations over isotropic tensors. Could the authors provide more\nintuitive explanations or mathematical nature of this theorem?\n\n• The authors state that parameterizing all permutation-invariant polyno-\nmial functions may be as challenging as solving the graph isomorphism\nproblem. Could the authors provide an estimate of the computational\ncomplexity or approximate methods for tackling this problem? can we\npush this into high-dimensional settings?\n\n• While MLPs are popular in approximating the polynomial functions, is\nthere a risk that they might not fully capture the necessary polynomial\nstructures or equivariance properties?\n\n• The learned SVH models perform better when stringent data assumptions\nare not met. Could the authors elaborate on why their models are more\nrobust in these scenarios? Please provide valuable insights.\n\n• The paper focuses on equivariance. Could you derive some examples for\nother groups, such as unitary groups or affine transformations?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":4},"strengths":{"value":"## Originality\nThe paper presents a novel and significant advancement in the field of equiv-\nariant machine learning. Additionally, it uniquely incorporates the indefinite\northogonal group O(s,k− s) and the symplectic group Sp(d), which are funda-\nmental in physics and other scientific domains but have been less explored in the\ncontext of physics informed machine learning. The application of these theo-\nretical developments to the sparse vector estimation problem further showcases\noriginality by demonstrating how algorithms can outperform state-of-the-art\nmethods in regimes not previously addressed.\n\n## Quality\nThe paper is of high quality, offering rigorous mathematical formulations and\nproofs that underpin the proposed methods. The experiments are well-designed,\nand the empirical results convincingly demonstrate the effectiveness.\n\n## Significance\nThe significance of the work is substantial. By providing explicit parameteriza-\ntions for equivariant functions of tensor inputs and outputs, the paper equips\nresearchers and practitioners with powerful tools. The successful application to\nsparse vector es"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"## Limited Scope of Experiments\nOne of the main weaknesses of the paper is the limited scope of the experimental\nevaluation. The experiments are primarily focused on the sparse vector estima-\ntion problem, which, while important, represents a rather narrow application\ndomain. The method should have broader applications, such as in physics and\ngeneral relativity, but these areas are not sufficiently discussed or explored.\n\n## Clarity\nThe paper is difficult to follow due to disorganized notation and unclear variable\ndefinitions. The use of coordinates is somewhat messy, which makes the math-\nematical developments hard to track. Variables are often introduced without\nproper explanation or context, causing confusion. For example:\n\n• In Definition 4, the indices need more explanation.\n\n• In Theorem 1, how each tensor alk contracts with certain indices (dimen-\nsions) of cl1 ,l2 ,...,lr should be explained more clearly.\n\n• In Example 1, there should be more explanation of why only the generic\nelements of the G4 group are used while others are contracted.\n\n• In Lemma 1, the meaning and constraints of ασ need clarification.\nMoreover, the progression from theoretical concepts to experimental appli-\ncations lacks smoothness. The presentation could be improved by organizing\nnotation more systematically and clearly defining all variables.\n\n## Comparative Analysis\nThere are other contemporary techniques for sparse vector recovery and tensor\nanalysis that are not considered. The lack of comparison with a wider array\nof methods makes it difficult to fully assess the advantages of the proposed\napproach."}},"nonreaders":[],"tmdate":1731428206528,"tcdate":1730172541758,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12216/Reviewer_dQUR"],"signatures":["ICLR.cc/2025/Conference/Submission12216/Reviewer_dQUR"],"forum":"kyVzYpDxHg","number":2,"license":"CC BY 4.0","cdate":1730172541758,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12216/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428206528,"domain":"ICLR.cc/2025/Conference","replyto":"kyVzYpDxHg","id":"DOWoQI3KZt","forumContent":{"TLDR":{"value":"We provide a characterization for equivariant tensor polynomials and use that for a machine learning approach to solving the sparse vector recovery problem in new settings."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["equivariant machine learning","tensors","sparse vector recovery"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This work characterizes equivariant polynomial functions from tuples of tensor inputs to tensor outputs. Loosely motivated by physics, we focus on equivariant functions with respect to the diagonal action of the orthogonal group on tensors. We show how to extend this characterization to other linear algebraic groups, including the Lorentz and symplectic groups. \n\nOur goal behind these characterizations is to define equivariant machine learning models. In particular, we focus on the sparse vector estimation problem. This problem has been broadly studied in the theoretical computer science literature, and explicit spectral methods, derived by techniques from sum-of-squares, can be shown to recover sparse vectors under certain assumptions. Our numerical results show that the proposed equivariant machine learning models can learn spectral methods that outperform the best theoretically known spectral methods in some regimes. The experiments also suggest that learned spectral methods can solve the problem in settings that have yet to be theoretically analyzed.\n\nThis is an example of a promising direction in which theory can inform machine learning models and machine learning models can inform theory."},"_bibtex":{"value":"@misc{\ngregory2025learning,\ntitle={Learning equivariant tensor functions with applications to sparse vector recovery},\nauthor={Wilson G. Gregory and Josu{\\'e} Tonelli-Cueto and Nicholas F. Marshall and Andrew S. Lee and Soledad Villar},\nyear={2025},\nurl={https://openreview.net/forum?id=kyVzYpDxHg}\n}"},"title":{"value":"Learning equivariant tensor functions with applications to sparse vector recovery"},"pdf":{"value":"/pdf/ea9caf2ec53dd9a01b9ca78d22ac86e0a1c01ee1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"gregory|learning_equivariant_tensor_functions_with_applications_to_sparse_vector_recovery"},"authorids":{"value":["~Wilson_G._Gregory1","~Josué_Tonelli-Cueto1","~Nicholas_F._Marshall1","~Andrew_S._Lee1","~Soledad_Villar2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wilson G. Gregory","Josué Tonelli-Cueto","Nicholas F. Marshall","Andrew S. Lee","Soledad Villar"]}},"version":2},{"content":{"summary":{"value":"A paper proposing an original solution to an important computational physics problem, namely computing free energy changes in non-equilibrium thermodynamics. The proposed solution uses flow matching techniques in an innovative way to suggest a more efficient/ accurate algorithm to sample Monte Carlo trajectories over which to estimate such free energy changes."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Is it conceivable that your algorithmic improvements to learning the escorting protocol could be extended to other common ML scenarios where distributions are to be matched, e.g. diffusion models?\n- Your approach is a practical solution to minimising the bias/ variance of the Monte-Carlo Jaczinsky estimator, does it come with guarantees/ special scenarios in which it could turn out to be exact?\n- I didn't get much rationale for the choice of flow architecture etc, which potentially could be a factor in determining efficiency gains. Did you just take an off-the-shelf approach or are there mileage in optimising that side?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"I really enjoyed reading this paper, albeit as a somewhat knowledgeable outsider. The authors make a good effort to provide a comprehensive background accessible to non-specialists (although I wonder how well it fares with non-physicists). The idea of using ML to minimise the Jaczinsky lower bound is elegant and the proposed improvements in terms of adopting larger time steps/ extending to multi-state estimation are potentially important for the community. The empirical validation is well carried out albeit not extensive (similar in that to physics papers)"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The main weakness to me is how much this paper could appeal outside of the computational physics community, which is but a (small) strand in the ICLR community. Some opportunities to broaden the appeal of the paper are listed below in the questions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942067065,"tcdate":1761465973439,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22109/Reviewer_X5Rb"],"signatures":["ICLR.cc/2026/Conference/Submission22109/Reviewer_X5Rb"],"forum":"Da8PJXp0js","number":1,"license":"CC BY 4.0","cdate":1761465973439,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22109/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942067065,"domain":"ICLR.cc/2026/Conference","replyto":"Da8PJXp0js","id":"Dj9c4APXBC","forumContent":{"TLDR":{"value":"Conditional Flow Networks and Density Matching for Multistate Escorted Free-Energy Estimation"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Free-Energy","Jarzynski Equality","Crooks Fluctuation Theorem","Non-Equilibrium","Transport","Thermodynamics","Stochastic Thermodynamics","Flow Matching"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Estimating relative free energy differences between multiple thermodynamic states lies at the core of numerous problems in computational biochemistry. Traditional estimators, such as Free Energy Perturbation and its non-equilibrium counterpart based on the Jarzynski equality, rely on defining a switching protocol between thermodynamic states and computing the free energy difference from the work performed during this process. In this work, we present a method for learning such switching protocols within the class of escorted protocols that combine deterministic and stochastic steps. For this purpose, we use Conditional Flow Matching, and  introduce Conditional Density Matching (CDM)  for the purpose of estimating the change in Free-Energy. We further reduce the variance in the multistate setting by coupling multiple flows between thermodynamic states into a Flow Graph, enforcing estimator consistency across different transition paths."},"_bibtex":{"value":"@inproceedings{\nholdijk2026learning,\ntitle={Learning Escorted Protocols For Multistate Free-Energy Estimation},\nauthor={Lars Holdijk and Nithishwer Mouroug Anand and Michael M. Bronstein and Max Welling},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Da8PJXp0js}\n}"},"title":{"value":"Learning Escorted Protocols For Multistate Free-Energy Estimation"},"pdf":{"value":"/pdf/61ab6b00dfebb9323d9270d07ea030531d8e4cdb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"holdijk|learning_escorted_protocols_for_multistate_freeenergy_estimation"},"authorids":{"value":["~Lars_Holdijk1","~Nithishwer_Mouroug_Anand1","~Michael_M._Bronstein1","~Max_Welling1"]},"authors":{"value":["Lars Holdijk","Nithishwer Mouroug Anand","Michael M. Bronstein","Max Welling"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Aletheia, a benchmark dataset designed for learning neural PDE solvers in the context of 3D NDT, specifically focusing on multi-frequency ECPT. ECPT involves coupled electromagnetic (Maxwell's equations) and thermal (Heat equation) physics, used to detect subsurface defects in conductive materials like rails.\n\nThe key contributions are:\n\n1. A High-Fidelity 3D Dataset: Aletheia comprises over 4,700 simulations generated using COMSOL, covering six distinct types of internal rail defects. It provides time-resolved volumetric heat source (q) and temperature (u) fields on both regular and irregular grids.\n\n2. Multi-Frequency Data: To address the ill-posedness of the inverse heat conduction problem, the dataset includes 10 excitation frequencies (1-100 kHz), leveraging the electromagnetic skin-depth effect to probe different material depths.\n\n3. Real-World Calibration: The simulations are calibrated using real infrared thermography data from physical rail specimens.\n\n4. Benchmark Suite: The authors define eight tasks spanning in-distribution and Out-of-Distribution (OOD, based on frequency shifts). These include forward modeling (Q2T) and challenging inverse tasks, notably Surface-to-Source reconstruction (S2Q).\n\n5. Baseline Evaluations: Several neural operators (FNO variants, Transolver, LNO, etc.) are benchmarked on a subset of the data.\n\nAletheia aims to bridge the gap between academic PDE benchmarks and realistic, 3D, multi-physics inverse problems relevant to industrial applications."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Q1: Why was the empirical evaluation restricted to only the Type I double-layer subset (600 samples)? The validity of the benchmark relies on evaluating the models across the diverse range of defect types provided. Can you provide results utilizing the full dataset or, at minimum, results for the other five defect types?\n\nQ2: a) How does the performance of the models on the S2Q task degrade when realistic levels of sensor noise are added to the surface temperature input? b) Were models trained on the synthetic data tested on the real experimental measurements? If so, what were the results regarding the sim-to-real gap?\n\nQ3: Given the relatively poor performance in deriving defect parameters from the reconstructed heat field Q (Table 7, Appendix D, RMSE ≈ 1mm for depth), how do you justify the use of Q reconstruction as the primary benchmark task for NDT? Is this accuracy sufficient for practical rail inspection?\n\nQ4: a) What is the justification for the aggressive downsampling to 8000 points and the use of batch size 1? b) Can you provide an experiment quantifying the improvement gained by using multi-frequency data compared to single-frequency data for the inverse tasks?\n\nQ5: The dataset size is 1.89 TB. What is the concrete strategy for hosting the data to ensure long-term, accessible availability, including mechanisms for downloading specific subsets?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The primary strength of this work lies in the dataset generation effort, which addresses significant limitations in the existing landscape of SciML benchmarks.\n\n1. Significance and Relevance: The creation of Aletheia is a substantial effort. It moves beyond standard academic benchmarks (e.g., Darcy Flow) towards a complex, industrially relevant NDT problem. As highlighted in Table 1, it uniquely combines 3D geometries, inverse problems, and partial observations in a multi-physics setting (coupled electromagnetic and thermal PDEs).\n\n2. Rigorous Data Generation Methodology: The approach of using high-fidelity multi-physics simulations (COMSOL) and, crucially, calibrating them with real experimental data (Section 3.1) is commendable. The inclusion of both regular and irregular grids is valuable for testing mesh invariance.\n\n3. Physically Motivated Design: The inclusion of multi-frequency excitations (1-100 kHz) is well-motivated. It leverages the electromagnetic skin-depth effect to provide depth-sensitive information, which is essential for mitigating the ill-posedness of the inverse thermal problem (Figure 1)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the dataset itself is a valuable contribution, the paper, presented as a benchmark study, suffers from significant weaknesses in its experimental design and analysis, which undermine the benchmark's utility and the conclusions drawn.\n\n1. Severely Limited Scope of Benchmark Evaluation (Critical Flaw): The empirical evaluation is conducted on a very small subset of the data. Section 4.2 states that experiments involved 600 samples from the \"Type I double-layer defect simulations.\" This represents only ~12.5% of the total 4,782 samples and only 1 out of the 6 defect types. A benchmark paper must establish baselines across the diversity of the data it introduces. The conclusions drawn regarding the relative performance of neural operators are therefore not substantiated for the vast majority of the dataset, particularly the more complex defect types (e.g., multi-layer or closed cracks).\n\n2. Failure to Evaluate Under Realistic Conditions (Noise and Sim-to-Real): The abstract claims the dataset addresses challenges involving \"sparse and noisy boundary observations.\" While the S2Q task addresses sparsity, the benchmark entirely ignores noise.\nReal infrared data is noisy. Robustness to sensor noise is crucial for ill-posed inverse problems but is not evaluated, and while the real data is used for calibration, the models trained on simulations are never evaluated on the real experimental data. Assessing the sim-to-real gap is essential for any simulation-based benchmark intended for real-world application.\n\n3. Weak Link Between Benchmark Tasks and NDT Goals: The primary inverse tasks (T2Q, S2Q) focus on reconstructing the heat source field q(x,t). The ultimate goal of NDT is defect characterization (geometry, depth). Appendix D attempts to bridge this gap by regressing defect parameters from q. However, the results in Table 7 are weak. For instance, the RMSE for crack depth is approximately 1.0 mm (MAE 0.70-0.91). Given that the defects themselves range from 0.2mm to ~4mm in depth (Table 6), an error of 1mm is very large relative to the defect scale. This casts doubt on whether optimizing for q reconstruction MSE is sufficient for the intended NDT application.\n\n4. Questionable Experimental Methodology and Data Usage Mismatch: The high-resolution unstructured data (50,000 points, Appendix B) is aggressively downsampled to 8000 points for the experiments (Sec 4.2). This may discard the fine-grained details necessary for accurate defect reconstruction. Also, the use of a batch size of 1 (Table 10) for all models is highly unusual for training large models and can lead to unstable training and poor generalization, potentially affecting the validity of the comparisons. Finally, while the motivation for multi-frequency data is clear (Fig 1), the benchmark does not empirically quantify this benefit. A crucial missing experiment is a comparison of inverse reconstruction performance using single-frequency vs. multi-frequency data.\n\n5. Limited Scope of OOD Generalization: The OOD tasks are exclusively focused on unseen frequencies. In NDT, generalization to unseen defect morphologies (Geometric OOD) is critical but is not evaluated using the diverse defect types available."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926809274,"tcdate":1761713581807,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16764/Reviewer_Cusi"],"signatures":["ICLR.cc/2026/Conference/Submission16764/Reviewer_Cusi"],"forum":"iexhzXxLV5","number":2,"license":"CC BY 4.0","cdate":1761713581807,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16764/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926809274,"domain":"ICLR.cc/2026/Conference","replyto":"iexhzXxLV5","id":"qXeAtu4KvY","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Partial Differential Equations","Nondestructive Testing","Neural Operators","Thermal holography"]},"supplementary_material":{"value":"/attachment/1eec09edc58baa180f13b5d35158c66b5a0ffcd0.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Learning neural solvers for spatiotemporal partial differential equations (PDEs) under real-world constraints remains a key challenge in scientific machine learning, especially for inverse tasks with sparse and noisy boundary observations. We present the **Aletheia** dataset, the first 3D benchmark for learning data-driven solvers in the context of **nondestructive testing (NDT)**. The dataset simulates eddy-current-induced heating in conductive solids and models the resulting transient heat propagation governed by the heat equation. Aletheia contains over 4,700 high-resolution samples across 10 excitation frequencies (1-100\\,kHz), each providing volumetric heat source and temperature fields over time. It supports both forward prediction of temperature evolution and inverse reconstruction of internal heat sources or defects from surface infrared measurements. Real infrared thermography data from cracked rail specimens are included for calibration and generalization studies. We define three canonical tasks on both regular and irregular grids and benchmark them using various neural operators. Aletheia establishes a unified platform for evaluating neural PDE solvers under realistic NDT conditions, enabling progress in reliable, data-driven inverse modeling."},"_bibtex":{"value":"@misc{\nsun2026aletheia,\ntitle={{ALETHEIA}: A Multi-Frequency Eddy Current Pulsed Thermography Dataset for Neural Operator Learning in Nondestructive Testing},\nauthor={Changbin Sun and Xiaojie and Xiaotian Chen and Yuankai Wu},\nyear={2026},\nurl={https://openreview.net/forum?id=iexhzXxLV5}\n}"},"title":{"value":"ALETHEIA: A Multi-Frequency Eddy Current Pulsed Thermography Dataset for Neural Operator Learning in Nondestructive Testing"},"pdf":{"value":"/pdf/1f5abfb1ff28c486737af9eb5fe73af2a4dac82a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sun|aletheia_a_multifrequency_eddy_current_pulsed_thermography_dataset_for_neural_operator_learning_in_nondestructive_testing"},"authorids":{"value":["~Changbin_Sun1","~Xiaojie1","~Xiaotian_Chen5","~Yuankai_Wu2"]},"authors":{"value":["Changbin Sun","Xiaojie","Xiaotian Chen","Yuankai Wu"]}},"version":2},{"content":{"comment":{"value":"Dear reviewer,\n\nWe are very grateful for your feedback and we would like to give more clarifications in complement to our last response.\n\n> A. Generality and task adaptability.\n\nWe acknowledge the significance of your point regarding the landscape of gradient optimization. It's crucial to note that even in scenarios involving contact-making and breaking, the trajectories remain differentiable, as demonstrated by frameworks such as PlasticineLab[1] and DiffTaichi[2], which effectively handle contact-rich tasks using differentiable physics. Antonova et al. [3] have highlighted the actual challenges of gradient-based trajectory optimization, particularly its susceptibility to rugged loss landscapes and local optima. In response, our method utilizes LLM to decompose complex tasks into shorter, more manageable substages. This decomposition simplifies the optimization landscape, making the trajectory optimization process smoother and more reliable, as ablated in Section 4.3. As you said, our method benefits from a favorable optimization landscape and should be taken as our advantage. \n\n[1] Huang, Zhiao, Yuanming Hu, Tao Du, Siyuan Zhou, Hao Su, Joshua B. Tenenbaum, and Chuang Gan. \"Plasticinelab: A soft-body manipulation benchmark with differentiable physics.\" arXiv preprint arXiv:2104.03311 (2021).\n\n[2] Hu, Yuanming, Luke Anderson, Tzu-Mao Li, Qi Sun, Nathan Carr, Jonathan Ragan-Kelley, and Frédo Durand. \"Difftaichi: Differentiable programming for physical simulation.\" arXiv preprint arXiv:1910.00935 (2019).\n\n[3] Antonova, Rika, Jingyun Yang, Krishna Murthy Jatavallabhula, and Jeannette Bohg. \"Rethinking optimization with differentiable simulation from a global perspective.\" In Conference on Robot Learning, pp. 276-286. PMLR, 2023.\n\nFurthermore, we argue that while learning-based methods like PASTA are valuable, they inherently struggle with tasks that fall outside their training data, limiting their applicability to new, complex tasks such as doughnut making. The sim-to-real gap further exacerbates this limitation, as their learned policies often fail to transfer effectively to the real world. This is also evidenced in both visualizations and quantitative comparisons in our paper. \n\n> B. Interpretability and quality of the results\n\n In our efforts to enhance interpretability, we have updated our supplementary materials to include videos and visualizations with overlaid subgoals. These improvements aim to provide a clearer and more intuitive understanding of how our planning framework guides the robotic manipulations towards the desired outcomes. We invite you to review these enhancements in our updated video and the visual representations found in Appendix A.9. As we explained in our previous response, PASTA uses heuristic policies (they are detailed in Section 3.1 of their appendix, one example: *The roll policy first moves the roller down to make contact with the dough. Then, based on the goal component’s length, the policy calculates the distance it needs to move the roller back and forth when making contact with the dough*.). While this can produce visually appealing results, it does not reflect the adaptive, generalized problem-solving our method aims to achieve. Our results may not look as perfect as theirs since our actions are directly derived from our optimizations rather than heuristics.\n\nWe hope that these updates address your concerns comprehensively. Your feedback has been invaluable in this process, and we look forward to any additional comments you may have."},"title":{"value":"More Clarifications (part 2)"}},"tmdate":1700624943391,"tcdate":1700624919975,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2006/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission2006/Authors"],"forum":"iTsHStJKcm","number":19,"license":"CC BY 4.0","cdate":1700624919975,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2006/-/Official_Comment"],"mdate":1700624943391,"domain":"ICLR.cc/2024/Conference","replyto":"2ST3WDm460","id":"hb4wCKLf5W","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Deformable Object Manipulation; Large Language Models; Long-horizon Task Planning"]},"supplementary_material":{"value":"/attachment/a1e1137c84fc6219201117f9f2bf8921577e6cec.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Deformable object manipulation stands as one of the most captivating yet formidable challenges in robotics. While previous techniques have predominantly relied on learning latent dynamics through demonstrations, typically represented as either particles or images, there exists a pertinent limitation: acquiring suitable demonstrations, especially for long-horizon tasks, can be elusive. Moreover, basing learning entirely on demonstrations can hamper the model's ability to generalize beyond the demonstrated tasks. In this work, we introduce a demonstration-free hierarchical planning approach capable of tackling intricate long-horizon tasks without necessitating any training. We employ large language models (LLMs) to articulate a high-level, stage-by-stage plan corresponding to a specified task. For every individual stage, the LLM provides both the tool's name and the Python code to craft intermediate subgoal point clouds. With the tool and subgoal for a particular stage at our disposal, we present a granular closed-loop model predictive control strategy. This leverages Differentiable Physics with Point-to-Point correspondence (DiffPhysics-P2P) loss in the earth mover distance (EMD) space, applied iteratively. Experimental findings affirm that our technique surpasses multiple benchmarks in dough manipulation, spanning both short and long horizons. Remarkably, our model demonstrates robust generalization capabilities to novel and previously unencountered complex tasks without any preliminary demonstrations. We further substantiate our approach with experimental trials on real-world robotic platforms."},"_bibtex":{"value":"@misc{\nyou2024make,\ntitle={Make a Donut: Language-Guided Hierarchical {EMD}-Space Planning for Zero-shot Deformable Object Manipulation},\nauthor={Yang You and Bokui Shen and Congyue Deng and Haoran Geng and He Wang and Leonidas Guibas},\nyear={2024},\nurl={https://openreview.net/forum?id=iTsHStJKcm}\n}"},"title":{"value":"Make a Donut: Language-Guided Hierarchical EMD-Space Planning for Zero-shot Deformable Object Manipulation"},"pdf":{"value":"/pdf/b956a05bf3eedf4288b402bcc03b7845c2391f79.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"you|make_a_donut_languageguided_hierarchical_emdspace_planning_for_zeroshot_deformable_object_manipulation"},"authorids":{"value":["~Yang_You2","~Bokui_Shen1","~Congyue_Deng1","~Haoran_Geng1","~He_Wang5","~Leonidas_Guibas1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yang You","Bokui Shen","Congyue Deng","Haoran Geng","He Wang","Leonidas Guibas"]}},"version":2},{"content":{"venue":{"value":"CoRR 2019"},"pdf":{"value":"https://arxiv.org/pdf/1904.05298v1"},"venueid":{"value":"dblp.org/journals/CORR/2019"},"paperhash":{"value":"li|cnm_an_interpretable_complexvalued_network_for_matching"},"authorids":{"value":["~Qiuchi_Li1","https://dblp.org/search/pid/api?q=author:Benyou_Wang:","https://dblp.org/search/pid/api?q=author:Massimo_Melucci:"]},"html":{"value":"http://arxiv.org/abs/1904.05298"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-1904-05298,\n  publtype={informal},\n  author={Qiuchi Li and Benyou Wang and Massimo Melucci},\n  title={CNM: An Interpretable Complex-valued Network for Matching},\n  year={2019},\n  cdate={1546300800000},\n  journal={CoRR},\n  volume={abs/1904.05298},\n  url={http://arxiv.org/abs/1904.05298}\n}\n"},"abstract":{"value":"This paper seeks to model human language by the mathematical framework of quantum physics. With the well-designed mathematical formulations in quantum physics, this framework unifies different linguistic units in a single complex-valued vector space, e.g. words as particles in quantum states and sentences as mixed systems. A complex-valued network is built to implement this framework for semantic matching. With well-constrained complex-valued components, the network admits interpretations to explicit physical meanings. The proposed complex-valued network for matching (CNM) achieves comparable performances to strong CNN and RNN baselines on two benchmarking question answering (QA) datasets."},"title":{"value":"CNM: An Interpretable Complex-valued Network for Matching"},"authors":{"value":["Qiuchi Li","Benyou Wang","Massimo Melucci"]}},"tmdate":1768539515218,"pdate":1546300800000,"externalIds":["dblp:journals/corr/abs-1904-05298"],"tcdate":1768539503022,"writers":["~"],"signatures":["~Qiuchi_Li1"],"forum":"GkTjXGfDGh","license":"CC BY-SA 4.0","number":767222,"cdate":1546300800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768539515218,"domain":"DBLP.org","id":"GkTjXGfDGh","version":2},{"content":{"summary":{"value":"The paper proposes a new tensor network operator system structure search approach. Inspired by symmetry breaking in physics, they propose a two phase optimization procedure: First, they run normal tensor networ structure search, then they add a regularizer in the structure search optimization problem, which encourages asymmetric task-specific tensor network structures. The regularizer takes the form of a simple core tensor masking. They show that this formulation yields significantly more compact tensor representations in three distinct tensor network settings: Tensor network decomposition, parameter-efficient fine-tuning, and quantum circuit optimization."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Instead of the proposed regularizer, could one just directly add the number of parameters into the optimization problem to incentivize more efficient solutions?  \n\nOverall, I am currently unsure about the paper, in particular the relevance for an ML conference, but if the questions are addressed satisfactorily, I am willing to increase my score."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"Tensor networks seem to be a general enough formulation that this might have a lot of use cases, although I am a bit unsure about it, see the weaknesses section\n\nThe results of the proposed algorithm look convincing, consistently yielding good performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I have two concerns with the current paper:\n\n1.I find the paper quite inaccessible in its current form for readers not already familiar with tensor networks. A better motivation of tensor networks and the structure search setup, with an example formulation for an ML application, would be helpful at the beginning or at least in the appendix of the paper. Something of the form: A tensor network is an expression of the form einsum(A_ij,B_jk,C_kl), with A,B,C being called core tensors and could stand for ... in < application>. \n\nThe metaphor with the Higgs potential also seems unhelpful to me; I don't see how the Higgs potential maps to tensor networks or the structure search problem. In my opinion, as someone not very familiar with this topic, it did not aid my understanding, and the space would be better used for more intuitive motivation and problem setup of tensor network structure search in general, and how it can be useful for an ML practitioner. \n\n\n2.The method is specifically designed for high-order tensor networks. While I have no doubt that they are common in computational physics, I am unsure how common these forms of tensor networks are in ML specifically. Could you give some more examples where these methods could be useful in ML? \n\nPEFT for LLMs is given as an ML example, but also prefaced that it is not intended as a new practical PEFT method. Could you expand on what's missing for a practical algorithm?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359992046,"tcdate":1761079071680,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15111/Reviewer_ADcS"],"signatures":["ICLR.cc/2026/Conference/Submission15111/Reviewer_ADcS"],"forum":"DZ76Xr7zct","number":1,"license":"CC BY 4.0","cdate":1761079071680,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15111/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359992046,"domain":"ICLR.cc/2026/Conference","replyto":"DZ76Xr7zct","id":"UmP5xydLkB","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["tensor decomposition","tensor networks."]},"supplementary_material":{"value":"/attachment/7acfa1f6dac717ce28fede29f115c1154a7d237c.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Tensor networks (TNs) offer a compact representation for high-dimensional operators in physics and machine learning. While TN structure search (TN-SS) has advanced model selection, prior work is limited to a single operator. Yet real systems, such as transformers and quantum circuits, would contain multiple coupled operators, where treating them independently or enforcing a single shared structure is fundamentally limiting. We introduce joint TN-SS, the first framework for multi-operator structure search. Our physics-inspired algorithm runs in two phases: a symmetry phase, where standard TN-SS finds a shared structure capturing common inductive bias; and a symmetry-breaking phase, where operator-specific diversity emerges through greedy core masking, guided by task-explainable loss tolerances. Across tensor decomposition, parameter-efficient fine-tuning of LLMs, and quantum circuit optimization, joint TN-SS delivers more compact representations with equal or better accuracy than state-of-the-arts, with affordable search cost. These results demonstrate that symmetry-driven diversification offers a simple, general, and scalable solution to TN structure selection in multi-operator systems."},"_bibtex":{"value":"@misc{\nhe2026joint,\ntitle={Joint Structure Search for Tensor Network Operators Inspired by Symmetry Breaking},\nauthor={Yicong He and Chao Li and Yuchen Cong and Tomonori Shirakawa and Seiji Yunoki and Qibin Zhao},\nyear={2026},\nurl={https://openreview.net/forum?id=DZ76Xr7zct}\n}"},"title":{"value":"Joint Structure Search for Tensor Network Operators Inspired by Symmetry Breaking"},"pdf":{"value":"/pdf/4cedc3fae5c5614c621a968506b40a2d5ff8241b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"he|joint_structure_search_for_tensor_network_operators_inspired_by_symmetry_breaking"},"authorids":{"value":["~Yicong_He1","~Chao_Li12","~Yuchen_Cong1","~Tomonori_Shirakawa1","~Seiji_Yunoki1","~Qibin_Zhao1"]},"authors":{"value":["Yicong He","Chao Li","Yuchen Cong","Tomonori Shirakawa","Seiji Yunoki","Qibin Zhao"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Megalodon, which introduces three advancements over Mega: complex EMA, timestep normalization, and normalized attention. These advancements address the limitations of chunk-wise attention and architecture divergence across different tasks and data types. The new model is evaluated alongside Llama-2, both trained on the same public dataset, and demonstrates competitive and superior performance across a wide range of benchmarks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Have the authors conducted ablation studies for small-scale models before moving to 7B, similar to the results in Appendix C? Including these ablation studies, if already performed, would help readers understand how the three designs impact performance."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Clear motivations: All three improvements directly target the limitations of Mega.\n2. The complex EMA is a novel approach.\n3. The authors provide efficient parallelism."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Recent theoretical work [1] has shown that efficient versions of attention (like the chunk-based method used in this paper) can limit the expressiveness of the model, particularly for reasoning tasks that involve long-range information. The paper evaluates Megalodon with long-context open-book QA tasks. How will Megalodon perform on PhoneBook lookup [2] with ICL, especially for phonebook lengths longer than 4K tokens? How will it perform on complex reasoning tasks requiring long context, such as 8-shot or 16-shot math problems with GSM8K or coding tasks on HumanEval?\n\nEmpirically, it is not clear how CEMA improves expressiveness. It would be most direct to compare using CEMA versus using EMA on these tasks.\n\nThe reviewer understands that the rebuttal period is short and is therefore not requiring most of these experiments to be added.\n\n[1] Yang, Kai, et al. \"Do Efficient Transformers Really Save Computation?\" Forty-first International Conference on Machine Learning.\n\n[2] Jelassi, Samy, David Brandfonbrener, and Sham M. Kakade. \"Repeat After Me: Transformers are Better than State Space Models at Copying.\" Forty-first International Conference on Machine Learning."},"limitations":{"value":"The paper does not discuss its limitations. The authors believe there is no negative societal impact. In the paper checklist, justifications are required for answers marked \"Yes,\" but the authors have deleted the justification for several items. For limitations, the authors claim they are discussed, but there is no justification provided.\n\nThe anonymous link is not working."}},"nonreaders":[],"tmdate":1730878960481,"tcdate":1721210458745,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission4897/Reviewer_U1i1"],"signatures":["NeurIPS.cc/2024/Conference/Submission4897/Reviewer_U1i1"],"forum":"XlAbMZu4Bo","number":4,"license":"CC BY 4.0","cdate":1721210458745,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission4897/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878960481,"domain":"NeurIPS.cc/2024/Conference","replyto":"XlAbMZu4Bo","id":"7RjB2wmrvl","forumContent":{"TLDR":{"value":"Megalodon: Efficient Long-Context LLM Pretraining and Inference with Unlimited Context Length"},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Mega","Efficient Architecture","Long Sequence Modeling","Unlimited Context Length"]},"primary_area":{"value":"deep_learning_architectures"},"abstract":{"value":"The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirically underperform Transformers in pretraining efficiency and downstream task accuracy. We introduce MEGALODON, an neural architecture for efficient sequence modeling with unlimited context length. MEGALODON inherits the architecture of MEGA (exponential moving average with gated attention), and further introduces multiple technical components to improve its capability and stability, including complex exponential moving average (CEMA), timestep normalization layer, normalized attention mechanism and pre-norm with two-hop residual configuration. In a controlled head-to-head comparison with LLAMA2, MEGALODON achieves better efficiency than Transformer in the scale of 7 billion parameters and 2 trillion training tokens. MEGALODON reaches a training loss of 1.70, landing mid-way between LLAMA2-7B (1.75) and LLAMA2-13B (1.67). This result is robust throughout a wide range of benchmarks, where MEGALODON consistently outperforms Transformers across different tasks, domains, and modalities."},"_bibtex":{"value":"@inproceedings{\nma2024megalodon,\ntitle={Megalodon: Efficient {LLM} Pretraining and Inference with Unlimited Context Length},\nauthor={Xuezhe Ma and Xiaomeng Yang and Wenhan Xiong and Beidi Chen and LILI YU and Hao Zhang and Jonathan May and Luke Zettlemoyer and Omer Levy and Chunting Zhou},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=XlAbMZu4Bo}\n}"},"title":{"value":"Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length"},"pdf":{"value":"/pdf/70aaca704207816c7c033948248607819f055288.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"ma|megalodon_efficient_llm_pretraining_and_inference_with_unlimited_context_length"},"authorids":{"value":["~Xuezhe_Ma1","~Xiaomeng_Yang1","~Wenhan_Xiong1","~Beidi_Chen1","~LILI_YU1","~Hao_Zhang2","~Jonathan_May1","~Luke_Zettlemoyer1","~Omer_Levy1","~Chunting_Zhou1"]},"authors":{"value":["Xuezhe Ma","Xiaomeng Yang","Wenhan Xiong","Beidi Chen","LILI YU","Hao Zhang","Jonathan May","Luke Zettlemoyer","Omer Levy","Chunting Zhou"]}},"version":2},{"content":{"summary":{"value":"This paper studies in-context learning with task descriptions. It uses a synthetic setup of linear regression, where the means from which the inputs are samples serve as the task description. Theory shows that a single-layer linear attention model can reach the optimal solution and the optimal solutions are characterized. Experiments show that a single-layer can indeed reach the optimal solution, but deeper models still perform better."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. It's quite puzzling to understanding the embedding matrix formulation of the problem. I'm used to seeing in-context examples as running examples in text, which are made of words that are embedded. But here, the embedding matrix itself is the object to learn. More specifically, given this embedding matrix, the objective is to learn the bottom-right entry, which contains the prediction y_query. What are the trainable parameters? Is E itself updated during training? How does this formulation relate to the usual ICL setup of running text? \n\n2. What are multiple tasks mentioned in equation 6 and does it make sense to take the expectation over all of them?  What is a task specifically?\n\n3. What are the different sequences near equation 7? Are there multiple sequences per task? One sequence per task? \n\n4. Can you motivate the initialization in page 4? any clearer motivation besides having two matrices have the same norm?\n\n5. I am possibly missing some background to understand this, but what do you mean by \"We run gradient flow on the population loss\"? How is the optimization done exactly? \n\n6. What is the significance of characterizing the optimal solutions in section 4? How should they be interpreted and what does it tell us about ICL with descriptions more broadly? \n\n7. If a single linear layer transformer can reach the optimal solution, then why do the experimental results show deeper models to perform better? Is it because of training difficulties with the single layer case? Is it just a sample complexity issue, with the single-layer model not having enough samples?\n\n8. Since there's no pre-training and fine-tuning going on here, it seems like all instances of \"pretraining\" could just be changed to \"training\"."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Important research question on the effect of task descriptions in ICL\n- Simple synthetic setup to study \n- Theory showing optimality of a specific model class"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The biggest issue I have is whether the synthetic setup is a good proxy to study ICL, and specifically task descriptors. I understand similar synthetic setups were used to study ICL. But why is giving the mean a good task description? How does it simulate the natural language case, for example the one cited from Brown et al.? \n- It would be helpful to explicitly highlight the insights drawn from the various lemmas and theorems throughout section 4. Unfortunately I had a hard time following this part and so I indicate this is a low-confidence review. \n- The experiments are a nice edition to the theory, but I'm confused about the single-layer being inferior to the deeper models. \n- See other comments and questions below."}},"nonreaders":[],"tmdate":1731428184731,"tcdate":1730023394778,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5317/Reviewer_UsGc"],"signatures":["ICLR.cc/2025/Conference/Submission5317/Reviewer_UsGc"],"forum":"lZNb1CVm5O","number":1,"license":"CC BY 4.0","cdate":1730023394778,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5317/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428184731,"domain":"ICLR.cc/2025/Conference","replyto":"lZNb1CVm5O","id":"HJ6JKUtQAH","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Transformer","in-context learning","linear regression","task descriptor"]},"primary_area":{"value":"learning theory"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large language models (LLMs) exhibit strong in-context learning (ICL) ability, which allows the model to make predictions on new examples based on the given prompt. Recently, a line of research (Von Oswald et al., 2023; Aky¨urek et al., 2023; Ahn et al., 2023; Mahankali et al., 2023; Zhang et al., 2024) considered ICL for a simple linear regression setting and showed that the forward pass of Transformers is simulating some variants of gradient descent (GD) algorithms on the in-context examples. In practice, the input prompt usually contains a task descriptor in addition to in-context examples. We investigate how the task description helps ICL in the linear regression setting. Consider a simple setting where the task descriptor describes the mean of input in linear regression. Our results show that gradient flow converges to a global minimum for a linear Transformer. At the global minimum, the Transformer learns to use the task descriptor effectively to improve its performance. Empirically, we verify our results by showing that the weights converge to the predicted global minimum and Transformers indeed perform better with task descriptors."},"_bibtex":{"value":"@inproceedings{\nhuang2025task,\ntitle={Task Descriptors Help Transformers Learn Linear Models In-Context},\nauthor={Ruomin Huang and Rong Ge},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=lZNb1CVm5O}\n}"},"title":{"value":"Task Descriptors Help Transformers Learn Linear Models In-Context"},"pdf":{"value":"/pdf/b63cd66ad856bd9f832f4c1b0649bc600513be78.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"huang|task_descriptors_help_transformers_learn_linear_models_incontext"},"authorids":{"value":["~Ruomin_Huang1","~Rong_Ge1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruomin Huang","Rong Ge"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (2) 2020"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-65351-4_42.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2020"},"paperhash":{"value":"abrahão|an_algorithmic_information_distortion_in_multidimensional_networks"},"authorids":{"value":["~Felipe_S._Abrahão1","https://dblp.org/search/pid/api?q=author:Klaus_Wehmuth:","~Hector_Zenil1","https://dblp.org/search/pid/api?q=author:Artur_Ziviani:"]},"html":{"value":"https://doi.org/10.1007/978-3-030-65351-4_42"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/AbrahaoWZZ20,\n  author={Felipe S. Abrahão and Klaus Wehmuth and Hector Zenil and Artur Ziviani},\n  title={An Algorithmic Information Distortion in Multidimensional Networks},\n  year={2020},\n  cdate={1577836800000},\n  pages={520-531},\n  url={https://doi.org/10.1007/978-3-030-65351-4_42},\n  booktitle={COMPLEX NETWORKS (2)},\n  crossref={conf/complexnetworks/2020-2}\n}\n"},"abstract":{"value":"Network complexity, network information content analysis, and lossless compressibility of graph representations have been played an important role in network analysis and network modeling. As multidimensional networks, such as time-varying, multilayer, or dynamic multilayer networks, gain more relevancy in network science, it becomes crucial to investigate in which situations universal algorithmic methods based on algorithmic information theory applied to graphs cannot be straightforwardly imported into the multidimensional case. In this direction, as a worst-case scenario of lossless compressibility distortion that increases linearly with the number of distinct dimensions, this article presents a counter-intuitive phenomenon that occurs when dealing with networks within non-uniform and sufficiently large multidimensional spaces. In particular, we demonstrate that the algorithmic information necessary to encode multidimensional networks that are isomorphic to logarithmically compressible monoplex networks may display exponentially larger distortions in the general case."},"title":{"value":"An Algorithmic Information Distortion in Multidimensional Networks"},"authors":{"value":["Felipe S. Abrahão","Klaus Wehmuth","Hector Zenil","Artur Ziviani"]}},"tmdate":1747011702861,"pdate":1577836800000,"tcdate":1746888761778,"writers":["~"],"signatures":["~Hector_Zenil1"],"forum":"bvT1xRVMFX","license":"CC BY-SA 4.0","number":415716,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747011702861,"domain":"DBLP.org","id":"bvT1xRVMFX","version":2},{"content":{"summary":{"value":"Large language models show strong capabilities in complex reasoning tasks. Training with difficult problems can further improve the model's performance. Based on this, the authors proposed a data synthesis method to generate more challenging problems. Specifically, they combined difficulty-aware graph sampling for prompts and difficulty-aware rejection fine-tuning to create high-difficulty training data. Training the model with this synthetic data led to some improvement in its performance."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"+ Please explain the selection strategy for the various models in Section 3.1.\n\n+ Please provide the performance of the problem generator on coding tasks.\n\n+ Why not use OCR-Qwen-7B-Instruct as the backbone model? It has stronger reasoning capabilities.\n\n+ How does QueST perform compared to TACO when the data volume is the same? If QueST cannot significantly outperform TACO, it cannot be concluded that QueST's problems are more difficult or of higher quality.\n\n+ The performance of the three models with RL in Table 4 shows no significant improvement and is unstable. Does this indicate that the quality of QueST's data is normal?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"+ The authors generated a set of training data related to coding tasks, totaling 100,000 examples.\n\n+ The authors trained the model using the synthetic data, which led to improvements."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"+ In Section 3.1 of the paper, the data synthesis process requires a large number of external models for assistance, but the authors do not explain the strategy for selecting these models, which makes it unconvincing.\n\n+ The authors cannot explain the reason for the low accuracy in the synthetic problems. On one hand, it may be due to higher difficulty, and on the other hand, it may be because the answers contain errors.\n\n+ How well does the problem generator perform on coding tasks? The paper does not explain this. If the trained model cannot outperform the problem generator, then the synthetic data is meaningless, especially since the authors claim in the abstract that \"QueST pushes the boundaries of reasoning abilities in large language models.\""}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941651040,"tcdate":1760935519091,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21237/Reviewer_HaRZ"],"signatures":["ICLR.cc/2026/Conference/Submission21237/Reviewer_HaRZ"],"forum":"HuFMP0R4DR","number":1,"license":"CC BY 4.0","cdate":1760935519091,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21237/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941651040,"domain":"ICLR.cc/2026/Conference","replyto":"HuFMP0R4DR","id":"cCaahsgRFk","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["reasoning","code generation","large language model"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large Language Models have achieved strong performance on reasoning tasks, solving competition-level coding and math problems. However, their scalability is limited by human-labeled datasets and the lack of large-scale, challenging coding problem training data. Existing competitive coding datasets contain only thousands to tens of thousands of problems. Previous synthetic data generation methods rely on either augmenting existing instruction datasets or selecting challenging problems from human-labeled data. In this paper, we propose QueST, a novel framework which combines difficulty-aware graph sampling for prompt and difficulty-aware rejection fine-tuning that directly optimizes specialized generators to create challenging coding problems. Our trained generators demonstrate superior capability at creating challenging problems compared to even proprietary models such as GPT-4o. We leverage this method to generate large-scale synthetic coding problems, which we then use to distill from long Chain-of-Thought (CoT) models or conduct reinforcement learning for smaller models, proving effective in both scenarios. Our distilled model achieves the best performance compared to similarly sized models trained on previous long CoT SFT datasets. By training generators to create more difficult problems, QueST pushes the boundaries of reasoning abilities in large language models."},"_bibtex":{"value":"@misc{\nhu2026quest,\ntitle={Que{ST}: Incentivizing {LLM}s to Generate Difficult Problems},\nauthor={Hanxu Hu and Xingxing Zhang and Jannis Vamvas and Rico Sennrich and Furu Wei},\nyear={2026},\nurl={https://openreview.net/forum?id=HuFMP0R4DR}\n}"},"title":{"value":"QueST: Incentivizing LLMs to Generate Difficult Problems"},"pdf":{"value":"/pdf/7d6156c70596fb3ecb230e49ca6d85758a638ac6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"hu|quest_incentivizing_llms_to_generate_difficult_problems"},"authorids":{"value":["~Hanxu_Hu1","~Xingxing_Zhang1","~Jannis_Vamvas1","~Rico_Sennrich1","~Furu_Wei1"]},"authors":{"value":["Hanxu Hu","Xingxing Zhang","Jannis Vamvas","Rico Sennrich","Furu Wei"]}},"version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2308.12939v1"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"fang|learning_only_on_boundaries_a_physicsinformed_neural_operator_for_solving_parametric_partial_differential_equations_in_complex_geometries"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zhiwei_Fang:","https://dblp.org/search/pid/api?q=author:Sifan_Wang:","~Paris_Perdikaris1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2308.12939"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2308-12939,\n  publtype={informal},\n  author={Zhiwei Fang and Sifan Wang and Paris Perdikaris},\n  title={Learning Only On Boundaries: a Physics-Informed Neural operator for Solving Parametric Partial Differential Equations in Complex Geometries},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2308.12939},\n  url={https://doi.org/10.48550/arXiv.2308.12939}\n}\n"},"abstract":{"value":"Recently deep learning surrogates and neural operators have shown promise in solving partial differential equations (PDEs). However, they often require a large amount of training data and are limited to bounded domains. In this work, we present a novel physics-informed neural operator method to solve parametrized boundary value problems without labeled data. By reformulating the PDEs into boundary integral equations (BIEs), we can train the operator network solely on the boundary of the domain. This approach reduces the number of required sample points from $O(N^d)$ to $O(N^{d-1})$, where $d$ is the domain's dimension, leading to a significant acceleration of the training process. Additionally, our method can handle unbounded problems, which are unattainable for existing physics-informed neural networks (PINNs) and neural operators. Our numerical experiments show the effectiveness of parametrized complex geometries and unbounded problems."},"title":{"value":"Learning Only On Boundaries: a Physics-Informed Neural operator for Solving Parametric Partial Differential Equations in Complex Geometries"},"authors":{"value":["Zhiwei Fang","Sifan Wang","Paris Perdikaris"]}},"tmdate":1727708546567,"pdate":1672531200000,"tcdate":1727708537970,"writers":["~"],"signatures":["~Paris_Perdikaris1"],"forum":"wGQBRfS97I","license":"CC BY-SA 4.0","number":119614,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727708546567,"domain":"DBLP.org","id":"wGQBRfS97I","version":2},{"content":{"summary":{"value":"The paper presents 3DPhysVideo, a training-free pipeline designed to generate physically realistic videos from a single image input. It reuses a pre-trained Image-to-Video (I2V) flow model across two stages.\n\nIn Stage 1: Single Image to 3D, the I2V model functions as a view synthesizer to reconstruct 360-degree 3D scene geometry.\n\nIn Stage 2: Simulation to Video, Material Point Method (MPM) physics simulation is applied to the geometry. The resulting simulated point trajectories, which support complex dynamics like fluids and viscous substances, then guide the same I2V model to synthesize the final photorealistic video.\n\nThe core mechanism, Consistency-Guided Flow SDE, adapts the I2V model for both 3D reconstruction and simulation-guided rendering."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.\tCould the authors elaborate on the empirical or theoretical rationale for entirely eliminating the denoising bias ?\n2.\tWhat is the measured reliability or accuracy of these automatically inferred physical parameters compared to manually specified inputs?\n3.\tSince the current SDE is heavily reliant on visual consistency, how would the core consistency metric and the model’s latent inputs need to be adapted or redefined to effectively enforce a non-visual inductive bias, such as alignment with a detailed text prompt, without requiring additional model training?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The 3DPhysVideo pipeline generates physically realistic videos from a single image using a training-free approach. It repurposes an off-the-shelf Image-to-Video (I2V) model in two stages.\n1. 3D Reconstruction: The I2V model first acts as a novel view synthesizer to reconstruct 360-degree 3D scene geometry.\n2. Physics Generation: The geometry undergoes Material Point Method (MPM) physics simulation. The resulting simulated dynamics then guide the same I2V model to synthesize the final photorealistic video.\nThis dual functionality is enabled by the Consistency-Guided Flow SDE, which adapts the pre-trained model for both geometry and dynamics synthesis. The method achieves good physical realism compared to baselines, especially in multi-object and fluid interaction scenarios, while offering user control over physical properties."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed method appears incremental, with limited distinction from prior work.\n\n2.Experiments are limited in scope; key baselines and datasets are missing.\n\n3. Core assumptions lack rigorous justification or mathematical support.\n\n4. Result interpretation is shallow; no discussion of failure cases or parameter sensitivity.\n\n5. Figures and explanations are sometimes unclear, reducing readability and impact."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915731385,"tcdate":1761933485028,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1309/Reviewer_DQZT"],"signatures":["ICLR.cc/2026/Conference/Submission1309/Reviewer_DQZT"],"forum":"8TgzLrWgrk","number":2,"license":"CC BY 4.0","cdate":1761933485028,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1309/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915731385,"domain":"ICLR.cc/2026/Conference","replyto":"8TgzLrWgrk","id":"pGo74UbIoB","forumContent":{"TLDR":{"value":"."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Generation","3D Reconstruction","Physically Plausible Video"]},"supplementary_material":{"value":"/attachment/92311b2b1f7b1e9d7f08aae7736b79c212281762.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in real-world physical dynamics. Recent works such as PhysGen3D tackle single image-to-3D physics through mesh reconstruction and Physically-Based Rendering, but challenges remain in modeling fluid dynamics and photorealism. This work introduces 3DPhysVideo, a novel training-free pipeline that generates physically realistic videos from a single image. We repurpose an off-the-shelf video model for two stages. First, we use it as a novel view synthesizer to reconstruct complete 360-degree 3D scene geometry by guiding the image-to-video (I2V) flow model with rendered point clouds derived from an initial 3D estimation. Second, after applying Material Point Method (MPM) physics simulation to this geometry, the simulated point cloud is used to guide the same I2V flow model to synthesize final, high-quality videos. Consistency-Guided Flow SDE, which decomposes the predicted velocity of the I2V flow model into denoising and consistency bias, allows us to effectively repurpose the model for both 3D reconstruction and simulation-guided video generation. Our method successfully bridges the gap from single-images to physically plausible videos  while remaining efficient to run on a single consumer gpu. In the extensive experiments, our approach outperforms state-of-the-art baselines on both GPT-based evaluations and VideoPhy physics-consistency benchmark, across diverse scenarios including single-object, multi-object, and fluid interaction sequences."},"_bibtex":{"value":"@misc{\nanonymous2026dphysvideo,\ntitle={3{DP}hysVideo: 3D Scene Reconstruction and Physical Animation Leveraging a Video Generation Model via Consistency-Guided Flow {SDE}},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=8TgzLrWgrk}\n}"},"title":{"value":"3DPhysVideo: 3D Scene Reconstruction and Physical Animation Leveraging a Video Generation Model via Consistency-Guided Flow SDE"},"pdf":{"value":"/pdf/1d37cb0c53318381e966555bf18592a09cb61d7a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"kim|3dphysvideo_3d_scene_reconstruction_and_physical_animation_leveraging_a_video_generation_model_via_consistencyguided_flow_sde"},"authorids":{"value":["~Hwidong_Kim1","~Yunho_Kim2","~Tae-Kyun_Kim2"]},"authors":{"value":["Hwidong Kim","Yunho Kim","Tae-Kyun Kim"]}},"version":2},{"content":{"comment":{"value":"> Re Q1.2: Can the authors give an example of the invariant feature of a REAL chaotic system?\n\n\nThe four features we adopted in the evaluation of our experiments have been widely used to characterize the chaotic systems in real applications, e.g. climate:\n\n1. The joint distribution of physics-informed summary statistics is a common choice for evaluating the quality of emulators for complex chaotic systems ([5]).\n\n2. The Fourier energy spectrum, which described the distribution of kinetic energy as a function of frequency, is a common high-dimensional invariant statistic used in the analysis of fluid systems ([6, 7]).\n\n3. The leading Lyapunov exponent (LLE) is a widely adopted measurement that characterizes how fast the dynamical states diverge with respect to time ([2, 8]). \n\n4. The fractal dimension ([1, 9]), as a measure of complexity, has been used to describe the spatiotemporal climate data.\n\n\n> The Lorenz-96 is a synthetic system with a known low-dimensional feature. If the invariant feature of a real chaotic system is too complex, or too hard to define, how to verify the model can be applied to a real system?\n\n\nOur proposed CL approach is exactly designed to deal with complex systems where there is no significant prior knowledge! While simulated, Lorenz-96 and Kuramoto–Sivashinsky represent *high-dimensional* chaotic attractors with origins as models for fluid turbulence, and our experiments show that our proposed approaches work well for training better emulators for such systems. We are not aware of a low-dimensional characterization of the Lorenz-96 or Kuramoto–Sivashinsky attractor.\n\n\n> Generally I kind of agree with Reviewer GF1X that '...to see how the methods perform on at least one empirical chaotic problem would have also been nice...'\n\n\nWe would also love to evaluate on empirical data! However, as we discussed in our response to reviewer GF1X, there is very little publicly available empirical data for systems as complex as the spatiotemporal chaos shown in our experiments. \nWe would love to see this change in the future. However, current spatiotemporal modeling papers are almost all evaluated on high-quality simulations. The one exception is weather data used for training emulators such as FourCastNet [10] and ClimaX [11], which are very large models that cost millions to train. We would love to see how well these approaches scale to such models but that is outside the scope of this work.\n\n\nThank you again for your time providing the feedback! If there are any further questions or details you’d like to discuss, we are here to assist. We look forward to hearing back from you.\n\n\n[1] Estimating the Dimensions of Weather and Climate Attractors. Fraedrich, Klaus. (1986)\n\n[2] Predicting uncertainty in forecasts of weather and climate, Palmer, T N. (2000)\n\n[3] Universal behavior of extreme value statistics for selected observables of dynamical systems. Lucarini, Valerio, et al. (2011)\n\n[4] A locally time-invariant metric for climate model ensemble predictions of extreme risk. Virdee, Mala, et al. (2023)\n\n[5] Multiscale Simulations of Complex Systems by Learning their Effective Dynamics. Vlachas, Pantelis R., et  al. (2021)\n\n[6] Characterization and prediction of runoff dynamics: a nonlinear dynamical view. Islam, M.N, Sivakumar, B. (2002)\n\n[7] Global energy spectrum of the general oceanic circulation. Storer, Benjamin A., et al. (2022)\n\n[8] Predictability of Weather and Climate. Krishnamurthy, V. (2019)\n\n[9] Estimating the Fractal Dimension and the Predictability of the Atmosphere. Zeng, X., et al. (1922)\n\n[10] FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators. Pathak, Jaideep et al. (2022)\n\n[11] ClimaX: A foundation model for weather and climate. Nguyen, Tung et al. (2023)\n"},"title":{"value":"Response to your questions (2/2)"}},"tmdate":1704272697663,"tcdate":1692299731991,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission9364/Authors"],"signatures":["NeurIPS.cc/2023/Conference/Submission9364/Authors"],"forum":"8xx0pyMOW1","number":8,"license":"CC BY 4.0","cdate":1692299731991,"mdate":1704272697663,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission9364/-/Official_Comment","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"ZlVQqemhel","id":"VkOLboBw9o","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Neural operators","contrastive learning","optimal transport","chaotic attractors","invariant measures"]},"supplementary_material":{"value":"/attachment/e7116ea62f1aaa9abeaef9c0014a17e33ceb316e.pdf"},"_bibtex":{"value":"@inproceedings{\njiang2023training,\ntitle={Training neural operators to preserve invariant measures of chaotic attractors},\nauthor={Ruoxi Jiang and Peter Y. Lu and Elena Orlova and Rebecca Willett},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=8xx0pyMOW1}\n}"},"title":{"value":"Training neural operators to preserve invariant measures of chaotic attractors"},"paperhash":{"value":"jiang|training_neural_operators_to_preserve_invariant_measures_of_chaotic_attractors"},"abstract":{"value":"Chaotic systems make long-horizon forecasts difficult because small perturbations in initial conditions cause trajectories to diverge at an exponential rate. In this setting, neural operators trained to minimize squared error losses, while capable of accurate short-term forecasts, often fail to reproduce statistical or structural properties of the dynamics over longer time horizons and can yield degenerate results. In this paper, we propose an alternative framework designed to preserve invariant measures of chaotic attractors that characterize the time-invariant statistical properties of the dynamics. Specifically, in the multi-environment setting (where each sample trajectory is governed by slightly different dynamics),  we consider two novel approaches to training with noisy data. First, we propose a loss based on the optimal transport distance between the observed dynamics and the neural operator outputs. This approach requires expert knowledge of the underlying physics to determine what statistical features should be included in the optimal transport loss. Second, we show that a  contrastive learning framework, which does not require any specialized prior knowledge, can preserve statistical properties of the dynamics nearly as well as the optimal transport approach. On a variety of chaotic systems, our method is shown empirically to preserve invariant measures of chaotic attractors."},"pdf":{"value":"/pdf/b14a5dfa81903acdb97fbc7b18430ff742f2e28f.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Ruoxi_Jiang1","~Peter_Y._Lu1","~Elena_Orlova1","~Rebecca_Willett1"]},"authors":{"value":["Ruoxi Jiang","Peter Y. Lu","Elena Orlova","Rebecca Willett"]}},"version":2},{"content":{"summary":{"value":"This paper studies whether the synthetic data can substitute the golden data when fine-tuning the generator in a RAG system. GPT-4o is applied to generated the synthetic HotpotQA like questions based on the same golden passage pairs as the HoptpotQA training set. Llama-3.1-8B-Instruct is selected as the generator and trained on golden training set and the synthetic training set separately. Metrics, like EM, F1, BERT Score and LLM-as-a-judge are utilized to compare the performances. The author concludes that the generator fine-tuned on synthetic data outperforms the counterpart fine-tuned on golden training set with the imperfect retriever, MPNet. In addition, the generator fine-tuned on the synthetic data has better generalization than its counter-part."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"The primary challenge in generating multi-hop questions lies in collecting the linked passages required to answer them. This paper bypasses that difficulty by directly using the same passage pairs as the HotpotQA dataset. Why not consider a more general synthetic data generation method?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. The author clearly describes the experimental setup, and the Figure 1 task diagram is helpful to understand the baselines.\n2. The synthetic QA dataset is released and may be helpful for the research in this area"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The synthetic questions are generated from the same golden passages as real HotpotQA, making the claims overestimated.\n2. The synthetic data generation and fine-tuning are only conducted on HotpotQA, raising the concerns about the generalization.\n3. The usefulness of synthetic data has been proved in many previous studies, and it is not supervise to observe the benefits of using the synthetic data to train the generator."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918643392,"tcdate":1762121965856,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6354/Reviewer_bDWW"],"signatures":["ICLR.cc/2026/Conference/Submission6354/Reviewer_bDWW"],"forum":"4UPTS9eA5E","number":4,"license":"CC BY 4.0","cdate":1762121965856,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6354/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918643392,"domain":"ICLR.cc/2026/Conference","replyto":"4UPTS9eA5E","id":"mrIF1HydDL","forumContent":{"TLDR":{"value":"We study the degree to which synthetic data can effectively substitute real data for RAG generator fine-tuning."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Question Answering","Synthetic Data","Retrieval Augmented Generation","Natural Language Processing"]},"supplementary_material":{"value":"/attachment/a6ad41acd5fa6e53cad0f6542b405079515544d0.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"To improve large language models (LLMs) for question answering (QA) tasks, system architects often look to retrieval-augmented generation (RAG) or fine-tuning approaches to increase a model's performance.  In many applications, however, there is a dearth of real data of sufficient quality to support model fine-tuning to improve RAG system performance for QA tasks.  In this work, we study the degree to which synthetic data can effectively substitute real data for RAG generator fine-tuning.  Using GPT-4o, we generate a synthetic version of the HotpotQA training set and fine-tune a Llama-3 generator separately on both real and synthetic data.  We evaluate our models with a range of metrics such as token-level F1, Bertscore, and LLM-as-a-judge. Across these metrics, model performance generally increases after fine-tuning primarily due to better conformity to the style of the answer distribution and secondarily due to improved use of retrieved contexts.  We observe that relative performance depends on the quality of the retriever, emphasizing the importance of the training data distribution in improving the model's reasoning over multiple contexts.  We further show that the fine-tuned model trained on synthetic data generalizes better to similar held-out QA tasks, outperforming an LLM fine-tuned on real data by 36% in LLM-judged correctness over the RepLiQA dataset.  These findings motivate a system-level analysis of the marginal benefits of generator fine-tuning in RAG pipelines, providing practical insights on the utility of synthetic training data for the benefit of both RAG systems engineers and future researchers."},"_bibtex":{"value":"@misc{\nturner2026improving,\ntitle={Improving {RAG} Question Answering Generation with Synthetic Data},\nauthor={Matthew J. Turner and Anhthy Ngo and Ben Wellner},\nyear={2026},\nurl={https://openreview.net/forum?id=4UPTS9eA5E}\n}"},"title":{"value":"Improving RAG Question Answering Generation with Synthetic Data"},"pdf":{"value":"/pdf/f954b48be8e44d426383b4273db785e09420e296.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"turner|improving_rag_question_answering_generation_with_synthetic_data"},"authorids":{"value":["~Matthew_J._Turner1","~Anhthy_Ngo1","~Ben_Wellner1"]},"authors":{"value":["Matthew J. Turner","Anhthy Ngo","Ben Wellner"]}},"version":2},{"content":{"comment":{"value":"**W.2.** The proposed method, while interesting, may lack sufficient complexity and novelty to meet ICLR's technical bar.\n\n**Answer:** We gently beg to differ. A complex architecture is likely not an ideal solution for this problem as complex architecture presents rigidity, which may be a barrier to model haphazard inputs. The simplicity of our architecture imparts agility and an elegant mechanism to scale up and down as needed. Our novel \"clean and intuitive design\" was also noted by reviewer 1. Also, our contribution to the \"dynamic handling of varying input dimensions through activation/deactivation of LSTMs\" is appreciated by reviewers 1 and 2. Further, the \"effective balance between local feature information and globally shared information\" in our architecture imparts the ability to learn without forgetting."},"title":{"value":"Reply to Weakness 2"}},"tmdate":1731782945355,"tcdate":1731782945355,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10272/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission10272/Authors"],"forum":"VzdycorGTt","number":20,"license":"CC BY 4.0","cdate":1731782945355,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10272/-/Official_Comment"],"mdate":1731782945355,"domain":"ICLR.cc/2025/Conference","replyto":"Q8ACARg375","id":"lxAzFz5Ynr","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Varying Input Dimension","Stremaing Data","Online Learning","Recurrent Neural Network","Catastrophic Forgetting"]},"supplementary_material":{"value":"/attachment/78273da7b324d0d5f8565fecfd45e0e91c261d6e.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We study the online learning problem characterized by the varying input feature space of streaming data. Although LSTMs have been employed to effectively capture the temporal nature of streaming data, they cannot handle the dimension-varying streams in an online learning setting. Therefore, we propose a dynamic LSTM-based novel method, packetLSTM, to model the dimension-varying streams. The packetLSTM's dynamic framework consists of an evolving packet of LSTMs, each dedicated to processing one input feature. Each LSTM retains the local information of its corresponding feature, while a shared common memory consolidates global information. This configuration facilitates continuous learning and mitigates the issue of forgetting, even when certain features are absent for extended time periods. The idea of utilizing one LSTM per feature coupled with a dimension-invariant operator for information aggregation enhances the dynamic nature of packetLSTM. This dynamic nature is evidenced by the model's ability to activate, deactivate, and add new LSTMs as required, thus seamlessly accommodating varying input dimensions. The packetLSTM achieves state-of-the-art results on five datasets, and its underlying principle is extended to other RNN types, like GRU and vanilla RNN."},"_bibtex":{"value":"@misc{\nagarwal2025packetlstm,\ntitle={packet{LSTM}: Dynamic {LSTM} Framework for Streaming Data with Varying Feature Space},\nauthor={Rohit Agarwal and Karaka Prasanth Naidu and Alexander Horsch and Krishna Agarwal and Dilip Prasad},\nyear={2025},\nurl={https://openreview.net/forum?id=VzdycorGTt}\n}"},"title":{"value":"packetLSTM: Dynamic LSTM Framework for Streaming Data with Varying Feature Space"},"pdf":{"value":"/pdf/f8818eb0b22a1e126eef200f0f993048776a4b2d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"agarwal|packetlstm_dynamic_lstm_framework_for_streaming_data_with_varying_feature_space"},"authorids":{"value":["~Rohit_Agarwal3","~Karaka_Prasanth_Naidu1","~Alexander_Horsch1","~Krishna_Agarwal1","~Dilip_K._Prasad1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rohit Agarwal","Karaka Prasanth Naidu","Alexander Horsch","Krishna Agarwal","Dilip Prasad"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"https://arxiv.org/pdf/2411.14951v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"li|morph_a_motionfree_physics_optimization_framework_for_human_motion_generation"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zhuo_Li:","~Mingshuang_Luo1","~Ruibing_Hou3","https://dblp.org/search/pid/api?q=author:Xin_Zhao:","https://dblp.org/search/pid/api?q=author:Hao_Liu:","~Hong_Chang1","https://dblp.org/search/pid/api?q=author:Zimo_Liu:","https://dblp.org/search/pid/api?q=author:Chen_Li:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2411.14951"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2411-14951,\n  publtype={informal},\n  author={Zhuo Li and Mingshuang Luo and Ruibing Hou and Xin Zhao and Hao Liu and Hong Chang and Zimo Liu and Chen Li},\n  title={Morph: A Motion-free Physics Optimization Framework for Human Motion Generation},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2411.14951},\n  url={https://doi.org/10.48550/arXiv.2411.14951}\n}\n"},"abstract":{"value":"Human motion generation has been widely studied due to its crucial role in areas such as digital humans and humanoid robot control. However, many current motion generation approaches disregard physics constraints, frequently resulting in physically implausible motions with pronounced artifacts such as floating and foot sliding. Meanwhile, training an effective motion physics optimizer with noisy motion data remains largely unexplored. In this paper, we propose \\textbf{Morph}, a \\textbf{Mo}tion-F\\textbf{r}ee \\textbf{ph}ysics optimization framework, consisting of a Motion Generator and a Motion Physics Refinement module, for enhancing physical plausibility without relying on expensive real-world motion data. Specifically, the motion generator is responsible for providing large-scale synthetic, noisy motion data, while the motion physics refinement module utilizes these synthetic data to learn a motion imitator within a physics simulator, enforcing physical constraints to project the noisy motions into a physically-plausible space. Additionally, we introduce a prior reward module to enhance the stability of the physics optimization process and generate smoother and more stable motions. These physically refined motions are then used to fine-tune the motion generator, further enhancing its capability. This collaborative training paradigm enables mutual enhancement between the motion generator and the motion physics refinement module, significantly improving practicality and robustness in real-world applications. Experiments on both text-to-motion and music-to-dance generation tasks demonstrate that our framework achieves state-of-the-art motion quality while improving physical plausibility drastically."},"title":{"value":"Morph: A Motion-free Physics Optimization Framework for Human Motion Generation"},"authors":{"value":["Zhuo Li","Mingshuang Luo","Ruibing Hou","Xin Zhao","Hao Liu","Hong Chang","Zimo Liu","Chen Li"]}},"tmdate":1784696954235,"pdate":1735603200000,"externalIds":["dblp:journals/corr/abs-2411-14951"],"tcdate":1770727297820,"writers":["~"],"signatures":["~Mingshuang_Luo1"],"forum":"RpH9XVkpq1","license":"CC BY-SA 4.0","number":821575,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1784696954235,"domain":"DBLP.org","id":"RpH9XVkpq1","version":2},{"content":{"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"joshi|adaptive_physicsguided_transfer_learning_model_for_predictive_maintenance"},"authorids":{"value":["~Twinkle_Joshi1"]},"html":{"value":"https://doi.org/10.1109/IACIS65746.2025.11210883"},"abstract":{"value":"Hybrid Digital Twins (HDT), which combine physics-based Reduced-Order Models (ROM), have shown strong promise for the predictive maintenance of rotating machinery by enabling realistic synthetic data generation and real-time fault and Remaining Useful Life (RUL) inference. However, existing HDT implementations often require extensive recalibration for each new asset, which limits the domain generalization and practical deployment across various motor types. Hence, this study proposes Adaptive Physics-Guided Transfer Learning (APG-TL) to enable physically consistent adaptation of HDT-based predictive maintenance. Initially, a compact physics ROM through Finite Element Analysis (FEA) and model-order reduction was developed to capture electromagnetic and mechanical dynamics. Subsequently, a One-Dimensional Convolutional Neural Network (1D-CNN) encoder was trained on combined ROM-generated synthetic\nruns and source hardware data. Furthermore, the ROM for a new target asset is calibrated using Bayesian parameter estimation to fit the physical parameters. Next, unsupervised Domain-Adversarial Adaptation (DANN) is employed to align the encoder features to target operations while enforcing a physics-consistency regularizer. Finally, online ROM parameter updates were deployed with occasional few-shot finetuning for the continuous real-time prediction of faults and loweffort adaptation. The proposed APG-TL achieved better\nresults in terms of accuracy (96.54%) than the existing DT with RUL estimation (DT-RUL)."},"title":{"value":"Adaptive Physics-Guided Transfer Learning Model for Predictive Maintenance"},"authors":{"value":["Twinkle Joshi"]}},"tmdate":1776282761019,"pdate":1764990000000,"tcdate":1776282649921,"writers":["~Twinkle_Joshi1"],"signatures":["~Twinkle_Joshi1"],"forum":"dBGzYEopNC","license":"CC BY 4.0","number":46974,"cdate":1776282649921,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1776282761019,"domain":"OpenReview.net/Archive","id":"dBGzYEopNC","version":2},{"content":{"review":{"value":"## Paper Summary\n\nThis paper studies the problem of estimating task-relevant latent dimensionality through a mutual-information-based framework. The authors formulate dimensionality estimation as a symmetric information-preservation problem and analyze how common separable or bilinear critics can inflate estimated dimensionality. To address this, they propose a hybrid critic architecture that decouples critic expressivity from bottleneck dimensionality and introduce a practical one-shot dimensionality estimation protocol based on participation ratios. The approach is validated on synthetic data and physics-inspired datasets, showing stable and interpretable dimensionality estimates across noisy settings.\n\nThe motivation of linking information bottleneck principles with latent dimensionality estimation is well justified, and the paper is generally well positioned within the mutual information and representation learning literature. The work clearly supports its claims through theoretical analysis and a broad set of experiments, demonstrating that the proposed hybrid critic improves robustness and mitigates dimensionality inflation observed with separable critics.\n\n\n## Strengths\n\n\n\n1. Clarity and presentation: the paper is well written, logically structured, and technically clear. The motivation, theoretical framework, and methodological design are presented in an accessible way.\n\n2. Strong methodological insight: the analysis of dimensionality inflation caused by separable critics is insightful, and the proposed hybrid critic offers a concrete and practical solution that improves the reliability of dimensionality estimation.\n\n4. Comprehensive empirical evaluation: the method is evaluated on synthetic experiments and real physics-inspired datasets, including noisy and controlled scenarios, which provides convincing empirical evidence supporting the claims.\n\n5. Practical contribution: the one-shot dimensionality estimation procedure is a useful practical tool that reduces the need for repeated training runs and makes the approach more applicable in practice, aligning with GRaM's themes of scale and simplicity.\n\n\n## Weaknesses\n\n1. Dependence on MI estimation quality: as acknowledged by the authors, performance still depends on the stability of neural mutual-information estimators, which may limit robustness in more complex settings.\n\n\n\n## Recommendation\n\nAccept\n\n### Key reasons:\n\nThe paper is relevant to GRaM and provides a clear, technically sound, and well-motivated contribution to understanding latent dimensionality from a geometric and information-theoretic perspective.\nThe hybrid critic design and one-shot estimator represent meaningful methodological improvements supported by extenditively conducted experiments."},"confidence":{"value":2},"rating":{"value":7},"title":{"value":"Well-motivated and technically solid contribution with convincing empirical validation"},"pmlr_suitability":{"value":"Yes"}},"parentInvitations":"ICLR.cc/2026/Workshop/GRaM/-/Official_Review","nonreaders":[],"tmdate":1772472130492,"tcdate":1771806025899,"writers":["ICLR.cc/2026/Workshop/GRaM","ICLR.cc/2026/Workshop/GRaM/Submission79/Reviewer_GsHu"],"signatures":["ICLR.cc/2026/Workshop/GRaM/Submission79/Reviewer_GsHu"],"forum":"B2WaoKzIyF","number":1,"license":"CC BY 4.0","cdate":1771806025899,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/GRaM/Submission79/-/Official_Review","ICLR.cc/2026/Workshop/GRaM/-/Edit"],"mdate":1772472130492,"domain":"ICLR.cc/2026/Workshop/GRaM","replyto":"B2WaoKzIyF","id":"LKwzaQ1S2j","forumContent":{"TLDR":{"value":"We develop a mutual-information–based method for accurately inferring task-relevant dimensionality while preserving latent geometry, validated on synthetic noisy high-dimensional data and applied to critical and chaotic physics systems."},"venue":{"value":"ICLR 2026 Workshop GRaM Poster"},"keywords":{"value":["Latent Dimensionality","Dimensionality estimation","Intrinsic Dimension","Mutual Information","MI Estimation"]},"abstract":{"value":"Estimating the dimensionality of the latent representation needed for prediction---the task-relevant dimension---is a difficult, largely unsolved problem with broad scientific applications. We cast it as an Information Bottleneck question: what embedding bottleneck dimension is sufficient to compress predictor and predicted views while preserving their mutual information (MI). This repurposes neural MI estimators for dimensionality estimation. We show that standard neural estimators with separable/bilinear critics systematically inflate the inferred dimension, and we address this by introducing a hybrid critic that retains an explicit dimensional bottleneck while allowing flexible nonlinear cross-view interactions, thereby preserving the latent geometry. We further propose a one-shot protocol that reads off the effective dimension from a single over-parameterized hybrid model, without sweeping over bottleneck sizes. We validate the approach on synthetic problems with known task-relevant dimension. We extend the approach to intrinsic dimensionality by constructing paired views of a single dataset, enabling comparison with classical geometric dimension estimators. In noisy regimes where those estimators degrade, our approach remains reliable. Finally, we demonstrate the utility of the method on multiple physics datasets."},"_bibtex":{"value":"@inproceedings{\ngulati2026mutual,\ntitle={Mutual Information and Task-Relevant Latent Dimensionality},\nauthor={Paarth Gulati and Eslam Abdelaleem and Audrey Sederberg and Ilya Nemenman},\nbooktitle={ICLR 2026 Workshop on Geometry-grounded Representation Learning and Generative Modeling},\nyear={2026},\nurl={https://openreview.net/forum?id=B2WaoKzIyF}\n}"},"title":{"value":"Mutual Information and Task-Relevant Latent Dimensionality"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"pdf":{"value":"/pdf/40e9b8ec90fefab65af092a97a092fab3748d0ce.pdf"},"venueid":{"value":"ICLR.cc/2026/Workshop/GRaM"},"paperhash":{"value":"gulati|mutual_information_and_taskrelevant_latent_dimensionality"},"authorids":{"value":["~Paarth_Gulati1","~Eslam_Abdelaleem1","~Audrey_Sederberg1","~Ilya_Nemenman1"]},"Track":{"value":"long paper (up to 8 pages)"},"authors":{"value":["Paarth Gulati","Eslam Abdelaleem","Audrey Sederberg","Ilya Nemenman"]}},"version":2},{"content":{"review":{"value":"The authors propose a method to train physics-informed radial basis function networks for solving forward and inverse problems describing processes in piecewise homogeneous media. The paper stressed that their significant contribution is the tuning of the parameters of the basis functions during network training. However, this method is not novel, see references [14], [16] in the same paper. I suggest putting this paper in the context of the work done in [14] and [16].\n\n[14] D.A. Stenkin V.I. Gorbachenko. Physics-informed radial basis function networks: Solving inverse problems for partial differential equations. Cyber-Physical Systems and Control II. CPSC 2021. Lecture Notes in Networks and Systems, 460:3–12, 2023.\n\n[16] D.A. Stenkin V.I. Gorbachenko. Physics-informed radial basis-function networks. Technical Physics, 68:151–157, 2023."},"confidence":{"value":2},"rating":{"value":6},"title":{"value":"Review of Physics-informed radial basis function networks as Kolmogorov-Arnold networks"}},"nonreaders":[],"tmdate":1765526092887,"tcdate":1740681935159,"writers":["mathai.club/MathAI/2025/Conference","mathai.club/MathAI/2025/Conference/Submission25/Reviewer_WiCP"],"signatures":["mathai.club/MathAI/2025/Conference/Submission25/Reviewer_WiCP"],"forum":"wBz7H1xBbV","number":2,"license":"CC BY 4.0","cdate":1740681935159,"readers":["everyone"],"invitations":["mathai.club/MathAI/2025/Conference/Submission25/-/Official_Review","mathai.club/MathAI/2025/Conference/-/Edit"],"mdate":1765526092887,"domain":"mathai.club/MathAI/2025/Conference","replyto":"wBz7H1xBbV","id":"ar2wO5ZFbS","forumContent":{"TLDR":{"value":"Solving differential equations on neural networks."},"venue":{"value":"MathAI 2025 Oral"},"pdf":{"value":"/pdf/60a381c348725964e7e99d84be537e70ab8143f0.pdf"},"keywords":{"value":["partial differential equations","radial basis function networks","Kolmogorov-Arnold networks","physics-informed neural networks","piecewise homogeneous medium","coefficient inverse problem","incorrect problem","Navier-Stokes equations","Kovasznay flow"]},"venueid":{"value":"mathai.club/MathAI/2025/Conference"},"paperhash":{"value":"|physicsinformed_radial_basis_function_networks_and_kolmogorovarnold_networks"},"authorids":{"value":["~Стенькин_Дмитрий_Александрович1","~Горбаченко_Владимир_Иванович1"]},"abstract":{"value":"Physics-informed neural networks are trained by minimizing the loss function, which is the sum of the squares of the residuals of the equation or system of equations being solved. Such networks do not require grid construction, which is especially important when solving inverse boundary value problems and problems with a complex solution domain. We use radial basis function networks with a Gaussian function. Physics-informed radial basis function networks are easier to train than fully connected networks. They allow one to analytically obtain formulas for the gradient of the loss function. A special feature of our approach to training networks based on radial basis functions is the adjustment of not only the weights, but also the parameters of the radial basis functions, which does not require the selection of parameters of the radial basis functions and accelerates the training process. Algorithms for solving direct and inverse boundary value problems, an algorithm for solving a system of differential equations for modeling the Kovasznay flow have been developed. Programs have been developed that use various algorithms for training physics-informed radial basis function networks."},"_bibtex":{"value":"@inproceedings{\n2025physicsinformed,\ntitle={{PHYSICS}-{INFORMED} {RADIAL} {BASIS} {FUNCTION} {NETWORKS} {AND} {KOLMOGOROV}-{ARNOLD} {NETWORKS}},\nauthor={{\\CYRS}{\\cyrt}{\\cyre}{\\cyrn}{\\cyrsftsn}{\\cyrk}{\\cyri}{\\cyrn} {\\CYRD}{\\cyrm}{\\cyri}{\\cyrt}{\\cyrr}{\\cyri}{\\cyrishrt} {\\CYRA}{\\cyrl}{\\cyre}{\\cyrk}{\\cyrs}{\\cyra}{\\cyrn}{\\cyrd}{\\cyrr}{\\cyro}{\\cyrv}{\\cyri}{\\cyrch} and {\\CYRG}{\\cyro}{\\cyrr}{\\cyrb}{\\cyra}{\\cyrch}{\\cyre}{\\cyrn}{\\cyrk}{\\cyro} {\\CYRV}{\\cyrl}{\\cyra}{\\cyrd}{\\cyri}{\\cyrm}{\\cyri}{\\cyrr} {\\CYRI}{\\cyrv}{\\cyra}{\\cyrn}{\\cyro}{\\cyrv}{\\cyri}{\\cyrch}},\nbooktitle={First Conference of Mathematics of AI},\nyear={2025},\nurl={https://openreview.net/forum?id=wBz7H1xBbV}\n}"},"title":{"value":"PHYSICS-INFORMED RADIAL BASIS FUNCTION NETWORKS AND KOLMOGOROV-ARNOLD NETWORKS"},"authors":{"value":["Стенькин Дмитрий Александрович","Горбаченко Владимир Иванович"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new method for synthesizing heterogeneous graphs called SynHING. SynHING generates synthetic heterogeneous graphs through modules such as Major Motif Generation, Base Subgraph Generation, Intra-Cluster Merge, Inter-Cluster Merge, and Node Feature Generation, allowing for flexible adjustment of the size of the HIN. Multiple experiments validate the effectiveness and scalability of the synthetic heterogeneous graphs."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. This paper proposes a new direction for creating synthetic datasets for heterogeneous graphs, which is scarce in this area.\n2. The methods in the paper are innovative, and each module is necessary and effective.\n3. The experiments in the paper are sufficient and validate the effectiveness of the proposed synthetic graphs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The images in the paper are too small and difficult to see clearly, especially Figure 2.\n2. The authors should provide a detailed introduction of when synthetic graphs need to approximate reference graphs and when they should differ from them. In my view, synthetic graphs should address some of the shortcomings of the reference graphs; otherwise, Section 5.4 lacks significance.\n3. I believe the authors need to conduct explainable experiments on the synthetic graphs to validate their effectiveness and to verify whether the ground truth is accurate. Potential models include: xPath[1] and HENCE-X[2].\n4. The proposed synthetic dataset seems targeted at node classification tasks; can it also be applied to graph classification tasks?\n\nReference:\n[1] Li, T., Deng, J., Shen, Y., Qiu, L., Yongxiang, H., & Cao, C. C. (2023, June). Towards fine-grained explainability for heterogeneous graph neural network. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 7, pp. 8640-8647).\n[2] Lv, G., Zhang, C. J., & Chen, L. (2023). HENCE-X: Toward Heterogeneity-Agnostic Multi-Level Explainability for Deep Graph Networks. Proceedings of the VLDB Endowment, 16(11), 2990-3003."}},"nonreaders":[],"tmdate":1733544594550,"tcdate":1730284345645,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8484/Reviewer_3db8"],"signatures":["ICLR.cc/2025/Conference/Submission8484/Reviewer_3db8"],"forum":"ZbHIDgDFN0","number":2,"license":"CC BY 4.0","cdate":1730284345645,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8484/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733544594550,"domain":"ICLR.cc/2025/Conference","replyto":"ZbHIDgDFN0","id":"xPQ2NllNrt","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["synthetic graph generation","heterogeneous information networks","graph neural networks","explainable artificial intelligence"]},"supplementary_material":{"value":"/attachment/1a596c472d696653ea9d00ed17f09eac6df0c5af.zip"},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Graph Neural Networks (GNNs) excel in modeling graph structures across diverse domains, such as community analysis and recommendation systems. As the need for GNN interpretability grows, there is an increasing demand for robust baselines and comprehensive graph datasets, especially within the realm of Heterogeneous Information Networks (HIN). To address this, we introduce SynHING, a framework for Synthetic Heterogeneous Information Network Generation designed to advance graph learning and explanation.\nAfter identifying key motifs in a target HIN, SynHING systematically employs a bottom-up generation process with intra-cluster and inter-cluster merge modules. This process, along with post-pruning techniques, ensures that the synthetic HIN accurately mirrors the structural and statistical properties of the original graph. The effectiveness of SynHING is validated using four datasets - IMDB, Recipe, ACM, and DBLP - spanning three distinct application categories, demonstrating both its generality and practicality.\nFurthermore, SynHING provides ground-truth motifs for evaluating GNN explainer models, establishing a new benchmark for explainable, synthetic HIN generation. This contributes significantly to advancing interpretable machine learning in complex network environments."},"_bibtex":{"value":"@misc{\nhong2025synhing,\ntitle={Syn{HING}: Synthetic Heterogeneous Information Network Generation for Graph Learning and Explanation},\nauthor={Ming-Yi Hong and Yi-Hsiang Huang and Shao-En Lin and You-Chen Teng and Chih-Yu Wang and Che Lin},\nyear={2025},\nurl={https://openreview.net/forum?id=ZbHIDgDFN0}\n}"},"title":{"value":"SynHING: Synthetic Heterogeneous Information Network Generation for Graph Learning and Explanation"},"pdf":{"value":"/pdf/6ef74ba4277eeea71e09dd3cf7b2e42402e1a696.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"hong|synhing_synthetic_heterogeneous_information_network_generation_for_graph_learning_and_explanation"},"authorids":{"value":["~Ming-Yi_Hong1","~Yi-Hsiang_Huang1","~Shao-En_Lin1","~You-Chen_Teng1","~Chih-Yu_Wang1","~Che_Lin1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ming-Yi Hong","Yi-Hsiang Huang","Shao-En Lin","You-Chen Teng","Chih-Yu Wang","Che Lin"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a unified and lightweight framework designed for real-time infrared and visible image fusion in environments characterized by complex interferences from a frequency domain perspective. It tackles the issue of complex scenes fusion problems, such as adverse weather, low-light environments, and noisy fusion. Authors introduce a multi-modality information interaction guidance module for multi-modality feature interaction and extraction. Extensive fusion experiments in four complex conditions: noise, rain, overexposure, and low-light, verified the effectiveness of the proposed method in dealing with interfering information."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"n/a"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"(1) This paper introduces a unified framework for real-time infrared and visible image fusion in different complex scenes, this is the first work of addressing complex scenes image fusion problems in frequency domain. \n\n(2) The paper proposes a multi-modality interactive guidance mechanism within the Fourier domain, which efficiently extracts and restores useful features from degraded pixels by leveraging the complementary strengths of different modalities.\n\n(3) The fusion performance of this work is very impressive. Extensive complex scenes fusion experiments cover rain, overexposure, low-light, and noisy demonstrate this method outperforms the state-of-the-art fusion methods in both subject and object evaluations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) In the caption of figure 5, it is recommended to add the words “complex scenes”.\n\n(2) Section 4.4 and 4.5 could be combined as one part.\n\n(3) The source code is suggested to be public."}},"nonreaders":[],"tmdate":1731428661050,"tcdate":1730548154598,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6459/Reviewer_5BEs"],"signatures":["ICLR.cc/2025/Conference/Submission6459/Reviewer_5BEs"],"forum":"RqJ0px8osW","number":2,"license":"CC BY 4.0","cdate":1730548154598,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6459/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428661050,"domain":"ICLR.cc/2025/Conference","replyto":"RqJ0px8osW","id":"7XRsk2OI0d","forumContent":{"TLDR":{"value":"We propose a unified lightweight network for infrared and visible image fusion designed for real-time processing in complex scenes, such as adverse weather, low-light, overexposure, and noise conditions."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Infrared and visible image fusion","Complex Scenes","Unified Network","Frequency domain","Real time"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Existing infrared and visible image fusion (IVIF) techniques typically integrate the useful information from different modalities within the ideal conditions. Nevertheless, current state-of-the-art IVIF methods are ineffective when facing complex scene interferences such as bad weather, low light, and high noise, and they typically need to be used in conjunction with other de-interference baselines, which inevitably resulting in the high memory costs and error accumulation, thus yielding sub-optimal fusion results. To address these challenges, We propose a unified lightweight real-time IVIF network for multiple complex scenes. We conducted a theoretically thorough analysis of modal degradations in the frequency domain, leveraging the complementary strengths of both modalities to enhance network learning. Our method facilitates the extraction of critical features even amidst significant pixel interference. For reconstructing fusion results, we introduce a spatial domain branching strategy which significantly improves the local detail resolution, thereby mitigating potential omissions from frequency domain analysis. Extensive qualitative and quantitative experiments demonstrate that our framework excels in handling multiple complex scenes, while maintaining real-time computational efficiency for prompt image processing applications."},"_bibtex":{"value":"@misc{\nli2025a,\ntitle={A unified lightweight complex scenes-oriented network for infrared and visible image fusion},\nauthor={Xilai Li and Xiaosong Li and Tianshu TAN and Wuyang Liu and ye tao},\nyear={2025},\nurl={https://openreview.net/forum?id=RqJ0px8osW}\n}"},"title":{"value":"A unified lightweight complex scenes-oriented network for infrared and visible image fusion"},"pdf":{"value":"/pdf/285c6cf86941647e7f162b085ef6e2d781ec5164.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|a_unified_lightweight_complex_scenesoriented_network_for_infrared_and_visible_image_fusion"},"authorids":{"value":["~Xilai_Li2","~Xiaosong_Li2","~Tianshu_TAN1","~Wuyang_Liu2","~ye_tao8"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xilai Li","Xiaosong Li","Tianshu TAN","Wuyang Liu","ye tao"]}},"version":2},{"content":{"venue":{"value":"ICCV 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10376473/10376477/10376505.pdf"},"venueid":{"value":"dblp.org/conf/ICCV/2023"},"paperhash":{"value":"dai|xvoe_measuring_explanatory_violation_of_expectation_in_physical_events"},"authorids":{"value":["~Bo_Dai5","https://dblp.org/search/pid/api?q=author:Linge_Wang:","https://dblp.org/search/pid/api?q=author:Baoxiong_Jia:","~Zeyu_Zhang4","~Song-Chun_Zhu1","~Chi_Zhang12","https://dblp.org/search/pid/api?q=author:Yixin_Zhu_0001:"]},"html":{"value":"https://doi.org/10.1109/ICCV51070.2023.00369"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iccv/0025WJ0Z0023,\n  author={Bo Dai and Linge Wang and Baoxiong Jia and Zeyu Zhang and Song-Chun Zhu and Chi Zhang and Yixin Zhu},\n  title={X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events},\n  year={2023},\n  cdate={1672531200000},\n  pages={3969-3979},\n  url={https://doi.org/10.1109/ICCV51070.2023.00369},\n  booktitle={ICCV},\n  crossref={conf/iccv/2023}\n}\n"},"abstract":{"value":"Intuitive physics is pivotal for human understanding of the physical world, enabling prediction and interpretation of events even in infancy. Nonetheless, replicating this level of intuitive physics in artificial intelligence (AI) remains a formidable challenge. This study introduces X-VoE, a comprehensive benchmark dataset, to assess AI agents’ grasp of intuitive physics. Built on the developmental psychology-rooted Violation of Expectation (VoE) paradigm, X-VoE establishes a higher bar for the explanatory capacities of intuitive physics models. Each VoE scenario within X-VoE encompasses three distinct settings, probing models’ comprehension of events and their underlying explanations. Beyond model evaluation, we present an explanation-based learning system that captures physics dynamics and infers occluded object states solely from visual sequences, without explicit occlusion labels. Experimental outcomes highlight our model’s alignment with human commonsense when tested against X-VoE. A remarkable feature is our model’s ability to visually expound VoE events by reconstructing concealed scenes. Concluding, we discuss the findings’ implications and outline future research directions. Through X-VoE, we catalyze the advancement of AI endowed with human-like intuitive physics capabilities."},"title":{"value":"X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events"},"authors":{"value":["Bo Dai","Linge Wang","Baoxiong Jia","Zeyu Zhang","Song-Chun Zhu","Chi Zhang","Yixin Zhu"]}},"tmdate":1747133123626,"pdate":1672531200000,"tcdate":1727599547102,"writers":["~"],"signatures":["~Chi_Zhang12"],"forum":"aDBqfKCKcv","license":"CC BY-SA 4.0","number":105562,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747133123626,"domain":"DBLP.org","id":"aDBqfKCKcv","version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2308.10441v1"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"dai|xvoe_measuring_explanatory_violation_of_expectation_in_physical_events"},"authorids":{"value":["~Bo_Dai5","https://dblp.org/search/pid/api?q=author:Linge_Wang:","https://dblp.org/search/pid/api?q=author:Baoxiong_Jia:","https://dblp.org/search/pid/api?q=author:Zeyu_Zhang_0001:","~Song-Chun_Zhu1","~Chi_Zhang12","https://dblp.org/search/pid/api?q=author:Yixin_Zhu_0001:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2308.10441"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2308-10441,\n  publtype={informal},\n  author={Bo Dai and Linge Wang and Baoxiong Jia and Zeyu Zhang and Song-Chun Zhu and Chi Zhang and Yixin Zhu},\n  title={X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2308.10441},\n  url={https://doi.org/10.48550/arXiv.2308.10441}\n}\n"},"abstract":{"value":"Intuitive physics is pivotal for human understanding of the physical world, enabling prediction and interpretation of events even in infancy. Nonetheless, replicating this level of intuitive physics in artificial intelligence (AI) remains a formidable challenge. This study introduces X-VoE, a comprehensive benchmark dataset, to assess AI agents' grasp of intuitive physics. Built on the developmental psychology-rooted Violation of Expectation (VoE) paradigm, X-VoE establishes a higher bar for the explanatory capacities of intuitive physics models. Each VoE scenario within X-VoE encompasses three distinct settings, probing models' comprehension of events and their underlying explanations. Beyond model evaluation, we present an explanation-based learning system that captures physics dynamics and infers occluded object states solely from visual sequences, without explicit occlusion labels. Experimental outcomes highlight our model's alignment with human commonsense when tested against X-VoE. A remarkable feature is our model's ability to visually expound VoE events by reconstructing concealed scenes. Concluding, we discuss the findings' implications and outline future research directions. Through X-VoE, we catalyze the advancement of AI endowed with human-like intuitive physics capabilities."},"title":{"value":"X-VoE: Measuring eXplanatory Violation of Expectation in Physical Events"},"authors":{"value":["Bo Dai","Linge Wang","Baoxiong Jia","Zeyu Zhang","Song-Chun Zhu","Chi Zhang","Yixin Zhu"]}},"tmdate":1747081098851,"pdate":1672531200000,"tcdate":1727599547015,"writers":["~"],"signatures":["~Chi_Zhang12"],"forum":"B1o8UekLuo","license":"CC BY-SA 4.0","number":105556,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747081098851,"domain":"DBLP.org","id":"B1o8UekLuo","version":2},{"content":{"summary":{"value":"This paper proposes MAD, an inference algorithm for diffusion generation under the setting where the diffusion model was trained on noisy data, but at inference time one wants to generate clean data. The proposed algorithm is based on the intuition that meaningful semantics lie on a lower dimensional manifold, while noise is orthogonal off the manifold. Based on this, at each denoising step, the proposed algorithm calculates the derivative of the score with respect to time, and updates the image using a weighted sum of score and the derivative. The paper provides motivational and intuitive analysis on synthetic settings and on applying the algorithm on diffusion models trained on clean real-world data, and performs denoising experiments on the intended setting on a synthetic dataset and a real dataset EMPIAR-11618."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"* It would be helpful to see quantitative evaluations, such as FID with the clean data\n* It would be helpful to see more discussion or analysis on the scope and applicability of the proposed method:\n    * How well does it scale with dataset complexity, e.g. for datasets that go beyond simple semantics and strong inductive bias like CIFAR, ImageNet?\n    * How does the performance look like when there are different levels of corruption/noise in the training data?\n    * What kinds of corruption is the method effective against? Is it primarily designed for general noise, or does it also work for other corruption types such as blurring, compression artifacts, or other distortions? e.g. in the synthetic setting, would the method still be applicable if only the first step of blurry blobs was applied without the second step of adding Gaussian noise?\n    * It would be useful to have some discussion or analysis on contextualizing the performance of this inference-time approach relative to training-time methods or finetuning with small amount of clean data methods. Is the proposed method's performance significantly worse, indicating a performance-compute tradeoff, or is it comparable? Additionally, a potential relevant inference-time baseline to consider for comparison is truncated sampling [1].\n\n[1] Daras, Giannis, Yeshwanth Cherapanamjeri, and Constantinos Daskalakis. \"How much is a noisy image worth? data scaling laws for ambient diffusion.\" ICLR 2025."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"* This paper looks at an interesting setting of obtaining clean generations from a diffusion model trained on noisy data, and proposes an inference only method that does not require retraining. \n* The proposed method is technically sound and theoretically grounded.\n* The illustrative examples and generations on FFHQ, AFHQV2, and ImageNet help gain intuition of the proposed algorithm's behavior.\n* Qualitative visualizations on the synthetic data and the EMPIAR-11618 dataset look promising."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The experimental evaluations are limited, relying primarily on qualitative visualizations. For the denoising experiments, it would be helpful to include quantitative metrics, such as FID scores computed against the clean data distribution.\n* The two datasets in the denoising experiments have very simple semantics, making it hard to assess the scope of general applicability of the proposed method. It would be helpful to provide evaluations on more complex datasets,  e.g. CIFAR or ImageNet. Although the visualizations in 5.2 involve complex data, they are not a direct application of MAD’s intended setting but used diffusion models trained on clean data instead. \n* The performance appears sensitive to the hyperparameters, and there is no obvious rule-of-thumb across different datasets. This could introduce non-trivial hyperparameter sweep overhead when applying MAD in practice, especially if, e.g. the dataset is complex and involve various sub-categories. \n* It would be helpful to provide some evaluations and analysis on the inference speed and potential speed-performance tradeoffs. By default, each denoising step of MAD invokes two forward passes of the scoring network, which could raise concerns about significant computational overhead. It would be helpful to see, e.g. if MAD can work well with sampling speedup strategies (e.g. the simplest one could be using fewer sampling steps)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927943093,"tcdate":1761948647015,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18196/Reviewer_LVkS"],"signatures":["ICLR.cc/2026/Conference/Submission18196/Reviewer_LVkS"],"forum":"ZrP2evfmhq","number":3,"license":"CC BY 4.0","cdate":1761948647015,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18196/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927943093,"domain":"ICLR.cc/2026/Conference","replyto":"ZrP2evfmhq","id":"JPrMh2dydT","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["manifold hypothesis","diffusion models","denoising"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Score-based diffusion models are a highly effective method for generating samples from a distribution of images. We consider scenarios where the training data comes from a noisy version of the target distribution, and present an efficiently implementable modification of the inference procedure to generate noiseless samples.\nOur approach is motivated by the manifold hypothesis, according to which meaningful data is concentrated around some low-dimensional manifold of a high-dimensional ambient space. The central idea is that noise manifests as low magnitude variation in off-manifold directions in contrast to the relevant variation of the desired distribution which is mostly confined to on-manifold directions. We introduce the notion of an extended score and show that, in a simplified setting, it can be used to reduce small variations to zero, while leaving large variations mostly unchanged. We describe how its approximation can be computed efficiently from an approximation to the standard score and demonstrate its efficacy on toy problems, synthetic data, and real data."},"_bibtex":{"value":"@misc{\nelbrachter2026mad,\ntitle={{MAD}: Manifold Attracted Diffusion},\nauthor={Dennis Elbr{\\\"a}chter and Giovanni S Alberti and Matteo Santacesaria},\nyear={2026},\nurl={https://openreview.net/forum?id=ZrP2evfmhq}\n}"},"title":{"value":"MAD: Manifold Attracted Diffusion"},"pdf":{"value":"/pdf/547dfe07a397a053adf72ce2b48a8be8f2768036.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"elbrächter|mad_manifold_attracted_diffusion"},"authorids":{"value":["~Dennis_Elbrächter1","~Giovanni_S_Alberti1","~Matteo_Santacesaria1"]},"authors":{"value":["Dennis Elbrächter","Giovanni S Alberti","Matteo Santacesaria"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel neural operator architecture called Nonlocal Attention Operator (NAO) for learning both forward and inverse problems in physical systems from data. The key contributions are:\n\n1. A new attention-based neural operator that can simultaneously learn the forward mapping (physics modeling) and inverse mapping (physics discovery) for PDE systems.\n2. Theoretical analysis showing how the attention mechanism provides a kernel map that explores the space of identifiability for kernels, helping to resolve ill-posedness in inverse problems.\n3. Empirical demonstration of NAO's advantages over baseline methods, especially for generalization to unseen systems and data-efficient learning in ill-posed inverse problems.\n4. Interpretability of the learned kernels, allowing insight into the discovered physical mechanisms.\n\nThe authors evaluate NAO on several PDE learning tasks, including radial kernel learning, solution operator learning for Darcy flow, and heterogeneous nonlinear material modeling."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Can you provide more insight into how the learned kernels can be interpreted to gain physical understanding? Perhaps a case study showing how NAO discovers meaningful structure in a physical system would be illuminating."},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper presents a novel neural operator architecture that leverages attention in a unique way for PDE learning. The idea of using attention to build a kernel map that can generalize across multiple physical systems is innovative.\n2. The theoretical analysis in Section 4 provides rigorous justification for how the attention mechanism helps address ill-posedness. The empirical evaluations are comprehensive, testing multiple aspects like generalization, data efficiency, and interpretability.\n3. The paper is generally well-written and clearly structured. The motivation, methodology, and results are presented logically.\n4. The ability to simultaneously learn forward and inverse mappings for PDE systems, while providing interpretable kernels, could have broad impact in scientific machine learning and physics-informed AI. The zero-shot generalization capability is particularly notable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The experimental evaluation is limited to relatively simple PDE systems. It's unclear how well NAO would scale to more complex, high-dimensional PDEs encountered in real-world applications.\n2. The interpretability claims could be further substantiated. While learned kernels are visualized, there's limited discussion on how these provide physical insights beyond matching ground truth."},"limitations":{"value":"Yes."}},"nonreaders":[],"tmdate":1730880176239,"tcdate":1722851869747,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission21440/Reviewer_eXH4"],"signatures":["NeurIPS.cc/2024/Conference/Submission21440/Reviewer_eXH4"],"forum":"uSKzEaj9zJ","number":4,"license":"CC BY 4.0","cdate":1722851869747,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission21440/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730880176239,"domain":"NeurIPS.cc/2024/Conference","replyto":"uSKzEaj9zJ","id":"hYcpDfnh2O","forumContent":{"venue":{"value":"NeurIPS 2024 spotlight"},"keywords":{"value":["Foundation Model","Neural Operators","Inverse PDE Problems","Physical Modeling"]},"primary_area":{"value":"machine_learning_for_physical_sciences"},"abstract":{"value":"Despite recent popularity of attention-based neural architectures in core AI fields like natural language processing (NLP) and computer vision (CV), their potential in modeling complex physical systems remains under-explored. Learning problems in physical systems are often characterized as discovering operators that map between function spaces based on a few instances of function pairs. This task frequently presents a severely ill-posed PDE inverse problem. In this work, we propose a novel neural operator architecture based on the attention mechanism, which we coin Nonlocal Attention Operator (NAO), and explore its capability towards developing a foundation physical model. In particular, we show that the attention mechanism is equivalent to a double integral operator that enables nonlocal interactions among spatial tokens, with a data-dependent kernel characterizing the inverse mapping from data to the hidden parameter field of the underlying operator. As such, the attention mechanism extracts global prior information from training data generated by multiple systems, and suggests the exploratory space in the form of a nonlinear kernel map. Consequently, NAO can address ill-posedness and rank deficiency in inverse PDE problems by encoding regularization and achieving generalizability. Lastly, we empirically demonstrate the advantages of NAO over baseline neural models in terms of the generalizability to unseen data resolutions and system states. Our work not only suggests a novel neural operator architecture for learning an interpretable foundation model of physical systems, but also offers a new perspective towards understanding the attention mechanism."},"_bibtex":{"value":"@inproceedings{\nyu2024nonlocal,\ntitle={Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics Discovery},\nauthor={Yue Yu and Ning Liu and Fei Lu and Tian Gao and Siavash Jafarzadeh and Stewart A Silling},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=uSKzEaj9zJ}\n}"},"title":{"value":"Nonlocal Attention Operator: Materializing Hidden Knowledge Towards Interpretable Physics Discovery"},"pdf":{"value":"/pdf/6cac3b157a00647405bf6878194d4c24024eac22.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"yu|nonlocal_attention_operator_materializing_hidden_knowledge_towards_interpretable_physics_discovery"},"authorids":{"value":["~Yue_Yu3","~Ning_Liu6","~Fei_Lu2","~Tian_Gao1","~Siavash_Jafarzadeh1","~Stewart_A_Silling1"]},"authors":{"value":["Yue Yu","Ning Liu","Fei Lu","Tian Gao","Siavash Jafarzadeh","Stewart A Silling"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a CQA method on KGs with numeric values and binary operations. This approach can effectively handles more than 100 types of complex numerical reasoning queries. On three public datasets, the proposed method CNR-NST demonstrates SOTA performance in complex numerical queries, achieving an average improvement of over 40% compared to existing methods."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1) Why LitCQD is mentioned but not compared? \n2) Why equation 9 is used, does it satisfy commutative, associative and distributive laws, and many others?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1) The first work considering complex queries involving binary numeric operations.\n2) Experimental improvements seem significant."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) A comparision with some naive baselines could significantly improve the perception of the experimental results. Especially for the query types with binary operations. The MRR numbers are very small and there is no baseline, and hence it is very hard to judge whether the results are good or not. For example, one could use some simple numeric rules mined from the training graph to derive answers. \n2) It seems that the techniqical contributions are two-fold: 1) Multi-ComplEx, which is a direct extension of ComplEx used in CQD to deal with numerical information; 2) The numerical computation framework. However,  it is unclear whether the numerical computation is a reasonable or not. Does it satisfy some laws like commutative, associative and distributive Laws? I see no discussion about this but I think this is the key which influences the generalization capability of the reasoning. \n3) The test queries are generated as \"hard queries\" in the sense that as least one missing link is in the test graph. However, it is unclear for a multi-hop query, how much percent of the links are seen in the training graph. Note that this is important, as if most of the links in a multi-hop queries are seen. Then the problem can be reduced to a link prediction problem."}},"nonreaders":[],"tmdate":1731429433878,"tcdate":1730469875789,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10446/Reviewer_NDES"],"signatures":["ICLR.cc/2025/Conference/Submission10446/Reviewer_NDES"],"forum":"1epaSm9QRs","number":3,"license":"CC BY 4.0","cdate":1730469875789,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10446/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429433878,"domain":"ICLR.cc/2025/Conference","replyto":"1epaSm9QRs","id":"Ar2tmq1Ac9","forumContent":{"TLDR":{"value":"We propose a Complex Numerical Reasoning with Numerical Semantic Pre-Training Framework, which can perform binary operations on numerical attributes within numerical knowledge graphs and supports complex numerical reasoning tasks."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Numerical Reasoning","Complex Query Answering","Knowledge Graph"]},"supplementary_material":{"value":"/attachment/6f562a3227dc9ec2774d579b998e954af480f6ea.zip"},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Multi-hop complex reasoning over incomplete knowledge graphs has been extensively studied, but research on numerical knowledge graphs remains relatively limited. Recent approaches focus on separately encoding entities and numerical values, using neural networks to process query encodings for reasoning. However, in complex multi-hop reasoning tasks, numerical values are not merely symbols; they carry specific semantics and logical relationships that must be accurately represented. Directly encoding numerical values often leads to the loss of such semantic information. In this work, we propose a Complex Numerical Reasoning with Numerical Semantic Pre-Training Framework CNR-NST. Specifically, we designed a joint link predictor to learn numerical semantics. The proposed framework is the first to enable binary operations on numerical attributes in numerical knowledge graphs, allowing new numerical attributes to be inferred from existing knowledge. The CNR-NST framework can perform binary operations on numerical attributes in numerical knowledge graphs, enabling it to infer new numerical attributes from existing knowledge. Our approach effectively handles up to 102 types of complex numerical reasoning queries. On three public datasets, CNR-NST demonstrates SOTA performance in complex numerical queries, achieving an average improvement of over 40\\% compared to existing methods. Notably, this work expands the range of query types for complex multi-hop numerical reasoning and introduces a new evaluation metric for numerical answers, which has been validated through comprehensive experiments."},"_bibtex":{"value":"@misc{\nzhang2025complex,\ntitle={Complex Numerical Computation  with Numerical Semantic Pre-training Framework},\nauthor={Jun Zhang and Haihong E and Tianyi Hu and Yifan Zhu and Meina Song},\nyear={2025},\nurl={https://openreview.net/forum?id=1epaSm9QRs}\n}"},"title":{"value":"Complex Numerical Computation  with Numerical Semantic Pre-training Framework"},"pdf":{"value":"/pdf/34ca584e29aa34f4d4f60797eac8ebe5370a2ab8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|complex_numerical_computation_with_numerical_semantic_pretraining_framework"},"authorids":{"value":["~Jun_Zhang41","~Haihong_E1","~Tianyi_Hu2","~Yifan_Zhu1","~Meina_Song1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jun Zhang","Haihong E","Tianyi Hu","Yifan Zhu","Meina Song"]}},"version":2},{"content":{"comment":{"value":">**Q4. How is the initial mesh \"extracted from the initial frame via 2D Gaussian Splatting\"?**\n\n**Replay：** Excellent question! Briefly, we first apply 2D Gaussian splatting to obtain the depth maps $D_1^{1:N}$ for the first frame from all camera views. We then use TSDF fusion and Marching Cubes implemented in Open3D [1] to reconstruct the initial mesh $M_1$ from these depth maps. The detailed procedure and algorithmic table have been added to **Appendix H.9** and highlighted in light blue.\n\n\n>**Q5. Does \"extracted from the initial frame\" rely on any pre-trained model?**\n\n**Replay：** Insightful question! Thanks for your careful reading. In fact, the extraction of the initial mesh **does not require any pre-trained model**. As discussed in Q4, 2D Gaussian splatting enables the recovery of depth information directly from images, which can then be used for downstream mesh construction. Therefore, initial mesh extraction can be performed without any pre-trained model. A similar strategy is also employed in Gaussian Garments[2], referenced by Reviewer kW74.\n\n\n\n>**Q6. Comparison to geometry-aware, unsupervised methods.**\n\n**Replay：** Thanks for this insightful comment! This suggestion is extremely valuable for improving our manuscript. To address your concern, we adapted the existing methods LIP [3] and CS [4] to the CDR task. LIP uses **point clouds** as the geometric representation, while CS models cloth using **meshes**. It is important to note that although LIP and CS are geometry-aware, they are **not fully unsupervised**: both rely on pretrained models, in contrast to our fully unsupervised CloDS setting.\n\nTo adapt these methods to the cloth dynamic grounding settings, we pretrain the Particle Posterior Estimator and the Probabilistic Particle Simulator in LIP on cloth point-cloud data. For CS, we pretrain the GNN on cloth-mesh data.\n\nThe predicted 3D geometry is visualized in **Figure 9a**. Videos are available in Part 1 at ***https://anonymous.4open.science/r/ICLR_rebuttal_video***. The quantitative results are presented in **Table A**.\n\n**Table A.** Average RMSE between predicted mesh nodes and ground truth.\n\n|Model|RMSE|\n|:-:|:-:|\n|**LIP**|1.975 $\\pm$ 0.538|\n|**CS**|0.7402 $\\pm$ 0.098|\n|**CloDS(ours)**|**0.1302 $\\pm$ 0.025**|\n\nWe observe that LIP, which uses point clouds as its geometric representation, suffers from **rapidly accumulating errors** during rollout, leading to severe distortions and eventual collapse of the cloth shape. In contrast, CS represents cloth with meshes, and the mesh connectivity provides additional constraints that help preserve the cloth shape better than LIP. Both CS and LIP perform substantially worse than CloDS, which further demonstrates the effectiveness of our approach. The newly added analysis and results have been highlighted in light blue in the revised manuscript.\n\n\n>**Q7. Result on DeepFashion3D V2 (mesh).**\n\n**Replay：** Excellent suggestion, and thank you for directing us to such a high-quality dataset! We construct a Real-Garment training and test dataset by performing physics-based simulations on high-quality garment meshes from the DeepFashion3D V2 dataset. CloDS is then trained and evaluated on this dataset, with the quantitative and qualitative results shown in **Figure 9c**. Videos are available in Part 3 at ***https://anonymous.4open.science/r/ICLR_rebuttal_video***. The quantitative results are presented in **Table B**. We observe that CloDS learns reliable dynamics even on realistic garment geometries, indicating strong generalization to complex real-world shapes.\n\n\n**Table B.** Average RMSE of CloDS between predicted mesh nodes and ground truth in DeepFashion3D V2 dataset.\n\n|Dataset|RMSE|\n|:-:|:-:|\n|**Pants**|0.065 $\\pm$ 0.002|\n|**Dress**|0.016 $\\pm$ 0.007|\n\n**Concluding remark:** We sincerely thank you for putting forward excellent comments. We hope the above responses are helpful to clarify your questions. We look forward to addressing any additional questions. Your consideration of improving the rating of our paper will be much appreciated!\n\n**References:**\n\n[1] Open3D: A modern library for 3D data processing\n\n[2] Gaussian Garments: Reconstructing Simulation-Ready Clothing with Photorealistic Appearance from Multi-View Video. 3DV 2025.\n\n[3] Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D Video, ICLR 2024\n\n[4] Cloth-splatting: 3d cloth state estimation from rgb supervision, CoRL 2024"},"title":{"value":"Rebuttal by Authors (2/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764159778844,"tcdate":1763708297968,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8344/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission8344/Authors"],"forum":"Wj5qJnS1zW","number":7,"license":"CC BY 4.0","cdate":1763708297968,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8344/-/Official_Comment"],"mdate":1764159778844,"domain":"ICLR.cc/2026/Conference","replyto":"NWSbSixHQD","id":"aQ4E0gPT9g","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Neural Dynamic Simulation; Visual Dynamics Grounding; Unsupervised Learning"]},"supplementary_material":{"value":"/attachment/93b90a081c6f4a1cbd806a964abaf3e40999cbc7.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Deep learning has demonstrated remarkable capabilities in simulating complex dynamic systems. However, existing methods require known physical properties as supervision or inputs, limiting their applicability under unknown conditions. To explore this challenge, we introduce Cloth Dynamics Grounding (CDG), a novel scenario for unsupervised learning of cloth dynamics from multi-view visual observations. We further propose Cloth Dynamics Splatting (CloDS), an unsupervised dynamic learning framework designed for CDG. CloDS adopts a three-stage pipeline that first performs video-to-geometry grounding and then trains a dynamics model on the grounded meshes. To cope with large non-linear deformations and severe self-occlusions during grounding, we introduce a dual-position opacity modulation that supports bidirectional mapping between 2D observations and 3D geometry via mesh-based Gaussian splatting in video-to-geometry grounding stage. It jointly considers the absolute and relative position of Gaussian components. Comprehensive experimental evaluations demonstrate that CloDS effectively learns cloth dynamics from visual data while maintaining strong generalization capabilities for unseen configurations. Our code is available at https://github.com/whynot-zyl/CloDS. Visualization results are available at https://github.com/whynot-zyl/CloDS_video."},"_bibtex":{"value":"@inproceedings{\nzhan2026clods,\ntitle={Clo{DS}: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions},\nauthor={Yu-Liang Zhan and Jian Li and Wenbing Huang and Yang Liu and Hao Sun},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Wj5qJnS1zW}\n}"},"title":{"value":"CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions"},"pdf":{"value":"/pdf/9bb2b44467c45d546b8b3ed392b5bd0a69a6238e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhan|clods_visualonly_unsupervised_cloth_dynamics_learning_in_unknown_conditions"},"authorids":{"value":["~Yu-Liang_Zhan1","~Jian_Li24","~Wenbing_Huang1","~Yang_Liu52","~Hao_Sun4"]},"authors":{"value":["Yu-Liang Zhan","Jian Li","Wenbing Huang","Yang Liu","Hao Sun"]}},"version":2},{"content":{"summary":{"value":"This paper introduces OmniChat, a spoken dialogue system using ShareChatX, a synthetic dataset covering emotional, audio, and musical contexts to enable nuanced, multi-modal interactions. With Mix-Former, a fusion module, OmniChat dynamically integrates features like emotion and background sounds, achieving strong performance on complex dialogues. OmniChat outperforms existing methods on DailyTalk and ShareChatX."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What are the model versions of the pre-trained audio encoders used, including Whisper, Emotion2Vec, and Beat?\n\n2. Could you elaborate on the methods for manual verification (line 183) and manual evaluation (line 322)?\n\n3. How do you overlay the music or audio onto the spoken dialogue speech data? What Signal-to-Noise Ratio (SNR) do you use, or are they simply added without adjustment?\n\n4. What is the evaluation prompt used for GPT-Eval?\n\n5. What parameters are included in the style settings? Is it gender, pitch, emotion, energy (line 197), or gender, pitch, speed, emotion (line 182)?\n\n6. Have you evaluated the accuracy of style parameters other than emotion?\n\n7. How many trainable parameters does OmniChat have? Including a comparison of trainable parameters in other models would be beneficial.\n\n8. Missing references: Existing works in audio/emotion dialogue systems include:\n   - *SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words*\n   - *EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions*\n\n9. What are the GPT and human evaluation results for ShareChatX?\n\n10. What is the inference cost of the proposed model?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"This paper introduces a new large-scale synthetic dataset, ShareChatX, covering different scenarios, including emotion, audio and music. Based on this dataset, this paper also presents a competitive spoken dialogue system, OmniChat, achieving SOTA performance on DailyTalk and ShareChatX. Overall, this paper is well-structured and could potentially contribute to the field of spoken dialogue systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The evaluation benchmark is not very comprehensive. Although the evaluation includes two datasets, one of the test sets is constructed similarly to the training set. This similarity may contribute to OmniChat’s superior performance, primarily due to in-domain training. Additionally, the evaluation datasets focus solely on chit-chat. It would be beneficial to include information-seeking instructions in the evaluation as well.\n\n2. The evaluation metrics primarily include reference-based methods (BLEU, ROUGE-L, BERTScore, F1). Previous studies have shown that these metrics may not fully capture conversational quality. Moreover, GPT and human evaluation results are only provided for DailyTalk.\n\n3. Although the paper emphasizes the importance of balancing synthetic and real data, it lacks a detailed analysis of how each synthetic data type (emotion, audio, music) individually impacts model performance.\n\n4. The paper does not provide sufficient details about the model and experimental setup to ensure reproducibility.\n\n5. Many existing voice assistants prioritize real-time settings, which are essential for practical applications. Given the model’s complexity, real-time latency and response times could present challenges."}},"nonreaders":[],"tmdate":1731427733326,"tcdate":1730656432226,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3090/Reviewer_Yjoc"],"signatures":["ICLR.cc/2025/Conference/Submission3090/Reviewer_Yjoc"],"forum":"cVgOIjcNoQ","number":3,"license":"CC BY 4.0","cdate":1730656432226,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3090/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427733326,"domain":"ICLR.cc/2025/Conference","replyto":"cVgOIjcNoQ","id":"ssmwdJPpkm","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"Enhancing Spoken Dialogue Systems with Scalable Synthetic Data"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Spoken Dialogue System","Synthetic Data","Multi-modal Large Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these systems still struggle to handle the full complexity of real-world conversations, including audio events, musical contexts, and emotional expressions, mainly because current dialogue datasets are constrained in both scale and scenario diversity. In this paper, we propose leveraging synthetic data to enhance the dialogue models across diverse scenarios. We introduce **ShareChatX**, the first comprehensive, large-scale dataset for spoken dialogue that spans diverse scenarios. Based on this dataset, we introduce **OmniChat**, a multi-turn dialogue system with a heterogeneous feature fusion module, designed to optimize feature selection in different dialogue contexts. In addition, we explored critical aspects of training dialogue systems using synthetic data. Through comprehensive experimentation, we determined the ideal balance between synthetic and real data, achieving state-of-the-art results on the real-world dialogue dataset DailyTalk. We also highlight the crucial importance of synthetic data in tackling diverse, complex dialogue scenarios, especially those involving audio and music. For more details, please visit our demo page at \\url{https://sharechatx.github.io/}."},"_bibtex":{"value":"@misc{\ncheng2024omnichat,\ntitle={OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios},\nauthor={Xize Cheng and Dongjie Fu and Xiaoda Yang and Minghui Fang and Ruofan Hu and Jingyu Lu and Bai Jionghao and Zehan Wang and Shengpeng Ji and Rongjie Huang and Linjun Li and Yu Chen and Tao Jin and Zhou Zhao},\nyear={2024},\nurl={https://openreview.net/forum?id=cVgOIjcNoQ}\n}"},"title":{"value":"OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios"},"pdf":{"value":"/pdf/18710de70969375ea1129596e69146d27b8f844c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"cheng|omnichat_enhancing_spoken_dialogue_systems_with_scalable_synthetic_data_for_diverse_scenarios"},"authorids":{"value":["~Xize_Cheng1","~Dongjie_Fu1","~Xiaoda_Yang1","~Minghui_Fang1","~Ruofan_Hu2","~Jingyu_Lu1","~Bai_Jionghao2","~Zehan_Wang2","~Shengpeng_Ji1","~Rongjie_Huang1","~Linjun_Li2","~Yu_Chen29","~Tao_Jin2","~Zhou_Zhao3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xize Cheng","Dongjie Fu","Xiaoda Yang","Minghui Fang","Ruofan Hu","Jingyu Lu","Bai Jionghao","Zehan Wang","Shengpeng Ji","Rongjie Huang","Linjun Li","Yu Chen","Tao Jin","Zhou Zhao"]}},"version":2},{"content":{"venue":{"value":"ACL (1) 2025"},"pdf":{"value":"https://aclanthology.org/2025.acl-long.811.pdf"},"venueid":{"value":"dblp.org/conf/ACL/2025"},"paperhash":{"value":"zhang|physreason_a_comprehensive_benchmark_towards_physicsbased_reasoning"},"authorids":{"value":["~Xinyu_Zhang26","~Yuxuan_Dong4","https://dblp.org/search/pid/api?q=author:Yanrui_Wu:","https://dblp.org/search/pid/api?q=author:Jiaxing_Huang:","https://dblp.org/search/pid/api?q=author:Chengyou_Jia:","https://dblp.org/search/pid/api?q=author:Basura_Fernando:","https://dblp.org/search/pid/api?q=author:Mike_Zheng_Shou:","https://dblp.org/search/pid/api?q=author:Lingling_Zhang:","https://dblp.org/search/pid/api?q=author:Jun_Liu_0036:"]},"html":{"value":"https://aclanthology.org/2025.acl-long.811/"},"_bibtex":{"value":"@inproceedings{DBLP:conf/acl/ZhangDWHJFSZ025,\n  author={Xinyu Zhang and Yuxuan Dong and Yanrui Wu and Jiaxing Huang and Chengyou Jia and Basura Fernando and Mike Zheng Shou and Lingling Zhang and Jun Liu},\n  title={PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning},\n  year={2025},\n  cdate={1735689600000},\n  pages={16593-16615},\n  url={https://aclanthology.org/2025.acl-long.811/},\n  booktitle={ACL (1)},\n  crossref={conf/acl/2025-1}\n}\n"},"abstract":{"value":"Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoning-based (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard requiring 15.6, reflecting the complexity of physics-based reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive step-level evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identified four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position PhysReason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models."},"title":{"value":"PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning"},"authors":{"value":["Xinyu Zhang","Yuxuan Dong","Yanrui Wu","Jiaxing Huang","Chengyou Jia","Basura Fernando","Mike Zheng Shou","Lingling Zhang","Jun Liu"]}},"tmdate":1771504064004,"pdate":1735689600000,"externalIds":["dblp:conf/acl/ZhangDWHJFSZ025"],"tcdate":1762419109152,"writers":["~"],"signatures":["~Yuxuan_Dong4"],"forum":"nFAn3qWPQi","license":"CC BY-SA 4.0","number":665158,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1771504064004,"domain":"DBLP.org","id":"nFAn3qWPQi","version":2},{"content":{"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"zhang|physreason_a_comprehensive_benchmark_towards_physicsbased_reasoning"},"authorids":{"value":["~Xinyu_Zhang26","~Yuxuan_Dong4","~Yanrui_Wu2","~Jiaxing_Huang4","~Chengyou_Jia1","~Basura_Fernando1","~Mike_Zheng_Shou1","~Lingling_Zhang1","~Jun_Liu10"]},"abstract":{"value":"Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoning-based (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard requiring 15.6, reflecting the complexity of physics-based reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive step-level evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identified four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position PhysReason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models."},"title":{"value":"PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning"},"authors":{"value":["Xinyu Zhang","Yuxuan Dong","Yanrui Wu","Jiaxing Huang","Chengyou Jia","Basura Fernando","Mike Zheng Shou","Lingling Zhang","Jun Liu"]}},"tmdate":1753161698229,"pdate":1747464300000,"tcdate":1753161698229,"writers":["~Xinyu_Zhang26","~Yuxuan_Dong4","~Yanrui_Wu2","~Jiaxing_Huang4","~Chengyou_Jia1","~Basura_Fernando1","~Mike_Zheng_Shou1","~Lingling_Zhang1","~Jun_Liu10"],"signatures":["~Xinyu_Zhang26"],"forum":"8txoMXcect","license":"CC BY-SA 4.0","number":37902,"cdate":1753161698229,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1753161698229,"domain":"OpenReview.net/Archive","id":"8txoMXcect","version":2},{"content":{"summary":{"value":"This paper tackles a fundamental issue in the optimization of Physics-Informed Neural Networks (PINNs): the imbalance between heterogeneous loss terms (e.g., PDE residual and boundary condition losses) with distinct curvature spectra. To address this, the paper introduces AutoBalance, a post-combine framework in which each loss term is assigned an independent adaptive optimizer (e.g., AdamW). The individually preconditioned updates are then aggregated, forming a unified update direction that achieves an emergent auto-balancing behavior without introducing extra hyperparameters."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Q1. How does AutoBalance differ from MultiAdam? \n\nQ2. For a fair comparison and to isolate the effect of AdamW, please include additional experiments using Adam (without weight decay) as the base optimizer. \n\nQ3. Could you report the mean ± standard deviation of performance across multiple random seeds, rather than only presenting the best-performing results?\n\nQ4. How is the “AutoAdam Preconditioned Hessian” formally defined? The definition provided in the Proof of Theorem 1 appears to be restricted to a quadratic toy example.\n\nI am open to raising my overall evaluation if the concerns outlined in the Weaknesses and Questions sections stem from a misunderstanding or are adequately addressed in the rebuttal."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"S1. **conceptual clarity**: The paper offers a clear conceptual reframing of the loss imbalance problem in PINNs. Identifying “spectral heterogeneity” in the Hessians of multi-objective losses as the core difficulty is insightful. The “pre-combine vs. post-combine” distinction provides an intuitive optimization perspective.\n\nS2. **emergent auto-balancing insight**: I think the observation in Section 3.3 is novel and offers a new insight into the optimization dynamics of PINNs. This phenomenon provides meaningful understanding of why the proposed method achieves stable convergence.\n\nS3. **strong empirical results and orthogonality**: The proposed method achieves consistent improvements over baselines. Furthermore, this paper highlights its orthogonality by showing that AutoBalance enhances the performance of advanced PINN architectures such as RBA-PINN and gPINN.\n\nS4. **simple and easy implementation**: The algorithm is simple to implement and does not rely on additional hyperparameters or tuning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1. **Insufficient literature review**: The paper overlooks several key prior studies that share the same fundamental problem. For example, see [1,2]. Beyond these, there are other closely related works that analyze or propose solutions to address similar issues in PINNs. \n\n[1] Hwang, Youngsik, and Dongyoung Lim. \"Dual cone gradient descent for training physics-informed neural networks.\" Advances in Neural Information Processing Systems 37 (2024): 98563-98595.\n\n[2] Yao, Jiachen, et al. \"Multiadam: Parameter-wise scale-invariant optimizer for multiscale training of physics-informed neural networks.\" International Conference on Machine Learning. PMLR, 2023.\n\nW2. **Limited novelty**: Although the paper frames AutoBalance as a new “post-combine” optimization paradigm,its algorithmic mechanism appears nearly identical to that of MultiAdam [2]. The distinction between the two methods is unclear.\n\nW3. **Theoretical Limitations**: While the simplified quadratic analysis in Section 3.2 offers some intuitive understanding, it falls short of generalizing to the highly non-convex loss landscape of practical PINNs. Consequently, the theoretical justification for why the proposed algorithm influence PINN training dynamics remains quite limited.\n\nIn particular, the use of the condition number as a central analytical tool is problematic since  it represents a globla property of the loss surface. Relying on such a global metric to analyze the behavior of adaptive preconditioned optimizer is fundamentally inappropriate. As a result, the extension of the current analysis to the non-convex setting is theoretically weak.\n\nW4. **benchmark scope**: The evaluation is restricted to classical 1D and 2D PDEs. Recent PINN research increasingly emphasizes stiff, high-dimensional domain problems. Demonstrating AutoBalance’s robustness on these more challenging cases would considerably strengthen the contribution and practical relevance.\n\nW5. **Unfair choice of base optimizer across baselines**: All baseline approaches employ Adam as their undlerying optimizer, whereas AutoBalance adopts AdamW as the base optimizer. This discrepancy raises a fairness issue in the comparative evaluation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933732549,"tcdate":1761878873945,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20240/Reviewer_iM8N"],"signatures":["ICLR.cc/2026/Conference/Submission20240/Reviewer_iM8N"],"forum":"WzivtOfk42","number":3,"license":"CC BY 4.0","cdate":1761878873945,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20240/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933732549,"domain":"ICLR.cc/2026/Conference","replyto":"WzivtOfk42","id":"sfnMmtD0bO","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Physics Informed Neural Networks","Partial Differential Equations","AI4Science","Multi-task Learning","Optimization"]},"supplementary_material":{"value":"/attachment/8cafa9a48e796dab1a6ef78cfaec60a1c0bb400e.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Physics-Informed Neural Networks (PINNs) provide a powerful and general framework for solving Partial Differential Equations (PDEs) by embedding physical laws into loss functions. However, training PINNs is notoriously difficult due to the need to balance multiple loss terms, such as PDE residuals and boundary conditions, which often have conflicting objectives and vastly different curvatures. Existing methods address this issue by manipulating gradients before optimization (a \"pre-combine\" strategy). We argue that this approach is fundamentally limited, as forcing a single optimizer to process gradients from spectrally heterogeneous loss landscapes disrupts its internal preconditioning. In this work, we introduce AutoBalance, a novel \"post-combine\" training paradigm. AutoBalance assigns an independent adaptive optimizer to each loss component and aggregates the resulting preconditioned updates afterwards. Extensive experiments on challenging PDE benchmarks show that AutoBalance consistently outperforms existing frameworks, achieving significant reductions in solution error, as measured by both the MSE and $L^{\\infty}$ norms. Moreover, AutoBalance is orthogonal to and complementary with other popular PINN methodologies, amplifying their effectiveness on demanding benchmarks."},"_bibtex":{"value":"@misc{\nan2026autobalance,\ntitle={AutoBalance: An Automatic Balancing Framework for Training Physics-Informed Neural Networks},\nauthor={Kang An and Chenhao Si and Ming Yan and Shiqian Ma},\nyear={2026},\nurl={https://openreview.net/forum?id=WzivtOfk42}\n}"},"title":{"value":"AutoBalance: An Automatic Balancing Framework for Training Physics-Informed Neural Networks"},"pdf":{"value":"/pdf/6c60ef494b7c81d290144a74666bd3cb3ebd5da5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"an|autobalance_an_automatic_balancing_framework_for_training_physicsinformed_neural_networks"},"authorids":{"value":["~Kang_An1","~Chenhao_Si1","~Ming_Yan1","~Shiqian_Ma3"]},"authors":{"value":["Kang An","Chenhao Si","Ming Yan","Shiqian Ma"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a benchmark suite for evaluating active learning strategies in autonomous laboratories. The work includes synthetic benchmark functions, two real-world tasks (protein design and electron microscopy), and introduces a new metric called landscape flatness to characterize objective functions. The authors evaluate 11 baseline methods on synthetic tasks and 4 methods on real-world applications."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"See Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"* The authors provide a broad set of tasks, from synthetic to complex real-world scenarios in biology and materials science. The authors explain clear difference from traditional optimization benchmarks.\n* The authors introduce the landscape flatness metric, which can quantify the complexity of the objective landscape.\n* The authors provide detailed experiments with 11 baseline methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The authors do not provide enough validation for the proposed landscape flatness metric. It would be better to have theoretical evidence to support the robustness of this metric across different tasks.\n* For experiments, the authors do not explain different numbers of trials (5 trials for synthetic tasks, 3 trials for real-world tasks). Results in Table 1 show high variance across trials, but the authors do not discuss this variation.\n* The authors highlight the scalability as a key contribution, but the paper's analysis of this aspect is limited. For example, there is no quantitative analysis of how computation time scales with dimensionality. And the maximum dimension is 100D.\n* Others:\n    * It is hard to understand the figures (e.g., Fig. 2, 6) due to small size, unclear labeling and limited context.\n    * The introduction contains redundant information about self-driving labs; Minor typo errors and inconsistencies are present in the paper."}},"nonreaders":[],"tmdate":1731428975141,"tcdate":1731111246938,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13952/Reviewer_ycRe"],"signatures":["ICLR.cc/2025/Conference/Submission13952/Reviewer_ycRe"],"forum":"PHkUNcno9n","number":3,"license":"CC BY 4.0","cdate":1731111246938,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13952/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428975141,"domain":"ICLR.cc/2025/Conference","replyto":"PHkUNcno9n","id":"ynypxk8WYu","forumContent":{"TLDR":{"value":"We introduce BALSA: a benchmark tailored to evaluate search algorithms in autonomous laboratories within the active learning framework."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["active learning","experimental design","AI for science"]},"supplementary_material":{"value":"/attachment/09cfe5b65ee8550d0dd14dc2fc8f9c669c015ea3.zip"},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Accelerating scientific discoveries holds significant potential to address some of the most pressing challenges facing society, from mitigating climate change to combating public health crises, such as the growing antibiotics resistance. The vast and complex nature of design parameter spaces makes identifying promising candidates both time-consuming and resource-intensive, rendering conventional exhaustive searches impractical. However, recent advancements in data-driven methods, particularly within the framework of \"active learning,\" have led to more efficient strategies for scientific discovery. By iteratively identifying and labeling the most informative data points, these methods function in a closed loop, guiding experiments or simulations to accelerate the identification of optimal candidates while reducing the demand for data labeling. Despite these advancements, the lack of standardized benchmarks in this emerging field of autonomous scientific discovery impedes progress and limits its potential translational impact. To address this, we introduce BALSA: a comprehensive benchmark specifically designed for evaluating various search algorithms applied in autonomous laboratories within the active learning framework. BALSA offers a standardized evaluation protocol, provides a metric to characterize high-dimensional objective functions, and includes reference implementations of recent methodologies, with a focus on minimizing the data required to reach optimal results. It provides not only a suite of synthetic functions or controlled simulators but also real-world active learning tasks in biology and materials science — each presenting unique challenges for autonomous laboratory tasks."},"_bibtex":{"value":"@misc{\ntung2025balsa,\ntitle={{BALSA}: Benchmarking Active Learning Strategies for Autonomous laboratories},\nauthor={Po-Yen Tung and Yangtao Chen and Peng Bo and Hao Zhang and Wenjie Du and Stefan Bauer and Ye Wei},\nyear={2025},\nurl={https://openreview.net/forum?id=PHkUNcno9n}\n}"},"title":{"value":"BALSA: Benchmarking Active Learning Strategies for Autonomous laboratories"},"pdf":{"value":"/pdf/1a6e468c50d6b6fbd49f3a0ca440f857083bd4c4.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"tung|balsa_benchmarking_active_learning_strategies_for_autonomous_laboratories"},"authorids":{"value":["~Po-Yen_Tung1","~Yangtao_Chen3","~Peng_Bo2","~Hao_Zhang85","~Wenjie_Du2","~Stefan_Bauer1","~Ye_Wei1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Po-Yen Tung","Yangtao Chen","Peng Bo","Hao Zhang","Wenjie Du","Stefan Bauer","Ye Wei"]}},"version":2},{"content":{"summary":{"value":"The paper presents a Physics-Based Interactive Motion Transition (PIMT) framework for generating seamless and natural character animations in real-time 3D interactive applications. The framework combines motion capture animation with physical simulation to connect motion clips smoothly. Key innovations include:  \n1.Self-Behavior Cloning (SBC): Enhances unsupervised reinforcement learning for motion transitions, leveraging exploration trajectories to improve the learning process.  \n2.Sophisticated Reward Function: Integrates objectives for transition accuracy and naturalness, balancing between diversity and stability.  \n3.Task Planning and Curriculum Learning: Employs a state machine and a curriculum learning strategy to plan and train the control policy efficiently."},"suitability":{"value":3},"strengths":{"value":"The framework allows arbitrary specification of transition moments and target motion rotations, providing high responsiveness and user control. Unlike interpolation methods, the physics-based approach ensures that transitions adhere to human body dynamics, avoiding visual artifacts like foot-skating."},"confidence":{"value":2},"rating":{"value":5},"limitations":{"value":"Without reference motion datasets, the method sometimes produces unnatural upper body movements.  The framework's inference and simulation times may not be sufficient for all real-time applications, and further efficiency improvements are needed."}},"nonreaders":[],"tmdate":1721657388620,"tcdate":1716556600865,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission4868/Reviewer_kEYm"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission4868/Reviewer_kEYm"],"forum":"2CUpjuBbtX","number":1,"license":"CC BY 4.0","cdate":1716556600865,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/Submission4868/-/Official_Review","acmmm.org/ACMMM/2024/Conference/-/Edit"],"mdate":1721657388620,"domain":"acmmm.org/ACMMM/2024/Conference","replyto":"2CUpjuBbtX","id":"BZXIPBKQTz","forumContent":{"venue":{"value":"MM2024 Poster"},"supplementary_material":{"value":"/attachment/bb9eb4601d4d8eb80cb1654a433d982604227ca6.zip"},"abstract":{"value":"Motion transitions, which serve as bridges between two sequences of character animation, play a crucial role in creating long variable animation for real-time 3D interactive applications. In this paper, we present a framework to produce hybrid character animation, which combines motion capture animation and physical simulation animation that seamlessly connects the front and back motion clips. In contrast to previous works using interpolation for transition, our physics-based approach inherently ensures physical validity, and both the transition moment of the source motion clip and the horizontal rotation of the target motion clip can be specified arbitrarily within a certain range, which achieves high responsiveness and wide latitude for user control. The control policy of character can be trained automatically using only the motion capture data that requires transition, and is enhanced by our proposed Self-Behavior Cloning (SBC), an approach to improve the unsupervised reinforcement learning of motion transition. We show that our framework can accomplish the interactive transition tasks from a fully-connected state machine constructed from nine motion clips with high accuracy and naturalness."},"relevance_to_conference":{"value":"The framework proposed in the paper combines two forms of character animation, motion capture animation controlled by joint position sequences and physical simulation animation controlled by joint motor torques (generated by the model), allowing real-time animation systems organized through state machine to generate corresponding transition animation between motion capture animations based on the control data input by the user (including the selected next motion, input time, and rotation angle)."},"_bibtex":{"value":"@inproceedings{\ndeng2024pimt,\ntitle={{PIMT}: Physics-Based Interactive Motion Transition for Hybrid Character Animation},\nauthor={Yanbin Deng and Zheng Li and Ning Xie and Wei Zhang},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=2CUpjuBbtX}\n}"},"title":{"value":"PIMT: Physics-Based Interactive Motion Transition for Hybrid Character Animation"},"secondary_subject_area":{"value":["[Experience] Art and Culture","[Generation] Generative Multimedia"]},"pdf":{"value":"/pdf/41b0497781cce294ffcdecb1e34a329a543ad294.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"deng|pimt_physicsbased_interactive_motion_transition_for_hybrid_character_animation"},"primary_subject_area":{"value":"[Experience] Interactions and Quality of Experience"},"authorids":{"value":["~Yanbin_Deng1","~Zheng_Li31","~Ning_Xie7","~Wei_Zhang110"]},"authors":{"value":["Yanbin Deng","Zheng Li","Ning Xie","Wei Zhang"]}},"version":2},{"content":{"summary":{"value":"The authors explore graph rewiring for GNN-based surrogates for CFD modeling. They note that previous rewiring schemes and PIORF' long-range connections ignore fluid dynamics and can violate physical principles of fluid dynamics. Authors propose their FLARE method (Flow alighment rewiring) that adds selected 2-hop neighbours and uses dot product between sender's velocity and the displacement vector to decide if a directed edge should be added. This methodology aligns flow with the velocity and graph is rewired at each time step. Experiments on three datasets show that simply adding all 2-hop neighbors already rivals PIORF, and FLARE offers gains. Ablations show that adding edges opposite to the flow decreases performance and 3/4-hop rewiring also degrades results."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Make more solid explanation of superiority of FLARE. See W1. Try not to use vague proves. Authors explain dataset dependency using this sentence \"It is worth noting that FLARE gains more significant improvements on CylinderFlow than Airfoil likely because the former has more dynamics and challenging  flow conditions, giving more room for FLARE to improve.\"  Could the authors provide a more detailed analysis of why FLARE benefits more from dynamic flows, and how the method might be adapted to perform better on simpler cases such as the Airfoil density prediction?\n2. Analyze flow-alighment threshold parameter sensitivity.\n3. Conduct runtime/memory cost analysis.\n4. Can you show rollout error vs. time for Airfoil and Tandem-Airfoil, not only CylinderFlow?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The idea is novel, the paper is well-written, good presentation, also:\n1. The physics-informed rewiring heuristic uses simple fluid-dynamics principles and often gives a gain in performance.\n2. The method is easy to implement.\n3. The authors made ablation study."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Gains of FLARE are dataset-dependent sometimes. But the paper draws broad conclusions. For example, \"By comparing FLARE with 2-HOP-ALL and PIORF,we can conclude the effectiveness of physics guided design. The directional and local connections determined by flow directions are essential to achieve these performance gains.\" \n2. Authors have not analyzed threshold parameter sensitivity. It is simply fixed to T=0 in main experiments as I understand. Therefore, ablation study is limited.\n3. No runtime/memory cost analysis was performed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358718810,"tcdate":1761312329122,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15260/Reviewer_ebA5"],"signatures":["ICLR.cc/2026/Conference/Submission15260/Reviewer_ebA5"],"forum":"izLvJEBkae","number":2,"license":"CC BY 4.0","cdate":1761312329122,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15260/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358718810,"domain":"ICLR.cc/2026/Conference","replyto":"izLvJEBkae","id":"NO9mSr80Pe","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed Graph Neural Networks; CFD; mesh; rewiring; unsteady flow"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"To overcome computation burden of traditional computational fluid dynamics (CFD) simulations, researchers have explored different architectures to develop physics-informed simulation methods. Among them, graph neural networks (GNN) are most suitable for adopting CFD meshes, which are extensively used in engineering and industrial applications. However, classical GNNs propagate information among neighbour nodes, which highly restrict information exchange within the network. To address this issue, graph rewiring methods have been developed for generic graph problems, but not particular for fluid simulation. PIORF, introducing edges connecting distant nodes, is the first graph rewiring method to do so, and previous experiments have demonstrated its effectiveness against state-of-the-art generic rewiring methods. Nevertheless, in this work, we found that simply connecting all 2-hop nodes can provide competitive performance with PIORF. This result raises three questions: 1) Is physics-informed rewiring really useful for improving flow predictions? 2) Should we consider just local connection, instead of connecting distant nodes? 3) Do we need to change the connections based on input flow for rollout simulations? By thoroughly adopting physical fluid principles, we propose a simple yet very efficient method, Flow Alignment Rewiring (FLARE) technique, which connects 2-hop nodes only when the node direction aligns with input flow direction. Hence, FLARE is a physics-informed local rewiring method, different from PIORF and well-aligned with fluid physics. Extensive numerical experiments on flows over a cylinder and single and tandem airfoil under different flow conditions and deep network architectures demonstrate that FLARE outperforms PIORF and various 2-hop rewiring approaches by a significant margin."},"_bibtex":{"value":"@misc{\nli2026graph,\ntitle={Graph Rewiring based on Flow Alignment for Improving Fluid Simulation},\nauthor={Zenong Li and Wei Xian Lim and Wai Lee Chan and Adams Wai-Kin Kong},\nyear={2026},\nurl={https://openreview.net/forum?id=izLvJEBkae}\n}"},"title":{"value":"Graph Rewiring based on Flow Alignment for Improving Fluid Simulation"},"pdf":{"value":"/pdf/fde6cfdb3e702fa96f577b9f6fb0cfb47669442b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|graph_rewiring_based_on_flow_alignment_for_improving_fluid_simulation"},"authorids":{"value":["~Zenong_Li1","~Wei_Xian_Lim2","~Wai_Lee_Chan1","~Adams_Wai-Kin_Kong1"]},"authors":{"value":["Zenong Li","Wei Xian Lim","Wai Lee Chan","Adams Wai-Kin Kong"]}},"version":2},{"content":{"summary":{"value":"Current Referring Expression Segmentation benchmarks mainly focus on either single targets with short queries or multiple targets from distinctly different queries on a single domain. To address this issue, this paper proposes WildRES, which incorporates long queries with diverse attributes and non-distinctive queries for multiple targets. Therefore, the proposal benchmark aims to deal with complex reasoning in a real-world setting. In light that existing models perform significantly worse on WildRES, this paper proposes SynRES, an automatic pipeline to generate paired synthetic data. SynRES shows promising performance gain for different models, not only on the WildRES dataset, but also on the classic RES benchmarks."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See above"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The WildRES is well curated, which complements the current limitation in the RES benchmarks in 1) providing more attributes for the  single object referring, 2) images featuring multiple objects sharing the same class but distinct attributes \n\n\nIt is relatively hard to obtain high-quality paired data for the RES tasks, and the motivation to leverage the automatic synthetic data makes sense. The synthetic data generation pipeline is helpful in generating data in a more controlled setting: many attributes/ identical objects with different attributes. The author proposes a solid pipeline to conduct the data generation: 1) captioning to generate diverse attributes and T2I generation for the synthetic images; 2) pseudo-label generation and filtering via group similarity; 3) RES-specific augmentation enhancement. Comprehensive ablation studies are provided to justify the key design choices in this synthetic data generation workflow.\n\nExperimental results across both conventional RES benchmarks and the proposed WildRES dataset—under in-domain and cross-domain settings—demonstrate strong performance and validate the effectiveness of the proposed approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The method part for the SynRES can be organized in a better way. In Section 4.2, it is good to have clear mathematical annotations, but it would be better to organize them in a more systematic manner. It is also suggested to have some high-level ideas before diving into the technical details. For example, what is the high-level idea and assumption for clustering according to expression pairs' similarity in a global manner (i.e., averaging over images)? Without it, the design choice would seem to be more random and not well-motivated. \n\nIn the Debiased Text Augment in Section 4.3, I understand that if you do the superclass replacement for the conventional RES dataset, there can be false negatives present. It also makes sense that we can avoid it with the synthetic data, but the way that it is achieved is not obvious enough. The author should make this part clearer.\n\nFrom the LISA baseline in Table one,  it seems that the improvement on the multiple objects with shared attributes is the most significant; this also applies to the LISA in Table 2 for the WildRES-DS setting. Adding the synthetic data training seems to hurt the many attribute performance a little bit in terms of gIoU in some settings. Can the Author comment on the possible reason?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924826241,"tcdate":1761674231204,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14417/Reviewer_yvJL"],"signatures":["ICLR.cc/2026/Conference/Submission14417/Reviewer_yvJL"],"forum":"cgr5OAXe3q","number":2,"license":"CC BY 4.0","cdate":1761674231204,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14417/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924826241,"domain":"ICLR.cc/2026/Conference","replyto":"cgr5OAXe3q","id":"3Snk6h1rDZ","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Referring Expression Segmentation","Synthetic data","Multimodal augmentation"]},"supplementary_material":{"value":"/attachment/c086636c0c2aeea809f4586753e9c5607fc0af65.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Despite the advances in Referring Expression Segmentation (RES) benchmarks, their evaluation protocols remain constrained, primarily focusing on either single targets with short queries (containing minimal attributes) or multiple targets from distinctly different queries on a single domain. This limitation significantly hinders the assessment of more complex reasoning capabilities in RES models.\n    We introduce  WildRES, a novel benchmark that incorporates long queries with diverse attributes and non-distinctive queries for multiple targets. This benchmark spans diverse application domains, thus enabling more rigorous evaluation of complex reasoning capabilities in real-world settings. \n    Our analysis reveals that existing RES models demonstrate substantial performance deterioration when evaluated on WIldRES. To address this challenge, we introduce SynRES, an automated pipeline generating densely paired compositional synthetic training data through three innovations: (1) a dense caption-driven synthesis for attribute-rich image-mask-expression triplets, (2) reliable semantic alignment mechanisms rectifying caption-pseudo mask inconsistencies via Image-Text Aligned Grouping, and (3) domain-aware augmentations incorporating mosaic composition and superclass replacement to emphasize generalization ability and distinguishing attributes over object categories.\n    Experimental results demonstrate that models trained with SynRES achieve consistent improvements on not only our complex WildRES benchmark but also classic RES benchmarks (e.g. RefCOCO/+/g).\nCode is available at https://anonymous.4open.science/r/SynRES-Review-4B1F.\nDataset will be available upon acceptance."},"_bibtex":{"value":"@misc{\nkim2026towards,\ntitle={Towards Robust Referring Expression Segmentation for Complex Reasoning in the Wild},\nauthor={Dong-Hee Kim and Hyunjee Song and Donghyun Kim},\nyear={2026},\nurl={https://openreview.net/forum?id=cgr5OAXe3q}\n}"},"title":{"value":"Towards Robust Referring Expression Segmentation for Complex Reasoning in the Wild"},"pdf":{"value":"/pdf/a05b72f8da22c8e836233b9c978c4462034241ec.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|towards_robust_referring_expression_segmentation_for_complex_reasoning_in_the_wild"},"authorids":{"value":["~Dong-Hee_Kim1","~Hyunjee_Song1","~Donghyun_Kim2"]},"authors":{"value":["Dong-Hee Kim","Hyunjee Song","Donghyun Kim"]}},"version":2},{"content":{"summary":{"value":"1. **Originality-wise**: the paper proposes a large-scale UAV event-based dataset for robust 3D reconstruction.\n2. **Quality-wise**: the proposed dataset has varied large-scale scenes including different flight path, scenarios types, illumination, and height settings.\n3. **Clarity-wise**: the manuscript is clearly written, with well-structured methodology, detailed explanations, and intuitive visualizations that enhance understanding."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please refer to Weaknesses part."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. A Pioneering and Highly Valuable Dataset: The most significant contribution of this work is the SkyEvents dataset itself. SkyEvents is not only large in scale (over 8 hours, covering 0.72 km²) but also rich in modalities (RGB, Event, LiDAR) and provides accurate 6-DoF poses, filling a major gap. \n2. Solves a Critical Practical Problem: The proposed GTA module directly addresses a very challenging yet common real-world problem: precise timestamp alignment between multiple sensors, especially in low-cost setups without hardware synchronization"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Regarding the GTA Module**: While the paper demonstrates the module's effectiveness indirectly through its positive impact on 3D reconstruction and other downstream tasks, it lacks a more direct and intuitive quantitative evaluation. To compellingly validate the module's performance, I suggest the authors conduct experiments on existing datasets with pre-aligned ground-truth timestamps (e.g., DSEC, MVSEC). This would allow for a direct measurement of the method's alignment error and a quantitative comparison of its effectiveness against other approaches.\n2. **Regarding Low-Light Image Generation**: The approach of generating synthetic low-light images is practical for settings where scenes cannot be re-captured. However, aligning event data captured directly under bright sunlight with these synthetically darkened images could introduce a significant data mismatch. Event cameras are known to exhibit different noise characteristics, such as leaky events, in bright conditions compared to true dark environments. This raises concerns about the fidelity of the alignment. I am curious if the authors implemented any specific procedures to address this potential distortion.\n3. **Regarding Evaluation Metrics**: I agree that using the same metrics as the original 3DGS is valid for the motion deblur task. However, for low-light or over-exposed conditions, these reference-based metrics may become ineffective due to the quality degradation of the ground-truth images themselves. An evaluation based on a flawed reference is not meaningful. I recommend that the authors incorporate no-reference image quality assessment metrics (e.g., BRISQUE, HyperIQA) to ensure the validity and intuitiveness of the evaluation results.\n4. The video visualization is lacking and this reduces the credibility of the real effect.\n\n---\nI have listed my concerns, and the score will be adjusted based on the author's response."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917579500,"tcdate":1760511604304,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4795/Reviewer_3NrW"],"signatures":["ICLR.cc/2026/Conference/Submission4795/Reviewer_3NrW"],"forum":"dxHPqQindP","number":1,"license":"CC BY 4.0","cdate":1760511604304,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4795/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917579500,"domain":"ICLR.cc/2026/Conference","replyto":"dxHPqQindP","id":"Ox5bTeIjOF","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Event","3D Scene Reconstruction"]},"supplementary_material":{"value":"/attachment/824802384d9550990a5b7b5e59edb1183e19bb14.pdf"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in large-scale 3D scene reconstruction using unmanned aerial vehicles (UAVs) have spurred increasing interest in neural rendering techniques. However, existing approaches with conventional cameras struggle to capture consistent multi-view images of scenes, particularly in extremely blurred and low-light environments, due to the inherent limitations in dynamic range caused by long exposure and motion blur resulting from camera motion. As a promising solution, bio-inspired event cameras exhibit robustness in extreme scenarios, due to their high dynamic range and microsecond-level temporal resolution. Nevertheless, dedicated event datasets specifically tailored for large-scale UAV 3D scene reconstruction remain limited. To bridge this gap, we introduce SkyEvents, a pioneering large-scale event-enhanced UAV dataset for 3D scene reconstruction, incorporating RGB, event, and LiDAR data. SkyEvents encompasses 45 sequences, spanning over 8 hours of video, captured across a diverse set of illumination conditions, scenarios, and flight altitudes. To facilitate the event-based 3D scene reconstruction with SkyEvents, we propose the Geometry-constrained Timestamp Alignment (GTA) module to align timestamps between the event and RGB cameras. Furthermore, we introduce a Region-wise Event Rendering (RER) loss for supervising the rendering optimization. With SkyEvents, we aim to motivate and equip researchers to advance large-scale 3D scene reconstruction in challenging environments, harnessing the unique strengths of event cameras. Dataset and code will be available at https://github.com/Anthony-ECPKN/SkyEvent."},"_bibtex":{"value":"@inproceedings{\nma2026skyevents,\ntitle={SkyEvents: A Large-Scale Event-enhanced {UAV} Dataset for Robust 3D Scene Reconstruction},\nauthor={Wenzong Ma and Zhuoxiao Li and Jinjing Zhu and Tongyan Hua and Kanghao Chen and Zidong Cao and Da Yang and Peilun Shi and Yibo Zhou and Wufan Zhao and Hui Xiong},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=dxHPqQindP}\n}"},"title":{"value":"SkyEvents: A Large-Scale Event-enhanced UAV Dataset for Robust 3D Scene Reconstruction"},"pdf":{"value":"/pdf/ee6d696358dea92d7216468ee90aaaaca157809d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ma|skyevents_a_largescale_eventenhanced_uav_dataset_for_robust_3d_scene_reconstruction"},"authorids":{"value":["~Wenzong_Ma1","~Zhuoxiao_Li3","~Jinjing_Zhu1","~Tongyan_Hua1","~Kanghao_Chen1","~Zidong_Cao1","~Da_Yang3","~Peilun_Shi1","~Yibo_Zhou4","~Wufan_Zhao1","~Hui_Xiong1"]},"authors":{"value":["Wenzong Ma","Zhuoxiao Li","Jinjing Zhu","Tongyan Hua","Kanghao Chen","Zidong Cao","Da Yang","Peilun Shi","Yibo Zhou","Wufan Zhao","Hui Xiong"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Symbolic-R1, an LLM-based symbolic regression method trained on SymbArena, a new benchmark with 148K synthetic equations. The approach uses instruction tuning followed by reinforcement learning (Form-GRPO) with structure-aware rewards, and employs iterative refinement (HER) during inference. The paper introduces a form-level consistency metric alongside numerical accuracy metrics. Results show improvements over baselines including PySR on the proposed benchmark."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. The reality-verification ensures similarity to known physics, and benchmarks include famous equations likely in pre-training data. How was contamination assessed? Could results be provided for: a) SRBench black-box problems (real-world data with no known equations) compared to traditional (PySR) and transformer-based methods (E2E, TPSR)? b) LLM-SRBench datasets - Table 1 tags these as \"LLM-based only\" but they can be used for any SR method, so this characterization seems inaccurate.\n\n2. How is the data generation different from Lample & Charton? Is the same code used? The claim that reality-checking makes it \"fairer than directly using real data...avoiding the possibility that LLM has seen them during pre-training\" while simultaneously ensuring similarity to existing equations needs clarification; this appears to increase contamination risk for LLMs.\n\n3. LLM-SR and SGA are designed for context-rich problems but evaluated on context-free data, so they don't seem like appropriate comparisons. Wouldn't LaSR (Grayeli et al.), which integrates with PySR, be more suitable? Could comparison be provided, along with qualitative analysis on Feynman equations to understand whether the model recovers structure or leverages memorization?\n\n4. Could you break down SRBench performance for Feynman equations, ODEs, and black-box problems separately, providing symbolic accuracy, numeric accuracy, complexity, and runtime for each category? Also, I would suggest to bring the results on standard benchmarks to the main body of the paper as they are more reliable than test set of synthetic SymbArena dataset, especially since the test data distribution could be very close to the training distribution as they are sampled from the same generator.\n\n5. Table 9 shows identical or very similar LLM-SR and SGA outputs across different LLMs. Given different models and frameworks, one wouldn't expect identical outputs.\n\n6. Given that various symbolic accuracy metrics have been studied (SymPy symbolic recovery, tree edit distance, LLM-as-judge), what specific advantage does the proposed form-level consistency metric provide? and why is it novel?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"* The core idea of fine-tuning LLMs specifically for symbolic regression is interesting and relatively underexplored compared to inference-time scaling approaches.\n* The GRPO training scheme with multiple reward types is a clever design that tries to balance structural correctness with numerical accuracy.\n* The reality-verification step for test equations (using LLM to check similarity to known physics) to ensure practical relevance of the benchmark seems novel and promising."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Problem formulation.** The paper focuses on traditional SR without relying on domain knowledge (finding equations from data only) but the positioning is confusing. LLM-SR and SGA are designed for context-rich problems and seem tested outside their scope here. If we consider the contribution as \"fine-tuning\" the model for SR, one naturally thinks of transformer-based methods like E2E which, while trained on much more data, struggle with symbolic recovery. So it is not clear if the advantage here is coming from LLM prior knowledge, the training method, or test set characteristics (see the next point)? The paper doesn't clearly establish what problem it's solving and against which baselines it should primarily be evaluated.\n\n**Contamination concerns.** The test set is \"reality-verified\" to be similar to known physics equations, which seems to increase rather than decrease contamination risk; if equations are similar to known physics, the pre-trained LLM likely learned these patterns. Evaluation on famous benchmarks (Nguyen, Feynman equations in SRBench) that are almost certainly in LLM pre-training data raises questions about whether improvements come from genuine SR capability or memorization. Contamination should be studied not only with respect to SymbArena's train set but also the LLM's pre-training corpus.\n\n**Evaluation setup.** Following the previous point, the benchmark datasets may not be suitable for proper evaluation of general SR methods (see questions section for some suggested datasets). For completeness, it would be good to also add relevant neural baselines like E2E and TPSR. Evaluating on complex, real-world problems without contamination concerns could better demonstrate performance.\n\n**Minor presentation issues.** Several typos exist (e.g., \"Defination\" in Figure 1, \"Hypothe- sis\" in third contribution bullet point). Table 8 doesn't specify which metric is reported."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916719481,"tcdate":1761882384792,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3423/Reviewer_6dC2"],"signatures":["ICLR.cc/2026/Conference/Submission3423/Reviewer_6dC2"],"forum":"OjRaJw4tnr","number":2,"license":"CC BY 4.0","cdate":1761882384792,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3423/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916719481,"domain":"ICLR.cc/2026/Conference","replyto":"OjRaJw4tnr","id":"dSYoUVNJdC","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Benchmark","Scientific Discovery","Large Language Models","Symbolic Regression"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Deriving governing equations from observational data, known as Symbolic Regression (SR), is a cornerstone of scientific discovery. \nLarge Language Models (LLMs) have shown promise in this task by leveraging their vast cross-disciplinary scientific knowledge. However, existing LLM-based methods primarily rely on direct inference or prompt engineering, often requiring excessive inference iterations to converge on correct formulas or failing to treating complex equation targets. These limitations in effectiveness and generalization stem from an inherent tension between pre-trained LLMs' proficiency in approximate reasoning and the high-precision demands of SR tasks. The underlying reason for this stems from a fundamental mismatch between the general-purpose pre-training of LLMs and the specialized nature of symbolic regression, a problem exacerbated by the scarcity of high-quality, task-specific data. To bridge this gap, we propose to fine-tune LLMs for enhanced SR capability. Yet, the absence of dedicated datasets for SR-oriented fine-tuning remains a critical barrier. We thus introduce SymbArena, specifically engineered to optimize LLMs for SR. This benchmark comprises 148,102 diverse equations formulated as corpora of 1.83 billion tokens for LLM utilization, enabling effective training and inference. Further, SymbArena proposes a heuristics metric to precisely quantify form-level consistency, going beyond existing SR numerical-oriented evaluation strategies. \n%is designed for the rigorous evaluation of both LLM-based and traditional SR methods and includes a novel, graded, and interpretable metric to precisely quantify structural similarity between expressions.  With this benchmark, we explore mainstream LLM fine-tuning techniques for SR tasks and establish SymbolicChat, a simple yet effective LLM-based SR strong baseline. Experimental results validate SymbolicChat as the first LLM to exceed traditional numerical methods in both numerical precision and symbolic form accuracy, outperforming  the second-best LLM baseline with improvements of 2-fold gains in R^2 score and 8.37% in form-level consistency score."},"_bibtex":{"value":"@misc{\nhua2025finetuning,\ntitle={Finetuning Large Language Model as an Effective Symbolic Regressor},\nauthor={Yingfan Hua and Ruikun Li and Jun Yao and Guohang Zhuang and SHIXIANG TANG and Bin Liu and Wanli Ouyang and Yan Lu},\nyear={2025},\nurl={https://openreview.net/forum?id=OjRaJw4tnr}\n}"},"title":{"value":"Finetuning Large Language Model as an Effective Symbolic Regressor"},"pdf":{"value":"/pdf/fd9582ca3f87537f5b3067f1ad00438dc4e17c10.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hua|finetuning_large_language_model_as_an_effective_symbolic_regressor"},"authorids":{"value":["~Yingfan_Hua1","~Ruikun_Li3","~Jun_Yao3","~Guohang_Zhuang1","~SHIXIANG_TANG1","~Bin_Liu5","~Wanli_Ouyang1","~Yan_Lu10"]},"authors":{"value":["Yingfan Hua","Ruikun Li","Jun Yao","Guohang Zhuang","SHIXIANG TANG","Bin Liu","Wanli Ouyang","Yan Lu"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026 Workshop SynData4CV"},"pdf":{"value":"/pdf/1e79053267b790d2c0336b1171d47a0b988fe156.pdf"},"keywords":{"value":["3D Vision","Synthetic Data","Tree Segmentation","LiDAR Simulation","AI for Science"]},"venueid":{"value":"thecvf.com/CVPR/2026/Workshop/SynData4CV"},"paperhash":{"value":"she|scaling_up_3d_forest_vision_with_synthetic_lidar","readers":["everyone"]},"authorids":{"readers":["everyone"],"value":["~Yihang_She1","~Andrew_Blake3","~David_Coomes1","~Srinivasan_Keshav1"]},"abstract":{"value":"Accurate tree segmentation from forest laser scans is essential to understanding ecosystem functions in carbon cycling and beyond. We have developed a new synthetic data generation pipeline to do this for forest vision tasks, integrating advances in game-engines with physics-based LiDAR simulation. As a result, we have produced a comprehensive, diverse, annotated 3D forest dataset on an unprecedented scale.\n\nExtensive experiments with a state-of-the-art tree segmentation algorithm and a popular real dataset show that our synthetic data can substantially reduce the need for labelled real data. After fine-tuning on just a single, real, forest plot of less than 0.1 hectare, the pretrained model achieves segmentations that are competitive with a model trained on the full scale real data. We have also identified critical factors for successful use of synthetic data: physics, diversity, and scale, paving the way for more robust 3D forest vision systems in the future."},"title":{"value":"Scaling Up 3D Forest Vision with Synthetic LiDAR"},"authors":{"readers":["everyone"],"value":["Yihang She","Andrew Blake","David Coomes","Srinivasan Keshav"]}},"tmdate":1782159697280,"tcdate":1773748344103,"writers":["thecvf.com/CVPR/2026/Workshop/SynData4CV","thecvf.com/CVPR/2026/Workshop/SynData4CV/Submission63/Authors"],"signatures":["thecvf.com/CVPR/2026/Workshop/SynData4CV/Submission63/Authors"],"forum":"MlqsJmdVTd","license":"CC BY 4.0","number":63,"cdate":1773748344103,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Workshop/SynData4CV/-/Submission","thecvf.com/CVPR/2026/Workshop/SynData4CV/-/Post_Submission","thecvf.com/CVPR/2026/Workshop/SynData4CV/-/Edit"],"mdate":1782159697280,"domain":"thecvf.com/CVPR/2026/Workshop/SynData4CV","id":"MlqsJmdVTd","version":2},{"content":{"venue":{"value":"2025 IEEE European Conference on Antenna and Propagation"},"pdf":{"value":"/pdf/d9c6ae91fb5f66e097259351d5b67193121254af.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"pan|physicsinformed_graph_neural_networks_for_the_inverse_design_of_ghz_reconfigurable_antenna"},"authorids":{"value":["~Cindy_Pan1","~Naveen_Verma1","~James_Sturm1"]},"abstract":{"value":"Reconfigurable antennas, as a subclass of meta-\nsurfaces, offer innovative and dynamic capabilities for wireless\ncommunication systems. Specifically, enabling radiation pat-\ntern reconfigurability allows for flexible beam steering through\nreverse-engineering of antenna parameters such as surface cur-\nrent distributions. In this work, we present a physics-informed\nmachine learning model, leveraging fundamental physics such as\nKirchhoff’s current Law, to predict the switch configurations of\n2-dimensional antenna arrays. We utilize a graph neural network\n(GNN) to effectively capture the spatial relationships between\nradio-frequency (RF) switches and antenna patches, closely\nemulating the antenna topology. Simulation results demonstrate\nthat our approach successfully predicts switch configurations\nneeded to generate complex far-field radiation patterns."},"title":{"value":"Physics-Informed Graph Neural Networks for the Inverse Design of GHz Reconfigurable Antenna"},"authors":{"value":["Cindy Pan","Naveen Verma","James Sturm"]}},"tmdate":1747946092222,"pdate":1747810800000,"tcdate":1745964638746,"writers":["~Cindy_Pan1","~Naveen_Verma1","~James_Sturm1"],"signatures":["~Cindy_Hsin_Pan1"],"forum":"7kh9q0U4PZ","license":"CC BY 4.0","number":35745,"cdate":1745964638746,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1747946092222,"domain":"OpenReview.net/Archive","id":"7kh9q0U4PZ","version":2},{"content":{"summary":{"value":"The paper introduces TRENDy, a framework designed for learning low-dimensional surrogates of complex dynamical systems where underlying PDEs are unknown, and data is noisy or limited. The main contributions of TRENDy include:\n\n1. Modeling Effective Dynamics: TRENDy maps underlying PDE using multiscale filtering into a reduced space and models the reduced representation with a neural ODE. This NODE captures the system's behavior based on its governing parameters, which enables TRENDy to predict system dynamics in new parameter spaces.\n\n2. Predicting Bifurcation: moreover, for a parameter-dependent system, the framework is able to predict bifurcations, where a sudden qualitative change in its behaviors. TRENDy also shows robustness to noise in bifurcation localization.\n\n3. Application to Real-World Data: authors also used the patterning in the ocellated lizard as an example to illustrate how the framework's latent space captures meaningful biological features."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Besides several concerns that mentioned in the weakness part, here are several questions regarding the paper details:\n\nFigure 1: what is $S_i(0)$ here? And why do they have different heights?\n\nEq 1: it is better to use $u(x, y)$ rather than $u(r)$, as you are talking about 2D space now.\n\nUsage of subscripts (Line 138 and other notations): the subscripts sometimes are very misleading. e.g. $u_{\\theta}$ and $u_0$.\n\nLine 139: for $u_0 \\notin D$, do you mean interpolation and/or exploration?\n\nLine 141: needs explanation of what $U$ is.\n\nLine 148: similar as what mentioned previously, why do you assume $\\Phi$ is hardwired and unlearned? Can the multiscale filtering parameters be learnable?\n\nFigure 2: this figure needs more details to explain. For example, you should say the inset squares means PDE solutions (otherwise it is misleading).\n\nLine 207: the approximately equal symbol here is incorrect. And moreover, what is SINDyCP? Formula of it? Does it have any assumptions? Have you cited it?\n\nLines 230-232: may need a figure to illustrate your conditions. For example, what is “patches”?\n\nLine 266: need to show why $S_{1, 2}$ almost equals to $<u>$ and $<v>$.\n\nLine 306: use def eq symbol “:=” here in $d_{\\gamma} (\\theta)$."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Originality: \n\nThe approach that combines scattering transform and neural ODEs to model the effective dynamics is novel, especially given its application to bifurcation prediction, a challenging task where data is limited and governing equations are unknown.\n\n2. Quality and Clarity: \n\nThe paper shows rigorous methodology and fruitful details in various experiments. Explanations on filtering operations, the NODE structure, and training details led the model's design to be crafty and reliable. \n\n3. Significance: \n\nTRENDy addresses a crucial question in modeling systems governed by unknown or complex PDEs, where direct analytic solutions are impractical to get. The framework’s adaptability to new parameter spaces may also have numerous applications in real-time system control and scenario exploration. In a nutshell, the authors have shown that TRENDy has the potential to significantly advance research in fields like synthetic biology, physics, climate change and ecology, where such questions regarding complex dynamical systems are pretty common."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Multiscale Filtering:\n\nThe use of multiscale filtering (e.g., scattering transforms) is central to TRENDy, while the specific choice and design of the filtering process are not fully explored in the paper. Authors should provide more why they prefer this type of dimension reduction technique rather than others (for example, do ablation studies on other type of techs and show the one you mentioned is the best). Moreover, compared with too  many experimental details in the main text (better go to supporting materials), it is necessary to say more on multiscale filtering details, e.g. effects of the choice on scattering coefficients. Such explanations / experiments are essential to keep novelty of the paper, since they are numerous papers working on PDE + DL topics (and some of them should be acknowledged, e.g. PDE-net by Long et al. [1], PINNs by Karniadakis et al. [2], and other papers focusing on effective dynamics, see [3] and [4]). \n\n2. Reconstructing State Space:\n\nJust like lifting and restriction in the equation-free approach, TRENDy should have the module which maps the latent dimensions back to the full PDE state space. Without such an explicit decoder, the ability to verify the reduced dynamics against full state predictions will be limited. It will also become an obstacle for researchers in other fields to explore the explainability by utilizing your model. It seems adding a mechanism for decoding reduced dynamics back into full spatial states or maybe explaining why this is not feasible in your scope is essential.\n\n3. Miscellaneous:\n\nI suggested the reviewers consider the following issues, and if time allows, do some elaboration.\n\na) Extending the experimental scope (e.g. systems with chaotic attractors, or discrete-time systems).\n\nb) Discussing the model’s performance on large datasets and its computational demands in both training and inference.\n\nc) Implementing interpretability techniques (e.g., parameter sensitivity, feature importance) to provide insights on multiscale filtering.\n\n[1] https://arxiv.org/pdf/1710.09668\n\n[2] https://arxiv.org/pdf/1711.10561 \n\n[3] https://www.nature.com/articles/s41467-024-48024-7 \n\n[4] https://pubs.aip.org/aip/cha/article-abstract/34/6/063128/3298062/Tipping-points-of-evolving-epidemiological?redirectedFrom=fulltext"}},"nonreaders":[],"tmdate":1731429027780,"tcdate":1730699534141,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9948/Reviewer_ahWW"],"signatures":["ICLR.cc/2025/Conference/Submission9948/Reviewer_ahWW"],"forum":"NvDRvtrGLo","number":3,"license":"CC BY 4.0","cdate":1730699534141,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9948/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429027780,"domain":"ICLR.cc/2025/Conference","replyto":"NvDRvtrGLo","id":"q1AQxuqOZS","forumContent":{"TLDR":{"value":"We learn reduced order models of PDEs for robust bifurcation prediction."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["dynamical systems; neural ODEs","representation learning"]},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Spatiotemporal dynamics pervade the natural sciences, from the morphogen dynamics underlying patterning in animal pigmentation to the protein waves controlling cell division. A central challenge lies in understanding how controllable parameters induce qualitative changes in system behavior called bifurcations. This endeavor is particularly difficult in realistic settings where governing partial differential equations (PDEs) are unknown and data is limited and noisy. To address this challenge, we propose TRENDy (Temporal Regression of Effective Nonlinear Dynamics), an equation-free approach to learning low-dimensional, predictive models of spatiotemporal dynamics. TRENDy first maps input data to a low-dimensional space of effective dynamics through a cascade of multiscale filtering operations. Our key insight is the recognition that these effective dynamics can be fit by a neural ordinary differential equation (NODE) having the same parameter space as the input PDE. The preceding filtering operations strongly regularize the phase space of the NODE, making TRENDy significantly more robust to noise compared to existing methods. We train TRENDy to predict the effective dynamics of synthetic and real data representing dynamics from across the physical and life sciences. We then demonstrate how we can automatically locate both Turing and Hopf bifurcations in unseen regions of parameter space. We finally apply our method to the analysis of spatial patterning of the ocellated lizard through development. We found that TRENDy's predicted effective state not only accurately predicts spatial changes over time but also identifies distinct pattern features unique to different anatomical regions, such as the tail, neck, and body--an insight that highlights the potential influence of surface geometry on reaction-diffusion mechanisms and their role in driving spatially varying pattern dynamics."},"_bibtex":{"value":"@inproceedings{\nricci2025trendy,\ntitle={{TREND}y: Temporal Regression of Effective Nonlinear Dynamics},\nauthor={Matt Ricci and Guy Pelc and Zoe Piran and Noa Moriel and Mor Nitzan},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=NvDRvtrGLo}\n}"},"title":{"value":"TRENDy: Temporal Regression of Effective Nonlinear Dynamics"},"pdf":{"value":"/pdf/79bb2478098db04bc2f1d00a919f80e55d70bc9d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"ricci|trendy_temporal_regression_of_effective_nonlinear_dynamics"},"authorids":{"value":["~Matt_Ricci1","~Guy_Pelc1","~Zoe_Piran1","~Noa_Moriel1","~Mor_Nitzan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Matt Ricci","Guy Pelc","Zoe Piran","Noa Moriel","Mor Nitzan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a method to stabilize the training of Mixture-of-Experts (MoE) models.\nIn MoE training, experts may compete excessively or fail to cooperate effectively. To address this issue, the authors introduce a probabilistic expert-assignment distribution that blends cooperative and competitive mechanisms. They propose an EM-based algorithm to optimize the parameters of this distribution and describe its efficient implementation on GPUs. Experiments are conducted on synthetic datasets and within the physics-informed neural network (PINN) framework, demonstrating the effectiveness of the proposed method."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. How does the computational cost of the proposed method scale with respect to data dimensionality, dataset size, and the depth of the MoE hierarchy?\n2. Are there any theoretical guarantees on the performance of the MoE trained by the proposed method? Even in a toy setting, it would be helpful to clarify under what conditions the proposed method outperforms existing approaches.\n3. Is there any comparison with naive MoE training or with prior work that also attempted to stabilize MoE learning?\n4. In line 281, it is stated that *“the hierarchical experts allow for an unsupervised concentration of resolution around the relevant feature of the problem.”* — how can this be inferred from Figure 2?\n5. In Figure 5, the input dimensionality is up to 128. Do you have experimental results for higher-dimensional settings?\n6. Why was PINN chosen as the application task? Could you also demonstrate the method on other types of problems?\n7. Are there any prior studies applying MoE to PINNs? If so, can you compare your results in Section 4.5 with those works?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- The proposed approach is motivated by a convincing idea — balancing cooperative and competitive terms in MoE learning.\n- The effectiveness of the method is evaluated from multiple perspectives, including convergence stability, performance on synthetic data, GPU-based speedup, and sensitivity to hyperparameters."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is no comparison with existing methods. Both the theoretical and experimental sections focus solely on the proposed method, without sufficient evaluation of how its convergence speed or training stability compares to prior work.\n2. The algorithm appears complex, and it is unclear whether it can be applied effectively to large-scale tasks. In addition, the only application shown experimentally is the PINN problem."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933960220,"tcdate":1761648377122,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20540/Reviewer_agLC"],"signatures":["ICLR.cc/2026/Conference/Submission20540/Reviewer_agLC"],"forum":"SbMWnYTJNj","number":1,"license":"CC BY 4.0","cdate":1761648377122,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20540/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933960220,"domain":"ICLR.cc/2026/Conference","replyto":"SbMWnYTJNj","id":"OzIYjyWxjy","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["hierarchical mixture of experts","expectation maximization","optimization"]},"supplementary_material":{"value":"/attachment/c726f1c475773057f5e3b72db53889e0b1dfc518.zip"},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"abstract":{"value":"We present a novel probabilistic expectation-maximization scheme for training hierarchical mixture-of-experts models that both exposes and exploits parallelism during training. By replacing the typical categorical distribution used in gating networks with a joint distribution blending cooperative and competitive mechanisms, we obtain a likelihood that encodes both global and local interactions between experts. The application of an M-splitting scheme reveals an M-step that enables the solution of localized, embarrassingly parallel subproblems governing local experts, with deferred corrections accounting for global coupling between experts. When combined with a hierarchical decomposition of nested networks, this yields a fast multi-level training scheme reminiscent of multigrid algorithms, which avoids under-utilization of experts, exposes further GPU parallelism and outperforms standard models on regression tasks. We provide experiments using a scalable GPU implementation that demonstrate rapid convergence and parallel scalability of the iterative scheme, as well as strong localization of the model for non-smooth, high-dimensional regression problems."},"_bibtex":{"value":"@misc{\nanonymous2026a,\ntitle={A scalable cooperative/competitive splitting scheme for mixture of experts models},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=SbMWnYTJNj}\n}"},"title":{"value":"A scalable cooperative/competitive splitting scheme for mixture of experts models"},"pdf":{"value":"/pdf/e1468f6e6b92d46c4eb13c3dba142fb92f8447a0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"nguyenvu|a_scalable_cooperativecompetitive_splitting_scheme_for_mixture_of_experts_models"},"authorids":{"value":["~Hien_T._Nguyen-Vu1","~Alexey_Voronin1","~Shyam_Sankaran1","~Eric_C_Cyr1","~Nathaniel_Trask2"]},"authors":{"value":["Hien T. Nguyen-Vu","Alexey Voronin","Shyam Sankaran","Eric C Cyr","Nathaniel Trask"]}},"version":2},{"content":{"comment":{"value":"Thank you very much for your detailed review and insightful comments. We would like to address your concerns.\n\n**Weakness: Synthetic setting and Question 2**\n>In the synthetic setting, the added context tokens actually depend on the specific inputs and functions used in the example. On the other hand, in the real-world setting, the additional context tokens are sourced from the educational prompts and randomly sampled without considering he contents of a particular example. To summarize, it appears to me that contextual fine-tuning in the synthetic experiments actually provides a significant source of additional supervision (reminiscent of COT approaches) whereas this is absent in real-world instantiation of contextual FT.\n\n>I would like it if the authors could further justify how the synthetic data setup should be viewed as comparable to the real-world setup. In particular, why are the additional tokens added in the synthetic settings input dependent while the contextual prompts used in real settings are randomly sampled independently of the contents of the document/example?\n\nOur current hypothesis for why our approach works is that gradients under prompts that contain semantic content relevant for learning serve to regularize the process of learning via fine-tuning. However testing this hypothesis directly is challenging since (a) different LLMs might interpret semantic information in a prompt differently (as a function of scale) and (b) it requires knowing which neurons are responsible for representing the inferred semantic information in the prompt -- an open problem in mechanistic interpretability.\n\nTo that end the primary objective of the synthetic experiment was to analyze how contextual prompts affect the gradients of transformer models during training in a controlled setting where we can describe the semantic information that is necessary for learning explicitly via text. The advantage of this is that it enables us to not worry about how the transformer encodes semantic information (thus enabling the study of this phenomenon on much smaller models) and consequently better understand what properties of the gradient enable this.\nTo expand on this further, the sequence of tokens we use in the synthetic data, by design, $(x_1,f(x_1),x_2,f(x_2),\\ldots,x_k,f(x_k))$ encode the semantic information necessary for learning this synthetic class of problem, facilitated by conditioning on the prompts.\n\nOur empirical results, presented in Appendix G, show that contextual fine-tuning is more effective for instruction-tuned and chat models compared to non-chat models. This observation suggests that models capable of following instructions are better at leveraging contextual prompts during fine-tuning, even when the prompts are not customized to each example. Our intention with the synthetic experiment was to provide insight into the potential mechanisms by which contextual prompts can enhance learning, acknowledging that direct analysis of gradients in large-scale language models is infeasible."},"title":{"value":"Response to Weakness: Synthetic Setting and Question 2"}},"tmdate":1732332970175,"tcdate":1732332970175,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11845/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission11845/Authors"],"forum":"FS2nukC2jv","number":4,"license":"CC BY 4.0","cdate":1732332970175,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11845/-/Official_Comment"],"mdate":1732332970175,"domain":"ICLR.cc/2025/Conference","replyto":"FN368F55fH","id":"ar9l6pAq2C","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large Language Models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Prompting Large Language Models (LLMs), or providing context on the expected model of operation, is an effective way to steer the outputs of such models to satisfy human desiderata after they have been trained. But in rapidly evolving domains, there is often need to fine-tune LLMs to improve either the kind of knowledge in their memory or their abilities to perform open ended reasoning in new domains. When human's learn new concepts, we often do so by linking the new material that we are studying to concepts we have already learned before. To that end, we ask, \"can prompting help us teach LLMs how to learn\". In this work, we study a novel generalization of instruction tuning, called contextual fine-tuning, to fine-tune LLMs. Our method leverages instructional prompts designed to mimic human cognitive strategies in learning and problem-solving to guide the learning process during training, aiming to improve the model’s interpretation and understanding of domain-specific knowledge. We empirically demonstrate that this simple yet effective modification improves the ability of LLMs to be fine-tuned rapidly on new datasets both within the medical and financial domains."},"_bibtex":{"value":"@inproceedings{\nchoi2025teaching,\ntitle={Teaching {LLM}s How to Learn with Contextual Fine-Tuning},\nauthor={Younwoo Choi and Muhammad Adil Asif and Ziwen Han and John Willes and Rahul Krishnan},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=FS2nukC2jv}\n}"},"title":{"value":"Teaching LLMs How to Learn with Contextual Fine-Tuning"},"pdf":{"value":"/pdf/8634ecdc5fa1f832a485080c138c1492fef4d902.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"choi|teaching_llms_how_to_learn_with_contextual_finetuning"},"authorids":{"value":["~Younwoo_Choi1","~Muhammad_Adil_Asif1","~Ziwen_Han1","~John_Willes2","~Rahul_G_Krishnan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Younwoo Choi","Muhammad Adil Asif","Ziwen Han","John Willes","Rahul Krishnan"]}},"version":2},{"content":{"summary":{"value":"The authors propose a framework to connect and unify several learning settings through the language of classical physics, specifically the principle of least action. The framework describes learning as the minimisation of a Lagrangian where, in place of physical notions of position and time, the authors instead consider a dataset that is iteratively added to. They provide three applications of this framework—to experimental design, reinforcement learning, and optimisation—recovering classical machine learning results such as A-optimality, Bellman’s optimality equation, and natural gradient descent."},"soundness":{"value":1},"confidence":{"value":3},"questions":{"value":"- Could the authors please provide a derivation for the equality on the right-hand side of equation (5)?\n- Are the authors familiar with the online-to-PAC framework (Lugosi and Neu, 2023)? It might be relevant to their work.\n- I found section 3.3 very difficult to follow. Could the authors please clearly explain this, taking greater care to state what things are assumed/hypothesized and which parts are mathematical deductions?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- The paper identifies a connection between physics and machine learning, two fields of interest to the conference.\n- The authors attempt to unify concepts in experimental design, reinforcement learning, and second-order optimisation. In this sense, the work is refreshingly ambitious.\n- The proposed formulation of active learning/experimental design via physical analogues appears to be a novel perspective, though I am not familiar enough with this specific literature to comment on its novelty with confidence."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper’s clarity is a significant concern. Many explanations, especially in Section 3, feel underdeveloped or are presented too hastily. The terms “postulate”, “conjecture”, and “hypothesise” are used frequently, often immediately before a strong claim is made, which blurs the line between assumption and proof.\n- The paper uses physics notation and ideas without sufficient exposition. A large proportion of the ICLR audience has a background in Mathematics and Computer Science; notation that is not standard in these fields should be clearly defined. I would encourage the authors to add a dedicated background/notation section.\n- The paper's motivation is somewhat puzzling. The authors claim a \"surprising link\" between physics and ML, yet also describe how important physics has been to ML's development. The fact that different learning problems can be cast as optimisation problems (minimising a Lagrangian) is not suprising given the role that classical physics has played in this development.\n- The contribution seems overstated on several occasions. The authors claim their work provides physical justification for the Adam optimiser, but the analysis is for natural gradient descent, which is theoretically distinct. Similarly, the claim of applications to generative models is justified by this same NGD analysis. This connection feels tenuous and appears to overstate the framework's relevance to contemporary generative models.\n- \"Insight No. 1\" (that learning is a decelerating process) is a well-known property of learning curves and may not constitute a novel insight of this framework.\n- The \"Main Postulate\" is conceptually challenging. It states that learning is a search for a stationary path (data sequence), but in most learning settings (outside of experimental design), the data path is fixed, and the search is over the model class.\n- The paper's contribution appears thin for a main-track submission. It lacks novel theoretical guarantees or new experimental results. The primary contribution is the proposal of a framework and various conjectures. However, the framework is presented in a way that is difficult to follow, which unfortunately detracts from its value as a contribution.\n- The equality on the right-hand side of (5) doesn’t appear to be correct (see questions)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928103594,"tcdate":1762299597960,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18398/Reviewer_Szcw"],"signatures":["ICLR.cc/2026/Conference/Submission18398/Reviewer_Szcw"],"forum":"FUEzlNM4jx","number":4,"license":"CC BY 4.0","cdate":1762299597960,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18398/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928103594,"domain":"ICLR.cc/2026/Conference","replyto":"FUEzlNM4jx","id":"CeeKw1jmIV","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"A physics perspective in efficient learning"},"keywords":{"value":["physics; learning; reinforcement learning; generative models; learning theory"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"We study the problem of building an efficient learning system. Efficient learning processes information in the least time, i.e., building a system that reaches a desired error threshold with the least number of observations. Building upon least action principles from physics, we derive classic learning algorithms, Bellman's optimality equation in reinforcement learning, and the Adam optimizer in generative models from first principles, i.e., the Learning $\\textit{Lagrangian}$. We postulate that learning searches for stationary paths in the Lagrangian, and learning algorithms are derivable by seeking the stationary trajectories."},"_bibtex":{"value":"@misc{\nguo2026physics,\ntitle={Physics of Learning: A Lagrangian perspective to different learning paradigms},\nauthor={Siyuan Guo and Bernhard Sch{\\\"o}lkopf},\nyear={2026},\nurl={https://openreview.net/forum?id=FUEzlNM4jx}\n}"},"title":{"value":"Physics of Learning: A Lagrangian perspective to different learning paradigms"},"pdf":{"value":"/pdf/d62dd55055d859ddad558fe8884679035dad36b9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"guo|physics_of_learning_a_lagrangian_perspective_to_different_learning_paradigms"},"authorids":{"value":["~Siyuan_Guo1","~Bernhard_Schölkopf1"]},"authors":{"value":["Siyuan Guo","Bernhard Schölkopf"]}},"version":2},{"content":{"summary":{"value":"The paper proposes two low level tests for intuitive physics in MLLMs, Next Frame Selection and Temporal Coherence Verification, focused mainly on fluids. It introduces Scene Dynamic Field, a simulator derived motion visualization used as a visual prompt within a multi task fine tuning scheme, and reports sizable gains on the proposed tests with some transfer to cloth, sand, smoke, and plasticine."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How does SDF compare to simple optical flow overlays, grayscale motion magnitude, or event frame stacks when training with the same budget?\n\n- Are results robust when distractors are generated with a feature space disjoint from any model under test?\n\n- Can you evaluate on external physics benchmarks without re curating the data to validate generality?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Clear problem decomposition toward low level dynamics rather than high level QA\n- Simple intermediate representation that is easy to plug into existing MLLMs\n- Ablations on stride, prompts, model scale, and expert vs self distilled data\n- Some transfer beyond fluids and an attention analysis that supports the claim that SDF shifts focus to earlier frames"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The weakest point is that there is no comparison to strong motion baselines. SDF is a velocity magnitude style visual prompt, which is conceptually close to optical flow magnitude, flow stacks, dynamic images, or even simple frame differencing. Without head to head baselines under the same training data and budget, the gains could come from adding any explicit motion cue rather than from SDF itself. This leaves the central claim unproven.\n\n- Benchmarks are author designed and multiple choice, so improvements may reflect distractor design rather than genuine physics understanding\n\n- Absolute accuracy remains low, so practical impact is unclear\n\n- Limited evaluation beyond fluids for rigid body scenes or causal reasoning tasks, so the title and claims feel broader than what is shown\n- Distractor pruning uses SigLIP embeddings which are related to encoders used by evaluated models, creating a risk of bias\n\nMy recommendation is reject. The idea is interesting and the empirical gains are clear on the authors benchmark, but the evaluation misses strong motion baselines, relies on potentially biased distractor construction, and the absolute performance and scope do not yet support the paper’s broad claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360114794,"tcdate":1761847611971,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7741/Reviewer_Vi2b"],"signatures":["ICLR.cc/2026/Conference/Submission7741/Reviewer_Vi2b"],"forum":"Ax02eR2c3d","number":2,"license":"CC BY 4.0","cdate":1761847611971,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7741/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360114794,"domain":"ICLR.cc/2026/Conference","replyto":"Ax02eR2c3d","id":"9jqfHq2fY4","forumContent":{"TLDR":{"value":"This paper introduces two low-level tasks to test intuitive physics understanding and proposes Scene Dynamic Field, a method to integrate visual representation from physics simulators to MLLMs while showcasing generalization."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multi-modal LLM","Intuitive Physics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. \nIn this work, we investigate the first step of physical reasoning, i.e., **intuitive physics understanding**, revealing substantial limitations in understanding the dynamics of continuum objects. \nTo isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. \nTo address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. \nSDF substantially improves performance, achieving up to $20.7\\%$ gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field."},"_bibtex":{"value":"@inproceedings{\nli2026beyond,\ntitle={Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models},\nauthor={Nanxi Li and Xiang Wang and Yuanjie Chen and Haode Zhang and Hong Li and Yong-Lu Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Ax02eR2c3d}\n}"},"title":{"value":"Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models"},"pdf":{"value":"/pdf/3b199b74bbdbddb8166b5e799fc00ddb71a413ae.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|beyond_static_vision_scene_dynamic_field_unlocks_intuitive_physics_understanding_in_multimodal_large_language_models"},"authorids":{"value":["~Nanxi_Li1","~Xiang_Wang19","~Yuanjie_Chen2","~Haode_Zhang2","~Hong_Li5","~Yong-Lu_Li1"]},"authors":{"value":["Nanxi Li","Xiang Wang","Yuanjie Chen","Haode Zhang","Hong Li","Yong-Lu Li"]}},"version":2},{"content":{"TLDR":{"value":"We present GPhyT, a transformer trained on 1.8TB of diverse physics simulations that can zero-shot generalize to entirely new physical systems by inferring governing dynamics from context alone."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics Foundation Model","Multi-physics Learning","In-context Learning","Zero-shot Generalization","Scientific Machine Learning","Physics-Aware Machine Learning","Spatiotemporal Transformers"]},"supplementary_material":{"value":"/attachment/bb7557fc3f8da9f1e6e373f392add528e405a51d.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Foundation models have revolutionized natural language processing through a ``train once, deploy anywhere'' paradigm, where a single pre-trained model adapts to countless downstream tasks without retraining. Access to a Physics Foundation Model (PFM) would be transformative - democratizing access to high-fidelity simulations, accelerating scientific discovery, and eliminating the need for specialized solver development. Yet current physics-aware machine learning approaches remain fundamentally limited to single, narrow domains and require retraining for each new system. We present the General Physics Transformer (GPhyT), trained on 1.8 TB of diverse simulation data, that demonstrates foundation model capabilities are achievable for physics. Our key insight is that transformers can learn to infer governing dynamics from context, enabling a single model to simulate fluid-solid interactions, shock waves, thermal convection, and multi-phase dynamics without being told the underlying equations. GPhyT achieves three critical breakthroughs: (1) superior performance across multiple physics domains, outperforming specialized architectures by more than 7x, (2) plausible zero-shot generalization to entirely unseen physical systems through in-context learning, and (3) more stable long-term predictions through long-horizon rollouts. By establishing that a single model can learn generalizable physical principles from data alone, this work opens the path toward a universal PFM that could transform computational science and engineering."},"_bibtex":{"value":"@misc{\nwiesner2026towards,\ntitle={Towards a Physics Foundation Model},\nauthor={Florian Wiesner and Matthias Wessling and Stephen Baek},\nyear={2026},\nurl={https://openreview.net/forum?id=q62POvqTLb}\n}"},"title":{"value":"Towards a Physics Foundation Model"},"pdf":{"value":"/pdf/a4f22c125a4e4ee33cdd9e42a7b76196da8bf16f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wiesner|towards_a_physics_foundation_model"},"authorids":{"value":["~Florian_Wiesner1","~Matthias_Wessling1","~Stephen_Baek1"]},"authors":{"value":["Florian Wiesner","Matthias Wessling","Stephen Baek"]}},"tmdate":1770804915768,"tcdate":1758092344077,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8602/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission8602/Authors"],"forum":"q62POvqTLb","license":"CC BY 4.0","number":8602,"cdate":1758092344077,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission8602/-/Full_Submission","ICLR.cc/2026/Conference/Submission8602/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770804915768,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"q62POvqTLb","version":2},{"content":{"TLDR":{"value":"We add computational chemistry data to SAIR, create various splits of the data pertinent to drug discovery campaign, and compare different model approaches to predict experimental binding affinities of protein-ligand systems."},"venue":{"value":"GEM 2026"},"keywords":{"value":["binding affinity prediction","protein-ligand interaction","molecular graphs","drug discovery","geometric deep learning"]},"Funding":{"value":"No, the presenting author of this submission does not fall under ICLR’s funding aims, or has sufficient alternate funding."},"abstract":{"value":"The success of deep learning binding affinity prediction models depends critically on expanding experimental data with reliable synthetic data. We extend the Structurally Augmented IC50 Repository (SAIR) with physics-based computations and present two distinct data splits, SAIR-FEP and SAIR-OOD. With SAIR-FEP, we perform $\\approx$80K absolute free energy perturbation calculations (AFEP) and curate two train/test splits to simulate realistic drug discovery scenarios. The free energy of binding and other physics-based computations are then used as either input features. We compare the performance of proteochemometric and state-of-the-art structure-based deep learning models and show that including physics-based features improves predictions, and that the quality of the structure plays a key role in their performance. For SAIR-OOD, we remove SAIR entries that overlap with complexes in public-facing benchmarks and demonstrate that simultaneous training on synthetic and experimental data improves performance on public-facing, experimental benchmarks."},"_bibtex":{"value":"@inproceedings{\nryczko2026on,\ntitle={On improving experimental binding affinity predictions with synthetic data},\nauthor={Kevin Ryczko and Phyo Phyo Kyaw Zin and Jordan Crivelli-Decker and Ly Le and Punit K Jha and Benjamin J. Shields and Pablo Lemos and Sasaank Bandi and Maarten Van Damme and Martin Ganahl and Andrea Bortolato},\nbooktitle={ICLR 2026 Workshop on Generative and Experimental Perspectives for Biomolecular Design},\nyear={2026},\nurl={https://openreview.net/forum?id=aiQyNhZ3s5}\n}"},"title":{"value":"On improving experimental binding affinity predictions with synthetic data"},"pdf":{"value":"/pdf/13a2129b6d74e0f06daf335749f31d3e94d9bb92.pdf"},"Format":{"value":"Yes, the presenting author will attend in person if this work is accepted to the workshop."},"Presenter":{"value":"~Kevin_Ryczko1"},"venueid":{"value":"ICLR.cc/2026/Workshop/GEM"},"paperhash":{"value":"ryczko|on_improving_experimental_binding_affinity_predictions_with_synthetic_data"},"authorids":{"value":["~Kevin_Ryczko1","~Phyo_Phyo_Kyaw_Zin1","~Jordan_Crivelli-Decker1","~Ly_Le1","~Punit_K_Jha1","~Benjamin_J._Shields1","~Pablo_Lemos1","~Sasaank_Bandi1","~Maarten_Van_Damme1","~Martin_Ganahl1","~Andrea_Bortolato2"]},"authors":{"value":["Kevin Ryczko","Phyo Phyo Kyaw Zin","Jordan Crivelli-Decker","Ly Le","Punit K Jha","Benjamin J. Shields","Pablo Lemos","Sasaank Bandi","Maarten Van Damme","Martin Ganahl","Andrea Bortolato"]}},"tmdate":1779832645907,"pdate":1772423420071,"tcdate":1770329237884,"writers":["ICLR.cc/2026/Workshop/GEM","ICLR.cc/2026/Workshop/GEM/Submission53/Authors"],"signatures":["ICLR.cc/2026/Workshop/GEM/Submission53/Authors"],"forum":"aiQyNhZ3s5","license":"CC BY 4.0","number":53,"cdate":1770329237884,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/GEM/-/Submission","ICLR.cc/2026/Workshop/GEM/-/Post_Submission","ICLR.cc/2026/Workshop/GEM/-/Edit"],"mdate":1779832645907,"odate":1772673011862,"domain":"ICLR.cc/2026/Workshop/GEM","id":"aiQyNhZ3s5","version":2},{"content":{"summary":{"value":"This paper tackles the challenge of understanding complex long videos with sparse task-relevant information. To improve AI reasoning, the paper propose CogniGPT, a framework inspired by human visual cognition. CogniGPT combines a perception agent (MPGA) to focus on task-relevant details, and a reflection agent (VERA)  to verify key information, reducing errors and improving efficiency."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weaknesses above."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"- The paper is well-written and easy to follow\n\n- The figures are intuitive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The proposed method is straightforward and intuitive. \n\n - As an LLM-based agent, its performance is inferior to single models like Qwen2.5-VL and significantly worse than VideoChat-A1[1], which uses VL models. This raises the question of why such a complex LLM agent is needed when simpler and more effective alternatives exist.\n\n- While it uses fewer frames, the claim of \"selecting key information\" is not convincingly supported by the results. \n\n- The table formatting in the paper is poorly organized.\n\n- The paper should use \\citep instead of \\cite for citation\n\n\n[1] Wang Z, Chen B, Yue Z, et al. VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning[J]. arXiv preprint arXiv:2506.06097, 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919069318,"tcdate":1761925524598,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6800/Reviewer_m7vg"],"signatures":["ICLR.cc/2026/Conference/Submission6800/Reviewer_m7vg"],"forum":"klRdsrja2h","number":4,"license":"CC BY 4.0","cdate":1761925524598,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6800/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919069318,"domain":"ICLR.cc/2026/Conference","replyto":"klRdsrja2h","id":"pfb90Sr2QQ","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Long Video Understanding","LLM Agents"]},"supplementary_material":{"value":"/attachment/9ea01fdfe779771cbf51b0c64958a076e87a83ed.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long videos, characterized by temporal complexity and sparse task-relevant information, pose significant reasoning challenges for AI systems. Although various Large Language Model (LLM)-based approaches have advanced long video understanding, they still struggle to achieve both completeness and efficiency in capturing task-critical information. Inspired by human progressive visual cognition, we propose CogniGPT, a framework that leverages an interactive loop between Multi-Granular Perception Agent (MGPA) and Verification-Enhanced Reflection Agent (VERA) for efficient and reliable long video understanding. Specifically, MGPA mimics human visual divergent and focused attention to capture task-related information, while VERA verifies perceived key clues to mitigate hallucination and optimize subsequent perception strategies. Through this interactive process, CogniGPT explores a minimal set of informative and reliable task-related clues.\nExtensive experiments on EgoSchema, Video-MME, NExT-QA, and MovieChat datasets demonstrate CogniGPT's superiority in both accuracy and efficiency. Notably, on EgoSchema, it surpasses existing training-free methods using only 11.2 frames and achieves performance comparable to Gemini 1.5-Pro."},"_bibtex":{"value":"@misc{\nli2026perceive,\ntitle={Perceive, Reflect and Understand Long Video: Progressive Multi-Granular Clue Exploration with Interactive Agents},\nauthor={Jiahua Li and Kun Wei and Zhe Xu and Zibo Su and Xu Yang and Cheng Deng},\nyear={2026},\nurl={https://openreview.net/forum?id=klRdsrja2h}\n}"},"title":{"value":"Perceive, Reflect and Understand Long Video: Progressive Multi-Granular Clue Exploration with Interactive Agents"},"pdf":{"value":"/pdf/8d03466ce072ac61b16318436d4360580f6f6948.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|perceive_reflect_and_understand_long_video_progressive_multigranular_clue_exploration_with_interactive_agents"},"authorids":{"value":["~Jiahua_Li1","~Kun_Wei1","~Zhe_Xu6","~Zibo_Su1","~Xu_Yang6","~Cheng_Deng2"]},"authors":{"value":["Jiahua Li","Kun Wei","Zhe Xu","Zibo Su","Xu Yang","Cheng Deng"]}},"version":2},{"content":{"TLDR":{"value":"GIFARC is an analogy-inspired ARC dataset synthesized from GIF images that provides explicit human-intuitive analogies, significantly enhancing AI systems' abstract reasoning capabilities and improving solver accuracy on the ARC benchmark."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Abstraction and Reasoning Corpus (ARC)","analogical reasoning","synthetic datasets","GIF images","benchmark improvement"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"The Abstraction and Reasoning Corpus (ARC) poses a stringent test of general AI capabilities, requiring solvers to infer abstract patterns from only a handful of examples. Despite substantial progress in deep learning, state-of-the-art models still achieve accuracy rates of merely 40–55% on the 2024 ARC Competition, indicative of a significant gap between their performance and human-level reasoning. In this work, we seek to bridge that gap by introducing an analogy-inspired ARC dataset, GIFARC. Leveraging vision-language models (VLMs), we synthesize new ARC-style tasks from a variety of GIF images that include analogies. Each new task is paired with ground-truth analogy, providing an explicit mapping between visual transformations and everyday concepts. By embedding robust human-intuitive analogies into ARC-style tasks, GIFARC guides AI agents to adopt analogical reasoning approaches, facilitating more concise and human-understandable solutions. We empirically demonstrate that GIFARC improves task-solving performance by aligning model reasoning with human analogical problem-solving strategies."},"_bibtex":{"value":"@misc{\nsim2026gifarc,\ntitle={{GIFARC}: Synthetic Dataset for Leveraging Human-Intuitive Analogies to Elevate {AI} Reasoning},\nauthor={Woochang Sim and HyunSeokRyu and Kyungmin Choi and Sungwon Han and Sundong Kim},\nyear={2026},\nurl={https://openreview.net/forum?id=LScx9M0nLk}\n}"},"title":{"value":"GIFARC: Synthetic Dataset for Leveraging Human-Intuitive Analogies to Elevate AI Reasoning"},"pdf":{"value":"/pdf/694a72a6c439d299af26e27750abaf8458ed9bf5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sim|gifarc_synthetic_dataset_for_leveraging_humanintuitive_analogies_to_elevate_ai_reasoning"},"authorids":{"value":["~Woochang_Sim1","~HyunSeokRyu1","~Kyungmin_Choi1","~Sungwon_Han1","~Sundong_Kim1"]},"authors":{"value":["Woochang Sim","HyunSeokRyu","Kyungmin Choi","Sungwon Han","Sundong Kim"]}},"tmdate":1770805086066,"tcdate":1758288144546,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18480/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission18480/Authors"],"forum":"LScx9M0nLk","license":"CC BY 4.0","number":18480,"cdate":1758288144546,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission18480/-/Full_Submission","ICLR.cc/2026/Conference/Submission18480/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770805086066,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"LScx9M0nLk","version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS 2021"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-93413-2_35.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"liu|multiple_role_discovery_in_complex_networks"},"html":{"value":"https://doi.org/10.1007/978-3-030-93413-2_35"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/LiuTNU21,\n  author={Shu Liu and Fujio Toriumi and Mao Nishiguchi and Shohei Usui},\n  title={Multiple Role Discovery in Complex Networks},\n  year={2021},\n  cdate={1609459200000},\n  pages={415-427},\n  url={https://doi.org/10.1007/978-3-030-93413-2_35},\n  booktitle={COMPLEX NETWORKS},\n  crossref={conf/complexnetworks/2021-2}\n}\n"},"abstract":{"value":"The role of a node in complex networks is the aggregation of structural features and functions. Role discovery is the field of mining the proper roles, and many methods have been proposed. Those methods mainly focus on discovering a single role for each node. However, in real-world networks, a node may have multiple roles. Therefore, we propose a multiple-role discovery framework by extending the single-role discovery framework. Furthermore, we also suggest a way to assign sub-networks divided by community extraction methods to the source network and the validation network to select pre-labeling nodes, which is a significant challenge for multiple-role discovery in real-world networks. To evaluate the accuracy of the proposed method, we conduct computational experiments for multiple-role discovery of the real-world Wikipedia network and Blogcatalog network. We show that the proposed method achieves higher accuracy and more stable results than conventional methods used for comparison."},"title":{"value":"Multiple Role Discovery in Complex Networks"},"authors":{"value":[{"fullname":"Shu Liu","username":""},{"fullname":"Fujio Toriumi","username":"~Fujio_Toriumi3"},{"fullname":"Mao Nishiguchi","username":""},{"fullname":"Shohei Usui","username":""}]}},"tmdate":1779771747939,"pdate":1640908800000,"externalIds":["dblp:conf/complexnetworks/LiuTNU21"],"tcdate":1779771733764,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~FUJIO_TORIUMI2"],"forum":"yZMzW89Qgd","license":"CC BY-SA 4.0","number":26296,"cdate":1609459200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1779771747939,"domain":"OpenReview.net/Public_Article","id":"yZMzW89Qgd","version":2},{"content":{"summary":{"value":"This paper presents a study on evaluating deep neural networks designed to forecast the evolution of stochastic complex systems. The authors identify a gap in traditional evaluation methods—such as threshold-based classification metrics and error-based scoring rules—which focus on a model's ability to replicate observed ground truth but fail to assess how well the model has learned the underlying stochastic process. To address this issue, they introduce a new property called Fidelity to Stochastic Process, representing the DNN's ability to predict the statistical ground truth of the stochastic process.\n\nThe paper proposes using the Expected Calibration Error (ECE) as an evaluation metric that satisfies the necessary conditions for assessing fidelity to statistical ground truth. This work underscores the importance of capturing the underlying stochastic processes in deep neural networks  evaluations for complex systems."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"L50: Is --> is (lowercase)\nFig1: no need to write the whole name, you can use acronyms because they're already defined in the text, however MSE is not defined at this point.\nL88: fidelity to realization --> F2R (it was already defined previously, so you can use the acronym)\nL99: the notation of the dimension of the real vector O_t is confusing, what is (R^n)^(H x W), is n = H x W? If so, make that explicit.\nTable 1: some rows end with full stop, other don't. Please make it consistent. Either all with or all without.\nI find it odd to place Figures in columns as Figure 1 (which has a large top white margin) and Figure 3. I would suggest column figures into one row figure with multiple subfigures as you did with Figure 2. \nL201: Isn't the indicator variable already defined as B_t in L99? Why defining again with different notation?\nL298: MSE already defined in text previously, no need to write the whole name again.\nL516: ECE already defined in text previously, no need to write the whole name again.\nTable 2 and Table 7: highlight the best performing DNNs."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"The paper makes a significant contribution by introducing the concept of Fidelity to Stochastic Process (F2SP), a novel evaluation criterion specifically designed to assess a DNN's ability to learn the underlying stochastic interactions in complex systems.\n\nThe authors provide a rigorous formalization of F2SP within a stochastic framework, establishing clear criteria for its valid measurement. The use of Expected Calibration Error (ECE) as an evaluation metric is well-justified."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I found it hard to read the paper because there was a lack of consistency in the acronyms, the authors would redefine them in several parts of the text again and again. I addressed my comments on text in the questions section. \n\nIn the tables, the best neural networks based on each criterion are not highlighted, which makes it difficult to the reader to infer and correlate the arguments in the text. I addressed my comments on text in the questions section. \n\nThe focus of the paper is primarily on binary or discrete prediction tasks, leaving out regression tasks where the definition of calibration is more complex. While the authors acknowledge this and suggest it as an area for future work, the current scope limits the immediate applicability of the findings to a broader range of problems involving continuous outcomes.\n\nAdditionally, the use of the NDWS dataset, which is restricted to next-day predictions, prevents the assessment of ECE over longer time horizons, which are common in many complex systems. Could you elaborate on how future work might address this limitation? \n\nThe paper highlights the lack of open-source complex system datasets as a barrier to broader validation. Are there any ongoing initiatives or plans to develop, collect, or standardize such datasets?"}},"nonreaders":[],"tmdate":1732346222240,"tcdate":1730338026030,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11418/Reviewer_sNm1"],"signatures":["ICLR.cc/2025/Conference/Submission11418/Reviewer_sNm1"],"forum":"2U8owdruSQ","number":3,"license":"CC BY 4.0","cdate":1730338026030,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11418/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732346222240,"domain":"ICLR.cc/2025/Conference","replyto":"2U8owdruSQ","id":"wFirUKJfZc","forumContent":{"TLDR":{"value":"A novel evaluation criterion to assess whether DNNs modeling stochastic complex systems have learnt the underlying stochastic process"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["evaluation","deep neural network","stochasticity","complex systems","forecasting"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper presents the first systematic study of evaluating Deep Neural Networks (DNNs) designed to forecast the evolution of stochastic complex systems. We show that traditional evaluation methods like threshold-based classification metrics and error-based scoring rules assess a DNN's ability to replicate the observed ground truth but fail to measure the DNN's learning of the underlying stochastic process. To address this gap, we propose a new evaluation criteria called _Fidelity to Stochastic Process (F2SP)_, representing the DNN's ability to predict the system property _Statistic-GT_—the ground truth of the stochastic process—and introduce an evaluation metric that exclusively assesses F2SP. We formalize F2SP within a stochastic framework and establish criteria for validly measuring it. We formally show that Expected Calibration Error (ECE) satisfies the necessary condition for testing F2SP, unlike traditional evaluation methods. Empirical experiments on synthetic datasets, including wildfire, host-pathogen, and stock market models, demonstrate that ECE uniquely captures F2SP. We further extend our study to real-world wildfire data, highlighting the limitations of conventional evaluation and discuss the practical utility of incorporating F2SP into model assessment. This work offers a new perspective on evaluating DNNs modeling complex systems by emphasizing the importance of capturing underlying the stochastic process."},"_bibtex":{"value":"@inproceedings{\nkumar2025has,\ntitle={Has the Deep Neural Network learned the Stochastic Process? An Evaluation Viewpoint},\nauthor={Harshit Kumar and Beomseok Kang and Biswadeep Chakraborty and Saibal Mukhopadhyay},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=2U8owdruSQ}\n}"},"title":{"value":"Has the Deep Neural Network learned the Stochastic Process? An Evaluation Viewpoint"},"pdf":{"value":"/pdf/b06ae0fc426f0b455068cfcf40a18ae1875a5d09.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"kumar|has_the_deep_neural_network_learned_the_stochastic_process_an_evaluation_viewpoint"},"authorids":{"value":["~Harshit_Kumar2","~Beomseok_Kang1","~Biswadeep_Chakraborty1","~Saibal_Mukhopadhyay2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Harshit Kumar","Beomseok Kang","Biswadeep Chakraborty","Saibal Mukhopadhyay"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a training-free defense framework to mitigate adversarial perturbations in deep learning based MRI reconstruction, particularly for physics-driven deep learning (PD-DL) networks such as MoDL.\n\nThe proposed defense exploits cyclic measurement consistency (CMC): by resimulating undersampled k-space measurements from model reconstructions and enforcing self-consistency, the method detects and corrects adversarial perturbations without retraining or parameter modification.\n\nThe method formulates a reverse projected gradient descent (PGD) optimization in the input space to find a “corrective perturbation” that restores CMC. It is evaluated on fastMRI knee (Cor-PD) and brain (Ax-FLAIR) datasets, showing improvements over adversarial training and SMUG (smoothed unrolling). The method also generalizes to various PD-DL architectures and remains effective under blind and adaptive attacks."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"na"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper introduces a novel use of cyclic measurement consistency as a defensive objective, not just a training or calibration tool. This idea is well grounded in MRI physics and elegantly bridges signal reconstruction principles with adversarial robustness.\n\nThe approach works with any trained PD-DL model. This is a major advantage over retraining-based defenses like AT or SMUG. It can be combined with existing robust training for further gains.\n\nEvaluated across multiple datasets, perturbation levels, and attack types. Also assessed on several architectures (MoDL, XPDNet, RIM, E2E-VarNet, Recurrent-VarNet), confirming generality.\n\nProvides a theoretical analysis (Theorem 1) relating k-space perturbations and error propagation, lending physics-based interpretability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The defense requires multiple forward passes through the reconstruction model (often dozens per iteration of reverse PGD). The runtime cost, though reported, may limit practical use in real-time MRI settings.\n\nWhile strong within PD-DL MRI, the paper could discuss whether the approach generalizes to other imaging modalities or to non-physics-driven DL pipelines."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931170119,"tcdate":1761857983805,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19159/Reviewer_miBj"],"signatures":["ICLR.cc/2026/Conference/Submission19159/Reviewer_miBj"],"forum":"vcdMBXeiet","number":1,"license":"CC BY 4.0","cdate":1761857983805,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19159/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931170119,"domain":"ICLR.cc/2026/Conference","replyto":"vcdMBXeiet","id":"AC1qwbAO6n","forumContent":{"TLDR":{"value":"We mitigate the effect of adversarial attacks on deep learning-based MRI reconstruction without re-training."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["adversarial attacks","computational imaging","magnetic resonance imaging","inverse problems"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Deep learning (DL) methods have become the state-of-the-art for reconstructing sub-sampled magnetic resonance imaging (MRI) data. However, studies have shown that these methods are susceptible to small adversarial input perturbations, or attacks, resulting in major distortions in the output images. Various strategies have been proposed to reduce the effects of these attacks, but they require retraining and may lower reconstruction quality for non-perturbed/clean inputs. In this work, we propose a novel approach for mitigating adversarial attacks on MRI reconstruction models without any retraining. Based on the idea of cyclic measurement consistency, we devise a novel mitigation objective that is minimized in a small ball around the attack input. Results show that our method substantially reduces the impact of adversarial perturbations across different datasets, attack types/strengths and PD-DL networks, and qualitatively and quantitatively outperforms conventional mitigation methods that involve retraining. We also introduce a practically relevant scenario for small adversarial perturbations that models impulse noise in raw data, which relates to herringbone artifacts, and show the applicability of our approach in this setting. Finally, we show our mitigation approach remains effective in two realistic extension scenarios: a blind setup, where the attack strength or algorithm is not known to the user; and an adaptive attack setup, where the attacker has full knowledge of the defense strategy."},"_bibtex":{"value":"@misc{\nsaberi2026trainingfree,\ntitle={Training-Free Defense Against Adversarial Attacks In Deep Learning {MRI} Reconstruction},\nauthor={Mahdi Saberi and Chi Zhang and Mehmet Akcakaya},\nyear={2026},\nurl={https://openreview.net/forum?id=vcdMBXeiet}\n}"},"title":{"value":"Training-Free Defense Against Adversarial Attacks In Deep Learning MRI Reconstruction"},"pdf":{"value":"/pdf/502767c941f492dd0a41ed2d83b628be2a1b747d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"saberi|trainingfree_defense_against_adversarial_attacks_in_deep_learning_mri_reconstruction"},"authorids":{"value":["~Mahdi_Saberi1","~Chi_Zhang65","~Mehmet_Akcakaya1"]},"authors":{"value":["Mahdi Saberi","Chi Zhang","Mehmet Akcakaya"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewer for their time and their helpful comments.\n\n> The idea of learning dynamics/physics simulators from videos is not particularly new (e.g., NeRF-dy, [A-B]), but the intro and related work are positioned in a way that appears those works are not relevant.\n\nThank you for the feedback. We certainly did not intend to make it seem as if these works are not relevant, and of course NeRF-dy in particular informed our main ablations. We have updated the introduction and related work to add more detailed discussion of these works and the differences between them and VPD.\n\n> To obtain better physics simulators (of system dynamics), what do we gain by learning from videos?\n\nWe gain two main things: (1) the possibility of customizing simulators to particular physics, (2) learning physics for systems that we do not understand sufficiently well to make analytic simulators for. For point (1): in theory, customization can be supported by system identification. However, previous work has suggested that it can be difficult to model certain real systems (such as those with friction or contact dynamics) with system identification (e.g. [1, 2])) especially when perception from a video is necessarily noisy and imperfect. Using a learned simulation technique instead can overcome these perceptual errors. For point (2): there are some physical dynamics that are extremely difficult to model analytically. Friction is a famous example – there are almost imperceptible ways in which the surfaces of objects end up affecting the way that their surface interact. Writing an analytic simulator is therefore not possible in these cases, so we may want to learn one from watching a video.\n\n> Specifically, the problems mentioned in the intro seem can be solved by system identification with a differentiable simulator, e.g., Taichi, Warp, Brax, dojo, which can find physical parameters (e.g., friction coefficient) from input videos without re-learn how to simulate.\n\nIndeed, SysID is a common technique to estimate parameters in combination with DiffSims. However, system identification from video poses additional challenges in obtaining 3D models of objects and tracking them accurately. For example, Le Cleac'h [3] and NeuPhys [A] and [B], perform sysID on a scene with a single object. Moreover, these often involve either a relatively reduced set of parameters to be estimated [A], or can require binary object masks, knowledge of geometry [B] or 3D models learned ahead of time from a collection of still images [3],  none of which are necessary for VPD. In contrast, we were motivated by prior work on learned simulation which found that graph network simulators can outperform analytic simulators with sysID on ground-truth states [1, 2]. This can be especially promising for rigid contacts, which as you observe require small timesteps for analytic simulators. \n\nWe have added further discussion of the papers referenced in this review to our related work.\n\n> If the goal is to speed up the simulation, it seems distilling physics-based simulators into a neural architecture is a strong competitor.\n\nIn effect, this is what Mesh Graph Networks [4] do (and other learned/hybrid simulators). They distill physics-based simulators into graph neural networks, which results in computational speed-ups of 100-1000x. In this paper, we show how to use similar graph neural network architectures for physical dynamics, but adapted to work from RGB-D videos. In principle, this then allows us to distill physics simulators for any system that we can capture RGB-D video from.\n\n> Overall, it would be great if the paper could give a compelling example where VPD outperforms (or has the hope to outperform) the physics-based simulator. One aspect might be the generality. Another case is when there are complex disturbances in the environment that cannot be modeled by the physics simulator.\n\nPrior work on learned simulation has found that graph network simulators can outperform physics-based simulators with sysID on ground-truth states for rigid body dynamics (specifically, cube-ground collisions, and planar pushing  [1, 2]). We demonstrate VPD on the same cube-ground collisions here, with the additional complexity that VPD learns from RGB-D videos rather than from state information. Given that the underlying dynamics architecture is similar to [1, 2], we expect that VPD would similarly outperform system identification with a physics-based simulator for these cases. Furthermore, relative to Brax and dojo, VPD is more general – it can operate over non-rigid dynamics, which we show in the deformable experiments."},"title":{"value":"Reply 1 / 2"}},"tmdate":1700669322106,"tcdate":1700669322106,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3677/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission3677/Authors"],"forum":"4rBEgZCubP","number":4,"license":"CC BY 4.0","cdate":1700669322106,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3677/-/Official_Comment"],"mdate":1700669322106,"domain":"ICLR.cc/2024/Conference","replyto":"YUFOaSRD8T","id":"DFVZz2b4ok","forumContent":{"TLDR":{"value":"Learned dynamics models that combine 3D representations and ray-based rendering."},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["simulation","dynamics","nerf","particle dynamics"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Realistic simulation is critical for applications ranging from robotics to animation. Traditional analytic simulators sometimes struggle to capture sufficiently realistic simulation which can lead to problems including the well known \"sim-to-real\" gap in robotics. Learned simulators have emerged as an alternative for better capturing real-world physical dynamics, but require access to privileged ground truth physics information such as precise object geometry or particle tracks. Here we propose a method for learning simulators directly from observations. Visual Particle Dynamics (VPD) jointly learns a latent particle-based representation of 3D scenes, a neural simulator of the latent particle dynamics, and a renderer that can produce images of the scene from arbitrary views. VPD learns end to end from posed RGB-D videos and does not require access to privileged information. Unlike existing 2D video prediction models, we show that VPD's 3D structure enables scene editing and long-term predictions. These results pave the way for downstream applications ranging from video editing to robotic planning."},"_bibtex":{"value":"@inproceedings{\nwhitney2024learning,\ntitle={Learning 3D Particle-based Simulators from {RGB}-D Videos},\nauthor={William F Whitney and Tatiana Lopez-Guevara and Tobias Pfaff and Yulia Rubanova and Thomas Kipf and Kim Stachenfeld and Kelsey R Allen},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=4rBEgZCubP}\n}"},"title":{"value":"Learning 3D Particle-based Simulators from RGB-D Videos"},"pdf":{"value":"/pdf/c084f0fff026d69efbf43d593934b7e30a668247.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"whitney|learning_3d_particlebased_simulators_from_rgbd_videos"},"authorids":{"value":["~William_F_Whitney1","~Tatiana_Lopez-Guevara1","~Tobias_Pfaff1","~Yulia_Rubanova2","~Thomas_Kipf2","~Kim_Stachenfeld1","~Kelsey_R_Allen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["William F Whitney","Tatiana Lopez-Guevara","Tobias Pfaff","Yulia Rubanova","Thomas Kipf","Kim Stachenfeld","Kelsey R Allen"]}},"version":2},{"content":{"summary":{"value":"This paper introduces CuCoDistill, a highly novel and complex framework for knowledge distillation (KD) in Hypergraph Neural Networks (HGNNs). The authors address the failure of existing HGNN attention mechanisms to handle hypergraph asymmetries and the limitations of standard KD in preserving higher-order structures . The framework's core innovations include: (1) a hypergraph-aware adaptive attention mechanism with provable spectral guarantees; (2) a unified co-evolutionary architecture where teacher and student models train simultaneously rather than sequentially ; and (3) a spectral curriculum scheduler that dynamically adjusts learning difficulty based on hypergraph properties. The paper theoretically and empirically demonstrates the counter-intuitive finding that, under certain conditions (e.g., noisy datasets), the compressed student model can systematically outperform the larger teacher model ."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The ablation study shows the \"Spectral Curriculum\" has the smallest individual impact (0.9-1.1%). Given its complexity (calculating dual difficulties, quantile thresholds), is this component truly necessary, or could a simpler regularization suffice?\n\n2. In the t-SNE analysis (Figure 4, ), the student embedding space for DBLP shows a worse silhouette score (0.327) than the teacher (0.614), yet the student model outperforms the teacher on the DBLP task (Table 1). This is counter-intuitive. Could the authors explain why degraded cluster quality in the embedding space leads to better classification accuracy in this case?\n\n\n3. There is a citation error in the baseline description (Section D.2.2). The text cites (Zhang et al., 2019b) for Hyper-SAGNN but then describes HyGCL-AdT (Qian et al., 2024) . This should be corrected."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The framework is innovative, particularly its \"co-evolutionary\" architecture and the theoretical demonstration that a student model can surpass its teacher. \n\n\n\n2. The work is theoretically deep, providing provable guarantees for its attention mechanism (Theorem 1) and formalizing the conditions for student superiority (Theorem 2) , lending rigor to its claims.\n\n\n\n\n\n3. The empirical results are good, showing state-of-the-art performance, efficiency gains (6.25x speedup, 10x memory reduction), and, crucially, validating the \"student surpasses teacher\" phenomenon on several large-scale, noisy datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The framework's complexity is extremely high, potentially hindering reproducibility and adoption. It integrates multiple complex components (multi-scale attention, co-evolution, spectral curriculum, multi-level KD losses ), creating a system that is very difficult to implement and tune.\n\n2. The claimed \"student superiority\" is highly conditional and not a general outcome. The results clearly show this phenomenon occurs only on large, noisy, or feature-redundant datasets (e.g., DBLP, IMDB, Yelp). On clean, well-structured datasets (e.g., CC-Cora), the teacher model remains superior, a critical nuance that limits the generality of the titular claim.\n\n3. The method introduces a very large number of new hyperparameters. The spectral curriculum (adaptive thresholds, loss weights $\\lambda(t)$) , attention mechanism (Top-K $\\alpha$) , and various loss component weights create a complex tuning space, even with the sensitivity analysis provided in the appendix."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920253545,"tcdate":1762161095244,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8334/Reviewer_MvEx"],"signatures":["ICLR.cc/2026/Conference/Submission8334/Reviewer_MvEx"],"forum":"GgBv8mo9Aw","number":5,"license":"CC BY 4.0","cdate":1762161095244,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8334/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920253545,"domain":"ICLR.cc/2026/Conference","replyto":"GgBv8mo9Aw","id":"Tkf7lkEQQj","forumContent":{"TLDR":{"value":"An asymmetric contrastive scheme where only the teacher processes both clean and perturbed views, fusing them via a learnable gating mechanism to produce high‐quality distillation targets."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Hypergraph Learning","Attention","knowledge distillation","Co-Distillation"]},"supplementary_material":{"value":"/attachment/1bd7f2736279c2ad4aa1659318cc0173945ee526.zip"},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Many real-world systems involve complex many-to-many relationships naturally represented as hypergraphs, from social networks to molecular interactions. While hypergraph neural networks (HGNNs) have shown promise, existing attention mechanisms fail to handle hypergraph-specific asymmetries between node-to-node, node-to-hyperedge, and hyperedge-to-node interactions, leading to suboptimal structural encoding. We introduce \\textbf{CuCoDistill}, a novel framework that challenges fundamental assumptions in knowledge distillation by demonstrating that student models can systematically outperform their teachers through hypergraph-aware adaptive attention with provable spectral guarantees. Our approach features: (1) set-aware attention fusion that handles variable-sized hyperedge sets with approximation error bounds of $\\epsilon\\sqrt{|\\mathcal{V}|}\\max_i|\\mathcal{E}_i|$; (2) co-evolutionary unified architecture where teacher and student jointly discover structural patterns in a single forward pass; and (3) theoretically-grounded curriculum distillation based on hypergraph spectral properties. We prove that when student's constrained attention aligns with the hypergraph's intrinsic spectral dimension, superior generalization emerges through beneficial regularization. Extensive experiments across nine benchmarks show our students achieve up to 1.8\\% higher accuracy than teachers while delivering 6.25× inference speedup and 10× memory reduction, consistently outperforming state-of-the-art methods and establishing new efficiency-performance frontiers for hypergraph learning."},"_bibtex":{"value":"@misc{\nforouzandeh2026when,\ntitle={When Students Surpass Teachers: Hypergraph-Aware Knowledge Distillation with Spectral Guarantees},\nauthor={Saman Forouzandeh and kamal berahmand and Parham Moradi and Mahdi Jalili},\nyear={2026},\nurl={https://openreview.net/forum?id=GgBv8mo9Aw}\n}"},"title":{"value":"When Students Surpass Teachers: Hypergraph-Aware Knowledge Distillation with Spectral Guarantees"},"pdf":{"value":"/pdf/aefb360b4c7652e5baa0cee7344e6461c6187db6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"forouzandeh|when_students_surpass_teachers_hypergraphaware_knowledge_distillation_with_spectral_guarantees"},"authorids":{"value":["~Saman_Forouzandeh1","~kamal_berahmand1","~Parham_Moradi2","~Mahdi_Jalili1"]},"authors":{"value":["Saman Forouzandeh","kamal berahmand","Parham Moradi","Mahdi Jalili"]}},"version":2},{"content":{"summary":{"value":"This paper proposes VideoJudge, a bootstrapped framework for training multimodal large language models to serve as automatic evaluators for video understanding tasks. By generating synthetic data through an iterative generator–evaluator process, the method enables scalable supervision without human labels. The resulting small models outperform larger MLLM judges in correlation with human judgments across multiple benchmarks."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"See weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.\tThe generator-evaluator cycle to produce synthetic training data for the judge is meaningful and well-motivated. It reduces reliance on human annotations and enables larger scale.\n\n2.\tThe authors show that a relatively small model (3B) can outperform much larger baselines when trained appropriately, which is a strong practical result."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThe bootstrapping relies on LLM-based generator and evaluator and that LLM may carry bias and errors. If the evaluator is weak, the whole pipeline could propagate flawed judgments. The paper could more deeply analyze potential bias or drift in this synthetic supervised signal.\n\n2.\tWhile the benchmarks used show strong correlation results, it is not entirely clear how robust the judge will be to entirely new video domains, tasks, or instruction types. The paper would benefit from a “domain shift” evaluation (unseen video types).\n\n3.\tWhile the correlation metrics are good, more detailed breakdowns of failure cases where the judge disagrees with humans would help understand the limitations. Are there types of errors the judge misses (e.g., subtle temporal reasoning, common sense)?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942200178,"tcdate":1761300941476,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22397/Reviewer_zfmP"],"signatures":["ICLR.cc/2026/Conference/Submission22397/Reviewer_zfmP"],"forum":"31CznLfRIS","number":1,"license":"CC BY 4.0","cdate":1761300941476,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22397/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942200178,"domain":"ICLR.cc/2026/Conference","replyto":"31CznLfRIS","id":"CzEwT16fQo","forumContent":{"TLDR":{"value":"Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Meta-evaluation","llm-as-judge","synthetic data","self-refinement","video understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Precisely evaluating video understanding models remains challenging: commonly used metrics such as BLEU, ROUGE, and BERTScore fail to capture the nuances of human judgment, while obtaining such judgments through manual evaluation is costly. Recent work has explored using large language models (LLMs) or multimodal LLMs (MLLMs) as evaluators, but their extension to video understanding remains relatively unexplored. In this work, we introduce VideoJudge, a 3B and 7B-sized MLLM judge specialized to evaluate outputs from video understanding models (\\textit{i.e.}, text responses conditioned on videos). To train VideoJudge, our recipe builds on the interplay between a generator and an evaluator: the generator is prompted to produce responses conditioned on a target rating, and responses not matching the evaluator's rating are discarded. Across three out of four meta-evaluation benchmarks, VideoJudge-7B outperforms or is on par with larger MLLM judge baselines such as Qwen2.5-VL (32B and 72B). Notably, we find that LLM judges (Qwen3) models perform worse than MLLM judges (Qwen2.5-VL), and long chain-of-thought reasoning does not improve performance, indicating that providing video inputs is crucial for the evaluation of video understanding tasks."},"_bibtex":{"value":"@inproceedings{\nwaheed2026videojudge,\ntitle={VideoJudge: Bootstrapping Enables Scalable Supervision of {MLLM}-as-a-Judge for Video Understanding},\nauthor={Abdul Waheed and Zhen Wu and Dareen Safar Alharthi and Seungone Kim and Bhiksha Raj},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=31CznLfRIS}\n}"},"title":{"value":"VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding"},"pdf":{"value":"/pdf/5bb508da58e73095fe0543c2bcf02662e637500b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"waheed|videojudge_bootstrapping_enables_scalable_supervision_of_mllmasajudge_for_video_understanding"},"authorids":{"value":["~Abdul_Waheed1","~Zhen_Wu5","~Dareen_Safar_Alharthi1","~Seungone_Kim1","~Bhiksha_Raj1"]},"authors":{"value":["Abdul Waheed","Zhen Wu","Dareen Safar Alharthi","Seungone Kim","Bhiksha Raj"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LaTo, a landmark-tokenized diffusion transformer for fine-grained, identity-preserving human face editing. \nLaTo quantizes landmark coordinates into discrete facial tokens, integrating them into diffusion transformers using location-mapped positional encoding. \nIt also introduces a landmark predictor based on a fine-tuned vision–language model with structured chain-of-thought reasoning, allowing intuitive user control without explicit landmark input. \nThe authors further curate HFL-150K, a large-scale dataset containing 150k real and synthetic face editing pairs with precise annotations. Extensive experiments demonstrate that LaTo achieves superior identity preservation and semantic consistency compared to prior state-of-the-art models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How are identity-dependent facial variations (e.g., different muscle or bone structures) handled when generating landmarks for the same expression?\n2. How are identity preservation metrics validated against human perceptual ratings? Is there a correlation analysis?\n3. What is the sensitivity of the model to noisy or inaccurate landmark predictions from the VLM?\n4. Can the proposed landmark predictor generate accurate and stable landmark coordinates for non-realistic, stylized cartoon faces? Does the full LaTo pipeline (landmark tokenization + DiT fusion) maintain editing quality when the input lies outside the natural-face manifold?\n5. Can the landmark tokenizer or location-mapping positional encoding be applied to other structured domains, such as full-body or hand motion editing?\n6. How does the landmark predictor perform under noisy or ambiguous instructions?\n7. Are there any plans to release HFL-150K in subsets (e.g., synthetic only)?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Direct landmark tokenization provides an efficient geometric prior for diffusion transformers.\n2. LaTo achieves consistent improvement across benchmarks (HFL-150K, ICE-Bench, GEdit-Bench).\n3. The effect of positional encoding and classifier-free guidance is well quantified.\n4. HFL-150K is a valuable, large-scale dataset that provides both real and synthetic samples with fine-grained instructions — a useful resource for the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The landmark predictor tends to produce nearly identical landmark configurations for the same expression instruction (e.g., “make him happy strongly”) across subjects, for example, the first four columns of Figure 1.\nThis over-templating reduces individual expressiveness and may lead to uniform, less personalized facial deformations.\nThe paper lacks analysis on inter-identity landmark variance.\n2. Figure 1 (second to last column) includes a cartoon-style input image as an example. In such cases, landmark prediction becomes substantially more difficult because the facial geometry deviates from human anatomy — e.g., mouths drawn as single lines, exaggerated eyes, missing noses, or highly stylized jawlines. \nThese features make standard 68-point facial landmark detection ill-defined.\nHowever, the paper does not clarify whether the proposed landmark predictor is trained or evaluated on such stylized data. \nGiven that HFL-150K is described as primarily real-human-face-based (Section 3.1), it is doubtful that the model has seen sufficient cartoon or synthetic faces during training.\n3. Although the paper promises public release “upon acceptance,” both LaTo’s code and the HFL-150K dataset are currently unavailable.\nFor a paper emphasizing both model and dataset contributions, this limits reproducibility."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920618796,"tcdate":1761970956009,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8855/Reviewer_Cy7B"],"signatures":["ICLR.cc/2026/Conference/Submission8855/Reviewer_Cy7B"],"forum":"7bv3jLhlYZ","number":4,"license":"CC BY 4.0","cdate":1761970956009,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8855/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920618796,"domain":"ICLR.cc/2026/Conference","replyto":"7bv3jLhlYZ","id":"smnswDCsQl","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Image Editing; Face Editing; Identity Preservation; Landmark-tokenized"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent multimodal models for instruction-based face editing enable semantic manipulation but still struggle with precise attribute control and identity preservation. Structural facial representations such as landmarks are effective for intermediate supervision, yet most existing methods treat them as rigid geometric constraints, which can degrade identity when conditional landmarks deviate significantly from the source (e.g., large expression or pose changes, inaccurate landmark estimates). To address these limitations, we propose LaTo, a landmark-tokenized diffusion transformer for fine-grained, identity-preserving face editing. Our key innovations include: (1) a landmark tokenizer that directly quantizes raw landmark coordinates into discrete facial tokens, obviating the need for dense pixel-wise correspondence; (2) a location-mapped positional encoding and a landmark-aware classifier-free guidance that jointly facilitate flexible yet decoupled interactions among instruction, geometry, and appearance, enabling strong identity preservation; and (3) a landmark predictor that leverages vision–language models to infer target landmarks from instructions and source images, whose structured chain-of-thought improves estimation accuracy and interactive control. To mitigate data scarcity, we curate HFL-150K, to our knowledge the largest benchmark for this task, containing over 150K real face pairs with fine-grained instructions. Extensive experiments show that LaTo outperforms state-of-the-art methods by 7.8% in identity preservation and 4.6% in semantic consistency. Code is available at https://github.com/alibaba/landmark-tokenized-dit."},"_bibtex":{"value":"@inproceedings{\nzhang2026lato,\ntitle={LaTo:  Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing},\nauthor={Zhenghao Zhang and Ziying Zhang and Junchao Liao and Xiangyu Meng and Qiang Hu and Siyu Zhu and Xiaoyun Zhang and Long Qin and Weizhi Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=7bv3jLhlYZ}\n}"},"title":{"value":"LaTo:  Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing"},"pdf":{"value":"/pdf/5943a9484fa89b924fb7e92be6a20839f9c23098.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|lato_landmarktokenized_diffusion_transformer_for_finegrained_human_face_editing"},"authorids":{"value":["~Zhenghao_Zhang3","~Ziying_Zhang1","~Junchao_Liao1","~Xiangyu_Meng2","~Qiang_Hu2","~Siyu_Zhu1","~Xiaoyun_Zhang1","~Long_Qin2","~Weizhi_Wang2"]},"authors":{"value":["Zhenghao Zhang","Ziying Zhang","Junchao Liao","Xiangyu Meng","Qiang Hu","Siyu Zhu","Xiaoyun Zhang","Long Qin","Weizhi Wang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes the VIPER-R1 framework, which aims to achieve automated discovery of physical laws by integrating visual perception (such as phase portraits and trajectory plots) with symbolic reasoning. The methodology involves a two-stage pipeline: first, hypothesis generation via Motion Structure Induction (MSI), followed by reinforcement learning-based optimization of the model's output through Reward-Guided Symbolic Calibration (RGSC). During inference, VIPER-R1 invokes an external symbolic regression tool to perform Symbolic Residual Realignment (SR^2), enhancing the consistency between hypotheses and empirical data. Experiments are conducted on the newly constructed PhysSymbol dataset, and comparisons with other Vision-Language Models (VLMs) are made based on structural matching and symbolic accuracy metrics."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"How are the weights (e.g., w_f, w_s, w_a) in the RGSC reward function determined? Is there experimental evidence demonstrating the dominant role of the structural reward (R_{\\text{structural}})?\n\nDoes the synthesis process of the PhysSymbol dataset consider real-world physical constraints (e.g., energy conservation)? How do you plan to extend it to real data in the future?\n\nIn the SR^2 stage, does the processing time of the symbolic regression tool become a bottleneck? How scalable is VIPER-R1 for complex systems (e.g., chaotic systems)?\n\nThe paper mentions that VIPER-R1 \"proactively invokes\" external tools during inference. Does this require manual intervention? What is the degree of automation of the framework?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper addresses an underexplored research area - multimodal physics formula discovery, attempting to establish connections between visual perception and symbolic reasoning. This direction distinguishes itself from methods relying solely on symbolic regression or pure textual reasoning approaches.\n\n2. A well-structured phased training pipeline is designed through the combination of MSI and RGSC. MSI serves for initial hypothesis generation, while RGSC performs structural optimization of the output via reinforcement learning. The introduction of SR² provides an adjustment mechanism to align theoretical models with empirical data.\n\n3. The paper establishes a systematic experimental foundation by constructing the PhysSymbol dataset containing various physical scenarios and proposing evaluation metrics such as structural score and accuracy score, providing both dataset resources and evaluation benchmarks for subsequent research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The core reinforcement learning framework of VIPER-RFT (the RGSC stage) shows significant similarity to Visual-RFT [1]. Both methods share common procedures:\n   • Generating multiple candidate responses from a policy model\n   • Employing rule-based reward functions (VIPER-RFT's structural reward vs. Visual-RFT's IoU/classification rewards) for output evaluation\n   • Optimizing policies through relative advantage normalization and KL regularization\n\n   Although VIPER-RFT customizes its reward function for physics formula structure, the high-level paradigm of \"verifiable reward-driven RL for multimodal tasks\" has been previously established by Visual-RFT, which somewhat diminishes the perceived innovativeness of the proposed framework.\n\n[1] Liu Z, Sun Z, Zang Y, et al. Visual-rft: Visual reinforcement fine-tuning. CVPR, 2025.\n\n2. The authors compare their fine-tuned VIPER-R1 (specifically adapted to the PhysSymbol dataset) against pre-trained, non-fine-tuned VLMs (e.g., GPT-4o, Gemini). This benchmark setup appears unfair since these base models lack task-specific adaptation. Given that large model fine-tuning is widely studied, a rigorous comparison should include:\n   • Other state-of-the-art VLMs fine-tuned on the same PhysSymbol dataset\n   • Ablation studies demonstrating the necessity of each component (MSI, RGSC) beyond simple baselines\n\n3. The PhysSymbol dataset is synthetic, comprising idealized trajectories and phase portraits. However, real-world physical data often involve noise, occlusions, and complex boundary conditions. The paper does not validate whether VIPER-R1's performance can generalize to noisy or real-world scenarios, raising concerns about its practical applicability.\n\n4. The SR^2 stage relies on external symbolic regression tools (e.g., PySR) for parameter refinement. This dependency may impact the method's reproducibility and scalability, particularly when the tools struggle with high-dimensional or noisy residuals. The paper lacks ablation studies analyzing the impact of different symbolic regression tools on final performance."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915943232,"tcdate":1761929120665,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1909/Reviewer_5ST1"],"signatures":["ICLR.cc/2026/Conference/Submission1909/Reviewer_5ST1"],"forum":"nRhHbKP1y9","number":2,"license":"CC BY 4.0","cdate":1761929120665,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1909/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915943232,"domain":"ICLR.cc/2026/Conference","replyto":"nRhHbKP1y9","id":"F3d3u3FNjY","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics Formula Discovery","Multimodal Scientific Reasoning","Vision-Language Models"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to uni-modal data and overlook the rich, visual phenomenological representations of motion that are indispensable to physicists. This \"sensory deprivation\" severely weakens their ability to interpret the inherent spatio-temporal patterns within dynamic phenomena. To address this gap, we propose VIPER-R1, a multimodal model that performs Visual Induction for Physics-based Equation Reasoning to discover fundamental symbolic formulas. It integrates visual perception, trajectory data, and symbolic reasoning to emulate the scientific discovery process. The model is trained via a curriculum of Motion Structure Induction (MSI), using supervised fine-tuning to interpret kinematic phase portraits and to construct hypotheses guided by a Causal Chain of Thought (C-CoT), followed by Reward-Guided Symbolic Calibration (RGSC) to refine the formula structure with reinforcement learning. During inference, the trained VIPER-R1 acts as an agent: it first posits a high-confidence symbolic ansatz, then proactively invokes an external symbolic regression tool to perform Symbolic Residual Realignment (SR^2). This final step, analogous to a physicist's perturbation analysis, reconciles the theoretical model with empirical data. To support this research, we introduce PhysSymbol, a new 10,000-instance multimodal corpus. Experiments show that VIPER-R1 consistently outperforms state-of-the-art VLM baselines in accuracy and interpretability, enabling more precise discovery of physical laws."},"_bibtex":{"value":"@misc{\nliu2026mimicking,\ntitle={Mimicking the Physicist's Eye: A {VLM}-centric Approach for Physics Formula Discovery},\nauthor={Jiaqi Liu and Songning Lai and Pengze Li and Di Yu and Zhou wenjie and Yiyang Zhou and Peng Xia and Zijun Wang and Xi Chen and SHIXIANG TANG and LEI BAI and Wanli Ouyang and Mingyu Ding and Huaxiu Yao and Aoran Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=nRhHbKP1y9}\n}"},"title":{"value":"Mimicking the Physicist's Eye: A VLM-centric Approach for Physics Formula Discovery"},"pdf":{"value":"/pdf/53e9a62249b142c46601ed29ca4cbddde40ac387.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"liu|mimicking_the_physicists_eye_a_vlmcentric_approach_for_physics_formula_discovery"},"authorids":{"value":["~Jiaqi_Liu7","~Songning_Lai1","~Pengze_Li3","~Di_Yu3","~Zhou_wenjie2","~Yiyang_Zhou1","~Peng_Xia1","~Zijun_Wang4","~Xi_Chen20","~SHIXIANG_TANG1","~LEI_BAI1","~Wanli_Ouyang1","~Mingyu_Ding1","~Huaxiu_Yao1","~Aoran_Wang1"]},"authors":{"value":["Jiaqi Liu","Songning Lai","Pengze Li","Di Yu","Zhou wenjie","Yiyang Zhou","Peng Xia","Zijun Wang","Xi Chen","SHIXIANG TANG","LEI BAI","Wanli Ouyang","Mingyu Ding","Huaxiu Yao","Aoran Wang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Mixture of Agents Alignment (MoAA), a scalable and diverse synthetic data recipe that leverages the strengths of various language models to provide high-quality data for model alignment. The method showed improvements over baselines on datasets such as Arena-Hard and AlpacaEval2."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.⁠ ⁠A simple and intuitive approach that leverages multiple open-source models for generating high-quality synthetic data for model alignment.\n2.⁠ ⁠The approach demonstrates performance gains over baselines across datasets.\n3.⁠ ⁠The paper provides a detailed analysis, offering insights into the impact of each training stage."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.⁠ ⁠Limited Novelty: The paper primarily extends existing methodologies, such as the Mixture of Agents (MoA) framework, for generating synthetic data for supervised fine-tuning and preference optimization. Its contributions largely revolve around practical refinements and adaptations of known techniques.\n2.⁠ ⁠Weak Baselines: Baselines only consist of open-source instruction-tuned models without comparing against stronger baselines such as [1,2].\n3.⁠ ⁠Limited Model Evaluation: Evaluation is limited to Llama 3.1 8B and Gemma 2 9B models. Expanding to a broader range of model sizes and architectures (Llama 3.2 1B and 3B, Qwen 2 0.5B and 1.5B) can provide a good insight into the generalizability of the method.\n\n[1] META-REWARDING LANGUAGE MODELS:\nSelf-Improving Alignment with LLM-as-a-Meta-Judge, https://arxiv.org/abs/2407.19594v2 \n[2] MAGPIE: ALIGNMENT DATA SYNTHESIS FROM SCRATCH\nBY PROMPTING ALIGNED LLMS WITH NOTHING. https://arxiv.org/abs/2406.08464v1"}},"nonreaders":[],"tmdate":1731428186995,"tcdate":1730595574500,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12530/Reviewer_n9xg"],"signatures":["ICLR.cc/2025/Conference/Submission12530/Reviewer_n9xg"],"forum":"lXFGpwtkRl","number":3,"license":"CC BY 4.0","cdate":1730595574500,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12530/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428186995,"domain":"ICLR.cc/2025/Conference","replyto":"lXFGpwtkRl","id":"wyr2MLiyjg","forumContent":{"TLDR":{"value":"We propose MoAA that leverages multiple language models to generate diverse, high-quality data for scalable model alignment."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Model Alignment","Multi-Agent Inference","Large Language Model"]},"supplementary_material":{"value":"/attachment/3688fb8b1a87cbf1250f5fec5c86d2868f170103.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building helpful and harmless large language models (LLMs) requires effective model alignment approach based on human instructions and feedback; this necessitates high-quality human-labeled data. Constructing such datasets is often expensive and not scalable, and may face potential bottleneck on diversity. To address these challenges, we introduce Mixture-of-Agent Alignment (MoAA), an effective approach that leverages the collective strengths of various language models to provide high-quality data for model alignment. By employing MoAA, we enhance both supervised fine-tuning (SFT) and preference optimization, leading to improved performance compared to using a single model alone, including the state-of-ther-art commercial model. This approach leads to an intriguing direction of model alignment through an scalable and diverse instruction data recipe based on open-sourced models."},"_bibtex":{"value":"@misc{\nwang2025improving,\ntitle={Improving Model Alignment Through Collective Intelligence of Open-Source Models},\nauthor={Junlin Wang and Roy Xie and Shang Zhu and Jue WANG and Ben Athiwaratkun and Bhuwan Dhingra and Shuaiwen Leon Song and Ce Zhang and James Zou},\nyear={2025},\nurl={https://openreview.net/forum?id=lXFGpwtkRl}\n}"},"title":{"value":"Improving Model Alignment Through Collective Intelligence of Open-Source Models"},"pdf":{"value":"/pdf/e044149835730abe29a22fa5fea70998c5bf5095.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|improving_model_alignment_through_collective_intelligence_of_opensource_models"},"authorids":{"value":["~Junlin_Wang1","~Roy_Xie1","~Shang_Zhu1","~Jue_WANG1","~Ben_Athiwaratkun1","~Bhuwan_Dhingra1","~Shuaiwen_Leon_Song1","~Ce_Zhang1","~James_Zou1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Junlin Wang","Roy Xie","Shang Zhu","Jue WANG","Ben Athiwaratkun","Bhuwan Dhingra","Shuaiwen Leon Song","Ce Zhang","James Zou"]}},"version":2},{"content":{"summary":{"value":"This paper introduces GovBench, a hierarchical benchmark for data governance automation. It also proposes DataGovAgent, an multi-agent workflow with a Planner-Executor-Evaluator architecture for translating natural language into verified governance pipelines. Experiments on GovBench show DataGovAgent outperforms SOTA models and general agent frameworks, boosting complex DAG-task ATS."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"See weakness. Overall, core components are presented as implementation steps, lacking insightful analysis, making the work read more like a tool-building project than a research paper."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. GovBench overcomes the limitations of existing snippet-focused benchmarks. The proposed hierarchical tasks (operator/DAG-level) and targeted noise injection simulate real-world data governance scenarios. \n2. DataGovAgent’s architecture is intuitive. Contract-guided planning, retrieval-augmented generation, and meta-cognitive debugging directly solve complex workflow decomposition and error-correction issues.\n3. Comprehensive experiments (vs. SOTA models, agent frameworks, humans) clearly demonstrate the benchmark’s challenge and the framework’s effectiveness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Motivation and workflow design lack theoretical rigor, leaning overly toward engineering. The paper frames its work primarily around \"fixing practical tool limitations\" but fails to anchor this in broader research gaps. While DAG construction (via LCS-aware algorithms) and noise injection (reverse-objective method) are technically detailed, the paper offers little analysis of their theoretical significance (e.g., why these designs effectively test model capabilities). The result is work that reads more like a tool-building project than a research contribution that advances conceptual understanding of data governance automation.\n2. Limited novel research insights. Core components like \"contract-guided planning\" and \"meta-cognitive debugging\" are presented as implementation steps (e.g., \"how to extract contracts\" or \"how to generate debug feedback\"). The paper doesn’t elaborate on why these methods outperform alternatives beyond experimental results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917033580,"tcdate":1761898817456,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3795/Reviewer_uuCg"],"signatures":["ICLR.cc/2026/Conference/Submission3795/Reviewer_uuCg"],"forum":"6CBLcRkuaN","number":2,"license":"CC BY 4.0","cdate":1761898817456,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3795/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917033580,"domain":"ICLR.cc/2026/Conference","replyto":"6CBLcRkuaN","id":"xKFos6po0v","forumContent":{"TLDR":{"value":"A benchmark and agentic system for automating data governance: from natural language instructions to executable data pipelines."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["benchmarks","agent","data governance","large language models"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Data governance is essential for scaling modern AI development. To automate data governance, numerous tools and models have emerged that translate user intent into executable governance code. However, the effectiveness of existing tools and models is largely unverified. The evaluation is severely hampered by the lack of a realistic, standardized, and quantifiable benchmark. This critical gap presents a significant obstacle to systematically evaluating utility and impedes further innovation in the field. To bridge this gap, we introduce GovBench, a benchmark featuring a diverse set of tasks with targeted noise to simulate real-world scenarios and standardized scoring scripts for reproducible evaluation. Our analysis reveals that current data governance tools and models struggle with complex, multi-step workflows and lack robust error-correction mechanisms. We therefore propose DataGovAgent, a novel framework for end-to-end data governance utilizing a Planner-Executor-Evaluator architecture. This design incorporates contract-guided planning, retrieval from a reliable operator library, and sandboxed meta-cognitive debugging. Experimental results validate our approach: DataGovAgent significantly boosts the Average Task Score (ATS) on complex Directed Acyclic Graph (DAG) tasks from 39.7 to 54.9 and reduces debugging iterations by over 77.9\\% compared to general-purpose agent frameworks, a step toward more reliable automation of data governance. Code is available at https://anonymous.4open.science/r/GovBench-F6C6."},"_bibtex":{"value":"@misc{\nliu2025govbench,\ntitle={GovBench: From Natural Language to Executable Pipelines, A New Benchmark for Data Governance Automation},\nauthor={Zhou Liu and ZhaoYang Han and Guochen Yan and Zeli Su and Bohan Zeng and Hao Liang and Xiaochen Ma and Yuanfeng SONG and Xing Chen and Wentao Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=6CBLcRkuaN}\n}"},"title":{"value":"GovBench: From Natural Language to Executable Pipelines, A New Benchmark for Data Governance Automation"},"pdf":{"value":"/pdf/af459d142ccef3a2aeb6c08152722cdc886f5786.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|govbench_from_natural_language_to_executable_pipelines_a_new_benchmark_for_data_governance_automation"},"authorids":{"value":["~Zhou_Liu6","~ZhaoYang_Han2","~Guochen_Yan1","~Zeli_Su1","~Bohan_Zeng1","~Hao_Liang7","~Xiaochen_Ma1","~Yuanfeng_SONG1","~Xing_Chen7","~Wentao_Zhang1"]},"authors":{"value":["Zhou Liu","ZhaoYang Han","Guochen Yan","Zeli Su","Bohan Zeng","Hao Liang","Xiaochen Ma","Yuanfeng SONG","Xing Chen","Wentao Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper presents HERON, a framework for long-horizon human–robot collaboration (HRC) that integrates large language models (LLMs), physics-based reasoning, and mixed-integer linear programming (MILP). HERON decomposes a natural language task into structured sub-tasks via an LLM-based task graph generator, estimates execution time and agent assignment using a physics-guided LLM, and computes an optimized task schedule with MILP. The system further monitors execution to handle human uncertainties such as performance variability, interruptions, and dynamic goal changes through re-planning. Experiments in simulated household tasks show improved task success rate and efficiency compared to existing LLM-based planners."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How robust is HERON if the environment description \\$\\mathcal{E}\\$ is noisy or incomplete — for instance, if certain object attributes or coordinates are missing? Could the Task Decomposition LLM still produce usable task graphs?\n\n2. Have you quantitatively evaluated the EET predictions against measured execution times in simulation? \n\n3. How costly is each re-planning cycle?\n\n4. Do you envision extending HERON to handle raw visual or multimodal inputs (e.g., VLM-based grounding) to overcome dependence on symbolic scene JSONs?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper formalizes LLM-based task decomposition into a structured graph representation, bridging natural language reasoning and symbolic planning.\n\n2. The physics-guided time estimation and MILP optimization provide a principled way to achieve physically feasible and efficient task allocation between humans and robots.\n\n3. The system’s ability to handle failures, interruptions, and goal changes through iterative re-planning is a meaningful advancement toward resilient HRC."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The task decomposition LLM relies on structured scene descriptions (object lists, coordinates) that must be provided explicitly. This limits HERON’s applicability to simulators or digital twins, since real-world visual perception rarely yields such clean symbolic input.\n\n2. The time estimation (EET) module’s accuracy and reliability are not evaluated. There is no ablation or comparison between predicted and actual execution times, making it unclear how critical or accurate this module is to system performance.\n\n3. The experiments are limited to simulation (AI2-THOR) with synthetic human models. The real-world robustness of HERON, especially given its reliance on structured metadata rather than sensory input, remains unproven."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921256284,"tcdate":1762237323056,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9765/Reviewer_Gs3j"],"signatures":["ICLR.cc/2026/Conference/Submission9765/Reviewer_Gs3j"],"forum":"QobJeymX6Z","number":3,"license":"CC BY 4.0","cdate":1762237323056,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9765/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921256284,"domain":"ICLR.cc/2026/Conference","replyto":"QobJeymX6Z","id":"C3tCxF49mi","forumContent":{"TLDR":{"value":"A framework that combines LLM-based task decomposition, physics-guided estimation, and MILP optimization to enable efficient and resilient human–robot collaboration under uncertainty."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Human-Robot Collaboration","Long-Horizon Planning","Task Scheduling"]},"supplementary_material":{"value":"/attachment/c0fa68634bf257dcd3b192ea6841103bcf745ee8.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"The integration of humans into long-horizon planning introduces unique challenges that extend beyond conventional robotic task planning. Unlike robots, humans exhibit inherent uncertainty in task execution, including variable performance, unexpected interruptions, and dynamic goal changes, all of which complicate efficient collaboration. To address these challenges, we propose Human-robot collaboration with Efficient and Resilient OptimizatioN for Long-horizon planning (HERON), a novel framework that combines large language models (LLMs), physics-guided reasoning, and optimization techniques. HERON leverages LLMs in two complementary roles: (i) decomposing natural language task descriptions into structured sub-tasks with agent assignments, and (ii) generating physics-guided execution time estimates and determining sub-task assignments for both human and robot agents based on physical constraints and complementarities. These outputs are incorporated into a mixed-integer linear programming scheduler, which dynamically re-schedules based on observed human uncertainties. This integration ensures that scheduling is not only feasible with respect to physical limitations but also robust to human unpredictability while maintaining efficiency in resource and time allocation. Experiments demonstrate that HERON enables resilient and adaptive human-robot collaboration, achieving more efficient scheduling and higher task success rates compared to existing LLM-based planning frameworks. Website at https://sites.google.com/view/heron-planner."},"_bibtex":{"value":"@misc{\nkim2026heron,\ntitle={{HERON}: Human-robot collaboration with Efficient and Resilient OptimizatioN for Long-horizon planning},\nauthor={Taehyeon Kim and Gyeongmin Kim and E. Cho Smith and Byung-Cheol Min},\nyear={2026},\nurl={https://openreview.net/forum?id=QobJeymX6Z}\n}"},"title":{"value":"HERON: Human-robot collaboration with Efficient and Resilient OptimizatioN for Long-horizon planning"},"pdf":{"value":"/pdf/50037fd37f8eac3749f9a003072abfa94ece03ef.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|heron_humanrobot_collaboration_with_efficient_and_resilient_optimization_for_longhorizon_planning"},"authorids":{"value":["~Taehyeon_Kim3","~Gyeongmin_Kim3","~E._Cho_Smith1","~Byung-Cheol_Min1"]},"authors":{"value":["Taehyeon Kim","Gyeongmin Kim","E. Cho Smith","Byung-Cheol Min"]}},"version":2},{"content":{"venue":{"value":"Computational Imaging 2022"},"pdf":{"value":"https://library.imaging.org/admin/apis/public/api/ist/website/downloadArticle/ei/34/14/COIMG-306"},"venueid":{"value":"dblp.org/conf/CIMAGING/2022"},"paperhash":{"value":"price|inferring_surface_properties_of_oscillating_fluids_from_video_by_inversion_of_physics_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Bob_Price:","https://dblp.org/search/pid/api?q=author:Svyatoslav_Korneev:","https://dblp.org/search/pid/api?q=author:Adrian_Lew:","https://dblp.org/search/pid/api?q=author:Christoforos_Somarakis:","~Raja_Bala3","https://dblp.org/search/pid/api?q=author:Jonathan_(Shengtai)_Ju:"]},"html":{"value":"https://doi.org/10.2352/EI.2022.34.14.COIMG-306"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cimaging/PriceKLSBJ22,\n  author={Bob Price and Svyatoslav Korneev and Adrian Lew and Christoforos Somarakis and Raja Bala and Jonathan Shengtai Ju},\n  title={Inferring surface properties of oscillating fluids from video by inversion of physics models},\n  year={2022},\n  cdate={1640995200000},\n  pages={1-7},\n  url={https://doi.org/10.2352/EI.2022.34.14.COIMG-306},\n  booktitle={Computational Imaging},\n  crossref={conf/cimaging/2022}\n}\n"},"abstract":{"value":"Abstract Measuring the shape, motion and physical properties of os- cillating fluids is critical for understanding the physics of fluidic systems, as well as optimizing and controlling such systems in real time. Conventional surface measurement techniques such as profile analysis or stereo reconstruction are not effective for mon- itoring fluids in industrial processes due to the presence of oc- cluding structures, extreme heat, and complex light interactions at the fluid surface. We propose a video-based method comprising forward and inverse transforms. The forward transform employs a physics-based fluid surface model combined with a ray-traced renderer to map shape and motion parameters to synthetic video frames. The inverse transform uses machine learning models to recover surface parameters from video. The inverse models are trained on synthetic data generated by the forward transform. We illustrate the method on an industrial 3D printer for which we recover the motion and surface of a molten aluminum alloy os- cillating inside a microscopic nozzle. The inverse transform is ill-posed, but can be regularized. We show that surface properties can be reliably inferred with either a suitably regularized non- parametric k-nearest neighbor regressor or a deep convolutional network whose results are less stable but faster to compute."},"title":{"value":"Inferring surface properties of oscillating fluids from video by inversion of physics models"},"authors":{"value":["Bob Price","Svyatoslav Korneev","Adrian Lew","Christoforos Somarakis","Raja Bala","Jonathan Shengtai Ju"]}},"tmdate":1762725871709,"pdate":1640995200000,"externalIds":["dblp:conf/cimaging/PriceKLSBJ22"],"tcdate":1762725856649,"writers":["~"],"signatures":["~Raja_Bala3"],"forum":"hIm3TwxlsY","license":"CC BY-SA 4.0","number":683785,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762725871709,"domain":"DBLP.org","id":"hIm3TwxlsY","version":2},{"content":{"summary":{"value":"The paper asks whether interaction helps VLMs acquire generalizable intuitive physics. Using Two TDW tower datasets and four tasks, the authors fine-tune Qwen2.5 VL with SFT or GRPO. Both reach near ceiling on the task they train on, yet show little transfer to related tasks. Linear probes can decode relevant physical quantities from activations, but this competence does not translate into zero-shot performance. Additional SFT on a new task learns faster than from the base model. Overall the study delivers careful negative results (which I appreciate a lot that the authors shared this.)"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- What happens with multi-step interaction and longer action horizons?\n- Does heavy domain randomization of textures, lighting, camera pose, and block size improve transfer?\n- How does joint multitask SFT across all four tasks compare to single task post-training?\n- Do larger backbones or different families change the outcome?\n- Can you test on a external stability dataset to validate generality?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Clear question and hypotheses grounded in cognitive science\n- Controlled comparison of SFT and RL with matched PEFT settings and budgets\n- Simple tasks with explicit rewards and prompt templates, plus training logs\n- Generalization matrix across all train and test task pairs\n- Decodability analysis that separates representation competence from output performance\n- Useful visualization of reward landscapes and attention maps\n- Negative results are reported transparently"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Very narrow scope. One model family at one size and one environment\n- Interaction is minimal. One step RL with short textual actions, not true multi-step closed loop control\n- Fixed camera and block sizes make pixel shortcuts likely, which undermines conclusions about physics learning\n- No baselines for multitask SFT, joint training across tasks, or auxiliary representation losses\n- Linear probe dataset is small and lacks controls such as image only probes or interventions"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943103829,"tcdate":1761848096568,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24501/Reviewer_Tj6t"],"signatures":["ICLR.cc/2026/Conference/Submission24501/Reviewer_Tj6t"],"forum":"XdLgOm5giq","number":2,"license":"CC BY 4.0","cdate":1761848096568,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24501/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943103829,"domain":"ICLR.cc/2026/Conference","replyto":"XdLgOm5giq","id":"3ry1VY4IN5","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Vision language models","Intuitive physics","Interaction","Cognitive Science","Computational Cognitive Science","Human-like machine learning"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"abstract":{"value":"Pre-trained vision language models do not have good intuitions about the physical world. Recent work has shown that supervised fine-tuning can improve model performance on simple physical tasks. However, fine-tuned models do not appear to learn robust physical rules that can generalize to new contexts. Based on research in cognitive science, we hypothesize that models need to interact with an environment to properly learn its physical dynamics. We train models that learn through interaction with the environment using reinforcement learning, as well as models that learn without interaction using supervised fine-tuning. While both reinforcement learning and supervised fine-tuning appear to improve within-task performance, they fail to produce models with generalizable physical intuitions. Models trained on one task do not reliably generalize to related tasks, even if they share visual statistics and physical principles, and regardless of whether they are trained through interaction."},"_bibtex":{"value":"@misc{\nbuschoff2026can,\ntitle={Can vision language models learn intuitive physics from interaction?},\nauthor={Luca M. Schulze Buschoff and Konstantinos Voudouris and Can Demircan and Eric Schulz},\nyear={2026},\nurl={https://openreview.net/forum?id=XdLgOm5giq}\n}"},"title":{"value":"Can vision language models learn intuitive physics from interaction?"},"pdf":{"value":"/pdf/1d06abe1f86256716c7d6a870bfac945f106010c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"buschoff|can_vision_language_models_learn_intuitive_physics_from_interaction"},"authorids":{"value":["~Luca_M._Schulze_Buschoff1","~Konstantinos_Voudouris1","~Can_Demircan1","~Eric_Schulz1"]},"authors":{"value":["Luca M. Schulze Buschoff","Konstantinos Voudouris","Can Demircan","Eric Schulz"]}},"version":2},{"content":{"summary":{"value":"Fair4Free introduces an innovative generative model designed to create fair synthetic data without accessing actual datasets. It leverages a technique called data-free distillation, where knowledge is transferred from a teacher model to a smaller student model using only noise as input. The approach is touted for its effectiveness in generating data that adheres to fairness, utility, and quality benchmarks, surpassing existing models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Enhanced Privacy Protection: The model's ability to operate without real data makes it highly suitable for environments with strict privacy regulations or sensitive data restrictions.\n2.  By employing a smaller model architecture, Fair4Free minimizes the computational demands, making it feasible for deployment on less powerful devices, including edge devices.\n3. It consistently outperforms other models in generating synthetic data that scores highly on fairness, utility, and quality metrics, as validated by rigorous experimental evaluations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  The model currently focuses on addressing bias with respect to single sensitive attributes, potentially overlooking complex bias scenarios involving multiple intersecting attributes.\n2. Scalability Concerns: While the model is efficient, scaling it to handle larger or more complex datasets without compromising performance remains a challenge.\n3. Can Fair4Free be adapted to efficiently manage multiple sensitive attributes to tackle intersectional biases more effectively?\n4. In terms of decision-making and predictive accuracy, how do the synthetic datasets generated by Fair4Free compare to those derived from traditional data generation methods?\n5. Adaptability to Data Shifts: What measures can be taken to enhance Fair4Free's robustness against dynamic changes in data distribution that are common in real-world settings?"}},"nonreaders":[],"tmdate":1731428297500,"tcdate":1730481823100,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6131/Reviewer_9nCi"],"signatures":["ICLR.cc/2025/Conference/Submission6131/Reviewer_9nCi"],"forum":"iRgzG5DKgA","number":3,"license":"CC BY 4.0","cdate":1730481823100,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6131/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428297500,"domain":"ICLR.cc/2025/Conference","replyto":"iRgzG5DKgA","id":"lSlqUhCSTQ","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["data fairness","fair generative models","knowledge distillation","latent space distillation","synthetic data","biased data"]},"supplementary_material":{"value":"/attachment/96403b6eb869aa033b81be0ee2f8d2ad8174de3a.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This work presents Fair4Free, a novel generative model to generate synthetic fair data using data-free distillation in the latent space. Fair4Free can work on the situation when the data is private or inaccessible.  In our approach, we first train a teacher model to create fair representation and then distil the knowledge to a student model (using a smaller architecture). The process of distilling the student model is data-free, i.e. the student model does not have access to the training dataset while distilling. After the distillation, we use the distilled model to generate fair synthetic samples. Our extensive experiments show that our synthetic samples outperform state-of-the-art models in all three criteria (fairness, utility and synthetic quality) with a performance increase of 5\\% for fairness, 8\\% for utility and 12\\% in synthetic quality for both tabular and image datasets."},"_bibtex":{"value":"@misc{\nsikder2025fairfree,\ntitle={Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data-Free Distillation},\nauthor={Md Fahim Sikder and Daniel de Leng and Fredrik Heintz},\nyear={2025},\nurl={https://openreview.net/forum?id=iRgzG5DKgA}\n}"},"title":{"value":"Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data-Free Distillation"},"pdf":{"value":"/pdf/bf09a39d8fc93c7601b5f410a98be727c91cb446.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"sikder|fair4free_generating_highfidelity_fair_synthetic_samples_using_datafree_distillation"},"authorids":{"value":["~Md_Fahim_Sikder1","~Daniel_de_Leng1","~Fredrik_Heintz1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Md Fahim Sikder","Daniel de Leng","Fredrik Heintz"]}},"version":2},{"content":{"venue":{"value":"British Machine Vision Conference 2025"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"xu|interactive_occlusion_boundary_estimation_through_exploitation_of_synthetic_data"},"authorids":{"value":["~Lintao_XU3","~Chaohui_Wang1"]},"abstract":{"value":"Occlusion boundaries (OBs) geometrically localize occlusion events in 2D images and provide critical cues for scene understanding. In this paper, we present the first systematic study of Interactive Occlusion Boundary Estimation (IOBE), introducing MS3PE – a novel multi-scribble-guided deep-learning framework that advances IOBE through two key innovations: (1) an intuitive multi-scribble interaction mechanism, and (2) a 3-encoding-path network enhanced with multi-scale strip convolutions. Our MS3PE surpasses adapted baselines from seven state-of-the-art interactive segmentation methods, and demonstrates strong potential for OB benchmark construction through our real-user experiment. Besides, to address the scarcity of well-annotated real-world data, we propose using synthetic data for training IOBE models, and developed Mesh2OB, the first automated tool for generating precise ground-truth OBs from 3D scenes with self- occlusions explicitly handled, enabling creation of the OB-FUTURE synthetic benchmark that facilitates generalizable training without domain adaptation. Finally, we introduce OB-LIGM – a high-quality real-world benchmark comprising 120 meticulously annotated high-resolution images advancing evaluation standards in OB research. Source code and resources are available at https://github.com/xul-ops/IOBE."},"title":{"value":"Interactive Occlusion Boundary Estimation through Exploitation of Synthetic Data"},"authors":{"value":["Lintao XU","Chaohui Wang"]}},"tmdate":1771815230212,"pdate":1764518400000,"tcdate":1771815230212,"writers":["~Lintao_XU3","~Chaohui_Wang1"],"signatures":["~Chaohui_Wang1"],"forum":"Bw9D549wtg","license":"CC BY 4.0","number":45278,"cdate":1771815230212,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1771815230212,"domain":"OpenReview.net/Archive","id":"Bw9D549wtg","version":2},{"content":{"comment":{"value":"---\n\n**Q1. Experiment (v) still outperforms the proposed one on the _Dynamic-U_ data and performs on par with the proposed one for _NAO robot_, but significantly degrades on the _Eigenmike_. Please expand on this discrepancy?**\n\n* Experiment (v) used the same logit-based initialization as the proposed method but did not update the LNuDFT parameters during training. Thus, its frequency-bin allocation remained fixed and could not adapt to unseen microphone geometries or acoustic conditions.\n* Synthetic domain (_Dynamic-U_).\n\n  Although _Dynamic-U_ contained unseen channel counts, its overall acoustic conditions were still highly consistent with the simulated training domain (i.e., identical RIR generation process). Under such controlled settings, even a fixed initialization remained effective, which explained why experiment (v) occasionally performed on par with — or slightly better than — the learned version.\n* Real domain (_Eigenmike_).\n\n  The Eigenmike recordings differed substantially from the training conditions in several aspects:\n\n  1. a different number of channels, with a relatively large 32-channel microphone array,\n  2. real-world factors such as sensor mismatch, microphone manufacturing tolerances, and non-ideal RIRs.\n\n  These mismatches degraded the performance of experiment (v), which could not adapt its frequency allocation to compensate for these discrepancies. In contrast, the proposed method learned and adjusted the LNuDFT parameters during training, allowing the model to refine its initial allocation and generalize more effectively to unseen geometries in real-world acoustics.\n\n---\n\n**Q2. In L70, “Both components facilitate the extraction of spatial representations with physics-based inductive biases”: Why are LNuDFT and rMPE considered as imposing (physics-based) inductive biases? To me, these are perceived as a process that introduces and makes learnable parameters for feature extraction more flexible, and I see it as a process that relieves the inductive bias.**\n\n* We intended to convey that LNuDFT and rMPE incorporate physics-based principles—Fourier analysis and inter-microphone phase relationships—into the feature extraction process. These domain principles guide the model toward acoustically meaningful spatial cues and thus serve as physics-informed inductive biases that improve data efficiency and generalization.\n\n  At the same time, certain parameters of LNuDFT are made learnable, which introduces flexibility and allows the model to adapt to the statistics of the training data. This design balances inductive structure with learnable adaptability, rather than removing inductive bias entirely.\n\n  To avoid misunderstanding, we revised the wording in Lines 70–73 of the manuscript as follows:\n\n  “Both components incorporate physics-based structural assumptions into the feature extraction process while still allowing adaptivity through training. This yields physics-informed inductive biases that guide learning toward acoustically meaningful representations.”\n\n---\n\n**Q3. Can the LNuDFT initialization or update scheme get stuck in poor local minima, e.g., if the initial frequency allocation is far from optimal? Have the authors tried more physically motivated or data-driven initializations beyond logit-based mapping?**\n\n* LNuDFT, like other gradient-based models, is in principle susceptible to local minima. In our experiments, however, both uniform initialization (equivalent to the standard DFT) and the proposed logit-based initialization consistently converged to similar solutions, allocating denser frequency bins in the mid-frequency range. This indicates that the optimization was relatively stable under the conditions we tested. We agree that extremely poor initialization may lead to suboptimal convergence, and a more systematic analysis of the optimization landscape, as well as the development of robust training strategies, would be an important direction for future work.\n* Regarding alternative initializations, we experimented with:\n\n  1. Uniform initialization, which reliably converged to mid-frequency–dense allocations (Appendix A.2), motivating the use of a physics-informed prior via logit-based initialization,\n  2. Warped DFT initialization [1], which can emphasize low-frequency regions with a nonlinear frequency warping. However, these low-frequency–focused allocations did not improve performance, as mid-frequency IPD cues were more informative for SSL in our target domain.\n\n  Therefore, among the examined schemes, the logit-based initialization provided the most effective inductive prior for learning LNuDFT parameters. Future work could explore more sophisticated data-driven or physics-informed initialization schemes to further enhance convergence and performance.\n\n\n  [1] A Makur and S K Mitra. Warped discrete-Fourier transform: Theory and applications. IEEE Transactions on Circuits and Systems I: Fundamental Theory and Applications, vol. 48, no. 9, pp. 1086-1093, 2001.\n\n---"},"title":{"value":"Response to Reviewer aRq1 (2/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763628115083,"tcdate":1763628115083,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10505/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission10505/Authors"],"forum":"bWXpJFesLS","number":2,"license":"CC BY 4.0","cdate":1763628115083,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10505/-/Official_Comment"],"mdate":1763628115083,"domain":"ICLR.cc/2026/Conference","replyto":"inIORM9Ln3","id":"Rz53yCQyBI","forumContent":{"TLDR":{"value":"This paper proposes audio-geometry-grid representation learning for grid-flexible and geometry-invariant sound source localization, leveraging learnable non-uniform discrete Fourier transform and relative microphone positional encoding."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Sound Source Localization","Geometry-Invariant","Grid-Flexible","Representation Learning","Physics-Informed Design","Learnable Non-uniform DFT","Relative Microphone Positional Encoding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Sound source localization (SSL) is a fundamental task in spatial audio understanding, yet most deep neural network-based methods are constrained by fixed array geometries and predefined directional grids, limiting generalizability and scalability. To address these issues, we propose _audio-geometry-grid representation learning_ (AGG-RL), a novel framework that jointly learns audio-geometry and grid representations in a shared latent space, enabling both geometry-invariant and grid-flexible SSL. Moreover, to enhance generalizability and interpretability, we introduce two physics-informed components: a _learnable non-uniform discrete Fourier transform_ (LNuDFT), which optimizes the dense allocation of frequency bins in a non-uniform manner to emphasize informative phase regions, and a _relative microphone positional encoding_ (rMPE), which encodes relative microphone coordinates in accordance with the nature of inter-channel time differences. Experiments on synthetic and real datasets demonstrate that AGG-RL achieves superior performance, particularly under unseen conditions. The results highlight the potential of representation learning with physics-informed design towards a universal solution for spatial acoustic scene understanding across diverse scenarios."},"_bibtex":{"value":"@inproceedings{\nbaek2026physicsinformed,\ntitle={Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization},\nauthor={Min-Sang Baek and Gyeong-Su Kim and Donghyun Kim and Joon-Hyuk Chang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=bWXpJFesLS}\n}"},"title":{"value":"Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization"},"pdf":{"value":"/pdf/dcd11d5d38c6a0bf65ded06ef46bf42c740dab7c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"baek|physicsinformed_audiogeometrygrid_representation_learning_for_universal_sound_source_localization"},"authorids":{"value":["~Min-Sang_Baek1","~Gyeong-Su_Kim1","~Donghyun_Kim20","~Joon-Hyuk_Chang1"]},"authors":{"value":["Min-Sang Baek","Gyeong-Su Kim","Donghyun Kim","Joon-Hyuk Chang"]}},"version":2},{"content":{"summary":{"value":"This paper presents a systematic analysis of how prompt complexity influences the quality, diversity, and consistency of synthetic data from text-to-image models. It establishes that increasing prompt complexity reduces diversity and consistency while improving the realism of generated images. The authors provide theoretical insight, showing that models struggle to generalize to general prompts, and introduce a comprehensive benchmarking framework for evaluation. Their large-scale empirical study identifies prompt expansion combined with advanced guidance like APG as the most effective approach for achieving optimal utility trade-offs in synthetic data generation."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"1.The authors rely on the DSG score for consistency evaluation. Could they clarify how they ensured this metric's validity for the very long, complex prompts from the DCI dataset? Providing a small-scale human evaluation to correlate with the DSG scores on these long prompts would help solidify the consistency findings.\n\n2.The authors note a drop in consistency for highly complex prompts. Could they comment on the semantic plausibility of these generations? Is the failure primarily in missing attributes, or also in generating globally incoherent scenes?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1.Introduces the novel problem of systematically evaluating \"prompt complexity\" in T2I models, combining theoretical insight (\"OR\" vs. \"AND\" generalization) with practical interventions.\n\n2.Methodologically rigorous, featuring a well-designed benchmarking framework, compelling synthetic experiments, and a large-scale, comprehensive evaluation across models and datasets.\n\n3.Exceptionally well-structured and clearly written, with a logical narrative flow and effective figures that make the complex study accessible and reproducible.\n\n4.Provides crucial guidance for synthetic data generation and model evaluation, revealing fundamental trade-offs and setting an important agenda for future T2I model development."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The paper relies on the DSG score for consistency evaluation but does not critically address its known limitations with very long, complex prompts (as in the DCI dataset). As prompt length increases, the VQA models underlying DSG may themselves fail, potentially conflating model inconsistency with VQA model error. A targeted analysis, such as a human evaluation on a subset of long prompts to verify DSG's correlation with human judgment in this regime, is needed to ensure the reported consistency drop is reliable.\n\n2.While prompt expansion is highlighted as a powerful intervention, the paper provides a limited discussion of its significant downside: it actively moves generations outside the support of the reference real data (as shown by reduced precision/density). This is a critical trade-off that is under-explored. The work would be improved by a deeper analysis of when this \"hallucination\" is beneficial (e.g., for data augmentation) versus detrimental (e.g., for faithful dataset replication), framing it not just as a boost in diversity but as a fundamental shift in the data distribution."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762939057121,"tcdate":1761892030532,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20963/Reviewer_PuV8"],"signatures":["ICLR.cc/2026/Conference/Submission20963/Reviewer_PuV8"],"forum":"RBIBMCdw7y","number":2,"license":"CC BY 4.0","cdate":1761892030532,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20963/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762939057121,"domain":"ICLR.cc/2026/Conference","replyto":"RBIBMCdw7y","id":"QvEQfe8H6Y","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["text-to-image models","prompt complexity","synthetic data"]},"supplementary_material":{"value":"/attachment/19f32551040643a1e2887ebe2e8ce7caa29f875a.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Text-to-image (T2I) models offer great potential for creating virtually limitless synthetic data, a valuable resource compared to fixed and finite real datasets. Previous works evaluate the utility of synthetic data from T2I models on three key desiderata: quality, diversity, and consistency. While prompt engineering is the primary means of interacting with T2I models, the systematic impact of prompt complexity on these critical utility axes remains underexplored. In this paper, we first conduct synthetic experiments to motivate the difficulty of generalization w.r.t. prompt complexity and explain the observed difficulty with theoretical derivations. Then, we introduce a new evaluation framework that can compare the utility of real data and synthetic data, and present a comprehensive analysis of how prompt complexity influences the utility of synthetic data generated by commonly used T2I models. We conduct our study across diverse datasets, including CC12M, ImageNet-1k, and DCI, and evaluate different inference-time intervention methods. Our synthetic experiments show that generalizing to more general conditions is harder than the other way round, since the former needs an estimated likelihood that is not learned by diffusion models. Our large-scale empirical experiments reveal that increasing prompt complexity results in lower conditional diversity and prompt consistency, while reducing the synthetic-to-real distribution shift, which aligns with the synthetic experiments. Moreover, current inference-time interventions can augment the diversity of the generations at the expense of moving outside the support of real data. Among those interventions, prompt expansion, by deliberately using a pre-trained language model as a likelihood estimator, consistently achieves the highest performance in both image diversity and aesthetics, even higher than that of real data. Combining advanced guidance interventions with prompt expansion results in the most appealing utility trade-offs of synthetic data."},"_bibtex":{"value":"@inproceedings{\nzhang2026the,\ntitle={The Intricate Dance of Prompt Complexity, Quality, Diversity and Consistency in T2I Models},\nauthor={Xiaofeng Zhang and Aaron Courville and Michal Drozdzal and Adriana Romero-Soriano},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=RBIBMCdw7y}\n}"},"title":{"value":"The Intricate Dance of Prompt Complexity, Quality, Diversity and Consistency in T2I Models"},"pdf":{"value":"/pdf/bafe0f0325f15c351788c224c0b6b3caaac71dcd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|the_intricate_dance_of_prompt_complexity_quality_diversity_and_consistency_in_t2i_models"},"authorids":{"value":["~Xiaofeng_Zhang2","~Aaron_Courville3","~Michal_Drozdzal1","~Adriana_Romero-Soriano1"]},"authors":{"value":["Xiaofeng Zhang","Aaron Courville","Michal Drozdzal","Adriana Romero-Soriano"]}},"version":2},{"content":{"summary":{"value":"1. The proposed direction-disentangled 3DGS (DDGS) method decomposes the radiosity contribution into isotropic and direction-dependent components, able to approximate complex anisotropic interactions without complex runtime simulations. specifically, it modeling isotropic and anisotropic contributions via distinct 3D Gaussians.\n\n2. The paper clearly mentioned that 3DGS-for-DRR solutions [30, 11] do not account for noise-inducing photon interactions (e.g., scattering) when applying the analytical methods [29, 5] to render their training data."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. what is the means of high-dimensional residual contribution？Is there a formal description? This concept is not explained in the full text.\n\n2. Before use the method, it need first get the 3D CT volume. The reconstruction is based on the projection obtained by the CT machine. However, there are errors and noise in the 3D CT volume (from the reconstruction of the projection noise) and also from the reconstruction errors caused by the polychromatic X-rays in medical CT machine. The paper is equivalent to reprojecting the 3D CT volume using gaussian splatting. How do the work ensure the reprojection process does not amplify this error from the 3D CT volume?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper has a clear definition of the problem to be solved and a good visualization.\n2. Experiments combining 2D/3D CT Image registration were also presented."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although Gaussian Splatting is widely used in natural color scenes, this paper does not provide a detailed mathematical description of the \"migrated version\" of Gaussian Splatting in combination with the physical scene of CT.  Instead, it directly use the mathematical description of Gaussian Splatting in natural color scenes (Equation (1)). In CT imaging, the contribution of a voxel to the final projection does not suffer from attenuation due to occlusion by the following voxels, but is instead Beer-Lambert law. Therefore, equation (1) no longer holds. This makes it difficult to highlight the motivation of this paper in terms of algorithm details.\n\n2.  Theoretical contributions appear to be minor. (Formula (3) is the theoretical contributions of this paper). This makes this paper more suitable for delivery to places that are more concerned with simulation applications.\n\n3. The paper needs to compare with Physics-based Monte-Carlo simulations [2, 15, 3 , 4 ]. Although these methods are time-consuming, they guarantee the quality of reprojection. The paper needs to clearly point out the difference in image quality with Physics-based Monte-Carlo simulations. The ground truth(GT) needs to select the real projection obtained by the CT machine. So the comparison table will be：\n______________________________________________________________________________________\n                         |    Physics-based Monte-Carlo simulations     |     3DGS   |  X-Gaussian  |   DDGS\n_____________________________________________________________________________________\nPSNR with GT |  \n______________________________________________________________________________________\nSSIM  with GT | \n______________________________________________________________________________________"},"limitations":{"value":"1. If this method is used in clinical practice, what problems will it cause without considering the time cost? For example, the appearance of artifacts, because once artifacts appear, it will seriously interfere with the doctor's clinical operation."}},"nonreaders":[],"tmdate":1730878854108,"tcdate":1720533212213,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission3496/Reviewer_g5Vp"],"signatures":["NeurIPS.cc/2024/Conference/Submission3496/Reviewer_g5Vp"],"forum":"mY0ZnS2s9u","number":1,"license":"CC BY 4.0","cdate":1720533212213,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission3496/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878854108,"domain":"NeurIPS.cc/2024/Conference","replyto":"mY0ZnS2s9u","id":"9KLbiJYzKy","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"A novel approach that balances realistic physics-inspired X-ray simulation with efficient, differentiable DRR generation using 3D Gaussian splatting (3DGS)"},"keywords":{"value":["3D Gaussian splatting","image registration","pose estimation"]},"primary_area":{"value":"machine_learning_for_healthcare"},"abstract":{"value":"Digitally reconstructed radiographs (DRRs) are simulated 2D X-ray images generated from 3D CT volumes, widely used in preoperative settings but limited in intraoperative applications due to computational bottlenecks. Physics-based Monte Carlo simulations provide accurate representations but are extremely computationally intensity. Analytical DRR renderers are much more efficient, but at the price of ignoring anisotropic X-ray image formation phenomena such as Compton scattering. We propose a novel approach that balances realistic physics-inspired X-ray simulation with efficient, differentiable DRR generation using 3D Gaussian splatting (3DGS). Our direction-disentangled 3DGS (DDGS) method decomposes the radiosity contribution into isotropic and direction-dependent components, able to approximate complex anisotropic interactions without complex runtime simulations. Additionally, we adapt the 3DGS initialization to account for tomography data properties, enhancing accuracy and efficiency. Our method outperforms state-of-the-art techniques in image accuracy and inference speed, demonstrating its potential for intraoperative applications and inverse problems like pose registration."},"_bibtex":{"value":"@inproceedings{\ngao2024ddgsct,\ntitle={{DDGS}-{CT}: Direction-Disentangled Gaussian Splatting for Realistic Volume Rendering},\nauthor={Zhongpai Gao and Benjamin Planche and Meng Zheng and Xiao Chen and Terrence Chen and Ziyan Wu},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=mY0ZnS2s9u}\n}"},"title":{"value":"DDGS-CT: Direction-Disentangled Gaussian Splatting for Realistic Volume Rendering"},"pdf":{"value":"/pdf/0c564dbb6edaaa1d191947e75b2e6803793b4e5d.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"gao|ddgsct_directiondisentangled_gaussian_splatting_for_realistic_volume_rendering"},"authorids":{"value":["~Zhongpai_Gao1","~Benjamin_Planche1","~Meng_Zheng1","~Xiao_Chen12","~Terrence_Chen4","~Ziyan_Wu2"]},"authors":{"value":["Zhongpai Gao","Benjamin Planche","Meng Zheng","Xiao Chen","Terrence Chen","Ziyan Wu"]}},"version":2},{"content":{"summary":{"value":"The paper explores the direction of injecting inductive biases (into Transformers) by modifying the data (particularly, through synthetic pre-training). The paper tries to inject Finite State Transducer (FST)-like biases by generating relevant synthetic pre-training data. The paper pre-trains a Transformer on the synthetic data and test it for OOD generalization during fine-tuning different FST tasks. The paper found better OOD generalization with FST pre-training and prefixes than other baselines. The paper also shows positive transfer on some natural language tasks."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1. Decent focused exploration on injection of inductive bias through synthetic pre-training. \n\n2. Shows the ability to demonstrate OOD generalizations (iteration generalization and systematic generalization) in FST-tasks from synthetic pre-training.\n\n3. Shows transfer from pre-training on FST to some specific natural language tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. If By-T5 is already pre-trained in natural data before the synthetic pre-training, I wonder how much of an influence there is from the \"pre-pre-training\" in enabling OOD generalization and such. \n\n2. The scope feels limited. We already have prior works showing the viability of synthetic pre-training and knowledge transfer from natural language tasks. It is not as clear what the motivation for exploring particularly FST-related tasks is. Transformers have been shown to underperform in OOD generalization on logical inference [1,2], ListOps [2], Flip-Flop languages [3], parity tasks/sensitive tasks [4], automata tasks [5], and others [6]. It would have been good to contrast the approach with some of such works, reconcile with them, and see if the synthetic pre-training proposed here can be used. \n\n[1] The Importance of Being Recurrent for Modeling Hierarchical Structure - Tran. et al. EMNLP 2018\n\n[2] Ordered Memory - Shen et al. NeurIPS 2019\n\n[3] Exposing Attention Glitches with Flip-Flop Language Modeling - Liu et al. NeurIPS 2023\n\n[4] Simplicity Bias in Transformers and their Ability to Learn Sparse Boolean Functions - Bhattamishra et al. ACL 2023\n\n[5] Transformers Learn Shortcuts to Automata  - Liu et al. ICLR 2023\n\n[6] Neural Networks and the Chomsky Hierarchy - Delétang et al. ICLR 2023"},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. What are the average and maximum sequence lengths in the pre-training data, training data, and iteration generalization data? \n\n2. Would it be possible to explore generalizations to higher lengths e.g. 100 or more?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636604121,"tcdate":1698732886208,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5755/Reviewer_sUW9"],"signatures":["ICLR.cc/2024/Conference/Submission5755/Reviewer_sUW9"],"forum":"Oashk4fDD9","number":1,"license":"CC BY 4.0","cdate":1698732886208,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5755/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636604121,"domain":"ICLR.cc/2024/Conference","replyto":"Oashk4fDD9","id":"1DHFOZ8NR2","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["systematic generalization","transformers","sequence modelling","finite state methods","natural language processing"]},"supplementary_material":{"value":"/attachment/e5d3f49dcab87ecb83f0f338dff7437a2dbea059.zip"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Strong inductive biases enable learning from little data and help generalization outside of the training distribution. Popular neural architectures such as Transformers lack strong structural inductive biases for seq2seq NLP tasks on their own. Consequently, they struggle with systematic generalization beyond the training distribution, e.g. with extrapolating to longer inputs, even when pre-trained on large amounts of text. We show how a structural inductive bias can be injected into a seq2seq model by pre-training it to simulate structural transformations on synthetic data. Specifically, we inject an inductive bias towards Finite State Transducers (FSTs) into a Transformer by pre-training it to simulate FSTs given their descriptions. Our experiments show that our method imparts the desired inductive bias, resulting in improved systematic generalization and better few-shot learning for FST-like tasks."},"_bibtex":{"value":"@misc{\nlindemann2024injecting,\ntitle={Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation},\nauthor={Matthias Lindemann and Alexander Koller and Ivan Titov},\nyear={2024},\nurl={https://openreview.net/forum?id=Oashk4fDD9}\n}"},"title":{"value":"Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation"},"pdf":{"value":"/pdf/5085e4e32763fbcbc9a1a7357d960d20761b1ae3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"lindemann|injecting_a_structural_inductive_bias_into_a_seq2seq_model_by_simulation"},"authorids":{"value":["~Matthias_Lindemann1","~Alexander_Koller2","~Ivan_Titov1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Matthias Lindemann","Alexander Koller","Ivan Titov"]}},"version":2},{"content":{"summary":{"value":"This paper identifies that large language models (LLMs) struggle with complex, interdependent rule systems, as standard methods like Chain-of-Thought (CoT) treat rules as unstructured text and suffer from error propagation. To address this, the authors propose the Dynamic Adjudication Template (DAT), a three-stage framework inspired by human expert reasoning: (1) Qualitative Assessment for an initial judgment, (2) Evidence Gathering using dynamic templates with placeholders for targeted rule verification, and (3) Adjudication to synthesize the validated evidence into a final decision. The authors also introduce an automated pipeline for generating, filtering, and selecting these templates using a global and local (DPO-trained) selector. Experiments on the EVADE e-commerce benchmark show that DAT enables smaller LLMs to match or exceed the performance of larger models using standard CoT."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Could the authors elaborate on the required effort to adapt this system to a new domain, such as tax law? Would the entire template library and DPO selector need to be rebuilt from scratch?\n\n2. What is the total computational cost of the DAT framework, including the one-time setup (generation, training) and the per-inference cost (selector + three-stage reasoning)? How does this compare to simply using a state-of-the-art model like GPT-4.1 with a more detailed multi-shot CoT prompt?\n\n3. The performance seems to rely on \"distilling\" reasoning patterns from very large models (Gemini-Pro, Qwen-14B) into templates for a 7B model. How much of the performance lift is from the DAT method itself, versus this implicit knowledge distillation?\n\n4. Can you explain the significant performance degradation for Qwen-VL-Max on the \"Body\" task? This seems to be a strong counter-example to the paper's core hypothesis.\n\n5. How robust is the system to errors from the template selector? What happens if the global-local selector, $S_{final}$, chooses a suboptimal template for a given query?"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"No"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. Clear Problem Definition: The paper addresses a well-defined and significant limitation of LLMs in applying complex, interdependent rules, which is a critical capability for domains like law, finance, and content moderation.\n\n2. Intuitive Framework: The proposed three-stage DAT framework (Qualitative Assessment, Evidence Gathering, Adjudication) is logical, intuitive, and well-grounded in how human experts might systematically approach such problems, offering a more structured alternative to free-form reasoning.\n\n3. Promising Efficiency: The central claim that a structured reasoning method can allow smaller, more efficient models to outperform larger, more general-purpose models (on specific tasks) is a valuable and compelling direction for research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Generalizability and Dataset Dependency: The entire DAT pipeline (template generation, filtering, DPO-trained selector, and evaluation) is developed and validated on a single, specific dataset (EVADE). It is highly questionable whether this complex system would generalize to other rule-intensive domains (e.g., legal text, medical guidelines) without a complete and costly re-generation and re-training process. The system risks overfitting to the specific structure and biases of this one benchmark.\n\n2. High System Complexity: The proposed solution is far from simple. It introduces a complex pipeline that relies on multiple large models: a powerful generator (Gemini 2.5-Pro) to create templates and another model (Qwen-2.5-14B) to train the DPO-based local selector. This \"scaffolding\" complexity arguably offsets the final-use benefit of running inference on a smaller model (e.g., Qwen-2.5-7B). The paper does not provide a clear analysis of the total computational cost.\n\n3. Weak Baseline Comparisons: The primary baseline is standard CoT prompting. While DAT shows clear improvements, CoT is a very general-purpose method. The comparison would be much stronger if it included other structured-reasoning baselines, such as verifier models, tool-use frameworks, or Retrieval-Augmented Generation (RAG) systems specifically designed to retrieve and apply rules.\n\n4. Unconvincing VLM Results: The preliminary results on Vision-Language Models (VLMs) in Table 3 are mixed and undermine the paper's claims. For instance, DAT significantly degrades the performance of Qwen-VL-Max on the \"Body\" task (from 80.75 to 65.42). The paper's hypothesis that this is due to weak instruction-following in smaller VLMs does not explain this failure in a large VLM, suggesting the DAT approach may not be robustly applicable to multimodal reasoning."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921786457,"tcdate":1761799279389,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10500/Reviewer_CRuu"],"signatures":["ICLR.cc/2026/Conference/Submission10500/Reviewer_CRuu"],"forum":"yd55HULWM1","number":2,"license":"CC BY 4.0","cdate":1761799279389,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10500/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921786457,"domain":"ICLR.cc/2026/Conference","replyto":"yd55HULWM1","id":"MbWpV40Uxv","forumContent":{"TLDR":{"value":"Our paper introduces the Dynamic Adjudication Template, a novel approach that enables structured reasoning in complex rule systems, empowering smaller models to outperform traditional methods such as CoT and to rival larger state-of-the-art LLMs."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Large Language Models","Complex Rule Systems","Reasoning Divergence","Rule-based Reasoning","Dynamic Adjudication Template"]},"supplementary_material":{"value":"/attachment/4620318b7bdc5cc5b4304e8fe6cd500a78d7685d.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large language models (LLMs) face significant challenges when processing complex rule systems, as they typically treat interdependent rules as unstructured textual data rather than as logically organized frameworks. This limitation results in reasoning divergence, where models often overlook critical rule dependencies essential for accurate interpretation. Although existing approaches such as Chain-of-Thought (CoT) reasoning have shown promise, they lack systematic methodologies for structured rule processing and are particularly susceptible to error propagation through sequential reasoning chains. To address these limitations, we propose the Dynamic Adjudication Template (DAT), a novel framework inspired by expert human reasoning processes. DAT structures the inference mechanism into three methodical stages: *qualitative analysis, evidence gathering, and adjudication*. During the *qualitative analysis* phase, the model comprehensively evaluates the contextual landscape. The subsequent *evidence gathering* phase involves the targeted extraction of pertinent information based on predefined template elements ([placeholder]), followed by systematic verification against applicable rules. Finally, in the *adjudication* phase, the model synthesizes these validated components to formulate a comprehensive judgment. Empirical results demonstrate that DAT consistently outperforms conventional CoT approaches in complex rule-based tasks. Notably, DAT enables smaller language models to match, and in some cases exceed, the performance of significantly larger LLMs, highlighting its efficiency and effectiveness in managing intricate rule systems."},"_bibtex":{"value":"@misc{\nyang2025structuring,\ntitle={Structuring Reasoning for Complex Rules Beyond Flat Representations},\nauthor={Zhihao Yang and Ancheng Xu and Jingpeng Li and Liang Yan and Jiehui Zhou and Zhen Qin and Hengyu Chang and Ahmadreza Argha and Hamid Alinejad-Rokny and Minghuan Tan and Yujun Cai and Min Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=yd55HULWM1}\n}"},"title":{"value":"Structuring Reasoning for Complex Rules Beyond Flat Representations"},"pdf":{"value":"/pdf/423091a168b65d947ec36e22a92c84c4064ce9ce.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yang|structuring_reasoning_for_complex_rules_beyond_flat_representations"},"authorids":{"value":["~Zhihao_Yang6","~Ancheng_Xu1","~Jingpeng_Li4","~Liang_Yan7","~Jiehui_Zhou1","~Zhen_Qin14","~Hengyu_Chang2","~Ahmadreza_Argha1","~Hamid_Alinejad-Rokny1","~Minghuan_Tan1","~Yujun_Cai1","~Min_Yang4"]},"authors":{"value":["Zhihao Yang","Ancheng Xu","Jingpeng Li","Liang Yan","Jiehui Zhou","Zhen Qin","Hengyu Chang","Ahmadreza Argha","Hamid Alinejad-Rokny","Minghuan Tan","Yujun Cai","Min Yang"]}},"version":2},{"content":{"summary":{"value":"The paper proposed InterfaceDiff to design protein-protein complexes with aware interface. It models the intra-chain graph and inter-chain interface graph, and applied diffusion model to achieve design. The proposed method is evaluated on the general protein-protein complex design and target protein binder complex design tasks."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Please see above weaknesses."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"The idea of co-designing protein-protein complex sequence and structure is interesting and useful."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The problem setting of this paper has some issues: The paper assume we already know the protein-protein complex interface, based on which it models the intra-chain graph and inter-chain interface graph. However, in reality, the interface is unknown. Basically, what we know might only be the target protein. If we already know the interface, the complex design task would not be challenging any more. The assumption of this paper is not practical in reality. \n\n2. The experimental setting also has some issues: In table1 and table2, almost all the baselines are inverse folding models, like ProteinMPNN, PiFold. Most of these models are neither sequence-structure co-design models nor protein-protein complex design models. The comparisons are totally unfair and meaning less.\n\n3. The metrics are not indicative enough. For co-design parts, there should be some consistency metrics between the designed sequence and structure. For the complex design parts, there should be something like binding affinity scores or AlphaFold3 ipTM scores. Metrics like the sequence recovery rate is not useful at all."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927598381,"tcdate":1762030310130,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17758/Reviewer_zMYU"],"signatures":["ICLR.cc/2026/Conference/Submission17758/Reviewer_zMYU"],"forum":"4m7Vkchn99","number":4,"license":"CC BY 4.0","cdate":1762030310130,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17758/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927598381,"domain":"ICLR.cc/2026/Conference","replyto":"4m7Vkchn99","id":"RxvsE4lT6T","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Protein-protein complex design","Diffusion models","Graph neural networks","Interface modeling","Sequence–structure co-design"]},"supplementary_material":{"value":"/attachment/b413a19f26be4bc5bb631caa04c2586487bc8ecf.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"The rational design of protein–protein complexes remains a fundamental challenge in synthetic biology and therapeutic development. Current generative methods often fall short in performing sequence–structure co-design, particularly in treating the functionally critical protein-protein interface as a first-class target. To bridge this gap, we present InterfaceDiff, a graph-based diffusion framework for interface-aware co-design of protein complexes. The complex is encoded by intra-chain graphs coupled through an explicit bipartite interface graph, concentrating modeling capacity on physically interacting residues. InterfaceDiff learns a joint distribution over discrete amino acid sequences and continuous local rigid frames (rotations and translations) by a simultaneous denoising process. To achieve this efficiently, we develop a novel graph neural network denoiser inspired by Invariant Point Attention, which performs message passing on the sparse graph representation while avoiding the computational overhead of fully SE(3)-equivariant networks. We evaluate InterfaceDiff across multiple design tasks, demonstrating its ability to generate diverse, high-quality, and physically plausible all-atom complexes. Our method achieves strong performance on key biophysical and geometric metrics, offering a scalable and geometrically efficient approach for controllable protein complex engineering. This work establishes a foundation for generative co-design of novel molecular interactions."},"_bibtex":{"value":"@misc{\nhe2025interfacediff,\ntitle={InterfaceDiff: Interface-Aware Sequence-Structure Co-Design of Protein Complexes with Graph-Based Diffusion},\nauthor={Jing He and Weiguo Zheng and Haoran Qian and Bo Fu},\nyear={2025},\nurl={https://openreview.net/forum?id=4m7Vkchn99}\n}"},"title":{"value":"InterfaceDiff: Interface-Aware Sequence-Structure Co-Design of Protein Complexes with Graph-Based Diffusion"},"pdf":{"value":"/pdf/cd7cbface97f3b16551a715796c1fbdad19d6b60.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"he|interfacediff_interfaceaware_sequencestructure_codesign_of_protein_complexes_with_graphbased_diffusion"},"authorids":{"value":["~Jing_He7","~Weiguo_Zheng1","~Haoran_Qian1","~Bo_Fu8"]},"authors":{"value":["Jing He","Weiguo Zheng","Haoran Qian","Bo Fu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces MOSIV, a new framework created to solve the problem of identifying the physical properties of multiple interacting objects simultaneously from a video. MOSIV works by using a differentiable simulator. It directly optimizes the specific, continuous material parameters for each object by trying to match the geometry observed in the video. MOSIV also presents a new synthetic benchmark with contact-rich, multi-object interactions."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- The method appears to be computationally expensive? It involves (1) optimizing a 4D Gaussian scene and assigning instance partitions, (2) converting object's reconstruction into simulation-ready continuum (3) running a differentiable MPM simulation to optimize per-object parameter vectors, and (4) optimizing this entire unrolled simulation. What are the typical run time and memory footprint? how does it compare with baseline methods?\n- what are the novel interactions in section 3.7? if my understanding is correctly, the physical motions are kept the same, you only change the materials?\n- what consitutive models have been used? how are they chosen, can you provide more details on that?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- the vision of learning physical properties and its interactions directly from videos seems appealing\n- The overall approach is logical and builds upon several state-of-the-art components, including the dynamic Gaussian Splatting for reconstruction and a differentiable MPM for physics-based parameter identification\n- the author introduces a new multi-object dataset with diverse geometry, materials properties and physical motions, that could be used from the community"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- the paper studies multi-object system interaction, but the proposed dataset only contains two-object interactions, also 30 frames of interactions seems quite short for evaluation, given the authors claim that the calibrated models generalizes to \"long-horizon predictions of complex multi-object dynamics\""}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916389857,"tcdate":1761996818233,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2816/Reviewer_yKeE"],"signatures":["ICLR.cc/2026/Conference/Submission2816/Reviewer_yKeE"],"forum":"0ylAe3Orfy","number":3,"license":"CC BY 4.0","cdate":1761996818233,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2816/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916389857,"domain":"ICLR.cc/2026/Conference","replyto":"0ylAe3Orfy","id":"BQSUJtbjqD","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Object Property Identification","Physics-based Modeling"]},"supplementary_material":{"value":"/attachment/63257c056e57720c17cab5e65c7505826802726e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We introduce the challenging problem of multi-object system identification from videos, for which prior methods are ill-suited due to their focus on single-object scenes or discrete material classification with a fixed set of material prototypes. To address this, we propose MOSIV, a new framework that directly optimizes for continuous, per-object material parameters using a differentiable simulator guided by geometric objectives derived from video. We also present a new synthetic benchmark with contact-rich, multi-object interactions to facilitate evaluation. On this benchmark, MOSIV substantially improves grounding accuracy and long-horizon simulation fidelity over adapted baselines, establishing it as a strong baseline for this new task. Our analysis shows that object-level fine-grained supervision and geometry-aligned objectives are critical for stable optimization in these complex, multi-object settings. The source code and dataset will be released."},"_bibtex":{"value":"@inproceedings{\nliu2026multiobject,\ntitle={Multi-Object System Identification from Videos},\nauthor={Chunjiang Liu and Xiaoyuan Wang and Qingran Lin and Albert Xiao and Haoyu Chen and Shizheng Wen and Hao Zhang and Lu Qi and Ming-Hsuan Yang and Laszlo A. Jeni and Min Xu and Yizhou Zhao},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=0ylAe3Orfy}\n}"},"title":{"value":"Multi-Object System Identification from Videos"},"pdf":{"value":"/pdf/c23f0dd11c0234fd016f0daf616563c593d3a411.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|multiobject_system_identification_from_videos"},"authorids":{"value":["~Chunjiang_Liu2","~Xiaoyuan_Wang1","~Qingran_Lin1","~Albert_Xiao1","~Haoyu_Chen15","~Shizheng_Wen1","~Hao_Zhang47","~Lu_Qi1","~Ming-Hsuan_Yang1","~Laszlo_Attila_Jeni1","~Min_Xu4","~Yizhou_Zhao2"]},"authors":{"value":["Chunjiang Liu","Xiaoyuan Wang","Qingran Lin","Albert Xiao","Haoyu Chen","Shizheng Wen","Hao Zhang","Lu Qi","Ming-Hsuan Yang","Laszlo A. Jeni","Min Xu","Yizhou Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper extends Sheaf Neural Networks (SNNs) to directed graphs by introducing the Directed Cellular Sheaf and the corresponding Directed Sheaf Laplacian (DSL), which explicitly encode edge orientation through complex-valued, direction-aware restriction maps. Building on this framework, the authors propose the Directed Sheaf Neural Network (DSNN), enabling principled learning on directed and heterophilic graphs. Experiments on synthetic and real-world datasets demonstrate that DSNN effectively captures directional dependencies and outperforms existing SNN and GNN models where edge directionality is crucial."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. The introduction should better articulate why extending SNNs remains valuable in 2025. Specifically, it should discuss the advantages of combining SNNs with GNNs for handling directed graphs and heterophily, and explain why these structural refinements are still meaningful in the era of more unified and general graph learning paradigms.\n\n2. The comparison set is mostly limited to models from 2020–2022, with only one 2024 method included. More recent GNNs addressing heterophily and oversmoothing (from 2024–2025) should be added to strengthen the empirical evaluation and demonstrate DSNN’s relevance against state-of-the-art models.\n\n3. The paper does not include baselines from graph prompting, graph foundation, or pre-trained graph models, which have recently become dominant in node classification and general graph learning. Including such comparisons would position the work more clearly within the modern landscape of graph representation learning."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper introduces a mathematically principled extension of SNNs to directed graphs through the Directed Cellular Sheaf and Directed Sheaf Laplacian.\n\n2. It effectively captures asymmetric and directional relationships while maintaining robustness to heterophily.\n\n3. The experiments demonstrate consistent performance gains on both synthetic and real-world directed graph datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  Most compared models are from 2020–2022, with only one from 2024. The evaluation lacks more recent direction-aware or topology-based GNNs, which weakens the empirical evidence for DSNN’s claimed advantages.\n\n2. The paper does not clearly justify why Sheaf Neural Networks are the right framework for addressing heterophily or extending to directed graphs. Given the recent shift toward graph foundation models and unified architectures, it remains unclear whether adapting SNNs is the most effective or timely direction, rather than developing more generalizable approaches."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917907118,"tcdate":1761470976165,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5143/Reviewer_uTjh"],"signatures":["ICLR.cc/2026/Conference/Submission5143/Reviewer_uTjh"],"forum":"iDiiETH7Qv","number":1,"license":"CC BY 4.0","cdate":1761470976165,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5143/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917907118,"domain":"ICLR.cc/2026/Conference","replyto":"iDiiETH7Qv","id":"OicNXhNKXa","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["directed sheaf neural network","directed graphs","directed cellular sheaves"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Sheaf Neural Networks (SNNs) are a powerful algebraic-topology generalization of Graph Neural Networks (GNNs), and have been shown to significantly improve our ability to model complex relational data. While the GNN literature proved that incorporating directionality can substantially boost performance in many real-world applications, no SNNs approaches are known with such a capability. To address this limitation, we introduce the Directed Cellular Sheaf, a generalized cellular sheaf designed to explicitly account for edge orientations. Building on it, we define a corresponding sheaf Laplacian, the Directed Sheaf Laplacian $L^{\\widetilde{\\mathcal{F}}}$, which exploits the sheaf's structure to capture both the graph’s topology and its directions. $L^{\\widetilde{\\mathcal{F}}}$ serves as the backbone of the Directed Sheaf Neural Network (DSNN), the first SNN model to embed a directional bias into its architecture. Extensive experiments on twelve real-world benchmarks show that DSNN consistently outperforms many baseline methods. The source\ncode can be found at https://github.com/hakanaktas0/DSNN."},"_bibtex":{"value":"@inproceedings{\nfiorini2026sheaves,\ntitle={Sheaves Reloaded: A Direction Awakening},\nauthor={Stefano Fiorini and Hakan Aktas and Iulia Duta and Pietro Morerio and Alessio Del Bue and Pietro Lio and Stefano Coniglio},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=iDiiETH7Qv}\n}"},"title":{"value":"Sheaves Reloaded: A Direction Awakening"},"pdf":{"value":"/pdf/57e845fda6d56b487ec365a89e5b7211f789a44c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"fiorini|sheaves_reloaded_a_direction_awakening"},"authorids":{"value":["~Stefano_Fiorini1","~Hakan_Aktas2","~Iulia_Duta1","~Pietro_Morerio1","~Alessio_Del_Bue2","~Pietro_Lio1","~Stefano_Coniglio2"]},"authors":{"value":["Stefano Fiorini","Hakan Aktas","Iulia Duta","Pietro Morerio","Alessio Del Bue","Pietro Lio","Stefano Coniglio"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2502.12054v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"zhang|physreason_a_comprehensive_benchmark_towards_physicsbased_reasoning"},"authorids":{"value":["~Xinyu_Zhang26","~Yuxuan_Dong4","https://dblp.org/search/pid/api?q=author:Yanrui_Wu:","https://dblp.org/search/pid/api?q=author:Jiaxing_Huang:","https://dblp.org/search/pid/api?q=author:Chengyou_Jia:","https://dblp.org/search/pid/api?q=author:Basura_Fernando:","https://dblp.org/search/pid/api?q=author:Mike_Zheng_Shou:","https://dblp.org/search/pid/api?q=author:Lingling_Zhang:","https://dblp.org/search/pid/api?q=author:Jun_Liu_0036:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2502.12054"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-12054,\n  publtype={informal},\n  author={Xinyu Zhang and Yuxuan Dong and Yanrui Wu and Jiaxing Huang and Chengyou Jia and Basura Fernando and Mike Zheng Shou and Lingling Zhang and Jun Liu},\n  title={PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.12054},\n  url={https://doi.org/10.48550/arXiv.2502.12054}\n}\n"},"abstract":{"value":"Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoning-based (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard requiring 15.6, reflecting the complexity of physics-based reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive step-level evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identified four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position PhysReason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models. Our code and data will be published at https:/dxzxy12138.github.io/PhysReason."},"title":{"value":"PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning"},"authors":{"value":["Xinyu Zhang","Yuxuan Dong","Yanrui Wu","Jiaxing Huang","Chengyou Jia","Basura Fernando","Mike Zheng Shou","Lingling Zhang","Jun Liu"]}},"tmdate":1771504064000,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2502-12054"],"tcdate":1762419109230,"writers":["~"],"signatures":["~Yuxuan_Dong4"],"forum":"9ozQAe2NgA","license":"CC BY-SA 4.0","number":665160,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1771504064000,"domain":"DBLP.org","id":"9ozQAe2NgA","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2411.00554v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"yang|differentiable_physicsbased_system_identification_for_robotic_manipulation_of_elastoplastic_materials"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xintong_Yang:","~Ze_Ji1","https://dblp.org/search/pid/api?q=author:Yu-Kun_Lai:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2411.00554"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2411-00554,\n  publtype={informal},\n  author={Xintong Yang and Ze Ji and Yu-Kun Lai},\n  title={Differentiable Physics-based System Identification for Robotic Manipulation of Elastoplastic Materials},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2411.00554},\n  url={https://doi.org/10.48550/arXiv.2411.00554}\n}\n"},"abstract":{"value":"Robotic manipulation of volumetric elastoplastic deformable materials, from foods such as dough to construction materials like clay, is in its infancy, largely due to the difficulty of modelling and perception in a high-dimensional space. Simulating the dynamics of such materials is computationally expensive. It tends to suffer from inaccurately estimated physics parameters of the materials and the environment, impeding high-precision manipulation. Estimating such parameters from raw point clouds captured by optical cameras suffers further from heavy occlusions. To address this challenge, this work introduces a novel Differentiable Physics-based System Identification (DPSI) framework that enables a robot arm to infer the physics parameters of elastoplastic materials and the environment using simple manipulation motions and incomplete 3D point clouds, aligning the simulation with the real world. Extensive experiments show that with only a single real-world interaction, the estimated parameters, Young's modulus, Poisson's ratio, yield stress and friction coefficients, can accurately simulate visually and physically realistic deformation behaviours induced by unseen and long-horizon manipulation motions. Additionally, the DPSI framework inherently provides physically intuitive interpretations for the parameters in contrast to black-box approaches such as deep neural networks. The project is fully open-sourced via https://ianyangchina.github.io/SI4RP-data/."},"title":{"value":"Differentiable Physics-based System Identification for Robotic Manipulation of Elastoplastic Materials"},"authors":{"value":["Xintong Yang","Ze Ji","Yu-Kun Lai"]}},"tmdate":1762710456324,"pdate":1704067200000,"externalIds":["dblp:journals/corr/abs-2411-00554"],"tcdate":1762710438165,"writers":["~"],"signatures":["~Ze_Ji1"],"forum":"zScfQ61L3V","license":"CC BY-SA 4.0","number":683286,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762710456324,"domain":"DBLP.org","id":"zScfQ61L3V","version":2},{"content":{"comment":{"value":"### Reviewer Comment 2:\n>Evaluation largely internal. The results rely heavily on the authors’ synthetic dataset and GPT-4o-based evaluation. No human evaluation or cross-dataset validation is performed. It’s unclear how well the model generalizes to unseen real-world 3D scenes.\n\n**Response**\n\n\nThanks for your concern about the internal evaluation. We have addressed the concerns regarding human verification, real-world generalization, and scene complexity with **three major new experiments** conducted during the rebuttal:\n\n**0. Alignment with Community Standards:** \n\nFirstly, we note that using synthetic datasets (e.g., 3D-FRONT) is the standard protocol for 3D layout generation research (e.g., _LayoutVLM[1]_, I-Design[2], Holddeck [3]). Our addition of rigorous Physical Metrics (Collision/Constraint Ratios) actually sets a **higher standard for validity** than prior works that rely solely on semantic scores.\n\n**1. Human & Cross-Model Verification (External Validation):**\n\nTo address the lack of human oversight, we conducted a blind user study with 10 experts.\n\n-   **Results:** As detailed in our **Response to Reviewer jKB9 (Comments 1 & 2)** and **Reviewer 563R**, humans consistently preferred our method (**Score 0.65**) over baselines like LayoutGPT (0.54).\n    \n-   **Consistency:** This preference is consistent across **Gemini 3** and **Qwen3-VL-30B** evaluations. Crucially, human ratings show a high Pearson correlation ($r > 0.85$) with our automated GPT-4o metrics, confirming that our internal metric is a reliable proxy for human preference.\n    \n\n**2. Generalization to Unseen Real-World Scenes:**\n\nTo prove our model works beyond our synthetic training distribution, we evaluated it on two external real-world benchmarks:\n\n-   **Open3DVQA (Realistic Urban Scenes):** As shown in **Appendix G**, our model achieves **73.5% accuracy** on complex spatial reasoning in realistic urban environments, significantly outperforming generalist models like **GPT-4o (58.7%)**.\n    \n-   **SpatialScore (Real-World Images):** We further stress-tested the model on real-world images using the SpatialScore benchmark (derived from VGBench). As detailed in **Response 5 to Reviewer o4o8**, our **MetaSpatial-7B** model (Score: **41.85**) remarkably outperforms the much larger **Qwen2.5-VL-72B** (Score: **40.20**) and achieves parity with GPT-4o, demonstrating efficient scaling for real-world spatial understanding.\n    \n\n**3. Robustness in Complex Scenarios (20+ Assets):**\n\nTo address concerns about the simplicity of synthetic scenes, we conducted a stress test on dense scenes with an average of 20+ assets (double the standard setting).\n\n-   **Results:** As reported in **Response 1 to Reviewer jKB9**, our method maintains high physical feasibility (14.2% Collision Rate) and semantic quality (Human Score 0.70) even in these highly complex environments, significantly outperforming baselines like LayoutGPT.\n    \n\n**Conclusion:** These extensive external validations—spanning human experts, real-world benchmarks (Open3DVQA/SpatialScore), and complex dense scenes—confirm that our RL-based training instills robust, generalized spatial reasoning that extends well beyond the internal synthetic dataset.\n\n\n[1] LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models\n\n[2] I-Design: Personalized LLM Interior Designer\n\n[3] Language Guided Generation of 3D Embodied AI Environments\n\n----------"},"title":{"value":"Response for Question 2"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764041371995,"tcdate":1764041371995,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14336/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission14336/Authors"],"forum":"EdQzLC0Zra","number":15,"license":"CC BY 4.0","cdate":1764041371995,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14336/-/Official_Comment"],"mdate":1764041371995,"domain":"ICLR.cc/2026/Conference","replyto":"jleXTUAzug","id":"jQSY5GpZPw","forumContent":{"TLDR":{"value":"MetaSpatial leverages reinforcement learning to enhance 3D spatial reasoning in vision-language models (VLMs), enabling more structured, realistic, and adaptive scene generation."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["spatial reasoning","vision language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present MetaSpatial, the first reinforcement learning (RL) framework for enhancing 3D spatial reasoning in vision-language models (VLMs), enabling real-time 3D scene layout generation without post-processing. MetaSpatial addresses two key challenges: (i) the need for extensive post-processing, as existing VLMs lack inherent 3D spatial reasoning to generate realistic layouts; and (ii) the inefficiency of supervised fine-tuning (SFT) for layout generation due to scarcity of perfect annotations. Our core contribution is the 3D Spatial Policy Optimization (3D-SPO) algorithm, which incorporates physics-aware modulation into advantage estimates at the object level and trajectory-level reward from a training-only multi-turn refinement pipeline. This design enhances temporal credit assignment and encourages spatially consistent policy learning. Empirical evaluations across models of varying scales demonstrate that MetaSpatial improves spatial coherence, physical plausibility, and formatting stability, leading to more realistic and functionally coherent object placements applicable to metaverse environments."},"_bibtex":{"value":"@inproceedings{\npan2026metaspatial,\ntitle={MetaSpatial: Reinforcing 3D Spatial Reasoning in {VLM}s for the Metaverse},\nauthor={Zhenyu Pan and Han Liu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=EdQzLC0Zra}\n}"},"title":{"value":"MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse"},"pdf":{"value":"/pdf/e9be0940303eba3e3eacf72042d92ce2235b2e49.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"pan|metaspatial_reinforcing_3d_spatial_reasoning_in_vlms_for_the_metaverse"},"authorids":{"value":["~Zhenyu_Pan1","~Han_Liu4"]},"authors":{"value":["Zhenyu Pan","Han Liu"]}},"version":2},{"content":{"summary":{"value":"This paper aims to solve dynamic reconstruction task. It proposes a new loss called ReMatching loss, in order to inject some physical priors into the learned deformation field. They also propose five categories of physics priors, and derive the corresponding forms of the RM loss respectively. The paper builds the model based on the deformable-3d-gaussian paper, and tests the performance for three of the five categories on both synthetic and real-world datasets, showing comparable performance to baselines."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Apart from the questions listed in weaknesses, here are some more questions:\n\n1. Why the baseline numbers in the table not the same as the published version? According to the published version of D3G and GA3D papers, the performance of this paper is actually on the same level. For example, on D-NeRF datasets, D3G has 40.43 PSNR on average, while two types of the model have 40.36 and 40.54 on average. (lego scene is deleted due to data discrepancy, details in next question)\n\n2. The lego scene is known that data discrepancy exists in the scenarios presented in the training and test sets. This can be substantiated by examining the flip angles of the Lego shovels. So some works, like D3G, try to evaluate the model on the validation split because evaluating on the test split is meaningless. Which split is evaluated here? \n\n3. GA3D has reported its performance on HyperNeRF dataset, so why this baseline is deleted from the HyperNeRF experiments?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"no"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"In presentation, this paper shows dedicatedly derived loss form and clear derivation for different kinds of velocity priors. For the idea, the loss proposed by this paper constrains the function shape of deformation field with some physics priors. Working similar as a PINN loss, this framework seems flexible to different deformation designs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In synthetic experiments, type IV and type V priors are selected for each scene. They should be thoroughly tested respectively to show the applicability under different scenarios. \n\n2. In real-world experiments, it is not clear which type of priors are used. And also, all of the priors should be tested to show the abilities. \n\n3. The velocity branches in the model only serve as a constraint part, but not used directly. The paper should demonstrate whether the learned velocity is meaningful or not. For example, using the velocity to advect the scene. \n\n4. The segmentation compared with SAM is not proper and confusing. Since SAM segments objects according to semantics but this paper segments according to motions, it’s not proper to punish SAM’s over-segmentation. Another question is about the bouncing ball example shown in the paper. As I know this scene includes some moving shadow on the ground, so the motion should not segment the shadow part as the same as the static ground. How is the segmentation applied?\n\n5. The performances are only comparable, it’s hard to tell whether the priors are really helpful or not. \n\n6. There is a lack of an illustrative figure to show the main architecture of the proposed method, so the reader can only know how the idea is working till the end of the paper.\n\n7. There are too many equations or derivations in the main text. Although they are clear and I appreciate the details, too many derivations in the main text make the paper lose its main focus. \n\n8. There are inconsistent baseline names in text and tables, for example, 3D Geometry-aware Deformable Gaussians is called DAG in main text but GA3D in the table."}},"nonreaders":[],"tmdate":1733137864510,"tcdate":1730386680600,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission655/Reviewer_GNax"],"signatures":["ICLR.cc/2025/Conference/Submission655/Reviewer_GNax"],"forum":"bwhI6bCGY1","number":2,"license":"CC BY 4.0","cdate":1730386680600,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission655/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733137864510,"domain":"ICLR.cc/2025/Conference","replyto":"bwhI6bCGY1","id":"B1rwCZCPry","forumContent":{"TLDR":{"value":"We introduce the ReMatching framework—a novel method for designing and incorporating deformation priors into dynamic reconstruction models, ensuring fidelity to input data while adhering to the specified priors."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Dynamic Reconstruction","Flow Modeling","Gaussian Splatting","Novel view synthesis"]},"supplementary_material":{"value":"/attachment/8b2f125fb7c14a5aeef230dd9538b494ebb66a76.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Reconstructing a dynamic scene from image inputs is a fundamental computer\nvision task with many downstream applications. Despite recent advancements, existing approaches still struggle to achieve high-quality reconstructions from unseen viewpoints and timestamps. This work introduces the ReMatching framework, designed to improve reconstruction quality by incorporating deformation priors into dynamic reconstruction models. Our approach advocates for velocity-field based priors, for which we suggest a matching procedure that can seamlessly supplement existing dynamic reconstruction pipelines. The framework is highly adaptable and can be applied to various dynamic representations. Moreover, it supports integrating multiple types of model priors and enables combining simpler ones to create more complex classes. Our evaluations on popular benchmarks involving both synthetic and real-world dynamic scenes demonstrate that augmenting current state-of-the-art methods with our approach leads to a clear improvement in reconstruction accuracy."},"_bibtex":{"value":"@inproceedings{\noblak2025rematching,\ntitle={ReMatching Dynamic Reconstruction Flow},\nauthor={Sara Oblak and Despoina Paschalidou and Sanja Fidler and Matan Atzmon},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=bwhI6bCGY1}\n}"},"title":{"value":"ReMatching Dynamic Reconstruction Flow"},"pdf":{"value":"/pdf/4f7782b721c23efbd060cb8c4012496b7fae11dd.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"oblak|rematching_dynamic_reconstruction_flow"},"authorids":{"value":["~Sara_Oblak1","~Despoina_Paschalidou1","~Sanja_Fidler1","~Matan_Atzmon1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sara Oblak","Despoina Paschalidou","Sanja Fidler","Matan Atzmon"]}},"version":2},{"content":{"summary":{"value":"This paper introduces the Causal Process Framework (CPF) and its neural implementation, the Causal Process Model (CPM), which reinterprets the attention mechanism of Transformer networks as a reinforcement learning problem for causal discovery. Instead of soft attention weights, the model employs two RL agents to decide which causal edges to instantiate between objects and forces over time. This allows the system to construct sparse, time-varying causal graphs that reflect active physical interactions rather than dense potential dependencies. Experiments in a synthetic physics environment show that CPM outperforms Graph Neural Networks (GNNs), Transformers, Recurrent Independent Mechanisms (RIMs), and modular networks in prediction accuracy, long-horizon generalization, and downstream RL performance."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Can the authors clarify how the learned rewards $R_O$ and $R_{O<->F}$ correspond to meaningful causal evaluation criteria?  \n  - Are the interaction-scope and effect-attribution agents trained jointly or alternately, and how stable is this process?  \n  - How sensitive are results to the inductive biases (e.g., pairwise force-object constraints)?  \n  - Could the authors show qualitative examples of inferred causal graphs during different physical interactions to support interpretability claims?  \n  - How does CPM scale with larger numbers of objects or higher-dimensional state representations?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":3},"strengths":{"value":"- **Novel conceptual link between attention and RL:** The paper offers a fresh theoretical view by reframing attention as a decision-making problem. This perspective connects causal discovery, reinforcement learning, and neural attention in an elegant way.  \n  - **Dynamic causal modeling:** The proposed Causal Process Framework explicitly models causal graphs that evolve over time, addressing a key limitation of static Structural Causal Models when applied to dynamic physical systems.  \n  - **Interpretability and sparsity:** The all-or-nothing edge construction naturally yields interpretable causal graphs, where connections correspond to actual interactions (e.g., collisions) rather than dense message passing.  \n  - **Strong empirical results:** The CPM demonstrates clear improvements over baselines in multi-object physical environments and provides consistent advantages in both observed and unobserved generalization settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Clarity and notation:** Section 3.1 is particularly dense and difficult to follow. Some key symbols (e.g., $J^t$) are used before being defined, and sets $N$ and $M$ are not introduced at all. The abundance of indices and nested distributions makes it hard to parse the formalism without additional diagrams or examples.  \n  - **Imprecise language:** The paper frequently uses vague phrasing that leaves important concepts underdefined. For example, the phrase “defining the causal chain of object-to-force and force-to-object connections” lacks a precise mathematical meaning and forces the reader to infer the intended interpretation.  \n  - **Limited evaluation diversity:** While the synthetic physics environment provides proof of concept, it remains relatively simple. There are no experiments on more complex, real-world settings or comparisons to recent structured causal transformers beyond Melnychuk et al. (2022).  \n  - **Reward learning unclear:** While Figure 5 reports mean reward values across object counts, it remains unclear how these rewards are linked to the CPM’s core training objective. The paper does not specify whether maximizing these rewards directly improves causal graph accuracy, predictive performance, or both. Moreover, the overall optimization landscape is ambiguous, what exactly constitutes the optimum? Is it a state where the RL agents select edges yielding minimal prediction loss, or where the learned reward MLPs stabilize under inverse RL? Without this connection between the agent-level rewards and the CPM loss, it’s difficult to interpret the learning dynamics or judge convergence.\n  - **Missing ablation or analysis of learned structure:** The discussion section mentions plans to analyze semantic sub-vectors (mutable, causal, controllable), but such analysis would have strengthened the current submission by demonstrating interpretability concretely."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358863035,"tcdate":1761840830740,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission25201/Reviewer_c1aN"],"signatures":["ICLR.cc/2026/Conference/Submission25201/Reviewer_c1aN"],"forum":"fY4proGNFD","number":1,"license":"CC BY 4.0","cdate":1761840830740,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission25201/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358863035,"domain":"ICLR.cc/2026/Conference","replyto":"fY4proGNFD","id":"Op41TgSCnl","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Causal World Models","Causal Reinforcement Learning","Causal Processes","Causal Representation Learning"]},"primary_area":{"value":"causal reasoning"},"abstract":{"value":"Most neural models of causality assume static causal graphs, failing to capture the dynamic and sparse nature of physical interactions where causal relationships emerge and dissolve over time. We introduce the Causal Process Framework and its neural implementation, Causal Process Models (CPMs), for learning sparse, time-varying causal graphs from visual observations. Unlike traditional approaches that maintain dense connectivity, our model explicitly constructs causal edges only when objects actively interact, dramatically improving both interpretability and computational efficiency. We achieve this by formulating causal discovery as a multi-agent reinforcement learning problem, where specialized agents sequentially decide which objects are causally connected at each timestep. Our key innovation is a structured representation that factorizes object and force vectors along three learned dimensions (mutability, causal relevance, and control relevance), enabling the automatic discovery of semantically meaningful encodings. We demonstrate that a CPM significantly outperforms dense graph baselines on physical prediction tasks, particularly for longer horizons and varying object counts."},"_bibtex":{"value":"@misc{\norujlu2025reframing,\ntitle={Reframing attention as a reinforcement learning problem for causal discovery},\nauthor={Turan Orujlu and Christian Gumbsch and Martin V. Butz and Charley M Wu},\nyear={2025},\nurl={https://openreview.net/forum?id=fY4proGNFD}\n}"},"title":{"value":"Reframing attention as a reinforcement learning problem for causal discovery"},"pdf":{"value":"/pdf/7bbe48e371a74719167670c33ecaea2cf55a08df.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"orujlu|reframing_attention_as_a_reinforcement_learning_problem_for_causal_discovery"},"authorids":{"value":["~Turan_Orujlu1","~Christian_Gumbsch1","~Martin_V._Butz2","~Charley_M_Wu1"]},"authors":{"value":["Turan Orujlu","Christian Gumbsch","Martin V. Butz","Charley M Wu"]}},"version":2},{"content":{"TLDR":{"value":"This work presents a contrastive learning-based method to improve privacy risk assessment in synthetic tabular data, addressing GDPR-defined \"singling out\" risks with efficient and effective metrics."},"venue":{"value":"RLGMSD 2024 Poster"},"pdf":{"value":"/pdf/4992c719600ec2633defa2e4d55a7ff6ff3199d1.pdf"},"keywords":{"value":["Contrastive Learning","Tabular Data","Privacy Metrics"]},"venueid":{"value":"ELLIS.eu/2024/Workshop/RLGMSD"},"paperhash":{"value":"palacios|contrastive_learningbased_privacy_metrics_in_tabular_synthetic_datasets"},"authorids":{"value":["~Milton_Nicolás_Plasencia_Palacios1"]},"abstract":{"value":"Synthetic data has garnered attention as a Privacy Enhancing Technology in sectors such as healthcare and finance. When using synthetic data in practical applications, it is important to provide protection guarantees. We introduce a contrastive method that improves privacy assessment of synthetic datasets by embedding the data in a more representative space. This overcomes obstacles surrounding the multitude of data types and attributes. It also makes the use of intuitive distance metrics possible for similarity measurements and as an attack vector. Our results show that relatively efficient, easy to implement privacy metrics can perform equally well as more advanced metrics explicitly modeling conditions for privacy referred to by the GDPR."},"_bibtex":{"value":"@inproceedings{\npalacios2025contrastive,\ntitle={Contrastive Learning-Based privacy metrics in Tabular Synthetic Datasets},\nauthor={Milton Nicol{\\'a}s Plasencia Palacios},\nbooktitle={ELLIS workshop on Representation Learning and Generative Models for Structured Data},\nyear={2025},\nurl={https://openreview.net/forum?id=PyjYYHDSx0}\n}"},"title":{"value":"Contrastive Learning-Based privacy metrics in Tabular Synthetic Datasets"},"authors":{"value":["Milton Nicolás Plasencia Palacios"]}},"tmdate":1740136236412,"pdate":1740136236328,"tcdate":1737318920267,"writers":["ELLIS.eu/2024/Workshop/RLGMSD","ELLIS.eu/2024/Workshop/RLGMSD/Submission4/Authors"],"signatures":["ELLIS.eu/2024/Workshop/RLGMSD/Submission4/Authors"],"forum":"PyjYYHDSx0","license":"CC BY 4.0","number":4,"cdate":1737318920267,"readers":["everyone"],"invitations":["ELLIS.eu/2024/Workshop/RLGMSD/-/Submission","ELLIS.eu/2024/Workshop/RLGMSD/-/Post_Submission","ELLIS.eu/2024/Workshop/RLGMSD/-/Edit"],"mdate":1740136236412,"odate":1740136188447,"domain":"ELLIS.eu/2024/Workshop/RLGMSD","id":"PyjYYHDSx0","version":2},{"content":{"review":{"value":"*Paper Summary & Main Contributions*. The paper addresses the scarcity of paired ULF–HF MRI data for supervised enhancement by (i) simulating realistic ULF images from HF pediatric scans via an explicit MR-physics degradation pipeline (bias field, T2*-related decay, B0 inhomogeneity-driven dephasing, k-space bandwidth limitation, undersampling, noise — Eqs. 1–3), and (ii) proposing a composite spatial-frequency loss (voxel L1 + band-weighted log-spectrum consistency + gradient regularization, Eqs. 4–7) applied on top of three backbone translation architectures (nnU-Net Translation, Pix2Pix, BBDM diffusion). Models are trained exclusively on the synthetic paired dataset D1 (833 pediatric 1.5T volumes) and evaluated both on synthetic held-out data (Table 1) and, for generalization, on real 64 mT Hyperfine ULF scans (D2) via downstream multiclass segmentation (Table 2) and a blinded radiologist reader study.\n\n*Key findings*: on synthetic data, nnU-NetT with the proposed loss achieves the best overall metrics (SSIM 0.9466, PSNR 28.76 dB), marginally above the same backbone without the loss (SSIM 0.9458, PSNR 28.69 dB). On real D2 data, nnU-NetT (trained on D1) yields the strongest downstream segmentation performance (DSC 0.7466 vs. 0.7366 baseline) and, together with GAMBAS, is preferred by radiologists over SynthSR/LoHiResGAN in the reader study.\n\n**Strengths (Pros)**\n- The physics-based degradation model (Eqs. 1–3) is well justified with cited literature-grounded parameter ranges (coil sensitivity fall-off, T2*/B0 sampling ranges, k-space bandwidth/undersampling factors) rather than being an ad hoc noise injection — this is a real methodological strength.\n\n- Evaluating the proposed loss across three structurally different backbones (encoder–decoder, GAN, diffusion) is good practice for demonstrating that a contribution is architecture-agnostic rather than backbone-specific.\n\n- Going beyond pixel-level metrics to downstream segmentation (hippocampus/basal ganglia) and a blinded radiologist reader study (with inter-rater agreement statistics, Friedman test, Bonferroni-corrected post-hoc comparisons) is exactly the kind of clinically grounded validation the workshop's \"Trustworthy AI\" and \"Clinical Relevance\" pillars call for.\n\n- The synthetic training corpus (833 volumes) is reasonably large, with sensible train/val/test partitioning.\n\n**Weaknesses & Critical Concerns (Cons)**\n\n- **The paper's own proposed loss contribution is not validated on real data**. Table 2 (downstream segmentation) and the reader study both use \"nnU-NetT\" — i.e., the backbone trained on the synthetic D1 pairs, but without the spatial-frequency composite loss — rather than \"nnU-NetT_ULF-Synth,\" which is the model the paper claims achieves best overall performance in Table 1. The abstract's claim that \"the resulting models generalize effectively to real 64 mT ULF acquisitions, improving downstream... segmentation and achieving higher radiologist preference\" is therefore substantiated only for the physics-based synthesis contribution, not for the spatial-frequency loss contribution, which is the paper's more novel component. This is a significant gap between what is claimed and what is shown.\n\n- **Marginal, statistically untested gains on synthetic data**. The improvement of the proposed loss over baseline in Table 1 is small (e.g., nnU-Net: +0.08 SSIM points, +0.07 dB PSNR) and in one case (BBDM) LPIPS actually worsens (0.058 vs. 0.042). No significance testing (paired t-test, Wilcoxon, or bootstrap CIs) is reported to establish these differences are reliable given the reported standard deviations.\n\n- **Small, restricted reader study**. Only 5 D2 subjects and 3 radiologists were used, and the study excludes the paper's own strongest baselines (BBDM, Pix2Pix) as well as any ULF-Synth loss variant, comparing only SynthSR, LoHiResGAN, GAMBAS, and nnU-NetT.\n\n- **Segmentation gains in Table 2 are small relative to variance**, and one metric (RVE) shows a non-monotonic, high-variance pattern (baseline 0.85±19.05; LF-SynthSR 3.25±14.83 — actually worse than baseline) that is not discussed or explained.\nSingle real dataset/field strength. Real-world validation uses only one 64 mT Hyperfine cohort; no multi-site or multi-scanner real ULF validation is presented, which is a meaningful limitation given the workshop's stated emphasis on multi-site validation and domain shift across diverse settings.\n\n- **Degradation-model calibration is only qualitatively validated** (\"empirically tuned through iterative qualitative assessment and radial power spectrum analysis\") with no quantitative domain-gap metric reported (e.g., distributional distance between synthetic and real ULF frequency/noise statistics), which weakens confidence that D1 truly matches real acquisition characteristics."},"confidence":{"value":4},"rating":{"value":5},"title":{"value":"A physics-grounded and architecture-agnostic synthetic-data framework for ULF MRI enhancement with a genuinely compelling degradation model, but whose central claim of real-world generalization is only demonstrated for baseline models trained on the synthetic data, not for the paper's own proposed spatial-frequency loss."}},"parentInvitations":"MICCAI.org/2026/Workshop/AFRICAI/-/Official_Review","nonreaders":[],"tmdate":1787850791253,"tcdate":1785002284945,"writers":["MICCAI.org/2026/Workshop/AFRICAI","MICCAI.org/2026/Workshop/AFRICAI/Submission19/Reviewer_k2Xu"],"signatures":["MICCAI.org/2026/Workshop/AFRICAI/Submission19/Reviewer_k2Xu"],"forum":"ebyw01oZpc","number":3,"license":"CC BY 4.0","cdate":1785002284945,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/AFRICAI/Submission19/-/Official_Review","MICCAI.org/2026/Workshop/AFRICAI/-/Edit"],"mdate":1787850791253,"domain":"MICCAI.org/2026/Workshop/AFRICAI","replyto":"ebyw01oZpc","id":"THeBaD3Pjq","forumContent":{"TLDR":{"value":"We show that MRI physics-guided simulation and degradation of high-field MRI, to create synthetic low-field pairs for supervised image enhancement is possible, and can effectively transfer to realistic ultra-low-field MRI for generative enhancement."},"venue":{"value":"AFRICAI 2026 Oral"},"keywords":{"value":["Ultra-low-field MRI","MRI Enhancement","Physics-Guided MRI Synthesis","Synthetic Data","k-Space"]},"venueid":{"value":"MICCAI.org/2026/Workshop/AFRICAI"},"paperhash":{"value":"musah|ulfsynth_physicsguided_ultralowfield_mri_enhancement_for_pediatric_neuroimaging"},"authorids":{"value":["~Toufiq_Musah1","~Salvatore_Calcagno1","~Federica_Proietto_Salanitri1","~Xiaomeng_Li1","~Maruf_Adewole2","~Marawan_Elbatel2"]},"_bibtex":{"value":"@inproceedings{\nmusah2026ulfsynth,\ntitle={{ULF}-Synth: Physics-Guided Ultra-Low-Field {MRI} Enhancement for Pediatric Neuroimaging},\nauthor={Toufiq Musah and Salvatore Calcagno and Federica Proietto Salanitri and Xiaomeng Li and Maruf Adewole and Marawan Elbatel},\nbooktitle={First Workshop on Advancing African Medical AI through Global Integration},\nyear={2026},\nurl={https://openreview.net/forum?id=ebyw01oZpc}\n}"},"title":{"value":"ULF-Synth: Physics-Guided Ultra-Low-Field MRI Enhancement for Pediatric Neuroimaging"},"authors":{"value":["Toufiq Musah","Salvatore Calcagno","Federica Proietto Salanitri","Xiaomeng Li","Maruf Adewole","Marawan Elbatel"]}},"version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["agentic AI","language models","evaluation"]},"_bibtex":{"value":"@inproceedings{\nxia2026buildarena,\ntitle={BuildArena: A Physics-Aligned Interactive Benchmark of {LLM}s for Engineering Construction},\nauthor={Tian Xia and Tianrun Gao and Wenhao Deng and Long Wei and Xiaowei Qian and Chenglei Yu and Tailin Wu},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=QAQKmIp3SZ}\n}"},"title":{"value":"BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction"},"paperhash":{"value":"xia|buildarena_a_physicsaligned_interactive_benchmark_of_llms_for_engineering_construction"},"originally_submitted_PDF":{"value":"/pdf/6099fc720f1450419c3bbab10f5aac5ca51b1511.pdf"},"primary_area":{"value":"applications->everything_else"},"abstract":{"value":"Engineering construction automation aims to transform natural language specifications into physically viable structures, requiring complex integrated reasoning under strict physical constraints. While modern LLMs possess broad knowledge and strong reasoning capabilities that make them promising candidates for this domain, their construction competencies remain largely unevaluated. To address this gap, we introduce BuildArena, the first physics-aligned interactive benchmark designed for language-driven engineering construction. Technically, it contributes to the community in two aspects: (1) an extendable task design strategy spanning static and dynamic mechanics across multiple difficulty tiers; (2) a 3D Spatial Geometric Computation Library for supporting construction based on language instructions.\nOn nine frontier LLMs and three additional open-weight models, BuildArena comprehensively evaluates their capabilities for language-driven and physics-grounded construction automation. We release the code at https://github.com/AI4Science-WestlakeU/BuildArena to benefit construction automation in engineering applications."},"link_to_code":{"value":"https://github.com/AI4Science-WestlakeU/BuildArena"},"pdf":{"value":"/pdf/239295bd579df4db711a991eecaa582d1a1bb7a6.pdf"},"lay_summary":{"value":"Engineering tasks like designing bridges, vehicles, and rockets require complex spatial reasoning and an understanding of physics. While AI language models have shown impressive abilities in writing, coding, and reasoning, it remains unclear whether they can actually design structures that work under real physical conditions.\nWe created BuildArena, a benchmark where LLMs receive plain-language instructions and construct 3D structures inside a physics simulator. The benchmark spans tasks of increasing difficulty across both static structures (like bridges) and moving objects (like vehicles and rockets), and automatically tests whether each design functions correctly under realistic physics.\nOur evaluation of twelve AI models reveals that while some can produce surprisingly sophisticated structures including designs resembling real-world engineering solutions like truss bridges, significant capability gaps remain, especially for more complex tasks. BuildArena gives the research community a standardized way to measure and improve LLMs' ability to turn language into physically viable constructions, an important step toward automating parts of the engineering design process."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Tian_Xia15","~Tianrun_Gao3","~Wenhao_Deng2","~Long_Wei1","~Xiaowei_Qian3","~Chenglei_Yu1","~Tailin_Wu1"]},"authors":{"value":["Tian Xia","Tianrun Gao","Wenhao Deng","Long Wei","Xiaowei Qian","Chenglei Yu","Tailin Wu"]}},"tmdate":1790067310068,"pdate":1777576261704,"tcdate":1768655747155,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission6579/Authors"],"signatures":["ICML.cc/2026/Conference/Submission6579/Authors"],"forum":"QAQKmIp3SZ","license":"CC BY 4.0","number":6579,"cdate":1768655747155,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission6579/-/Reciprocal_Reviewing_Correction","ICML.cc/2026/Conference/Submission6579/-/Full_Submission","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/Submission6579/-/Camera_Ready_Revision"],"mdate":1790067310068,"odate":1782341922356,"domain":"ICML.cc/2026/Conference","id":"QAQKmIp3SZ","version":2},{"content":{"venue":{"value":"DASFAA (2) 2018"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-91458-9_16.pdf"},"venueid":{"value":"dblp.org/conf/DASFAA/2018"},"paperhash":{"value":"xin|efficient_complex_social_eventparticipant_planning_based_on_heuristic_dynamic_programming"},"authorids":{"value":["~Junchang_Xin1","~Mo_Li5","https://dblp.org/search/pid/api?q=author:Wangzihao_Xu:","https://dblp.org/search/pid/api?q=author:Yizhu_Cai:","https://dblp.org/search/pid/api?q=author:Minhua_Lu:","https://dblp.org/search/pid/api?q=author:Zhiqiong_Wang:"]},"html":{"value":"https://doi.org/10.1007/978-3-319-91458-9_16"},"_bibtex":{"value":"@inproceedings{DBLP:conf/dasfaa/XinLXCLW18,\n  author={Junchang Xin and Mo Li and Wangzihao Xu and Yizhu Cai and Minhua Lu and Zhiqiong Wang},\n  title={Efficient Complex Social Event-Participant Planning Based on Heuristic Dynamic Programming},\n  year={2018},\n  cdate={1514764800000},\n  pages={264-279},\n  url={https://doi.org/10.1007/978-3-319-91458-9_16},\n  booktitle={DASFAA (2)},\n  crossref={conf/dasfaa/2018-2}\n}\n"},"abstract":{"value":"To manage the Event Based Social Networks (EBSNs), an important task is to solve the Global Event Planning with Constraints (GEPC) problem, which arranges suitable social events to target users. Existing studies are not efficient enough because of the two-step framework. In this paper, we propose a more efficient method, called Heuristic-DP, which asynchronously considers all the constraints together. Using this method, we improve the computational complexity from \\(O(|E|^2 + |U||E|^2)\\) to O(|U||E|), where |U| is the number of users and |E| is the number of events in an EBSN platform. We also propose an improved heuristic strategy in one function of the heuristic-DP algorithm, which slightly increases the time cost, but can obtain a more accurate result. Finally, we verify the effectiveness and efficiency of our proposed algorithms through extensive experiments over real and synthetic datasets."},"title":{"value":"Efficient Complex Social Event-Participant Planning Based on Heuristic Dynamic Programming"},"authors":{"value":["Junchang Xin","Mo Li","Wangzihao Xu","Yizhu Cai","Minhua Lu","Zhiqiong Wang"]}},"tmdate":1744074689394,"pdate":1514764800000,"tcdate":1739176426190,"writers":["~"],"signatures":["~Mo_Li5"],"forum":"J97YrBL0pW","license":"CC BY-SA 4.0","number":314879,"cdate":1514764800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1744074689394,"domain":"DBLP.org","id":"J97YrBL0pW","version":2},{"content":{"summary":{"value":"The paper addresses a novel task, SVG extraction that combines vision and structured graphics, which is timely given the rise of multimodal LLMs. The authors introduce WildSVG, the first benchmark for extracting vector graphics from real images. They construct two complementary datasets: Natural WildSVG and Synthetic WildSVG. The evaluation protocol is carefully designed with multiple metrics to capture different aspects of output quality. Overall, framing SVG extraction is a well-motivated problem with clear definitions."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"See weakness section."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The authors evaluate a wide range of state-of-the-art vision-language models (Qwen, Gemini, Claude, GPT, StarVector, GLM) on both Natural and Synthetic WildSVG test sets.\n\n2. They use four complementary metrics (L2, SSIM for pixel fidelity; LPIPS, DINO for perceptual/semantic similarity), which is an appropriate choice to capture different aspects of the generated SVG."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. How often did the VLLM-based SVG matching fail or produce incorrect logo–SVG pairs in Natural WildSVG? Were any manual checks done, and how sensitive are the results to these mismatches?\n\n2. Can you clarify how the “focus prompt” is formulated and used? If the prompt is ambiguous or generic, how does it affect the model’s output?\n\n3. Can you provide more details on the synthetic data creation? How diverse are the embedded SVG contexts (lighting, occlusion, styles)?\n\n4. How do you ensure that the chosen pixel/semantic metrics correlate with true SVG fidelity?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762940253672,"tcdate":1761995537197,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21011/Reviewer_9H4j"],"signatures":["ICLR.cc/2026/Conference/Submission21011/Reviewer_9H4j"],"forum":"Tab9dmIGRg","number":3,"license":"CC BY 4.0","cdate":1761995537197,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21011/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762940253672,"domain":"ICLR.cc/2026/Conference","replyto":"Tab9dmIGRg","id":"Ye8ir6jG5p","forumContent":{"TLDR":{"value":"We present a new dataset and benchmark for a new SVG generation task"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["SVG","VLLM","LLM"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We introduce SVG extraction, the task of translating specific visual inputs into scalable vector graphics. Existing multimodal models such as StarVector achieve strong results when generating SVGs from clean renderings or textual descriptions, but they fall short in real-world scenarios where natural images introduce noise, clutter, and domain shifts. To address this gap, we extend StarVector’s capabilities toward robust vision-to-SVG translation in the wild. A central challenge in this direction is the lack of suitable benchmarks. To fill this need, we develop two complementary datasets: Natural WildSVG, consisting of real-world images paired with SVG annotations, and Synthetic WildSVG, which integrates complex and elaborate SVG designs into real-life scenarios to simulate challenging conditions. Together, these resources provide the first foundation for systematic benchmarking SVG extraction. Building on them, we benchmark StarVector and related models. Our study establishes SVG extraction as a new problem domain, introduces datasets and evaluation protocols for its study, taking initial steps toward extending multimodal LLMs to handle reliable SVG generation in complex, natural scenes."},"_bibtex":{"value":"@misc{\nrodriguez2025wildsvg,\ntitle={Wild{SVG}: Towards reliable {SVG} generation under Real-Word conditions},\nauthor={Marco Terral Rodriguez and David Vazquez and Juan A. Rodriguez},\nyear={2025},\nurl={https://openreview.net/forum?id=Tab9dmIGRg}\n}"},"title":{"value":"WildSVG: Towards reliable SVG generation under Real-Word conditions"},"pdf":{"value":"/pdf/f35098ef94fe2341669e3ee2ae1d684076331407.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"rodriguez|wildsvg_towards_reliable_svg_generation_under_realword_conditions"},"authorids":{"value":["~Marco_Terral_Rodriguez1","~David_Vazquez1","~Juan_A._Rodriguez1"]},"authors":{"value":["Marco Terral Rodriguez","David Vazquez","Juan A. Rodriguez"]}},"version":2},{"content":{"venue":{"value":"Communications Physics"},"pdf":{"value":"https://www.nature.com/articles/s42005-026-02627-2.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"baretz|towards_aiassisted_neutrino_flavor_theory_design"},"html":{"value":"https://doi.org/10.1038/s42005-026-02627-2"},"abstract":{"value":"Particle physics theories, such as those which explain neutrino flavor mixing, arise from a vast landscape of model-building possibilities. A model’s construction typically relies on the intuition of theorists. It also requires considerable effort to identify appropriate symmetry groups, assign field representations, and extract predictions for comparison with experimental data. We develop Autonomous Model Builder (AMBer), a framework in which a reinforcement learning agent interacts with a streamlined physics software pipeline to search these spaces efficiently. AMBer selects symmetry groups, particle content, and group representation assignments to construct models while minimizing the number of free parameters introduced. We validate our approach in well-studied regions of theory space and extend the exploration to a previously unexamined symmetry group. While demonstrated in the context of neutrino flavor theories, this approach of reinforcement learning with physics software feedback may be extended to other theoretical model-building problems in the future. The vast landscape of particle physics model-building, particularly for neutrino flavor mixing, presents challenges in identifying appropriate symmetry groups and minimizing theoretical complexity. Here, the authors develop an Autonomous Model Builder (AMBer) using reinforcement learning to efficiently explore theory spaces, offering a novel approach that could transform theoretical model-building across physics."},"title":{"value":"Towards AI-assisted neutrino flavor theory design"},"authors":{"value":[{"fullname":"Jason Benjamin Baretz","username":"https://orcid.org/orcid-search/search?searchQuery=Jason%20Benjamin%20Baretz"},{"fullname":"Max Fieg","username":"https://orcid.org/orcid-search/search?searchQuery=Max%20Fieg"},{"fullname":"Vijay Ganesh","username":"https://orcid.org/orcid-search/search?searchQuery=Vijay%20Ganesh"},{"fullname":"Aishik Ghosh","username":"https://orcid.org/orcid-search/search?searchQuery=Aishik%20Ghosh"},{"fullname":"V. Knapp-Pérez","username":"~V._Knapp-Pérez1"},{"fullname":"Jake Rudolph","username":"~Jake_Rudolph1"},{"fullname":"Daniel Whiteson","username":"https://orcid.org/orcid-search/search?searchQuery=Daniel%20Whiteson"}]}},"tmdate":1777852779031,"pdate":1777593600000,"externalIds":["doi:10.1038/s42005-026-02627-2"],"tcdate":1777829501669,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Jake_Rudolph1"],"forum":"mE4g5IrNnp","license":"CC BY-SA 4.0","number":65844,"cdate":1777602714650,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1777852779031,"domain":"OpenReview.net/Public_Article","id":"mE4g5IrNnp","version":2},{"content":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Withdrawn by Authors"},"abstract":{"value":"We present, \\textbf{AdaFNIO} - Adaptive Fourier Neural Interpolation Operator, a neural operator-based architecture to perform synthetic frame generation. Current deep learning-based methods rely on local convolutions for feature learning and suffer from not being scale-invariant, thus requiring training data to be augmented through random flipping and re-scaling. On the other hand, \\textbf{AdaFNIO} leverages the principles of physics to learn the features in the frames, independent of input resolution, through token mixing and global convolution in the Fourier spectral domain by using Fast Fourier Transform (FFT). We show that \\textbf{AdaFNIO} can produce visually smooth and accurate results. To evaluate the visual quality of our interpolated frames, we calculate the structural similarity index (SSIM) and Peak Signal to Noise Ratio (PSNR) between the generated frame and the ground truth frame. We provide the quantitative performance of our model on Vimeo-90K dataset, DAVIS, UCF101 and DISFA+ dataset. Lastly, we apply the model to in-the-wild videos such as photosynthesis, brain MRI recordings and red blood cell animations"},"_bibtex":{"value":"@article{\nanonymous2024adafnio,\ntitle={Ada{FNIO}: A Physics-Informed Adaptive Fourier Neural In- terpolation Operator for Synthetic Frame Generation},\nauthor={Anonymous},\njournal={Submitted to Transactions on Machine Learning Research},\nyear={2024},\nurl={https://openreview.net/forum?id=Yi9mtDeWL7},\nnote={Withdrawn}\n}"},"title":{"value":"AdaFNIO: A Physics-Informed Adaptive Fourier Neural In- terpolation Operator for Synthetic Frame Generation"},"pdf":{"value":"/pdf/b2ca34827e7bb85e884885f3fcba769b75806b5e.pdf"},"venueid":{"value":"TMLR/Withdrawn_Submission"},"assigned_action_editor":{"value":"~Yanwei_Fu2"}},"tmdate":1726599816398,"tcdate":1715898252515,"writers":["TMLR","TMLR/Paper2704/Authors"],"signatures":["TMLR/Paper2704/Authors"],"forum":"Yi9mtDeWL7","license":"CC BY 4.0","number":2704,"cdate":1715898252515,"readers":["everyone"],"invitations":["TMLR/-/Submission","TMLR/-/Edit","TMLR/-/Under_Review","TMLR/-/Withdrawn"],"mdate":1726599816398,"odate":1715996747518,"domain":"TMLR","id":"Yi9mtDeWL7","version":2},{"content":{"summary":{"value":"This paper introduces an end-to-end framework, Predictive Inverse Dynamics Models (PIDM), which predicts actions using inverse dynamics models conditioned on the robot’s forecasted visual states. By integrating vision and action in a closed-loop system, the end-to-end PIDM serves as a scalable and effective action learner. The approach demonstrates improved performance over state-of-the-art methods, both in simulations and on a real robot."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The joint prediction of inverse dynamics (actions) and forward dynamics (next observation) reminds me of RPT [1], where it performs random masking of action and latent representations of observations and pre-trained via mask reconstruction. The key difference between this approach and RPT is that RPT doesn’t predict exact forward dynamics due to the random mask pattern, so it can learn shortcuts by copying from other future frames. Is there some way to evaluate the importance of predicting forward dynamics? Specifically, for table 3(b): is it possible to pre-train with only inverse dynamics?\n2. For finetuning the model, are all weights considered, or is LoRA applied to the model? It would be interesting to know how this method's performance scales with the amount of parameters used for fine-tuning. \n\n[1] Radosavovic, Ilija, Baifeng Shi, Letian Fu, Ken Goldberg, Trevor Darrell, and Jitendra Malik. \"Robot learning with sensorimotor pre-training.\" In Conference on Robot Learning, pp. 683-693. PMLR, 2023.\n\n[2] Hu, Edward J., Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. \"Lora: Low-rank adaptation of large language models.\" arXiv preprint arXiv:2106.09685 (2021)."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Thorough ablation studies and experiments that demonstrate the effectiveness of jointly predicting forward and inverse dynamics as a pre-training task for robot learning\n2. Evaluation in both simulation and in real on the effectiveness of the method\n3. Shows scalability of the proposed method by pre-training on DROID."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Learning inverse dynamics models from visual inputs has been explored in the past (i.e. [1,2]). It would be good to discuss these papers in the context of this paper. \n2. The paper shows scalability in the direction of pre-training and finetuning data. To fully demonstrate scalability, it would be good to demonstrate the scalability in the model capacity axis as well.\n3. The current formulation of the model does not seem to take account of the history (past observations). This makes it challenging to extend to more complex environments where stronger task planning is needed. \n\n[1] Agrawal, Pulkit, Ashvin V. Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine. \"Learning to poke by poking: Experiential learning of intuitive physics.\" Advances in neural information processing systems 29 (2016).[1] Agrawal, Pulkit, Ashvin V. Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine. \"Learning to poke by poking: Experiential learning of intuitive physics.\" Advances in neural information processing systems 29 (2016).\n\n[2] Brandfonbrener, David, Ofir Nachum, and Joan Bruna. \"Inverse dynamics pretraining learns good representations for multitask imitation.\" Advances in Neural Information Processing Systems 36 (2024)."}},"nonreaders":[],"tmdate":1731428147954,"tcdate":1730336838794,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7358/Reviewer_zoME"],"signatures":["ICLR.cc/2025/Conference/Submission7358/Reviewer_zoME"],"forum":"meRCKuUpmc","number":4,"license":"CC BY 4.0","cdate":1730336838794,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7358/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428147954,"domain":"ICLR.cc/2025/Conference","replyto":"meRCKuUpmc","id":"QO2L5iTJWm","forumContent":{"venue":{"value":"ICLR 2025 Oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Robotic Manipulation ; Pre-training ; Visual Foresight ; Inverse Dynamics ; Large-scale robot dataset"]},"supplementary_material":{"value":"/attachment/b0834376a4280c3d172cd07cd3b254c105e2645d.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on \"action,\" which involves behavior cloning from extensive collections of robotic data, while the other emphasizes \"vision,\" enhancing model generalization by pre-training representations or generative models, also referred to as world models, using large-scale visual datasets. This paper presents an end-to-end paradigm that predicts actions using inverse dynamics models conditioned on the robot's forecasted visual states, named Predictive Inverse Dynamics Models (PIDM). By closing the loop between vision and action, the end-to-end PIDM can be a better scalable action learner. In practice, we use Transformers to process both visual states and actions, naming the model Seer. It is initially pre-trained on large-scale robotic datasets, such as DROID, and can be adapted to real-world scenarios with a little fine-tuning data. Thanks to large-scale, end-to-end training and the continuous synergy between vision and action at each execution step, Seer significantly outperforms state-of-the-art methods across both simulation and real-world experiments. It achieves improvements of 13% on the LIBERO-LONG benchmark, 22% on CALVIN ABC-D, and 43% in real-world tasks. Notably, it demonstrates superior generalization for novel objects, lighting conditions, and environments under high-intensity disturbances. Code and models will be publicly available."},"_bibtex":{"value":"@inproceedings{\ntian2025predictive,\ntitle={Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation},\nauthor={Yang Tian and Sizhe Yang and Jia Zeng and Ping Wang and Dahua Lin and Hao Dong and Jiangmiao Pang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=meRCKuUpmc}\n}"},"title":{"value":"Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation"},"pdf":{"value":"/pdf/91a1d5e7a6417130767d8f9fb4625586f6f9a16f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"tian|predictive_inverse_dynamics_models_are_scalable_learners_for_robotic_manipulation"},"authorids":{"value":["~Yang_Tian1","~Sizhe_Yang3","~Jia_Zeng2","~Ping_Wang7","~Dahua_Lin1","~Hao_Dong3","~Jiangmiao_Pang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yang Tian","Sizhe Yang","Jia Zeng","Ping Wang","Dahua Lin","Hao Dong","Jiangmiao Pang"]}},"version":2},{"content":{"summary":{"value":"The paper targets efficient training of a class of flow models called Rectified Flow (RF) trained using flow matching objective. The paper has two broad contributions: (1) justification of 2-RF (‘reflow’-ed once) being close to optimal, (2) and using those findings to improve training of 2-RF.\n\nAuthors argued that, in practical scenario, the pairs generated by optimal 1-RF model is ‘crossing-free’, i.e. the stochastic interpolating paths rarely cross each other. They provided an intuitive and empirical evidence of the same. They claimed this to be a motivating factor for two improvements of 2-RF model training — a new time-step distribution $p_t$ & the  a new distance measure for the regression task of diffusion objective.\n\nGood empirical performance is shown in terms of FID on several datasets and models. An ablation study is also done for the relevant part of the finding."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Authors said “*any other premetric .. deviates from the posterior expectation*” — but is it really true ? Can you mathematically show why ?\n2. Is it possible to have trajectory crossings *at all* when samples are from $p_{xz}^2$ ? If no, can you prove it formally ?\n3. L99: “.. use a specific non-linear interpolation” — what is that exactly ? I thought the interpolation is still linear.\n4. Eq.4: How did you decompose the loss — can you show the steps ? And what is $\\bar{\\mathcal{L}}$ ?\n5. L65: “*training loss of 2-RF has zero lower bound*” — please clarify: Isn’t it true that *any* L2 loss has zero lower bound ? How does it matter whether trajectories cross or not ? Even the FM loss with independent coupling $p_{\\mathbf{xz}}^1$ has zero lower bound — no ?\n6. I think it is unclear which quantity the authors are arguing to be zero when trajectories do not cross. The term ‘curvature’ was used some times (L124) — what does that mean ? Can you write this object in mathematical terms ? Just curious, what is the equation that needs to be proved if one wants to show that no trajectories from 1-RF coupling cross each other (ideal case) ? (Related to Q2 above)."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The problem statement chosen and the arguments provided are quite credible. It is well known that, for RF to work well, one requires several reflows. Reflows are expensive, hence the contributions of the paper (if credible) can be significant.\n- A good intuitive analysis is done by the authors to justify their argument of 2-RF being near-optimal."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I like the overall outcome of the paper in terms of empirical performance. I also like the problem statement. However, some concerns remain. Following points explain the issues and points to some more questions in question section.\n\n### Major concerns\n\n- This is a major concern for me. The paper has two parts: (1) justifying that 2-RF is near optimal & providing intuitive and empirical evidence; (2) proposing two new measures for training improvement. I think, (1) & (2) both are overall correct on their own. I just don’t think that (1) is the right motivator for (2).\n- Continuing the above point, I felt that the rationale behind both ‘improvements’ are weakly connected to the empirical observations (in section 3). After all, what the authors ultimately proposed are ad-hoc and exists in the literature. (L162) “.. *focusing on the tasks where training loss is high ..*” is a known technique which authors admitted themselves. Using non-L2 loss is also not unheard of [1]. More importantly, I don’t think (not) using L2 loss has anything to do with your observation is section 3 (as argued in L195). The L2 loss has its origin in score-matching which yields theoretical benefits, but I suspect it is not necessary. Q1 is related to this.\n- Let’s talk about the observation of section 3 itself (i.e. 2-RF being optimal). Authors must be clear about whether they are making a theoretical assertion (L64-65) or just an empirical observation (they used words like “under *realistic* setting”, “*rarely* intersect each other”). If you are providing a *guarantee*, you must provide better formal proofs. To clarify, I think the argument provided in section 3 is indeed correct and seems reasonable. But if it is a guarantee, empirical evidence is not enough. Q2 is a related question.\n\n[1] “Improving Diffusion Models's Data-Corruption Resistance using Scheduled Pseudo-Huber Loss”, Kharpov et al., arXiv 2403.16728.\n\n### Presentations\n\n- The notation used in section 2 & 3 are confusing sometimes. There are three pairs of notation — $(\\mathbf{x}, \\mathbf{z})$, $(\\mathbf{x}_0, \\mathbf{z}_0)$ \\& $(\\mathbf{x}_1, \\mathbf{z}_1)$. I am confused about which is what. I would recommend authors to follow the same notation as the original RF paper by Liu et al.\n- Notations like $\\mathbf{z}\\_0 \\sim p\\_{\\mathbf{x}}$ (L86) are very confusing.\n- The factor $\\frac{t}{1-t}$ in Fig 2(a) (at the top left) should be $\\frac{1-t}{t}$, right ?\n- Fig. 2 caption says $\\mathbf{z}'' = \\mathbf{z} + (\\mathbf{x}' - \\mathbf{x}'')$ — where is the factor $\\frac{1-t}{t}$ ? Did you assume $t = 0.5$ ?\n- Again, it is hard to parse notations like $\\mathbf{x}, \\mathbf{x}', \\mathbf{x}''$ or $\\mathbf{z}, \\mathbf{z}', \\mathbf{z}''$.\n- Eq.4: Can you please denote the suffix of the $\\mathbb{E}_{??}\\left[ \\cdot \\right]$ properly ? It is hard to read otherwise.\n\n### Results/Experiments\n\nResult section is okay-ish. The following are some comments/questions.\n\n- Table.1: Need clarification: ‘Base (A)’ is a 2-RF, and the other ones are written as “(A) + <something>” — does that mean they are 3-RF (meaning 2 reflows) ?\n- Section 5.2 seems totally unnecessary. That has nothing to do with the core contributions of the paper.\n- Section 5.3 is also very much unnecessary. I even doubt its correctness. You seem to be proposing a new sampler/solver with a very weak (intuitive) motivation. Designing a solver require a lot more than that. And then “*.. detailed analysis is provided in appendix E*” — appendix E barely has any details !  Also, obvious question, why are FIDs going up (fig.4) with higher NFE ? Does that even make sense ?\n- Fig.5(b): The inverted noise norm distribution still looks quite different (higher variance) from the true noise. Just having the norm closer to the truth isn’t necessarily making a good case for your method."},"limitations":{"value":"Some limitation are mentioned, which are reasonable."}},"nonreaders":[],"tmdate":1730879395369,"tcdate":1719167648343,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission10536/Reviewer_aFYc"],"signatures":["NeurIPS.cc/2024/Conference/Submission10536/Reviewer_aFYc"],"forum":"mSHs6C7Nfa","number":2,"license":"CC BY 4.0","cdate":1719167648343,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission10536/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879395369,"domain":"NeurIPS.cc/2024/Conference","replyto":"mSHs6C7Nfa","id":"fVRoSovWKc","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"We propose improved techniques for training rectified flows, allowing them to compete with knowledge distillation methods even in the low NFE setting"},"keywords":{"value":["generative modeling","rectified flow","diffusion model"]},"primary_area":{"value":"generative_models"},"abstract":{"value":"Diffusion models have shown great promise for image and video generation, but sampling from state-of-the-art models requires expensive numerical integration of a generative ODE.\n    One approach for tackling this problem is rectified flows, which iteratively learn smooth ODE paths that are less susceptible to truncation error.\n    However, rectified flows still require a relatively large number of function evaluations (NFEs).\n    In this work, we propose improved techniques for training rectified flows, allowing them to compete with knowledge distillation methods even in the low NFE setting.\n    Our main insight is that under realistic settings, a single iteration of the Reflow algorithm for training rectified flows is sufficient to learn nearly straight trajectories; hence, the current practice of using multiple Reflow iterations is unnecessary.\n    We thus propose techniques to improve one-round training of rectified flows, including a U-shaped timestep distribution and LPIPS-Huber premetric.\n    With these techniques, we improve the FID of the previous 2-rectified flow by up to 75\\% in the 1 NFE setting on CIFAR-10.\n    On ImageNet 64$\\times$64, our improved rectified flow outperforms the state-of-the-art distillation methods\n    such as consistency distillation and progressive distillation in both one-step and two-step settings and rivals the performance of improved consistency training (iCT) in FID.\n    Code is available at https://github.com/sangyun884/rfpp."},"_bibtex":{"value":"@inproceedings{\nlee2024improving,\ntitle={Improving the Training of Rectified Flows},\nauthor={Sangyun Lee and Zinan Lin and Giulia Fanti},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=mSHs6C7Nfa}\n}"},"title":{"value":"Improving the Training of Rectified Flows"},"pdf":{"value":"/pdf/876809b80692e3d3bb48e5861b766c0b86adece6.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"lee|improving_the_training_of_rectified_flows"},"authorids":{"value":["~Sangyun_Lee1","~Zinan_Lin1","~Giulia_Fanti1"]},"authors":{"value":["Sangyun Lee","Zinan Lin","Giulia Fanti"]}},"version":2},{"content":{"summary":{"value":"The authors propose a method for open surface reconstruction from 3D point clouds. They train a network to predict unsigned distance functions (UDFs) from point cloud patches using only synthetic data of quadratic surfaces. Evaluation shows that the trained network generalizes well to other complex patterns and is more resilient to noise when reconstructing 3D surfaces from point clouds."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Main questions I would like the answers for:\n- Analyses on using the synthetic patches to approximate the real local 3D geometries.\n- Speed test.\n- Quantitative results on real-world scanned data, such as 3D-Scene dataset."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The idea of training a UDF regression network using only synthetic data of quadratic surfaces is quite intriguing. This approach allows for a controlled and systematic way to generate training data, which can be more consistent and free from the imperfections and variability found in real-world data. \n\nUsing quadratic surfaces as a basis for synthetic data is quite interesting where the authors argue that quadratic surfaces can approximate various local geometries. I have some doubts about this but it is reasonable and novel to a certain extend."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I have three main concerns: potential biases with the synthetic training data, the evaluation scheme, and the applicability in practice due to the patch radius.\n\n\n## Potential Biases with Synthetic Training Data: \n\nThe use of primitive geometrical patches as training data for a UDF regressor might introduce biases. The observation that any local geometries can be approximated by quadratic surfaces is only valid at a very fine resolution, which requires dense point clouds to observe reliably. It is unclear how many patches have been synthesized and whether they provide a good approximation of universal geometrical primitives. Some analyses would be helpful here: for example, using the ShapeNet car dataset, cropping all local patches from each car, and finding the closest synthesized one to check for approximation errors. Are there any patterns with sufficiently high approximation errors? Can we perform these analyses at different resolutions and see how they correlate with surface reconstruction?\n\n## Evaluation Scheme\n\nThe testing data is simulated to match the scenarios the method is designed for: the point cloud is quite dense, and artificial noise is added similarly to the training data. There is no quantitative evaluation for real-world scanned data. \n\n## How sensitive is the method to different patch radii?\n\nThe value of 0.018 is oddly specific. I suspect that with a larger value of r, the method will generate overly smooth surfaces (as shown in Figure 6 - right), and with a smaller value of r, it will generate holes due to the point cloud not being dense enough. Overall, there is an inherent issue with this trade-off that may not be resolvable with this approach. Detailed experiments varying the patch radius and analyzing the impact on reconstruction quality would be helpful.\n\n## Additional Concerns\n\n- Ablation studies on the network architecture are missing. I am unsure about the roles of the two branches and the cross-attention mechanism. There is a potential issue with the reliance on Point-Net for embedding computation. Is this network pre-trained on other datasets? If so, we should be careful with the claim of using only synthetic data, as pre-training on real data could influence the results.\n\n- It also seems that the method could be quite slow. The authors should include a speed test to provide insights into the computational efficiency of the proposed approach. Evaluating the method's runtime on different hardware setups and for varying point cloud sizes would give a clearer picture of its practical applicability."},"limitations":{"value":"The authors acknowledged that the proposed method cannot handle incomplete point cloud. However, it is unclear what how resilient it is when dealing with this."}},"nonreaders":[],"tmdate":1730879057531,"tcdate":1720799463498,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission6150/Reviewer_AbEM"],"signatures":["NeurIPS.cc/2024/Conference/Submission6150/Reviewer_AbEM"],"forum":"7nbAots3f8","number":4,"license":"CC BY 4.0","cdate":1720799463498,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission6150/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879057531,"domain":"NeurIPS.cc/2024/Conference","replyto":"7nbAots3f8","id":"gjgKPaBtCG","forumContent":{"venue":{"value":"Submitted to NeurIPS 2024"},"keywords":{"value":["Surface Reconstruction","Implicit Fields","Unsigned Distance Fields"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Unsigned distance fields (UDFs) provide a versatile framework for representing a diverse array of 3D shapes, encompassing both watertight and non-watertight geometries. Traditional UDF learning methods typically require extensive training on large datasets of 3D shapes, which is costly and often necessitates hyperparameter adjustments for new datasets. This paper presents a novel neural framework, LoSF-UDF, for reconstructing surfaces from 3D point clouds by leveraging local shape functions to learn UDFs. We observe that 3D shapes manifest simple patterns within localized areas, prompting us to create a training dataset of point cloud patches characterized by mathematical functions that represent a continuum from smooth surfaces to sharp edges and corners. Our approach learns features within a specific radius around each query point and utilizes an attention mechanism to focus on the crucial features for UDF estimation. This method enables efficient and robust surface reconstruction from point clouds without the need for shape-specific training. Additionally, our method exhibits enhanced resilience to noise and outliers in point clouds compared to existing methods. We present comprehensive experiments and comparisons across various datasets, including synthetic and real-scanned point clouds, to validate our method's efficacy."},"_bibtex":{"value":"@misc{\nanonymous2024learning,\ntitle={Learning Unsigned Distance Fields from Local Shape Functions for 3D Surface Reconstruction},\nauthor={Anonymous},\nyear={2024},\nurl={https://openreview.net/forum?id=7nbAots3f8}\n}"},"title":{"value":"Learning Unsigned Distance Fields from Local Shape Functions for 3D Surface Reconstruction"},"pdf":{"value":"/pdf/95b69584cccda3dda361615635fc9f18580876e9.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"hu|learning_unsigned_distance_fields_from_local_shape_functions_for_3d_surface_reconstruction"},"authorids":{"value":["~Jiangbei_Hu1","~Yanggeng_Li1","~Fei_Hou1","~Junhui_Hou2","~Zhebin_Zhang2","~Shengfa_Wang1","~Na_Lei1","~Ying_He1"]},"authors":{"value":["Jiangbei Hu","Yanggeng Li","Fei Hou","Junhui Hou","Zhebin Zhang","Shengfa Wang","Na Lei","Ying He"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Physics-Inspired Neural Compression (PINC), a method that integrates physics-informed losses to compress massive gyrokinetic plasma simulation data by up to 70,000×. It preserves key spatial and temporal turbulence characteristics that conventional compression methods may fail to maintain."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Have the authors considered combining VQ-VAE with an entropy codec (e.g., arithmetic or Huffman coding) for improved hybrid compression efficiency?\n2. What is the codebook utilization of the VQ-VAE component, and how does it affect compression quality and efficiency?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Significance\nThis paper compares compression methods for high-dimensional scientific simulations, including neural field and vector quantized approaches in plasma turbulence, and presents solid experiments in a field lacking established benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Originality\n\n1. The proposed Physics-Inspired Neural Compression (PINC) lacks methodological novelty, as similar physics-informed compression frameworks have been explored in prior works such as GINN [1] and VQ-VAE for tubulence [2]. The paper mainly adopts existing  method on the plasma turbulence dataset without introducing a new architectural or algorithmic contribution.\n\nQuality and Clarity\n\n2. The paper does not clearly explain how the proposed physics-informed loss differs from or improves upon prior physics-guided neural compression works [1][2].\n\nSignificance \n\n3. The experimental results (Table 1) indicate a trade-off between PSNR and L1 error when incorporating physics-informed losses, suggesting that improved physical fidelity comes at the cost of reconstruction accuracy.\n4. While the evaluation pipeline is comprehensive, the method itself is incremental, and its advantages over previous approaches are not sufficiently quantified.\n5. The novelty and generalizability of the proposed method are limited, which may reduce its long-term impact compared to its benchmarking contributions.\n\n[1] Geometry-Informed Neural Networks  \n[2] A Physics-Informed Vector Quantized Autoencoder for Data Compression of Turbulent Flow"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927981640,"tcdate":1761214283332,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18255/Reviewer_WgAS"],"signatures":["ICLR.cc/2026/Conference/Submission18255/Reviewer_WgAS"],"forum":"fixbsplpdw","number":1,"license":"CC BY 4.0","cdate":1761214283332,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18255/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927981640,"domain":"ICLR.cc/2026/Conference","replyto":"fixbsplpdw","id":"2vzHcbaoAa","forumContent":{"TLDR":{"value":"Neural compression methods enable extreme compression of plasma turbulence simulation data while maintaining low reconstruction error and preserving key physical characteristics."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["physics-inspired","turbulence","plasma","neural compression","autoencoders","neural fields"]},"supplementary_material":{"value":"/attachment/0814372529d7d0e61c5f895f3e8f4b70b3af92df.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"High-fidelity scientific simulations are now producing unprecedented amounts of data, creating a storage and analysis bottleneck. A single simulation can generate tremendous data volumes, often forcing researchers to discard valuable information. A prime example of this is plasma turbulence described by the Gyrokinetic equations: nonlinear, multiscale, and 5D in phase space. They represent one of the most computationally demanding frontiers of modern science, with runs taking weeks and resulting in tens of terabytes of data dumps.\nThe increasing storage demands underscore the importance of compression, however, compressed snapshots might not preserve essential physical characteristics after reconstruction. To assess whether such characteristics are captured, we propose a spatiotemporal evaluation pipeline which accounts for structural phenomena and multi-scale transient fluctuations. Indeed, we find that various compression techniques lack preservation of temporal turbulence characteristics. Therefore, we explore Physics-Informed Neural Compression (PINC), which incorporates physics-informed losses tailored to gyrokinetics and enables extreme compressions of over 100000x. This direction provides a viable and scalable solution to the prohibitive storage demands of gyrokinetics, enabling post-hoc analyses that were previously infeasible."},"_bibtex":{"value":"@misc{\ngalletti2026physicspreserving,\ntitle={Physics-Preserving Compression of High-Dimensional Plasma Turbulence Simulations},\nauthor={Gianluca Galletti and Gerald Gutenbrunner and Fabian Paischer and Sandeep Suresh Cranganore and William Hornsby and Naomi Carey and Lorenzo Zanisi and Stanislas Pamela and Johannes Brandstetter},\nyear={2026},\nurl={https://openreview.net/forum?id=fixbsplpdw}\n}"},"title":{"value":"Physics-Preserving Compression of High-Dimensional Plasma Turbulence Simulations"},"pdf":{"value":"/pdf/88f0fd80bf4066acbbabbaac93b2beb84178ea60.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"galletti|physicspreserving_compression_of_highdimensional_plasma_turbulence_simulations"},"authorids":{"value":["~Gianluca_Galletti1","~Gerald_Gutenbrunner1","~Fabian_Paischer1","~Sandeep_Suresh_Cranganore1","~William_Hornsby1","~Naomi_Carey1","~Lorenzo_Zanisi1","~Stanislas_Pamela1","~Johannes_Brandstetter1"]},"authors":{"value":["Gianluca Galletti","Gerald Gutenbrunner","Fabian Paischer","Sandeep Suresh Cranganore","William Hornsby","Naomi Carey","Lorenzo Zanisi","Stanislas Pamela","Johannes Brandstetter"]}},"version":2},{"content":{"summary":{"value":"This paper studies 3D-Aware Classification and multiclass 3D pose estimation. It introduces a new framework CIDA-3D, which leverages unsupervised domain adaptation into these two tasks. The proposed method doesn't requires  3D data or object labels and only use unlabeled image in the target domain, allowing for generalized feature updates across multiple classes, even in noisy target domains. Via experiments, the proposed method is demonstrated to be able to adapt from synthetic to complex real-world target domains."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- The local parts robustness is reasonable, but this paper does not explore how different levels of local feature robustness influence performance. For example, many objects may exhibit different robustness levels in different local parts. A more fine-grained analysis could be very insightful.\n- Since Conformal Prediction is the main contribution of this work, I'm wondering if it's possible to demonstrate the Interpretability of this module, to help us better understand how well does this module work. \n- In L270-282 the authors explains on why conformal prediction was chosen over alternative uncertainty management methods. I'm wondering if there are empirical evidence to support these claims. \n- Is it possible to provide a way or metric to evaluate local part robustness and plurality in different dataset?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"The paper template was modified:\n- The font size of the figure caption is changed to be smaller\n- In all the tables the citations become numbers rather than author names. \n\nI have read the author guidance and cannot determine the severeness of above things. They probably are not a big deal but I raise it here for reference."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The strengths of this work can be summarized as below:\n- CIDA-3D introduces unsupervised domain adaptation into the tasks of 3D-aware classification and multiclass 3D pose estimation.\n- In target domain, 3D data or object labels are needed, and only unlabeled images are used.\n- The local part plurality and robustness hypothesis is well motivated and sound. The model builds on previous method 3DUDA by utilizing local part plurality and robustness to enhance the adaptation process, allowing the transfer of features across different classes in noisy or occluded environments.\n- CIDA-3D uses conformal prediction to manage covariate shifts between source and target domains, which reduces computational overhead while maintaining confidence in predictions.\n- The experiments demonstrate that CIDA-3D enables synthetic-to-real domain adaptation, bridging the sim2real domain gap and enhancing robustness against complex real-world nuisances such as texture, shape, and environmental variations."},"flag_for_ethics_review":{"value":["Yes, Other reasons (please specify below)"]},"weaknesses":{"value":"The weaknesses of this work are summarized below:\n- CIDA-3D relies on a domain classifier to distinguish two domain. This is practical but could introduce dependency on the classifier’s accuracy and increase complexity. The whole system is vulnerable to misclassification, and potentially lead to less effective conformal prediction. It's unclear how robust and what's the capacity of this domain classifier, or if it's possible to take alternative classifier-free UDA solutions. \n- The ablation studies lack depth regarding key components like the choice of neural mesh and conformal prediction. A detailed analysis of the important design choices, parameters, alternative solutions could help demonstrate the effectiveness of these models.\n- Since this paper tackles classification, to fully establish CIDA-3D’s generalizability it's crucial to test on large-scale classification datasets. For example, test on imagenet or openimage dataset where more object categories are presented. The current scope may limit the reader’s ability to assess performance in diverse applications, which would be especially important for real-world implementations."}},"nonreaders":[],"tmdate":1731428171842,"tcdate":1730917188577,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4571/Reviewer_k5Un"],"signatures":["ICLR.cc/2025/Conference/Submission4571/Reviewer_k5Un"],"forum":"XYFBmp08sP","number":4,"license":"CC BY 4.0","cdate":1730917188577,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4571/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428171842,"domain":"ICLR.cc/2025/Conference","replyto":"XYFBmp08sP","id":"P8OEHPmNHO","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"Image Only UDA for 3D Aware Classification and MultiClass 3D Pose Estimation"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["UDA","3D pose estimation","3D-Aware classification","occlusion","robustness"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Cognitive Science studies show that human perception becomes robust to occlusions and other nuisances due to internal 3D representations of objects. This idea has been incorporated into computer vision models to improve their ability to understand and reason about the 3D world. However, collecting 3D annotations in vision datasets is expensive. This makes the robustness of the perception model to distribution shifts challenging. We introduce Conformal Inference aided unsupervised Domain Adaptation (CIDA)-3D for the complex setting of multiclass pose estimation. Our method adapts category level pose estimation (3D) models in nuisance ridden target domains directly from images without class label information, by harnessing uncertainty in model predictions (using conformal sets). This allows for significantly better and computationally efficient adaptation to target domains with synthetic and real-world noise. We also show a robust adaptation from fully synthetic data to complex real-world domains. To the best of our knowledge, this method is the first to attempt unsupervised domain adaptation for robust 3D-aware classification and multiclass pose estimation in real-world scenarios by adapting models trained on procedurally generated synthetic data."},"_bibtex":{"value":"@misc{\nkaushik2024cidad,\ntitle={{CIDA}3D: Conformal Inference aided unsupervised Domain Adaptation for 3D-Aware Classification},\nauthor={Prakhar Kaushik and Aayush Mishra and Anqi Liu and Adam Kortylewski and Alan Yuille},\nyear={2024},\nurl={https://openreview.net/forum?id=XYFBmp08sP}\n}"},"title":{"value":"CIDA3D: Conformal Inference aided unsupervised Domain Adaptation for 3D-Aware Classification"},"pdf":{"value":"/pdf/6d7d30a5dc27f28d2e911b78a635e0b3cc8d1bbb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"kaushik|cida3d_conformal_inference_aided_unsupervised_domain_adaptation_for_3daware_classification"},"authorids":{"value":["~Prakhar_Kaushik1","~Aayush_Mishra1","~Anqi_Liu2","~Adam_Kortylewski1","~Alan_Yuille1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Prakhar Kaushik","Aayush Mishra","Anqi Liu","Adam Kortylewski","Alan Yuille"]}},"version":2},{"content":{"summary":{"value":"The paper introduces TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation. The dataset is generated by a complex, two-stage \"agentic pipeline\" involving four distinct Large Language Models (LLMs) : a Profile LLM (to create a listener profile), a Goal LLM (to set a conversation goal), a Listener LLM (to simulate the user), and a Recsys LLM (to act as the recommender). This pipeline is grounded in real music data from the LFM-2b dataset and augmented with multimodal information (audio, images) from Spotify. The primary claims are that this multi-agent, multimodal pipeline produces more diverse and \"natural\" conversations than prior single-LLM approaches , which is validated through LLM-as-a-judge and a small human evaluation."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- How can the authors justify claims of \"high naturalness\" (4.15/5) for the entire 16.5K dataset based on a human evaluation with only 26 raters? This sample is less than 0.2% of the test set conversations.\n- The validation relies heavily on LLM-as-a-judge (Table 6), using Gemini 2.5 Pro to judge Gemini 2.5 Flash. How do you account for the well-known self-enhancement bias where LLMs favorably rate outputs from their own architectural family?\n- The core contributions over TalkPlayData 1 are (1) multimodality and (2) a multi-agent pipeline. Both are existing, known techniques. What is the fundamental research contribution of this paper beyond a simple (and very complex) engineering integration?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The agentic design correctly identifies a flaw in single-LLM generation, which is the tendency to \"cheat\" by seeing all information. The separation of information (e.g., Recsys LLM not seeing the goal) is a good design choice.\n- The ablation study (Table 8) effectively demonstrates that within their own complex pipeline, the Goal and Profile agents are necessary for achieving topical diversity."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Critically Lacking Novelty: The paper is highly incremental. It combines TalkPlayData 1 with known multi-agent techniques and off-the-shelf multimodal LLMs.\n- Fundamentally Unsound Evaluation: The work relies almost entirely on LLM-as-a-judge validation, which is self-referential. The human validation is too small to be statistically valid.\n- Compounding Error: The 4-agent pipeline  is a complex, brittle system where errors in early stages (e.g., a bad profile) cannot be caught and will poison the final generated conversation.\n- Cost and Complexity: The proposed pipeline is incredibly complex and expensive (reports $109.08 for 1,000 conversations ) and it is not clear that this massive increase in complexity yields a meaningfully better dataset than the simpler, single-LLM approaches it aims to replace, especially given the weak validation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918358380,"tcdate":1762884203483,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5932/Reviewer_x9zc"],"signatures":["ICLR.cc/2026/Conference/Submission5932/Reviewer_x9zc"],"forum":"Phv5G3rxOx","number":4,"license":"CC BY 4.0","cdate":1762884203483,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5932/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918358380,"domain":"ICLR.cc/2026/Conference","replyto":"Phv5G3rxOx","id":"UwpVxVcbo9","forumContent":{"TLDR":{"value":"Use multiple, multimodal LLMs with specialized roles and prompts to synthesize a dataset for conversational music recommendation"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["music","recsys","dataset","conversational recommendation","LLM"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In the proposed pipeline, multiple large language model (LLM) agents are created under various roles with specialized prompts and access to different parts of information, and the chat data is acquired by logging the conversation between the Listener LLM and the Recsys LLM. To cover various conversation scenarios, for each conversation, the Listener LLM is conditioned on a finetuned conversation goal. Finally, all the LLMs are multimodal with audio and images, allowing a simulation of multimodal recommendation and conversation. In the LLM-as-a-judge and subjective evaluation experiments, TalkPlayData 2 achieved the proposed goal in various aspects related to training a generative recommendation model for music."},"_bibtex":{"value":"@misc{\nchoi2026talkplaydata,\ntitle={TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal  Conversational Music Recommendation},\nauthor={Keunwoo Choi and SeungHeon Doh and Juhan Nam},\nyear={2026},\nurl={https://openreview.net/forum?id=Phv5G3rxOx}\n}"},"title":{"value":"TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal  Conversational Music Recommendation"},"pdf":{"value":"/pdf/e50311560ba5fd8b18d1480815010c3404c7a0cc.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"choi|talkplaydata_2_an_agentic_synthetic_data_pipeline_for_multimodal_conversational_music_recommendation"},"authorids":{"value":["~Keunwoo_Choi1","~SeungHeon_Doh1","~Juhan_Nam1"]},"authors":{"value":["Keunwoo Choi","SeungHeon Doh","Juhan Nam"]}},"version":2},{"content":{"summary":{"value":"The paper proposes training RNNs to reproduce behavioral distributions including characteristic error modes than optimizing task performance. They do so by first using a generative model to produce synthetic data for training RNNs, then adapting diffusion training procedures for RNNs to produce realistic behavioral distributions. The authors tested their method on the visual working memory task, and found that the DDPM-trained RNNs, instead of the task-optimized RNNs, produces realistic swap errors and neural representations."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Does well does the method generalize to a more complex versions of the task, for example, with more items?\n2. How does the method apply to tasks where an output is required at every time step of the trial? \n3. Fig 3C: for the brown-border model, shouldn’t close probe feature values have more swap error? (It looks like it's the opposite in the figure)."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Training RNNs to produce realistic and complex behavioral distributions is a very important and timely research direction. \n2. This paper is very well-written and easy to follow. \n3. The experiments are well designed and the authors provided analysis at both the behavioral and representational level."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The BNS model used to generate synthetic data still requires hand-tuning, so “automatic discovery” feels overstated. Could the authors comment on this limitation?\n2. In Figure 3A, the correspondence between border colors and RNN models is unclear. Please state this explicitly in the caption (or add a legend).\n3. On line 203, (F) has not been introduced, and the network equation is missing. It would help to present Equation 3 (or at least the network update equation within it) earlier.\n4. On line 252, I think the intended expression is (x = G(r) = W_x r) (the r appears to be missing).\n5. The approach depends on a generative model to create synthetic behavioral data for training the RNNs, which raises several concerns:\n\n   * For more complex tasks, well-founded hypotheses or statistical/generative models of the behavioral process may be unknown or unavailable.\n   * The method assumes the generative model captures key characteristics of behavior, yet the paper provides no quantitative evidence for this beyond showing that swap error increases when the probe is closer to the cued item.\n   * Lines 146–149 claim that models listed in §1 “can be used to faithfully generate synthetic behavioural data,” but this is not supported with empirical evidence or citations to prior work.\n\nIf these concerns are properly addressed, I may consider increasing my score."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359626888,"tcdate":1760553953234,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14123/Reviewer_YEvN"],"signatures":["ICLR.cc/2026/Conference/Submission14123/Reviewer_YEvN"],"forum":"NGThArVrD3","number":1,"license":"CC BY 4.0","cdate":1760553953234,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14123/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359626888,"domain":"ICLR.cc/2026/Conference","replyto":"NGThArVrD3","id":"KCW30RJOLZ","forumContent":{"TLDR":{"value":"Training RNNs to reproduce realistic error patterns (rather than optimal performance) produces networks that better mimic biological neural computation, demonstrated through a working memory task where networks were taught to make swap errors."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["neuroscience","working memory","recurrent neural networks","diffusion models","behavioral modeling"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"abstract":{"value":"Discovering the neural mechanisms underpinning cognition is one of the grand challenges of neuroscience. However, previous approaches for building models of recurrent neural network (RNN) dynamics that explain behaviour required iterative refinement of architectures and/or optimisation objectives, resulting in a piecemeal, and mostly heuristic, human-in-the-loop process. Here, we offer an alternative approach that automates the discovery of viable RNN mechanisms by explicitly training RNNs to reproduce behaviour, including the same characteristic errors and suboptimalities, that humans and animals produce in a cognitive task. Achieving this required two main innovations. First, as the amount of behavioural data that can be collected in experiments is often too limited to train RNNs, we use a non-parametric generative model of behavioural responses to produce surrogate data for training RNNs. Second, to capture all relevant statistical aspects of the data, rather than a limited number of hand-picked low-order moments as in previous moment-matching-based approaches, we developed a novel diffusion model-based approach for training RNNs. To showcase the potential of our approach, we chose a visual working memory task as our test-bed, as behaviour in this task is well known to produce response distributions that are patently multimodal (due to so-called swap errors). The resulting network dynamics correctly predicted previously reported qualitative features of neural data recorded in macaques. Importantly, these results were not possible to obtain with more traditional approaches, i.e., when only a limited set of behavioural signatures (rather than the full richness of behavioural response distributions) were fitted, or when RNNs were trained for task optimality (instead of reproducing behaviour). Our approach also yields novel predictions about the mechanism of swap errors, which can be readily tested in experiments. These results suggest that fitting RNNs to rich patterns of behaviour provides a powerful way to automatically discover the neural network dynamics supporting important cognitive functions."},"_bibtex":{"value":"@inproceedings{\nradmard2026setting,\ntitle={Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors},\nauthor={Puria Radmard and Paul M. Bays and M{\\'a}t{\\'e} Lengyel},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=NGThArVrD3}\n}"},"title":{"value":"Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors"},"pdf":{"value":"/pdf/8a4c81e360a6555c93c2ae8b7bccb45eab515dfb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"radmard|setting_up_for_failure_automatic_discovery_of_the_neural_mechanisms_of_cognitive_errors"},"authorids":{"value":["~Puria_Radmard1","~Paul_M._Bays1","~Máté_Lengyel1"]},"authors":{"value":["Puria Radmard","Paul M. Bays","Máté Lengyel"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a global physics-informed loss function that avoids the need for higher-order automatic differentiation, which is costly and challenging for QNNs. The paper shows that the proposed method can reduce the number of shots required to train QNNs for solving ODEs by an order of magnitude compared to existing q-PINN methods."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The authors propose a novel method for solving differential equations using quantum neural networks on near-term quantum hardware, which can reduce the number of shots required to train QNNs by an order of magnitude compared to existing methods.\n\n2. It introduces a global physics-informed loss function that avoids the need for higher-order automatic differentiation, which is costly and challenging for QNNs.  The numerical experiments on simulated quantum circuits demonstrate the effectiveness and efficiency of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks a clear motivation and contribution to the field of quantum machine learning. It does not explain why solving differential equations with quantum neural networks is important or novel, and how it compares to existing methods in classical or quantum computing.\n\n2. It does not provide sufficient theoretical analysis or justification for the proposed method of global physics-informed losses. It does not show how the method is derived from general principles, what assumptions or limitations it has, and how it guarantees the convergence or accuracy of the solution.\n\n3. The paper does not present any rigorous experimental results or benchmarks to demonstrate the effectiveness or efficiency of the proposed method. It only shows some qualitative plots of the solutions for three simple ODEs, without any quantitative metrics, error analysis, or comparison with other methods especially the classical learning-based methods for solving ODEs. It also does not report any details on the implementation, such as the number of qubits, circuit depth, optimization algorithm, hyperparameters, etc."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Please see the weaknesses."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636675457,"tcdate":1698050823328,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6201/Reviewer_Nx77"],"signatures":["ICLR.cc/2024/Conference/Submission6201/Reviewer_Nx77"],"forum":"2EamGPuWSc","number":1,"license":"CC BY 4.0","cdate":1698050823328,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6201/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636675457,"domain":"ICLR.cc/2024/Conference","replyto":"2EamGPuWSc","id":"uQ23p8yZOJ","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Variational Quantum Algorithms","Physics-Informed Machine Learning","Quantum Computing"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Physics-informed regularisation on quantum neural networks provides a promising means for solving differential equations on near-term quantum computers.\nHowever, most demonstrations of this technique assume idealised simulated quantum circuits where the respective expectations  are available.\nIn real quantum hardware, such ideal expectations are not accessible and must be averaged over many shots, introducing additional computations, the cost of which  has not been considered in the majority of the preceding studies.\nThe requirements of higher-order derivatives for physics-informed regularisers are especially high in terms of circuit repetitions (shots) compared to lower-order derivatives required for supervised learning.\nWe demonstrate how to construct a global formulation of physics-informed losses especially amenable to solve ordinary differential equations on near-term quantum computers in a shot-efficient manner.\nThe resulting approach can reduce the order of derivatives required to calculate a loss compared to Physics-informed Neural Networks (PINNs). \nIn the case of initial value problems in ordinary differential equations (ODEs) and some partial differential equations (PDEs), our method removes completely the need for higher-order automatic differentiation,\nthus providing an $\\mathcal{O}(N)$ improvement in shot-efficiency, where $N$ is the number of data-encodings of the quantum neural network.\nOur formulation naturally incorporates boundary conditions and physics-informed losses into a single optimisation term.\nNumerical experiments demonstrate favourable empirical performance, in terms of both shot-efficiency and error, on (simulated) quantum circuits compared to existing quantum methodologies.\nWe demonstrate that the relative performance of quantum neural network algorithms in the infinite shot limit does not necessarily correspond to relative performance in the finite shot limit.\nWe hope this works provides insights on how to efficiently design schemes that will reduce the shot requirements and will become the basis for further developing efficient quantum algorithms for the solution of differential equations."},"_bibtex":{"value":"@misc{\nghosh2024a,\ntitle={A Shot-Efficient Differential Equation Integrator using Quantum Neural Networks},\nauthor={Atiyo Ghosh and Gergana V. Velikova and Panagiotis Barkoutsos and Vincent Emanuel Elfving},\nyear={2024},\nurl={https://openreview.net/forum?id=2EamGPuWSc}\n}"},"title":{"value":"A Shot-Efficient Differential Equation Integrator using Quantum Neural Networks"},"pdf":{"value":"/pdf/3b3e754f117a640081bbfa3fabb6e9326ab3dc75.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"ghosh|a_shotefficient_differential_equation_integrator_using_quantum_neural_networks"},"authorids":{"value":["~Atiyo_Ghosh1","~Gergana_V._Velikova1","~Panagiotis_Barkoutsos1","~Vincent_Emanuel_Elfving1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Atiyo Ghosh","Gergana V. Velikova","Panagiotis Barkoutsos","Vincent Emanuel Elfving"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026 Findings"},"abstract":{"value":"3D Gaussian Splatting (3DGS) has achieved state-of-the-art photorealistic rendering, but the representation gap prevents these assets from being physically interactive. Production-grade physics engines do not understand the 3DGS representation, while prior physics-for-3DGS methods are monolithic silos. These prior works are fundamentally limited, demonstrating only object-centric physics in isolated environments, such as on an ideal plane — they are incapable of interacting with complex static collision geometry or heterogeneous assets. We propose a novel framework that, for the first time, bridges this gap by enabling 3DGS assets to participate in scene-level, heterogeneous, multi-solver physical simulations. Our core contribution is a Representation Abstraction Framework that ``translates\" all diverse assets---including 3DGS, virtual meshes, and fluids---into a unified physical particle set. This abstraction is key to enabling complex behaviors, such as the non-rigid deformation of 3DGS assets, within a unified physics pipeline. This particle set, along with the static scene collision boundaries derived from scene capture, is processed within a solver-agnostic physics kernel. The physical results are then mapped back to drive each asset's specific visual reconstruction. This architecture unlocks capabilities impossible with prior art. We demonstrate complex, two-way interactions between deformable 3DGS assets, standard CG assets (like fluids and meshes), and large-scale captured static environments, showcasing realistic coupled phenomena that were previously unattainable."},"_bibtex":{"value":"@inproceedings{\nliu2026scenelevel,\ntitle={Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats},\nauthor={Xiaoyang Liu and Shangzhe Wu and Kai Han},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=uWggfyoHrC}\n}"},"title":{"value":"Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026F/papers/Liu_Scene-Level_Heterogeneous_Physics_Simulation_with_3D_Gaussian_Splats_CVPRF_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"liu|scenelevel_heterogeneous_physics_simulation_with_3d_gaussian_splats"},"authorids":{"value":["~Xiaoyang_Liu5","~Shangzhe_Wu2","~Kai_Han1"]},"authors":{"value":["Xiaoyang Liu","Shangzhe Wu","Kai Han"]}},"tmdate":1790016658525,"pdate":1790016608351,"tcdate":1765223051205,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission42831/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission42831/Authors"],"forum":"uWggfyoHrC","license":"CC BY 4.0","number":42831,"cdate":1765223051205,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission42831/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/Submission42831/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Findings_Flag","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1790016658525,"odate":1790016608351,"domain":"thecvf.com/CVPR/2026/Conference","id":"uWggfyoHrC","version":2},{"content":{"comment":{"value":"> In Section 4.4, the authors observed that Human300 outperforms Human300+MG60, and they hypothesized that it is because “adding MimicGen data dilutes the contribution of other human datasets”. However, there is not support for this hypothesis.\n\nIndeed, our experiments do not currently show strong positive benefits from incorporating MimicGen data during pre-training improves performance. Although MimicGen allows us to generate large quantities of synthetic trajectories at scale, we find that these are mixed quality demonstrations. We have included examples of MimicGen demonstrations in the supplementary materials for reference. Such mixed quality data can yield suboptimal performance when using standard imitation learning algorithms.\n\nOur work does not make a definitive statement on the value of synthetic demonstration data. We believe that the utility of large-scale synthetic data is dependent on the training algorithm and policy architecture, and our work only explores a small set of current policy learning methods. By releasing this dataset, we hope to enable future work that can more effectively leverage large-scale data of mixed quality and further explore how synthetic demonstrations can complement human demonstrations.\n\n> Related work section lacks full discussion on benchmarks for generalist robots outside of simulation\n\nThank you for the suggestions! We’ve added the suggested references to the manuscript.\n\n> There is no discussion of the underlying simulation framework and how it compares to other simulation frameworks: how fast is it? Is GPU rendering supported?\n\nRoboCasa365 is built on top of RoboSuite, which uses the MuJoCo physics engine. While MuJoCo’s core physics computations are CPU-based, we leverage GPU-based rendering. RoboCasa365 simulates at 20 Hz, with the simulation running approximately in real time—slightly faster or slower depending on scene complexity and hardware specifications. Multiple asynchronous environments can be run in parallel, allowing overall throughput to scale with the number of available CPU cores and GPUs. We have added this discussion to Appendix I.1.\n\n> Can you say more about how the MG60 dataset is generated and how diverse it is?\n\nWe use the open-source MimicGen codebase released by RoboCasa (Nasiriany et al., 2024), which supports 24 atomic tasks, and extend it to cover a total of 60 atomic tasks. For each task, we define object-centric subtasks. For example, the CloseBlenderLid task includes two stages: (1) the robot reaches the lid, and (2) the robot places the lid on the counter.\nFor each of the 60 MimicGen-supported tasks, we start with 100 human source demonstrations and generate 10,000 synthetic demonstrations. These 10,000 demonstrations are distributed across 2,500 kitchen scenes. Example videos illustrating the diversity and quality of these demonstrations are provided in the supplementary materials.\n\n> How optimal are the human teleoperated demonstration data?\n\nQuantifying the optimality of teleoperation data is challenging, but we take steps to ensure it is near-optimal. We discard episodes that involve dropped objects, jerky robot motions, excessive grasping attempts, or collisions with the environment. The supplementary material (zip file in the submission) has been updated to include videos of the teleoperation demonstrations for reference."},"title":{"value":"Author Response to Reviewer tQ4v (Part 2/2)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763949794049,"tcdate":1763949794049,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22486/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission22486/Authors"],"forum":"tQJYKwc3n4","number":5,"license":"CC BY 4.0","cdate":1763949794049,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22486/-/Official_Comment"],"mdate":1763949794049,"domain":"ICLR.cc/2026/Conference","replyto":"crT97Q3coV","id":"nfxkJQBF2K","forumContent":{"TLDR":{"value":"RoboCasa365 is a large-scale benchmark of 365 everyday tasks that advances the study and evaluation of generalist robots across diverse environments and data."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Robot Datasets and Benchmarking","Vision-Language-Action Models","Robot Simulation"]},"supplementary_material":{"value":"/attachment/17863d93ce39c19692c5912fe70f8e990fed84b5.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in robot learning have accelerated progress toward generalist robots that can perform everyday tasks in human environments. Yet it remains difficult to gauge how close we are to this vision. The field lacks a reproducible, large-scale benchmark for systematic evaluation. To fill this gap, we present RoboCasa365, a comprehensive simulation benchmark for household mobile manipulation. Built on the RoboCasa platform, RoboCasa365 introduces 365 everyday tasks across 2,500 diverse kitchen environments, with over 600 hours of human demonstration data and over 1600 hours of synthetically generated demonstration data---making it one of the most diverse and large-scale resources for studying generalist policies. RoboCasa365 is designed to support systematic evaluations for different problem settings, including multi-task learning, robot foundation model training, and lifelong learning. We conduct extensive experiments on this benchmark with state-of-the-art methods and analyze the impacts of task diversity, dataset scale, and environment variation on generalization. Our results provide new insights into what factors most strongly affect the performance of generalist robots and inform strategies for future progress in the field."},"_bibtex":{"value":"@inproceedings{\nnasiriany2026robocasa,\ntitle={RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots},\nauthor={Soroush Nasiriany and Sepehr Nasiriany and Abhiram Maddukuri and Yuke Zhu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=tQJYKwc3n4}\n}"},"title":{"value":"RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots"},"pdf":{"value":"/pdf/9f0ef5cd7a44233fc06f8d5bbacfba0cbe04e008.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"nasiriany|robocasa365_a_largescale_simulation_framework_for_training_and_benchmarking_generalist_robots"},"authorids":{"value":["~Soroush_Nasiriany1","~Sepehr_Nasiriany1","~Abhiram_Maddukuri1","~Yuke_Zhu1"]},"authors":{"value":["Soroush Nasiriany","Sepehr Nasiriany","Abhiram Maddukuri","Yuke Zhu"]}},"version":2},{"content":{"summary":{"value":"Authors enhance MeshGraphNets by physics-informed loss function based on finite-volume method. The approach significantly reduces convergence time by up to 33% and improves predictive accuracy by up to 7.4% on dataset generated by authors using OpenFOAM.  It also allows for training with smaller datasets. The authors test these benefits across various data scales and demonstrate that the method is able to create more efficient and physically consistent GNN-based surrogate models."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Please fix the title in pdf or in openreview.\n2. Provide more scenarios where you train and evaluate your approach. It is needed to better show generalization ability.\n3. See W3. Explain what does \"no physics\" mean. Is it optimized with the same SOAP optimizer or it is vanilla MeshGraphNets.\n4. Provide exact values of coefficients a,b,c in equation 5 with the procedure how they were determined. Ideally, do ablation study with different values of these coefficients.\n5. Provide more explanation for this statement: \"Despite these advances, the trade-offs between data-only training and physics-informed training remain poorly characterized\". Now it is a little bit vague.\n6. Authors say \"In conjunction with standard practice, the equations are nondimensionalized to discard the effects of physical units e.g. density and viscosity\". Probably, \"discard\" is wrong word.  In nondimensionalization, the effects of density and viscosity are not discarded. Instead, they are encapsulated within dimensionless numbers like the Reynolds number.\n7. See W5. Please explain more about dataset pruning procedure.\n8. Figure 3 shows comparison at a \"representative timestep of the simulation\". But the title of the picture above sounds like \"Comparison for u at timestep 0 | MSE: 0.000000...\". Authors should provide explanation and fix this issue. \n9. Justify choosing SOAP optimizer (why not to choose any other optimizer?).\n10. Probably a typo on page 8: \"Figure 3 presents both the lowest data loss achieved and the average loss across epochs. \" But there are not such values represented on Figure 3.\n11. See W7."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Authors integrated FVM residuals into the GNN loss function and achieved improved predictive accuracy and more physically consistent simulations.\n2. Convergence time was reduced by up to 33%, it is beneficial in data-limited scenarios.\n3. Less training data is needed.\n4. Avoids numerical errors from data interpolation by using cell-centered discretization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Mostly, the paper feels largely incomplete. \n\n1. Firstly, the title of paper in pdf is different from openreview: \"Effects of soft physics constraints on graph neural network-based fluid mechanics modeling\" vs \"From Numerical Solvers to Graph Surrogates: Physics-Informed Losses for Data-Efficient CFD Modeling\".  Although the meaning is probably the same, it is a bit confusing.\n2. Generalization ability is limited. The entire study focuses on incompressible flow over a cylinder with fixed Reynolds number. The paper acknowledges that \"The largest remaining errors occur in the turbulent wake regions\", which are inherently difficult to model. This raises questions about how well the method would generalize to more complex scenarios. \n3. \"No physics\" baseline definition. While \"no physics\" is clear in principle, the authors need to explain whether the baseline models were optimized with the same SOAP optimizer and learning rate scheduler. If it is not true, some of the performance gains attributed to physics could be partially influenced by optimization techniques.\n4. The paper states that the loss function is \"simply a linear weighting\" and defines the coefficients a, b, c (equation 5). However, there's no discussion on how these weights were determined (and their values are not mentioned), if they were tuned, or their sensitivity to different flow conditions / dataset sizes. The choice of these weights can significantly impact the balance between data fidelity and physical consistency.\n5. We don't know exactly how dataset pruning was done. Do authors drop samples randomly? Or do they drop some simulations based on some criteria? This is not stated on paper.\n6. Strong justification of choosing SOAP optimizer is needed.\n7. Provide explanation about why there are not direct comparisons with some previous studies mentioned in \"Related work\" section."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943197802,"tcdate":1761558016601,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24787/Reviewer_4MZT"],"signatures":["ICLR.cc/2026/Conference/Submission24787/Reviewer_4MZT"],"forum":"iBLHGdBImw","number":1,"license":"CC BY 4.0","cdate":1761558016601,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24787/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943197802,"domain":"ICLR.cc/2026/Conference","replyto":"iBLHGdBImw","id":"dPqkEqaSEB","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics Informed Neural Networks; Graph Neural Networks; Fluid Dynamics Surrogate"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Graph neural networks (GNN) represent a promising method for creating robust and physically interpretable surrogate models for fluid dynamics. These surrogates offer a significant advantage over traditional computational fluid dynamics (CFD) solvers based on numerical methods because they require much less computational cost. In a GNN designed as a surrogate model for spatio-temporal partial differential equations, message passing can be interpreted as the propagation of physical quantities such as velocity, pressure, and temperature. The complexity of the Navier-Stokes equations, however, can limit the generalizability of existing models and lead to long training times.\nWe show that including a physics-informed loss function based on the numerical methods used to generate the training data, specifically the finite volume method, can reduce the amount of data needed to train an accurate physics-informed surrogate compared with a purely data-driven baseline. By reducing the dataset size by 20\\% and applying this approach, we achieved a 33\\% reduction in convergence time. For larger datasets, model accuracy improved by up to 7.4\\% within the same timeframe. Our method also avoids interpolation between cell centers and vertices, which can introduce errors from numerical discretization. Applying this soft constraint during training can support the development of future CFD surrogate GNN models that perform well even with smaller datasets."},"_bibtex":{"value":"@misc{\ngarcia2026from,\ntitle={From Numerical Solvers to Graph Surrogates: Physics-Informed Losses for Data-Efficient {CFD} Modeling},\nauthor={Michael Lawrence Garcia and Chen Zhuang and Junya Onishi and Peng Chen and Makoto Tsubokura and Mohamed Wahib},\nyear={2026},\nurl={https://openreview.net/forum?id=iBLHGdBImw}\n}"},"title":{"value":"From Numerical Solvers to Graph Surrogates: Physics-Informed Losses for Data-Efficient CFD Modeling"},"pdf":{"value":"/pdf/d9030d18b4c981fb65fb7231aa580cea92efbdd6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"garcia|from_numerical_solvers_to_graph_surrogates_physicsinformed_losses_for_dataefficient_cfd_modeling"},"authorids":{"value":["~Michael_Lawrence_Garcia1","~Chen_Zhuang2","~Junya_Onishi1","~Peng_Chen16","~Makoto_Tsubokura1","~Mohamed_Wahib1"]},"authors":{"value":["Michael Lawrence Garcia","Chen Zhuang","Junya Onishi","Peng Chen","Makoto Tsubokura","Mohamed Wahib"]}},"version":2},{"content":{"summary":{"value":"This paper presents TDDPM, a model for generating high-fidelity, privacy-preserving trajectory data in complex environments. TDDPM leverages a denoising diffusion approach to deaggregate spatial data into individual trajectories, allowing realistic time-series generation that generalizes to unseen areas. By conditioning on spatial aggregates, it achieves strong out-of-distribution performance. The model also uses k-anonymity for privacy and introduces a new benchmark for evaluating synthetic trajectory data, making it a valuable tool for urban planning and autonomous driving applications."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"See weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"(1) The demonstration of proposed method TDDPM is detailed. \n(2) The authors describe the motivation and background of generating out-of-distribution trajectories in detail."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The contributions of TDDPM are unclear. The authors should clearly claim the contributions in the end of introduction.\n2. In Table 1 and Table 2, standard deviation of KL and JS divergence are not reported.\n3. This paper lacks an ablation study part. Experiments should be added to verify the effectiveness of proposed two steps mentioned in Section 4.\n4. The experiment setting is confusing. The authors claim that TDDPM could achieve out-of-distribution generalization in Abstract. More analysis should be added to demonstrate the diffirence between the synthetic dataset and Geolife/Porto.\n5. Line 308 to line 325 seems making no sense."}},"nonreaders":[],"tmdate":1732531945651,"tcdate":1730564718906,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7080/Reviewer_jUbV"],"signatures":["ICLR.cc/2025/Conference/Submission7080/Reviewer_jUbV"],"forum":"dDdxbdhMsY","number":2,"license":"CC BY 4.0","cdate":1730564718906,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7080/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732531945651,"domain":"ICLR.cc/2025/Conference","replyto":"dDdxbdhMsY","id":"dFHkryqFrn","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["generative learning","mobility data","denoising diffusion probablistic models"]},"supplementary_material":{"value":"/attachment/7b3532070ca43ee34262c5b334e276c2a21cf6db.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Access to spatio-temporal trajectory data is essential for improving infrastructure, preventing the spread of disease and for building autonomous vehicles. However, it remains underutilized due to limited availability, as it cannot be shared publicly due privacy concerns or other sensitive attributes. Generative time-series models have shown promise in generating non-sensitive data, but show poor performance for large-scale and complex environments. In this paper we propose a spatio-temporal generative model for trajectories, TDDPM, which outperforms and scales substantially better than state-of-the-art. The focus is primarily on trajectories of peoples' movement in cities. We propose a conditional distribution approach which unlock out-of-distribution generalization, such as to city-areas not trained on, from a spatial aggregate prior. We also show that data can be generated in a privacy-preserving manner using $k$-anonymity. Further, we propose a new comprehensive benchmark across several standard datasets, and evaluation measures, considering key distribution properties."},"_bibtex":{"value":"@misc{\nbergstrom2025deep,\ntitle={Deep Temporal Deaggregation: Large-Scale Spatio-Temporal Generative Models},\nauthor={David Bergstr{\\\"o}m and Mattias Tiger and Fredrik Heintz},\nyear={2025},\nurl={https://openreview.net/forum?id=dDdxbdhMsY}\n}"},"title":{"value":"Deep Temporal Deaggregation: Large-Scale Spatio-Temporal Generative Models"},"pdf":{"value":"/pdf/bf78c32713f182284f3fcbfe23f9dad8b71fb7aa.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"bergström|deep_temporal_deaggregation_largescale_spatiotemporal_generative_models"},"authorids":{"value":["~David_Bergström2","~Mattias_Tiger1","~Fredrik_Heintz1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["David Bergström","Mattias Tiger","Fredrik Heintz"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Information-preserving Graph Neural Simulators (IGNS), a graph-based neural simulator that improves modeling of complex physical systems. IGNS enforces Hamiltonian dynamics to preserve long-range interactions and extends to non-conservative systems. It includes warmup initialization, geometric encoding, and multi-step training to enhance stability. Evaluated on new benchmarks with long-range dependencies and external forces, IGNS outperforms state-of-the-art methods in accuracy and robustness for dynamic systems."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In L201-203 and in Appendix D, the paper states that $\\gamma_\\theta(t)$ and $\\tau_\\theta(t)$ are time-varying coefficient vectors produced by MLPs with parameters θ. Could the authors clarify what exactly is used as the input t to these MLPs? Is t normalized to a fixed range (e.g., [0, 1]) or directly represented as the raw timestep index? Additionally, if the model is trained on trajectories with 400 steps, can the learned time-dependent MLPs generalize to longer rollouts (e.g., 1000 steps) without retraining, or does the model rely on an absolute temporal scale?\n2. L231-233 states: \"Thanks to the energy conserving core of IGNS, this globally informed latent state is preserved throughout the rollout, rather than being dissipated.\" Could the authors clarify why this happens? Can you provide qualitative or quantitative analysis showing how the latent state is preserved over time？\n3. Regarding the multi-step loss: how many time steps were included in the loss computation? Are the reported results based on single-step MSE or on the rollout of the entire sequence? Please include these experiment details.\n4. Why does WaveBall not require a warmup phase?\n5.  In the supplementary code (`igns.py`), L651: `x = self.one_step(x, edge_index, edge_weight, batch, t=i)` passes `i` as `t`. Here, `i` corresponds to the layer index rather than the time step. Can the authors explain why this is done?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The proposed Information-preserving Graph Neural Simulator introduces a principled integration of port-Hamiltonian dynamics into graph-based simulators, marking a significant step beyond existing message-passing and oscillatory GNN frameworks.\n- Theoretical analyses are thorough and provide a clear justification for the model’s ability to capture complex and long-range physical interactions. \n- Experimental evaluation is comprehensive, spanning six datasets and consistently demonstrating the superior accuracy and stability of IGNS compared to strong baselines.\n- The paper is clearly written and well structured: the motivation for each component (port-Hamiltonian core, warmup phase, geometric encoding, and multi-step loss) is clearly articulated, and the accompanying figures effectively convey both the methodology and the empirical findings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although the theoretical analysis establishes information preservation and universality, it remains largely qualitative in linking these properties to the observed empirical improvements. A more quantitative or ablation-based verification (e.g., measuring gradient norms or energy conservation over rollouts) would provide stronger evidence for the theoretical claims.\n- The training/testing computational overhead of the port-Hamiltonian formulation and the warmup phase is not explicitly analyzed; reporting runtime or memory costs relative to standard GNSs would clarify the practical trade-offs. \n- Although the benchmarks are diverse, most tasks are synthetic or controlled simulations. It would strengthen the paper’s significance to include or discuss applications in more realistic or large-scale physical systems.\n- The geometric encoding used to map edges to features follows the same formulation as previous works (e.g., MGN) and therefore cannot be considered a novel contribution.\n- Several related approaches are not cited or compared [1–4], which limits the contextual positioning of this work within recent advances in graph-based physical simulation. \n- Introducing \"warmup phase\" in GNNs is not novel. Eagle [2] employs a warmup-like phase in its encoder, using multiple message-passing blocks to aggregate local and global context before rollout, which parallels the proposed initialization strategy.\n- The separation of state variables into coordinates and momenta, as well as the coordinate–momentum supervision in Eq. (10), are established techniques already used in [2–3].\n\n[1] EvoMesh: Adaptive Physical Simulation with Hierarchical Graph Evolutions. ICML 2025\n\n[2] Eagle: Large-Scale Learning of Turbulent Fluid Dynamics with Mesh Transformers. ICLR 2023\n\n[3] Efficient Learning of Mesh-Based Physical Simulation with BSMS-GNN. ICML 2023\n\n[4] Physics meets Topology: Physics-informed topological neural networks for learning rigid body dynamics"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920758315,"tcdate":1760847700775,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9040/Reviewer_S8R7"],"signatures":["ICLR.cc/2026/Conference/Submission9040/Reviewer_S8R7"],"forum":"x66u6TEDUw","number":1,"license":"CC BY 4.0","cdate":1760847700775,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9040/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920758315,"domain":"ICLR.cc/2026/Conference","replyto":"x66u6TEDUw","id":"HcyQAmUy8g","forumContent":{"TLDR":{"value":"We propose a novel Graph Neural Simulator that preserves information during propagation, enabling it to model complex physical dynamical systems with long-range dependencies."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Graph Neural Simulators","Long-range interactions","Learning Simulators","AI4Science"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Learning to simulate complex physical systems from data has emerged as a promising way to overcome the limitations of traditional numerical solvers, which often require prohibitive computational costs for high-fidelity solutions. Recent Graph Neural Simulators (GNSs) accelerate simulations by learning dynamics on graph-structured data, yet often struggle to capture long-range interactions and suffer from error accumulation under autoregressive rollouts. To address these challenges, we propose Information-preserving Graph Neural Simulators (IGNS), a graph-based neural simulator built on the principles of Hamiltonian dynamics. This structure guarantees preservation of information across the graph, while extending to port-Hamiltonian systems allows the model to capture a broader class of dynamics, including non-conservative effects. IGNS further incorporates a warmup phase to initialize global context, geometric encoding to handle irregular meshes, and a multi-step training objective that facilitates PDE matching, where the trajectory produced by integrating the port-Hamiltonian core aligns with the ground-truth trajectory, thereby reducing rollout error. To evaluate these properties systematically, we introduce new benchmarks that target long-range dependencies and challenging external forcing scenarios. Across all tasks, IGNS consistently outperforms state-of-the-art GNSs, achieving higher accuracy and stability under challenging and complex dynamical systems. Our project page: https://thobotics.github.io/neural_pde_matching."},"_bibtex":{"value":"@inproceedings{\nhoang2026improving,\ntitle={Improving Long-Range Interactions in Graph Neural Simulators via Hamiltonian Dynamics},\nauthor={Tai Hoang and Alessandro Trenta and Alessio Gravina and Niklas Freymuth and Philipp Becker and Davide Bacciu and Gerhard Neumann},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=x66u6TEDUw}\n}"},"title":{"value":"Improving Long-Range Interactions in Graph Neural Simulators via Hamiltonian Dynamics"},"pdf":{"value":"/pdf/3e3b5f42da70e6b5dba62ba0faf5208abf2511fc.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"hoang|improving_longrange_interactions_in_graph_neural_simulators_via_hamiltonian_dynamics"},"authorids":{"value":["~Tai_Hoang1","~Alessandro_Trenta2","~Alessio_Gravina1","~Niklas_Freymuth1","~Philipp_Becker1","~Davide_Bacciu1","~Gerhard_Neumann2"]},"authors":{"value":["Tai Hoang","Alessandro Trenta","Alessio Gravina","Niklas Freymuth","Philipp Becker","Davide Bacciu","Gerhard Neumann"]}},"version":2},{"content":{"TLDR":{"value":"We introduce PLAID: a flexible framework for representing and sharing datasets of physics simulations, six carefully crafted datasets in PLAID (covering structural mechanics and computational fluid dynamics) and corresponding baseline benchmarks."},"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Physics Learning","Data Model","Standardization","Benchmark","MeshGraphNet","MMGP","Vision Transformer","DAFNO","MARIO"]},"primary_area":{"value":"AL/ML_data_and_benchmarks_for_physics"},"abstract":{"value":"Machine learning-based surrogate models have emerged as a powerful tool to accelerate simulation-driven scientific workflows. However, their widespread adoption is hindered by the lack of large-scale, diverse, and standardized datasets tailored to physics-based simulations. While existing initiatives provide valuable contributions, many are limited in scope-focusing on specific physics domains, relying on fragmented tooling, or adhering to overly simplistic datamodels that restrict generalization. To address these limitations, we introduce PLAID (Physics-Learning AI Datamodel), a flexible and extensible framework for representing and sharing datasets of physics simulations. PLAID defines a unified standard for describing simulation data and is accompanied by a library for creating, reading, and manipulating complex datasets across a wide range of physical use cases~(\\href{https://gitlab.com/drti/plaid}{gitlab.com/drti/plaid}). We release six carefully crafted datasets under the PLAID standard, covering structural mechanics and computational fluid dynamics, and provide baseline benchmarks using representative learning methods. Benchmarking tools are made available on Hugging Face, enabling direct participation by the community and contribution to ongoing evaluation efforts (\\href{https://huggingface.co/PLAIDcompetitions}{huggingface.co/PLAIDcompetitions})"},"_bibtex":{"value":"@misc{\ncasenave2026physicslearning,\ntitle={Physics-Learning {AI} Datamodel ({PLAID}) datasets: a collection of physics simulations for machine learning},\nauthor={Fabien Casenave and Xavier Roynard and Brian Staber and William PIAT and Michele Alessandro Bucci and Nissrine Akkari and Abbas Kabalan and Xuan Minh Vuong Nguyen and Luca Saverio and Raphael Carpintero Perez and Anthony Kalaydjian and Samy Fouch{\\'e} and Thierry Gonon and Ghassan Najjar and Emmanuel Menier and Matthieu Nastorg and Giovanni Catalani and Christian Rey},\nyear={2026},\nurl={https://openreview.net/forum?id=0vryNakJYC}\n}"},"title":{"value":"Physics-Learning AI Datamodel (PLAID) datasets: a collection of physics simulations for machine learning"},"pdf":{"value":"/pdf/86189b4af8fbb726d63b826a0ff12a7ad4fa36ed.pdf"},"croissant_file":{"value":"/attachment/f30c9d969b80f29bfae50ea26fe845fa8d043723.zip"},"code_URL":{"value":"https://gitlab.com/drti/plaid"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"paperhash":{"value":"casenave|physicslearning_ai_datamodel_plaid_datasets_a_collection_of_physics_simulations_for_machine_learning"},"authorids":{"value":["~Fabien_Casenave1","~Xavier_Roynard1","~Brian_Staber1","~William_PIAT1","~Michele_Alessandro_Bucci2","~Nissrine_Akkari1","~Abbas_Kabalan1","~Xuan_Minh_Vuong_Nguyen1","~Luca_Saverio1","~Raphael_Carpintero_Perez2","~Anthony_Kalaydjian1","~Samy_Fouché1","~Thierry_Gonon1","~Ghassan_Najjar1","~Emmanuel_Menier1","~Matthieu_Nastorg1","~Giovanni_Catalani1","~Christian_Rey2"]},"dataset_URL":{"value":"https://huggingface.co/collections/PLAID-datasets/plaid-datasets-6674504b8a183bfe993a85d8"},"authors":{"value":["Fabien Casenave","Xavier Roynard","Brian Staber","William PIAT","Michele Alessandro Bucci","Nissrine Akkari","Abbas Kabalan","Xuan Minh Vuong Nguyen","Luca Saverio","Raphael Carpintero Perez","Anthony Kalaydjian","Samy Fouché","Thierry Gonon","Ghassan Najjar","Emmanuel Menier","Matthieu Nastorg","Giovanni Catalani","Christian Rey"]}},"tmdate":1776956124592,"tcdate":1746473339437,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission547/Authors"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission547/Authors"],"forum":"0vryNakJYC","license":"CC BY-SA 4.0","number":547,"cdate":1746473339437,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Submission","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Post_Submission","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission547/-/Full_Submission","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1776956124592,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","id":"0vryNakJYC","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2505.02974v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"casenave|physicslearning_ai_datamodel_plaid_datasets_a_collection_of_physics_simulations_for_machine_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Fabien_Casenave:","https://dblp.org/search/pid/api?q=author:Xavier_Roynard:","https://dblp.org/search/pid/api?q=author:Brian_Staber:","https://dblp.org/search/pid/api?q=author:William_Piat:","https://dblp.org/search/pid/api?q=author:Michele_Alessandro_Bucci:","https://dblp.org/search/pid/api?q=author:Nissrine_Akkari:","https://dblp.org/search/pid/api?q=author:Abbas_Kabalan:","https://dblp.org/search/pid/api?q=author:Xuan_Minh_Vuong_Nguyen:","https://dblp.org/search/pid/api?q=author:Luca_Saverio:","https://dblp.org/search/pid/api?q=author:Raphaël_Carpintero_Perez:","https://dblp.org/search/pid/api?q=author:Anthony_Kalaydjian:","https://dblp.org/search/pid/api?q=author:Samy_Fouché:","https://dblp.org/search/pid/api?q=author:Thierry_Gonon:","https://dblp.org/search/pid/api?q=author:Ghassan_Najjar:","https://dblp.org/search/pid/api?q=author:Emmanuel_Menier:","https://dblp.org/search/pid/api?q=author:Matthieu_Nastorg:","~Giovanni_Catalani1","https://dblp.org/search/pid/api?q=author:Christian_Rey:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2505.02974"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2505-02974,\n  publtype={informal},\n  author={Fabien Casenave and Xavier Roynard and Brian Staber and William Piat and Michele Alessandro Bucci and Nissrine Akkari and Abbas Kabalan and Xuan Minh Vuong Nguyen and Luca Saverio and Raphaël Carpintero Perez and Anthony Kalaydjian and Samy Fouché and Thierry Gonon and Ghassan Najjar and Emmanuel Menier and Matthieu Nastorg and Giovanni Catalani and Christian Rey},\n  title={Physics-Learning AI Datamodel (PLAID) datasets: a collection of physics simulations for machine learning},\n  year={2025},\n  month={May},\n  cdate={1746057600000},\n  journal={CoRR},\n  volume={abs/2505.02974},\n  url={https://doi.org/10.48550/arXiv.2505.02974}\n}\n"},"abstract":{"value":"Machine learning-based surrogate models have emerged as a powerful tool to accelerate simulation-driven scientific workflows. However, their widespread adoption is hindered by the lack of large-scale, diverse, and standardized datasets tailored to physics-based simulations. While existing initiatives provide valuable contributions, many are limited in scope-focusing on specific physics domains, relying on fragmented tooling, or adhering to overly simplistic datamodels that restrict generalization. To address these limitations, we introduce PLAID (Physics-Learning AI Datamodel), a flexible and extensible framework for representing and sharing datasets of physics simulations. PLAID defines a unified standard for describing simulation data and is accompanied by a library for creating, reading, and manipulating complex datasets across a wide range of physical use cases (gitlab.com/drti/plaid). We release six carefully crafted datasets under the PLAID standard, covering structural mechanics and computational fluid dynamics, and provide baseline benchmarks using representative learning methods. Benchmarking tools are made available on Hugging Face, enabling direct participation by the community and contribution to ongoing evaluation efforts (huggingface.co/PLAIDcompetitions)."},"title":{"value":"Physics-Learning AI Datamodel (PLAID) datasets: a collection of physics simulations for machine learning"},"authors":{"value":["Fabien Casenave","Xavier Roynard","Brian Staber","William Piat","Michele Alessandro Bucci","Nissrine Akkari","Abbas Kabalan","Xuan Minh Vuong Nguyen","Luca Saverio","Raphaël Carpintero Perez","Anthony Kalaydjian","Samy Fouché","Thierry Gonon","Ghassan Najjar","Emmanuel Menier","Matthieu Nastorg","Giovanni Catalani","Christian Rey"]}},"tmdate":1756285847897,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2505-02974"],"tcdate":1756285842869,"writers":["~"],"signatures":["~Giovanni_Catalani1"],"forum":"yW1YI3f7hO","license":"CC BY-SA 4.0","number":618575,"cdate":1746057600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1756285847897,"domain":"DBLP.org","id":"yW1YI3f7hO","version":2},{"content":{"venue":{"value":"IEEE Trans. Neural Networks Learn. Syst. 2013"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/5962385/6632919/06573411.pdf"},"venueid":{"value":"dblp.org/journals/TNN/2013"},"paperhash":{"value":"tengtrairat|singlechannel_blind_separation_using_pseudostereo_mixture_and_complex_2d_histogram"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Naruephorn_Tengtrairat:","https://dblp.org/search/pid/api?q=author:Bin_Gao_0003:","~Wai_Lok_Woo1","https://dblp.org/search/pid/api?q=author:Satnam_Singh_Dlay:"]},"html":{"value":"https://doi.org/10.1109/TNNLS.2013.2258680"},"_bibtex":{"value":"@article{DBLP:journals/tnn/TengtrairatGWD13,\n  author={Naruephorn Tengtrairat and Bin Gao and Wai Lok Woo and Satnam Singh Dlay},\n  title={Single-Channel Blind Separation Using Pseudo-Stereo Mixture and Complex 2-D Histogram},\n  year={2013},\n  cdate={1356998400000},\n  journal={IEEE Trans. Neural Networks Learn. Syst.},\n  volume={24},\n  number={11},\n  pages={1722-1735},\n  url={https://doi.org/10.1109/TNNLS.2013.2258680}\n}\n"},"abstract":{"value":"A novel single-channel blind source separation (SCBSS) algorithm is presented. The proposed algorithm yields at least three benefits of the SCBSS solution: 1) resemblance of a stereo signal concept given by one microphone; 2) independent of initialization and a priori knowledge of the sources; and 3) it does not require iterative optimization. The separation process consists of two steps: 1) estimation of source characteristics, where the source signals are modeled by the autoregressive process and 2) construction of masks using only the single-channel mixture. A new pseudo-stereo mixture is formulated by weighting and time-shifting the original single-channel mixture. This creates an artificial mixing system whose parameters will be estimated through our proposed weighted complex 2-D histogram. In this paper, we derive the separability of the proposed mixture model. Conditions required for unique mask construction based on maximum likelihood are also identified. Finally, experimental testing on both synthetic and real-audio sources is conducted to verify that the proposed algorithm yields superior performance and is computationally very fast compared with existing methods."},"title":{"value":"Single-Channel Blind Separation Using Pseudo-Stereo Mixture and Complex 2-D Histogram"},"authors":{"value":["Naruephorn Tengtrairat","Bin Gao","Wai Lok Woo","Satnam Singh Dlay"]}},"tmdate":1753446139837,"pdate":1356998400000,"tcdate":1753446058622,"writers":["~"],"signatures":["~Wai_Lok_Woo1"],"forum":"2QMxXsM21x","license":"CC BY-SA 4.0","number":588618,"cdate":1356998400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1753446139837,"domain":"DBLP.org","id":"2QMxXsM21x","version":2},{"content":{"summary":{"value":"The authors propose a method using soft physics-based regularization to predict the distribution of brain tumor. The main contribution is a novel discretization scheme of physics equations to model brain tumor."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Is there any bias in the evaluation metrics since there are no ground truth labels and authors use recurrence coverage?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This work is an improvement for physics-based approaches for medical imaging \n\n2. The proposed method of discretizing the physics residuals is novel and potentially applicable to related problems and of interest to the research community \n\n3. The paper is clearly written and the technical details are sound and sufficient"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The main weakness is the lack of ground truth data to train and evaluate the model performance. The evaluation is performed not on the original task of brain tumor detection but rather on a downstream task of “recurrence coverage”, which measures the percentage of the tumor detected in follow-up MRIs (rather than the original). \n\n2. Related to above, the authors only use one metric to evaluate their model, which may be insufficient, especially since there is no ground truth labels"},"limitations":{"value":"1. The authors should discuss the limitations of their methods in more details, especially related to the lack of ground truth labels to train the model"}},"nonreaders":[],"tmdate":1730879913023,"tcdate":1720761571444,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission17469/Reviewer_ZDBA"],"signatures":["NeurIPS.cc/2024/Conference/Submission17469/Reviewer_ZDBA"],"forum":"YfVMcbcDqo","number":2,"license":"CC BY 4.0","cdate":1720761571444,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission17469/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879913023,"domain":"NeurIPS.cc/2024/Conference","replyto":"YfVMcbcDqo","id":"FUR8SCLuHT","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"We propose a physics-regularized learning approach on dynamic discrete meshes to address the complex inverse problem of tumor localization."},"keywords":{"value":["Inverse Problems","System Identification","Physics-Informed","Biomechanical Modeling","Tumor Growth"]},"primary_area":{"value":"machine_learning_for_healthcare"},"flagged_for_ethics_review":{"value":true},"abstract":{"value":"Physical models in the form of partial differential equations serve as important priors for many under-constrained problems. One such application is tumor treatment planning, which relies on accurately estimating the spatial distribution of tumor cells within a patient’s anatomy. While medical imaging can detect the bulk of a tumor, it cannot capture the full extent of its spread, as low-concentration tumor cells often remain undetectable, particularly in glioblastoma, the most common primary brain tumor. Machine learning approaches struggle to estimate the complete tumor cell distribution due to a lack of appropriate training data. Consequently, most existing methods rely on physics-based simulations to generate anatomically and physiologically plausible estimations. However, these approaches face challenges with complex and unknown initial conditions and are constrained by overly rigid physical models. In this work, we introduce a novel method that integrates data-driven and physics-based cost functions, akin to Physics-Informed Neural Networks (PINNs). However, our approach parametrizes the solution directly on a dynamic discrete mesh, allowing for the effective modeling of complex biomechanical behaviors. Specifically, we propose a unique discretization scheme that quantifies how well the learned spatiotemporal distributions of tumor and brain tissues adhere to their respective growth and elasticity equations. This quantification acts as a regularization term, offering greater flexibility and improved integration of patient data compared to existing models. We demonstrate enhanced coverage of tumor recurrence areas using real-world data from a patient cohort, highlighting the potential of our method to improve model-driven treatment planning for glioblastoma in clinical practice."},"_bibtex":{"value":"@inproceedings{\nbalcerak2024physicsregularized,\ntitle={Physics-Regularized Multi-Modal Image Assimilation for Brain Tumor Localization},\nauthor={Michal Balcerak and Tamaz Amiranashvili and Andreas Wagner and Jonas Weidner and Petr Karnakov and Johannes C. Paetzold and Ivan Ezhov and Petros Koumoutsakos and Benedikt Wiestler and bjoern menze},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=YfVMcbcDqo}\n}"},"title":{"value":"Physics-Regularized Multi-Modal Image Assimilation for Brain Tumor Localization"},"pdf":{"value":"/pdf/d652352e6b706e6110a4a40f48dfc2eec542165a.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"balcerak|physicsregularized_multimodal_image_assimilation_for_brain_tumor_localization"},"authorids":{"value":["~Michal_Balcerak1","~Tamaz_Amiranashvili1","~Andreas_Wagner1","~Jonas_Weidner1","~Petr_Karnakov1","~Johannes_C._Paetzold1","~Ivan_Ezhov1","~Petros_Koumoutsakos1","~Benedikt_Wiestler1","~bjoern_menze1"]},"authors":{"value":["Michal Balcerak","Tamaz Amiranashvili","Andreas Wagner","Jonas Weidner","Petr Karnakov","Johannes C. Paetzold","Ivan Ezhov","Petros Koumoutsakos","Benedikt Wiestler","bjoern menze"]}},"version":2},{"content":{"summary":{"value":"This paper investigates a critical, often-overlooked assumption in mechanistic interpretability: that causal interventions (like activation patching) produce latent representations that are ”faithful” or ”in-distribution” for the\ntarget model. The authors argue that interventions can create ”divergent,” out-of-distribution states, which may\nundermine the validity of the resulting explanations.\nThe paper’s contributions are threefold:\n\n- It empirically demonstrates that common intervention methods, including Mean Difference Vector Patching, Sparse Autoencoders (SAEs), and Distributed Alignment Search (DAS), do create representations that\ndiverge significantly from the model’s natural distribution.\n\n-  It provides a theoretical framework for classifying these divergences into ”harmless” (e.g., in the null-space\nor within existing decision boundaries) and ”pernicious” (e.g., activating hidden, non-native computational\npathways or causing dormant behavioral changes). The authors use synthetic examples to show how pernicious divergences can lead to misleadingly ”affirming” results.\n\n- To mitigate this issue, the authors propose an adapted Counterfactual Latent (CL) loss (based on Grant\n(2025)). This loss regularizes interventions by encouraging the intervened latent state to remain close to an\naverage of native states that share the same intended causal properties . In a synthetic setting, they show this\nmethod reduces divergence and improves out-of-distribution (OOD) generalization."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"-  The CL loss mitigation seems promising but raises practical questions. How do you\nenvision applying this to large language models where the ground-truth causal variables (needed to find\nmatching native states for $h_{CL}$) are not known a priori? Doesn’t this create a circular dependency where\nyou need the causal abstraction to find the CL vectors, but the intervention (which you’re trying to fix) is\nwhat you use to find the abstraction?\n\n- You note that the CL loss is a ”broad-stroke” solution. Do you\nhave any initial thoughts on how one might distinguish pernicious from harmless divergence, perhaps without a full causal abstraction? For instance, could one use a measure of ”local faithfulness” or analyze\nactivation changes in all output dimensions (as hinted at in 4.2.3 ) to flag interventions that recruit ”hidden”\npathways?\n\n- The concept of ”dormant behavioral changes” (Section 4.2.3, Appendix A.2) is particularly concerning. You mention that an ”infinitely expansive” dataset could detect them. In practice, how\ncould a researcher gain any confidence that their intervention hasn’t created such a dormant vulnerability?\nDoes the CL loss’s reduction in OOD error (Fig 3d) suggest it is mitigating these, or is that a separate issue?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"-  By questioning the\n”faithfulness” of intervened states, the work highlights a critical potential failure mode for many mechanistic\ninterpretability claims.\n•  The theoretical categorization of divergences into ”harmless” (e.g., nullspace, within-boundary covariance) and ”pernicious” (e.g., hidden pathways, dormant behavior) is clear,\nintuitive, and  valuable.\n-  The synthetic examples in Section 4.2 are particularly strong. They offer\nconcrete, simple-to-understand illustrations of how an intervention can be misleading; for instance, by\nbreaking a ”balanced subspace” or activating a pathway via a mean-difference vector that is not used by\nany native data point . The concept of ”dormant behavioral changes” is also a very insightful and worrying\nfailure mode.\n-  It also proposes a concrete (if\npreliminary) mitigation strategy. The adaptation of the CL loss to ”anchor” interventions to the native data\nmanifold is a good approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper successfully categorizes divergences but does not offer a method\nto detect or classify whether a divergence observed in a practical (non-synthetic) setting is harmless or\npernicious. This makes it difficult to know when to be concerned about the results of their\ninterventions.\n-  The primary weakness of the proposed solution, acknowledged by the authors,\nis that the CL loss is a ”broad-stroke” approach. It penalizes all divergence, rather than selectively targeting only ”pernicious” divergence. This might be overly restrictive, as some ”harmless” divergences (e.g.,\ninterventions that explore the null-space) could be desirable for making certain causal claims.\n\n- The empirical validation for the CL loss mitigation is confined to a synthetic, small-scale MLP setting. It is unclear how this approach would scale to modern,\nlarge-scale models. For example, generating the ”Counterfactual Latent” (CL) vectors requires averaging\nover a pre-recorded set of native states with specific causal properties. This seems computationally challenging and, more importantly, dependent on having a correct, known causal abstraction, which is often\nwhat is being searched for in the first place."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942498800,"tcdate":1762050682269,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23064/Reviewer_YbiP"],"signatures":["ICLR.cc/2026/Conference/Submission23064/Reviewer_YbiP"],"forum":"cZrTMqYVL6","number":2,"license":"CC BY 4.0","cdate":1762050682269,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23064/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942498800,"domain":"ICLR.cc/2026/Conference","replyto":"cZrTMqYVL6","id":"sWSio5GcE7","forumContent":{"TLDR":{"value":"We show empirical representational divergence between native and causally intervened latent states, we show that this can be pernicious and propose a solution."},"venue":{"value":"ICLR 2026 Oral"},"keywords":{"value":["activation patching","mech interp","DAS","representational divergence","faithfulness"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"A common approach to mechanistic interpretability is to causally manipulate model representations via targeted interventions in order to understand what those representations encode. Here we ask whether such interventions create out-of-distribution (divergent) representations, and whether this raises concerns about how faithful their resulting explanations are to the target model in its natural state. First, we demonstrate theoretically and empirically that common causal intervention techniques often do shift internal representations away from the natural distribution of the target model. Then, we provide a theoretical analysis of two cases of such divergences: \"harmless\" divergences that occur in the behavioral null-space of the layer(s) of interest, and \"pernicious\" divergences that activate hidden network pathways and cause dormant behavioral changes. Finally, in an effort to mitigate the pernicious cases, we apply and modify the Counterfactual Latent (CL) loss from Grant (2025) allowing representations from causal interventions to remain closer to the natural distribution, reducing the likelihood of harmful divergences while preserving the interpretive power of the interventions. Together, these results highlight a path towards more reliable interpretability methods."},"_bibtex":{"value":"@inproceedings{\ngrant2026addressing,\ntitle={Addressing divergent representations from causal interventions on neural networks},\nauthor={Satchel Grant and Simon Jerome Han and Alexa R. Tartaglini and Christopher Potts},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=cZrTMqYVL6}\n}"},"title":{"value":"Addressing divergent representations from causal interventions on neural networks"},"pdf":{"value":"/pdf/f608b52d1638ab06c9c0f72be594048e574c9484.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"grant|addressing_divergent_representations_from_causal_interventions_on_neural_networks"},"authorids":{"value":["~Satchel_Grant1","~Simon_Jerome_Han1","~Alexa_R._Tartaglini1","~Christopher_Potts1"]},"authors":{"value":["Satchel Grant","Simon Jerome Han","Alexa R. Tartaglini","Christopher Potts"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Gen-LRA, a novel membership inference attack (MIA) methodology for evaluating privacy risks in synthetic tabular data. The authors propose a hypothesis testing framework that computes a likelihood ratio specifically targeted at identifying any local overfitting of the target record. The method requires minimal assumptions, just access to the released synthetic dataset and a reference dataset. They find their method to outperform baselines from the literature across 15 datasets. They further find their method to be particularly successful against outliers, in contrast with other MIAs from the literature."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Can you expand the related work to also include the shadow-modeling based MIAs? \n\n- To truly understand the contribution, could you implement the shadow-modeling based MIAs [1,2,3] as well and report their results? Right now, the Gen-LRA method seems to be better than the prior work you consider, and does so with limited assumptions for the attacker and with limited computational cost. How does this change when  the attacker now (i) has knowledge of the training algorithm and (ii) has the computational resources to train shadow models? Could authors implement these shadow-model MIAs and report the results alongside Gen-LRA? This would help to position the method and its results in the literature, giving a clear understanding of the impact of certain assumptions and computational cost on the MIA results. \n\n- Similarly, the work on shadow modeling MIAs also discusses disparate vulnerability of outliers [1,2,3]. Stadler et al [1] finds outliers to be more vulnerable than randomly selected records, while Meeus et al [3] proposes a method to identify more vulnerable records. Could authors have more elaborate results for the outlier discussion (e.g. show MIA results for outliers vs random points across datasets) and relate these findings to prior work? While the fact that Gen-LRA focuses on outliers is distinct from distance-based methods, these findings might not be very different than the ones in shadow-modeling based MIAs."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Technically novel, and interesting, way to compute the membership inference inference signal from synthetic data. The method is theoretically grounded, computationally efficient and relies on limited assumptions for the attacker. \n- They show the method to outperform a range of MIAs from the literature\n- Comprehensive evaluation of the attack across 15 datasets\n- Authors include intuitive examples (eg Fig 1 and Sec 6.2) that are well explained and help the understanding of the paper."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(More details see questions)\n\n- My main concern comes down to a lack of related work being discussed. A range of important works have studied MIAs against synthetic tabular data using shadow modeling [1,2,3]. While I understand that these works are computationally more expensive and additionally rely on the attacker's knowledge of the training algorithm, I find these works to be very relevant to position this paper and its findings. \n- Limited secondary insights with experimental depth. For instance, to make the claim that the method works better for outliers (especially compared to other methods), section 5.3 is mostly anecdotal. \n\n[1] Stadler, T., Oprisanu, B., & Troncoso, C. (2022). Synthetic data–anonymisation groundhog day. In 31st USENIX Security Symposium (USENIX Security 22) (pp. 1451-1468).\n\n[2] Houssiau, F., Jordon, J., Cohen, S. N., Daniel, O., Elliott, A., Geddes, J., ... & Szpruch, L. TAPAS: a Toolbox for Adversarial Privacy Auditing of Synthetic Data. In NeurIPS 2022 Workshop on Synthetic Data for Empowering ML Research.\n\n[3] Meeus, M., Guepin, F., Creţu, A. M., & de Montjoye, Y. A. (2023, September). Achilles’ heels: vulnerable record identification in synthetic data publishing. In European Symposium on Research in Computer Security (pp. 380-399). Cham: Springer Nature Switzerland."}},"nonreaders":[],"tmdate":1731429674314,"tcdate":1730286651189,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12851/Reviewer_6TYk"],"signatures":["ICLR.cc/2025/Conference/Submission12851/Reviewer_6TYk"],"forum":"02DCEU6vSU","number":2,"license":"CC BY 4.0","cdate":1730286651189,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12851/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429674314,"domain":"ICLR.cc/2025/Conference","replyto":"02DCEU6vSU","id":"IDT940ZREW","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Privacy","Membership Inference Attacks","Generative Models"]},"supplementary_material":{"value":"/attachment/50c96fb68049a4bec3f129b7c7f85b812793218e.pdf"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Evaluating the potential privacy leakage of synthetic data is an important but unresolved problem. Most existing adversarial auditing frameworks for synthetic data rely on heuristics and unreasonable assumptions to attack the failure modes of generative models, exhibiting limited capability to describe and detect the privacy exposure of training data. In this paper, we study designing Membership Inference Attacks (MIAs) that specifically exploit the observation that generative models tend to memorize certain data points in their training sets, leading to significant local overfitting. Here, we propose Generative Likelihood Ratio Attack (Gen-LRA), a novel, computationally efficient shadow-box MIA that, with no assumption of model knowledge or access, attacks the generated synthetic dataset by conducting a hypothesis test that it is locally overfit to potential training data. Assessed over a comprehensive benchmark spanning diverse datasets, model architectures, and attack parameters, we find that Gen-LRA consistently dominates other MIAs for generative models across multiple performance metrics. These results underscore Gen-LRA's effectiveness as an interpretable and robust privacy auditing tool, highlighting the significant privacy risks posed by generative model overfitting in real-world applications"},"_bibtex":{"value":"@misc{\nward2024genlra,\ntitle={Gen-{LRA}: Towards a Principled Membership Inference Attack for Generative Models},\nauthor={Joshua Ward and Chi-Hua Wang and Guang Cheng},\nyear={2024},\nurl={https://openreview.net/forum?id=02DCEU6vSU}\n}"},"title":{"value":"Gen-LRA: Towards a Principled Membership Inference Attack for Generative Models"},"pdf":{"value":"/pdf/bcad18f87958725e9b50970906e168913dcdf521.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"ward|genlra_towards_a_principled_membership_inference_attack_for_generative_models"},"authorids":{"value":["~Joshua_Ward1","~Chi-Hua_Wang1","~Guang_Cheng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Joshua Ward","Chi-Hua Wang","Guang Cheng"]}},"version":2},{"content":{"summary":{"value":"The paper studies how models learn when trained on mixtures of real and synthetic data. It assumes that the underlying knowledge follows a long-tailed distribution, with synthetic data representing a truncated portion of the real distribution. Under this setting, the authors show that training progresses through three distinct phases, corresponding to how the model successively acquires head and tail knowledge.\n\nBuilding on this perspective, the paper derives a generalisation bound that describes learning behaviour under real–synthetic mixtures. Using the theoretical results, the authors propose a practical, retraining-free data valuation method that estimates the relative importance of different data subsets.\n\nExperiments on long-tailed benchmarks provide partial empirical support for the theoretical analysis and show that the proposed valuation method can identify informative and high-impact data across multiple tasks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. **About Lemma 1 and the loss definition:**  \n   In Lemma 1, it appears that $\\mathcal{L}_{\\text{test}}$ should correspond to the *error rate* (as implied by the derivation of lemma1), rather than an arbitrary loss function. However, in both the appendix and the main text, the derivation seems to treat it as a general loss. Could the authors clarify which interpretation is correct, and whether the asymptotic equivalence used in the lemma depends on this choice?\n\n2. **Assumption on $\\lambda_{\\min}(\\Theta_0) > 0$ in Theorem 1:**  \n   While $\\Theta_0$ is symmetric (hence has real eigenvalues), it is only strictly positive definite if the gradient feature vectors of all samples are linearly independent. In practice, however, when some data points are identical or highly similar, these gradients can become correlated, making $\\Theta_0$ singular or ill-conditioned. Could the authors discuss how realistic this assumption is for practical networks and whether any regularisation (e.g., adding $\\lambda I$) was used to ensure stability or invertibility in their implementation?\n\n3. **Fitting of $w_1, w_2, w_3, w_4$:**  \n   In the experiments, the weights are learned via linear regression using the average of the empirical loss and the MMD score as the target. Given that this target is directly related to the first two components of the value function, do the fitted results make $w_3$ and $w_4$ nearly zero? Could the authors clarify the rationale for this design choice and whether this regression target effectively serves as a baseline, since it is already part of the objective?\n\n4. **Concept of data value and dependency among contributors:**  \n   The paper defines data value based on the generalisation bound but does not explicitly account for *dependencies among contributors* or for what knowledge the model has already learned. Is this aspect ignored because the bound is not tight, or is there an implicit assumption of independence between subsets?\n\nIn summary, while the paper is clear and conceptually interesting, the justification of key assumptions and the breadth of empirical validation are limited. These issues make the work less convincing as a practical contribution, though it remains theoretically insightful.\nIf the authors can convincingly address these weaknesses and questions, I believe the paper would be strong enough for acceptance."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The paper is clearly written and presents a well-structured, well-defined problem formulation.\n\n* The problem formulation is interesting and original, offering a novel way to analyse how real and synthetic data interact during training.\n\n* It tackles a timely and relevant problem, contributing to our understanding of synthetic data in large-scale model development.\n\n* The theoretical analysis is intuitive and well interpreted, giving clear insights into the learning dynamics under long-tailed knowledge distributions.\n\n* The paper includes some empirical evidence on real-world datasets that provides partial support for the theoretical claims and the proposed data-valuation method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The paper does not provide sufficient analysis to justify the central hypothesis that knowledge in real–synthetic mixtures follows a long-tailed distribution and that synthetic data represent a truncated version of it. In real-world settings, the distribution of knowledge can vary substantially across modalities and tasks (e.g. datasets such as CIFAR-100 are relatively balanced). Moreover, model-based synthetic data generation methods often suffer from hallucination or misalignment with real data, and their outputs can be highly sensitive to the training distribution or the prompt used for generation. It would also be valuable to examine, either empirically or theoretically, how sensitive the proposed theoretical results are to deviations from this idealised assumption. For instance, if the synthetic distribution exhibits different levels of noise or if the real data distribution departs from a pure long-tail form. A deeper empirical and theoretical investigation of this assumption would significantly strengthen the paper’s applicability and credibility.\n\n* The analysis of the proposed data-valuation method is limited. The score involves four components with associated weights, yet there is no ablation or sensitivity study showing how the valuation results change when these weights are adjusted or when specific terms are removed. It remains unclear which components are most critical for stable performance.\n\n* The NTK-based term is computed at random initialization, which may exhibit large variance across random seeds or architectures. The paper does not discuss whether this variability could affect the reliability or reproducibility of the valuation results.\n\n\n* The work focuses mainly on data-attribution–style approaches and overlooks recent data-weighting frameworks that address similar goals. In particular, methods such as [1] and [2] also aim to evaluate or adjust the contribution of data mixtures (at domain or instance level). Interestingly, the first two terms of the proposed score relate to distribution alignment, and the third term addresses redundancy, concepts that also appear in these weighting-based methods. A deeper comparison or discussion would clarify the novelty and positioning of this work.\n\n[1] Xie et al., DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining. NeurIPS 2023\n\n[2] Kuo et al., Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification. ICLR 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923938704,"tcdate":1760832678328,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13262/Reviewer_KWsB"],"signatures":["ICLR.cc/2026/Conference/Submission13262/Reviewer_KWsB"],"forum":"3iXyRG2nzT","number":1,"license":"CC BY 4.0","cdate":1760832678328,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13262/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923938704,"domain":"ICLR.cc/2026/Conference","replyto":"3iXyRG2nzT","id":"x832OwpRZC","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Data Valuation","LLM","Scaling Dynamics"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"The rapid progress of large language models (LLMs) is fueled by the growing reliance on datasets that blend real and synthetic data. While synthetic data offers scalability and cost-efficiency, it often introduces systematic distributional discrepancies, particularly underrepresenting long-tail knowledge due to truncation effects from data generation mechanisms like top-$p$ sampling, temperature scaling, and finite sampling. These discrepancies pose fundamental challenges in characterizing and evaluating the utility of mixed real-synthetic datasets. In this paper, we identify a three-phase scaling behavior characterized by two breakpoints that reflect transitions in model behavior across learning head and tail knowledge. We further derive an LLM generalization bound designed for real and synthetic mixtures, revealing several key factors that govern their generalization performance. Building on our theoretical findings, we propose an effective yet efficient data valuation method that scales to large-scale datasets. Comprehensive experiments across four tasks, including image classification, sentiment classification, instruction following, and complex reasoning, demonstrate that our method surpasses state-of-the-art baselines in data valuation with significantly low computational cost."},"_bibtex":{"value":"@misc{\nwang2025data,\ntitle={Data Value in the Age of Scaling: Understanding {LLM} Scaling Dynamics Under Real{\\textendash}Synthetic Data Mixtures},\nauthor={Haohui Wang and Jingyuan Qi and Jianpeng Chen and Jun Wu and Lifu Huang and Lecheng Zheng and Kevin Choi and Balaji Veeramani and Edward Bowen and Alison Hu and Tyler Cody and Dawei Zhou},\nyear={2025},\nurl={https://openreview.net/forum?id=3iXyRG2nzT}\n}"},"title":{"value":"Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real–Synthetic Data Mixtures"},"pdf":{"value":"/pdf/5aa1e4438f7169afa72c9a2942a6b8153822c667.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|data_value_in_the_age_of_scaling_understanding_llm_scaling_dynamics_under_realsynthetic_data_mixtures"},"authorids":{"value":["~Haohui_Wang1","~Jingyuan_Qi1","~Jianpeng_Chen1","~Jun_Wu3","~Lifu_Huang1","~Lecheng_Zheng1","~Kevin_Choi1","~Balaji_Veeramani1","~Edward_Bowen1","~Alison_Hu1","~Tyler_Cody1","~Dawei_Zhou1"]},"authors":{"value":["Haohui Wang","Jingyuan Qi","Jianpeng Chen","Jun Wu","Lifu Huang","Lecheng Zheng","Kevin Choi","Balaji Veeramani","Edward Bowen","Alison Hu","Tyler Cody","Dawei Zhou"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"li|freegave_3d_physics_learning_from_dynamic_videos_by_gaussian_velocity"},"authorids":{"value":["~Jinxi_Li1","https://dblp.org/search/pid/api?q=author:Ziyang_Song:","https://dblp.org/search/pid/api?q=author:Siyuan_Zhou:","https://dblp.org/search/pid/api?q=author:Bo_Yang_0027:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2506.07865"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2506-07865,\n  publtype={informal},\n  author={Jinxi Li and Ziyang Song and Siyuan Zhou and Bo Yang},\n  title={FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity},\n  year={2025},\n  month={June},\n  cdate={1748736000000},\n  journal={CoRR},\n  volume={abs/2506.07865},\n  url={https://doi.org/10.48550/arXiv.2506.07865}\n}\n"},"abstract":{"value":"In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing works often fail to learn complex physical motions at boundaries or require object priors such as masks or types. In this paper, we propose FreeGave to learn the physics of complex dynamic 3D scenes without needing any object priors. The key to our approach is to introduce a physics code followed by a carefully designed divergence-free module for estimating a per-Gaussian velocity field, without relying on the inefficient PINN losses. Extensive experiments on three public datasets and a newly collected challenging real-world dataset demonstrate the superior performance of our method for future frame extrapolation and motion segmentation. Most notably, our investigation into the learned physics codes reveals that they truly learn meaningful 3D physical motion patterns in the absence of any human labels in training."},"title":{"value":"FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity"},"authors":{"value":["Jinxi Li","Ziyang Song","Siyuan Zhou","Bo Yang"]}},"tmdate":1762939991373,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2506-07865"],"tcdate":1762935949460,"writers":["~"],"signatures":["~Jinxi_Li2"],"forum":"PJD5CIkD6a","license":"CC BY-SA 4.0","number":691136,"cdate":1748736000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762939991373,"domain":"DBLP.org","id":"PJD5CIkD6a","version":2},{"content":{"summary":{"value":"The paper investigates the relationship between language model size and its capacity to store factual knowledge, quantified in bits per parameter. The authors introduce a framework that measures a model’s knowledge based on tuple-based information (e.g., (Entity, Relation, Attribute)) and propose that language models, after sufficient training, achieve an approximate capacity of 2 bits per parameter. The study extends this analysis across various factors such as model architecture, quantization levels, sparsity, and training data quality. Using synthetic and controlled datasets, the authors examine how these elements influence the knowledge storage capacity of language models.\n\nThe technical claims are generally supported by experiments; however, there are concerns about the robustness of the findings. The reliance on synthetic data raises questions about the applicability of the results to real-world scenarios. Additionally, the paper does not thoroughly explore the impact of quantization during training.\n\nThe paper has several formatting issues that hinder its readability. Notably, it lacks a conclusion section. Some figures are not clear. The organization of the paper could be improved drastically."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1.\tHow would the proposed 2-bit/parameter scaling law hold up when applied to language models trained on real-world, diverse datasets?\n\n2.\tCould incorporating quantization into the training process mitigate the reduction in capacity observed with int4 quantization?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"•\tOriginality: Introducing a framework to measure language model capacity in bits per parameter is a novel approach that adds a quantitative dimension to model evaluation.\n\n•\tMethodology: The use of controlled synthetic datasets allows for the isolation of specific variables, providing clarity in the analysis of different factors affecting knowledge capacity."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"•\tFormatting Issues: The absence of a conclusion section and unclear figures detract from the overall quality of the paper and impede the reader’s understanding.\n\n•\tGeneralization to Real-world Data: The heavy reliance on synthetic data limits the applicability of the findings to natural language processing tasks involving complex and diverse datasets.\n\n•\tIncomplete Exploration of Quantization: The paper does not investigate quantization during training, which could provide insights into mitigating the observed decrease in capacity with int4 quantization.\n\n•\tLimited Architectural Diversity: The study does not explore a wide range of model architectures, such as encoder-only or decoder-only models, which could affect the generalizability of the proposed scaling law."}},"nonreaders":[],"tmdate":1731429304987,"tcdate":1730696001925,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13207/Reviewer_Z8yV"],"signatures":["ICLR.cc/2025/Conference/Submission13207/Reviewer_Z8yV"],"forum":"FxNNiUgtfa","number":4,"license":"CC BY 4.0","cdate":1730696001925,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13207/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429304987,"domain":"ICLR.cc/2025/Conference","replyto":"FxNNiUgtfa","id":"sjf86z98gr","forumContent":{"venue":{"value":"ICLR 2025 Spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["scaling laws","knowledge capacity","language models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Scaling laws describe the relationship between the size of language models and their capabilities. Unlike prior studies that evaluate a model's capability via loss or benchmarks, we estimate information-theoretically the number of knowledge \\emph{bits} a model stores. We focus on factual knowledge represented as tuples, such as (USA, capital, Washington D.C.) from a Wikipedia page. Through multiple controlled datasets, we establish that language models can and only can store \\emph{2 bits of knowledge per parameter, even when quantized to int8}, and such knowledge can be flexibly extracted for downstream applications. \n\nMore broadly, we present 12 results on how (1) training duration, (2) model architecture, (3) quantization, (4) sparsity constraints such as MoE, and (5) data signal-to-noise ratio affect a model's knowledge storage capacity."},"_bibtex":{"value":"@inproceedings{\nallen-zhu2025physics,\ntitle={Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws},\nauthor={Zeyuan Allen-Zhu and Yuanzhi Li},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=FxNNiUgtfa}\n}"},"title":{"value":"Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws"},"pdf":{"value":"/pdf/7a9985d2cec78e1746e3dc81372bed15021471b0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"allenzhu|physics_of_language_models_part_33_knowledge_capacity_scaling_laws"},"authorids":{"value":["~Zeyuan_Allen-Zhu1","~Yuanzhi_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyuan Allen-Zhu","Yuanzhi Li"]}},"version":2},{"content":{"summary":{"value":"The manuscript provides training guarantees for physics-informed neural networks in an overparametrized setting with NTK-scale parametrization, both for first-order gradient descent as well as a second-order natural gradient descent. For gradient descent, the allowed step size of $O(1/\\lambda_\\max)$ is different compared to prior works with $O(\\lambda_\\min)$ step size on optimization in PINNs. For natural gradient descent, a convergence rate independent of the condition number of the problem is derived. Additionally, computational experiments comparing first-order to second-order optimizers in PINNs are provided."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"+ Can you elaborate why you regard your $O(1/\\lambda_\\max)$ as an improvement over $O(\\lambda_\\min)$? \n+ Can you elaborate on the relation of the natural gradient defined in Section 4 to previously proposed methods like energy natural gradient? In case of equivalence, can you adapt your presentation accordingly, see also my discussion above. \n+ Can you comment why you make the overparametrization assumption? Also, a discussion of the overparametrization assumption inside the manuscript seems reasonable. \n+ Can you comment on the different condition numbers of supervised learning and PINNs? Can you elaborate why you only expect a bad condition number for complex PDEs (see line 251)? This could be added in the manuscript as a motivation for second order and natural gradient methods. \n+ Can you validate the theoretical convergence guarantees within the computational experiments? In particular: \n   + Compare the loss curves to the linear convergence guarantees from Theorem 3.7 and 4.7. \n   + Compare the generalization error for different network size and optimizers. As the convergence guarantees are merely given with respect to the loss function, it is informative to contrast it with the relative L^2 error. \n   + Improve readability of the plot. \n\n**Minor comments:**\n1. In the abstract, it is stated that, *However, the learning rate of GD for training two-layer neural networks exhibits poor dependence on the sample size and the Gram matrix*. Note that the term learning rate usually refers to the step size in machine learning, not the convergence rate of the learning process. \n2. It is stated in several places that the learning rate can be O(1). However, it is shown that the convergence rate is O(1). \n3. In the limitations section on scalability of natural gradient methods in these scenarios, the following directly relevant works are not referenced: \n+ Kronecker-Factored Approximate Curvature for Physics-Informed Neural Networks, Dangel et al. 2024\n+ Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization, Dangel et al. 2025"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"+ Overall, the paper is well written and easy to navigate and the main contributions are clearly described. \n+ The paper addresses a timely and important question of the convergence of optimization physics-informed models. \n+ The guarantees go beyond the use of first-order methods and treat natural gradient, which seems to be the arguably most efficient optimizer in PINNs. \n+ The theoretical analysis uses timely methods from supervised learning and rigorously transfers them to PINNs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"+ Discussion of related works: \n    + Use of *our NGD*: At multiple places in the manuscript, *our NGD* is used to refer to the natural gradient method that is being analyzed. Also, in Remark 4.2 it is claimed that the method studied is different from previously proposed natural gradient or Gauss-Newton methods for PINNs. The iteration of energy natural gradient (Müller and Zeinhofer, 2023), which reduces to a Gauss-Newton iteration for PINNs, is given by $$ w(k+1)=w(k)-\\eta (J(k)^T J(k))^+J(k)^T \\binom{s(k)}{h(k)}, $$ where $A^+$ denotes an arbitrary pseudo-inverse of $A$. By choosing the Moore-Penrose inverse and considering an overparametrized setting in which $J(k)J(k)^T$ has full rank, this specializes to $$w(k+1)=w(k)-\\eta J(k)^T (J(k)J(k)^T)^{-1} \\binom{s(k)}{h(k)},$$ which agrees with the iteration studied in the manuscript. Hence, from my understanding, the manuscript is studying the convergence properties of the optimizer proposed by Müller and Zeinhofer (2023), however, this is not mentioned at any place. Rather, the impression is conveyed that the manuscript proposed a new variant of a natural gradient method, see Remark 4.2 as well as line 408 and Table 1. This \n   + There are certain works on optimization in PINNs missing in the discussion of related works in Subsection 1.2., e.g., \n      + AN OPERATOR PRECONDITIONING PERSPECTIVE ON TRAINING IN PHYSICS-INFORMED MACHINE LEARNING, Tim De Ryck, Florent Bonnet, Siddhartha Mishra, Emmanuel de Bézenac; in particular, Theorem 2.3 shows that for the linearization problem with preconditioning achieves a converge rate given by the condition number and that the natural gradient preconditioning achieves an optimal condition number of 1. As such, this seems to be a linearized (hence easier) version of the main result in the manuscript. \n      + Convergence of Stochastic Gradient Methods for Wide Two-Layer Physics-Informed Neural Networks, BANGTI JIN AND LONGJUN WU, 2025; provide an exponential convergence rate of SGD under overparametrization\n      + Non-Asymptotic Analysis of Projected Gradient Descent for Physics-Informed Neural Networks, JONAS NIESEN, AND JOHANNES MÜLLER, 2025; provide a sublinear convergence guarantee for arbitrary sampled (S)GD without assumption on the network size \n+ Overparametrization assumption: The manuscript is making the global assumption of overparametrization. Where I understand that this allows the use of an established machinery, I do not believe that this is the setting that PINNs are used in. In particular, note that in PINNs, the data points are synthetically sampled integration points of the computational domain rather than data points like in supervised learning. As such, the problem in PINNs is rather an optimization than a statistical problem. In practice, new data points are sampled continuously throughout optimization. Hence, it is not clear how practically relevant the setting of overparametrization is. However, this assumption is not uncommon and the only work I am aware not making an assumption on the size of the PINN is by Niessen and Müller (2025). \n+ Difference of PINNs to supervised learning: In the introduction, it is mentioned that the reason why NGD is not used in supervised learning is due to its high computational cost. However, I believe that there another, arguably more important reason: The loss function of PINNs has a much worse conditionining due to the appearence of the PDE operator. It is stated in line 251 that the conditioning can be really bad for complex PDEs. From my understanding, this can also be the case for simple PDEs. \n+ Experiments: The experiments compare natural gradient to SGD, Adam and L-BFGS. However, this comparison has been made at several places, in particular, by Müller and Zeinhofer (2023). I think, it would be much more informative to not repeat this comparison, but to see, how the empirically observed convergence rates relate to the theoretical guarnatees. Further, in relation to the overparametrization assumption, the influence on the network size and optimizer on the generalization error is not studied empirically."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359726063,"tcdate":1761571368368,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10662/Reviewer_jYGf"],"signatures":["ICLR.cc/2026/Conference/Submission10662/Reviewer_jYGf"],"forum":"KWWfLgkySm","number":2,"license":"CC BY 4.0","cdate":1761571368368,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10662/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359726063,"domain":"ICLR.cc/2026/Conference","replyto":"KWWfLgkySm","id":"BqtfxNIsiR","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["natural gradient descent","over-parameterization","physics-informed neural networks","neural tangent kernel"]},"supplementary_material":{"value":"/attachment/217bfd79d386772fa22f1c50cfd58f0cb5117391.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"In the context of over-parameterization, there is a line of work demonstrating that randomly initialized (stochastic) gradient descent (GD) converges to a globally optimal solution at a linear convergence rate for the quadratic loss function. However, the convergence rate of GD for training two-layer neural networks exhibits poor dependence on the sample size and the Gram matrix, leading to a slow training process. In this paper, we show that for training two-layer $\\text{ReLU}^3$ Physics-Informed Neural Networks (PINNs), the learning rate can be improved from the smallest eigenvalue of the limiting Gram matrix to the reciprocal of the largest eigenvalue, implying that GD actually enjoys a faster convergence rate. Despite such improvements, the convergence rate is still tied to the least eigenvalue of the Gram matrix, leading to slow convergence. We then develop the positive definiteness of Gram matrices with general smooth activation functions and provide the convergence analysis of natural gradient descent (NGD) in training two-layer PINNs, demonstrating that the maximal learning rate can be $\\mathcal{O}(1)$ and at this rate, the convergence rate is independent of the Gram matrix. In particular, for smooth activation functions, the convergence rate of NGD is quadratic. Numerical experiments are conducted to verify our theoretical results."},"_bibtex":{"value":"@inproceedings{\nxu2026fast,\ntitle={Fast Convergence of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks},\nauthor={Xianliang Xu and Wang Kong and JiahengM and Zhongyi Huang and Ye Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KWWfLgkySm}\n}"},"title":{"value":"Fast Convergence of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks"},"pdf":{"value":"/pdf/68c7c9ed57158d4d6105099fef3519ce7167320b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"xu|fast_convergence_of_natural_gradient_descent_for_overparameterized_physicsinformed_neural_networks"},"authorids":{"value":["~Xianliang_Xu1","~Wang_Kong1","~JiahengM1","~Zhongyi_Huang2","~Ye_Li6"]},"authors":{"value":["Xianliang Xu","Wang Kong","JiahengM","Zhongyi Huang","Ye Li"]}},"version":2},{"content":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["representation learning","physics identification","orthogonality"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Accurately identifying the underlying physical laws in complex systems is vital for effective control and interpretation. However, many systems are governed by a combination of known physical principles and unobservable or poorly understood components. Traditional model-based methods like Kalman filters and state-space models often rely on oversimplified assumptions, while modern data-driven approaches, such as physics-informed neural networks (PINNs), can suffer from overfitting or lack theoretical guarantees in recovering true physical dynamics. We propose the Orthogonal Deep Neural Network (ODNN) architecture to address these limitations. ODNN disentangles known physical components from unobservable or poorly understood components by imposing orthogonal constraints on the deep neural network. Unlike additive regularization methods, ODNN converts the physical constraints directly into the network structure, ensuring that the DNN focuses on capturing the unknown or complex dynamics without overfitting. This novel approach leverages both explicit orthogonality (e.g., zero inner product) and implicit orthogonality (e.g., contrasting convexity, periodicity, or symmetry) between physical laws and unknown components. Theoretically, we prove that ODNN provides strong guarantees for accurate system identification under mild orthogonality assumptions, building on the universal approximation theorem. Empirically, ODNN is evaluated across eight synthetic and real-world datasets, showcasing its ability to recover governing physical equations with high accuracy and interpretability. Our results demonstrate that ODNN offers significant advantages in terms of generalizability and robustness, making it a valuable framework for physics-based model identification in complex systems."},"_bibtex":{"value":"@misc{\nxiao2025orthogonal,\ntitle={Orthogonal Deep Neural Networks ({ODNN}): Uncovering Hidden Physics in Partially Observable Systems},\nauthor={CHENHAN XIAO and Yang Weng},\nyear={2025},\nurl={https://openreview.net/forum?id=ZujMVRn7Md}\n}"},"title":{"value":"Orthogonal Deep Neural Networks (ODNN): Uncovering Hidden Physics in Partially Observable Systems"},"pdf":{"value":"/pdf/a0db04dd97c2cd93dbed6db975b85ec6e11eff6b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"xiao|orthogonal_deep_neural_networks_odnn_uncovering_hidden_physics_in_partially_observable_systems"},"authorids":{"value":["~CHENHAN_XIAO1","~Yang_Weng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["CHENHAN XIAO","Yang Weng"]}},"tmdate":1738190682962,"tcdate":1727233768907,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4095/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission4095/Authors"],"forum":"ZujMVRn7Md","license":"CC BY 4.0","number":4095,"cdate":1727233768907,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/Submission4095/-/Full_Submission","ICLR.cc/2025/Conference/Submission4095/-/Rebuttal_Revision","ICLR.cc/2025/Conference/-/Edit","ICLR.cc/2025/Conference/-/Withdrawn_Submission"],"mdate":1738190682962,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"ZujMVRn7Md","version":2},{"content":{"summary":{"value":"This paper introduces Squirrel Benchmark, a synthetic benchmark for evaluating LLMs on enterprise-scale SQL debugging rather than generation. The authors construct large, complex SQL queries via LLM-based synthesis, inject bugs using a taxonomy derived from enterprise logs, and evaluate ~30 LLMs. They also compare SFT-based methods and an agentic debugging loop. Results show state-of-the-art models still struggle on SQL correction tasks, particularly semantic debugging."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Q1. How can we validate that LLM-synthesized “enterprise SQL” truly reflects ETL pipelines, cross-db joins, dbt lineage trees, UDFs, or vendor dialects?\n\nQ2. What percentage of bugs in the taxonomy were seen in actual production logs vs synthetic augmentation?\n\nQ3. Given that the benchmark is not publicly available, can you release examples of real enterprise-authored SQL and bug metadata (even if anonymized / masked) in addition to the info described in Section 4?\n\nQ4. How stable is the benchmark to using different seed models, or is it biased toward the Claude generator?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"S1. The paper addresses an emerging and important problem, using LLMs for iterative SQL debugging. It is not only relevant to enterprise text-to-SQL but also important as a standalone problem.\n\nS2. The newly created Squirrel benchmark is significantly longer and more complex than Spider/BIRD, with >140 lines and richer nesting, aiming to reflect real ETL workloads.\n\nS3. The evaluation demonstrates that even top LLMs (Claude, GPT-5, Qwen3) have difficulty, particularly on semantic bugs, reinforcing the difficulty of enterprise SQL debugging.\n\nS4. The work describes a systematic reverse-engineering-style pipeline for bug injection and adversarial filtering."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1. The primary contribution is a benchmark, not new methods. Most components—synthetic SQL generation, taxonomy-guided corruption, execution-free matching—are incremental extensions of existing ideas (e.g., BIRD-Critic reverse debugging, Spider2.0 industrial focus). Core technical innovation is modest.\n\nW2. Despite cross-checking with some enterprise logs, the majority of data is still LLM-generated, raising concerns about over-alignment with prompting style of the generator model, divergence from messy, schema-sprawling real SQL found in enterprises.\n\nW3. The work largely evaluates the performance of out-of-the-box LLMs. Additional baselines (e.g., NL2SQL, SQL error detect/correction) should be included in the evaluation.\n\nW4. The graph-match and modify-better metrics may favor structurally close but semantically incorrect SQL, especially in complex and large query DAGs. This undermines the enterprise debugging claim unless validated carefully with execution traces. Also what \n\nW6. The agent & SFT experiments are weak. While results are promising, but the paper does not provide enough details to fully understand the pros and cons of these two methodologies in the context of SQL debugging. Related to W3, certain NL2SQL solutions, such as DIN-SQL, SQLens, Reforce, Arctic-Text2SQL-R1, adopted agent framework or SFT/PT, which can be better candidates to answer the two questions(SFT/agent methods solving the SQL debugging)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915640046,"tcdate":1762210469536,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission916/Reviewer_z7M7"],"signatures":["ICLR.cc/2026/Conference/Submission916/Reviewer_z7M7"],"forum":"8Fm6OKFuRv","number":3,"license":"CC BY 4.0","cdate":1762210469536,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission916/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915640046,"domain":"ICLR.cc/2026/Conference","replyto":"8Fm6OKFuRv","id":"R3DthSdiS1","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["SWE","LLM","benchmark","SQL","Bugfixing","Agent"]},"supplementary_material":{"value":"/attachment/789847c44546dd17125942d333f31e123bb6ff93.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"SQL is the core of data analysis and engineering across industries, powering large-scale workflows for data extraction, transformation, and loading. However, in enterprise-level scenarios, it is challenging to generate fully correct SQL code in a single attempt—even for experienced developers or advanced \\ttsql LLMs. Multiple iterations of debugging are usually required, yet LLMs often get lost in multi-turn correction.\nTo address this gap, we introduce \\ourbench, the first benchmark designed for enterprise-level SQL reasoning and debugging. Our benchmark is built upon two key innovations: (1) an automated construction workflow that employs reverse engineering to systematically inject realistic bugs into large-scale SQL code, enabling scalable and diverse benchmark generation; and (2) an \\textbf{execution-free evaluation framework} tailored for enterprise settings, providing fast, accurate, and resource-efficient assessment.\n\\ourbench comprises 469 \\ourbenchsyn queries featuring syntax errors with explicit error messages, and 516 \\ourbenchsem queries targeting semantic errors where SQL fails to meet the requirement. The SQLs are substantially complex, averaging over 140 lines with abstract syntax trees of high complexity (average width >11, depth >8.7).\nWe evaluate nearly 30 LLMs on \\ourbench. Even state-of-the-art reasoning models struggle: Claude-4-Sonnet achieves only 36.46\\% success on \\ourbenchsyn and 32.17\\% on \\ourbenchsem. Most models fail to reach 20\\% success, underscoring the significant gap between current LLM capabilities and the demands of enterprise SQL debugging.\nTo bridge this gap, we systematically \\textbf{explore four potential solution strategies and conduct extensive experiments} to evaluate and compare their effectiveness.  Our experiments not only highlight the challenges but also identify effective strategies for advancing SQL debugging with LLMs."},"_bibtex":{"value":"@misc{\nye2026beyond,\ntitle={Beyond Text-to-{SQL}: Can {LLM}s Really Debug Enterprise {SQL}?},\nauthor={Jing Ye and Yiwen Duan and Yonghong Yu and Victor Ma and Gaoyang and Xing Chen},\nyear={2026},\nurl={https://openreview.net/forum?id=8Fm6OKFuRv}\n}"},"title":{"value":"Beyond Text-to-SQL: Can LLMs Really Debug Enterprise SQL?"},"pdf":{"value":"/pdf/23eda7f901edf103a08cab6d4bb87cb164773da8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"ye|beyond_texttosql_can_llms_really_debug_enterprise_sql"},"authorids":{"value":["~Jing_Ye3","~Yiwen_Duan1","~Yonghong_Yu2","~Victor_Ma1","~Gaoyang1","~Xing_Chen7"]},"authors":{"value":["Jing Ye","Yiwen Duan","Yonghong Yu","Victor Ma","Gaoyang","Xing Chen"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (2) 2018"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-05414-4_9.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2018"},"paperhash":{"value":"dugué|bringing_a_feature_selection_metric_from_machine_learning_to_complex_networks"},"authorids":{"value":["~Nicolas_Dugué1","~Jean-Charles_Lamirel1","https://dblp.org/search/pid/api?q=author:Anthony_Perez_0001:"]},"html":{"value":"https://doi.org/10.1007/978-3-030-05414-4_9"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/DugueL018,\n  author={Nicolas Dugué and Jean-Charles Lamirel and Anthony Perez},\n  title={Bringing a Feature Selection Metric from Machine Learning to Complex Networks},\n  year={2018},\n  cdate={1514764800000},\n  pages={107-118},\n  url={https://doi.org/10.1007/978-3-030-05414-4_9},\n  booktitle={COMPLEX NETWORKS (2)},\n  crossref={conf/complexnetworks/2018-2}\n}\n"},"abstract":{"value":"Introduced in the context of machine learning, the Feature F-measure is a statistical feature selection metric without parameters that allows to describe classes through a set of salient features. It was shown efficient for classification, cluster labeling and clustering model quality measurement. In this paper, we introduce the Node F-measure, its transposition in the context of networks, where it can by analogy be applied to detect salient nodes in communities. This approach benefits from the parameter-free system of Feature F-Measure, its low computational complexity and its well-evaluated performance. Interestingly, we show that in addition to these properties, Node F-measure is correlated with certain centrality measures, and with measures designed to characterize the community roles of nodes. We also observe that the usual community roles measures are strongly dependent from the size of the communities whereas the ones we propose are by definition linked to the density of the community. This hence makes their results comparable from one network to another. Finally, the parameter-free selection process applied to nodes allows for a universal system, contrary to the thresholds previously defined empirically for the establishment of community roles. These results may have applications regarding leadership in scientific communities or when considering temporal monitoring of communities."},"title":{"value":"Bringing a Feature Selection Metric from Machine Learning to Complex Networks"},"authors":{"value":["Nicolas Dugué","Jean-Charles Lamirel","Anthony Perez"]}},"tmdate":1773669339545,"pdate":1514764800000,"tcdate":1752303162451,"writers":["~"],"signatures":["~Jean-Charles_Lamirel1"],"forum":"DYf6WGu2pd","license":"CC BY-SA 4.0","number":567025,"cdate":1514764800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1773669339545,"domain":"DBLP.org","id":"DYf6WGu2pd","version":2},{"content":{"summary":{"value":"This paper proposes Physics-Manifold Flow Matching (PMFM), a generative framework for PDE simulation that enforces hard physical constraints by restricting the entire generative trajectory to a physical manifold. The approach addresses a key limitation of existing generative models for PDEs: they often violate fundamental conservation laws and boundary conditions. PMFM introduces two main innovations: (1) a projection mechanism that constrains trajectories to lie on a manifold where all states are physically valid by construction, combined with a Geometric Guidance Mechanism (GGM) that recovers high-frequency information typically lost during projection; (2) an Adaptive Constraint Projection Framework that dynamically selects and parameterizes active physical laws for complex multi-physics problems. The method is validated on several benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. **On the train-test gap**: The physical consistency is structurally guaranteed by the projection operator during the inference phase. However, during training, the interpolation path $u_t = (1-t)u_0 + t u_1$ does not lie on the manifold $\\mathcal{M}$ since $u_0 \\notin \\mathcal{M}$ and convex combinations generally leave $\\mathcal{M}$. Can you provide theoretical or empirical analysis on why learning $v_\\theta$ on off-manifold states transfers well to inference on manifold-constrained trajectories?\n\n2. **On convergence guarantees**: Theorem 1 assumes $u_0 \\in \\mathcal{M}$, but in practice inference starts from Gaussian noise $u_0 \\sim \\mathcal{N}(0,I)$ which violates constraints. Do you have convergence results for the projected ODE $\\dot{u} = \\Pi_{T_u\\mathcal{M}} v_\\theta(u,t)$ starting from $u_0 \\notin \\mathcal{M}$?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"**Novel approach to physical consistency:** The paper presents a principled method to enforce hard physical constraints in generative models through manifold projection, ensuring physical validity by construction rather than through soft penalties.\n    \n**Comprehensive experimental validation:** The method is tested on diverse PDE benchmarks covering different physical phenomena (conservation laws, shocks, incompressible flows), demonstrating consistent improvements over strong baselines including FNO, WNO, and diffusion models.\n    \n**Long-term stability:** Results show that PMFM maintains accuracy over extended temporal rollouts, addressing a critical limitation of purely data-driven approaches that accumulate errors during long-term prediction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"## Training-Inference Distribution Mismatch\n\nA critical weakness lies in the inconsistency between training and inference distributions. During training (Algorithm 1), the framework uses straight-line interpolation $u_t = (1-t)u_0 + t u_1$ where $u_0 \\sim p_0$ (Gaussian noise) and $u_1 \\sim p_{\\text{data}}$ (physical solutions). Since the physical manifold $\\mathcal{M}$ is generally non-convex and $u_0 \\notin \\mathcal{M}$, intermediate states $u_t$ for $t \\in (0,1)$ typically violate physical constraints, i.e., $C(u_t) \\neq 0$. However, during inference, trajectories are strictly constrained to $\\mathcal{M}$ via projection $\\Pi_{T_u\\mathcal{M}}$ at every integration step. This creates several issues:\n\n- **Distribution shift**: The network is trained on off-manifold states but deployed exclusively on manifold-constrained trajectories, potentially leading to suboptimal generalization.\n\n- **Theoretical inconsistency**: Theorem 1 assumes $u_0 \\in \\mathcal{M}$ to guarantee manifold invariance, but this condition is systematically violated. The tangent space $T_{u_t}\\mathcal{M}$ at non-manifold points $u_t$ lacks rigorous geometric interpretation.\n\nThe paper does not discuss this mismatch or provide ablation studies comparing alternative training strategies, such as geodesic interpolation on $\\mathcal{M}$ or explicit projection of $u_t$ at each training step. The strong empirical results may rely on the interpolation path remaining \"close\" to $\\mathcal{M}$ due to data smoothness, but this lacks formal justification.\n\n## Unclear Presentation of Core Mechanisms\n\nWhile the mathematical formulation is sound, the presentation of the two key innovations: the Geometric Guidance Mechanism (GGM) and the Adaptive Constraint Projection Framework, lacks clarity. For GGM (Sec. 4.1), the paper does not provide sufficient intuition for why encoding the residual $r$ into a latent code $z$ is necessary, or how the three components ($E_\\phi$, $B_\\theta$, $\\alpha_\\phi$) collaborate during training versus inference. The training-inference gap (where $z=0$ at inference) is mentioned but not justified. For the Adaptive Framework (Sec. 4.3), the distinction between \"analytical structure\" and \"learnable parameters\" in the constraint library is abstract, and critical training details (e.g., how the gating network $G_\\psi$ is trained, what happens if constraints conflict) are missing.\n\n## Insufficient Ablation Studies\n\nTable 2 only compares the full PMFM against a variant without physical constraints, which does not isolate the contributions of GGM and the Adaptive Framework. Key ablations are missing: (1) vanilla projection vs. +GGM, (2) +GGM vs. +GGM+Adaptive, (3) the role of the residual encoder $E_\\phi$, and (4) quantitative analysis of gating accuracy (e.g., how often does $G_\\psi$ select the correct active constraints?). Also, there are several previous work discussing how to use penalty loss in generative models to generate more accurate PDE solutions (Riemannian Score-Based Generative Modelling, Physics-Informed Diffusion Models, Generating Physical Dynamics under Priors). Authors should consider including these baselines.\n\n## Minors\n\n**Typos:**\nline 034, 038, 040, 046, 082, 103, 124, 190 (empty line), 209, 322, 334, 351, 352, 359 (capital), 403, 453, 465, 465.\n\n**Missing related works:** Riemannian Score-Based Generative Modelling, Physics-Informed Diffusion Models, Generating Physical Dynamics under Priors.\n\nD-Flow is not mentioned in the main text."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358588933,"tcdate":1761633707687,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4855/Reviewer_Lt8W"],"signatures":["ICLR.cc/2026/Conference/Submission4855/Reviewer_Lt8W"],"forum":"lRGAMx3f6N","number":1,"license":"CC BY 4.0","cdate":1761633707687,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4855/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358588933,"domain":"ICLR.cc/2026/Conference","replyto":"lRGAMx3f6N","id":"Q2itOi6Gj4","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Operator Learning","Hard Constraint","Flow Matching"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Simulating physical systems governed by partial differential equations (PDEs) is crucial across science and engineering. Recently, generative models—exemplified by Flow Matching—have emerged as a highly competitive approach due to their ability to effectively model high-dimensional solution distributions. However, these models often struggle to ensure physical consistency, frequently violating fundamental conservation laws or boundary conditions. In this work, we propose Physics-Manifold Flow Matching (PMFM), a novel generative framework for PDE simulation that directly addresses this challenge. PMFM introduces two key innovations. First, it enforces strict, hard physical constraints by restricting the entire generative trajectory to a physical manifold defined by analytical equations, while employing a Geometric Guidance Mechanism (GGM) to maintain high-fidelity solutions. Second, to handle complex multi-physics problems, we introduce an Adaptive Constraint Projection Framework that learns to dynamically select and parameterize the currently active physical laws. We validate PMFM on several challenging systems that are highly sensitive to physical constraints, and the results show that our framework is significantly superior to state-of-the-art physics-informed generative models in producing physically valid, long-term-stable simulations."},"_bibtex":{"value":"@misc{\nsun2026flowbased,\ntitle={Flow-based Automatic Neural Operator with Hard Physical Constraints},\nauthor={Li Sun and Hongbo Lv and Yunhui Xu and Yutong Ye and Peng Tang and Zhongtian Sun and Philip S. Yu},\nyear={2026},\nurl={https://openreview.net/forum?id=lRGAMx3f6N}\n}"},"title":{"value":"Flow-based Automatic Neural Operator with Hard Physical Constraints"},"pdf":{"value":"/pdf/03e0b044a067178cd46ddcdc997c77007581662d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"sun|flowbased_automatic_neural_operator_with_hard_physical_constraints"},"authorids":{"value":["~Li_Sun4","~Hongbo_Lv1","~Yunhui_Xu4","~Yutong_Ye1","~Peng_Tang7","~Zhongtian_Sun1","~Philip_S._Yu1"]},"authors":{"value":["Li Sun","Hongbo Lv","Yunhui Xu","Yutong Ye","Peng Tang","Zhongtian Sun","Philip S. Yu"]}},"version":2},{"content":{"venue":{"value":"xAI (2) 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-032-08324-1_17.pdf"},"venueid":{"value":"dblp.org/conf/XAI/2025"},"paperhash":{"value":"naujoks|leveraging_influence_functions_for_resampling_data_in_physicsinformed_neural_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jonas_R._Naujoks:","https://dblp.org/search/pid/api?q=author:Aleksander_Krasowski:","~Moritz_Weckbecker1","https://dblp.org/search/pid/api?q=author:Galip_Ümit_Yolcu:","https://dblp.org/search/pid/api?q=author:Thomas_Wiegand_0001:","https://dblp.org/search/pid/api?q=author:Sebastian_Lapuschkin:","https://dblp.org/search/pid/api?q=author:Wojciech_Samek:","~René_P._Klausen1"]},"html":{"value":"https://doi.org/10.1007/978-3-032-08324-1_17"},"_bibtex":{"value":"@inproceedings{DBLP:conf/xai/NaujoksKWY0LSK25,\n  author={Jonas R. Naujoks and Aleksander Krasowski and Moritz Weckbecker and Galip Ümit Yolcu and Thomas Wiegand and Sebastian Lapuschkin and Wojciech Samek and René P. Klausen},\n  title={Leveraging Influence Functions for Resampling Data in Physics-Informed Neural Networks},\n  year={2025},\n  cdate={1735689600000},\n  pages={383-395},\n  url={https://doi.org/10.1007/978-3-032-08324-1_17},\n  booktitle={xAI (2)},\n  crossref={conf/xai/2025-2}\n}\n"},"abstract":{"value":"Physics-informed neural networks (PINNs) offer a powerful approach to solving partial differential equations (PDEs), which are ubiquitous in the quantitative sciences. Applied to both forward and inverse problems across various scientific domains, PINNs have recently emerged as a valuable tool in the field of scientific machine learning. A key aspect of their training is that the data—spatio-temporal points sampled from the PDE’s input domain—are readily available. Influence functions, a tool from the field of explainable AI (XAI), approximate the effect of individual training points on the model, enhancing interpretability. In the present work, we explore the application of influence function-based sampling approaches for the training data. Our results indicate that such targeted resampling based on data attribution methods has the potential to enhance prediction accuracy in physics-informed neural networks, demonstrating a practical application of an XAI method in PINN training."},"title":{"value":"Leveraging Influence Functions for Resampling Data in Physics-Informed Neural Networks"},"authors":{"value":["Jonas R. Naujoks","Aleksander Krasowski","Moritz Weckbecker","Galip Ümit Yolcu","Thomas Wiegand","Sebastian Lapuschkin","Wojciech Samek","René P. Klausen"]}},"tmdate":1769356635931,"pdate":1767139200000,"externalIds":["dblp:conf/xai/NaujoksKWY0LSK25"],"tcdate":1769351858101,"writers":["~"],"signatures":["~Moritz_Weckbecker1"],"forum":"o8RJ4OlI5x","license":"CC BY-SA 4.0","number":798397,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769356635931,"domain":"DBLP.org","id":"o8RJ4OlI5x","version":2},{"content":{"TLDR":{"value":"A physics-informed latent network with adaptive weighting and a learned weak-form measure enables robust single-sensor forecasting under sparse sensing, outperforming SOTA models"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Physics-informed learning","conservation laws","adaptive loss weighting","latent field","monotone neural mapping","time-series forecasting"]},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Forecasting conservation-governed dynamics is often constrained by sparse sensing: in practice, we may have only a single boundary sensor and noisy exogenous variables. In this work we design an Adaptive Physics-Informed Latent Network (APILaNet) that learns a latent field and enforces 1D-conservation of physics law in the weak form using a learned, normalized space--time measure. Normalization makes physics enforcement insensitive to quadrature resolution and concentrates it on transient violations. A monotone, Lipschitz measurement layer maps latent variables to observed targets, improving identifiability from a single sensor. An adaptive, bounded scheduler scales the physics and smoothness loss terms with meaningful representations, emphasizing conservation of physics laws during events while preserving training stability. Learning a space-time measure for weak-form enforcement, combined with a monotone mapping and adaptive scheduling, enables accurate, data-efficient single-sensor forecasting in physics-governed systems. We evaluate APILaNet through a synthetic and hydrological case study, APILaNet outperforms strong sequence baselines and reduces MSE during extreme events, while improving Nash--Sutcliffe efficiency. Code will be released upon acceptance."},"_bibtex":{"value":"@misc{\nkucia2026apilanet,\ntitle={{APIL}aNet: Adaptive Physics-Informed Latent Network for Single-Sensor Forecasting},\nauthor={Adrian Kucia and Edward Rollason and Wai Lok Woo},\nyear={2026},\nurl={https://openreview.net/forum?id=VScnURO2g1}\n}"},"title":{"value":"APILaNet: Adaptive Physics-Informed Latent Network for Single-Sensor Forecasting"},"pdf":{"value":"/pdf/8cb8df5af300f217d6fe04bca6fa0d4678477421.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kucia|apilanet_adaptive_physicsinformed_latent_network_for_singlesensor_forecasting"},"authorids":{"value":["~Adrian_Kucia1","~Edward_Rollason1","~Wai_Lok_Woo1"]},"authors":{"value":["Adrian Kucia","Edward Rollason","Wai Lok Woo"]}},"tmdate":1770805126795,"tcdate":1758308345050,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20627/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission20627/Authors"],"forum":"VScnURO2g1","license":"CC BY 4.0","number":20627,"cdate":1758308345050,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission20627/-/Full_Submission","ICLR.cc/2026/Conference/Submission20627/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770805126795,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"VScnURO2g1","version":2},{"content":{"summary":{"value":"This paper presents Matting Anything 2, a versatile video matting model designed to overcome the domain-specificity (e.g., human-centric) and restrictive first-frame mask requirements of existing methods. The core technical contributions are twofold. First, a Promptable Dual-mode Decoder (PDD) that jointly predicts segmentation masks and high-quality trimaps, leveraging trimap-based guidance for generalization. Second, a Memory-Separable Siamese (MSS) mechanism that recurrently isolates trimap prediction from interfering mask memory, crucially improving temporal consistency for challenging transparent objects. To validate these contributions, the authors introduce the new, diverse Natural Video Matting (NVM) dataset. Experiments demonstrate that MAM2 significantly outperforms state-of-the-art methods on both diverse natural scenes and human portraits, accepting flexible prompts like points or boxes."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"The paper presents an extension of SAM 2 to the matting domain, and the quantitative and qualitative results shown are excellent. My primary question, however, concerns the validation scope. All experiments were conducted on synthetic (composited) videos. This raises a question about the method's true capability to video matting \"anything\". To fully substantiate the paper's strong claims, I recommend that the authors provide qualitative results (and quantitative results if possible) on real-world video clips, such as those sourced from YouTube, to demonstrate the model's robustness to non-synthetic artifacts."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The paper demonstrates compelling quantitative and qualitative results, significantly outperforming previous state-of-the-art methods.\n2. The paper well extends SAM2's promptable, generalist architecture to handle the distinct and more complex task of alpha matting. \n3. The paper is well-written, clearly organized, and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The name Natural Video Matting is confusing. In matting literature, \"natural\" typically implies real-world, non-composited videos. Since NVM is synthetic (composited from assets), this name is a misnomer and should be revised to avoid ambiguity.\n2. All experiments are conducted exclusively on synthetic (composited) videos. This leaves a significant gap in evaluation, as performance on real-world videos that contain artifacts like complex lighting, sensor noise, and motion blur remains unproven. The matting \"anything\" claim is therefore not fully substantiated.\n3. The paper lacks a dedicated limitations. There is no discussion of potential failure cases."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915487116,"tcdate":1761931395900,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission293/Reviewer_Qe9K"],"signatures":["ICLR.cc/2026/Conference/Submission293/Reviewer_Qe9K"],"forum":"6K08FPo2cf","number":3,"license":"CC BY 4.0","cdate":1761931395900,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission293/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915487116,"domain":"ICLR.cc/2026/Conference","replyto":"6K08FPo2cf","id":"yKmsiaw4VG","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Matting"]},"supplementary_material":{"value":"/attachment/e730155445628f4f98577db17e816ae46162e870.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video matting is a crucial task for many applications, but existing methods face significant limitations. They are often domain-specific, focusing primarily on human portraits, and rely on the mask of first frame that is challenging to acquire for transparent or intricate objects like fire or smoke. To address these challenges, we introduce Matting Anything 2 (MAM2), a versatile and robust video matting model that handles diverse objects using flexible user prompts such as points, boxes, or masks. We first propose Promptable Dual-mode Decoder (PDD), an effective structure that simultaneously predicts a segmentation mask and a corresponding high-quality trimap, leveraging trimap-based guidance to improve generalization. To tackle prediction instability for transparent objects across video frames, we further propose a Memory-Separable Siamese (MSS) mechanism. MSS employs a recurrent approach that isolates trimap prediction from potentially interfering mask memory, significantly enhancing temporal consistency. To validate our method's performance on diverse objects, we introduce the Natural Object Video Matting dataset, a new benchmark with substantially greater diversity. Extensive experiments show that MAM2 possesses exceptional matting accuracy and generalization capabilities. We believe MAM2 demonstrates a significant leap forward in creating a video matting method for anything."},"_bibtex":{"value":"@inproceedings{\nzhang2026matting,\ntitle={Matting Anything 2:  Towards Video Matting for Anything},\nauthor={Chenyi Zhang and Yiheng Lin and Yunchao Wei and Hongsong Wang and Caifeng Shan and Fang Zhao},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=6K08FPo2cf}\n}"},"title":{"value":"Matting Anything 2:  Towards Video Matting for Anything"},"pdf":{"value":"/pdf/0d14fdd0597751cfa86b9b08681f6ce8cffe4b65.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|matting_anything_2_towards_video_matting_for_anything"},"authorids":{"value":["~Chenyi_Zhang4","~Yiheng_Lin2","~Yunchao_Wei1","~Hongsong_Wang2","~Caifeng_Shan2","~Fang_Zhao1"]},"authors":{"value":["Chenyi Zhang","Yiheng Lin","Yunchao Wei","Hongsong Wang","Caifeng Shan","Fang Zhao"]}},"version":2},{"content":{"venue":{"value":"Journal of High Energy Physics"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"cho|optimass_a_package_for_the_minimization_of_kinematic_mass_functions_with_constraints"},"abstract":{"value":"Reconstructed mass variables, such as M <SUB>2</SUB>, M <SUB>2 C </SUB>, M <SUB> T </SUB> <SUP>*</SUP> , and M <SUB> T2</SUB> <SUP> W </SUP> , play an essential role in searches for new physics at hadron colliders. The calculation of these variables generally involves constrained minimization in a large parameter space, which is numerically challenging. We provide a C++ code, O ptimass, which interfaces with the M inuit library to perform this constrained minimization using the Augmented Lagrangian Method. The code can be applied to arbitrarily general event topologies, thus allowing the user to significantly extend the existing set of kinematic variables. We describe this code, explain its physics motivation, and demonstrate its use in the analysis of the fully leptonic decay of pair-produced top quarks using M <SUB>2</SUB> variables...."},"title":{"value":"OPTIMASS: a package for the minimization of kinematic mass functions with constraints"},"authors":{"value":[{"fullname":"Won Sang Cho","username":"https://orcid.org/orcid-search/search?searchQuery=Won%20Sang%20Cho"},{"fullname":"James S. Gainer","username":"https://orcid.org/orcid-search/search?searchQuery=James%20S.%20Gainer"},{"fullname":"Doojin Kim","username":"https://orcid.org/orcid-search/search?searchQuery=Doojin%20Kim"},{"fullname":"Sung Hak Lim","username":"~Lim_Sung_Hak1"},{"fullname":"Konstantin T. Matchev","username":"https://orcid.org/orcid-search/search?searchQuery=Konstantin%20T.%20Matchev"},{"fullname":"Filip Moortgat","username":"https://orcid.org/orcid-search/search?searchQuery=Filip%20Moortgat"},{"fullname":"Luc Pape","username":"https://orcid.org/orcid-search/search?searchQuery=Luc%20Pape"},{"fullname":"Myeonghun Park","username":"https://orcid.org/orcid-search/search?searchQuery=Myeonghun%20Park"}]}},"tmdate":1788846521277,"pdate":1451606400000,"externalIds":["doi:10.1007/jhep01(2016)026"],"tcdate":1788846520986,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Sung_Hak_Lim1"],"forum":"62UEZoPsep","license":"CC BY-SA 4.0","number":107898,"cdate":1673380703720,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1788846521277,"domain":"OpenReview.net/Public_Article","id":"62UEZoPsep","version":2},{"content":{"summary":{"value":"This article introduces physical forces as novel conditioning signals for video generation models. The core insight is that pretrained video models possess latent intuitive physics priors that can be elicited via fine-tuning on limited synthetic data. There are several innovations. Physics-based control encoding, Sim2Real generalization and emergent mass understanding. Human evaluations demonstrate superior force adherence and realism over text-conditioned baselines and trajectory-based methods."},"code_of_conduct_acknowledgement":{"value":"Yes"},"confidence":{"value":4},"originality":{"value":2},"questions":{"value":"Please refer to weakness."},"rating":{"value":4},"paper_formatting_concerns":{"value":"n/a"},"final_justification":{"value":"The rebuttal addressed my concerns, I will raise the rating to borderline accept."},"strengths_and_weaknesses":{"value":"Strength\n\n1. The authors carried out a novel and interesting task, and it achieved certain results. Additionally, they provide codes in supplementary materials.\n\n2. The experimental results are very abundant, with a wide variety of objects, and the experimental data and network parameters descriptions are also very detailed.\n\n3. Human evaluation shown in Table 1 reflects the effectiveness of proposed model.\n\nWeakness\n\n1. In some cases, such as in the example of a scene where the wind blows, usually only the target object moves while the background remains stationary. This does not conform to the principles of physics.\n\n2. Figure 3 is hard to follow. What do the blue lines in the last two columns mean?\n\n3. The presented model has poor performance on real physics, such as fluid dynamics. As shown in Figure 4, the diffusion of smoke affected by the wind is very unrealistic. The smoke emerging from the cigarette end cannot be straight, and the inertia is not so significant."},"quality":{"value":3},"significance":{"value":2},"responsible_reviewing_acknowledgement":{"value":"Yes"},"clarity":{"value":3},"limitations":{"value":"Please refer to weakness."},"ethical_concerns":{"value":["NO or VERY MINOR ethics concerns only"]}},"parentInvitations":"NeurIPS.cc/2025/Conference/-/Official_Review","nonreaders":[],"tmdate":1761706394134,"tcdate":1751720072384,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission3817/Reviewer_fCVZ"],"signatures":["NeurIPS.cc/2025/Conference/Submission3817/Reviewer_fCVZ"],"forum":"eX5aXfJQZc","number":4,"license":"CC BY 4.0","cdate":1751720072384,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/Submission3817/-/Official_Review","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission3817/Official_Review4/-/Review_Revision"],"mdate":1761706394134,"domain":"NeurIPS.cc/2025/Conference","replyto":"eX5aXfJQZc","id":"65ngGwBV0l","forumContent":{"venue":{"value":"NeurIPS 2025 poster"},"TLDR":{"value":"Video generalization models can learn physics-based controls from synthetic data and generalize this control across diverse geometries, settings, and materials."},"keywords":{"value":["controllable video generation","physical forces","synthetic data"]},"supplementary_material":{"value":"/attachment/857b716cd4bf6141e347e5786eb0f97ee6cc6501.zip"},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments.\nWhile navigation has been well-explored, physically meaningful interactions that mimic real-world forces remain largely understudied. \nIn this work, we investigate using physical forces as a control signal for video generation and propose force prompts which enable users to interact with images through both localized point forces, such as poking a plant, and global wind force fields, such as wind blowing on fabric. We demonstrate that these force prompts can enable videos to respond realistically to physical control signals by leveraging the physical prior in the original pretrained model, without using any 3D asset or physics simulator at inference. The primary challenge of force prompting is the difficulty in obtaining high quality paired force-video training data, both in the real world due to the difficulty of obtaining force signals, and in synthetic data due to limitations in the visual quality and domain diversity of physics simulators. Our key finding is that video generation models can *generalize* remarkably well when adapted to follow physical force conditioning from videos synthesized by Blender, even with limited demonstrations of few objects (e.g., flying flags, rolling balls, etc.). Our method can generate videos which simulate forces across diverse geometries, settings, and materials. We also try to understand the source of this generalization and perform ablations on the training data that reveal two key elements: visual diversity and the use of specific text keywords during training. Our approach is trained on only around 15k training examples for a single day on four A100 GPUs, and outperforms existing methods on force adherence and physics realism, bringing world models closer to real-world physics interactions. All datasets, code, and model weights will be open-sourced. Video examples can be found at https://sites.google.com/view/force-prompting-neurips2025"},"_bibtex":{"value":"@inproceedings{\ngillman2025force,\ntitle={Force Prompting: Video Generation Models Can Learn And Generalize Physics-based Control Signals},\nauthor={Nate Gillman and Charles Herrmann and Michael Freeman and Daksh Aggarwal and Evan Luo and Deqing Sun and Chen Sun},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=eX5aXfJQZc}\n}"},"title":{"value":"Force Prompting: Video Generation Models Can Learn And Generalize Physics-based Control Signals"},"pdf":{"value":"/pdf/83f23aff6d59602e22432a48c3869c9f56b4bba9.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"gillman|force_prompting_video_generation_models_can_learn_and_generalize_physicsbased_control_signals"},"authorids":{"value":["~Nate_Gillman1","~Charles_Herrmann1","~Michael_Freeman1","~Daksh_Aggarwal1","~Evan_Luo2","~Deqing_Sun2","~Chen_Sun1"]},"authors":{"value":["Nate Gillman","Charles Herrmann","Michael Freeman","Daksh Aggarwal","Evan Luo","Deqing Sun","Chen Sun"]}},"version":2},{"content":{"summary":{"value":"### Summary\n\n1. Extends Eigenvalue Analysis to LPV Models:\n\n   The paper builds on earlier work that analyzed eigenvalues in Linear Time-Invariant (LTI) State Space Models, where eigenvalues capture memory, stability, and selectivity. Eigenvalues near 1 mean good long-term memory, while those near 0 indicate forgetting or gating. \n\n   The authors extend this idea to Linear Parameter-Varying (LPV) systems using the Dynamical Systems Framework (DSF), which writes models like Attention and Mamba-2 as dynamical systems:\n   \n   $h_i = \\Lambda_i h_{i-1} + B_i u_i, \\quad y_i = C_i h_i + D_i u_i$\n\n   Here, the parameters change with input, allowing a similar eigenvalue analysis across model types.\n\n2. Analysis:\n\n   On memory, authors show the intuitive result that models that remember well keep eigenvalues close to 1. On language modeling (WikiText), the authors find that “gating” or eigenvalues near zero help model selectively store tokens.\n\n3. Architectural Modifications Based on the Analysis:\n\n   The paper tries to use these insights to improve models with changes including: Gating, Short Convs, Varying the number of layers etc"},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper’s central idea, that is extending eigenvalue analysis from LTI SSMs to LPV systems and attention, is natural and easy to follow.\n2. The connection between eigenvalues and memory retention/selectivity is intuitive and aligns with established understanding.\n3. The paper conducts experiments across a breadth of models (S4, Mamba-2, Self-Attention, Linear Attention etc) and tasks (LRA, MQAR, WikiText)"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### Weaknesses\n\n1. **On memory tasks**\n\n   1. The analysis of memory is intuitive and largely reiterates known understanding—that eigenvalues near one preserve information and those near zero lead to forgetting.\n   2. The attention results, though potentially interesting, are underexplained. The paper attributes attention’s poor performance on LRA to having both very low and very high eigenvalues, but I suspect this explanation may be correlational rather than causal. Attention is known to perform strongly on other memory-intensive retrieval tasks (e.g., NIAH [1]). It would have been useful if the paper discussed this discrepancy and explored why attention struggles on LRA despite its well-documented ability to retain long-term information in other settings.\n\n2. **On gating/selectivity tasks**\n\n   1. The interpretation that “gating” corresponds to zeroing out eigenvalues in WikiText models like Mamba-2 is reasonable but not novel as it directly follows from the Mamba-2's design of selection mechanism.\n   2. The follow-up “add gate” experiment on attention has a conceptual mismatch---The gating mechanism used in the experiments differs from Mamba-2’s within-sequence gating, and is instead applied per-token AFTER sequence mixing, making the analogy weak.\n   3. The task choice and conclusions seem inconsistent: gating is shown to help on IMDb, a memory-heavy task where (if i understand correctly) it should theoretically hurt (or at-least not help). Furthermore, ListOps, which is also memory intensive, unexpectedly develops gating, which remains unexplained. \n\n       I expected to see a gating-improved task, like WikiText, show improvement when gating is added to attention.\n\n3. **On convolution and varying number of layers**\n\n   1. These experiments are not motivated by eigenvalue analysis and are instead motivated from Mamba-2's strong performance which feels disconnected from the main argument.\n   2. The claim that short convolutions “take over long-range memory” seems to be incorrect—short convolutions are \"short\" (of size 4) and cannot provide such capacity. The paper employs this argument to justify why a one layer attention model with convolution could solve MQAR. I believe this justification is incorrect. Prior work [1] has shown that MQAR performance requires \"induction-head formation\": first sequence mixer mixes keys and values and the second sequence mixer retrieves the correct value. Convolution helps because it performs the first task, not because it replaces long-range retrieval.\n\n4. **On evaluation scope**\n\n   1. For the proposed architectural changes, the experiments are limited to small-scale settings and on synthetic tasks.\n   2. It would be more convincing to test these modifications on language modeling and downstream evals at across multiple scales (e.g., 125M, 350M, 750M, 1.3B) to assess their importance.\n\n5. **On scope and completeness**\n\n   1. The paper briefly mentions that eigenvalues evolve during training, suggesting potential task-dependent initialization, but this idea is never explored.\n   2. The architectural additions should be tested on language modeling for validation.\n   3. Overall, the paper reads as an incremental extension of prior DSF analyses, with several unexplained experimental results.\n\n-------\n[1]: Mechanistic evaluation of Transformers and state space models. Aryaman Arora, Neil Rathi, Nikil Roashan Selvam, Róbert Csordás, Dan Jurafsky, Christopher Potts"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762940622778,"tcdate":1762122283750,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21047/Reviewer_N9x6"],"signatures":["ICLR.cc/2026/Conference/Submission21047/Reviewer_N9x6"],"forum":"ALf0xHCCaP","number":2,"license":"CC BY 4.0","cdate":1762122283750,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21047/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762940622778,"domain":"ICLR.cc/2026/Conference","replyto":"ALf0xHCCaP","id":"QsCkLROH9C","forumContent":{"TLDR":{"value":"This paper shows that eigenvalues offer a metric for memory retention and selective forgetting in sequence models, guiding model design choices to meet task requirements."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Sequence models","dynamical systems","memory dynamics"]},"supplementary_material":{"value":"/attachment/a34ab7106844c8b1974675fdf300085c394ec367.pdf"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Although softmax attention drives state-of-the-art performance for sequence models, its quadratic complexity limits scalability, motivating linear alternatives such as state space models (SSMs). While these alternatives improve efficiency, their fundamental differences in information processing remain poorly understood. In this work, we leverage the recently proposed dynamical systems framework to represent softmax, norm and linear attention as dynamical systems, enabling a structured comparison with SSMs by analyzing their respective eigenvalue spectra. Since eigenvalues capture essential aspects of dynamical system behavior, we conduct an extensive empirical analysis across diverse sequence models and benchmarks. We first show that eigenvalues influence essential aspects of memory and long-range dependency modeling, revealing spectral signatures that align with task requirements. Building on these insights, we then investigate how architectural modifications in sequence models impact both eigenvalue spectra and task performance.  This correspondence further strengthens the position of eigenvalue analysis as a principled metric for interpreting, understanding, and ultimately improving the capabilities of sequence models."},"_bibtex":{"value":"@misc{\nrickenbach2026tasklevel,\ntitle={Task-Level Insights from Eigenvalues across Sequence Models},\nauthor={Rahel Rickenbach and Jelena Trisovic and Alexandre Didier and Jerome Sieber and Melanie Zeilinger},\nyear={2026},\nurl={https://openreview.net/forum?id=ALf0xHCCaP}\n}"},"title":{"value":"Task-Level Insights from Eigenvalues across Sequence Models"},"pdf":{"value":"/pdf/ef15be1e4ca44ec34adb76dbbbb5e30448f4997c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"rickenbach|tasklevel_insights_from_eigenvalues_across_sequence_models"},"authorids":{"value":["~Rahel_Rickenbach1","~Jelena_Trisovic1","~Alexandre_Didier1","~Jerome_Sieber1","~Melanie_Zeilinger1"]},"authors":{"value":["Rahel Rickenbach","Jelena Trisovic","Alexandre Didier","Jerome Sieber","Melanie Zeilinger"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (3) 2023"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-53472-0_7.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2023"},"paperhash":{"value":"bougiatiotis|efficient_complex_network_representation_using_prime_numbers"},"authorids":{"value":["~Konstantinos_Bougiatiotis1","https://dblp.org/search/pid/api?q=author:Georgios_Paliouras:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-53472-0_7"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/BougiatiotisP23,\n  author={Konstantinos Bougiatiotis and Georgios Paliouras},\n  title={Efficient Complex Network Representation Using Prime Numbers},\n  year={2023},\n  cdate={1672531200000},\n  pages={75-86},\n  url={https://doi.org/10.1007/978-3-031-53472-0_7},\n  booktitle={COMPLEX NETWORKS (3)},\n  crossref={conf/complexnetworks/2023-3}\n}\n"},"abstract":{"value":"In this work, we propose a novel representation of complex networks, which is compact and enables very efficient network analysis. Multi-relational networks capture complex data relationships and have a wide range of applications. As they get to be used with ever larger quantities of data, it is crucial to find efficient ways to represent and analyse them. This paper introduces the concept of Prime Adjacency Matrices (PAMs), which utilize prime numbers, to represent the relations of the network. Due to the Fundamental Theorem of Arithmetic, this allows for a lossless, compact representation of a complete multi-relational graph, using a single adjacency matrix. Moreover, this representation enables the fast computation of multi-hop adjacency matrices, which can be useful for a variety of downstream tasks. We illustrate the benefits of using the proposed approach through various network analysis tasks."},"title":{"value":"Efficient Complex Network Representation Using Prime Numbers"},"authors":{"value":["Konstantinos Bougiatiotis","Georgios Paliouras"]}},"tmdate":1759174680492,"pdate":1672531200000,"externalIds":["dblp:conf/complexnetworks/BougiatiotisP23"],"tcdate":1759174664727,"writers":["~"],"signatures":["~Konstantinos_Bougiatiotis1"],"forum":"trpmtKRmUR","license":"CC BY-SA 4.0","number":632073,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1759174680492,"domain":"DBLP.org","id":"trpmtKRmUR","version":2},{"content":{"summary":{"value":"This paper addresses bilevel optimization problems where both the upper and lower level objective functions are expensive black-box functions.  The authors are the first to introduce information-theoretic principles to this domain, proposing a new acquisition function named BLJES (Bilevel optimization via Lower-bound based Joint Entropy Search). The core idea is to shift the optimization goal from directly seeking higher function values to maximizing the bilevel information gain regarding the optimal solutions $(x^\\*, \\theta^\\*)$ and their corresponding optimal values $(f^\\*, g^\\*)$ for both levels. This unified decision criterion balances exploration across both levels of the problem. As direct computation of this information gain (which is the mutual information) is intractable, the paper derives a computable variational lower bound. A key technical contribution is the creative extension of the \"truncation approximation\" concept, which is mature in single-level BO, to the complex structure of bilevel problems. Furthermore, the paper demonstrates the flexibility of this information-theoretic framework, showing its natural extension to the decoupled setting where observations can be separated and to more complex scenarios involving expensive black-box constraints. Experimental results show that BLJES achieves competitive performance against the current state-of-the-art method, BILBO, on various synthetic datasets and benchmark problems."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The core methodology relies on a variational lower bound to approximate the intractable mutual information. The quality of this bound, and the validity of maximizing it as a proxy for the true information gain, critically depends on how well the chosen variational distribution $q$ approximates the true posterior $p$. The paper mainly justifies this approach by citing its empirical success in prior work, but does not provide a theoretical analysis of the bound's tightness in this more complex bilevel optimization setting.\n\nCan the authors provide more theoretical results regarding the tightness of this bound? For example, is it possible to theoretically analyze the error induced by this approximation, or under what conditions this lower bound becomes a closer approximation to the true mutual information? Providing analysis in this area would significantly strengthen the theoretical completeness of the proposed method."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- This work is the first to establish a rigorous information-theoretic framework for bilevel Bayesian optimization. The proposed bilevel information gain concept offers a novel perspective for tackling this complex problem.\n\n- The paper successfully transforms a theoretically ideal yet computationally intractable objective (mutual information) into an engineering-feasible algorithm through a series of well-founded approximations, including the variational lower bound, truncation approximation, and Monte Carlo sampling. The extension of the truncation approximation from the single-level to the bilevel case appears intuitively sound."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- A significant shortcoming of this paper is the absence of theoretical analysis for the proposed algorithm. The authors note in the introduction that the SOTA comparison, BILBO, possesses a theoretical regret guarantee, yet this paper provides no equivalent theoretical support for BLJES.  Although the experimental section uses bilevel simple regret as an evaluation metric and demonstrates good empirical convergence, this is not a substitute for formal theoretical proof. The absence of theoretical results such as a regret bound leaves the algorithm’s sample complexity and convergence behavior under finite budgets uncharacterized.  This is a considerable theoretical gap for a Bayesian optimization method, where sample efficiency is central to its purpose.\n\n- All experiments are conducted in extremely low-dimensional spaces (both upper and lower dimensions $d_X, d_\\Theta \\leq 2$).  This severely limits the generality of the results. It is well known that BO methods suffer from the curse of dimensionality, and the nested nature of bilevel optimization likely worsens this challenge.  Thus, success on low-dimensional synthetic problems does not demonstrate scalability or robustness of BLJES in realistic, high-dimensional cases.\n\n-  Although one experiment includes a simulator-based energy market problem, the vast majority still rely on synthetic functions sampled from GP priors and standard academic benchmarks.  To convincingly demonstrate the method’s practical value, validation on real-world black-box problems (e.g., engineering design, hyperparameter tuning) is needed.  The current experimental setup feels thin and does not fully exhibit the method’s ability to handle complex, noisy, or constrained real-world tasks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926017278,"tcdate":1761640767082,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15784/Reviewer_k4Ay"],"signatures":["ICLR.cc/2026/Conference/Submission15784/Reviewer_k4Ay"],"forum":"39GLKT8ZBy","number":3,"license":"CC BY 4.0","cdate":1761640767082,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15784/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926017278,"domain":"ICLR.cc/2026/Conference","replyto":"39GLKT8ZBy","id":"fNRPZjgX1t","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Bilevel optimization","Bayesian optimization"]},"supplementary_material":{"value":"/attachment/b31c9ba45734ca1b963653aa9a0e1b58bf75d47c.zip"},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"abstract":{"value":"A bilevel optimization problem consists of two optimization problems nested as an upper- and a lower-level problem, in which the optimality of the lower-level problem defines a constraint for the upper-level problem. This paper considers Bayesian optimization (BO) for the case that both the upper- and lower-levels involve expensive black-box functions. Because of its nested structure, bilevel optimization has a complex problem definition and, compared with other standard extensions of BO such as multi-objective or constraint settings, it has not been widely studied. We propose an information-theoretic approach that considers the information gain of both the upper- and lower-optimal solutions and values. This enables us to define a unified criterion that measures the benefit for both level problems, simultaneously. Further, we also show a practical lower bound based approach to evaluating the information gain. We empirically demonstrate the effectiveness of our proposed method through several benchmark datasets."},"_bibtex":{"value":"@misc{\nkanayama2026informationtheoretic,\ntitle={Information-Theoretic Bayesian Optimization for Bilevel Optimization Problems},\nauthor={Takuya Kanayama and Yuki Ito and Tomoyuki Tamura and Masayuki Karasuyama},\nyear={2026},\nurl={https://openreview.net/forum?id=39GLKT8ZBy}\n}"},"title":{"value":"Information-Theoretic Bayesian Optimization for Bilevel Optimization Problems"},"pdf":{"value":"/pdf/d8f778ea56b20c3cb26a711f2c2e263cc9db6225.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kanayama|informationtheoretic_bayesian_optimization_for_bilevel_optimization_problems"},"authorids":{"value":["~Takuya_Kanayama1","~Yuki_Ito1","~Tomoyuki_Tamura1","~Masayuki_Karasuyama1"]},"authors":{"value":["Takuya Kanayama","Yuki Ito","Tomoyuki Tamura","Masayuki Karasuyama"]}},"version":2},{"content":{"summary":{"value":"This paper presents PFGS (Pose-Fused 3D Gaussian Splatting), a pose-aware framework for reconstructing complete 3D objects from multi-pose image captures. Unlike conventional 3DGS approaches that assume a static object pose, PFGS handles scenarios where an object must be reoriented to reveal occluded regions.\n\nThe method iteratively fuses each auxiliary-pose sequence into a unified 3DGS representation of a main pose through three stages:\n\n1. Global registration via a mixed-pose subset processed by foundation models (Fast3R, VGGT) and refined through a silhouette-consensus alignment.\n2. Gradient-based local registration using silhouette and RGB consistency.\n3. 3DGS completion through balanced sampling and fine-tuning.\n\nQuantitative results on synthetic and real datasets demonstrate substantial improvements in pose accuracy (orders of magnitude lower error than VGGT/Fast3R) and enhanced 3D reconstruction."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How does the varying illumination by the pose change the impact the reconstruction result iof PFGS?\n2. Why do VGGT and Fast3R perform much worse on masked (foreground-only) images?\n3. How does PFGS handle cases where auxiliary poses differ by extreme rotations (e.g., >120°)?\n4. Runtime comparison: How does the 22.6 min average compare to pure COLMAP or foundation-model-only approaches under the same GPU setup?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Proposing problem “Multi-pose reconstruction” seems challenging and interesting research direction.\n- Novel framework to unify the poses of multiple sets of frames with efficient usage of foundation models\n- Evaluations on well-curated synthetic and real datasets followed by ablation studies on each components."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Illumination inconsistency ignored:** When an object’s pose changes, surrounding illumination directions and shading patterns also change. Fusing such data into a single 3DGS inevitably introduces appearance inconsistency. However, the method focuses on “photorealistic, view-consistent” results without modeling this radiometric gap. \n2. **Insufficient analysis of baseline performance drop:** Table 1 shows VGGT and Fast3R failing dramatically on masked objects, but the authors do not analyze why (e.g., loss of background context or inconsistent normalization). A discussion on this could provide insight into the necessity of the fusion method.\n3. **Limited scalability evidence:** While synthetic and real datasets are informative, all objects are relatively small-scale with moderate pose changes. It remains unclear whether PFGS generalizes to complex objects (e.g., articulated or deformable subjects) with extreme pose changes."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928117037,"tcdate":1761725628250,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18415/Reviewer_Df9k"],"signatures":["ICLR.cc/2026/Conference/Submission18415/Reviewer_Df9k"],"forum":"8bJs2ElHGM","number":1,"license":"CC BY 4.0","cdate":1761725628250,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18415/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928117037,"domain":"ICLR.cc/2026/Conference","replyto":"8bJs2ElHGM","id":"ISEnhwkMoq","forumContent":{"TLDR":{"value":"3D Gaussian Splatting for Complete Multi-Pose Object Reconstruction"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["3D Gaussian Splatting","Object Reconstruction","Multi-Pose Object Capture"]},"supplementary_material":{"value":"/attachment/8873d5fde5f026a6b0c378a9dd714f21a5df9acd.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent advances in 3D Gaussian Splatting (3DGS) have enabled high-quality, real-time novel-view synthesis from multi-view images. However, most existing methods assume the object is captured in a single, static pose, resulting in incomplete reconstructions that miss occluded or self-occluded regions. We introduce PFGS, a pose-aware 3DGS framework that addresses the practical challenge\nof reconstructing complete objects from multi-pose image captures. Given images of an object in one main pose and several auxiliary poses, PFGS iteratively fuses each auxiliary set into a unified 3DGS representation of the main pose. Our pose-aware fusion strategy combines global and local registration to merge views effectively and refine the 3DGS model. While recent advances in 3D foundation models have improved registration robustness and efficiency, they remain limited by high memory demands and suboptimal accuracy. PFGS overcomes these challenges by incorporating them more intelligently into the registration process: it leverages background features for per-pose camera pose estimation and employs foundation models for cross-pose registration. This design captures the best of both approaches while resolving background inconsistency issues. Experimental results demonstrate that PFGS consistently outperforms strong baselines in both qualitative and quantitative evaluations, producing more complete reconstructions and higher-fidelity 3DGS models."},"_bibtex":{"value":"@misc{\nyen2025pfgs,\ntitle={{PFGS}: Pose-Fused 3D Gaussian Splatting for Complete Multi-Pose Object Reconstruction},\nauthor={Ting-Yu Yen and Yu-Sheng Chiu and Shih-Hsuan Hung and Peter Wonka and Hung-Kuo Chu},\nyear={2025},\nurl={https://openreview.net/forum?id=8bJs2ElHGM}\n}"},"title":{"value":"PFGS: Pose-Fused 3D Gaussian Splatting for Complete Multi-Pose Object Reconstruction"},"pdf":{"value":"/pdf/4e421293462119be3510b809dd5e99c257a6e9ab.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yen|pfgs_posefused_3d_gaussian_splatting_for_complete_multipose_object_reconstruction"},"authorids":{"value":["~Ting-Yu_Yen3","~Yu-Sheng_Chiu1","~Shih-Hsuan_Hung1","~Peter_Wonka1","~Hung-Kuo_Chu2"]},"authors":{"value":["Ting-Yu Yen","Yu-Sheng Chiu","Shih-Hsuan Hung","Peter Wonka","Hung-Kuo Chu"]}},"version":2},{"content":{"summary":{"value":"This paper studies the in-context learning capabilities of MLPs and mixer-MLP models, comparing them to Transformers on tasks such as synthetic regression and classification. The models are evaluated on limited training data and larger test sets to determine when they shift from in-weight learning to in-context learning. The authors also introduce unique relational tasks—match-to-sample, sphere oddball, and line oddball—revealing that MLPs and relationally bottlenecked MLPs outperform Transformers on these tasks. They suggest that the results may stem from the inductive biases of these architectures."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Can you provide more insights on why transformers are failing in relational tasks?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. Every experiment in the paper is designed thoroughly. \n2. This is the first work encountered that explores the ICL capabilities of MLPs, which could be relevant to the literature on foundation models, especially in time series.\n3. The addition of relational tasks to the existing synthetic regression and classification experiments contributes valuable insights into Transformer limitations. Transformers perform poorly when test exemplars differ significantly from the training data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper could have included real regression data. Most existing literature focuses on synthetic tasks, and exploring real data (even simple regression datasets) with somewhat complex underlying distributions would have added valuable insights."}},"nonreaders":[],"tmdate":1731427411499,"tcdate":1730722429795,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1336/Reviewer_2Mnt"],"signatures":["ICLR.cc/2025/Conference/Submission1336/Reviewer_2Mnt"],"forum":"MbX0t1rUlp","number":3,"license":"CC BY 4.0","cdate":1730722429795,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1336/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427411499,"domain":"ICLR.cc/2025/Conference","replyto":"MbX0t1rUlp","id":"c84mAV7Bw3","forumContent":{"TLDR":{"value":"On a range of widely studied synthetic in-context learning tasks, we find that MLPs perform comparably with Transformers under the same compute budget."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["In-context learning","relational reasoning","synthetic tasks","MLP","MLP-Mixer","Transformer"]},"supplementary_material":{"value":"/attachment/dff9b6374ad3a5441a8531987e2d47dd94293259.zip"},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In-context learning (ICL), the remarkable ability to solve a task from only input exemplars, is often assumed to be a unique hallmark of Transformer models. By examining commonly employed synthetic ICL tasks, we demonstrate that multi-layer perceptrons (MLPs) can also learn in-context. Moreover, MLPs, and the closely related MLP-Mixer models, learn in-context comparably with Transformers under the same compute budget in this setting. We further show that MLPs outperform Transformers on a series of classical tasks from psychology designed to test relational reasoning, which are closely related to in-context classification. These results underscore a need for studying in-context learning beyond attention-based architectures, while also challenging prior arguments against MLPs' ability to solve relational tasks. Altogether, our results highlight the unexpected competence of MLPs in a synthetic setting, and support the growing interest in all-MLP alternatives to Transformer architectures. It remains unclear how MLPs perform against Transformers at scale on real-world tasks, and where a performance gap may originate. We encourage further exploration of these architectures in more complex settings to better understand the potential comparative advantage of attention-based schemes."},"_bibtex":{"value":"@inproceedings{\ntong2025mlps,\ntitle={{MLP}s Learn In-Context on Regression and Classification Tasks},\nauthor={William Lingxiao Tong and Cengiz Pehlevan},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=MbX0t1rUlp}\n}"},"title":{"value":"MLPs Learn In-Context on Regression and Classification Tasks"},"pdf":{"value":"/pdf/13cb6a3404468e0575bca172720e7ed9f428dfbc.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"tong|mlps_learn_incontext_on_regression_and_classification_tasks"},"authorids":{"value":["~William_Lingxiao_Tong1","~Cengiz_Pehlevan2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["William Lingxiao Tong","Cengiz Pehlevan"]}},"version":2},{"content":{"venue":{"value":"Eng. Appl. Artif. Intell. 2022"},"venueid":{"value":"dblp.org/journals/EAAI/2022"},"paperhash":{"value":"li|consensus_reaching_model_for_counterintuitive_in_ds_evidence_theory_and_application_under_2tuple_linguistic_representation"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Xiaobing_Yu:"]},"html":{"value":"https://doi.org/10.1016/j.engappai.2022.104832"},"_bibtex":{"value":"@article{DBLP:journals/eaai/LiY22,\n  author={Chenliang Li and Xiaobing Yu},\n  title={Consensus reaching model for counter-intuitive in D-S evidence theory and application under 2-tuple linguistic representation},\n  year={2022},\n  cdate={1640995200000},\n  journal={Eng. Appl. Artif. Intell.},\n  volume={112},\n  pages={104832},\n  url={https://doi.org/10.1016/j.engappai.2022.104832}\n}\n"},"abstract":{"value":"In the information fusion field, Dempster–Shafer (D–S) evidence theory is a multi-source technology to solve the uncertain problems. With the aim of improving the decision accuracy, D–S evidence theory can make full use of information from different sources which are redundant and complementary. However, the influence caused by conflicting evidence during information fusion, named the counter-intuitive result, will confuse the selection of decision-makers (DMs). Inspired by the distance-based uncertainty measure, which is a typical technique used to manage uncertain information, the consensus reaching model is applied to overcome the influence caused by conflicting evidence in this paper. Considering the hesitant and uncertain of cognition, 2-tuple linguistic representation method is introduced to model and manage this vague decision information given by DMs. Finally, a consensus reaching model for counter-intuitive result in D–S evidence theory is put forward. To verify the effectiveness of the proposed method, the selection of plant protection machine suppliers is modeled as a multi-criteria decision-making (MCDM) problem. According to the decision-making process, the best option for plant protection machine suppliers is obtained. The decision result indicates a strong correlation between the consistency value and conflicting value of evidence."},"title":{"value":"Consensus reaching model for counter-intuitive in D-S evidence theory and application under 2-tuple linguistic representation"},"authors":{"value":["Chenliang Li","Xiaobing Yu"]}},"tmdate":1741016957797,"pdate":1640995200000,"tcdate":1741016562811,"writers":["~"],"signatures":["~Chenliang_Li4"],"forum":"x2JtGN9N6v","license":"CC BY-SA 4.0","number":346416,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1741016957797,"domain":"DBLP.org","id":"x2JtGN9N6v","version":2},{"content":{"summary":{"value":"In this paper, the authors propose to adapt reinforcement learning with physical feedback to fine-tune diffusion models, in order to produce equilibrium structures adhering to physical principles. The method extends DDPO to the molecular domain, employing a physics-informed reward function derived from force fields evaluations, that guides the generation toward meaningful configurations. The authors demonstrate the effectiveness of the approach through various experiments on QM9 and GEOM-drug datasets, and using two different pretrained diffusion models."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How does the number of sampled trajectories affect the performance?\n- Could you report the molecule stability and uniqueness of the generated samples for GEOM-drug (Table 2)? \n- How does the proposed RLPF approach affect different properties of the generated samples, such as QED, SA and LogP, as measured in [2]?\n\n[1] Black, Kevin, et al. \"Training diffusion models with reinforcement learning.\" arXiv preprint arXiv:2305.13301 (2023).\n[2] Hoogeboom, Emiel, et al. \"Equivariant diffusion for molecule generation in 3d.\" International conference on machine learning. PMLR, 2022."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- This paper is well-written and well-structured.\n- The authors incorporate physics-informed reinforcement learning to guide diffusion models to generate physically plausible outputs. \n- The authors validate the effectiveness of the proposed approach by showing an improved performance across different datasets, and combined with different pretrained diffusion models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper offers limited methodological novelty. It primarily builds on top of prior work [1], that adapts DDPO to optimize pre-trained diffusion models for various downstream objectives. This paper simply applies the same methodology to the molecular domain, using a different objective function defined over generated 3D molecular structures. \n- The impact of the masking mechanism introduced in Equation 13 on performance remains unclear. Could the authors provide additional experiments to support the claim that this masking stabilizes training?\n- What is the rationale behind using the reward function in Equation 12? Have you explored alternative physics-informed reward functions? And if so, how do they compare?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923160627,"tcdate":1760975498011,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12211/Reviewer_ygra"],"signatures":["ICLR.cc/2026/Conference/Submission12211/Reviewer_ygra"],"forum":"eurcml8JFs","number":1,"license":"CC BY 4.0","cdate":1760975498011,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12211/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923160627,"domain":"ICLR.cc/2026/Conference","replyto":"eurcml8JFs","id":"0Lpp8A7izS","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Molecule Generation，Equivariant Diffusion model，Reinforcement Learning"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Generating physically realistic 3D molecular structures remains a core challenge in molecular generative modeling. While diffusion models equipped with equivariant neural networks have made progress in capturing molecular geometries, they often struggle to produce equilibrium structures that adhere to physical principles such as force field consistency. To bridge this gap, we propose Reinforcement Learning with Physical Feedback (RLPF), a novel framework that extends Denoising Diffusion Policy Optimization to 3D molecular generation. RLPF formulates the task as a Markov decision process and applies proximal policy optimization to fine-tune equivariant diffusion models. Crucially, RLPF introduces reward functions derived from force-field evaluations, providing direct physical feedback to guide the generation toward energetically stable and physically meaningful structures. Experiments on the QM9 and GEOM-drug datasets demonstrate that RLPF significantly improves molecular stability compared to existing methods. These results highlight the value of incorporating physics-based feedback into generative modeling."},"_bibtex":{"value":"@misc{\nzhou2026guiding,\ntitle={Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation},\nauthor={Zhijian Zhou and Junyi An and Zongkai Liu and Yun-Fei Shi and Fenglei Cao and Xuan Zhang and Wenli Wang and Chao Qu and Yuan Qi},\nyear={2026},\nurl={https://openreview.net/forum?id=eurcml8JFs}\n}"},"title":{"value":"Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation"},"pdf":{"value":"/pdf/edce68a8115ca6d15a0b3d1a51286a159d375607.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhou|guiding_diffusion_models_with_reinforcement_learning_for_stable_molecule_generation"},"authorids":{"value":["~Zhijian_Zhou2","~Junyi_An1","~Zongkai_Liu1","~Yun-Fei_Shi1","~Fenglei_Cao1","~Xuan_Zhang31","~Wenli_Wang3","~Chao_Qu3","~Yuan_Qi2"]},"authors":{"value":["Zhijian Zhou","Junyi An","Zongkai Liu","Yun-Fei Shi","Fenglei Cao","Xuan Zhang","Wenli Wang","Chao Qu","Yuan Qi"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the long-standing \"performance-interpretability trade-off\" in machine learning—where simple, interpretable models (e.g., decision trees, physics-based formulas) lack expressive power, while complex black-box models (e.g., neural networks) sacrifice transparency. It proposes a novel Tutor-Pupil augmentation framework to resolve this trade-off by leveraging \"minimal input-space corrections\" rather than output adjustments, enabling both performance gains and enhanced interpretability."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"see weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Unlike prior work that corrects outputs (e.g., residual networks, ensemble stacking), this paper corrects inputs, preserving the Pupil’s interpretability.\n\n- The paper’s originality lies in redefining the paradigm of model augmentation, removing limitations of prior work, and creating novel links between data-driven learning and theoretical insight—all of which challenge long-standing practices in interpretable AI.\n\n- The paper does not limit the Tutor-Pupil framework to a single task type but adapts it to three distinct domains—a creative extension that proves its generality."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper strictly adopts a \"train Pupil first, then train Tutor\" serial paradigm (Pupil parameters are frozen during Tutor training but fails to explore joint training—a critical gap that limits the framework’s ability to fully leverage synergies between the two models and may amplify Pupil’s inherent flaws.\n\n- Novelty Gap: “Input-space correction” is not new. Position the paper as “systematic, global counterfactuals for interpretable models” rather than a brand-new paradigm and provide a taxonomy table that shows how Tutor-Pupil differs from (i) local counterfactuals, (ii) adversarial examples, (iii) data-augmentation policies on objectives, constraints, and evaluation metrics.\n\n- The paper validates the framework exclusively with interpretable Pupils (decision trees, logistic regression, ideal gas law but fails to test black-box Pupils (e.g., ResNet, Transformer)—a critical gap, as many real-world systems rely on complex models that need interpretive tools (e.g., medical image classifiers using CNNs)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359355253,"tcdate":1760501750329,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21311/Reviewer_4yEq"],"signatures":["ICLR.cc/2026/Conference/Submission21311/Reviewer_4yEq"],"forum":"TvP90DWijM","number":1,"license":"CC BY 4.0","cdate":1760501750329,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21311/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359355253,"domain":"ICLR.cc/2026/Conference","replyto":"TvP90DWijM","id":"E9IohE0sTy","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Model Augmentation","Machine learning for physical sciences"]},"supplementary_material":{"value":"/attachment/3c3be36a327b5f49d0f782e7dccf6b3533fe3df5.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"State-of-the-art machine learning models often incorporate prior knowledge or structural information about the task or data distribution. In some tasks, such knowledge may arise from first principles or emerge as simplified, learned functions that distill essential aspects of the data distribution. Model augmentation has emerged as a strategy to leverage this structured knowledge by coupling it with an auxiliary model to improve predictive performance, while preserving the interpretability offered by the simpler component. In this work, we present a new augmentation framework called the Tutor-Pupil scheme, which is designed to enhance both performance and interpretability. The Pupil is a fixed model, structurally designed for the core task, while the Tutor is a more flexible model trained to apply minimal input-level corrections to improve the Pupil’s performance on the modified input. This strict separation of roles enables the Tutor not only to compensate for the Pupil’s limitations but also to act as a diagnostic instrument. By examining the Tutor’s targeted interventions, we can identify failure modes, detect regions where the Pupil struggles to generalize, and uncover residual patterns or higher-order structures in the data not captured by the original model."},"_bibtex":{"value":"@inproceedings{\nbiparva2026the,\ntitle={The Tutor-Pupil Augmentation: Enhancing Learning and Interpretability via Input Corrections},\nauthor={Darya Biparva and Maarten Schoukens and Donatello Materassi},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=TvP90DWijM}\n}"},"title":{"value":"The Tutor-Pupil Augmentation: Enhancing Learning and Interpretability via Input Corrections"},"pdf":{"value":"/pdf/2d39911ef1d849cb36990ea563092b423907555a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"biparva|the_tutorpupil_augmentation_enhancing_learning_and_interpretability_via_input_corrections"},"authorids":{"value":["~Darya_Biparva1","~Maarten_Schoukens1","~Donatello_Materassi2"]},"authors":{"value":["Darya Biparva","Maarten Schoukens","Donatello Materassi"]}},"version":2},{"content":{"summary":{"value":"This paper introduces AC-PKAN to fit functions and solve PDEs. The proposed AC-PKAN method employs a Residual Gradient Attention (RGA) mechanism to dynamically adjusts loss term weights, wavelet-activated MLPs with learnable parameters, and Chebyshev polynomial\nBased KANs to improve training efficiency and prediction accuracy. The paper is well-written. The method is effective compared to existing methods on five benchmarks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See in the weakness"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The method AC-RKAN is original. Ablation studies show the necessity of each module of AC-RKAN.\n2. The paper is well-written. The theoretical insight is interesting and the experiments are rich, demonstrating the effective of the method.\n3. The AC-RKAN method can make a timely contribution to the community of physics-informed machine learning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The experiments and the comparisons are not challenging. For the 1D-wave case, the author claimed PINNsFormer has a relative l2 norm 0.32 in Table 2, but Wang 2022 (fig.6 of their paper) has trained PINN for this case to achieve a relative l2 norm 1.7e-3, which is the same order to AC-KAN. The other PDEs are also simple 2d Poisson eq. Although with Heterogeneous Problem and Complex Geometry, the solution is smooth, and are easy to solve by simple traditional numerical methods such as finite element methods. Challenging PINN problems Wang 2023 such as the Kuramoto–Sivashinsky equation with chaotic behavior，Lid-driven cavity flow (Re=3200)， Navier–Stokes flow in a torus or around a cylinder, are not reported in the paper.\n\n[a]Wang S, Yu X, Perdikaris P. When and why PINNs fail to train: A neural tangent kernel perspective[J]. Journal of Computational Physics, 2022, 449: 110768.\n[b]Wang S, Sankaran S, Wang H, et al. An expert's guide to training physics-informed neural networks[J]. arXiv preprint arXiv:2308.08468, 2023.\n\n2.The method introduction is expatiatory, from page3-7. A concentration of the method introduction is recommended and more theoretical insights and experiments can be discussed in the main text."}},"nonreaders":[],"tmdate":1731427586562,"tcdate":1729392787939,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2375/Reviewer_w3dV"],"signatures":["ICLR.cc/2025/Conference/Submission2375/Reviewer_w3dV"],"forum":"kqdNvAhJrJ","number":1,"license":"CC BY 4.0","cdate":1729392787939,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2375/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427586562,"domain":"ICLR.cc/2025/Conference","replyto":"kqdNvAhJrJ","id":"cfTUC3hP9w","forumContent":{"TLDR":{"value":"We introduce AC-PKAN, a novel framework that enhances Physics-Informed Neural Networks (PINNs) with Chebyshev polynomials and attention mechanisms to improve efficiency and accuracy in solving complex partial differential equations."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-Informed Neural Networks","Kolmogorov–Arnold Networks","Attention Mechanism","PDEs","Chebyshev Polynomials"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper introduces AC-PKAN, an advanced framework for Physics-Informed Neural Networks (PINNs) that integrates Kolmogorov–Arnold Networks (KANs) with Chebyshev Type-I polynomials and incorporates both internal and external attention mechanisms. Traditional PINNs based on Multilayer Perceptrons (MLPs) encounter challenges when handling complex partial differential equations (PDEs) due to vanishing gradients, limited interpretability, and computational inefficiency. To address these issues, we enhance the model from both external and internal perspectives. Externally, we propose a novel Residual Gradient Attention (RGA) mechanism that dynamically adjusts loss term weights based on gradient norms and residuals, thereby mitigating gradient stiffness and residual imbalance. Internally, AC-PKAN employs point-wise Chebyshev polynomial-based KANs, wavelet-activated MLPs with learnable parameters, and internal attention mechanisms. These integrated components improve both training efficiency and prediction accuracy. We provide mathematical proofs demonstrating that AC-PKAN can theoretically solve any finite-order PDE. Experimental results from five benchmark tasks across three domains show that AC-PKAN consistently outperforms or matches state-of-the-art models such as PINNsFormer, establishing it as a highly effective tool for solving complex real-world engineering problems."},"_bibtex":{"value":"@misc{\nzhang2025acpkan,\ntitle={{AC}-{PKAN}: Attention-Enhanced and Chebyshev Polynomial-Based Physics-Informed Kolmogorov{\\textendash}Arnold Networks},\nauthor={Hangwei Zhang and Yan Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=kqdNvAhJrJ}\n}"},"title":{"value":"AC-PKAN: Attention-Enhanced and Chebyshev Polynomial-Based Physics-Informed Kolmogorov–Arnold Networks"},"pdf":{"value":"/pdf/5607e4b5a49b4930999f11adf7aff12a98865897.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|acpkan_attentionenhanced_and_chebyshev_polynomialbased_physicsinformed_kolmogorovarnold_networks"},"authorids":{"value":["~Hangwei_Zhang2","~Yan_Wang12"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hangwei Zhang","Yan Wang"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Pattern Anal. Mach. Intell. 2026"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/34/11372200/11278613.pdf"},"venueid":{"value":"dblp.org/journals/PAMI/2026"},"paperhash":{"value":"wu|physicsinformed_matrix_factorization_operator"},"authorids":{"value":["","","","~Lingling_Li1","","","~Wenping_Ma2",""]},"html":{"value":"https://doi.org/10.1109/TPAMI.2025.3640844"},"_bibtex":{"value":"@article{DBLP:journals/pami/WuTJLLLMY26,\n  author={Wenming Wu and Chenxi Tian and Licheng Jiao and Lingling Li and Xu Liu and Fang Liu and Wenping Ma and Shuyuan Yang},\n  title={Physics-Informed Matrix Factorization Operator},\n  year={2026},\n  month={March},\n  cdate={1772323200000},\n  journal={IEEE Trans. Pattern Anal. Mach. Intell.},\n  volume={48},\n  number={3},\n  pages={3556-3570},\n  url={https://doi.org/10.1109/TPAMI.2025.3640844}\n}\n"},"abstract":{"value":"Matrix factorization is a fundamental characterization model in machine learning and is usually solved using mathematical decomposition reconstruction loss. However, matrix factorization is a data-driven model whose results depend on data quality, making it susceptible to noise. Inspired by physics, the law of conservation of energy is used to introduce physical laws into matrix factorization, which is called Physics-informed Matrix Factorization operator (PiMF). The PiMF operator uses the heat conduction equation to construct the energy objective function for matrix factorization, thereby retaining the mathematical model’s decomposition meaning and satisfying the interpretability of physics. The PiMF follows the physical laws, thereby suppressing irregular or sudden noise signals that violate these physical principles. The solutions of the PiMF operator include more comprehensive knowledge of mathematics and physics, which improves the ability to generalize complex data, especially for noisy data. We demonstrate the consistency of the energy objective function and the mathematical model, which verifies the feasibility of matrix factorization using physical energy laws. In addition, the physical interpretability of the PiMF operator is proved from the perspective of energy decline. This study proposes two practical algorithms for PiMF in classification and clustering tasks, enhancing the practicability of matrix factorization by incorporating task-specific prior information constraints. The experimental results of PiMF for classification and clustering demonstrate the advantages of the proposed operator. The importance of physics-informed matrix factorization is verified, especially for noisy data."},"title":{"value":"Physics-Informed Matrix Factorization Operator"},"authors":{"value":["Wenming Wu","Chenxi Tian","Licheng Jiao","Lingling Li","Xu Liu","Fang Liu","Wenping Ma","Shuyuan Yang"]}},"tmdate":1774445947335,"pdate":1798675200000,"externalIds":["dblp:journals/pami/WuTJLLLMY26"],"tcdate":1774445763470,"writers":["~"],"signatures":["~Wenping_Ma2"],"forum":"DrjIx3ZTmV","license":"CC BY-SA 4.0","number":859448,"cdate":1772323200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1774445947335,"domain":"DBLP.org","id":"DrjIx3ZTmV","version":2},{"content":{"summary":{"value":"This paper presents EM-GANSim, a learning-based approach for real-time electromagnetic (EM) propagation simulation in indoor environments. The core technical contribution is a modified conditional GAN architecture that incorporates both geometric information and transmitter location to predict power distribution heatmaps while adhering to electromagnetic propagation principles. The authors propose a physically-inspired learning framework that integrates direct propagation, reflection, and diffraction effects through specialized loss terms in the GAN's objective function.\n\nThe method claims to achieve comparable accuracy to traditional ray tracing-based simulators while offering significant speed improvements (reported as 5X faster). The authors evaluate their approach on 15 indoor scenes and provide ablation studies examining the impact of noise and physical constraints. They also introduce a dataset comprising over 2,000 indoor scene models with corresponding EM simulation heatmaps.\n\nWhile I am not an expert in electromagnetic propagation simulation and wireless communications, the paper appears to address an important practical challenge in real-time EM simulation. However, there is some ambiguity in how the method handles true 3D environments versus 2D representations, and the room generation and data preparation processes could benefit from clearer documentation. The paper presents an interesting application of deep learning to physics-based simulation, though both its theoretical foundations and physical accuracy need closer examination."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"- Could you clarify how the method handles true 3D propagation versus 2D layout information? The current results only show 2D heatmaps. Could you provide vertical propagation results at different heights? How does the network architecture specifically process and maintain height information?\n- Please describe in detail how the 2K+ room models were created/sourced. What is the distribution of room types, sizes, and configurations in your dataset? How to ensure the synthetic scenes are physically realistic? How are different materials modeled and validated?\n- How to determine the weights (α, β, γ) in the physics loss function? What measures are taken to ensure training stability? How to handle varying room sizes in the network?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper presents an interesting application of conditional GANs to EM simulation. While both GANs and EM simulation are established fields, their combination for real-time indoor propagation simulation represents a fresh approach to an important practical problem.\n- The method achieves notable acceleration (reported 5X speedup) compared to traditional ray tracing methods. If these results can be thoroughly validated, this could be valuable for real-time applications.\n- The attempt to incorporate electromagnetic principles through specialized loss terms (direct propagation, reflection, and diffraction) shows thoughtful consideration of the physics involved, though the theoretical guarantees need more examination.\n- While the dataset generation process needs better documentation, the collection of indoor scenes and EM simulation results could be useful for future research in this direction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- A weakness is the unclear treatment of \"3D\" simulation. While the paper claims to handle \"3D indoor environments,\" the evidence presented is primarily 2D heatmaps. There's no clear explanation of how height information is processed in the network, no visualization of vertical propagation effects, and no analysis of height-dependent signal variations. Table 2 only specifies area (square meter) without height information. The paper needs to either demonstrate true 3D capability or clarify that it's a 2.5D approach.\n- Critical details about the \"2K+ models and 64M heatmaps\" are missing. The paper doesn't explain how these indoor scenes were generated, validated, or processed. Without this information, readers cannot assess data quality or reproduce the results.\n- The method description lacks important specifics. The GAN architecture details, training process, and hyperparameter selection are not fully described. The physics-based loss weights lack justification, and there's minimal discussion of training stability.\n- The experimental validation relies mainly on MSE comparisons. The performance measurements lack important context - hardware specifications, memory requirements, and preprocessing costs are not reported. The gap between training (3 dbm²) and testing (8.5 dbm²) MSE also needs explanation."}},"nonreaders":[],"tmdate":1731428375125,"tcdate":1730533293678,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5331/Reviewer_2hyN"],"signatures":["ICLR.cc/2025/Conference/Submission5331/Reviewer_2hyN"],"forum":"29JDZxRgPZ","number":3,"license":"CC BY 4.0","cdate":1730533293678,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5331/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428375125,"domain":"ICLR.cc/2025/Conference","replyto":"29JDZxRgPZ","id":"sBfCkGsH65","forumContent":{"TLDR":{"value":"A novel GAN-based approach for real-time 3D indoor electromagnetic simulation, drastically reducing computation time while maintaining accuracy."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Generative Adversarial Networks (GAN)","Electromagnetic Propagation","Real-time Simulation","3D Indoor Environments"]},"supplementary_material":{"value":"/attachment/4ba4ac8ffde4dcec176ff3b6f9fa566e06700166.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present a novel machine-learning (ML) approach  (EM-GANSim) for real-time electromagnetic (EM) propagation that is used for wireless communication simulation in 3D indoor environments. Our approach uses a modified conditional Generative Adversarial Network (GAN) that incorporates encoded geometry and transmitter location while adhering to the electromagnetic propagation theory. The overall physically-inspired learning is able to predict the power distribution in 3D scenes, which is represented using heatmaps.  Our overall accuracy is comparable to ray tracing-based EM simulation, as evidenced by lower mean squared error values. Furthermore, our GAN-based method drastically reduces the computation time, achieving a 5X speedup on complex benchmarks. In practice, it can compute the signal strength in a few milliseconds on any location in 3D indoor environments. We also present a large dataset of 3D models and EM ray tracing-simulated heatmaps. To the best of our knowledge, EM-GANSim is the first real-time algorithm for EM simulation in complex 3D indoor environments. We plan to release the code and the dataset."},"_bibtex":{"value":"@misc{\nwang2025emgansim,\ntitle={{EM}-{GANS}im: Real-time and Accurate {EM} Simulation Using Conditional {GAN}s for 3D Indoor Scenes},\nauthor={Ruichen Wang and Dinesh Manocha},\nyear={2025},\nurl={https://openreview.net/forum?id=29JDZxRgPZ}\n}"},"title":{"value":"EM-GANSim: Real-time and Accurate EM Simulation Using Conditional GANs for 3D Indoor Scenes"},"pdf":{"value":"/pdf/14594ce9613897746e50c37f5ce7df92ad457b46.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|emgansim_realtime_and_accurate_em_simulation_using_conditional_gans_for_3d_indoor_scenes"},"authorids":{"value":["~Ruichen_Wang4","~Dinesh_Manocha3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruichen Wang","Dinesh Manocha"]}},"version":2},{"content":{"summary":{"value":"The paper proposed a model which integrate the Latent Diffusion Model (LDM) with physics-based domain knowledge. A set of loss functions were designed. The paper proposed a blending algorithm to improve the accuracy of inpainting task. The proposed method is work for simulated parallel projection data (not real-world data)."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1 The training and inference time comparison is necessary for sampling is time-consuming.\n2 Notation is not defined clearly, e.g. equation (1) sg. and z_qn never used.\n3 Can the method extend to 2D fan-beam and 3D Cone beam reconstruction\n4 Visual comparison of reconstructed image from recovered sinogram also needed for that a minor error of sinogram may lead to streaky artifacts in image."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"(1)\tdomain specific physics knowledge of CT image formation for inpainting sinograms taking into account both measurement and reconstruction domains.\n(2)\tRecover sinogram with different masks and different sampling ratio. \n(3)\tSuitable for Sparse data acquisition and Missing Wedge Problem."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1)\tThe loss is complex and too many parameters. Ablation of every part is necessary.\n(2)\tThe paper uses the simulated projection data other than real word data.\n(3)\tThe proposed method is only work with Parallel beam projection geometry with is xxxxx\nin real application.\n\n(4)\tThe downstream task of the method is image reconstruction. Comparison with reconstruction method for sparse view reconstruction and limit view reconstruction is necessary, such as dual domain reconstruction.\n[1] W. Wu, D. Hu, C. Niu, H. Yu, V. Vardhanabhuti and G. Wang, \"DRONE: Dual-Domain Residual-based Optimization NEtwork for Sparse-View CT Reconstruction,\" in IEEE Transactions on Medical Imaging, vol. 40, no. 11, pp. 3002-3014, Nov. 2021, doi: 10.1109/TMI.2021.3078067\n[2] Ding, Q., Ji, H., Gao, H., Zhang, X. (2021). Learnable Multi-scale Fourier Interpolation for Sparse View CT Image Reconstruction.\n(5) The literature review is limited, many CT reconstruction works, such as iterative reconstruction and deep learning reconstruction (image domain, unrolling (ADMM-Net), and plug-&play method) are not given"}},"nonreaders":[],"tmdate":1731429219260,"tcdate":1730639125500,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13542/Reviewer_tPkr"],"signatures":["ICLR.cc/2025/Conference/Submission13542/Reviewer_tPkr"],"forum":"IfPfUHRowT","number":3,"license":"CC BY 4.0","cdate":1730639125500,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13542/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429219260,"domain":"ICLR.cc/2025/Conference","replyto":"IfPfUHRowT","id":"J9QT5ojSdW","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Sinogram Inpainting","Physics","Latent Diffusion Model","X-ray Imaging"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Computed Tomography (CT) is a widely used non-invasive imaging technique for materials at microscopic or sub-microscopic length scales in synchrotron radiation facilities. Typically, the object is rotated relative to the X-ray beam, and 2D projection images are recorded by the detector at different rotation angles. The 3D object is then reconstructed by combining these projections and solving a computationally demanding inverse problem. The quality of the reconstructed image is critical for scientific analysis and is influenced by various factors, including the number of projections, exposure time or dose, and the reconstruction algorithm. In this work, we develop a foundation model by integrating a Generative AI-based Latent Diffusion Model (LDM) with physics-based domain knowledge. Specifically, we first incorporate a set of loss functions into our LDM that accurately capture the physical properties of the CT data acquisition process. We demonstrate that addition of these loss functions aids in stable training of the autoencoder in the LDM and improves its accuracy. The autoencoder and the Diffusion model of the LDM is trained with real-world experimental data. Collecting real world experimental data from Synchrotron beamlines is often time-consuming and challenging. We demonstrate that the autoencoder trained with a combination of real world experimental data and phantom shapes features also performs comparable to the autoencoder trained with real world data. Second, we introduce a novel image blending method to combine the LDM’s generated output with the original, extremely sparse sinogram data. Since our model integrates physics-guided loss functions focused on CT data acquisition, it simplifies the creation of downstream tasks and facilitates the adaptation of new features from different experiments. Our experimental evaluation demonstrates improvements of upto 23.5 % in SSIM for sinogram quality and 13.8 % for reconstructed image quality compared to state-of-the-art techniques."},"_bibtex":{"value":"@misc{\nbanerjee2025inpainting,\ntitle={Inpainting the Sinogram from Computed Tomography using Latent Diffusion Model and Physics},\nauthor={Srutarshi Banerjee and Jiaze E and Bin Ren and Tekin Bicer},\nyear={2025},\nurl={https://openreview.net/forum?id=IfPfUHRowT}\n}"},"title":{"value":"Inpainting the Sinogram from Computed Tomography using Latent Diffusion Model and Physics"},"pdf":{"value":"/pdf/865420cfd0809612b7cf6b28ef1996be447748d2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"banerjee|inpainting_the_sinogram_from_computed_tomography_using_latent_diffusion_model_and_physics"},"authorids":{"value":["~Srutarshi_Banerjee1","~Jiaze_E1","~Bin_Ren1","~Tekin_Bicer1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Srutarshi Banerjee","Jiaze E","Bin Ren","Tekin Bicer"]}},"version":2},{"content":{"venue":{"value":"VISIGRAPP (2): VISAPP 2025"},"venueid":{"value":"dblp.org/conf/VISIGRAPP/2025"},"paperhash":{"value":"lens|conditioned_generative_ai_for_synthetic_training_of_6d_object_pose_detection"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Mathijs_Lens:","https://dblp.org/search/pid/api?q=author:Aaron_Van_Campenhout:","~Toon_Goedemé1"]},"html":{"value":"https://doi.org/10.5220/0013130600003912"},"_bibtex":{"value":"@inproceedings{DBLP:conf/visigrapp/LensCG25,\n  author={Mathijs Lens and Aaron Van Campenhout and Toon Goedemé},\n  title={Conditioned Generative AI for Synthetic Training of 6D Object Pose Detection},\n  year={2025},\n  cdate={1735689600000},\n  pages={324-331},\n  url={https://doi.org/10.5220/0013130600003912},\n  booktitle={VISIGRAPP (2): VISAPP},\n  crossref={conf/visigrapp/2025-2}\n}\n"},"abstract":{"value":"In this paper, we propose a method to generate synthetic training images for a more complex computer vision task compared to image classification, specifically 6D object pose detection. We demonstrate that conditioned diffusion models can generate unlimited training images for training an object pose detection model for a custom object type. Moreover, we investigate the potential of (automatically) filtering out ill-produced images in the dataset, which increases the quality of the image dataset, and show the importance of finetuning the trained model with a limited amount of real-world images to bridge the remaining sim2real domain gap. We demonstrate our pipeline in the use case of parcel box detection for the automation of delivery vans. All code is publicly available on our GitLab https://gitlab.com/EAVISE/avc/generative-ai-synthetic-training-pose-detection."},"title":{"value":"Conditioned Generative AI for Synthetic Training of 6D Object Pose Detection"},"authors":{"value":["Mathijs Lens","Aaron Van Campenhout","Toon Goedemé"]}},"tmdate":1758954169409,"pdate":1735689600000,"externalIds":["dblp:conf/visigrapp/LensCG25"],"tcdate":1758954158676,"writers":["~"],"signatures":["~Toon_Goedemé1"],"forum":"4yWioJ077K","license":"CC BY-SA 4.0","number":630768,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1758954169409,"domain":"DBLP.org","id":"4yWioJ077K","version":2},{"content":{"summary":{"value":"The manuscript presents a novel method for community detection called CLANN. This method addresses challenges in semi-supervised community detection, particularly issues with community core consistency and scalability. The authors draw an analogy between community formation and crystallization kinetics, introducing two main components: the Nucleus Proposer and the Transitive Annealer. These components use cliques as starting points and leverage physics-based principles to optimize and expand communities effectively. The paper claims that CLANN surpasses existing state-of-the-art methods in accuracy and efficiency based on extensive experiments across various datasets."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1:Can you provide more intuitive examples or visual explanations of how crystallization kinetics translate to the community detection process? \n\n2:How does CLANN compare with simpler heuristic-based community detection methods in terms of performance and resource consumption?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1:The analogy with crystallization kinetics introduces a novel physics-grounded perspective for community detection, enriching the field with fresh conceptual insights.\n\n2:Empirical evaluations show that CLANN outperforms established methods in both single and hybrid dataset scenarios, demonstrating its robustness and superior scalability.\n\n3:The integration of the Nucleus Proposer and Transitive Annealer simplifies the growth process without relying on computationally heavy methods like GANs or reinforcement learning.\n\n4:The paper provides extensive tests, including ablation studies and adaptability analyses, solidifying the validity of its contributions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1:While the physics-based analogy is intriguing, it may be difficult for readers without a background in crystallization kinetics to grasp fully. More simplified explanations or visual aids could make this clearer.\n\n2:Although the method improves scalability over GAN-based approaches, the process of clique enumeration can still be computationally expensive, potentially limiting applicability to extremely large graphs.\n\n3:The experiments, though diverse, may benefit from including more real-world networks with varying characteristics to generalize the method's effectiveness further.\n\n4:The paper does not thoroughly discuss the practical implementation details, such as computational resources required for different dataset sizes, which may affect the adoption of the model in real-world applications."}},"nonreaders":[],"tmdate":1731428263437,"tcdate":1730646831745,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6488/Reviewer_4nXW"],"signatures":["ICLR.cc/2025/Conference/Submission6488/Reviewer_4nXW"],"forum":"jQ5T1Pbnx7","number":4,"license":"CC BY 4.0","cdate":1730646831745,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6488/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428263437,"domain":"ICLR.cc/2025/Conference","replyto":"jQ5T1Pbnx7","id":"x4Wmmnj3TL","forumContent":{"TLDR":{"value":"semi-supervised community detection approach premised on crystallization kinetics"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Semi-supervised Community Detection","Clique","Annealing","Crystallization Kinetics"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Semi-supervised community detection methods are widely used for identifying specific communities due to the label scarcity. Existing semi-supervised community detection methods typically involve two learning stages \\ie, learning in both initial identification and subsequent adjustment, which often starts from an unreasonable community core candidate.\nMoreover, these methods encounter scalability issues because they depend on reinforcement learning and generative adversarial networks, leading to higher computational costs and restricting the selection of candidates. \nTo address these limitations, we draw a parallel between crystallization kinetics and community detection to integrate the spontaneity of the annealing process into community detection.\nSpecifically, we liken community detection to identifying a crystal subgrain (core) that expands into a complete grain (community) through a process similar to annealing. Based on this finding, we propose CLique ANNealing (CLANN), which applies kinetics concepts to community detection by integrating these principles into the optimization process to strengthen the consistency of the community core. Subsequently, a learning-free Transitive Annealer was employed to refine the first-stage candidates by merging neighboring cliques and repositioning the community core, enabling a spontaneous growth process that enhances scalability.\nExtensive experiments on diverse community detection datasets demonstrate that CLANN outperforms state-of-the-art methods across multiple real-world datasets, showcasing its exceptional efficacy and efficiency in community detection."},"_bibtex":{"value":"@misc{\ncheng2025new,\ntitle={New Recipe for Semi-supervised Community Detection: Clique Annealing under Crystallization Kinetics},\nauthor={ling Cheng and Jiashu Pu and Ruicheng Liang and Qian Shao and Hezhe Qiao and Feida Zhu},\nyear={2025},\nurl={https://openreview.net/forum?id=jQ5T1Pbnx7}\n}"},"title":{"value":"New Recipe for Semi-supervised Community Detection: Clique Annealing under Crystallization Kinetics"},"pdf":{"value":"/pdf/dfd4411dcbaa795727f424bd95ed21e87d906591.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"cheng|new_recipe_for_semisupervised_community_detection_clique_annealing_under_crystallization_kinetics"},"authorids":{"value":["~ling_Cheng1","~Jiashu_Pu1","~Ruicheng_Liang1","~Qian_Shao1","~Hezhe_Qiao1","~Feida_Zhu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["ling Cheng","Jiashu Pu","Ruicheng Liang","Qian Shao","Hezhe Qiao","Feida Zhu"]}},"version":2},{"content":{"summary":{"value":"The authors propose Fair4Free, a novel generative approach for generating high-fidelity synthetic fair data using knowledge distillation.\nTheir proposed method consists of three stages: 1) train a VAE on the biased dataset to learn a fair representation, 2) use it as a teacher model and distill the fair representation to a student model, and 3) use the trained VAE decoder and student model to reconstruct high-fidelity fair synthetic samples. They show experiment results on tabular and image datasets."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"* The paper needs to improve the clarity of implementation details and evaluation protocols mentioned above.\n* The paper needs an ablation study to justify the design choices of the method."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"* The paper addresses an important problem of generating synthetic fair data.\n* The experiments include a wide range of evaluation metrics to thoroughly evaluate the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The contributions of this paper are unclear. Although the authors claim the method is *data-free*, it still relies on a biased dataset to train the teacher model in the first stage. Furthermore, the method combines the student encoder with the teacher decoder to generate synthetic samples. Then why not simply use the teacher model directly? What advantages are gained by distilling the teacher encoder into a smaller student encoder only to recombine it with the teacher decoder? The paper also lacks an ablation study to justify the design choices in the method.\n* Overall, the paper needs substantial improvement in writing quality and clarity. Implementation details are severely lacking, e.g., hyperparameters used for training the proposed and compared methods, details on synthetic data generation during evaluation, how the random forest model is trained, and explanation of the evaluation metrics (e.g., DPR, EOR).\n* The method demonstrates minimal performance gains over the compared methods across downstream tasks (Table 2,3).\n* They should additionally report FID for synthetic data quality."}},"nonreaders":[],"tmdate":1731428297532,"tcdate":1730709233366,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6131/Reviewer_LGEY"],"signatures":["ICLR.cc/2025/Conference/Submission6131/Reviewer_LGEY"],"forum":"iRgzG5DKgA","number":4,"license":"CC BY 4.0","cdate":1730709233366,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6131/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428297532,"domain":"ICLR.cc/2025/Conference","replyto":"iRgzG5DKgA","id":"IeDY2M1uU5","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["data fairness","fair generative models","knowledge distillation","latent space distillation","synthetic data","biased data"]},"supplementary_material":{"value":"/attachment/96403b6eb869aa033b81be0ee2f8d2ad8174de3a.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This work presents Fair4Free, a novel generative model to generate synthetic fair data using data-free distillation in the latent space. Fair4Free can work on the situation when the data is private or inaccessible.  In our approach, we first train a teacher model to create fair representation and then distil the knowledge to a student model (using a smaller architecture). The process of distilling the student model is data-free, i.e. the student model does not have access to the training dataset while distilling. After the distillation, we use the distilled model to generate fair synthetic samples. Our extensive experiments show that our synthetic samples outperform state-of-the-art models in all three criteria (fairness, utility and synthetic quality) with a performance increase of 5\\% for fairness, 8\\% for utility and 12\\% in synthetic quality for both tabular and image datasets."},"_bibtex":{"value":"@misc{\nsikder2025fairfree,\ntitle={Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data-Free Distillation},\nauthor={Md Fahim Sikder and Daniel de Leng and Fredrik Heintz},\nyear={2025},\nurl={https://openreview.net/forum?id=iRgzG5DKgA}\n}"},"title":{"value":"Fair4Free: Generating High-fidelity Fair Synthetic Samples using Data-Free Distillation"},"pdf":{"value":"/pdf/bf09a39d8fc93c7601b5f410a98be727c91cb446.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"sikder|fair4free_generating_highfidelity_fair_synthetic_samples_using_datafree_distillation"},"authorids":{"value":["~Md_Fahim_Sikder1","~Daniel_de_Leng1","~Fredrik_Heintz1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Md Fahim Sikder","Daniel de Leng","Fredrik Heintz"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a diffusion-based method for reconstructing high-fidelity physics simulation data given low-fidelity input. The paper also proposes to incorporate prior knowledge via the gradient of PDE residual and a new weighting scheme based on multi-resolution analysis (Wavelet transform) for diffusion loss. The numerical experiments on several 2D flow problem showcase that the proposed method has better L2 accuracy and lower residuals over other baseline methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* Following point 2, what is the residual of the reference DNS simulation?\n\n* The authors state that the high fidelity data for 2D Kolmogorov flow is derived by running DNS on a 256 x 256 grid (feel free to correct me if I’m wrong). However, under a similar setting (Re=1000), prior works like Shu et al. [1] and Kochkov et al. [2] have used a much finer discretization, i.e. 2048 x 2048. \n\n* (Minor) In algorithm 1, what is the rationale for using Adam instead of simpler gradient descent? In addition, is Adam re-initialized after very DDIM step?\n\n[1] Shu, D., Li, Z., & Farimani, A. B. (2023). A physics-informed diffusion model for high-fidelity flow field reconstruction. Journal of Computational Physics, 478, 111972.\n\n[2] Kochkov, Dmitrii, et al. \"Machine learning–accelerated computational fluid dynamics.\" Proceedings of the National Academy of Sciences 118.21 (2021): e2101784118."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Upsampling and reconstructing under-resolved physics is important for building hybrid solver and inverse problem in PDE applications.\n\n2. A new spatial weighting scheme based on Wavelet transformation which modulate the loss based on the spectrum of signal. The new scheme is technically sound and experiments show that it consistently improves model performance.\n\n3. The introduction to the proposed method is clear and easy-to-follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Applying diffusion model with guidance based on target constraint for the inverse problem is not an entirely new technique (for example, DPS: https://arxiv.org/pdf/2209.14687 and On conditional diffusion models for PDE simulations: https://arxiv.org/abs/2410.16415 has explored similar technique)\n\n2. The evaluation of model’s prediction in terms of physics coherence is relatively vague. First, all the reported residuals are quite large and there is no reference showing what is a reasonable scale. While the average error of different frequency components in the wavelet domain is shown, there is no information on the spectrum of different wavelength.\n\n3. The authors run the solver on a coarse grid to get the “low fidelity” data and then run the solver on fine grid to get “high fidelity” data, instead of artificially downsampling the data. The low-fidelity simulation will deviate from the high-fidelity one as time evolves due to the under-resolved error. Yet looking at the figure comparing data trajectory of different fidelities, the general structures of the vortices are very similar across different fidelities (for example, Figure 8) in my opinion. I hope the authors can provide more clarification and analysis regarding the dataset, such as the spectrum of different simulation and perhaps a simple study of mesh convergence."}},"nonreaders":[],"tmdate":1732637786839,"tcdate":1730436075104,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4262/Reviewer_ZSxU"],"signatures":["ICLR.cc/2025/Conference/Submission4262/Reviewer_ZSxU"],"forum":"EaiU4F5pwn","number":2,"license":"CC BY 4.0","cdate":1730436075104,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4262/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732637786839,"domain":"ICLR.cc/2025/Conference","replyto":"EaiU4F5pwn","id":"W0XMubHmvE","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed Neural Networks","Computational Fluid Dynamics"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Machine learning (ML) models are increasingly explored in fluid dynamics as a promising way to generate high-fidelity computational fluid dynamics data more efficiently. A common strategy is to use low-fidelity data as computational-efficient inputs, and employ ML techniques to reconstruct high-fidelity flow fields. However, existing work typically assumes that low-fidelity data is artificially downsampled from high-fidelity sources, which limits model performance. In real-world applications, low-fidelity data is generated directly by numerical solvers with a lower initial state resolution, resulting in large deviations from high-fidelity data. To address this gap, we propose PG-Diff, a novel diffusion model for reconstructing high-fidelity flow fields, where both low- and high-fidelity data are generated from numerical solvers. Our experiments reveal that state-of-the-art models struggle to recover fine-grained high-fidelity details when using solver-generated low-fidelity inputs, due to distribution shift. To overcome this challenge, we introduce an \\textit{Importance Weight} strategy during training as self-guidance and a training-free \\textit{Residual Correction} method during inference as physical inductive bias, guiding the diffusion model toward higher-quality reconstructions. Experiments on four 2D turbulent flow datasets demonstrate the effectiveness of our proposed method."},"_bibtex":{"value":"@misc{\nli2025physicsinformed,\ntitle={Physics-Informed Self-Guided Diffusion Model for High-Fidelity Simulations},\nauthor={Ruoyan Li and Zijie Huang and Yizhou Sun and Wei Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=EaiU4F5pwn}\n}"},"title":{"value":"Physics-Informed Self-Guided Diffusion Model for High-Fidelity Simulations"},"pdf":{"value":"/pdf/41a8baae3bbf4450e9eab5f3ce90c57260b3836a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|physicsinformed_selfguided_diffusion_model_for_highfidelity_simulations"},"authorids":{"value":["~Ruoyan_Li1","~Zijie_Huang1","~Yizhou_Sun1","~Wei_Wang13"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruoyan Li","Zijie Huang","Yizhou Sun","Wei Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a novel method for representation learning of geometric simplicial complexes, in which a differential form parameterized by a multilayer perceptron is integrated over the simplices of an embedded complex, and these integrals are then read out into a representation of the complex that is not dependent on the size or dimension of the complex. This method is considered in a variety of tasks for processing simplicial complexes and graphs with geometric information."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The approach proposed by the authors is novel, yet simple, and rooted in a geometric view of learning for simplicial complex data.\n2. The proposed method is able to apply one learned function (k-form) to tasks over many geometric simplicial complexes, rather than only learning a function for a given complex."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This paper suffers from vagueness in some parts, particularly in the experiments section. This is my reason for rating the presentation as \"fair.\" Otherwise, the writing of the paper is good.\n2. Compared to message-passing schemes, the learned k-forms appear to be highly dependent on the particular geometric embedding of the complex, instead of the intrinsic topology/geometry.\n3. The type of tasks considered by the authors are not common ones in the literature, so I am not sure if there is a regime in which the proposed methods can be compared to existing simplicial neural networks, which are largely based on combinatorial, rather than geometric, information. However, there are message-passing graph neural networks that are designed to handle geometric information as well, which the authors did not compare to."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. My primary suggestions relate to the writing of the experiments section. There is not sufficient detail in Section 5 to understand each of the experiments. I think it would be better to shorten some of the background material in order to have space for better descriptions of your experiments. In particular:\n\na. Path classification: are the \"paths\" that you classify themselves simplicial complexes? I.e., is the idea to represent each path as a 1-d simplicial complex embedded in $\\mathbb{R}^2$, and then integrate against the simplices of the path? I see this to be the case in the appendix, but these details should really be in the main body of the paper for the purpose of readability.\n\nb. Visualizing simplicial Laplacian 1-eigenvectors: the description of this is far too short to understand what is going on here, and the utility of this is questionable. Comparing to the eigenvectors of the simplicial Laplacian is reasonable for the purpose of identifying (co)homological features of the complex, but wouldn't it be better to use neural k-forms to get a proxy for the intrinsic eigenfunctions of some manifold embedded in $R^n$ from which a complex is formed? This experiment needs a lot more motivation.\n\nc. Synthetic surface classification: again, having a description of the dataset in the main body of the paper is needed for readability here. It is fine to leave some details in the appendix, but the body of the paper should be sufficient to get a good understanding of what is going on.\n\n2. I have some questions about invariants that can be incorporated into these networks. As I commented in the weaknesses section, the function learned by neural k-forms is highly dependent on particular geometric embeddings of the simplicial complexes in Euclidean space, as opposed to the intrinsic geometry of the complex. For instance, if I apply a translation to the embedding of the complex, I am likely to get a completely different result when integrating the neural k-form against it. Do you have any comments or insights on handling issues like this? This relates also to your experiment on classifying molecular graphs: the graphs can be geometrically embedded in $\\mathbb{R}^3$, but is there a canonical rotation that these embeddings should have? Two identical molecules might be embedded in different ways that yield different results by the network, which is not a property that a message-passing network suffers from. Perhaps a group symmetry should be incorporated for this issue. See the following reference, for instance:\n\nHan, J., Rong, Y., Xu, T., & Huang, W. (2022). Geometrically equivariant graph neural networks: A survey. arXiv preprint arXiv:2202.07230.\n\nI think comparison between neural k-forms and geometric GNNs should be made, in order to yield a fairer picture of where your method stands for graph classification tasks.\n\n3. Can you please comment on whether or not simplicial complexes with non-geometric features could be incorporated into the method you propose? For instance, if the geometric complex not only has an embedding, but also a particular (co)chain supported on it supplied, how would one incorporate that information into a learning pipeline?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700715438408,"tcdate":1698776912968,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5340/Reviewer_gTeC"],"signatures":["ICLR.cc/2024/Conference/Submission5340/Reviewer_gTeC"],"forum":"Djw0XhjHZb","number":2,"license":"CC BY 4.0","cdate":1698776912968,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5340/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700715438408,"domain":"ICLR.cc/2024/Conference","replyto":"Djw0XhjHZb","id":"dZ5e5Zgxyy","forumContent":{"TLDR":{"value":"We learn differential $k$-forms on embedded graphs, leveraging a connection to singular cochains to obtain efficient, interpretable representations."},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["geometric deep learning","differential forms","representation learning","graph learning","geometry","topology"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Geometric deep learning extends deep learning to incorporate information about the geometry and topology data, especially in complex domains like graphs. Despite the popularity of message passing in this field, it has limitations such as the need for graph rewiring, ambiguity in interpreting data, and over-smoothing. In this paper, we take a different approach, focusing on leveraging geometric information from simplicial complexes embedded in $\\mathbb{R}^n$ using node coordinates. We use differential $k$-forms in $\\mathbb{R}^n$ to create representations of simplices, offering interpretability and geometric consistency without message passing. This approach also enables us to apply differential geometry tools and achieve universal approximation. Our method is efficient, versatile, and applicable to various input complexes, including graphs, simplicial complexes, and cell complexes. It outperforms existing message passing neural networks in harnessing information from geometrical graphs with node features serving as coordinates."},"_bibtex":{"value":"@inproceedings{\nmaggs2024simplicial,\ntitle={Simplicial Representation Learning with Neural \\$k\\$-Forms},\nauthor={Kelly Maggs and Celia Hacker and Bastian Rieck},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Djw0XhjHZb}\n}"},"title":{"value":"Simplicial Representation Learning with Neural $k$-Forms"},"pdf":{"value":"/pdf/09536d303a7f5b86bdae3f3e16a185c2eb2368ae.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"maggs|simplicial_representation_learning_with_neural_kforms"},"authorids":{"value":["~Kelly_Maggs1","~Celia_Hacker1","~Bastian_Rieck1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kelly Maggs","Celia Hacker","Bastian Rieck"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a novel method for explaining graph neural networks. The authors focus on counterfactual explanations and propose the UNR-Explainer, which aims to identify subgraphs that, when perturbed, lead to significant changes in node embeddings. The paper evaluates various explanation methods in unsupervised settings using synthetic and real-world datasets. The proposed method leverages the Monte Carlo Tree Search (MCTS) for efficient traversal in large search spaces. The paper also provides a theoretical analysis of the upper bound of Importance and discusses the algorithm for calculating Importance."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1)\tThe paper tackles CF reasoning in unsupervised settings, a relatively unexplored area potential implications for explainability in graph neural networks and unsupervised learning.\n2)\tThe paper leverages the Monte Carlo Tree Search (MCTS), a technique from reinforcement learning, to efficiently traverse the search space of potential subgraphs. MCTS is known for its effectiveness in large search spaces, making it a suitable choice for this problem.\n3)\tThe paper clearly defines the counterfactual property for unsupervised representation learning models, providing a solid foundation for their method.\n4)\tThe paper includes a theoretical analysis of the upper bound of Importance for GraphSAGE, adding a rigorous foundation to their empirical findings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1)\tWhile the paper does evaluate on both synthetic and real-world datasets, it might benefit from testing on more diverse datasets, especially those from different domains or with different characteristics. Information on how the method scales with larger datasets or more complex graphs, and its computational efficiency, would be valuable.\n2)\tI believe the paper would greatly benefit from additional visual illustrations or diagrams to depict the proposed method. Visual aids can provide a clearer understanding and offer readers an intuitive grasp of the methodology. Given the complexity and novelty of the approach, diagrams or flowcharts could enhance comprehension and make the content more accessible to a broader audience."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1)\tHow does the method scale with larger and more complex graphs? Are there any computational or memory constraints that might limit its applicability to very large datasets?\n2)\tHow sensitive is the method to the degree of perturbation applied to the subgraph? Would minor changes in perturbation lead to significantly different results in the algorithm of importance?\n3)\tGiven the contrastive approach employed by DGI and the inductive learning capability of GraphSAGE, how might these characteristics influence the types of counterfactual explanations generated? Furthermore, how would the proposed counterfactual explanation method adapt and perform when integrated with generative models such as GraphGAE or S2GAE?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637088885,"tcdate":1698756015955,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8687/Reviewer_8JCL"],"signatures":["ICLR.cc/2024/Conference/Submission8687/Reviewer_8JCL"],"forum":"0j9ZDzMPqr","number":3,"license":"CC BY 4.0","cdate":1698756015955,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8687/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637088885,"domain":"ICLR.cc/2024/Conference","replyto":"0j9ZDzMPqr","id":"VeDgHeveH3","forumContent":{"TLDR":{"value":"UNR-Explainer aims to provide counterfactual explanations for a single target node in unsupervised node representation models."},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["XAI","Unsupervised node representation learning","Counterfactual Explanations"]},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Node representation learning, such as Graph Neural Networks (GNNs), has become one of the important learning methods in machine learning, and the demand for reliable explanation generation is growing. Despite extensive research on explanation generation for supervised node representation learning, explaining unsupervised models has been less explored. To address this gap, we propose a method for generating counterfactual (CF) explanations in unsupervised node representation learning, aiming to identify the most important subgraphs that cause a significant change in the $k$-nearest neighbors of a node of interest in the learned embedding space upon perturbation. The $k$-nearest neighbor-based CF explanation method provides simple, yet pivotal, information for understanding unsupervised downstream tasks, such as top-$k$ link prediction and clustering. Furthermore, we introduce a Monte Carlo Tree Search (MCTS)-based explainability method for generating expressive CF explanations for **U**nsupervised **N**ode **R**epresentation learning methods, which we call **UNR-Explainer**. The proposed method demonstrates improved performance on six datasets for both unsupervised GraphSAGE and DGI."},"_bibtex":{"value":"@inproceedings{\nkang2024unrexplainer,\ntitle={{UNR}-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models},\nauthor={Hyunju Kang and Geonhee Han and Hogun Park},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=0j9ZDzMPqr}\n}"},"title":{"value":"UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning Models"},"pdf":{"value":"/pdf/790d3e0525600daa0b02aecf21fda646b3197859.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"kang|unrexplainer_counterfactual_explanations_for_unsupervised_node_representation_learning_models"},"authorids":{"value":["~Hyunju_Kang1","~Geonhee_Han1","~Hogun_Park2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hyunju Kang","Geonhee Han","Hogun Park"]}},"version":2},{"content":{"summary":{"value":"This paper introduces GraphPFN, a prior-data fitted network for node-level prediction that builds graph foundation models (GFMs) through synthetic pretraining. It generates diverse, realistic graphs using a hybrid of stochastic block models and preferential attachment, combined with graph-aware causal models for attributes and targets. By augmenting the tabular foundation model with attention-based neighborhood aggregation and training on these synthetic graphs, GraphPFN achieves strong performance."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"1. What is the initial motivation of applying synthetic data pretraining? Can it be replaced by real-world data, or a mix of both? \n2. How does the properties and amounts of pretraining data affect performance?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. This paper studies an interesting and practical research question of applying synthetic data in large-scale pretraining.\n2. The proposed data generation pipeline seems reasonable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The writing of this paper is confusing and the core contribution of the work is not well presented. In the introduction (line 066), the authors mentioned \"but such models still depend heavily on hand-crafted features and, as a result, are limited in their ability to capture complex graph patterns.\" This challenge is addressed by the proposed message-passing layer which replaces the structural features. However, the motivation of utilizing synthetic data pretraining is not discussed, which has no clear connection to the challenges mentioned. \n2. Efficiency concern: the proposed method builds on existing TFMs and requires post-training involving massive compute resources(8*A100*6days). At inference time, it involves dataset-specific fine-tuning to achieve the claimed SOTA performance. Given the lack of efficiency comparison, I doubt the real gains of the proposed method over baselines. \n3. Misleading experimental results: Table 2 compares GraphPFN with baselines. However, the gains of GraphPFN are inconsistent, with OOM issues which were not observed in baselines. Since a core edge of GraphPFN is that it does not rely on hand-crafted features, it is not reasonble to include the results for the LapPE enhanced version. As a result, the original version of the proposed method is less competitive, especially when compared with G2T. \n4. The data generation process involves many hyperparameters. Though the authors claim that they can be sampled from some distribution, the details are missing and the method still heavily relies on trial-and-error, limiting its real-world applicability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923033986,"tcdate":1761973895720,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12060/Reviewer_P18Y"],"signatures":["ICLR.cc/2026/Conference/Submission12060/Reviewer_P18Y"],"forum":"BLJ5DsJ0i6","number":4,"license":"CC BY 4.0","cdate":1761973895720,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12060/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923033986,"domain":"ICLR.cc/2026/Conference","replyto":"BLJ5DsJ0i6","id":"NeRpi4IM3t","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["graph foundation models","tabular foundation models","LimiX","graph neural network","graph machine learning"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Graph foundation models face several fundamental challenges including transferability across datasets and data scarcity, which calls into question the feasibility of graph foundation models at all.\nHowever, despite similar challenges, the tabular domain has recently witnessed the emergence of the first successful foundation models such as TabPFNv2 or LimiX.\nMany of these models are based on the prior-data fitted networks (PFN) framework, in which models are pretrained on carefully designed synthetic datasets to make predictions in an in-context learning regime.\nRecently, G2T-FM has made the first step towards adopting PFNs for graph tasks, yet it is limited to hand-crafted features and was never pretrained on graph data.\nIn this work, we make the next step by proposing GraphPFN, a PFN-based model designed and pretrained specifically for graphs.\nFollowing the PFN framework, we first design a prior distribution of synthetic attributed graphs by using a novel combination of multiple stochastic block models and a preferential attachment process for structure generation and graph-aware structured causal models for attribute generation.\nThen, we augment the tabular foundation model LimiX with attention-based graph neighborhood aggregation layers and train it on synthetic graphs sampled from our prior.\nOn diverse real-world graph datasets with up to $50{,}000$ nodes, GraphPFN shows strong in-context learning performance and achieves state-of-the-art results after finetuning, outperforming both G2T-FM and task-specific GNNs trained from scratch on most datasets.\nMore broadly, we hope that GraphPFN shows the potential of PFN-based models for building graph foundation models."},"_bibtex":{"value":"@misc{\neremeev2026graphpfn,\ntitle={Graph{PFN}: A Prior-Data Fitted Graph Foundation Model},\nauthor={Dmitry Eremeev and Oleg Platonov and Gleb Bazhenov and Artem Babenko and Liudmila Prokhorenkova},\nyear={2026},\nurl={https://openreview.net/forum?id=BLJ5DsJ0i6}\n}"},"title":{"value":"GraphPFN: A Prior-Data Fitted Graph Foundation Model"},"pdf":{"value":"/pdf/f4c614fb8e7378f47b8c42e087580845aadc49c0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"eremeev|graphpfn_a_priordata_fitted_graph_foundation_model"},"authorids":{"value":["~Dmitry_Eremeev1","~Oleg_Platonov1","~Gleb_Bazhenov1","~Artem_Babenko1","~Liudmila_Prokhorenkova1"]},"authors":{"value":["Dmitry Eremeev","Oleg Platonov","Gleb Bazhenov","Artem Babenko","Liudmila Prokhorenkova"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LoCoT2V-Bench, a new benchmark designed specifically for evaluating long video generation under complex input conditions.\nThe authors argue that existing benchmarks often rely on simplified prompts and focus on low-level metrics, neglecting higher-level dimensions such as narrative coherence and thematic expression. To address this gap, LoCoT2V-Bench presents two main contributions: (1) A suite of longer and more complex prompts derived from real-world videos, incorporating elements like scene transitions and event dynamics. (2) A multi-dimensional evaluation framework that, in addition to traditional metrics, proposes novel dimensions such as event-level alignment, content clarity, and HERD."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"**1. Well-motivated and Significant Problem**\n\nAs video generation models advance, the evaluation of longer, more complex videos has become a critical challenge. This paper accurately identifies the limitations of existing benchmarks and proposes a solution tailored for long video under complex prompts, which is crucial for advancing the field.\n\n**2. Comprehensive Evaluation Dimensions**\n\nThe evaluation is very comprehensive.  It not only covers fundamental metrics like Static Quality, Text-Video Alignment, and Temporal Quality but also innovatively introduces higher-level dimensions like Content Clarity and the Human Expectation Realization Degree. HERD, in particular, is a valuable exploration in video generation evaluation as it attempts to quantify abstract concepts like emotion, narrative, and character development.\n\n**3. Extensive Experiments and In-depth Analysis**\n\nThe authors have conducted a thorough evaluation of existing open-source LVG models. The evaluation not only reports overall performance but also delves into deeper exploration (e.g. Sec 4.2~4.4)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Limited Scale of the Benchmark**\n\nThe benchmark consists of 240 samples distributed across 18 themes. My main concern is that this sample size may be insufficient to draw robust evaluation about model capabilities.\n\n**2. Evaluation Reliability of HERD**\n\nThe proposed automated evaluation pipeline, especially the HERD metric, is heavily dependent on multiple third-party LLM/MLLMs. The generation and evaluation process for HERD involves a long chain of model calls. How do the authors evaluate and control the accumulated errors in this chain? How sensitive is the final metric to a potential failure at any stage of this pipeline?\n\n**3. Handling the Errors during Evaluation**\n\nFor fine-grained metrics like \"event-level temporal consistency,\" the calculation seems to presuppose that the corresponding events and subjects can be accurately located in the video. I think this assumption is questionable in complex scenarios. Also, there might be multiple subjects in the same video clip, or even in the same video frame. Does subject extraction confuse different subjects in these complex settings? \nIt is also unclear how the system robustly handles intermittent subject presence, such as when a subject is occluded or temporarily leaves the frame and then reappears."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943260136,"tcdate":1761033169928,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24955/Reviewer_RBq7"],"signatures":["ICLR.cc/2026/Conference/Submission24955/Reviewer_RBq7"],"forum":"YeWsA0VFZ5","number":1,"license":"CC BY 4.0","cdate":1761033169928,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24955/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943260136,"domain":"ICLR.cc/2026/Conference","replyto":"YeWsA0VFZ5","id":"1K7anfy7CP","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"LoCoT2V-Bench is a new benchmark for long-form text-to-video generation that uses complex prompts and multi-dimensional metrics."},"keywords":{"value":["Video Generation Benchmark; Text-to-Video Generation; Long-Form Video Evaluation; Multi-Dimensional Assessment"]},"supplementary_material":{"value":"/attachment/d4c5f787a3635c38c0e9337d5b475c0441f266a8.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recently text-to-video generation has made impressive progress in producing short, high-quality clips, but evaluating long-form outputs remains a major challenge especially when processing complex prompts. Existing benchmarks mostly rely on simplified prompts and focus on low-level metrics, overlooking fine-grained alignment with prompts and abstract dimensions such as narrative coherence and thematic expression. To address these gaps, we propose LoCoT2V-Bench, a benchmark specifically designed for long video generation (LVG) under complex input conditions. Based on various real-world videos, LoCoT2V-Bench introduces a suite of realistic and complex prompts incorporating elements like scene transitions and event dynamics. Moreover, it constructs a multi-dimensional evaluation framework that includes our newly proposed metrics such as event-level alignment, fine-grained temporal consistency, content clarity, and the Human Expectation Realization Degree (HERD) that focuses on more abstract attributes like narrative flow, emotional response, and character development. Using this framework, we conduct a comprehensive evaluation of nine representative LVG models, finding that while current methods perform well on basic visual and temporal aspects, they struggle with inter-event consistency, fine-grained alignment, and high-level thematic adherence, etc. Overall, LoCoT2V-Bench provides a comprehensive and reliable platform for evaluating long-form complex text-to-video generation and highlights critical directions for future method improvement."},"_bibtex":{"value":"@misc{\nzheng2026locotvbench,\ntitle={LoCoT2V-Bench: A Benchmark for Long-Form and Complex Text-to-Video Generation},\nauthor={Xiangqing Zheng and Chengyue Wu and Kehai Chen and Min Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=YeWsA0VFZ5}\n}"},"title":{"value":"LoCoT2V-Bench: A Benchmark for Long-Form and Complex Text-to-Video Generation"},"pdf":{"value":"/pdf/f90977ee1bad05e0857ccd66c143a3876f869feb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zheng|locot2vbench_a_benchmark_for_longform_and_complex_texttovideo_generation"},"authorids":{"value":["~Xiangqing_Zheng1","~Chengyue_Wu1","~Kehai_Chen2","~Min_Zhang9"]},"authors":{"value":["Xiangqing Zheng","Chengyue Wu","Kehai Chen","Min Zhang"]}},"version":2},{"content":{"summary":{"value":"This work presents a methodology design for looking into intuitive physics engine. A pouring-marble task is designed with various conditions and the results show some interesting behavior in cognitive strategies. Inspired by this, a framework called SHM is proposed for human mental simulation that aligns more precisely with human behavior."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"The modeling approach aims for a systematic methodology. How general is this model? Can this model handle some scenarios that the boundary cannot be described with a single parameter?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"The research topic is interesting and, compared with previous work, the scenario is more complicated and the experiment shows the effectiveness of the new modeling approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"An important contribution claimed in this paper is that, compared with previous works that mainly focused on a single task, this work provides a systematic methodology for learning heuristics. However, there is only one task in this paper although with varied conditions. I recommend adding another task with similar settings to show the general utility."}},"nonreaders":[],"tmdate":1731427290960,"tcdate":1730951368729,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission634/Reviewer_EBaN"],"signatures":["ICLR.cc/2025/Conference/Submission634/Reviewer_EBaN"],"forum":"BkeJro1xps","number":3,"license":"CC BY 4.0","cdate":1730951368729,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission634/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427290960,"domain":"ICLR.cc/2025/Conference","replyto":"BkeJro1xps","id":"gzlYSkcwm3","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Intuitive physics","physical reasoning","mental simulation","heuristic model"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The role of mental simulation in human behavior for various physical tasks is widely acknowledged, attributed to the generality of Intuitive Physics Engine (IPE). However, it remains unclear whether mental simulation is consistently employed across scenarios of different simulation costs and where its boundary is. Moreover, cognitive strategies beyond these boundaries have not been thoroughly investigated. Here, we adopted a pouring-marble task containing various conditions to study IPE's limits and strategies beyond. A human study revealed two distinct error patterns in predicting the pouring angle, differentiated by the simulation time using a boundary. This suggests a possible switching of the underlying reasoning strategies. Our initial experiment on IPE showed that its correlation with human judgments diminished in scenarios requiring extended time of simulation. This observation prompted the exploration of an alternative mechanism based on heuristics for intuitive physics. We uncovered that a linear heuristic model, relying exclusively on empirical data, replicated human prediction more accurately when the simulation time exceeded a certain boundary. Motivated by these observations, we propose a new framework, Simulation-Heuristics Model (SHM), which conceptualizes intuitive physics as a dual process: IPE is predominant only in short-time simulation, whereas a heuristics-based approach is applied as IPE's simulation time extends beyond the simulation boundary. The SHM model aligns more precisely with human behavior across various scenarios and demonstrates superior generalization capabilities under different conditions. Crucially, SHM integrates computational methods previously viewed as separate into a unified model, quantitatively studying their switching mechanism."},"_bibtex":{"value":"@misc{\nli2024a,\ntitle={A simulation-heuristics dual-process model for intuitive physics},\nauthor={Shiqian Li and Yuxi Ma and Bo Dai and Yujia Peng and Chi Zhang and Yixin Zhu},\nyear={2024},\nurl={https://openreview.net/forum?id=BkeJro1xps}\n}"},"title":{"value":"A simulation-heuristics dual-process model for intuitive physics"},"pdf":{"value":"/pdf/93856d41c3aa9841ddf51a3fb5488577d1272d6a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|a_simulationheuristics_dualprocess_model_for_intuitive_physics"},"authorids":{"value":["~Shiqian_Li1","~Yuxi_Ma2","~Bo_Dai5","~Yujia_Peng1","~Chi_Zhang12","~Yixin_Zhu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shiqian Li","Yuxi Ma","Bo Dai","Yujia Peng","Chi Zhang","Yixin Zhu"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the inefficiency of applying dense, frame-like processing to sparse event-based MLLMs. It makes two main contributions: (1) EventMind, a new 500k-sample instruction dataset for event vision, enabling a short-to-long curriculum learning strategy; and (2) EventFlash, an efficient MLLM using adaptive temporal (ATWA) and sparse spatial (SDGA) token sparsification(2). Experiments show EventFlash achieves a 12.4x throughput gain over its non-sparse baseline and can process much longer event sequences (1,000 bins)."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"See weaknesses above."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The creation of the 500k-sample EventMind dataset is a major contribution that addresses a critical resource gap for training and benchmarking event-based MLLMs.\n2. The paper targets the correct bottleneck: the inefficiency of applying dense methods to sparse data. The proposed spatiotemporal sparsification modules (ATWA and SDGA) are an intuitive and direct solution to this problem."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Insufficient comparison to SOTA event-based models: The paper fails to benchmark EventFlash against its direct competitors. On its new EventMind dataset, it only compares against frame-based models (Table 1). On the existing EventChat-Sub dataset, it only compares against EventGPT, omitting other SOTA event models like EventVL mentioned in the related work. This makes the SOTA performance claims unsubstantiated.\n2. Missing Key Methodological Ablations: The core Adaptive Temporal Window Aggregation (ATWA) module is a complex two-stage process (a spike-based merge followed by a semantic-based merge. However, the ablation study (Table 2) only validates the entire \"+T\" (Temporal) block at once. It never justifies the necessity of this complex two-stage design over a simpler, single-stage alternative."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925425297,"tcdate":1761986631932,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15105/Reviewer_UMdA"],"signatures":["ICLR.cc/2026/Conference/Submission15105/Reviewer_UMdA"],"forum":"QuvGqzLwf6","number":3,"license":"CC BY 4.0","cdate":1761986631932,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15105/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925425297,"domain":"ICLR.cc/2026/Conference","replyto":"QuvGqzLwf6","id":"jgmYj3fdwJ","forumContent":{"TLDR":{"value":"Event-Based Vision"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Event-Based Vision","Event-Language Alignment"]},"supplementary_material":{"value":"/attachment/7e8135b1e84d90445805cc4c4dfe02df8019de9a.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, current event-based MLLMs often rely on dense image-like processing paradigms, overlooking the spatiotemporal sparsity of event streams and resulting in high computational cost. In this paper, we propose EventFlash, the first efficient MLLM to explore spatiotemporal token sparsification for reducing data redundancy and accelerating inference. Technically, we first build EventMind, a large-scale and scene-diverse dataset with over 500k instruction sets, providing both short and long event stream sequences to support our curriculum training strategy. Then, we present the adaptive temporal window aggregation module for efficient temporal sampling, which adaptively compresses temporal tokens while retaining key temporal cues. Finally, the sparse density-guided attention module is designed to improve spatial token efficiency by selecting informative regions and suppressing empty or sparse areas. Experimental results show that EventFlash achieves a 12.4x throughput improvement over the baseline (EventFlash-Zero) while maintaining comparable performance. It supports long-range event stream processing with up to 1,000 bins, significantly outperforming EventGPT’s 5-bin limit. We believe EventFlash serves as an efficient foundation model for event-based vision."},"_bibtex":{"value":"@inproceedings{\nliu2026eventflash,\ntitle={EventFlash: Towards Efficient {MLLM}s for Event-Based Vision},\nauthor={Shaoyu Liu and Jianing Li and guanghui zhao and Yunjian Zhang and Wen Jiang and Ming Li and Xiangyang Ji},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=QuvGqzLwf6}\n}"},"title":{"value":"EventFlash: Towards Efficient MLLMs for Event-Based Vision"},"pdf":{"value":"/pdf/77ea58a96e431bd21922f488c307e8cc1e488ae6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|eventflash_towards_efficient_mllms_for_eventbased_vision"},"authorids":{"value":["~Shaoyu_Liu2","~Jianing_Li4","~guanghui_zhao1","~Yunjian_Zhang1","~Wen_Jiang9","~Ming_Li21","~Xiangyang_Ji1"]},"authors":{"value":["Shaoyu Liu","Jianing Li","guanghui zhao","Yunjian Zhang","Wen Jiang","Ming Li","Xiangyang Ji"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a Diffusion Transformer (DiT) based framework for pose-guided human image animation, specifically targeting complex, high-dynamic \"Hypermotions.\" The core novelty lies in the Spatial Low-Frequency Enhanced Rotary Positional Embedding (SLF-ROPE), designed to mitigate structural degradation during extreme movements. The authors also contribute a new dataset and benchmark, Open-HyperMotionX and HyperMotionX Bench. The proposed method shows impressive qualitative results and strong quantitative performance. The idea of explicitly addressing the stability issue in complex movements is valuable. However, significant concerns regarding the fairness of the comparison and the evaluation protocol must be addressed."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"In Table 3, the proposed method does not achieve best performance in terms of FID, VFID, and FVD. Could the authors explain this phenomenon?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.\tThe paper correctly identifies a major failure mode of current human animation models and discovers their inability to maintain fidelity and consistency during complex, non-standard movements. The goal is clear and impactful.\n2.\tSLF-ROPE and a new benchmark named Open-HyperMotionX are proposed.\n3.\tEstablishing a new benchmark for complex motion is necessary to drive future research in this domain."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThe primary concern is the potential unfairness of the comparison. Existing state-of-the-art methods are generally trained on large, general-purpose datasets (e.g., videos of daily life, simple movements), which predominantly feature simple motions. The proposed method is specifically trained on the newly introduced HyperMotionX dataset, which focuses on complex motions. If the videos used for evaluation (e.g., the supplementary videos) share a similar distribution or style to the videos in the HyperMotionX training set, the comparison is severely biased. The proposed method is essentially specialized for the test domain, while the baselines are being tested out-of-distribution (OOD) for complex motions.\n2.\tThe paper focuses heavily on complex motions. It is essential to demonstrate that the SLF-ROPE modification does not degrade performance or introduce artifacts when applied to the simple, common motions where existing methods already perform well. A comprehensive evaluation on a standard, general animation benchmark (e.g., TikTok, PATD) is mandatory.\n3.\tThe quality of the supplementary material is inconsistent. The IDs in the supplementary material is not consistently preserved well."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915718239,"tcdate":1761626198353,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1255/Reviewer_QtmM"],"signatures":["ICLR.cc/2026/Conference/Submission1255/Reviewer_QtmM"],"forum":"pzRDviUmbH","number":3,"license":"CC BY 4.0","cdate":1761626198353,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1255/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915718239,"domain":"ICLR.cc/2026/Conference","replyto":"pzRDviUmbH","id":"FfwnmTC4Ll","forumContent":{"TLDR":{"value":"DiT-based Pose-Guided Human  Image Animation"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Generation","Human Image Animation"]},"supplementary_material":{"value":"/attachment/914017d4ac5575fe6b734ac0142694d716dcd19c.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent animation sequences in regular motions and static scenes, there are still obvious limitations when facing complex human body motions that contain highly dynamic, non-standard motions, and the lack of a high-quality benchmark for evaluation of complex human motion animations. To address this challenge, we propose a simple yet powerful DiT-based video generation baseline and design spatial low-frequency enhanced RoPE, a novel module that selectively enhances low-frequency spatial feature modeling by introducing learnable frequency scaling. Furthermore, we introduce the Open-HyperMotionX Dataset and HyperMotionX Bench, which provide high-quality human pose annotations and curated video clips for evaluating and improving pose-guided human image animation models under complex human motion conditions. Our method significantly improves structural stability and appearance consistency in highly dynamic human motion sequences.  Extensive experiments demonstrate the effectiveness of our dataset and proposed approach in advancing the generation quality of complex human motion image animations. The codes and dataset will be made publicly available."},"_bibtex":{"value":"@misc{\nxu2026hypermotion,\ntitle={HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions},\nauthor={Shuolin Xu and Siming Zheng and Ziyi Wang and HC Yu and Jinwei Chen and Huaqi Zhang and Daquan Zhou and Bo Li and Peng-Tao Jiang},\nyear={2026},\nurl={https://openreview.net/forum?id=pzRDviUmbH}\n}"},"title":{"value":"HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions"},"pdf":{"value":"/pdf/605138d2cf34ffc4989039184adfaeb72620903f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"xu|hypermotion_ditbased_poseguided_human_image_animation_of_complex_motions"},"authorids":{"value":["~Shuolin_Xu1","~Siming_Zheng3","~Ziyi_Wang27","~HC_Yu1","~Jinwei_Chen3","~Huaqi_Zhang1","~Daquan_Zhou1","~Bo_Li20","~Peng-Tao_Jiang1"]},"authors":{"value":["Shuolin Xu","Siming Zheng","Ziyi Wang","HC Yu","Jinwei Chen","Huaqi Zhang","Daquan Zhou","Bo Li","Peng-Tao Jiang"]}},"version":2},{"content":{"summary":{"value":"This work introduces ESNv2, revisits ESN with a diagonal complex-valued recurrence to enable parallelism, but offers limited real innovation beyond reinterpreting existing state-space and recurrent ideas. \n\nThe main contribution includes: 1. using a diagonal complex-valued linear recurrence that allows parallel computation. 2. adding a nonlinear mixing layer for expressivity while keeping the readout as the only trainable part."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.  Please use more solid and credible experiments to demonstrate the effectiveness of the proposed solution. The currently used dataset is too old and too small. It's even a benchmark Lorenz96 published in 1996. While the largest classification dataset is MINST. This cannot represent the latest research progress at all.\n\n2. These mini datasets cannot demonstrate the efficient and parallelized effects that the author claims to have proposed. In the current situation, I suggest the author consider ImageNet[1] as a baseline for classification task and PILE[2] for sequence modelling. \n\n\n[1] Deng, Jia, et al. \"Imagenet: A large-scale hierarchical image database.\" 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009.\n\n[2] Gao, Leo, et al. \"The pile: An 800gb dataset of diverse text for language modeling.\" arXiv preprint arXiv:2101.00027 (2020)."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This work has a clear motivation. The author figure out the the lack of parallelism in RCs and attempt to address it.\n\n2. This work is well structured and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the method achieves faster training by eliminating sequential dependencies, this improvement comes at the expense of reduced model expressiveness, since the diagonal recurrence structure limits the network’s ability to capture complex temporal interactions.\n\n\nThe main problem is that the experiments only demonstrated mainly on simplistic synthetic benchmarks such as memory and forecasting tasks; moreover, comparisons with modern deep models like Transformers or LRUs appear shallow and unconvincing, as ESNv2 lacks the flexibility and scalability required for real-world applications. \n\nSince the author has been benchmarking against the concept of SSM, I suggest the author test the real capability on text datasets.\n\nTherefore, I believe this work does not meet the acceptance criteria"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359634105,"tcdate":1760533817929,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7872/Reviewer_FkkR"],"signatures":["ICLR.cc/2026/Conference/Submission7872/Reviewer_FkkR"],"forum":"N6G2Mmz8qs","number":1,"license":"CC BY 4.0","cdate":1760533817929,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7872/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359634105,"domain":"ICLR.cc/2026/Conference","replyto":"N6G2Mmz8qs","id":"dZwlwiLxIP","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"Introducting a novel framework for the construction of efficient, randomized RNNs based on diagonal linear recurrence in the complex space."},"keywords":{"value":["reservoir compting; echo state networks; recurrent neural networks;"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Reservoir Computing (RC) has established itself as an efficient paradigm for temporal processing, yet its scalability remains severely constrained by the necessity of processing temporal data sequentially. In this work, we revisit RC through the lens of structured operators and state space modeling, introducing Parallel Echo State Network (ParalESN), a framework that enables the construction of efficient reservoirs with diagonal linear recurrence in the complex space that can be parallelized during training. We provide a theoretical analysis demonstrating that ParalESN preserves the Echo State Property and the universality guarantees of classical Echo State Networks while admitting an equivalent representation of arbitrary linear reservoirs in the complex diagonal form. Empirically, ParalESN attains comparable predictive accuracy to traditional RC on memory and forecasting benchmarks, while delivering substantial gains in training efficiency. On 1-D pixel-level classification tasks, the model achieves competitive accuracy with fully trainable networks, reducing computational costs and energy consumption. Overall, ParalESN offers a promising, scalable, and principled pathway for integrating RC within the deep learning landscape."},"_bibtex":{"value":"@misc{\nlagomarsini2026esnv,\ntitle={{ESN}v2: Resurrecting Reservoir Computing in the Deep Learning era},\nauthor={Giacomo Lagomarsini and Matteo Pinna and Andrea Ceni and Claudio Gallicchio},\nyear={2026},\nurl={https://openreview.net/forum?id=N6G2Mmz8qs}\n}"},"title":{"value":"ESNv2: Resurrecting Reservoir Computing in the Deep Learning era"},"pdf":{"value":"/pdf/e5ded58a3f760d3c2ab7d6fe259661595c31c6d3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"lagomarsini|esnv2_resurrecting_reservoir_computing_in_the_deep_learning_era"},"authorids":{"value":["~Giacomo_Lagomarsini1","~Matteo_Pinna1","~Andrea_Ceni1","~Claudio_Gallicchio1"]},"authors":{"value":["Giacomo Lagomarsini","Matteo Pinna","Andrea Ceni","Claudio Gallicchio"]}},"version":2},{"content":{"TLDR":{"value":"We provide BuildArena, a physics‑aligned interactive benchmark that tests the engineering construction capabilities of frontier LLMs."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Engineering construction","LLM","benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Engineering construction automation aims to transform natural language specifications into physically viable structures, requiring complex integrated reasoning under strict physical constraints. While modern LLMs possess broad knowledge and strong reasoning capabilities that make them promising candidates for this domain, their construction competencies remain largely unevaluated. To address this gap, we introduce BuildArena, the first physics-aligned interactive benchmark designed for language-driven engineering construction. It takes a first step towards engineering automation using LLMs. Technically, it contributes to the community in two aspects: (1) an extendable task design strategy spanning static and dynamic mechanics across multiple difficulty tiers; (2) a 3D Spatial Geometric Computation Library for supporting construction based on language instructions. On eight frontier LLMs, BuildArena comprehensively evaluates their capabilities for language-driven and physics-grounded construction automation. We release the code at https://anonymous.4open.science/r/BuildArena-9B7B/ to benefit construction automation in engineering applications."},"_bibtex":{"value":"@misc{\nanonymous2026buildarena,\ntitle={BuildArena: A Physics\\nobreakdash-Aligned Interactive Benchmark of {LLM}s for Engineering Construction},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=oml3PWSYcc}\n}"},"title":{"value":"BuildArena: A Physics‑Aligned Interactive Benchmark of LLMs for Engineering Construction"},"pdf":{"value":"/pdf/26c6163a776c85a6701b6693a31ffcf1c6edc130.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"xia|buildarena_a_physicsaligned_interactive_benchmark_of_llms_for_engineering_construction"},"authorids":{"value":["~Tian_Xia15","~Tianrun_Gao3","~Wenhao_Deng2","~Long_Wei1","~Xiaowei_Qian3","~Jiang_Yixian1","~Chenglei_Yu1","~Tailin_Wu1"]},"authors":{"value":["Tian Xia","Tianrun Gao","Wenhao Deng","Long Wei","Xiaowei Qian","Jiang Yixian","Chenglei Yu","Tailin Wu"]}},"tmdate":1770907888874,"tcdate":1757235317644,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2752/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission2752/Authors"],"forum":"oml3PWSYcc","license":"CC BY 4.0","number":2752,"cdate":1757235317644,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission2752/-/Full_Submission","ICLR.cc/2026/Conference/Submission2752/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Desk_Rejected_Submission","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770907888874,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"oml3PWSYcc","version":2},{"content":{"summary":{"value":"The paper introduces JustLogic, a benchmark designed to test LLMs' deductive reasoning without the bias of prior knowledge. It highlights the shortcomings of the exitsting logical benchmarks based on complexity and error analysis -- showing some of the sythetic datasets, while may not rely on prior knowledge, are not complex enough. Other datasets, are complex but may be solve using prior human knowledge. Authors curate a large number of logical derivations, using a mix of LLM and human generated templates and propositions from GenericsKB, authors create a logical benchmark, which is synthetic, (possibly) not true in the real world and complex.  Despite improved performance by models like OpenAI o1-preview, they still lag behind human reasoning. JustLogic aims to drive deeper evaluation and enhancement of LLM capabilities."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Context-ndependent is mentioned in abstract/intro many times. It is unclear how the authors ensured. Seemed like nothing special has been done.\n2. Why is the average human performance so low? If you recruit more trained puzzle solvers, such as say linguistic olympiad participants or winners, I believe they may do much better, In that case, this false supremacy of o1 should not be there.\n3. L100: Why is it novel? I did not understand what is claimed to be novel here at all.\n4. Tab 1: It will be better to clarify what does complexity mean in statistical terms, number of clauses, words etc.\n5. L290: This Step 3 is written in a very convouted way or is it the process? Are we saying we have a conclusion and we do not know using logic, whether the conclusion is entailed or not. This can not be true.\nThen why a paragraph is randomly assigned an answer? Please explain clearly.\n6. L317: How is the number of domains measured? Is it from GenericsKB? \nThe vocab seems to have been created by a mix of chosen templates and words from GenericsKB. Hard to say that it is significantly complex -- give GKB sentences are quite simple and there are limited templates authors curated.\n7. L333: This may not be that simple. One may be able to prove the negation in a smaller set of steps. Is that considered as well? Or, the dataset does not have such examples?\n8. L350-353: Very interesting paragraph. Again how do authors ensure the factual incorrectness of the conclusions? \n9. L442:Why is the performance on CLUTTRR so low compared to JustLogic?\n10. L514: For the apparently long tail argument forms, have you qualitatively evaluated the CoT or some other ways to actually see what reasoning is being done by the LLMs?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper highlights flaws in current reasoning benchmarks: low complexity, prior knowledge dependence, and limited error analysis.\n2. It addresses the flaws highlighted in point 1 while constructing the new dataset JustLogic that have been unaddressed in datasets like FOLIO, ProofWriter, etc.\n3. The solution to the flaws like focusing on argument complexity and Prior knowledge independence looks interesting.\n4. Very minimal manual effort is required to construct the dataset.\n\nAuthors also aspire to make it more future-proof by leveraging synthetic creation abilities."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There are several points which are claimed but do not have enough justification, for example, it may not be naturally true that a \"conclusion\" is untrue in the real world just because it is synthetic. No effort have been made to enforce factual inaccuracies -- they should also compare with PrOntoQA, which created a \"Fictonal\" ontology and a \" false\" ontology for this purpose. I believe, such a process is required, if you really want to ensure that prior knowledge will not affect the results.\n2. Overall the conclusion that every LLM except o1 is lagging -- seems to not add much to the ongoing body of knowledge. Why are some of the derivations difficult vs easy. Only 1/2 line of explanations are given, that too in a hand-wavy manner (L513)\n3. There are some issues with authors arguments about \"flaws\" in the benchmarks. See questions."}},"nonreaders":[],"tmdate":1732456309117,"tcdate":1730703781257,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6386/Reviewer_NSeq"],"signatures":["ICLR.cc/2025/Conference/Submission6386/Reviewer_NSeq"],"forum":"THSm9HyCKo","number":3,"license":"CC BY 4.0","cdate":1730703781257,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6386/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732456309117,"domain":"ICLR.cc/2025/Conference","replyto":"THSm9HyCKo","id":"mkwzljC81d","forumContent":{"TLDR":{"value":"We present JustLogic, a benchmark to measure deductive reasoning capabilities of LLMs, that is more challenging, reliable, and insightful than existing benchmarks."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["benchmark","logical reasoning","LLM","natural language processing (NLP)","propositional logic"]},"supplementary_material":{"value":"/attachment/6a5873f5b0206f9cb1b80e039529d0be7ca696dc.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Logical reasoning is a critical component of Large Language Models (LLMs), and substantial research efforts in recent years have aimed to enhance their deductive capabilities. However, existing deductive reasoning benchmarks, which are crucial for evaluating and advancing LLMs, are inadequate due to their lack of task complexity, presence of prior knowledge as a confounder, and superficial error analysis. To address these deficiencies, we introduce JustLogic, a synthetically generated deductive reasoning benchmark designed for rigorous evaluation of LLMs. JustLogic is (i) highly complex, capable of generating a diverse range of linguistic patterns, vocabulary, and argument structures; (ii) context-independent, eliminating the advantage of models possessing prior knowledge and ensuring that only deductive reasoning is used to answer questions; and (iii) capable of in-depth error analysis on the heterogeneous effects of reasoning depth and argument form on model accuracy. Our experimental results on JustLogic reveal that the performance of most state-of-the-art (SOTA) LLMs, specifically Llama3-8B (57.8\\%), Llama3-70B (64.6\\%), and GPT-4o (65.6\\%), is significantly worse than the average human performance (73.0\\%). A recently released reasoning model, OpenAI o1-preview, performed substantially better, with an accuracy of 81.0\\%. However, it still lags behind the human ceiling of 100.0\\%. These results demonstrate that the JustLogic benchmark is realistic and achievable for both humans and models and that there is still substantial room for improvement in the deductive reasoning capabilities of LLMs. We posit that the use of context-dependent and relatively simplistic benchmarks has misrepresented the reasoning abilities of many SOTA models. We release our open-source dataset to provide accurate evaluations of model performance in deductive reasoning and to facilitate LLM advancement through in-depth error analysis."},"_bibtex":{"value":"@misc{\nchen2025justlogic,\ntitle={JustLogic: A benchmark for natural language deductive reasoning},\nauthor={Michael K. Chen and Xikun ZHANG and Dacheng Tao},\nyear={2025},\nurl={https://openreview.net/forum?id=THSm9HyCKo}\n}"},"title":{"value":"JustLogic: A benchmark for natural language deductive reasoning"},"pdf":{"value":"/pdf/7c281a6490f4adeba36db64d521506dee459fa74.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|justlogic_a_benchmark_for_natural_language_deductive_reasoning"},"authorids":{"value":["~Michael_K._Chen1","~Xikun_ZHANG2","~Dacheng_Tao1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Michael K. Chen","Xikun ZHANG","Dacheng Tao"]}},"version":2},{"content":{"summary":{"value":"This paper studies an iterative approach to synthetic data generation for fine-tuning small language models (SLMs). Instead of generating a large static dataset from a teacher model at once, the authors propose a closed-loop scheme where the student model’s current state guides which examples are selected for further data generation by the teacher.\nThe work benchmarks several existing data selection strategies from active learning—uncertainty sampling (high-loss), reward-based scoring, LLM-as-a-judge, and BADGE diversity selection—and claims that simple heuristics like high-loss sampling outperform more expensive LLM-judge–based methods. Experiments are conducted on four reasoning datasets (GSM8K, Math1–3, ProntoQA, Game of 24) using various small instruction-tuned models. Results suggest that simple heuristics such as high-loss selection outperform more complex and expensive methods like LLM-as-a-judge, offering improved data efficiency under a fixed compute budget."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.\tHow do you ensure that the iterative selection process does not bias the dataset, e.g., easier questions?\n2.\tWhy does the performance of your proposed method have such a better performance, even better than the teacher model, on the Game of 24 dataset? \n3.\tHow would the method behave with a larger teacher model?\n4.\tIn your comparison, you do not control the SFT dataset size. Will that cause unfair comparison among different methods?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"•\tThe authors provide a solid benchmark showing that simple, inexpensive heuristics (e.g., high-loss selection) can outperform more complex LLM-as-a-judge strategies.\n•\tThe work provides practical guidance for synthetic data generation under constrained compute budgets, which can be valuable for practitioners training SLMs.\n•\tThe paper is clearly written and easy to reproduce.\n•\tIt contributes to the empirical understanding of how different data-selection heuristics impact fine-tuning performance and efficiency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"•\tLimited novelty: The core idea—iterative, student-aware synthetic data generation—has been explored in multiple prior works. This paper mainly repackages it under the active learning perspective.\n•\tLack of theoretical or conceptual insight: The paper does not explain why the compared heuristics differ or what properties (difficulty, diversity, informativeness) they capture.\n•\tMarginal performance gains: Improvements are small or inconsistent. For GSM8K and ProntoQA, performance remains below or comparable to prior SFT results; only one dataset (Game of 24) shows notable gains.\n•\tScope limitation: All experiments are conducted on small 7–8B models; scalability to larger models or broader domains is untested.\n•\tPotential bias amplification: Since selection is based on student performance, the loop can reinforce sampling bias (e.g., favoring easy samples), which the paper neither analyzes nor mitigates.\n•\tUnclear takeaway: The results show minor absolute improvements, so the main claimed advantage, data efficiency, needs stronger quantitative justification."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924201783,"tcdate":1761673857481,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13622/Reviewer_3RBg"],"signatures":["ICLR.cc/2026/Conference/Submission13622/Reviewer_3RBg"],"forum":"U0I590wrsm","number":1,"license":"CC BY 4.0","cdate":1761673857481,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13622/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924201783,"domain":"ICLR.cc/2026/Conference","replyto":"U0I590wrsm","id":"2ndv0NO3rd","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"Generating synthetic data conditioned on the student model enables more data efficient supervised finetuning."},"keywords":{"value":["Synthetic data generation","active learning","language models","supervised finetuning"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"A common and effective means for improving language model capabilities involves finetuning a “student” language model’s parameters on generations from a more proficient “teacher” model. Termed “synthetic data”, these generations are often produced before any student finetuning, but some work has considered generating new synthetic samples as training progresses. This paper studies and advocates for the latter case, where data are generated in an iterative, closed-loop fashion that is guided by the current state of the student model. For a fixed budget of generated samples, or a budget in terms of compute spent querying a teacher, we show that this curation of finetuning data affords improved student performance over static generation. Further, while there have been several LLM-specific methods proposed that operate in this regime, we find that simple, inexpensive selection criteria from the active learning literature tend to be most performant. We validate these claims across four mathematical and logical reasoning datasets using four different small language models."},"_bibtex":{"value":"@misc{\nkessler2026towards,\ntitle={Towards Active Synthetic Data Generation for Finetuning Language Models},\nauthor={Samuel Kessler and Menglin Xia and Daniel Madrigal and Dongge Han and Helia Hashemi and Saravan Rajmohan and Victor R{\\\"u}hle and Jordan T. Ash},\nyear={2026},\nurl={https://openreview.net/forum?id=U0I590wrsm}\n}"},"title":{"value":"Towards Active Synthetic Data Generation for Finetuning Language Models"},"pdf":{"value":"/pdf/e926528dddf04a02d5114c5187c9af59858f0b3c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kessler|towards_active_synthetic_data_generation_for_finetuning_language_models"},"authorids":{"value":["~Samuel_Kessler1","~Menglin_Xia1","~Daniel_Madrigal1","~Dongge_Han1","~Helia_Hashemi1","~Saravan_Rajmohan2","~Victor_Rühle1","~Jordan_T._Ash1"]},"authors":{"value":["Samuel Kessler","Menglin Xia","Daniel Madrigal","Dongge Han","Helia Hashemi","Saravan Rajmohan","Victor Rühle","Jordan T. Ash"]}},"version":2},{"content":{"summary":{"value":"The authors propose ModelBench, a benchmark for evaluating end-to-end AI-based extraction of physical models from physics literature. In this first version of ModelBench, the authors choose a subset of 20 papers in photonics, from which experts extracted gold models (Python scripts) reproducing the experiments in these papers. A candidate submission consists of a model implementation (compared to the reference one), plots of the results and goodness fit metrics. For each paper, the benchmark contains the gold model along with a set of evaluation criteria (weighted, hierarchical rubrics with binary responses) that stress-test a candidate submission's correctness.\nModelBench is model-agnostic and can work using LLM-as-a-judge evaluation as well as human feedback to produce an evaluation score. The authors plan on extending ModelBench with more models from various domains of physics."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Was a consistent team of experts involved in validating the rubrics? If not, was any cross-validation process involved to ensure that different experts would create similar rubrics for the same paper?\n- As mentioned above, scaling ModelBench up seems like a major hurdle. Do the authors plan to address this in future releases, and how? In particular with respect to the rubric generation mentioned above.\n- Why were closed papers chosen as part of the initial release? Was this done for realism reasons (as much of physics literature is locked behind paywalls)? Are those papers fundamentally more relevant?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- This paper's main contribution is the evaluation gap it is filling. As argued by the authors, expert-based model extraction from physics literature is slow and costly, and while the physics community has steadily been adopting AI-based extraction tools for this purpose, the lack of gold standard evaluation benchmarks makes it hard to quantify the reliability of these methods. ModelBench aims to fill this gap by proposing such a benchmark.\n- ModelBench can be seen as a harder, physics-oriented version of PaperBench. ModelBench is completely end-to-end, requiring models to work from unannotated physics papers. In this, it mirrors the task of human researchers who need to infer real-world parameters from the often incomplete/implicit descriptions found in the literature.\n- The well-structured and model-agnostic nature of the benchmark ensures that evaluation can be systematically performed across various models and systems.\n- A considerable amount of work has been put into extracting gold models from articles and carefully designing relevant evaluation points for each of them, making the data of the benchmark itself a valuable resource.\n- ModelBench does not claim to solve the reproducibility crisis, but is a pragmatic approach to leveraging the implicit assumptions used in research articles. Tools like ModelBench could eventually contribute to solving that root cause by validating that a future AI model is able to fully leverage those hidden assumptions and parameters.\n- Among the evaluation criteria are explicit checks that the produced model follows physical constraints, such as energy conservation, and can reproduce experimental results. This is missing in previous works, such as PaperBench (notably due to the inherent differences in the targeted domains)\n- The experimental results on GPT-5 and Claude Opus 4.1 provide some initial insights into the limitations of current LLMs for scientific modeling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- In the initial release, 5 out of 21 papers are not available in open-access. While 75% of the benchmark remains freely available, this could constitute a significant hurdle that makes the benchmark harder to use, and makes its installation non-automatable, and more time-consuming for future releases. In addition, building a benchmark for reproducibility that relies on non-reproducible (inaccessible) sources is contradictory and runs counter to the increasingly strong movement towards open science. There are (valid) justifications for this choice, but it is hard to understand why this limitation is not mentioned or justified in the paper.\n- While being model-agnostic is a strength in terms of flexibility, it also comes with several issues: The evaluation process requires a judge, which can either be an automated script or another LLM, to assign binary scores to rubric items. The authors correctly acknowledge that judge variance remains an open issue. A more thorough investigation of the reliability and biases of using an LLM as a judge is needed. The authors use human audits as a first good step, but a more detailed analysis of inter-judge agreement and potential calibration methods would strengthen the evaluation protocol.\n- Creating gold-standard models and detailed rubrics by domain experts is very labor-intensive. The resulting benchmark is of high quality, but this also makes scaling the benchmark to a large number of tasks and domains challenging, which is not discussed in the paper. A discussion on potential strategies to scale this process in the future would be beneficial to the paper (partial automation? streamlining the process using a crowdsourced platform? etc.).\n- The hierarchical rubric provides a structured evaluation, and different criteria seem to be used for different papers. However, as the benchmark scales up, there is a risk that future AI systems could be specifically optimized to perform well on the rubric's criteria without necessarily achieving a deeper scientific understanding, which may eventually pose issues. The authors could consider incorporating more open-ended evaluation metrics or human-in-the-loop assessments to mitigate this risk.\n- The initial release of ModelBench is focused on photonic integrated circuits, with a dataset of 20 tasks. This limitation is acknowledged by the authors, who mention future expansion plans. However, the current narrow scope may limit the generalizability of the findings to other domains of physics. Future versions should prioritize a broader range of physics and engineering problems to demonstrate the benchmark's versatility."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920986277,"tcdate":1761928350583,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9367/Reviewer_rfsS"],"signatures":["ICLR.cc/2026/Conference/Submission9367/Reviewer_rfsS"],"forum":"Bw9LCBz9KW","number":1,"license":"CC BY 4.0","cdate":1761928350583,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9367/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920986277,"domain":"ICLR.cc/2026/Conference","replyto":"Bw9LCBz9KW","id":"i0fTq2kpPN","forumContent":{"TLDR":{"value":"ModelBench is a benchmark for testing whether AI systems can read physics papers and produce executable, physics-based models."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Scientific AI benchmarks","Physics","LLM-as-judge","Rubric-based evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We introduce **ModelBench**, a benchmark for evaluating whether AI systems can extract\nexecutable physics-based models from scientific literature. ModelBench couples\n(i) gold-standard reference models,\n(ii) a hierarchical, weighted binary rubric covering physics correctness, code quality, and reproduction quality, and\n(iii) a judge protocol that produces pass/fail scores at rubric leaves.\nUnlike code-generation benchmarks that test function-level correctness, ModelBench\ntargets the end-to-end task of reconstructing physically grounded models from incomplete and underspecified scientific descriptions.\nWe release the benchmark specification, rubric generator and judge prompts,\nand an initial set of 20 gold models within the field of photonic integrated circuits, alongside scripts for fully reproducible evaluation.\nCandidate systems are required to produce a Python implementation of the model,\na plot of the fitted results, and evaluate MSE and $R^2$ metric of the fit.\nUsing general-purpose LLMs as neutral baselines, we report aggregate scores and case studies that reveal common failure modes\n(e.g., constraint violations, phenomenological overfitting) and show how rubric structure aids in diagnostic evaluation.\nWe discuss limitations (judge variance, dataset breadth, implicit-knowledge gaps) and outline a roadmap to expand domains,\ntighten constraint checking, and support multiple valid solutions. ModelBench provides a transparent platform\nfor tracking scientific modeling capabilities in AI under physical and empirical constraints."},"_bibtex":{"value":"@misc{\nschoolkate2026modelbench,\ntitle={ModelBench: A Benchmark for Extracting Executable, Physics-Based Models from Scientific Literature},\nauthor={Pim Schoolkate and Patrick Huembeli and Frank Sch{\\\"a}fer and Krystian Nowakowski and Carlos Arribalzaga Jov{\\'e} and Frank Koppens and Dirk Englund and Jacob M. Taylor},\nyear={2026},\nurl={https://openreview.net/forum?id=Bw9LCBz9KW}\n}"},"title":{"value":"ModelBench: A Benchmark for Extracting Executable, Physics-Based Models from Scientific Literature"},"pdf":{"value":"/pdf/76d73ba2e8b90fc035df3773654a2a2f446bc09b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"schoolkate|modelbench_a_benchmark_for_extracting_executable_physicsbased_models_from_scientific_literature"},"authorids":{"value":["~Pim_Schoolkate1","~Patrick_Huembeli1","~Frank_Schäfer1","~Krystian_Nowakowski1","~Carlos_Arribalzaga_Jové1","~Frank_Koppens1","~Dirk_Englund1","~Jacob_M._Taylor1"]},"authors":{"value":["Pim Schoolkate","Patrick Huembeli","Frank Schäfer","Krystian Nowakowski","Carlos Arribalzaga Jové","Frank Koppens","Dirk Englund","Jacob M. Taylor"]}},"version":2},{"content":{"summary":{"value":"In this paper, the authors extend the contextual dueling bandit formulation to neural bandit settings, where the reward function is arbitrary and does not assume linearity. To address this task, they propose two methods based on the existing NeuralUCB and NeuralTS algorithms. Both theoretical and empirical analyses on synthetic datasets are provided."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to \"Weaknesses\" above."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The application of neural bandits to solve real-world dueling bandit problems is intuitive, as existing linear contextual dueling bandit algorithms may fail to capture complex reward relationships.\n\nAdditionally, the analysis of neural bandits with binary feedback may be of particular interest to the bandit research community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Firstly, my concerns are with the problem formulation and theoretical analysis. I will list a few below. For instance:\n\nIf my understanding is correct, the problem formulation heavily relies on the link function $\\mu$, whose properties significantly influence the regret bound, particularly through terms such as $\\kappa\\_{\\mu}$ and $L\\_{\\mu}$. Furthermore, unlike conventional stochastic bandit works, this paper assumes the reward feedback associates with the sub-Gaussian noise, which in this case, depends on the chosen link function. Typically in stochastic bandit settings, sub-Gaussian noise is independent of the learning formulation and is determined solely by environmental settings, which makes this formulation unnatural.\n\nAdditionally, in Assumption 1, $\\mathcal{X}$ is not clearly defined. Does $\\mathcal{X}$ represent the entire arm context space with $\\mathcal{X}\\_t \\subset \\mathcal{X}$? If so, how would the learner possess prior knowledge of $\\mathcal{X}$ to select the optimal $\\kappa\\_{\\mu}$ value for exploration (i.e., the $\\beta\\_T$ term in Equation 3)? A similar issue also exists for $\\tilde{d}$, as it is unrealistic to assume the learner has prior knowledge of this parameter.\n\nThe theoretical analysis appears to rely heavily on the proof framework of NeuralUCB and NeuralTS, adapting assumptions like the positive definite NTK matrix for arm context. The authors also mention the $c_0$ term, representing gradient differences between arm contexts in Theorem 2. This term should be clearly shown in the regret bound, as its data dependence could lead it to grow with $T$ and should not be hidden by big-O notation.\n\nAll experiments are conducted on synthetic datasets with predefined configurations. While the authors claim the paper is motivated by dueling bandit applications in LLM tuning, no experiments are included to support this claim. Furthermore, the theoretical analysis may not align with real RLHF settings, particularly given the strong over-parameterization assumption, which may not be realistic. In practice, each arm could be a generated response from an LLM, with high-dimensional vector representations and complex reward mappings. While RLHF methods like DPO employ powerful neural architectures, the analysis in this paper assumes an MLP, which may be insufficient. Therefore, I believe real dataset experiments relevant to LLM tuning, as indicated by the authors, are necessary."}},"nonreaders":[],"tmdate":1731428771999,"tcdate":1730581550115,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13697/Reviewer_uMSZ"],"signatures":["ICLR.cc/2025/Conference/Submission13697/Reviewer_uMSZ"],"forum":"VELhv9BBfn","number":3,"license":"CC BY 4.0","cdate":1730581550115,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13697/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428771999,"domain":"ICLR.cc/2025/Conference","replyto":"VELhv9BBfn","id":"sH3GMy8G3h","forumContent":{"TLDR":{"value":"We study contextual dueling bandits problem and propose upper confidence bound- and Thompson sampling-based algorithms that use a neural network to estimate the reward function using human preference feedback and have sub-linear regret guarantees."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Contextual Dueling Bandits","Preferences Learning","Human Feedback","Neural Bandits","Thompson Sampling"]},"supplementary_material":{"value":"/attachment/19ed90ef9b8272e3593b14e3ac79cbdd8d5f4e22.zip"},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Contextual dueling bandit is used to model the bandit problems, where a learner's goal is to find the best arm for a given context using observed noisy human preference feedback over the selected arms for the past contexts. However, existing algorithms assume the reward function is linear, which can be complex and non-linear in many real-life applications like online recommendations or ranking web search results. To overcome this challenge, we use a neural network to estimate the reward function using preference feedback for the previously selected arms. We propose upper confidence bound- and Thompson sampling-based algorithms with sub-linear regret guarantees that efficiently select arms in each round. We also extend our theoretical results to contextual bandit problems with binary feedback, which is in itself a non-trivial contribution. Experimental results on the problem instances derived from synthetic datasets corroborate our theoretical results."},"_bibtex":{"value":"@inproceedings{\nverma2025neural,\ntitle={Neural Dueling Bandits: Preference-Based Optimization with Human Feedback},\nauthor={Arun Verma and Zhongxiang Dai and Xiaoqiang Lin and Patrick Jaillet and Bryan Kian Hsiang Low},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=VELhv9BBfn}\n}"},"title":{"value":"Neural Dueling Bandits: Preference-Based Optimization with Human Feedback"},"pdf":{"value":"/pdf/875613f857a562bc6de9f80ec0421b5e179060b8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"verma|neural_dueling_bandits_preferencebased_optimization_with_human_feedback"},"authorids":{"value":["~Arun_Verma1","~Zhongxiang_Dai1","~Xiaoqiang_Lin1","~Patrick_Jaillet1","~Bryan_Kian_Hsiang_Low1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Arun Verma","Zhongxiang Dai","Xiaoqiang Lin","Patrick Jaillet","Bryan Kian Hsiang Low"]}},"version":2},{"content":{"comment":{"value":"We clarify that the statement “when the local data heterogeneity is severe, the model learning should rely more on the centralized data” is on the assumption  that “ synthetic and the extended KD datasets are similar to the global ones” Thus the LHS of Eq(9) will be dominated by $ \\epsilon_{{P}^T}(f^{Syn})$ and $ \\epsilon_{{P}^T}(f_{\\text{KD}}^{Syn})$. It is intuitive that if synthetic data is ideal aligned with global distribution and we fully rely on synthetic data in local training, the performance gain between training on synthetic data vs on local data is expected to be larger when local data is more heterogeneous. Also, as noted in Theorem 1, the performance is not solely related to hyperparameters $\\lambda_{REG}$ and $\\lambda_{KD}$. It will be also affected (or dominated) by the fourth term of Eq 8, when we rely on sufficient number of ideal synthetic data (aligned with global distribution)."},"title":{"value":"Clarification of Proposition 2"}},"tmdate":1700471266166,"tcdate":1700471266166,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission545/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission545/Authors"],"forum":"PcBJ4pA6bF","number":27,"license":"CC BY 4.0","cdate":1700471266166,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission545/-/Official_Comment"],"mdate":1700471266166,"domain":"ICLR.cc/2024/Conference","replyto":"kHapeueync","id":"MGFH3n97lw","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Federated Learning","Data Heterogeneity","Model Heterogeneity"]},"supplementary_material":{"value":"/attachment/9f59eb9a0983d84551769b886e45eac851ffafb7.zip"},"primary_area":{"value":"general machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Conventional Federated Learning (FL) involves collaborative training of a global model by multiple client local models. In this emerging paradigm, the central server assumes a critical role in aggregating local models and maintaining the global model. However, it encounters various challenges, including scalability, management, and inefficiencies arising from idle client devices. \nRecently, studies on serverless decentralized FL have shown advantages in overcoming these challenges, enabling clients to own different local models and separately optimize local data. Despite the promising advancements in decentralized FL, it is crucial to thoroughly investigate the implications of data and model heterogeneity, which pose unique challenges that must be overcome. Therefore, the research question to be answered in this study is: How can every client's local model learn generalizable representation?\nTo address this question, we propose a novel Decentralized FL technique by introducing Synthetic Anchors, dubbed as DeSA. Inspired by the theory of domain adaptation and Knowledge distillation (KD), we leverage the synthetic anchors to design two effective regularization terms for local training: 1) anchor loss that matches the distribution of the client's latent embedding with an anchor and 2) KD loss that enables clients learning from others. \nIn contrast to previous KD-based heterogeneous FL methods, we don’t presume access to real public or a global data generator. \nDeSA enables each client's model to become robust to distribution shift across different client-domains. Through extensive experiments on diverse client data distributions, we showcase the effectiveness of \\ours{} in enhancing both inter and intra-domain accuracy of each client."},"_bibtex":{"value":"@misc{\nhuang2024overcoming,\ntitle={Overcoming Data and Model heterogeneities in Decentralized Federated Learning via Synthetic Anchors},\nauthor={Chun-Yin Huang and Kartik Srinivas and Xin Zhang and Xiaoxiao Li},\nyear={2024},\nurl={https://openreview.net/forum?id=PcBJ4pA6bF}\n}"},"title":{"value":"Overcoming Data and Model heterogeneities in Decentralized Federated Learning via Synthetic Anchors"},"pdf":{"value":"/pdf/aa9370a559289b8e6c23498cbecfb3d977d6d6e3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"huang|overcoming_data_and_model_heterogeneities_in_decentralized_federated_learning_via_synthetic_anchors"},"authorids":{"value":["~Chun-Yin_Huang1","~Kartik_Srinivas1","~Xin_Zhang16","~Xiaoxiao_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Chun-Yin Huang","Kartik Srinivas","Xin Zhang","Xiaoxiao Li"]}},"version":2},{"content":{"comment":{"value":"**W2. \"The method is applied to very simple cases. As discussed by the authors themselves, applications in more complex scenarios would be great to have in the future.\"**\n\nGood comment! To address it, we now consider an experiment with a significantly more complex permeability. In this experiment, we generate permeability based on a log-uniform distribution while excluding permeability values that are greater than 0.01 and less than 100.\n\n[Complex initial data, exp.1](https://drive.google.com/file/d/17-c_VcNSLmM_m1v6iNwXZmfm8T_xtCCM/view?usp=sharing)\n\nIn these experiments, we will take a grid consisting of 900 points and demonstrate coarsening to 90 points. First, let's demonstrate the results of our algorithm in the zero sink. We use this hyperparameters: \n\n- Time Step: 0.000005 \n- Number of Epochs: 900 \n- Learning Rate: 0.01\n- Sigmoid Weight: 15\n- Physics Loss Weight: 25\n\n[Results of applying our method to the complex scenario, exp.1](https://drive.google.com/file/d/14wB4rV6N3lyAvMu2AiIBZiAi85JvttSg/view?usp=sharing)\n\nIt can be seen that in this example of the algorithm, more steps are required to converge to a solution, and it is also seen that our loss function has many perturbations. Now let's look at the algorithm from the  [Shumilin et.al.]. In this case, at one of the steps of the algorithm, the scheme of the explicit solver diverges, which leads to the algorithm stopping. The algorithm we have demonstrated provides stability loss, which prevents the discrepancy of the explicit solution scheme:\n\n[Results of applying the method from Shumilin et al., 2024, exp.1](https://drive.google.com/file/d/1tVkxoFl_wTIw4KB52Ui5AG6uWaaBQ48k/view?usp=sharing)\n\nNext, we will demonstrate a slightly different experiment, where we reduce the range of permeability values.  This is our permeability field:\n\n[Complex initial data, exp. 2](https://drive.google.com/file/d/1_Two_9boHfBYfEDE4lat2IYC2dsjwq5f/view?usp=sharing)\n\nNow we will leave only permeabilities with values in the range from 0.1 to 10. Here are the results of our algorithm in the zero sink. We use this hyperparameters:\n\n- Time Step: 0.000005 \n- Number of Epochs: 900 \n- Learning Rate: 0.012\n- Sigmoid Weight: 15\n- Physics Loss Weight: 25\n\n[Results of applying our method to the complex scenario, exp. 2](https://drive.google.com/file/d/10lGMnoOVBpoD3wSuM9uQqa_NsDjGxsSi/view?usp=sharing)\n\nIt can be seen that in this experiment with a simpler permeability distribution, our algorithm converges faster. \n\n[Results of applying the method from Shumilin et al., 2024, exp. 2](https://drive.google.com/file/d/1MtrUZ2Eo8jCOdKHVdnYuNOQAQHBZ90t1/view?usp=sharing)\n\nIn this case, the algorithm from the  [Shumilin et.al.] got a more accurate result but also broke down at one of the steps. These experiments demonstrate the importance of stability loss and the fact that the algorithm we have demonstrated is capable of solving more complex problems."},"title":{"value":"Answer to Reviewer wVwb Part 2"}},"tmdate":1732308076131,"tcdate":1732308076131,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11649/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission11649/Authors"],"forum":"TSTgP4W3ga","number":5,"license":"CC BY 4.0","cdate":1732308076131,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11649/-/Official_Comment"],"mdate":1732308076131,"domain":"ICLR.cc/2025/Conference","replyto":"TE8iwYK3DV","id":"kiHjYeSGXR","forumContent":{"TLDR":{"value":"We developed a fully differentiable framework for unstructured grid coarsening, driven by the underlying physics simulation and its stability requirements.."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Graph Neural Networks","differentiable solvers","numerical modelling","grid coarsening","upscaling"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Efficient simulations of complex physical systems described by partial differential equations (PDE) require computational methods that can reduce the resource demands without sacrificing the accuracy. Traditionally, this is achieved by ``upscaling'' the simulation grids or by aggregating cells based on a priori information. Here, we introduce a novel framework based on graph neural networks (GNN) for learnable self-supervised differentiable coarsening of unstructured computational grids. We leverage graph-based representation of the physical system and offer a graph coarsening method which preserves the underlying physical properties together with the stability of the chosen numerical scheme. This is achieved by minimizing the error between the output of the simulations using coarsened and original graph. We demonstrate the approach on several example differential equations, modeling sub-surface flow and wave propagation. We demonstrate that the model exhibits ability to maintain high fidelity in simulation outputs even after 95\\% reduction on the nodes, significantly reducing computational overhead. We also show that the model exhibits generalizability to unseen scenarios, thereby outperforming the baselines. Thus, the developed approach demonstrates the ability to accelerate simulation without comprising accuracy and hence has potential for accelerating physical simulations in various domains."},"_bibtex":{"value":"@misc{\nryabov2025learnable,\ntitle={Learnable Stability-Aware Unstructured Grid Coarsening Using Graph Neural Networks for Accelerated Physics Simulations},\nauthor={Alexander Ryabov and Sergei Shumilin and Viacheslav Naumov and Nikolay Yavich and Sayan Ranu and N M Anoop Krishnan and Evgeny Burnaev and Vladimir Vanovskiy},\nyear={2025},\nurl={https://openreview.net/forum?id=TSTgP4W3ga}\n}"},"title":{"value":"Learnable Stability-Aware Unstructured Grid Coarsening Using Graph Neural Networks for Accelerated Physics Simulations"},"pdf":{"value":"/pdf/7337abbebaa58ef6b472912fb5859788a19f1931.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ryabov|learnable_stabilityaware_unstructured_grid_coarsening_using_graph_neural_networks_for_accelerated_physics_simulations"},"authorids":{"value":["~Alexander_Ryabov1","~Sergei_Shumilin1","~Viacheslav_Naumov1","~Nikolay_Yavich1","~Sayan_Ranu2","~N_M_Anoop_Krishnan1","~Evgeny_Burnaev1","~Vladimir_Vanovskiy1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Alexander Ryabov","Sergei Shumilin","Viacheslav Naumov","Nikolay Yavich","Sayan Ranu","N M Anoop Krishnan","Evgeny Burnaev","Vladimir Vanovskiy"]}},"version":2},{"content":{"summary":{"value":"The paper revisits the scaling laws of large language model (LLM) pretraining by introducing an explicit, dimensionless data quality parameter Q, extending the traditional analysis of model size and data volume to a joint framework that incorporates data quality. The authors propose a quality-aware scaling law. They systematically inject synthetic noise and vary data coverage in neural machine translation and causal language modeling tasks. Experimental results show that high-quality data can significantly reduce loss for a given model size, and under high-quality data conditions, smaller models with less computational resources can achieve strong performance."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please refer to the relevant points in the Weaknesses section. If the authors can provide clarification and improvements, I would be very happy to raise my score."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The work innovatively incorporates data quality as a single parameter into the scaling law, theoretically demonstrating a strong correlation between model performance and data quality.\n2. Controlled experiments are conducted across multiple experimental settings, validating the practical applicability of the quality-aware scaling law."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although synthetic noise offers strong controllability, it may not capture the highly complex and heavy-tailed distribution of low-quality data in real training datasets. How can the current single-parameter quantification method be extended to more realistic data scenarios?\n\n2. The authors use a fixed learning rate to investigate the scaling law. During training, could different learning rates have additional effects on the conclusions regarding the data-quality scaling law?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933969737,"tcdate":1761881275828,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20553/Reviewer_BKxM"],"signatures":["ICLR.cc/2026/Conference/Submission20553/Reviewer_BKxM"],"forum":"x54wwB6QvL","number":2,"license":"CC BY 4.0","cdate":1761881275828,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20553/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933969737,"domain":"ICLR.cc/2026/Conference","replyto":"x54wwB6QvL","id":"P17n6r6jOK","forumContent":{"TLDR":{"value":"We extend the Chinchilla scaling law by introducing a data‐quality measure that predicts how noise and coverage affect loss, and validate it on machine translation and autoregressive modeling tasks."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["quality ware scaling laws","scaling laws","data quality","LLM pretraining"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Scaling laws for language model training traditionally characterize how performance scales with model size and dataset volume. Prior work has explored architecture variants and data treatments such as dataset filtering and noise injection in language model pretraining; however, these studies have not formalized data quality within a principled scaling law. We introduce a dimensionless data-quality parameter Q, and propose a quality-aware scaling law extending the Chinchilla framework to predict loss as a joint function of model size, data volume, and data quality. The law is motivated by an effective-sample-size and information-theoretic view of noisy or redundant corpora, and it admits two practical estimators for Q: (i) a corruption rate proxy and (ii) a deficiency measure. Through synthetic experiments in neural machine translation and autoregressive modeling--where we systematically control data quality via multiple levels of noise injection variation--we show that loss scales predictably with data quality and that higher-quality data can substantially reduce model size and hence compute requirements. Our results demonstrate a sublinear decay of effective data with quality and robustness to moderate data corruption; out-of-sample evaluations further validate the predictive form of the law. Unlike prior empirical analyses, our work establishes an explicit, generalizable law for data quality, offering concrete guidance for balancing data curation effort and model scale in large-scale pretraining."},"_bibtex":{"value":"@inproceedings{\nsubramanyam2026scaling,\ntitle={Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining},\nauthor={Anirudh Subramanyam and Yuxin Chen and Robert L. Grossman},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=x54wwB6QvL}\n}"},"title":{"value":"Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining"},"pdf":{"value":"/pdf/84d2b007d5109b66762818426b58c936984d2acc.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"subramanyam|scaling_laws_revisited_modeling_the_role_of_data_quality_in_language_model_pretraining"},"authorids":{"value":["~Anirudh_Subramanyam1","~Yuxin_Chen1","~Robert_L._Grossman2"]},"authors":{"value":["Anirudh Subramanyam","Yuxin Chen","Robert L. Grossman"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2509.11201v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"she|scaling_up_forest_vision_with_synthetic_data"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yihang_She:","https://dblp.org/search/pid/api?q=author:Andrew_Blake_0004:","https://dblp.org/search/pid/api?q=author:David_Coomes:","~Srinivasan_Keshav1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2509.11201"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2509-11201,\n  publtype={informal},\n  author={Yihang She and Andrew Blake and David Coomes and Srinivasan Keshav},\n  title={Scaling Up Forest Vision with Synthetic Data},\n  year={2025},\n  month={September},\n  cdate={1756684800000},\n  journal={CoRR},\n  volume={abs/2509.11201},\n  url={https://doi.org/10.48550/arXiv.2509.11201}\n}\n"},"abstract":{"value":"Accurate tree segmentation is a key step in extracting individual tree metrics from forest laser scans, and is essential to understanding ecosystem functions in carbon cycling and beyond. Over the past decade, tree segmentation algorithms have advanced rapidly due to developments in AI. However existing, public, 3D forest datasets are not large enough to build robust tree segmentation systems. Motivated by the success of synthetic data in other domains such as self-driving, we investigate whether similar approaches can help with tree segmentation. In place of expensive field data collection and annotation, we use synthetic data during pretraining, and then require only minimal, real forest plot annotation for fine-tuning. We have developed a new synthetic data generation pipeline to do this for forest vision tasks, integrating advances in game-engines with physics-based LiDAR simulation. As a result, we have produced a comprehensive, diverse, annotated 3D forest dataset on an unprecedented scale. Extensive experiments with a state-of-the-art tree segmentation algorithm and a popular real dataset show that our synthetic data can substantially reduce the need for labelled real data. After fine-tuning on just a single, real, forest plot of less than 0.1 hectare, the pretrained model achieves segmentations that are competitive with a model trained on the full scale real data. We have also identified critical factors for successful use of synthetic data: physics, diversity, and scale, paving the way for more robust 3D forest vision systems in the future. Our data generation pipeline and the resulting dataset are available at https://github.com/yihshe/CAMP3D.git."},"title":{"value":"Scaling Up Forest Vision with Synthetic Data"},"authors":{"value":["Yihang She","Andrew Blake","David Coomes","Srinivasan Keshav"]}},"tmdate":1761295958703,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2509-11201"],"tcdate":1761295939818,"writers":["~"],"signatures":["~Srinivasan_Keshav1"],"forum":"f0sjCrna9q","license":"CC BY-SA 4.0","number":648227,"cdate":1756684800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1761295958703,"domain":"DBLP.org","id":"f0sjCrna9q","version":2},{"content":{"comment":{"value":"## (Part 2) Advantages of iRGB over others\nAs Reviewer QqAh also raised the same point regarding iRGB's advantages over other representations, we have now listed them below:   \n\n**Intuitive Mapping of Color to Complex Plane:** Our R2C method provides a more intuitive and geometrically grounded mapping of colors to the complex domain. By identifying two distinct Argand planes in the RGB color space itself, our method captures both the color relationship with respect to the grayscale (grayline) and the spatial positioning of the color within the RGB cube, resulting in a complex-valued color model called iRGB. We hardly have any complex-valued color models except iHSV (proposed in FCCNs paper [1]), and our experiments have shown better performance of our iRGB over iHSV.     \n\n**Richer & Dual Complex Representations:** This approach enables us to represent each color as two complex numbers, which provides a richer representation than typical methods which only focus on one color channel (e.g., Hilbert transform) at a time or a single transformation (e.g., using quaternions, where a 4D complex number [$0 + iR + jG + kB$] is employed without considering the inherent color space geometry). The two complex numbers derived from the Argand planes allow our color model (iRGB) to preserve more nuanced information about color differences and the relationship between color and luminance. Through this dual complex representation, our method captures the perceptual relationship between luminance and chrominance in a way that is not explicitly addressed by the quaternion or Hilbert transform approaches. This makes our R2C method applicable to a broader range of image processing algorithms.\n\n**Computational Efficiency:** The R2C method transforms real-valued images to complex-valued images by processing each pixel independently, resulting in a computational complexity of $O(n)$ and making it highly parallelizable. In contrast, methods like the Hilbert transform have higher computational costs, at least $O(n\\log n)$. This efficiency makes R2C well-suited for large-scale image processing tasks.\n\n**Flexibility in Applications:** While methods like quaternion representation or complex logarithmic transformations typically focus on specific kinds of data (e.g., 3D rotations or magnitude-phase representations), our method is more flexible in mapping any color into a pair of complex numbers, making it a general framework for complex-valued image generation.\n\nIn summary, while quaternion representations and other well-established transformations have their place in the broader context of complex-valued image processing, the R2C method offers a unique advantage by providing an intuitive, dual-complex, efficient representation of color that better captures the geometry of the RGB color space, while leading to wider application domains and applications in various image processing algorithms. We believe this contribution significantly advances the ability to generate complex-valued images from real-valued ones.\n\n\n**Advantages of iRGB over simple [R+iG,G+iB] representation**\n\nThe representation $[R+iG,G+iB]$ is a relatively naive approach and disregards key principles of color model design. Below, we highlight three specific advantages of our iRGB model:\n\n*Independence of Color Components*: In $[R+iG,G+iB]$, the imaginary part of the first component ($G$) is tied to the real part of the second component ($G$). This dependency inhibits the independent processing of color components. In contrast, our iRGB representation utilizes two fully independent components, $||v||e^{i\\theta}$ and $||u||e^{i\\phi}$, allowing greater flexibility and more effective color representation.\n\n*Unbiased Representation*: The $[R+iG,G+iB]$ approach inherently biases the green ($G$) color channel due to its repetition, potentially leading to suboptimal results. In comparison, iRGB avoids such biases, providing a balanced and unbiased representation across color channels.\n\n*Intuitive Meaning*: The $[R+iG,G+iB]$ representation lacks a clear and intuitive physical or conceptual interpretation. On the other hand, iRGB directly encodes both intensity and color information, offering a more meaningful and interpretable representation as detailed earlier.\n\nWe hope this clarifies various advantages our iRGB offers over others.  \n\n[1] Saurabh Yadav; Koteswar Rao Jerripothula, FCCNs: Fully Complex-valued Convolutional Networks using Complex-valued Color Model and Loss Function. ICCV 2023."},"title":{"value":"Author Response to additional comments of Reviewer F89z (part 2)"}},"tmdate":1733201474880,"tcdate":1733170555224,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1276/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission1276/Authors"],"forum":"9hmDl8fFDs","number":28,"license":"CC BY 4.0","cdate":1733170555224,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1276/-/Official_Comment"],"mdate":1733201474880,"domain":"ICLR.cc/2025/Conference","replyto":"c03jsl6RYW","id":"P6WRmX51tR","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"A robust complex-valued approach in Spatio-spectral domain for multiple tasks on both real and complex data."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Deep Complex Newtworks","Complex-valued color transformation"]},"supplementary_material":{"value":"/attachment/d3c1bebc77a7f4edd23a402e6580595b34f472e4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex-valued neural networks have attracted growing attention for their ability to handle complex-valued data with enhanced representational capacity. However, their potential in computer vision remains relatively untapped. \nIn this paper, we introduce Deep Complex Spatio-Spectral Network (DCSNet), a fully complex-valued token-based, end-to-end neural network designed for binary segmentation tasks. Additionally, our DCSNet encoder can be used for image classification in the complex domain. We also propose an invertible real-to-complex (R2C) transform, which generates two complex-valued input channels, complex intensity and complex hue, while producing complex-valued images with distinct real and imaginary components.\nDCSNet operates in both spatial and spectral domains by leveraging complex-valued inputs and complex Fourier transform.\nAs a result, the complex-valued representation is maintained throughout DCSNet, and we avoid the information loss typically associated with Real$\\leftrightarrow$Complex transformations. Extensive experiments show that DCSNet surpasses existing complex-valued methods across various tasks on both real and complex-valued data and achieves competitive performance compared to existing real-valued methods, establishing a robust framework for handling both data types effectively."},"_bibtex":{"value":"@misc{\nyadav2025deep,\ntitle={Deep Complex Spatio-Spectral Networks with Complex Visual Inputs},\nauthor={Saurabh Yadav and Koteswar Rao Jerripothula},\nyear={2025},\nurl={https://openreview.net/forum?id=9hmDl8fFDs}\n}"},"title":{"value":"Deep Complex Spatio-Spectral Networks with Complex Visual Inputs"},"pdf":{"value":"/pdf/78685d9476ab42033341d680f3e952693b0cbc64.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yadav|deep_complex_spatiospectral_networks_with_complex_visual_inputs"},"authorids":{"value":["~Saurabh_Yadav2","~Koteswar_Rao_Jerripothula3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Saurabh Yadav","Koteswar Rao Jerripothula"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Noise-Conditioned Energy-based Annealed Rewards (NEAR), a novel framework for imitation learning from observation using energy-based generative models. NEAR leverages denoising score matching to learn smooth representations of the expert's motion distribution and uses these energy functions as rewards. Unlike adversarial imitation learning approaches, NEAR avoids unstable min-max optimization, achieving smoother and more stable reward signals. Additionally, an annealing strategy progressively transitions between energy functions to provide more refined guidance for the agent’s policy. NEAR is evaluated on complex humanoid tasks, showing promising results when compared to state-only adversarial imitation learning baselines like Adversarial Motion Priors (AMP) in terms of motion quality, stability, and imitation accuracy."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Could you provide a comparison of NEAR’s computational requirements (e.g., training time, GPU hours) relative to other baselines like AMP?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. NEAR performs well on a range of complex motion tasks, including stylized walking, running, and martial arts. The results demonstrate competitive imitation accuracy and smoothness compared to AMP, particularly in complex tasks where AMP struggles with stability.\n2. The paper includes ablation studies to explore the impact of key components."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. NEAR’s effectiveness is primarily evaluated in humanoid tasks, which are continuous and physics-driven. The framework’s applicability in other types of imitation learning tasks, especially those with discrete actions or diverse goal-oriented, is not fully explored.\n2. NEAR requires training a noise-conditioned energy model, which can be computationally intensive. A detailed comparison of training costs relative to other methods, particularly in terms of time and resources, would be beneficial.\n3. NEAR is only compared against one baseline - AMP and it doesn't seem to always be the winner (but has higher variance in most cases) despite the additional complexity of learning an energy network."}},"nonreaders":[],"tmdate":1732552888890,"tcdate":1730648010853,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10326/Reviewer_PK39"],"signatures":["ICLR.cc/2025/Conference/Submission10326/Reviewer_PK39"],"forum":"DL9txImSzm","number":3,"license":"CC BY 4.0","cdate":1730648010853,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10326/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732552888890,"domain":"ICLR.cc/2025/Conference","replyto":"DL9txImSzm","id":"sLwh2RqDwK","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Proposes a new algorithm for imitation learning for observation that uses score-based models to learn the expert distribution's energy function and uses these learnt energy functions as reward functions."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["imitation learning","energy based generative models","reinforcement learning","imitation from observation"]},"supplementary_material":{"value":"/attachment/58a4a6887ac00b91c040a1985a542a0577d92d37.zip"},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper introduces a new imitation learning framework based on energy-based generative models capable of learning complex, physics-dependent, robot motion policies through state-only expert motion trajectories. Our algorithm, called Noise-conditioned Energy-based Annealed Rewards (NEAR), constructs several perturbed versions of the expert's motion data distribution and learns smooth, and well-defined representations of the data distribution's energy function using denoising score matching. We propose to use these learnt energy functions as reward functions to learn imitation policies via reinforcement learning. We also present a strategy to gradually switch between the learnt energy functions, ensuring that the learnt rewards are always well-defined in the manifold of policy-generated samples. We evaluate our algorithm on complex humanoid tasks such as locomotion and martial arts and compare it with state-only adversarial imitation learning algorithms like Adversarial Motion Priors (AMP). Our framework sidesteps the optimisation challenges of adversarial imitation learning techniques and produces results comparable to AMP in several quantitative metrics across multiple imitation settings."},"_bibtex":{"value":"@inproceedings{\ndiwan2025noiseconditioned,\ntitle={Noise-conditioned Energy-based Annealed Rewards ({NEAR}): A Generative Framework for Imitation Learning from Observation},\nauthor={Anish Abhijit Diwan and Julen Urain and Jens Kober and Jan Peters},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=DL9txImSzm}\n}"},"title":{"value":"Noise-conditioned Energy-based Annealed Rewards (NEAR): A Generative Framework for Imitation Learning from Observation"},"pdf":{"value":"/pdf/13098c5b4053f6e6b5487444ed2a1aa1501ca321.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"diwan|noiseconditioned_energybased_annealed_rewards_near_a_generative_framework_for_imitation_learning_from_observation"},"authorids":{"value":["~Anish_Abhijit_Diwan1","~Julen_Urain2","~Jens_Kober1","~Jan_Peters3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Anish Abhijit Diwan","Julen Urain","Jens Kober","Jan Peters"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a method called Test-Time Self-Improvement (TT-SI), enabling an LLM agent to refine itself during inference. The core pipeline is:\n\n1. Uncertainty estimation : flag test samples on which the model is “uncertain”;\n2. Data synthesis : automatically generate synthetic training data from these samples;\n3. Test-time tuning : perform a lightweight LoRA fine-tune on the model with the synthetic data, then infer on the same sample.\n\nExperiments show that the proposed method improves the model’s performance on the test set."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. The core of TT-SI is to generate *k* synthetic training examples that resemble the model-uncertain test point, fine-tune on them, and then re-predict the very same input. This means every uncertain test sample is—directly or indirectly—present in the training signal. Could the authors clarify how TT-SI is fundamentally different from ordinary “test-set leakage”?\n2. Even if TT-SI is not identical to classic leakage, performing one parameter update per uncertain question is impractical.\n3. Can the authors provide evidence that TT-SI yields better generalization than fine-tuning on the *full* training set, rather than merely over-fitting to the single “uncertain” example?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper validates the method’s effectiveness on three test sets.\n2. The proposed TT-D variant shows how synthetic-data quality influences performance.\n3. The pipeline cleanly modularizes into uncertainty estimation, data synthesis, and test-time fine-tuning, so it is readily extensible."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The core of TT-SI is to generate *k* synthetic training examples that resemble the model-uncertain test point, fine-tune on them, and then re-predict the very same input. This means every uncertain test sample is—directly or indirectly—present in the training signal. \n2. Despite the authors' characterization of the method as \"test-time adaptation,\" its central conclusion—that generating training data resembling the test sample and fine-tuning on it boosts performance—is mere common sense and offers no novelty."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941968824,"tcdate":1761389973229,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21886/Reviewer_R7kG"],"signatures":["ICLR.cc/2026/Conference/Submission21886/Reviewer_R7kG"],"forum":"M1zSTXY1xr","number":2,"license":"CC BY 4.0","cdate":1761389973229,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21886/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941968824,"domain":"ICLR.cc/2026/Conference","replyto":"M1zSTXY1xr","id":"Q8TGLb6Djb","forumContent":{"TLDR":{"value":"We introduce a test-time self-improvement algorithm where agents detect uncertain test samples they struggle with, generate new examples from them, and use these at test-time fine-tuning, achieving higher accuracy with far fewer samples."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Test-Time Training","Self-Improvement","Language Agents"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"One paradigm of language model (LM) fine-tuning relies on creating large training datasets, under the assumption that high quantity and diversity will enable models to generalize to novel tasks after post‑training. In practice, gathering large sets of data is inefficient, and training on them is prohibitively expensive; worse, there is no guarantee that the resulting model will handle complex scenarios or generalize better. Moreover, existing techniques rarely assess whether a training sample provides novel information or is redundant with the knowledge already acquired by the model, resulting in unnecessary costs. In this paper, we explore a new test-time self-improvement method to create more effective and generalizable agentic LMs *on-the-fly*. The proposed algorithm can be summarized in three steps: (i) first it identifies the samples that the model struggles with by using an uncertainty function (self-awareness), (ii) then generates similar examples from the detected uncertain samples (self-data augmentation), and (iii) uses these newly generated samples at test-time fine-tuning (self-learning). We study two variants of this approach: *Test-Time Self-Improvement* (TT-SI), where the same model generates additional training examples from its own uncertain cases and then learns from them, and contrast this approach with *Test-Time Distillation* (TT-D), where a stronger model generates similar examples for those same uncertain cases, enabling the student to adapt using distilled supervision. Empirical evaluations across different agent benchmarks demonstrate that TT-SI surpasses other standard learning methods with +5.36% absolute gain in average accuracy, yet trains using 68x less training samples and TT-D further improves performance in harder scenarios that require diverse training signals. Our findings highlight the promise of TT-SI and limitations in current learning frameworks regarding cost and generalizability, demonstrating the potential of self-evolving LMs at test-time as a new paradigm for building more capable agents on complex scenarios."},"_bibtex":{"value":"@misc{\nacikgoz2025selfimproving,\ntitle={Self-Improving {LLM} Agents at Test-Time},\nauthor={Emre Can Acikgoz and Cheng Qian and Heng Ji and Dilek Hakkani-T{\\\"u}r and Gokhan Tur},\nyear={2025},\nurl={https://openreview.net/forum?id=M1zSTXY1xr}\n}"},"title":{"value":"Self-Improving LLM Agents at Test-Time"},"pdf":{"value":"/pdf/eba64dd8508e0180488e29df6065733c65524622.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"acikgoz|selfimproving_llm_agents_at_testtime"},"authorids":{"value":["~Emre_Can_Acikgoz1","~Cheng_Qian4","~Heng_Ji3","~Dilek_Hakkani-Tür1","~Gokhan_Tur2"]},"authors":{"value":["Emre Can Acikgoz","Cheng Qian","Heng Ji","Dilek Hakkani-Tür","Gokhan Tur"]}},"version":2},{"content":{"comment":{"value":"**Re: (4) The inclusion of an “anti-physics” category is similar to “counterfactual prompts.”**\n\nThank you for raising this point. While both our anti-physics category and the counterfactual prompts in T2VPhysBench involve physics-violating scenarios, they differ significantly in motivation, design, and evaluation goals.\n\n**1. Different purpose**  \n- *Counterfactual prompts* (T2VPhysBench) test whether models will **follow an impossible instruction**.  \n- Our *anti-physics category* evaluates whether a model can **maintain internal logical consistency**—motion continuity, object interactions, visual stability—*after* an impossible event occurs.  \n  - Our goal is **robustness under unphysical conditions**, not instruction obedience.\n\n**2. Different design**  \n- T2VPhysBench includes isolated counterfactual prompts targeting individual laws.  \n- PhyWorldBench defines a structured and systematic taxonomy:  \n  **5 anti-physics subcategories**  \n  (defying gravity, perpetual motion, object phasing, time reversal, infinite duplication)  \n  × **7 scenarios**  \n  × **3 prompt types**,  \n  providing substantially broader and more diverse coverage.\n\n**3. Different evaluation criteria**  \n- T2VPhysBench evaluates simply whether the model complies with or violates the counterfactual instruction.  \n- Our benchmark uses **Basic Standards + Key Standards** to judge whether the video remains coherent and physically stable *within* an intentionally impossible scenario—an aspect not captured by counterfactual prompts.\n\nIn summary, while both involve physics-violating setups, our anti-physics category is **not equivalent** to counterfactual prompts. It provides a **richer, more systematic framework** for analyzing how generation models behave under intentionally unphysical conditions.\n\n---\n\n**Re: (5) Quantifying the “rationalization” failure mode**\n\nThank you for the suggestion. We analyzed **600 failure videos** (50 per model) across all 12 video generation systems to measure how often models “rationalize” physics violations. The results are summarized below in a two-row table:\n\n| Pika | Sora | Kling | Luma | Gen-3 | Wanx | Hunyuan | Step-Video-T2V | Open-Sora | CogVideo | Open-Sora-Plan | LTX-Video |\n|------|------|--------|-------|--------|--------|-----------|------------------|------------|------------|------------------|-------------|\n| 34% | 24% | 24% | 22% | 14% | 28% | 20% | 18% | 4% | 16% | 6% | 4% |\n\nWe found that **higher-performing models tend to have a higher proportion of rationalization errors**. Lower-performing models typically fail earlier on semantic or visual fidelity issues (e.g., object distortion), whereas stronger models increasingly “solve” violations by generating visually polished but physically incorrect static scenes. This confirms rationalization as an important emergent failure mode.\n\n---\n\n**Re: (6) Difficulty ratings**\n\nThis is an excellent idea. We will include difficulty analysis in the next revision by:\n\n1. Measuring relative model performance per prompt, and  \n2. Having human annotators label difficulty to validate alignment.\n\nThis will help researchers focus on the most challenging scenarios."},"title":{"value":"Rebuttal 2/2"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764146110901,"tcdate":1764146110901,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8464/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission8464/Authors"],"forum":"rlZeILv3fm","number":6,"license":"CC BY 4.0","cdate":1764146110901,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8464/-/Official_Comment"],"mdate":1764146110901,"domain":"ICLR.cc/2026/Conference","replyto":"nCKBEAQoa2","id":"l699dOh5dF","forumContent":{"venue":{"value":"ICLR 2026 Oral"},"TLDR":{"value":"Large-scale, multidimensional video generation for physics"},"keywords":{"value":["Video Generation","Video Evaluation"]},"supplementary_material":{"value":"/attachment/bd8cdaa60661c40d3e703b3dd09ff6689653ba72.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. This paper presents $PhyWorldBench$\n, a comprehensive benchmark designed to evaluate video generation models based on their adherence to the laws of physics. The benchmark covers multiple levels of physical phenomena, ranging from fundamental principles like object motion and energy conservation to more complex scenarios involving rigid body interactions and human or animal motion. Additionally, we introduce a novel \"Anti-Physics\" category, where prompts intentionally violate real-world physics, enabling the assessment of whether models can follow such instructions while maintaining logical consistency. Besides large-scale human evaluation, we also design a simple yet effective method that could utilize current MLLM to evaluate the physics realism in a zero-shot fashion. We evaluate 10 state-of-the-art text-to-video generation models, including five open-source and five proprietary models, with a detailed comparison and analysis. we identify pivotal challenges models face in adhering to real-world physics. Through systematic testing of their outputs across 1,050 curated prompts—spanning fundamental, composite, and anti-physics scenarios—we identify pivotal challenges these models face in adhering to real-world physics. We then rigorously examine their performance on diverse physical phenomena with varying prompt types, deriving targeted recommendations for crafting prompts that enhance fidelity to physical principles."},"_bibtex":{"value":"@inproceedings{\ngu2026phyworldbench,\ntitle={\\$PhyWorldBench\\$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models},\nauthor={Jing Gu and Xian Liu and Yu Zeng and Ashwin Nagarajan and Fangrui Zhu and Daniel Hong and Yue Fan and Qianqi Yan and Kaiwen Zhou and Ming-Yu Liu and Xin Eric Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=rlZeILv3fm}\n}"},"title":{"value":"$PhyWorldBench$: A Comprehensive Evaluation of Physical Realism in Text-to-Video Models"},"pdf":{"value":"/pdf/6138b5d05836a6d9ee27ae5fc6f3bbbe3667ff02.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gu|phyworldbench_a_comprehensive_evaluation_of_physical_realism_in_texttovideo_models"},"authorids":{"value":["~Jing_Gu2","~Xian_Liu1","~Yu_Zeng1","~Ashwin_Nagarajan1","~Fangrui_Zhu1","~Daniel_Hong1","~Yue_Fan3","~Qianqi_Yan1","~Kaiwen_Zhou3","~Ming-Yu_Liu1","~Xin_Eric_Wang2"]},"authors":{"value":["Jing Gu","Xian Liu","Yu Zeng","Ashwin Nagarajan","Fangrui Zhu","Daniel Hong","Yue Fan","Qianqi Yan","Kaiwen Zhou","Ming-Yu Liu","Xin Eric Wang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2506.23135v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"shang|roboscape_physicsinformed_embodied_world_model"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yu_Shang:","https://dblp.org/search/pid/api?q=author:Xin_Zhang_0123:","https://dblp.org/search/pid/api?q=author:Yinzhou_Tang:","https://dblp.org/search/pid/api?q=author:Lei_Jin:","~Chen_Gao3","https://dblp.org/search/pid/api?q=author:Wei_Wu_0021:","~Yong_Li7"]},"html":{"value":"https://doi.org/10.48550/arXiv.2506.23135"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2506-23135,\n  publtype={informal},\n  author={Yu Shang and Xin Zhang and Yinzhou Tang and Lei Jin and Chen Gao and Wei Wu and Yong Li},\n  title={RoboScape: Physics-informed Embodied World Model},\n  year={2025},\n  month={June},\n  cdate={1748736000000},\n  journal={CoRR},\n  volume={abs/2506.23135},\n  url={https://doi.org/10.48550/arXiv.2506.23135}\n}\n"},"abstract":{"value":"World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models exhibit limited physical awareness, particularly in modeling 3D geometry and motion dynamics, resulting in unrealistic video generation for contact-rich robotic scenarios. In this paper, we present RoboScape, a unified physics-informed world model that jointly learns RGB video generation and physics knowledge within an integrated framework. We introduce two key physics-informed joint training tasks: temporal depth prediction that enhances 3D geometric consistency in video rendering, and keypoint dynamics learning that implicitly encodes physical properties (e.g., object shape and material characteristics) while improving complex motion modeling. Extensive experiments demonstrate that RoboScape generates videos with superior visual fidelity and physical plausibility across diverse robotic scenarios. We further validate its practical utility through downstream applications including robotic policy training with generated data and policy evaluation. Our work provides new insights for building efficient physics-informed world models to advance embodied intelligence research. The code is available at: https://github.com/tsinghua-fib-lab/RoboScape."},"title":{"value":"RoboScape: Physics-informed Embodied World Model"},"authors":{"value":["Yu Shang","Xin Zhang","Yinzhou Tang","Lei Jin","Chen Gao","Wei Wu","Yong Li"]}},"tmdate":1768965172322,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2506-23135"],"tcdate":1768608292856,"writers":["~"],"signatures":["~Yong_Li7"],"forum":"Jo8PhYaCFt","license":"CC BY-SA 4.0","number":769218,"cdate":1748736000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768965172322,"domain":"DBLP.org","id":"Jo8PhYaCFt","version":2},{"content":{"summary":{"value":"This paper presents METALINT, a framework for teaching LLMs to perform code quality analysis by following natural language instructions. Its core \"easy-to-hard\" generalization idea involves training (via SFT and DPO) on synthetic data from simple linter rules. The authors show this enables a small 4B model to generalize to complex semantic PEP idioms, achieving SOTA detection recall on a custom benchmark."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.  The non-CoT model achieved the highest recall, while the CoT model had much higher precision. Does this suggest a fundamental precision/recall trade-off? Was an ensemble of the two models considered to potentially achieve the best of both?\n2.  Table 10 shows a trade-off based on the fraction of \"NO VIOLATIONS\" (NV) data in DPO training. The 0% NV model had the highest recall on the Ruff test set. What was the performance (P/R/F1) of this 0% NV model on the \"hard\" PEP benchmark? Is it possible the main SOTA recall claim is actually an underestimate?\n3.  The meta-task definition includes both a description ($D_I$) and examples ($E_I$). How sensitive is the model's performance to the quality and quantity of these examples? An ablation on few-shot vs. zero-shot (description only) instructions would be insightful."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.  The paper is well-written and addresses the critical problem of LLM-based code quality analysis, rightly identifying that models struggle to adapt to evolving best practices.\n2.  The proposed \"easy-to-hard\" generalization framework is novel and clever. It leverages instruction tuning on linter-generated synthetic data, avoiding the need for expensive manual annotation.\n3.  Using the linter itself as a \"verifiable reward model\" for preference optimization (RS-DPO) is a strong methodological contribution that provides a data-efficient path to generalization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  The \"meta-linting\" formulation appears to evaluate only one idiom specification at a time. This is a potential departure from real-world linters, which must check hundreds of rules simultaneously. It would be beneficial to explore how METALINT's performance scales when many idiom specifications are provided in a single prompt.\n2.  The paper convincingly shows that DPO enables generalization from \"easy\" to \"hard\" idioms, but the underlying why remains an interesting open question. It would be valuable to further investigate what the model is learning: a generalized concept of \"code quality,\" or a more general \"instruction-following\" capability. Further analysis here could strengthen this compelling hypothesis.\n3.  The SOTA claims on the \"hard\" PEP benchmark are very promising. However, the benchmark's current size (536 examples) is relatively modest. Expanding this benchmark in future work could help further solidify the robustness of these strong results.\n4.  The finding that the CoT model underperforms (lower recall) is counter-intuitive. The paper attributes this to \"overthinking.\" An alternative hypothesis worth exploring is that the RS-SFT data collection (filtering for perfect rewards) may have inadvertently trained the model to be overly conservative when facing ambiguity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917942635,"tcdate":1761995647309,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5203/Reviewer_ZLtP"],"signatures":["ICLR.cc/2026/Conference/Submission5203/Reviewer_ZLtP"],"forum":"Ue4PoLitpp","number":4,"license":"CC BY 4.0","cdate":1761995647309,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5203/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917942635,"domain":"ICLR.cc/2026/Conference","replyto":"Ue4PoLitpp","id":"wrfVabBjJ4","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["LLM4Code","Code Generation","Code Quality","Static Analysis","Transfer Learning","Post Training"]},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"abstract":{"value":"Large Language Models, though successful in code generation, struggle with code quality analysis because they are limited by static training data and can’t easily adapt to evolving best practices. We introduce MetaLint, an instruction-following framework that formulates code quality analysis as the task of detecting and fixing problematic semantic code fragments or code idioms based on high-level specifications. Unlike conventional approaches that train models on static code quality conventions, MetaLint employs instruction tuning on synthetic linter-generated data with dynamic conventions to support easy-to-hard generalization, enabling models to adapt to novel or complex code patterns without retraining. \nTo evaluate this, we construct a benchmark of challenging idioms inspired by real-world coding standards such as Python Enhancement Proposals (PEPs) \nand assess whether MetaLint-trained models reason adaptively or simply memorize. \nOur results show that MetaLint training improves generalization to unseen idioms. Qwen3-4B attains a 70.37% F-score on a manually curated and challenging PEP idiom detection benchmark, achieving the highest recall (70.43%) among all evaluated models. For localization, it reaches 26.73%, which is a strong outcome for its 4B parameter size and comparable to larger state-of-the-art models such as o3-mini, highlighting its potential for future-proof code quality analysis. Furthermore, MetaLint training enables generalization in idiom detection across model families, model scales, synthetic data from diverse linters, and Java idioms, demonstrating the general applicability of our approach."},"_bibtex":{"value":"@misc{\nnaik2026metalint,\ntitle={MetaLint: Generalizable Idiomatic Code Quality Analysis Through Instruction-Following and Easy-to-Hard Generalization},\nauthor={Atharva Naik and Lawanya Baghel and Dhatchinamoorthi Kunde Govindarajan and Darsh Agrawal and Daniel Fried and Carolyn Rose},\nyear={2026},\nurl={https://openreview.net/forum?id=Ue4PoLitpp}\n}"},"title":{"value":"MetaLint: Generalizable Idiomatic Code Quality Analysis Through Instruction-Following and Easy-to-Hard Generalization"},"pdf":{"value":"/pdf/d10d47ed2c277e0412eca1b0c71e368ca6da5458.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"naik|metalint_generalizable_idiomatic_code_quality_analysis_through_instructionfollowing_and_easytohard_generalization"},"authorids":{"value":["~Atharva_Naik1","~Lawanya_Baghel1","~Dhatchinamoorthi_Kunde_Govindarajan1","~Darsh_Agrawal1","~Daniel_Fried1","~Carolyn_Rose1"]},"authors":{"value":["Atharva Naik","Lawanya Baghel","Dhatchinamoorthi Kunde Govindarajan","Darsh Agrawal","Daniel Fried","Carolyn Rose"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a pixel-dependent noise variance stabilization method for X-ray imaging systems with radially symmetric beam profiles. The method extends standard noise variance stabilization (e.g., GAT) by incorporating spatially varying exposure due to the beam shape in flat-panel detector systems. The goal is to improve robustness and generalization of U-Net-based denoisers under non-uniform noise conditions. The approach is motivated by physical modeling of X-ray photon generation and detector response, where noise variance depends on both electronic and quantum components. A radial correction is applied to stabilize noise across the image before training a U-Net denoiser. The authors demonstrate improved uniformity of noise distribution and better qualitative denoising performance on phantom abdominal images."},"review":{"value":"- This is a well-motivated and practical work grounded in imaging physics, targeting a real issue in X-ray guided imaging systems: spatially varying noise due to beam geometry and flat-panel detector response. The idea of extending noise variance stabilization to account for radial beam profiles is reasonable and naturally follows from the limitations of standard global transforms like GAT.\n\n- The main strength of the paper is its strong physical intuition. The formulation of noise components (quantum + electronic) and the connection to beam-dependent intensity variation are clearly described. The proposed modification is also simple and computationally cheap, which is valuable for real-world deployment in medical imaging pipelines.\n\n- However, the core idea is essentially a spatially varying extension of existing noise stabilization techniques rather than a fundamentally new learning or modeling approach. While the physics motivation is strong, the actual algorithmic contribution is incremental.\n\n- The denoiser (U-Net) is treated as a fixed downstream model, and there is no exploration of whether the stabilization improves training dynamics, convergence, or generalization in a systematic way. The evaluation is also limited to qualitative and descriptive observations, with no quantitative comparison against standard preprocessing methods or ablations isolating the effect of radial correction."},"strengths":{"value":"- Strong physics-driven motivation based on realistic X-ray imaging systems, and addresses an important practical issue on spatially varying noise in MV imaging\n- Simple and computationally efficient preprocessing idea, and clear connection to established noise models (quantum + electronic noise)\n- Potentially useful for improving robustness of existing U-Net denoisers and well-aligned with real-world deployment constraints in medical imaging"},"weaknesses":{"value":"- Limited methodological novelty (extension of existing noise variance stabilization), and lack of quantitative comparison against standard preprocessing baselines\n- No ablation studies isolating contribution of radial correction, also the evaluation is mostly qualitative and phantom-based, I hope the authors addressed this in final version\n- Limited analysis of downstream impact on denoiser training and generalization\n- Algorithmic description of radial stabilization is not fully precise\n- Weak exploration of robustness across scanners, dose levels, or anatomies"},"confidence":{"value":4},"rating":{"value":4},"justification_of_rating":{"value":"The paper is strong in physical motivation and practical relevance, and the idea is useful for real X-ray denoising systems. However, from a methodological standpoint, it is an incremental extension of known noise stabilization techniques with limited experimental depth. It would be stronger with rigorous quantitative evaluation, clearer algorithmic formalization, and stronger comparison against existing preprocessing methods."},"title":{"value":"Physics-informed pixel-dependent noise stabilization for X-ray denoising; practical but limited methodological novelty"}},"parentInvitations":"MIDL.io/2026/Short_Papers/-/Official_Review","nonreaders":[],"tmdate":1778421298862,"tcdate":1778083431157,"writers":["MIDL.io/2026/Short_Papers","MIDL.io/2026/Short_Papers/Submission104/Reviewer_RDiN"],"signatures":["MIDL.io/2026/Short_Papers/Submission104/Reviewer_RDiN"],"forum":"7aQnZUhdD7","number":1,"license":"CC BY 4.0","cdate":1778083431157,"readers":["everyone"],"invitations":["MIDL.io/2026/Short_Papers/Submission104/-/Official_Review","MIDL.io/2026/Short_Papers/-/Edit"],"mdate":1778421298862,"domain":"MIDL.io/2026/Short_Papers","replyto":"7aQnZUhdD7","id":"HW0NB9zLBC","forumContent":{"TLDR":{"value":"Pixel-dependent Noise Varaince Stabilization to improve deep-learning based denoising of x-ray systems"},"venue":{"value":"MIDL 2026 - Short Papers Poster"},"keywords":{"value":["Image denoising","UNETs","Convolutional Neural Networks"]},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"X-ray guided medical procedures may expose the patient to a non-negligible amount of radiation dose. To mitigate the dose and reduce the risk of potentially correlated health issues, it is important to optimize the radiation exposure for both the patients and clinical staff. This means that the applied radiation dose should be as low as reasonably achievable while ensuring that the required image quality is reached. UNETs have become the state-of-the-art denoising algorithms. In this article we show a preprocessing algorithm, which improves the robustness and generalization of denosing models of systems whose images have position-dependent x-ray intensity, while also improving signal–noise discrimination."},"_bibtex":{"value":"@inproceedings{\narroyo2026pixeldependent,\ntitle={Pixel\\nobreakdash-Dependent Noise Variance Stabilization for Learning\\nobreakdash-Based Denoising in Imaging Systems with Radially Symmetric Beams},\nauthor={Pablo Corral Arroyo and Fasil Gadjimuradov and Sai Gokul Hariharan},\nbooktitle={Medical Imaging with Deep Learning - Short Papers},\nyear={2026},\nurl={https://openreview.net/forum?id=7aQnZUhdD7}\n}"},"title":{"value":"Pixel‑Dependent Noise Variance Stabilization for Learning‑Based Denoising in Imaging Systems with Radially Symmetric Beams"},"pdf":{"value":"/pdf/8766d5ad3735ac021d398449434b04e56d351606.pdf"},"visa":{"value":"Yes"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Short_Papers"},"paperhash":{"value":"arroyo|pixeldependent_noise_variance_stabilization_for_learningbased_denoising_in_imaging_systems_with_radially_symmetric_beams"},"authorids":{"value":["~Pablo_Corral_Arroyo1","~Fasil_Gadjimuradov1","~Sai_Gokul_Hariharan1"]},"registration":{"value":"Yes"},"authors":{"value":["Pablo Corral Arroyo","Fasil Gadjimuradov","Sai Gokul Hariharan"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"venue":{"value":"CoRR 2022"},"pdf":{"value":"http://arxiv.org/pdf/2212.04092v1"},"venueid":{"value":"dblp.org/journals/CORR/2022"},"paperhash":{"value":"dua|successive_prompting_for_decomposing_complex_questions"},"authorids":{"value":["~Dheeru_Dua1","~Shivanshu_Gupta2","https://dblp.org/search/pid/api?q=author:Sameer_Singh_0001:","https://dblp.org/search/pid/api?q=author:Matt_Gardner_0001:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2212.04092"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2212-04092,\n  publtype={informal},\n  author={Dheeru Dua and Shivanshu Gupta and Sameer Singh and Matt Gardner},\n  title={Successive Prompting for Decomposing Complex Questions},\n  year={2022},\n  cdate={1640995200000},\n  journal={CoRR},\n  volume={abs/2212.04092},\n  url={https://doi.org/10.48550/arXiv.2212.04092}\n}\n"},"abstract":{"value":"Answering complex questions that require making latent decisions is a challenging task, especially when limited supervision is available. Recent works leverage the capabilities of large language models (LMs) to perform complex question answering in a few-shot setting by demonstrating how to output intermediate rationalizations while solving the complex question in a single pass. We introduce ``Successive Prompting'', where we iteratively break down a complex task into a simple task, solve it, and then repeat the process until we get the final solution. Successive prompting decouples the supervision for decomposing complex questions from the supervision for answering simple questions, allowing us to (1) have multiple opportunities to query in-context examples at each reasoning step (2) learn question decomposition separately from question answering, including using synthetic data, and (3) use bespoke (fine-tuned) components for reasoning steps where a large LM does not perform well. The intermediate supervision is typically manually written, which can be expensive to collect. We introduce a way to generate a synthetic dataset which can be used to bootstrap a model's ability to decompose and answer intermediate questions. Our best model (with successive prompting) achieves an improvement of ~5% absolute F1 on a few-shot version of the DROP dataset when compared with a state-of-the-art model with the same supervision."},"title":{"value":"Successive Prompting for Decomposing Complex Questions"},"authors":{"value":["Dheeru Dua","Shivanshu Gupta","Sameer Singh","Matt Gardner"]}},"tmdate":1752709054702,"pdate":1640995200000,"tcdate":1747599759866,"writers":["~"],"signatures":["~Dheeru_Dua2"],"forum":"vNZ5Rnts3i","license":"CC BY-SA 4.0","number":522741,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1752709054702,"domain":"DBLP.org","id":"vNZ5Rnts3i","version":2},{"content":{"summary":{"value":"This paper addresses the critical yet under-explored issue of logical consistency in Large Vision-Language Models (LVLMs). The authors argue that while LVLMs have strong perceptual abilities, they often fail at complex visual reasoning tasks and produce contradictory answers to logically equivalent questions. To tackle this, the paper presents two main contributions:\n\n1. ConVBench: A new vision-centric benchmark designed to rigorously evaluate both complex reasoning and logical consistency. Each image in ConVBench is paired with two logically equivalent questions across six reasoning categories. The benchmark introduces two novel metrics: logical consistency and robust accuracy.\n\n2. ConVLM: A framework for improving LVLM reasoning by enforcing consistency. The method uses GRPO with a novel dual-reward mechanism. This reward combines a standard accuracy signal with a new consistency reward, which encourages the model to produce agreeing outputs for logically equivalent question pairs. Notably, the training data for this process is generated automatically by a powerful LVLM  and then validated by humans, making the approach scalable.\n\nThe authors demonstrate through extensive experiments that their ConVLM-7B model achieves state-of-the-art results among open-source models on ConVBench, significantly outperforming strong baselines. Furthermore, the model shows excellent generalization capabilities on other challenging benchmarks like V*Bench and InfoVQA, indicating that the consistency training imparts a more robust and generalizable reasoning ability."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The ablation study in Table 2 shows that training with only the consistency reward (w/o-Acc) leads to a dramatic improvement in both consistency and accuracy over the baseline. This is a fascinating result. Could you elaborate on why enforcing consistency provides such a strong implicit signal for accuracy?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper tackles a fundamental flaw in current LVLMs. Logical consistency is a cornerstone of reliable and trustworthy AI, and this work provides a formal framework to measure and improve it.\n\n2. ConVBench is a valuable asset for the field. Its design principles—focusing on vision-centric tasks, complex reasoning, and logically equivalent question pairs—fill an important gap in existing evaluation suites.\n\n3. The ConVLM framework is elegant and well-motivated. The key novelty lies in the dual-reward design that explicitly optimizes for consistency alongside accuracy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the concept is powerful, the examples shown primarily involve rephrasing or direct one-step implications (e.g., hitting a ball implies it's moving away). The paper could benefit from a more detailed discussion on the diversity and complexity of the logical relationships present in ConVBench.\n\n2. Proving the method's efficacy on STEM-related data, such as visual mathematics or physics problems (e.g., on benchmarks like MathVista), would greatly strengthen the paper's claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918575632,"tcdate":1762078362779,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6260/Reviewer_rAAu"],"signatures":["ICLR.cc/2026/Conference/Submission6260/Reviewer_rAAu"],"forum":"OoChIYXsfA","number":3,"license":"CC BY 4.0","cdate":1762078362779,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6260/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918575632,"domain":"ICLR.cc/2026/Conference","replyto":"OoChIYXsfA","id":"Umg13kfreu","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Robustness","Consistency","Large Vision Language Models","Multimodal"]},"primary_area":{"value":"generative models"},"abstract":{"value":"While Large Vision-Language Models (LVLMs) exhibit strong perceptual capabilities, they remain vulnerable in visual reasoning tasks. Existing benchmarks largely focus on symbolic mathematical or scientific problems and simple vision-centric tasks, offering limited assessment of complex visual reasoning and logical consistency, a critical requirement for reliable reasoning systems. We introduce ConVBench, a complex vision-centric reasoning benchmark where each image is paired with two logically equivalent questions across six categories: action and state, complex counting, spatial reasoning, causal and intent understanding, commonsense reasoning, and temporal perception. To complement this benchmark, we define two evaluation metrics, logical consistency and robust accuracy, that jointly assess both correctness and consistency of model responses. We further present ConVLM, which improves LVLM reasoning through Group Relative Policy Optimization (GRPO)-based reinforcement learning with novel consistency reward. This method leverages automatically generated logically equivalent question–answer pairs and a dual reward design combining accuracy- and consistency-based signals, encouraging agreement between paired responses. The framework functions effectively with or without strict answer supervision. On our ConVBench, ConVLM-7B achieves 73.36% logical consistency and 66.83% robust accuracy, setting a new state of the art among open-source models, and generalizes strongly to V*Bench (84.90% accuracy) and InfoVQA-test (81.90 ANLS)."},"_bibtex":{"value":"@misc{\njing2026be,\ntitle={Be Consistent! Enhancing Robust Visual Reasoning in {LVLM}s with Consistency Constraints},\nauthor={Liqiang Jing and Xiong Zhou and Siddharth Varia and Neha Anna John and Xinya Du and Vassilis N. Ioannidis},\nyear={2026},\nurl={https://openreview.net/forum?id=OoChIYXsfA}\n}"},"title":{"value":"Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints"},"pdf":{"value":"/pdf/580f7dfa6a06aed287f8fc9ad6ec006851a1b67c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"jing|be_consistent_enhancing_robust_visual_reasoning_in_lvlms_with_consistency_constraints"},"authorids":{"value":["~Liqiang_Jing1","~Xiong_Zhou2","~Siddharth_Varia2","~Neha_Anna_John1","~Xinya_Du1","~Vassilis_N._Ioannidis1"]},"authors":{"value":["Liqiang Jing","Xiong Zhou","Siddharth Varia","Neha Anna John","Xinya Du","Vassilis N. Ioannidis"]}},"version":2},{"content":{"venue":{"value":"ISBI 2007"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/4193196/4193197/04193515.pdf"},"venueid":{"value":"dblp.org/conf/ISBI/2007"},"paperhash":{"value":"hamarneh|quantification_and_visualization_of_localized_and_intuitive_shape_variability_using_a_novel_medialbased_shape_representation"},"authorids":{"value":["~Ghassan_Hamarneh1","https://dblp.org/search/pid/api?q=author:Aaron_D._Ward:","https://dblp.org/search/pid/api?q=author:Richard_Frank:"]},"html":{"value":"https://doi.org/10.1109/ISBI.2007.357081"},"_bibtex":{"value":"@inproceedings{DBLP:conf/isbi/HamarnehWF07,\n  author={Ghassan Hamarneh and Aaron D. Ward and Richard Frank},\n  title={Quantification and Visualization of Localized and Intuitive Shape Variability Using a Novel Medial-Based Shape Representation},\n  year={2007},\n  cdate={1167609600000},\n  pages={1232-1235},\n  url={https://doi.org/10.1109/ISBI.2007.357081},\n  booktitle={ISBI},\n  crossref={conf/isbi/2007}\n}\n"},"abstract":{"value":"Quantification and visualization of anatomical shape variability in different populations is essential for diagnosis and tracking progression of diseases. We present a new 3D medial-based shape representation method capable of analysis and visualization of 3D anatomy and demonstrate its ability to quantify and highlight shape variability in an intuitive manner. 3D shapes are represented via orientations and elongations of one or more medial sheets, along with thickness values encoding the distances to the shape surface. Two parameters traverse each medial sheet and are mapped to orientation, elongation, and thickness values; we call this map a medial patch. Shape variability is decomposed intuitively into bend, stretch, or bulge deformations, via operators acting on the components of the medial patch. In a simple manner, the location, extent, type, and amplitude of the deformation operators can be specified to capture local and global intuitive shape variability. We demonstrate the capabilities and intuitiveness of this approach through synthetic 3D shape deformations, as well as deformations that capture the 3D shape of an anatomical structure. We demonstrate the ability to highlight regions containing specific types of intuitive changes in anatomy."},"title":{"value":"Quantification and Visualization of Localized and Intuitive Shape Variability Using a Novel Medial-Based Shape Representation"},"authors":{"value":["Ghassan Hamarneh","Aaron D. Ward","Richard Frank"]}},"tmdate":1727454652301,"pdate":1167609600000,"tcdate":1727453020945,"writers":["~"],"signatures":["~Ghassan_Hamarneh1"],"forum":"dDF15vlcT5","license":"CC BY-SA 4.0","number":86414,"cdate":1167609600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727454652301,"domain":"DBLP.org","id":"dDF15vlcT5","version":2},{"content":{"summary":{"value":"The paper presents a novel method for extracting per point topological features - TOPF. The method builds on previous results in topological data analysis which described a shape or a point cloud with a single global feature, by generating per-point topologically-aware features. The paper presents a quantitative evaluation and comparison of the proposed method with prior art on a new benchmark consisting of several synthetic examples, evaluates the robustness of the proposed method under noise, as well as presents qualitative examples of its performance on synthetic and real work data."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"* What is the runtime of the proposed method? How does it change with the point cloud size and does it have limitations on the size of point cloud that can be processed with it?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"* The paper is well written and easy to follow. Prior art and the proposed algorithm description is detailed and comprehensive.\n* To my understanding, the paper describes a novel method for per-point feature extraction based on topological information contained in a point cloud, and describes theoretical guarantees for its correctness on point clouds sampled from multiple n-spheres.\n* The paper describes a new topological point clustering benchmark dataset consisting of seven synthetic point clouds with up to 5 labels, and evaluate the proposed and existing methods on this dataset showing that the proposed method outperforms existing methods in most cases."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The paper lists common machine learning applications requiring point level features as a motivation for the proposed method. However, only quantitative experiments for point cloud clustering on a set of synthetic examples, and anecdotal evidence of performance on real world data, were presented. In order to fully understand the potential of the proposed approach to be applied beyond synthetic data, it would be beneficial to include additional evaluation, qualitative and quantitative, on real-world data and additional applications, e.g. as described in lines 304-307.\n* Specifically, it would be interesting to see experiments on non-synthetic datasets with topological structure mentioned in line 266.\n* Additionally, comparison with other well performing modern machine learning methods, such as graph neural networks for point cloud clustering, needs to be discussed, for completeness."},"limitations":{"value":"The authors adequately addressed the limitations and impact of the proposed approach."}},"nonreaders":[],"tmdate":1730879851302,"tcdate":1722153148406,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission16537/Reviewer_mNMA"],"signatures":["NeurIPS.cc/2024/Conference/Submission16537/Reviewer_mNMA"],"forum":"jwE1dgOox1","number":4,"license":"CC BY 4.0","cdate":1722153148406,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission16537/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879851302,"domain":"NeurIPS.cc/2024/Conference","replyto":"jwE1dgOox1","id":"n20y5o4d4l","forumContent":{"venue":{"value":"Submitted to NeurIPS 2024"},"keywords":{"value":["Topological Data Analysis","TDA","Hodge Laplacian","Higher-Order Networks","Simplicial Complexes","Algebraic Topology","Differential Geometry","Point Clouds","Persistent Homology"]},"primary_area":{"value":"other"},"abstract":{"value":"Topological Data Analysis (TDA) allows us to extract powerful topological, and higher-order information on the global shape of a data set or point cloud. Tools like Persistent Homology or the Euler Transform give a single complex description of the global structure of the point cloud. However, common machine learning applications like classification require point-level information and features to be available. In this paper, we bridge this gap and propose a novel method to extract node-level topological features from complex point clouds using discrete variants of concepts from algebraic topology and differential geometry. We verify the effectiveness of these topological point features (TOPF) on both synthetic and real-world data and study their robustness under noise."},"_bibtex":{"value":"@misc{\nanonymous2024nodelevel,\ntitle={Node-Level Topological Representation Learning on Point Clouds},\nauthor={Anonymous},\nyear={2024},\nurl={https://openreview.net/forum?id=jwE1dgOox1}\n}"},"title":{"value":"Node-Level Topological Representation Learning on Point Clouds"},"pdf":{"value":"/pdf/b32be92f73e27fca80ea7684793aa63ae77f9a83.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"grande|nodelevel_topological_representation_learning_on_point_clouds"},"authorids":{"value":["~Vincent_Peter_Grande1","~Michael_T_Schaub1"]},"authors":{"value":["Vincent Peter Grande","Michael T Schaub"]}},"version":2},{"content":{"summary":{"value":"This paper critically re-evaluates the utility of synthetic data generated by driving world models for downstream perception tasks. The authors argue that previous works often rely on an unfair evaluation protocol where models trained on hybrid (real + synthetic) data use twice the training epochs of real-data-only baselines. They demonstrate that when training epochs are matched, the benefits of existing synthetic data augmentation methods become negligible or even negative.\n\nTo address this, the paper introduces **Dream4Drive**, a novel framework for generating high-quality, 3D-aware synthetic data. Instead of relying on sparse conditioning signals like BEV maps, Dream4Drive first decomposes a real video into a set of dense, 3D-aware guidance maps (depth, normal, edge, etc.). It then renders 3D assets into these maps and fine-tunes a Diffusion Transformer-based world model to generate photorealistic, multi-view videos that incorporate these new objects. This approach allows for precise, instance-level control and ensures geometric and visual consistency.\n\nFurthermore, the authors contribute **DriveObj3D**, a large-scale 3D asset dataset tailored for driving scenarios, along with an automated pipeline for its creation. Through extensive experiments on detection and tracking tasks, the paper shows that augmenting the training set with a very small fraction (<2%) of data generated by Dream4Drive consistently and significantly improves perception performance, even under fair, epoch-matched comparisons."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Regarding the `DriveObj3D` creation pipeline: Could you quantify the robustness of this pipeline? What is the approximate failure rate, and what are the common failure modes? How much manual filtering or intervention was required to curate the final high-quality dataset?\n2. The scene selection for insertion is described as choosing frames \"where no other vehicles appear along the insertion trajectory.\" How is this selection performed (manually or automatically)? Could this process introduce a bias, for example, by preferentially selecting less cluttered scenes, thereby limiting the complexity of the generated data?\n3. The paper highlights improved realism via shadows and reflections. Does the generative model also learn to synthesize more complex physical interactions? For example, does an inserted vehicle generate splashes when moving through a puddle, or create dust on a dirt road? To what extent can the model handle nuanced lighting effects beyond direct shadows, such as colored light from traffic signals reflecting on the vehicle?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. **Important and Timely Critique:** The central argument regarding the unfair evaluation of synthetic data is a significant contribution. By demonstrating that simply increasing training epochs on real data can match or exceed the performance of hybrid-data training, the paper forces the community to reconsider how the value of synthetic data is measured. This provides a strong and compelling motivation for the proposed work.\n2. **Comprehensive and Convincing Experiments:** The experimental validation is thorough and directly supports the paper's claims.\n    - The head-to-head comparisons under 1x, 2x, and 3x epoch settings provide clear evidence for the effectiveness of Dream4Drive over baselines.\n    - The ablation studies are insightful, systematically analyzing the impact of insertion position, distance, 3D asset source, and different components of the guidance maps. This provides a deeper understanding of what makes the synthetic data effective.\n    - The evaluation on both detection and tracking tasks, at multiple resolutions, demonstrates the general applicability of the generated data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited Scope of Generated Corner Cases:** The paper defines \"corner cases\" primarily as the insertion of new objects into a scene. While this is an important class of long-tail events, it does not cover other critical scenarios, such as adverse weather conditions (heavy rain, snow, fog), unusual lighting (lens flare, low sun), complex multi-agent interactions, or environmental changes (e.g., road construction). The framework's applicability to these other types of corner cases is not explored.\n2. **Potential Bottlenecks in the Asset Generation Pipeline:** The `DriveObj3D` pipeline is a cascade of multiple sophisticated models (segmentation, multi-view generation, mesh generation). The final asset quality is contingent on the successful execution of every step. This pipeline may be brittle; for instance, a failure in multi-view generation could lead to an incomplete or distorted 3D mesh. A discussion on the robustness, failure modes, and potential need for manual curation of this pipeline would strengthen the paper.\n3. **Limited Technical Contribution.** While the specific application and the system-level design are effective in research application, the work does not introduce a new core generative modeling technique or a new research framework. For some reviewers, the novelty might be perceived as incremental, lying more in the clever combination of existing parts than in foundational innovation.\n4. **Computational Cost Analysis is Missing:** The proposed data generation process is multi-staged and involves fine-tuning a large Diffusion Transformer model. This suggests a significant computational cost. While the paper argues for fairness in terms of training *epochs*, a discussion on fairness in terms of total *compute* (cost of data synthesis + cost of training) would provide a more complete picture. For instance, how does the cost of generating 420 synthetic samples compare to the cost of training the perception model for an additional epoch?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915650464,"tcdate":1760583275330,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission971/Reviewer_yxwm"],"signatures":["ICLR.cc/2026/Conference/Submission971/Reviewer_yxwm"],"forum":"z3cFADf6zZ","number":1,"license":"CC BY 4.0","cdate":1760583275330,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission971/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915650464,"domain":"ICLR.cc/2026/Conference","replyto":"z3cFADf6zZ","id":"XNRUbwVNji","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Autonomous Driving","Driving World Model","Perception Tasks","Synthetic Data"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Recent advancements in driving world models enable controllable generation of high-quality RGB videos or multimodal videos. \nExisting methods primarily focus on metrics related to generation quality and controllability. \nHowever, they often overlook the evaluation of downstream perception tasks, which are {\\bf really crucial} for the performance of autonomous driving. \nExisting methods usually leverage a training strategy that first pretrains on synthetic data and finetunes on real data, resulting in twice the epochs compared to the baseline (real data only). \nWhen we double the epochs in the baseline, the benefit of synthetic data becomes negligible.\nTo thoroughly demonstrate the benefit of synthetic data, we introduce Dream4Drive, a novel synthetic data generation framework designed for enhancing the downstream perception tasks.\nDream4Drive first decomposes the input video into several 3D-aware guidance maps and subsequently renders the 3D assets onto these guidance maps.\nFinally, the driving world model is fine-tuned to produce the edited, multi-view photorealistic videos, which can be used to train the downstream perception models.\nDream4Drive enables unprecedented flexibility in generating multi-view corner cases at scale, significantly boosting corner case perception in autonomous driving. \nTo facilitate future research, we also contribute a large-scale 3D asset dataset named DriveObj3D, covering the typical categories in driving scenarios and enabling diverse 3D-aware video editing.\nWe conduct comprehensive experiments to show that Dream4Drive can effectively boost the performance of downstream perception models under various training epochs. \nProject website: \\url{https://wm-research.github.io/Dream4Drive/}."},"_bibtex":{"value":"@inproceedings{\nzeng2026rethinking,\ntitle={Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks},\nauthor={Kai Zeng and Zhanqian Wu and Kaixin Xiong and Xiaobao Wei and Xiangyu Guo and Zhenxin Zhu and Kalok Ho and Lijun Zhou and Bohan Zeng and Ming Lu and Haiyang Sun and BING WANG and Guang Chen and Hangjun Ye and Wentao Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=z3cFADf6zZ}\n}"},"title":{"value":"Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks"},"pdf":{"value":"/pdf/71bda5c5ed57972d48727bc4da82c6856d114ee8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zeng|rethinking_driving_world_model_as_synthetic_data_generator_for_perception_tasks"},"authorids":{"value":["~Kai_Zeng7","~Zhanqian_Wu1","~Kaixin_Xiong1","~Xiaobao_Wei1","~Xiangyu_Guo5","~Zhenxin_Zhu1","~Kalok_Ho1","~Lijun_Zhou1","~Bohan_Zeng1","~Ming_Lu2","~Haiyang_Sun2","~BING_WANG23","~Guang_Chen2","~Hangjun_Ye1","~Wentao_Zhang1"]},"authors":{"value":["Kai Zeng","Zhanqian Wu","Kaixin Xiong","Xiaobao Wei","Xiangyu Guo","Zhenxin Zhu","Kalok Ho","Lijun Zhou","Bohan Zeng","Ming Lu","Haiyang Sun","BING WANG","Guang Chen","Hangjun Ye","Wentao Zhang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces EnvSocial-Diff, a diffusion-based crowd simulation model informed by social physics and augmented with two novel modules:\n(1) a structured environmental conditioning mechanism that explicitly encodes obstacles, objects of interest (OOI), and lighting conditions, and\n(2) an Individual–Group Interaction (IGI) module that captures both fine-grained interpersonal dynamics and group-level conformity via graph neural networks.\n\nThe model extends the Social Physics Informed Diffusion Model (SPDiff, Chen et al., AAAI 2024) by incorporating richer environmental signals and multi-level social reasoning into the generative diffusion process.\nExperiments on the GC and UCY datasets demonstrate state-of-the-art performance across multiple trajectory prediction metrics (MAE, FDE, OT, MMD, DTW, and collision count). Ablation studies confirm the contributions of environmental factors and the IGI module, while qualitative visualizations support the interpretability and realism of simulated trajectories."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"> Could the authors clarify whether the environmental encoders are trained jointly with the diffusion network or frozen (especially the ResNet and BERT backbones)?\n\n> How does EnvSocial-Diff perform under dynamic environments (e.g., moving obstacles)?\n\n> Have the authors considered testing the model’s capacity for long-horizon rollouts (>5s) to evaluate accumulation of social-environmental errors?\n\n> Could the lighting conditioning be replaced or augmented by other perceptual features (e.g., crowd density maps, saliency maps)?\n\n> Is the approach compatible with large-scale agent-based simulations (e.g., thousands of pedestrians), or does the GNN limit scalability?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"+ Novel Integration: Elegant fusion of social physics and diffusion modeling with explicit environmental conditioning.\n\n+ Interpretability: Maintains physically grounded meaning for forces and accelerations.\n\n+ Comprehensive Evaluation: Multiple datasets, metrics, and ablations validate both performance and generalization.\n\n+ General Applicability: Applicable to domains such as simulation, safety planning, and digital twin environments.\n\n+ Strong Theoretical Foundation: Builds directly on the Social Force Model while extending its scope through learnable conditioning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Computational Complexity: The paper does not report training/inference times or resource comparisons versus SPDiff or data-driven baselines. This limits understanding of scalability in real-time simulation.\n\n- Limited Dataset Diversity: Experiments rely mainly on GC and UCY datasets. These are standard but relatively small; inclusion of additional or synthetic datasets (e.g., ETH, SDD) would strengthen generalization claims.\n\n- Lighting Factor Validation: The contribution of the lighting module is modest and potentially dataset-specific. A more detailed justification (e.g., psychophysical rationale or ablation under controlled illumination changes) would improve the argument.\n\n- Minor Clarity Issues: Some notation (e.g., the dual use of f_{\\text{light}} and \\tilde{f}_{\\text{light}}) could be clarified (e.g., f_{\\text{light}}^{\\text{raw}} and f_{\\text{light}}^{\\text{enc}}) for better readability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921948314,"tcdate":1761960568235,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10714/Reviewer_DkCA"],"signatures":["ICLR.cc/2026/Conference/Submission10714/Reviewer_DkCA"],"forum":"2XBAm3Dbnt","number":2,"license":"CC BY 4.0","cdate":1761960568235,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10714/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921948314,"domain":"ICLR.cc/2026/Conference","replyto":"2XBAm3Dbnt","id":"rl7NS8NZis","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"A diffusion-based crowd simulation model with environmental conditioning and individual-group interaction."},"keywords":{"value":["Crowd simulation","Social physics force","Diffusion model"]},"supplementary_material":{"value":"/attachment/b5e0a6a16bde36ba1f6b943f605e484f6315ea8e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Modeling realistic pedestrian trajectories requires accounting for both social interactions and environmental context, yet most existing approaches largely emphasize social dynamics. We propose EnvSocial-Diff: a diffusion-based crowd simulation model informed by social physics and augmented with environmental conditioning and individual-group interaction. Our structured environmental conditioning module explicitly encodes obstacles, objects of interest, and lighting levels, providing interpretable signals that capture scene constraints and attractors. In parallel, the individual-group interaction module goes beyond individual-level modeling by capturing both fine-grained interpersonal relations and group-level conformity through a graph-based design. Experiments on multiple benchmark datasets demonstrate that EnvSocial-Diff outperforms the latest state-of-the-art methods, underscoring the importance of explicit environmental conditioning and multi-level social interaction for realistic crowd simulation."},"_bibtex":{"value":"@inproceedings{\nzhao2026envsocialdiff,\ntitle={EnvSocial-Diff: A Diffusion-Based Crowd Simulation Model with Environmental Conditioning and Individual-Group Interaction},\nauthor={Bingxue Zhao and Qi Zhang and Hui Huang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=2XBAm3Dbnt}\n}"},"title":{"value":"EnvSocial-Diff: A Diffusion-Based Crowd Simulation Model with Environmental Conditioning and Individual-Group Interaction"},"pdf":{"value":"/pdf/6be61c23c2e7def47a9666339c640097edb54588.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhao|envsocialdiff_a_diffusionbased_crowd_simulation_model_with_environmental_conditioning_and_individualgroup_interaction"},"authorids":{"value":["~Bingxue_Zhao1","~Qi_Zhang11","~Hui_Huang3"]},"authors":{"value":["Bingxue Zhao","Qi Zhang","Hui Huang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Physics-Informed Distillation of Diffusion Models (PIDDM), a two-stage framework that enforces PDE constraints on generated samples rather than on posterior means, addressing the *Jensen’s gap* in existing methods. It claims to achieve *both physical fidelity and generative quality* while enabling one or a few-step inference. The authors demonstrate broad applicability across forward, inverse, and reconstruction tasks in PDE generation settings."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"**Q1**: Has the method been evaluated under out-of-distribution constraint conditions, where the enforced constraints differ from those observed in the training distribution?\n\n**Q2.** Can the authors disentangle the contributions of the physics-informed residual loss and the distillation process itself? For instance, how does a pure distillation baseline (without residual loss) compare to PIDDM in terms of residual error and generative quality?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"**S1.** The paper presents a clear and well-motivated solution to the Jensen’s gap problem in physics-informed diffusion models through a principled post-hoc distillation framework that enforces constraints directly on generated samples.\n\n**S2.** The approach is conceptually simple yet effective, showing strong and consistent performance across diverse PDE benchmarks and tasks such as forward, inverse, and reconstruction problems, outperforming competitive baselines.\n\n**S3.** The experiments include well-designed ablations and illustrative toy studies, and the paper is clearly written and easy to follow, making the contributions accessible and convincing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**W1.** The theoretical and methodological novelty of the work is incremental. While the paper presents a clear and convincing empirical illustration of the Jensen’s gap, the idea of distillation in constrained generative modeling has already been explored in frameworks such as rectified flows and consistency models. The proposed post-hoc distillation strategy largely extends existing one-step distillation approaches rather than introducing a fundamentally new formulation. It also remains unclear whether the observed gains stem primarily from the physics-informed residual loss or simply from the benefits of distillation itself. Furthermore, the stronger-performing variant, PIDDM-ref, achieves its improvements through gradient-based noise refinement that closely mirrors the optimization procedure of D-Flow, making the most effective component of the method derivative of established approaches.\n\n\n**W2.** The datasets used in the paper are relatively simple and low-dimensional, which limits the strength and generalizability of the empirical conclusions. In several benchmarks, both the teacher and the ECI baseline already achieve very low PDE residuals, often near numerical precision, making the reported improvements marginal in absolute terms. The paper should clarify the dataset composition, including training and validation sizes, and evaluate whether PIDDM maintains its advantages on more challenging settings. Since the distillation formulation effectively optimizes the residual loss as a constraint, it would be valuable to test its performance on harder PDEs where optimizing PINN objectives is known to exhibit failure modes, such as those described in [1, 2].\n\n\n**W3.** The paper would benefit from a more detailed computational analysis. Although PIDDM-1 and PIDDM-ref are distinguished conceptually, the work does not present explicit runtime or wall-clock comparisons. Since PIDDM-ref incorporates gradient-based refinement at inference, while other baselines rely on multi-step sampling with varying computational costs, providing wall-clock times would help clarify the true efficiency trade-offs and strengthen the empirical claims. Additionally, the related work section could be more focused; emphasizing the most directly comparable methods such as [3] and [4] would help clarify the paper’s position and distinct contributions.\n\n\n**References**\n\n[1] Krishnapriyan, Aditi, et al. *Characterizing Possible Failure Modes in Physics-Informed Neural Networks.* *Advances in Neural Information Processing Systems* 34 (2021): 26548–26560.  \n\n[2] Rathore, Pratik, et al. *Challenges in Training PINNs: A Loss Landscape Perspective.* arXiv preprint arXiv:2402.01868 (2024).  \n\n[3] Utkarsh, U., et al. *Physics-Constrained Flow Matching: Sampling Generative Models with Hard Constraints.* arXiv preprint   arXiv:2506.04171 (2025).  \n\n[4] Christopher, J. K., S. Baek, and N. Fioretto. *Constrained Synthesis with Projected Diffusion Models.* *Advances in Neural Information Processing Systems* 37 (2024): 89307–89333."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923038368,"tcdate":1761936337906,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12068/Reviewer_Fu51"],"signatures":["ICLR.cc/2026/Conference/Submission12068/Reviewer_Fu51"],"forum":"hW7P3x9W8A","number":4,"license":"CC BY 4.0","cdate":1761936337906,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12068/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923038368,"domain":"ICLR.cc/2026/Conference","replyto":"hW7P3x9W8A","id":"AK04h4RfQZ","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Diffusion","Physical Sciences"]},"supplementary_material":{"value":"/attachment/bdf75a4d6c308d3a5a9ea981e5401828c7091e06.zip"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Modeling physical systems in a generative manner offers several advantages, including the ability to handle partial observations, generate diverse solutions, and address both forward and inverse problems. Recently, diffusion models have gained increasing attention in the modeling of physical systems, particularly those governed by partial differential equations (PDEs). However, diffusion models only access noisy data $\\boldsymbol{x}_t$ at intermediate steps, making it infeasible to directly enforce constraints on the clean sample $\\boldsymbol{x}_0$ at each noisy level. As a workaround, constraints are typically applied to the expectation of clean samples $\\mathbb{E}[\\boldsymbol{x}_0|\\boldsymbol{x}_t]$, which is estimated using the learned score network. However, imposing PDE constraints on the expectation does not strictly represent the one on the true clean data, known as Jensen's Gap. This gap creates a trade-off: enforcing PDE constraints may come at the cost of reduced accuracy in generative modeling. To address this, we propose a simple yet effective post-hoc distillation approach, where PDE constraints are not injected directly into the diffusion process, but instead enforced during a post-hoc distillation stage. We term our method as Physics-Informed Distillation of Diffusion Models (PIDDM). This distillation not only facilitates single-step generation with improved PDE satisfaction, but also support both forward and inverse problem solving and reconstruction from randomly partial observation. Extensive experiments across various PDE benchmarks demonstrate that PIDDM significantly both improves PDE satisfaction and generative modeling over several recent and competitive baselines, such as PIDM, DiffusionPDE, and ECI-sampling, while achieving lower computational overhead and avoiding extensive hyperparameter tuning. Our approach can shed light on more efficient and effective strategies for incorporating physical constraints into diffusion models."},"_bibtex":{"value":"@misc{\nzhang2025physicsinformed,\ntitle={Physics-Informed Distillation of Diffusion Models for {PDE}-Constrained Generation},\nauthor={Yi Zhang and Difan Zou},\nyear={2025},\nurl={https://openreview.net/forum?id=hW7P3x9W8A}\n}"},"title":{"value":"Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation"},"pdf":{"value":"/pdf/17d887d71f2708c8581d2d7408e3bf48c01acea0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|physicsinformed_distillation_of_diffusion_models_for_pdeconstrained_generation"},"authorids":{"value":["~Yi_Zhang94","~Difan_Zou1"]},"authors":{"value":["Yi Zhang","Difan Zou"]}},"version":2},{"content":{"venue":{"value":"CVPR 2025"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2025/papers/Li_FreeGave_3D_Physics_Learning_from_Dynamic_Videos_by_Gaussian_Velocity_CVPR_2025_paper.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2025"},"paperhash":{"value":"li|freegave_3d_physics_learning_from_dynamic_videos_by_gaussian_velocity"},"authorids":{"value":["~Jinxi_Li1","https://dblp.org/search/pid/api?q=author:Ziyang_Song:","https://dblp.org/search/pid/api?q=author:Siyuan_Zhou:","https://dblp.org/search/pid/api?q=author:Bo_Yang_0027:"]},"html":{"value":"https://openaccess.thecvf.com/content/CVPR2025/html/Li_FreeGave_3D_Physics_Learning_from_Dynamic_Videos_by_Gaussian_Velocity_CVPR_2025_paper.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/LiSZY25,\n  author={Jinxi Li and Ziyang Song and Siyuan Zhou and Bo Yang},\n  title={FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity},\n  year={2025},\n  cdate={1735689600000},\n  pages={12433-12443},\n  url={https://openaccess.thecvf.com/content/CVPR2025/html/Li_FreeGave_3D_Physics_Learning_from_Dynamic_Videos_by_Gaussian_Velocity_CVPR_2025_paper.html},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2025}\n}\n"},"abstract":{"value":"In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing works often fail to learn complex physical motions at boundaries or require object priors such as masks or types. In this paper, we propose FreeGave to learn physics of complex dynamic 3D scenes without needing any object priors. The key to our approach is to introduce a physics code followed by a carefully designed divergence-free module for estimating a per-Gaussian velocity field, without relying on the inefficient PINN losses. Extensive experiments on three public datasets and a newly collected challenging real-world dataset demonstrate the superior performance of our method for future frame extrapolation and motion segmentation. Most notably, our investigation into the learned physics codes reveals that they truly learn meaningful 3D physical motion patterns in the absence of any human labels in training. Our code and data are available at https://github.com/vLAR-group/FreeGave."},"title":{"value":"FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity"},"authors":{"value":["Jinxi Li","Ziyang Song","Siyuan Zhou","Bo Yang"]}},"tmdate":1762939418808,"pdate":1735689600000,"externalIds":["dblp:conf/cvpr/LiSZY25"],"tcdate":1762935949154,"writers":["~"],"signatures":["~Jinxi_Li2"],"forum":"cHOenCvb03","license":"CC BY-SA 4.0","number":691134,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762939418808,"domain":"DBLP.org","id":"cHOenCvb03","version":2},{"content":{"summary":{"value":"The paper proposes PhySTA, a physics-inspired framework that unifies continuous operator learning with graph-based spatio-temporal modeling. It introduces two main modules:\n(1) a Graph–Time Fourier Neural Operator (GT-FNO) equipped with Time-Gated Spectral Segmentation Perception (TGSSP) for modeling continuous spectral dynamics on graphs, and\n(2) an Adaptive Multi-Scale Interaction (AMI) mechanism that captures multi-scale node–edge relationships via coupled convolution and hierarchical graph construction.\nA Continuity–Discreteness Interaction Module (CDIM) further fuses both continuous and discrete predictions for arbitrary inference in unobserved regions.\nExperiments on large-scale traffic and air-quality datasets demonstrate strong accuracy, robustness, and efficiency compared with several state-of-the-art baselines."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Novel integration of physics-inspired operator learning and GNNs:\nThe proposed GT-FNO extends Fourier Neural Operators to non-Euclidean graphs, enabling continuous modeling over directed graphs—a clear conceptual innovation.\n\n2. Multi-scale and coupled graph design:\nThe AMI module effectively captures long-range, multi-level dependencies within a single layer, addressing over-smoothing and inefficiency issues seen in deep GNNs.\n\n3. Strong empirical performance and efficiency:\nPhySTA achieves consistent improvements across datasets with fewer parameters and memory cost (≈74% FLOP reduction), showing excellent trade-offs between accuracy and scalability.\n\n4. Comprehensive experiments and ablation analysis:\nThe inclusion of multiple datasets, mask ratios, and detailed component ablations provides good evidence of robustness and interpretability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Limited comparison to recent operator-based or physics-informed baselines:\nThe paper mainly compares against classical and GNN-based methods (STGCN, DGCRN, etc.), but omits recent neural operator or PDE-based baselines such as Graph Neural Operator (Li et al., 2023) or Geo-FNO (Li et al., 2024). These would strengthen the claim of operator generalization.\n\n* Writing quality and presentation:\nThe exposition is heavy and sometimes unclear, especially in the methodology section. Some mathematical notations are inconsistent, and figures (e.g., Fig. 2, Fig. 3) are not fully self-explanatory. The authors could simplify and streamline the presentation for readability.\n\n* Ablation and interpretability could be expanded:\nAlthough the ablation table is informative, qualitative insights on how each frequency band or subgraph level contributes to the final prediction are missing. Visualizations of spectral energy distribution or temporal gating behavior would enhance interpretability.\n\n* Scalability limitation not sufficiently addressed:\nThe reliance on magnetic Laplacian spectral decomposition may hinder scalability for very large graphs. While this is briefly mentioned in the limitations, empirical evaluation on larger graphs would make the claim more convincing.\n\n* Incomplete baseline coverage:\nSome recent transformer-based and neural-operator hybrid methods (e.g., Graphormer, SpaceTimeFormer) are missing from comparison, which may weaken the “state-of-the-art” claim."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928097841,"tcdate":1761794529668,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18391/Reviewer_sBzk"],"signatures":["ICLR.cc/2026/Conference/Submission18391/Reviewer_sBzk"],"forum":"b6Py2zy0fK","number":3,"license":"CC BY 4.0","cdate":1761794529668,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18391/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928097841,"domain":"ICLR.cc/2026/Conference","replyto":"b6Py2zy0fK","id":"0NUzDvo5nC","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Neural operators","Spatio-temporal systems","Graph neural networks","Data mining"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Modern spatio-temporal learning techniques usually exploit sampled discrete observations to foresee the future. Actually, spatio-temporal dynamics are continuous and evolve continuously across time and space, thus modeling  spatio-temporal dynamics in a continuous space can be a long-standing challenge. Existing deep learning architectures often fail to generalize to unseen regions and new graph topologies, while many physics-driven approaches are confined to Euclidean grids and  poorly scale to complex graph structures. To address this gap, we propose PhySTA, a physics-inspired spatio-temporal learning framework designed for efficient and scalable arbitrary inference over graph-structured data. PhySTA integrates two key modules: (1) Continuous Operator-based Spectrum-Temporal Learning (CoSTL), which leverages a Graph-Time Fourier Neural Operator combined with Time-Gated Spectral Segmentation Perception to model continuous dynamics in operator space, and (2) Adaptive Multi-scale Interaction (AMI) that constructs multi-scale subgraphs and introduces node-edge coupled convolution to capture discrete interaction patterns and refine continuous predictions. By bridging operator learning with node-edge-graph interaction, PhySTA achieves both continuity-aware dynamic modeling and hierarchical interactive refinement. Extensive experiments across large-scale benchmarks demonstrate that PhySTA attains state-of-the-art accuracy while reducing computation cost and lowering parameter overhead."},"_bibtex":{"value":"@inproceedings{\nge2026enabling,\ntitle={Enabling arbitrary inference in spatio-temporal dynamic systems: A physics-inspired perspective},\nauthor={Yan Ge and Zhengyang Zhou and Qihe Huang and Yuxuan Liang and Yang Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=b6Py2zy0fK}\n}"},"title":{"value":"Enabling arbitrary inference in spatio-temporal dynamic systems: A physics-inspired perspective"},"pdf":{"value":"/pdf/16958b5e7cceb167f7c5695f4723e40d08a60268.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ge|enabling_arbitrary_inference_in_spatiotemporal_dynamic_systems_a_physicsinspired_perspective"},"authorids":{"value":["~Yan_Ge5","~Zhengyang_Zhou1","~Qihe_Huang2","~Yuxuan_Liang1","~Yang_Wang32"]},"authors":{"value":["Yan Ge","Zhengyang Zhou","Qihe Huang","Yuxuan Liang","Yang Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes AutoDrive-R^2, a novel VLA framework designed to enhance the reasoning and self-reflection capabilities of autonomous driving systems, addressing the limitations of existing methods such as physically infeasible trajectory generation and inadequate reasoning for complex scenarios. The framework adopts a two-stage training approach: in the first stage, a CoT dataset named nuScenesR^2-6K (with 6,000 image-trajectory pairs) is constructed for SFT, which guides the model through a four-step logical chain to build cognitive connections between input information and output trajectories. In the second stage, a physics-grounded reward framework integrated with GRPO is employed for RL, incorporating spatial alignment, vehicle dynamics, and temporal smoothness constraints to ensure trajectory feasibility. Experimental results on nuScenes and Waymo datasets demonstrate that AutoDrive-R^2 achieves state-of-the-art  performance and robust zero-shot generalization."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"For the dataset: How were the manual annotations of the CoT reasoning process validated? How to ensure the consistency of reasoning steps? Additionally, since the dataset is derived from nuScenes, does it inherit the scene bias of the original dataset (e.g., urban road dominance), and if so, how does this affect the model's generalization to non-urban scenarios (e.g., highways or rural roads)?\n\nThe paper sets all weight coefficients (λ_pos, λ_ste, λ_vel, λ_tem) to 1 in experiments. Have you tested the impact of different weight combinations on performance? Besides, have you analyzed the impact of GRPO hyperparameters (e.g., number of candidate responses, beta in KL-divergence) on trajectory prediction accuracy and training stability?\n\nFor the zero-shot performance on Waymo, this paper attributes the excellent zero-shot generalization to the model's structured reasoning capabilities, but it does not compare with other methods that also claim zero-shot adaptation. Could you explain why your method has a more significant zero-shot advantage, and whether the dataset contains scene features that are common to both nuScenes and Waymo?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"nuScenesR^2-6K is the first dataset in autonomous driving that integrates self-reflection for validation, providing detailed reasoning chains to bridge input information and output trajectories, which enhances model interpretability.\n\nThe combination of SFT with structured CoT reasoning and RL with physics-grounded rewards effectively addresses both reasoning inadequacy and physical infeasibility of trajectories\n\nAutoDrive-R² outperforms state-of-the-art methods on nuScenes and Waymo datasets, with significant error reductions and robust zero-shot capabilities, validating its effectiveness and generality."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The RL stage with GRPO requires generating multiple candidate responses, which may introduce high computational overhead, but the manuscript does not discuss inference speed or real-time deployment feasibility.\n\nAlthough the paper emphasizes the role of self-reflection in the four-step reasoning chain, it does not clearly explain how the model corrects inconsistent trajectories during self-reflection. For example, there is no detailed description of the decision rules (e.g., threshold for determining trajectory inconsistency) or the specific adjustment strategies (e.g., how to modify velocity or steering angle) adopted in the self-reflection stage.\n\nThis paper does not analyze the computational overhead (e.g., inference time per trajectory, GPU memory usage) or compare it with lightweight baseline methods. This may restrict the practical deployment on edge devices with limited computing resources.\n\nTable 2 appears to be outside the page content"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916499962,"tcdate":1761442702038,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3009/Reviewer_TqXQ"],"signatures":["ICLR.cc/2026/Conference/Submission3009/Reviewer_TqXQ"],"forum":"KVWaCzJrrq","number":2,"license":"CC BY 4.0","cdate":1761442702038,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3009/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916499962,"domain":"ICLR.cc/2026/Conference","replyto":"KVWaCzJrrq","id":"2Z8edGwrcE","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Applications","Robots","Vision–Language–Action Models"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Vision–Language–Action (VLA) models in autonomous driving systems have recently demonstrated transformative potential by integrating multimodal perception with decision-making capabilities. However, the interpretability and coherence of the decision process and the plausibility of action sequences remain largely underexplored. To address these issues, we propose AutoDrive-R², a novel VLA framework that enhances both reasoning and self-reflection capabilities of autonomous driving systems through chain-of-thought (CoT) processing and reinforcement learning (RL). Specifically, we first propose an innovative CoT dataset named nuScenesR²-6K for supervised fine-tuning, which effectively builds cognitive bridges between input information and output trajectories through a four-step logical chain with self-reflection for validation. Moreover, to maximize both reasoning and self-reflection during the RL stage, we further employ the Group Relative Policy Optimization (GRPO) algorithm within a physics-grounded reward framework that incorporates spatial alignment, vehicle dynamic, and temporal smoothness criteria to ensure reliable and realistic trajectory planning. Extensive evaluation results across both nuScenes and Waymo datasets demonstrates the state-of-the-art performance and robust generalization capacity of our proposed method."},"_bibtex":{"value":"@inproceedings{\nyuan2026autodriver,\ntitle={AutoDrive-R{\\texttwosuperior}: Incentivizing Reasoning and Self-Reflection Capacity for {VLA} Model in Autonomous Driving},\nauthor={Zhenlong Yuan and Chengxuan Qian and Jing Tang and Rui Chen and Zijian Song and Lei Sun and Xiangxiang Chu and Yujun Cai and Dapeng Zhang and Shuo Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KVWaCzJrrq}\n}"},"title":{"value":"AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving"},"pdf":{"value":"/pdf/d23c679c86e3ecf9f93074146b4bc68130b86425.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|autodriver_incentivizing_reasoning_and_selfreflection_capacity_for_vla_model_in_autonomous_driving"},"authorids":{"value":["~Zhenlong_Yuan1","~Chengxuan_Qian1","~Jing_Tang8","~Rui_Chen40","~Zijian_Song4","~Lei_Sun14","~Xiangxiang_Chu1","~Yujun_Cai1","~Dapeng_Zhang6","~Shuo_Li3"]},"authors":{"value":["Zhenlong Yuan","Chengxuan Qian","Jing Tang","Rui Chen","Zijian Song","Lei Sun","Xiangxiang Chu","Yujun Cai","Dapeng Zhang","Shuo Li"]}},"version":2},{"content":{"summary":{"value":"The paper presents ClimateLLM, a frequency-domain foundation model for global weather forecasting built on a new architecture called SAED-Former (Scale-Aware Entangled Dynamics Transformer). The key insight is that traditional spectral or deep learning models treat complex-valued frequency representations as flat vectors, conflating amplitude (energy evolution) and phase (spatial propagation). ClimateLLM explicitly decouples these two dynamics into dual states and models their interactions through:\n\nA phase-centric propagation kernel that governs interactions solely based on phase (spatial propagation),\n\nA scale-aware evolution module (SAEM) that applies band-conditioned projections to encode wave-number–dependent physics,\n\nA dual-state representation for amplitude and phase evolution.\n\nThe model operates autoregressively in the frequency domain, predicting future weather states using FFT-based embeddings and reconstructing the spatial field via inverse FFT."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"How does ClimateLLM handle localized, non-periodic boundary conditions where FFT-based global modes may be inefficient or physically inconsistent?\n\nCan the phase-centric kernel be related to or derived from wave propagation equations (e.g., Helmholtz or Navier–Stokes dispersion relations)?\n\nHow does performance scale with resolution—does the efficiency advantage persist for 0.25° or higher-resolution grids?\n\nCould the authors compare to physics-informed hybrid models that incorporate conservation laws directly in spectral space?\n\nAre there stability concerns when phase wrapping is handled via the angular wrap() function for long temporal horizons?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Novel spectral formulation: The decoupling of amplitude and phase in a dual-state architecture is a clear conceptual and mathematical innovation, providing a more physically interpretable model of atmospheric dynamics.\n\nPhysics-aligned inductive biases: The phase-centric propagation kernel and scale-conditional evolution module are elegant, well-motivated, and contribute to both interpretability and efficiency.\n\nStrong empirical performance: The model demonstrates consistent or superior results to leading deep learning weather models (FourCastNet, ClimODE) on ERA5 benchmarks, with massive reductions in compute and memory.\n\nZero-/few-shot generalization: The cross-variable transfer results are particularly impressive, hinting at reusable representations across meteorological variables.\n\nThorough evaluation: Ablation, sensitivity, and real-case analyses (Ahvaz heat event) are well-presented, reinforcing claims about physical fidelity and robustness.\n\nClarity and reproducibility: The paper provides detailed algorithmic pseudocode and well-documented baselines, enhancing reproducibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited scale of validation: Experiments are limited to low-resolution (5.625°) ERA5 grids; generalization to higher resolutions or local/regional models is untested.\n\nTheoretical justification remains heuristic: While the phase–amplitude separation is intuitively justified, the physics–mathematics linkage (especially in phase-centric attention) lacks formal derivation or stability analysis.\n\nDependence on FFT assumptions: The method relies heavily on global Fourier bases, which might underperform for non-stationary, non-periodic regional phenomena (e.g., tropical convection, localized storms).\n\nComparison scope: While comparisons to neural operators and weather transformers are included, it lacks benchmarking against emerging 2025–2026 hybrid or foundation models (e.g., EarthGPT, DeepClimaNet).\n\nTerminological overreach: Calling ClimateLLM an “LLM” may be overstated—while GPT-style architecture is used, there is no natural language interface or token-level semantic modeling."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924982379,"tcdate":1761375943088,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14599/Reviewer_MJbq"],"signatures":["ICLR.cc/2026/Conference/Submission14599/Reviewer_MJbq"],"forum":"MGy6FHMqnd","number":1,"license":"CC BY 4.0","cdate":1761375943088,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14599/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924982379,"domain":"ICLR.cc/2026/Conference","replyto":"MGy6FHMqnd","id":"VC42AHCFmX","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Spatiotemporal Modeling","Mixture-of-Experts","Weather Forecasting"]},"supplementary_material":{"value":"/attachment/3dba4712f144cc231dd70fd576bfc1c961796947.pdf"},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Recent progress in deep learning has advanced global weather forecasting, with larger and higher-resolution models steadily improving skill. In parallel, spectral methods provide an efficient basis for global dynamics. Yet most spectral approaches treat the complex spectrum as generic features, conflating the distinct physics encoded in amplitude (energy evolution) and phase (spatial propagation). **We propose ClimateLLM, a physics-aligned, frequency-domain forecasting framework powered by SAED-Former.** At its core, **SAED-Former** explicitly separates these two processes via a *dual-state representation*, computes interactions through a *phase-centric propagation kernel*, and injects wave-number–aware priors using *scale-conditional projection*. This physics-aligned design yields compact, robust frequency-domain representations. On standard reanalysis benchmarks, ClimateLLM matches or exceeds state-of-the-art accuracy across short- and medium-range horizons while training on a single GPU within hours. Moreover, the model supports *cross-variable transference*: networks trained on data-rich variables produce robust zero-shot forecasts for data-scarce variables. By elevating spectral structure to first-class status, ClimateLLM improves forecast quality, efficiency, and generalization."},"_bibtex":{"value":"@misc{\nli2026climatellm,\ntitle={Climate{LLM}: Efficient Weather Forecasting via Frequency-Aware Large Language Models},\nauthor={Shixuan Li and Wei Yang and Peiyu Zhang and Xiongye Xiao and Defu Cao and Yuehan Qin and Xiaole Zhang and Yue Zhao and Paul Bogdan},\nyear={2026},\nurl={https://openreview.net/forum?id=MGy6FHMqnd}\n}"},"title":{"value":"ClimateLLM: Efficient Weather Forecasting via Frequency-Aware Large Language Models"},"pdf":{"value":"/pdf/f458c457f45f967c29de9318d5c73f02c6b02af6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|climatellm_efficient_weather_forecasting_via_frequencyaware_large_language_models"},"authorids":{"value":["~Shixuan_Li2","~Wei_Yang21","~Peiyu_Zhang1","~Xiongye_Xiao1","~Defu_Cao1","~Yuehan_Qin1","~Xiaole_Zhang1","~Yue_Zhao13","~Paul_Bogdan1"]},"authors":{"value":["Shixuan Li","Wei Yang","Peiyu Zhang","Xiongye Xiao","Defu Cao","Yuehan Qin","Xiaole Zhang","Yue Zhao","Paul Bogdan"]}},"version":2},{"content":{"summary":{"value":"This paper presents PIA-MBRL (Physics-Informed Augmentation with Model-Based RL), a framework for rapidly adapting reinforcement learning agents to new tasks by integrating lightweight analytical models into a model-based RL (MBRL) pipeline. The key insight is to use stable ODE-based vehicle models to generate physics-grounded synthetic rollouts, which are merged with limited high-fidelity simulator data (Assetto Corsa Gym) to improve sample efficiency and generalization. The approach builds on TD-MPC2 and the offline-to-online adaptation paradigm of FOWM. Tested in diverse conditions within the autonomous car racing scenario, the method achieves faster convergence and improved task performance. \n\nExperiments in autonomous racing show that physics-informed augmentation enables agents to adapt to unseen tracks and surface conditions (e.g., dusty, low-grip) within a few episodes. The authors release over 1000 hours of racing data and ODE rollouts, along with an enhanced version of ACGym that supports dynamic model swapping and JAX-based real-time control."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weaknesses above. I will be happy to see the paper further improved."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Justified design choice: Uses analytic ODE models as reliable, long-horizon data generators for MBRL. This idea is well-suited for the racing task.\n\n- Engineering contribution: Extends ACGym with Linux support, modular dynamics swapping, and real-time TD-MPC2 in JAX.\n\n- Extensive eval Thorough experiments across 15 tracks, 3 surface conditions, and multiple baselines (SAC, IQL, FOWM, TD-MPC2)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited novelty in algorithmic form: Builds directly upon existing frameworks (TD-MPC2, FOWM), with the main innovation being the physics-based data source.\n- Also, the methodology is only tested on a single task in a simulation environment. Based on the limited domain demonstrated, the paper's title should also incorporate the specific domain, e.g., add \"...for autonomous racing\"\n- Another disadvantage of using a strong model prior assumption is the unmodeled effects. ODE models use fixed parameters and do not adapt to varying surface frictions or uncertainties. Despite being more efficient in terms of learning efficiency and even final performance, I do see that this approach has rather limited expressiveness when encountering more complex models.\n- Generalization claims: While racing is a strong testbed, broader applicability (e.g., robotics or UAVs) remains unverified.\n\nApart from the methodology section,\n- Sim2real verification: An alternative way to significantly improve the paper is to demonstrate that the approach can perform some sim2real verification to demonstrate the approach's effectiveness (not even necessarily for car racing at all, it can be even simpler tasks). This will strongly support the model learned, which can handle real-world scenarios."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917801771,"tcdate":1761662509747,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4974/Reviewer_mYmD"],"signatures":["ICLR.cc/2026/Conference/Submission4974/Reviewer_mYmD"],"forum":"himcPrg6sS","number":2,"license":"CC BY 4.0","cdate":1761662509747,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4974/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917801771,"domain":"ICLR.cc/2026/Conference","replyto":"himcPrg6sS","id":"T3yvmFA0wt","forumContent":{"TLDR":{"value":"We show that physics-informed data augmentation with lightweight models improves sample efficiency, safety, and generalization in model-based RL for autonomous racing."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Reinforcement learning; Model-based reinforcement learning; Autonomous racing; Physics-informed learning; Data augmentation"]},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"A central challenge in reinforcement learning (RL) is achieving agents that generalize and adapt to new tasks and conditions. Many works address this via offline RL which is constrained by dataset coverage, or online RL which requires costly and potentially unsafe exploration. We propose a framework for rapid adaptation of RL agents by augmenting model-based RL with physics-informed data augmentation. Specifically, we use lightweight analytical models to generate stable, physics-grounded rollouts that complement real interaction data and allows the model-based RL agent to adapt in just a few trials. We validate our approach in autonomous racing, an extreme testbed with fast dynamics and strict safety constraints, using Assetto Corsa paired with lightweight vehicle models for data augmentation. Across diverse tracks and surfaces, our method achieves faster convergence, lower lap times, and fewer incidents than a set of strong baselines.\nAlthough demonstrated in racing, our framework is domain-agnostic, offering a practical path to data-efficient control wherever simple models exist as priors."},"_bibtex":{"value":"@misc{\nremonda2026leveraging,\ntitle={Leveraging Physics-Based Models for Rapid Adaptation in Reinforcement Learning},\nauthor={Adrian Remonda and Jiajun Xi and Nicklas Hansen and Marcus Greiff and John Talbot and John Subosits and Xiaolong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=himcPrg6sS}\n}"},"title":{"value":"Leveraging Physics-Based Models for Rapid Adaptation in Reinforcement Learning"},"pdf":{"value":"/pdf/1febfa5147226206dba256ea2825066ccfecd370.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"remonda|leveraging_physicsbased_models_for_rapid_adaptation_in_reinforcement_learning"},"authorids":{"value":["~Adrian_Remonda1","~Jiajun_Xi1","~Nicklas_Hansen1","~Marcus_Greiff1","~John_Talbot1","~John_Subosits1","~Xiaolong_Wang3"]},"authors":{"value":["Adrian Remonda","Jiajun Xi","Nicklas Hansen","Marcus Greiff","John Talbot","John Subosits","Xiaolong Wang"]}},"version":2},{"content":{"summary":{"value":"This paper studies NLOS vision by classifying hand gestures from indirect wall shadows. It proposes RacoNet, a physics-guided model combining light-transport constraints (RCLT) and a geometry-recovery module (GIAO), fused via KA-ELNR. Across three simulated sign-language datasets (plus one small measured setup), RacoNet beats general-purpose vision backbones, suggesting that physics-informed modeling helps on this synthetic shadow-classification task."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"Please refer to weakness for more details."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"This paper propose a physics-aware deep architecture that explicitly models light transport and incorporates geometric reasoning, which is conceptually sound. This reflects a thoughtful attempt to bridge physical modeling with neural network design."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper frames the work as NLOS decoding, but the core evaluation is closed-set classification of static hand gestures. NLOS problems typically prioritize reconstructing occluded geometry/appearance and then reasoning on top. Without any reconstruction, pose recovery, or even an interpretable intermediate, it’s hard to claim broader NLOS competence. A simple multi-task variant or silhouette/pose proxies would strengthen the case.\n\n- Most baselines are general-purpose vision backbones that aren’t designed for NLOS. The paper should compare against passive NLOS pipelines that first reconstruct an image/silhouette/volume from wall observations and then classify. A two-step reconstruct-then-classify baseline might be competitive and would reveal whether RacoNet’s end-to-end approach is truly advantageous for NLOS.\n\n- The main results rely on shadows simulated from sign-language datasets; the single measured set appears to be captured in a controlled testbed with fixed geometry. There is no evidence of performance in realistic conditions. Claims about cross-room communication remain speculative without demonstrations under varied, messy setups.\n\n- Apart from limited lighting variation within the simulator, the paper does not stress test common NLOS failure modes: ambient flicker, dynamic illuminants, motion blur, defocus, rolling-shutter distortions, low-bit-depth quantization, mixed wall materials, or geometry shift. A principled robustness suite and cross-setup generalization study are needed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921709511,"tcdate":1761809173437,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10388/Reviewer_tf9f"],"signatures":["ICLR.cc/2026/Conference/Submission10388/Reviewer_tf9f"],"forum":"AlqrnU93o7","number":2,"license":"CC BY 4.0","cdate":1761809173437,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10388/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921709511,"domain":"ICLR.cc/2026/Conference","replyto":"AlqrnU93o7","id":"XFh6k4Qki3","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Gesture recognition","Computer Vision","Deep Learning","Pattern recognition"]},"supplementary_material":{"value":"/attachment/271ea332c6e17de92d5de6b15f01e4e1127e02f4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Accurately decoding hidden information in dynamic shadows for Non-Line-of-Sight (NLOS) imaging enables us to overcome visual occlusions and perceive or reconstruct obscured targets. This breakthrough holds significant potential for real-world applications such as disaster rescue, autonomous driving, and security surveillance. Conventional algorithms struggle to model the physical propagation of light in space. Furthermore, the signal distortions introduced by nonlinear transformations incur the loss of geometric information about the source scene, limiting sensitivity to subtle shadow variations. To overcome these challenges, we present Radiation-constraint Network (RacoNet) that marries physical propagation simulation with geometric-information recovery to interpret minute gesture signals embedded in dynamic shadows. In RacoNet, Radiance-Constrained Light-Transportation (RCLT) optical propagation is proposed to capture complete light-space information. Meanwhile, Geometric Information Aliment Operation (GIAO) restores source-scene geometry lost in the modulated shadow through layer-by-layer refined prior attention. Moreover, Kolmogorov-Arnold Enhanced Layerwise Nonlinear Reorganization (KA-ELNR) fuses light-space and geometric cues to produce the final decoded output. Extensive experiments show that RacoNet markedly surpasses existing approaches in both accuracy and robustness for dynamic-shadow decoding, confirming the possibility of gesture-based information interaction via shadows."},"_bibtex":{"value":"@misc{\nzheng2026shadowspeak,\ntitle={ShadowSpeak: Is It Possible to Communicate Cross-Room Solely by Decoding Gesture Shadows?},\nauthor={Zhiwen Zheng and Yubo Chen and Shaowei Jiang and Huiyu Zhou and Zhao Huang and Tao Zhang and Jin Liu and Guangyuan Zhang and Xiaoshuai Zhang and Xingru Huang},\nyear={2026},\nurl={https://openreview.net/forum?id=AlqrnU93o7}\n}"},"title":{"value":"ShadowSpeak: Is It Possible to Communicate Cross-Room Solely by Decoding Gesture Shadows?"},"pdf":{"value":"/pdf/1b1a3d75d4826dc6ce191c5e822701a9f6d2d043.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zheng|shadowspeak_is_it_possible_to_communicate_crossroom_solely_by_decoding_gesture_shadows"},"authorids":{"value":["~Zhiwen_Zheng1","~Yubo_Chen4","~Shaowei_Jiang1","~Huiyu_Zhou3","~Zhao_Huang2","~Tao_Zhang5","~Jin_Liu22","~Guangyuan_Zhang1","~Xiaoshuai_Zhang2","~Xingru_Huang1"]},"authors":{"value":["Zhiwen Zheng","Yubo Chen","Shaowei Jiang","Huiyu Zhou","Zhao Huang","Tao Zhang","Jin Liu","Guangyuan Zhang","Xiaoshuai Zhang","Xingru Huang"]}},"version":2},{"content":{"comment":{"value":"#### Nitpicks and Misc\n\n**C.1 - Round 29.97 FPS to 30 FPS**\n\nBefore starting this project we were not aware of the subtleties of FPS in recording videos but there is, unfortunately, a meaningful practical difference between 29.97 FPS and 30 FPS that can cause considerable pain when writing code to preprocess videos. To ensure this doesn't happen to others we have leaned to being overly explicit here.\n\n**C.2 - Confusing hypothetical**\n\nWe will make the phrasing more explicit.\n\n**C.3 - Cut or move Section 3.4**\n\nWe will substantially reduce the size of this section and move what remains to the appendix.\n\n**C.4 - Dedicated intuitive-physics models baseline**\n\nWe looked at a number of these as potential baselines but generally have found that they either require ground-truth or non-RGB sensor data, do not have publicly available (or runnable) code, or are not trained using natural videos. For instance, DensePhysNet takes depth maps as input rather than RGB images while PHYRE is a 2D-only benchmark and we suspect that retraining any model designed for that dataset on realistic images would be its own contribution.\n\n**C.5 - Appreciate inclusion of statistical significance**\n\nThank you!\n\n**C.6 & C.7 - Move Table 1 and reduce limitations section**\n\nWe will do this.\n"},"title":{"value":"Response to Reviewer EhCD (2/2)"}},"tmdate":1661988427058,"tcdate":1661988427058,"writers":["TMLR","TMLR/Paper301/Authors"],"signatures":["TMLR/Paper301/Authors"],"forum":"9NjqD9i48M","number":6,"cdate":1661988427058,"mdate":1661988427058,"readers":["everyone"],"invitations":["TMLR/Paper301/-/Official_Comment"],"domain":"TMLR","replyto":"8JKYTq8k6y","id":"hHMBuexZzp","forumContent":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://allenai.org/project/inflevel/home"},"abstract":{"value":"To what extent do modern AI systems comprehend the physical world? We introduce the open-access Infant-Level Physical Reasoning Benchmark (InfLevel) to gain insight into this question. We evaluate ten neural-network architectures developed for video understanding on tasks designed to test these models' ability to reason about three essential physical principles which researchers have shown to guide human infants' physical understanding. We explore the sensitivity of each AI system to the continuity of objects as they travel through space and time, to the solidity of objects, and to gravity. We find strikingly consistent results across 60 experiments with multiple systems, training regimes, and evaluation metrics: current popular visual-understanding systems are at or near chance on all three principles of physical reasoning. We close by suggesting some potential ways forward."},"_bibtex":{"value":"@article{\nweihs2022benchmarking,\ntitle={Benchmarking Progress to Infant-Level Physical Reasoning in {AI}},\nauthor={Luca Weihs and Amanda Yuile and Ren{\\'e}e Baillargeon and Cynthia Fisher and Gary Marcus and Roozbeh Mottaghi and Aniruddha Kembhavi},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2022},\nurl={https://openreview.net/forum?id=9NjqD9i48M},\nnote={}\n}"},"title":{"value":"Benchmarking Progress to Infant-Level Physical Reasoning in AI"},"certifications":{"value":[]},"changes_since_last_submission":{"value":"### Additional context regarding metrics\n\nWe have clarified our position regarding our metrics in Section 3.4 where we now say:\n\n> For this reason, InfLevel is an evaluation-only benchmark: no training on InfLevel is allowed. This has an unfortunate implication we must overcome: models that wish to report scores on InfLevel must be able to provide a scalar “surprise” score for every input video. This requirement, used by other benchmarks (Riochet et al., 2018), is limiting as it requires that anyone wishing to evaluate on InfLevel to train a special “surprise” decoder on their model using some external data source. To circumvent this problem, we model surprise as out-of-domain (OOD) detection using the intuition that models with sufficient physical understanding should consider physically implausible events more out-of-domain than physically plausible ones. In Sec. 4, we propose several OOD metrics, each of which takes a representation of a video and returns a scalar quantifying how out-of-domain the video is. While we show, in Sec. 4, that this OOD approach is empirically promising, it is easy to show that there is no surprise metric which can be used to detect physical understanding for all possible models (see App. E.2). Given this, anyone evaluating on InfLevel is free to define their own surprise metrics so long as: (1) the same metric is used across all subsets of InfLevel (i.e. there should not be one metric for Continuity and another for Gravity) and (2) these surprise metrics are neither trained on InfLevel nor designed explicitly to exploit regularities in InfLevel data (which would be an implicit form of training). Going forward, we hope that researchers will continue to improve and refine the model-agnostic OOD surprise measures we propose.\n\n### Typos and minor changes\n\nWe have fixed some typos and added one additional concurrent work to our related work section."},"license":{"value":"Creative Commons Attribution 4.0 International (CC BY 4.0)"},"pdf":{"value":"/pdf/a388498ddf8e5e9f1b744459576c8da685c4498e.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"weihs|benchmarking_progress_to_infantlevel_physical_reasoning_in_ai","readers":["everyone"]},"authorids":{"readers":["everyone"],"value":["~Luca_Weihs1","~Amanda_Yuile1","~Renée_Baillargeon1","~Cynthia_Fisher1","~Gary_Marcus1","~Roozbeh_Mottaghi1","~Aniruddha_Kembhavi1"]},"assigned_action_editor":{"value":"~Josh_Merel1"},"authors":{"readers":["everyone"],"value":["Luca Weihs","Amanda Yuile","Renée Baillargeon","Cynthia Fisher","Gary Marcus","Roozbeh Mottaghi","Aniruddha Kembhavi"]}},"version":2},{"content":{"venue":{"value":"UAI 2023"},"pdf":{"value":"https://proceedings.mlr.press/v216/kreacic23a/kreacic23a.pdf"},"venueid":{"value":"dblp.org/conf/UAI/2023"},"paperhash":{"value":"kreacic|differentially_private_synthetic_data_using_kdtrees"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Eleonora_Kreacic:","https://dblp.org/search/pid/api?q=author:Navid_Nouri:","~Vamsi_K._Potluru1","https://dblp.org/search/pid/api?q=author:Tucker_Balch:","https://dblp.org/search/pid/api?q=author:Manuela_Veloso:"]},"html":{"value":"https://proceedings.mlr.press/v216/kreacic23a.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/uai/KreacicNPBV23,\n  author={Eleonora Kreacic and Navid Nouri and Vamsi K. Potluru and Tucker Balch and Manuela Veloso},\n  title={Differentially private synthetic data using KD-trees},\n  year={2023},\n  cdate={1672531200000},\n  pages={1143-1153},\n  url={https://proceedings.mlr.press/v216/kreacic23a.html},\n  booktitle={UAI},\n  crossref={conf/uai/2023}\n}\n"},"abstract":{"value":"Creation of a synthetic dataset that faithfully represents the data distribution and simultaneously preserves privacy is a major research challenge. Many space partitioning based approaches have emerged in recent years for answering statistical queries in a differentially private manner. However, for synthetic data generation problem, recent research has been mainly focused on deep generative models. In contrast, we exploit space partitioning techniques together with noise perturbation and thus achieve intuitive and transparent algorithms. We propose both data independent and data dependent algorithms for $\\epsilon$-differentially private synthetic data generation whose kernel density resembles that of the real dataset. Additionally, we provide theoretical results on the utility-privacy trade-offs and show how our data dependent approach overcomes the curse of dimensionality and leads to a scalable algorithm. We show empirical utility improvements over the prior work, and discuss performance of our algorithm on a downstream classification task on a real dataset."},"title":{"value":"Differentially private synthetic data using KD-trees"},"authors":{"value":["Eleonora Kreacic","Navid Nouri","Vamsi K. Potluru","Tucker Balch","Manuela Veloso"]}},"tmdate":1736515127460,"pdate":1672531200000,"tcdate":1736515116274,"writers":["~"],"signatures":["~Vamsi_K._Potluru1"],"forum":"BxZ7FW2i7d","license":"CC BY-SA 4.0","number":262530,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1736515127460,"domain":"DBLP.org","id":"BxZ7FW2i7d","version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2306.13211v1"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"kreacic|differentially_private_synthetic_data_using_kdtrees"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Eleonora_Kreacic:","https://dblp.org/search/pid/api?q=author:Navid_Nouri:","~Vamsi_K._Potluru1","https://dblp.org/search/pid/api?q=author:Tucker_Balch:","https://dblp.org/search/pid/api?q=author:Manuela_Veloso:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2306.13211"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2306-13211,\n  publtype={informal},\n  author={Eleonora Kreacic and Navid Nouri and Vamsi K. Potluru and Tucker Balch and Manuela Veloso},\n  title={Differentially Private Synthetic Data Using KD-Trees},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2306.13211},\n  url={https://doi.org/10.48550/arXiv.2306.13211}\n}\n"},"abstract":{"value":"Creation of a synthetic dataset that faithfully represents the data distribution and simultaneously preserves privacy is a major research challenge. Many space partitioning based approaches have emerged in recent years for answering statistical queries in a differentially private manner. However, for synthetic data generation problem, recent research has been mainly focused on deep generative models. In contrast, we exploit space partitioning techniques together with noise perturbation and thus achieve intuitive and transparent algorithms. We propose both data independent and data dependent algorithms for $\\epsilon$-differentially private synthetic data generation whose kernel density resembles that of the real dataset. Additionally, we provide theoretical results on the utility-privacy trade-offs and show how our data dependent approach overcomes the curse of dimensionality and leads to a scalable algorithm. We show empirical utility improvements over the prior work, and discuss performance of our algorithm on a downstream classification task on a real dataset."},"title":{"value":"Differentially Private Synthetic Data Using KD-Trees"},"authors":{"value":["Eleonora Kreacic","Navid Nouri","Vamsi K. Potluru","Tucker Balch","Manuela Veloso"]}},"tmdate":1736515122374,"pdate":1672531200000,"tcdate":1736515116073,"writers":["~"],"signatures":["~Vamsi_K._Potluru1"],"forum":"EZy3qb7o1k","license":"CC BY-SA 4.0","number":262522,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1736515122374,"domain":"DBLP.org","id":"EZy3qb7o1k","version":2},{"content":{"summary":{"value":"This proposes a method to train large language models to perform continuous black-box optimization by fine-tuning them on millions of synthetic optimization trajectories generated from various Bayesian optimization algorithms."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"If computation allowed, I'd be curious to see how larger models with reasoning ability can improve the optimization performance."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"The paper’s main contribution is the formulation of black-box optimization as a reasoning and sequence prediction problem for large language models, enabling optimization to be approached through text-based decision-making rather than analytical computation. It introduces GPTOpt, a fine-tuned LLM trained on millions of synthetic optimization trajectories generated by diverse Bayesian optimization methods, allowing it to learn generalizable optimization strategies."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The approach of teaching an LLM to perform numerical optimization is conceptually questionable, as language models are not designed for precise arithmetic or quantitative reasoning.\n\n2. The experiments are limited to low-dimensional problems (up to 10D), raising concerns about the method’s scalability and effectiveness in higher-dimensional or more complex optimization tasks.\n\n3. Consequently, the general applicability of GPTOpt to real-world, high-dimensional optimization scenarios remains uncertain."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933985820,"tcdate":1761990194161,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20574/Reviewer_ZSq7"],"signatures":["ICLR.cc/2026/Conference/Submission20574/Reviewer_ZSq7"],"forum":"aFJc2POtEQ","number":4,"license":"CC BY 4.0","cdate":1761990194161,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20574/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933985820,"domain":"ICLR.cc/2026/Conference","replyto":"aFJc2POtEQ","id":"8WehnNzNxz","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"We teach LLMs to perform black-box optimization through fine-tuning on synthetic datasets."},"keywords":{"value":["Black-box optimization","Bayesian optimization","Large language models"]},"primary_area":{"value":"optimization"},"abstract":{"value":"Global optimization of expensive, derivative-free black-box functions demands extreme sample efficiency. Classical methods such as Bayesian Optimization (BO) can be effective, but they often require careful parameter tuning to each application domain. At the same time, Large Language Models (LLMs) have shown broad capabilities, yet state-of-the-art models remain limited in solving continuous black-box optimization tasks. We introduce GPTOpt, an LLM-based optimization method that equips LLMs with continuous black-box optimization capabilities. By fine-tuning large language models on extensive synthetic datasets derived from diverse BO parameterizations, GPTOpt leverages LLM pre-training to generalize across optimization tasks. On a variety of black-box optimization benchmarks, GPTOpt surpasses traditional optimizers, highlighting the capacity of LLMs for advanced numerical reasoning and introducing a flexible framework for global optimization without parameter tuning."},"_bibtex":{"value":"@misc{\nmeindl2026gptopt,\ntitle={{GPTO}pt: Towards Efficient {LLM}-based Black-Box Optimization},\nauthor={Jamison Meindl and Yunsheng Tian and Tony Cui and Veronika Thost and Zhang-Wei Hong and Jie Chen and Wojciech Matusik and Mina Konakovic Lukovic},\nyear={2026},\nurl={https://openreview.net/forum?id=aFJc2POtEQ}\n}"},"title":{"value":"GPTOpt: Towards Efficient LLM-based Black-Box Optimization"},"pdf":{"value":"/pdf/6a6f369345eececd9626f2dd3fcf1515e4eac323.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"meindl|gptopt_towards_efficient_llmbased_blackbox_optimization"},"authorids":{"value":["~Jamison_Meindl1","~Yunsheng_Tian1","~Tony_Cui1","~Veronika_Thost1","~Zhang-Wei_Hong1","~Jie_Chen1","~Wojciech_Matusik2","~Mina_Konakovic_Lukovic1"]},"authors":{"value":["Jamison Meindl","Yunsheng Tian","Tony Cui","Veronika Thost","Zhang-Wei Hong","Jie Chen","Wojciech Matusik","Mina Konakovic Lukovic"]}},"version":2},{"content":{"summary":{"value":"Good paper on a useful physics dataset. Multiphase flows represent a frontier domain in flow physics. Time-series datasets are also useful across different ML domain including video modeling, etc.\n\nI only have questions that would help clarify some context and experimental choices for better presentation.\n\nEdit 1: Concerns have been addressed. Raising score to 8 to recommend for strong acceptance."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. In appendix A, can you briefly describe the Allen Cahn equations a bit more for the readers? Specifically on big picture descriptions on how close is this to direct numerical simulation of Navier Stokes?\n2. In appendix A, what's the benefit of Lattice Boltzman methods vs conventional interface-capturing Finite Volume Solvers? Are there any cost-accuracy tradeoffs with your simulation approach? This could be useful for readers to know as well.\n3. Since Section 4, line 365. How were hyper parameters chosen?\n4. Section 4 and 5 -- How many train/val/test splits?\n5. Can you spend a paragraph or 2 explaining the broader applications of this dataset and importance of sequence to field and sequence predictions benchmarks in the intro?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Time-dependent multiphase data -- rich dataset.\n2. Extensive model evaluation\n3. Good lit review of previous work\n4. Lattice boltzmann solvers are high fidelity\n5. 4000 GPU hours is substantial\n6. Good Qualititative demonstration of ML predictions\n7. Solid Appendix\n8. Good reproducibility efforts."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Applications of this dataset are not obvious -- could be emphasized more in introduction or via eval demonstrations\n2. Description of physics methods requires a bit more clarity for non-physics readers in this general ML venue.\n3. Connection to anonymous repo had 522 timeout when I clicked -- I assume that this will be fixed after double blind review."}},"nonreaders":[],"tmdate":1733186966506,"tcdate":1730498022934,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12095/Reviewer_beKU"],"signatures":["ICLR.cc/2025/Conference/Submission12095/Reviewer_beKU"],"forum":"QPVK1ne9gI","number":2,"license":"CC BY 4.0","cdate":1730498022934,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12095/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733186966506,"domain":"ICLR.cc/2025/Conference","replyto":"QPVK1ne9gI","id":"D7X3a2juSZ","forumContent":{"TLDR":{"value":"SciML benchmark with 11000 two-phase flow LBM simulations"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Scientific Machine Learning (SciML)","Multiphase Flow","Complex Physics Simulation","Lattice Boltzmann Method (LBM)","Droplet Dynamics","Bubble Dynamics"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Multiphase fluid dynamics, such as falling droplets and rising bubbles, are critical to many industrial applications. However, simulating these phenomena efficiently is challenging due to the complexity of instabilities, wave patterns, and bubble breakup. This paper investigates the potential of scientific machine learning (SciML) to model these dynamics using neural operators and foundation models. We apply sequence-to-sequence techniques on a comprehensive dataset generated from 11,000 simulations, comprising 1 million time snapshots, produced with a well-validated Lattice Boltzmann method (LBM) framework. The results demonstrate the ability of machine learning models to capture transient dynamics and intricate fluid interactions, paving the way for more accurate and computationally efficient SciML-based solvers for multiphase applications."},"_bibtex":{"value":"@misc{\nshadkhah2025mpfbench,\ntitle={{MPFB}ench: A Large Scale Dataset for Sci{ML} of Multi-Phase-Flows: Droplet and Bubble Dynamics},\nauthor={Mehdi Shadkhah and Ronak Tali and Ali Rabeh and Ethan Herron and Cheng-Hau Yang and Abhisek Upadhyaya and Adarsh Krishnamurthy and Chinmay Hegde and Aditya Balu and Baskar Ganapathysubramanian},\nyear={2025},\nurl={https://openreview.net/forum?id=QPVK1ne9gI}\n}"},"title":{"value":"MPFBench: A Large Scale Dataset for SciML of Multi-Phase-Flows: Droplet and Bubble Dynamics"},"pdf":{"value":"/pdf/a35588c0c129db17a4e8b7c466717b766d2626c1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"shadkhah|mpfbench_a_large_scale_dataset_for_sciml_of_multiphaseflows_droplet_and_bubble_dynamics"},"authorids":{"value":["~Mehdi_Shadkhah1","~Ronak_Tali1","~Ali_Rabeh1","~Ethan_Herron1","~Cheng-Hau_Yang1","~Abhisek_Upadhyaya1","~Adarsh_Krishnamurthy1","~Chinmay_Hegde1","~Aditya_Balu1","~Baskar_Ganapathysubramanian1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mehdi Shadkhah","Ronak Tali","Ali Rabeh","Ethan Herron","Cheng-Hau Yang","Abhisek Upadhyaya","Adarsh Krishnamurthy","Chinmay Hegde","Aditya Balu","Baskar Ganapathysubramanian"]}},"version":2},{"content":{"comment":{"value":"### **(Continue)**\n\n---\n\n### **Physics failure** \n\n**Definition**. The reconstructed object has a movement over 0.05 of the scene unit length or rotates over 10 degrees in the simulator after applying gravity in the composed scene (the same threshold as we evaluate stability in paper). \n\n**Reason and Phenomenon**. Physics failures are mainly caused by:\n\n* *Invalid shapes* — Missing or incomplete surfaces make objects incompatible with the physics simulator.\n* *Overlooked contact pairs* — Some methods (e.g., PhyRecon) focus only on object-ground contacts (as stated in Supp D.1: `focused solely on object-ground support for simplicity and training efficiency`), ignoring object-object interactions. This weakens physical supervision (PhyRecon Sec. 3.2).\n* *Penetration* — Inter-object geometry penetration causes repelling forces during simulation, leading to instability.\n* *Drifting* — Inaccurate shape modeling at contact points (even with minimal penetration) can cause gradual drifting from the original position.\n\n\n**Evaluation**. We evaluate physics failures across **all datasets** used in the paper: Replica, ScanNet++, and iGibson. The table below summarizes the `physics failure rate` for each dataset.\n\n|  Method   | Replica % &darr; | Scannet++ % &darr; | iGibson % &darr; |\n|:---------:|:----------------:|:------------------:|:----------------:|\n| ObjSDF++  |       60.6       |        71.8        |       53.6       |\n| PhyRecon  |       94.4       |        90.6        |       66.0       |\n| DP-Recon  |       91.5       |        90.6        |       72.2       |\n| HoloScene |     **18.3**     |      **29.4**      |     **24.7**     |\n\nTo better understand the sources of failure, we also break down the failure counts by underlying cause:\n\n|  Method   | Total object # | Failure # &darr; | Failure rate % &darr; | Invalid shapes # | Overlooked contact pairs # | Penetration # | Drifting # |\n|:---------:|:--------------:|:----------------:|:---------------------:|:-------:|:-------------------:|:-----------:|:--------:|\n| ObjSDF++  |      253       |       166        |         65.6          |   38    |          0          |     76      |    52    |\n| PhyRecon  |      253       |       236        |         93.3          |   58    |         121         |     33      |    24    |\n| DP-Recon  |      253       |       235        |         92.9          |   64    |          0          |     102     |    69    |\n| HoloScene |      253       |      **66**      |       **26.1**        |    0    |          0          |      6      |    60    |\n\nThe results reveal several key insights into physics failure across methods:\n- **Failure Tied to Physics Modeling**:  \n  ObjSDF++ and DP-Recon often fail due to *invalid shapes* and *penetrations*, stemming from limited or no physics modeling. PhyRecon includes physics but suffers from high failure rates by considering only object-ground contact, neglecting object-object interactions (PhyRecon Supp D.1).\n- **Shape Priors vs. Physical Plausibility**:  \n  DP-Recon emphasizes visual completeness with shape priors but lacks physical stability, leading to drifting and penetration. PhyRecon ensures ground contact but doesn't generalize to more complex interactions.\n- **HoloScene’s Integrated Solution**:  \n  HoloScene combines *generative priors*, *scene graph reasoning*, and a *simulator-as-critic* to model object-object contact and enforce physical plausibility. This reduces penetration failures and eliminates invalid or oversimplified geometry.\n\nThese findings suggest that physical consistency should not be treated as an afterthought—whether through *post-hoc correction* or *surrogate loss terms*—but instead integrated directly into the generative reconstruction process via *hard constraints* and *physical validation*. Joint reasoning over structure, geometry, and physics emerges as a promising direction for building more robust, simulation-ready digital assets. We thank the reviewer again for the nice suggestion."},"title":{"value":"(2/2) Response to Reviewer BhAa's Follow-Up Comments"}},"parentInvitations":"NeurIPS.cc/2025/Conference/-/Official_Comment","tmdate":1761724959609,"tcdate":1754675662835,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission632/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission632/Authors"],"forum":"BOwPpmRgmW","number":10,"license":"CC BY 4.0","cdate":1754675662835,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/Submission632/-/Official_Comment","NeurIPS.cc/2025/Conference/-/Edit"],"mdate":1761724959609,"domain":"NeurIPS.cc/2025/Conference","replyto":"LKNW7mpsXD","id":"kPqUiu9TSw","forumContent":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["computer vision","digital twin","simulation"]},"supplementary_material":{"value":"/attachment/e62cc4f86a0871ad55cf500e711a175265549196.zip"},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Digitizing the physical world into accurate simulation‑ready virtual environments offers significant opportunities in a variety of fields such as augmented and virtual reality, gaming, and robotics. However, current 3D reconstruction and scene-understanding methods commonly fall short in one or more critical aspects, such as geometry completeness, object interactivity, physical plausibility, photorealistic rendering, or realistic physical properties for reliable dynamic simulation. To address these limitations, we introduce HoloScene, a novel interactive 3D reconstruction framework that simultaneously achieves these requirements. HoloScene leverages a comprehensive interactive scene-graph representation, encoding object geometry, appearance, and physical properties alongside hierarchical and inter-object relationships. Reconstruction is formulated as an energy-based optimization problem, integrating observational data, physical constraints, and generative priors into a unified, coherent objective. Optimization is efficiently performed via a hybrid approach combining sampling-based exploration with gradient-based refinement. The resulting digital twins exhibit complete and precise geometry, physical stability, and realistic rendering from novel viewpoints. Evaluations conducted on multiple benchmark datasets demonstrate superior performance, while practical use-cases in interactive gaming and real-time digital-twin manipulation illustrate HoloScene's broad applicability and effectiveness."},"_bibtex":{"value":"@inproceedings{\nxia2025holoscene,\ntitle={HoloScene: Simulation\\nobreakdash-Ready Interactive 3D Worlds from a Single Video},\nauthor={Hongchi Xia and Chih-Hao Lin and Hao-Yu Hsu and Quentin Leboutet and Katelyn Gao and Michael Paulitsch and Benjamin Ummenhofer and Shenlong Wang},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=BOwPpmRgmW}\n}"},"title":{"value":"HoloScene: Simulation‑Ready Interactive 3D Worlds from a Single Video"},"pdf":{"value":"/pdf/7d96141820ef7a6f8f2329e784d111789a8eceac.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"xia|holoscene_simulationready_interactive_3d_worlds_from_a_single_video"},"authorids":{"value":["~Hongchi_Xia1","~Chih-Hao_Lin1","~Hao-Yu_Hsu1","~Quentin_Leboutet1","~Katelyn_Gao1","~Michael_Paulitsch1","~Benjamin_Ummenhofer1","~Shenlong_Wang1"]},"authors":{"value":["Hongchi Xia","Chih-Hao Lin","Hao-Yu Hsu","Quentin Leboutet","Katelyn Gao","Michael Paulitsch","Benjamin Ummenhofer","Shenlong Wang"]}},"version":2},{"content":{"summary":{"value":"The paper works on the zero-shot video generation task and proposes,  MotionCraft. It uses physics simulations to generate optical flow that follows physical dynamics. Then, optical flow is applied to warp the noise in the latent space with the stable diffusion model. This approach ensures coherent motion application and consistent scene evolution, avoiding artefacts and missing content typical in pixel space flow applications. Compared to the state-of-the-art Text2Video-Zero, MotionCraft shows both qualitative and quantitative improvements in generating videos with complex motion dynamics."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"How do you specify the region for simulating physical dynamics? \n\nWill the type of dynamic physics simulator affect the quality of the generated video?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is well-motivated and well-written. \n2. The idea of using physics simulator to generate the optical flow which is then applied in latent space is very interesting. \n3. The qualitative results are impressive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. More quantitative comparison with baselines. Table 1 only reports the comparison with T2V0 on the generated videos. However, it is not clear which benchmark it is. Is it possible to compare with other baselines on more benchmarks, like MUG, MHAD?\n\n2. The method seems limited by specific types of dynamics, like fluid dynamics. It is not clear how to generate more general dynamics in real world. This may limit the potential application of the proposed method. \n\n3. The method is assumes that \"Optical Flow is preserved in the Latent Space of Stable Diffusion\" based on the observation of average correlations 0.727 between optical flows estimated in the RGB and noise latent spaces. Does this assumption hold true for generating realistic, pixel-wise precise motion in video with only 0.72 cosine simlarity?"},"limitations":{"value":"The method is limited to specific types of dynamic simulators and may be hard to apply to generic real-world video generation."}},"nonreaders":[],"tmdate":1730879884969,"tcdate":1720955386363,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission16937/Reviewer_ypn7"],"signatures":["NeurIPS.cc/2024/Conference/Submission16937/Reviewer_ypn7"],"forum":"lvcWA24dxB","number":3,"license":"CC BY 4.0","cdate":1720955386363,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission16937/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879884969,"domain":"NeurIPS.cc/2024/Conference","replyto":"lvcWA24dxB","id":"eQIG8RDoLc","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Zero-shot video generation","diffusion model","physics-based video generation"]},"supplementary_material":{"value":"/attachment/1c486da7f9ad3348d6c4d46255d7d6ac422054c7.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Generating videos with realistic and physically plausible motion is one of the main recent challenges in computer vision. \nWhile diffusion models are achieving compelling results in image generation, video diffusion models are limited by heavy training and huge models, resulting in videos that are still biased to the training dataset. In this work we propose MotionCraft, a new zero-shot video generator to craft physics-based and realistic videos. MotionCraft is able to warp the noise latent space of an image diffusion model, such as Stable Diffusion, by applying an optical flow derived from a physics simulation. We show that warping the noise latent space results in coherent application of the desired motion while allowing the model to generate missing elements consistent with the scene evolution, which would otherwise result in artefacts or missing content if the flow was applied in the pixel space.\nWe compare our method with the state-of-the-art Text2Video-Zero reporting qualitative and quantitative improvements, demonstrating the effectiveness of our approach to generate videos with finely-prescribed complex motion dynamics."},"_bibtex":{"value":"@inproceedings{\nmontanaro2024motioncraft,\ntitle={MotionCraft: Physics-Based Zero-Shot Video Generation},\nauthor={Antonio Montanaro and Luca Savant Aira and Emanuele Aiello and Diego Valsesia and Enrico Magli},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=lvcWA24dxB}\n}"},"title":{"value":"MotionCraft: Physics-Based Zero-Shot Video Generation"},"pdf":{"value":"/pdf/204eafaa11af77e60f19e38ab6fac1c10fb66f5d.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"montanaro|motioncraft_physicsbased_zeroshot_video_generation"},"authorids":{"value":["~Antonio_Montanaro1","~Luca_Savant_Aira1","~Emanuele_Aiello1","~Diego_Valsesia1","~Enrico_Magli1"]},"authors":{"value":["Antonio Montanaro","Luca Savant Aira","Emanuele Aiello","Diego Valsesia","Enrico Magli"]}},"version":2},{"content":{"summary":{"value":"This paper introduces SIPDO (Self-Improving Prompts through Data-Augmented Optimization), a closed-loop framework for prompt optimization that links a synthetic data generator with a prompt optimizer. The system iteratively generates challenging synthetic examples tailored to a prompt’s current weaknesses; prompt updates are then performed in response to observed failures. The framework integrates dynamic difficulty adjustment and uses synthetic data as a feedback signal, moving beyond static prompt optimization. Across multiple reasoning benchmarks (e.g., MMLU, BIG-Bench, ProofWriter), SIPDO is shown to outperform existing prompt optimization approaches, demonstrating strong generalization and robustness improvements."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"see weeknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"SIPDO introduces a closed-loop optimization framework that couples a synthetic data generator with a prompt optimizer through a dual-agent collaboration mechanism. The generator dynamically produces challenging samples targeting the current prompt’s weaknesses, while a progressive difficulty parameter ccc enables a curriculum learning strategy from simple to complex tasks. Ablation results demonstrate that this difficulty-gradient method improves average performance by 17.3%–24.3% compared to one-shot extreme sampling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although Table 2 provides a comprehensive overview of prompt optimization baselines, the current comparisons mainly cover works from 2022–2024 and lack the inclusion of more recent 2025 methods. In particular, direct comparisons with the latest closed-loop or iterative prompt optimization approaches are missing. Most existing baselines used in this paper focus on heuristic or search-based prompt engineering rather than a fully integrated feedback loop.\n2. While the related work section (pp. 2–3) is thorough, it does not sufficiently engage with progress in EM-like optimization procedures or Bayesian optimization–based feedback mechanisms. This omission weakens the paper’s positioning and makes SIPDO’s core feedback-loop concept appear more novel than it actually is, as similar closed-loop designs have recently emerged.\n3. Section 3.1 describes sampling from a synthetic generator regularized via KL divergence to mitigate label imbalance and mode collapse. However, key implementation details—such as how the generator is parameterized, instantiated, and updated, especially for more challenging tasks—remain underexplored and would benefit from further clarification."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926566011,"tcdate":1761931142038,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16457/Reviewer_pPqD"],"signatures":["ICLR.cc/2026/Conference/Submission16457/Reviewer_pPqD"],"forum":"kpT5rbbLdY","number":2,"license":"CC BY 4.0","cdate":1761931142038,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16457/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926566011,"domain":"ICLR.cc/2026/Conference","replyto":"kpT5rbbLdY","id":"jSFUpBcsjR","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Synthetic Data","Prompt Optimization"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Prompt quality plays a critical role in the performance of large language models (LLMs), motivating a growing body of work on prompt optimization. Most existing methods optimize prompts over a fixed dataset, assuming static input distributions and offering limited support for iterative improvement. We introduce SIPDO (Self-Improving Prompts through Data-Augmented Optimization), a closed-loop framework for prompt learning that integrates synthetic data generation into the optimization process. SIPDO couples a synthetic data generator with a prompt optimizer, where the generator produces new examples that reveal current prompt weaknesses and the optimizer incrementally refines the prompt in response. This feedback-driven loop enables systematic improvement of prompt performance without assuming access to external supervision or new tasks. Experiments across question answering and reasoning benchmarks show that SIPDO outperforms standard prompt tuning methods, highlighting the value of integrating data synthesis into prompt learning workflows."},"_bibtex":{"value":"@inproceedings{\nyu2026sipdo,\ntitle={{SIPDO}: Closed-Loop Prompt Optimization via Synthetic Data Feedback},\nauthor={Yaoning Yu and Ye Yu and Peiyan Zhang and Kai Wei and Haojing Luo and Haohan Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=kpT5rbbLdY}\n}"},"title":{"value":"SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback"},"pdf":{"value":"/pdf/3638d86d10dff1fe67a1a82961b4345ecd3b15dd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yu|sipdo_closedloop_prompt_optimization_via_synthetic_data_feedback"},"authorids":{"value":["~Yaoning_Yu1","~Ye_Yu4","~Peiyan_Zhang1","~Kai_Wei6","~Haojing_Luo1","~Haohan_Wang1"]},"authors":{"value":["Yaoning Yu","Ye Yu","Peiyan Zhang","Kai Wei","Haojing Luo","Haohan Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces ShareChatX, a large-scale synthetic spoken dialogue dataset covering diverse scenarios (emotion, audio events, music), and OmniChat, a multi-turn spoken dialogue system with a heterogeneous feature fusion module (Mix-Former). The authors argue that synthetic data can address the limitations of existing spoken dialogue datasets in terms of scale and diversity. They present extensive experiments, including ablations on data scale and synthetic/real data ratios, and claim state-of-the-art results on the DailyTalk dataset."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Can you provide more evidence (e.g., human evaluation, qualitative analysis) that models trained on ShareChatX generalize to real-world spoken dialogue, beyond the limited automatic metrics on DailyTalk?\n\n2. How do you ensure that the synthetic dialogues (especially for complex scenarios like music and audio events) are realistic and representative of real conversations?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper presents ShareChatX, a large and diverse synthetic spoken dialogue dataset, which could be a valuable resource for the community if released.\n2. The authors conduct a wide range of experiments, including ablation studies on data scale, synthetic/real data ratios, and feature fusion strategies."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed module and overall system architecture are incremental over existing multi-modal fusion approaches. The technical novelty is limited, and the paper does not provide sufficient analysis to justify the need for the new module.\n2. The evaluation on real-world data is limited in scope and depth. The evaluation primarily uses the DailyTalk dataset, which is not a standard benchmark for AudioLLM models. For a fair and meaningful comparison with other AudioLLM approaches, it would be more appropriate to use widely adopted public datasets. This limits the credibility and generalizability of the reported results.\n3. The paper’s main contribution is based on synthetic data, but it does not convincingly demonstrate that models trained on such data generalize well to real-world scenarios. The improvements on real datasets (e.g., DailyTalk) are marginal and may be due to overfitting to synthetic patterns."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942709121,"tcdate":1761865847090,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23551/Reviewer_6SJm"],"signatures":["ICLR.cc/2026/Conference/Submission23551/Reviewer_6SJm"],"forum":"vlx35uFkEK","number":3,"license":"CC BY 4.0","cdate":1761865847090,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23551/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942709121,"domain":"ICLR.cc/2026/Conference","replyto":"vlx35uFkEK","id":"RCfT8QcHab","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["emotional dialogue"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these systems still struggle to handle the full complexity of real-world conversations, including audio events, musical contexts, and emotional expressions, mainly because current dialogue datasets are constrained in both scale and scenario diversity. In this paper, we propose leveraging synthetic data to enhance the dialogue models across diverse scenarios. We introduce ShareChatX, the first comprehensive, large-scale dataset for spoken dialogue that spans diverse scenarios. Based on this dataset, we introduce OmniChat, a multi-turn dialogue system with a heterogeneous feature fusion module, designed to optimize feature selection in different dialogue contexts. In addition, we explored critical aspects of training dialogue systems using synthetic data. Through comprehensive experimentation, we determined the ideal balance between synthetic and real data, achieving state-of-the-art results on the real-world dialogue dataset DailyTalk. We also highlight the crucial importance of synthetic data in tackling diverse, complex dialogue scenarios, especially those involving audio and music. For more details, please visit our demo page at \\url{this https URL}."},"_bibtex":{"value":"@misc{\ncheng2026omnichat,\ntitle={OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios},\nauthor={Xize Cheng and Dongjie Fu and Xiaoda Yang and Minghui Fang and Ruofan Hu and Jingyu Lu and Bai Jionghao and Zehan Wang and Shengpeng Ji and Rongjie Huang and Linjun Li and Zhangzhenhua and Tao Jin and Zhou Zhao},\nyear={2026},\nurl={https://openreview.net/forum?id=vlx35uFkEK}\n}"},"title":{"value":"OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios"},"pdf":{"value":"/pdf/a3b5bfb14654fab7f5a61154883406dfbc156d5a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"cheng|omnichat_enhancing_spoken_dialogue_systems_with_scalable_synthetic_data_for_diverse_scenarios"},"authorids":{"value":["~Xize_Cheng1","~Dongjie_Fu1","~Xiaoda_Yang1","~Minghui_Fang1","~Ruofan_Hu2","~Jingyu_Lu1","~Bai_Jionghao1","~Zehan_Wang2","~Shengpeng_Ji1","~Rongjie_Huang1","~Linjun_Li2","~Zhangzhenhua1","~Tao_Jin2","~Zhou_Zhao3"]},"authors":{"value":["Xize Cheng","Dongjie Fu","Xiaoda Yang","Minghui Fang","Ruofan Hu","Jingyu Lu","Bai Jionghao","Zehan Wang","Shengpeng Ji","Rongjie Huang","Linjun Li","Zhangzhenhua","Tao Jin","Zhou Zhao"]}},"version":2},{"content":{"summary":{"value":"The manuscript presents a model for handling complex distribution shifts in graph data. The proposed GraphMETRO model employs a mixture-of-experts (MoE) architecture, with a gating model to identify distributional shifts and multiple expert models to generate shift-invariant representations. This method aims to generalize graph neural networks (GNNs) to non-synthetic distribution shifts occurring naturally in real-world graph data, achieving state-of-the-art results on several datasets."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Could you provide detailed theoretical analyses to demonstrate why the model can excellently address the problem and perform better than the baselines?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper studies an interesting research problem that is handling complex distribution shifts in graph data. This research problem has been very hot recently.\n2. The model design is easy to understand. The paper provides a detailed explanation of the proposed model.\n3. The experiment demonstrates the effectiveness of the model. The performance improvement on some comparisons seems to be significant."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the motivation is clear to me, the novelty is one of my concerns since explicitly considering instance heterogeneity for graph OOD problems is not new. I think the authors could address my concern by presenting more differences with related works clearly.\n2. There are no detailed theoretical analyses to demonstrate why the model can well address the problem and perform better than the baselines.\n3. The reviews of the related works are not enough. Please refer to the paper Out-Of-Distribution Generalization on Graphs: A Survey for a more comprehensive discussion of the related works on graph OOD problems. For example, the discussions on invariant learning in section 2 are too brief. The main claim “the standard invariant learning approaches are not well-equipped to mitigate the complex distribution shifts” is a little arbitrary considering the graph OOD literature."},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1730879007740,"tcdate":1720815661713,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission5517/Reviewer_ymW6"],"signatures":["NeurIPS.cc/2024/Conference/Submission5517/Reviewer_ymW6"],"forum":"QtYg4g3Deu","number":1,"license":"CC BY 4.0","cdate":1720815661713,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission5517/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879007740,"domain":"NeurIPS.cc/2024/Conference","replyto":"QtYg4g3Deu","id":"9PWhPbRme1","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"GraphMETRO utilizes a mixture-of-experts architecture to effectively handle complex distribution shifts in graph data, achieving state-of-the-art results on benchmark datasets."},"keywords":{"value":["Graph Neural Network","Distribution Shifts","Generalization","Mixture-of-expert model"]},"primary_area":{"value":"graph_neural_networks"},"abstract":{"value":"Graph data are inherently complex and heterogeneous, leading to a high natural diversity of distributional shifts. However, it remains unclear how to build machine learning architectures that generalize to the complex distributional shifts naturally occurring in the real world. Here, we develop GraphMETRO, a Graph Neural Network architecture that models natural diversity and captures complex distributional shifts. GraphMETRO employs a Mixture-of-Experts (MoE) architecture with a gating model and multiple expert models, where each expert model targets a specific distributional shift to produce a referential representation w.r.t. a reference model, and the gating model identifies shift components. Additionally, we design a novel objective that aligns the representations from different expert models to ensure reliable optimization. GraphMETRO achieves state-of-the-art results on four datasets from the GOOD benchmark, which is comprised of complex and natural real-world distribution shifts, improving by 67% and 4.2% on the WebKB and Twitch datasets. Code and data are available at https://github.com/Wuyxin/GraphMETRO."},"_bibtex":{"value":"@inproceedings{\nwu2024graphmetro,\ntitle={Graph{METRO}: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts},\nauthor={Shirley Wu and Kaidi Cao and Bruno Ribeiro and James Zou and Jure Leskovec},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=QtYg4g3Deu}\n}"},"title":{"value":"GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts"},"pdf":{"value":"/pdf/a59c39b6e2e8c57881013980e1abeffcb307c42e.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"wu|graphmetro_mitigating_complex_graph_distribution_shifts_via_mixture_of_aligned_experts"},"authorids":{"value":["~Shirley_Wu1","~Kaidi_Cao1","~Bruno_Ribeiro1","~James_Zou1","~Jure_Leskovec1"]},"authors":{"value":["Shirley Wu","Kaidi Cao","Bruno Ribeiro","James Zou","Jure Leskovec"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2503.11801v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"huang|diffusecloc_guided_diffusion_for_physicsbased_character_lookahead_control"},"authorids":{"value":["~Xiaoyu_Huang1","https://dblp.org/search/pid/api?q=author:Takara_Truong:","https://dblp.org/search/pid/api?q=author:Yunbo_Zhang:","https://dblp.org/search/pid/api?q=author:Fangzhou_Yu:","https://dblp.org/search/pid/api?q=author:Jean-Pierre_Sleiman:","https://dblp.org/search/pid/api?q=author:Jessica_K._Hodgins:","https://dblp.org/search/pid/api?q=author:Koushil_Sreenath:","https://dblp.org/search/pid/api?q=author:Farbod_Farshidian:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2503.11801"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2503-11801,\n  publtype={informal},\n  author={Xiaoyu Huang and Takara Truong and Yunbo Zhang and Fangzhou Yu and Jean-Pierre Sleiman and Jessica K. Hodgins and Koushil Sreenath and Farbod Farshidian},\n  title={Diffuse-CLoC: Guided Diffusion for Physics-based Character Look-ahead Control},\n  year={2025},\n  month={March},\n  cdate={1740787200000},\n  journal={CoRR},\n  volume={abs/2503.11801},\n  url={https://doi.org/10.48550/arXiv.2503.11801}\n}\n"},"abstract":{"value":"We present Diffuse-CLoC, a guided diffusion framework for physics-based look-ahead control that enables intuitive, steerable, and physically realistic motion generation. While existing kinematics motion generation with diffusion models offer intuitive steering capabilities with inference-time conditioning, they often fail to produce physically viable motions. In contrast, recent diffusion-based control policies have shown promise in generating physically realizable motion sequences, but the lack of kinematics prediction limits their steerability. Diffuse-CLoC addresses these challenges through a key insight: modeling the joint distribution of states and actions within a single diffusion model makes action generation steerable by conditioning it on the predicted states. This approach allows us to leverage established conditioning techniques from kinematic motion generation while producing physically realistic motions. As a result, we achieve planning capabilities without the need for a high-level planner. Our method handles a diverse set of unseen long-horizon downstream tasks through a single pre-trained model, including static and dynamic obstacle avoidance, motion in-betweening, and task-space control. Experimental results show that our method significantly outperforms the traditional hierarchical framework of high-level motion diffusion and low-level tracking."},"title":{"value":"Diffuse-CLoC: Guided Diffusion for Physics-based Character Look-ahead Control"},"authors":{"value":["Xiaoyu Huang","Takara Truong","Yunbo Zhang","Fangzhou Yu","Jean-Pierre Sleiman","Jessica K. Hodgins","Koushil Sreenath","Farbod Farshidian"]}},"tmdate":1747298288877,"pdate":1735689600000,"tcdate":1747298276998,"writers":["~"],"signatures":["~Xiaoyu_Huang1"],"forum":"7vmCZK6pDr","license":"CC BY-SA 4.0","number":479610,"cdate":1740787200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747298288877,"domain":"DBLP.org","id":"7vmCZK6pDr","version":2},{"content":{"summary":{"value":"This paper introduces COTCAgent, a novel agent framework for proactive medical consultation, designed to address the limitations of current healthcare AI systems in dynamic temporal reasoning and hallucination mitigation. The framework consists of two synergistic modules: a Time Series Analysis (TSA) module that programmatically extracts clinical trends from longitudinal EHR data, and a Probabilistic Chain-of-Thought Completion (COTC) module. The COTC module calculates disease risks using a novel Inverse Disease Frequency weighting scheme, identifies gaps in the reasoning chain for high-probability diseases, and proactively queries the user to complete the chain. Extensive experiments show that COTCAgent achieves state-of-the-art performance on a medical record prediction task and on the HealthBench sequential diagnosis benchmark. The work represents a significant step towards building more reliable and interactive agent-aided clinical systems."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Regarding the Database: Could the authors elaborate on the validation process for the Symptom-Trend-Disease Database? How was the agreement among the 16 validating clinicians measured? More importantly, how does the COTCAgent handle \"out-of-knowledge-base\" scenarios where a patient's symptoms or trends do not have a match in the database? Does it gracefully degrade, or does it fail?\n\n2. Regarding the Proactive Consultation Logic: The proactive questioning is a key feature. How does the system decide which and how many questions to ask? For instance, is there a confidence threshold after which the agent stops inquiring, or a limit to the number of interaction turns?\n\n3. Regarding the TSA Module's Generality: The mathematical formulations in the appendix are impressive, but how are they operationalized? For a novel user query, how does the system select the appropriate statistical test (e.g., choose a mixed-effects model over a simple regression)? Is this selection process itself automated and robust?\n\n4. Regarding Baselines: In Tables 2 and 3, the paper compares against TimeCAP and Google's Agent. Could the authors briefly characterize the architectures of these baselines to better contextualize why they fall short? Specifically, do they also employ an agent-based framework but lack the programmatic TSA module or the probabilistic COTC logic?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Novel and Well-Motivated Framework: The core idea of \"Probabilistic Chain-of-Thought Completion\" is a conceptual advance over standard CoT. The framework more faithfully mimics clinical reasoning than existing approaches. The decoupling of quantitative time-series analysis from probabilistic reasoning is a good architectural choice.\n\n\n2. Addressing a Critical and High-Impact Problem: The paper tackles two of the most pressing challenges for LLMs in healthcare: reasoning over dynamic, longitudinal data and mitigating harmful hallucinations. \n\n\n3. Enhanced Interpretability and Trustworthiness: By design, the agent's reasoning is not a black box. The explicit, multi-turn dialogue to resolve uncertainties provides a clear audit trail of its diagnostic process."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Dependence on a Custom, Partially Synthetic Knowledge Base: The framework's performance heavily relies on the \"Symptom-Trend-Disease Database.\" While the inclusion of temporal trends is innovative, the database's construction involves LLM-based augmentation, and its size is modest (23,456 entities). This raises concerns about its completeness, potential biases inherited from the augmenting LLM, and the generalizability of the system. The model's performance on diseases or symptom trends not present in this specific knowledge base is unclear.\n\n2. Simplicity of the Probabilistic Heuristic: The Inverse Disease Frequency (IDF) weighting is an intuitive and clever heuristic, but it is a simplified proxy for true diagnostic probability. It does not account for symptom co-occurrence, disease prevalence, or more complex conditional dependencies that are central to differential diagnosis. The paper would be strengthened by a discussion against more established probabilistic graphical models (e.g., Bayesian Networks)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923009098,"tcdate":1761384068444,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12026/Reviewer_cBZg"],"signatures":["ICLR.cc/2026/Conference/Submission12026/Reviewer_cBZg"],"forum":"Uo9xZ2kePW","number":1,"license":"CC BY 4.0","cdate":1761384068444,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12026/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923009098,"domain":"ICLR.cc/2026/Conference","replyto":"Uo9xZ2kePW","id":"RDh0GHJKlj","forumContent":{"TLDR":{"value":"The first to propose a query method for completing the Chain-of-Thought (CoT) based on time-series health check-up data, which significantly improves the accuracy and satisfaction of personalized proactive consultation and risk prediction"},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Chain-of-Thought，LLM Agent，Healthcare，Time Series Indicators，Risk Quantification"]},"supplementary_material":{"value":"/attachment/d34bdfc3e0f59172adccf4f5b75641e87e25fdf7.zip"},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Current agent-based healthcare AI systems struggle with dynamic temporal reasoning and hallucination mitigation, limiting their preventive care utility. We introduce COTCAgent, a proactive consultation framework featuring a novel Probabilistic Chain-of-Thought completion mechanism. Our approach integrates two synergistic modules: a Time Series Analysis Module that extracts clinical trends from longitudinal EHRs, and a COTC Module that calculates disease risks via Inverse Disease Frequency weighting and completes reasoning chains through targeted questioning. Extensive evaluations demonstrate state-of-the-art performance, with COTCAgent achieving 89.2\\% accuracy in medical risk prediction (vs. 77.9-80.2\\% for baselines) and 69.8\\% on the challenging HealthBench sequential diagnosis benchmark. This work bridges temporal analysis with probabilistic reasoning to enable truly personalized preventive care."},"_bibtex":{"value":"@misc{\nanonymous2026cotcagent,\ntitle={{COTCA}gent: Preventive Care Proactive Consultation Driven by a Probabilistic Chain-of-Thought Completion Framework},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=Uo9xZ2kePW}\n}"},"title":{"value":"COTCAgent: Preventive Care Proactive Consultation Driven by a Probabilistic Chain-of-Thought Completion Framework"},"pdf":{"value":"/pdf/22a1e1860d773dd11b3af28b82ce55fcaa411ec2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"dengzihan|cotcagent_preventive_care_proactive_consultation_driven_by_a_probabilistic_chainofthought_completion_framework"},"authorids":{"value":["~DengZihan1","~XiaozhenZhong1","~Le_Yu8","~Guoliang_Li6"]},"authors":{"value":["DengZihan","XiaozhenZhong","Le Yu","Guoliang Li"]}},"version":2},{"content":{"TLDR":{"value":"Physics-regulated DRL"},"venue":{"value":"ICLR 2024 spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed deep reinforcement learning","Safety-critical autonomous systems"]},"supplementary_material":{"value":"/attachment/289020ab57bf742ddf172c907d97beacd14306b7.zip"},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This paper proposes the Phy-DRL: a physics-regulated deep reinforcement learning (DRL) framework for safety-critical autonomous systems. The Phy-DRL has three distinguished invariant-embedding designs: i) residual action policy (i.e., integrating data-driven-DRL action policy and physics-model-based action policy), ii) automatically constructed safety-embedded reward, and iii) physics-model-guided neural network (NN) editing, including link editing and activation editing. Theoretically, the Phy-DRL exhibits 1) a mathematically provable safety guarantee and 2) strict compliance of critic and actor networks with physics knowledge about the action-value function and action policy. Finally, we evaluate the Phy-DRL on a cart-pole system and a quadruped robot. The experiments validate our theoretical results and demonstrate that Phy-DRL features guaranteed safety compared to purely data-driven DRL and solely model-based design while offering remarkably fewer learning parameters and fast training towards safety guarantee."},"_bibtex":{"value":"@inproceedings{\ncao2024physicsregulated,\ntitle={Physics-Regulated Deep Reinforcement Learning: Invariant Embeddings},\nauthor={Hongpeng Cao and Yanbing Mao and Lui Sha and Marco Caccamo},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=5Dwqu5urzs}\n}"},"title":{"value":"Physics-Regulated Deep Reinforcement Learning: Invariant Embeddings"},"pdf":{"value":"/pdf/da97cca0d38172e4f655bd23891b3c1509513545.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"cao|physicsregulated_deep_reinforcement_learning_invariant_embeddings"},"authorids":{"value":["~Hongpeng_Cao1","~Yanbing_Mao1","~Lui_Sha1","~Marco_Caccamo2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hongpeng Cao","Yanbing Mao","Lui Sha","Marco Caccamo"]}},"tmdate":1711320560098,"pdate":1705410829781,"tcdate":1695117176058,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1772/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission1772/Authors"],"forum":"5Dwqu5urzs","number":1772,"cdate":1695117176058,"mdate":1711320560098,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/-/Submission","ICLR.cc/2024/Conference/-/Post_Submission","ICLR.cc/2024/Conference/Submission1772/-/Revision","ICLR.cc/2024/Conference/Submission1772/-/Rebuttal_Revision","ICLR.cc/2024/Conference/-/Edit","ICLR.cc/2024/Conference/Submission1772/-/Camera_Ready_Revision"],"odate":1697213872796,"domain":"ICLR.cc/2024/Conference","id":"5Dwqu5urzs","version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Existing feed-forward networks excel at predicting a single set of physical properties from visual appearance, but this point-estimate paradigm fundamentally fails to capture the real world’s inherent physical ambiguity. We address this by reframing physics prediction as a task of learning a controllable, continuous distribution of material properties. We introduce UNIPIXIE, a framework trained to predict a continuous and parameterized path of physically plausible material properties from a single visual input. By learning a direct mapping along an object’s softest-to-stiffest spectrum on our PIXIEMULTIVERSE dataset, UNIPIXIE allows for controllable generation of diverse, physically-valid material fields via a single intuitive parameter. Crucially, UNIPIXIE introduces a novel unified architecture to produce simulation-ready parameters for diverse physics solvers, including continuum-based Material Point Method (MPM), reduced-order deformation based on Linear Blend Skinning (LBS), and anchor-based Spring-Mass systems, addressing a key portability issue in prior work. Experiments show our approach not only generates a rich variety of plausible dynamics but also reduces Young’s Modulus prediction error by over 50% against the strongest deterministic baseline, bridging the gap between static point-estimates and the continuous nature of physical reality."},"_bibtex":{"value":"@inproceedings{\nhuang2026unipixie,\ntitle={UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching},\nauthor={Qilin Huang and Quynh Anh Huynh and Long Le and Chen Wang and Chuhao Chen and Ryan Lucas and Eric Eaton and Lingjie Liu},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=P5gNBztFer}\n}"},"title":{"value":"UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Huang_UniPixie_Unified_and_Probabilistic_3D_Physics_Learning_via_Flow_Matching_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"huang|unipixie_unified_and_probabilistic_3d_physics_learning_via_flow_matching"},"authorids":{"value":["~Qilin_Huang1","~Quynh_Anh_Huynh1","~Long_Le1","~Chen_Wang13","~Chuhao_Chen1","~Ryan_Lucas1","~Eric_Eaton2","~Lingjie_Liu1"]},"authors":{"value":["Qilin Huang","Quynh Anh Huynh","Long Le","Chen Wang","Chuhao Chen","Ryan Lucas","Eric Eaton","Lingjie Liu"]}},"tmdate":1789656586889,"pdate":1789656437888,"tcdate":1765222188058,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission39186/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission39186/Authors"],"forum":"P5gNBztFer","license":"CC BY 4.0","number":39186,"cdate":1765222188058,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission39186/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/Submission39186/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656586889,"odate":1789656437888,"domain":"thecvf.com/CVPR/2026/Conference","id":"P5gNBztFer","version":2},{"content":{"summary":{"value":"This paper introduces WEBDART, a framework that enables large language model (LLM) agents to handle complex web tasks that require long-horizon reasoning and structured exploration. It dynamically decomposes each task into three subtasks: (1) navigation, (2) information extraction, and (3) execution, allowing the model to focus on one ability at a time. During navigation, the agent adaptively re-plans its strategy when new filters or interface shortcuts appear, reducing redundant actions and improving efficiency. This modular and adaptive design enhances task completion and robustness in complex web environments while maintaining strong performance on simpler tasks. Overall, WEBDART demonstrates that dynamic decomposition and real-time re-planning can significantly improve the reasoning and adaptability of LLM-based web agents."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- In Section 4.2, the authors claim that the observed improvements “highlight the advantage of shifting constraint handling to the data analysis stage.” However, it is unclear how the empirical results in Table 1 specifically support this interpretation. Could the authors clarify what evidence connects the performance gains to this design choice? \n\n- Table 3 reports the Results on the WebArena benchmark and includes additional baselines such as HybridAgent [1] and WebPilot [2], which show competitive performance. How do these baselines perform on WebChoreArena, and were they excluded due to reproducibility constraints or unavailability of results? \n\n- As an ablation, how does performance change when the routing module is disabled, particularly on the WebArena benchmark? It would be helpful to know how much accuracy drops and what types of routing errors occur (e.g., skipping extraction when it is actually required). Additionally, could the authors provide a brief analysis of the common failure cases in WEBDART?\n\n[1] Song, et al. \"Beyond browsing: Api-based web agents.\" arXiv preprint arXiv:2410.16464 (2024).\n\n[2] Zhang, et al. \"Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration.\" AAAI 2025"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- **Well-Justified Motivation**: The paper effectively addresses the importance of long-horizon web tasks as a fundamental challenge in current web-agent research.\n\n- **Clear Writing and Organization**: The paper is well-written and easy to follow, with a well-organized structure and clear presentation of the proposed approach.\n\n- **Simple Yet Effective Design**:This paper employs an intuitive three-stage decomposition that mirrors how humans naturally approach complex web tasks, resulting in a method that is both easy to understand and practically effective."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Lack of empirical justification for the conservative decomposition scheme**: The paper adopts the conservative scheme (deferring constraint handling to later stages) as the default strategy. However, this design choice is not supported by any preliminary analysis, empirical comparison, or prior evidence—for example, there is no ablation or user study contrasting conservative versus tightly coupled decompositions. Given that the efficiency of each scheme “hinges on site features” (line 204), a fixed conservative default appears heuristic rather than data-driven, and its general validity across domains remains unclear. A short pilot experiment or reference to earlier literature on adaptive task partitioning would strengthen this methodological decision and clarify why the conservative bias is justified beyond intuition.\n\n- **Heuristic Nature of Information Extraction**: The information extraction pipeline is heuristic, relying on LLM prompts to select relevant pages and extract fields without any quantitative validation or ablation. The paper explains that the model “returns an index set that marks the pages most likely to contain the required information,” yet provides no concrete mechanism or evidence to show how reliable this selection is. Furthermore, the dismissal of the LLM-generated parser baseline is entirely qualitative, lacking any comparative results or failure statistics. Overall, the decision to rely solely on prompt-based extraction appears intuitive rather than experimentally justified, leaving uncertainty about its robustness and reproducibility across diverse web structures.\n\n- **Lack of In-depth Performance Analysis**: While the paper reports overall success rates on the WebChoreArena benchmark [1], it does not provide finer-grained analyses that could strengthen its empirical claims. In the original benchmark, performance is typically broken down by cross-site domains as well as by task types such as Calculate, Long-Term Memory, Massive Memory, and Other. However, WEBDART’s results are aggregated, making it unclear which categories drive the observed improvements. The absence of such detailed breakdowns limits the interpretability of the reported gains and prevents deeper insights into where the proposed method truly excels or struggles.\n\n[1] Miyai, Atsuyuki, et al. \"WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks.\" arXiv preprint arXiv:2506.01952 (2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942206210,"tcdate":1761722253332,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22407/Reviewer_RREB"],"signatures":["ICLR.cc/2026/Conference/Submission22407/Reviewer_RREB"],"forum":"RmZuzuOu9g","number":2,"license":"CC BY 4.0","cdate":1761722253332,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22407/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942206210,"domain":"ICLR.cc/2026/Conference","replyto":"RmZuzuOu9g","id":"DdAIlm2W7K","forumContent":{"TLDR":{"value":"WebDART tackles complex web tasks by splitting navigation, extraction, and analysis, and adding dynamic re-planning, reducing navigation burden and boosting success by +13.7% on WebChoreArena."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["LLM","Agent","Web Agent","Task Decomposition"]},"supplementary_material":{"value":"/attachment/aea490857515bd1208af33fc1e2793a22c70f5a2.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large-language-model (LLM) agents are becoming competent at straightforward web tasks, such as opening an item page or submitting a form, but still struggle with objectives that require long-horizon navigation, large-scale information extraction, and reasoning under constraints. We present WebDART, a general framework that enables a single LLM to handle such complex chores. WebDART (i) dynamically decomposes each objective into three focused subtasks—navigation, information extraction, and execution—so the model concentrates on one skill at a time, and (ii) continuously re-plans the decomposition as new webpages are revealed, taking advantage of newly discovered filters or shortcuts and avoiding redundant exploration. Evaluated on WebChoreArena, WebDART lifts end-to-end success rates by up to 13.7 percentage points over previous state-of-the-art agents, while matching their performance on the easier WebArena suite and completing tasks with up to 14.7 fewer navigation steps. Code will be publicly available."},"_bibtex":{"value":"@misc{\nyang2026webdart,\ntitle={Web{DART}: Dynamic Decomposition and Re-planning for Complex Web Tasks},\nauthor={Jingbo Yang and Bairu Hou and Wei Wei and Yujia Bao and Shiyu Chang},\nyear={2026},\nurl={https://openreview.net/forum?id=RmZuzuOu9g}\n}"},"title":{"value":"WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks"},"pdf":{"value":"/pdf/ed3698230f03da949d06299e7e424582925355f7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yang|webdart_dynamic_decomposition_and_replanning_for_complex_web_tasks"},"authorids":{"value":["~Jingbo_Yang3","~Bairu_Hou2","~Wei_Wei15","~Yujia_Bao1","~Shiyu_Chang2"]},"authors":{"value":["Jingbo Yang","Bairu Hou","Wei Wei","Yujia Bao","Shiyu Chang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"https://arxiv.org/pdf/2310.12181v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"zhu|precise_influence_evaluation_in_complex_networks"},"html":{"value":"https://doi.org/10.48550/arXiv.2310.12181"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2310-12181,\n  publtype={informal},\n  author={Bingyu Zhu and Qingyun Sun and Jianxin Li and Daqing Li},\n  title={Precise influence evaluation in complex networks},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2310.12181},\n  url={https://doi.org/10.48550/arXiv.2310.12181}\n}\n"},"abstract":{"value":"Evaluating node influence is fundamental for identifying key nodes in complex networks. Existing methods typically rely on generic indicators to rank node influence across diverse networks, thereby ignoring the individualized features of each network itself. Actually, node influence stems not only from general features but the multi-scale individualized information encompassing specific network structure and task. Here we design an active learning architecture to predict node influence quantitively and precisely, which samples representative nodes based on graph entropy correlation matrix integrating multi-scale individualized information. This brings two intuitive advantages: (1) discovering potential high-influence but weak-connected nodes that are usually ignored in existing methods, (2) improving the influence maximization strategy by deducing influence interference. Significantly, our architecture demonstrates exceptional transfer learning capabilities across multiple types of networks, which can identify those key nodes with large disputation across different existing methods. Additionally, our approach, combined with a simple greedy algorithm, exhibits dominant performance in solving the influence maximization problem. This architecture holds great potential for applications in graph mining and prediction tasks."},"title":{"value":"Precise influence evaluation in complex networks"},"authors":{"value":[{"fullname":"Bingyu Zhu","username":""},{"fullname":"Qingyun Sun","username":"~Qingyun_Sun2"},{"fullname":"Jianxin Li","username":""},{"fullname":"Daqing Li","username":"~Daqing_Li1"}]}},"tmdate":1790772025065,"pdate":1703980800000,"externalIds":["dblp:journals/corr/abs-2310-12181"],"tcdate":1784621876028,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Qingyun_Sun2"],"forum":"Su9DCgzixf","license":"CC BY-SA 4.0","number":80856,"cdate":1672531200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1790772025065,"domain":"OpenReview.net/Public_Article","id":"Su9DCgzixf","version":2},{"content":{"summary":{"value":"This paper investigates how prompt complexity affects synthetic data utility from T2I models across quality, diversity, and consistency axes. The authors (1) conduct synthetic experiments on Gaussian mixtures with theoretical derivations, (2) introduce an evaluation framework comparing real vs. synthetic data, and (3) perform large-scale empirical analysis across CC12M, ImageNet-1k, and DCI datasets with multiple T2I models and inference-time interventions. Key findings: increasing prompt complexity reduces diversity and consistency but narrows the real-synthetic gap; prompt expansion consistently improves diversity and quality; combining advanced guidance with prompt expansion yields optimal trade-offs"},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- **Can you characterize discarded images in the alignment step?** Specifically, do they have different mean aesthetic quality, object count, scene complexity, or semantic diversity compared to retained images? This would help assess whether the observed trends are artifacts of selection bias.\n- **On Theoretical Assumptions:** Could you discuss the validity of the conditional independence assumption (Appendix A.2) in the context of compositional text encoders like CLIP? How might the violation of this assumption, where concepts like \"red\" and \"car\" are highly correlated and compositionally represented, affect the interpretation that generalizing to general prompts is inherently \"harder\" than generalizing to fine-grained ones? Could this theory-practice gap also explain some of the divergent behaviors observed between the CC12M and ImageNet experiments?\n- **Can you provide confidence intervals or significance tests for main trends?** With 5,000 prompts per complexity, bootstrapped confidence intervals for Vendi, aesthetic, and FDD would strengthen claims about trends and help assess how statistically significant differences between adjacent complexity levels are."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- **Important and underexplored research question.** This is the first systematic study of how prompt complexity affects T2I synthetic data utility. Given the widespread use of T2I models for data generation and the common practice of training on synthetic captions, understanding this relationship is timely and valuable.\n    \n- **Well-designed evaluation framework.** The 5-step pipeline (captioning, pairing, alignment, sampling, generation) enables fair comparisons between real and synthetic data across prompt complexities, a non-trivial methodological contribution that could benefit future work.\n    \n- **Comprehensive empirical evaluation.** The study covers multiple datasets (CC12M, ImageNet-1k, DCI), models (LDMv1.5/XL/v3.5M/v3.5L, Flux, Infinity), inference methods (CFG, CADS, Interval, APG, prompt expansion), and metrics (reference-free: Vendi, aesthetic, DSG; reference-based: FDD, precision, density, coverage), providing thorough coverage.\n    \n- **Human validation strengthens metric choice.** The human evaluation (Appendix E) shows Vendi score strongly correlates with human-perceived diversity, validating the automatic diversity metric.\n    \n- **Illustrative toy example provides intuition.** The synthetic Gaussian mixture experiments (Section 2) with mathematical derivations (Appendix A.2) offer clear intuition for why generalizing to general prompts is harder, nicely complementing the empirical findings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Independence assumptions in theoretical derivations may not hold for real text encoders.** The synthetic experiments (Section 2, Appendix A.2) derive score functions assuming conditional independence of text concepts. Specifically, Equations 5-7 assume that for fine-grained conditioning $c_f$ composed of general concepts ${c^i_g}$, we have $p(x_t, c^1_g, c^2_g, ..., c^K_g) = p(x_t) ∏_{i} p(c^i_g|x_t)$. However, real text encoders (CLIP, T5) do not treat concepts as independent: (1) they encode compositional phrases where \"white\" modifying \"dog\" produces entangled representations rather than separable \"white\" and \"dog\" signals, (2) they learn correlations from training data where certain concept combinations (e.g., \"white dog\") appear frequently, making $p(c^1_g|x_t, c^2_g) ≠ p(c^1_g|x_t)$, and (3) concept embeddings are context-dependent and non-orthogonal in the latent space. Similarly, the OR operator derivation (Equation 1) requires fine-grained categories to be exhaustive and mutually exclusive, which may not align with how models internally represent general concepts. The Gaussian mixture model provides valuable intuition, but the theoretical predictions should be interpreted cautiously when applied to real T2I models. The authors should discuss this gap and consider how it might affect the interpretation of their results. \n- **The alignment step lacks transparency and may introduce selection bias.** Algorithm 1 iteratively removes images not shared across complexities, but provides no analysis of what is discarded. Table 1 shows complexity-4 prompts cover only 46,066 images vs. 61,334 for complexity-1, a 25% reduction. What visual or semantic characteristics differentiate retained vs. discarded images? If alignment preferentially keeps images easier to describe with detailed prompts (e.g., clear objects vs. abstract scenes), this biases the evaluation set. Do discarded images have different mean aesthetic quality, diversity, or complexity than retained ones? Without this analysis, it is difficult to disentangle the genuine effects of prompt complexity from potential artifacts introduced by the evaluation pipeline itself."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762939056004,"tcdate":1762444442306,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20963/Reviewer_QZqM"],"signatures":["ICLR.cc/2026/Conference/Submission20963/Reviewer_QZqM"],"forum":"RBIBMCdw7y","number":4,"license":"CC BY 4.0","cdate":1762444442306,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20963/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762939056004,"domain":"ICLR.cc/2026/Conference","replyto":"RBIBMCdw7y","id":"uMwGgF430R","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["text-to-image models","prompt complexity","synthetic data"]},"supplementary_material":{"value":"/attachment/19f32551040643a1e2887ebe2e8ce7caa29f875a.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Text-to-image (T2I) models offer great potential for creating virtually limitless synthetic data, a valuable resource compared to fixed and finite real datasets. Previous works evaluate the utility of synthetic data from T2I models on three key desiderata: quality, diversity, and consistency. While prompt engineering is the primary means of interacting with T2I models, the systematic impact of prompt complexity on these critical utility axes remains underexplored. In this paper, we first conduct synthetic experiments to motivate the difficulty of generalization w.r.t. prompt complexity and explain the observed difficulty with theoretical derivations. Then, we introduce a new evaluation framework that can compare the utility of real data and synthetic data, and present a comprehensive analysis of how prompt complexity influences the utility of synthetic data generated by commonly used T2I models. We conduct our study across diverse datasets, including CC12M, ImageNet-1k, and DCI, and evaluate different inference-time intervention methods. Our synthetic experiments show that generalizing to more general conditions is harder than the other way round, since the former needs an estimated likelihood that is not learned by diffusion models. Our large-scale empirical experiments reveal that increasing prompt complexity results in lower conditional diversity and prompt consistency, while reducing the synthetic-to-real distribution shift, which aligns with the synthetic experiments. Moreover, current inference-time interventions can augment the diversity of the generations at the expense of moving outside the support of real data. Among those interventions, prompt expansion, by deliberately using a pre-trained language model as a likelihood estimator, consistently achieves the highest performance in both image diversity and aesthetics, even higher than that of real data. Combining advanced guidance interventions with prompt expansion results in the most appealing utility trade-offs of synthetic data."},"_bibtex":{"value":"@inproceedings{\nzhang2026the,\ntitle={The Intricate Dance of Prompt Complexity, Quality, Diversity and Consistency in T2I Models},\nauthor={Xiaofeng Zhang and Aaron Courville and Michal Drozdzal and Adriana Romero-Soriano},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=RBIBMCdw7y}\n}"},"title":{"value":"The Intricate Dance of Prompt Complexity, Quality, Diversity and Consistency in T2I Models"},"pdf":{"value":"/pdf/bafe0f0325f15c351788c224c0b6b3caaac71dcd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|the_intricate_dance_of_prompt_complexity_quality_diversity_and_consistency_in_t2i_models"},"authorids":{"value":["~Xiaofeng_Zhang2","~Aaron_Courville3","~Michal_Drozdzal1","~Adriana_Romero-Soriano1"]},"authors":{"value":["Xiaofeng Zhang","Aaron Courville","Michal Drozdzal","Adriana Romero-Soriano"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a new knowledge graph complex query answering method called NLISA, featured by computationally fast. It is a neural-symbolic framework with local and global constraints. Experiments on the benchmarks show that the NLISA framework could reduce computation by 90% with a minimal performance loss."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. In Figure2, should the query graph represent the query “Fine someone who is married to a person who graduated from the same institution”? Since the in the triples (x1, Graduate, y) and (x1, Graduate, x2), the head entity refers to the same variable x1. \n2. The $\\top$ in Equation (3) and (4) is hard to understand. Does the $\\top$ represents operations between the truth values?  How should we interpret the Equation (3) and (4)? \n3. What are the entity embedding and relation embedding from equation in line 244 from? is it pre-trained by some methods? \n4. How the hyper-network are train? What data are used for the training?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper studies an important problem of fast knowledge graph complex query answering, which could benefit the development of the knowledge graph complex query answering methods. \n2. The proposed fast knowledge graph complex query answering method is efficient that could reduce the computation by 90%."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The main method is hard to understand. After reading, I find it is not clear how the neural logical indices are achieved. \n2. Some method and experiment details are missing(refer to question 3 and 4 ). \n3. Though the NLISA is efficient but the complex query answering performances of NLISA are suboptimal. Table 1 shows the NLISA(local) is more efficient that achieves the best efficiency results. But most of the task performance of NLISA(local) in Table 2 are worse than FIT method. Especially the AVG(P) results on FB15k datasets is significantly worse."}},"nonreaders":[],"tmdate":1731428854966,"tcdate":1730720475747,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7276/Reviewer_p8gq"],"signatures":["ICLR.cc/2025/Conference/Submission7276/Reviewer_p8gq"],"forum":"PbxKOPtoEE","number":3,"license":"CC BY 4.0","cdate":1730720475747,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7276/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428854966,"domain":"ICLR.cc/2025/Conference","replyto":"PbxKOPtoEE","id":"g5gX04WJGa","forumContent":{"TLDR":{"value":"We present an efficient symbolic search framework."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["complex logical query","knowledge graph","query embedding","neural link predictor"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Complex query answering (CQA) over knowledge graphs is a crucial multi-hop reasoning task aimed at addressing first-order logical queries within large and incomplete knowledge graphs. Direct traversal search methods rely solely on graph topology and often miss answers due to the incompleteness of the graph, thus neural models have been proposed to generalize the neglected answers from observed facts. There are primarily two lines of research tackling the challenges of CQA. Query embedding models learn representations for complex queries, offering fast speed but often providing only generic performance. In contrast,  neural symbolic search methods deliver better performance, although they tend to be computationally more expensive. In this paper, we propose an efficient and scalable search framework that combines the precision of symbolic methods with the speed of embedding techniques. Our model utilizes embedding methods to compute Neural Logical Indices (NLI) to reduce the search domain for each variable in advance, followed by an approximate symbolic search for fine ranking. The search is precise for tree-structured queries and approximates cyclic queries (which are NP-complete) in quadratic complexity concerning the search domain, matching the complexity of tree-form queries. Experiments on various CQA benchmarks show that our framework reduces computation by 90% with a minimal performance loss, alleviating both efficiency and scalability issues."},"_bibtex":{"value":"@misc{\nfei2025neural,\ntitle={Neural Logical Index for Fast Knowledge Graph Complex Query Answering},\nauthor={Weizhi Fei and Zihao Wang and Hang Yin and Shukai Zhao and Wei Zhang and Hanghang Tong and Yangqiu Song},\nyear={2025},\nurl={https://openreview.net/forum?id=PbxKOPtoEE}\n}"},"title":{"value":"Neural Logical Index for Fast Knowledge Graph Complex Query Answering"},"pdf":{"value":"/pdf/05060ba715c702da2701cac44ed0df5b9e90b69a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"fei|neural_logical_index_for_fast_knowledge_graph_complex_query_answering"},"authorids":{"value":["~Weizhi_Fei1","~Zihao_Wang11","~Hang_Yin3","~Shukai_Zhao1","~Wei_Zhang126","~Hanghang_Tong3","~Yangqiu_Song1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Weizhi Fei","Zihao Wang","Hang Yin","Shukai Zhao","Wei Zhang","Hanghang Tong","Yangqiu Song"]}},"version":2},{"content":{"comment":{"value":"**5) Have you incorporated real-world anomalies or edge cases into your data generation process?**\n\n- Yes, the injected trends mostly reflect real-world anomalies and edge cases, incorporated with guidance from domain experts, including complex seasonal patterns. \n\n- For example, in Datasets 14 and 6, the average Time to Resolution (TTR) steadily increases despite adequate staffing. The LLM Agent must analyze further to uncover that this isn’t due to ineffective support staff but rather because some staff are on leave, causing others to be overloaded.\n\n**6) Have you validated the synthetic data against any real-world datasets?**\n\n- We have evaluated on real ServiceNow datasets and we have found consistent relative performance results (w.r.t. L3-Eval metric) between AgentPoirot and PandasAgent on both our synthetic and real world data. The real-world data includes CSM, ITSM, and HR. \n\n|        Method    | Our Synthetic Data | Real World Data |\n|--------------------|---------------------------|------------------------|\n| PandasAgent |           0.54               |          0.58           |\n|  AgentPoirot   |           0.60               |          0.64           |\n\n- We unfortunately cannot include real enterprise data due to its proprietary nature.\n- Please note that our synthetic data was vetted by 12 ServiceNow experts that we partnered with that have deep knowledge of actual enterprise data to ensure it accurately represents real-world nuances.  Additionally, we have conducted a human evaluation study (detailed in Appendix B.4) where enterprise data specialists rated our synthetic data highly for its quality, coverage, and relevance. \n\n\n(2/2) End of Rebuttal"},"title":{"value":"Rebuttal (2/2) for Reviewer nFSq on InsightBench"}},"tmdate":1732129544566,"tcdate":1732129544566,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11745/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission11745/Authors"],"forum":"ZGqd0cbBvm","number":7,"license":"CC BY 4.0","cdate":1732129544566,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11745/-/Official_Comment"],"mdate":1732129544566,"domain":"ICLR.cc/2025/Conference","replyto":"IJjK6rk5eH","id":"TGoy7Z8084","forumContent":{"TLDR":{"value":"We propose a comprehensive benchmark to evaluate LLM-based agents on their ability to perform multi-step data analysis and discover interesting insights in data."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Automated Data Analysis","Data Analytics Benchmark","LLM agents","Code Generation","LLM Evaluation"]},"supplementary_material":{"value":"/attachment/5ea652062cd4788c28d4c40a13dc39357f950f1f.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Data analytics is essential for extracting valuable insights from data that can assist organizations in making effective decisions. We introduce InsightBench, a benchmark dataset with three key features. First, it consists of 100 datasets representing diverse business use cases such as finance and incident management, each accompanied by a carefully curated set of insights planted in the datasets. Second, unlike existing benchmarks focusing on answering single queries, InsightBench evaluates agents based on their ability to perform end-to-end data analytics, including formulating questions, interpreting answers, and generating a summary of insights and actionable steps. Third, we conducted comprehensive quality assurance to ensure that each dataset in the benchmark had clear goals and included relevant and meaningful questions and analysis. Furthermore, we implement a two-way evaluation mechanism using LLaMA-3 as an effective, open-source evaluator to assess agents’ ability to extract insights. We also propose AgentPoirot, our baseline data analysis agent capable of performing end-to-end data analytics. Our evaluation on InsightBench shows that AgentPoirot outperforms existing approaches (such as Pandas Agent) that focus on resolving single queries. We also compare the performance of open- and closed-source LLMs and various evaluation strategies. Overall, this benchmark serves as a testbed to motivate further development in comprehensive automated data analytics and can be accessed here: https://github.com/ServiceNow/insight-bench."},"_bibtex":{"value":"@inproceedings{\nsahu2025insightbench,\ntitle={InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation},\nauthor={Gaurav Sahu and Abhay Puri and Juan A. Rodriguez and Amirhossein Abaskohi and Mohammad Chegini and Alexandre Drouin and Perouz Taslakian and Valentina Zantedeschi and Alexandre Lacoste and David Vazquez and Nicolas Chapados and Christopher Pal and Sai Rajeswar and Issam H. Laradji},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=ZGqd0cbBvm}\n}"},"title":{"value":"InsightBench: Evaluating Business Analytics Agents Through Multi-Step Insight Generation"},"pdf":{"value":"/pdf/be19956312bdbfe85a7b95921ff9bf089650fc68.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"sahu|insightbench_evaluating_business_analytics_agents_through_multistep_insight_generation"},"authorids":{"value":["~Gaurav_Sahu2","~Abhay_Puri1","~Juan_A._Rodriguez1","~Amirhossein_Abaskohi1","~Mohammad_Chegini1","~Alexandre_Drouin2","~Perouz_Taslakian1","~Valentina_Zantedeschi2","~Alexandre_Lacoste1","~David_Vazquez1","~Nicolas_Chapados1","~Christopher_Pal1","~Sai_Rajeswar2","~Issam_H._Laradji1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Gaurav Sahu","Abhay Puri","Juan A. Rodriguez","Amirhossein Abaskohi","Mohammad Chegini","Alexandre Drouin","Perouz Taslakian","Valentina Zantedeschi","Alexandre Lacoste","David Vazquez","Nicolas Chapados","Christopher Pal","Sai Rajeswar","Issam H. Laradji"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel framework named AFL (Agentic Framework with LLMs), which aims to fully automate the solving of complex Vehicle Routing Problems (VRPs) using Large Language Models (LLMs). It decomposes the complex solving pipeline into three subtasks: problem description, code generation, and solution derivation. It also introduces four specialized LLM Agents—Generation, Critique, Revision, and Error Analysis—to collaborate, significantly improving the reliability of the generated code and the feasibility of the final solution."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"see in weakness"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. AFL achieves full end-to-end automation, from the VRP instance file to the final solution, without requiring human intervention during execution or relying on external solvers or predefined code libraries.\n2. Decomposing the complex task and introducing multiple collaborating LLM Agents with specialized roles (Generation, Critique, Revision, Error Analysis) is an effective strategy for enhancing the reliability of LLMs in complex programming and reasoning tasks. This approach is generalizable and could potentially be applied to other optimization problems.\n3. The experiments not only include standard VRPs but also specifically test complex variants more common in real-world scenarios, such as Electric VRPs (EVRPs) with multiple combined constraints, demonstrating the framework's effectiveness and generality in handling complex constraints."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. It is highly dependent on the powerful code generation, comprehension, and reasoning capabilities of the LLM used (GPT-4.1). How the framework performs on less capable, open-source LLMs, and its sensitivity to different LLMs, warrants further investigation.\n2. Excessive time consumption is a potential issue, as shown in Table 3. How do the authors view this trade-off of sacrificing time for accuracy? In scenarios requiring rapid feedback, is this algorithm still viable?\n\nSince I am not an expert in this field,  I don't know if the novelty of this kind of prompt engineering paper is sufficient."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926416352,"tcdate":1761609144215,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16267/Reviewer_KZcf"],"signatures":["ICLR.cc/2026/Conference/Submission16267/Reviewer_KZcf"],"forum":"BMOgYw4EhQ","number":2,"license":"CC BY 4.0","cdate":1761609144215,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16267/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926416352,"domain":"ICLR.cc/2026/Conference","replyto":"BMOgYw4EhQ","id":"iHn4wodo3Q","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We propose a fully automation LLM agent for solving complex vehicle routing problem"},"keywords":{"value":["Vehicle Routing Problems","Agent","LLM"]},"primary_area":{"value":"optimization"},"abstract":{"value":"Complex vehicle routing problems (VRPs) remain a fundamental challenge, demanding substantial expert effort for intent interpretation and algorithm design. While large language models (LLMs) offer a promising path toward automation, current approaches still rely on external intervention, which restrict autonomy and often lead to execution errors and low solution feasibility. To address these challenges, we propose an Agentic Framework with LLMs (AFL) for solving complex vehicle routing problems, achieving full automation from problem instance to solution. AFL directly extracts knowledge from raw inputs and enables self-contained code generation without handcrafted modules or external solvers. To improve trustworthiness, AFL decomposes the overall pipeline into three manageable subtasks and employs four specialized agents whose coordinated interactions enforce cross-functional consistency and logical soundness. Extensive experiments on 60 complex VRPs, ranging from standard benchmarks to practical variants, validate the effectiveness and generality of our framework, showing comparable performance against meticulously designed algorithms. Notably, it substantially outperforms existing LLM-based baselines in both code reliability and solution feasibility, achieving rates close to 100% on the evaluated benchmarks."},"_bibtex":{"value":"@inproceedings{\nzhang2026an,\ntitle={An Agentic Framework with {LLM}s for Solving Complex Vehicle Routing Problems},\nauthor={Ni Zhang and Zhiguang Cao and Jianan Zhou and Cong Zhang and Yew-Soon Ong},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=BMOgYw4EhQ}\n}"},"title":{"value":"An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems"},"pdf":{"value":"/pdf/851ae6360b1ec96e42a4dd82e5862259a62d6748.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|an_agentic_framework_with_llms_for_solving_complex_vehicle_routing_problems"},"authorids":{"value":["~Ni_Zhang6","~Zhiguang_Cao1","~Jianan_Zhou1","~Cong_Zhang3","~Yew-Soon_Ong1"]},"authors":{"value":["Ni Zhang","Zhiguang Cao","Jianan Zhou","Cong Zhang","Yew-Soon Ong"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a variation of the large reconstruction models for reconstructing objects (geometry and albedo) from single or multi-view input images. By incorporating physically based rendering pipeline for image synthesis in training, as well as ground truth images rendered under sampled BRDF and lighting conditions, the model is able to utilize additional supervision based on diffuse and specular color maps, in an attempt to improve the generalization ability of the model to images of diverse and complex materials and lighting combinations. The model is evaluated against baselines include InstantMesh, and additional ablation is provided on losses related to PBR, as well as robustness to materials, number of views and field of view."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see Weakness section for comments and questions."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"[1] The paper is able to improve on a line of work around LRMs, by incorporating PBR related insights into training, and showcases by incorporating ground truth and supervision from sampled complex materials and lighting conditions, the model better handles images of those conditions, and additionally gains improvement on geometry estimation due to the improved modeling of physics.\n\n[2] The speed up of PBR rendering with split sum approximation enables efficient on-the-fly view synthesis and ground truth generation in a large parameter space of materials and lighting. Mostly as a technical contribution, but it will enable efficient data generation and augmentation on-the-fly when modeling complex parameter space of PBR.\n\n[3] Extensive evaluation. The model is able to compare on standard benchmarks against baseline models in this task, but additionally provides extensive ablation study to justify the design choices by ablating the PBR-related losses, as well as robustness to number of input views and FOVs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"[1] Clarification. Several details need to be clarified to better understand the model and the training strategy.\n\n(a) What does the model estimate w.r.t. the PBR parameters? Does it only estimate albedo? \n\n(b) With sampled metallic, roughness and lighting envmaps, do we apply the metallic and roughness globally? If yes: 1)what happens if the original CAD model is already associated with spatially-varying (SV) BRDF maps? 2) And if applied globally, does this strategy diminish the model's generation ability towards real images with complex SV materials? 3) Given a good portion of Objaverse models are assigned with PBR materials, does it benefit the training to also predict ground truth BRDF (roughness, metallic) without manually sampling and enforcing global roughness and metallic?\n\n(c) Is the split-sum approximation only applied to synthesizing estimated image from estimated representations, or it is also used to render ground truth images? Are ground truth images rendered on-the-fly for each batch in training?\n\n[2] Writing. Language issues are abundant and need to be fixed for a polished version. Examples:\n\n(a)  L015: for what purposes?\n\n(b) L020: Need to introduce the full name of PBR before first use of the abbreviation.\n\n(c) L050, L235: Need to clarify 'dependence on images rendered under fixed and simple lighting conditions' of previous methods. Mostly previous methods use PBR materials and envmap base lighting similar to this paper, so it would be important to clarify this assertion.\n\n(d) L186: functionalities -> downstream applications of ...\n\n(e) L283: what is 'a richer set of equations'?\n\n[3] Additional evaluation results on images of complex lighting and materials. The paper is able to showcase the robustness towards complex lighting and materials in Fig. 9, however one scene is too few, and comparison with baselines on this setting is necessary to further justify the claim."}},"nonreaders":[],"tmdate":1731427578555,"tcdate":1730973205551,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2330/Reviewer_Y5in"],"signatures":["ICLR.cc/2025/Conference/Submission2330/Reviewer_Y5in"],"forum":"AkL2ID5rRV","number":4,"license":"CC BY 4.0","cdate":1730973205551,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2330/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427578555,"domain":"ICLR.cc/2025/Conference","replyto":"AkL2ID5rRV","id":"VLUvd9Fxkd","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["3D reconstruction","feed-fowared reconstruction model","photometric stereo"]},"supplementary_material":{"value":"/attachment/16ab34b9a316225c18059548105a47123caa94ce.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We propose PRM, a novel photometric stereo based large reconstruction model to reconstruct high-quality meshes with fine-grained local details.\nUnlike previous large reconstruction models that prepare images under fixed and simple lighting as both input and supervision, PRM renders photometric stereo images by varying materials and lighting for the purposes, which not only improves the precise local details by providing rich photometric cues but also increases the model’s robustness to variations in the appearance of input images. \nTo offer enhanced flexibility of images rendering, we incorporate a real-time rendering method and mesh rasterization for online images rendering.\nMoreover, in employing an explicit mesh as our 3D representation, PRM ensures the application of differentiable PBR, which supports the utilization of multiple photometric supervisions and better models the specular color for high-quality geometry optimization.\nOur PRM leverages  photometric stereo images to achieve high-quality reconstructions with fine-grained local details, even amidst sophisticated image appearances. Extensive experiments demonstrate that PRM significantly outperforms other models."},"_bibtex":{"value":"@misc{\nge2025prm,\ntitle={{PRM}:  Photometric Stereo based Large Reconstruction Model},\nauthor={Wenhang Ge and Jiantao Lin and Guibao Shen and Jiawei Feng and Tao Hu and Xinli Xu and Ying-Cong Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=AkL2ID5rRV}\n}"},"title":{"value":"PRM:  Photometric Stereo based Large Reconstruction Model"},"pdf":{"value":"/pdf/e237a847ff393de9618a8770734f680f8ebb6ee5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ge|prm_photometric_stereo_based_large_reconstruction_model"},"authorids":{"value":["~Wenhang_Ge1","~Jiantao_Lin2","~Guibao_Shen1","~Jiawei_Feng1","~Tao_Hu1","~Xinli_Xu1","~Ying-Cong_Chen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wenhang Ge","Jiantao Lin","Guibao Shen","Jiawei Feng","Tao Hu","Xinli Xu","Ying-Cong Chen"]}},"version":2},{"content":{"summary":{"value":"This work is sloving a interesting task, which infers the downstream analysis posterior using synthetic data. The work proved that the Bernstein-von Mises theroy applies, the method can converage to the ture posterio as the number of synthetic datasets. The experimental settings are under two examples, i.e. non-private univariate Gaussian\n56 mean estimation and differentially private Bayesian logistic regression."},"soundness":{"value":"2 fair"},"confidence":{"value":"1: Your assessment is an educated guess. The submission is not in your area or the submission was difficult to understand. Math/other details were not carefully checked."},"questions":{"value":"See weaknesses."},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"strengths":{"value":"1, The work is trying to solve an interesting task, which is infering the downstream analysis posterior using synthetic data.\n\n2, The paper is well-writen and presented. \n\n3, The code is provided. So it will be helpful for the following work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Since synthetic data is generated by models which are trained using real data. So why synthetic data can improve the consisten bayesian inference is not clear. I think the paper needs more discussion about differences bewteen the real data and synthetic data.\n\n2, The synthetic data is a big topic. In the work, for me, it is not clear which synthetic data methods are used and how the synthetic data method is trained using real data.\n\n3, The applications are missing. Is it possible to extend the proposed method or therory to some kind of real application."},"limitations":{"value":"See weaknesses."}},"nonreaders":[],"tmdate":1702411083797,"tcdate":1690525464083,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission6883/Reviewer_z4aQ"],"signatures":["NeurIPS.cc/2023/Conference/Submission6883/Reviewer_z4aQ"],"forum":"Y8Jfbqx0bA","number":5,"license":"CC BY 4.0","cdate":1690525464083,"mdate":1702411083797,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission6883/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"Y8Jfbqx0bA","id":"b6sjFDZd0a","forumContent":{"venue":{"value":"Submitted to NeurIPS 2023"},"keywords":{"value":["synthetic data","Bayesian inference","Bernstein-von Mises theorem","differential privacy"]},"supplementary_material":{"value":"/attachment/8c6d55f0c99bc25b1fca2e0eb327e00eaa0cf7fd.zip"},"_bibtex":{"value":"@misc{\nr{\\\"a}is{\\\"a}2023on,\ntitle={On Consistent Bayesian Inference from Synthetic Data},\nauthor={Ossi R{\\\"a}is{\\\"a} and Joonas J{\\\"a}lk{\\\"o} and Antti Honkela},\nyear={2023},\nurl={https://openreview.net/forum?id=Y8Jfbqx0bA}\n}"},"title":{"value":"On Consistent Bayesian Inference from Synthetic Data"},"paperhash":{"value":"räisä|on_consistent_bayesian_inference_from_synthetic_data"},"TLDR":{"value":"We prove that consistent Bayesian inference from synthetic data reusing existing samplers is possible with multiple large synthetic datasets."},"abstract":{"value":"Generating synthetic data, with or without differential privacy, has attracted significant attention as a potential solution to the dilemma between making data easily available, and the privacy of data subjects. Several works have shown that consistency of downstream analyses from synthetic data, including accurate uncertainty estimation, requires accounting for the synthetic data generation. There are very few methods of doing so, most of them for frequentist analysis. In this paper, we study how to perform consistent Bayesian inference from synthetic data. We prove that mixing posterior samples obtained separately from multiple large synthetic datasets converges to the posterior of the downstream analysis under standard regularity conditions when the analyst's model is compatible with the data provider's model. We show experimentally that this works in practice, unlocking consistent Bayesian inference from synthetic data while reusing existing downstream analysis methods.\n"},"pdf":{"value":"/pdf/e17083fa57e2bafbb1fbbe6d274f192daa73fdcd.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference/Rejected_Submission"},"authorids":{"value":["~Ossi_Räisä1","~Joonas_Jälkö1","~Antti_Honkela1"]},"authors":{"value":["Ossi Räisä","Joonas Jälkö","Antti Honkela"]}},"version":2},{"content":{"comment":{"value":"Thank you for your suggestions. We have added new experiments with 10-15 latent variables under both sparse and dense causal graphs, which is shown in the below Table 2. The results show that accuracy remains stable as the system size increases, while runtime scales approximately linearly with the number of causal edges $|E|$.  In addition to scaling the number of latent variables, our synthetic setup already includes heterogeneous observation conditions (VFO/LFO/LPO), multiple noise regimes (which is shown in Table 3 in my answer to question 3 below), and causal-graph perturbation experiments (parent removal, edge reversal, spurious edges, which is shown in Table 1 in my previous answer). These complementary synthetic tests further demonstrate that SVGD maintains stable inference behavior under broader forms of complexity beyond system size alone.\n\nTable 2. Synthetic experiments for scalability analysis. Mean $\\pm$ s.d. aggregated over all latent variables. Runtime is measured relative to the 3-variable VFO baseline.\n\n| Setting (N, graph)   | NRMSE       | MAPE (\\%)   | Runtime ($\\times$ baseline) |\n| -------------------- | ----------- | ---------- | -------------------- |\n| VFO (10, sparse) | 0.11 $\\pm$ 0.02 | 9.3 $\\pm$ 1.4  | 3.3 $\\times$                 |\n| LFO (10, sparse) | 0.14 $\\pm$ 0.02 | 11.2 $\\pm$ 1.5 | 3.6 $\\times$                 |\n| LPO (10, sparse) | 0.19 $\\pm$ 0.05 | 14.7 $\\pm$ 1.9 | 3.9 $\\times$                 |\n| VFO (10, dense)  | 0.12 $\\pm$ 0.02 | 10.1 $\\pm$ 1.6 | 5.6 $\\times$                 |\n| VFO (12, sparse) | 0.12 $\\pm$ 0.02 | 9.8 $\\pm$ 1.5  | 4.1 $\\times$                 |\n| LFO (12, sparse) | 0.15 $\\pm$ 0.04 | 11.9 $\\pm$ 1.7 | 4.4 $\\times$                 |\n| LPO (12, sparse) | 0.20 $\\pm$ 0.04 | 15.3 $\\pm$ 2.1 | 4.6 $\\times$                 |\n| VFO (12, dense)  | 0.13 $\\pm$ 0.02 | 10.6 $\\pm$ 1.6 | 6.3 $\\times$                 |\n| VFO (15, sparse) | 0.13 $\\pm$ 0.02 | 10.2 $\\pm$ 1.6 | 5.0 $\\times$                 |\n| LFO (15, sparse) | 0.16 $\\pm$ 0.03 | 12.6 $\\pm$ 1.8 | 5.4 $\\times$                 |\n| LPO (15, sparse) | 0.21 $\\pm$ 0.04 | 15.9 $\\pm$ 2.2 | 5.7 $\\times$                 |\n| VFO (15, dense)  | 0.14 $\\pm$ 0.02 | 11.0 $\\pm$ 1.6 | 9.8 $\\times$                 |"},"title":{"value":"Reply to weakness 2: \"Limited Synthetic Experiment Complexity\""}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1763689667553,"tcdate":1763689667553,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11905/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission11905/Authors"],"forum":"M4Z2A1jYpU","number":18,"license":"CC BY 4.0","cdate":1763689667553,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11905/-/Official_Comment"],"mdate":1763689667553,"domain":"ICLR.cc/2026/Conference","replyto":"O7UZLHjInQ","id":"KPBzrTXGBX","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Causal Score Conditioning","Variational causal inference","Probabilistic graphical models","Multi-resolution observations","Score-based diffusion models"]},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"abstract":{"value":"Complex causal systems with interdependent variables require inference from heterogeneous observations that vary in spatial resolution, temporal frequency, and noise characteristics due to data acquisition constraints. Existing multi-modal fusion approaches assume uniform data quality or complete observability -- assumptions often violated in real-world applications. Current methods face three limitations: they treat causally-related variables independently, failing to exploit causal relationships; they cannot integrate multi-resolution observations effectively; and they lack theoretical frameworks for cascaded approximation errors. We introduce the Score-based Variational Graphical Diffusion Model (SVGDM), which integrates score-based diffusion within causal graphical structures for inference under heterogeneous incomplete observations. SVGDM introduces causal score decomposition enabling information propagation across causally-connected variables while preserving original observation characteristics. Diffusion provides a natural way to model scale-dependent sensing noise, which is common in remote-sensing, climate, and physical measurement systems, while the causal graph encodes well-established mechanistic dependencies between latent processes. We provide theoretical analysis and demonstrate superior performance on both synthetic and real-world datasets compared to relevant baselines."},"_bibtex":{"value":"@inproceedings{\nli2026causal,\ntitle={Causal Score Conditioning for Multi-Resolution Latent Systems},\nauthor={Xuechun Li and Shan Gao and Susu Xu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=M4Z2A1jYpU}\n}"},"title":{"value":"Causal Score Conditioning for Multi-Resolution Latent Systems"},"pdf":{"value":"/pdf/1bf4acd35b0ab017ed8bc36988184f8eb65769e0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|causal_score_conditioning_for_multiresolution_latent_systems"},"authorids":{"value":["~Xuechun_Li2","~Shan_Gao6","~Susu_Xu2"]},"authors":{"value":["Xuechun Li","Shan Gao","Susu Xu"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on the topic of synthetic network generation, aiming to address a key challenge: how to generate graphs that preserve the structural properties of a real-world network while remaining computationally efficient. To tackle this, the authors propose SyNGLER, a two-stage framework. First, a latent space network model is used to learn low-dimensional node embeddings, then, a distribution-free generator is trained in this latent space to produce new node embeddings, from which synthetic graphs are generated. The paper provides theoretical guarantees via KL divergence decomposition and consistency analysis, and empirically shows on synthetic and real datasets that SyNGLER better preserves structural statistics while being more computationally efficient than existing methods."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I was curious to see what would be the practical impact of the SyNGLER generated graphs, and how to validate these practical impacts. \n\n1. For the practical impact, the method only focuses on generating network structure (adjacency), without modeling node or edge attributes. In most of the real-world applications, graphs are attributed, and structure alone is insufficient. Ignoring attributes may lead to significant information loss, and it is unclear how the generated synthetic graphs could be used for tasks that rely on feature–structure interactions (e.g., GNN training, attributed diffusion). I wonder if the authors has any empirical or insights on how SyNGLER could be extended to attributed graphs?\n\n2. For validating the practical impact, the paper demonstrates that the generated networks match several structural statistics of the input graph. However, structural similarity alone does not guarantee that synthetic networks can serve as valid surrogates for scientific analysis or downstream decision-making. How can we validate functional equivalence of the input graph and the synthetic graph?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses an important problem — generating synthetic graphs that preserve structural properties while being computationally efficient.\n2. Provides theoretical analysis.\n3. The experiments cover both synthetic and multiple real-world networks (YouTube, DBLP, Yelp, PolBlogs), and evaluate diverse structural metrics (degree, clustering, eigenvalues, triangle density)"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While synthetic graph generation is indeed a meaningful and important research direction, this paper focuses only on generating non-attributed graphs that replicate structural statistics (e.g., degree, clustering, spectrum) of a real network. First of all, in modern applications, most graphs are attributed, and the interaction between attributes and structure is often essential to the task, how could the proposed SyNGLER benefit such graphs remain undiscussed. Moreover, the paper does not demonstrate or validate any downstream usage, so it remains unclear what practical purpose these synthetic structure-only graphs serve beyond matching statistical properties.\n\n2. Would be better if code could be provided."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933915537,"tcdate":1761958400976,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20479/Reviewer_DqnC"],"signatures":["ICLR.cc/2026/Conference/Submission20479/Reviewer_DqnC"],"forum":"JtL7kCe32S","number":3,"license":"CC BY 4.0","cdate":1761958400976,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20479/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933915537,"domain":"ICLR.cc/2026/Conference","replyto":"JtL7kCe32S","id":"AJDV1U3Ze5","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["social networks","latent space models","structured data generation"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Network data are ubiquitous across the social sciences, biology, and information systems. Generating realistic synthetic network data has broad applications, from network simulation to scientific  discovery. However, many existing black-box approaches for network generation tend to overfit observed data while overlooking characteristic network structure, and incur substantial computational overhead at scale. These practical challenges call for synthetic network generation methods that are both efficient and capable of capturing structural properties of networks. In this paper, we introduce Synthetic Network Generation via Latent Embedding Reconstruction (SyNGLER), a general and efficient framework for synthetic network generation that builds on latent space network models. Given an observed network, SyNGLER first learns low-dimensional latent node embeddings via a latent space network model and then reconstructs the latent space by building a distribution-free generator over these embeddings. For generation, SyNGLER first samples (or resamples) node embeddings from the generator in the latent space and then produces synthetic networks using the latent space network model. Through the latent space framework, SyNGLER preserves unique characteristics in networks such as sparsity and node degree heterogeneity, while allowing for efficient training with lower computational cost than many existing deep architectures. We provide theoretical guarantees by developing consistency results regarding the distance between the true and synthetic edge distributions. Empirical studies further demonstrate the effectiveness of SyNGLER, where SyNGLER efficiently produces networks that better preserve key network characteristics such as network moments and degree distributions compared with existing approaches."},"_bibtex":{"value":"@misc{\njiang2026efficient,\ntitle={Efficient Synthetic Network Generation via Latent Embedding Reconstruction},\nauthor={Feifan Jiang and Yinan Bu and Shihao Wu and Gongjun Xu and Ji Zhu},\nyear={2026},\nurl={https://openreview.net/forum?id=JtL7kCe32S}\n}"},"title":{"value":"Efficient Synthetic Network Generation via Latent Embedding Reconstruction"},"pdf":{"value":"/pdf/91f16e8593fae2323473ee36ac686506349017d2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"jiang|efficient_synthetic_network_generation_via_latent_embedding_reconstruction"},"authorids":{"value":["~Feifan_Jiang1","~Yinan_Bu1","~Shihao_Wu3","~Gongjun_Xu1","~Ji_Zhu2"]},"authors":{"value":["Feifan Jiang","Yinan Bu","Shihao Wu","Gongjun Xu","Ji Zhu"]}},"version":2},{"content":{"venue":{"value":"SynS & ML @ ICML2023"},"pdf":{"value":"/pdf/e8f3b5c8b5f4934eabf7a60e14d90399b3c6723b.pdf"},"keywords":{"value":["Microstructure Dataset of Titanium Alloy","Synthetic Microstructure Generation","Physics-Based Generative Models"]},"venueid":{"value":"ICML.cc/2023/Workshop/SynS_and_ML"},"paperhash":{"value":"jangid|titanium_3d_microstructure_for_physicsbased_generative_models_a_dataset_and_primer"},"authorids":{"value":["~Devendra_Kumar_Jangid1","nbrodnik@ucsb.edu","mechlin@ucsb.edu","~Samantha_Daly1","~Tresa_Pollock1","~B.S._Manjunath1"]},"abstract":{"value":"When engineers design components, they rely on accurate property descriptions of the materials being used to predict performance. Most materials used for engineering applications are composed of an arrangement of atomic constituents into crystalline phases, which control the properties of that material. The crystal orientations embedded in this microstructural information differ from the information in conventional light optical images, and are critical for developing and designing materials for a range of applications. However, collecting microstructure information through experimental methods is expensive and time-consuming, especially when 3D information is needed. In order to model material properties under different material processing conditions (resulting in different microstructural arrangements), physics-based generative models are needed to create realistic synthetic microstructures. This research releases microstructural data of a titanium alloy, Ti-6Al-4V, and discusses their information modalities and the physics needed to be incorporated to enable the design of physics-based generative models for generating synthetic microstructures."},"_bibtex":{"value":"@inproceedings{\njangid2023titanium,\ntitle={Titanium 3D Microstructure for Physics-based Generative Models: A Dataset and Primer},\nauthor={Devendra Kumar Jangid and Neal R Brodnik and McLean P Echlin and Samantha Daly and Tresa Pollock and B.S. Manjunath},\nbooktitle={1st Workshop on the Synergy of Scientific and Machine Learning Modeling @ ICML2023},\nyear={2023},\nurl={https://openreview.net/forum?id=uFSNgD9Msk}\n}"},"title":{"value":"Titanium 3D Microstructure for Physics-based Generative Models: A Dataset and Primer"},"authors":{"value":["Devendra Kumar Jangid","Neal R Brodnik","McLean P Echlin","Samantha Daly","Tresa Pollock","B.S. Manjunath"]}},"tmdate":1690516851420,"pdate":1690516851407,"tcdate":1685296703158,"writers":["ICML.cc/2023/Workshop/SynS_and_ML","ICML.cc/2023/Workshop/SynS_and_ML/Submission61/Authors"],"signatures":["ICML.cc/2023/Workshop/SynS_and_ML/Submission61/Authors"],"forum":"uFSNgD9Msk","number":61,"cdate":1685296703158,"mdate":1690516851420,"readers":["everyone"],"invitations":["ICML.cc/2023/Workshop/SynS_and_ML/-/Submission","ICML.cc/2023/Workshop/SynS_and_ML/-/Post_Submission","ICML.cc/2023/Workshop/SynS_and_ML/Submission61/-/Camera_Ready","ICML.cc/2023/Workshop/SynS_and_ML/-/Edit"],"odate":1690516851407,"domain":"ICML.cc/2023/Workshop/SynS_and_ML","id":"uFSNgD9Msk","version":2},{"content":{"summary":{"value":"The paper \"ToolACE: Enhancing Function Calling with Accuracy, Complexity, and Diversity\" presents a novel data generation pipeline for function-calling tasks LLMs. The approach leverages a tool self-evolution synthesis module, a self-guided dialog generation module, and a dual-layer verification module to create accurate, complex, and diverse tool-calling scenarios. ToolACE aims to improve LLMs' zero-shot function-calling capabilities by generating comprehensive training data that is validated through rule-based and model-based checks. The experiments show promising results, particularly with the ToolACE-8B model, which outperforms several existing LLMs."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please refer to weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The introduction of ToolACE's multi-step data generation, including evolutionary diversity and self-guided complexity, provides an innovative solution for generating complex and diverse function-calling data.\n\n- The DLV system, combining rule-based and model-based checks, enhances the reliability of the generated data. This is a strong point, as it helps maintain data quality, which is critical for training LLMs effectively.\n\n- The paper provides an extensive set of experiments, including comparisons with state-of-the-art models and an ablation study to assess the contribution of different components like accuracy, complexity, and diversity in the dataset. These experiments illustrate the potential benefits of the proposed pipeline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The evaluation scenarios are limited to synthetic function-calling tasks and benchmarks like BFCL and APIBank. The paper would benefit from more realistic evaluations or applications in real-world tool usage scenarios. This would better demonstrate ToolACE’s utility beyond controlled benchmark settings.\n\n- The self-guided dialog generation process heavily relies on the LLM being trained to evaluate the complexity of generated data. This creates a circular dependency where the model is used both as a learner and an evaluator, which may introduce bias in the complexity estimation. More external validation or use of independent evaluators would make the results more robust.\n\n- The use of complexity-based sampling to dynamically adjust dialog difficulty has merit but may lead to unintended biases, as data that is either too simple or too complex is filtered out. The approach may fail to fully explore the impact of diverse and extreme cases, leading to gaps in the model’s capabilities in certain contexts.\n\n- While the paper compares ToolACE to several other function-calling models, the comparison is often superficial. The benefits of using ToolACE versus simpler data augmentation techniques are not well articulated, and it is unclear how much of the improvement can be attributed to the synthesis method versus the increased volume of data.\n\n- The paper claims that ToolACE-8B is competitive with GPT-4 series models. However, it does not fully address the limitations of ToolACE-8B in terms of generalization and applicability to a broader range of tasks beyond function calling. A more detailed discussion of these limitations would provide a more balanced perspective.\n\n- The font size in Figures is too small, which is unclear for readers."}},"nonreaders":[],"tmdate":1732514775810,"tcdate":1730185429328,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2471/Reviewer_ViF9"],"signatures":["ICLR.cc/2025/Conference/Submission2471/Reviewer_ViF9"],"forum":"8EB8k6DdCU","number":2,"license":"CC BY 4.0","cdate":1730185429328,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2471/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732514775810,"domain":"ICLR.cc/2025/Conference","replyto":"8EB8k6DdCU","id":"g8cgECC9aZ","forumContent":{"TLDR":{"value":"Improving the function calling capability of large language models with data accuracy, diversity, and complexity."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Tool leaning","Function calling","Large language models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Function calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pipelines often lack coverage and accuracy. In this paper, we present ToolACE, an automatic agentic pipeline designed to generate accurate, complex, and diverse tool-learning data, specifically tailored to the capabilities of LLMs. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, under the guidance of a complexity evaluator. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. We demonstrate that models trained on our synthesized data---even with only 8B parameters---achieve state-of-the-art performance, comparable to the latest GPT-4 models. Our model and a subset of the data are publicly available at https://huggingface.co/Team-ACE."},"_bibtex":{"value":"@inproceedings{\nliu2025toolace,\ntitle={Tool{ACE}: Winning the Points of {LLM} Function Calling},\nauthor={Weiwen Liu and Xu Huang and Xingshan Zeng and xinlong hao and Shuai Yu and Dexun Li and Shuai Wang and Weinan Gan and Zhengying Liu and Yuanqing Yu and Zezhong WANG and Yuxian Wang and Wu Ning and Yutai Hou and Bin Wang and Chuhan Wu and Wang Xinzhi and Yong Liu and Yasheng Wang and Duyu Tang and Dandan Tu and Lifeng Shang and Xin Jiang and Ruiming Tang and Defu Lian and Qun Liu and Enhong Chen},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=8EB8k6DdCU}\n}"},"title":{"value":"ToolACE: Winning the Points of LLM Function Calling"},"pdf":{"value":"/pdf/9c7c53cc6199348d235063c044442216d84429c4.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"liu|toolace_winning_the_points_of_llm_function_calling"},"authorids":{"value":["~Weiwen_Liu1","~Xu_Huang2","~Xingshan_Zeng1","~xinlong_hao1","~Shuai_Yu3","~Dexun_Li1","~Shuai_Wang34","~Weinan_Gan1","~Zhengying_Liu2","~Yuanqing_Yu1","~Zezhong_WANG1","~Yuxian_Wang1","~Wu_Ning1","~Yutai_Hou1","~Bin_Wang12","~Chuhan_Wu2","~Wang_Xinzhi1","~Yong_Liu14","~Yasheng_Wang1","~Duyu_Tang3","~Dandan_Tu1","~Lifeng_Shang1","~Xin_Jiang1","~Ruiming_Tang2","~Defu_Lian1","~Qun_Liu1","~Enhong_Chen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Weiwen Liu","Xu Huang","Xingshan Zeng","xinlong hao","Shuai Yu","Dexun Li","Shuai Wang","Weinan Gan","Zhengying Liu","Yuanqing Yu","Zezhong WANG","Yuxian Wang","Wu Ning","Yutai Hou","Bin Wang","Chuhan Wu","Wang Xinzhi","Yong Liu","Yasheng Wang","Duyu Tang","Dandan Tu","Lifeng Shang","Xin Jiang","Ruiming Tang","Defu Lian","Qun Liu","Enhong Chen"]}},"version":2},{"content":{"summary":{"value":"This paper introduces APEX, a framework for enhancing LLM’s physical reasoning capabilities for real-time task planning by integrating with simulations of physical engines. Given some prompt, the system is able to generate qualitative and quantitative predictions of the scene in text. Specifically, given two frames, APEX first creates a scene graph that captures object relations and interactions, then passes the current relational state to a physics engine to simulate possible futures under different actions. The resulting information is incorporated into a unified prompt for the LLM, which then selects the optimal action.\n\nThe paper introduces three benchmarks and evaluates APEX-enhanced LLMs with vanilla LLMs, and shows that APEX-enhanced LLMs outperforms vanilla LLMs on physics reasoning, Tetris, and obstacle avoidance."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"- What are the assumptions on the scene, dynamics, and action space?\n- The method only uses $G_t$ and $G_{t + \\Delta t}$ as context. How does this capture second-order dynamics information such as acceleration? If such information must be provided through text, this imposes an additional requirement on the system that these values be known, yet information such as object velocity or acceleration is typically unavailable in real-world decision-making.\n- Why use scene graphs? Although Appendix 6.7.2 provides some motivation in terms of interaction and temporal saliency filtering, these considerations mainly arise from converting everything into text. Why is this representation preferable to, for example, a latent world model, where such filtering would not be necessary?\n- In the Tetris experiment, only APEX and vanilla LLMs are evaluated. What would the oracle performance be, e.g. from a classical planning algorithm or human players?\n- While the paper provides some prompt and task examples in the appendix, what would a complete input–output pipeline look like for a specific task (e.g., Tetris)? What are the exact inputs and outputs of the Graphormer encoder and the physics engine? How are the entities represented and instantiated within the physics engine?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- The paper attempts to address the physical reasoning limitations of LLMs by making use of physics engines.\n- The paper provides some concrete examples of tasks and model outputs in the appendix."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Overall, the introduced pipeline is very limiting. The framework makes many assumptions, but none are explicitly described in the paper. For example, the scene graph formulation assumes the scene contains mostly distinctive and rigid objects. Since the predictions are simulated with a physical engine, the method also assumes those objects are compatible with the simulator. What constraints are imposed on the acceptable objects in the scene? Moreover, the decision-making procedure requires a finite, enumerable set of actions, as the algorithm simulates each action one by one. This is largely infeasible in real decision-making tasks.\n- The paper lacks baselines. It motivates the work by describing limitations of latent world models, RL, etc., but only compares APEX-enhanced LLMs with vanilla LLMs. What about latent world models, RL, or other methods as baselines? If they are not compatible, why motivate the paper based on their limitations?\n- There is no description of how the physical engine is designed or used. How is the simulated scene created based on the current relation state? The paper only mentions MuJoCo in Figure 2, but what are the assumptions on the scene to make it simulatable by MuJoCo?\n- The proposed pipeline appears to bottleneck information through text. If the system already has access to input frames, created scene graphs, and a physics engine, why is text used as the medium for reasoning and decision-making? Wouldn’t this cause loss of critical information, and isn’t it the case that not all scenes can be described accurately through text?\n- The paper’s writing could be improved. It presents irrelevant information while missing some critical details about the method. E.g. Some works in the related work section appear misplaced,  for instance, it is unclear why RL is discussed, or why several non-RL papers are included in Section 2.3. Similarly, it is not clear why R3M is categorized under world models in Section 2.2. \n- Some of the task figures should be moved to the main body. Currently, it is difficult to ground the evaluations, even though the paper provides some visuals in the appendix."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924233554,"tcdate":1762030903400,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13665/Reviewer_Z9SE"],"signatures":["ICLR.cc/2026/Conference/Submission13665/Reviewer_Z9SE"],"forum":"ROB3ALLKIX","number":3,"license":"CC BY 4.0","cdate":1762030903400,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13665/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924233554,"domain":"ICLR.cc/2026/Conference","replyto":"ROB3ALLKIX","id":"odPEfkNmwU","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We build a framework that helps LLMs quantify when a cat will collide with them and assess the physical outcomes of different escape routes (e.g., whether they’ll crash into a table) by feeding them physics-based simulations."},"keywords":{"value":["Physics-Enhanced LLMs","Graph-Based Perception","Task Planning","Predictive Simulation"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Large Language Models (LLMs) demonstrate strong reasoning and task planning capabilities but remain fundamentally limited in physical interaction modeling. Existing approaches integrate perception via Vision-Language Models (VLMs) or adaptive decision-making through Reinforcement Learning (RL), but they fail to capture dynamic object interactions or require task-specific training, limiting their real-world applicability.\nWe introduce APEX (Anticipatory Physics-Enhanced Execution), a framework that equips LLMs with physics-driven foresight for real-time task planning. APEX constructs structured graphs to identify and model the most relevant dynamic interactions in the environment, providing LLMs with explicit physical state updates. Simultaneously, APEX provides low-latency forward simulations of physically feasible actions, allowing LLMs to select optimal strategies based on predictive outcomes rather than static observations.\nWe evaluate APEX on three benchmarks designed to assess perception, prediction, and decision-making: (1) Physics Reasoning Benchmark, testing causal inference and object motion prediction; (2) Tetris, evaluating whether physics-informed prediction enhances decision-making performance in long-horizon planning tasks; (3) Dynamic Obstacle Avoidance, assessing the immediate integration of perception and action feasibility analysis. APEX significantly outperforms standard LLMs and VLM-based models, demonstrating the necessity of explicit physics reasoning for bridging the gap between language-based intelligence and real-world task execution."},"_bibtex":{"value":"@misc{\nhuang2025apex,\ntitle={{APEX}: Empowering {LLM}s with Physics-Based Task Planning for Real-time Insight},\nauthor={Wanjing Huang and Weixiang Yan and Zhen Zhang and Ambuj Singh},\nyear={2025},\nurl={https://openreview.net/forum?id=ROB3ALLKIX}\n}"},"title":{"value":"APEX: Empowering LLMs with Physics-Based Task Planning for Real-time Insight"},"pdf":{"value":"/pdf/0529797af939423f17dcaccea10909babb2a5c4d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|apex_empowering_llms_with_physicsbased_task_planning_for_realtime_insight"},"authorids":{"value":["~Wanjing_Huang1","~Weixiang_Yan1","~Zhen_Zhang16","~Ambuj_Singh1"]},"authors":{"value":["Wanjing Huang","Weixiang Yan","Zhen Zhang","Ambuj Singh"]}},"version":2},{"content":{"venue":{"value":"COMPLEX NETWORKS (1) 2020"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-030-65347-7_21.pdf"},"venueid":{"value":"dblp.org/conf/COMPLEXNETWORKS/2020"},"paperhash":{"value":"bloemheuvel|graph_signal_processing_on_complex_networks_for_structural_health_monitoring"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Stefan_Bloemheuvel:","https://dblp.org/search/pid/api?q=author:Jurgen_van_den_Hoogen:","~Martin_Atzmueller1"]},"html":{"value":"https://doi.org/10.1007/978-3-030-65347-7_21"},"_bibtex":{"value":"@inproceedings{DBLP:conf/complexnetworks/BloemheuvelHA20,\n  author={Stefan Bloemheuvel and Jurgen van den Hoogen and Martin Atzmueller},\n  title={Graph Signal Processing on Complex Networks for Structural Health Monitoring},\n  year={2020},\n  cdate={1577836800000},\n  pages={249-261},\n  url={https://doi.org/10.1007/978-3-030-65347-7_21},\n  booktitle={COMPLEX NETWORKS (1)},\n  crossref={conf/complexnetworks/2020-1}\n}\n"},"abstract":{"value":"In this work, we demonstrate the application of a framework targeting Complex Networks and Graph Signal Processing (GSP) for Structural Health Monitoring (SHM). By modeling and analyzing a large bridge equipped with strain and vibration sensors, we show that GSP is capable of selecting the most important sensors, investigating different optimization techniques for selection. Furthermore, GSP enables the detection of graph signal patterns (mode shapes), grasping the physical function of the sensors in the network. Our results indicate the efficacy of GSP on complex sensor data modeled in complex networks."},"title":{"value":"Graph Signal Processing on Complex Networks for Structural Health Monitoring"},"authors":{"value":["Stefan Bloemheuvel","Jurgen van den Hoogen","Martin Atzmueller"]}},"tmdate":1747311227865,"pdate":1577836800000,"tcdate":1747310938147,"writers":["~"],"signatures":["~Martin_Atzmueller1"],"forum":"lLFU7ncmTM","license":"CC BY-SA 4.0","number":485244,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747311227865,"domain":"DBLP.org","id":"lLFU7ncmTM","version":2},{"content":{"summary":{"value":"The paper formalizes a new inverse problem termed the Ensemble Inverse Problem (EIP), where an additional set of observations is incorporated into the inversion process. The authors propose to address EIP using conditional generative models, specifically diffusion and flow-matching–based approaches (EI-DDPM and EI-FM). The model is evaluated on three tasks—synthetic 2D Gaussian problem, particle-physics data unfolding, and MNIST image inversion—and demonstrates superior performance."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Can authors provide more concrete examples of how EIP corresponds to a real imaging scenarios (e.g. MRI[1], deblurring[2], etc) as well as experimental evidence? \n- Why do the authors tailor the algorithms specifically for DDPM and FM? It's known that diffusion model is equivalent to FM up to a simple reparameterization for Gaussian prior setting [3]. Are there specific empirical motivations for retaining both?\n\n[1]: Sriram, Anuroop, et al. \"End-to-end variational networks for accelerated MRI reconstruction.\" _International conference on medical image computing and computer-assisted intervention_. Cham: Springer International Publishing, 2020.\n\n[2]: Mardani, Morteza, et al. \"A variational perspective on solving inverse problems with diffusion models.\" _arXiv preprint arXiv:2305.04391_ (2023).\n\n[3]: [Diffusion Models and Gaussian Flow Matching: Two Sides of the Same Coin](https://d2jud02ci9yv69.cloudfront.net/2025-04-28-diffusion-flow-173/blog/diffusion-flow/)"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The paper provides a clean formalization of the ensemble inverse problem, emphasizing inference across multiple priors with a shared but unknown forward operator."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While the EIP formulation is interesting conceptually, its practical relevance to real-world inverse problems is not clear. The authors mention applications in high-energy physics and inverse imaging, but for readers unfamiliar with high-energy physics, the motivation in that domain is difficult to assess. For inverse imaging problems, it is unclear to me how the EIP problem setting arises in practice. \n- Despite the new terminology, Algorithms 1 and 2 follow standard conditional diffusion training and sampling procedures. The only notable modification is the addition of a Set Transformer encoder for $\\mathcal{Y}$, and even this encoder is used only in one experiment (the particle-physics data unfolding). \n- The \"high-dimensional\" inverse imaging experiment only uses MNIST (28x28 dimensions), which is far too simple to demonstrate the proposed model’s utility for serious inverse imaging problems. \n- The proposed method requires retraining for each new forward model, as it depends on paired truth–observation datasets. In many practical settings, generating such datasets or retraining large diffusion models may be infeasible."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931318626,"tcdate":1762131433415,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19400/Reviewer_RwgK"],"signatures":["ICLR.cc/2026/Conference/Submission19400/Reviewer_RwgK"],"forum":"aVXXZAp41g","number":3,"license":"CC BY 4.0","cdate":1762131433415,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19400/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931318626,"domain":"ICLR.cc/2026/Conference","replyto":"aVXXZAp41g","id":"nYbpJcoxp0","forumContent":{"TLDR":{"value":"We introduce the ensemble inverse problem and propose a posterior sampling method based on generative models to solve it."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Inverse problems","conditional generative models","posterior sampling","permutation invariant neural network"]},"supplementary_material":{"value":"/attachment/c25721b5053dc87abe458c387415015bf452f63c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We introduce a new multivariate statistical problem that we refer to as the Ensemble Inverse Problem (EIP). The aim of EIP is to invert for an ensemble that is distributed according to the pushforward of a prior under a forward process. In high energy physics (HEP), this is related to a widely known problem called unfolding, which aims to reconstruct the true physics distribution of quantities, such as momentum and angle, from measurements that are distorted by detector effects. In recent applications, the EIP also arises in inverse imaging with unknown priors. We propose non-iterative inference-time methods that construct posterior samplers based on a new class of conditional generative models, which we call  ensemble inverse generative models. For the posterior modeling, these models additionally use the ensemble information contained in the observation set on top of single measurements.  Unlike existing methods, our proposed methods avoid explicit and iterative use of the forward operator at inference time via training across several sets of truth-observation pairs that are consistent with the same forward operator, but originate from a wide range of priors. We demonstrate that this training procedure implicitly encodes the likelihood model. The use of ensemble information helps posterior inference and enables generalization to unseen priors. We benchmark the proposed method on several synthetic and real datasets in HEP and inverse imaging."},"_bibtex":{"value":"@misc{\nhuan2026the,\ntitle={The Ensemble Inverse Problem: Applications and Methods},\nauthor={Zhengyan Huan and Camila Pazos and Martin Klassen and Vincent Croft and Pierre-Hugues Beauchemin and Shuchin Aeron},\nyear={2026},\nurl={https://openreview.net/forum?id=aVXXZAp41g}\n}"},"title":{"value":"The Ensemble Inverse Problem: Applications and Methods"},"pdf":{"value":"/pdf/7efc21d8ad2fc1796d9ea8a75ce49b47e338d44b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"huan|the_ensemble_inverse_problem_applications_and_methods"},"authorids":{"value":["~Zhengyan_Huan2","~Camila_Pazos1","~Martin_Klassen1","~Vincent_Croft1","~Pierre-Hugues_Beauchemin1","~Shuchin_Aeron2"]},"authors":{"value":["Zhengyan Huan","Camila Pazos","Martin Klassen","Vincent Croft","Pierre-Hugues Beauchemin","Shuchin Aeron"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2402.00326v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"wang|piratenets_physicsinformed_deep_learning_with_residual_adaptive_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Sifan_Wang:","https://dblp.org/search/pid/api?q=author:Bowen_Li:","https://dblp.org/search/pid/api?q=author:Yuhan_Chen:","~Paris_Perdikaris1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2402.00326"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2402-00326,\n  publtype={informal},\n  author={Sifan Wang and Bowen Li and Yuhan Chen and Paris Perdikaris},\n  title={PirateNets: Physics-informed Deep Learning with Residual Adaptive Networks},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2402.00326},\n  url={https://doi.org/10.48550/arXiv.2402.00326}\n}\n"},"abstract":{"value":"While physics-informed neural networks (PINNs) have become a popular deep learning framework for tackling forward and inverse problems governed by partial differential equations (PDEs), their performance is known to degrade when larger and deeper neural network architectures are employed. Our study identifies that the root of this counter-intuitive behavior lies in the use of multi-layer perceptron (MLP) architectures with non-suitable initialization schemes, which result in poor trainablity for the network derivatives, and ultimately lead to an unstable minimization of the PDE residual loss. To address this, we introduce Physics-informed Residual Adaptive Networks (PirateNets), a novel architecture that is designed to facilitate stable and efficient training of deep PINN models. PirateNets leverage a novel adaptive residual connection, which allows the networks to be initialized as shallow networks that progressively deepen during training. We also show that the proposed initialization scheme allows us to encode appropriate inductive biases corresponding to a given PDE system into the network architecture. We provide comprehensive empirical evidence showing that PirateNets are easier to optimize and can gain accuracy from considerably increased depth, ultimately achieving state-of-the-art results across various benchmarks. All code and data accompanying this manuscript will be made publicly available at \\url{https://github.com/PredictiveIntelligenceLab/jaxpi}."},"title":{"value":"PirateNets: Physics-informed Deep Learning with Residual Adaptive Networks"},"authors":{"value":["Sifan Wang","Bowen Li","Yuhan Chen","Paris Perdikaris"]}},"tmdate":1727708546245,"pdate":1704067200000,"tcdate":1727708537971,"writers":["~"],"signatures":["~Paris_Perdikaris1"],"forum":"EO9wyFsXGh","license":"CC BY-SA 4.0","number":119615,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727708546245,"domain":"DBLP.org","id":"EO9wyFsXGh","version":2},{"content":{"venue":{"value":"AI4Math@ICML25 Poster"},"TLDR":{"value":"SymbolicVision"},"pdf":{"value":"/pdf/474e3a5131889d684306f84f56c2bee9faadd22f.pdf"},"keywords":{"value":["Symbolic Regression","Remote Sensing","Vision Transformer","Physics-aware","Multi-spectral Imagery"]},"venueid":{"value":"ICML.cc/2025/Workshop/AI4MATH"},"paperhash":{"value":"yu|physicsconstrained_symbolic_regression_from_imagery"},"authorids":{"value":["~Zhenyu_Yu1","~MOHD_YAMANI_IDNA_IDRIS1","~Pei_Wang7"]},"abstract":{"value":"We propose *SymbolicVision*, a physics-constrained symbolic regression framework that derives interpretable mathematical expressions directly from multi-spectral remote sensing imagery. Unlike black-box deep models, *SymbolicVision* combines a Vision-based image encoder with a Transformer-based symbolic decoder to enable cross-modal learning between visual features and symbolic formulas. A hybrid loss design ensures both numerical accuracy and physical plausibility. Evaluated on symbolic benchmarks (SRBench) and real satellite datasets (Open-Canopy), *SymbolicVision* achieves high predictive accuracy (\\$R^2>0.99$\\), and robust performance on geospatial tasks. This work highlights the potential of interpretable, physics-aware models for scientific remote sensing."},"_bibtex":{"value":"@inproceedings{\nyu2025physicsconstrained,\ntitle={Physics-Constrained Symbolic Regression from Imagery},\nauthor={Zhenyu Yu and MOHD YAMANI IDNA IDRIS and Pei Wang},\nbooktitle={2nd AI for Math Workshop @ ICML 2025},\nyear={2025},\nurl={https://openreview.net/forum?id=3qcKSA7ibv}\n}"},"title":{"value":"Physics-Constrained Symbolic Regression from Imagery"},"authors":{"value":["Zhenyu Yu","MOHD YAMANI IDNA IDRIS","Pei Wang"]}},"tmdate":1753454773865,"pdate":1752039856551,"tcdate":1750401047599,"writers":["ICML.cc/2025/Workshop/AI4MATH","ICML.cc/2025/Workshop/AI4MATH/Submission92/Authors"],"signatures":["ICML.cc/2025/Workshop/AI4MATH/Submission92/Authors"],"forum":"3qcKSA7ibv","license":"CC BY-NC-SA 4.0","number":92,"cdate":1750401047599,"readers":["everyone"],"invitations":["ICML.cc/2025/Workshop/AI4MATH/-/Submission","ICML.cc/2025/Workshop/AI4MATH/-/Post_Submission","ICML.cc/2025/Workshop/AI4MATH/-/Edit","ICML.cc/2025/Workshop/AI4MATH/Submission92/-/Revision"],"mdate":1753454773865,"odate":1752626829243,"domain":"ICML.cc/2025/Workshop/AI4MATH","id":"3qcKSA7ibv","version":2},{"content":{"summary":{"value":"This paper explores using synthetic data to improve visual grounding in vision-and-language models. The authors present SynGround, a pipeline that generates synthetic image-text-box triplets by combining advances in text-to-image generation, language models, and object detection. They compare synthetic data with real and web-crawled data on RefCOCO+ and Flickr30k benchmarks. Results show SynGround enhances localization in ALBEF and BLIP models, outperforming web-crawled data and offering potential for infinite data generation."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Have you considered evaluating SynGround with more recent and state-of-the-art visual grounding models?\n\nCould you elaborate on the computational resources required for generating and utilizing the synthetic data, especially in the context of scaling up to larger datasets?\n\nHave you observed any limitations or saturation points when increasing the scale of synthetic data used for training?\n\nCould you discuss the potential impact of biases present in the source data (e.g., caption descriptions) on the generated synthetic data and downstream visual grounding performance?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Systematic exploration: The paper systematically explores different strategies for generating synthetic image-text and image-text-box data, providing valuable insights into the factors influencing performance. The paper compares the performance of models trained on synthetic data with models trained on real and web-crawled data.\n\n2. Pipeline for synthetic data generation: The proposed SynGround pipeline offers a structured approach for creating synthetic data for visual grounding, combining several advanced techniques.\n\n3. Outperforming web-crawled data: The finding that synthetic data outperforms web-crawled data is a notable strength, suggesting the potential for creating more tailored and effective training datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Use of older models: The paper relies on ALBEF and BLIP, which are relatively older models in the rapidly evolving field of vision and language. The performance in Experiment 1 does not compare to any of the models in the papersincode leaderboard (e.g., https://paperswithcode.com/sota/referring-expression-comprehension-on-refcoco-1). Evaluating SynGround with more recent and state-of-the-art models would significantly strengthen the claims.\n\n2. Limited performance gains:  While improvements are reported, the absolute gains from using synthetic data, especially when combined with real data, are relatively modest and may not be statistically significant.  Error bars or further statistical analysis should be provided to support the claims of improvement.\n\n3. Clarity and organization: The presentation of experiments could be improved.  The motivation and reasoning behind each experiment could be more clearly articulated.  Consolidating related experiments (like the BLIP experiments) into fewer tables would enhance readability.  The paper would benefit from focusing on the key findings, such as the comparison with web-crawled data, earlier in the presentation.\n\n4. Lack of analysis on scaling limitations: While the paper mentions the potential for infinite data generation, it does not discuss or analyze potential limitations or saturation points in scaling up the use of synthetic data."}},"nonreaders":[],"tmdate":1731427464044,"tcdate":1729271887952,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1676/Reviewer_4VzW"],"signatures":["ICLR.cc/2025/Conference/Submission1676/Reviewer_4VzW"],"forum":"EuoHhIqvRD","number":1,"license":"CC BY 4.0","cdate":1729271887952,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1676/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427464044,"domain":"ICLR.cc/2025/Conference","replyto":"EuoHhIqvRD","id":"9RLsH48SiY","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Visual Grounding","Referring Expression Comprehension","Learning from Models","Synthetic Data"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper extensively investigates the effectiveness of synthetic training data to improve the capabilities of vision-and-language models for grounding textual descriptions to image regions. We explore various strategies to best generate image-text pairs and image-text-box triplets using a series of pretrained models under different settings and varying degrees of reliance on real data. Through comparative analyses with synthetic, real, and web-crawled data, we identify factors that contribute to performance differences, and propose SynGround, an effective pipeline for generating useful synthetic data for visual grounding. Our findings show that SynGround can improve the localization capabilities of off-the-shelf vision-and-language models and offers the potential for infinite data generation. Particularly, SynGround improves the pointing game accuracy of pretrained ALBEF and BLIP models by 4.81% and 17.11% absolute percentage points, respectively, across the RefCOCO+ and the Flickr30k benchmarks."},"_bibtex":{"value":"@misc{\nhe2024is,\ntitle={Is Synthetic Data Ready for Improving Visual Grounding?},\nauthor={Ruozhen He and Ziyan Yang and Paola Cascante-Bonilla and Alexander C. Berg and Vicente Ordonez},\nyear={2024},\nurl={https://openreview.net/forum?id=EuoHhIqvRD}\n}"},"title":{"value":"Is Synthetic Data Ready for Improving Visual Grounding?"},"pdf":{"value":"/pdf/0451ed16c34402e803305d2c3eed2bcf1e789c97.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"he|is_synthetic_data_ready_for_improving_visual_grounding"},"authorids":{"value":["~Ruozhen_He1","~Ziyan_Yang1","~Paola_Cascante-Bonilla1","~Alexander_C._Berg1","~Vicente_Ordonez2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruozhen He","Ziyan Yang","Paola Cascante-Bonilla","Alexander C. Berg","Vicente Ordonez"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a physics guided deep learning solution for modeling under data paucity and coarse-grained data. Essentially, the paper employs a super-resolution approach to generate fine-grained data from coarse-grained data employing supervised losses at the coarse-grained scale while employing physics losses (conservation conditions) between successive time-steps at the predicted fine-grained scale. Specifically, the proposed architecture comprises a sort of self-supervised task wherein the low-dimensional data is input into an encoder module which produces the corresponding high-dimensional output (predicted). This predicted high-dimensional output is passed into a transition module which predicts the high-dimensional output at the next time step. This high-dimensional output at the successive time-step is downsampled (by a deterministic function) and compared with the ground-truth low-dimensional data using a data-driven loss."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"- The proposed solution is (somewhat) novel and is a creative way to effectively employ coarse-grained data and physics to perform super-resolution. \n \n\n- The results are extensive (although not entirely convincing) and have been performed on multiple important PDE contexts."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Overall, the novelty in the paper is somewhat limited and analyses of the drawbacks of the proposed finite-difference based physics encoding method have not been fully carried out. Specifically, discussions regarding where the proposed method might lack, how fine-grained data can be incorporated (when available) will be helpful additions to the narrative. \n \n\n2. Some results indicate that baselines outperform the proposed method. E.g., Table 1 NSWE indicate that PINO* has lower re-construction errors. Why is this? \n \n\n3. The paper methodology is hard to understand and needs to be significantly improved. The reviewer feels the entire methodology can be explained in 1 – 2 paragraphs (half a page) but is needlessly convoluted and interspersed with details making it hard to get a high-level idea.  \n \n\n4.There are many ambiguous phrases/ design decisions that have been made without explanation. \n\n    a. Why has the U-Net architecture been employed for the encoder while UNet++ [1] , Transformer [2] and many newer image encoding / SR architectures superior to UNet have been proposed more than 2 – 3 years ago? \n \n\n    b. What does “hard encoding” $\\tilde{o}$ into the corresponding $\\hat{u}_t$ mean? \n\n        i. Does it mean that assuming the low-res data was n/2 X h/2 and high-res data was n x h , that every 4th pixel in  the high-res data would have the corresponding $\\tilde{o}$ value? Or does it mean something else? \n\n        ii. If it means the same as <4.b.i>, would this design decision not overtly couple the high-res and low-res solutions? How might the high-res solution significantly improve upon the low-res solution with this constraint? \n \n\n5. Results don't seem practically usable. It is important to comment on this owing to the context (i.e., mapping from low-res to high-res with predominantly low-res training data). In most real-world scientific simulations, physical consistency / errors are assumed to in the range `1e^-5 – 1e^-7` . A discussion about the practicality of the obtained results and the usability of the proposed method is required but missing. \n\n \n\nReferences: \n\n1. Zhou, Zongwei, et al. \"Unet++: A nested u-net architecture for medical image segmentation.\" Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support: 4th International Workshop, DLMIA 2018, and 8th International Workshop, ML-CDS 2018, Held in Conjunction with MICCAI 2018, Granada, Spain, September 20, 2018, Proceedings 4. Springer International Publishing, 2018. \n \n\n2. Dosovitskiy, Alexey, et al. \"An image is worth 16x16 words: Transformers for image recognition at scale.\" arXiv preprint arXiv:2010.11929 (2020)."},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. Why has the U-Net architecture been employed for the encoder while UNet++ [1] , Transformer [2] and many newer image encoding / SR architectures superior to UNet have been proposed more than 2 – 3 years ago? \n \n\n2. What does “hard encoding” $\\tilde{o}$ into the corresponding $\\hat{u}_t$ mean? \n\n    a. Does it mean that assuming the low-res data was n/2 X h/2 and high-res data was n x h , that every 4th pixel in  the high-res data would have the corresponding $\\tilde{o}$ value? Or does it mean something else? \n\n    b. If it means the same as <2.a>, would this design decision not overtly couple the high-res and low-res solutions? How might the high-res solution significantly improve upon the low-res solution with this constraint? \n \n\n3. Additionally, the function of $f_\\theta$ in equation (1) is described as “imitating the implementation of higher-order finite difference to leverage abundant temporal feature of $\\{ \\tilde{o}_{t-i}\\}_{i=0}^n$. ” \n\n    a. What exactly does “imitating the implementation of higher-order FD” mean? Is there a FD operator that has been embedded into $f_\\theta$? Or is there something special (I.e., some special input transformation) that has been applied to the inputs of $f_\\theta$ that makes it “immitate” an FD operator?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700594975754,"tcdate":1698648129859,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission76/Reviewer_iagL"],"signatures":["ICLR.cc/2024/Conference/Submission76/Reviewer_iagL"],"forum":"Dw6y6bEtXm","number":2,"license":"CC BY 4.0","cdate":1698648129859,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission76/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700594975754,"domain":"ICLR.cc/2024/Conference","replyto":"Dw6y6bEtXm","id":"60yIfSJJ6F","forumContent":{"TLDR":{"value":"Modeling physical systems face two challenges: insufficient data and coarse-grained data quality. We propose a novel PICL framework that reconstructs the learnable fine-grained state and enhances the predictive ability in a physics-informed manner."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physics-informed machine learning","Coarse-grained data","PDEs","Neural operator"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Physics-informed machine learning has emerged as a promising approach for modeling physical systems. However, two significant challenges limit its real-world applicability. First, most realistic scenarios allow only coarse-grained measurements due to sensor limitations, making the use of physics loss based on finite dimensional approximations infeasible. Second, the high cost of data acquisition impedes the model's predictive ability. To address these challenges, we introduce a novel framework called Physics-Informed Coarse-grained data Learning (PICL) that incorporates physics information via the learnable fine-grained state representation from coarse-grained data. This framework effectively integrates data-driven methods with physics-informed objectives, thereby significantly improving the predictive ability of the model. The PICL framework comprises two modules: the encoding module, responsible for generating the learnable fine-grained state, and the transition module, used for predicting the subsequent state. To train these modules, we employ a base-training period followed by a two-stage fine-tuning period. The key idea behind this training strategy is that we can leverage physics loss to enhance the reconstruction ability of the encoding module and the generalization ability of the transition module, using both labeled and unlabeled data. In the base-training period, we train both modules collaboratively using data loss and physics loss. In the two-stage fine-tuning period, we first tune the transition module with physics loss using unlabeled data and then tune the encoding module with data loss using labeled data to propagate the information from the transition module to the encoding module. We demonstrate that PICL exhibits superior predictive ability across modeling various PDE-governed physical systems. Code is available on GitHub: https://github.com/PI-CL/PICL."},"_bibtex":{"value":"@misc{\nfeng2024picl,\ntitle={{PICL}: Incorporating Coarse-Grained Data and Physics Information for Superior Physical Systems Modeling},\nauthor={Haodong Feng and Yue Wang and Dixia Fan},\nyear={2024},\nurl={https://openreview.net/forum?id=Dw6y6bEtXm}\n}"},"title":{"value":"PICL: Incorporating Coarse-Grained Data and Physics Information for Superior Physical Systems Modeling"},"pdf":{"value":"/pdf/fd290bb4e74affbfe746cc3ffa1a425c682eb242.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"feng|picl_incorporating_coarsegrained_data_and_physics_information_for_superior_physical_systems_modeling"},"authorids":{"value":["~Haodong_Feng1","~Yue_Wang15","~Dixia_Fan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haodong Feng","Yue Wang","Dixia Fan"]}},"version":2},{"content":{"summary":{"value":"This work targets the dynamical system identification using observation data, which is a hot topic and essential application. The key differences with respect to existing works are clearly stated: 1) identifying unknown physics rather than predefined symbolics, 2) considering the coupling behaviors in the system or interactions when the problem covers multiple domains rather than single mechanical systems."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the bullet points in Weaknesses."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"•\tAs the reviewer summarizes above, key limitations are correctly identified such that the contributions of this work are clear. To the best of the reviewer’s knowledge, this work is new.\n\n•\tThe theoretical analysis is solid, where the definitions and theorems clearly show how the Dirac structure encapsulates internal and external component couplings. Moreover, the corresponding examples of different dynamical systems are well-explained to differentiate the proposed method from existing methods like HNN and NODE.\n\n•\tThe results on multiple systems look promising, where various experiment scenarios and evaluation metrics are comprehensive to validate PoDiNNs’ capabilities."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Some technical details, as well as the claimed capabilities, are unclear, which might be because the reviewer is unfamiliar with all kinds of multi-domain dynamical systems. The confusions are listed below.\n\n•\tThe capability to deal with multi-physics problems is claimed several times in the paper. Specifically, Remark 1 explains the representation of inter-coupling using a bivector element (which is also an NN, right?) Remark 4 with Table 2 briefly demonstrates the scenarios to capture coupled physics in multiple domains. The reviewer would like to know how PoDiNNs represent such interactions.  Is it the same way as the traditional simulation tool, e.g., through iterative refinement of two (or more) simulations of a single domain or system? \n\n•\tThe proposed work targets unknown physics/dynamics, which is quite challenging as there are no predefined physical symbolics in PINN-alike works. How to ensure the PoDiNNs capture the correct physics without causing overfitting problem or continuous good performance in extrapolation?\n\n•\tMoreover, PoDiNNs focus on behaviors that may not be captured by generic models, e.g., NODE, which seem to need intensive resources. Especially, the couplings in multi-physics usually require heavy computation in traditional simulations. The more fine-grained, the heavier. What is PoDiNNs’ capability in this aspect?\n\n•\tFor the last paragraph of Sec. 3.4, an example of electric circuit is used. Could the reviewer further explain with more details: why ODEs or using NODE alone cannot capture the current flow and balanced voltage level? Subsequently, how does PoDiNN mitigate the issue of limited representation? A toy example with mathematical derivations or diagram will be helpful."}},"nonreaders":[],"tmdate":1733276646915,"tcdate":1730603117885,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4280/Reviewer_zHG9"],"signatures":["ICLR.cc/2025/Conference/Submission4280/Reviewer_zHG9"],"forum":"U1DjXQeJRx","number":4,"license":"CC BY 4.0","cdate":1730603117885,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4280/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733276646915,"domain":"ICLR.cc/2025/Conference","replyto":"U1DjXQeJRx","id":"8X4rTrGVUo","forumContent":{"TLDR":{"value":"Poisson-Dirac formulation with ports enables neural networks to model various dynamical systems across domains and identify their internal structures."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["neural ordinary differential equations","coupled system","Poisson system","Dirac structure"]},"supplementary_material":{"value":"/attachment/376c68399daaa997b7f3994fa1de6490412f27aa.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Deep learning has achieved great success in modeling dynamical systems, providing data-driven simulators to predict complex phenomena, even without known governing equations. However, existing models have two major limitations: their narrow focus on mechanical systems and their tendency to treat systems as monolithic. These limitations reduce their applicability to dynamical systems in other domains, such as electrical and hydraulic systems, and to coupled systems. To address these limitations, we propose Poisson-Dirac Neural Networks (PoDiNNs), a novel framework based on the Dirac structure that unifies the port-Hamiltonian and Poisson formulations from geometric mechanics. This framework enables a unified representation of various dynamical systems across multiple domains as well as their interactions and degeneracies arising from couplings. Our experiments demonstrate that PoDiNNs offer improved accuracy and interpretability in modeling unknown coupled dynamical systems from data."},"_bibtex":{"value":"@inproceedings{\nkhosrovian2025poissondirac,\ntitle={Poisson-Dirac Neural Networks for Modeling Coupled Dynamical Systems across Domains},\nauthor={Razmik Arman Khosrovian and Takaharu Yaguchi and Hiroaki Yoshimura and Takashi Matsubara},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=U1DjXQeJRx}\n}"},"title":{"value":"Poisson-Dirac Neural Networks for Modeling Coupled Dynamical Systems across Domains"},"pdf":{"value":"/pdf/cbe5869996e1da69833fe07385317929fb66b5f9.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"khosrovian|poissondirac_neural_networks_for_modeling_coupled_dynamical_systems_across_domains"},"authorids":{"value":["~Razmik_Arman_Khosrovian1","~Takaharu_Yaguchi1","~Hiroaki_Yoshimura1","~Takashi_Matsubara1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Razmik Arman Khosrovian","Takaharu Yaguchi","Hiroaki Yoshimura","Takashi Matsubara"]}},"version":2},{"content":{"venue":{"value":"CoRR 2021"},"pdf":{"value":"https://arxiv.org/pdf/2109.01621v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"oleary|stochastic_physicsinformed_neural_networks_spinn_a_momentmatching_framework_for_learning_hidden_physics_within_stochastic_differential_equations"},"html":{"value":"https://arxiv.org/abs/2109.01621"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2109-01621,\n  publtype={informal},\n  author={Jared O'Leary and Joel A. Paulson and Ali Mesbah},\n  title={Stochastic Physics-Informed Neural Networks (SPINN): A Moment-Matching Framework for Learning Hidden Physics within Stochastic Differential Equations},\n  year={2021},\n  cdate={1609459200000},\n  journal={CoRR},\n  volume={abs/2109.01621},\n  url={https://arxiv.org/abs/2109.01621}\n}\n"},"abstract":{"value":"Stochastic differential equations (SDEs) are used to describe a wide variety of complex stochastic dynamical systems. Learning the hidden physics within SDEs is crucial for unraveling fundamental understanding of these systems' stochastic and nonlinear behavior. We propose a flexible and scalable framework for training artificial neural networks to learn constitutive equations that represent hidden physics within SDEs. The proposed stochastic physics-informed neural ordinary differential equation framework (SPINODE) propagates stochasticity through the known structure of the SDE (i.e., the known physics) to yield a set of deterministic ODEs that describe the time evolution of statistical moments of the stochastic states. SPINODE then uses ODE solvers to predict moment trajectories. SPINODE learns neural network representations of the hidden physics by matching the predicted moments to those estimated from data. Recent advances in automatic differentiation and mini-batch gradient descent with adjoint sensitivity are leveraged to establish the unknown parameters of the neural networks. We demonstrate SPINODE on three benchmark in-silico case studies and analyze the framework's numerical robustness and stability. SPINODE provides a promising new direction for systematically unraveling the hidden physics of multivariate stochastic dynamical systems with multiplicative noise."},"title":{"value":"Stochastic Physics-Informed Neural Networks (SPINN): A Moment-Matching Framework for Learning Hidden Physics within Stochastic Differential Equations"},"authors":{"value":[{"fullname":"Jared O'Leary","username":""},{"fullname":"Joel A. Paulson","username":""},{"fullname":"Ali Mesbah","username":"~Ali_Mesbah1"}]}},"tmdate":1784557056798,"pdate":1640908800000,"externalIds":["dblp:journals/corr/abs-2109-01621"],"tcdate":1784557051269,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Ali_Mesbah1"],"forum":"OJo8Ap2ihC","license":"CC BY-SA 4.0","number":74500,"cdate":1609459200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784557056798,"domain":"OpenReview.net/Public_Article","id":"OJo8Ap2ihC","version":2}],"count":10000}