Intern-S2-Preview: Scientific Agentic Foundation Model

📄 arXiv: 2608.13505v1 📥 PDF

作者: Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, Erfei Cui, Xuanlang Dai, Shengyuan Ding, Shangheng Du, Yanhui Duan, Yue Fan, Youqing Fang, Quan Gan, Yuanyuan Gao, Jiaye Ge, Lixin Gu, Yuzhe Gu, Qipeng Guo, Junjun He, Xin Hong, Ming Hu, Zhouqi Hua, Haian Huang, Junhao Huang, Zixian Huang, Minxi Jin, Lingkai Kong, Alexander Lam, Zehao Li, Zonglin Li, Tianhao Liang, Dahua Lin, Junyao Lin, Tianyang Lin, Zhouhan Lin, Jiangning Liu, Jin Liu, Kuikun Liu, Wenran Liu, Yifei Liu, Yuhong Liu, Yuhong Liu, Zhoumianze Liu, Ziyan Liu, Ziyu Liu, Haijun Lv, Han Lv, Chengqi Lyu, Le Ma, Ningsheng Ma, Zerun Ma, Haoyang Peng, Runyu Peng, Jifei Shan, Zixin Shang, Kou Shi, Xiang Shi, Qisheng Su, Xuerui Su, Hao Sun, Xiao Sun, Yanan Sun, Yu Sun, Huanze Tang, Yinghao Tang, Wenhui Tian, Zhongbo Tian, Bingli Wang, Haomin Wang, Jiarui Wang, Jingzhi Wang, Rui Wang, Xiquan Wang, Yi Wang, Zhecan Wang, Ziyi Wang, Zun Wang, Rubin Wei, Lianyi Wu, Wen Wu, Yue Wu, Yuhan Wu, Zhenyu Wu, Zijian Wu, Shuhao Xing, Jun Xu, Xingle Xu, Xuenan Xu, Xiangchao Yan, Ziang Yan, Bowen Yang, Danni Yang, Lin Yang, Zhiqi Yang, Qian Yao, Haochen Ye, Peng Ye, Jinhui Yin, Jiashuo Yu, Dingbo Yuan, Fei Yuan, Yuhang Zang, Bo Zhang, Chao Zhang, Chen Zhang, Hongjie Zhang, Junming Zhang, Wenlong Zhang, Wenwei Zhang, Yiming Zhang, Zhuo Zhang, Ziyang Zhang, Haiteng Zhao, Penghao Zhao, Yibo Zhao, Zhonghan Zhao, Zhihang Zhong, Bowen Zhou, Peiheng Zhou, Xin Zhou, Xinyu Zhou, Yunhua Zhou, Dongsheng Zhu, Yicheng Zou

分类: cs.LG, cs.CL, cs.CV

发布日期: 2026-08-13

备注: 35 pages, 12 figures


💡 一句话要点

提出Intern-S2-Preview以解决科学发现中的多模态理解问题

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多模态理解 科学发现 强化学习 时间序列建模 记忆增强 长时间任务 数据分析 智能决策

📋 核心要点

  1. 现有方法在科学发现中面临多模态理解和长时间任务的挑战,难以有效整合异构数据。
  2. 论文提出的Intern-S2-Preview通过多模态预训练和强化学习等技术,旨在提升科学理解和推理能力。
  3. 实验结果表明,Intern-S2-Preview-397B在多个基准测试中表现优异,尤其在科学信号理解和预测方面取得显著提升。

📝 摘要(中文)

科学发现日益依赖于能够对异构模态的科学证据进行推理、与科学工具和环境互动并在长任务周期内持续进展的AI系统。本文提出了Intern-S2-Preview,一系列旨在支持多模态科学理解、推理、生成和长时间任务的科学代理基础模型。训练流程从科学多模态预训练开始,涵盖渲染的科学文档、交错的图像-文本数据和多样的科学语料库。在预训练检查点的基础上,应用统一的后训练流程,包括监督微调、可扩展的多任务强化学习、黑白盒代理强化学习和在线策略蒸馏。评估结果显示,Intern-S2-Preview-397B在多个科学和多模态基准上实现了竞争性或领先的结果。

🔬 方法详解

问题定义:论文旨在解决科学发现中对异构模态数据的理解和推理不足的问题。现有方法在处理长时间任务时,往往无法有效整合和利用多种数据源。

核心思路:Intern-S2-Preview通过建立一个多模态的科学代理基础模型,结合预训练和后训练的策略,提升模型在科学任务中的推理和生成能力。这样的设计使得模型能够在复杂的科学环境中进行有效的互动和学习。

技术框架:整体架构包括科学多模态预训练、监督微调、可扩展的多任务强化学习、黑白盒代理强化学习和在线策略蒸馏等多个阶段。每个阶段都针对不同的任务需求进行优化,以确保模型的稳定性和效率。

关键创新:最重要的技术创新在于引入了时间序列建模和记忆解码器,使得模型能够在长序列理解和数值预测上表现出色。这与现有方法的主要区别在于其对长时间依赖关系的处理能力。

关键设计:在参数设置上,采用了适应性长度正则化和在线推测解码等技术细节,以提高训练的稳定性和效率。此外,模型的多任务优化和经验组装策略也为代理任务的执行提供了支持。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,Intern-S2-Preview-397B在多个科学和多模态基准测试中取得了竞争性或领先的成绩。特别是在SciTS数据集上,时间序列模块显著提升了科学信号的理解和预测能力。此外,Intern-MemDec-4B扩展在Biology-Instructions任务中的平均得分从56.92提升至60.32,显示出显著的性能改进。

🎯 应用场景

该研究的潜在应用领域包括科学研究、数据分析和智能决策支持等。通过提升AI在科学领域的理解和推理能力,Intern-S2-Preview有望加速科学发现的进程,并在实际应用中提供更为精准的分析和预测。未来,该模型可能在生物医学、环境科学等多个领域发挥重要作用。

📄 摘要(原文)

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.