🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-17
- 类型
- ai-daily
- 字数
- 3546
- 阅读时长
- 17 min
2026-07-17 AI日更 | Kimi K3 推高开源规模战,AI 治理转向“可追溯、可审计” 链接到标题
今日主线一明一暗:Kimi K3 以超大规模引发开源模型竞争再升温,但真实能力与成本仍待验证;另一边,安全访问、生物韧性、数据溯源与前提审计成为治理重点。智能体也在补齐长记忆、自我改进和协作基础设施。
📖 本期 Watch List 深度导读 链接到标题
今天最值得连读的第一条主线,是“AI 走向安全可用”的治理议题:从青少年应当获得受保护的 AI 访问权,到 DeepMind 的 bioresilience,再到 OriginBlame 的训练数据可追溯和 Grounding Audits 的前提依赖审计,几篇文章都在回答同一个问题——AI 能否在开放使用的同时保持可控、可删、可验。第二条主线是智能体基础设施化,Oracle Agent Memory、Self-Improvements survey、Networked Intelligence 和 SPINE 连起来看,很清楚地描绘出长记忆、自我改进、多人协作与机器人落地正在从原型走向系统工程。第三条则偏方法论,关于安全行为学习、神经符号推理与自动微分的几篇新作,适合研发团队补课。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Moonshot AI Launches Kimi K3, Largest Open Frontier Model at 2.8 Trillion Parameters 链接到标题
- 分类:AI · News
- 概况:热度时间:10 hours ago,相关帖子数:11000
- 是什么事:Moonshot AI 发布了 Kimi K3,称其为一款拥有 2.8 万亿参数的开源前沿大模型。
- 为什么重要:这意味着开源大模型在规模、能力和竞争力上继续向顶级闭源模型逼近,也可能推动行业对超大模型架构、训练成本和开源生态的重新评估。
- 讨论概况:X 上主要在讨论它的真实能力是否配得上“前沿模型”定位、2.8 万亿参数的技术含义与训练/推理成本,以及它相较其他头部模型和开源方案的实际优势。
话题 2:Cerebras Unveils Blueprint for Internal Knowledge Base Handling 15,000 Daily Queries 链接到标题
- 分类:AI · News
- 概况:热度时间:2 hours ago,相关帖子数:124
- 是什么事:Cerebras公布了一套企业内部知识库方案,宣称可支持每天约1.5万次查询。
- 为什么重要:该方案显示AI基础设施公司正从训练和推理硬件扩展到企业知识管理场景,反映大模型在内部检索、问答和工作流自动化中的落地需求。
- 讨论概况:X上的讨论主要集中在其性能与成本是否优于现有RAG方案、能否在企业数据安全和权限管理上可靠落地,以及Cerebras是否借此强化其AI推理生态。
话题 3:Moonshot AI Launches Kimi K3, Matching Top U.S. Models 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:15000
- 是什么事:月之暗面(Moonshot AI)发布新一代模型 Kimi K3,宣称其性能已接近或匹配美国顶尖 AI 模型。
- 为什么重要:这被视为中国大模型在推理、代码和通用能力上继续缩小与美国领先模型差距的信号,也加剧了全球 AI 竞争与开源/闭源路线之争。
- 讨论概况:X 上讨论集中在 Kimi K3 的真实基准表现、成本与可用性是否足以挑战 OpenAI、Anthropic 和 Google;支持者认为这是中国模型快速追赶的证据,质疑者则关注评测透明度、实际应用稳定性以及是否存在夸大宣传。
话题 4:OpenAI Launches $230 Codex Micro Macro Pad, Sells Out in Hours 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:24000
- 是什么事:OpenAI 推出了一款售价 230 美元的 Codex Micro Macro Pad,并在数小时内售罄。
- 为什么重要:这反映出 AI 品牌与开发者周边硬件的市场号召力,也显示出围绕 AI 工具生态、开发者工作流和品牌延伸的新商业化路径。
- 讨论概况:X 上的讨论主要集中在这款产品是否只是高溢价周边、是否真正实用,以及 OpenAI 借助限量发售制造热度和稀缺性的营销策略是否奏效。
话题 5:Researcher Adds Vision to Top Open AI Model with Just 50 Million Parameters 链接到标题
- 分类:AI · News
- 概况:热度时间:17 hours ago,相关帖子数:191
- 是什么事:有研究者通过仅新增约5000万参数,就为一款顶级开源 AI 模型补上了视觉能力。
- 为什么重要:这表明多模态能力未必只能依赖大幅扩参或重训,若方法有效,将有助于降低模型升级成本、提升开源模型的可扩展性,并推动更轻量的视觉-语言整合方案。
- 讨论概况:X 上的讨论主要集中在这种“低参数增量实现视觉能力”的效率有多高、是否足以接近原生多模态模型,以及这类方法对开源模型生态和算力门槛意味着什么。
话题 6:Basis AI’s Ode to Accounting Film Traces 8000-Year Legacy 链接到标题
- 分类:AI · Entertainment
- 概况:热度时间:,相关帖子数:41
- 是什么事:Basis AI 发布了一部以会计为主题的影片,回顾会计从古代记账到现代智能化工具的约 8000 年演进。
- 为什么重要:这体现了 AI 公司正尝试用生成式内容和叙事方式包装垂直行业应用,突出 AI 在财务、审计和企业运营自动化中的潜在价值。
- 讨论概况:X 上的讨论主要集中在影片创意是否新颖、会计行业是否会被 AI 深度改造,以及这种品牌叙事究竟是有效科普还是营销包装。
话题 7:MoonPay and Venice AI Launch Lumara Film Festival for AI-Generated Shorts 链接到标题
- 分类:AI · Other
- 概况:热度时间:,相关帖子数:157
- 是什么事:MoonPay 与 Venice AI 联合推出 Lumara Film Festival,面向由 AI 生成的短片作品征集和展示。
- 为什么重要:该活动体现了 AI 生成视频正在从技术演示走向创意产业应用,也显示加密支付、创作者经济与生成式 AI 的进一步交汇。
- 讨论概况:X 上的讨论主要集中在 AI 电影节是否能降低创作门槛、推动独立创作者曝光;同时也有人质疑 AI 生成内容的原创性、版权归属,以及其对传统影视从业者的冲击。
今日 X 上的 AI 舆情小结 链接到标题
今天 X 上的舆论主线,是“AI 能力继续升级,但市场更关心真假实力与落地价值”。不少人形成共识:无论是 Kimi K3 这类超大开源模型,还是企业知识库、轻量多模态补强和生成式内容活动,都说明 AI 正从单点技术演示走向规模化应用,且开源阵营和中国模型都在加速逼近头部闭源模型。分歧主要集中在“是否真的前沿”——Kimi K3 的参数、基准和成本是否配得上宣传,企业方案和周边产品到底是实用创新还是营销包装,也有人质疑低参数增量、多模态补丁和 AI 电影节这类案例的实际含金量。潜在风险则在于评测不透明、过度营销抬高预期,以及企业数据安全、版权归属和创作者/从业者被替代等问题可能在热度之下被低估。
💡 大佬观点(Influencer Insights) 链接到标题
日度 AI 大佬观点速报 链接到标题
基于过去 24 小时内多位 AI 领域 Influencers 的推文动态,我们为你梳理出昨日最受关注的行业脉动。
1. 技术趋势与产品热点 链接到标题
🧠 新模型激战:Kimi K3 成为国产黑马 链接到标题
昨天最热的话题之一是 Kimi K3 的发布。
- @Pluvio9yte 首先爆料“kimi k3出了,传闻超过了opus…”。
- @vista8 随后给出极高评价,称其为“国产模型第一名”,并展示了其惊艳的前端代码生成能力,如一句 prompt 复刻多个风格的高质量网站,赞叹其“美感”。他还预告将发布详细评测。
💻 AI 编程生态:Codex 持续爆发与 Grok Build 开源 链接到标题
OpenAI Codex 和 Grok Build 是编程工具侧的两大热点。
- @dotey 密集播报了 Codex 的最新动态:用户数突破 800 万并再次重置了使用额度;其配套硬件键盘
kbd-1.0-codex-micro正式亮相,设计酷炫。他详细分享了自己基于 Claude Code 和 Codex 的“设计-开发”工作流闭环,并推荐其开发的视频剪辑 SkillBaoCut。 - @vista8 指出 马斯克已将 Grok Build 正式开源,并分享了由 AI 分析该源码后的完整文档。此举被 @vista8 解读为“这是在打 OpenAI 的脸吗?”,凸显了模型工具开源与闭源的路线之争。
🌍 “世界模型”成为新焦点 链接到标题
@Pluvio9yte 发布了一篇深度长文,提出 “AI 正在从生成视频进化到生成世界”。他详细阐述了“世界模型(World Model)”的概念,并重点介绍了开源项目 Alaya World:它能让用户通过文本、图片或视频实时生成可自由探索、交互的 3D 环境,支持 720P 24FPS 的流式生成和大于一分钟的稳定探索。他认为这是 AI 从工具转变为环境的关键一步。
🏠 端侧模型(On-Device)与轻量化 链接到标题
- @zhixianio 持续关注端侧模型,并发布了一期播客专题讨论。他此前测评了 MiniCPM-o 4.5 的全双工音视频能力,对 9B 模型的效果表示满意。同时,他认同 @geekbb 提出的 “Model-Pak”(大模型卡带化)畅想,预示着端侧模型分发和使用的未来形态。
- 行业关注 量化感知训练 (QAT),@zhixianio 转发了 Google 的 QAT 模型,认为这是在训练阶段就为量化“特化”,能更好地让模型在消费级硬件上运行。
2. 独特观点与行业前瞻 链接到标题
- “Vibe Coding”的陷阱与出路(@gefei55):针对越来越多开发者沉迷于用 AI 快速生成 App 的现象,@gefei55 尖锐地指出:“以前一个月写一个没人用的 App,现在一个月花一万块钱 Token 写了 37 个没人用的 App”,生产效率的提高并不意味着能赚到钱。他强调,开发者必须跳出舒适区,学会挖掘用户需求、做宣传推广,否则只是浪费 Token。
- AI 编程工作流的最佳实践(@dotey):@dotey 分享了其高效的开发 Loop:先用 AI 生成和打磨设计原型(Claude Opus/Fable 优于 GPT),再让 AI 基于设计稿 1:1 还原实现,最后由 Agent 自动完成测试与发布。他强调“人”的核心作用是提出想法并最终验证,而非替代 AI。
- AI 时代的新面试题(@ruanyf):他抛出深刻问题:“如果未来代码都是 AI 写的,我们如何招聘程序员?” 指出未来的面试官需要考察的是候选人驾驭 AI 的能力,而不再是手写代码本身,但如何考察尚无定论。
- 学习“预测”以理解本质(@lijigang):@lijigang 提出一个洞察:LLM 靠预测下一个 token 学会语言结构,人也应该通过预测所在领域的下一步,来逼迫自己看清该领域的“生成机制”。
- AI 体感下的“反向消费”(@AI_Jasonyu & @zhixianio):@AI_Jasonyu 分享了从抖音看到的高赞 prompt 技巧,包括问 AI“眼下你最没把握的事是什么”和“我最大的遗漏是什么”,以此让 AI 进行更彻底的“自我审查”,提升输出质量。
3. 推荐的工具与资源 链接到标题
🛠️ AI 智能体工具 & Skills 链接到标题
grillme(Skill):@Pluvio9yte 强烈推荐。它能通过高强度、穷举式的问答,在开发规划阶段帮你榨干所有需求细节,避免后期返工。rn-wechat-extract(Skill):@Pluvio9yte 开源。解决了 Agent 无法读取微信公众号文章的痛点,通过模拟微信 UA 绕过封锁,一键将文章转为 Markdown。AnySearch(AI 搜索基础设施):@Pluvio9yte 分享。专为 Agent 设计的搜索工具,支持金融、学术等垂直领域搜索,可直接将网页转为结构化 Markdown,极大提升 Agent 调研效率。qiaomu-tiny-gif(Skill):@vista8 开源。用于 GIF 压缩,解决公众号 GIF 大小限制问题。REDSkill社区:@ruanyf 观察到小红书开始内测 “REDSkill”,允许用户上传、分享 Skill 文件,试图将社媒平台与 Skill Hub 结合,打造“Skill 的 GitHub”。- 前端交互设计名词库:@vista8 推荐了一个整理 Web/App 常见组件、动效、设计风格名称的网站,方便在 Vibe Coding 时用精确的术语与 AI 沟通。
📚 观点与报告 链接到标题
- Grok Build 源码解析文档:@vista8 将 AI 对 Grok Build 源码的分析文档公开,为开发者学习提供了宝贵材料。
- Ben’s GPT 5.6 模型搭配指南:@vista8 转述了知名博主 Ben 的经验:开发创意用 Sol,复杂任务用 Ultra,日常对话用 Luna。
以上内容由 AI 行业分析师基于 @zhixianio, @Pluvio9yte, @dotey, @vista8, @ruanyf, @gefei55, @lijigang, @AI_Jasonyu 等博主观点整理。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 34 条更新
OpenAI Blog (A_full) 链接到标题
Why teens deserve access to safe AI
- 发布时间:2026-07-17 00:00 北京时间
- 摘要:- 青少年是伴随人工智能成长的第一代人,这项技术将在很大程度上塑造他们的未来。
- 如今,ChatGPT 上近十分之九的青少年在一周内使用它来学习、获取信息、培养技能或提高生产力。
- 这就是为什么我们认为青少年接触人工智能至关重要。
- 阻止青少年在成年之前使用它就像要求上一代人在 18 岁之前避免使用互联网或搜索引擎一样,让他们不太准备好使用他们那个时代的决定性技术之一。
- 但访问必须与专为青少年设计的保护措施相结合。
- EN 要点:
- Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partnerships.
How Cars24 scales conversations and builds faster with OpenAI
- 发布时间:2026-07-16 08:00 北京时间
- 摘要:- Cars24 在印度运营着世界上最大的人工智能原生汽车生态系统之一,用于购买和销售汽车,并在阿联酋和澳大利亚开展了其他业务。
- 该公司支持整个汽车拥有过程,从发现和融资到转售和购后服务,同时通过更高效、更方便的二手车生态系统帮助延长车辆的生命周期,而市场上大多数交易仍然是手动、受监管和分散的。
扩展复杂的、对话驱动的市场。 链接到标题
- 与传统电子商务不同,印度的汽车买卖很少发生在一次交易中。
- 大部分流程发生在应用程序之外,包括通话、文件检查和后续行动,可能需要数天或数周的时间。
- EN 要点:
- Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams acr…
Google DeepMind Blog (A_full) 链接到标题
- Our approach to bioresilience
- 发布时间:2026-07-16 17:30 北京时间
- 摘要:- Google DeepMind 和 Isomorphic Labs 正在分享我们在生物弹性和人工智能模型方面的联合方法。
- 这篇来自 Google DeepMind 博客的文章解释了我们的生物弹性方法如何塑造更广泛的人工智能和基础设施景观。
- 它还为遵循我们的生物弹性方法的创始人、运营商和投资者带来了实际影响。
- EN 要点:
- Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
Two Minute Papers (B_intro+search) 链接到标题
- Claude Just Revealed AI’s Biggest Problem
- 发布时间:2026-07-16 23:39 北京时间
- 摘要:- ❤️ 在这里查看 Lambda 并注册他们的 GPU Cloud:。
- 📝 该论文可在此处获取:.
- Adam Bridges、Benji Rabhan、B Shang、Cameron Navor、Charles Ian Norman Venn、Christian Ahlin、Eric T、Fred R、Gordon Child、Juan Benet、Michael Tedder、Owen Skarpness、Richard Sundvall、Ryan Stankye、Shawn Becker、Steef、Taras Bobrovytsky、Tazaur Sagenclaw、Tybie Fitzhugh、Ueli Gallizzi。
- 克劳德刚刚揭示了人工智能最大的问题。
- EN 要点:
- ❤️ Check out Lambda here and sign up for their GPU Cloud:
- 📝 The paper is available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Ska…
ArXiv cs.AI (B_intro+search) 链接到标题
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13037v1 公告类型:新。
- 摘要:当数据贡献者请求删除时,模型训练者面临一个实际差距:遗忘算法需要遗忘集,但没有工具可以找到哪些训练记录属于给定作者。
- 现有的来源系统在文件或数据集级别运行,导致灾难性的过度删除。
- 我们提出了 ob,一个记录级和令牌级数据来源系统,它通过数据处理管道传播作者身份,并通过确定性查询将撤销请求解析为精确的遗忘集。
- EN 要点:
- arXiv:2607.13037v1 Announce Type: new
- Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate whic…
- Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion
- We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation reques…
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13049v1 公告类型:新。
- 摘要:基础模型为机器人提供了复杂决策的复杂大脑,但将这种智能部署到物理平台中仍然需要繁琐的、专家驱动的校准。
- 这种部署差距,即机器人的脊髓,仍然是可扩展的嵌入式人工智能的主要瓶颈。
- 因此,我们提出 SPINE(具有ageNtic 专业知识的可扩展物理集成):一种代理框架,用于系统地调试和部署具有最少机器人专业知识的双手机器人。
- EN 要点:
- arXiv:2607.13049v1 Announce Type: new
- Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still…
- This deployment gap, the robot’s spinal cord, remains a primary bottleneck to scalable Embodied AI
- Hence, we propose SPINE (Scalable Physical Integration with ageNtic Expertise): an agentic framework for systematically debugging and deploying bimanual robots…
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13069v1 公告类型:新。
- 摘要:大型语言模型产生的思想链 (CoT) 推理看似逻辑合理,但可能并不真正依赖于其规定的前提。
- 我们引入介入性基础审计,即前提依赖性的黑盒、步骤级测试:我们通过用新符号替换其目标谓词来干预单个前提,重新运行模型,并检查每个推理步骤的规范化结论(规范谓词形式)是否发生变化。
- 我们在 ProntoQA 上进行评估,这是一种具有黄金证明树的合成多跳演绎推理基准,其中步骤级前提依赖性是已知的。
- EN 要点:
- arXiv:2607.13069v1 Announce Type: new
- Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises
- We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target pr…
- We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap’s Typed Intensional FOL
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13073v1 公告类型:新。
- 摘要:基于 $IFOL_B$ 的神经符号人工智能是一种将神经学习和符号推理相结合的方法,以克服纯神经系统的局限性(例如缺乏可解释性和逻辑结构),并具有用于自我参考的形式逻辑机制。
- 在本文中,我们基于 Nilsson 的 $IFOL_B$ 概率结构,通过对当前未知句子进行概率计算来扩展 $IFOL_B$ 的认知能力。
- 我们引入了保留当前知识数据库和逻辑推导的全局对称变换,以及用于对仅涉及 $IFOL_B$ 谓词的非常严格子集的具体(子)问题进行实时决策的局部对称变换。
- EN 要点:
- arXiv:2607.13073v1 Announce Type: new
- Abstract: Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like…
- In this paper we expand the cognitive power of $IFOL_B$ by using the probability computation for the currently unknown sentences, based on Nilsson’s probability…
- We introduce the global symmetry transformation that preserves the current knowledge database and logical deduction, and the local one used for real-time decisi…
Self-Improvements in Modern Agentic Systems: A Survey
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13104v1 公告类型:新。
- 摘要:自我改进的自主代理正在从研究原型转向部署系统。
- 主要目标是通过最少甚至没有人类输入的经验进行可控进化或适应。
- 这项调查将现代自我改进代理视为自适应系统,将经验转化为累积的能力增益。
- EN 要点:
- arXiv:2607.13104v1 Announce Type: new
- Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems
- The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input
- This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13115v1 公告类型:新。
-摘要:小语言模型(SLM)已显示出从 SMILES 字符串进行零样本分子属性预测的希望,但由于序列表示未指定关键的图拓扑线索,它们经常遭受结构失明的困扰。
- 我们提出了一个模块化的上下文增强提示框架,可以在推理时使用代理工具:经过训练的 GNN 专家模型可以自信地提供预测提示,并且 GNN 提取特定于实例的解释子图(例如,子图 SMILES 和随附的解释性段落)。
- 我们在五种提示配置(从仅 SMILES 到使用手头所有可用工具)下评估 MUTAG 和 Tox21 上的三种常用 SLM。
- EN 要点:
- arXiv:2607.13115v1 Announce Type: new
- Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural…
- We propose a modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hin…
- We evaluate three commonly used SLMs on MUTAG and Tox21 under five prompting configurations ranging from SMILES-only to using all available tools at hand
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13157v1 公告类型:新。
- 摘要:智能体记忆是长视野智能体的一个系统问题。
- 实际部署需要在扩展对话中保留任务状态,在会话中恢复用户特定的事实和偏好,以及从先前结果中积累程序知识。
- 这些要求超出了文档检索的范围:内存层必须确定哪些交互成为持久状态、如何确定该状态的范围、如何在延迟限制下检索它以及如何随着时间的推移对其进行修改或删除。
- EN 要点:
- arXiv:2607.13157v1 Announce Type: new
- Abstract: Agent memory is a systems problem for long-horizon agents
- Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumu…
- These requirements extend beyond document retrieval: a memory layer must determine which interactions become durable state, how that state is scoped, how it is…
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13172v1 公告类型:新。
- 摘要:我们解决了在环境动态未知且没有合适的奖励函数可用的情况下安全训练代理策略并部署良好且安全的策略的问题。
- 在安全关键环境的背景下,我们认为传统的强化学习不切实际,并求助于人力输入资源。
- 我们推出 DROPJ,一种以人为本的安全培训和部署方法。
- EN 要点:
- arXiv:2607.13172v1 Announce Type: new
- Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown…
- In the context of safety-critical environments, we consider traditional reinforcement learning impractical and resort to the resource of human input
- We introduce DROPJ, a human-centred method for both safe training and deployment
CayleyR: Solving the TopSpin puzzle via cycle intersection
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13219v1 公告类型:新。
- 摘要:我们提出了 cayleyR,这是一个 R 包,用于通过检测凯莱图中的循环交点来解决排列难题。
- 核心算法执行迭代双向搜索:从初始排列状态和目标排列状态,随机操作序列在对称群Sn的凯莱图中生成循环;它们的交叉点产生了一条连接路径。
- 当没有找到直接交叉点时,距离引导的桥梁选择会缩小差距,然后重复该过程。
- EN 要点:
- arXiv:2607.13219v1 Announce Type: new
- Abstract: We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs
- The core algorithm performs an iterative bidirectional search: from both the initial and target permutation states, random operation sequences generate cycles i…
- When no direct intersection is found, a distance-guided bridge selection narrows the gap, and the process repeats
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13220v1 公告类型:新。
- 摘要:大多数人工智能科学系统专注于通过更好的模型、更大的上下文窗口、长视野代理执行或与一个主要用户合作的数字联合科学家来扩展单一推理过程。
- 然而,具有挑战性的科学问题很少仅靠一个推理者就能解决。
- 这些问题是由团队成员解决的,这些团队的成员具有不同的先验、实验背景、隐性知识和领域训练的直觉。
- EN 要点:
- arXiv:2607.13220v1 Announce Type: new
- Abstract: Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic execution,…
- However, challenging scientific problems are rarely solved by one reasoner alone
- They are solved by teams whose members bring different priors, experimental backgrounds, tacit knowledge, and domain-trained intuitions
ArXiv cs.CL (B_intro+search) 链接到标题
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13035v1 公告类型:新。
- 摘要:云服务经常发生事件,需要快速诊断和解决。
- 故障排除指南可帮助工程师做出一致的响应,但手动创建它们是劳动密集型的,导致覆盖不完整和过时的文档。
- 我们推出 FixItFlow,这是一个自动化系统,可以使用大型语言模型根据历史事件数据生成故障排除指南。
- EN 要点:
- arXiv:2607.13035v1 Announce Type: new
- Abstract: Cloud services experience frequent incidents that require rapid diagnosis and resolution
- Troubleshooting guides help engineers respond consistently, but creating them manually is labor-intensive, resulting in incomplete coverage and outdated documen…
- We present FixItFlow, an automated system that generates troubleshooting guides from historical incident data using large language models
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13036v1 公告类型:新。
- 摘要:大语言模型 (LLM) 越来越多地用于医疗保健领域的决策支持,但临床证据往往不完整或不断变化。
- 当现有信息不足以支持可靠的答案时,模型应要求澄清或弃权,而不是提供不受支持的答案。
- 然而,现有的医疗基准通常假设预先可以获得完整的信息。
- EN 要点:
- arXiv:2607.13036v1 Announce Type: new
- Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving
- When the available information is insufficient to support a reliable answer, models should request clarification or abstain rather than provide unsupported resp…
- Existing medical benchmarks, however, typically assume that complete information is available upfront
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13044v1 公告类型:新。
- 摘要:欧洲专利局 (EPO) 报告了 2025 年的创纪录申请,2026 年 EPO 指南要求申请人根据第 83 条和第 42 条对法学硕士辅助内容严格负责,这给对可疑的人工智能生成的专利文本进行分类造成了压力。
- 有两个限制使这变得困难。
- 首先,现实的起诉设置通常只有具有约 8 GB VRAM 的消费级 GPU,而不是数据中心级评分堆栈。
- EN 要点:
- arXiv:2607.13044v1 Announce Type: new
- Abstract: The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted co…
- Two constraints make this hard
- First, realistic prosecution settings often have only consumer GPUs with about 8 GB VRAM, not datacenter-class scoring stacks
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13158v1 公告类型:新。
- 摘要:同步语音翻译 (SimulST) 需要在严格的延迟限制下进行增量翻译,但由于上下文有限和跨语言重新排序,对于仅解码器的 LLM 系统来说仍然具有挑战性。
- 最近的方法经常引入架构更改或显式读/写策略来控制输出时序,这在分段边界不明确的会话语音中可能很脆弱。
- 我们提出了一个简单的数据驱动替代方案:用于累积流解码的固定长度块,具有基于倒回的提交前缀,以及教师标记的前缀到前缀(P2P)目标,具有有限等待微调,产生 CSSEL-P2P,其中 CSSEL 是我们提出的分块流语音编码器 LLM。
- EN 要点:
- arXiv:2607.13158v1 Announce Type: new
- Abstract: Simultaneous speech translation (SimulST) requires incremental translation under strict latency constraints, yet remains challenging for decoder-only…
- Recent approaches often introduce architectural changes or explicit read/write policies to control output timing, which can be brittle in conversational speech…
- We present a simple data-driven alternative: fixed-length chunks for cumulative streaming decoding with a rewind-based committed prefix, and teacher-labeled pre…
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13162v1 公告类型:新。
- 摘要:语言模型会做什么和不会做什么很大程度上是在训练后设定的,但它表达、隐藏或抵制哪些行为并不能仅通过提示来揭示。
- 角色向量,激活空间中的行为方向,可以探测这个组织,但之前的工作只涵盖了少数特征。
- 我们首次在这种规模上系统地应用人物角色向量,编制了跨越四个行为不同领域的 53 个特征清单,并将两个开放权重模型中的每个特征标记为自然(在基线上表达)、可操纵的潜在但可放大或难以处理(难以标准提取)。
- EN 要点:
- arXiv:2607.13162v1 Announce Type: new
- Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by…
- Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits
- We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and label…
Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13164v1 公告类型:新。
- 摘要:手语是数百万聋人和听力障碍人士的主要交流渠道,但文本到手语者的视频生成成本仍然很高,因为视频传播模型的训练和评估成本很高。
- 本文介绍了 Text2Sign,这是一种在单个 NVIDIA L4 GPU 上运行的短手语剪辑的文本条件扩散模型。
- 它将冻结视觉语言文本编码器与 3D 编码器解码器和分解时空注意力相结合,以降低全视频注意力的成本,同时保持运动连贯性。
- EN 要点:
- arXiv:2607.13164v1 Announce Type: new
- Abstract: Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly b…
- This paper presents Text2Sign, a text-conditioned diffusion model for short sign-language clips that runs on a single NVIDIA L4 GPU
- It combines a frozen vision-language text encoder with a 3D encoder-decoder and factorized spatiotemporal attention to reduce the cost of full-video attention w…
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13189v1 公告类型:新。
- 摘要:我们介绍了 RAGthoven,这是我们用于 SemEval-2026 任务 1 (MWAHAHA)、子任务 A(英语、西班牙语和中文的多语言受限幽默生成)的系统。
- RAGthoven 将创造性文本生成分解为多阶段大型语言模型 (LLM) 管道(规划器、Best-of-N Writer、自我批评反射器、LLM-as-a-judge Judge),该管道以计算幽默理论(良性违规理论、基于脚本的幽默语义理论)为基础,并通过十个实验进行完善。
- 在我们的最终配置中,我们通过来自精选笑话语料库的检索增强生成(RAG)来增强规划器,并使用不同的笑话机制播种生成。
- EN 要点:
- arXiv:2607.13189v1 Announce Type: new
- Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chin…
- RAGthoven decomposes creative text generation into a multi-stage large language model (LLM) pipeline (Planner, Best-of-N Writer, Reflector for self-critique, LL…
- In our final configuration, we augment the Planner with retrieval-augmented generation (RAG) from a curated joke corpus, seeding generation with diverse joke me…
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13205v1 公告类型:新。
- 摘要:基于注意力的 KV 缓存驱逐(H2O 及其后代)通过对累积的注意力质量(此处被视为信号能量)并保持最重的令牌进行排序来压缩长上下文模型的内存受限状态。
- 在模式密集的输入流(例如嵌套 JSON)上,此分数充当不成比例地保留噪声的非平稳过滤器:非内容接收器角色(分隔符或空格)比任何内容角色携带的能量多一个数量级,并且结构 KEY 标记以大约 1.8 倍于携带答案的 VALUE 标记的速率过度保留,将精确匹配精度从 88% 降至 0%,精度为 5%预算随着保留状态的信噪比降低而减少。
- 反事实实验表明,抑制 KEY 令牌是最好的可部署过滤器。
- EN 要点:
- arXiv:2607.13205v1 Announce Type: new
- Abstract: Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking tokens on accum…
- On schema-dense input streams such as nested JSON, this score acts as a non-stationary filter that disproportionately retains noise: a non-content sink role (de…
- A counterfactual experiment establishes that suppressing KEY tokens is the best deployable filter
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13248v1 公告类型:新。
- 摘要:大型语言模型(LLM)中数学推理的评估主要集中在英语等高资源语言上。
- 这对人工智能在孟加拉国等语言多样化地区的公平开发和部署造成了重大障碍,该地区有超过 2.3 亿人讲孟加拉语。
- 尽管具有全球意义,但孟加拉语数学推理方面的先前工作很少,并且没有现有的研究系统地对扰动的孟加拉语数学数据集进行基准测试,从而在评估模型的稳健性和模式识别之外的真实理解方面留下了关键的空白。
- EN 要点:
- arXiv:2607.13248v1 Announce Type: new
- Abstract: The evaluation of mathematical reasoning in large language models (LLMs) has predominantly focused on high-resource languages like English
- This has created a significant barrier to the equitable development and deployment of AI in linguistically diverse regions such as Bangladesh, where over 230 mi…
- Despite this global significance, there has been minimal prior work on mathematical reasoning in Bengali and no existing research that systematically benchmarks…
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13260v1 公告类型:新。
- 摘要:政策文件塑造治理结果,但其推理往往是隐含的。
- 参与承诺和管理控制通常在同一文本中共存,而它们之间的紧张关系很少被直接说明。
- 现有的政策话语计算方法无法表达驱动这些紧张局势的框架中介关系,其中一种论点缩小或工具化了另一种论点,而不是拒绝它。
- EN 要点:
- arXiv:2607.13260v1 Announce Type: new
- Abstract: Policy documents shape governance outcomes, but their reasoning is often implicit
- Participatory commitments and managerial control routinely coexist in the same text, and the tensions between them are rarely stated directly
- Existing computational approaches to policy discourse cannot express the frame-mediated relations that drive these tensions, where one argument narrows or instr…
ArXiv cs.LG (B_intro+search) 链接到标题
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13042v1 公告类型:新。
- 摘要:本文用明确的数值追踪了 PyTorch 的自动微分 (AD) 引擎如何计算物理信息神经网络 (PINN) 训练的梯度——这种设置需要两个级别的微分:通过网络计算物理导数 $\hat{y}’(t)=d\hat{y}/dt$,并计算本身依赖的损失的参数梯度 $\nabla_\theta L$ $\hat{y}’(t)$。
- 使用 1-3-3-1 多层感知器和初始值问题 $y’(t)+y(t)=0$、$y(0)=1$,我们跟踪每个节点的完整管道:在正向传递过程中构建的计算图、在单次传递中计算所有 22 个参数梯度的反向模式反向遍历,以及 \texttt{create_graph=True} 通过图上图机制实现正确区分物理信息残差。
- 每个伴随值都根据 Tahimi (2026) 的手推导进行验证,将 $P/Q$ 灵敏度框架连接到 PyTorch 的 autograd 引擎使用的向量雅可比乘积。
- EN 要点:
- arXiv:2607.13042v1 Announce Type: new
- Abstract: This paper traces, with explicit numerical values, how PyTorch’s automatic differentiation (AD) engine computes gradients for Physics-Informed Neural…
- Using a 1-3-3-1 multilayer perceptron and the initial value problem $y’(t)+y(t)=0$, $y(0)=1$, we trace the complete pipeline at every node: the computational gr…
- Every adjoint value is verified against the hand derivations of Tahimi (2026), connecting the $P/Q$ sensitivity framework to the vector–Jacobian products used…
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13043v1 公告类型:新。
- 摘要:深度学习模型实现了最先进的图像分类,但由于计算成本和能源需求而面临部署挑战。
- 我们提出了一种轻量级训练策略,使模型的归一化层适应新领域,并将特征提取与分类器优化解耦,通过仅预计算一次特征来减少开销。
- 重新设计的分类器头具有基于边际的加权损失,进一步最大限度地减少了模糊性,而无需端到端反向传播。
- EN 要点:
- arXiv:2607.13043v1 Announce Type: new
- Abstract: Deep learning models achieve state-of-the-art image classification but face deployment challenges due to computational costs and energy demands
- We propose a lightweight training strategy that adapts normalization layers of the model to the new domain and decouples feature extraction from classifier opti…
- A redesigned classifier head with margin-based weighted loss further minimizes ambiguity without end-to-end backpropagation
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13045v1 公告类型:新。
- 摘要:联邦学习(FL)已成为跨分布式和异构数据源的隐私保护协作模型训练的关键范例。
- 通过将原始数据保留在本地,FL 解决了数据机密性问题,但它并没有解决现代机器学习模型的不透明性。
- 与此同时,可解释人工智能(XAI)因提高透明度、信任和问责制而受到关注,特别是在高风险领域。
- EN 要点:
- arXiv:2607.13045v1 Announce Type: new
- Abstract: Federated Learning (FL) has emerged as a key paradigm for privacy-preserving collaborative model training across distributed and heterogeneous data so…
- By keeping raw data local, FL addresses data confidentiality concerns, yet it does not resolve the opacity of modern machine learning models
- In parallel, Explainable Artificial Intelligence (XAI) has gained attention for improving transparency, trust, and accountability, particularly in high-stakes d…
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13046v1 公告类型:新。
- 摘要:我们为机器学习模型丢弃的信息开发了一个框架,其输入带有李群动作。
- 给定空间 $V$ 上李群 $G$ 的表示 $\pi$ 和学习函数 $f\colon V \to \mathbb{R}$,我们定义两个测量 $f$ 不可见的对称性的对象。
- V$ 中 $x \in 点的零纤维是 $N_G(f,x) = {g \in G : f(\pi(g^{-1}) \cdot x) = f(x)}$ 的群元素的集合,其对 $x$ 的逆作用是 $f$ 无法检测到的。
- EN 要点:
- arXiv:2607.13046v1 Announce Type: new
- Abstract: We develop a framework for the information discarded by machine learning models whose inputs carry a Lie group action
- Given a representation $\pi$ of a Lie group $G$ on a space $V$ and a learned function $f\colon V \to \mathbb{R}$, we define two objects measuring the symmetry i…
- The null fiber at a point $x \in V$ is the set $N_G(f,x) = {g \in G : f(\pi(g^{-1}) \cdot x) = f(x)}$ of group elements whose inverse action on $x$ is undetec…
Targeted Recovery of Weight-Space Mechanisms From Neural Networks
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13047v1 公告类型:新。
- 摘要:参数分解(PD)将神经网络分解为可解释的计算组件,忠实地反映原始网络的操作。
- 然而,将 PD 扩展到大型模型需要大量计算,这使其成为一项成本高昂且风险很大的工作。
- 在这里,我们提出有针对性的 PD (tPD),它通过引入处理所有非目标数据的高级包罗万象的组件,仅识别处理感兴趣的特定输入的组件 - 从孤立的提示到大型子任务。
- EN 要点:
- arXiv:2607.13047v1 Announce Type: new
- Abstract: Parameter decomposition (PD) decomposes neural networks into interpretable computational components that faithfully reflect the original network’s ope…
- However, scaling PD to large models requires vast compute, making it a costly and risky endeavor
- Here we propose targeted PD (tPD), which identifies only the components that process specific inputs of interest – from isolated prompts to large subtasks – b…
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13048v1 公告类型:新。
- 摘要:流式推理管道越来越多地将轻量级快速模型与大型语言模型 (LLM) 结合起来,以大量成本提供丰富的语义理解。
- 何时启用法学硕士这一核心问题得到了有限的正式处理。
- 我们将其视为基于风险的顺序停止问题,当观察历史记录中的风险函数超过阈值时,触发策略就会触发。
- EN 要点:
- arXiv:2607.13048v1 Announce Type: new
- Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at…
- The central question of when to invoke the LLM has received limited formal treatment
- We cast this as a risk-based sequential stopping problem, where a trigger policy fires when a risk functional over the observation history exceeds a threshold
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13101v1 公告类型:新。
- 摘要:全球站天气预报(GSWF)对于关键地区的局部和极端天气预测至关重要。
- 尽管努力利用回溯窗口,但现有方法的精度增益有限,并且难以应对极端事件和错误累积。
- 这些限制源于对短期模式的过度依赖,这些模式不足以捕捉混乱的天气动态,特别是在部分观测下。
- EN 要点:
- arXiv:2607.13101v1 Announce Type: new
- Abstract: Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions
- Despite efforts to exploit look-back windows, existing methods show limited accuracy gains and struggle with extreme events and error accumulation
- These limitations stem from overreliance on short-term patterns, which are insufficient to capture chaotic weather dynamics, especially under partial observatio…
Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13103v1 公告类型:新。
- 摘要:知识追踪(KT)旨在通过对历史交互中不断变化的知识状态进行建模来预测学生的未来表现。
- 现有的KT方法通常将原始交互序列视为统一的行为过程,忽视了学习行为的阶段特定性。
- 我们的初步观察表明,经过充分的练习,学生更有可能正确回答以前失败的知识概念,这表明学生从能力培养向熟练学习转变。
- EN 要点:
- arXiv:2607.13103v1 Announce Type: new
- Abstract: Knowledge tracing (KT) aims to predict students’ future performance by modeling their evolving knowledge states from historical interactions
- Existing KT methods usually treat the raw interaction sequence as a unified behavioral process, overlooking the phase-specific nature of learning behaviors
- Our preliminary observations show that students are more likely to correctly answer previously failed knowledge concepts after sufficient practice, suggesting a…
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13108v1 公告类型:新。
-摘要:现实世界的交通数据表现出异构的空间相关性和非线性的时间动态,给准确的时空预测带来了巨大的挑战。
- 现有方法已经开发出越来越复杂的图形、注意力和分解架构,而底层非线性函数逼近器的影响受到的关注相对较少。
- 在这项工作中,我们提出了 STKAN,一种时空预测架构,它将泰勒多项式 Kolmogorov-Arnold 网络模块引入到时空令牌混合中。
- EN 要点:
- arXiv:2607.13108v1 Announce Type: new
- Abstract: Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-…
- Existing approaches have developed increasingly sophisticated graph, attention, and decomposition architectures, while the influence of the underlying nonlinear…
- In this work, we propose STKAN, a spatio-temporal forecasting architecture that introduces Taylor-polynomial Kolmogorov–Arnold Network modules into spatial and…
A Hybrid Mamba for Audio-Visual Navigation
- 发布时间:2026-07-16 12:00 北京时间
- 摘要:- arXiv:2607.13110v1 公告类型:新。
- 摘要:自 2020 年以卷积神经网络和循环架构为中心的范式建立以来,视听导航的基础骨干网络在五年多的时间里没有发生本质的变化,使得它们不足以支持动态多模态序列的有效表示。
- 本文提出了 Samba(用于视听导航的混合曼巴)。
- 它使用支持自适应选择的 Mamba 状态编码器 (M-SE) 来代替传统的 GRU 进行时间聚合,并构建音频 Mamba 编码器 (AME) 来弥补卷积算子在捕获频谱图中的全局时频依赖性方面的局限性。
- EN 要点:
- arXiv:2607.13110v1 Announce Type: new
- Abstract: Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone networks fo…
- This paper proposes Samba(A Hybrid Mamba for Audio-Visual Navigation)
- It uses the adaptive selection-enabled Mamba State Encoder (M-SE) to replace conventional GRUs for temporal aggregation, and constructs an Audio Mamba Encoder (…