🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-18
- 类型
- ai-daily
- 字数
- 3520
- 阅读时长
- 17 min
2026-07-18 AI日更 | 智能体进入基建战:本地助理、百万沙箱与世界模型浮出水面 链接到标题
今日主线是智能体工程化:本地优先助理、长程记忆、沙箱并发与机器人部署共同补齐执行层。模型竞争仍在代码能力和底层架构延续,但企业更关注投入产出、数据溯源与可靠性评测。
📖 本期 Watch List 深度导读 链接到标题
今天最值得先读的是“智能体落地”这一组:OpenClaw 的本地助理实践、Oracle 的长程智能体记忆、SPINE 面向双手机器人的部署框架,以及自我改进智能体综述,共同指向一个趋势——智能体竞争正从模型能力转向记忆、工具、物理执行和持续适应。
第二条主线是“AI 价值与治理”。《A scorecard for the AI age》从 CFO 视角重提 AI 投入产出,OriginBlame 则把训练数据溯源推进到记录和 token 级,适合关注企业合规、遗忘请求和成本核算的团队细读。
最后推荐几篇评测与安全论文:CoT 前提依赖审计、VLM 多轮反驳稳定性测试、基于人类偏好和 world model 的安全训练,都是判断模型“是否真的可靠”的实用方法论。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Fake Claim of Chinese AI Distilling Anthropic Model Spreads on X 链接到标题
- 分类:AI · News
- 概况:热度时间:,相关帖子数:121
- 是什么事:X 上流传关于“中国 AI 公司蒸馏 Anthropic 模型”的说法,但该说法被指缺乏可靠证据或存在误导。
- 为什么重要:这反映出大模型竞争中关于数据来源、模型蒸馏和知识产权的争议日益敏感,也凸显 AI 领域信息误传可能影响企业声誉和监管讨论。
- 讨论概况:讨论焦点集中在该指控是否有实证依据、模型相似性是否足以证明蒸馏,以及西方 AI 公司与中国 AI 公司之间的技术竞争是否被政治化;也有人呼吁在传播相关指控前提供可验证证据。
话题 2:Moonshot AI’s Kimi K3 Tops Frontend Code Arena Leaderboard 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:114000
- 是什么事:Moonshot AI 的 Kimi K3 在 Frontend Code Arena 排行榜上登顶,引发开发者社区关注。
- 为什么重要:这显示中国大模型在前端代码生成、界面实现和工程化任务上的竞争力提升,也反映代码能力正成为评估 AI 模型实用价值的重要指标。
- 讨论概况:X 上的讨论主要集中在 Kimi K3 的真实编码能力、排行榜评测是否能代表实际开发体验,以及它与 Claude、GPT、Gemini 等模型在前端生成质量、稳定性和可用性上的差距。
话题 3:Boris Cherny Maps Path to 1,000-Agent AI Coding Teams 链接到标题
- 分类:AI · News
- 概况:热度时间:4 hours ago,相关帖子数:306
- 是什么事:Boris Cherny提出了将AI编程协作扩展到“1000个智能体团队”的路线图,设想大量代码智能体并行完成软件开发任务。
- 为什么重要:这反映了AI编程从单一助手走向多智能体协作的趋势,可能改变软件工程的组织方式、开发效率和工程管理流程。
- 讨论概况:X上的讨论集中在多智能体编程是否真正可扩展、如何解决代码质量与协调成本、以及这类系统会增强还是替代人类开发者等问题。
话题 4:Moonshot AI’s Zhilin Yang Returns to China After CMU PhD, Sparks Talent Debate 链接到标题
- 分类:AI · News
- 概况:热度时间:17 hours ago,相关帖子数:6300
- 是什么事:Moonshot AI(月之暗面)联合创始人杨植麟在完成卡内基梅隆大学博士学习后回到中国,引发关于顶尖 AI 人才流向的讨论。
- 为什么重要:这一动向被视为中国大模型创业公司吸引全球高端研究人才、提升基础模型研发能力的信号,也反映中美 AI 竞争中人才与科研生态的重要性。
- 讨论概况:X 上讨论集中在中国 AI 公司是否具备留住顶尖人才的环境、海外博士回国是否会改变技术竞争格局,以及人才流动应被视为正常职业选择还是地缘科技竞争的一部分。
话题 5:Modal Launches Sandbox Platform for 1 Million Concurrent AI Environments 链接到标题
- 分类:AI · News
- 概况:热度时间:22 hours ago,相关帖子数:274
- 是什么事:云计算平台 Modal 推出面向 AI 工作负载的 Sandbox 平台,宣称可同时运行 100 万个隔离 AI 环境。
- 为什么重要:这表明 AI 应用正在从单一模型调用走向大规模并发、可隔离执行的代理和代码运行场景,对推理基础设施、弹性调度和安全沙箱能力提出更高要求。
- 讨论概况:X 上讨论集中在其并发规模是否真实可用、成本和延迟表现、与现有云服务和容器平台的差异,以及该能力是否会加速 AI Agent、自动化编程和批量实验的落地。
话题 6:Higgsfield AI Open-Sources Prompts for Seedance 2.0 Cinematic Videos 链接到标题
- 分类:AI · News
- 概况:热度时间:22 hours ago,相关帖子数:1500
- 是什么事:Higgsfield AI 开源了一组用于 Seedance 2.0 生成电影感视频的提示词模板。
- 为什么重要:这有助于降低高质量 AI 视频生成的门槛,让创作者更容易复现镜头语言、运动控制和电影化风格,也推动提示词工程在视频生成领域的共享与标准化。
- 讨论概况:X 上的讨论主要集中在这些提示词是否能稳定产出专业级效果、Seedance 2.0 与 Runway、Kling、Veo 等模型的表现对比,以及开源提示词会提升创作效率还是加剧同质化内容的问题。
话题 7:Decart AI Launches Lucy 2.5 for Real-Time Video Edits 链接到标题
- 分类:AI · News
- 概况:热度时间:22 hours ago,相关帖子数:3000
- 是什么事:Decart AI 发布 Lucy 2.5,声称可在视频直播过程中通过提示词实时完成特效、换装、换场景和风格化等编辑。
- 为什么重要:如果实时生成式视频编辑达到稳定可用,将使视频从固定成品变为可交互、可动态改写的内容层,可能影响直播、电商、广告、社交、游戏和流媒体等 AI 应用场景。
- 讨论概况:X 上讨论集中在 Lucy 2.5 是否标志着“实时世界模型”走向产品化,以及它对创作者工作流和商业内容定制的影响;同时也有人质疑演示是否经过筛选、实际延迟和一致性是否足以支撑大规模生产使用。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线围绕“AI 能力快速产品化”与“竞争叙事加剧”展开:从 Kimi K3 登顶代码榜、多智能体编程、百万级沙箱,到实时视频编辑和电影感提示词,社区普遍认同 AI 正从单点模型能力走向工程化、规模化和创作流程重塑。共识在于,中国 AI 公司在代码、视频和人才吸引力上的存在感明显上升,基础设施和 Agent 化执行环境也正成为下一阶段竞争焦点。分歧则集中在这些进展的真实性和可迁移性:排行榜是否代表真实开发体验,演示是否经过筛选,百万并发是否可用,多智能体是否会被协调成本抵消,以及模型相似性是否足以支持“蒸馏”指控。潜在风险包括未经证实的技术指控被地缘政治化并伤害企业声誉,AI 视频与实时编辑带来内容同质化、真实性和滥用问题,以及大规模 Agent/沙箱系统在成本、安全隔离和代码质量控制上仍可能暴露新型工程与治理挑战。
💡 大佬观点(Influencer Insights) 链接到标题
好的,基于过去24小时的推文数据,以下是为你准备的 AI 行业资讯日报。
AI 行业日报:大佬观点速览 (2026年7月17日) 链接到标题
1. 核心趋势:模型竞赛白热化,Agent 工程化成为焦点 链接到标题
今日的讨论核心围绕着两大巨头的激烈角逐和 Agent 生态的快速基建。
Kimi K3 发布,剑指 Claude Fable 5 Kimi K3 的发布无疑是今日引爆社区的最大热点。多位博主将其视作国产模型的里程碑,并直接对标甚至声称超越了 Claude Fable 5。
- @Pluvio9yte 直接指出:“kimi k3出了,传闻超过了opus…”,并大胆预测 Anthropic 可能因此再次延长 Fable 5 的付费访问权限,将其视作一次市场防御行为。
- @vista8 在快速测试后,盛赞 Kimi K3 的“美感”和前端代码生成能力,称“目前是国产模型第一名”,并展示了其“一句话复刻网站”的强大功能。
- 竞争带来的用户红利:模型能力竞争直接带来了更优的套餐和策略。@Pluvio9yte 和 @dotey 都观察到,Claude 和 Codex 在彼此新模型发布前后频繁进行“额度重置”,@vista8 将此现象总结为:“果然需要反垄断啊,有竞争群众才有利。”
Agent 框架与开发者工具的“战国时代” Agent 不再仅仅是模型能力的展示,一套完整的开发者工具链和生态系统正在形成。
- Grok Build 开源:马斯克的 xAI 开源其终端 AI 编码 Agent grok-build。@AI_Jasonyu 和 @vista8 都第一时间关注了此事。@AI_Jasonyu 点评其功能齐全,“在 Claude Code 里用的那套它基本都有”,包括 MCP、Skills、沙箱等,并且是用 Rust 编写的。这为 Agent 赛道注入了强大的开源竞争力。
- Codex 生态的完善与外延:OpenAI 的 Codex 依然是讨论中心,但焦点转向了其生态。官方推出了 Codex Micro 物理键盘(@dotey 关注),社区则出现了为其更换主题皮肤(@vista8)和用百元手柄平替 Codex Micro 的开源项目(@dotey 转发),显示出开发者对这一产品的狂热和高度参与。
2. 独特观点与行业前瞻 链接到标题
AI 编程的边界:从“代码生成”到“世界模型”的演进
- @Pluvio9yte 发表长文,深刻阐述了 AI 视频模型的下一步趋势:不是生成更长的视频,而是构建可实时交互的「世界模型」(World Models)。他以新近开源的 Alaya World 为例,指出 AI 正在从生成内容的工具演变为可探索、可交互的环境,这将对游戏、具身智能等领域产生根本性影响。“AI Is Moving Beyond ‘Generating Videos’ — Toward ‘Generating Worlds’。”
- 与此同时,@Pluvio9yte 也分享了自己“前端先行”(Frontend-First)的 AI 开发策略:先通过 AI 生成前端界面和前后端契约,待 UX 和交互完全确认后,再根据已确定的契约补齐后端,避免反复修改的高昂成本。这是一个务实的工程化方法。
Kimi K3 背后的技术深度:重思 Transformer 地基
- @dotey 详细解读了月之暗面 CEO 杨植麟的演讲,指出其核心思想是“把 AI 训练中三个沿用了近十年的基础组件,重新做一遍”:
- 优化器:用 MuonClip 替代 Adam,挖掘每个 Token 的价值。
- 注意力机制:用 Kimi Linear (KDA) 替代 Full Attention,提升长上下文效率。
- 残差连接:提出 Attention Residue 作为残差连接的下一代方案,让模型主动从历史层中选择信息。
- @dotey 认为,这指明了在数据墙和成本压力下,重新设计底层架构 是仅次于堆参数的重要突破方向。
- @dotey 详细解读了月之暗面 CEO 杨植麟的演讲,指出其核心思想是“把 AI 训练中三个沿用了近十年的基础组件,重新做一遍”:
端侧模型与硬件的务实探讨
- @ruanyf 提出了一个反直觉的判断:对于本地运行 AI,“板载芯片组的迷你 PC 很多时候比独立显卡更好”。他指出,搭载 128GB 统一内存的 AMD Strix Halo 平台,能运行 5090 显卡因 32G 显存限制而无法承载的大模型,且成本更低。这为开发者选择本地设备提供了全新视角。
- @zhixianio 分享了测 Gemma 4 12B Coder 的挫败感,指出其 12B 的体量天花板导致其难以处理“长篇、有状态、一次成型”的复杂程序,如俄罗斯方块。这强调了在模型体量和任务复杂度之间做匹配的重要性。
- @dotey 则强烈建议,开发 Agent 应用首选 Mac,因为其生态支持最好,并特别强调 “内存一定要高,16 G 都小了”。
AI 时代的人类竞争力反思
- @dotey 针对 GPT 5.6 设计能力差的讨论,点明:“AI 能提升普通人的下限,但是提升不了上限;AI 能放大和加速专业能力。” 品味和设计能力依然由人主导。
- @gefei55 批评了那种用 AI 一个月写 37 个无人问津 App 的行为,认为 AI 提高了生产效率,但创业者若只停留在编码,而不去 “先调研再做 App” 和 “去宣传推广”,将陷入“Token 陷阱”,浪费钱和时间。他呼吁开发者练习营销能力,成为产销一条龙的超级个体。
3. 值得尝试的工具与资源 链接到标题
- Grok Build: 马斯克 xAI 开源的终端 AI 编码 Agent,Rust 编写,功能全面,支持 MCP/Skills/沙箱。GitHub 地址见 @AI_Jasonyu 或 @vista8 推文。
- rn-wechat-extract (Skill): 由 @Pluvio9yte 开发的开源 Claude Code/Codex 技能。解决了 Agent 无法直接读取微信公众号文章内容的问题,通过模拟微信 UA 获文章内容并转为 Markdown。安装命令:
npx -y skills add Pluviobyte/rnskill --skill rn-wechat-extract。 - AnySearch: 由 @Pluvio9yte 推荐的为 AI Agent 设计的搜索基础设施,支持金融、学术等垂直领域搜索,并将结果结构化输出,极大提升 Agent 调研效率。
- grillme (Skill): @Pluvio9yte 强烈推荐的项目规划 Skill,在代码编写前通过问答形式“榨干”所有不确定的需求,生成事无巨细的规划。“规划阶段多花20分钟完善方案比代码完成之后多花60分钟改模块效率要高的多。”
- Qwen3 ASR: 由 @dotey 推荐,配合 Qwen3-ForcedAligner 可实现高精度的词级时间戳转写,0.6B 模型即可本地运行,是 OpenAI Whisper 的优秀开源替代品,特别适合字幕制作。
- Emil 的 UI 设计 Skill: @vista8 推荐的“如果只能推荐一个去 AI 味设计 Skill”,由设计师 emil 创建,能生成高审美、带精美动效的 UI 界面。安装指令:
npx skills add emilkowalski/skill。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 33 条更新
Y Combinator Podcast (B_intro+search) 链接到标题
- World Models, Explained
- 发布时间:2026-07-17 22:43 北京时间
- 摘要:- 您可能已经听说过 OpenClaw(以前称为 Clawdbot/Moltbot)。
- 引起轰动的开源人工智能助手可以在您自己的设备上运行,与您已经使用的消息应用程序连接,并且超越聊天功能,实际执行管理电子邮件、日历、文件、工作流程等任务。
- 现在来认识一下它背后的人。
- YC 的 Raphael Schaad 与 OpenClaw 的创始人 Peter Steinberger 坐下来讨论病毒式个人 AI 代理背后的“顿悟”时刻、为什么本地优先代理可以取代当今的许多应用程序,以及个人代理将如何重塑软件的未来。
- EN 要点:
- Why do even our best AI models need tens of thousands of examples to learn skills that a human picks up in a handful of tries
- Solving this problem is one of the great open challenges in modern AI
- World models, which give AI an internal simulation of its environment, are one of the most promising paths forward.In this episode of Decoded, YC’s Ankit Gupta…
- Full Transcript:
Stratechery by Ben Thompson (A_full) 链接到标题
- 2026.29: Mainframes and Main Characters
- 发布时间:2026-07-18 01:00 北京时间
- 摘要:- 左为 Jean Pigozzi,来自 Andy Hertzfeld,右为 GenAI,来自 Muse Image。
- 欢迎回到本周的Stratechery!
- 提醒一下,每周、每周五,我们都会发送 Stratechery 捆绑包中的内容概述;突出显示的链接对所有人免费。
- 此外,您可以完全控制我们发送给您的内容。
- 就此而言,这是本周我们最喜欢的一些。
- EN 要点:
- Jean Pigozzi via Andy Hertzfeld on left, GenAI via Muse Image on right
- Welcome back to This Week in Stratechery
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
- Additionally, you have complete control over what we send to you
OpenAI Blog (A_full) 链接到标题
- A scorecard for the AI age
- 发布时间:2026-07-17 18:00 北京时间
- 摘要:- 我从各地首席财务官那里听到的问题很简单:我们如何从人工智能支出中获得更多价值?
- 多年来,市场通过采用来衡量软件的成功:购买席位、活跃用户、更新许可证。
- 了解人工智能的价值需要更强有力的衡量标准:工作完成情况。
- 首席财务官和其他企业领导者面临的基本经济问题是,人工智能完成的工作的价值增长速度是否快于生产成本的增长速度。
- 回答这个问题需要比每个代币成本等指标更深入地研究。
- EN 要点:
- Sarah Friar, CFO of OpenAI, introduces a practical AI scorecard to measure ROI through useful work, cost per successful task, dependability, and return on compu…
ArXiv cs.AI (B_intro+search) 链接到标题
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13037v1 公告类型:新。
- 摘要:当数据贡献者请求删除时,模型训练者面临一个实际差距:遗忘算法需要遗忘集,但没有工具可以找到哪些训练记录属于给定作者。
- 现有的来源系统在文件或数据集级别运行,导致灾难性的过度删除。
- 我们提出了 ob,一个记录级和令牌级数据来源系统,它通过数据处理管道传播作者身份,并通过确定性查询将撤销请求解析为精确的遗忘集。
- EN 要点:
- arXiv:2607.13037v1 Announce Type: new
- Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate whic…
- Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion
- We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation reques…
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13049v1 公告类型:新。
- 摘要:基础模型为机器人提供了复杂决策的复杂大脑,但将这种智能部署到物理平台中仍然需要繁琐的、专家驱动的校准。
- 这种部署差距,即机器人的脊髓,仍然是可扩展的嵌入式人工智能的主要瓶颈。
- 因此,我们提出 SPINE(具有ageNtic 专业知识的可扩展物理集成):一种代理框架,用于系统地调试和部署具有最少机器人专业知识的双手机器人。
- EN 要点:
- arXiv:2607.13049v1 Announce Type: new
- Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still…
- This deployment gap, the robot’s spinal cord, remains a primary bottleneck to scalable Embodied AI
- Hence, we propose SPINE (Scalable Physical Integration with ageNtic Expertise): an agentic framework for systematically debugging and deploying bimanual robots…
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13069v1 公告类型:新。
- 摘要:大型语言模型产生的思想链 (CoT) 推理看似逻辑合理,但可能并不真正依赖于其规定的前提。
- 我们引入介入性基础审计,即前提依赖性的黑盒、步骤级测试:我们通过用新符号替换其目标谓词来干预单个前提,重新运行模型,并检查每个推理步骤的规范化结论(规范谓词形式)是否发生变化。
- 我们在 ProntoQA 上进行评估,这是一种具有黄金证明树的合成多跳演绎推理基准,其中步骤级前提依赖性是已知的。
- EN 要点:
- arXiv:2607.13069v1 Announce Type: new
- Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises
- We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target pr…
- We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap’s Typed Intensional FOL
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13073v1 公告类型:新。
- 摘要:基于 $IFOL_B$ 的神经符号人工智能是一种将神经学习和符号推理相结合的方法,以克服纯神经系统的局限性(例如缺乏可解释性和逻辑结构),并具有用于自我参考的形式逻辑机制。
- 在本文中,我们基于 Nilsson 的 $IFOL_B$ 概率结构,通过对当前未知句子进行概率计算来扩展 $IFOL_B$ 的认知能力。
- 我们引入了保留当前知识数据库和逻辑推导的全局对称变换,以及用于对仅涉及 $IFOL_B$ 谓词的非常严格子集的具体(子)问题进行实时决策的局部对称变换。
- EN 要点:
- arXiv:2607.13073v1 Announce Type: new
- Abstract: Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like…
- In this paper we expand the cognitive power of $IFOL_B$ by using the probability computation for the currently unknown sentences, based on Nilsson’s probability…
- We introduce the global symmetry transformation that preserves the current knowledge database and logical deduction, and the local one used for real-time decisi…
Self-Improvements in Modern Agentic Systems: A Survey
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13104v1 公告类型:新。
- 摘要:自我改进的自主代理正在从研究原型转向部署系统。
- 主要目标是通过最少甚至没有人类输入的经验进行可控进化或适应。
- 这项调查将现代自我改进代理视为自适应系统,将经验转化为累积的能力增益。
- EN 要点:
- arXiv:2607.13104v1 Announce Type: new
- Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems
- The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input
- This survey frames modern self-improving agents as adaptive systems that convert experience into accumulated capability gains
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13115v1 公告类型:新。
-摘要:小语言模型(SLM)已显示出从 SMILES 字符串进行零样本分子属性预测的希望,但由于序列表示未指定关键的图拓扑线索,它们经常遭受结构失明的困扰。
- 我们提出了一个模块化的上下文增强提示框架,可以在推理时使用代理工具:经过训练的 GNN 专家模型可以自信地提供预测提示,并且 GNN 提取特定于实例的解释子图(例如,子图 SMILES 和随附的解释性段落)。
- 我们在五种提示配置(从仅 SMILES 到使用手头所有可用工具)下评估 MUTAG 和 Tox21 上的三种常用 SLM。
- EN 要点:
- arXiv:2607.13115v1 Announce Type: new
- Abstract: Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from structural…
- We propose a modular Context-Augmented Prompting framework that enables agentic tool use at inference time: a trained GNN expert model provides a predictive hin…
- We evaluate three commonly used SLMs on MUTAG and Tox21 under five prompting configurations ranging from SMILES-only to using all available tools at hand
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13157v1 公告类型:新。
- 摘要:智能体记忆是长视野智能体的一个系统问题。
- 实际部署需要在扩展对话中保留任务状态,在会话中恢复用户特定的事实和偏好,以及从先前结果中积累程序知识。
- 这些要求超出了文档检索的范围:内存层必须确定哪些交互成为持久状态、如何确定该状态的范围、如何在延迟限制下检索它以及如何随着时间的推移对其进行修改或删除。
- EN 要点:
- arXiv:2607.13157v1 Announce Type: new
- Abstract: Agent memory is a systems problem for long-horizon agents
- Practical deployments require retention of task state across extended conversations, recovery of user-specific facts and preferences across sessions, and accumu…
- These requirements extend beyond document retrieval: a memory layer must determine which interactions become durable state, how that state is scoped, how it is…
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13172v1 公告类型:新。
- 摘要:我们解决了在环境动态未知且没有合适的奖励函数可用的情况下安全训练代理策略并部署良好且安全的策略的问题。
- 在安全关键环境的背景下,我们认为传统的强化学习不切实际,并求助于人力输入资源。
- 我们推出 DROPJ,一种以人为本的安全培训和部署方法。
- EN 要点:
- arXiv:2607.13172v1 Announce Type: new
- Abstract: We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics are unknown…
- In the context of safety-critical environments, we consider traditional reinforcement learning impractical and resort to the resource of human input
- We introduce DROPJ, a human-centred method for both safe training and deployment
CayleyR: Solving the TopSpin puzzle via cycle intersection
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13219v1 公告类型:新。
- 摘要:我们提出了 cayleyR,这是一个 R 包,用于通过检测凯莱图中的循环交点来解决排列难题。
- 核心算法执行迭代双向搜索:从初始排列状态和目标排列状态,随机操作序列在对称群Sn的凯莱图中生成循环;它们的交叉点产生了一条连接路径。
- 当没有找到直接交叉点时,距离引导的桥梁选择会缩小差距,然后重复该过程。
- EN 要点:
- arXiv:2607.13219v1 Announce Type: new
- Abstract: We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs
- The core algorithm performs an iterative bidirectional search: from both the initial and target permutation states, random operation sequences generate cycles i…
- When no direct intersection is found, a distance-guided bridge selection narrows the gap, and the process repeats
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.13220v2 公告类型:新。
- 摘要:大多数人工智能科学系统专注于通过使用更好的模型、更大的上下文窗口、长视野代理执行或与一个主要用户合作的数字联合科学家来扩展单一推理过程。
- 然而,具有挑战性的科学问题很少仅靠一个推理者就能解决。
- 它们由团队成员解决,团队成员具有不同的先验、实验背景、隐性知识和领域训练的直觉。
- EN 要点:
- arXiv:2607.13220v2 Announce Type: new
- Abstract: Most AI-for-science systems focus on scaling a single reasoning process by using better models, larger context windows, long-horizon agentic execution…
- However, challenging scientific problems are rarely solved by one reasoner alone
- They are solved by teams whose members carry different priors, experimental background, tacit knowledge, and domain-trained intuitions
ArXiv cs.CL (B_intro+search) 链接到标题
Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14099v1 公告类型:新。
- 摘要:在现实环境中部署视觉语言模型(VLM)不仅需要强大的视觉推理能力,还需要在持续的对话压力下保持稳定性。
- 我们引入了 Just Keep Prompting (JKP),这是一种多轮评估框架,当用户反复挑战、质疑或反驳模型的答案时,可以衡量 VLM 认知稳定性。
- JKP 使用三种策略探测模型最多 10 个后续回合:对抗性否定(反复拒绝)、纯粹苏格拉底式审讯(反复要求重新评估确定性)和上下文感知苏格拉底式总结(在要求重新考虑之前反映模型的先前基本原理)。
- EN 要点:
- arXiv:2607.14099v1 Announce Type: new
- Abstract: Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained conversati…
- We introduce Just Keep Prompting (JKP), a multi-turn evaluation framework that measures VLM epistemic stability when users repeatedly challenge, question, or co…
- JKP probes models for up to 10 follow-up turns using three strategies: Adversarial Negation (repeated rejection), Pure Socratic Interrogation (repeated calls to…
Quantum Compositional NLP for Arabic: Grammar, Morphology, and Word Sense in Circuit Topology
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14100v1 公告类型:新。
- 摘要:我们首次将基于预组语法的量子组合自然语言处理(QNLP)应用于阿拉伯语;一种形态丰富、自由字序的语言,其结构复杂性为量子电路中的意义构成理论提供了一个独特且苛刻的测试平台。
- 我们的系统将阿拉伯语句子转换为量子电路,其拓扑反映了语法结构:主语、动词和宾语成为量子门,它们之间的类型依赖性(预组语法)决定了这些门如何连接在一起。
- 我们进行了三项涵盖词序、形态时态和动词意义消歧的受控实验,将量子电路方法与经典基线(包括 AraVec(阿拉伯语词嵌入)和 AraBERT(预先训练的阿拉伯语转换器))进行比较。
- EN 要点:
- arXiv:2607.14100v1 Announce Type: new
- Abstract: We present the first application of pregroup grammar-based quantum compositional natural language processing (QNLP) to Arabic; a morphologically rich,…
- Our system converts Arabic sentences into quantum circuits whose topology mirrors grammatical structure: subjects, verbs, and objects become quantum gates, and…
- We conduct three controlled experiments spanning word order, morphological tense, and verb sense disambiguation, comparing quantum circuit methods against class…
LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14101v1 公告类型:新。
-摘要:在硬标签场景中,以低查询预算生成高质量的对抗性文本仍然是一个具有挑战性的问题。
- 大多数现有方法依赖于贪婪算法,其中选择文本中的一个位置进行替换,然后替换其他位置。
- 这种本地搜索方法可能无法发现高质量的对抗性示例,并且常常导致查询成本过高。
- EN 要点:
- arXiv:2607.14101v1 Announce Type: new
- Abstract: Generating high-quality adversarial texts with low query budgets remains a challenging problem in the hard-label scenario
- Most existing approaches rely on greedy algorithms, where one position in the text is selected for substitution, followed by the substitutions of other position…
- This local search approach may fail to discover high-quality adversarial examples and often leads to excessive query costs
UniSAGE: Unifying Static and Dynamic Attributes with Hyper-Structure
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14102v1 公告类型:新。
- 摘要:随着数字数据的快速增长,现实世界的应用程序越来越多地涉及将静态属性与动态记录相结合的分层信息。
- 以统一且可概括的方式对此类异构数据进行建模仍然具有挑战性。
- 现有方法通常依赖于大量的手动设计,与特定的数据模式紧密耦合,并且通常单独处理静态和动态属性,从而忽略了它们隐式的交互。
- EN 要点:
- arXiv:2607.14102v1 Announce Type: new
- Abstract: With the rapid growth of digital data, real-world applications increasingly involve hierarchical information that combines static attributes with dyna…
- Modeling such heterogeneous data in a unified and generalizable manner remains challenging
- Existing approaches often rely on extensive manual design, are tightly coupled to specific data schemas, and typically process static and dynamic attributes in…
Latent Communication Between Language Model Agents: Channels, Alignment, and the Limits of Text
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14103v1 公告类型:新。
- 摘要:多智能体系统 (MAS) 在许多环境和许多行业中得到应用。
- 这些 MAS 依赖于代理间通信,通常通过明文消息传递来实现。
- 我们假设,当需要传达复杂的概念时,大型语言模型可能拥有一个超出文本表达能力的世界模型。
- EN 要点:
- arXiv:2607.14103v1 Announce Type: new
- Abstract: Multi-agent systems (MAS) are utilized in many contexts and many professions
- Those MAS rely on inter-agent communication, usually implemented by clear-text message passing
- We hypothesize that Large Language Models may have a world model at their disposal that exceeds expressibility in text when complex concepts need to be communic…
UzWordnet and Generative AI for Learning Uzbek by Game Playing
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14104v1 公告类型:新。
- 摘要:本文提出了一种教育系统架构,使学习者能够通过游戏来练习乌兹别克语。
- 该架构集成了 UzWordnet 和目前最大的乌兹别克语拼字词典作为核心词汇资源,以及生成人工智能作为学习支持的基本组成部分。
- 我们设计了四种教育游戏来促进乌兹别克语学习,并提出了一种基于游戏的方法来改进 UzWordnet,作为游戏动态的直接副产品。
- EN 要点:
- arXiv:2607.14104v1 Announce Type: new
- Abstract: This paper presents an educational system architecture that enables learners to practice the Uzbek language through game-playing
- The architecture integrates UzWordnet and the largest currently available orthographic dictionary for Uzbek as core lexical resources, together with generative…
- We design four educational games to facilitate Uzbek language learning and propose a game-based methodology for improving UzWordnet as a direct by-product of ga…
Automatically Evolving Prompt Guidelines for Task-Specific Optimization
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14105v1 公告类型:新。
- 摘要:为了使大型语言模型可靠地回答用户查询,用户必须明确指定需求、上下文和约束。
- 然而,在实践中,用户查询通常未指定,迫使模型推断出可能与实际用户意图不一致的未声明的假设。
- 现有的即时工程指南旨在缓解这个问题,它们通常是通用的且与任务无关,限制了它们的实际效用。
- EN 要点:
- arXiv:2607.14105v1 Announce Type: new
- Abstract: For Large Language Models to reliably answer user queries, users must clearly specify requirements, context, and constraints
- In practice, however, user queries are often underspecified, forcing models to infer unstated assumptions that may misalign with the actual user intent
- Existing prompt engineering guidelines aim to mitigate this issue, they are typically generic and task-agnostic, limiting their practical utility
Token Time Continuous Diffusion for Language Modeling
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14106v1 公告类型:新。
-摘要:在本文中,我们介绍了令牌时间连续扩散(TTCD),这是一种新的扩散语言模型,它(a)在连续空间中运行,确定性地将高斯噪声映射到最终令牌画布,无需进一步采样,并且至关重要的是(b)结合了每个令牌时间的新概念,一些令牌以比其他令牌更快的速度从噪声到令牌。
- 连续空间建模有助于 TTCD 避免多个标记的并行采样,这是纯粹在离散空间中迭代的模型高速加速时不准确的一个关键来源。
- 每个令牌时间的概念有助于 TTCD 更好地模拟条件生成,允许更确定的令牌以更快的速度进行,并允许在细化过程中区分令牌间的影响。
- EN 要点:
- arXiv:2607.14106v1 Announce Type: new
- Abstract: In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, deterministi…
- Continuous space modeling helps TTCD avoid the parallel sampling of multiple tokens, which is a key source of inaccuracy at high speedups for models that iterat…
- The notion of per-token times helps TTCD to better model conditional generation, allows for more sure tokens to proceed at a faster rate, and allows for differe…
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14107v1 公告类型:新。
- 摘要:扩散大语言模型 (dLLM) 的推理效率受到两个挑战的限制:双向注意力阻碍了有效的 KV 缓存重用,而使用静态置信阈值增加解码并行性可能会损害生成质量。
- 我们观察到这两个挑战都源于一个共同的现象:当令牌被解码时,它们通过双向注意力的上下文整合导致令牌表示在解码步骤中漂移(进化)。
- 这种见解激发了 Polestar,这是一个免训练的推理框架,它使用令牌表示漂移作为统一信号来共同应对这两个挑战。
- EN 要点:
- arXiv:2607.14107v1 Announce Type: new
- Abstract: The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-c…
- We observe that both challenges arise from a shared phenomenon: as tokens are decoded, their contextual integration through bidirectional attention causes token…
- This insight motivates Polestar, a training-free inference framework that uses token representation drift as a unified signal to jointly address both challenges
Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14108v1 公告类型:新。
- 摘要:本文介绍了工具效率,这是一种新的定量指标,用于评估 LLM 代理轨迹中有用工具调用率。
- 为了确保工具效率得到明确定义,我们还引入了边际工具效用,这是每个工具调用定义的新定量指标,指示工具是否有用或是否可以安全地从工具套件中删除而不影响准确性,同时提高工具效率;在本文中,我们使用 LLM-as-a-Judge 确定轨迹中每个工具调用的边际工具效用符号。
- 虽然之前已经做了很多工作来开发改进法学硕士工具使用的技术,并设计使用准确性作为代理间接测量效率的评估方法,但我们的工作重点是通过本文在事后轨迹分析中提出的定量指标直接测量效率。
- EN 要点:
- arXiv:2607.14108v1 Announce Type: new
- Abstract: This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory
- To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined per tool call indicating whether a to…
- While much prior work has been done to develop techniques that improve tool use by LLMs and design evaluation methods measuring efficiency indirectly using accu…
ArXiv cs.LG (B_intro+search) 链接到标题
Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14123v1 公告类型:新。
- 摘要:尽管可解释人工智能(XAI)技术不断涌现——从特征归因到稀疏自动编码器——解释很少影响现实世界的工作流程。
- 在实践中,它们常常在没有指导有意义的行动的情况下产生和丢弃。
- 这一差距反映了根本性的缺陷:研究尚未建立将解释整合到端到端、人机交互系统中的方法。
- EN 要点:
- arXiv:2607.14123v1 Announce Type: new
- Abstract: Despite the proliferation of Explainable AI (XAI) techniques – from feature attributions to sparse autoencoders – explanations rarely influence real…
- In practice, they are often generated and discarded without guiding meaningful action
- This gap reflects foundational shortcomings: research has not yet established methodologies for integrating explanations into end-to-end, human-in-the-loop syst…
CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Black-Box Vision-Language Models
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14125v1 公告类型:新。
-摘要:预训练的视觉语言模型(VLM)通过计算图像和文本描述之间的相似性分数来实现零样本图像分类,通常通过将类标签(例如“猫”)插入到提示(例如“a的照片”)中来形成。
- 由于给定图像类对的分数对提示的选择很敏感,因此现有研究使用加权向量来集成多个提示,以汇总不同提示的分数。
- 然而,在当前的策略中,分配给每个提示的权重向量在所有类之间共享,隐含地假设提示有条件地独立于类,这在实践中通常不成立,因为像“鸟瞰图”这样的提示可能适合“机场”,但不适合“苹果”。
- EN 要点:
- arXiv:2607.14125v1 Announce Type: new
- Abstract: Pre-trained vision-language models (VLMs) enable zero-shot image classification by computing the similarity score between an image and textual descrip…
- Since the score for a given image-class pair is sensitive to the choice of prompt, existing studies ensemble multiple prompts using a weighting vector to aggreg…
- Yet, in current strategies, the weighting vector assigned to each prompt is shared across all classes, implicitly assuming that prompts are conditionally indepe…
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14127v1 公告类型:新。
- 摘要:代表性杂波高度 (RCH) 是无线电传播和干扰分析中的关键参数,因为它捕获了导致终端杂波损耗的局部障碍物的主要高度。
- 目前的做法通常依赖于ITU-R P.452-18建议书中分配给土地使用类别的固定杂波高度,但这在类别变化范围内有所遗漏,并可能导致保守的禁区以及低地球轨道地面站选址和频谱协调的站点排名不佳。
- 我们提出了一个可解释的、可在全球部署的机器学习框架,用于根据开放地理空间数据预测 RCH。
- EN 要点:
- arXiv:2607.14127v1 Announce Type: new
- Abstract: Representative clutter height (RCH) is a key parameter in radio propagation and interference analysis because it captures the dominant height of local…
- Current practice often relies on fixed clutter heights assigned to land use classes in Recommendation ITU-R P.452-18, but this misses within class variation and…
- We present an interpretable, globally deployable machine learning framework for predicting RCH from open geospatial data
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14157v1 公告类型:新。
-摘要:对混合多个领域的语料库进行检索通常会返回相关但错误的领域证据,这些证据表明排名指标缺失,并且保形风险控制范围仅很小,无法覆盖最差的领域。
- 这项工作引入了 C3R,一个嵌入式控制层,它根据推断的域后验和无查询时间标签,在可行的情况下验证每个域的污染预算,否则放弃而不是默默地违反;在最困难的领域,它保证减少,而不是严格限制。
- 核心是基于风险控制预测集的二分方案,其有限样本转移边界从推断域跨越到具有完全可估计松弛的真实域,支持异构预算和部署反转。
- EN 要点:
- arXiv:2607.14157v1 Announce Type: new
- Abstract: Retrieval over corpora that mix several domains often returns relevant but wrong-domain evidence that ranking metrics miss and that conformal risk con…
- This work introduces C3R, a drop-in control layer that, from an inferred domain posterior and no query-time label, certifies a per-domain contamination budget w…
- The core is a two-split scheme built on risk-controlling prediction sets, whose finite-sample transfer bound crosses from the inferred to the true domain with f…
QFireNet: A Quantum-Enhanced U-Net for Wildfire Segmentation from Sentinel-2 Imagery
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14160v1 公告类型:新。
-摘要:卫星图像野火检测是一个语义图像分割问题,由于类别不平衡、特征复杂性和大气干扰等挑战,该问题已被证明是困难的。
- 在本文中,我们在基本的 U-Net 图像分割模型的基础上开发了一种量子混合解决方案,希望能够更有效地对 Sen2Fire 数据集的高维光谱特征空间进行建模。
- 我们在 U-Net 的瓶颈部分注入变分量子电路,特别是 QuFeX 和 QB-Net ansatze。
- EN 要点:
- arXiv:2607.14160v1 Announce Type: new
- Abstract: Wildfire detection from satellite imagery is a semantic image segmentation problem that has proven to be difficult due to challenges such as class imb…
- In this paper, we build on the foundational U-Net image segmentation model to develop a quantum-hybrid solution in hopes of more effectively modeling the high-d…
- We inject a variational quantum circuit in the bottleneck portion of U-Net, specifically the QuFeX and QB-Net ansatzes
Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14171v1 公告类型:新。
- 摘要:强化学习已成为训练与可执行沙箱交互的大型语言模型(LLM)代理的主导范例。
- 最先进的算法,如 PPO、RLOO 和 GRPO 继承了 RLHF 的推出拓扑:对于每个提示,从初始状态采样 N 个独立轨迹,并通过减去组基线来计算优势。
- 此设计忽略了代理沙箱的定义属性。
- EN 要点:
- arXiv:2607.14171v1 Announce Type: new
- Abstract: Reinforcement learning has emerged as the dominant paradigm for training large language model (LLM) agents that interact with executable sandboxes
- State-of-the-art algorithms such as PPO, RLOO, and GRPO inherit their rollout topology from RLHF: for each prompt, N independent trajectories are sampled from t…
- This design ignores a defining property of agent sandboxes
How Much of a 10-K Matters? Aggregation-Dependent Value of Full-Text versus Risk-Factor Sentiment
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14174v1 公告类型:新。
- 摘要:金融情绪提取在很大程度上依赖于新闻文本和仅针对退货标签的监督提取,而 10-K 归档和波动性、目标风险披露可以说是最适合通知的相对未经探索。
- 我们将监督词典学习方法扩展到 10-K 文件及其第 1A 项风险因素部分,在三个聚合级别(行业、投资组合和个体公司)针对回报和波动性标签训练情绪得分。
- 通过来自 94 家纳斯达克 100 科技成分股的 1,383 份申请(2006–2023 年),我们评估了最终的 12 种情绪指标,包括分类准确性、与已实现的市场结果的相关性以及定性词汇内容。
- EN 要点:
- arXiv:2607.14174v1 Announce Type: new
- Abstract: Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings – and vola…
- We extend a supervised lexicon-learning approach to 10-K filings and their Item 1A risk-factor sections, training sentiment scores against both return and volat…
- Across 1,383 filings from 94 Nasdaq-100 technology constituents (2006–2023), we evaluate the resulting twelve sentiment metrics on classification accuracy, cor…
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14176v1 公告类型:新。
- 摘要:可靠、低延迟的上行链路连接是密集城市环境中 C-V2X 网络的关键要求,在这些环境中,快速的信道变化和阻塞往往会降低车辆到基础设施的直接链路。
- 多跳中继可以恢复覆盖范围,但在无线电、容量和路由约束下的中继链路激活会导致 NP 难优化问题,通常通过混合整数线性规划 (MILP) 来解决,其运行时间随图形大小的扩展很差。
- 本文介绍了一种用于实时继电器选择的边缘感知学习优化框架。
- EN 要点:
- arXiv:2607.14176v1 Announce Type: new
- Abstract: Reliable, low-latency uplink connectivity is a key requirement for C-V2X networks in dense urban environments, where fast channel variations and block…
- Multi-hop relaying can restore coverage, but relay-link activation under radio, capacity, and routing constraints results in an NP-hard optimisation problem, ty…
- This paper introduces an edge-aware Learning-to-Optimise framework for real-time relay selection
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14180v1 公告类型:新。
- 摘要:世界模型广泛应用于离线强化学习(RL)中,以提高样本效率并生成超出固定数据集的经验。
- 然而,它们很容易受到数据覆盖范围薄弱的模型利用的影响。
- 先前的工作通过收集更多的专家演示来解决这个问题,这通常是昂贵的、不安全的或不可用的,或者通过避免不确定区域的保守算法来解决,这限制了泛化。
- EN 要点:
- arXiv:2607.14180v1 Announce Type: new
- Abstract: World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset
- However, they are vulnerable to model exploitation where data coverage is thin
- Prior work addresses this either by collecting more expert demonstrations, which is often expensive, unsafe, or unavailable, or by conservative algorithms that…
Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape
- 发布时间:2026-07-17 12:00 北京时间
- 摘要:- arXiv:2607.14185v1 公告类型:新。
-摘要:反馈驱动的循环支持大型语言模型、强化学习和自主发现的迭代改进,但在重复的内部反馈下,它们的收益往往会减少。
- 我们研究为什么闭环知识系统会饱和,以及哪些外部信息可以使它们超越当前的吸引子。
- 我们引入了一个三级操作框架,其中知识状态 $x_t$ 通过由结构参数 $\theta$ 索引的转换内核 $K_{\theta}$ 演化。
- EN 要点:
- arXiv:2607.14185v1 Announce Type: new
- Abstract: Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often…
- We study why closed-loop knowledge systems saturate and what external information can move them beyond their current attractors
- We introduce a three-level operational framework in which knowledge states $x_t$ evolve through transition kernels $K_{\theta}$ indexed by a structural paramete…