🤖 AI 速览

今天的主线是 OpenAI 将 GPT-5.6 接入 Microsoft 365 Copilot,并推出 ChatGPT Work,模型能力开始直接嵌入文档、表格、幻灯片等企业工作流。与此同时,智能体落地的关注点从“能否完成任务”转向执行轨迹、权限边界、编排成本与可验证性。开发者角色也在变化:更像工程经理,而不是单纯写代码的人。
📋 文章元数据
发布时间
2026-07-10
类型
ai-daily
字数
3531
阅读时长
17 min

2026-07-10 AI日更 | OpenAI 把 GPT-5.6 推进 Office:企业 AI 竞争进入工作流入口 链接到标题

今天的主线是 OpenAI 将 GPT-5.6 接入 Microsoft 365 Copilot,并推出 ChatGPT Work,模型能力开始直接嵌入文档、表格、幻灯片等企业工作流。与此同时,智能体落地的关注点从“能否完成任务”转向执行轨迹、权限边界、编排成本与可验证性。开发者角色也在变化:更像工程经理,而不是单纯写代码的人。

📖 本期 Watch List 深度导读 链接到标题

今天最值得先看的是 OpenAI 这一组更新:GPT-5.6 进入 Microsoft 365 Copilot,同时 ChatGPT Work 被定位为可跨应用交付文档、表格、幻灯片和 Web 应用的工作代理。这意味着“模型升级”正在迅速落到企业生产力入口,建议产品和工程团队重点关注其工作流边界、权限与交付质量。

第二条主线是智能体评估与成本控制。AgentLens 强调不要只看任务是否通过,而要审查完整执行轨迹;“Harness Effect”则提醒企业级 Agent 的真正成本杠杆在编排层,而非单纯等待 token 降价。这两篇适合正在做 Agent 落地的团队深读。

研究侧可关注推理与工具增强:上下文搜索理论、ARC 低成本 Agent、SageMath 增强数学代理,都在回答同一个问题——如何用更少试错、更可靠反馈获得可验证能力。另有 GPT-5.5 Bio Bug Bounty,值得安全与治理团队跟进。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:SpaceXAI Launches Grok 4.5 for Coding and Engineering Tasks 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:159000
  • 是什么事:X 上热议称 SpaceXAI 发布面向编码和工程任务的旗舰模型 Grok 4.5,主打高速推理、复杂软件开发能力和较低调用价格。
  • 为什么重要:如果属实,这将加剧 AI 编程助手和企业级模型市场竞争,尤其是在成本、速度、工程任务自动化和多模型平台整合方面,对 OpenAI、Anthropic 等厂商形成压力。
  • 讨论概况:讨论焦点集中在 Grok 4.5 的实际编程能力是否匹配宣传、低价策略能否改变开发者选择、AI 基础设施投入是否会拖累财务表现,以及相关并购和公司架构说法的真实性。

话题 2:Anthropic Adds /checkup Command to Clean Up Claude Code 链接到标题

  • 分类:AI · News
  • 概况:热度时间:22 hours ago,相关帖子数:2300
  • 是什么事:Anthropic 为 Claude Code 新增了“/checkup”命令,用于检查项目状态、清理上下文并帮助开发者整理代码工作流。
  • 为什么重要:这体现了 AI 编程工具正从单次代码生成走向持续协作与工程维护,强调上下文管理、代码质量和长期项目可用性。
  • 讨论概况:X 上的讨论主要集中在该功能能否减少 Claude Code 使用中的上下文混乱、提升大型项目开发效率;也有人质疑其实际效果、自动清理是否可能遗漏重要信息,以及与其他 AI 编程工具相比是否具备明显优势。

话题 3:OpenAI Launches GPT-Live for Natural Voice Conversations 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:55000
  • 是什么事:OpenAI 推出 GPT-Live 语音模型,使 ChatGPT 支持更自然的实时全双工语音对话,可边听边说、处理中断并提供实时翻译等功能。
  • 为什么重要:这标志着 AI 助手从文本问答进一步转向自然语音交互,可能改变用户与 AI 的使用方式,并推动语音模型、实时推理和多模态体验成为竞争重点。
  • 讨论概况:X 上讨论主要集中在 GPT-Live 是否真正接近人与人对话体验、全双工语音对客服和个人助理场景的影响、免费版与付费版能力差异,以及 OpenAI 在语音 AI 赛道是否会进一步扩大领先优势。

话题 4:OpenClaw Foundation Launches to Secure Open-Source AI Forever 链接到标题

  • 分类:AI · News
  • 概况:热度时间:20 hours ago,相关帖子数:691
  • 是什么事:OpenClaw Foundation 宣布成立,目标是通过基金会治理和长期资源支持,保障开源 AI 项目持续开放与可用。
  • 为什么重要:随着 AI 基础模型和工具链越来越集中于少数公司,开源 AI 的治理、资金、安全与长期维护成为行业关键议题,该基金会的出现被视为对开放生态可持续性的探索。
  • 讨论概况:X 上讨论主要集中在基金会是否能真正防止开源 AI 被商业化或封闭化、其治理结构是否透明可信,以及开放模型在安全风险与创新自由之间应如何平衡。

话题 5:ICML 2026 Spotlights Agentic AI Breakthroughs in Seoul 链接到标题

  • 分类:AI · News
  • 概况:热度时间:15 hours ago,相关帖子数:88
  • 是什么事:ICML 2026将在首尔举行,并将重点展示智能体式AI(Agentic AI)相关研究与突破。
  • 为什么重要:这表明AI研究重心正从单一模型能力扩展到具备规划、工具使用、协作与自主决策能力的系统,对下一代AI应用和安全治理具有重要影响。
  • 讨论概况:X上的讨论主要集中在Agentic AI是否会成为ICML 2026的核心方向、首尔承办对亚洲AI生态的意义,以及智能体系统在可靠性、评测标准和安全风险方面仍存在哪些挑战。

今日 X 上的 AI 舆情小结 链接到标题

今天的舆论主线是,AI 竞争正从“单一模型能力”转向更贴近真实使用场景的系统化能力:代码工程、实时语音、智能体协作、上下文维护和开源治理都成为焦点。共识在于,开发者与用户越来越看重低成本、高速度、长期可用性和自然交互体验,AI 工具正在从一次性生成走向持续协作和自主执行。分歧主要集中在各家发布或传闻中的能力是否名副其实,例如 Grok 4.5 的真实性与实际编程水平、Claude Code 新功能的工程价值、GPT-Live 是否真正接近人类对话,以及开源基金会能否保持透明和独立。潜在风险则包括过度宣传导致预期泡沫、智能体系统可靠性和安全评测不足、自动化工具遗漏关键上下文,以及开源生态在商业化、安全监管和长期资金之间难以平衡。

💡 大佬观点(Influencer Insights) 链接到标题

AI 行业每日情报速递 (07/09-07/10) 链接到标题

1. 今日核心热点:OpenAI 全系发布与模型军备竞赛白热化 链接到标题

今天的讨论几乎被 OpenAI 占据,导火索是 GPT-5.6 全系模型(Sol/Terra/Luna)正式公开上线,同时伴随着 ChatGPT 与 Codex 应用的大一统

  • 模型矩阵与平台大一统: 综合 @dotey 的深度分析,OpenAI 此次发布的三个模型中,Sol 为旗舰,主打复杂推理与自主工作;Terra 主打性价比;Luna 为轻量快速版。伴随模型发布的还有 ChatGPT Work 功能,让 AI 从聊天助手真正进化为跨应用执行任务的 Agent,能够连接 Google Drive、Slack 等工具。此举被 @dotey 评价为 OpenAI 向“超级应用”战略迈进的关键一步,旨在对标 Google Workspace 和 Microsoft 365,为即将到来的 IPO 讲好企业级故事。同时,GPT-Live 全双工语音模式的推出,将语音交互从“对讲机”升级为更自然的打断与并行处理,但网友们 @dotey 反馈其实测“mhmm”等回应过于频繁,稍显聒噪。

  • 头部竞赛进入“开大”阶段: @dotey 爆料并汇总了 GPT-6 即将于一个月内发布 的消息,指出这是 OpenAI 为回应 Anthropic 的 Mythos 模型 而直接跳过小版本迭代的举措。与此同时,Meta 的 Muse Spark 1.1 和 xAI 的 Grok 4.5 相继刷屏。

    • Grok 4.5 实测:@vista8 (@vista8) 给出了初评,认为其在独立 CLI 开发中不够全面,但在前端审美上优于 Codex,作为 Premium + 订阅的附加价值,含金量进一步提升。
    • Fable 5 的突然重置:@Pluvio9yte (@Pluvio9yte) 和多位博主发现,伴随 GPT-5.6 的上线,Claude Code 的 Fable 5 周额度突然被重置,被调侃为 Anthropic 的“防守型操作”。
  • 端侧与开源模型的暗流涌动: 在巨头打核战争的同时,端侧模型仍有突破。@zhixianio (@zhixianio) 深度测试了 Gemma 4 12B Coder 并与 Qwen 3.6 35B 对比,得出结论:虽然 12B 小模型在微调后效率极高,但受限于参数量,在处理“长程序、有状态、一次成型”的复杂程序时存在天花板。此外,中国的 腾讯混元 Hy3 模型(295B MoE)备受关注,@ruanyf 和 @vista8 均指出其以较小的参数体积达到了接近 GLM 5.1 的水平,凭借低成本优势适合日常高频使用。

2. 独特观点与行业前瞻 链接到标题

  • Vibe Coding 的新角色:工程经理: @dotey (@dotey) 提出,用 Coding Agent 开发时,开发者的角色已经从程序员转变为了工程经理(EM)。开发者需要负责拆解需求、分配任务和验收。如果你不审查代码,只关注功能,就像是一个不称职的 EM。同时,他建议采用 “持续集成”的思路做 Vibe Coding,一次只让 AI 做一个小功能点,方便验证和纠错,而不是一次性生成屎山。

  • Skill 管理的科学与非科学: @dotey 总结了一套 Skill 管理法则,引用 SkillsBench 的数据指出:AI 自生成的 Skill 甚至比完全不用 Skill 效果更差,只有人类专家引导产出的 Skill 才有价值;同时,大而全的 Skill 不如聚焦核心的小 Skill;软件工程类 Skill 对模型提升收效甚微,因为模型早已被喂饱,反而在医疗等长尾领域 Skill 提升巨大。

  • 努力需用对地方: 针对 AI 时代独立开发者,@gefei55 (@gefei55) 火力全开,指出以前一个月写一个没人用的 App,现在 AI 辅助下能写 37 个,但这只是“烧 Token 的陷阱”。真正缺乏的是需求调研、推广营销和敢于离开舒适区的勇气

  • 小红书与 Skill 分发: @ruanyf 发现 小红书正在内测 REDSkill 社区,允许用户将 AI Agent 的 Skill 文件分发至平台,试图结合社交媒体与 Skill Hub,打造“Skill 的 GitHub”。这为开发者接触海量非技术用户提供了全新渠道。

3. 推荐的工具与资源 链接到标题

  • 开源项目与发布

    • RN-Skill 合集 (@Pluvio9yte 开源):一套涵盖写作、视频制作与质检的 AI Agent Skill,特别适合处理 AI 文章的去味精修与动效视频导演。
    • 乔木 RSS 阅读器 (@vista8 开源):主打 AI 自动翻译与重写 Newsletter,整合了 Hacker News 等优质信源,适合缓解信息过载。
    • Topview 3D Shot Composer (@AI_Jasonyu 推荐):AI 视频创作中解决构图难题的新工具,支持先在 3D 空间摆好角色站位、机位,再由 AI 生成。
  • 开发辅助与插件

    • Obsidian → X 长文插件 (@kaitoxhacker 开发,@AI_Jasonyu 推荐):解决本地 Markdown 笔记一键发布到 X 平台的长文章痛点。
    • 公众号批量下载 Skill:一款基于 Python 标准库的工具,支持自动将微信公众号文章转为 Markdown、下载配图并建立索引。
  • 设计美学参考

    • Apple-Design Skill (@emilkowalski 发布,@vista8 推荐):总结了 17 条 Apple WWDC 视频中的设计与动效原则,能有效提升 AI 生成前端界面的审美。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 36 条更新

Y Combinator Podcast (B_intro+search) 链接到标题

  • How To Better Understand Your Users
    • 发布时间:2026-07-10 04:25 北京时间
    • 摘要:- 您可能已经听说过 OpenClaw(以前称为 Clawdbot/Moltbot)。
      • 引起轰动的开源人工智能助手可以在您自己的设备上运行,与您已经使用的消息应用程序连接,并且超越聊天功能,实际执行管理电子邮件、日历、文件、工作流程等任务。
      • 现在来认识一下它背后的人。
      • YC 的 Raphael Schaad 与 OpenClaw 的创始人 Peter Steinberger 坐下来讨论病毒式个人 AI 代理背后的“顿悟”时刻、为什么本地优先代理可以取代当今的许多应用程序,以及个人代理将如何重塑软件的未来。
    • EN 要点:
      • Most founders obsess over dashboards and aggregate metrics, but some of the best product insights come from understanding how individual users actually use thei…
      • In this episode of Startup School, YC’s David Lieb walks through one of his favorite tools for better understanding your users, the dot plot
      • It’s a simple two-dimensional grid that reveals usage patterns no aggregate chart can show you
      • He’ll cover why it gives founders a better sense of product health, what patterns to look for, and real-world exam

Stratechery by Ben Thompson (A_full) 链接到标题

  • Muse Image, Grok 4.5, Alex Karp on CNBC
    • 发布时间:2026-07-09 18:00 北京时间
    • 摘要:- 从 Meta 到 Grok 再到前沿实验室,可验证数据的争夺者越来越多地定义了人工智能竞赛。
      • 15 美元/月150 美元/年。
      • 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
      • 策略采访
      • 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
    • EN 要点:
      • The batter for verifiable data is increasingly defining the AI race, from Meta to Grok to the frontier labs.

OpenAI Blog (A_full) 链接到标题

  • GPT-5.6 is now the preferred model in Microsoft 365 Copilot

    • 发布时间:2026-07-09 21:00 北京时间
    • 摘要:- 今天,OpenAI 发布了 GPT‑5.6,它将成为 Microsoft 365 Copilot(Word、Excel、PowerPoint、Chat 和 Cowork)中的新首选模型。
      • 对于 Microsoft 365 客户,此次更新将 OpenAI 的最新旗舰模型系列引入人们日常使用的生产力工具中,帮助他们在工作流中利用更强大的 AI 辅助工具进行创建、分析和协作。
      • GPT-5.6是OpenAI的最新旗舰模型系列,它可以从每个代币中提供更多有用的工作,具有更强的性价比和针对最复杂任务的按需能力。
      • 借助 GPT‑5.6,Microsoft 365 用户将能够在他们已经依赖的应用程序中以更少的精力创建更高质量的工作产品:
        • 在 Word 中,GPT-5.6 可以帮助人们在更少的提示下起草、编辑和完善文档。
    • EN 要点:
      • Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, higher-quality work.
  • ChatGPT is now a partner for your most ambitious work

    • 发布时间:2026-07-09 18:00 北京时间
    • 摘要:- 隆重推出 ChatGPT Work,这是 ChatGPT 中的一个代理,可帮助您承担更艰巨的任务。
      • 它可以跨应用程序和工作流程收集信息,以创建工作表、幻灯片、文档和 Web 应用程序等成品材料,并通过将复杂的项目分解为更小的步骤并独立完成,在数小时内处理它们。
      • 借助内置的 Codex 技术,ChatGPT 现在不仅可以回答问题,还可以跨网络、移动设备和桌面完成实际工作。
      • 每周有超过 500 万人使用 Codex。
      • 尽管它最初是作为开发人员的编码代理,但现在有超过 100 万人使用它进行软件开发之外的工作,这表明它的功能如何支持更广泛的任务。
    • EN 要点:
      • ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.
  • GPT-5.5 Bio Bug Bounty

    • 发布时间:2026-07-09 18:00 北京时间
    • 摘要:- 有关 OpenAI Bio 赏金计划的详细信息。
      • OpenAI 博客的这篇文章解释了 GPT-5.5 Bio Bug Bounty 如何塑造更广泛的人工智能和基础设施景观。
      • GPT-5.5 Bio Bug Bounty 还为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • Details about the OpenAI Bio Bounty program
  • GPT-5.6: Frontier intelligence that scales with your ambition

    • 发布时间:2026-07-09 18:00 北京时间
    • 摘要:- 每个代币都具有更多智能,每美元的性能更强,并且为您最艰苦的工作提供更多所需的功能。
      • OpenAI 博客中的这篇文章解释了 GPT-5.6:随着您的雄心壮志而扩展的前沿智能如何塑造更广泛的人工智能和基础设施格局。
      • 它还为遵循 GPT-5.6 的创始人、运营商和投资者提供了实际影响:随着您的雄心壮志而扩展的前沿智能。
    • EN 要点:
      • More intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work.

ArXiv cs.AI (B_intro+search) 链接到标题

  • AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06624v1 公告类型:新。
      • 摘要:我们推出 AgentLens,这是交互式代码代理的生产评估基准。
      • 大多数代码代理基准测试将运行次数减少到一位 - 任务通过了吗?
        • 但实际使用这些代理的人会经历整个轨迹:代理如何遵循指令,使用其工具,验证自己的工作,从错误中恢复,并一路与他们交谈。
    • EN 要点:
      • arXiv:2607.06624v1 Announce Type: new
      • Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents
      • Most code-agent benchmarks reduce a run to a single bit – did the task pass
      • – but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, rec…
  • When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06720v1 公告类型:新。
      • 摘要:通过扩展推理训练大型语言模型 (LLM) 实现了上下文搜索,其中模型迭代地生成、批判和修改解决方案尝试。
      • 我们通过将上下文搜索建模为推理轨迹上的近似推理来提供上下文搜索的理论分析,其中基本模型定义先验,自我反思为后验更新提供反馈,并研究由此产生的推理时间采样复杂性 - 实现高成功概率所需的连续尝试次数。
      • 我们表明,当反射可靠地定位早期错误时,上下文搜索可以在基本模型上产生指数级改进,仅使用多项式数量的连续尝试来解决指数小零样本通过率的问题,而当此属性失败时,对过去尝试的条件与并行采样相比没有渐近优势。
    • EN 要点:
      • arXiv:2607.06720v1 Announce Type: new
      • Abstract: Training large language models (LLMs) with extended reasoning has enabled in-context search, in which models iteratively generate, critique, and revis…
      • We provide a theoretical analysis of in-context search by modeling it as approximate inference over reasoning traces, where the base model defines a prior and s…
      • We show that when reflections reliably localize early mistakes, in-context search can yield exponential improvements over the base model, solving problems with…
  • LLM-powered reasoning in agent-based modeling

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06757v1 公告类型:新。
      • 摘要:基于代理的建模(ABM)能够对数百万个人及其交互进行建模,这对于政策制定非常有用。
      • 然而,ABM 传统上依赖于静态先验,这阻碍了模型适应实时变化。
      • 我们的研究提供了一种解决这一信息差距的新方法。
    • EN 要点:
      • arXiv:2607.06757v1 Announce Type: new
      • Abstract: Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making
      • However, ABMs have traditionally relied on static prior, which prevents the models from adapting to real-time changes
      • Our research provides a novel approach to addressing this information gap
  • QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06760v1 公告类型:新。
      • 摘要:部分可观察性下的自治系统根据信念而不是原始传感器事件起作用。
      • QANTIS 将量子处理器视为该循环中的校准信念更新服务:它接收先验模型和观察模型,估计罕见事件证据项,并将普通后验返回给经典规划器。
      • 本文询问该服务是否可以在当前 IBM Heron 硬件上的顺序 Tiger POMDP 范围内重用,而不会破坏面向规划器的后验。
    • EN 要点:
      • arXiv:2607.06760v1 Announce Type: new
      • Abstract: Autonomous systems under partial observability act on beliefs, not raw sensor events
      • QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event e…
      • This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing p…
  • Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06764v1 公告类型:新。
      • 摘要:ARC-AGI-1 所公开架构的最新进展主要来自两个方面:前沿模型上的大量测试时计算(进化搜索、穷举采样、扩展思想链),或针对特定基准的训练,其中小模型在 ARC 数据上进行微调,通常采用特定于任务的架构。
      • 我们研究第三种制度:严格预算下的非思维模式下的开放权重模型(DeepSeek V3.2),没有特定于 ARC 的微调。
      • 我们研究仅通过架构可恢复的内容,构建显式分解模式发现和程序综合阶段的代理工具。
    • EN 要点:
      • arXiv:2607.06764v1 Announce Type: new
      • Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionar…
      • We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning
      • We study what is recoverable through architecture alone, building agentic harnesses that decompose pattern-discovery and program-synthesis stages explicitly
  • Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06820v1 公告类型:新。
      • 摘要:数学人工智能的最新进展主要集中在自动形式化和定理证明上,而计算机代数系统(CAS)在代理法学硕士工作流程中的作用尚未得到充分探索。
      • 我们提出了一种 ReAct 风格的代理设置,将 LLM 推理与来自 SageMath 的可验证反馈相结合,以及用于最新文档的 Context7。
      • 我们跨前沿模型评估这种代理设置,以在模拟计算数学研究循环的设置中解决 RealMath 基准的研究级数学问题。
    • EN 要点:
      • arXiv:2607.06820v1 Announce Type: new
      • Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS…
      • We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentati…
      • We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a…
  • The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06906v1 公告类型:新。
      • 摘要:今天的代理人工智能开发以代币最大化为基础:用代币购买能力——更长的推理轨迹、更多的回合、更广泛的工具有效负载、更大的重播上下文——因此每个任务的代币增长速度快于任务价值。
      • 每个代币价格的下跌掩盖了这一模式;总支出无论如何都会增加。
      • 我们认为,反对代币最大化的决定性杠杆是工具:编排层,它组装上下文、公开工具、顺序轮转、委派工作并承载企业可观察性和治理。
    • EN 要点:
      • arXiv:2607.06906v1 Announce Type: new
      • Abstract: Agentic AI development today runs on token maxing: buying capability with tokens – longer reasoning traces, more turns, wider tool payloads, bigger r…
      • Falling per-token prices mask the pattern; total spend rises anyway
      • We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work,…
  • Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06925v1 公告类型:新。
      • 摘要:以语言目标为条件的紧凑世界模型承诺使用一组稀疏的显式\emph{参考锚}来实现诸如“将红色块放在蓝色块左侧”等基础关系。
      • 我们询问此类参考何时真正建立关系,并识别出一个陷阱:目标条件预测器达到惊人的 0.90 美元关系读出精度,但这只是 \emph{指令转录},而不是感知。
      • 保留目标会使其崩溃($0.90!\to!0.27$,三个种子),并且反事实指令使预测的锚点遵循 \emph{false} 指令 $94.5%$(真实场景 $2.3%$;$N{=}256$)。
    • EN 要点:
      • arXiv:2607.06925v1 Announce Type: new
      • Abstract: Compact world models that condition on a language goal promise to ground relations such as ``put the red block left of the blue block’’ using a sparse…
      • We ask when such references actually ground a relation, and identify a trap: a goal-conditioned predictor reaches a striking $0.90$ relation-readout accuracy, y…
      • Withholding the goal collapses it to chance ($0.90\
  • Large Behavior Model: A Promptable Digital Twin of the Retail Customer

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06993v1 公告类型:新。 -摘要:客户行为建模是推荐、营销和决策支持的基础,但现有方法要么在不解释决策的情况下优化预测准确性,要么在不以真实行为数据为基础的情况下模拟用户。
      • 我们提出了大型行为模型(LBM),该模型通过统一的人环境公式直接从大规模零售交易中学习客户决策。
      • 客户状态由源自历史购买的行为档案来表示,而产品上下文则通过检索增强生成来合并。
    • EN 要点:
      • arXiv:2607.06993v1 Announce Type: new
      • Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy with…
      • We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environm…
      • Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented gener…
  • Learning social norms enhances compatibility in dynamic human-AI coordination

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.07021v1 公告类型:新。
      • 摘要:人类在动态互动中不断地与他人协调,通常是通过隐含的、难以量化的社会规范,这些社会规范充当交互主体之间共享的默认期望。
      • 随着包括大型语言模型 (LLM) 在内的人工智能代理融入日常生活,它们越来越多地参与此类交互并重塑社交交互结构。
      • 然而,它们常常无法以有效、体贴和自然的方式与人类协调。
    • EN 要点:
      • arXiv:2607.07021v1 Announce Type: new
      • Abstract: Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit exp…
      • As AI agents, including large language models (LLMs), become embedded in daily life, they increasingly participate in such interactions and reshape social inter…
      • Yet they often fail to coordinate with humans in an effective, considerate, and natural manner

ArXiv cs.CL (B_intro+search) 链接到标题

  • Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06611v1 公告类型:新。
      • 摘要:自动识别语音中的积极或消极情绪是一项具有挑战性的任务,需要分析声音变化和解释说出的单词。
      • 最近的解决方案依赖于音频基础模型来解决任务,但尚不清楚此类模型是否可以考虑所有方面。
      • 为此,我们提出了一种多模态解决方案,通过跨模态转换器集成音频和文本信息,其中文本转录通过自动语音识别(ASR)工具自动生成。
    • EN 要点:
      • arXiv:2607.06611v1 Announce Type: new
      • Abstract: Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and…
      • Recent solutions rely on audio foundation models to solve the task, but it remains unclear if such models can take all aspects into account
      • To this end, we propose a multimodal solution that integrates audio and text information via cross-modal transformers, where text transcripts are automatically…
  • Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06641v1 公告类型:新。 -摘要:大型语言模型(LLM)在医学问答基准上取得了有希望的结果,但它们在公共卫生中的使用受到幻觉和官方指导快速发展的限制。
      • 检索增强生成(RAG)通过将响应建立在明确维护的语料库中来减轻这些风险,但端到端性能主要取决于检索配置和超越多项选择格式的评估。
      • 我们将 PubHealthBench(源自英国政府公共卫生指南的 7,929 个问题的问答 (QA) 基准)扩展到检索增强环境中,并系统地评估检索和生成选择。
    • EN 要点:
      • arXiv:2607.06641v1 Announce Type: new
      • Abstract: Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hall…
      • Retrieval-Augmented Generation (RAG) mitigates these risks by grounding responses in an explicitly maintained corpus, but end-to-end performance depends critica…
      • We extend PubHealthBench, a question answering (QA) benchmark of 7,929 questions derived from UK Government public health guidance, into a retrieval-augmented s…
  • Ad Headline Generation using Self-Critical Masked Language Model

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06818v1 公告类型:新。
      • 摘要:对于任何电子商务网站来说,建立吸引购物者的持久广告都是一个不小的问题。
      • 很难通过网站的创意质量标准,尤其是在大型网站上。
      • 因此,我们提出了一种程序化解决方案,使用零售内容生成产品广告标题。
    • EN 要点:
      • arXiv:2607.06818v1 Announce Type: new
      • Abstract: For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers
      • It is hard to pass the creative quality bar of the website, especially at a large scale
      • We thus propose a programmatic solution to generate product advertising headlines using retail content
  • Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06831v1 公告类型:新。
      • 摘要:语音到文本对齐意味着找到音频中每个单词的时间边界。
      • 有些模型直接提供这样的对齐,而其他模型则不提供。
      • 连接主义时间分类(CTC)和转换器模型在构造上具有对齐,而基于注意力的编码器解码器(AED)和语音大语言模型(LLM)则没有,并且它们的单词计时通常是从注意力权重中读取的。
    • EN 要点:
      • arXiv:2607.06831v1 Announce Type: new
      • Abstract: Speech-to-text alignment means finding the temporal boundaries of each word in the audio
      • Some models provide such an alignment directly and others do not
      • Connectionist temporal classification (CTC) and transducer models have an alignment by construction, whereas attention-based encoder-decoders (AED) and speech l…
  • LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06845v1 公告类型:新。
      • 摘要:非裔美国英语 (AAE) 是一种受规则管辖的方言,有超过 3000 万人使用,经常被大型语言模型 (LLM) 误解和“纠正”。
      • 在六个指令调整的法学硕士(14B 到 70B)中,我们表明最先进的模型系统地更喜欢标准美式英语(SAE)延续,即使前面的上下文是 AAE,有效地将 AAE 重写为 SAE。
      • 我们提出了一个端到端框架来审核和减轻这种偏见。
    • EN 要点:
      • arXiv:2607.06845v1 Announce Type: new
      • Abstract: African American English (AAE), a rule-governed dialect spoken by over 30 million people, is routinely misinterpreted and “corrected” by large languag…
      • Across six instruction-tuned LLMs (14B to 70B), we show that state-of-the-art models systematically prefer Standard American English (SAE) continuations even wh…
      • We present an end-to-end framework to audit and mitigate this bias
  • Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06940v1 公告类型:新。
      • 摘要:大型语言模型(LLM)在语言任务中的出色表现凸显了对其响应质量进行综合评估的迫切需要。
      • 流行的方法通常局限于单一维度,无法捕获模型功能的全部范围。
      • 本研究引入了多因素评分范式,集成了准确性、简洁性、事实一致性、可读性和连贯性,并辅以用于可视化结果的图形用户界面 (GUI)。
    • EN 要点:
      • arXiv:2607.06940v1 Announce Type: new
      • Abstract: The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their respon…
      • Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities
      • This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a grap…
  • MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06974v1 公告类型:新。
      • 摘要:大型语言模型(LLM)通过额外的计算在测试时日益提高其推理能力,但大多数现有的工作都是孤立地处理每个问题的。
      • 当问题依次出现时,积累可重用的经验可以进一步提高性能。
      • 现有的基于内存的方法要么存储对新问题泛化能力较差的整体解决方案模板,要么使用未针对最终答案正确性进行优化的启发式步骤级别选择。
    • EN 要点:
      • arXiv:2607.06974v1 Announce Type: new
      • Abstract: Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem…
      • When problems arrive sequentially, accumulating reusable experience across them can further improve performance
      • Existing memory-based methods either store whole-solution templates that generalize poorly to novel problems or use heuristic step-level selection that is not o…
  • Riemannian Geometry for Pre-trained Language Model Embeddings

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.07047v1 公告类型:新。
      • 摘要:了解预训练语言模型嵌入的几何结构对于可解释性和安全性至关重要。
      • 我们询问句子级分类信号是否存在于上下文标记嵌入的黎曼几何中,并通过从学习编码器的分析雅可比行列式中提取每个标记的回拉度量并将它们与对称正定(SPD)流形上的 Fr’echet 均值聚合来探测它;我们将此过程称为黎曼均值池化 (RMP)。
      • 在具有重要语言结构的三个数据集(CoLA、CREAK、RTE)中,RMP 优于欧几里德均值池,而在 FEVER-Symmetric(为消除注释驱动的词汇伪影而构建的基准)上,该方法正确地保持了偶然性。
    • EN 要点:
      • arXiv:2607.07047v1 Announce Type: new
      • Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety
      • We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullba…
      • Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark co…
  • Behavior Leverage Imbalance in Multi-Teacher On-Policy Distillation

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.07050v1 公告类型:新。
      • 摘要:代理语言模型必须学习何时调用工具、何时使用工具响应以及何时直接回答。
      • 这使得多教师同策略蒸馏成为一种自然的培训策略:一名教师可以专门研究工具调用,另一名教师可以专门研究直接响应,学生可以从两者中学习。
      • 它自己生成的发行版。
    • EN 要点:
      • arXiv:2607.07050v1 Announce Type: new
      • Abstract: Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly
      • This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize in tool calls, another in direct responses, and the stud…
      • its own generated distribution
  • From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.07141v1 公告类型:新。
      • 摘要:新开发的项目通常必须在了解其心理测量特性之前进行现场测试,这给项目校准带来了冷启动问题。
      • 根据特征预测项目参数是一个长期存在的测量问题,可以追溯到线性逻辑测试模型;现代文本嵌入现在可以自动化传统上手动指定的设计矩阵。
      • 我们提出了一个评估框架,结合了项目文本嵌入的正则化回归、重复交叉验证的 R 平方报告及其重采样标准差以及两个性能上限:从参数标准误差得出的可靠性上限,以及从基于模拟的功率校准得出的设计上限。
    • EN 要点:
      • arXiv:2607.07141v1 Announce Type: new
      • Abstract: Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibrati…
      • Predicting item parameters from features is a long standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings now automa…
      • We propose an evaluation framework combining regularized regression on item text embeddings, repeated cross validated R squared reported with its resampling sta…

ArXiv cs.LG (B_intro+search) 链接到标题

  • TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06601v1 公告类型:新。
      • 摘要:条件计算可以将语言模型质量与每个标记的推理成本解耦,但领先的技术单独作用于单轴:专家混合 (MoE) 稀疏 FFN,深度混合 (MoD) 跳过整个转换器块,KV 缓存量化压缩注意力记忆。
      • 我们认为这三个决策(注意力分辨率、专家选择和缓存位宽)是强耦合的,应该联合做出:一个足够稀有、足以保证充分关注的令牌可能也需要高精度缓存,无论哪个专家处理它。
      • 我们引入了 TriRoute,一个跨所有三个轴共享的单个轻量级控制器,对于每一层的每个令牌,发出协调的策略:(i)注意力模式(跳过/本地/完整),(ii)一组稀疏的 FFN 专家(具有恢复 MoD 的空专家),以及(iii)KV 缓存位宽。
    • EN 要点:
      • arXiv:2607.06601v1 Announce Type: new
      • Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: M…
      • We argue these three decisions (attention resolution, expert selection, and cache bit-width) are strongly coupled and should be made jointly: a token rare enoug…
      • We introduce TriRoute, a single lightweight controller shared across all three axes that, for every token at every layer, emits a coordinated policy: (i) an att…
  • A Quiet Failure in Calibrated Virtual Screening: Marginal Conformal Prediction Under-Covers the Minority Class, and a Class-Conditional Fix Recovers It

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06605v1 公告类型:新。
      • 摘要:药物发现中采用保形预测来对模型可靠性给出一个诚实的数字:选择一个错误率 alpha,该方法以至少 1 - alpha 的概率返回包含真实标签的预测集。
      • 我们证明这种保证对于不平衡的数据集可能是危险的。
      • 在四个数据集中,标准(边际)适形预测达到了其全球 90% 的覆盖率目标,同时使少数类别严重暴露:实现的少数类别在血脑屏障渗透方面的覆盖率下降至 64.8%,在临床试验毒性方面下降至 4.2%,其中稀有类别几乎被放弃。
    • EN 要点:
      • arXiv:2607.06605v1 Announce Type: new
      • Abstract: Conformal prediction is being adopted in drug discovery to put an honest number on model reliability: pick an error rate alpha, and the method returns…
      • We show this guarantee can be dangerous on imbalanced datasets
      • Across four datasets, standard (marginal) conformal prediction hits its global 90% coverage target while leaving the minority class badly exposed: realized mino…
  • NEST: Tackling Dataset-Level Distribution Shifts via Regime-Oriented Mixture-of-Experts

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06607v1 公告类型:新。
      • 摘要:复杂系统中准确的长期预测经常受到数据集级别分布变化的影响,其中不同的潜在行为模式和不断发展的系统状态驱动动态多元时间序列。
      • 虽然现有方法主要关注局部时间变化,但它们无法明确模拟全球结构性挑战,其中数据集是不同操作机制的组合。
      • 在本文中,我们提出了 NEST,这是一个专门的框架,旨在通过两相密集 MoE 架构对这些不断演变的结构进行建模和重构。
    • EN 要点:
      • arXiv:2607.06607v1 Announce Type: new
      • Abstract: Accurate long-term forecasting in complex systems is frequently compromised by dataset-level distribution shifts, where diverse underlying behavioral…
      • While existing methods predominantly focus on local temporal shifts, they fail to explicitly model the global structural challenge where datasets are composites…
      • In this paper, we propose NEST, a specialized framework designed to model and recompose these evolving structures through a two-phase dense MoE architecture
  • D2PO: Optimizing Diffusion Samplers via Dynamic Preference

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06609v1 公告类型:新。 -摘要:我们提出了 D2PO(动态直接偏好优化),这是一个用于优化关于时间步计划和无分类器指导(CFG)权重的扩散采样策略的原则框架。
      • 我们的工作是由现有学生-教师回归框架的基本限制所推动的;低 NFE 学生采样器被训练来模仿高 NFE 教师,通常会牺牲高频纹理保真度,同时保留粗糙的全局结构,从而使采样器与感知质量错位。
      • D2PO 通过利用直接偏好优化 (DPO) 框架,将采样器优化重新表述为基于偏好的对齐问题,从而解决了这一挑战。
    • EN 要点:
      • arXiv:2607.06609v1 Announce Type: new
      • Abstract: We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to timestep s…
      • Our work is motivated by a fundamental limitation of existing student-teacher regression frameworks; low-NFE student samplers are trained to mimic high-NFEteach…
      • D2PO addresses this challenge by reformulating sampler optimization as a preference-based alignment problem, leveraging the Direct Preference Optimization (DPO)…
  • Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06610v1 公告类型:新。
      • 摘要:不确定性下的投资组合优化本质上是一个多目标决策问题,涉及回报、风险、市场动态和实际投资约束之间复杂的相互作用。
      • 现有的基于可靠性的投资组合优化方法主要依赖于静态优化框架,通常无法捕捉顺序决策、尾部风险和交易成本等市场摩擦。
      • 为了解决这些限制,我们提出了一种基于多目标可靠性的投资组合优化的深度强化学习框架(MORP-DRL)。
    • EN 要点:
      • arXiv:2607.06610v1 Announce Type: new
      • Abstract: Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dy…
      • Existing reliability based portfolio optimization approaches primarily rely on static optimization frameworks and often fail to capture sequential decision maki…
      • To address these limitations, we propose a deep reinforcement learning framework for multi-objective reliability based portfolio optimization (MORP-DRL)
  • STAGformer: A Spatio-temporal Agent Graph Transformer for Micro Mobility Demand Forecasting

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06614v1 公告类型:新。
      • 摘要:准确的站点级需求预测对于共享单车系统的高效运行至关重要,但由于复杂的时空依赖性和大规模的城市网络,它仍然具有挑战性。
      • 本文提出了 STAGformer,一种时空代理图转换器,可实现具有线性计算复杂度的高效全局建模。
      • 该模型引入了两步代理注意力机制,其中一小组可学习的空间和时间代理令牌首先聚合全局信息,然后将其广播回各个站点和时间步,有效捕获远程交互,同时降低标准自注意力 O(NT) 的二次成本。
    • EN 要点:
      • arXiv:2607.06614v1 Announce Type: new
      • Abstract: Accurate station-level demand forecasting is essential for the efficient operation of bike-sharing systems, yet it remains challenging due to complex…
      • This paper presents STAGformer, a Spatio-Temporal Agent Graph Transformer that achieves efficient global modeling with linear computational complexity
      • The model introduces a two-step agent attention mechanism, where a small set of learnable spatial and temporal agent tokens first aggregate global information a…
  • WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06616v1 公告类型:新。
      • 摘要:联邦学习(FL)中的标签偏差会导致客户端漂移并降低全局准确性。
      • 合成数据增强可以减少这种不平衡;然而,全类平衡需要大量的计算成本。
      • 我们提出 FedEAS,这是一种为每个客户端分配根据其本地标签分布计算的熵自适应每类生成预算的策略。
    • EN 要点:
      • arXiv:2607.06616v1 Announce Type: new
      • Abstract: Label skew in federated learning (FL) causes client drift and degrades global accuracy
      • Synthetic data augmentation can reduce this imbalance; however, full class balancing requires substantial computation cost
      • We propose FedEAS, a policy that assigns each client an entropy-adaptive per-class generation budget computed from its local label distribution
  • Inertia-1: An Open Exploration of Wearable Motion Foundation Models

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06617v1 公告类型:新。 -摘要:可穿戴运动传感为人类行为和健康提供了一个连续且可扩展的窗口,使其非常适合基础模型,但其预训练和扩展原理仍然知之甚少。
      • 先前的工作研究了孤立的设计选择,例如传感器放置或采样频率,通常是在固定设置和狭窄的下游任务下,无法捕获现实世界的传感多样性。
      • 我们推出 Inertia-1,这是对可穿戴运动基础模型的完全开放探索。
    • EN 要点:
      • arXiv:2607.06617v1 Announce Type: new
      • Abstract: Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation models, yet i…
      • Prior work studies isolated design choices, such as sensor placement or sampling frequency, often under fixed settings and narrow downstream tasks that fail to…
      • We introduce Inertia-1, a fully open exploration of wearable motion foundation models
  • Fingerprint, Not Blueprint: How Positional Schemes Set the Default Spectral Algebra of Attention

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06621v1 公告类型:新。
      • 摘要:注意力头的 pre-softmax 分数是学习算子 $M = W_q^T W_k$ 中的双线性形式 $score(i,j) = x_i^T M x_j$。
      • 由于 M 通常是非对称的,因此是非正态的,因此它具有复杂的特征谱和非正交特征向量,即应用非厄米特和随机矩阵工具的机制。
      • 我们询问该频谱在前代令牌和感应电路的三个级别上编码什么。
    • EN 要点:
      • arXiv:2607.06621v1 Announce Type: new
      • Abstract: The pre-softmax score of an attention head is a bilinear form $score(i,j) = x_i^T M x_j$ in a learned operator $M = W_q^T W_k$
      • Because M is generally non-symmetric, hence non-normal, it has a complex eigenspectrum and non-orthogonal eigenvectors, the regime where non-Hermitian and rando…
      • We ask what this spectrum encodes, at three levels for previous-token and induction circuits
  • LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting

    • 发布时间:2026-07-09 12:00 北京时间
    • 摘要:- arXiv:2607.06623v1 公告类型:新。
      • 摘要:流程工业依靠时间序列预测和软传感来估计难以在线测量的质量变量。
      • 标记数据稀缺,操作机制频繁变化,并且针对每个场景重新训练模型或重建对齐管道成本高昂。
      • 此类设置通常提供变量表和过程文档,记录变量名称、单位、物理含义和过程角色。
    • EN 要点:
      • arXiv:2607.06623v1 Announce Type: new
      • Abstract: Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online
      • Labeled data are scarce, operating regimes change frequently, and retraining models or rebuilding alignment pipelines for each scenario is costly
      • Such settings often provide variable tables and process documents that record variable names, units, physical meanings, and process roles