🤖 AI 速览

Google DeepMind同日发布Gemma 4 12B统一多模态模型与Gemini 3.5 Live实时翻译,端侧AI能力大幅超出预期;Anthropic推出Claude Fable 5,通过安全分类器将更高能力层级向公众分层开放,标志着模型部署正围绕“能力许可边界”构建新范式。同时,开发者社区转向以“结果工程”驱动AI编程,人与代理的协作正从提示词调试向目标定义升级。
📋 文章元数据
发布时间
2026-06-10
类型
ai-daily
字数
3744
阅读时长
18 min

2026-06-10 AI日更 | 从Gemma 4到Claude Fable,AI部署进入安全分级与端侧爆发双轨时代 链接到标题

Google DeepMind同日发布Gemma 4 12B统一多模态模型与Gemini 3.5 Live实时翻译,端侧AI能力大幅超出预期;Anthropic推出Claude Fable 5,通过安全分类器将更高能力层级向公众分层开放,标志着模型部署正围绕“能力许可边界”构建新范式。同时,开发者社区转向以“结果工程”驱动AI编程,人与代理的协作正从提示词调试向目标定义升级。

📖 本期 Watch List 深度导读 链接到标题

今天 Google DeepMind 连发三篇重磅内容,从 Gemini 3.5 Live Translate 近乎实时的自然语音翻译,到 Gemma 4 12B 这一统一无编码器的多模态模型,再到欧洲机器人技术的前瞻布局,非常清晰地在描绘 AI 基础设施与全球化落地的下一层拼图,建议生态参与者逐篇细读。

与此同时,Codex 正在重塑软件工程的底层逻辑。Nextdoor 和 Notion 的工程团队都给出了极具启发的实践复盘:他们不再把精力花在反复调试提示词上,而是转向“结果工程”——工程师直接定义想要达成的结果,与代理协作设计实现路径,甚至多年不写生产代码的管理者也重新回到了代码库。这几乎是开发者角色向上迁移的强烈信号,推荐所有工程团队关注。

最后,一篇关于 iPhone 最后堡垒的思考值得产品人注意。它一针见血地指出,Agent 的终局不是帮人操作计算机,而是让请求到结果之间的全部过程对用户不可见。当交互实现“隐身”,增长与体验的规则将被重写,值得结合安永对 Agentic AI 投资规则变化的洞察一同咀嚼。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Anthropic Launches Claude Fable 5 as Most Capable Public AI Model 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:107000
  • 是什么事:Anthropic 被曝发布 Claude Fable 5,定位为面向公众的首个 Mythos 级模型,能力层级高于 Opus,并内置针对双重用途网络威胁的安全护栏。
  • 为什么重要:此事标志着前沿 AI 模型不再只比拼智能高低,更开始围绕“能力许可边界”进行分层:将原先受限的顶级能力通过安全蒸馏和策略控制开放给公众,这或许定义了未来模型部署的新范式。
  • 讨论概况:讨论焦点集中在“护栏税”上——即公众版 Fable 为安全而牺牲了多少原始 Mythos 级的长程任务执行、编码和自主能力;用户还关注其能否在阻断恶意网络滥用的同时精准保留合法的安全研究功能,以及它与 Opus、受限 Mythos 的定位关系是否会重塑 Claude 的产品层级。

话题 2:Kimi AI Swarm Predicts Spain and France to Win 2026 World Cup 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:7 hours ago,相关帖子数:161
  • 是什么事:Kimi AI Swarm 系统预测西班牙和法国将赢得2026年世界杯冠军。
  • 为什么重要:此次预测展示了AI在复杂体育赛事预测中的潜力,推动数据分析与群体智能在非传统领域的应用,可能影响未来博彩、球队策略和公共观点形成。
  • 讨论概况:讨论焦点包括AI预测的准确性、是否过分依赖历史数据、西班牙与法国真实实力的评估,以及AI群体决策是否优于人类专家判断。

话题 3:Developers Debate Codex vs Claude Code AI Agents 链接到标题

  • 分类:AI · News
  • 概况:热度时间:9 hours ago,相关帖子数:877
  • 是什么事:开发者社区正在激烈辩论 OpenAI 的 Codex 与 Anthropic 的 Claude Code 哪一个是更优的 AI 编程代理工具,并热议两者的升级传闻。
  • 为什么重要:AI 编程代理是 AI 应用最成熟的领域之一,两大巨头的竞争直接塑造开发者工作流和工具标准,标志着 AI 从代码补全向深度工程协作演变。
  • 讨论概况:核心分歧在于:1) Claude Code 的类 Agent 能力在处理复杂工程任务上是否优于 Codex/Copilot;2) Codex 的高性价比与 Claude Code 高端配置下的性能优势如何取舍;3) 应该选择单一工具还是组合使用(如 Cursor 编排、Claude Code 深度执行、Codex 工作流审查)。

话题 4:Tesla’s FSD Supervised Gains Approval in Denmark, Fourth European Country 链接到标题

  • 分类:AI · News
  • 概况:热度时间:19 hours ago,相关帖子数:34000
  • 是什么事:特斯拉的全自动驾驶(监督版)功能在丹麦获得监管批准,成为继荷兰、挪威和瑞典之后第四个在欧洲获批上路的美国市场。
  • 为什么重要:这表明基于纯视觉与端到端神经网络的辅助驾驶方案正在突破以严格著称的欧洲监管壁垒,为 AI 在安全关键场景中的实际部署积累了合规经验,对全球自动驾驶商用化具有风向标意义。
  • 讨论概况:讨论焦点集中在:获批所依据的安全验证是否充分、‘完全自动驾驶’命名与实际能力之间的落差引发的误导争议、以及丹麦的决定是否会加快德国等主要欧盟市场的审批进程,也出现了对本地道路适配表现的期待与担忧。

话题 5:HeyGen Brings Video Creation to Claude AI with Hyperframes 链接到标题

  • 分类:AI · News
  • 概况:热度时间:,相关帖子数:985
  • 是什么事:HeyGen推出Hyperframes技术,将AI视频创作功能集成至Anthropic的Claude AI平台。
  • 为什么重要:此举标志着大型语言模型与AI视频生成工具的进一步融合,拓展了Claude的多模态内容生产能力,可能重塑创意工作流。
  • 讨论概况:X平台用户关注Hyperframes的技术实现、与Sora等竞品的差异、对视频创作者的影响,存在关于其实际效果和商业化潜力的不同看法。

话题 6:The Sandbox Launches AI Engine for Rapid Game Creation 链接到标题

  • 分类:AI · News
  • 概况:热度时间:,相关帖子数:102
  • 是什么事:元宇宙平台The Sandbox发布了一款AI引擎,支持用户快速生成游戏内容。
  • 为什么重要:这体现了生成式AI在交互式3D内容创作中的落地,有望大幅降低游戏开发门槛,加速虚拟世界构建。
  • 讨论概况:社区关注该引擎的实际生成质量、是否保留创作自由度,以及如何与现有创作者经济结合;部分讨论对其在NFT和元宇宙生态中的真正价值存有分歧。

话题 7:World Labs Launches ICARE: Browser-Based 3D Adventure with Genius Icons 链接到标题

  • 分类:AI · News
  • 概况:热度时间:5 hours ago,相关帖子数:75
  • 是什么事:World Labs 发布了一款名为 ICARE 的浏览器端 3D 冒险体验,用户可在其中与历史上多位“天才”形象的化身互动。
  • 为什么重要:这展示了空间智能与生成式 AI 结合在实时 3D 环境中的新可能,为沉浸式教育、AI 驱动的叙事和轻量化浏览器 3D 应用提供了前沿范例。
  • 讨论概况:X 上的讨论主要围绕其技术实现是否真正依赖生成式 3D、体验的趣味性与教育价值,以及 World Labs(李飞飞团队)在空间智能赛道的差异化路径;部分用户对“天才图标”的交互深度和浏览器的性能表现存在分歧。

今日 X 上的 AI 舆情小结 链接到标题

今日的舆论主线在于AI行业正从单纯比拼智力转向构建“能力许可边界”,通过Anthropic将更高能力层级的模型以安全蒸馏方式向公众开放,以及特斯拉FSD再突破欧洲严格监管等事件,显示出尖端模型在大众化过程中必须捆绑精密安全护栏的新范式。社区共识在于AI编程代理、多模态生成与预测分析等领域的实用化水平显著提升,然而对Codex与Claude Code谁更胜任复杂工程、Kimi世界杯预测是否过度依赖历史数据等技术路径的分歧,揭示出工具标准化远未形成。更深层的分歧集中在“护栏税”究竟牺牲了多少核心能力、“完全自动驾驶”命名是否构成误导,以及生成式3D/视频内容的真实质量与创作自由度的权衡上。这些争论背后潜藏着多维风险:过度安全限制可能削弱模型实用价值,验证不足的命名和低质生成内容会加剧信任赤字,而编程工具碎片化竞争与预测模型被不当依赖,则可能在安全关键场景和公共舆论中造成错判和监管反弹。

💡 大佬观点(Influencer Insights) 链接到标题

AI 行业动态日报(2026-06-09) 链接到标题

一、今日共同关注的技术趋势与产品热点 链接到标题

1. Claude Fable 5 发布:Anthropic 的"安全版 Mythos" 链接到标题

@dotey 详细解读了 Anthropic 同日发布的两款新模型:

  • Claude Fable 5:面向普通用户,内置安全分类器,检测到敏感内容时自动降级至 Opus 4.8 处理
  • Claude Mythos 5:去除部分安全限制,仅开放给 Project Glasswing 网络安全合作伙伴

关键数据:定价较 Mythos Preview 降低 60%(输入 $10/百万 token,输出 $50/百万),但仍是 Opus 4.8 的 2 倍。Pro/Max/Team/企业用户可在 6 月 22 日前免费使用,之后需购买 usage credits。

政策变化:Mythos 级别模型流量强制保留 30 天用于安全监控,打破此前零留存承诺。

2. 端侧模型爆发:Gemma 4 与本地 AI 生态成熟 链接到标题

@zhixianio 密集测试了 Google 的端侧布局:

  • Gemma 4 12B:统一多模态模型,在 M5 Max 128G 上通过 mlx-vlm 运行,英语/日语语音识别"秒出",中文表现欠佳
  • Gemma 4 E4B + MTP:日文邮件解析分类任务性能优异
  • QAT(量化感知训练):@zhixianio 指出这是"训练时就假定会被量化"的优化思路,Google 对端侧的重视预示"Android 快能用上自带模型"

实测反馈:@zhixianio 使用 Qwen3.6-35B-A3B-oQ6-fp16-mtp 在 oMLX 上运行,“响应速度比远程 LLM 快,智商在线”,端侧模型能力"大大超出预期"。

3. AI Agent 浏览器与自动化工具 链接到标题

@vista8 深度体验了 @okasupportgroup 开发的 Aye 浏览器

  • 基于 Chromium,完全 AI 模拟真人操作(非 CLI/插件,避免账号异常检测)
  • 内置 Skill 录制与定时执行,支持自动回复小红书评论、转写文章到多平台等
  • 集成 RSS 阅读器、广告拦截、视频翻译/下载

@vista8 建议:需支持 Chrome 账号迁移、插件生态,并尽快明确付费计划。


二、独特观点与行业前瞻 链接到标题

1. “Contract First”——Vibe Coding 的最佳实践 链接到标题

@Pluvio9yte 从安全从业者转型全栈开发的经验总结:

“Vibe Coding 的最佳实践其实并不是 Requirement First 或者 Code First,而是 Contract First。没有定义好契约,其他一切都是空谈。”

他基于 OpenSpec 二开了一套开发框架,将"容易漂移的上下文外化成契约",让人和 AI 都有稳定参照物。

2. 微信 AI 生态的"创新者窘境" 链接到标题

@dotey 连续发声批评微信的 AI 策略:

“微信总以为自己是 OS,但它不是,只是寄生在手机系统上的庞然大物…未来微信的入口属性会越来越少,以后的年轻人不会再去打开微信,只会问自己的 Agent。”

他建议微信应"用和目前团队在物理和财务上隔绝的团队,去做一个接近完全独立的 AI app"。

3. AI 编程的成本悖论 链接到标题

@ruanyf 引用 OpenClaw 创始人的 Token 消耗数据:

“一个月 6030 亿 Token,价值 130 万美元…就算改用便宜模型,一年也要 200-300 万人民币。公司会发现,如果无限量使用,AI 编程比真人程序员昂贵多了。”

4. 技能包(Skill)的封装形态探索 链接到标题

@lijigang 提出 LLM 发展的两条路径:

  • 往下走:原子化,拆分为具体任务的技能包
  • 往上走:组件化,封装场景最佳实践(workflow、节点优化、技能包)

并追问:“浏览器的扩展机制是不是可能的答案?”

5. “影子之书"阅读法 链接到标题

@lijigang 提出 AI 时代独有的阅读方式:

“印刷时代我们只能读作者写下的这一本书;AI 时代独有的动作,是把那些影子之书读出来——读到任何一个论断,立刻调用 AI 分析:三个反对它的学派、作者略过的前提、思想传承于谁…”


三、推荐工具与资源 链接到标题

开发工具 链接到标题

工具推荐人说明
oMLX v0.4.0@zhixianio原生 Swift macOS 端侧模型运行框架,支持 Native MTP
Owlia Nest@zhixianio为 PA(如 OpenClaw)设计的文件浏览工具,支持 Tailscale 内网访问、PWA、Markdown 在线编辑
Aye 浏览器@vista8AI Agent 专用浏览器,支持 Skill 录制与自动化网页操作
Glaze@vista8@raycast 新作,“一句话生成 Mac 软件并发布上架”,10 分钟开发音乐电台 App 实测
baoyu-design skill@dotey支持导入 Design System,保留 Claude Design 原始工作流

效率工具 链接到标题

工具推荐人说明
Bartender 6@Pluvio9yteMac 状态栏整理工具,$20 买断
Maccy@Pluvio9yte开源粘贴板工具
Screen Studio@Pluvio9yte录屏工具,带缩放动画(提示:闲鱼有低价)
Mos@Pluvio9yte鼠标滚动方向转换,开源免费
Perculia@vista8免费蓝牙管理工具,Menu Bar 一键切换设备
OpenWiki@AI_Jasonyu自动整理剪贴板内容为 Wiki,生成知识图谱,支持 MCP 连接 Claude Desktop

内容创作 链接到标题

工具推荐人说明
视频翻译工具(焚决)@Pluvio9yte本地视频一条龙处理:下载→转写→翻译→润色→烧字幕
抖音合规自检 Skill@Pluvio9yte (via @Zesee)检测"Claude Code"“GitHub 下载"等抖音违规词
书籍口播解读 Skill@vista8生成书籍口播脚本(待开源)

模型与 API 链接到标题

资源推荐人说明
DeepSeek-V4-Pro@zhixianio“成本和效果都有惊喜”,OpenClaw 场景推荐
MiniCPM5-1B@zhixianio (via @OpenBMB)2B 以下最强开源基座模型,AA 指数 17.9 超 Qwen3.5-2B
Codex 9.9 元试用@ruanyf150 美元 API 月度使用量,需境外网络

四、值得关注的数据点 链接到标题

  • GitHub 提交量:今年前三个月是去年同期 14 倍(@ruanyf)
  • 智谱市值:已等于小米,约两个京东,“世界市值最高的开源软件公司”(@ruanyf)
  • SEO 价值:某站点日自然搜索流量 1.1 万,月省 30 万营销费用(@gefei55)
  • 域名投资:.com 域名 10 几刀注册,半月后 1000 刀成交(@gefei55)

报告基于 2026-06-09 前后 24 小时 X 平台公开推文整理

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 37 条更新

All-In Podcast (A_full) 链接到标题

  • Bill Maris: How Google Could Crush AI Competitors, Why Small Funds Win, and AI’s Atari Stage
    • 发布时间:2026-06-09 23:07 北京时间
    • 摘要:- 安永 - Agentic AI 正在引入新的投资规则。
      • 随着人工智能转向基于消费的模型,安永将支出与企业价值联系起来。
      • 纽约证券交易所 - 感谢我们的合作伙伴纽约证券交易所 - 一个致力于建设未来的现代化市场和交易所。
      • Plaud,我们在 All-In Liquidity Summit 上的官方可穿戴人工智能笔记合作伙伴,捕捉到了每一个见解。
      • Bill Maris:谷歌如何碾压 AI 竞争对手、小基金为何获胜以及 AI 的雅达利舞台。
    • EN 要点:
      • (0:00) Bill Maris joins the Besties
      • (0:33) Four critical lessons from a career in technology
      • (5:58) Building Google Ventures with data and machine learning
      • (9:51) Why small VC funds beat big ones on average

Stratechery by Ben Thompson (A_full) 链接到标题

  • The iPhone’s Last Stand
    • 发布时间:2026-06-09 18:00 北京时间
    • 摘要:- 听这个帖子**:**。
      • 多年来,苹果粉丝都会嘲笑微软谈论可能会或可能不会发布的产品的倾向,嘲笑它们是雾件。 -> 当你考虑人工智能的下一波浪潮:代理时,这一点就更清楚了。
      • 代理的目的不是为您使用计算机;是为了完成某项特定任务。
      • 至少在理论上,请求和结果之间的所有内容都应该对用户不可见。
    • EN 要点:
      • Listen to this post :
      • Log in to listen
      • Apple fans would, for years and years, sneer at Microsoft’s penchant for talking about products that may or may not ship, deriding them as vaporware
      • After Apple’s bungled 2024 launch of Apple Intelligence and new Siri , however, vaporware is fair game, and just in time for this Article

OpenAI Blog (A_full) 链接到标题

  • How engineers at Nextdoor use Codex to build without limits

    • 发布时间:2026-06-09 20:00 北京时间
    • 摘要:- 像 Nextdoor 这样的产品为 11 个国家/地区的超过 1.1 亿用户提供服务,这对平台团队提出了很多要求。
      • 对于工程主管 Cory Dolphin 来说,Codex 代表了一个重要的转变:“从迭代地提示代理转向结果工程,工程师开始思考他们想要看到的结果,并与代理合作来设计该结果。”
      • 这意味着个别工程师在堆栈中向上移动 - 不再被锁定为某个系统或框架的专家,他们能够或多或少地拥有端到端的产品体验,甚至跨多个平台。
      • 生产力的提速如此之快,以至于瓶颈不再是工程,而是关于下一步要构建什么的棘手战略问题。 ->“Codex 从根本上改变了我们对工程的看法,以至于我们甚至无法想象没有它的工程。”。
    • EN 要点:
      • How engineers at Nextdoor use Codex with GPT-5.5 to investigate hard-to-reproduce issues, build across platforms, and focus on product outcomes.
  • What Codex unlocks for Notion

    • 发布时间:2026-06-09 18:00 北京时间
    • 摘要:- 在 Notion,Codex 正在改变工程师的构建方式。
      • 该公司正在重新考虑其构建的软件原语和抽象,以便代理可以使用它们。
      • 当将新工程师引入团队时,他们是出于好奇心和开放思想而招聘的,因为该领域通常需要的多年经验尚不存在。
      • 多年没有编写生产代码的经理又回到了代码库,与他们的团队一起发布。
      • Ryan Nystrom 在 Notion 负责人工智能产品工程。
    • EN 要点:
      • How Notion uses Codex to one-shot specs, build AI Voice Input for the web, and multiply engineering power across small teams.

Google DeepMind Blog (A_full) 链接到标题

  • Fluid, natural voice translation with Gemini 3.5 Live Translate

    • 发布时间:2026-06-09 23:16 北京时间
    • 摘要:- Gemini 3.5 Live Translate 为 Google AI Studio、Google Translate 和 Google Meet 带来近乎实时、自然的语音翻译。
      • 这篇来自 Google DeepMind 博客的文章解释了 Gemini 3.5 Live Translate 的流畅、自然的语音翻译如何塑造更广泛的人工智能和基础设施格局。
      • 使用 Gemini 3.5 Live Translate 进行流畅、自然的语音翻译后,它还为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.
  • Introducing Gemma 4 12B: a unified, encoder-free multimodal model

    • 发布时间:2026-06-09 22:10 北京时间
    • 摘要:- 隆重推出 Gemma 4 12B:统一、无编码器的多模态模型。
      • Google DeepMind 博客中的这篇文章解释了 Gemma 4 12B 简介:一个统一的、无编码器的多模态模型如何塑造更广泛的人工智能和基础设施景观。
      • 在推出 Gemma 4 12B:一个统一的、无编码器的多模态模型之后,它还为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • Introducing Gemma 4 12B: a unified, encoder-free multimodal model
  • Powering the future of robotics in Europe

    • 发布时间:2026-06-09 22:02 北京时间
    • 摘要:- 为欧洲机器人技术的未来提供动力。
      • 这篇来自 Google DeepMind 博客的文章解释了为欧洲机器人技术的未来提供动力如何塑造更广泛的人工智能和基础设施格局。
      • 在“为欧洲机器人技术的未来提供动力”之后,它还为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • Powering the future of robotics in Europe

ArXiv cs.AI (B_intro+search) 链接到标题

  • PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07549v1 公告类型:新。 -摘要:多模态大语言模型(MLLM)和代理工作流程的最新进展为计算病理学展现了巨大的前景,但可靠的补丁级推理仍然具有挑战性。
      • 端到端病理学 MLLM 通常会产生形态特征的幻觉,而最近的代理系统通常将工具输出和检索到的知识合并到共享上下文中,从而使决策容易受到相互矛盾的证据和上下文污染的影响。
      • 我们提出 PathoSage,一个三阶段框架,明确区分补丁级病理学多模态推理的知识检索、证据收集和证据判定。
    • EN 要点:
      • arXiv:2606.07549v1 Announce Type: new
      • Abstract: Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise for computational pathology, yet reliable pa…
      • End-to-end pathology MLLMs often hallucinate morphological features, while recent agentic systems usually merge tool outputs and retrieved knowledge into a shar…
      • We propose PathoSage, a three-stage framework that explicitly separates knowledge retrieval, evidence collection, and evidence adjudication for patch-level path…
  • OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07577v1 公告类型:新。 -摘要:视听大语言模型(LLM)对长视频理解有着巨大的希望,但它们的长视频推理从根本上受到视频令牌和键值(KV)缓存的线性增长的限制。
      • 我们推出 OmniMem,这是一种专为视听法学硕士设计的内存高效流框架。
      • 与统一处理所有标记的现有压缩方法不同,OmniMem 引入了一种模态感知内存分配策略,该策略单独管理视觉和音频上下文,解决两种模态之间严重的标记不平衡问题。
    • EN 要点:
      • arXiv:2606.07577v1 Announce Type: new
      • Abstract: Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limit…
      • We present OmniMem, a memory-efficient streaming framework designed specifically for audio-visual LLMs
      • Unlike existing compression methods that treat all tokens uniformly, OmniMem introduces a modality-aware memory allocation strategy that separately manages visu…
  • Syll: Open-Source Personal Automation with Cross-Surface Execution

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07594v1 公告类型:新。
      • 摘要:个人人工智能代理必须越来越多地跨 API、shell、Web 界面和桌面 GUI 进行操作,但许多系统仍然调整为单一界面,并对用户教学和可审核性提供有限的支持。
      • 我们推出了 Syll,一种开源、自托管的多模式代理工具,它将 MCP/API 工具、CLI 执行和可视化 GUI 控制统一在模块化运行时中,使代理能够跨异构接口协调计算机使用,同时简化用户和代理交换信息的方式。
      • Syll的核心是双向用户代理交互层:用户通过直接演示来教授程序,Syll将其编译成可重用的技能;代理执行被转换回多模式证据(日志、关键帧和批准检查点)以进行检查和控制。
    • EN 要点:
      • arXiv:2606.07594v1 Announce Type: new
      • Abstract: Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface…
      • We present Syll, an open-source, self-hosted multimodal agent harness that unifies MCP/API tools, CLI execution, and visual GUI control in a modular runtime, en…
      • At the core of Syll is a bidirectional user-agent interaction layer: users teach procedures through direct demonstration, which Syll compiles into reusable skil…
  • A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07718v1 公告类型:新。 -摘要:代理人工智能工具为自动化科学研究管道中的软件开发瓶颈提供了一条有希望的途径,特别是对于领域专家需要数天到数月才能构建的阶段,在这些阶段,科学家关心的是正确性和稳健性,而不是实现细节。
      • 我们提出了对飞行光遗传学数据到发现管道上的通用编码剂的实证研究。
      • 我们评估代理的任务远大于现有基准,数据集大几个数量级,评估标准基于领域专家标准。
    • EN 要点:
      • arXiv:2606.07718v1 Announce Type: new
      • Abstract: Agentic AI tools offer a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages that…
      • We present an empirical study of general-purpose coding agents on a fly optogenetics data-to-discovery pipeline
      • We assess agents on tasks substantially larger than existing benchmarks, datasets orders of magnitude bigger, and evaluation criteria grounded in domain expert…
  • Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07720v1 公告类型:新。
      • 摘要:大型语言模型(LLM)在数学和多跳规划任务上表现出了卓越的推理能力。
      • CoCoNuT(连续思维链)范式~\cite{hao2024coconut}通过使模型能够在潜在空间中进行推理来扩展这一点,同时探索多个推理路径,而不是早期致力于单个链。
      • 然而,我们发现了一个限制,我们称之为\textbf{概念瓶颈}。
    • EN 要点:
      • arXiv:2606.07720v1 Announce Type: new
      • Abstract: Large language models (LLMs) have demonstrated remarkable reasoning abilities on mathematical and multi-hop planning tasks
      • The CoCoNuT (Chain of Continuous Thought) paradigm~\cite{hao2024coconut} extends this by enabling models to reason in latent space, exploring multiple reasoning…
      • However, we identify a limitation we term the \textbf{concept bottleneck}
  • Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07721v1 公告类型:新。
      • 摘要:目标:从自由文本放射学报告中自动提取数据可以实现大规模研究,但很少有研究评估大型语言模型 (LLM) 在荷兰神经放射学报告中的性能。
      • 方法:我们分析了来自三级记忆诊所(2016-2021 年)的 947 份脑部 MRI 报告,这些报告由神经放射学家顾问撰写。
      • 受过训练的医学生注释了三十个变量; 100 份报告经过双重注释,以评估评估者间的可靠性。
    • EN 要点:
      • arXiv:2606.07721v1 Announce Type: new
      • Abstract: Objectives: Automatic data extraction from free-text radiology reports enables large-scale research, but few studies assessed the performance of large…
      • Methods: We analyzed 947 brain MRI reports from a tertiary memory clinic (2016-2021), authored by consultant neuroradiologists
      • Trained medical students annotated thirty variables; 100 reports were double-annotated to assess inter-rater reliability
  • Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07722v1 公告类型:新。
      • 摘要:本文提供了在讨论与其解决方案相关的问题时,聊天机器人作为真正的对话伙伴的本质的观点。
      • 聊天机器人能做什么,不能做什么,如何解释?
      • 我们的论点借鉴了聚合动力学、认知语言学、神经心理学和心理学。
    • EN 要点:
      • arXiv:2606.07722v1 Announce Type: new
      • Abstract: This article offers a perspective on the nature of chatbots as genuine conversation partners when discussing problems in relation to their solutions
      • What can chatbots do and what can’t they do, and how can this be explained
      • Our argument draws on Aggregation Dynamics, Cognitive Linguistics, Neuropsychology and Psychology
  • Land cover and flood type govern the detection limits of satellite-based flood mapping across diverse global flood events

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07780v1 公告类型:新。
      • 摘要:洪水是最具破坏性的自然灾害之一,气候变化导致洪水发生频率不断增加,使得基于卫星的洪水测绘对于灾害应对至关重要。
      • 在卫星档案上预先训练的地理空间基础模型提供了地理可转移性,但它们在各种未曾见过的事件中的运行可靠性仍然没有得到表征。
      • 在这里,我们在跨越六大洲、八个气候区和六种洪水机制的 19 个分布外洪水事件(2017-2025 年)中部署了 Prithvi-EO-2.0,并针对两个独立参考产品进行了验证。
    • EN 要点:
      • arXiv:2606.07780v1 Announce Type: new
      • Abstract: Floods are among the most destructive natural hazards, and their increasing frequency under climate change makes satellite-based inundation mapping es…
      • Geospatial foundation models pretrained on satellite archives offer geographic transferability, but their operational reliability across diverse, unseen events…
      • Here we deploy Prithvi-EO-2.0 across 19 out-of-distribution flood events (2017-2025) spanning six continents, eight climate zones, and six flood mechanisms, val…
  • Reconstructing and forecasting disease trajectories of patients with Alzheimer’s disease using routine data in resource-constrained settings

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07798v1 公告类型:新。
      • 摘要:阿尔茨海默病是一种进行性神经退行性疾病,其进展情况因患者而异。
      • 现有工作旨在预测患者未来的认知状态,而很少关注从过去的就诊中重建状态。
      • 此外,在当前的研究中,量化预测不确定性仍然没有得到充分探索,并且依赖于昂贵的模式,例如 MRI、PET 和 CSF,限制了它们在资源有限的环境中的部署。
    • EN 要点:
      • arXiv:2606.07798v1 Announce Type: new
      • Abstract: Alzheimer’s disease is a progressive neurodegenerative disorder, and its progression varies substantially across patients
      • Existing work aims to forecast patients’ future cognitive state, with minimal focus on reconstructing the state from past visits
      • Furthermore, in current research, quantifying predictive uncertainty remains underexplored and relies on costly modalities such as MRI, PET, and CSF, limiting t…
  • Improving Multimodal Reasoning via Worst Dimension Optimization

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07801v1 公告类型:新。
      • 摘要:多模态推理需要一条在从视觉基础到逻辑一致性的各种约束下保持完整性的路径。
      • 然而,当前的过程奖励模型侧重于同等权衡这些因素的启发式定义的奖励,这可能导致主导因素掩盖个体维度的失败,而不保证推理过程总体的有效性。
      • arXiv:2606.07801v1 公告类型:新 摘要:多模态推理需要一条在从视觉基础到逻辑一致性的各种约束条件下保持完整性的路径。然而,当前的过程奖励模型侧重于同等权衡这些因素的启发式定义的奖励,这可能会导致个体的隐藏……。
    • EN 要点:
      • arXiv:2606.07801v1 Announce Type: new
      • Abstract: Multimodal reasoning requires a path that retains integrity over a wide range of constraints, from visual grounding to logic consistency
      • However, the current Process Reward Models focus on heuristically defined rewards that equally weigh these factors, which may lead to the concealment of individ…

ArXiv cs.CL (B_intro+search) 链接到标题

  • Bidirectional Small-Granularity Search between Code and Text

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07519v1 公告类型:新。 -摘要:我们介绍了代码和文本之间双向小粒度搜索的新颖任务,其中查询是文本或代码的小片段,结果也是相反模态(即代码或文本)的小片段。
      • 该任务在科学出版物中的文本和相应的代码段之间建立直接链接,以支持更好更快地理解科学方法。
      • 我们为所提出的任务引入了一个大型数据集,其中包括一个训练分区,其中包含使用 GPT-4 自动生成的代码的文本描述,以及三个测试分区,一个域内和两个域外 (OOD),其中包含手动注释的数据以及来自其他域的材料。
    • EN 要点:
      • arXiv:2606.07519v1 Announce Type: new
      • Abstract: We introduce the novel task of bidirectional small-granularity search between code and text, where the queries are small snippets of text or code and…
      • This task establishes direct links between text in scientific publications and corresponding code segments, in support of better and faster understanding of sci…
      • We introduce a large dataset for the proposed task that includes a training partition with textual descriptions of code generated automatically using GPT-4, and…
  • TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist Ensembles

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07520v1 公告类型:新。
      • 摘要:指令跟随(IF)是法学硕士的核心能力,需要严格遵守各种约束,从可验证的(例如输出长度)到不可验证的(例如语气)。
      • 具有可验证奖励的强化学习已成为 IF 任务的范例,利用法学硕士作为法官来评估不可验证的约束。
      • 然而,我们凭经验发现这种方法仍然是一个重大瓶颈,遭受严重的奖励黑客攻击和更高的计算开销。
    • EN 要点:
      • arXiv:2606.07520v1 Announce Type: new
      • Abstract: Instruction Following (IF) is a core capability of LLMs, requiring strict adherence to diverse constraints, ranging from verifiable ones (e.g., output…
      • Reinforcement learning with verifiable rewards has emerged as a paradigm for IF tasks, leveraging LLM-as-a-judge to assess unverifiable constraints
      • However, we empirically find that this approach remains a significant bottleneck, suffering from severe reward hacking and higher computational overhead
  • Evaluating Hallucinations in Domain-Adapted Large Language Models

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07521v1 公告类型:新。
      • 摘要:本研究调查了领域适应大型语言模型 (LLM) 中的幻觉现象,重点是使用 Lamini 数据集对 Llama-2 模型进行微调。
      • 幻觉,或者法学硕士生成无意义或不忠实的内容,构成了重大挑战,特别是当这些模型使用特定领域的数据进行微调时。
      • 我们的方法涉及一系列实验,测试经过微调的法学硕士的记忆、回忆和推理能力,比较其在新颖的问答对和特定领域信息上的表现。
    • EN 要点:
      • arXiv:2606.07521v1 Announce Type: new
      • Abstract: This study investigates the phenomenon of hallucinations in domain-adapted Large Language Models (LLMs), focusing on the fine-tuning of the Llama-2 mo…
      • Hallucinations, or the generation of nonsensical or unfaithful content by LLMs, pose a significant challenge, especially when these models are fine-tuned with d…
      • Our methodology involves a series of experiments testing memorization, recall, and reasoning capabilities of the fine-tuned LLM, comparing its performance on no…
  • Community-Specific Slang and Entity Detection via Semantic Shift in Fine-Tuned Language Models

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07522v1 公告类型:新。
      • 摘要:我们提出了一种无监督方法,通过隔离词典中语义转移程度最高的单词来解析在线社区中的俚语、独特实体和民间传说。
      • 语义转变被定义为单词编码表示的演变,这是在社区特定文本语料库上微调预训练的大型语言模型 (LLM) 的结果。
      • 该值与单词的基本模型的编码表示和微调模型的编码表示之间的余弦相似度成反比。
    • EN 要点:
      • arXiv:2606.07522v1 Announce Type: new
      • Abstract: We propose an unsupervised method of resolving slang, unique entities, and folklore from online communities by isolating words in the lexicon that hav…
      • Semantic shift is defined as the evolution of a word’s encoded representation as a result of fine-tuning a pretrained Large Language Model (LLM) on a community-…
      • This value is inversely proportional to the cosine similarity between the base model’s encoded representation of a word, and a fine-tuned model’s encoded repres…
  • Retrieval Augmented Generation Framework for the Nepali Legal Domain Question Answering

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07523v1 公告类型:新。
      • 摘要:英语等高资源语言的法律领域已广泛采用人工智能进行法律问答。
      • 然而,尼泊尔语等资源匮乏语言的数据稀缺限制了尼泊尔法律文本大型语言模型的训练。
      • 这项研究首次应用基于检索增强生成的模型来回答尼泊尔法律问题,该模型使用从尼泊尔 Kanun Patrika 数字档案中提取的判例法。
    • EN 要点:
      • arXiv:2606.07523v1 Announce Type: new
      • Abstract: Legal domains in high-resource languages like English have widely adopted artificial intelligence for legal question answering
      • However, data scarcity in low resource languages such as Nepali has limited the training of large language models on Nepali legal texts
      • This study presents the first application of a Retrieval Augmented Generation based model for Nepali legal question answering using case laws extracted from the…
  • ABLE: Representing and Mapping LLMs via Attribution-Based Large-model Embedding

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07524v1 公告类型:新。 -摘要:大型语言模型(LLM)的爆炸性增长创造了一个异构且文档匮乏的生态系统,使得系统模型比较对于出处审计、安全分析和模型选择变得越来越重要。
      • 现有的表示方法很难有效地解决这个问题。
      • 当架构兼容时,分析内部参数的方法非常强大,但在结构异构性下面临可扩展性障碍,而依赖外部输出的方法可能会将具有相似行为的模型混为一谈,并且难以在不同分词器之间的更丰富的输出空间中对齐。
    • EN 要点:
      • arXiv:2606.07524v1 Announce Type: new
      • Abstract: The explosive growth of large language models (LLMs) has created a heterogeneous and poorly documented ecosystem, making systematic model comparison i…
      • Existing representation methods struggle to address this setting efficiently
      • Approaches analyzing internal parameters are powerful when architectures are compatible, but face scalability barriers under structural heterogeneity, while met…
  • Implicit Causal Graph Construction in Text via Chain Discovery

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07525v1 公告类型:新。
      • 摘要:文本中的因果图通常由可观察的预定义事件填充。
      • 相比之下,我们通过将每个描述的因果对视为潜在潜在因果图的开始和终点,并使用大型语言模型(LLM)来推断中间因果事件,来研究从文本构建隐式因果图。
      • 我们将端到端图构建与将任务框架为因果链发现的方法进行比较。
    • EN 要点:
      • arXiv:2606.07525v1 Announce Type: new
      • Abstract: Causal graphs in text are typically populated by observable, predefined events
      • In contrast, we study implicit causal graph construction from text by treating each described cause-effect pair as the begin- and endpoint of an underlying late…
      • We compare end-to-end graph construction with methods that frame the task as causal chain discovery
  • GraphLoRA: Structure-Aware Low-Rank Adaptation for Large Language Model Recommendation

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07526v1 公告类型:新。 -摘要:大型语言模型(LLM)由于其强大的推理和泛化能力而显示出强大的推荐潜力(LLMRec)。
      • 然而,有效地将法学硕士建模的文本语义与协作信号保持一致仍然是一个关键挑战。
      • 现有方法要么将协作信息转换为文本提示,要么将预先训练的嵌入注入到 LLM 中,这两种方法都将结构信息视为静态输入,并且无法捕获高阶关系依赖性。
    • EN 要点:
      • arXiv:2606.07526v1 Announce Type: new
      • Abstract: Large Language Models (LLMs) have shown strong potential for recommendation (LLMRec) due to their powerful reasoning and generalization abilities
      • However, effectively aligning the textual semantics modeled by LLMs with the collaborative signals remains a key challenge
      • Existing methods either translate collaborative information into textual prompts or inject pre-trained embeddings into the LLM, both of which treat structural i…
  • Post-training is (Massive) Supervised Learning

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07527v1 公告类型:新。
      • 摘要:LLM 培训的主流范式已经发展为依赖于由 SFT 和 RL 组成的大规模培训后阶段。
      • 在这篇立场文件中,我们认为这种方法有效地标志着 BERT 时代“预训练然后微调”方法的回归,明确地根据所需的行为和评估模型的具体基准来定制模型。
      • 我们首先回顾法学硕士的历史,描述法学硕士发展的不同阶段。
    • EN 要点:
      • arXiv:2606.07527v1 Announce Type: new
      • Abstract: The prevailing paradigm for training LLMs has evolved to rely on a massive post-training phase consisting of SFT and RL
      • In this position paper, we argue that this methodology effectively marks a reversion to the ``pre-train then fine-tune’’ approach of the BERT era, explicitly ta…
      • We begin with a historical overview of LLMs, describing the different phases of the LLM evolution
  • BEACON: Behavioral Entropy Aggregation for Cross-Model Hallucination Detection in Large Language Models

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07528v1 公告类型:新。
      • 摘要:大语言模型 (LLM) 中的幻觉(定义为生成事实上不正确或不受支持的内容)仍然是可靠部署的关键障碍。
      • 我们提出了 BEACON(跨模型幻觉检测的行为熵聚合),这是一种黑盒幻觉检测框架,纯粹基于模型输出运行,无需访问内部表示或外部知识库。
      • BEACON 从结构化多通道生成中提取 31 维特征向量,集成基于 NLI 的语义熵、嵌入几何、思想链一致性和释义稳定性信号。
    • EN 要点:
      • arXiv:2606.07528v1 Announce Type: new
      • Abstract: Hallucination in large language models (LLMs), defined as the generation of factually incorrect or unsupported content, remains a critical barrier to…
      • We present BEACON (Behavioral Entropy Aggregation for Cross-model hallucination detectiON), a black-box hallucination detection framework that operates purely o…
      • BEACON extracts a 31-dimensional feature vector from structured multi-pass generation, integrating NLI-based semantic entropy, embedding geometry, chain-of-thou…

ArXiv cs.LG (B_intro+search) 链接到标题

  • Offline Reinforcement Learning for Plasma Control in Nuclear Fusion: Codebase and Benchmark

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07550v1 公告类型:新。 -摘要:离线强化学习(RL)为根据历史托卡马克数据开发等离子体控制器提供了一条有前途的途径,因为在真实设备上进行在线试错成本高昂且风险很大。
      • 然而,由于缺乏针对核聚变中实际多执行器、长视界等离子体控制问题的标准化离线强化学习基准,这一方向的进展仍然难以衡量。
      • 我们推出了 RL4F,一种核聚变等离子体控制的离线强化学习基准,提供闭环评估环境和四个全轮廓跟踪任务的基线比较:旋转、密度、温度和压力。
    • EN 要点:
      • arXiv:2606.07550v1 Announce Type: new
      • Abstract: Offline reinforcement learning (RL) offers a promising route for developing plasma controllers from historical tokamak data, since online trial-and-er…
      • However, progress in this direction remains difficult to measure due to the lack of a standardized offline RL benchmark for realistic multi-actuator, long-horiz…
      • We introduce RL4F, an Offline Reinforcement Learning Benchmark for Plasma Control in Nuclear Fusion, providing closed-loop evaluation environments and baseline…
  • MedicalRec: Medical recommender system for image classification without retraining

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07553v1 公告类型:新。
      • 摘要:机器学习和深度学习的出现彻底改变了医疗保健中诊断、治疗和管理系统的效率。
      • 然而,这种快速采用的代价是需要大量的计算能力和能源消耗,以及电子废物处理和碳排放。
      • 这些模型的挑战之一是为分类任务选择正确的模型。
    • EN 要点:
      • arXiv:2606.07553v1 Announce Type: new
      • Abstract: The emergence of machine learning and deep learning has revolutionized the efficiency of diagnostic, therapeutic, and administrative systems in health…
      • However, this rapid adoption has come at the cost of requiring significant computing power and energy consumption, as well as e-waste disposal and carbon emissi…
      • One of the challenges of these models is choosing the right model for classification tasks
  • SPIN: Decentralized Swarm Control via Tensorized Policy Coordination

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07557v1 公告类型:新。
      • 摘要:资源受限的边缘平台上的分散式多智能体群体协调仍然受到联合行动空间的指数级扩展和高延迟通信开销的根本瓶颈。
      • 本文介绍了群体策略干扰网络(SPIN)框架,这是一种通过将群体拓扑建模为压缩张量网络来绕过这些限制的架构范例。
      • 我们将本地多智能体派系的联合策略张量分解为矩阵乘积状态(MPS)链,将评估的计算复杂度从指数 $O(n^m)$ 墙降低到严格线性 $O(m \cdot n \cdot \chi^2)$ 约束。
    • EN 要点:
      • arXiv:2606.07557v1 Announce Type: new
      • Abstract: Decentralized multi-agent swarm coordination on resource-constrained edge platforms remains fundamentally bottlenecked by the exponential scaling of j…
      • This paper introduces the Swarm Policy Interference Network (SPIN) framework, an architectural paradigm that bypasses these limitations by modeling swarm topolo…
      • We factorize the joint policy tensors of local multi-agent cliques into Matrix Product State (MPS) chains, reducing the computational complexity of evaluation f…
  • Boundary Variance Inflation Causes Acquisition Bias in Gaussian Processes

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07561v1 公告类型:新。
      • 摘要:有界域上具有固定核的高斯过程在边界附近表现出膨胀的后验方差。
      • 尽管边界引起的采集偏差是地质统计学中长期公认的伪影,也是贝叶斯优化中过度探索的根源,但边界引起的采集偏差的原因和影响尚未得到充分研究。
      • 我们将根本原因追溯到一个简单的几何机制:域边界处核相关邻域的截断产生了与观察无关的失真,这种失真随着维度的增加而恶化。
    • EN 要点:
      • arXiv:2606.07561v1 Announce Type: new
      • Abstract: Gaussian processes with stationary kernels on bounded domains exhibit inflated posterior variance near the boundary
      • Despite being a long-recognized artifact in geostatistics and a source of over-exploration in Bayesian optimization, the causes and effects of boundary-induced…
      • We trace the root cause to a simple geometric mechanism: the truncation of the kernel correlation neighborhood at the domain boundary creates an observation-ind…
  • Emergence via Phase Transitions: Mechanism Landscapes and Universal Convergence Across Complex Systems

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07563v1 公告类型:新。
      • 摘要:在机器学习、生物学和物理学中,独立进化的系统通常会趋向于惊人相似的高级结构,尽管微观细节截然不同。
      • Grokking 回路在随机种子上收敛,进化谱系重新发现相似的代谢解决方案,并且重整化流接近共同的固定点。
      • 我们提出分层涌现框架(HEF)作为此类融合现象的候选普遍性框架。
    • EN 要点:
      • arXiv:2606.07563v1 Announce Type: new
      • Abstract: Across machine learning, biology, and physics, independently evolving systems often converge toward strikingly similar high-level structures despite r…
      • Grokking circuits converge across random seeds, evolutionary lineages rediscover similar metabolic solutions, and renormalization flows approach common fixed po…
      • We propose the Hierarchical Emergence Framework (HEF) as a candidate universality framework for such convergence phenomena
  • STARIXNet: Multivariate and Multi-attribute Deep Learning Approach to Real-Time Resource Allocation in Cloud Platforms

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07565v1 公告类型:新。
      • 摘要:云平台中微服务的智能扩展对于缓解不断上升的计算成本同时避免服务中断至关重要。
      • 当前的解决方案仅限于单变量空间,通常仅关注 CPU 使用情况来驱动扩展决策。
      • 此外,他们将问题作为纯粹的预测任务来解决,重点关注预测精度,而忽略了低估和系统响应延迟的更大风险。
    • EN 要点:
      • arXiv:2606.07565v1 Announce Type: new
      • Abstract: Intelligent scaling of microservices in cloud platforms is crucial for mitigating escalating compute costs while avoiding service disruptions
      • Current solutions are limited to the univariate space, typically focusing on CPU usage alone to drive scaling decisions
      • Moreover, they address the problem as a purely forecasting task, focusing on prediction precision while neglecting the greater risks of underestimation and dela…
  • TriHead-GAN: A Generative Adversarial Network with Triple-Head Discriminator for Carbon Emission Time Series Generation

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07569v1 公告类型:新。
      • 摘要:准确的碳排放监测对于气候政策和欧盟碳边界调整机制等新兴监管机制至关重要,但城市级高频监测数据仍然极其稀缺,严重限制了需要大量数据的深度学习模型。
      • 时间序列生成是一种自然的补救措施,但现有的 GAN 和基于扩散的生成器通常对碳排放数据的域结构提供有限的显式监督:它们可能匹配边际分布统计数据,同时不足以保留 CO$_2$ 与共同排放的污染物和气象因素之间的跨变量相关性,并且往往会破坏大气测量的一阶差分统计数据,产生平均平滑但缺乏基础信号的实际逐步变化的序列。
      • 我们提出 TriHead-GAN,一种基于 Transformer 的对抗框架,其三头判别器共同监督联合分布的三个互补方面:通过 Wasserstein 批评家的分布真实性、通过目标变量的无泄漏回归的跨变量依赖性以及通过相邻差异预测的逐步时间平滑性。
    • EN 要点:
      • arXiv:2606.07569v1 Announce Type: new
      • Abstract: Accurate carbon emission monitoring is critical for climate policy and emerging regulatory mechanisms such as the EU Carbon Border Adjustment Mechanis…
      • Time series generation is a natural remedy, but existing GAN and diffusion-based generators often provide limited explicit supervision for the domain structure…
      • We propose TriHead-GAN, a Transformer-based adversarial framework whose triple-head discriminator jointly supervises three complementary aspects of the joint di…
  • Enabling KV Caching of Shared Prefix for Diffusion Language Models

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07571v1 公告类型:新。
      • 摘要:共享前缀的键值 (KV) 缓存对于高吞吐量大型语言模型 (LLM) 服务至关重要,但它在新兴的扩散语言模型 (DLM) 中面临着严峻的挑战。
      • 在 DLM 中,双向注意力意味着更新任何 token 都会动态改变整个上下文及其相应的 KV。
      • 因此,为 LLM 开发的现有缓存技术假设 KV 一旦计算后保持不变,会破坏共享前缀 KV。
    • EN 要点:
      • arXiv:2606.07571v1 Announce Type: new
      • Abstract: Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical challenges in em…
      • In DLMs, bidirectional attention means that updating any token dynamically alters the entire context and its corresponding KVs
      • Thus, existing caching techniques developed for LLMs, which assume that KVs remain invariant once computed, corrupt the shared prefix KVs
  • When Should an AI Scientist Stop? Verifiable Experiment Steering and Refusal for Autonomous Discovery

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07576v1 公告类型:新。
      • 摘要:我们提出了 CARTOGRAPH,这是人工智能科学家的一个验证层,它将未解决的子空间实验引导(选择)、显式歧义闭合(解决)和基于残差的库不足检测(拒绝)结合起来。
      • 在局部线性高斯桥下,原始未解析投影是各向同性未解析 Fisher 信息迹,而 CARTOGRAPH-A 是精确的未解析 A 最优规则;封闭式 EIG 和 Box-Hill 是作为局部比较器而不是全局比较器出现的。
      • 在五个测试平台上,CARTOGRAPH-A 在复制的结构化级联中在 d = 8 (p < 10^-21) 时击败原始投影 129W/0T/15L。
    • EN 要点:
      • arXiv:2606.07576v1 Announce Type: new
      • Abstract: We present CARTOGRAPH, a verification layer for AI scientists that couples unresolved-subspace experiment steering (select), explicit ambiguity closur…
      • Under a local linear-Gaussian bridge, raw unresolved projection is the isotropic unresolved Fisher-information trace, while CARTOGRAPH-A is the exact unresolved…
      • Across five testbeds, CARTOGRAPH-A beats raw projection 129W/0T/15L at d = 8 (p < 10^-21) in a replicated structured cascade
  • MST-Direct at Scale: Multivariate and Conditional Geostatistical Simulation via Sinkhorn Optimal Transport

    • 发布时间:2026-06-09 12:00 北京时间
    • 摘要:- arXiv:2606.07578v1 公告类型:新。
      • 摘要:本文扩展了 MST-Direct(一种用于多元地质统计模拟的通过 Sinkhorn 传输匹配的方法),从原始的双变量、无条件、小网格公式扩展到多元、条件和大网格设置。
      • 我们解决了原始工作中确定的三个主要限制:(i)通过具有 O(nC) 内存复杂度的稀疏、候选限制的 Sinkhorn 匹配器,可扩展性超过数千个节点; (ii) 通过将目标值元组匹配到独立的 FFT-MA 高斯主干上来扩展至多个变量,该主干可再现指定的变差函数; (iii) 硬数据调节,通过将观察到的数据元组固定在其空间位置,同时通过克里金法调节主干网。
      • 因为传输计划仍然是目标元组的排列,所以多元联合分布被准确地保留。
    • EN 要点:
      • arXiv:2606.07578v1 Announce Type: new
      • Abstract: This paper extends MST-Direct, a Matching-via-Sinkhorn-Transport approach for multivariate geostatistical simulation, from the original bivariate, unc…
      • We address the three main limitations identified in the original work: (i) scalability beyond a few thousand nodes through a sparse, candidate-restricted Sinkho…
      • Because the transport plan remains a permutation of the target tuples, the multivariate joint distribution is preserved exactly