🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-30
- 类型
- ai-daily
- 字数
- 3453
- 阅读时长
- 17 min
2026-06-30 AI日更 | Agent 走向长期规划,AI 竞争进入可验证系统阶段 链接到标题
今天的主线不再是单纯追逐更大模型,而是 Agent 的长期规划、多智能体协作与可信部署。前沿模型仍在加速迭代,但受限访问、安全解禁、开放权重争议和真实场景反馈,正在共同重塑 AI 的竞争边界。
📖 本期 Watch List 深度导读 链接到标题
今天的 Watch List 建议重点看三条线:其一是智能体“长期规划”的底层能力,多篇 arXiv 都在讨论世界模型、符号反馈、自我迭代与幻觉传播,说明 Agent 研究正从提示工程转向可验证、可度量的规划机制。其二是多智能体与模型网络,从人格组合到 AI-Model Network,值得关注大模型协作形态如何影响任务绩效与系统架构。其三是垂直场景里的可信 AI:事实核查、法律适配、临床训练、阅读障碍学习者支持等工作,都在把模型能力落到高风险、低资源或强解释需求的场景中。总体看,今天的深读关键词不是“更大模型”,而是“可控、可验证、可部署”。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Zhipu AI Seeks Community Input on GLM-5.3 Vision Features 链接到标题
- 分类:AI · News
- 概况:热度时间:19 hours ago,相关帖子数:1300
- 是什么事:智谱 AI 正在向社区征集 GLM-5.3 视觉能力的功能反馈与改进建议。
- 为什么重要:这表明多模态大模型竞争正从单纯参数和榜单表现,转向围绕真实使用场景、产品体验和开发者需求进行迭代。
- 讨论概况:X 上讨论主要集中在 GLM-5.3 是否应强化图像理解、视觉推理、OCR、视频能力和智能体工作流集成;也有人关注社区反馈能否真正影响产品路线,以及国产模型在多模态领域与 Claude、GPT、Gemini 等模型的差距。
话题 2:DeepSeek V4 Models Launch Officially in Mid-July with Peak Pricing 链接到标题
- 分类:AI · News
- 概况:热度时间:23 hours ago,相关帖子数:2800
- 是什么事:DeepSeek 据称将在7月中旬正式推出 V4-Pro 和 V4-Flash 模型,主打 1M token 上下文、开源权重、更快推理和较低输出价格。
- 为什么重要:如果相关性能、开源许可和价格信息属实,DeepSeek V4 将进一步推动大模型在长上下文、代码能力、推理效率和低成本部署上的竞争,并对 OpenAI、Google、Anthropic、xAI 等闭源模型厂商形成价格与开放生态压力。
- 讨论概况:X 上讨论集中在 DeepSeek 是否会以 MIT 许可真正开放高性能权重、1.6 万亿参数和 SWE-bench 成绩的可信度、低价是否可持续,以及其对全球 AI 模型定价和中美 AI 竞争格局的影响。
话题 3:Flexion Robotics Unveils Reflect v1.0 for Long-Horizon Robot Autonomy 链接到标题
- 分类:AI · News
- 概况:热度时间:14 hours ago,相关帖子数:249
- 是什么事:Flexion Robotics 发布了 Reflect v1.0,主打面向长时程任务的机器人自主能力。
- 为什么重要:这表明机器人领域正从短指令执行迈向更复杂的持续规划、环境适应与错误恢复能力,是通用具身智能落地的重要方向。
- 讨论概况:X 上的讨论主要集中在 Reflect v1.0 是否能显著提升真实场景中的稳定性与泛化能力,也有人关注其与现有机器人基础模型、仿真训练和端到端控制方案的差异及商业化前景。
话题 4:Dario Amodei’s 2023 Open-Source AI Warning Resurfaces Amid Model Surge 链接到标题
- 分类:AI · Other
- 概况:热度时间:1 day ago,相关帖子数:21000
- 是什么事:Anthropic CEO Dario Amodei 在 2023 年关于开源 AI 可能带来安全风险的警告,因近期开源与开放权重模型快速增长而在 X 上重新引发关注。
- 为什么重要:这反映出 AI 领域在模型能力扩散、透明度、创新速度与安全治理之间的核心矛盾,尤其是在高性能模型越来越容易被获取和部署的背景下。
- 讨论概况:X 上的讨论主要围绕开源 AI 是推动竞争与民主化,还是会降低滥用门槛;支持者强调开放生态有助于审计和创新,批评者则担心强模型被用于网络攻击、欺诈或绕过监管。
话题 5:xAI Launches Grok 4.5 Private Beta at SpaceX and Tesla 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:43000
- 是什么事:马斯克称 xAI 的 Grok 4.5 已在 SpaceX 和 Tesla 内部进入私人测试,早期表现据称接近或超过 Claude Opus 的部分评测,公开发布仍待进一步验证。
- 为什么重要:这件事的重要性不只在于单个模型性能,而在于 xAI 可能正在形成从算力、基础模型、编码代理、企业内部部署到真实任务反馈和快速再训练的闭环,体现前沿 AI 正从单次模型发布转向高频迭代的“模型工厂”模式。
- 讨论概况:X 上讨论焦点集中在 Grok 4.5 是否真的具备 Opus 级能力、所谓 1.5T 参数和 Cursor 相关训练数据的可信度与数据权利问题,以及 Tesla、SpaceX 内部场景是否能成为高质量强化学习反馈来源;分歧在于支持者认为 xAI 的发布速度和垂直整合会挑战 OpenAI、Anthropic,质疑者则强调缺乏公开、可复现、同等工具环境下的第三方评测。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线是,AI 竞争正在从单点模型性能转向“能力、成本、开放生态与真实场景反馈”的综合较量:多模态、长上下文、机器人长时程任务和企业内部闭环迭代都成为焦点。共识在于,模型厂商必须更贴近实际使用场景,通过开发者反馈、低成本部署、持续训练和产品化能力来形成优势,而不是只依赖榜单宣传。分歧主要集中在开放权重与安全治理之间的取舍,以及 DeepSeek、xAI 等新模型的参数规模、评测成绩、许可承诺和内部测试结果是否可信。潜在风险则包括高性能开源模型降低滥用门槛、模型宣传缺乏可复现验证、训练数据权利争议,以及机器人和智能体系统在真实环境中稳定性不足带来的安全与商业化落差。
💡 大佬观点(Influencer Insights) 链接到标题
好的,基于对过去 24 小时内各位 AI 大佬推文的分析,以下是您的行业简报。
AI 行业大佬动态日报:前沿模型加速迭代,AI 融入工作流成主旋律 链接到标题
1. 核心技术与产品热点:前沿模型密集发布,端侧与 Agent 基建成为新战场 链接到标题
今日的讨论热点集中在基础模型的迭代、开发工具的快速进化,以及 AI 如何更深入地嵌入组织工作流。
OpenAI GPT-5.6 系列发布与受限访问引发热议 这是今天最热门的话题。OpenAI 发布了新一代模型 GPT-5.6,包含旗舰级的 Sol、均衡型的 Terra 和经济型的 Luna 三个版本。
- @dotey 详细解读了该模型信息:Sol 支持
max(深度推理)和ultra(多 Agent 并行)模式,在编程基准测试中表现优异并重点介绍了其安全机制。但最引人注目的是,该模型目前仅向约 20 家经美国政府审批的合作伙伴开放,普通用户和开发者暂时无法使用。@dotey 认为,这标志着美国政府对前沿 AI 模型的发布前审查正在“从个案变成惯例”。 - 与此同时,社区已在探索灰度测试的迹象。@dotey 和 @Pluvio9yte 等博主分享了通过“Juice 测试 Prompt”和 Codex Analytics 面板验证自己的账户是否被灰度到 GPT-5.6 Sol 的方法,引发了大量讨论。
- @dotey 详细解读了该模型信息:Sol 支持
Claude 生态的关键动态:Mythos 5 解禁与 Claude Tag 范式
- @dotey 报道了 Anhtropic 的 Mythos 5 在因安全原因被封禁两周后,获得美国政府的部分解禁令,将被重新部署给约 100 家美国机构用于网络安全防御。@dotey 点评道:“一个模型因为太危险而被下架,又因为太有用而被请回来”,深刻反映了前沿模型在安全与实用性之间的博弈。
- 与此同时,Anthropic 发布的 Claude Tag 功能引发了关于 AI 与团队协作模式的深度探讨。@dotey 转述了 @GergelyOrosz 的理解,强调其突破点不在于 Slack,而在于“一个云端 AI 被接入了公司内部系统后开箱即用”。@dotey 也引用了 @karpathy 的观点,将此称为继网页版和 App 版之后,LLM 交互方式的“第三次重大重新设计”。
端侧模型与应用的高速发展
- 模型新范式:@zhixianio 持续关注端侧模型,设想了“Model-Pak”概念,即模型像卡带一样即插即用,并认为这代表了未来方向。
- 性能实测:@zhixianio 实测了 MiniCPM-o 4.5 的音视频全双工能力,认为效果满意;同时,他对 Google Gemma 4 系列(12B Coder)进行了详尽的本地测试。测试结果显示,虽然 12B 模型在特定任务上表现出色,但面对“俄罗斯方块”这类需要长篇、有状态、一次成型的复杂程序,与 35B 的 Qwen 模型仍有显著差距,指出模型“天花板”依然存在。
非侵入式脑机接口取得突破 @dotey 分享了 Meta 发布的 Brain2Qwerty v2,它能够利用 MEG(脑磁图)设备,将大脑信号实时解码为句子,平均单词准确率达 61%,表现最好者可达 78%。@dotey 强调,这一非侵入式技术的突破,证明了“不开刀也有可能做到接近开刀的效果”,为失去沟通能力的人带来新希望。
2. 独特观点与行业前瞻 链接到标题
关于 AI 编程的真实成本与价值:
- @ruanyf 从 OpenClaw 创始人的天价 Token 消耗(一个月价值 130 万美元)案例出发,冷静指出:如果无限量使用顶级模型,AI 编程的 Token 成本可能比真人程序员更昂贵。
- @gefei55 则提出另一种视角:“Token 是无限的,时间精力是有限的”,警示大家不要迷失在“什么都能做的 Token 陷阱”里,要避免无意义的消耗生命。
- 福特公司的案例也被 @dotey 拿来做反例:该公司因 AI 质检未达预期,重新雇佣 350 名资深工程师来“调教” AI,并拿下新车质量榜第一。这表明,AI 的好坏取决于训练它的人和数据的质量,人机协作仍是现实。
AI 产品与商业模式的思考:
- @zhixianio 对 Fable 印象深刻,称其 40 分钟内完成了一个 Demo 70% 的工作,不仅实现功能,还指出了原设计的不足并采用更好的方案。这代表了 AI Agent 从被动执行到主动参与设计的范式转变。
- @Pluvio9yte 分享了火山引擎的 Coding Plan 以 9.9 元/月的超低价杀入市场,感叹市场竞争的“抽象”与激烈。
AI 时代的社会与人文洞察:
- @ruanyf 敏锐地提出了一个深刻问题:“如果未来的代码都是 AI 写的,那么我们怎么招聘程序员呢?”,指出考察候选人是否会使用 AI 的能力变得比考察代码能力本身更重要。
- @lijigang 则持续输出哲思,提出“品味是一个人的损失函数”、“数字世界,人以点赞为食”等金句,并从认知科学角度观察:重度使用某个模型会染上它的语言风格,这与大脑“吃”context 的特性有关。
对“开源”的重新定义: @ruanyf 引用了 Anthropic 创始人的观点,指出 AI 模型公开权重不等于传统开源,因为你无法看到内部运作或参与开发,这更像“开放权重”。
3. 推荐的工具与资源 链接到标题
AI 开发与生产力工具:
- RepoPrompt (社区版):由 @dotey 推荐。这款用于“上下文工程”的工具已开源,其新架构支持用推理模型做任务分解,然后分发给不同 Agent 并行执行。
- Codex 高效使用技巧:@Pluvio9yte 开源了一个 Skill 管理工具,可扫描并统计所有已安装 Skills 的使用频率,帮助“断舍离”。@dotey 则分享了
/btw和fork等会话管理技巧。 - EdgeOne Makers:由 @vista8 推荐。腾讯云推出的 Agent 开发平台,只需三行命令即可部署 AI Agent 开发框架,并解决了沙箱环境、持久化、可观测性等上线难题,适合希望快速落地的开发者。
内容创作与 SEO 工具:
- 视频制作 Skills 库:@Pluvio9yte 开源了一整套视频制作 Skills,可直接复刻特定风格的视频,降低了视频创作门槛。
- GEO (生成式引擎优化) 资料包:@vista8 在开展 GEO 公开课后,慷慨地分享了全套资料,包括操作手册、系统研究报告、提示词和 Skill 等,对于希望在 AI 搜索时代做好内容优化的人极具价值。
- 推特热点挖掘工具:@gefei55 开源了一款小工具,通过调用 X API 扫描高互动带链接推文,并反查域名流量,帮助用户在信息差中比 Google Trends 更快地发现新词和热点。
实用资源与学习资料:
- Codex 学习宝典:@AI_Jasonyu 和 @dotey 共同推荐。@bozhou_ai 和 @dotey 分别开源了系统学习 Codex 的“橙皮书”和“Claude Code From Scratch”电子书,后者用 4300 行复现了 Claude Code 核心架构,是深入理解 Coding Agent 的绝佳资源。
- ChatGPT 管理插件:@Pluvio9yte 自己开发并推荐了一款 Chrome 插件,可批量删除或归档 ChatGPT 网页版聊天记录,解决了用户的一大痛点。
总结:过去 24 小时,行业火力全开。一方面以 GPT-5.6 和 Claude 5 系列为代表的前沿模型在能力、定价和安全合规上展开激烈竞速;另一方面,AI 正在从独立的工具转变为嵌入组织、融入工作流(如 Claude Tag)、赋能个体(如端侧模型和自动化 Skills)的基础设施。大佬们的讨论已经从单纯的模型能力比拼,转向了成本、效率、合规以及如何真正将 AI 转化为生产力的深度思考。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 33 条更新
All-In Podcast (A_full) 链接到标题
- Nate Silver Predicts: Democrats Take the House, Newsom Is Fading & AOC Might Win It All in 2028
- 发布时间:2026-06-29 23:47 北京时间
- 摘要:- 西北注册代理——创业?
- 西北注册代理为您提供构建完整企业身份所需的一切,包括免费工具和内置隐私。
- 请访问 Northwestregisteredagent.com/ALLINFREE 了解更多信息。
- PLAUD - 如果您的工作依赖于对话 - 会议、交易流程、访谈、客户电话 - Plaud 可以帮助您通过高度准确的 AI 生成的笔记来捕获和组织所有内容,这些笔记不仅仅是简单的摘要,还可以突出显示痛点、关键决策、后续步骤和可自定义的摘要模板。
- 在 Plaud.ai/allin 上查看 Plaud,并使用代码 ALLIN 享受高达 20% 的折扣!
- EN 要点:
- (0:00) Nate Silver joins the pod
- (10:02) California’s ballot counting problem: Raman’s late-mail surge, ballot harvesting claims, and why the US counts slower than India
- (25:18) Democrats’ three-way civil war: The left, the abundance libs, and Newsom’s “resistance lib” base
- (34:48) The winning 2028 playbook: Anti-oligarch messaging, why young men want control, and immigrants fleeing the Dems
Stratechery by Ben Thompson (A_full) 链接到标题
- Summer Break: Week of June 29
- 发布时间:2026-06-29 18:00 北京时间
- 摘要:- Stratechery 于 6 月 29 日这一周放暑假。
- 不会有每周文章或更新。
- 下一次更新将于 7 月 6 日星期一进行。
- Dithering、夏普科技和夏普中国将……。
- EN 要点:
- Stratechery is on summer break the week of June 29
- There will be no Weekly Article or Updates
- The next Update will be on Monday, July 6
- Dithering , Sharp Tech , and Sharp China will also return the week of July 6
OpenAI Blog (A_full) 链接到标题
- Mapping Europe’s AI Workforce Opportunity
- 发布时间:2026-06-29 15:00 北京时间
- 摘要:- AI能力可以快速跨越国界。
- 工作的变化不会如此顺利。
- 工作是由许可制度、当地机构以及提供护理、教育、司法、公共服务和其他形式的人力支持的实际情况决定的。
- 这些系统很重要,因为它们有助于确定人工智能是否以及如何改变劳动力市场。
- 人工智能对劳动力市场有何影响?
- EN 要点:
- A new OpenAI report maps how AI could reshape jobs across the EU, highlighting which occupations may face automation, growth, or workflow changes.
ArXiv cs.AI (B_intro+search) 链接到标题
AI-Model Network: Concept, Current State and Future
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27382v1 公告类型:新。
- 摘要:计算机的主要功能在于计算和处理,而互联网的核心价值则植根于共享和协作。
- 计算机创造互联网,互联网赋予计算机价值。
- 互联网、云计算、大数据的快速发展正在推动人工智能进入大模型(LM)时代。
- EN 要点:
- arXiv:2606.27382v1 Announce Type: new
- Abstract: While the primary function of computers lies in computation and processing, the core value of the Internet is rooted in sharing and collaboration
- Computers create the Internet, and the Internet empowers the value of computers
- The rapid development of the Internet, cloud computing, and big data is pushing artificial intelligence into the era of large models (LMs)
When Does Personality Composition Matter for Multi-Agent LLM Teams?
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27443v1 公告类型:新。
- 摘要:个性提示决定了大型语言模型的沟通方式,但这些行为转变是否影响客观任务结果仍有待探索。
- 先前的研究表明,低宜人性提示的智能体会产生对抗性语言,而高宜人性提示的智能体会变得合作,但沟通方式与任务绩效之间的关系尚未在多个领域进行系统地检验。
- 在这项工作中,我们通过在结构化编码、开放式研究合作和竞争性谈判这三个任务领域操纵前沿法学硕士的人格特质,研究人格构成是否对多智能体团队绩效重要。
- EN 要点:
- arXiv:2606.27443v1 Announce Type: new
- Abstract: Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-e…
- Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but…
- In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating personality traits across frontier LLMs on…
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27483v1 公告类型:新。
-摘要:大语言模型(LLM)智能体在顺序决策方面表现出了强大的能力,但它们在长期任务中仍然基本上是反应性的。
- 与在承诺之前使用“假设”推理来评估潜在计划的人类不同,标准代理缺乏内部世界模型来模拟未来的结果。
- 因此,我们建议通过训练单个自回归模型来内部化未来感知规划,以表达预期状态的推出和计划条件的成功估计(Q 值的文本模拟)。
- EN 要点:
- arXiv:2606.27483v1 Announce Type: new
- Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-h…
- Unlike humans who employ “what-if” reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outc…
- Therefore, we propose to internalize future-aware planning by training a single autoregressive model to verbalize both a prospective state rollout and a plan-co…
Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27593v1 公告类型:新。
-摘要:我们引入了一个名为 ODYSSEY 的分类框架,用于构建可验证的、本地保真基础模型作为铸造厂的组合:指定本地上下文、本地表示族、限制图、粘合规则、阻塞策略、更新义务和面向人类的视图的构建块架构组件。
- 铸造厂是一个有组织的知识库,其中包含论证成分。
- 混凝土铸造厂是根据通用铸造厂建造的,例如证据/论证、运营决策、机构/财务、市场意义、科学挑战、研究计划、辅助建造和评估装备铸造厂。
- EN 要点:
- arXiv:2606.27593v1 Announce Type: new
- Abstract: We introduce a categorical framework called ODYSSEY for constructing verifiable, local truth-preserving foundation models as compositions of foundries…
- A foundry is an organized sheaf of knowledge that carries within it an argumentation component
- Concrete foundries are built from generic foundries such as evidence/argument, operational decision, institutional/financial, market meaning, scientific challen…
DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27619v1 公告类型:新。
- 摘要:阅读障碍学习者越来越多地使用人工智能 (AI) 工具来支持阅读、写作、组织和学习相关任务。
- 然而,他们使用这些工具的生活经验在很大程度上仍未得到充分检验。
- 本文提出了 DysLexLens,这是一种低资源 LLM 框架,旨在通过在线论坛讨论来分析阅读障碍学习者使用 AI 的体验。
- EN 要点:
- arXiv:2606.27619v1 Announce Type: new
- Abstract: Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks
- However, their lived experiences with these tools remain largely underexamined
- This paper proposes DysLexLens, a low-resource LLM framework, designed to analyse dyslexic learners experience with AI through online forum discussions
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27652v1 公告类型:新。
- 摘要:我们发现显式推理不一定会转化为更好的多模态情感识别(MER)准确性,尽管它使预测更容易解释。
- 具体来说,对于基于推理的 MLLM,通过触发直接答案的快速思考通常胜过深思熟虑推理后的缓慢思考。
- 我们的实证分析表明,快速思考可以通过更广泛、更自信的预测来提高回忆能力,而缓慢思考则可以通过保守地过滤不正确的类别来提高精确度。
- EN 要点:
- arXiv:2606.27652v1 Announce Type: new
- Abstract: We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predict…
- Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberative reasoning
- Our empirical analyses show that fast thinking improves recall with broader and more confident predictions, whereas slow thinking favors precision through conse…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27736v1 公告类型:新。
-摘要:假新闻的快速传播对信息生态系统构成了越来越大的威胁,特别是在生成引擎优化(GEO)中毒下人工智能生成的错误信息使得检索系统系统性地呈现出敌对制作的内容,从而污染了法学硕士的推理。
- 在本文中,我们提出了证据树(ToE),这是一种用于自动事实检查的分层证据推理框架,将每个主张建模为动态扩展的论证树。
- ToE集成了强化学习驱动的多源检索代理、证据评估代理和论据树聚合算法,通过可解释的证据链迭代分解、检索和验证主张。
- EN 要点:
- arXiv:2606.27736v1 Announce Type: new
- Abstract: The rapid spread of fake news poses increasing threats to information ecosystems, especially as AI-generated misinformation under Generative Engine Op…
- In this paper, we propose Tree of Evidence (ToE), a hierarchical evidence reasoning framework for automated fact-checking that models each claim as a dynamicall…
- ToE integrates a reinforcement learning-driven multi-source retrieval agent, an evidence evaluation agent, and an argument tree aggregation algorithm to iterati…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27757v1 公告类型:新。
-摘要:大型语言模型(LLM)已引起学术界和工业界的广泛关注,但其部署引发了有关鲁棒性和可靠性的关键安全问题。
- 规划是智能行为的核心组成部分,对于法学硕士来说仍然具有挑战性,由于固有的复杂性,他们经常在长期决策任务中产生不可行或不正确的解决方案。
- 在本文中,我们提出了一种符号反馈驱动的迭代自我完善框架,以增强法学硕士在长期规划中的稳健性和可靠性。
- EN 要点:
- arXiv:2606.27757v1 Announce Type: new
- Abstract: Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns re…
- Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision…
- In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-horizon plann…
Understanding Rollout Error in Graph World Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27780v1 公告类型:新。
- 摘要:世界模型通常用于通过向前滚动学习动态来进行规划。
- 然而,许多规划环境不是矢量或图像;它们是代理、工具、技能、路线和依赖关系的图表。
- 在这些设置中,局部预测误差可能保留在局部或在图形中传播,并且当边缘被预测而不是固定时,故障模式会再次发生变化。
- EN 要点:
- arXiv:2606.27780v1 Announce Type: new
- Abstract: World models are often used for planning by rolling learned dynamics forward
- Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and dependencies
- In these settings, a local prediction error may stay local or spread through the graph, and the failure mode changes again when edges are predicted rather than…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27806v1 公告类型:新。
- 摘要:语言代理的世界模型有两种有用的形式。
- 基于代理的世界模型调用LLM API并用语言灵活地进行推理,但其错误表现为幻觉的状态变化,很难用普通的回归损失来评分。
- 参数化世界模型是经过训练的转换预测器;它的错误更容易用 NodeMSE、增量准确度和有效性准确度等数量来衡量,但作为独立规划器,它通常较弱。
- EN 要点:
- arXiv:2606.27806v1 Announce Type: new
- Abstract: World models for language agents come in two useful forms
- An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with or…
- A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity ac…
ArXiv cs.CL (B_intro+search) 链接到标题
Generating in the Limit with Infinitely Many Hallucinations
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28354v1 公告类型:新。
- 摘要:极限语言识别的经典范式将学习建模为对手(揭示未知目标语言的字符串)和学习者(负责识别该语言)之间的游戏。
- 最近引入的极限语言生成框架改变了目标,以更好地反映现代语言建模,要求学习者从目标语言中生成有效的、看不见的字符串。
- 相关工作强调了一种根本性的紧张关系:目标的广泛覆盖往往是以有效性为代价的。
- EN 要点:
- arXiv:2606.28354v1 Announce Type: new
- Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown targe…
- The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to…
- Related work highlighted a fundamental tension: a broad coverage of the target often comes at the cost of validity
Extracting Knowledge from an Arabic-English Machine-Readable Dictionary Using Information Extraction
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28457v1 公告类型:新。
- 摘要:自然语言处理(NLP)应用需要大量且丰富的语言知识。
- 此外,词典、百科全书和语料库等电子语言资源也变得可用。
- 因此,出现了自动方法来从这些来源中提取词汇信息,以克服知识获取瓶颈。
- EN 要点:
- arXiv:2606.28457v1 Announce Type: new
- Abstract: Natural language processing (NLP) applications need large and rich amount of linguistic knowledge
- Furthermore, electronic language sources such as dictionaries, encyclopedia, and corpora became available
- So, automatic methods are emerged to extract lexical information from those sources to overcome the knowledge acquisition bottleneck
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28524v1 公告类型:新。
- 摘要:最近的工作表明,大型语言模型(LLM)对文本描述的主体的信念状态敏感,如通过错误信念任务(FBT)测量的那样,但对结构有效性的持续关注仍然存在。
- 我们采用发展的视角,在 Olmo2 和 Pythia 语言模型套件的多个训练阶段中追踪心理状态推理行为的模式,以及这种行为可能的先决条件。
- 我们发现,高于机会的 FBT 表现取决于模型大小和足够的训练量,在预训练中出现相对较晚,并且在最容易诊断心智化(错误信念、隐式)的情况下通过训练后干预(SFT、DPO)得到最大改善。
- EN 要点:
- arXiv:2606.28524v1 Announce Type: new
- Abstract: Recent work suggests that Large Language Models (LLMs) are sensitive to the belief states of agents described by text, as measured by the false belief…
- We adopt a developmental perspective, tracing the pattern of mental state reasoning behavior – and likely preconditions for this behavior – across mul…
- We find that above-chance FBT performance depends both on model size and sufficient training volume, emerges relatively late in pretraining, and is most improve…
A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28526v1 公告类型:新。
-摘要:医学生的临床和沟通技能通常通过客观结构化临床考试(OSCE)进行评估,其中包括对医患互动的简短场景驱动模拟。
- 然而,培训往往受到人类标准化患者的可用性较低的限制,这激励了现实虚拟患者(VP)的开发。
- 为了解决这一差距,我们引入了法国 OSCE 对话数据集,其中包含 240 个学生与患者的培训互动。
- EN 要点:
- arXiv:2606.28526v1 Announce Type: new
- Abstract: The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which cons…
- However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs)
- To address this gap, we introduce a French OSCE dialogue dataset comprising 240 student-patient training interactions
Legal Domain Adaptation of Modern BERT Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28538v1 公告类型:新。
- 摘要:我们研究现代 BERT 模型在法律领域的领域适应性。
- 我们使用屏蔽语言建模目标进一步针对所有美国法院意见对 ModernBERT 进行预训练。
- 尽管 ModernBERT 的训练数据比原始 BERT 多了大约 500 倍,但我们仍然发现该模型受益于法律领域的进一步预训练和领域适应:我们报告称,与普通的 ModernBERT 相比,在与美国法院意见相关的所有数据集上都有显着改进。
- EN 要点:
- arXiv:2606.28538v1 Announce Type: new
- Abstract: We investigate domain adaptation of modern BERT models in the legal domain
- We further pre-train ModernBERT on all US court opinions using the masked language modeling objective
- Although ModernBERT has been trained on roughly 500x more data than original BERT, we still find that this model benefits from further pre-training and domain a…
Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28548v1 公告类型:新。
- 摘要:稀疏自动编码器(SAE)已成为提取语言模型中可解释特征的有用工具。
- 然而,标准 SAE 架构对单个令牌激活进行操作,这意味着活动特征的数量与上下文长度成线性比例,并且研究长模型转录本变得困难。
- 我们引入了回合平均 SAE,它通过学习重建整个回合的平均模型激活来表示具有固定数量特征的单个人类或助理回合。
- EN 要点:
- arXiv:2606.28548v1 Announce Type: new
- Abstract: Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models
- However, standard SAE architectures operate on individual token activations, meaning that the number of active features scales linearly with context length, and…
- We introduce turn-averaged SAEs, which represent a single Human or Assistant turn with a fixed number of features by learning to reconstruct the average model a…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28560v1 公告类型:新。
- 摘要:我们研究稀疏自注意力,其中每个查询关注一个密集的局部窗口加上一组斐波那契间隔的偏移量,并使用每层标量 alpha 来压缩或扩展间隔。
- 在一个匹配配方(60M 参数、512 个隐藏层、16 层、426M 标记)下训练的 21 个语言模型中,我们比较了四种设置跨深度 alpha 的方法:固定、每层学习、静态线性交错、该交错的 coprime(反网格)重新分配,以及范围匹配的 2 幂控制。
- 首先,静态的每层交错改善了固定和学习的 alpha 的困惑度,并且增益与基数无关:将相同的交错应用于 2 的幂基数将其提升到固定斐波那契之上,并与学习的斐波那契注意力相当。
- EN 要点:
- arXiv:2606.28560v1 Announce Type: new
- Abstract: We study sparse self-attention in which each query attends to a dense local window plus a set of Fibonacci-spaced offsets, with a per-layer scalar alp…
- Across 21 language models trained under one matched recipe (60M parameters, 512 hidden, 16 layers, 426M tokens), we compare four ways of setting alpha across de…
- Three results stand out
SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28562v1 公告类型:新。
- 摘要:在策略蒸馏(OPD)具有离线蒸馏和强化学习所没有的属性:教师监督质量取决于学生的能力。
- 不连贯的推出会产生嘈杂的梯度;已经掌握的令牌会产生多余的令牌。
- 这在三个层面(令牌、训练阶段和提示)造成了浪费,但现有方法进行统一监督。
- EN 要点:
- arXiv:2606.28562v1 Announce Type: new
- Abstract: On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence
- Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones
- This creates waste at three scales (tokens, training phases, and prompts) yet existing methods supervise uniformly
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28574v1 公告类型:新。
- 摘要:当大型语言模型 (LLM) 像人类注释者一样对文本中的结构进行编码时,该协议使 LLM 成为可靠的编码器。
- 然而,可靠性并未影响结构有效性。
- 该工具可能是理论幼稚的,通过不满足构造理论提出的任何要求的相关性到达代码,并且除了真正的测量之外,没有当前的方法可以说明这一点。
- EN 要点:
- arXiv:2606.28574v1 Announce Type: new
- Abstract: When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder
- Yet reliability leaves construct validity untouched
- The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct’s theory makes, and no current method tel…
Phonological Perception of Sign Language Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28667v1 公告类型:新。
- 摘要:手语是一种组合系统,其意义是通过结合词汇下的语音参数(例如手形、位置和动作)而产生的。
- 虽然手语识别 (SLR) 深度学习模型在翻译基准上取得了更高的性能,但仍不清楚这些模型是区分抽象语音特征还是仅仅依赖于低级统计相关性。
- 这项工作通过使用最小对探索语音敏感性并评估与人类行为数据的表征一致性,评估了接受美国手语 (ASL) 训练的 SLR 模型的语音感知。
- EN 要点:
- arXiv:2606.28667v1 Announce Type: new
- Abstract: Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and moveme…
- While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these m…
- This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs an…
ArXiv cs.LG (B_intro+search) 链接到标题
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28406v1 公告类型:新。
- 摘要:文本到图像和多模态生成模型越来越多地用于生成科学图形,例如机制图、实验设计示意图、概念框架和图形摘要。
- 然而现有的图像生成基准(例如 GenEval、T2I-CompBench、DPG-Bench)评估自然图像并测量构图、对象计数或照片写实度。
- 它们都没有衡量生成的科学图形的可用因素:正确且清晰的文本标签、对实体及其关系的忠实描述、连贯的图表结构以及对学科绘图惯例的遵守。
- EN 要点:
- arXiv:2606.28406v1 Announce Type: new
- Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch…
- Yet existing image-generation benchmarks (e.g., GenEval, T2I-CompBench, DPG-Bench) evaluate natural images and measure compositionality, object counting, or pho…
- None of them measure what makes a generated scientific figure usable: correct and legible text labels, faithful depiction of entities and their relations, coher…
On the Necessity of a Liquid Substrate for Mesh Intelligence
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28413v1 公告类型:新。
- 摘要:主权代理的网格没有中心:没有共享时钟,没有共享模型,也没有收集数据或重新训练的协调器。
- 它的能力取决于每个智能体将其同伴发出的预测折叠成一个在线的单一内部状态,这些预测来自不规则、未安排时间的观察,在其权重无法重新训练的基底上。
- 这些约束中的任何一个都可以单独处理;一次最佳地折叠在所有三个之下并非如此。
- EN 要点:
- arXiv:2606.28413v1 Announce Type: new
- Abstract: A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain
- Its competence rests on each agent folding the projections its peers emit into a single internal state, online, from observations that arrive at irregular, unsc…
- Any one of these constraints is tractable on its own; folding optimally under all three at once is not
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28433v1 公告类型:新。
- 摘要:强化学习 (RL) 研究的目标之一是了解通用顺序决策,使用基准模拟器作为部署设置中学习的代理。
- 然而,在运行实验时,在模拟器中实现高性能的目标可能会转变为专注于解决模拟器问题。
- 为了获得高分,研究人员可以采用专门用于解决模拟器问题的解决方案,而不是在代理部署在模拟器外部时进行学习。
- EN 要点:
- arXiv:2606.28433v1 Announce Type: new
- Abstract: One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy fo…
- When running experiments, however, the goal of achieving high performance in the simulator can mutate into focusing exclusively on solving the simulator
- To achieve high scores, researchers may adopt solutions exclusively meant for solving simulators, rather than learning while the agent is deployed outside a sim…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28441v1 公告类型:新。
-摘要:在线潜在状态估计构成了人工智能领域的一项基本挑战,作为各种应用的基础工具,包括顺序决策、异常和变化点检测。
- 在本文中,提出了一种新颖的在线分布式传感框架,其中代理协作并交换信息以执行潜在状态估计。
- 所提出的估计器将可用的部分领域知识与深度神经网络的表示能力相结合。
- EN 要点:
- arXiv:2606.28441v1 Announce Type: new
- Abstract: Online latent state estimation constitutes a fundamental challenge within the artificial intelligence field, serving as a foundational tool for divers…
- In this paper, a novel online distributed sensing framework, where agents collaborate and exchange information to perform latent state estimation, is presented
- The proposed estimator combines available partial domain knowledge with the representation capabilities of deep neural networks
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28444v1 公告类型:新。
- 摘要:经典的万能逼近定理建立了 S 形多层感知器的表达能力,但它们没有规定初始权重应如何编码数据分布的几何形状。
- 我们提出了 S-GAI,一种用于单隐藏层 sigmoidal MLP 的光谱几何感知初始化框架。
- 从 sigmoid 单元可以充当平滑半空间门的建设性想法出发,我们从手工指定的平面几何图形转向从图像数据估计的类级谱几何图形。
- EN 要点:
- arXiv:2606.28444v1 Announce Type: new
- Abstract: Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial w…
- We propose S-GAI, a spectral geometry-aware initialization framework for one-hidden-layer sigmoidal MLPs
- Starting from the constructive idea that sigmoid units can act as smooth half-space gates, we move from hand-specified planar geometry to class-wise spectral ge…
scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28459v1 公告类型:新。
- 摘要:单细胞 RNA 测序 (scRNA-seq) 聚类对于识别细胞类型至关重要,但高维、稀疏、丢失和技术噪声阻碍了稳健的表达表示和细胞图构建。
- 现有的掩码自动编码器主要使用表达恢复来进行特征重建,而图聚类方法通常依赖于固定的KNN图,并且不会将恢复的表达反馈回图优化中。
- 我们提出了 scKDGM,一种用于 scRNA-seq 聚类的 KAN 引导的动态图屏蔽学习框架。
- EN 要点:
- arXiv:2606.28459v1 Announce Type: new
- Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical n…
- Existing masked autoencoders mainly use expression recovery for feature reconstruction, while graph clustering methods usually depend on fixed KNN graphs and do…
- We propose scKDGM, a KAN-guided dynamic graph masked learning framework for scRNA-seq clustering
Counterfactual Residual Data Augmentation for Regression
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28460v1 Announce Type: new.
-摘要:现实世界回归任务中的数据驱动建模通常会受到训练样本有限、收集成本高和观察结果嘈杂的影响。
- 受到数据增强对视觉和语言的影响的启发,我们提出了一种用于表格回归的新颖的反事实残差数据增强(CRDA)技术。
- 我们的主要见解是,一旦回归器对数据的系统组成部分进行了建模,剩余的噪声就可以被视为不变的残差,在精心选择的特征的小扰动下保持稳定。
- EN 要点:
- arXiv:2606.28460v1 Announce Type: new
- Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations
- Inspired by the impact of data augmentation in vision and language, we propose a novel Counterfactual Residual Data Augmentation (CRDA) technique for tabular re…
- Our key insight is that once a regressor has modeled the systematic component of the data, the remaining noise can be viewed as an invariant residual that remai…
Singular Learning and Occam’s Razor in Deep Monomial Networks
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28464v1 公告类型:新。
- 摘要:在神经网络的优化中,梯度动态受到模型架构产生的临界点的影响。
- 这些临界点出现在模型参数化的雅可比行列式缺乏秩的地方,并且是奇异学习理论中研究的最明显的奇异点。
- 我们通过多项式代数工具(例如梅森定理)研究具有单项式激活的深度全连接网络中的此类点。
- EN 要点:
- arXiv:2606.28464v1 Announce Type: new
- Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model’s architecture
- These critical points occur where the Jacobian of the model’s parametrization is rank-deficient, and are the most pronounced singularities studied in Singular L…
- We investigate such points in deep fully-connected networks with monomial activations via tools from polynomial algebra such as Mason’s Theorem
An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28467v1 公告类型:新。
- 摘要:办公楼中的设备级能源监控会产生噪音警报,非专业设施管理人员很难使用。
- 本文提出了一种端到端的代理管道,它结合了深度时间序列预测、变分异常检测和基于 LLM 的推理,以生成优先的、可操作的维护建议。
- 该系统使用混合奇异谱分析 (SSA) 和长短期记忆 (LSTM) 预测模型来跟踪七种办公设备,并应用每个设备的 LSTM 变分自动编码器 (VAE),重点关注标记异常的日常消费事件。
- EN 要点:
- arXiv:2606.28467v1 Announce Type: new
- Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use
- This paper proposes an end-to-end agentic pipeline that combines deep time-series forecasting, variational anomaly detection, and LLM-based reasoning to generat…
- The system tracks seven office appliances using a hybrid Singular Spectrum Analysis (SSA) and Long Short-Term Memory (LSTM) forecasting model, and applies a per…
Modelling Emotional Memory in Children with Tensor Networks
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28470v1 公告类型:新。
- 摘要:我们展示了情感效价如何影响儿童认知记忆的顺序依赖结构:对一系列情感效价玩具的正确回忆不仅取决于给定玩具本身的效价,还取决于它之前和之后展示的玩具的效价。
- 虽然标准心理模型确认顺序依赖性在事件(按顺序显示的一组玩具)中有所不同,但准确性较低,并且该模型无法反映对情感对象的记忆如何影响该组中的其他对象。
- 考虑化合价的经典张量网络模型在对研究结果进行建模时能够达到 77.98% 的准确度。
- EN 要点:
- arXiv:2606.28470v1 Announce Type: new
- Abstract: We demonstrate how emotional valence influences the order-dependent structure of children’s recognition memory: correct recall of a sequence of emotio…
- Whilst standard psychological models confirm that order-dependence differs across an event (a set of toys shown in sequence), accuracy is low and the model does…
- A classical tensor network model factoring in valence is able to achieve a 77.98% accuracy in modelling the results of the study