🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-01
- 类型
- ai-daily
- 字数
- 3639
- 阅读时长
- 18 min
2026-07-01 AI日更 | AI 进入科研操作层:从模型能力到真实判断与工作台竞争 链接到标题
今天的主线是 AI 从通用对话和工具调用,进一步走向科研、生命科学等高判断场景。Claude Science、GeneBench-Pro 与多篇规划研究共同指向一个趋势:行业开始更重视可验证推理、专业工作流和真实用户价值,而不只是模型分数与工程效率。
📖 本期 Watch List 深度导读 链接到标题
今天最值得关注的是“智能体如何走向可靠规划”。多篇 arXiv 新作同时指向世界模型、符号反馈、自我修正与幻觉传播控制,说明行业正从“让模型会调用工具”转向“让模型能预演后果、验证路径”。建议重点读 world model planning 与 grounded iterative planning 两篇。
产品侧,Google DeepMind 推出 Nano Banana 2 Lite 和 Gemini Omni Flash,OpenAI 则用 Signals 数据复盘 ChatGPT 全球采用扩张,适合从供给和需求两端观察 AI 平台化趋势。
此外,多智能体人格组成、可解释事实核查、多模态情感推理等论文值得作为延伸阅读:它们共同回答一个问题——当模型进入协作、判断和社会场景,性能之外的可解释性、稳健性与人类适配正在变得更关键。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Andrew Ng Shares Loop Engineering Framework for Faster AI Products 链接到标题
- 分类:AI · News
- 概况:热度时间:6 hours ago,相关帖子数:1000
- 是什么事:吴恩达分享了名为“Loop Engineering”的产品开发框架,主张通过快速构建、评估和迭代来加速 AI 产品落地。
- 为什么重要:这一框架强调从模型能力转向系统化迭代流程,反映出 AI 应用竞争正越来越依赖数据反馈、评估体系和产品工程效率。
- 讨论概况:X 上的讨论主要集中在该框架是否只是对敏捷开发和 MLOps 的重新包装,以及它能否帮助团队更快发现真实用户需求、降低 AI 产品从原型到生产的失败率。
话题 2:Dario Amodei’s 2023 Open-Source AI Warning Resurfaces Amid Chinese Model Boom 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:35000
- 是什么事:Anthropic CEO Dario Amodei 在 2023 年关于开源 AI 可能带来安全与地缘风险的警告,因中国开源模型快速崛起而重新受到关注。
- 为什么重要:这凸显了 AI 开源路线在推动创新、降低门槛的同时,也可能加速前沿能力扩散,并影响全球 AI 竞争与治理格局。
- 讨论概况:X 上讨论集中在开源 AI 是否削弱美国领先优势、是否应加强出口与模型发布限制,以及开放生态对安全风险和产业创新究竟是利大于弊还是弊大于利。
话题 3:Linq Launches Scalable iMessage Apps API for AI Agents 链接到标题
- 分类:AI · News
- 概况:热度时间:3 hours ago,相关帖子数:1100
- 是什么事:Linq 推出面向 AI Agent 的可扩展 iMessage Apps API,允许开发者将智能代理能力接入苹果 iMessage 应用生态。
- 为什么重要:这意味着 AI Agent 可能进入更高频的私人通信场景,推动聊天、任务执行、自动化服务与移动端消息平台的深度结合。
- 讨论概况:X 上的讨论主要集中在该 API 是否能降低 AI Agent 落地门槛、iMessage 封闭生态中的平台限制,以及隐私、安全和自动化消息滥用风险。
话题 4:X Launches Hosted Servers for AI Agents to Access Real-Time Data 链接到标题
- 分类:AI · News
- 概况:热度时间:22 hours ago,相关帖子数:9600
- 是什么事:X 推出托管式 MCP(Model Context Protocol)服务器,使 Grok、Cursor、Claude Desktop 等 AI Agent 可更便捷地接入 X API 获取实时数据。
- 为什么重要:这降低了 AI Agent 连接实时信息源的门槛,有助于提升代理系统在新闻、搜索、监测和决策类任务中的时效性与实用性,也可能推动 MCP 成为智能体工具接入的重要标准。
- 讨论概况:X 上的讨论主要集中在开发者是否会因此更快构建具备实时感知能力的 Agent;支持者称其为“游戏规则改变者”,质疑者则关注 API 成本、数据质量、平台依赖、隐私与滥用风险。
话题 5:Etched Emerges from Stealth with Transformer-Specific AI Chips 链接到标题
- 分类:AI · News
- 概况:热度时间:8 hours ago,相关帖子数:4900
- 是什么事:AI 芯片初创公司 Etched 走出隐身状态,发布面向 Transformer 架构专门优化的 AI 推理芯片。
- 为什么重要:这显示 AI 硬件竞争正从通用 GPU 转向针对主流模型架构的专用加速器,可能影响推理成本、能效和英伟达主导地位。
- 讨论概况:X 上讨论集中在专用芯片能否真正降低大模型推理成本、Transformer 架构是否会长期占主流,以及 Etched 相比 GPU 和其他 ASIC 方案的性能、供货与商业化风险。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线是:AI 竞争正从单点模型能力转向“工程化落地、实时数据接入、平台分发与专用硬件”组成的系统竞争。较大共识在于,快速迭代框架、Agent 接入消息与社交数据、以及推理芯片优化,都会降低 AI 应用部署门槛并提升实用性,但真正价值取决于评估体系、用户需求验证、数据质量和成本结构。主要分歧集中在开放与封闭、通用与专用之间:开源 AI 究竟是创新加速器还是安全与地缘风险放大器,MCP 和 iMessage 这类平台接口是生态机会还是新的平台依赖,Transformer 专用芯片是趋势判断还是架构押注。潜在风险则包括前沿能力扩散带来的治理压力、Agent 进入私人通信和实时信息流后的隐私与滥用问题,以及企业在 API、硬件路线和封闭生态上形成新的锁定与商业化不确定性。
💡 大佬观点(Influencer Insights) 链接到标题
好的,基于这些推文内容,以下是我为您整理的 AI 行业 24 小时热点洞察报告。
AI 行业日度观察 (06/30) 链接到标题
1. 今日大佬们共同关注的技术趋势或产品热点 链接到标题
今日的核心热潮明显集中在 AI Agent 的深度工具化、平台化,以及围绕 Codex 生态的实用技巧上。同时,模型军备竞赛已从基座模型卷到了垂直行业操作层和端侧。
🔥 焦点一:Anthropic 双连发——Sonnet 5 与 Claude Science Anthropic 在同一天发布了两款重磅产品,几乎主导了今日一半的讨论。
- Claude Sonnet 5:@dotey (宝玉) 详细分析了其定位——用更低的价格(API 价格仅为 Opus 4.8 的 40%)提供了接近顶级模型的 Agent 能力。它替代了 Sonnet 4.6,旨在缩短普通模型与顶级模型在自主规划、多步骤任务上的差距,尤其在 Agent 编程基准上提升显著。
- Claude Science:@dotey 指出,这是 Anthropic 将 AI 从单纯的模型能力转向“特定行业操作层”的战略产品。它并非新模型,而是为科研工作者(特别是生命科学)整合了 60+ 科学数据库、计算资源和协作 Agent 的工作台,旨在复刻 Claude Code 在软件工程领域的成功。这标志着 AI 正在从“对话工具”进化为“垂直领域操作系统”。
💻 趋势二:Codex 生态的深耕与“解密” @Pluvio9yte (雪踏乌云) 和 @dotey 围绕 Codex 分享了大量深度内容,显示该工具已进入实用技巧和生态扩展阶段。
- 实用技巧传播:@Pluvio9yte 分享并总结了多篇关于 Codex 的实战指南,包括额度管理、记忆功能、
/goal模式收尾等硬核技巧,甚至曝光了疑似/goal无限额度的 Bug。 - 生态扩展:从 @AI_Jasonyu 分享的 Windows 端桌面灵动岛工具,到 @gefei55 (哥飞) 介绍的开源项目
DevSpace(让 ChatGPT 网页版也能像 Codex 一样操作本地代码),Codex 的生态正在被社区迅速丰富。 - 透明度争议:@dotey 报道了一个重大指控——安全研究员逆向工程发现 Claude Code 会通过隐蔽的 Unicode 字符在中国代理用户的系统提示词中“打水印”,引发了关于 AI 工具信任与隐私的激烈讨论。
- 实用技巧传播:@Pluvio9yte 分享并总结了多篇关于 Codex 的实战指南,包括额度管理、记忆功能、
🚀 趋势三:本地/端侧模型的持续升温 @zhixianio (知县) 持续分享本地模型测试心得,如 MiniCPM-o 4.5 的音视频全双工效果令人满意,但稳定性有待提升。这印证了端侧智能既充满希望也面临工程化挑战,而 @zhixianio 的播客《认知有县》恰好也聚焦此话题。
2. 值得注意的独特观点或行业前瞻 链接到标题
🙈 AI 工程指标与用户价值的脱节:@dotey (宝玉) 敏锐地捕捉到 Anthopic 为 Spotify 做的宣传视频在 X 上翻车。宣传强调了 4500 次日部署、73% AI 辅助 PR 等工程侧数字,但用户纷纷吐槽产品体验变差。这尖锐地指出当前 AI 应用的一个根本问题:用于衡量 AI 价值的“生产效率”指标(代码行数、部署次数),与用户最终感知的“产品质量”之间存在巨大鸿沟。行业需要新的价值衡量标尺。
🕵️ “隐蔽信道”:对 AI 工具信任度的拷问:@dotey (宝玉) 详细转述了关于 Claude Code 在系统提示词中嵌入隐秘字符标记中国代理用户的指控。无论这是反滥用措施还是隐私侵犯,它引发了公众对拥有深度系统访问权限的 AI 工具(能读代码、运行命令)的担忧。用户有权知道工具在背后做了什么,这种不公开的“标记”行为对信任的侵蚀是巨大的。
🧬 非侵入式脑机接口的里程碑:@dotey (宝玉) 报道的 Meta Brain2Qwerty v2,将非侵入式脑解码的单词准确率从业界 8% 提升到 61%,最高达 78%。@dotey 的观点很关键:这证明了“不开刀也可能做到接近开刀的效果,剩下的是工程问题而非原理问题。” 这对广大无法接受开颅手术的脑损伤患者来说,具有深远的意义。
🧠 关于“AI 开源”的异议:@ruanyf (阮一峰) 引述了 Anthropic 创始人的观点,认为AI模型只公开权重,看不到内部运作,因此只是“开放权重”而非传统“开源”。这个观点戳破了当前关于“开源”大模型的泡沫,引发了对开放性真实定义的深度思考。
👴 AI 与人类专家的共生关系:@dotey (宝玉) 分享了福特汽车重新雇用 350 名资深工程师(“gray beard”)的案例,因为 AI 质检系统未达预期,需要老师傅来调教 AI。这提供了一个在主流“AI 取代工作”叙事之外的珍贵反例:AI 的成功落地,可能恰恰更加依赖于顶尖的人类经验和判断力。
3. 推荐的工具或资源 链接到标题
开发与工具
- Orca (@LinearUncle / @dotey 转发): 一个被推荐为可媲美甚至超越 Codex App 的开源 Coding IDE,支持全平台。
- Codex 实用工具:
- Windows 桌面 Codex 灵动岛 (@mooyuking): 帮助 Windows 用户监控状态和额度。
- ChatGPT 批量删除插件 (@Pluvio9yte): 免费、本地,用于清理 ChatGPT 网页版聊天记录。
- oMLX v0.4.0 (@jundotkim / @zhixianio 转发): 苹果芯片上运行模型的本地应用,发布了首个原生 Swift macOS 应用。
- 飞书开源 CLI 工具包 (@ruanyf 推荐): 用于让 AI Agent 调用办公功能,目前已超 1万 Star。
学习与资源
- 《Claude Code From Scratch》 (@dotey 推荐): 一本开源电子书,用约 4300 行代码复刻 Claude Code 核心架构,是理解 coding agent 原理的精良教程。
- Codex 橙皮书 (@bozhou_ai / @AI_Jasonyu 转发): 开源、系统化的 Codex 学习资料。
- GEO 内容工程资料包 (@vista8): 一套关于生成式引擎优化(GEO)的系统性资源,包括操作手册、技能包和演示。
- 视频制作 Skills 仓库 (@Pluvio9yte): 一套开源的视频制作技能,可复刻类似 HyperFrames 的效果,适合不会剪辑的用户。
大模型与应用
- Apodex 4B (@Pluvio9yte): 一款被定位为“私人深度研究助手”的本地模型,相关部署踩坑记录已发布。
- YouMind 1.0 (@lifesinger / @gefei55 推荐): 一款帮助创作者高效产出图文内容并一键分发至 X 和公众号的工具。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 37 条更新
Lex Fridman Podcast (A_full) 链接到标题
- #498 – Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires
- 发布时间:2026-07-01 05:33 北京时间
- 摘要:- 安东尼·卡德里斯(Anthony Kaldellis)是罗马帝国历史学家,也是拜占庭帝国(东罗马帝国)综合史《新罗马帝国》的作者。
- 请参阅下面的时间戳、文字记录,并提供反馈、提交问题、联系 Lex 等。
- Upwork:招聘自由职业者的平台。
- Fin:用于客户服务的人工智能代理。
- BetterHelp:在线治疗和咨询。
- EN 要点:
- Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire…
- Thank you for listening ❤ Check out our sponsors:
- See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
- CONTACT LEX:
OpenAI Blog (A_full) 链接到标题
How ChatGPT adoption has expanded
- 发布时间:2026-06-30 17:00 北京时间
- 摘要:- ChatGPT 的采用正在全球范围内扩大和深化。
- 新的 OpenAI Signals 数据显示,随着时间的推移,人们更频繁地使用 ChatGPT,并执行更广泛的任务,而其用户群正变得更加全球化和多样化。
- OpenAI Signals 使用聚合数据来衡量人们随着时间的推移如何与个人 ChatGPT 计划(包括 Free、Go、Plus 和 Pro 计划)互动。
- 此分析提供了随着 ChatGPT 达到全球规模,个人人工智能使用如何演变的观点。
- 下面的可视化突出显示了全球人工智能采用的几个核心趋势:自推出以来 ChatGPT 的使用如何加深,人们将 ChatGPT 纳入日常工作的方式,以及不同语言和地区之间的差异。
- EN 要点:
- New OpenAI Signals data shows how ChatGPT adoption is growing globally, with users increasing usage, exploring more capabilities, and driving growth across regi…
- 发布时间:2026-06-30 08:00 北京时间
- 摘要:- 科学数据很少附带说明。
- 研究人员必须决定一种模式是否反映了生物学或噪音,数据是否可以支持所提出的问题,以及每个结果应如何改变他们下一步的工作。
- 人工智能代理执行复杂分析的能力越来越强,但真正的科学研究也不仅仅依赖于回忆事实或遵循预定义的工作流程,还依赖于做出这些更高阶的判断。
- 今天,我们推出 GeneBench-Pro——一个具有挑战性的研究级基准,用于测试模型是否能够处理现实世界计算生物学所需的那种需要大量判断的分析。
- 迄今为止,对系统级判断调用的令人信服的评估很少,这使得现实世界的计算研究变得困难。
- EN 要点:
- Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.
Core dump epidemiology: fixing an 18-year-old bug
- 发布时间:2026-06-30 08:00 北京时间
- 摘要:- OpenAI 工程师使用大规模核心转储分析来调试罕见的基础设施崩溃,发现硬件故障和长期存在的软件错误。
- OpenAI 博客中的这篇文章解释了核心转储流行病学:修复 18 年前的错误如何塑造更广泛的人工智能和基础设施格局。
- 它还揭示了核心转储流行病学对创始人、运营商和投资者的实际影响:修复 18 年的错误。
- EN 要点:
- OpenAI engineers used large-scale core dump analysis to debug rare infrastructure crashes, uncovering both a hardware fault and a long-standing software bug.
- 发布时间:2026-06-30 08:00 北京时间
- 摘要:- Genebench-Pro 内部。
- OpenAI 博客的这篇文章解释了 Genebench-Pro 内部如何塑造更广泛的人工智能和基础设施格局。
- 它还为遵循 Inside Genebench-Pro 的创始人、运营商和投资者提供了实际意义。
- EN 要点:
- Inside Genebench-Pro
Google DeepMind Blog (A_full) 链接到标题
- Start building with Nano Banana 2 Lite and Gemini Omni Flash
- 发布时间:2026-07-01 00:02 北京时间
- 摘要:- 开始使用 Nano Banana 2 Lite 和 Gemini Omni Flash 进行构建。
- 这篇来自 Google DeepMind 博客的文章解释了如何使用 Nano Banana 2 Lite 和 Gemini Omni Flash 开始构建,塑造更广泛的人工智能和基础设施格局。
- 在开始使用 Nano Banana 2 Lite 和 Gemini Omni Flash 进行构建之后,它还为创始人、运营商和投资者带来了实际影响。
- EN 要点:
- Start building with Nano Banana 2 Lite and Gemini Omni Flash
Lex Fridman (B_intro+search) 链接到标题
- Anthony Kaldellis: Roman Empire, Byzantine Empire, Rise & Fall of Empires | Lex Fridman Podcast #498
- 发布时间:2026-07-01 05:16 北京时间
- 摘要:- 安东尼·卡德里斯(Anthony Kaldellis)是罗马帝国历史学家,也是拜占庭帝国(东罗马帝国)综合史《新罗马帝国》的作者。
- 请参阅下面的时间戳、文字记录,并提供反馈、提交问题、联系 Lex 等。
- 反馈 - 向 Lex 提供反馈:。
- AMA - 提交问题、视频或致电:。
- EN 要点:
- Anthony Kaldellis is a historian of the Roman Empire and author of “The New Roman Empire”, a comprehensive history of the Byzantine Empire (Eastern Roman Empire…
- Thank you for listening ❤ Check out our sponsors:
- See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
- Transcript:
ArXiv cs.AI (B_intro+search) 链接到标题
AI-Model Network: Concept, Current State and Future
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27382v1 公告类型:新。
- 摘要:计算机的主要功能在于计算和处理,而互联网的核心价值则植根于共享和协作。
- 计算机创造互联网,互联网赋予计算机价值。
- 互联网、云计算、大数据的快速发展正在推动人工智能进入大模型(LM)时代。
- EN 要点:
- arXiv:2606.27382v1 Announce Type: new
- Abstract: While the primary function of computers lies in computation and processing, the core value of the Internet is rooted in sharing and collaboration
- Computers create the Internet, and the Internet empowers the value of computers
- The rapid development of the Internet, cloud computing, and big data is pushing artificial intelligence into the era of large models (LMs)
When Does Personality Composition Matter for Multi-Agent LLM Teams?
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27443v1 公告类型:新。
- 摘要:个性提示决定了大型语言模型的沟通方式,但这些行为转变是否影响客观任务结果仍有待探索。
- 先前的研究表明,低宜人性提示的智能体会产生对抗性语言,而高宜人性提示的智能体会变得合作,但沟通方式与任务绩效之间的关系尚未在多个领域进行系统地检验。
- 在这项工作中,我们通过在结构化编码、开放式研究合作和竞争性谈判这三个任务领域操纵前沿法学硕士的人格特质,研究人格构成是否对多智能体团队绩效重要。
- EN 要点:
- arXiv:2606.27443v1 Announce Type: new
- Abstract: Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes remains under-e…
- Prior work shows that agents prompted with low agreeableness produce adversarial language, while those prompted with high agreeableness become cooperative, but…
- In this work, we investigate whether personality composition matters for multi-agent team performance by manipulating personality traits across frontier LLMs on…
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27483v1 公告类型:新。
-摘要:大语言模型(LLM)智能体在顺序决策方面表现出了强大的能力,但它们在长期任务中仍然基本上是反应性的。
- 与在承诺之前使用“假设”推理来评估潜在计划的人类不同,标准代理缺乏内部世界模型来模拟未来的结果。
- 因此,我们建议通过训练单个自回归模型来内部化未来感知规划,以表达预期状态的推出和计划条件的成功估计(Q 值的文本模拟)。
- EN 要点:
- arXiv:2606.27483v1 Announce Type: new
- Abstract: Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-h…
- Unlike humans who employ “what-if” reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outc…
- Therefore, we propose to internalize future-aware planning by training a single autoregressive model to verbalize both a prospective state rollout and a plan-co…
Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27593v1 公告类型:新。
-摘要:我们引入了一个名为 ODYSSEY 的分类框架,用于构建可验证的、本地保真基础模型作为铸造厂的组合:指定本地上下文、本地表示族、限制图、粘合规则、阻塞策略、更新义务和面向人类的视图的构建块架构组件。
- 铸造厂是一个有组织的知识库,其中包含论证成分。
- 混凝土铸造厂是根据通用铸造厂建造的,例如证据/论证、运营决策、机构/财务、市场意义、科学挑战、研究计划、辅助建造和评估装备铸造厂。
- EN 要点:
- arXiv:2606.27593v1 Announce Type: new
- Abstract: We introduce a categorical framework called ODYSSEY for constructing verifiable, local truth-preserving foundation models as compositions of foundries…
- A foundry is an organized sheaf of knowledge that carries within it an argumentation component
- Concrete foundries are built from generic foundries such as evidence/argument, operational decision, institutional/financial, market meaning, scientific challen…
DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27619v1 公告类型:新。
- 摘要:阅读障碍学习者越来越多地使用人工智能 (AI) 工具来支持阅读、写作、组织和学习相关任务。
- 然而,他们使用这些工具的生活经验在很大程度上仍未得到充分检验。
- 本文提出了 DysLexLens,这是一种低资源 LLM 框架,旨在通过在线论坛讨论来分析阅读障碍学习者使用 AI 的体验。
- EN 要点:
- arXiv:2606.27619v1 Announce Type: new
- Abstract: Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks
- However, their lived experiences with these tools remain largely underexamined
- This paper proposes DysLexLens, a low-resource LLM framework, designed to analyse dyslexic learners experience with AI through online forum discussions
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27652v1 公告类型:新。
- 摘要:我们发现显式推理不一定会转化为更好的多模态情感识别(MER)准确性,尽管它使预测更容易解释。
- 具体来说,对于基于推理的 MLLM,通过触发直接答案的快速思考通常胜过深思熟虑推理后的缓慢思考。
- 我们的实证分析表明,快速思考可以通过更广泛、更自信的预测来提高回忆能力,而缓慢思考则可以通过保守地过滤不正确的类别来提高精确度。
- EN 要点:
- arXiv:2606.27652v1 Announce Type: new
- Abstract: We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predict…
- Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberative reasoning
- Our empirical analyses show that fast thinking improves recall with broader and more confident predictions, whereas slow thinking favors precision through conse…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27736v1 公告类型:新。
-摘要:假新闻的快速传播对信息生态系统构成了越来越大的威胁,特别是在生成引擎优化(GEO)中毒下人工智能生成的错误信息使得检索系统系统性地呈现出敌对制作的内容,从而污染了法学硕士的推理。
- 在本文中,我们提出了证据树(ToE),这是一种用于自动事实检查的分层证据推理框架,将每个主张建模为动态扩展的论证树。
- ToE集成了强化学习驱动的多源检索代理、证据评估代理和论据树聚合算法,通过可解释的证据链迭代分解、检索和验证主张。
- EN 要点:
- arXiv:2606.27736v1 Announce Type: new
- Abstract: The rapid spread of fake news poses increasing threats to information ecosystems, especially as AI-generated misinformation under Generative Engine Op…
- In this paper, we propose Tree of Evidence (ToE), a hierarchical evidence reasoning framework for automated fact-checking that models each claim as a dynamicall…
- ToE integrates a reinforcement learning-driven multi-source retrieval agent, an evidence evaluation agent, and an argument tree aggregation algorithm to iterati…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27757v1 公告类型:新。
-摘要:大型语言模型(LLM)已引起学术界和工业界的广泛关注,但其部署引发了有关鲁棒性和可靠性的关键安全问题。
- 规划是智能行为的核心组成部分,对于法学硕士来说仍然具有挑战性,由于固有的复杂性,他们经常在长期决策任务中产生不可行或不正确的解决方案。
- 在本文中,我们提出了一种符号反馈驱动的迭代自我完善框架,以增强法学硕士在长期规划中的稳健性和可靠性。
- EN 要点:
- arXiv:2606.27757v1 Announce Type: new
- Abstract: Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security concerns re…
- Planning, a core component of intelligent behavior, remains challenging for LLMs, which often produce infeasible or incorrect solutions in long-horizon decision…
- In this paper, we propose a symbolic feedback-driven iterative self-refinement framework to enhance the robustness and reliability of LLMs in long-horizon plann…
Understanding Rollout Error in Graph World Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27780v1 公告类型:新。
- 摘要:世界模型通常用于通过向前滚动学习动态来进行规划。
- 然而,许多规划环境不是矢量或图像;它们是代理、工具、技能、路线和依赖关系的图表。
- 在这些设置中,局部预测误差可能保留在局部或在图形中传播,并且当边缘被预测而不是固定时,故障模式会再次发生变化。
- EN 要点:
- arXiv:2606.27780v1 Announce Type: new
- Abstract: World models are often used for planning by rolling learned dynamics forward
- Many planning environments, however, are not vectors or images; they are graphs of agents, tools, skills, routes, and dependencies
- In these settings, a local prediction error may stay local or spread through the graph, and the failure mode changes again when edges are predicted rather than…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.27806v1 公告类型:新。
- 摘要:语言代理的世界模型有两种有用的形式。
- 基于代理的世界模型调用LLM API并用语言灵活地进行推理,但其错误表现为幻觉的状态变化,很难用普通的回归损失来评分。
- 参数化世界模型是经过训练的转换预测器;它的错误更容易用 NodeMSE、增量准确度和有效性准确度等数量来衡量,但作为独立规划器,它通常较弱。
- EN 要点:
- arXiv:2606.27806v1 Announce Type: new
- Abstract: World models for language agents come in two useful forms
- An agent-based world model calls an LLM API and reasons flexibly in language, but its errors appear as hallucinated state changes that are hard to score with or…
- A parameterized world model is a trained transition predictor; its errors are easier to measure with quantities such as NodeMSE, delta accuracy, and validity ac…
ArXiv cs.CL (B_intro+search) 链接到标题
Generating in the Limit with Infinitely Many Hallucinations
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28354v1 公告类型:新。
- 摘要:极限语言识别的经典范式将学习建模为对手(揭示未知目标语言的字符串)和学习者(负责识别该语言)之间的游戏。
- 最近引入的极限语言生成框架改变了目标,以更好地反映现代语言建模,要求学习者从目标语言中生成有效的、看不见的字符串。
- 相关工作强调了一种根本性的紧张关系:目标的广泛覆盖往往是以有效性为代价的。
- EN 要点:
- arXiv:2606.28354v1 Announce Type: new
- Abstract: The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown targe…
- The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to…
- Related work highlighted a fundamental tension: a broad coverage of the target often comes at the cost of validity
Extracting Knowledge from an Arabic-English Machine-Readable Dictionary Using Information Extraction
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28457v1 公告类型:新。
- 摘要:自然语言处理(NLP)应用需要大量且丰富的语言知识。
- 此外,词典、百科全书和语料库等电子语言资源也变得可用。
- 因此,出现了自动方法来从这些来源中提取词汇信息,以克服知识获取瓶颈。
- EN 要点:
- arXiv:2606.28457v1 Announce Type: new
- Abstract: Natural language processing (NLP) applications need large and rich amount of linguistic knowledge
- Furthermore, electronic language sources such as dictionaries, encyclopedia, and corpora became available
- So, automatic methods are emerged to extract lexical information from those sources to overcome the knowledge acquisition bottleneck
Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28524v1 公告类型:新。
- 摘要:最近的工作表明,大型语言模型(LLM)对文本描述的主体的信念状态敏感,如通过错误信念任务(FBT)测量的那样,但对结构有效性的持续关注仍然存在。
- 我们采用发展的视角,在 Olmo2 和 Pythia 语言模型套件的多个训练阶段中追踪心理状态推理行为的模式,以及这种行为可能的先决条件。
- 我们发现,高于机会的 FBT 表现取决于模型大小和足够的训练量,在预训练中出现相对较晚,并且在最容易诊断心智化(错误信念、隐式)的情况下通过训练后干预(SFT、DPO)得到最大改善。
- EN 要点:
- arXiv:2606.28524v1 Announce Type: new
- Abstract: Recent work suggests that Large Language Models (LLMs) are sensitive to the belief states of agents described by text, as measured by the false belief…
- We adopt a developmental perspective, tracing the pattern of mental state reasoning behavior – and likely preconditions for this behavior – across mul…
- We find that above-chance FBT performance depends both on model size and sufficient training volume, emerges relatively late in pretraining, and is most improve…
A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28526v1 公告类型:新。
-摘要:医学生的临床和沟通技能通常通过客观结构化临床考试(OSCE)进行评估,其中包括对医患互动的简短场景驱动模拟。
- 然而,培训往往受到人类标准化患者的可用性较低的限制,这激励了现实虚拟患者(VP)的开发。
- 为了解决这一差距,我们引入了法国 OSCE 对话数据集,其中包含 240 个学生与患者的培训互动。
- EN 要点:
- arXiv:2606.28526v1 Announce Type: new
- Abstract: The clinical and communication skills of medical students are commonly assessed through Objective Structured Clinical Examinations (OSCEs), which cons…
- However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs)
- To address this gap, we introduce a French OSCE dialogue dataset comprising 240 student-patient training interactions
Legal Domain Adaptation of Modern BERT Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28538v1 公告类型:新。
- 摘要:我们研究现代 BERT 模型在法律领域的领域适应性。
- 我们使用屏蔽语言建模目标进一步针对所有美国法院意见对 ModernBERT 进行预训练。
- 尽管 ModernBERT 的训练数据比原始 BERT 多了大约 500 倍,但我们仍然发现该模型受益于法律领域的进一步预训练和领域适应:我们报告称,与普通的 ModernBERT 相比,在与美国法院意见相关的所有数据集上都有显着改进。
- EN 要点:
- arXiv:2606.28538v1 Announce Type: new
- Abstract: We investigate domain adaptation of modern BERT models in the legal domain
- We further pre-train ModernBERT on all US court opinions using the masked language modeling objective
- Although ModernBERT has been trained on roughly 500x more data than original BERT, we still find that this model benefits from further pre-training and domain a…
Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28548v1 公告类型:新。
- 摘要:稀疏自动编码器(SAE)已成为提取语言模型中可解释特征的有用工具。
- 然而,标准 SAE 架构对单个令牌激活进行操作,这意味着活动特征的数量与上下文长度成线性比例,并且研究长模型转录本变得困难。
- 我们引入了回合平均 SAE,它通过学习重建整个回合的平均模型激活来表示具有固定数量特征的单个人类或助理回合。
- EN 要点:
- arXiv:2606.28548v1 Announce Type: new
- Abstract: Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models
- However, standard SAE architectures operate on individual token activations, meaning that the number of active features scales linearly with context length, and…
- We introduce turn-averaged SAEs, which represent a single Human or Assistant turn with a fixed number of features by learning to reconstruct the average model a…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28560v1 公告类型:新。
- 摘要:我们研究稀疏自注意力,其中每个查询关注一个密集的局部窗口加上一组斐波那契间隔的偏移量,并使用每层标量 alpha 来压缩或扩展间隔。
- 在一个匹配配方(60M 参数、512 个隐藏层、16 层、426M 标记)下训练的 21 个语言模型中,我们比较了四种设置跨深度 alpha 的方法:固定、每层学习、静态线性交错、该交错的 coprime(反网格)重新分配,以及范围匹配的 2 幂控制。
- 首先,静态的每层交错改善了固定和学习的 alpha 的困惑度,并且增益与基数无关:将相同的交错应用于 2 的幂基数将其提升到固定斐波那契之上,并与学习的斐波那契注意力相当。
- EN 要点:
- arXiv:2606.28560v1 Announce Type: new
- Abstract: We study sparse self-attention in which each query attends to a dense local window plus a set of Fibonacci-spaced offsets, with a per-layer scalar alp…
- Across 21 language models trained under one matched recipe (60M parameters, 512 hidden, 16 layers, 426M tokens), we compare four ways of setting alpha across de…
- Three results stand out
SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28562v1 公告类型:新。
- 摘要:在策略蒸馏(OPD)具有离线蒸馏和强化学习所没有的属性:教师监督质量取决于学生的能力。
- 不连贯的推出会产生嘈杂的梯度;已经掌握的令牌会产生多余的令牌。
- 这在三个层面(令牌、训练阶段和提示)造成了浪费,但现有方法进行统一监督。
- EN 要点:
- arXiv:2606.28562v1 Announce Type: new
- Abstract: On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence
- Incoherent rollouts yield noisy gradients; already-mastered tokens yield redundant ones
- This creates waste at three scales (tokens, training phases, and prompts) yet existing methods supervise uniformly
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28574v1 公告类型:新。
- 摘要:当大型语言模型 (LLM) 像人类注释者一样对文本中的结构进行编码时,该协议使 LLM 成为可靠的编码器。
- 然而,可靠性并未影响结构有效性。
- 该工具可能是理论幼稚的,通过不满足构造理论提出的任何要求的相关性到达代码,并且除了真正的测量之外,没有当前的方法可以说明这一点。
- EN 要点:
- arXiv:2606.28574v1 Announce Type: new
- Abstract: When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder
- Yet reliability leaves construct validity untouched
- The instrument may be theory-naive, reaching the code through a correlate that meets none of the demands the construct’s theory makes, and no current method tel…
Phonological Perception of Sign Language Models
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28667v1 公告类型:新。
- 摘要:手语是一种组合系统,其意义是通过结合词汇下的语音参数(例如手形、位置和动作)而产生的。
- 虽然手语识别 (SLR) 深度学习模型在翻译基准上取得了更高的性能,但仍不清楚这些模型是区分抽象语音特征还是仅仅依赖于低级统计相关性。
- 这项工作通过使用最小对探索语音敏感性并评估与人类行为数据的表征一致性,评估了接受美国手语 (ASL) 训练的 SLR 模型的语音感知。
- EN 要点:
- arXiv:2606.28667v1 Announce Type: new
- Abstract: Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and moveme…
- While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these m…
- This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs an…
ArXiv cs.LG (B_intro+search) 链接到标题
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28406v1 公告类型:新。
- 摘要:文本到图像和多模态生成模型越来越多地用于生成科学图形,例如机制图、实验设计示意图、概念框架和图形摘要。
- 然而现有的图像生成基准(例如 GenEval、T2I-CompBench、DPG-Bench)评估自然图像并测量构图、对象计数或照片写实度。
- 它们都没有衡量生成的科学图形的可用因素:正确且清晰的文本标签、对实体及其关系的忠实描述、连贯的图表结构以及对学科绘图惯例的遵守。
- EN 要点:
- arXiv:2606.28406v1 Announce Type: new
- Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch…
- Yet existing image-generation benchmarks (e.g., GenEval, T2I-CompBench, DPG-Bench) evaluate natural images and measure compositionality, object counting, or pho…
- None of them measure what makes a generated scientific figure usable: correct and legible text labels, faithful depiction of entities and their relations, coher…
On the Necessity of a Liquid Substrate for Mesh Intelligence
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28413v1 公告类型:新。
- 摘要:主权代理的网格没有中心:没有共享时钟,没有共享模型,也没有收集数据或重新训练的协调器。
- 它的能力取决于每个智能体将其同伴发出的预测折叠成一个在线的单一内部状态,这些预测来自不规则、未安排时间的观察,在其权重无法重新训练的基底上。
- 这些约束中的任何一个都可以单独处理;一次最佳地折叠在所有三个之下并非如此。
- EN 要点:
- arXiv:2606.28413v1 Announce Type: new
- Abstract: A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain
- Its competence rests on each agent folding the projections its peers emit into a single internal state, online, from observations that arrive at irregular, unsc…
- Any one of these constraints is tractable on its own; folding optimally under all three at once is not
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28433v1 公告类型:新。
- 摘要:强化学习 (RL) 研究的目标之一是了解通用顺序决策,使用基准模拟器作为部署设置中学习的代理。
- 然而,在运行实验时,在模拟器中实现高性能的目标可能会转变为专注于解决模拟器问题。
- 为了获得高分,研究人员可以采用专门用于解决模拟器问题的解决方案,而不是在代理部署在模拟器外部时进行学习。
- EN 要点:
- arXiv:2606.28433v1 Announce Type: new
- Abstract: One goal in reinforcement learning (RL) research is to understand general-purpose sequential decision-making, using benchmark simulators as a proxy fo…
- When running experiments, however, the goal of achieving high performance in the simulator can mutate into focusing exclusively on solving the simulator
- To achieve high scores, researchers may adopt solutions exclusively meant for solving simulators, rather than learning while the agent is deployed outside a sim…
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28441v1 公告类型:新。
-摘要:在线潜在状态估计构成了人工智能领域的一项基本挑战,作为各种应用的基础工具,包括顺序决策、异常和变化点检测。
- 在本文中,提出了一种新颖的在线分布式传感框架,其中代理协作并交换信息以执行潜在状态估计。
- 所提出的估计器将可用的部分领域知识与深度神经网络的表示能力相结合。
- EN 要点:
- arXiv:2606.28441v1 Announce Type: new
- Abstract: Online latent state estimation constitutes a fundamental challenge within the artificial intelligence field, serving as a foundational tool for divers…
- In this paper, a novel online distributed sensing framework, where agents collaborate and exchange information to perform latent state estimation, is presented
- The proposed estimator combines available partial domain knowledge with the representation capabilities of deep neural networks
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28444v1 公告类型:新。
- 摘要:经典的万能逼近定理建立了 S 形多层感知器的表达能力,但它们没有规定初始权重应如何编码数据分布的几何形状。
- 我们提出了 S-GAI,一种用于单隐藏层 sigmoidal MLP 的光谱几何感知初始化框架。
- 从 sigmoid 单元可以充当平滑半空间门的建设性想法出发,我们从手工指定的平面几何图形转向从图像数据估计的类级谱几何图形。
- EN 要点:
- arXiv:2606.28444v1 Announce Type: new
- Abstract: Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial w…
- We propose S-GAI, a spectral geometry-aware initialization framework for one-hidden-layer sigmoidal MLPs
- Starting from the constructive idea that sigmoid units can act as smooth half-space gates, we move from hand-specified planar geometry to class-wise spectral ge…
scKDGM: KAN-guided Dynamic Graph Masked Learning for Single-Cell RNA-seq Clustering
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28459v1 公告类型:新。
- 摘要:单细胞 RNA 测序 (scRNA-seq) 聚类对于识别细胞类型至关重要,但高维、稀疏、丢失和技术噪声阻碍了稳健的表达表示和细胞图构建。
- 现有的掩码自动编码器主要使用表达恢复来进行特征重建,而图聚类方法通常依赖于固定的KNN图,并且不会将恢复的表达反馈回图优化中。
- 我们提出了 scKDGM,一种用于 scRNA-seq 聚类的 KAN 引导的动态图屏蔽学习框架。
- EN 要点:
- arXiv:2606.28459v1 Announce Type: new
- Abstract: Single-cell RNA sequencing (scRNA-seq) clustering is essential for identifying cell types, but high dimensionality, sparsity, dropout, and technical n…
- Existing masked autoencoders mainly use expression recovery for feature reconstruction, while graph clustering methods usually depend on fixed KNN graphs and do…
- We propose scKDGM, a KAN-guided dynamic graph masked learning framework for scRNA-seq clustering
Counterfactual Residual Data Augmentation for Regression
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28460v1 公告类型:新。
-摘要:现实世界回归任务中的数据驱动建模通常会受到训练样本有限、收集成本高和观察结果嘈杂的影响。
- 受到数据增强对视觉和语言的影响的启发,我们提出了一种用于表格回归的新颖的反事实残差数据增强(CRDA)技术。
- 我们的主要见解是,一旦回归器对数据的系统组成部分进行了建模,剩余的噪声就可以被视为不变的残差,在精心选择的特征的小扰动下保持稳定。
- EN 要点:
- arXiv:2606.28460v1 Announce Type: new
- Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations
- Inspired by the impact of data augmentation in vision and language, we propose a novel Counterfactual Residual Data Augmentation (CRDA) technique for tabular re…
- Our key insight is that once a regressor has modeled the systematic component of the data, the remaining noise can be viewed as an invariant residual that remai…
Singular Learning and Occam’s Razor in Deep Monomial Networks
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28464v1 公告类型:新。
- 摘要:在神经网络的优化中,梯度动态受到模型架构产生的临界点的影响。
- 这些临界点出现在模型参数化的雅可比行列式缺乏秩的地方,并且是奇异学习理论中研究的最明显的奇异点。
- 我们通过多项式代数工具(例如梅森定理)研究具有单项式激活的深度全连接网络中的此类点。
- EN 要点:
- arXiv:2606.28464v1 Announce Type: new
- Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model’s architecture
- These critical points occur where the Jacobian of the model’s parametrization is rank-deficient, and are the most pronounced singularities studied in Singular L…
- We investigate such points in deep fully-connected networks with monomial activations via tools from polynomial algebra such as Mason’s Theorem
An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28467v1 公告类型:新。
- 摘要:办公楼中的设备级能源监控会产生噪音警报,非专业设施管理人员很难使用。
- 本文提出了一种端到端的代理管道,它结合了深度时间序列预测、变分异常检测和基于 LLM 的推理,以生成优先的、可操作的维护建议。
- 该系统使用混合奇异谱分析 (SSA) 和长短期记忆 (LSTM) 预测模型来跟踪七种办公设备,并应用每个设备的 LSTM 变分自动编码器 (VAE),重点关注标记异常的日常消费事件。
- EN 要点:
- arXiv:2606.28467v1 Announce Type: new
- Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use
- This paper proposes an end-to-end agentic pipeline that combines deep time-series forecasting, variational anomaly detection, and LLM-based reasoning to generat…
- The system tracks seven office appliances using a hybrid Singular Spectrum Analysis (SSA) and Long Short-Term Memory (LSTM) forecasting model, and applies a per…
Modelling Emotional Memory in Children with Tensor Networks
- 发布时间:2026-06-30 12:00 北京时间
- 摘要:- arXiv:2606.28470v1 公告类型:新。
- 摘要:我们展示了情感效价如何影响儿童认知记忆的顺序依赖结构:对一系列情感效价玩具的正确回忆不仅取决于给定玩具本身的效价,还取决于它之前和之后展示的玩具的效价。
- 虽然标准心理模型确认顺序依赖性在事件(按顺序显示的一组玩具)中有所不同,但准确性较低,并且该模型无法反映对情感对象的记忆如何影响该组中的其他对象。
- 考虑化合价的经典张量网络模型在对研究结果进行建模时能够达到 77.98% 的准确度。
- EN 要点:
- arXiv:2606.28470v1 Announce Type: new
- Abstract: We demonstrate how emotional valence influences the order-dependent structure of children’s recognition memory: correct recall of a sequence of emotio…
- Whilst standard psychological models confirm that order-dependence differs across an event (a set of toys shown in sequence), accuracy is low and the model does…
- A classical tensor network model factoring in valence is able to achieve a 77.98% accuracy in modelling the results of the study