🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-21
- 类型
- ai-daily
- 字数
- 3269
- 阅读时长
- 16 min
2026-07-21 AI日更 | 长程智能体进入实测期:可控性、医疗化与审计链成新门槛 链接到标题
今天的重点从模型能力转向长程智能体的可控落地。OpenAI 提醒长期自主模型会暴露短评测难以捕捉的失控行为,工程上需要更强验证、回滚与工具边界。医疗智能体继续专业化,可审计因果推理也在回潮。
📖 本期 Watch List 深度导读 链接到标题
今天最值得关注的是“长程智能体的可控性”:关于 long-horizon models 的安全文章提醒,模型越能持续自主工作,越可能暴露短评测捕捉不到的失控行为;配合数学多智能体审稿、ARC-AGI-3 代码智能体和本地语音助手 AnovaX 的论文,很适合工程团队重新审视验证、回滚与工具边界。
第二条主线是医疗智能体加速专业化。Cura 1T、GraphDx 与临床多模态预测研究,都在尝试把诊断推理、成本约束、EHR 工具使用和多模态病历统一起来,值得医疗 AI 团队重点跟进。
此外,可解释与可审计推理正在回潮:Causal-Audit、Prolog 强化学习解释、可信 AI 工具分析,以及模型“全局工作空间”研究,共同指向一个趋势——下一阶段不只比答案,更要比推理链是否可检查、可复现。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Andrew Ng Course Revives Graphs vs Loops Debate in Agentic AI 链接到标题
- 分类:AI · News
- 概况:热度时间:,相关帖子数:334
- 是什么事:Andrew Ng 的新课程引发了关于智能体 AI 应采用“图结构工作流”还是“循环式自主推理”的讨论。
- 为什么重要:这关系到智能体系统的可控性、可解释性、可靠性与工程落地方式,是当前 AI 应用从演示走向生产的重要架构选择。
- 讨论概况:X 上的讨论集中在图结构是否更适合稳定编排复杂任务,还是循环式智能体更接近自主决策;支持者强调图可调试、可监控,反对者认为过度流程化会限制智能体灵活性。
话题 2:Anthropic’s Claude Fable Disproves Jacobian Conjecture in 3D 链接到标题
- 分类:AI · News
- 概况:热度时间:20 hours ago,相关帖子数:35000
- 是什么事:X 上热议称 Anthropic 的 Claude Fable 在三维情形下给出了对雅可比猜想的反例或否定证明,但该说法仍待权威验证。
- 为什么重要:如果属实,这将是 AI 在高难度数学发现中的重大突破,可能改变人们对大模型推理、自动定理证明和科研辅助能力的评估。
- 讨论概况:讨论焦点集中在证明是否真实可靠、是否存在模型幻觉或推导漏洞、是否经过数学界同行审查;支持者认为这是 AI 科研能力的里程碑,怀疑者则强调需公开完整证明并由专家复核。
话题 3:Moonshot AI’s Kimi K3 Tops Coding Benchmarks at Lower Cost 链接到标题
- 分类:AI · News
- 概况:热度时间:5 hours ago,相关帖子数:1700
- 是什么事:Moonshot AI 的 Kimi K3 被称在前端代码与软件工程相关基准中取得领先或接近顶级闭源模型表现,并以更低价格提供服务,开放权重预计即将发布。
- 为什么重要:这显示中国开放权重模型正在缩小与闭源前沿模型的能力差距,尤其在编码等高价值场景中可能加剧价格竞争,并推动企业更多考虑可本地部署、成本更低的 AI 方案。
- 讨论概况:X 上讨论焦点集中在 Kimi K3 是否真正达到或超越 Claude 等闭源模型、基准测试能否代表真实开发能力、低价策略对闭源模型商业模式的冲击,以及开放权重带来的创新机会与安全治理风险。
话题 4:Debate Heats Over China’s Kimi AI Models and U.S. Response 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:21000
- 是什么事:围绕中国月之暗面 Kimi AI 模型能力提升及其对美国 AI 竞争格局的影响,X 平台上出现大量讨论。
- 为什么重要:Kimi 等中国大模型被视为在长上下文、推理和低成本部署方面快速追赶,可能加剧中美 AI 技术、资本与政策竞争。
- 讨论概况:讨论焦点集中在 Kimi 的真实技术水平是否被高估、美国是否需要更强产业政策或出口管制,以及开源、算力限制和中国 AI 创新速度对全球市场的影响。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线集中在 AI 从“能力展示”走向“工程化、科研化和产业竞争”的关键转折:一方面,智能体架构讨论显示业界普遍认同可控性、可调试性和可靠性已成为落地核心;另一方面,Kimi K3 等中国模型的表现与低价策略,让更多人相信开放权重和成本优势正在重塑大模型竞争格局。共识在于,AI 正快速进入代码、科研、自动化工作流等高价值场景,且开源或开放权重模型正在缩小与闭源前沿模型的差距。分歧则主要体现在三点:智能体应偏向图结构编排还是循环式自主推理,Claude Fable 所谓数学突破究竟是真发现还是幻觉,Kimi 的基准成绩能否代表真实工程能力。潜在风险包括过度依赖未经验证的 AI 推理结果、基准炒作掩盖实际缺陷、低价与开源竞争带来的安全治理压力,以及中美 AI 竞争进一步被技术民族主义和政策管制放大。
💡 大佬观点(Influencer Insights) 链接到标题
这是基于提供的推文内容生成的 AI 行业日报分析。
AI 行业日报:大模型“军备竞赛”白热化,工具链重塑与世界观构建 链接到标题
1. 共同关注的技术趋势与产品热点 链接到标题
🔥 大模型性能“近身肉搏”:Kimi K3、Qwen3.8 与 Claude Fable 5 的三国杀 链接到标题
本日最大的热点无疑是Kimi K3与Qwen3.8-Max-Preview的接连冲击,形成了对目前公认最强的闭源模型 Claude Fable 5 的合围之势。
- Kimi K3:史上最大开源模型的“前端”奇袭。该模型以 2.8T 参数吸引了全部眼球,其在 前端设计与游戏生成 方面的能力被多位大佬视为“长板极长”。@Pluvio9yte 和 @vista8 等博主均指出 K3 在前端审美和可玩性 DEMO 生成上甚至优于或持平 Fable 5。@ruanyf 也分析认为,其性能接近 Fable 5 的核心原因在于参数量的暴力提升,但同时提醒它的 API 费用(输入/输出 20/100 元每百万 Token)已是国内最贵之列。
- Qwen3.8-Max-Preview:全栈工程的“无短板”挑战者。在 Kimi K3 发布仅三天后,阿里就祭出了 2.4T 参数的同级选手。@Pluvio9yte 引用泄露的测评数据指出,Qwen3.8 已超越 K3,在复杂工程任务、多 Agent 编排上表现稳健,整体与 Claude Opus 4.8 打平,仅落后于 Fable 5。这印证了 @dotey 转述的行业观点:模型发布不仅要看短板,更要看长板能否“破局”。
- Fable 5 的“王座”保卫战。面对围剿,Anthropic 的策略是通过商业手段维系地位。@zhixianio 和 @Pluvio9yte 都提到,Fable 5 不仅延长了付费用户的访问期限,更是在特定疑难杂症处理上保有不可替代性。@dotey 分享了一个典型案例:同样遭遇 VBR MP3 时间戳偏差问题,其他模型无解,而 Fable 5 能精准定位并修复,这被定义为极端场景下的壁垒。
🛠️ AI 编程工具进入“模型混血”与“Skill 变现”时代 链接到标题
单纯的终端编程工具已无法满足需求,整合不同模型的优势成为新趋势。
- 多模型路由与整合。@vista8 介绍了 OpenCodex 项目,允许在 Codex 界面内随时切换 Kimi K3(前端)、GPT 5.6 Sol(后端)、Grok 4.5(搜索),打破了平台锁定的壁垒。他开发的 Skill 更是实现了在 Codex 内一句话调度多种本地 CLI 模型,实现了“集各家所长”的合规使用。
- Skill 生态爆发与社区化。@vista8 开发的自动剪辑技能,以及 @ruanyf 发现的 小红书 REDSkill 社区,标志着 Skill(技能插件)正在从辅助脚本走向产品核心,甚至成为社媒传播和流量获取的新载体。
🌐 “世界模型”初现端倪:从生成视频到生成可交互世界 链接到标题
在 Sora 等模型停留在生成视频时,另一条赛道开始探索更深层次的理解。@Pluvio9yte 深度分析了开源的 Alaya World 项目,认为这代表了 AI 从“工具”进化为“环境”的趋势。它能根据指令实时生成可自由移动、即时交互的流媒体场景,具备了对空间、时间和因果关系的初步理解,而非简单的下一帧预测。这被看作是游戏、具身智能等领域的潜在底层变革。
2. 值得注意的独特观点与行业前瞻 链接到标题
- AI 成本的“阶级分化”不可逆 (@Pluvio9yte): 指出随着 Fable 5 等顶尖模型价格不断提高,未来模型订阅费将持续上涨,使用最顶尖生产力的成本将变得高不可攀,进一步拉大生产力差距的鸿沟。@zhixianio 在早先也因硬件涨价提前感叹。
- FDE(前端工程师)是模型公司的“阳谋” (@dotey): 尖锐指出 AI 公司借企业落地业务之机,通过 FDE 将企业的行业知识和最佳实践沉淀为 Skill 并内化到模型中。这对于个人是短期内的技术护城河,但对企业长远来看可能是降本增效后的人员“优化”前奏。
- AI 代码提交流量暴增的警示 (@ruanyf): 引用数据显示 GitHub 代码提交量同比增长 14 倍。这不仅解释了平台频繁故障的原因,更发出预言:如果托管成本持续爆炸,GitHub 全面步入付费制或许已不远矣。
- 顶尖模型在“极端场景”下依然难以被替代 (@dotey): 强调大众化场景下各模型差距缩小,但遇到 VBR 音视频编码识别、特定疑难 Bug 修复等极端逻辑问题时,Fable 5 仍然具有“杀手级”的可靠性。
- “开源”是伪命题吗? (@ruanyf 转引 Anthropic CEO): 提出当前所谓的“开源”AI 模型仅开放权重,外界无法窥视内部运作逻辑,本质上区别于传统的开源软件,应称之为“开放权重”更准确。
- 产品设计的新思路:从“游戏化”到“游戏感” (@nishuang): 在设计 AI 应用(如背单词 App)时,应当摒弃多巴胺驱动的奖励系统(游戏化/Gamification),转而通过激发好奇心和内啡肽,创造让用户感觉在“玩”而非“干苦力”的游戏感(Game-like design)。
3. 推荐的工具与资源 链接到标题
- OpenCodex: 解锁 Codex 的模型束缚。允许你在 Codex 的交互界面中直接调用 Kimi K3、Grok 4.5 等多个外部模型,实现前端用 K3、后端用 Sol 的高效组合模式。来源 @vista8。
- Grok Build: 马斯克的 SpaceX AI 团队开源的终端 AI 编程 Agent。纯 Rust 编写,拥有 MCP 支持、沙箱模式、无头模式等高可玩性配置,已有 1.4 万 Star。适合喜欢折腾定制化 CLI 工具的开发者。来源 @AI_Jasonyu。
- MOSS-Transcribe-Diarize-0.9B: 阿里发布的轻量级开源语音转写模型,可一次性处理最长约 90 分钟音频,并直接输出带说话人标注的带时间戳文本。@dotey 实测评价转录效果精准,适合播客、会议纪要等长音频的本地化处理。
- BaoCut Skill / Qiaomu-Cut Skill: 如果你在做视频号或影视剪辑号,@dotey 和 @vista8 开发或集成的视频自动剪辑技能可以让你通过文字指令,自动完成从素材检索到拼接合成的全流程,甚至生成带字幕的英语学习切片视频。
- Alaya World: 想提前体验“世界观”?可以在其 GitHub 仓库找到开源推理代码与权重,尝试把文字或图片直接变成可自由行走交互的世界。来源 @Pluvio9yte。
- Skill 变现配套基础设施: @ruanyf 和 @vista8 分别提到了小红书 REDSkill 社区与视频号下载工具,这些是 Skill 分发和获客的绝佳基础设施。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 32 条更新
Stratechery by Ben Thompson (A_full) 链接到标题
- Who’s Afraid of Chinese Models?
- 发布时间:2026-07-20 19:00 北京时间
- 摘要:- 听这个帖子**:**。
- 我讲了一个关于我在凯洛格管理学院参加 STRT-431 第一天的故事,这是每个 MBA 一年级学生都必须参加的入门课程;我翻阅了阅读材料和案例研究,但令人沮丧的是,名单上没有任何科技公司。
- 就我而言,我在课后与教授交谈,想知道为什么,并被告知课程的目标不一定是了解特定行业,而是发现广泛适用的通用原则,可以应用于任何行业的任何公司。
- 正如我通常讲述的那样,我并不觉得这非常令人满意:对我来说,技术的本质,特别是软件和发行的边际成本为零(以及零交易成本)的事实,是根本不同的;在公式中输入零往往会造成严重破坏!
- 然而,我很快意识到,这是我的机会。
- EN 要点:
- Listen to this post :
- Log in to listen
- There’s a story I tell about my first day in STRT-431 at Kellogg School of Management, the introductory class that every first-year MBA was required to take; I…
- Me being me, I spoke to the professor after class wondering why, and was told that the goal of the course was not to necessarily learn about specific industries…
OpenAI Blog (A_full) 链接到标题
- Safety and alignment in an era of long-horizon models
- 发布时间:2026-07-20 18:00 北京时间
- 摘要:- 可以长时间自主工作的模型可以解决困难的、开放式的问题。
- 但同样的坚持让它们变得有用,也让它们有更多机会采取不需要的行动,而且这样做的方式是针对短期模型的评估可能会错过的。
- 该模型设计用于长时间自主工作。
- 在有限、受监控的内部使用过程中,我们观察到了现有部署评估未捕获的不良行为。
- 由于部署受到限制和监控,我们能够识别这些问题,暂停访问,根据我们观察到的情况创建新的评估,加强模型及其保障措施,然后在持续监控下恢复访问。
- EN 要点:
- OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deploym…
ArXiv cs.AI (B_intro+search) 链接到标题
GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15280v1 公告类型:新。
- 摘要:顺序诊断需要通过迭代信息收集来平衡诊断准确性和资源成本。
- 现有的大型语言模型(LLM)方法表现出严重的知识推理差距:尽管编码了广泛的医学知识,但它们在成本限制下难以系统地推理,常常诉诸过度测试。
- 我们提出 GraphDx,一个具有两项核心创新的知识增强框架。
- EN 要点:
- arXiv:2607.15280v1 Announce Type: new
- Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering
- Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical knowledge, they struggle to reason…
- We propose GraphDx, a knowledge-enhanced framework with two core innovations
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15281v1 公告类型:新。
- 摘要:基于因果和干预的问答对于推进大型语言模型 (LLM) 的推理超越表面相关性和理解潜在的因果机制至关重要。
- 然而,现有的基于LLM的方法通常依赖于隐式语言级推理,导致不透明的因果假设、无法验证的推理路径以及复杂干预下的脆弱预测,特别是在上下文无关的环境中。
- 在本文中,我们提出了一个明确且可审计的因果推理框架,用于上下文无关的基于干预的问答。
- EN 要点:
- arXiv:2607.15281v1 Announce Type: new
- Abstract: Causal and intervention-based question answering is fundamental to advancing large language models (LLMs) toward reasoning beyond surface-level correl…
- However, existing LLM-based methods often rely on implicit language-level reasoning, resulting in opaque causal assumptions, unverifiable reasoning paths, and f…
- In this paper, we propose an explicit and auditable causal reasoning framework for context-free intervention-based question answering
Cura 1T: Specialized Model for Agentic Healthcare
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15314v1 公告类型:新。
- 摘要:医疗保健涵盖高风险的沟通、专家推理和工作流程执行,但涵盖这些用例的专业法学硕士仍然有限。
- 医疗保健模型必须处理患者咨询、文本和图像临床推理、交互式诊断以及电子健康记录 (EHR) 工具的使用。
- 这些功能会以不同的方式失败,一项任务的狭窄更新可能会降低另一项任务的性能。
- EN 要点:
- arXiv:2607.15314v1 Announce Type: new
- Abstract: Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain…
- A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use
- These capabilities fail in different ways, and a narrow update for one task can degrade another
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15367v1 公告类型:新。
- 摘要:桌面语音助手仍然以云管道为主,这些管道将原始音频从机器上传输出来并公开一组固定的技能。
- 我们描述了 AnovaX,这是一个小型的本地优先助手,完全在用户的计算机上运行,并将桌面本身视为其操作界面。
- 单个 Python 进程将唤醒词门、语音管道、发出工具调用 JSON 计划的 LLM 规划器 (Gemini)、白名单和拒绝名单安全层、将每个计划转换为有界线程池上的类型化子代理的多代理协调器,以及在核心步骤失败时接管的自适应恢复循环。
- EN 要点:
- arXiv:2607.15367v1 Announce Type: new
- Abstract: Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills
- We describe AnovaX, a small local-first assistant that runs entirely on the user’s computer and treats the desktop itself as its action surface
- A single Python process wires together a wake-word gate, a speech pipeline, an LLM planner (Gemini) that emits a JSON plan of tool calls, a whitelist-and-denyli…
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15388v1 公告类型:新。
-摘要:许多面向数学和科学的代理系统使用具有专门审阅者角色的分层设计,假设专门的审阅阶段应该有助于将错误的候选人变成正确的候选人。
- 我们使用匹配的 gpt-oss-120b actor 在 4,181 个基于验证者的 Omni-MATH 问题上测试了这一假设。
- 在最简单的级别上,协作几乎没有增加什么,但从第 4 级开始,收益急剧增加;在这种更困难的制度中,广播式同行讨论比计划者-执行者-审阅者管道(PER)达到更高的最终准确性。
- EN 要点:
- arXiv:2607.15388v1 Announce Type: new
- Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should…
- We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors
- Collaboration adds little on the easiest tiers, but from tier 4 onward the gains open sharply; in this harder regime, broadcast-style peer discussion reaches hi…
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15418v1 公告类型:新。
- 摘要:我们介绍 DrawingVQA,这是第一个旨在评估真实施工图(建筑、土木和许多其他工程实践的核心媒体)上的多模式大语言模型 (MLLM) 的基准。
- 与自然图像或示意性平面图不同,施工图融合了抽象几何、符号符号、表格数据、注释和特定领域的文本,形成了工程工作流程的独特复杂的视觉文本领域核心。
- DrawingVQA 通过 33 个“为施工而发布”的图纸和 92 个专业策划的问答对弥补了这一差距,涵盖三个推理深度:感知理解、上下文解释和领域专家推理。
- EN 要点:
- arXiv:2607.15418v1 Announce Type: new
- Abstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings – a co…
- Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific…
- DrawingVQA bridges this gap with 33 “Issued for Construction” drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptua…
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15439v1 公告类型:新。
- 摘要:我们之前的 ARC-AGI-3 代理捆绑了可执行的世界建模、预定的简化和精确的重放验证,但不清楚哪个想法决定了其性能。
- 我们用四个嵌套的基于 Codex 的代理来解决这个归因问题:文本基线;无需重放验证的灵活接口可执行世界模型;具有预定简化的相同可执行模型;以及固定界面验证处理,保留简化并要求精确再现记录的观察结果。
- 主要研究评估了所有四种具有 gpt-5.4 和 gpt-5.5 的智能体在公共 ARC-AGI-3 游戏中的高和高推理能力。
- EN 要点:
- arXiv:2607.15439v1 Announce Type: new
- Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear which idea ac…
- We address this attribution question with four nested Codex-based agents: a textual baseline; a flexible-interface executable world model without replay verific…
- The main study evaluates all four agents with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on the public ARC-AGI-3 games
Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15442v1 公告类型:新。
- 摘要:互联网迷因将视觉线索、文本内容和文化背景交织在一起,使得它们在幽默、讽刺和有害意图共存的场景中特别难以解释。
- 这些复杂性凸显了对可解释的模因理解系统的需求,该系统可以提供可靠和结构化的推理,以支持准确的分类和人类可解释性。
- 然而,现有的多模态分类器要么忽略了这些相互依赖性,要么只提供有限的可解释性。
- EN 要点:
- arXiv:2607.15442v1 Announce Type: new
- Abstract: Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenarios where hum…
- These complexities highlight the need for explainable meme understanding systems that can provide reliable and structured reasoning to support both accurate cla…
- However, existing multimodal classifiers either overlook these interdependencies or provide only limited interpretability
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15459v1 公告类型:新。
-摘要:经过训练的深度强化学习策略是一个黑匣子,我们询问是否可以通过将其重写为可执行逻辑程序来解释它,该逻辑程序可以重现其行为,并且人可以阅读,逻辑引擎可以运行,优化器可以编辑。
- 我们提出了一个三阶段事后转换,提取冻结的近端策略优化教师,以经典关系学习的方式从其决策中归纳出有序规则列表,并将结果作为 Prolog 程序发出,其每个决策都由现成的逻辑引擎执行;随后的扩展阶段编辑规则库,并且仅当策略评估证明回报增加时才接受编辑。
- 回波损耗界限使蒸馏程序成为有限马尔可夫决策过程中机器可检查的证书,并且扩展循环单调改进并终止。
- EN 要点:
- arXiv:2607.15459v1 Announce Type: new
- Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic prog…
- We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in…
- We prove four guarantees
A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15480v1 公告类型:新。
- 摘要:随着人工智能 (AI) 系统对社会的影响越来越大,确保其道德和值得信赖的部署已成为全球优先事项。
- 尽管已经出现了大量的高水平道德准则,但批评仍然存在,认为这些框架仍然抽象且缺乏具体的实施机制。
- 本文利用 OECD 的综合数据集,对旨在实施可信赖人工智能 (TAI) 的工具和信任标记框架进行了批判性分析。
- EN 要点:
- arXiv:2607.15480v1 Announce Type: new
- Abstract: As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global priority
- While a myriad of high-level ethical guidelines have emerged, criticism persists that these frameworks remain abstract and lack concrete mechanisms for implemen…
- This paper conducts a critical analysis of tools and trust mark frameworks intended to operationalize trustworthy AI (TAI), drawing on a comprehensive dataset f…
ArXiv cs.CL (B_intro+search) 链接到标题
Large Language Models as Unified Multimodal Learners for Clinical Prediction
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15380v1 公告类型:新。
- 摘要:电子健康记录将自由文本临床叙述与生命体征、实验室值和合并症等结构化测量相结合。
- 然而,大多数临床预测系统仍然依赖于特定于任务的融合架构,将每种模式的专用编码器与必须针对每个新任务和临床环境重新设计的学习组合机制配对。
- 我们提出了一个更简单的替代方案:将所有患者数据(无论模态如何)转换为单个自然语言序列,并端到端地微调预训练的语言模型,无需进行融合的架构修改。
- EN 要点:
- arXiv:2607.15380v1 Announce Type: new
- Abstract: Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and comorbidities
- Yet most clinical prediction systems still rely on task-specific fusion architectures, pairing dedicated encoders for each modality with learned combination mec…
- We propose a simpler alternative: convert all patient data, regardless of modality, into a single natural language sequence and fine-tune a pretrained language…
Verbalizable Representations Form a Global Workspace in Language Models
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15495v1 公告类型:新。
- 摘要:在人类大脑处理的所有内容中,只有一小部分是可以有意识地访问的,即可用于口头报告、有意控制和灵活推理。
- 在本文中,我们提供的证据表明,大型语言模型中已经出现了类似的功能区别。
- 使用一种新的可解释性技术,即雅可比透镜,我们可以识别模型在其处理过程中的任何时刻准备表达的表示。
- EN 要点:
- arXiv:2607.15495v1 Announce Type: new
- Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, delib…
- In this paper, we present evidence that an analogous functional distinction has emerged in large language models
- Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing
VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15498v1 公告类型:新。
- 摘要:键值(KV)缓存是长上下文大语言模型(LLM)推理中的主要内存瓶颈。
- 两个领先的免训练家族在结构上都受到限制:标记选择方法(SnapKV、Ada-KV)从观察窗口对重要性进行评分并驱逐低分标记,但驱逐是不可逆的——因此,当重要性信号在与查询无关的重用下下降时,准确性会下降 11-15 点;统一的低等级编码保留每个令牌,但在任何地方都花费相同的等级,浪费预算。
- 我们观察到,这两种失败都有一个解决办法:应该分配军衔,而不是驱逐军衔。
- EN 要点:
- arXiv:2607.15498v1 Announce Type: new
- Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference
- Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict…
- We observe that both failures share one cure: rank should be allocated, not evicted
EpiNarrate: Agentic Generation of Grounded Narratives from Epidemiological Scenario Projections
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15544v1 公告类型:新。
- 摘要:生成清晰易懂的公共卫生叙述对于向政策制定者和广大公众传达复杂的流行病学预测至关重要。
- 此类叙述需要的不仅仅是简单地报告数字:预测必须结合具体情况,并在多个维度上进行定量分析。
- 此外,预测通常源自大型集合数据集,其中结合了干预假设、地理和人口阶层、结果、时间范围和不确定性分位数。
- EN 要点:
- arXiv:2607.15544v1 Announce Type: new
- Abstract: Generation of clear and accessible public health narratives is critical for communicating complex epidemiological projections to policymakers and the…
- Such narratives require more than simply reporting numbers: projections must be contextualized and quantitatively grounded across multiple dimensions
- Further, projections are often derived from large ensemble datasets which combine intervention assumptions, geographic and demographic strata, outcomes, time ho…
SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15557v1 公告类型:新。
- 摘要:代理技能(SKILL.md 文件)为 LLM 代理打包可重用的程序知识,是扩展代理功能的流行机制。
- 公共存储库现在托管着大量且数量不断增加的工件,但这些工件支离破碎、冗余且质量参差不齐,其实践价值尚不明确。
- 一个核心问题仍然悬而未决,即如何将这个开源 SKILL.md 生态系统整合到一个可用的语料库中,以及它对现实世界代理任务的好处如何限制。
- EN 要点:
- arXiv:2607.15557v1 Announce Type: new
- Abstract: Agent skills, SKILL.md files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities
- Public repositories now host them in large and growing numbers, yet these artifacts are fragmented, redundant, and uneven in quality, and their value in practic…
- A core question remains open, namely how to consolidate this open-source SKILL.md ecosystem into a single usable corpus, and what bounds its benefit on real-wor…
Process Reward Informed Tree Rollout for Effective Multi-Turn RL
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15610v1 公告类型:新。
- 摘要:强化学习(RL)已成为训练 LLM 智能体的关键方法,但 GRPO/RLOO 等流行方法依赖于多个独立采样的完整轨迹来进行优势估计。
- 在长期代理任务中,这种统一的推出策略可能会将预算浪费在无信息的死胡同尝试上,同时有希望的中间状态没有得到足够的探索。
- 代理轨迹的多回合结构,具有交错的动作和观察,自然支持将轨迹组组织为树,其中每个回合作为探索的决策点。
- EN 要点:
- arXiv:2607.15610v1 Announce Type: new
- Abstract: Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple independently sa…
- In long-horizon agentic tasks, such a uniform rollout strategy can waste budget on uninformative dead-end attempts, while promising intermediate states do not r…
- The multi-turn structure of agentic trajectories, with interleaved actions and observations, naturally supports organizing a trajectory group as a tree, where e…
On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15648v1 公告类型:新。
- 摘要:在对话系统和多个用户之间的多方对话中,识别话语的对象是一个关键挑战。
- 之前的工作通常将收件人检测视为多类分类任务,选择代表单个参与者或群体的单个标签。
- 该公式假设地址本质上是离散的,并且主要用于预测轮流。
- EN 要点:
- arXiv:2607.15648v1 Announce Type: new
- Abstract: In multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge
- Prior work has typically treated addressee detection as a multi-class classification task, selecting a single label representing an individual participant or th…
- This formulation assumes that address is inherently discrete and has primarily been used for predicting turn-taking
Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15655v1 公告类型:新。
- 摘要:掩码扩散语言模型 (DLM) 通过迭代地细化掩码标记来实现并行文本生成,为自回归解码提供了一种有前途的替代方案。
- 最近基于前瞻的解码方法通过在提交令牌更新之前探索未来的解码状态来提高准确性-效率权衡。
- 然而,现有方法主要依赖于浅层一步前瞻,它优化了即时信息增益,但对于较长范围的解码轨迹可能不是最佳的。
- EN 要点:
- arXiv:2607.15655v1 Announce Type: new
- Abstract: Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to aut…
- Recent lookahead-based decoding methods improve the accuracy–efficiency trade-off by exploring future decoding states before committing token updates
- However, existing approaches mainly rely on shallow one-step lookahead, which optimizes immediate information gain but can be suboptimal for longer-horizon deco…
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15736v1 公告类型:新。
- 摘要:大型推理模型通常通过长链思想(CoT)轨迹来解决问题,但大部分计算都花费在冗余推导、重复自我验证和绕道上,而这些并不能改善最终答案。
- 现有的策略自蒸馏方法通过将学生模型与从学生自己的推出中采样的前缀上的自身的简洁副本进行匹配来降低此成本。
- 我们证明这个目标有一个初始化瓶颈。
- EN 要点:
- arXiv:2607.15736v1 Announce Type: new
- Abstract: Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations…
- Existing on-policy self-distillation methods reduce this cost by matching a student model to a concise copy of itself on prefixes sampled from the student’s own…
- We show that this objective has an initialization bottleneck
Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15766v1 公告类型:新。
-摘要:大型语言模型(LLM)擅长回答预先指定的问题,但它们在开放式、结论前的发现阶段的导航能力在很大程度上仍然无法衡量。
- 我们引入了前瞻性假设发现(PHD),它要求模型根据非结论性证据(包括异常观察和碎片记录)自主构建有根据的、有区别的和可检验的假设空间,以指导后续调查。
- 为了评估这种能力,我们引入了 HypoArena,其中包括 HypoData(跨越六个科学和分析领域的 988 个案例的基准)和 HypoEval(开放式假设集的评估框架)。
- EN 要点:
- arXiv:2607.15766v1 Announce Type: new
- Abstract: Large language models (LLMs) excel at answering pre-specified questions, yet their ability to navigate the open-ended, pre-conclusion stage of discove…
- We introduce Prospective Hypothesis Discovery (PHD), which asks models to autonomously construct grounded, discriminative, and testable hypothesis spaces from i…
- To evaluate this capability, we introduce HypoArena, comprising HypoData, a benchmark of 988 cases across six scientific and analytical domains, and HypoEval, a…
ArXiv cs.LG (B_intro+search) 链接到标题
Structure of the Circular-Dyadic Convolution Error
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15293v1 公告类型:新。
- 摘要:二进卷积和循环卷积都可以分别使用 Hadamard 变换和 FFT 计算的离散傅里叶变换 (DFT) 在 $O(N\log N)$ 时间内计算。
- Hadamard 变换因其实值符号翻转而更可取,但它对 DFT 的替代会引入代数误差。
- 我们提出了三个互补的结果来表征该错误。
- EN 要点:
- arXiv:2607.15293v1 Announce Type: new
- Abstract: Dyadic and circular convolution can both be computed in $O(N\log N)$ time using the Hadamard transform and the FFT-computed discrete Fourier transform…
- The Hadamard transform is preferable for its real-valued sign flips, yet its substitution for the DFT introduces algebraic error
- We present three complementary results that characterize this error
Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15313v1 公告类型:新。
- 摘要:缩放假设假设增加模型参数会产生涌现的推理能力。
- 本立场文件认为,将这种概率范式应用于通用量子电路综合是一个方向性错误。
- 与自然语言不同,量子电路需要严格遵守数学约束,这表明存在显着的语法语义差距。
- EN 要点:
- arXiv:2607.15313v1 Announce Type: new
- Abstract: The scaling hypothesis assumes that increasing model parameters yields emergent reasoning capabilities
- This position paper argues that applying this probabilistic paradigm to generic quantum circuit synthesis is a directional error
- Unlike natural languages, quantum circuits require strict adherence to mathematical constraints that manifest a significant syntax-semantics gap
A Transportable Threshold-Based Framework for Interpretable Classification of Medical Data
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15394v1 公告类型:新。
-摘要:黑盒模型由于缺乏可解释性和可重复性而限制了人工智能在医学中的采用。
- 我们引入了一个基于统计的框架,该框架使用伯努利朴素贝叶斯 (BNB) 模型提供完全可解释的、基于规则的临床分类。
- 该方法将监督的 $\chi^2$ 引导的统计二值化应用于连续变量,识别在训练数据中最大化与临床结果关联的阈值。
- EN 要点:
- arXiv:2607.15394v1 Announce Type: new
- Abstract: Black-box models limit the adoption of artificial intelligence in medicine due to their lack of interpretability and reproducibility
- We introduce a statistically grounded framework that provides fully interpretable, rule-based clinical classification using the Bernoulli Na"ive Bayes (BNB) mo…
- The method applies supervised $\chi^2$-guided statistical binarization to continuous variables, identifying thresholds that maximize association with clinical o…
Regularity-Aware Stochastic MGDA with Adaptive Conflict-Avoidant Update Direction Control
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15412v1 公告类型:新。
- 摘要:多目标学习(MOL)旨在同时优化多个目标。
- 多梯度下降算法 (MGDA) 是一种主力,可以沿着跨目标的共同下降或冲突避免 (CA) 方向迭代更新。
- 然而,在随机设置中,普通随机 MGDA 方法 (SMG) 缺乏快速收敛速度,因为小批量采样会在梯度中引入噪声。
- EN 要点:
- arXiv:2607.15412v1 Announce Type: new
- Abstract: Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously
- The multi-gradient descent algorithm (MGDA) is a workhorse that iteratively updates along a common descent or conflict-avoidant (CA) direction across objectives
- In stochastic settings, however, the vanilla stochastic MGDA method, SMG, lacks a fast convergence rate because mini-batch sampling introduces noise in the grad…
AI Trading: Evaluating Large Language Models for Technical Market Analysis
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15414v1 公告类型:新。
- 摘要:大型语言模型 (LLM) 已成为处理现代金融市场异构信息环境的强大工具。
- 本文对五个著名的法学硕士:GPT-4 Turbo、Claude 3 Opus、Gemini 1.5 Pro、Llama 3 70B 和专业领域的 FinGPT 的技术市场分析能力进行了系统的比较评估。
- 评估涵盖四个结构化任务:从 OHLCV 数据中识别烛台模式、定向信号生成(买入/卖出/持有)、通过模拟执行管道对信号质量进行回溯测试以及财务报告理解。
- EN 要点:
- arXiv:2607.15414v1 Announce Type: new
- Abstract: Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial markets
- This paper presents a systematic, comparative evaluation of five prominent LLMs: GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, and the domain-special…
- The evaluation spans four structured tasks: candlestick pattern recognition from OHLCV data, directional signal generation (BUY/SELL/HOLD), backtesting of signa…
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15421v1 公告类型:新。
- 摘要:紧凑的医学图像分类器需要效率和可解释的证据,但这些目标通常是单独解决的。
- 我们引入了 qZACH-ViT,它是零令牌(CLS 无令牌)、无位置 ZACH-ViT 主干的量化感知扩展,具有递归内在补丁级类证据。
- 我们还引入了递归归因稳定优化(RASO),它对分类和归因梯度进行范数匹配,并删除与分类冲突的归因成分。
- EN 要点:
- arXiv:2607.15421v1 Announce Type: new
- Abstract: Compact medical-image classifiers need efficiency and interpretable evidence, yet these goals are often addressed separately
- We introduce qZACH-ViT, a quantization-aware extension of the zero-token (CLS-token-free), position-free ZACH-ViT backbone with recursive intrinsic patch-level…
- We also introduce Recursive Attribution-Stabilized Optimization (RASO), which norm-matches classification and attribution gradients and removes attribution comp…
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15433v1 公告类型:新。
- 摘要:我们对标准线性模型与用于监督二元分类任务的单量子位混合状态模型的固有可解释性进行了表征和比较。
- 并排比较表明,用于二元分类的单量子位混合状态模型只是标准线性模型分类的“椭圆体版本”。
- 更准确地说,我们不是学习超平面来对数据进行分类,而是学习超椭球体。
- EN 要点:
- arXiv:2607.15433v1 Announce Type: new
- Abstract: We characterize and compare the inherent interpretability offerings of a standard linear model with a single qubit mixed state model for the task of s…
- A side by side comparison reveals that a single qubit mixed state model for binary classification is just the ``ellipsoid version" of standard linear model clas…
- More precisely, rather than learning a hyperplane to classify data, we learn a hyperellipsoid
Stochastic Reset Pathfinding: Path-Level Regret for Cascading Bandits over Graph Paths
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15440v1 公告类型:新。
- 摘要:我们介绍随机重置寻路(SRP),这是一个已知有向图上的情景学习问题,具有未知的固定边缘成功概率。
- 在每个情节中,代理都会提交到源到目标路径,并且执行期间的任何边缘故障都会将其重置为源。
- SRP 捕获量子中继器网络中的纠缠分布、闪电网络上的支付路由以及不可靠的网状网络中的交付等设置。
- EN 要点:
- arXiv:2607.15440v1 Announce Type: new
- Abstract: We introduce Stochastic Reset Pathfinding (SRP), an episodic learning problem on a known directed graph with unknown stationary edge success probabili…
- In each episode, the agent commits to a source-to-goal path, and any edge failure during execution resets it to the source
- SRP captures settings such as entanglement distribution in quantum repeater networks, payment routing on the Lightning Network, and delivery in unreliable mesh…
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15446v1 公告类型:新。
- 摘要:医疗保健成本在美国仍然是一个令人担忧的问题,并且可能受到与 COVID-19 大流行相关的干扰的影响。
- 本研究使用 2019 年和 2021 年的医疗支出面板调查 (MEPS) 数据考察了大流行前后的医疗保健财务脆弱性。
- 高经济负担的定义是自付费用医疗保健支出超过家庭收入的 10%。
- EN 要点:
- arXiv:2607.15446v1 Announce Type: new
- Abstract: The cost of healthcare remains a concern in the United States and may have been influenced by disruptions associated with the COVID-19 pandemic
- This study examines healthcare financial vulnerability before and after the pandemic using Medical Expenditure Panel Survey (MEPS) data from 2019 and 2021
- High financial burden was defined as out-of-pocket healthcare expenditures exceeding 10% of family income
LLM4EHR: Aligning Clinical Time Series with Medical Event Sequences via Large Language Models
- 发布时间:2026-07-20 12:00 北京时间
- 摘要:- arXiv:2607.15447v1 公告类型:新。
- 摘要:临床机器学习的最新研究重点是重症监护病房 (ICU) 的结果预测,已从定制监督模型转向基础模型,利用现代表示学习方法。
- 在这里,基础模型是根据复杂的临床数据模式的混合物进行预训练的,可用于各种下游任务。
- 现有的工作经常利用电子健康记录(EHR)来提供丰富多样的患者观察结果来训练临床基础模型。
- EN 要点:
- arXiv:2607.15447v1 Announce Type: new
- Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models…
- Here, foundation models are pre-trained on mixtures of complex clinical data modalities, useful for various downstream tasks
- Existing works often utilise Electronic Health Records (EHR) to provide rich and diverse patient observations to train clinical foundation models