🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-24
- 类型
- ai-daily
- 字数
- 3363
- 阅读时长
- 16 min
2026-06-24 AI日更 | GPT-5 破解免疫学难题,AI 评估与端侧模型同步升温 链接到标题
今天的核心信号来自两端:GPT-5 Pro 进入免疫学实验判断环节,显示 AI for Science 正从资料辅助走向科研推理;OpenAI 推动先进 AI 共享标准,行业治理与能力评估更趋制度化。与此同时,端侧模型、AI 编程工具和企业协作助手继续加速落地,成本、可控性与真实工作流价值成为竞争焦点。
📖 本期 Watch List 深度导读 链接到标题
今天最值得把注意力放在三条线上:首先,GPT-5 帮助免疫学家破解三年谜题,是“AI for Science”从辅助检索走向实验判断的典型案例,建议科研与医疗团队重点读。其次,OpenAI 推动先进 AI 共享标准,叠加 VeriBound、FirstPass 等论文,显示行业正在把模型能力评估、同行评审、过程奖励模型纳入更严肃的治理框架。第三,端侧 RAG 压缩、语言 steering、多智能体后训练差异、情感 TTS 等研究,提醒我们:模型竞争不只在参数规模,也在可控性、成本和交互细节。另可关注内存芯片与中国模型议题,AI 基础设施与地缘供应链仍是长期变量。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Z.ai’s GLM-5.2 Tops Open AI Leaderboards, Matches Closed Models at Lower Cost 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:11000
- 是什么事:Z.ai 发布的 GLM-5.2 在多个公开 AI 基准榜单上名列前茅,并被称在更低成本下接近或匹配部分闭源模型表现。
- 为什么重要:这显示开源或开放权重模型在能力与成本效率上继续逼近闭源前沿模型,可能加速企业采用、模型商品化以及全球 AI 竞争格局变化。
- 讨论概况:X 上讨论集中在 GLM-5.2 的榜单成绩是否能代表真实应用能力、低成本优势是否可持续,以及开放模型对 OpenAI、Anthropic 等闭源厂商的竞争压力;也有人质疑基准测试可被优化,呼吁更多独立评测。
话题 2:Anthropic Launches Claude Tag for Slack Workspaces 链接到标题
- 分类:AI · News
- 概况:热度时间:6 hours ago,相关帖子数:7100
- 是什么事:Anthropic 推出面向 Slack 工作区的 Claude Tag 功能,让团队可在 Slack 中直接调用 Claude 协助处理对话和工作流。
- 为什么重要:这显示 AI 助手正进一步嵌入企业协作场景,从独立聊天工具转向与日常办公平台深度集成,可能影响企业知识管理、自动化协作和 AI 助手采用率。
- 讨论概况:X 上的讨论主要集中在其能否提升团队效率、与 Slack 原生 AI 和其他办公助手的竞争关系,以及企业数据隐私、权限控制和误用风险等问题。
话题 3:OpenAI Engineer Draws Late-Night Codex Projects from Devs Worldwide 链接到标题
- 分类:AI · News
- 概况:热度时间:14 hours ago,相关帖子数:418
- 是什么事:一名 OpenAI 工程师在 X 上发起深夜征集,邀请全球开发者展示他们用 Codex 构建的项目,引发数百条回应。
- 为什么重要:这反映出 AI 编程工具正从演示阶段进入真实开发场景,开发者生态和实际用例成为衡量模型能力与产品价值的重要指标。
- 讨论概况:讨论焦点集中在 Codex 是否真正提升开发效率、哪些项目最能体现其能力,以及 AI 编程工具会成为开发者助手还是逐步替代部分基础编码工作。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线是 AI 正从模型能力竞赛转向真实场景落地:GLM-5.2 代表开放模型在性能与成本上继续追赶闭源前沿,Claude Tag 体现助手深入企业协作流,Codex 讨论则显示 AI 编程工具正在进入实际开发生态。共识在于,AI 的价值越来越取决于能否以更低成本嵌入工作流、提升生产效率,并形成可持续的开发者和企业使用案例。分歧主要集中在公开基准是否等同于真实能力、低成本优势能否长期维持,以及 AI 助手究竟是增强人类工作还是替代部分岗位。潜在风险则包括基准被优化导致预期失真、企业数据隐私与权限边界不清,以及过度依赖 AI 工具带来的误用、责任归属和安全隐患。
💡 大佬观点(Influencer Insights) 链接到标题
好的,基于过去24小时内的推文汇总,以下是 AI 行业动态分析报告。
1. 今日大佬们共同关注的技术趋势或产品热点 链接到标题
核心趋势:端侧模型(On-device Models)与 AI 编程工具的激烈竞争已进入白热化阶段,成本与效率成为核心衡量指标。
端侧模型能力成熟,迎来爆发前夜: 多位大佬对本地运行的小模型能力表示惊叹。
- @zhixianio 在本地测试了 MiniCPM-o 4.5 的音视频全双工效果,认为其质量令人满意,对于一个 9B 模型来说“很难想象”。他还分享了自己强迫使用本地模型(Qwen3.6-35B-A3B)进行编程和日常任务,结果“速度和智商在线,多模态体验甚至比远程模型更爽”,并深度评测了 Google 的 Gemma 4 系列(包括 12B 多模态和 E4B)在端侧的潜力。
- 他提出的“Model-Pak”(大模型卡带)概念也引发了对未来端侧模型分发形态的关注。
AI 编程工具成为“血海”战场,Codex 风头正劲:
- 新模型发布与对比:Google Gemma 4 12B Coder 的发布引发热议。@zhixianio 对其进行了深度对比评测,结论是:虽然在简单任务上与 35B MoE 模型打平,但在复杂、有状态的程序生成上,12B 体量存在明显能力天花板,他的“甜点”依然是 Qwen 35B。
- OpenAI Codex 生态与能力扩展:@vista8、@AI_Jasonyu 等博主密集讨论了 Codex 的新功能,如 Record & Replay(操作录制回放,被视为超级 RPA)、内嵌浏览器实现的无限制画布生图(Cowart 项目)、以及 Codex CLI 的更新。@dotey 还提到有开源项目让 ChatGPT 也能获得类似 Codex 的本地代码操作能力。
- 工具间的选择:@ruanyf 直接发起关于“Codex vs Claude Code”的投票,反映了开发者在二者间的摇摆。
AI 生成的推理成本成为显性焦点:
- @ruanyf 转发了 OpenAI 员工一个月消耗预估价值 130 万美元 Token 的数据,引发了对 AI 比真人程序员更昂贵的讨论。@vista8 和 @Pluvio9yte 也提到 Codex/Fable 等高级模型的消耗速度极快,促使社区寻找更经济的 API 或重置方案。
2. 值得注意的独特观点或行业前瞻 链接到标题
从“相关性”到“因果性”:AI 发展的下一个关键变量: @Pluvio9yte 深度评述了黄碧薇教授创办的 Aether AI 及其“因果大模型”理念。他指出当前大模型基于相关性预测的局限,例如无法理解“往有洞的杯子倒水会漏”,并认为引入因果推理是解决具身智能、科学发现等复杂问题的关键,这代表了技术演进的一个重要探索方向。观点来源: @huang_biwei
AI 对劳动价值的重塑与反思: @ruanyf 分享了黑客新闻关于“AI 提高效率后能否放假”的讨论,引出一个深刻问题:如果 AI 仅提高了生产效率而未增加员工福利,那么对个体的意义何在?这触及了 AI 时代生产关系变革的核心矛盾。
“测试是新的护城河”,代码本身的壁垒在消失: @ruanyf 观察到有工程师用 1100 美元成本的 AI 复刻了 Next.js,借此指出代码的护城河已不复存在。未来软件的真正壁垒在于详尽的测试用例和领域 know-how,而非代码本身。这是一个对软件行业有深远影响的判断。
端侧模型专属的注意力机制技术创新: @vista8 深度解析了 百度 Unlimited OCR 模型的技术原理,其借鉴人类抄书方式的“滑动注意力窗口”机制,成功解决了长文档处理的 KV 缓存膨胀问题。这是工程技术创新提升模型应用效果的绝佳案例。
企业内部的 AI 创新困境: @dotey 详细报道了 Google 工程师因开发了广受欢迎的 Google Workspace CLI 而被开除的事件。这个案例深刻揭示了大公司内部 AI agent 战略与既有管理层利益之间的潜在冲突,以及官僚主义对创新的压制。
自媒体流量方法论与 AI 结合: @gefei55 分享了利用 AI agent 监控 Twitter 带链接的高互动推文,从而在 Google Trends 反应之前发现新词、新项目,实现“抢时间差”的出海 SEO 策略。@vista8 甚至用 Codex 开发了一个模拟知名媒体“新智元”标题风格的 Skill,展示了 AI 在特定内容生产模式上的解构能力。
3. 推荐的工具或资源 链接到标题
AI 编程与开发工具
- Gemma 4 12B Coder: Google 发布的最新本地代码生成模型 (@HuggingModels, 评测 @zhixianio)。
- Cowart: Codex + 无限画布工具插件,已开源,支持更直觉的图片标注和修改方式 (@zhongerxin, 推荐 @vista8)。
- DevSpace: 一个开源 MCP 服务器,能让 ChatGPT 网页端获得类似 Codex 的本地代码操作能力 (@gefei55)。
- getdesign.md: 收集了 Linear, Vercel, Notion 等真实品牌的
DESIGN.md设计系统文件。将其提供给 AI 编程工具,可大幅提升生成 UI 的专业感和一致性 (@Pluvio9yte)。 - 腾讯云 EdgeOne Makers: 新发布的一站式 Agent 托管平台,解决本地 demo 上线时遇到的并发、存储、安全等痛点,提供免费额度 (@AI_Jasonyu)。
AI 效率与自动化
- Fable-5 (体验技巧): 使用
/effort命令切换至max强度,可解锁更强大的能力 (@Pluvio9yte)。 - Claude Tag: Anthropic 发布的 Slack 新功能,让 Claude 以同事身份常驻频道,可被 @分配任务,支持多人协作、持续学习和主动推送 (@claudeai, 解读 @dotey)。
- 公众号转 PPT 工具: @vista8 正在开发中的开源项目,可将公众号文章一键转为可修改的 PPTX 文件。
- YouMind: 被 @AI_Jasonyu 和 @gefei55 推荐的内容创作工具,尤其擅长与各发布平台(如 X Article)的排版结合。
- Fable-5 (体验技巧): 使用
开源模型与知识库
- 百度 Unlimited OCR: 开源的长文档 OCR 模型,采用滑动注意力机制,在超长 PDF 处理上效果出色,仅 1.5MB (@Pluvio9yte, @vista8)。
- WeKnora: 腾讯悄悄开源的企业级知识平台,集成了 RAG 问答、ReAct Agent 和自维护 Wiki/知识图谱功能 (@Pluvio9yte)。
- 《Deep Agents in Action》: @zhanghaili0610 开源的第三本关于 Agent 开发的新书 (@vista8)。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 35 条更新
All-In Podcast (A_full) 链接到标题
- GameStop CEO Ryan Cohen’s $56B Plan to Take Over eBay
- 发布时间:2026-06-24 04:52 北京时间
- 摘要:- AppLovin Ads — AppLovin 的 AI 广告平台在移动游戏领域覆盖超过 10 亿每日活跃用户。
- 全屏视频广告,观看时间中位数为 35 秒。
- 广告商每天花费数十万美元获利,并且广告商访问仍处于封闭测试阶段。
- 纳斯达克 - 纳斯达克定位于技术和资本市场的纽带,以无与伦比的技术、见解和市场专业知识为全球资本市场及其他市场提供一流的平台和服务。
- GameStop 首席执行官 Ryan Cohen 的 560 亿美元接管 eBay 的计划。
- EN 要点:
- (0:00) David Friedberg intros GameStop CEO Ryan Cohen
- (1:56) Building and selling Chewy for $3.35B, how to compete with Amazon in e-commerce
- (11:58) Post-Chewy life, activist investing, the road to GameStop CEO, expanding into collectibles
- (26:39) Why he wants to buy eBay for $56B: Massive potential, poor execution (slow growth, rising expenses, seller relationship failure)
Stratechery by Ben Thompson (A_full) 链接到标题
- Memory Chips and China, Microsoft and Chinese Models
- 发布时间:2026-06-23 18:00 北京时间
- 摘要:- 三大内存制造商可能会后悔向中国内存制造商敞开大门;与此同时,微软非常积极地使用中国模式。
- 15 美元/月或150 美元/年。
- 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
- 策略采访。
- 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
- EN 要点:
- The big three memory makers may come to regret opening up the door to Chinese memory makers; Microsoft, meanwhile, is very incentivized to use Chinese models.
OpenAI Blog (A_full) 链接到标题
How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery
- 发布时间:2026-06-24 01:00 北京时间
- 摘要:- 医生兼免疫学家 Derya Unutmaz 多年来一直对人工智能感兴趣。
- 但他的“顿悟”时刻出现在 2025 年底,当时 GPT-5 Pro 帮助他和他的实验室重新审视了一个长达三年的谜题,该谜题的核心是一种特殊类型的免疫细胞,可以帮助人体对抗癌症和其他疾病。
- 这个谜团集中在免疫学中一个基本但重要的问题上:葡萄糖如何影响 T 细胞的发育和特化方式?
- T 细胞是免疫细胞,可以帮助身体对抗病毒、杀死癌细胞、对某些细菌和寄生虫做出反应,以及区分健康细胞和威胁。
- 随着他们的成长,他们承担不同的工作,包括可能导致癌症、自身免疫性疾病和感染的角色。
- EN 要点:
- GPT-5 Pro helped solve a 3-year-old immunology mystery, offering insights into T cell behavior
- The breakthrough could support cancer and autoimmune research.
Helping build shared standards for advanced AI
- 发布时间:2026-06-23 21:00 北京时间
- 摘要:- 能力日益增强的模型可以加强网络防御、加速科学发现并扩大获得专业知识的机会。
- 但如果他们的能力被误解、保障措施不足或政府缺乏应对所需的信息,他们也可能造成安全风险。
- 为了安全、自信地实现效益,社会将需要具有技术和治理能力的机构来评估、保护和治理能力日益增强的系统。
- Appia 将开发开放式模块化规范,旨在将国际标准和既定框架转化为整个人工智能价值链的实用评估标准。
- 它的工作可以帮助开发一个关键的缺失信任层,第三方可以通过该层检查是否符合标准,当不同组织开发模型、基础设施和应用程序时,生成更清晰、更可重用的证据。
- EN 要点:
- OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the Appia Foundation.
How Omio is building the future of conversational travel
- 发布时间:2026-06-23 08:00 北京时间
- 摘要:- 了解 Omio 如何使用 OpenAI 来增强对话式旅行体验、加速产品开发并转型为一家 AI 原生公司。
- OpenAI 博客的这篇文章解释了 Omio 如何构建对话式旅行的未来,塑造更广泛的人工智能和基础设施格局。
- 继 Omio 如何构建对话式旅行的未来之后,它还为创始人、运营商和投资者带来了实际意义。
- EN 要点:
- Discover how Omio uses OpenAI to power conversational travel experiences, accelerate product development, and transform into an AI-native company.
ArXiv cs.AI (B_intro+search) 链接到标题
On the Identifiability of User Adaptation in Co-Adaptive Neural Interfaces
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20569v1 公告类型:新。
- 摘要:我们分析了自适应人机系统中的可识别性。
- 我们表明闭环编码器估计不能唯一地识别用户适应,而是反映联合系统的属性。
- 我们讨论解释行为适应的影响并提出识别条件。
- EN 要点:
- arXiv:2606.20569v1 Announce Type: new
- Abstract: We analyze identifiability in co-adaptive human-machine systems
- We show that closed-loop encoder estimates do not uniquely identify user adaptation, but instead reflect properties of the joint system
- We discuss implications for interpreting behavioral adaptation and propose conditions for identification.
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20599v1 公告类型:新。
- 摘要:思想树(ToT)搜索已成为提高大型语言模型推理能力的一个有前途的方向,但在实践中部署这些方法提出了一个很少受到系统关注的问题:不同的搜索策略在不同的计算预算、模型大小和问题难度下表现如何?
- 在这项工作中,我们评估了两种代表性的 ToT 方法; DPTS(一种基于蒙特卡罗树搜索的方法)和 SSDP(一种基于语义重复数据删除的方法)跨越两个数学推理基准(Math500 和 GSM8K)、两个模型规模(Llama-3B 和 Llama-8B)以及四个令牌预算(3k–10k)。
- 我们的分析表明,这两种方法表现出相反方向的局限性。
- EN 要点:
- arXiv:2606.20599v1 Announce Type: new
- Abstract: Tree of Thought (ToT) search has become a promising direction for improving the reasoning capabilities of large language models, but deploying these m…
- In this work, we evaluate two representative ToT methods; DPTS, a Monte Carlo tree search based approach, and SSDP, a semantic deduplication based approach, acr…
- Our analysis reveals that the two methods exhibit limitations that pull in opposite directions
The New Associationism: Lessons from Deep Learning
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20600v1 公告类型:新。
- 摘要:现代人工智能的成功可以告诉我们有关人类如何学习的哪些信息?
- 本文认为,认真对待人工智能作为人类学习的模型支持适度但真正的联想主义。
- 核心发现是监督学习(由评估反馈驱动的学习)是令人惊讶的广泛当代人工智能系统的基础,从大型语言模型到游戏代理,其主要区别在于生成相关反馈信号需要多少工作。
- EN 要点:
- arXiv:2606.20600v1 Announce Type: new
- Abstract: What can the success of modern AI tell us about how humans learn
- This paper argues that taking AI seriously as a model of human learning supports a modest but genuine associationism
- The central finding is that supervised learning – learning driven by evaluative feedback – underlies a surprisingly wide range of contemporary AI systems, fro…
Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20615v1 公告类型:新。
- 摘要:人工智能代理现在作为一流的团队成员参与整个软件开发生命周期,但不存在规范语言来表达这种协作所需的人类代理责任边界、审批门和治理约束。
- 现有方法在代理提示中对流程进行编码(可能会发生偏差)、针对相邻域(工作流程管理、业务流程)或仅处理片段(访问控制、审批门)。
- 我们提出了一种特定于领域的语言,用于将 AI-SDLC 流程指定为协议,具有形式语法、格式良好的条件、操作语义和强制不变量。
- EN 要点:
- arXiv:2606.20615v1 Announce Type: new
- Abstract: AI agents now participate as first-class team members across the software development lifecycle, yet no specification language exists for expressing t…
- Existing approaches encode process in agent prompts (subject to drift), target adjacent domains (workflow management, business processes), or address only fragm…
- We propose a domain-specific language for specifying AI-SDLC processes as protocols, with formal syntax, well-formedness conditions, operational semantics, and…
PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20621v1 公告类型:新。
-摘要:多智能体辩论通过迭代同行批评提高了大型语言模型(LLM)的可靠性。
- 然而,固定拓扑通常会引入持续的位置偏差,放大不可靠的代理,并导致对角色分配的高度敏感。
- 我们引入了 \textit{排列等变自适应路由多代理辩论(PEAR)},这是一种推理时间协议,可以在连续的辩论轮次中动态地重新配置通信角色和稀疏拓扑。
- EN 要点:
- arXiv:2606.20621v1 Announce Type: new
- Abstract: Multi-agent debate improves the reliability of large language models (LLMs) through iterative peer critiques
- However, fixed topologies often introduce persistent positional biases, amplify unreliable agents, and cause high sensitivity to role assignments
- We introduce \textit{Permutation-Equivariant Adaptive Routing Multi-Agent Debate (PEAR)}, an inference-time protocol that dynamically reconfigures communication…
Darwin Mobile Agent: A Roadmap for Self-Evolution
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20622v1 公告类型:新。
- 摘要:人工智能的目标是创建能够在开放环境中进行通用、自适应行为的代理。
- 在“惨痛的教训”的指导下,我们认为实现这一目标的最有效途径是系统地消除人类的先验,并让智能通过与比智能体本身复杂几个数量级的“大世界”交互而自然地出现。
- 我们提出移动图形用户界面(GUI)作为这样一个世界的实用代理,并引入 Darwin Mobile Agent,这是一个开源基础设施,旨在作为该领域自主强化学习的基础。
- EN 要点:
- arXiv:2606.20622v1 Announce Type: new
- Abstract: The goal of artificial intelligence is to create agents capable of general, adaptive behaviour in open-ended environments
- Guided by the “Bitter Lesson”, we argue that the most effective path toward this goal is to systematically remove human priors and allow intelligence to natural…
- We propose the mobile Graphical User Interface (GUI) as a practical proxy for such a world and introduce Darwin Mobile Agent, an open-source infrastructure desi…
Path-dependent program induction under resource constraints explains human sequence learning
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20623v1 公告类型:新。
- 摘要:人们如何在有限的认知资源下从连续的经验中构建抽象的、可重用的知识?
- 为了回答这个问题,我们将率失真理论与程序归纳的最新进展相结合,以描述先验知识如何塑造哪些未来结构的编码成本低且易于发现。
- 我们将其形式化为分层适配器语法(HAG),具有不同的本地(任务内)和全局(跨任务)库,由内存和计算的约束共同管理。
- EN 要点:
- arXiv:2606.20623v1 Announce Type: new
- Abstract: How do people build abstract, reusable knowledge from sequential experience under bounded cognitive resources
- To answer this question, we integrate rate-distortion theory with recent advances in program induction to describe how prior knowledge shapes which future struc…
- We formalize this in a hierarchical Adaptor Grammar (HAG) with distinct local (within-task) and global (across-task) libraries, governed jointly by constraints…
In LLM Reasoning, there is Irrationality on top of Value Misalignment
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20624v1 公告类型:新。
- 摘要:在使法学硕士与目标价值函数保持一致方面取得了重大进展。
- 我们认为,即使法学硕士在(后)培训中得到了很好的协调,它仍然可能无法最大化推理中的协调价值。
- 我们在数学上将这种差距形式化为理性价值风险:模型部署的推理策略与其理性对应策略之间的效用差异,其定义为在最陡方向上最大化预期效用的响应。
- EN 要点:
- arXiv:2606.20624v1 Announce Type: new
- Abstract: Significant progress has been made in aligning LLMs with target value functions
- We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning
- We mathematically formalise this gap as rational value risk: the utility discrepancy between a model’s deployed reasoning strategy and its rational counterpart,…
AlphaMemo: Structured Search-Process Memory for Self-Evolving Alpha Mining Agents
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20625v1 公告类型:新。
- 摘要:LLM 代理有望通过结合财务先验、符号推理、可执行因子生成和反馈驱动的细化来进行 alpha 挖掘。
- 然而,他们面临着组合搜索空间、嘈杂的非平稳反馈、冗余发现以及天真地重用过去成功的过度拟合风险。
- 为了应对这些挑战,我们提出了 AlphaMemo,一种具有结构化搜索过程内存的自我进化 alpha 挖掘代理。
- EN 要点:
- arXiv:2606.20625v1 Announce Type: new
- Abstract: LLM agents are promising for alpha mining via combining financial priors, symbolic reasoning, executable factor generation, and feedback-driven refine…
- Yet, they face a combinatorial search space, noisy non-stationary feedback, redundant discoveries, and overfitting risks from naively reusing past successes
- To address these challenges, we propose AlphaMemo, a self-evolving alpha mining agent with Structured Search-Process Memory
Latent Goal Prediction from Language for Model-Based Planning
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20627v1 公告类型:新。
- 摘要:使用世界模型进行规划因复合预测误差和定义可优化目标的难度而受到瓶颈。
- 视觉目标提供精确的局部梯度,但远距离指导较差,而语言虽然灵活,但受到嘈杂的跨模式对齐或对大型生成模型的依赖的限制,不适合基于模型的规划的高采样性质。
- 为了应对这些挑战,我们引入了来自语言的潜在目标预测(LAGO),这是一个框架,可以根据语言指令和动作条件推出来预测中间目标状态序列,所有这些都在同一潜在空间内。
- EN 要点:
- arXiv:2606.20627v1 Announce Type: new
- Abstract: Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals
- Visual targets provide precise local gradients but poor distant guidance, while language is flexible yet limited by noisy cross-modal alignment or dependence on…
- To address these challenges, we introduce Latent Goal Prediction from Language (LAGO), a framework that predicts both sequences of intermediate goal states from…
ArXiv cs.CL (B_intro+search) 链接到标题
Less is More: Lightweight Prompt Compression for Question Answering Applications on Edge Devices
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20571v1 公告类型:新。
- 摘要:在代理驱动的问答 (QA) 应用中,通常引入检索增强生成 (RAG),通过提供额外的上下文来提高大型语言模型 (LLM) 的响应准确性。
- 由于检索结果中固有的噪声和文档级检索的粗粒度,检索到的上下文通常包含大量冗余信息。
- 在此设置中,由用户查询和关联的检索到的上下文组成的代理提示会在 LLM 推理期间导致不必要的计算开销。
- EN 要点:
- arXiv:2606.20571v1 Announce Type: new
- Abstract: In agent-driven question answering (QA) applications, retrieval-augmented generation (RAG) is commonly introduced to enhance the response accuracy of…
- Due to the inherent noise in retrieval results and the coarse granularity of document-level retrieval, the retrieved context often contains substantial redundan…
- In this setting, the agent prompt, consisting of the user query and the associated retrieved context, leads to unnecessary computational overhead during LLM inf…
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20572v1 公告类型:新。
- 摘要:实现对大型语言模型 (LLM) 的可靠控制需要精确、可扩展地理解它们如何解释语言线索。
- 我们引入了一个严格的框架,使用 Shapley 值来量化单个形容词对模型性能的影响,超越轶事启发法到原则归因。
- 将此方法应用于 MMLU 基准上各种模型(包括 o3、gpt-4o-mini、phi-3、llama-3-70b 和 deepseek-r1)中的 100 个形容词,我们发现了 AI 对齐的几个关键发现。
- EN 要点:
- arXiv:2606.20572v1 Announce Type: new
- Abstract: Achieving reliable control of Large Language Models (LLMs) requires a precise, scalable understanding of how they interpret linguistic cues
- We introduce a rigorous framework using Shapley values to quantify the steering effect of individual adjectives on model performance, moving beyond anecdotal he…
- Applying this method to 100 adjectives across a diverse suite of models (including o3, gpt-4o-mini, phi-3, llama-3-70b, and deepseek-r1) on the MMLU benchmark,…
Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20632v1 公告类型:新。
- 摘要:多法学硕士系统使用多种语言模型来审议、判断彼此的输出或作为代理进行协调。
- 它们的价值取决于模型在给予相同输入时产生明显不同的对话行为。
- 之前的离线研究建议为每个家庭绘制一个行为多样性模型,因为法学硕士在单独评价彼此时更喜欢来自自己家庭的输出。
- EN 要点:
- arXiv:2606.20632v1 Announce Type: new
- Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other’s outputs, or coordinate as agents
- Their value depends on the models producing measurably different conversational behaviors when given the same input
- Prior offline studies recommend drawing one model per family for behavioral diversity, because LLMs prefer outputs from their own family when rating one another…
EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20650v1 公告类型:新。
- 摘要:基于指令的可控语音合成使用户能够通过自然语言指定情感。
- 然而,现有的方法通常依赖于粗略的情感标签,并且缺乏细粒度强度的明确建模。
- 我们提出了 EmoInstruct-TTS,一种用于情感语音合成的双路径指令引导框架。
- EN 要点:
- arXiv:2606.20650v1 Announce Type: new
- Abstract: Instruction-based controllable speech synthesis enables users to specify emotions through natural language
- However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained intensity
- We propose EmoInstruct-TTS, a dual-path instruction-guided framework for emotional speech synthesis
Specific Domain Ontology Construction Using Large Language Models
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20691v1 公告类型:新。
- 摘要:本体是组织和维护人类和系统都可以理解的信息的有用结构。
- 然而,由于他们的手工制作是一项艰巨的任务,许多特定领域缺乏参考本体。
- 大型语言模型(LLM)所表现出的理解自然语言的杰出能力促使其应用到各种领域,包括本体开发。
- EN 要点:
- arXiv:2606.20691v1 Announce Type: new
- Abstract: Ontologies are useful structures to organize and maintain information that can be understood both by humans and systems
- However, since their manual crafting is a laborious task, many specific domains lack reference ontologies
- The outstanding ability for understanding natural language demonstrated by the Large Language Models (LLMs) has motivated their application to aid on a variety…
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20696v1 公告类型:新。
- 摘要:由于缺乏明显的语言输出、有限的训练数据和较大的主体间变异性,从非侵入性大脑信号中解码内部语音仍然是一个基本挑战。
- 现有的大脑到文本的方法通常依赖于特定于任务的解码器微调,这限制了可扩展性并使对新参与者的适应变得复杂。
- 我们提出了 MindAlign,一种解耦的两阶段大脑到语言框架,可以从 fMRI 信号生成开放式文本,而无需修改底层语言模型。
- EN 要点:
- arXiv:2606.20696v1 Announce Type: new
- Abstract: Decoding inner speech from non-invasive brain signals remains a fundamental challenge due to the absence of overt linguistic output, limited training…
- Existing brain-to-text approaches often rely on task-specific decoder fine-tuning, which restricts scalability and complicates adaptation to new participants
- We propose MindAlign, a decoupled two-stage brain-to-language framework that enables open-ended text generation from fMRI signals without modifying the underlyi…
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20740v1 公告类型:新。
- 摘要:过程奖励模型 (PRM) 为大型语言模型 (LLM) 推理提供了步骤级验证,但其训练数据获取仍然是一个瓶颈:人工注释成本高昂,而且蒙特卡洛推出估计存在噪声。
- 最近的方法 FOVER 在由 Z3 和 Isabelle 等形式验证工具自动注释的步骤级错误标签上训练 PRM,并凭经验观察从符号任务到不同推理基准的跨任务泛化。
- 然而,这种泛化现象缺乏任何理论解释,并且此类 PRM 的泛化误差、样本复杂性、收敛速度或下游 Best-of-K 性能不存在正式界限。
- EN 要点:
- arXiv:2606.20740v1 Announce Type: new
- Abstract: Process Reward Models (PRMs) provide step-level verification for Large Language Model (LLM) reasoning, yet their training data acquisition remains a b…
- A recent approach, FOVER, trains PRMs on step-level error labels automatically annotated by formal verification tools such as Z3 and Isabelle, and empirically o…
- However, this generalization phenomenon lacks any theoretical explanation, and no formal bounds exist on the generalization error, sample complexity, convergenc…
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20751v1 公告类型:新。
- 摘要:先进空中机动(AAM)是一种新兴的低空航空运输系统,其成功部署不仅取决于技术进步,还取决于公众的接受度。
- 这种接受将推动政府支持、法规、噪音标准和飞行意愿,进而提高 AAM 的整体商业可行性。
- 因此,了解公众对 AAM 的情绪对于识别其社会障碍并为其采用策略提供信息至关重要。
- EN 要点:
- arXiv:2606.20751v1 Announce Type: new
- Abstract: Advanced Air Mobility (AAM) is an emerging low-altitude air transportation system whose successful deployment depends not only on technological advanc…
- This acceptance will drive government support, regulations, noise standards, and willingness to fly, and in turn the overall commercial viability of AAM
- Understanding public sentiment toward AAM is therefore essential for identifying its societal barriers and informing its adoption strategies
FirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20769v1 公告类型:新。
- 摘要:用于同行评审的人工智能系统在三个方面失败:它们仅在计算机科学和机器学习场所进行训练,忽略验证科学的迭代对话,并根据风格模仿而不是真正的编辑判断进行评估。
- 我们引入了 FirstPass,一个解决所有这三个问题的数据集和微调模型。
- 策划《自然通讯》跨五个科学领域(生物学、化学、神经科学、物理学和地球科学)的 3,668 场完整的多轮同行评审对话,我们利用强制性透明同行评审(2022 年 11 月开始)并通过自动审核验证 100% 内容完整性。
- EN 要点:
- arXiv:2606.20769v1 Announce Type: new
- Abstract: AI systems for peer review fail on three fronts: they train on Computer Science and Machine Learning venues alone, ignore the iterative dialogue that…
- We introduce FirstPass, a dataset and fine-tuned model that addresses all three
- Curating 3,668 complete multi-round peer-review dialogues from Nature Communications across five scientific domains (biology, chemistry, neuroscience, physics,…
Beyond ‘One Language, One Script’: Quantifying Orthographic Bias in Multilingual VLMs with PuMVR
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20770v1 公告类型:新。
- 摘要:当前的视觉语言模型 (VLM) 因其多语言功能而闻名,但它们在一个有缺陷的假设下运行:一种语言对应于一种书写系统。
- 这忽略了旁遮普语、塞尔维亚语、印地语-乌尔都语、库尔德语等多文字语言的数十亿用户,对他们来说,模型的能力可能会因拼写偏差而受到影响。
- 我们推出 PuMVR(旁遮普多模态视觉推理),这是第一个基准测试,旨在通过旁遮普语三种活跃文字(Gurmukhi、Shahmukhi、Roman)中的 375 个基于文化的图像推理任务来量化依赖于文字的偏见。
- EN 要点:
- arXiv:2606.20770v1 Announce Type: new
- Abstract: Current Vision-Language Models (VLMs) are celebrated for their multilingual capabilities, yet they operate under a flawed assumption: that one languag…
- This overlooks billions of users of multi-script languages like Punjabi, Serbian, Hindi-Urdu, Kurdish, among many others, for whom a model’s capability may be f…
- We introduce PuMVR (Punjabi Multimodal Visual Reasoning), the first benchmark designed to quantify script-dependent bias through 375 culturally grounded image-r…
ArXiv cs.LG (B_intro+search) 链接到标题
Towards CSI-Native Foundation Models: A Channel-Adaptive Roadmap for 6G
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20670v1 公告类型:新。
- 摘要:无线基础模型为第六代 (6G) 系统提供了一条通往可重用信道状态信息 (CSI) 智能的途径。
- 然而,现有的通用骨干适应和 CSI 预训练方法通常将 CSI 视为任务张量而不是传播条件信道响应,从而无法捕获无线环境的固有时频空间几何形状。
- 本文提出了面向 CSI 原生基础模型的通道自适应路线图,提出了一个统一的框架,将预训练、位置建模和注意力控制与三个通道要求保持一致:尺度感知异构暴露、物理时频天线坐标和相关性限制令牌交互。
- EN 要点:
- arXiv:2606.20670v1 Announce Type: new
- Abstract: Wireless foundation models offer a path toward reusable channel state information (CSI) intelligence for sixth-generation (6G) systems
- However, existing generic-backbone adaptation and CSI pretraining methods often treat CSI as task tensors rather than propagation-conditioned channel responses,…
- This paper presents a channel-adaptive roadmap toward CSI-native foundation models, proposing a unified framework that aligns pretraining, positional modeling,…
NeuroShield: A Device-Agnostic Foundation Model for EEG Authentication
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20673v1 公告类型:新。
- 摘要:脑电图验证的一个核心挑战是模型通常与训练它们的采集设置相关联。
- 特别是,耳机硬件、通道布局和信号持续时间的变化会产生现有模型无法处理的异构录音,从而导致每个新耳机或数据集被视为单独的模型开发问题。
- 这种碎片化限制了多数据集学习,阻碍了知识转移,并降低了模型的可重用性。
- EN 要点:
- arXiv:2606.20673v1 Announce Type: new
- Abstract: A central challenge in EEG authentication is that models are typically tied to the acquisition settings in which they are trained
- In particular, variations in headset hardware, channel layout, and signal duration create heterogeneous recordings that existing models are not designed to hand…
- This fragmentation limits multi-dataset learning, hinders knowledge transfer, and reduces model reusability
Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20743v1 公告类型:新。
- 摘要:经过训练的 Transformer 可靠地开发大量激活,即少量隐藏维度,其大小远高于中值,并且集中在序列开始标记上。
- 这些异常值是否是残留流超载读写角色的可移除工件,或者是功能上的必需品,目前正在积极争论。
- 我们通过架构干预直接测试工件假设。
- EN 要点:
- arXiv:2606.20743v1 Announce Type: new
- Abstract: Trained transformers reliably develop massive activations, a small number of hidden dimensions whose magnitude is far above the median and which conce…
- Whether these outliers are a removable artifact of the residual stream’s overloaded read and write role, or instead a functional necessity, is actively debated
- We test the artifact hypothesis directly, with an architectural intervention
CIExplainer++: Generating Causal and Interpretable Explanations for Graph Neural Networks
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20747v1 公告类型:新。
- 摘要:可解释的人工智能旨在通过以人类可理解的方式呈现导致模型输出的元素,使黑盒模型更值得信赖。
- 这涉及(i)识别对输出具有真正因果影响的组成部分和联系,以及(ii)将此类结构转化为可解释的表示。
- 对于前者,我们引入 CIExplainer,这是一种基于因果推理的新颖的基于扰动的方法,用于解释图神经网络(GNN)。
- EN 要点:
- arXiv:2606.20747v1 Announce Type: new
- Abstract: Explainable Artificial Intelligence aims to make black-box models more trustworthy by presenting, in a human-understandable manner, the elements that…
- This involves both (i) identifying components and connections with genuine causal influence on outputs and (ii) translating such structures into an interpretabl…
- For the former, we introduce CIExplainer, a novel perturbation-based method grounded in causal inference for explaining Graph Neural Networks (GNNs)
Evidential Fusion Network for Multimodal Survival Prediction under Missing Modalities
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20757v1 公告类型:新。
- 摘要:最近的多模式生存预测模型通过利用跨模式的互补信息表现出了强大的预测性能。
- 然而,此类模型通常假设数据完整性,并且对缺失模式表现出有限的鲁棒性,这在现实世界的临床环境中经常遇到。
- 我们提出了证据缺失模态生存融合(EMMS)模型,用于缺失模态下的多模态生存预测。
- EN 要点:
- arXiv:2606.20757v1 Announce Type: new
- Abstract: Recent multimodal survival prediction models have demonstrated strong predictive performance by leveraging complementary information across modalities
- However, such models generally assume data completeness and exhibit limited robustness toward missing modalities, which are frequently encountered in real-world…
- We propose the Evidential Missing Modality Survival Fusion (EMMS) model for multimodal survival prediction under missing modalities
ELADO: Elliptic PDE Assessment Datasets for Operator Learning
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20771v1 公告类型:新。
- 摘要:我们介绍 ELADO(用于算子学习的椭圆偏微分方程评估数据集),这是一个系统基准套件,用于在学习椭圆偏微分方程解算子时显示和量化神经算子架构的故障模式。
- 虽然现有数据集的基准侧重于平均情况性能,但 ELADO 数据集的构建是为了突出椭圆偏微分方程问题中自然出现的挑战。
- 特别是,我们构建了几个围绕泊松方程和亥姆霍兹方程构建的数据集,每个数据集都具有非常数系数。
- EN 要点:
- arXiv:2606.20771v1 Announce Type: new
- Abstract: We introduce ELADO (Elliptic PDE Assessment Datasets for Operator Learning), a systematic benchmark suite constructed to show and quantify failure mod…
- While the benchmarks of existing datasets focus on average case performance, the ELADO datasets are constructed to highlight challenges that arise naturally in…
- In particular, we construct several datasets built around Poisson’s equation and the Helmholtz equation, each with non-constant coefficients
B[FM]$^2$: Brain Foundation Model via Flow Matching with SplitUNet
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20812v1 公告类型:新。
- 摘要:脑电图基础模型可以从大规模脑电图语料库中学习通用表示,从而实现跨不同临床和脑机接口任务的单主干传输。
- 现有模型通常将连续的多通道脑电图波形离散化为补丁或码本标记,并训练具有屏蔽自我监督的变压器。
- 认识到这种离散化会破坏连续的大脑节律并掩盖细粒度的时间动态,我们提出了 B[FM]$^2$(通过流匹配的大脑基础模型),其归纳偏差通过使用连续时间流匹配直接在原始信号上进行预训练而与数据对齐,无需补丁、标记化或掩蔽。
- EN 要点:
- arXiv:2606.20812v1 Announce Type: new
- Abstract: EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical…
- Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervis…
- Recognizing that this discretization fragments continuous brain rhythms and obscures fine-grained temporal dynamics, we present B[FM]$^2$(Brain Foundation Model…
CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20820v1 公告类型:新。
- 摘要:我们可以相信评估分数能够反映法学硕士真实的现实表现吗?
- 可认证的评估通过为LLM评估提供保证来回答这个问题。
- 特别是,现有方法顺序地策划评估样本并不断更新以高概率(例如,95%)覆盖真实性能的置信区间(CI),直到满足某些条件,例如 CI 宽度达到目标精度。
- EN 要点:
- arXiv:2606.20820v1 Announce Type: new
- Abstract: Can we trust evaluation scores to capture an LLM’s true real-world performance
- Certifiable evaluation answers this question by providing guarantee for LLM evaluation
- In particular, existing methods sequentially curate evaluation samples and keep updating confidence intervals (CIs) that cover the true performance with high pr…
Evolutionary Discovery of Developmental Reward Schedules in Deep Reinforcement Learning
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20858v1 公告类型:新。
- 摘要:强化学习(RL)中奖励构成的时间结构通常是手工设计的,并在整个训练过程中保持固定,而动机优先事项的进展在很大程度上未被探索。
- 在这项工作中,我们提出了一个用于发现发展奖励计划的进化框架,其中三个不同的受生物学启发的动机成分——代理性、新颖性和反应性——通过在训练过程中动态变化的时变权重结合起来。
- 对两个稀疏奖励 MiniGrid 任务:DoorKey-6x6 和 KeyCorridorS3R1 进行评估,我们的框架将四种进化算法:CMA-ES、xNES、DE 和 L-SHADE 与外部动机基线(我们的主要比较点)和另外三种手工设计方法的通用性进行了比较。
- EN 要点:
- arXiv:2606.20858v1 Announce Type: new
- Abstract: The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the…
- In this work, we propose an evolutionary framework for discovering developmental reward schedules, in which three distinct biologically inspired motivational co…
- Evaluated on two sparse-reward MiniGrid tasks: DoorKey-6x6 and KeyCorridorS3R1, our framework compares the generalizability of four evolutionary algorithms: CMA…
Machine Learning Classification of Cryopathy Syndromes: A Comprehensive Comparative Study
- 发布时间:2026-06-23 12:00 北京时间
- 摘要:- arXiv:2606.20874v1 公告类型:新。
- 摘要:冷冻病综合征很难分类,因为实验室模式经常在诊断类别之间重叠,而某些诊断却很少见。
- 这使得冷球蛋白相关测试的常规解释具有挑战性,并增加了对专家判断的依赖。
- 这项研究的目的是开发和比较机器学习方法,用于根据实验室数据对冷冻病综合征进行自动分类,并确定临床决策支持的实用策略。
- EN 要点:
- arXiv:2606.20874v1 Announce Type: new
- Abstract: Cryopathy syndromes are difficult to classify because laboratory patterns often overlap across diagnostic categories, while some diagnoses are rare
- This makes routine interpretation of cryoglobulin-related tests challenging and increases dependence on expert judgment
- The aim of this study was to develop and compare machine learning approaches for automated classification of cryopathy syndromes from laboratory data and to ide…