🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-16
- 类型
- ai-daily
- 字数
- 3451
- 阅读时长
- 17 min
2026-06-16 AI日更 | 智能体进入真实工作流,企业开始重估“Token 资本” 链接到标题
今天的主线是 AI Agent 从演示走向真实生产:办公代理完成率显著提升,软件工厂、BI、多模态编排等场景加速落地。与此同时,评测、安全、拒答控制和企业专有知识沉淀成为新基础设施,端侧代码模型也开始接受更严苛的生产力检验。
📖 本期 Watch List 深度导读 链接到标题
今天最值得细读的是“智能体走向真实工作流”这条线:WorkBench 两年回访显示,顶级办公代理任务完成率从 43% 提升到 89%,误操作显著下降;同时 TwinBI、WebDecept、Orchestra-o1 分别把代理推进到 BI 分析、多模态编排和电商安全测试,适合产品与平台团队系统跟进。
第二条是“评测与安全正在变得更细”。Anthropic 安全叙事、LLM-as-a-Judge 稳定性研究、拒绝机制干预、样本选择偏差导致模型崩溃等更新,共同提醒我们:模型能力扩张之后,可靠评估、拒答控制和数据闭环反而成为基础设施问题。
最后,创意 AI 值得关注 Ideogram 的开放权重图像模型访谈。它把焦点从“生成质量”转向文本、布局、可控编辑与设计工作流,代表图像模型进入更工程化的产品竞争阶段。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:AI Coding Agents Speed Past Human Review Limits 链接到标题
- 分类:AI · News
- 概况:热度时间:,相关帖子数:75
- 摘要:AI Coding Agents Speed Past Human Review Limits: A.I News A.I Def: artificial intelligence (AI), the ability of a digital computer or computer-controlled robot to perform tasks commonly associated with intelligent beings. The term is frequently applied to the project of developing systems endowed w…
话题 2:Factory AI Launches Factory 2.0 as Fully Autonomous Software Factories 链接到标题
- 分类:AI · News
- 概况:热度时间:5 hours ago,相关帖子数:1600
- 是什么事:Factory AI 发布 Factory 2.0,宣称将软件开发流程升级为可自主运行的“软件工厂”。
- 为什么重要:这反映出 AI Agent 正从辅助编程工具走向端到端软件交付系统,可能改变研发效率、团队分工和软件创业模式。
- 讨论概况:X 上的讨论集中在其自主开发能力是否真正可靠、DAA 等新指标能否衡量 Agent 产出,以及这类平台是否会推动硅谷投资逻辑从人力密集型团队转向高度自动化的软件工厂。
话题 3:Vercel Extends Serverless Functions to 30 Minutes for AI Workloads 链接到标题
- 分类:AI · News
- 概况:热度时间:5 hours ago,相关帖子数:184
- 是什么事:Vercel 将 Serverless Functions 的最长执行时间延长至 30 分钟,以更好支持长时间运行的 AI 推理、生成和后台任务。
- 为什么重要:这降低了开发者在 Vercel 上部署 AI 应用的架构复杂度,使长文本生成、Agent 工作流、多步骤推理等任务不必过早拆分到队列或独立后端服务。
- 讨论概况:X 上讨论主要集中在这是否会让 Vercel 更适合 AI 原型和生产部署;支持者认为它简化了全栈 AI 应用开发,质疑者则关注成本、冷启动、稳定性以及相比专用推理平台或传统后端的性价比。
话题 4:Le Chaton Fat: Mistral AI’s Massive Meme Hoax 链接到标题
- 分类:AI · Entertainment
- 概况:热度时间:22 hours ago,相关帖子数:10000
- 是什么事:X 上热传一个疑似冒用或戏仿 Mistral AI 的“Le Chaton Fat”话题,内容被普遍认为是围绕 AI 新模型/产品的梗图式恶作剧或骗局。
- 为什么重要:它反映出 AI 行业高度关注下,模型发布、品牌消息和技术传闻很容易被 meme 化并快速扩散,凸显信息核验与平台传播机制的重要性。
- 讨论概况:讨论焦点集中在这是无害娱乐还是误导性谣言;有人把它当作 AI 圈过度炒作的讽刺,也有人担心此类假消息会混淆真实发布、损害公司声誉并误导投资者和用户。
话题 5:AI-Coded WoW Clone Draws 12,000 Players in Days 链接到标题
- 分类:AI · Entertainment
- 概况:热度时间:17 hours ago,相关帖子数:501
- 是什么事:一款由 AI 辅助编写的《魔兽世界》风格克隆游戏在数天内吸引约 1.2 万名玩家。
- 为什么重要:这显示生成式 AI 正在降低游戏开发门槛,并可能加速从原型设计到上线运营的周期,对内容生产、独立开发和游戏行业分工产生影响。
- 讨论概况:X 上的讨论主要集中在 AI 是否真正提升了开发效率、这类作品的原创性与版权风险、游戏质量能否长期留住玩家,以及 AI 工具会强化个人开发者还是进一步冲击传统游戏岗位。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线是,AI 正从“辅助工具”继续向可独立完成复杂流程的生产系统扩展,软件开发、应用部署和游戏制作都在被重新想象为更自动化、更低门槛的流程。共识在于,Agent、长时运行 AI 工作流和生成式开发工具确实正在降低原型到上线的成本,并可能改变创业团队规模、研发分工和内容生产速度。分歧主要集中在这些能力是否已经足够可靠、是否具备真实生产价值,以及平台化方案相比传统后端、专用推理服务或人工团队是否真的更划算。潜在风险则包括过度炒作导致的预期泡沫、AI 生成内容的版权与原创性争议、长任务部署的成本和稳定性问题,以及假消息或 meme 化传播混淆真实产品发布、误导用户和投资者。
💡 大佬观点(Influencer Insights) 链接到标题
AI 行业前沿洞察日报 (2026-06-15) 链接到标题
基于过去 24 小时内多位 AI 大佬在 X 平台的发言,以下是今日份的行业动态速览与深度总结。
1. 今日技术趋势与产品热点 链接到标题
🔥 热点聚焦:端侧模型的“终极对决”与能力边界 链接到标题
今日最火的技术讨论集中在代码模型的实际落地能力上,尤其是在本地硬件环境下的性能对比。
Gemma 4 12B Coder 的实战检验: 谷歌新发布的 Gemma 4 12B Coder 引起了广泛关注,但其表现颇具争议。@zhixianio 在 M5 Max 上进行了严格测评,对比对象是“日常甜点” Qwen3.6-35B-A3B MoE。结果显示:
- 简单任务(Matplotlib绘图):双方打平。
- 复杂长程任务(Three.js 星系/俄罗斯方块):12B 的 Gemma 出现了黑屏、逻辑漏洞等严重问题,完全败给 35B 的 Qwen。@zhixianio 指出,12B 体量本身无法支撑“长篇、有状态、一次成型”的复杂程序,社区微调可以提高效率,但无法提升天花板。
- 亮点:他观察到 Gemma Coder 微调版学会了“想一下就动手”,收敛速度优于原版;但原版 Gemma 4 在开启思考模式后,出现了 12000 token 全在“想”而一行代码未生成的极端现象。
多模态与端侧应用的新边疆: @zhixianio 还深度测试了 Gemma 4 的 QAT (量化感知训练) 版本和多模态能力。在 M5 Max 上,其英文和日文识别效果极佳,速度很快,但中文识别“驴唇不对马嘴”。他高度评价 Google 的端侧策略,认为通过 QAT 让模型“原生适配量化”能极大降低内存占用,预示着 Android 设备很快能流畅运行自带的高性能模型。
“苦行僧”式本地化实践: @zhixianio 分享了自己全面转向本地模型(Qwen3.6-35B-A3B)的“修行”体验,结果令人震惊:在 PA 和 Coding 场景下,响应速度比远程 LLM 更快,智商在线,原生多模态体验甚至让他觉得“比 DSV4 Pro 还要爽”。这标志着高端消费级硬件的端侧智能已正式迈入生产力门槛。
🛠️ 工具与平台:Claude Design 引发的设计范式变革 链接到标题
Claude Design 的深度解析: @dotey (宝玉) 发布了一篇长文,深入剖析了为何 GPT-5.5 (Codex) 目前还做不出类似 Claude Design 的产品。核心观点在于,Claude Design 交付的不只是 UI 设计图,而是融合了完整数据架构与状态管理逻辑的高保真可交互原型。这要求模型在动手前,必须将整个交互系统设计妥帖,目前只有 Claude Opus 4.8 做到了。
Design-to-Code 的最佳实践: @dotey 现场演示了如何将 Claude Design 融入开发流:在设计稿中修改 UI,下载后通过
git diff查看变更,再由 Claude Code 自动同步 Swift 代码。这一流程彻底改变了传统前端开发的沟通模式,让设计师变成了“设计经理”,通过自然语言指挥 Agent 执行。Fable 5 的测评与泄露风波:
- 两极分化的评价:@Pluvio9yte (雪踏乌云) 给出了不同于主流的结论——速度极慢,但思维边界和架构能力极强,像是 Claude Opus 4.6++ 与 GPT-5.5++ 的结合,且消耗速度并不可怕。而 @zhixianio 则惊叹于 Fable 的自主性,称其 40 分钟内完成 70% 的开发工作,还指出了原方案的缺陷。
- 传说中的 System Prompt 泄露:@Pluvio9yte 分享了一份声称是 Fable 5 系统提示词的文档,指出这份样本对理解 AI Agent 设计极具价值。
2. 独特观点与行业前瞻 链接到标题
🧠 开创新概念:“Token资本” (Token Capital) 链接到标题
微软 CEO Satya Nadella 提出的这一概念被 @dotey 重点转述。其核心论点是,企业未来不仅有人力资本,还必须有Token 资本——即公司自己构建并沉淀在 AI 系统中的专有知识与经验。
- 关键检验标准:能否随时替换底层通用大模型,而不丢失公司积累的经验?如果不能,就是在“租用”智能。
- 学习飞轮 / 复利效应:企业最大的危机不是技术落后,而是将“学习能力”外包。Nadella 警告,不要让少数模型垄断所有行业价值,导致产业空心化,就像早期的全球化外包一样。
📉 反直觉观察与深度反思 链接到标题
- AI 编程比人更贵?:@ruanyf 通过计算 OpenAI 员工的 Token 消耗量(单人月耗 130 万美元等值 Token)指出,如果完全放开使用顶级模型,对公司的费用极其昂贵。哪怕换用国内便宜 30-50 倍的模型,一年仍要数百万人民币。
- Vibe Coding 的工程化困境:@Pluvio9yte 以亲身经历分享了自己从“Vibe Coder”成长为规范工程师的过程。他提出 “Contract First” 原则:在没有预先定义好 API 契约、数据模型的情况下直接让 AI 写代码,整个项目就会是一堵“透风的墙”。
- AI 时代的团队文化:@dotey 引述 Lovable 设计负责人的七条经验,其中 “让资深的人重新动手” 是最有启发性的一点。AI 让经验丰富但长期脱离一线的管理者,重新获得了个体贡献者的巨大杠杆效应。
- Skill 的进化方向:@lijigang (李继刚) 提出了对未来 AI 生态的思考,认为技能 (Skill) 的发展路径分为两条:一是 “向下原子化” ,将人的能力拆解为专项技能包;二是 “向上组件化” ,将完整的场景最佳实践封装起来。
3. 推荐的工具与资源 链接到标题
- 设计与原型:
- Claude Design:由 @dotey 和 @vista8 强力推荐,用于一句话生成高精度交互原型。
- 开发者工具与 Skill 框架:
- 数据分析与效率:
- AppStore 评论分析工具:@vista8 开源的工具,可自动抓取 App Store 评论并通过 DeepSeek 做情感和产品机会分析。 (GitHub)
- Sitedata 插件:由 @AI_Jasonyu 群友分享,可通过 Google AdSense ID 反向挖掘同一广告主下的其他网站。
- macOS 效率工具:@Pluvio9yte 推荐了 Maccy (开源剪贴板)、Mos (鼠标滚轮平滑)、Screen Studio (录屏) 等工具。
- 其他酷炫应用:
- YouMind 1.0:@vista8 和 @gefei55 共同推荐的创作工具,适合 PPT 和脑图制作。
- 世界杯观赛日历:@vista8 用 Codex + Goal Skill 仅耗时 24 分钟开发完成,支持个性化赛程日历订阅。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 32 条更新
a16z Podcast (A_full) 链接到标题
- Ideogram’s Open-Weights Image Model and the Future of AI Design
- 发布时间:2026-06-15 23:40 北京时间
- 摘要:- Yoko Li 和 Justine Moore 与 Ideogram 创始人兼首席执行官 Mohammad Norouzi 讨论图像生成模型、设计工作流程以及人工智能与创意工作之间不断发展的关系。
- 对话涵盖了 Ideogram 发布开放权重模型的决定、在图像中生成文本和布局的挑战,以及为什么可控性已成为越来越重要的研究领域。
- 他们讨论提示、定制、编辑以及通用模型和针对特定创意任务优化的系统之间的权衡。
- 在此过程中,Norouzi 分享了他对开源人工智能、设计工具、代理工作流程以及随着创作者和企业寻求对其输出的更大控制而图像生成模型如何发展的看法。
- 在 Apple 播客上收听 a16z 播客:。
- EN 要点:
- Yoko Li and Justine Moore speak with Ideogram founder and CEO Mohammad Norouzi about image generation models, design workflows, and the evolving relationship be…
- The conversation covers Ideogram’s decision to release an open-weight model, the challenges of generating text and layouts within images, and why controllabilit…
- They discuss prompting, customization, editing, and the tradeoffs between general-purpose models and systems optimized for specific creative tasks
- Along the way, Norouzi shares his views on open-source AI, design tools, agentic workflows, and how image generation models may evolve as creators and enterpris…
Stratechery by Ben Thompson (A_full) 链接到标题
- Anthropic’s Safety Superpower
- 发布时间:2026-06-15 18:00 北京时间
- 摘要:- 听这个帖子**:**。
- 我对那些愤世嫉俗的人表示同情,他们一贯将 Anthropic 的公开声明(尤其是围绕其模型发布的声明)描述为出于营销目的而散布恐慌。
- 就在两个月前,Anthropic 宣布了 Mythos Preview,他们认为该模型太危险,无法公开,特别是由于其先进的网络安全功能。
- 然后,两个月后,该公司公开发布了《Fable》,这是带有各种安全护栏的 Mythos 版本。
- 以我有限的经验来看,《神鬼寓言》是一个非常令人印象深刻的模型。
- EN 要点:
- Listen to this post :
- Log in to listen
- I’m sympathetic to the cynics who consistently characterize Anthropic’s public statements, particularly those surrounding their model releases, as scare-mongeri…
- It was only two months ago that Anthropic announced Mythos Preview, a model that they said was too dangerous to make publicly available, thanks in particular to…
ArXiv cs.AI (B_intro+search) 链接到标题
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13682v1 公告类型:新。
-摘要:开放车间调度问题(OSSP)出现在许多工业和服务环境中,但随着作业和机器数量的增加,计算仍然具有挑战性。
- 虽然精确的方法很快就会变得棘手,但经典的调度规则和元启发法可能需要大量调整才能维持大规模的解决方案质量。
- 这项研究使用具有多头注意力的编码器-解码器架构,为 OSSP 开发了一种基于 Transformer 的调度策略。
- EN 要点:
- arXiv:2606.13682v1 Announce Type: new
- Abstract: The open shop scheduling problem (OSSP) arises in many industrial and service settings but remains computationally challenging as the number of jobs a…
- While exact methods quickly become intractable, classical dispatching rules and metaheuristics may require substantial tuning to maintain solution quality at la…
- This study develops a Transformer-based scheduling policy for OSSP using an encoder-decoder architecture with multi-head attention
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13683v1 公告类型:新。
-摘要:为了解决当前对话策略规划方法难以动态适应不同用户特征的挑战,本文提出了一种具有大型语言模型的基于用户画像的嵌套推出策略适应(UP-NRPA)在线框架。
- 与依赖模型训练并需要针对用户组的离线强化学习策略模型的传统方法相比,UP-NRPA 通过自适应机制实现对话策略的动态定制。
- 这是通过利用实时用户反馈以及从当前用户画像映射的个性、偏好和目标来实现的,从而无需离线强化学习即可适应用户特征。
- EN 要点:
- arXiv:2606.13683v1 Announce Type: new
- Abstract: To address the challenge that current dialogue policy planning methods struggle to dynamically adapt to diverse user characteristics, this paper propo…
- In contrast to conventional approaches dependent on model training and require offline reinforcement learning policy models for user groups, UP-NRPA enables dyn…
- This is achieved by leveraging real-time user feedback alongside personality, preferences, and objectives mapped from the current user portrait, thereby adaptin…
History of the Muddy Children Puzzle
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13703v1 公告类型:新。
- 摘要:泥泞儿童之谜是一个关于知识和无知的谜题,一直激励着认知逻辑的发展。
- 我们通过过去两个世纪的逻辑和文学出版物追溯了泥泞儿童拼图的起源。
- 这个谜题激发了许多变化,例如涉及数字或彩色帽子。
- EN 要点:
- arXiv:2606.13703v1 Announce Type: new
- Abstract: The Muddy Children Puzzle is a puzzle about knowledge and ignorance that has been inspiring for the development of epistemic logic
- Who came up with it first
- This is unclear
Orchestra-o1: Omnimodal Agent Orchestration
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13707v1 公告类型:新。
-摘要:代理群最近的成功将基于大语言模型(LLM)的代理范式从单代理工作流程转变为多代理系统,凸显了代理编排对于任务分解和协作的重要性。
- 然而,现有的编排框架仅限于一组狭窄的模式,并且很难推广到异构模式共存和交互的更复杂的环境。
- 这种限制在全模态场景中变得尤为明显,其中任务需要对文本、图像、音频和视频等不同输入进行统一理解和协调。
- EN 要点:
- arXiv:2606.13707v1 Announce Type: new
- Abstract: The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-agent workflows to multi-agent syst…
- However, existing orchestration frameworks are limited to a narrow set of modalities and struggle to generalize to more complex settings where heterogeneous mod…
- This limitation becomes particularly pronounced in omnimodal scenarios, where tasks require the unified understanding and coordination of diverse inputs such as…
Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13710v1 公告类型:新。
- 摘要:深入研究和代理进化是人工智能代理在通用人工智能的现实应用中的实际任务。
- 前者能够在开放环境中自主检索和集成信息,以解决开放式研究任务,但它受到代理系统静态参数化深度研究能力的限制。
- 后者允许代理自主地与环境交互以获得发展模型功能的经验。
- EN 要点:
- arXiv:2606.13710v1 Announce Type: new
- Abstract: Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence
- The former enables autonomous retrieval and integration of information in open-ended environments to tackle open-ended research tasks, yet it is constrained by…
- The latter allows agents to autonomously interact with the environment to gain experiences that evolve model capabilities
WorkBench Revisited: Workplace Agents Two Years On
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13715v1 公告类型:新。
- 摘要:2024 年 3 月 WorkBench 上的最佳代理 GPT-4 完成了 43% 的任务,并对其中 26% 的任务采取了无意的有害操作,例如向错误的人发送电子邮件。
- 我们在 2026 年 6 月重新访问基准,发现迄今为止最好的代理 Claude Opus 4.8 完成了 89%,并对 2.5% 采取了意想不到的有害操作。
- 除了边境特工绩效的显着进步之外,还有三件事值得注意。
- EN 要点:
- arXiv:2606.13715v1 Announce Type: new
- Abstract: The best agent on WorkBench in March 2024, GPT-4, completed 43% of tasks and took an unintended harmful action, such as emailing the wrong person, on…
- We re-visit the benchmark in June 2026 and find that the best agent to date, Claude Opus 4.8, completes 89% and takes an unintended harmful action on 2.5%
- Aside from this considerable progress in frontier agent performance, three things stand out
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13720v1 公告类型:新。
- (2024) 表明,安全微调聊天模型中的拒绝是由残余流中的单个线性方向介导的,可通过有害和无害激活的均值差异 (DiM) 来恢复。
- 我们将基于 DiM 的干预措施(激活添加和定向消融)与源自迭代零空间投影(INLP)的两种干预措施(零空间投影和反事实翻转)在五个开放权重聊天模型上进行比较,询问 INLP 是否可以在转向拒绝方面与 DiM 相匹配,以及其更丰富的参数化是否会产生更多可调整的干预措施。
- INLP 反事实翻转在拒绝抑制方面与 DiM 定向消融具有竞争力,而零空间投影始终较弱。
- EN 要点:
- arXiv:2606.13720v1 Announce Type: new
- Abstract: Arditi et al
- (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual stream, recoverable by a difference-in-m…
- We compare DiM-based interventions (activation addition and directional ablation) with two interventions derived from Iterative Nullspace Projection (INLP) – n…
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13722v1 公告类型:新。
- 摘要:本文介绍了 YeasierAgent,一种基于共生代理、叙事世界和场景感知交互的应用程序构建范例。
- 它通过将应用程序重新定义为用户、代理和世界之间的协作空间,挑战了传统的设备耦合软件模型。
- 我们提出了一个系统架构,它实现了两个主要贡献:(1)通过利用与平台无关的交互单元(代理、场景、对话)而不是固定的图形布局,能够快速、跨平台地构建代理本机应用程序; (2)将智能代理的情感陪伴和实用工具执行属性统一在单个体验沙箱中。
- EN 要点:
- arXiv:2606.13722v1 Announce Type: new
- Abstract: This paper introduces YeasierAgent, an application-building paradigm based on symbiotic agents, narrative worlds, and scene-aware interaction
- It challenges the conventional device-coupled model of software by redefining applications as collaborative spaces among users, agents, and worlds
- We present a system architecture that achieves two primary contributions: (1) enabling the rapid, cross-platform construction of agent-native applications by ut…
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13731v1 公告类型:新。
- 摘要:商业智能 (BI) 越来越多地将仪表板交互与基于 LLM 的帮助相结合,但这两种模式在多步骤分析过程中常常不同步。
- 当用户在直接仪表板操作和自然语言查询之间切换时,跨过滤器、层次结构、指标和图表上下文保持一致的分析状态变得困难。
- 我们推出 TwinBI,这是一个代理数字孪生框架,它将基于 LLM 的代理系统与可执行的 BI 仪表板状态相结合。
- EN 要点:
- arXiv:2606.13731v1 Announce Type: new
- Abstract: Business intelligence (BI) increasingly combines dashboard interaction with LLM-based assistance, but these two modes often fall out of sync during mu…
- As users switch between direct dashboard manipulation and natural-language queries, it becomes difficult to preserve a consistent analytical state across filter…
- We present TwinBI, an agentic digital-twin framework that couples an LLM-based agent system with an executable BI dashboard state
When Sample Selection Bias Precipitates Model Collapse
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13732v1 公告类型:新。
- 摘要:对合成数据进行递归训练的激增可以缓解数据稀缺性,但存在模型崩溃的风险,其中重复训练会侵蚀分布尾部并使输出均质化。
- 数据选择被广泛视为一种补救措施,但其可靠性关键取决于验证者使用的参考分布。
- 我们表明,在低资源验证机制中,每个验证者仅观察目标流形的一小部分、碎片化且有偏差的部分,选择本身就会产生偏差。
- EN 要点:
- arXiv:2606.13732v1 Announce Type: new
- Abstract: The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes distrib…
- Data selection is widely viewed as a remedy, yet its reliability depends critically on the reference distribution used by the verifier
- We show that in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection its…
ArXiv cs.CL (B_intro+search) 链接到标题
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13685v1 公告类型:新。
-摘要:LLM-as-a-Judge 现在被广泛用于对模型输出进行排名、训练奖励模型以及填充公共排行榜,但其运行间的可靠性仍然未被充分表征。
- 我们使用两个 OpenAI 判断模型(GPT-4o-mini 和 GPT-4.1-mini)研究对跨越 10 个类别的 29 项任务的重复相同评估,每个问题进行 50 次配对试验和 50 次逐点试验,并辅以温度和提示敏感性消融。
- 在评委中,成对偏好的翻转率平均为 13.6%,其中 28% 的问题翻转率超过 20%,其中一个问题达到 56%。
- EN 要点:
- arXiv:2606.13685v1 Announce Type: new
- Abstract: LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains…
- We study repeated identical evaluations on 29 tasks spanning 10 categories using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini), with 50 pairwise trials…
- Across judges, pairwise preferences flip on average 13.6% of the time, with 28% of questions exceeding a 20% flip rate and one question reaching 56%
Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13686v1 公告类型:新。
- 摘要:随着越来越多的自主网络代理被部署来执行现实世界的任务,确保其安全已成为一个关键问题。
- 在这项工作中,我们研究电子商务领域中真实欺骗性界面下的网络代理行为。
- 我们引入了 WebDecept,这是一个轻量级且可配置的插件框架,可以将欺骗性界面模式受控地注入到现有的 Web 环境中。
- EN 要点:
- arXiv:2606.13686v1 Announce Type: new
- Abstract: As autonomous web agents are increasingly deployed to perform real-world tasks, ensuring their safety has become a critical concern
- In this work, we study web agent behavior under realistic deceptive interfaces in the e-commerce domain
- We introduce WebDecept, a lightweight and configurable plugin framework that enables controlled injection of deceptive interface patterns into existing web envi…
Which Models Perform Better in Inheritance Reasoning?
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13751v1 公告类型:新。
- 摘要:本文介绍了 PSL 团队参与 QIAS 2026 阿拉伯伊斯兰继承推理共享任务的情况。
- 该任务评估大型语言模型解决需要法律解释、多步推理和精确数值计算的继承案件的能力。
- 我们在统一的提示策略下比较 \textit{commercial} 和 \textit{open-source} 模型,以评估它们在结构化法律推理中的有效性,并尽可能减少特定任务的适应。
- EN 要点:
- arXiv:2606.13751v1 Announce Type: new
- Abstract: This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning
- The task evaluates the ability of large language models to solve inheritance cases that require legal interpretation, multi-step reasoning, and precise numerica…
- We compare \textit{commercial} and \textit{open-source} models under a unified prompting strategy to assess their effectiveness in structured legal reasoning wi…
QIAS 2026: Overview of the Shared Task on Islamic Inheritance Reasoning
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13756v1 公告类型:新。
- 摘要:本文全面概述了 QIAS 2026 共享任务,该任务是 OSACT7 研讨会的一部分,与 LREC 2026 同期举办。
- 共享任务旨在评估大型语言模型在伊斯兰继承的宗教和法律领域执行复杂推理的能力。
- 与传统问答基准不同,QIAS 2026 侧重于自然语言案例的端到端推理,要求系统执行完整的继承计算过程,从识别合格继承人到为每个受益人分配正确的份额。
- EN 要点:
- arXiv:2606.13756v1 Announce Type: new
- Abstract: This paper presents a comprehensive overview of the QIAS 2026 shared task, organized as part of the OSACT7 Workshop and co-located with LREC 2026
- The shared task was designed to evaluate the ability of large language models to perform complex reasoning in the religious and legal domain of Islamic inherita…
- Unlike conventional question-answering benchmarks, QIAS 2026 focuses on end-to-end reasoning from natural language cases, requiring systems to perform the full…
The Culture Funnel: You Can’t Align What isn’t in the Data
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13808v1 公告类型:新。
- 摘要:当前的文化对齐方法侧重于推理时间干预,假设模型已经包含足够的文化知识。
- 我们认为现代法学硕士管道受到文化数据漏斗的影响。
- 使用跨预训练、微调、对齐和推理数据集的多维标记框架,我们发现显性文化信号在训练后急剧下降,而地理上集中的任务专用数据占主导地位。
- EN 要点:
- arXiv:2606.13808v1 Announce Type: new
- Abstract: Current cultural alignment approaches focus on inference-time interventions, assuming models already contain sufficient cultural knowledge
- We argue modern LLM pipelines suffer from a cultural data funnel
- Using a multidimensional tagging framework across pretraining, fine-tuning, alignment, and reasoning datasets, we show explicit cultural signals decline sharply…
When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13835v1 公告类型:新。
-摘要:基于法学硕士的生成代理越来越多地用于城市模拟器,但目前尚不清楚它们是再现经验上现实的人类流动模式还是仅仅生成合理的流动叙述。
- 我们引入了一个验证框架,用于根据现实世界的移动数据评估基于法学硕士的城市模拟器的生成代理的移动性。
- 为此,我们使用移动法则、时间节律、网络主题、语义活动转换和行为移动配置文件。
- EN 要点:
- arXiv:2606.13835v1 Announce Type: new
- Abstract: LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobil…
- We introduce a validation framework for evaluating the mobility of generative agents of LLM-based urban simulators against real-world mobility data
- For this, we use mobility laws, temporal rhythms, network motifs, semantic activity transitions, and behavioral mobility profiles
Hybrid Classical-Quantum Variational Autoencoder for Neural Topic Modeling
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13852v1 公告类型:新。
- 摘要:神经主题模型能够实现可扩展的语义发现,但它们与量子硬件的集成在很大程度上仍未得到探索。
- 我们提出了一种用于主题建模的概念验证混合经典量子变分自动编码器(VAE),将参数化量子电路嵌入 VAE 推理网络中,同时保留经典主题词解码器。
- 为了解决量子硬件的资源限制,我们提出了一种改进的高斯 Softmax 后验,它将潜在空间维度与要提取的主题数量解耦,使模型能够在低资源 10 量子位量子设备上运行。
- EN 要点:
- arXiv:2606.13852v1 Announce Type: new
- Abstract: Neural topic models enable scalable semantic discovery, but their integration with quantum hardware remains largely unexplored
- We present a proof-of-concept hybrid classical-quantum variational autoencoder (VAE) for topic modeling, embedding parameterized quantum circuits within the VAE…
- To address the resource constraints of quantum hardware, we propose a modified Gaussian Softmax posterior that decouples latent space dimensionality from the nu…
SANA: What Matters for QA Agents over Massive Data Lakes?
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13904v1 公告类型:新。
- 摘要:数据湖上的探索性问答 (EQA) 需要 LLM 代理发现相关来源、分析检索到的数据并根据中间结果调整其操作。
- 端到端的准确性本身无法区分搜索、规划、数据分析或代理的行动策略中的失败:它决定下一步做什么以及何时提交答案。
- 我们提出了 SANA(搜索代理导航消融框架),这是一种诊断消融框架,可将 EQA 任务转换为包含黄金源序列、经过净化的子问题和执行记录的运行时配置文件。
- EN 要点:
- arXiv:2606.13904v1 Announce Type: new
- Abstract: Exploratory question answering (EQA) over data lakes requires an LLM agent to discover relevant sources, analyze retrieved data, and adapt its actions…
- End-to-end accuracy alone cannot distinguish failures in search, planning, data analysis, or the agent’s Action Policy: its decisions about what to do next and…
- We present SANA (Search Agent Navigation Ablation framework), a diagnostic ablation framework that transforms EQA tasks into runtime profiles containing gold so…
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13931v1 公告类型:新。
- 摘要:律师与委托人咨询是法律服务的重要起点。
- 有效的法律援助取决于从客户那里获取充分和真实的信息,以便制定最能保护其利益的策略。
- 这项任务需要大型语言模型(LLM)不仅能够执行稳健的法律推理,而且还能够通过多轮交互战略性地引出重要事实,并有效地指导具有不同性格的客户。
- EN 要点:
- arXiv:2606.13931v1 Announce Type: new
- Abstract: Lawyer-client consultation is a critical starting point for legal services
- Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect their interests
- This task requires Large Language Models (LLMs) not only to perform robust legal reasoning, but also to strategically elicit material facts through multi-turn i…
Can Post-Training Turn LLMs into Good Medical Coders? An Empirical Study of Generative ICD Coding
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13940v1 公告类型:新。
- 摘要:自动国际疾病分类 (ICD) 编码是计费、流行病学和临床决策支持的核心医疗编码任务。
- 生成式大语言模型 (LLM) 通常被认为是较弱的医学编码器,但这一发现主要来自推理时间设置,例如提示、检索、重新排序或工具使用,从而导致特定任务的训练后的作用尚未得到充分探索。
- 我们提出了一项针对生成式 ICD 编码训练后的受控实证研究,将区分基线与 LLM 编码员在通用协议和指标集下的提示、监督微调和强化学习方面进行比较。
- EN 要点:
- arXiv:2606.13940v1 Announce Type: new
- Abstract: Automated International Classification of Diseases (ICD) coding is a core medical-coding task for billing, epidemiology, and clinical decision support
- Generative large language models (LLMs) are often reported as weak medical coders, but this finding mainly comes from inference-time settings such as prompting,…
- We present a controlled empirical study of post-training for generative ICD coding, comparing discriminative baselines with LLM coders across prompting, supervi…
ArXiv cs.LG (B_intro+search) 链接到标题
Can Editing 1 Neuron Fix Repetition Loops in LLMs?
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13705v1 公告类型:新。
- Gemma 4 指令调整模型有一个可重现的失败:在较长的事实枚举提示下,例如列出电视剧的每一集、88 个 IAU 星座或 151 个原始 Pokemon,它们会陷入重复,要么是严格的逐字循环,要么是条目衰减为单个答案的列表。
- 这些循环的发生率高达 95%,并且能够在及时的改写、推理引擎更改和大多数采样调整中幸存下来。
- 在本文中,我们探讨了这种行为是否足够本地化,可以通过权重编辑来删除。
- EN 要点:
- arXiv:2606.13705v1 Announce Type: new
- Abstract: Yes
- Can it cure doom loops
- Probably not
Efficient On-Device Diffusion LLM Inference with Mobile NPU
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13740v1 公告类型:新。
- 摘要:扩散大型语言模型 (dLLM) 通过并行对多个标记进行去噪来加速生成,这使得它们对于延迟敏感的移动推理具有吸引力。
- 然而,重复降噪会在智能手机上引入大量计算。
- 移动神经处理单元 (NPU) 提供高吞吐量密集矩阵计算,但有效利用它们仍然具有挑战性:令牌承诺缩小了每个块的有效工作负载,令牌修订使 KV 缓存重用变得复杂,有限的 NPU 可见地址空间会导致昂贵的重新映射和数据传输开销。
- EN 要点:
- arXiv:2606.13740v1 Announce Type: new
- Abstract: Diffusion large language models (dLLMs) accelerate generation by denoising multiple tokens in parallel, making them attractive for latency-sensitive m…
- However, repeated denoising introduces substantial computation on smartphones
- Mobile neural processing units (NPUs) offer high-throughput dense matrix computation, but efficiently exploiting them remains challenging: token commitment shri…
High-Frequency Pricing at Scale for E-Commerce
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13741v1 公告类型:新。
- 摘要:本文介绍了一种专门用于时尚电子商务销售活动的预测然后优化算法定价工具的设计、开发和实现。
- 销售活动给定价带来了独特的挑战,包括不稳定的需求模式、快速的定价决策以及平衡短期收入与长期盈利能力的需要。
- 我们描述了我们的方法,该方法将使用梯度提升树的每日分辨率需求预测与多目标优化框架相结合,该框架可最大限度地提高超过 500 万件商品的长期利润和净商品价值。
- EN 要点:
- arXiv:2606.13741v1 Announce Type: new
- Abstract: This paper presents the design, development, and implementation of a specialized forecast-then-optimize algorithmic pricing tool for sales campaigns i…
- Sales events present unique challenges for pricing including volatile demand patterns, rapid pricing decisions, and the need to balance short-term revenue with…
- We describe our approach combining daily-resolution demand forecasting using gradient-boosted trees with a multi-objective optimization framework that maximizes…
A fully GPU-based workflow for building physics emulators of hypersonic flows
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13742v1 公告类型:新。
- 摘要:以高保真度和低计算成本解决复杂物理现象的能力是解决现代工程关键挑战的核心。
- 一个典型的例子是高超音速流,其中完整流场拓扑的精确预测,特别是关于冲击波位置和强度的预测至关重要。
- 然而,超音速和高超音速流仍然是传统降阶模型和神经模拟器的绊脚石,这些模型和神经模拟器很难在工业相关应用中捕获具有物理一致性的流动状态的陡峭梯度。
- EN 要点:
- arXiv:2606.13742v1 Announce Type: new
- Abstract: The ability to resolve complex physical phenomena with high fidelity and at low computational cost is central to addressing key challenges in modern e…
- A prime example lies in hypersonic flows, where the precise prediction of the full flowfield topology, in particular with respect to shock wave location and int…
- Yet supersonic and hypersonic flows continue to be a stumbling block for traditional reduced-order models and neural emulators that struggle to capture steep gr…
FedSPC: Shared Parameter Correction for Personalized Federated Learning
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13748v1 公告类型:新。
-摘要:个性化联邦学习(PFL)是联邦学习中解决统计异质性同时实现客户特定适应的重要方法之一。
- 许多 PFL 方法将模型分为共享参数和个性化参数,这些参数在每个客户端上联合训练。
- 然而,这会产生一个优化问题:共享参数由优化不同本地目标的客户端更新,这可能导致共享更新不一致并削弱共享表示。
- EN 要点:
- arXiv:2606.13748v1 Announce Type: new
- Abstract: Personalized federated learning (PFL) is one of the important approaches in federated learning for addressing statistical heterogeneity while enabling…
- Many PFL methods split the model into shared and personalized parameters, which are jointly trained on each client
- However, this creates an optimization issue: shared parameters are updated by clients optimizing different local objectives, which can lead to inconsistent shar…
The Weight Norm Sets the Grokking Timescale: A Causal Delay Law
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13753v1 公告类型:新。
- 摘要:Grokking 是神经网络泛化的延迟开始,在神经网络拟合训练数据很久之后才出现。
- 体重标准是否会导致这种延迟存在争议:一些研究报告了过渡时的关键标准,而另一些研究则观察到摸索时根本没有固定的标准。
- 我们通过在训练期间干预规范而不是仅仅遵守规范来解决这个问题。
- EN 要点:
- arXiv:2606.13753v1 Announce Type: new
- Abstract: Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data
- Whether the weight norm causes this delay is disputed: some studies report a critical norm at the transition, others observe grokking with no fixed norm at all
- We settle this by intervening on the norm during training rather than only observing it
D2H-AD: A Hybrid Model Utilizing Hyperdimensional Computing for Advanced Anomaly Detection
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13754v1 公告类型:新。
- 摘要:异常检测是智能系统的基本组成部分,应用于医疗保健、网络安全、智能电网和物联网环境。
- 尽管传统的机器学习和深度学习方法已证明在识别异常方面有效,但它们通常依赖于大型标记数据集,会产生高昂的计算成本,并且在边缘和高维设置中面临可扩展性挑战。
- 本文提出了 D2H-AD,一种基于超维计算 (HDC) 的新型异常检测框架,超维计算是一种使用高维分布式向量表示信息的大脑启发范例。
- EN 要点:
- arXiv:2606.13754v1 Announce Type: new
- Abstract: Anomaly detection is a fundamental component of intelligent systems with applications in healthcare, cybersecurity, smart grids, and IoT environments
- Although conventional machine learning and deep learning methods have demonstrated effectiveness in identifying anomalies, they often rely on large labeled data…
- This paper presents D2H-AD, a novel anomaly detection framework based on Hyperdimensional Computing (HDC), a brain-inspired paradigm that represents information…
Beyond LoRA: Is Sparsity-Induced Adaptation Better?
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13767v1 公告类型:新。
- 摘要:低秩适应(LoRA)及其变体为预训练模型的全面微调提供了一种内存和计算效率高的替代方案。
- 然而,关于这些方法的相对普遍性以及对低秩更新的结构限制如何保持有效的适应性能仍然存在问题。
- 我们提出了一个历史框架,涵盖过去(完全微调和原始 LoRA)、现在(LoRA 的不同变体),并通过在现有 LoRA 变体中引入稀疏性来提出更简单、更便宜、参数高效的扩展:便宜的 LoRA (cLA),用另一个固定的(确定性的或随机变体中的随机变体)和链式循环变体 ${c}^3$LA 训练单个低秩因子。
- EN 要点:
- arXiv:2606.13767v1 Announce Type: new
- Abstract: Low-rank adaptation (LoRA) and its variants provide a memory- and compute-efficient alternative to full fine-tuning of pre-trained models
- However, questions remain about the comparative generalizability of these approaches and how the structural restrictions on low-rank updates preserve effective…
- We present a historical framing, covering the past (full fine-tuning and original LoRA), the present (different variants of LoRA), and propose simpler, cheaper,…
Diffusion Policy Optimization without Drifting Apart
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13795v1 公告类型:新。
- 摘要:RL 后训练对于改进扩散策略变得越来越关键,但现有的扩散策略梯度方法往往不稳定,无法实现可靠的策略改进。
- 我们将原因确定为双漂移现象:优化变分代理可以让 ELBO 与真实的对数似然分离,从而使生成的代理策略梯度与预期回报的真实策略梯度不一致。
- 我们提出 \textbf{DiPOD},一种扩散策略优化框架,通过将自蒸馏与策略改进梯度更新交织在一起,在整个训练过程中保持紧束缚行为。
- EN 要点:
- arXiv:2606.13795v1 Announce Type: new
- Abstract: RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable a…
- We identify the cause as the double-drift phenomenon: optimizing a variational surrogate can let the ELBO separate from the true log-likelihood, which then make…
- We propose \textbf{DiPOD}, a diffusion policy optimization framework that maintains tight-bound behavior throughout training by interleaving self-distillation w…
Neural Variability Enhances Artificial Network Robustness
- 发布时间:2026-06-15 12:00 北京时间
- 摘要:- arXiv:2606.13801v1 公告类型:新。
- 摘要:皮层的神经反应在对重复刺激的反应中表现出显着的试验变异性,而周围感觉神经元的反应则更加一致,这导致许多人怀疑随机性是否可能具有意义。
- 现有的工作认为,噪声和信号相关性可以针对动物的区分进行优化,而人工神经网络 (ANN) 研究表明,噪声在机器学习任务中具有类似的好处,尽管大多数 ANN 工作都忽略了相关性的影响。
- 在这里,我们研究相关噪声是否提高了人工神经网络对对抗性攻击和自然图像修改的鲁棒性。
- EN 要点:
- arXiv:2606.13801v1 Announce Type: new
- Abstract: Neural responses in cortex exhibit substantial trial-to-trial variability in response to repeated stimuli, while peripheral sensory neurons respond fa…
- Existing work has argued that noise and signal correlations may be optimized for discrimination in animals, whereas artificial neural network (ANN) studies have…
- Here we investigate whether correlated noise improves the robustness of artificial neural networks to adversarial attacks and naturalistic image modifications