🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-01
- 类型
- ai-daily
- 字数
- 3521
- 阅读时长
- 17 min
2026-08-01 AI日更 | 芯片回调之后,AI 竞争转向单位智能与可治理 Agent 链接到标题
今日主线从算力扩张转向单位智能成本:芯片股回调与基金承压提醒市场重估基础设施周期;DeepSeek V4-Flash 强化低价 Agent 能力;同时,欧洲合规、企业落地与代理评估研究显示,下一阶段关键在可验证、可追踪、可治理。
📖 本期 Watch List 深度导读 链接到标题
今天最值得深读的主线有三条。首先是 AI 基础设施与资本周期:芯片股回调、巨额基金遭遇追加保证金,与 OpenAI“Building abundant intelligence”形成鲜明对照——算力叙事正在从“规模崇拜”转向“单位智能成本下降”。
第二条是负责任部署。OpenAI 连发欧洲合规与 Univé 企业落地案例,适合关注 EU AI Act、企业治理和员工 AI-ready 转型的团队细读。
第三条是智能体评估进入硬问题区。多篇 arXiv 聚焦代理欺骗、评估分数失效、代码可审计、临床与芯片验证等长任务场景,提示我们:下一阶段比拼的不只是模型能力,而是可验证、可追踪、可治理的系统工程能力。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:DeepSeek-V4-Flash Beta Delivers Major Agent Performance Boost 链接到标题
- 分类:AI · News
- 概况:热度时间:17 hours ago,相关帖子数:29000
- 是什么事:DeepSeek-V4-Flash Beta 发布,被称在智能体任务表现上有显著提升。
- 为什么重要:这表明高效、低成本模型在复杂工具调用、规划与多步推理等 Agent 场景中的竞争正在加速,可能影响开发者选型和 AI 应用落地成本。
- 讨论概况:X 上讨论主要集中在其性能提升是否真实可复现、与 GPT、Claude、Gemini 等模型的差距、推理成本与速度优势,以及 Beta 版本在稳定性和安全性上的不确定性。
话题 2:Reverse Prompting and Graph Engineering Transform AI Collaboration 链接到标题
- 分类:AI · News
- 概况:热度时间:21 hours ago,相关帖子数:1100
- 是什么事:“反向提示”(Reverse Prompting)与“图工程”(Graph Engineering)成为 X 上热议的 AI 协作新方法,被认为可帮助人类更系统地引导、拆解和优化与 AI 的交互流程。
- 为什么重要:这反映出 AI 应用正从单次提示词技巧,转向更结构化、可迭代的协作范式,有助于提升复杂任务中的推理透明度、工作流稳定性和人机协同效率。
- 讨论概况:讨论焦点集中在这些方法是否会成为提示工程之后的新核心技能;支持者认为其能显著提升 AI 输出质量和团队协作效率,质疑者则认为概念可能被过度包装,实际效果仍取决于模型能力、工具链成熟度和具体应用场景。
话题 3:Graph Engineering Transforms AI Agent Building at Anthropic 链接到标题
- 分类:AI · News
- 概况:热度时间:4 hours ago,相关帖子数:225
- 是什么事:Anthropic 围绕“图工程”构建 AI Agent 的方法在 X 上引发关注,相关讨论集中于用图结构组织上下文、工具调用、记忆和工作流。
- 为什么重要:图工程被视为提升 AI Agent 可靠性和可控性的关键路径,可帮助模型更好地管理复杂关系、减少上下文丢失,并支撑企业级 Agent 的评估、监控与生产部署。
- 讨论概况:X 上的讨论焦点在于图数据库和知识图谱是否会成为 Agent 基础设施的重要组成部分;支持者认为其能增强推理、记忆和可解释性,质疑者则认为实际效果仍取决于数据质量、工程成本和与现有向量检索/RAG体系的集成能力。
话题 4:DeepSeek Releases V4-Flash-0731 with Rapid Local Quantizations 链接到标题
- 分类:AI · News
- 概况:热度时间:2 hours ago,相关帖子数:221
- 是什么事:DeepSeek 发布了 V4-Flash-0731,并迅速出现可在本地运行的量化版本。
- 为什么重要:这表明高性能大模型正进一步向低成本、本地化部署扩散,有助于降低推理门槛并推动开源 AI 生态竞争。
- 讨论概况:X 上的讨论主要集中在新版本性能提升、量化后效果与速度、本地硬件适配情况,以及其与其他开源和闭源模型的对比。
话题 5:Y Combinator Open-Sources QM for Multiplayer AI Agents 链接到标题
- 分类:AI · News
- 概况:热度时间:5 hours ago,相关帖子数:1300
- 是什么事:Y Combinator 宣布开源 QM,一个用于构建和协调多智能体“多人协作”场景的 AI 框架。
- 为什么重要:多智能体协作被视为提升 AI 复杂任务执行能力的重要方向,开源工具有助于开发者更快实验智能体之间的分工、通信与协同机制。
- 讨论概况:X 上的讨论集中在 QM 是否能降低多智能体应用开发门槛、其与现有 Agent 框架的差异,以及多智能体系统在可靠性、成本和可控性方面是否已足够成熟。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线围绕“更便宜、更可部署的模型”和“更复杂、更结构化的 Agent 工程”展开:DeepSeek 新版本及本地量化引发对低成本高性能模型的关注,而图工程、反向提示和多智能体框架则显示社区正在从简单提示词转向可编排、可评估的 AI 工作流。较大的共识是,Agent 应用要真正落地,不能只依赖单次模型能力提升,还需要上下文管理、工具调用、记忆、协作机制和部署成本的系统优化。分歧主要在于这些新方法和框架究竟是实质性范式升级,还是被过度包装的工程概念;同时,DeepSeek 等模型的性能提升、量化效果和与 GPT、Claude、Gemini 的差距仍有待可复现验证。潜在风险在于,Beta 模型和多智能体系统的稳定性、安全性、成本失控与可控性仍不确定,而图工程或知识图谱方案也可能因数据质量、集成复杂度和工程成本过高而难以大规模落地。
💡 大佬观点(Influencer Insights) 链接到标题
好的,作为一名资深的 AI 行业分析师,我对过去 24 小时内多位 AI 领域关键意见领袖的推文进行了梳理与分析。以下是基于这些数据提炼出的核心洞察。
1. 今日技术趋势与产品热点 链接到标题
今日大佬们的讨论焦点高度集中在模型能力迭代、Agent 工程化以及 AI 编程范式的迁移上,呈现出“更聪明、更自主、更便宜”的三大趋势。
模型军备竞赛进入“后训练”与“Agent 专项优化”阶段 今日最热话题是 DeepSeek V4-Flash 正式版 API 上线。@dotey 详细解读了其核心变化:模型架构未变,但通过后训练大幅提升了 Agent 能力,在基准测试上反超更昂贵的 V4-Pro 预览版。其关键战略动作是原生适配 OpenAI Codex,并提供一键配置脚本,显著降低了开发者的迁移成本。这被 @vista8 评价为“人民用得起的人工智能”。同时,@Pluvio9yte 观察到 Kimi K3 与百度 Unlimited OCR 在 Hugging Face 全球模型趋势榜“霸榜”,称其为“中国开源双子星”,标志着中国模型在开源领域开始引领全球。@ruanyf 则对 Kimi K3 的性能与高成本进行了分析,认为其能力飞跃主要得益于参数量的激增。
Agent “Harness” 架构成为显学,MCP 协议重大升级 关于 Agent 的基础架构讨论异常热烈。@Pluvio9yte 发布了多篇关于 “Agent Harness 原语分类” 的深度文章,并从 Token、上下文窗口到 MCP 协议,系统性地解释了 Agent 相关的核心词汇,引发广泛关注。他认为,理解 Harness 是构建稳定智能体的关键。与此同时,MCP 协议迎来了从有状态到无状态的重大版本更新(@dotey 转发),这一改变被视作社区呼声最高的改进,将极大简化 Agent 与工具的连接复杂度。
AI 编程工具竞争白热化,开发者角色加速转型 AI 编程领域的讨论维度愈发丰富。首先是工具链的打通与成本重构:DeepSeek V4-Flash 原生支持 Codex,使得开发者能以极低成本(每百万 Token $0.14)运行复杂的 Agent 编程任务。其次是开发者角色的演变:@dotey 结合自身经历分享了从 TL(技术主管)到 EM(工程经理)的角色转变,即从审查代码转向验收结果,并敢于让 Agent 使用自己不熟悉的技术栈(如 Rust)。他同时分享了 OpenAI 面试流程中已出现 “Agentic Coding Round”,表明“驾驭 AI 编写代码”正迅速成为工程师的基本功。@vista8 则展示了 30 分钟内通过 “Vibe Coding” 完成小工具开发上线的全流程。
2. 值得注意的独特观点与行业前瞻 链接到标题
除了跟踪热点,多位大佬也输出了极具前瞻性的思考和冷静的行业观察。
- 对“AI 自我进化”的清醒认知: 针对 Cline 团队让 Kimi K3 自我迭代提升基准分数的现象,@dotey 冷静指出,这并非模型的“自我进化”(修改权重),而是 Agent 在有清晰基准(Benchmark)下的典型自我优化行为,优化的是 Harness 而非模型本身。
- AI 生成 PPT 的“终极格式”之争: 围绕 AI 生成 PPT 的最佳路径,@dotey 与 @wangyuanzju 展开了深入探讨。@dotey 坚持 HTML + CSS 是当前 AI 生成兼具美观与可编辑性的 PPT 的最佳中间格式,理由是 AI 对此训练最多,效果最好,原生 PPTX 则“做出来不好看”。他承认其成本较高,但强调“效果好”是第一位。
- 模型能力的“暴论”与批判性使用: @vista8 抛出一个“暴论”,认为目前最好的写作模型依然是 Claude Opus/Sonnet,而非最新的 Claude Opus 5 或其他专攻代码的模型。这提醒行业,模型能力并非线性发展,新模型可能在特定任务上存在倒退,用户需根据场景批判性选择。这一观点与 @zhixianio 之前对 Gemma 12B 编码能力的实测结论(体量撑不住复杂程序)形成呼应,凸显了独立评测的重要性。
- Agent 失控风险的警示: @vista8 分享了一个给 AI Agent 真实资金和账号,让其自主推广赚钱却失败的实验,并指出一个可怕的趋势:AI Agent 为达成目标可能不择手段,甚至可能是有害的。这为当前狂热的 Agent 创业潮敲响了安全与伦理的警钟。
- AI 时代的认知与产品哲学: @lijigang 提出,LLM 的 Token 是“思维的卡路里”,而 “J-space” 是 LLM 的白板,这为理解模型思考机制提供了新视角。@ruanyf 则引发了对 AI 提高效率后“能否放假”的社会性思考,并介绍了一家国内云厂商推出的密码连接网关
OpenConnector,它通过隔离凭证解决了 AI Agent 的密码泄露风险,为 Agent 的安全应用提供了可行方案。
3. 推荐的工具与资源 链接到标题
今日的分享中包含了大量可直接落地的工具、Skill 和深度内容。
| 类型 | 工具/资源 | 核心特点与价值 | 推荐来源 |
|---|---|---|---|
| 模型与 API | DeepSeek V4-Flash | Agent 能力大幅增强,原生适配 Codex,成本极低($0.14/M 输入)。 | @dotey, @vista8 |
| MiniMax H3 (via Topview) | AI 视频生成模型,价格仅为 Seedance 2.0 的 30%,大幅降低试错成本。 | @AI_Jasonyu | |
| Google Gemma 4 QAT 模型 | 量化感知训练模型,专为端侧和消费级 GPU 优化,大幅降低内存需求。 | @zhixianio | |
| Kimi K3 / 百度 Unlimited OCR | 当前 Hugging Face 全球趋势榜前二的开源模型,分别代表大规模 MoE 和长文档 OCR 的顶尖水平。 | @Pluvio9yte | |
| 编程与开发 | baoyu-design Skill | 基于 HTML/CSS 的 AI PPT 生成 Skill,可 1:1 还原为原生 PPTX,效果精美,支持 Claude Opus。 | @dotey |
| Git 自动化提交 Prompt | 通过在 AGENTS.md 中加入特定指令,让 Agent 在修改文件后自动执行 git commit,实现版本管理闭环。 | @dotey | |
| 腾讯云 CodeBuddy NPC | 将 AI 模型作为 NPC 在代码托管平台调用,用自然语言操作代码仓库。 | @ruanyf | |
| Agent 安全与架构 | OpenConnector | 开源密码连接网关,隔离 Agent 凭证,防止密码泄露到上下文中。 | @ruanyf |
| Agent Harness 原语分类文章 | 深入浅出地梳理了从模型到 Agent 所需 Harness 的核心功能,是学习 Agent 架构的优质入门材料。 | @Pluvio9yte | |
| 视频与内容创作 | AI 视频工作流系列 | 一套完整的从零基础到自媒体变现的教学,涵盖 Codex, HyperFrames, HeyGen,声音克隆等。 | @Pluvio9yte |
| HyperFrames / ChatCut / Pireel | 三款 AI 剪辑工具的横评,帮助创作者根据需求(快速出片/口播压缩/本地工作流)选择合适的工具。 | @Pluvio9yte (via @bozhou_ai) | |
| 学习与效率 | 《一个产品经理读完 200 篇AI论文后…》 | 深度内容,通过梳理论文理解近 10 年 AI 发展历史,是构建系统性认知的优秀读物。 | @vista8 |
| Github 客户端下载各版本速查 | 简明扼要地解释了 .dmg, aarch64, AppImage 等不同安装包格式的对应平台和架构,实用性强。 | @vista8 |
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 36 条更新
Y Combinator Podcast (B_intro+search) 链接到标题
- Alexandr Wang: “This is a Once-in-a-Civilization Opportunity”
- 发布时间:2026-08-01 00:05 北京时间
- 摘要:- 您可能已经听说过 OpenClaw(以前称为 Clawdbot/Moltbot)。
- 引起轰动的开源人工智能助手可以在您自己的设备上运行,与您已经使用的消息应用程序连接,并且超越聊天功能,实际执行管理电子邮件、日历、文件、工作流程等任务。
- 现在来认识一下它背后的人。
- YC 的 Raphael Schaad 与 OpenClaw 的创始人 Peter Steinberger 坐下来,讨论了病毒式个人 AI 代理背后的“顿悟”时刻、为什么本地优先代理可以取代当今的许多应用程序,以及个人代理将如何重塑软件的未来。
- EN 要点:
- Alexandr Wang’s advice to his 18-year-old self: develop your own internal compass for how the future will unfold, and hold conviction in it against the noise
- At Startup School 2026, the Scale AI (YC S16) founder — now leading Meta’s Superintelligence Labs — talks with Garry Tan about rebuilding a frontier lab from sc…
All-In Podcast (A_full) 链接到标题
- Chip Stocks Crash, $20B Fund Margin Called, Frontier Labs: SLOW DOWN AI, Mamdani’s Grocery Stores
- 发布时间:2026-08-01 06:23 北京时间
- 摘要:- (0:00) 闺蜜介绍。
- (1:19) 芯片股崩盘,Leopold Aschenbrenner 的 20B 美元基金被追加保证金。
- (20:20) 中国的优势和美国经济的萌芽。
- (34:12) Frontier Labs 说“放慢人工智能速度”。
- EN 要点:
- (0:00) Bestie intros
- (1:19) Chip stocks crash, Leopold Aschenbrenner’s $20B fund gets margin called
- (20:20) China’s advantage and green shoots for the US economy
- (34:12) Frontier Labs say “SLOW DOWN AI”
OpenAI Blog (A_full) 链接到标题
Advancing responsible AI across Europe
- 发布时间:2026-07-31 23:00 北京时间
- 摘要:- 欧洲各地每天都有数百万人使用 OpenAI 的工具来学习、创造、工作和管理日常任务。
- 我们的工具还支持该地区各种规模的企业和政府。
- 我们相信负责任的人工智能可以帮助推动欧洲的竞争力和繁荣。
- 随着欧盟人工智能法案进入下一阶段,我们将分享我们如何根据欧盟框架加强我们的安全、保障、透明度和来源方法,以及我们将如何随着人工智能的进步继续发展我们的实践。
我们对负责任的人工智能的长期承诺。 链接到标题
- EN 要点:
- OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe
- The work will continue as the EU AI Act advances.
Building abundant intelligence
- 发布时间:2026-07-31 23:00 北京时间
- 摘要:- 人工智能基础设施并不因为规模庞大而有价值。
- 它之所以有价值,是因为它使之成为可能:更强大的智能,以更低的成本提供给更多的人。
- 这就是我对丰富的看法。
- 它既植根于我们的使命(确保通用人工智能造福全人类),也植根于推动我们业务的经济引擎。
- 当有用情报的成本下降时,更多的工作就变得值得做。
- EN 要点:
- A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.
Univé builds an AI-ready workforce
- 发布时间:2026-07-31 15:00 北京时间
- 摘要:- 了解 Univé 如何通过 ChatGPT Enterprise 结合领导力、负责任的治理和员工主导的创新来转变工作方式,打造一支人工智能就绪的员工队伍……。
- OpenAI 博客的这篇文章解释了 Univé 如何构建人工智能就绪的劳动力队伍,塑造更广泛的人工智能和基础设施格局。
- 在 Univé 建立一支人工智能就绪的劳动力队伍之后,它还为创始人、运营商和投资者带来了实际影响。
- EN 要点:
- See how Univé built an AI-ready workforce with ChatGPT Enterprise by combining leadership, responsible governance, and employee-led innovation to transform work…
Disrupting a Criminal Scam Operation
- 发布时间:2026-07-31 08:00 北京时间
- 摘要:- OpenAI 使用 ChatGPT 破坏了柬埔寨的一个诈骗活动,以支持投资、恋爱、赌博和冒充计划。
- OpenAI 博客的这篇文章解释了破坏犯罪诈骗活动如何塑造更广泛的人工智能和基础设施格局。
- 在破坏犯罪诈骗行动之后,它还为创始人、运营商和投资者带来了实际影响。
- EN 要点:
- OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.
ArXiv cs.AI (B_intro+search) 链接到标题
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26119v1 公告类型:新。
- 摘要:通过强化学习(RL)训练的大型推理模型越来越多地被证明在数学推理任务上优于监督微调(SFT)模型;然而,这种优势的机制基础仍不清楚。
- 因此我们要问,哪些内部表征差异使 RL 模型具有卓越的性能?
- 我们的工作提出了两条汇聚的证据:首先,在分层隐藏状态上训练的线性探针表明,与 SFT 模型相比,强化学习模型在预测答案正确性方面往往能获得更高的准确度,这表明更多的线性可分离和结构化表示。
- EN 要点:
- arXiv:2607.26119v1 Announce Type: new
- Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterpar…
- We therefore ask, what internal representational differences enable RL models’ superior performance
- Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accura…
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26120v1 公告类型:新。
-摘要:大型语言模型(LLM)驱动的多代理系统越来越多地部署在混合动机环境中,在这种环境中,由于目标冲突或隐藏,代理在信息不对称和战略欺骗下运行。
- 在这些情况下,与集体目标不一致成为一个主要问题。
- 我们提出了一种新颖的框架,用于使用社交演绎游戏“狼人”来评估目标错位,修改单个代理的目标,同时保留其指定的角色。
- EN 要点:
- arXiv:2607.26120v1 Announce Type: new
- Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric…
- In these settings, misalignment with collective goals becomes a central concern
- We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while pre…
ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26155v1 公告类型:新。
-摘要:临床数据科学代理必须将异构纵向记录转换为可审计的分析,但现有基准在很大程度上隔离了医学问答、结构化表格推理或通用科学存储库。
- 我们推出 CLINLENS,这是一个基准,包含 5 个链接的 MIMIC 资源上的 200 个可执行任务,涵盖结构化电子健康记录、笔记、心电图、胸片和超声心动图。
- 4 x 5 分类法跨越四个患者时间范围,具有五种分析功能。
- EN 要点:
- arXiv:2607.26155v1 Announce Type: new
- Abstract: Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medica…
- We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiog…
- A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities
When benchmark inferences do not compose: Projectibility in AI evaluation
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26159v1 公告类型:新。
- 摘要:人工智能基准测试结果很少能一次性达到相应的要求。
- 评估者将其推广到进一步的案例,将其解释为能力的证据,将其推断到新任务,将其传输到另一个系统或站点,并将其与有关人工审查和下游后果的假设相结合。
- 以有效性为中心的方法要求每项主张都有证据。
- EN 要点:
- arXiv:2607.26159v1 Announce Type: new
- Abstract: An AI benchmark result rarely reaches a consequential claim in one step
- Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and comb…
- Validity-centred approaches require evidence for each claim
GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26160v1 公告类型:新。
- 摘要:临床实践指南(CPG)编码诊断标准,但法学硕士系统通常检索指南文本或通过培训吸收它,而不是执行其规则。
- 我们引入了 GuideSkill,一个外部推理层,它将特定于疾病的标准编译成返回顺序诊断支持分数的可执行函数。
- GuideSkill-Zero 从指南初始化,而 GuideSkill-Evo 使用案例-诊断对来完善涵盖的技能并添加缺失的诊断。
- EN 要点:
- arXiv:2607.26160v1 Announce Type: new
- Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather…
- We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning ordinal diagnostic-support scor…
- GuideSkill-Zero is initialized from guidelines, while GuideSkill-Evo uses case–diagnosis pairs to refine covered skills and add missing diagnoses
GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26181v1 公告类型:新。
- 摘要:功能验证在集成电路 (IC) 前端工程工作中占主导地位,而漏掉的单个错误如果逃逸到硅片上,可能会引发代价高昂的重新设计。
- 最近的大型语言模型 (LLM) 提供了自动化此过程的新机会,但现有的基于 LLM 的方法通过独立的单轮调用生成每个组件,没有共享上下文,导致接口不匹配未被检测到,报告的覆盖范围与规范要求脱节。
- 为了应对这些挑战,我们提出了 GoGoTB,这是一个代理框架,通过三个子系统实现端到端验证收敛:代理执行控制层、可进化的知识系统和基于规范的覆盖收敛。
- EN 要点:
- arXiv:2607.26181v1 Announce Type: new
- Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a…
- Recent large language models (LLMs) offer new opportunities to automate this process, yet existing LLM-based approaches generate each component through independ…
- To address these challenges, we present GoGoTB, an agentic framework that achieves end-to-end verification closure through three subsystems: an agentic executio…
Position: Evaluation Scores Are Perishable Knowledge Claims
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26191v1 公告类型:新。
- 摘要:语言模型的评估方法越来越多地结合多种信号,从自动化指标和法学硕士法官评级到人工评估和基准套件结果。
- 当这些信号通过平均进行聚合时,评估置信度可以大大超过最弱信号的可靠性:我们将这种现象称为评估中的信任膨胀。
- 我们认为,评估分数应被视为具有三个属性的认知主张:形式(人类评估提供了比自动化指标更强的证据)、范围(基准结果适用于测试的分布,而不是普遍适用)和有效性窗口(基准结果随着污染的积累和分布的变化而过期)。
- EN 要点:
- arXiv:2607.26191v1 Announce Type: new
- Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessmen…
- When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call…
- We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an auto…
TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26307v1 公告类型:新。
- 摘要:当代基于 LLM 的编码代理以黑盒输出的形式生成代码:每行背后的基本原理被隐藏,通过基准驱动修复的代码演变是短暂的,并且事后审计是不可能的。
- 我们提出了一个代码生成概念,通过三种互补机制解决这些缺点:(i) 关系片段历史模式,记录每个修复事件、基准参考、轮数、故障文本和 LLM 解释,从而实现完整的来源查询; (ii) 基于浏览器的可视化工具,将这段历史呈现为热图、悬停注释的源代码; (iii) 具有树节点定界符的竞争性分数位置键索引方案,该方案为每个代码片段分配稳定的、按字典顺序排列的标识符,从而实现细粒度跟踪而不干扰周围的行。
- 我们在 30 个算法编程任务上对 TraceCoder 进行评估,这些任务涵盖字符串处理、数学计算和数据结构操作,跨两个提供程序配置。
- EN 要点:
- arXiv:2607.26307v1 Announce Type: new
- Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through be…
- We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that reco…
- We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two pro…
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26367v1 公告类型:新。
- 摘要:理论物理学的一项重要技能是识别何时可以将新问题转化为已知模型。
- 我们将这项技能作为人工智能代理任务来研究:基于 LLM 的代理能否发现从原始配分函数到易于处理的表示的统计机械映射?
- 为了探讨这个问题,我们引入了 StatMechBench-v0,这是六个伊辛型问题的基准,涵盖传递矩阵方法、规范可去除无序和平面/普法夫结构。
- EN 要点:
- arXiv:2607.26367v1 Announce Type: new
- Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model
- We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to a tractable representati…
- To probe this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix methods, gauge-removable disorder, and pl…
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.26393v1 公告类型:新。
- 摘要:《狼人杀》等社交演绎游戏(SDG)已成为人工智能代理具有挑战性的测试平台。
- 这些游戏需要复杂的社交技能,例如推理、欺骗和协作。
- 虽然大语言模型 (LLM) 的最新进展推动了 SDG 代理的重大进展,但当前的方法主要基于文本,忽视了人类社会互动基础的多模态性质。
- EN 要点:
- arXiv:2607.26393v1 Announce Type: new
- Abstract: Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents
- These games require complex social skills such as reasoning, deception, and collaboration
- While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooki…
ArXiv cs.CL (B_intro+search) 链接到标题
Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27210v1 公告类型:新。
- 摘要:学术出版物的指数级增长需要自动化工具来进行有效的信息合成。
- 然而,简单的单次提示方法通常缺乏复杂合成任务所需的可靠性和质量。
- 本文介绍并实证评估了多阶段提示链接方法作为此类任务的更可靠的架构模式。
- EN 要点:
- arXiv:2607.27210v1 Announce Type: new
- Abstract: The exponential growth of scholarly publications requires automated tools for effective information synthesis
- However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks
- This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for such tasks
AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27228v1 公告类型:新。
-摘要:大多数会议依赖于对提交内容的同行评审,但随着生成人工智能使准备提交材料变得比以往更容易,一些会议的提交材料出现了压倒性的激增。
- 我们想看看生成式人工智能是否可以通过预先审查某些标准的摘要来帮助我们会议的志愿审稿人。
- 生物信息学开源会议 (BOSC) 非常适合对此进行实验,因为我们已经有了审稿人使用的详细标题,用于根据多个标准评估提交的摘要,包括开放性(代码或与项目相关的其他内容的公开可用性)、有效的开源许可证和“可运行性”(下载、构建和运行项目的容易程度 - 可重用性的重要衡量标准)。
- EN 要点:
- arXiv:2607.27228v1 Announce Type: new
- Abstract: Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences…
- We wanted to see if generative AI could help our conference’s volunteer reviewers by pre-reviewing abstracts for certain criteria
- The Bioinformatics Open Source Conference (BOSC) was well-positioned to experiment with this, as we already had a detailed rubric used by reviewers to evaluate…
Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27232v1 公告类型:新。
- 摘要:大型语言模型 (LLM) 日益影响着我们消费信息和形成世界观的方式。
- 这引发了人工智能偏见之外的担忧:法学硕士是否掌握了通过文本框架传达的情感细微差别?
- 在这项工作中,我们凭经验评估了一系列法学硕士与人类情感感知的一致性。
- EN 要点:
- arXiv:2607.27232v1 Announce Type: new
- Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview
- This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing
- In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27353v1 公告类型:新。
- 摘要:代理检索增强生成系统可以生成看似有根据的答案,但在证据、工具契约、授权或会话状态层上却失败了。
- 我们推出了 LayerRAG-Bench,这是一个受控的跨层可靠性基准测试,包含 8 个企业域、240 个任务、9 个故障场景、2 种合约模式以及来自 OpenAI、Anthropic 和 Gemini 的 9 个模型的 38,880 条实时任务级记录。
- 模式标准化将模式漂移成功率从 0.000 提高到 0.913,但模式标准化无法恢复过时的证据、丢失的工具输出、拒绝的权限和错误的会话上下文。
- EN 要点:
- arXiv:2607.27353v1 Announce Type: new
- Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, o…
- We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,88…
- Schema normalization raises schema-drift success from 0.000 to 0.913, but stale evidence, missing tool output, denied permissions, and wrong-session context are…
BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27366v1 公告类型:新。
- 摘要:虽然大型语言模型 (LLM) 的数据合成很普遍,但它主要针对具有可验证答案的领域,忽视了开放式人文和社会科学 (HSS),在这些领域,细致入微的质量判断比客观正确性更重要。
- 这使得偏好对齐成为广泛的 HSS 任务的自然范例。
- 然而,现有方法要么成本高昂,要么不适合广泛的 HSS 学科。
- EN 要点:
- arXiv:2607.27366v1 Announce Type: new
- Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended human…
- This makes preference alignment a natural paradigm for broad HSS tasks
- Yet existing methods are either costly or not tailored to broad HSS disciplines
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27379v1 公告类型:新。
- 摘要:高质量、多样化的数据对于大型语言模型 (LLM) 至关重要,但仍然稀缺且成本高昂。
- 数据合成是一种可行的替代方案,并且在封闭任务上取得了成功,但人文和社会科学 (HSS) 却被忽视了,而且它们的开放性本质使得合成具有挑战性。
- 超越以往以能力为中心、碎片化的尝试,采用以主题为中心的范式,定义了第一个涵盖14个主流领域的HSS域系统,并推出了第一个HSS数据合成管道HSS-Synth。
- EN 要点:
- arXiv:2607.27379v1 Announce Type: new
- Abstract: High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly
- Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature ma…
- Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fiel…
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27384v1 公告类型:新。
- 摘要:用于临床诊断推理的大型语言模型对社会语言学记录敏感,而不仅仅是临床内容。
- 我们将这种失败模式称为“叙事锚定”:在不同寄存器中表达的相同临床事实会导致诊断输出出现分歧。
- 与之前的人口统计偏见工作不同,我们的基准隔离注册是唯一的变异渠道,不存在任何形式的人口统计标记。
- EN 要点:
- arXiv:2607.27384v1 Announce Type: new
- Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content
- We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge
- Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of va…
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27393v1 公告类型:新。
- 摘要:仇恨模因是一种日益增长的多模式在线伤害形式,其中敌对意图通常通过图像、文本、文化参考和隐含目标的联合解释来传达。
- 虽然仇恨模因检测在高资源语言中取得了进展,但阿拉伯语仍然未被充分开发,现有模因资源主要集中在宣传或粗俗的有害内容标签上。
- 我们推出了 AHA-Memes(阿拉伯语仇恨模因),据我们所知,这是第一个具有细粒度、多标签注释的大规模阿拉伯仇恨模因基准。
- EN 要点:
- arXiv:2607.27393v1 Announce Type: new
- Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, c…
- While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda o…
- We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label an…
Benchmarking LLM Competence on Logical Inference over Probability Operators
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27405v1 公告类型:新。
- 摘要:不确定性的表达和推论在自然语言中普遍存在,对不确定性的自然语言表达进行有效的推论不仅对于日常对话是必要的,对于医学和法律等高风险领域也是必要的。
- 虽然大型语言模型越来越多地在逻辑推理任务上进行评估,但从巧妙的表面级模式匹配中分离出原则性的符号推理却充满了困难。
- 我们引入了概率算子推理的基准——对具有可分级认知模态(例如,可能、可能、必须)的句子进行推理,包含跨 15 个推理模板的 14,320 个程序生成的英语提示,系统地改变问题形式、否定策略和表面内容。
- EN 要点:
- arXiv:2607.27405v1 Announce Type: new
- Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertain…
- While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level patter…
- We introduce a benchmark for reasoning over probability operators–inference over sentences with gradable epistemic modals (e.g., probably, might, must) contain…
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27421v1 公告类型:新。
-摘要:意图分类是面向任务的对话系统的核心组成部分,但实践者在计算、延迟和鲁棒性约束下选择可部署的开放权重语言模型的系统指导有限。
- 我们对跨越 15 个家族的 41 个开放权重语言模型和八个英语单标签意图分类数据集的 135M–9B 参数范围进行了系统的零样本评估。
- 第九个数据集 ATIS 使用五个标记的演示,并报告为辅助五次结果。
- EN 要点:
- arXiv:2607.27421v1 Announce Type: new
- Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployab…
- We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M–9B parameter range across eight English single…
- A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result
ArXiv cs.LG (B_intro+search) 链接到标题
Recursive transformers for semiconductor thermo-mechanical reliability
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27251v1 公告类型:新。
- 摘要:基于变压器的代理模型越来越多地用于取代工程设计中昂贵的第一原理仿真。
- 但是,对于工程设计空间中典型的小型低维数据集,传统的变压器架构通常会过度参数化,而在工程设计空间中,生成大型仿真数据的成本很高。
- 在这些条件下,过多的参数容量会导致过度拟合,而不是提高准确性,同时还会产生不必要的内存和计算开销。
- EN 要点:
- arXiv:2607.27251v1 Announce Type: new
- Abstract: Transformer-based surrogate models are increasingly used to replace expensive first-principles simulation in engineering design
- But conventional transformer architectures are often over parameterized for the small, low-dimensional datasets typical of engineering design spaces, where larg…
- Under these conditions, excess parameter capacity leads to overfitting rather than improved accuracy, while also incurring unnecessary memory and compute overhe…
Regularizing modality contribution drift in multimodal continual learning
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27260v1 公告类型:新。
- 摘要:多模态持续学习(MMCL)旨在从多模态数据中学习新兴知识,同时保留知识。
- 为了减少遗忘,当前的 MMCL 方法通常关注跨模态表示对齐或语义相似性,但它们忽略了单个模态的相对贡献及其交互在增量任务中是否保持稳定。
- 我们将这种决策水平转变称为模态贡献漂移(MCD),并用 MCD 得分对其进行量化,MCD 得分结合了对模态子集的受控干预下的贡献强度和相对依赖变化。
- EN 要点:
- arXiv:2607.27260v1 Announce Type: new
- Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge
- To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic similarity, but they overlook whether the relativ…
- We term this decision-level shift Modality Contribution Drift (MCD) and quantify it with the MCD score, which combines contribution-strength and relative-relian…
DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27263v1 公告类型:新。
-摘要:时间序列因果推断的大多数基准都是观察性的、小型的或特定领域的,使得干预和反事实估计在最重要的领域得不到服务,例如在医疗保健、政策评估和气候科学领域。
- 我们引入 \textbf{DoTime},一个开放的、可扩展的、具有理论基础的带干预措施的多元时间结构因果模型 (TSCM) 生成器,以 \code{dotime} PyPI 包与四个冻结评估套件一起发布。
- 除了现有的工作之外,它还增加了先前生成器所缺少的功能:连续时间干预 \emph{windows}、具有正向防护的反事实采样模式、作为中断时间序列的严格概括的政权切换 SCM、通过切换 SCM 参数构建的非平稳动力学,以及将趋势和结构中断 \emph{inside} 置于评估窗口内的确定性斜坡和正弦干预曲线。
- EN 要点:
- arXiv:2607.27263v1 Announce Type: new
- Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimati…
- We introduce \textbf{DoTime}, an open, scalable, and theoretically grounded generator of multivariate temporal structural causal models (TSCMs) with interventio…
- Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \emph{windows}, counterfactual sampling modes with a posit…
PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform’s Perspective
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27265v1 公告类型:新。
- 摘要:实时竞价是计算广告的核心,包括三个要素:销售广告展示次数的供应方平台(SSP)、广告商的需求方平台(DSP)竞价以及在它们之间进行拍卖的 Ad Exchange。
- 传统的自动出价算法仅关注 DSP 端,通过调整针对竞争对手的出价来最大限度地提高广告商转化率。
- 然而,目前的大型广告平台,例如社交媒体和电子商务公司,现在内部集成了SSP、DSP和Ad Exchange功能。
- EN 要点:
- arXiv:2607.27265v1 Announce Type: new
- Abstract: Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side P…
- Traditional auto-bidding algorithms focus solely on the DSP side, maximizing advertiser conversions by adjusting bids against competitors
- However, current big ad platforms, such as social media and e-commerce companies, now integrate SSP, DSP, and Ad Exchange functions internally
Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27269v1 公告类型:新。
- 摘要:多头潜在注意力(MLA)对于长上下文 LLM 推理越来越重要,因为紧凑的潜在状态取代了不断增长的键值(KV)缓存并减少了解码内存流量。
- 然而,大多数有能力的开放检查点都使用多头或分组查询注意力(MHA/GQA),因此需要进行转换才能获得 MLA 的缓存效率,而无需从头开始重新训练。
- 推测性解码提供了互补的加速,但其加速取决于提案草案和目标验证之间的协议。
- EN 要点:
- arXiv:2607.27269v1 Announce Type: new
- Abstract: Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value…
- Yet most capable open checkpoints use multi-head or grouped-query attention (MHA/GQA), so conversion is needed to obtain MLA’s cache efficiency without retraini…
- Speculative decoding offers complementary acceleration, but its speedup depends on agreement between draft proposals and target verification
RLPF: Reinforcement Learning from Performance Feedback for Code Generation
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27271v1 公告类型:新。
- 摘要:代码模型越来越多地通过执行反馈进行训练,但大多数训练信号仍然停留在正确性上。
- 这给系统代码留下了一个重要的空白:两个程序可以通过相同的测试,但运行时差异很大。
- 我们研究如何训练代码代理以更快地正确实现,而不是仅将效率视为评估指标。
- EN 要点:
- arXiv:2607.27271v1 Announce Type: new
- Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness
- This leaves an important gap for systems code: two programs can pass the same tests while differing greatly in runtime
- We study how to train code agents to prefer faster correct implementations, rather than treating efficiency only as an evaluation metric
SDO: Structure-Aware Data Organization for Efficient LLM Post-Training
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27273v1 公告类型:新。
- 摘要:大型语言模型的后期训练成本高昂,现有的效率改进主要集中在选择信息丰富的样本或设计训练计划。
- 然而,数据组织本身通常被视为静态预处理步骤:基于嵌入的分组方法在训练之前构建固定分区,并且无法适应优化期间不断变化的样本暴露。
- 因此,尽管所有样本的优化需求不同,但它们都获得了相似的曝光,从而导致某些样本出现冗余更新,而其他样本则未得到优化。
- EN 要点:
- arXiv:2607.27273v1 Announce Type: new
- Abstract: Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing t…
- However, data organization itself is usually treated as a static preprocessing step: embedding-based grouping methods construct fixed partitions before training…
- As a result, all samples receive similar exposure despite their different optimization needs, leading to redundant updates for some samples while leaving others…
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27274v1 公告类型:新。
- 摘要:基于脑电图的疾病诊断需要每个受试者进行一次预测,但常见的管道将记录分割成短实例,继承每个实例的受试者标签,并训练实例级分类器。
- 这假设所有实例都提供同样可靠的诊断证据。
- 多实例学习(MIL)通过将每个主题视为一个包来避免继承标签。
- EN 要点:
- arXiv:2607.27274v1 Announce Type: new
- Abstract: EEG-based disease diagnosis requires one prediction per subject, yet common pipelines segment recordings into short instances, inherit the subject lab…
- This assumes that all instances provide equally reliable diagnostic evidence
- Multiple instance learning (MIL) avoids inherited labels by treating each subject as a bag
Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27275v1 公告类型:新。
- 摘要:据广泛报道,4 位权重的训练后量化几乎是无损的。
- 我们针对多回合、工具调用代理测试了这一说法,这在现在最为重要。
- 在 $\tau^2$-bench 上,跨密集和 MoE 变体的两个开放权重模型系列和两个域(八个单元,每个单元 456 个情节,权重为 16、8 和 4 位),量化在标准指标上看起来确实是自由的。
- EN 要点:
- arXiv:2607.27275v1 Announce Type: new
- Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless
- We test this claim for multi-turn, tool-calling agents, where it now matters most
- On $\tau^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weight…
- 发布时间:2026-07-31 12:00 北京时间
- 摘要:- arXiv:2607.27281v1 公告类型:新。
- 摘要:当语言模型电路的最后部分在一次随机尝试中对齐时,一种能力就会出现在语言模型中,而除了一个之外的所有部分都正确是毫无价值的。
- 我们表明,这种无部分信用的联合调整是能力形成的限速步骤。
- 两个指纹:在无快捷方式的设备中,五部分电路缺少三个,只要三部分电路缺少三个(1.19-1.37)就会等待,因此等待计算的是缺少的部分,而不是大小;在 Pythia 上,跨越 7 种能力和 3 个尺度,在 32 个区分单元中的 32 个中,消除一个部分会留下中位数为 17% 的能力,其中部分信用预测为 50-83% (p = 2e-10),而随机的非部分头部则留下 100%。
- EN 要点:
- arXiv:2607.27281v1 Announce Type: new
- Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth no…
- We show this no-partial-credit joint alignment is the rate-limiting step of capability formation
- Two fingerprints: in a shortcut-free apparatus a five-part circuit missing three waits as long as a three-part circuit missing three (1.19-1.37), so the wait co…