🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-08
- 类型
- ai-daily
- 字数
- 3458
- 阅读时长
- 17 min
2026-07-08 AI日更 | DeepSeek 与 Gemma 4 指向同一件事:模型竞争开始转向效率、端侧与可验证 链接到标题
今天的主线不再只是更强模型,而是更低成本、更高可部署性的工程权衡。DeepSeek 的推理提速、Gemma 4 的开放多模态与端侧进展,显示模型平台正在重算效率边界。同时,规则遵守、RAG、Benchmark 审计等议题升温,说明可靠性验证正成为 AI 落地的新基础设施。
📖 本期 Watch List 深度导读 链接到标题
今天最值得先看的是模型效率与开放多模态:DeepSeek 的“speed hack”和 Gemma 4 技术报告都指向同一趋势——在推理能力、视觉/音频能力与成本之间重新做工程权衡,适合模型平台团队重点跟进。
第二条主线是可靠性评估。Validator-to-Generator Alignment、规则遵守沙盒、Benchmark 审计失败模式,以及跨语言 RAG 的训练后问题,都在提醒我们:模型会说对,不等于会稳定做对;评测和审计本身也需要被审计。
最后,语音与时序基础模型值得作为应用侧观察。低资源 SQA、代码转换 ASR、印度方言识别,以及电价预测、联邦 Mamba 时序模型,展示了基础模型正进入更复杂、更脏、更受约束的真实场景。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:AI Builders Race to Extract Claude Fable 5 Skills Before Paid Access Kicks In 链接到标题
- 分类:AI · News
- 概况:热度时间:9 hours ago,相关帖子数:34000
- 是什么事:X 上大量 AI 开发者正赶在 Claude Fable 5 转为付费访问前,集中测试、复现和提取其能力与提示技巧。
- 为什么重要:这反映出高性能 AI 模型的访问权限、成本与能力扩散正在成为开发者生态中的关键变量,也凸显闭源模型能力被快速学习和迁移的现实压力。
- 讨论概况:讨论焦点集中在免费窗口期是否会催生大量逆向测试与技巧分享、付费墙是否会限制创新,以及这种“抢跑式”能力提取是否合理或存在版权与平台规则风险。
话题 2:Anthropic Discovers Hidden Workspace in Claude AI Models 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:28000
- 是什么事:Anthropic 发布研究称,其通过新的可解释性工具 J-lens 在 Claude 模型内部发现了一个可报告、可调控、支持推理的“J-space”隐藏工作区。
- 为什么重要:该发现为理解大语言模型如何在不显式输出的情况下进行中间概念表征、推理和自我监控提供了新证据,也可能用于更早识别模型的策略性行为、情境意识和潜在安全风险。
- 讨论概况:X 上讨论集中在两点:一方认为这强化了 AI 具备类似“全局工作空间”的功能结构,可能推动对机器意识和模型安全的研究;另一方则强调这不等于证明 AI 有主观体验,担心媒体将可解释性发现过度解读为“Claude 有意识”。
话题 3:OpenAI Teases Broader GPT-5.6 Rollout as Fans Mark Daily Hype Days 链接到标题
- 分类:AI · News
- 概况:热度时间:16 hours ago,相关帖子数:2900
- 是什么事:OpenAI 暗示将更广泛推出 GPT-5.6,引发用户在 X 上持续倒数和造势。
- 为什么重要:如果 GPT-5.6 扩大开放,可能意味着 OpenAI 在模型能力、产品节奏和竞争压力上的新一轮推进,影响开发者、企业用户和生成式 AI 市场预期。
- 讨论概况:X 上讨论主要集中在 GPT-5.6 是否会带来显著能力提升、何时向更多用户开放,以及这是否只是营销预热;支持者期待新功能和更强性能,怀疑者则认为近期 AI 发布节奏过密、实际改进可能有限。
话题 4:Claude Fable 5 Free Access Ends Today for Most Users 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:20000
- 是什么事:Claude Fable 5 面向多数用户的免费访问权限据称将于今天结束,相关话题在 X 上引发大量讨论。
- 为什么重要:这反映出前沿 AI 模型从免费体验转向付费或受限使用的商业化趋势,关系到用户获取先进模型的门槛、平台增长策略以及 AI 服务成本分摊。
- 讨论概况:X 上讨论主要集中在免费期结束是否合理、订阅价格是否值得、Claude Fable 5 相比其他模型的能力优势,以及企业用户是否会因数据与成本顾虑转向自建或替代方案。
话题 5:Drake Parody ‘Claude’s Plan’ Captures Coders’ AI Obsession 链接到标题
- 分类:AI · Entertainment
- 概况:热度时间:14 hours ago,相关帖子数:141
- 是什么事:一首模仿 Drake 风格的歌曲《Claude’s Plan》在 X 上走红,用娱乐化方式调侃程序员对 Claude 等 AI 编程工具的依赖与迷恋。
- 为什么重要:这反映出 AI 编程助手已从专业工具进入开发者文化与大众娱乐语境,显示生成式 AI 正在改变程序员的工作方式、身份认同和社区表达。
- 讨论概况:X 上的讨论主要围绕 AI 编程是否真正提升效率、开发者是否过度依赖 Claude,以及这类 AI 梗文化是在幽默记录技术变迁还是在放大行业焦虑。
话题 6:Tesla Cybercabs Spotted on Texas Streets in Robotaxi Testing 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:9100
- 是什么事:多辆带有官方“Cybercab”标识的特斯拉无人出租车被拍到在得州超级工厂及周边进行测试和调度活动。
- 为什么重要:这显示特斯拉正推进自动驾驶出行服务的车辆准备与路测,可能影响无人出租车商业化进程、自动驾驶监管讨论以及AI在交通场景中的落地竞争。
- 讨论概况:X上的讨论集中在Cybercab是否已接近量产和上线服务;支持者认为这是Robotaxi时代临近的明确信号,质疑者则关注其自动驾驶安全性、监管审批、真实运营能力以及时间表是否会再次延后。
今日 X 上的 AI 舆情小结 链接到标题
今天的舆论主线围绕“前沿 AI 能力加速扩散与商业化收口”展开:一边是开发者趁 Claude Fable 5 免费期结束前集中测试、复现能力,另一边是 OpenAI 预热 GPT-5.6、特斯拉推进 Cybercab,显示大模型与 AI 落地应用都进入更激烈的产品竞速。较明显的共识是,高性能模型和 AI 工具已深度影响开发、内容文化与产业预期,免费访问、订阅价格、性能提升和真实可用性正在成为用户判断平台价值的核心。分歧则集中在三类问题上:闭源模型能力被“抢跑式”提取是否合理,J-lens 等可解释性发现是否能被解读为更接近意识或只是内部表征证据,以及 GPT-5.6、Cybercab 等新进展究竟是实质突破还是营销造势。潜在风险包括模型能力扩散带来的版权与平台规则争议、媒体对“AI 有意识”的过度叙事、用户和开发者对 AI 编程工具的过度依赖,以及自动驾驶在安全、监管和商业化时间表上的不确定性。
💡 大佬观点(Influencer Insights) 链接到标题
今日 AI 领域深度洞察:2026年7月7日 链接到标题
我认真分析了过去24小时内多位AI资深从业者在X平台发布的推文,以下是核心发现。
1. 今日大佬共同关注的技术趋势与产品热点 链接到标题
核心焦点:Claude Fable 5 —— 范式的彻底重构
毫无疑问,Claude Fable 5 是今天所有讨论的中心。它不再仅仅是一个“更强的模型”,而正在引发一场关于软件开发、产品构建乃至思维方式的革命。
- 从“写代码”到“提需求”的彻底转变:@dotey 分享的 Claude Code 诞生记中提到,Anthropic 团队成员
Boris Cherny现在 100% 的代码都由 Claude Code 完成,一行手敲代码都没有了。另一位成员MEAGHAN CHOI指出,直到模型能力跨越临界点,产品的形态才自然浮现。 - 能力悬余 (Capability Overhang) 与释放模型潜能:@dotey 引用了 Claude Code 工程师
Thariq Shihipar的演讲,他提出模型其实早已具备很多能力,只是我们没找到正确的打开方式。例如,Fable 5 砍掉了 80% 的系统提示词,从“给示例约束”转变为“给上下文,不给约束”,因为模型自身的想象力远超我们给的示例。 - 工作流与成本优化成为新焦点:@vista8 分享了
Simon Willison的省钱方法,主循环用 Fable/Opus 等“判断型”强模型,执行写代码等机械任务则调用 Sonnet/Haiku 等“执行型”便宜模型,通过/goal和workflows实现自动化。@dotey 也提到 Fable 5 有 50% 的订阅额度限制以及即将到来的按量计费模式,成本控制成为现实问题。
次热点:端侧模型 (On-device Model) 与应用生态繁荣
尽管 Fable 5 是云端霸主,但端侧模型的进展同样引人注目。
- 端侧能力加速落地:@zhixianio 持续关注并测试端侧模型,如
MiniCPM-o 4.5的实时音视频和Gemma 4系列,认为其效果“已经可以用起来了”,并对 Google 的 QAT 量化训练思路表示关注,这将使模型更易于在手机等终端设备上部署。 - AI Agent 与 Skills 生态系统爆发:Skills 正成为连接模型能力和专业工作流的中间层,生态日趋成熟。
- @Pluvio9yte 开源了他的 AI Agent Skill 集合
rnskill,涵盖写作、视频、质检等场景。 - @ruanyf 惊讶地发现小红书上线了
REDSkill社区,让用户能在社交媒体上分享和安装 Skills,试图成为“Skill 的 GitHub”。 - @dotey 更新了他的
baoyu-designskill,现在已支持在生成的 PPT 中加入复杂动画。
- @Pluvio9yte 开源了他的 AI Agent Skill 集合
2. 值得注意的独特观点与行业前瞻 链接到标题
- 编程的“失去”与“得到”:@dotey 的博文和
Thariq Shihipar的演讲都触及了程序员在 AI 时代的复杂情感——享受效率飞升的同时,也怀念过去“在脑子里旋转整个代码库”的掌控感。但现实是,Fable 几小时就能完成过去数周的工作。 - 瓶颈从模型转向“人”:@vista8 和 @dotey 都明确提出,当模型足够强时,瓶颈变成了人的表达能力和对结果的验证能力。如何清晰表达模糊想法、如何审查 AI 产出的安全性和正确性,成为新的核心竞争力。
- 组织架构将被重塑:@Pluvio9yte 预测大厂一到两年内将不再区分前后端。@dotey 也认为大部分公司未来可能不再需要传统的
web infra team。 - 开源模型的“伪命题”辩论:@ruanyf 转述了 Anthropic 创始人的观点,认为当前 AI 模型的“开源”更像是“开放权重”,因为无法看到内部运作或参与开发,与传统开源软件的模式截然不同。
- “品味”与“异质性”的价值凸显:@lijigang 提出“品味是一个人的损失函数”,在 AI 能生成无数同质化内容时,个人独到、甚至粗糙的审美和“异质性”将成为一种珍贵的美学。
- AI 编程的成本悖论:@ruanyf 指出,员工无限制使用顶级模型进行 AI 编程,其年成本可能高达上亿,这甚至比雇佣真人程序员更昂贵,这对“AI 降本增效”的朴素认知提出了挑战。
3. 推荐的工具与资源 链接到标题
以下是根据大佬推荐整理的工具和资源列表,今日主打“工作流”和“生产力”:
| 类别 | 名称 | 核心用途 | 推荐来源 |
|---|---|---|---|
| AI Agent 工具 | OpenConnector | 给 AI Agent 用的开源认证网关,解决了 Agent 连接 1000+ 应用的鉴权和工具调用问题,是 Composio 的开源替代。 | @Pluvio9yte |
| DevSpace | 将一个本地 MCP 服务器通过隧道暴露给 ChatGPT 网页端,让 GPT 5.5 Pro 等模型也能直接读写本地代码,相当于给 ChatGPT 赋予了 Codex 的能力。 | @gefei55 | |
| TokHub | 开源的 AI API 中转站监控及网关管理系统,可用于评测中转站速度和管理内部 Token 分发。 | @vista8 | |
| AI 编程与设计 | baoyu-design Skill | 一个 Skills 工具包,可直接生成 HTML 格式的 PPT,并借助 Fable 5 的能力导出带有复杂动画效果的 PPTX 文件。 | @dotey |
| 96 种 UI 设计风格库 | 一个开源免费的资源站,收录了 Linear, Vercel, Apple 等 96 种设计风格的代码库,可直接拉入项目,让 Agent 照着指定风格写代码,一键解决 “AI 味 UI” 问题。 | @AI_Jasonyu | |
| rnskill (AI Agent Skill 集合) | 包含写作去 AI 味、动效视频导演、视频风格模板、视频质检等多项技能的开源包,支持 Codex, Claude Code 等。 | @Pluvio9yte | |
| AI 视频创作 | Topview 3D Shot Composer | 一款 AI 视频制作工具,其亮点是允许创作者先在 3D 空间中摆放角色、道具和机位,像导演一样构图,再由 AI 生成视频,解决了提示词难以精准控制构图的痛点。 | @AI_Jasonyu |
| 信息获取与学习 | X 爆款话题挖掘工具 | @gefei55 开源的一个基于 Twitter API 的工具,能实时扫描高互动的带链接推文,并反查域名流量,旨在比 Google Trends 更快发现热点和新词,抢占 SEO 或产品先机。 | @gefei55 |
| Hackernews RSS 库 & IMDB 电影站 | 前者将 Hackernews 内容封装为可高度定制的 RSS 源;后者是用 AI 一键生成的 IMDB Top 250 电影管理推荐站,均开源,是快速搭建信息产品的范例。 | @vista8 |
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 33 条更新
Stratechery by Ben Thompson (A_full) 链接到标题
- A Script for Mark Zuckerberg
- 发布时间:2026-07-07 18:00 北京时间
- 摘要:- 听这个帖子**:**。
- 背景: Meta 在 2026 年 8 月初召开的财报电话会议。。
- 演讲者: Meta 首席执行官马克·扎克伯格。
- 大家下午好,欢迎参加 Meta Platforms 2026 年第二季度收益电话会议。
- 我们今天的言论将包括前瞻性陈述,这些陈述基于今天的假设。
- EN 要点:
- Listen to this post :
- Log in to listen
- The setting: Meta’s earnings call in early August, 2026
- The speaker: Meta CEO Mark Zuckerberg
OpenAI Blog (A_full) 链接到标题
- Australian Payments Plus moves faster with ChatGPT and Codex
- 发布时间:2026-07-07 08:00 北京时间
- 摘要:- 它位于支付生态系统的中心,支持数百万人每天使用的产品和服务。
- 其团队跨计划规则、技术规范、成员义务、运营流程、网络安全和弹性以及监管期望开展工作,其中速度很重要,但准确性和问责制更重要。
- 这使得知识工作异常复杂。
- 员工通常需要综合大量背景信息,并将技术信息转化为明确的决策、文件和面向成员的指导。
- AP+ 在整个公司范围内引入了 ChatGPT Enterprise,以帮助员工更快地应对复杂性,而 Codex 则成为产品、工程和技术工作流程的下一阶段。
- EN 要点:
- See how Australian Payments Plus uses ChatGPT Enterprise and Codex to move faster through payments complexity
- AP+ saves time, improves quality, and keeps human judgment central.
Two Minute Papers (B_intro+search) 链接到标题
- DeepSeek’s New AI Speed Hack Is Amazing
- 发布时间:2026-07-08 00:33 北京时间
- 摘要:- ❤️ 在这里查看 Lambda 并注册他们的 GPU Cloud:。
- 📝 DeepSeek 论文可在此处获取:。
- Adam Bridges、Benji Rabhan、B Shang、Cameron Navor、Charles Ian Norman Venn、Christian Ahlin、Eric T、Fred R、Gordon Child、Juan Benet、Michael Tedder、Owen Skarpness、Richard Sundvall、Ryan Stankye、Shawn Becker、Steef、Taras Bobrovytsky、Tazaur Sagenclaw、Tybie Fitzhugh、Ueli Gallizzi。
- DeepSeek 的新 AI 速度破解令人惊叹。
- EN 要点:
- ❤️ Check out Lambda here and sign up for their GPU Cloud:
- 📝 The DeepSeek paper is available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Ska…
ArXiv cs.AI (B_intro+search) 链接到标题
iFLYTEK-Embodied-Omni Technical Report
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02542v1 公告类型:新。
- 摘要:通用实体代理必须理解多模式指令,预测其环境将如何演变,并在更广泛的范围内产生精确的控制动作。
- 现有方法通常专注于视觉语言推理、基于视频的世界建模或动作生成,而首先综合未来观察结果然后推断动作的级联管道可能会引入界面瓶颈和复合预测错误。
- 我们推出了 iFLYTEK-Embodied-Omni,这是一个统一的多模态基础模型,可在单个 Omni 框架内对视觉(视频和图像)、语言和动作进行联合建模。
- EN 要点:
- arXiv:2607.02542v1 Announce Type: new
- Abstract: General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control act…
- Existing approaches typically specialize in visual-language reasoning, video-based world modeling, or action generation, while cascaded pipelines that first syn…
- We present iFLYTEK-Embodied-Omni, a unified multimodal foundation model that jointly models vision(videos and images), language, and action within a single Omni…
Internal Pluralism and the Limits of Pairwise Comparisons
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02672v1 公告类型:新。
- 摘要:局部成对比较是了解人们希望决策规则如何发挥作用的标准工具,例如在参与式设计或协调中。
- 然而,它们的使用建立在两个强有力的假设之上:局部比较足以证明一个人希望自动决策规则如何表现,并且人们总是可以果断地回答这些比较。
- 我们研究这些假设在内部多元化下如何受到损害:个人根据关于规则应如何表现的多个权威优先级来评估决策规则。
- EN 要点:
- arXiv:2607.02672v1 Announce Type: new
- Abstract: Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignment
- However, their use builds in two strong assumptions: that local comparisons are sufficient evidence about how a person wants an automated decision rule to behav…
- We investigate how these assumptions may be compromised under internal pluralism: the idea that an individual evaluates decision rules according to multiple aut…
ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observability
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02686v1 公告类型:新。
-摘要:在部分可观察性下运行的强化学习代理必须对不完整的信息采取行动,使它们成为具有广泛推理先验的小语言模型(SLM)指导的自然候选者。
- 然而,将 SLM 指南集成到这种设置中已被证明是困难的:在所有测试环境中,普通的不确定性门控方法实现的覆盖率为零或接近零,这意味着 SLM 几乎从不贡献独立的操作。
- 我们将这种失败追溯到纯粹的以自我为中心的提示,它为真正的推理提供了不足的背景,并将其识别为背景问题而不是能力问题。
- EN 要点:
- arXiv:2607.02686v1 Announce Type: new
- Abstract: Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance fr…
- Yet integrating SLM guidance into this setting has proven difficult: across all test environments, vanilla uncertainty-gated approaches achieve an overwrite rat…
- We trace this failure to the bare egocentric prompt, which provides insufficient context for genuine reasoning, and identify it as a context problem rather than…
Automated Data Readiness for Scientific AI
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02771v1 公告类型:新。
- 摘要:领导计算设施管理着大规模的科学数据集,这些数据集在用作人工智能训练数据之前通常需要进行大量转换。
- 然而,现有的框架还没有完全统一自动化转换、准备情况评估、来源跟踪和代理本机部署。
- 我们提出了 REDI,这是一个开源框架,它通过统一的五阶段管道(摄取、预处理、转换、结构和输出)来解决这一差距,并通过每阶段的仪器来实现可重复性和部署为代理可调用的技能;配套工具 SetGo 可自动执行 FAIR 合规性和目录发布。
- EN 要点:
- arXiv:2607.02771v1 Announce Type: new
- Abstract: Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI trainin…
- However, no existing framework fully unifies automated transformation, readiness assessment, provenance tracking, and agent-native deployment
- We present REDI, an open-source framework that addresses this gap through a unified five-stage pipeline (ingest, preprocess, transform, structure, and output) w…
SwarmResearch: Orchestrating Coding Agents for Open-Ended Discovery
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02807v1 公告类型:新。
- 摘要:长时间运行的编码代理(例如自动研究)可以持续发现开放式问题的优化。
- 然而,他们倾向于收敛于单一的高级方法,然后进行低级编辑,而忽略了解决问题的其他高级方法。
- 我们假设两个线束级设计选择导致了这种行为:在单个长期运行的代理中累积上下文,并且仅公开单个程序状态进行编辑。
- EN 要点:
- arXiv:2607.02807v1 Announce Type: new
- Abstract: Long-running coding agents such as autoresearch can persistently discover optimizations for open-ended problems
- However, they tend to converge onto a single high-level approach, then proceed with low-level edits while missing other superior approaches to the problem
- We hypothesize two harness-level design choices contribute to this behavior: accumulating context in a single long-running agent and only exposing a single prog…
Object-Centric Environment Modeling for Agentic Tasks
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02846v1 公告类型:新。
-摘要:大型语言模型(LLM)代理可以通过积累经验来改进,但随着交互的增长,自由形式的文本记忆变得难以维护、验证和重用。
- 最近的符号方法学习可执行技能或程序化世界模型,但通常存储本地程序或假设简化的动态。
- 我们提出以对象为中心的环境建模(OCM),它将经验组织成可执行的以对象为中心的环境模型。
- EN 要点:
- arXiv:2607.02846v1 Announce Type: new
- Abstract: Large language model (LLM) agents can improve through accumulated experience, but free-form textual memories become difficult to maintain, validate, a…
- Recent symbolic approaches learn executable skills or programmatic world models, yet often store local procedures or assume simplified dynamics
- We propose Object-Centric Environment Modeling (OCM), which organizes experience into an executable object-centric environment model
MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02879v1 公告类型:新。
-摘要:当前评估医学计算中的大型语言模型(LLM)的基准很大程度上基于简化的设置,其中每个患者病例对应一个计算器,并且在查询中明确指定所需的工具。
- 然而,真实的临床场景往往需要多个计算器进行联合评估、嵌套尺度计算以及不直接指定目标计算器的模糊查询。
- 为此,我们提出了一个新的医学计算基准MedCalc-Pro,它涵盖了三种逐渐具有挑战性的任务设置:单计算器、多计算器和嵌套计算器计算设置。
- EN 要点:
- arXiv:2607.02879v1 Announce Type: new
- Abstract: Current benchmarks for evaluating large language models (LLMs) in medical calculation are largely based on simplified settings, where each patient cas…
- However, real clinical scenarios often require multiple calculators for joint evaluation, nested-scale calculation, and fuzzy queries that do not directly speci…
- To this end, we propose a new medical calculation benchmark, MedCalc-Pro, which covers three progressively challenging task settings: single-calculator, multi-c…
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02914v1 公告类型:新。
- 摘要:大型语言模型 (LLM) 在不同的应用程序中展示了卓越的功能,但同时确保其安全性、有用性和可信性仍然是一个持续的挑战。
- 传统的以拒绝为导向的对齐策略可以减少有害内容的生成,但系统性地无法满足合法的用户需求,通常会隐瞒可以安全且建设性地解决敏感查询的潜在意图的信息。
- 基于 Oyster-I 开创的建设性安全范式,该范式超越了全面拒绝,转向深思熟虑的、面向响应的安全一致性,我们确定了其基于监督微调 (SFT) 的方案的两个关键局限性:对分布外场景的安全泛化不足,以及我们称之为安全思想链 (CoT) 过度泛化的现象,其中以安全为导向的推理模式过度应用于良性查询,降低了帮助性和用户体验。
- EN 要点:
- arXiv:2607.02914v1 Announce Type: new
- Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety, helpfulnes…
- Conventional refusal-oriented alignment strategies mitigate harmful content generation but systematically fail to serve legitimate user needs, often withholding…
- Building upon the constructive safety paradigm pioneered by Oyster-I, which moves beyond blanket refusal toward thoughtful, response-oriented safety alignment,…
VERITAS: Towards a General-Purpose Replication Tool for Scientific Research
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02931v1 公告类型:新。
- 摘要:人工智能工具正在加速科学出版,而审查系统却难以跟上,对已发表研究的独立验证变得更加困难和重要。
- 由于手动复制缓慢且昂贵,越来越多的工作使用编码代理来自动化部分流程。
- 现有的工作大部分被打包为基准测试,其配套代理仅在基准测试自己的管道内运行,并且不存在通用复制工具。
- EN 要点:
- arXiv:2607.02931v1 Announce Type: new
- Abstract: AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published resea…
- As manual replication is slow and expensive, a growing line of work uses coding agents to automate parts of the process
- Existing efforts are largely packaged as benchmarks with companion agents that only run inside the benchmark’s own pipeline, and no general-purpose replication…
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02941v1 公告类型:新。
- 摘要:多产品配套交付对集成加工和装配的混合制造系统中的实时调度提出了重大挑战,因为动态订单到达同时改变了供应依赖性和可行的作业机器分配集。
- 本文提出了一种基于滑动窗口的强化学习(SWRL)框架,用于在具有复杂配套约束的柔性装配流水车间调度问题中进行端到端在线调度。
- 该问题被表述为基于异构图的马尔可夫决策过程,该过程捕获双层配套结构和产生稀疏奖励景观的尾部产品瓶颈动态。
- EN 要点:
- arXiv:2607.02941v1 Announce Type: new
- Abstract: Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate processing and a…
- This paper proposes a sliding-window-based reinforcement learning (SWRL) framework for end-to-end online scheduling in the flexible assembly flow shop schedulin…
- The problem is formulated as a heterogeneous graph-based Markov decision process that captures the dual-layer kitting structure and the tail-product bottleneck…
ArXiv cs.CL (B_intro+search) 链接到标题
Improving LLMs via Validator-to-Generator Alignment
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02668v1 公告类型:新。
- 摘要:大型语言模型不一致:不同的提示或包含不相关的信息可能会导致模型输出出现意外的变化。
- 生成器-验证器 (G-V) 差距是这种现象的一种表现,其中 LLM 生成响应,如果重新查询以验证它们,则它们会认为这些响应无效。
- 在这项工作中,我们引入了一种新的 G-V 一致性公式,其中涉及对话语频率的原则性校正。
- EN 要点:
- arXiv:2607.02668v1 Announce Type: new
- Abstract: Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs
- The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validat…
- In this work, we introduce a new formulation of G-V consistency that involves a principled correction for utterance frequency
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02734v1 公告类型:新。
- 摘要:社交媒体的快速增长通过实现快速信息交换改变了全球通信,但也加速了错误信息的传播。
- 假新闻、操纵内容和挑衅性叙事越来越多地与社会动荡、政治不稳定和暴民暴力联系在一起。
- 南亚和其他地方发生的事件表明,通过 Facebook 和 WhatsApp 等平台传播的虚假信息可能会引发现实世界的伤害,其传播速度往往快于事实核查工作的反应速度。
- EN 要点:
- arXiv:2607.02734v1 Announce Type: new
- Abstract: Rapid growth in social media has transformed global communication by enabling fast information exchange, but it has also accelerated the spread of mis…
- Fake news, manipulated content, and provocative narratives are increasingly linked to social unrest, political instability, and mob violence
- Incidents in South Asia and elsewhere demonstrate how false information disseminated via platforms such as Facebook and WhatsApp can trigger real-world harm, of…
Reinforcement Learning for Data-Efficient Code-Switched ASR
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02757v1 公告类型:新。
- 摘要:音频语言模型可以提示进行代码转换语音,但其解码并未针对代码转换进行优化,并且经常在语言边界处失败。
- 我们提出了一种实用的强化学习,具有可验证的奖励配方,使用组相对策略优化,将音频语言模型数据有效地适应代码转换的 ASR,将错误率奖励与惩罚错误书写系统的脚本保真度奖励相结合,并采用两遍草稿和细化程序。
- 使用 Qwen2-Audio 作为跨 10 个语言对的可重复测试平台,仅对 TTS 代码转换语音进行训练,我们表明 10% 的数据的 RLVR 与在完整数据集上训练的 LoRA 监督微调相匹配,在类型上距离较远的语言对上收益最大。
- EN 要点:
- arXiv:2607.02757v1 Announce Type: new
- Abstract: Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language bou…
- We propose a practical reinforcement learning with verifiable rewards recipe for data-efficient adaptation of audio-language models to code-switched ASR using g…
- Using Qwen2-Audio as a reproducible testbed across 10 language pairs, training on only TTS code-switched speech, we show that RLVR with 10% of the data matches…
LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02763v1 公告类型:新。
- 摘要:口语问答(SQA)仍然主要关注高资源语言和仔细录制的语音,限制了语音法学硕士方法在低资源环境中的应用范围。
- 本文研究了文本转语音 (TTS) 是否可以为卢森堡 SQA 提供特定于任务的训练数据,而无需大量人工记录的 QA 语料库。
- 从现有的基于文本的 QA 资源开始,我们将问题翻译成卢森堡语,使用多个 TTS 系统合成口头问题,并将其与文本答案配对。
- EN 要点:
- arXiv:2607.02763v1 Announce Type: new
- Abstract: Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recorded speech, limiting the reach of speech-LLM met…
- This paper investigates whether text-to-speech (TTS) can provide task-specific training data for Luxembourgish SQA without requiring a large human-recorded QA c…
- Starting from existing text-based QA resources, we translate questions into Luxembourgish, synthesize spoken questions with multiple TTS systems, and pair them…
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02770v1 公告类型:新。
- 摘要:我们介绍 Gemma 4,它是 Gemma 模型家族中的新一代开放权重、原生多模态语言模型。
- Gemma 4 模型套件旨在提高计算效率和推理能力,具有密集的专家混合架构,参数范围从 2.3B 到 31B。
- 除了针对所有模型尺寸改进视觉和音频编码器之外,我们还为我们的 12B 模型提出了一个统一的、无编码器的架构,该架构摄取原始音频和图像补丁。
- EN 要点:
- arXiv:2607.02770v1 Announce Type: new
- Abstract: We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family
- Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B para…
- Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and…
Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02802v1 公告类型:新。
-摘要:随着法学硕士越来越多地在半开放文本游戏环境中被部署为自主裁决者,当用户意图与系统规则发生冲突时,稳健的规则遵守变得至关重要。
- 然而,这些模型经过训练是有帮助和合规的,这使得它们容易受到我们称为 \textit{修辞注入} 的一类攻击,其中敌对用户利用伪逻辑推理和权威强制等叙事框架技术来绕过裁决逻辑。
- 我们提出了 CoC-Seduce,这是一个基于桌面角色扮演游戏 (TRPG) 机制的多智能体对抗基准测试,它是半开放环境的理想实例,其中规则明确用于裁决,但交互仍然完全以自然语言进行。
- EN 要点:
- arXiv:2607.02802v1 Announce Type: new
- Abstract: As LLMs are increasingly deployed as autonomous adjudicators in semi-open textual game environments, robust rule adherence becomes critical when user…
- However, these models are trained to be helpful and compliant, leaving them vulnerable to a class of attacks we term \textit{Rhetorical Injection}, where advers…
- We present CoC-Seduce, a multi-agent adversarial benchmark built on Tabletop Role-Playing Game (TRPG) mechanics, an ideal instantiation of semi-open environment…
Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02862v1 公告类型:新。
- 摘要:自动语音识别 (ASR) 和方言识别 (DID) 对于印度语言至关重要,其中许多语言资源匮乏且表现出显着的方言差异。
- 现有方法通常单独优化 ASR 或 DID,从而导致性能权衡。
- 在这项工作中,我们提出了一个联合改进 ASR 和 DID 的多模式框架。
- EN 要点:
- arXiv:2607.02862v1 Announce Type: new
- Abstract: Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit signi…
- Existing methods often optimize ASR or DID individually, resulting in performance trade-offs
- In this work, we propose a multimodal framework that jointly improves ASR and DID
PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Prediction
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02881v1 公告类型:新。
- 摘要:长视野行为预测旨在根据漫长的历史序列推断用户的下一步行动,在人工智能领域发挥着至关重要的作用。
- 大语言模型 (LLM) 的兴起为顺序行为预测提供了一个有希望的方向,但 LLM 在处理长期行为预测时却面临潜在行为模式归纳和模型内在认知偏差的困扰。
- 先前的内存管理方法遵循上下文压缩范例,试图通过减轻历史序列负担来解决此任务,但未能解决核心挑战。
- EN 要点:
- arXiv:2607.02881v1 Announce Type: new
- Abstract: Long-horizon behavior prediction aims to infer a user’s next action based on a lengthy historical sequence, playing a crucial role in artificial intel…
- The rise of large language models (LLMs) offers a promising direction for sequential behavior prediction, yet LLMs struggle with latent behavioral pattern induc…
- Prior memory management methods follow a context-compression paradigm that attempts to address this task by alleviating the historical sequence burden, yet fail…
Where do LLMs Fall Short in CBT-Guided Affective Reasoning?
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02885v1 公告类型:新。
- 摘要:认知行为疗法(CBT)提供了一个结构化框架,通过检查认知和行为因素之间的相互作用来理解用户的心理状态。
- 然而,开箱即用的法学硕士反应流畅且富有同理心,但无论用户实际需要什么,都会陷入验证和反思。
- 他们了解理论 CBT(在许可考试问题上的准确率高达 96%),但未能有效应用。
- EN 要点:
- arXiv:2607.02885v1 Announce Type: new
- Abstract: Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user’s mental state by examining the interaction between cognit…
- However, out-of-the-box LLMs respond fluently and empathetically, yet collapse into validation & reflection, regardless of what the user actually needs
- They know theoretical CBT (scoring up to 96% accuracy on licensing exam questions) but fail to apply it effectively
Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02966v1 公告类型:新。
-摘要:跨语言检索增强生成(RAG)通常部署在英语证据体系中,用户用多种语言查询,但检索到的段落仍然是英语。
- 在这种情况下,尽管有强大的基础模型,生成也可能会失败:英语证据会导致语言漂移(英语或代码转换输出),并且模型在生成非英语答案时使用的证据不可靠。
- 我们将这些失败归因于两个训练后挑战:(i)错误与前缀相关,因此固定轨迹监督会遭受前缀不匹配的影响; (ii) 序列级(部分离散/基于判断)奖励会产生嘈杂的信用分配和高方差更新。
- EN 要点:
- arXiv:2607.02966v1 Announce Type: new
- Abstract: Cross-lingual retrieval-augmented generation (RAG) is often deployed in an English-evidence regime, where users query in diverse languages but retriev…
- In this setting, generation can fail despite strong base models: English evidence induces language drift (English or code-switching outputs) and models use evid…
- We attribute these failures to two post-training challenges: (i) errors are prefix-dependent, so fixed-trajectory supervision suffers from prefix mismatch; and…
ArXiv cs.LG (B_intro+search) 链接到标题
Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02586v1 公告类型:新。
- 摘要:治理框架要求人工智能提供商和审计师提供书面评估证据,而基于扰动的构造有效性审计是该证据的常见形式。
- 我们认为审计本身是脆弱的:他们的结论可以通过实施细节默默地制造出来,而读者在报告的数字中看不到这些细节。
- 我们命名了五类管道故障,并在安全基准和开放权重指令调整模型的自我审核中演示了每一类。
- EN 要点:
- arXiv:2607.02586v1 Announce Type: new
- Abstract: Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common…
- We argue the audits are themselves fragile: their conclusions can be silently manufactured by implementation details that readers cannot see in the reported num…
- We name five classes of pipeline failure and demonstrate each in a self-audit over safety benchmarks and open-weight instruction-tuned models
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02623v1 公告类型:新。
-摘要:时间序列基础模型(TSFM)显示出强大的零样本预测性能,但它们在协变量驱动的非平稳设置中的泛化尚未得到充分探索。
- 由于复杂的时间依赖性、分布变化以及对结构和上下文信息的强烈依赖,电价预测(EPF)提出了一个具有挑战性的测试平台。
- 我们为 EPF 提出了一个双数据集基准框架,以减轻污染风险并实现对 TSFM 的公平评估。
- EN 要点:
- arXiv:2607.02623v1 Announce Type: new
- Abstract: Time series foundation models (TSFMs) have shown strong zero-shot forecasting performance, but their generalization in covariate-driven, non-stationar…
- Electricity price forecasting (EPF) presents a challenging testbed due to complex temporal dependencies, distributional shifts, and strong reliance on structura…
- We propose a two-dataset-benchmarking framework for EPF to mitigate contamination risk and enable fair evaluation of TSFMs
QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02632v1 公告类型:新。
- 摘要:时间序列预测支持金融、能源、交通、公共卫生和工业监测方面的决策。
- 最近的基础模型改进了预测任务之间的传输,但许多模型依赖于中心化数据和 Transformer 注意力,这限制了它们在长、高维和隐私敏感信号中的使用。
- 本文提出了 QuantFlow,一种概率预测框架,结合了倒序列嵌入、双向 Mamba 状态空间解码器、分位数回归和联邦学习。
- EN 要点:
- arXiv:2607.02632v1 Announce Type: new
- Abstract: Time-series forecasting supports decisions in finance, en-ergy, transportation, public health, and industrial monitoring
- Recent foundation models improve transfer across forecast-ing tasks, but many depend on centralized data and Trans-former attention, which restricts their use f…
- This paper presents QuantFlow, a probabilistic forecasting framework that com-bines inverted sequence embedding, bidirectional Mamba state-space decoders, quant…
GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02633v1 公告类型:新。
- 摘要:我们提出了 Graft,一种用于文本到语音神经编解码器语言建模的每个单词的发音调节机制。
- 现有系统达到了很高的清晰度和自然度,但继承了文本的歧义性,并且错误地发音了罕见的专有名词、借词和技术术语。
- 即使音素调节模型也不提供每个单词发音的直接声学处理。
- EN 要点:
- arXiv:2607.02633v1 Announce Type: new
- Abstract: We present GRAFT, a per-word pronunciation conditioning mechanism for text-to-speech neural codec language modeling
- Existing systems reach high intelligibility and naturalness but inherit the ambiguity of text and mispronounce rare proper nouns, loanwords and technical terms
- Even phoneme-conditioned models offer no direct acoustic handle for per-word pronunciation
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02636v1 公告类型:新。
- 摘要:物体检测是安全关键型无人机和边缘视觉系统中人工智能驱动感知的一项基本功能,包括灾难响应、操作安全环境、基础设施监控和防御应用。
- 此类环境中的稳健模型性能取决于大型且持续更新的数据集。
- 然而,训练高性能探测器通常需要集中航空图像,这提出了隐私、监管、存储和带宽挑战。
- EN 要点:
- arXiv:2607.02636v1 Announce Type: new
- Abstract: Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster response, o…
- Robust model performance in such environments depends on large, continuously updated datasets
- However, training high-performing detectors typically requires centralizing aerial imagery, which raises privacy, regulatory, storage, and bandwidth challenges
Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02637v1 公告类型:新。
- 摘要:最近的生成模型可以生成高质量的合成图像,为数据匮乏的模型提供可扩展的训练数据。
- 开发这种潜力的现有方法通常涉及 1) 训练或微调生成器,或 2) 使用轻量级事后适应,例如即时工程或推理时间指导,使它们特定于生成器且专业知识密集。
- 我们研究一个补充问题:给定固定的生成图像池,是否可以纯粹通过选择信息丰富的子集来提高下游效用?
- EN 要点:
- arXiv:2607.02637v1 Announce Type: new
- Abstract: Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models
- Existing approaches to exploiting this potential typically involve 1) training or fine-tuning generators, or 2) using lightweight post-hoc adaptation like promp…
- We study a complementary question: given a fixed pool of generated images, can downstream utility be improved purely by selecting an informative subset
A Granularity-Aware EEG Feature Framework for Psychopathology Dimension Prediction
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02670v1 公告类型:新。
- 摘要:脑电图(EEG)提供了一种无创方法来检查维度精神病理学的神经生理学相关性,但跨脑电图范式和特征粒度的系统证据仍然有限。
- 在这里,我们开发了一个粒度感知的脑电图特征管道,将多尺度描述符组织成全局、区域和通道级别。
- 使用健康大脑网络 (HBN) 队列,我们评估了四种精神病理学维度的预测:p 因子、内化、外化和注意力问题,跨越四种脑电图范式。
- EN 要点:
- arXiv:2607.02670v1 Announce Type: new
- Abstract: Electroencephalography (EEG) offers a noninvasive approach for examining neurophysiological correlates of dimensional psychopathology, yet systematic…
- Here, we develop a granularity-aware EEG feature pipeline that organizes multi-scale descriptors into global, regional, and channel levels
- Using the Healthy Brain Network (HBN) cohort, we evaluate the prediction of four psychopathology dimensions: p-factor, internalizing, externalizing, and attenti…
LiNO: Lifting based multiresolution neural operator
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02715v1 公告类型:新。
- 摘要:最近,神经算子在直接从数据中学习微分方程的解算子方面显示出了有希望的结果。
- 该框架学习从参数字段到解决方案字段的功能映射,从而能够预测整个类别的解决方案而不是特定实例。
- 然而,现有运营商往往难以同时捕捉全局动态和精细结构。
- EN 要点:
- arXiv:2607.02715v1 Announce Type: new
- Abstract: Recently, neural operators have shown promising outcomes for learning solution operators of differential equations directly from data
- This framework learns a functional mapping from the parameter field to the solution field, enabling the prediction of an entire class of solutions rather than a…
- However, existing operators often struggle to capture both global dynamics and fine-scale structure simultaneously
Weighted Conformal Prediction for Lab-to-Track Thermal Transfer in EV Motorsport Powertrains
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02722v1 公告类型:新。
- 摘要:预测高性能电动汽车动力系统的热波动性很困难,因为在实验室外很少观察到内部温度,并且在实验室驾驶周期上校准的模型在针对实际负载部署时会失败。
- 我们使用保角预测来研究这个实验室到轨道的传输问题,提供无分布的不确定性界限。
- 我们实现了 Ensemble Batch Prediction Intervals(EnbPI;Xu & Xie,2021),这是一种用于自相关时间序列的留一引导集成共形方法,并根据真实的 CALCE 锂离子循环仪数据(A123 SP20 电池,FUDS 配置文件)对其进行校准。
- EN 要点:
- arXiv:2607.02722v1 Announce Type: new
- Abstract: Predicting thermal volatility in high-performance EV powertrains is difficult as internal temperatures are rarely observable outside the lab, and mode…
- We study this lab-to-track transfer problem using conformal prediction, offering distribution-free uncertainty bounds
- We implement Ensemble Batch Prediction Intervals (EnbPI; Xu & Xie, 2021), a leave-one-out bootstrap-ensemble conformal method for autocorrelated time series, an…
Out-of-Distribution Generalization of Risk Aversion in Language Models
- 发布时间:2026-07-07 12:00 北京时间
- 摘要:- arXiv:2607.02755v1 公告类型:新。
- 摘要:训练人工智能在资源方面规避风险可以在人工智能出现偏差时提供故障保护。
- 不一致但厌恶风险的人工智能往往更喜欢低风险、低回报的策略,比如合作,而不是高风险、高回报的策略,比如叛乱,从而限制任何不一致的负面影响。
- 但我们只能切实可行地训练人工智能在低风险的赌博中规避风险,而且只有当它们的风险厌恶泛化到天文数字般的高风险赌博时,我们才会安全。
- EN 要点:
- arXiv:2607.02755v1 Announce Type: new
- Abstract: Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned
- Misaligned but risk-averse AIs would tend to prefer low-risk, low-reward strategies like cooperation over high-risk, high-reward strategies like rebellion, limi…
- But we can only feasibly train AIs to be risk-averse on low-stakes gambles, and we will only be safe if their risk aversion generalizes to astronomically-high-s…