🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-29
- 类型
- ai-daily
- 字数
- 6065
- 阅读时长
- 29 min
2026-08-29 AI日更 | 当 Agent 获得一台真实电脑:AI 入口迁移进入执行环境竞争 链接到标题
Grok Bot 以原生虚拟机和完整电脑操作能力,展示了 Agent 从对话工具走向执行环境的新阶段。与此同时,组织上下文正在成为企业护城河,但高昂的模型调用成本、源码泄露与凭证安全,也提醒行业:Agent 落地的关键不只在能力,更在可控性与投入产出比。
📖 本期 Watch List 深度导读 链接到标题
今天最值得盯住的,是三条线。第一,AI 正从“会答”走向“可控、可评、可落地”:TreeGraft、Can a Model Catch Its Own Hallucinations for Free?、ElementCheck 和 NeuronFuzz 都在补推理效率、幻觉识别与安全评测这几块底座,适合技术团队重点读。第二,Agent 正在从概念走向真实场景,OpenClaw 的端侧助手、TelecomGPT-R1 的行业推理、以及 Natural-Language Policies to Executable Decisions,都是把自然语言接到具体业务决策上的尝试。第三,Stratechery 对“互联网热度与现实变化”的复盘,和 Susan Kare 的经典设计回顾放在一起看,很能帮助判断:真正留下来的,往往不是最喧闹的技术,而是能被稳定产品化、并进入日常工作流的能力。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:Z.ai Reveals Ox Alpha as Powerful GLM-5.3-Flash Model 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:43000
- 是什么事:Z.ai 公布了 Ox Alpha,并将其定位为一款性能较强的 GLM-5.3-Flash 模型。
- 为什么重要:这表明中文大模型阵营仍在通过新版本和新命名继续推进性能、速度与部署成本之间的竞争,这对模型迭代节奏和产品化落地都有参考意义。
- 讨论概况:X 上主要围绕它的真实能力、与同类轻量模型的对比、是否接近或超越主流开源/闭源方案,以及名称和版本关系是否足够透明展开讨论。
话题 2:Lindsay Clancy Murder Trial Jury Pauses Without Verdict 链接到标题
- 分类:AI · Other
- 概况:热度时间:,相关帖子数:5700
- 是什么事:美国林赛·克兰西谋杀案审判中,陪审团在未作出裁决的情况下暂停审议。
- 为什么重要:这类高关注案件虽不直接涉及 AI 技术,但对 AI 新闻理解、舆情监测和敏感事件摘要能力很重要,尤其考验系统对司法程序与事实边界的把握。
- 讨论概况:X 上主要在讨论陪审团为何未能达成一致、案件责任认定与精神健康因素的影响,以及审判进展对被告和受害者家庭意味着什么。
话题 3:Yen Weakens Toward 160 Despite Japan’s Record $96 Billion Intervention 链接到标题
- 分类:AI · Other
- 概况:热度时间:13 hours ago,相关帖子数:7800
- 是什么事:日元在日本创纪录动用约960亿美元干预后,仍继续逼近160兑1美元关口。
- 为什么重要:这件事反映出汇率干预在高利差和强美元环境下的边际效果有限,也会影响全球资产定价、跨境资金流动和AI产业所依赖的进口成本与融资环境。
- 讨论概况:X上的讨论主要围绕日本干预是否只是短期托底、美国利率预期和日美息差才是决定因素,以及日本是否会继续出手、160是否会成为新的心理防线。
话题 4:OpenAI Resets ChatGPT Work and Codex Quotas for Plus Users 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:1300
- 是什么事:OpenAI 为 ChatGPT Plus 用户重置了 Codex 和 ChatGPT Work 的使用额度,恢复了 5 小时日限和完整周配额。
- 为什么重要:这反映了 AI 产品在算力成本、配额管理和付费分层上的现实约束,也影响开发者对工具稳定性和工作流连续性的依赖。
- 讨论概况:X 上主要在讨论这次重置是否缓解了使用焦虑,以及日限是否过于严格;支持者认为有助于控制算力消耗,批评者则认为这会打断生产力,高价 Pro 用户暂时免受日限也引发了订阅公平性的争议。
话题 5:Salesforce Stock Soars on Earnings Beat and Anthropic AI Partnership 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:11000
- 是什么事:Salesforce 发布超预期财报,并宣布与 Anthropic 达成 AI 合作后,股价大幅上涨。
- 为什么重要:这反映出企业软件厂商正把生成式 AI 深度嵌入核心业务流程,且市场会直接奖励具备明确 AI 落地路径和商业化能力的公司。
- 讨论概况:X 上主要在讨论这笔 Anthropic 合作对 Salesforce AI 战略的实际价值、能否转化为持续收入增长,以及财报超预期究竟更多来自基本面改善还是 AI 叙事带来的估值重估。
话题 6:Grok Bot Users Can Now Share Custom AI Agent Templates 链接到标题
- 分类:AI · News
- 概况:热度时间:16 hours ago,相关帖子数:12000
- 是什么事:xAI 的 Grok Bot 现在支持用户分享自定义 AI 代理模板,便于他人直接复用或改造现成的代理配置。
- 为什么重要:这意味着 AI 代理的创建、传播和复用门槛进一步降低,可能加速生态扩散、应用落地和模板化分发,同时也放大安全、滥用和提示词泄露风险。
- 讨论概况:X 上的讨论主要集中在这项功能会提升效率还是带来新的治理问题:支持者看重模板共享带来的复用和协作,质疑者则担心劣质模板扩散、越权行为、提示注入以及被用于生成不受控代理。
话题 7:Tencent Releases Hy4 Preview, Top Open-Source AI for Coding and Productivity 链接到标题
- 分类:AI · News
- 概况:热度时间:17 hours ago,相关帖子数:6100
- 是什么事:腾讯混元发布了 Hy4 Preview,这是一款面向编程和生产力场景的开源模型预览版,并被部分讨论认为处于开源模型第一梯队。
- 为什么重要:这件事重要在于它反映了大厂继续加码高性能开源模型,尤其是代码与效率工具方向,这会直接影响开发者生态、开源模型竞争格局和企业落地选择。
- 讨论概况:X 上的焦点主要集中在它的编程能力、与现有开源模型的横向对比、是否真的达到“顶级”水平,以及它在实际推理速度、成本和可用性上的表现是否足以支撑宣传。
话题 8:City Cruise to 4-1 Win Over Palace with Haaland and Cherki Braces 链接到标题
- 分类:AI · Sports
- 概况:热度时间:8 hours ago,相关帖子数:72000
- 是什么事:曼城以4比1击败水晶宫,哈兰德和谢尔基各进两球完成梅开二度。
- 为什么重要:这是一场足球赛事结果,本身并不直接涉及人工智能;其进入AI分类更可能反映了平台对体育热搜的自动归类或标签错误。
- 讨论概况:X上的讨论主要集中在曼城的进攻表现、哈兰德与谢尔基的进球效率,以及这场胜利对球队排名和争冠形势的影响;由于未提供代表性推文,无法确认更具体的分歧。
话题 9:Ronaldo’s Winner Lifts Al-Nassr to Perfect 2-1 Win Over Al Taawoun 链接到标题
- 分类:AI · Sports
- 概况:热度时间:4 hours ago,相关帖子数:52000
- 摘要:Ronaldo’s Winner Lifts Al-Nassr to Perfect 2-1 Win Over Al Taawoun:
话题 10:Chelsea and Aston Villa Complete Martínez-Jackson Goalkeeper-Striker Swap 链接到标题
- 分类:AI · Sports
- 概况:热度时间:1 day ago,相关帖子数:195000
- 是什么事:X 上热议一则关于切尔西与阿斯顿维拉完成“Martínez-Jackson”门将与前锋互换的消息。
- 为什么重要:这类高热体育话题体现了 AI 驱动的信息聚合、标题生成和舆情放大的传播特征,也会影响 AI 在体育新闻摘要与事实核验中的可靠性评估。
- 讨论概况:讨论焦点主要集中在消息是否属实、这笔互换是否只是讽刺性标题,以及两队阵容调整对赛季表现的影响。
话题 11:Man Charged in $1.3M Romance Scam Posing as 49ers Player 链接到标题
- 分类:AI · Other
- 概况:热度时间:2 days ago,相关帖子数:109000
- 是什么事:一名男子被指控冒充旧金山49人队球员实施恋爱诈骗,涉案金额约130万美元。
- 为什么重要:该事件凸显了生成式AI和深度伪造技术可能放大身份冒充、情感操控和网络诈骗风险,推动平台身份验证、反欺诈检测和用户保护机制受到更多关注。
- 讨论概况:X上的讨论主要集中在受害者为何会被骗、名人身份冒充诈骗的普遍性、平台和执法部门应承担的责任,以及AI工具是否正在让此类骗局更难识别。
话题 12:Chelsea in Talks with Monaco for Lamine Camara Midfield Move 链接到标题
- 分类:AI · Sports
- 概况:热度时间:4 hours ago,相关帖子数:19000
- 是什么事:切尔西与摩纳哥就中场拉米内·卡马拉的转会进行了接触,外界将其视为球队中场补强及恩佐·费尔南德斯去留不确定下的潜在替代方案。
- 为什么重要:这类高关注转会会影响豪门阵容配置、球员估值和夏窗交易链条,也反映出俱乐部在中场重建中的决策节奏与市场博弈。
- 讨论概况:X 上主要在讨论卡马拉是否真会成为恩佐·费尔南德斯的替代者、转会费是否匹配他的能力,以及摩纳哥在留人和出售之间的权衡;同时也有人关注这笔交易是否会引发英超与法甲之间的连锁转会。
话题 13:In Love Forever Episode 11 Delivers Emotional Highs for Fans 链接到标题
- 分类:AI · Entertainment
- 概况:热度时间:2 hours ago,相关帖子数:13000
- 是什么事:《In Love Forever》第11集引发粉丝强烈情绪共鸣,被认为是本季的情感高点之一。
- 为什么重要:这类围绕热门剧集的讨论能体现 AI 对娱乐舆情、粉丝情绪和内容热度的识别能力,也有助于理解跨领域话题如何在平台上形成传播峰值。
- 讨论概况:X 上主要在讨论这一集的情感冲击、角色关系走向以及是否为后续剧情埋下关键转折;分歧集中在剧情是否足够动人、节奏是否合理,以及部分观众对角色选择的不同解读。
话题 14:Tesla Adds 79 Model Ys to Texas Robotaxi Fleet in One Day 链接到标题
- 分类:AI · News
- 概况:热度时间:1 day ago,相关帖子数:5100
- 是什么事:特斯拉被指在一天内向德州的 Robotaxi 车队新增了 79 辆 Model Y,引发外界对其自动驾驶运营进展的关注。
- 为什么重要:这件事关系到自动驾驶商业化的落地速度、车队规模化调度能力,以及特斯拉在 AI 驱动出行服务上的真实推进程度。
- 讨论概况:X 上的讨论主要集中在这是否意味着 Robotaxi 已进入更实质的试运营阶段、79 辆车的新增是否具有运营意义,还是更多只是车辆配置和登记层面的变化。
话题 15:Wang Yibo Hits Shanghai Track in Custom Alo Yoga Livery 链接到标题
- 分类:AI · Entertainment
- 概况:热度时间:20 hours ago,相关帖子数:3900
- 是什么事:王一博在上海赛道相关活动中,出现了定制 Alo Yoga 涂装的内容并在 X 上引发关注。
- 为什么重要:这类话题显示 AI 与娱乐、品牌营销和内容分发的结合正在放大明星事件的传播效率,也反映了生成式内容和推荐机制在跨圈层扩散中的作用。
- 讨论概况:X 上的讨论主要集中在活动本身的视觉效果、王一博与品牌合作的商业价值,以及这类明星体育化、品牌化内容是否只是营销包装,还是能带来更强的互动与传播。
💡 大佬观点(Influencer Insights) 链接到标题
AI X 平台情报简报 链接到标题
以下是基于 2026 年 8 月 28 日前后 AI 领域 Influencers 动态的深度分析。
1. 今日核心关注:Agent 执行环境的军备竞赛与“计算即接口” 链接到标题
过去 24 小时,大佬们的注意力从单纯的模型能力转向了 “AI 如何真正操作世界”。核心战场聚焦于 Agent 的操作系统(OS)与浏览器沙箱。
Grok Bot 成为现象级 Agent 硬件: 这无疑是今日最大的热点。随着 X Premium+ 放开权限,多位 Influencer 第一时间深度体验。
- @vista8 和 @Pluvio9yte 给予了极高评价,认为其是目前体验最好的 Computer Use 产品。其核心优势在于原生级的虚拟机环境(Debian 13, 8核, 16GB 内存)。@vista8 指出,Grok Bot 拥有完整的 GUI 环境和软件安装权限,本质上是给了 Agent “一台真实的电脑”,这使得榜单监控、自动化登录测试等复杂长程任务变得极其丝滑。
- 技术亮点:@Pluvio9yte 强调了其 “防沉迷式决策” (不确定时停下来给选项)和密钥安全管理(不暴露给模型)。@vista8 则利用其创建“Bot 工厂”,安排 Bot 创建更多 Bot 并组织群聊协作。
- 争议与漏洞:@dotey 转发了源代码在打包时因 Source Maps 未关闭而被完整逆向还原的突发新闻,这给 Grok Bot 的安全性带来了短暂的阴云。
- 浏览器自动化新宠 ego lite:@Pluvio9yte 提到,在轻量级浏览器自动化方面,ego lite 因可以迁移本地 Chrome 登录态(Cookies/密码等)且允许 Agent 在独立空间干活而变得更丝滑,优于传统的 Agent-Browser。
Codex 与 PC 的深度融合:@vista8 体验了 Codex 的“电脑使用回顾”功能,AI 基于一周的屏幕操作生成了极为精准的人格化 Recap 报告,显示 Agent 正在成为个人行为的精密记录者与分析者。
2. 独特观点与行业前瞻 链接到标题
Agent 重塑企业护城河与组织架构(@dotey 的深度洞察)
- 入口迁移论:@dotey 在分析“豆包工作”及飞书融合时提出,应用正退居后台,Agent 成为工作流的新入口。过去是“人找应用”,现在是“Agent 调用应用,人验收结果”。
- 上下文即壁垒:通用 Agent 能力易被追平,但组织长期沉淀的上下文(会议、文档、脉络)是别人拿不走的护城河。谁能访问最完整的组织数据,谁的 Agent 就最懂业务。
- 外科手术团队复辟:结合《人月神话》的“外科手术团队”模式,@dotey 指出当前 “1个决策者 + 多个Agent” 的模式正在复刻这一经典。人类负责定义问题和做判断,AI 负责外围执行,这可能是当前架构下的效率极限。
AI 编程投入产出比的冷静审视(@ruanyf)
- 尽管 AI 编程火爆,@ruanyf 换算了一笔账:若像 OpenAI 员工一样放开使用顶级模型,单人年成本可达 1 亿人民币,即使换用便宜的国产开源模型仍需二三百万。这暴露出无限量使用 AI 编程远比真人员工昂贵的现实。
- @gefei55 也抛出犀利观点:“让程序员指挥 AI 写代码,可能是人类走过的一段弯路”,暗示未来应用的产品经理(懂业务逻辑者)直接驱动 AI 生成,而非由传统程序员做中转。
国产模型大爆发与“数据脏活”论(@vista8)
- 面对国产模型的“寒武纪大爆发”,@vista8 引用腾讯混元 Hy4 preview 的博文指出,真相在于 “把数据搞脏”(深度参与高质量专家数据共建)。比起算法上的傲慢,一线标注和高质数据才是模型灵魂。
产品设计的 Game-like 趋势(@nishuang)
- 在设计心理学上,@nishuang 区分了 “游戏感(Game-like design)” 与 “游戏化(Gamification)”。前者激发内啡肽产生持久的愉悦感,后者通过多巴胺刺激行苦役。他指出聪明的新一代 AI 产品会趋向于游戏感,因为用户已对简单的奖励系统脱敏。
3. 推荐工具与前沿资源 链接到标题
AI 视频与营销内容制作
- Topview Motion Studio(驱动:Seedance 2.5):@Pluvio9yte 和 @AI_Jasonyu 强推。主打极低成本产出高品质动效(例如 $3 产出 $3000 效果的 AE 动画)。@AI_Jasonyu 用其花了 1.5 小时制作出千万播放量的“孙哥瓜”视频。
- 本地视频工作流:@Pluvio9yte 推荐使用 ComfyUI MCP + Codex 搭建完全本地且免费的 Agent 工作流,用自然语言操控 ComfyUI,平替商业化平台(如小云雀)。
Agent 开发与基础设施
- Grok Bot:仅限 X Premium+ 会员,拥有原生虚拟机环境的 Computer Use Agent。
- ego lite:目前最丝滑的网页自动化工具,完美迁移本地缓存,支持 Codex/Claude Code 调用。
- OpenConnector(密码连接网关):@ruanyf 推荐,防止 Agent 泄露密码到上下文,统一管理 10000+ 个应用的授权,解决企业级 Agent 落地的安全痛点。
**AI “乐高/模版” **
- AI 游戏生成*:@gearzero_alaya 的 Gear Zero @AI_Jasonyu 和 @Pluvio9yte 用其实现了“一句话做游戏”,尤其以用皮影戏游戏解释“什么是 Harness/Agent/Skil”的案例极具创意。
- AI 语言学习 CapWords:@nishuang 强推,用“宝可梦搜集宝贝”的游戏感方式学单词,结合了 AI 抠图与情景记忆。
开源与基础模型
- RedSkill / 小红技能市场:@ruanyf 指出,小红书开始允许上传 AI Skill,并支持一键复制,试图打造“宠物用的 GitHub + 生活社区”。
- TTS 模型聚合器:@vista8 推荐了一家 YC 投资的 TTS 领域 OpenRouter,注册即送 100 美元,适合寻找免费语音合成方案的开发者。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 34 条更新
Y Combinator Podcast (B_intro+search) 链接到标题
- Susan Kare: Designing Icons & Graphics For the Original Mac
- 发布时间:2026-08-29 01:58 北京时间
- 摘要:【待翻译】- You’ve probably already heard all about OpenClaw (formerly Clawdbot/Moltbot).
- The viral sensation is an open-source AI assistant that runs on your own device, connects with messaging apps you already use, and goes beyond chat to actually execute tasks like managing your email, calendars, files, workflows, and more.
- Now meet the man behind it.
- YC’s Raphael Schaad sat down with Peter Steinberger, the creator of OpenClaw, to discuss the “aha” moment behind the viral personal AI agent, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
- EN 要点:
- Susan Kare joined Apple as an art history PhD who barely knew anything about computers
- She went on to design many of the icons, typefaces, and symbols that helped make the original Macintosh feel understandable and human, defining a visual languag…
Stratechery by Ben Thompson (A_full) 链接到标题
- 2026.35: Internet Hype and Real World Change
- 发布时间:2026-08-29 01:00 北京时间
- 摘要:【待翻译】- (Photo by Natalie Behring/Getty Images).
- Welcome back to This Week in Stratechery!
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone.
- Additionally, you have complete control over what we send to you.
- On that note, here were a few of our favorites this week.
- EN 要点:
- (Photo by Natalie Behring/Getty Images)
- Welcome back to This Week in Stratechery
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
- Additionally, you have complete control over what we send to you
OpenAI Blog (A_full) 链接到标题
- Supporting Thailand’s next generation of AI startups
- 发布时间:2026-08-28 10:00 北京时间
- 摘要:【待翻译】- Today in Bangkok, OpenAI and Thailand’s Ministry of Higher Education, Science, Research and Innovation (MHESI) announced a new accelerator to help Thai startups turn promising prototypes into products ready for real-world use and growth.
- The OpenAI x MHESI AI Accelerator brings together ten startups working across health, wellness, and education.
- It marks OpenAI’s first public-private partnership with the Thai government focused on supporting local startups, and is being delivered with partners including the National Innovation Agency (NIA), Mahidol University, and Techsauce.
- A compelling AI demonstration is only the beginning.
- Building a product that people can rely on is much harder.
- EN 要点:
- OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into trusted products.
Two Minute Papers (B_intro+search) 链接到标题
- This Free AI Just Caught The Billion Dollar Giants
- 发布时间:2026-08-28 17:44 北京时间
- 摘要:【待翻译】- ❤️ Check out Weights & Biases and sign up for a free demo here:.
- 📝 The paper and Qwen3.8-Flash-Next are available here:.
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
- This Free AI Just Caught The Billion Dollar Giants.
- EN 要点:
- ❤️ Check out Weights & Biases and sign up for a free demo here:
- 📝 The paper and Qwen3.8-Flash-Next are available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…
ArXiv cs.AI (B_intro+search) 链接到标题
EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26107v1 Announce Type: new.
- Abstract: Predicting students’ academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes.
- However, existing models often suffer from limited early detection capability and insufficient interpretability, leading to a “black-box” trust crisis that hinders their adoption in real-world pedagogical settings.
- To address these challenges, we propose EduRiskX, a neuro-symbolic framework that integrates a temporal Transformer-based predictor with F-Logic symbolic reasoning.
- EN 要点:
- arXiv:2608.26107v1 Announce Type: new
- Abstract: Predicting students’ academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes
- However, existing models often suffer from limited early detection capability and insufficient interpretability, leading to a “black-box” trust crisis that hind…
- To address these challenges, we propose EduRiskX, a neuro-symbolic framework that integrates a temporal Transformer-based predictor with F-Logic symbolic reason…
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26109v1 Announce Type: new.
- Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use.
- Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension because they separate data interpretation, guideline checking, and final explanation.
- This revised feasibility study preserves the original standalone-versus-agentic comparison while making the main clinical findings more explicit.
- EN 要点:
- arXiv:2608.26109v1 Announce Type: new
- Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for b…
- Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension because they separate data interpretation, guidelin…
- This revised feasibility study preserves the original standalone-versus-agentic comparison while making the main clinical findings more explicit
Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26111v1 Announce Type: new.
- Abstract: Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batteries across electric vehicles, grid storage, and consumer electronics.
- Conventional BPHM approaches, including physics-based models and task-centric deep learning methods, face challenges in computational efficiency and parameterization, cross-domain generalization, dependence on extensive labeled run-to-failure data, and model interpretability.
- Recent Large Models (LMs), built upon Transformer architectures and self-supervised pre-training, offer a transformative new paradigm to overcome these long-standing bottlenecks.
- EN 要点:
- arXiv:2608.26111v1 Announce Type: new
- Abstract: Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batteries across electri…
- Conventional BPHM approaches, including physics-based models and task-centric deep learning methods, face challenges in computational efficiency and parameteriz…
- Recent Large Models (LMs), built upon Transformer architectures and self-supervised pre-training, offer a transformative new paradigm to overcome these long-sta…
PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26113v1 Announce Type: new.
- Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications.
- PICasso couples a structured NL -> YAML -> GDS generation pipeline with PDK aware knowledge injection, automated placement and routing, DRC/LVS validation, and SAX-based photonic simulation.
- To systematically evaluate AI-driven photonic design, we introduce PIC-Set, a benchmark of 36 parameterized PIC design tasks spanning core photonic primitives and multi-component circuits.
- EN 要点:
- arXiv:2608.26113v1 Announce Type: new
- Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natur…
- PICasso couples a structured NL -> YAML -> GDS generation pipeline with PDK aware knowledge injection, automated placement and routing, DRC/LVS validation, and…
- To systematically evaluate AI-driven photonic design, we introduce PIC-Set, a benchmark of 36 parameterized PIC design tasks spanning core photonic primitives a…
CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26114v1 Announce Type: new.
- Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-based constraints.
- Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce numerically incorrect yet plausible answers when solving multi-step financial calculations.
- To address this limitation, we introduce CIFQA (Calculation-Intensive Financial Query Answering), a deterministic tool-grounded multi-agent LLM framework for financial question answering.
- EN 要点:
- arXiv:2608.26114v1 Announce Type: new
- Abstract: Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas, and rule-b…
- Although Large Language Models (LLMs) perform strongly on natural language tasks, they often produce numerically incorrect yet plausible answers when solving mu…
- To address this limitation, we introduce CIFQA (Calculation-Intensive Financial Query Answering), a deterministic tool-grounded multi-agent LLM framework for fi…
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26116v1 Announce Type: new.
- Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full simulation, and observe the outcome, without intervening during execution.
- We introduce a closed-loop framework based on autotelic reinforcement learning, in which an agent autonomously samples diverse goals and learns a goal-conditioned policy to intervene in a complex system through minimal, local perturbations.
- We instantiate this framework on Lenia, a continuous cellular automaton known for life-like self-organizing patterns, in an agentic system we call CARL, and demonstrate three capabilities.
- EN 要点:
- arXiv:2608.26116v1 Announce Type: new
- Abstract: Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execute a full si…
- We introduce a closed-loop framework based on autotelic reinforcement learning, in which an agent autonomously samples diverse goals and learns a goal-condition…
- We instantiate this framework on Lenia, a continuous cellular automaton known for life-like self-organizing patterns, in an agentic system we call CARL, and dem…
The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26134v1 Announce Type: new.
- Abstract: Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equally to on-device forecasting for mission-critical edge environments, including military systems.
- However, this paper identifies the Accuracy-Efficiency Paradox: high-precision energy forecasting models can ironically trigger a net energy deficit.
- This stems from both edge AI’s inference energy consumption and battery aging.
- EN 要点:
- arXiv:2608.26134v1 Announce Type: new
- Abstract: Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equally to on-device fore…
- However, this paper identifies the Accuracy-Efficiency Paradox: high-precision energy forecasting models can ironically trigger a net energy deficit
- This stems from both edge AI’s inference energy consumption and battery aging
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26145v1 Announce Type: new.
- Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing.
- Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions.
- Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards.
- EN 要点:
- arXiv:2608.26145v1 Announce Type: new
- Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the…
- Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions
- Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26149v1 Announce Type: new.
- Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems.
- Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality categorical features, complex inter-table dependencies, and repeated temporal observations.
- We introduce the Relational Hypergraph Transformer (RHT), a unified architecture that represents relational databases as hypergraphs, learns pentadimensional embeddings (PentE), and performs sparse relational attention with complexity proportional to the average relational degree rather than the square of the number of entities.
- EN 要点:
- arXiv:2608.26149v1 Announce Type: new
- Abstract: Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems
- Relational data combine several sources of complexity, including large data volume, high-dimensional variables, high-cardinality categorical features, complex i…
- We introduce the Relational Hypergraph Transformer (RHT), a unified architecture that represents relational databases as hypergraphs, learns pentadimensional em…
Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26150v1 Announce Type: new.
- Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs).
- This study reports an LLM pipeline development for extracting model-relevant information from 536 peer-reviewed agent-based modeling papers.
- We compare the results with those of a human-conducted SLR.
- EN 要点:
- arXiv:2608.26150v1 Announce Type: new
- Abstract: Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, inc…
- This study reports an LLM pipeline development for extracting model-relevant information from 536 peer-reviewed agent-based modeling papers
- We compare the results with those of a human-conducted SLR
ArXiv cs.CL (B_intro+search) 链接到标题
TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26112v1 Announce Type: new.
- Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm.
- Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length.
- However, existing tree-structured methods use a single drafter for all drafting steps, creating a dilemma: a smaller drafter is fast but yields lower-quality trees, whereas a larger drafter improves tree quality but suffers from high latency.
- EN 要点:
- arXiv:2608.26112v1 Announce Type: new
- Abstract: Speculative decoding accelerates large language model inference through a draft-then-verify paradigm
- Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, increasing the accepted length
- However, existing tree-structured methods use a single drafter for all drafting steps, creating a dilemma: a smaller drafter is fast but yields lower-quality tr…
ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26118v1 Announce Type: new.
- Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline.
- However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results.
- We propose ElementCheck, a complexity-aware framework that verifies long-form outputs via sentence elements.
- EN 要点:
- arXiv:2608.26118v1 Announce Type: new
- Abstract: Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline
- However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results
- We propose ElementCheck, a complexity-aware framework that verifies long-form outputs via sentence elements
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26119v1 Announce Type: new.
- Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this behavior, has received less attention than the related question of detecting fallacies in existing text.
- We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red herring), seven prompt framings, and 80 claims spanning four controversy levels.
- Refusal is governed primarily by request structure rather than claim content.
- EN 要点:
- arXiv:2608.26119v1 Announce Type: new
- Abstract: Whether large language models can be prompted to generate rhetorical fallacies on demand, and whether current safety post-training constrains this beh…
- We close this gap with DeflectBench, evaluating 23,990 generations from four frontier models across three deflection strategies (whataboutism, ad hominem, red h…
- Refusal is governed primarily by request structure rather than claim content
Recipes for Steering and Scaling LLMs via Sampling
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26120v1 Announce Type: new.
- Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization.
- While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient.
- In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling.
- EN 要点:
- arXiv:2608.26120v1 Announce Type: new
- Abstract: Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization
- While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient
- In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26121v1 Announce Type: new.
- Abstract: Large language models state false facts as fluently as true ones, yet a model often “knows” internally when it is on shaky ground: the probability it assigns to its own answer tends to dip on the facts it gets wrong.
- The usual way to act on this, teaching a model to abstain rather than guess, requires a labelled dataset of right and wrong answers.
- We ask whether the model’s own confidence, which is free and needs no labels, can do that job instead.
- EN 要点:
- arXiv:2608.26121v1 Announce Type: new
- Abstract: Large language models state false facts as fluently as true ones, yet a model often “knows” internally when it is on shaky ground: the probability it…
- The usual way to act on this, teaching a model to abstain rather than guess, requires a labelled dataset of right and wrong answers
- We ask whether the model’s own confidence, which is free and needs no labels, can do that job instead
Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26123v1 Announce Type: new.
- Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype.
- We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales.
- We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 generation requests spanning three prompt types per tradition - generic, culturally specific, and regional-language.
- EN 要点:
- arXiv:2608.26123v1 Announce Type: new
- Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising con…
- We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic,…
- We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 gene…
Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26124v1 Announce Type: new.
- Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly evolving, and inherently open-ended.
- Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM agents lack the reliability and auditability required for financial decisions.
- We present a production-grade LLM-powered pricing system with a strict decision boundary: LLMs perform structured extraction and bounded policy/path selection, while all numeric pricing, including total-price computation, is executed deterministically.
- EN 要点:
- arXiv:2608.26124v1 Announce Type: new
- Abstract: Pricing automation in large-scale tourism is challenging because travel orders are highly unstructured, while pricing policies are complex, rapidly ev…
- Traditional rule engines are brittle and costly to maintain, whereas unconstrained LLM agents lack the reliability and auditability required for financial decis…
- We present a production-grade LLM-powered pricing system with a strict decision boundary: LLMs perform structured extraction and bounded policy/path selection,…
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26125v1 Announce Type: new.
- Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation.
- Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context.
- We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability.
- EN 要点:
- arXiv:2608.26125v1 Announce Type: new
- Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation
- Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context
- We propose a \emph{training-time} explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performanc…
TelecomGPT-R1: A Unified Open-Source Reasoner for the Telecom Stack
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26126v1 Announce Type: new.
- Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint grounding in normative specifications, operational telemetry, vendor-specific fault evidence, and exact RF/network calculations.
- However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while domain-specific telecom LLMs remain limited in structured, multi-step reasoning.
- To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard.
- EN 要点:
- arXiv:2608.26126v1 Announce Type: new
- Abstract: Telecommunications is a high-leverage domain for large language model (LLM)-based reasoning because routine engineering workflows require joint ground…
- However, current LLM integration in telecom remains bottlenecked by a two-sided capability gap: generic reasoners often lack telecom-specific grounding, while d…
- To bridge this gap, we release TelecomGPT-R1-9B, a unified open-source telecom reasoner that ranks top-performing on the GSMA open telco leaderboard
FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26129v1 Announce Type: new.
- Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments.
- We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal.
- Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific domains (biology, chemistry, neuroscience, physics, and earth science), capturing the full iterative structure of scientific validation: initial referee reports, author point-by-point responses, and updated reviewer assessments.
- EN 要点:
- arXiv:2608.26129v1 Announce Type: new
- Abstract: Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ab…
- We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal
- Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific doma…
ArXiv cs.LG (B_intro+search) 链接到标题
SLM-Conditioned Hierarchical Relation Routing for Labeled Property Graph Learning
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26132v1 Announce Type: new.
- Abstract: Labeled property graphs combine relational structure with heterogeneous textual and categorical properties attached to both nodes and relationships.
- Conventional graph neural networks typically represent these properties as static feature vectors, limiting their ability to determine which semantic evidence should influence message propagation for a particular prediction target.
- We propose SLM-Conditioned Hierarchical Relation Routing, an architecture that integrates a small language model directly into graph message selection.
- EN 要点:
- arXiv:2608.26132v1 Announce Type: new
- Abstract: Labeled property graphs combine relational structure with heterogeneous textual and categorical properties attached to both nodes and relationships
- Conventional graph neural networks typically represent these properties as static feature vectors, limiting their ability to determine which semantic evidence s…
- We propose SLM-Conditioned Hierarchical Relation Routing, an architecture that integrates a small language model directly into graph message selection
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26222v1 Announce Type: new.
- Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks.
- Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model response to evaluate its attack effectiveness.
- This process is expensive and, more importantly, provides only sparse guidance on strongly aligned models, where most candidates are rejected with the same failure outcome.
- EN 要点:
- arXiv:2608.26222v1 Announce Type: new
- Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks
- Existing automated testing methods, however, largely rely on response-level feedback: each candidate prompt typically requires generating a target-model respons…
- This process is expensive and, more importantly, provides only sparse guidance on strongly aligned models, where most candidates are rejected with the same fail…
Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26233v1 Announce Type: new.
- Abstract: Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating deployment on constrained edge hardware with field-programmable gate arrays (FPGAs) and microcontrollers.
- Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill-suited to binarized representations and rarely translate into meaningful hardware savings.
- We introduce a PyTorch-based, research-oriented framework that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networks.
- EN 要点:
- arXiv:2608.26233v1 Announce Type: new
- Abstract: Extreme compression of deep neural networks, up to full binarization, dramatically reduces memory footprint and arithmetic complexity, facilitating de…
- Although combining binarization with pruning promises additional efficiency gains, existing pruning strategies are ill-suited to binarized representations and r…
- We introduce a PyTorch-based, research-oriented framework that incorporates freezing and pruning mechanisms for designing and optimizing binarized neural networ…
Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26288v1 Announce Type: new.
- Abstract: Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogonalizing its momentum with a few Newton-Schulz iterations.
- Existing theory either replaces this iteration with the exact polar factor it approximates, or treats its finite depth as an approximation error, and thus the iteration Muon actually runs can only hurt the guarantees.
- We show that finite Newton-Schulz can instead be beneficial for nonsmooth nonconvex optimization.
- EN 要点:
- arXiv:2608.26288v1 Announce Type: new
- Abstract: Muon has emerged as a strong optimizer for the matrix-valued parameters in large language model pretraining, approximately orthogonalizing its momentu…
- Existing theory either replaces this iteration with the exact polar factor it approximates, or treats its finite depth as an approximation error, and thus the i…
- We show that finite Newton-Schulz can instead be beneficial for nonsmooth nonconvex optimization
Algebraic Multigrid Acceleration for Efficient Label Spreading
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26309v1 Announce Type: new.
- Abstract: Modern machine learning models rely on large amounts of labeled data.
- However, manual annotation of large-scale datasets is expensive and time-consuming.
- Label spreading is a semi-supervised learning technique that addresses this challenge by propagating information from a few labeled examples to a larger pool of unlabeled data.
- EN 要点:
- arXiv:2608.26309v1 Announce Type: new
- Abstract: Modern machine learning models rely on large amounts of labeled data
- However, manual annotation of large-scale datasets is expensive and time-consuming
- Label spreading is a semi-supervised learning technique that addresses this challenge by propagating information from a few labeled examples to a larger pool of…
Privacy Without Regret: Differentially Private Inference-Time Alignment
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26324v1 Announce Type: new.
- Abstract: Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward hacking, in which the selected response exploits errors in the proxy reward model, and the absence of any privacy protection for the sensitive human preference data used to train that reward model.
- We show that a single intervention-adding calibrated noise to reward scores before selection-resolves both.
- Our first result, Private Best-of-N (PrivBoN), establishes that Gumbel noise at an appropriate scale simultaneously provides $\epsilon$-differential privacy and implements KL-regularized alignment.
- EN 要点:
- arXiv:2608.26324v1 Announce Type: new
- Abstract: Best-of-N (BoN) sampling is the simplest and most widely deployed inference-time alignment strategy, but it suffers from two distinct problems: reward…
- We show that a single intervention-adding calibrated noise to reward scores before selection-resolves both
- Our first result, Private Best-of-N (PrivBoN), establishes that Gumbel noise at an appropriate scale simultaneously provides $\epsilon$-differential privacy and…
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26332v1 Announce Type: new.
- Abstract: Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capability benchmarks that reveal little about operational behavior after deployment.
- We present Operational Embedding (OpEmbed), a framework for learning compact operational fingerprints of LLM cloud services from structured, privacy-preserving support-case metadata, without using case text.
- OpEmbed aggregates model–time windows into an eight-channel operational signature and learns a low-dimensional representation via temporal contrastive learning, cross-view reconstruction, and generational-ordinality regularization.
- EN 要点:
- arXiv:2608.26332v1 Announce Type: new
- Abstract: Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capability benchmarks tha…
- We present Operational Embedding (OpEmbed), a framework for learning compact operational fingerprints of LLM cloud services from structured, privacy-preserving…
- OpEmbed aggregates model–time windows into an eight-channel operational signature and learns a low-dimensional representation via temporal contrastive learning…
CG4AI: A Column Generation Framework for Training AI Models Under Constraints
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26375v1 Announce Type: new.
- Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will satisfy predefined rules or constraints on its outputs.
- In many real-world applications, ranging from autonomous systems to network routing, such guarantees are essential.
- We propose CG4AI, a framework that builds a convex combination of AI models while enforcing linear constraints on the combined output.
- EN 要点:
- arXiv:2608.26375v1 Announce Type: new
- Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will satisfy predefined r…
- In many real-world applications, ranging from autonomous systems to network routing, such guarantees are essential
- We propose CG4AI, a framework that builds a convex combination of AI models while enforcing linear constraints on the combined output
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26423v1 Announce Type: new.
- Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier’s confident decisions can be trusted.
- This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, (ii) locating a relatively small set of latent support vectors (~ 29% of total training examples) representing influential prompts for identifying tokens that alter the classifier’s predicted labels, and (iii) utilizing such tokens and their associated attack magnitudes for constructing a diagnostic taxonomy.
- This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier’s decision; flag Heuristic Bias and Heuristic Override cases; route Insufficient Context cases for further human/safety review.
- EN 要点:
- arXiv:2608.26423v1 Announce Type: new
- Abstract: This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies whic…
- This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is emp…
- This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier’s decision; flag…
FedCMAPSS: A Benchmark for Federated Learning in Remaining Useful Life Estimation
- 发布时间:2026-08-28 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.26433v1 Announce Type: new.
- Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL) estimation models is often limited by the scarcity of run-to-failure data.
- While federated learning offers a promising paradigm to collaboratively train predictive models without sharing sensor data, research efforts have operated so far in the absence of a common evaluation framework.
- To address this gap, this paper introduces FedCMAPSS, a benchmark for federated RUL estimation based on the commonly-used NASA C-MAPSS dataset.
- EN 要点:
- arXiv:2608.26433v1 Announce Type: new
- Abstract: Data-driven prognostics and health management has emerged as a key enabler for Industry 4.0, yet the development of robust remaining useful life (RUL)…
- While federated learning offers a promising paradigm to collaboratively train predictive models without sharing sensor data, research efforts have operated so f…
- To address this gap, this paper introduces FedCMAPSS, a benchmark for federated RUL estimation based on the commonly-used NASA C-MAPSS dataset