🤖 AI 速览

今天的主线是智能体从概念走向系统化落地:Gemini 将 computer use 下放到轻量模型,评测与安全框架同步升温;OpenAI 与 Broadcom 推理芯片显示模型竞争正深入硅层。端侧模型、Coding Agent 和组织协作工具也在加速成熟。
📋 文章元数据
发布时间
2026-06-25
类型
ai-daily
字数
3530
阅读时长
17 min

2026-06-25 AI日更 | 智能体进入基础设施竞争:从浏览器操作到推理芯片 链接到标题

今天的主线是智能体从概念走向系统化落地:Gemini 将 computer use 下放到轻量模型,评测与安全框架同步升温;OpenAI 与 Broadcom 推理芯片显示模型竞争正深入硅层。端侧模型、Coding Agent 和组织协作工具也在加速成熟。

📖 本期 Watch List 深度导读 链接到标题

今天最值得关注的主线,是“智能体从概念走向基础设施”。Mirendil 关于自我加速 AI 的访谈,适合关注前沿研究组织形态的人读;Google DeepMind 将 computer use 引入 Gemini 3.5 Flash,则显示浏览器/桌面操作能力正在被下放到更轻量模型。与之呼应,RIFT-Bench、AgenticInterpBench 和多篇关于 agent 边界、安全与可解释性的论文,都在提醒我们:智能体能力扩张越快,评测、红队和责任边界越需要系统化。

第二条线是“AI 全栈化继续加速”。OpenAI 与 Broadcom 推出面向 LLM 推理的芯片,且强调九个月完成设计到生产,值得基础设施团队重点关注:模型公司正在把竞争延伸到硅层、数据中心和能耗效率。

最后,推理与强化学习仍是研究热点。从 SGPO 的策略蒸馏,到有益模型的长期对齐、多目标推荐和安全多智能体 RL,今天的论文共同指向一个问题:AI 不只要会完成任务,还要在复杂约束下稳定、可解释、可泛化地行动。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Loop Engineering Emerges as New Way to Automate AI Coding Agents 链接到标题

  • 分类:AI · Other
  • 概况:热度时间:13 hours ago,相关帖子数:3600
  • 是什么事:X 平台热议“Loop Engineering”这一新思路,即通过循环式反馈、评估与修正机制来自动化 AI 编码代理的工作流程。
  • 为什么重要:它的重要性在于可能显著提升 AI 编码代理的稳定性、可控性和任务完成率,推动从“单次生成”走向“可持续执行”的工程化智能体。
  • 讨论概况:当前讨论主要集中在这种循环框架是否真的能提升复杂编码任务表现、与传统 agent 工作流相比的优势,以及它会不会成为未来 AI 编程自动化的标准范式。

话题 2:Developer Migrates $36K AI Agents from OpenClaw to Hermes for Better Reliability 链接到标题

  • 分类:AI · News
  • 概况:热度时间:11 hours ago,相关帖子数:83
  • 摘要:Developer Migrates $36K AI Agents from OpenClaw to Hermes for Better Reliability:

话题 3:Anthropic Accuses Alibaba-Linked Group of Massive Claude AI Distillation Attack 链接到标题

  • 分类:AI · News
  • 概况:热度时间:7 hours ago,相关帖子数:4700
  • 是什么事:Anthropic 指控一个与阿里巴巴有关联的组织大规模利用 Claude 输出进行模型蒸馏,以复制或提升自身 AI 系统能力。
  • 为什么重要:此事凸显前沿大模型的知识产权、模型安全与 API 滥用风险,也可能加剧 AI 公司对训练数据、访问控制和跨境技术竞争的监管压力。
  • 讨论概况:X 上讨论集中在 Anthropic 指控证据是否充分、模型蒸馏是否应被视为违规或行业常态、美国与中国 AI 竞争背景下是否存在政治化解读,以及平台应如何平衡开放访问与防止能力复制。

话题 4:China’s GLM-5.2 Tops Coding Benchmarks at Fraction of U.S. AI Costs 链接到标题

  • 分类:AI · News
  • 概况:热度时间:9 hours ago,相关帖子数:2500
  • 是什么事:中国的 GLM-5.2 被报道在代码基准测试中表现领先,同时训练或使用成本仅为美国同类 AI 的一小部分。
  • 为什么重要:这表明在 AI 尤其是编程模型领域,性能优势不一定只依赖更高算力和更高成本,可能改变行业对模型效率、研发投入和竞争格局的判断。
  • 讨论概况:X 上的讨论焦点主要集中在其基准成绩是否足以代表真实编程能力、成本优势是否可复制,以及这是否意味着中美 AI 竞争正在从“堆算力”转向“拼效率”与工程化能力。

话题 5:Users Build AI-Powered Second Brains with Claude and Obsidian 链接到标题

  • 分类:AI · News
  • 概况:热度时间:8 hours ago,相关帖子数:144
  • 是什么事:X 上围绕 Claude 与 Obsidian 构建“AI 第二大脑”的教程、插件和工作流讨论升温,用户分享无需编程即可让 AI 整理、关联和检索个人知识库的方法。
  • 为什么重要:这反映出大模型正从问答工具转向个人知识管理和生产力基础设施,使非技术用户也能低门槛构建可持续积累的智能工作流。
  • 讨论概况:讨论焦点集中在 Obsidian 笔记如何结构化以便 AI 识别模式、Claude 插件和 GitHub 项目的实用性,以及这种工具是否真正提升思考能力,还是只是制造新的效率幻觉和信息整理负担。

今日 X 上的 AI 舆情小结 链接到标题

今天 X 上的 AI 讨论主线,整体从“模型能力本身”转向“如何把模型变成可持续、可控、可落地的生产力系统”:一边是 Loop Engineering 这类循环式反馈框架,强调让编码代理从一次性生成走向反复评估与修正;另一边是 Claude + Obsidian 的“第二大脑”工作流,推动 AI 深入个人知识管理。共识是大家都认可 AI 正在从聊天工具演进为工程化和日常化基础设施,尤其在编程、检索和整理任务上潜力很大,但分歧在于这些方案到底是真提升了稳定性与思考质量,还是主要制造了效率幻觉。与此同时,围绕 Anthropic 指控蒸馏与 GLM-5.2 低成本高性能的争论,也把焦点拉回到模型竞争的底层逻辑:到底是靠算力、数据和封闭访问壁垒,还是靠效率、工程化和工作流创新取胜。潜在风险则集中在三个方面:模型能力被复制或滥用引发的知识产权与安全问题、基准成绩与真实场景脱节导致的误判,以及过度依赖 AI 工具可能放大信息组织成本和认知外包。

💡 大佬观点(Influencer Insights) 链接到标题

基于过去 24 小时内多位 AI 领域影响力人士(Influencers)在 X 平台的动态汇总,以下是核心观察与深度分析报告。


📊 AI 行业每日速递:端侧模型崛起、Coding Agent 生态与监管风暴 链接到标题

1. 🔥 今日焦点:技术趋势与产品热点 链接到标题

端侧模型(On-device Models)进入实用阶段 链接到标题

本地部署不再只是极客的玩具,而是正成为高生产力场景的主流选择。

  • Qwen 系列稳坐“甜点”宝座:@zhixianio 经过“苦行僧”式的测试后,表示在 M5 Max 上运行 Qwen3.6-35B-A3B (oMLX) 的体验极佳,响应速度比远程 API 更快,且在多模态个人助理(OpenClaw)和编程(PI-Mono)场景下“智商在线”。其甚至认为端侧表现已超越 DSV4 Pro。
  • 小模型的惊人潜力与天花板:@zhixianio 对 Gemma 4 12B 及其编码微调版进行了残酷的代码生成实测。结果表明,尽管 12B 模型在处理日常脚本时表现不错,但在生成“俄罗斯方块”等复杂有状态程序时,受限于体量,会频繁出现逻辑崩溃。这证明了 12B 是目前的“入门级编程模型”,但无法替代 35B+ 的复杂工程能力
  • Google 的端侧野望:@zhixianio 关注到 Google 发布的 QAT(量化感知训练)模型,认为这是 Google 为 Android 端侧 AI 铺路的关键一步,通过训练即优化量化的思路,试图让 AI 在手机上无缝运行。

Coding Agent 生态:从“能用”到“极致成本与体验” 链接到标题

AI 编程的讨论点已从“能不能写”转向了成本控制、Skill 工作流管理与跨模型调度

  • Codex 的用量焦虑:多位博主(@Pluvio9yte, @msjiaozhu)反馈 Codex 的 Token 消耗速度急剧增加,同样的任务消耗量是以前的数倍。@Pluvio9yte 作为 200 美元订阅用户,表达了强烈的不安。这暗示了 OpenAI 可能在调整后台计费权重或模型推理消耗。
  • 国产模型的 Coding 奇袭:@Pluvio9yte 和 @vista8 强力推荐字节跳动的 Doubao-Seed-2.1-Pro 模型。测试显示其前端与编码能力进步显著,且@Pluvio9yte 提供了在 Claude Code 中接入该模型的具体教程。多模型切换、寻找最高性价比的 API 成为了开发者的必备技能
  • Fable 的颠覆性生产力:@zhixianio 分享了他使用 Fable 的震撼体验,该 Agent 在 40 分钟内完成了 70% 的 Demo 工作,甚至指出了人类设计的不合理之处并自行优化。这标志着 AI 编程正从“执行者”向“架构审查者”演进。

AI 与办公软件的深层次融合 链接到标题

  • Claude Tag 发布:@dotey 详细介绍了 Anthropic 推出的 Claude Tag,它不再是简单的聊天机器人,而是作为“同事”常驻 Slack,具备上下文记忆、主动推送(Ambient 模式)和跨线程协作能力。@dotey 指出,Anthropic 内部已有 65% 的代码由其生成,这是 Agent 进入组织协作深水区的信号。

2. 🧠 独特观点与行业前瞻 链接到标题

  • 关于“因果大模型”的未来(@Pluvio9yte 引述 @huang_biwei): 当下的 LLM 是“相关性预测”,无法理解“往有破洞的杯子里倒水会漏”的物理因果。Aether AI 提出的因果世界模型(Causal World Models),试图让 AI 从数据拟合走向机制理解。这在具身智能与科学研发领域,被视为打破 Scaling Law 瓶颈的另一条路径。

  • Token 异质性与新生产要素(@lijigang & @vista8):

    • @lijigang 提出了一个极有深度的洞察:电力是同质的,但 Token 是异质的(不同模型的 Token 价值不同)。这种差异性容易形成“税收”式的供给瓶颈,而非简单的商品化基建。
    • @vista8 则认为,Agent 是一种趋近免费的数字化劳动力,但人的“注意力、信任和品牌”不仅不会贬值,反而因为信息过载而愈发珍贵。自媒体将成为新时代的“可信上下文入口”。
  • 警惕“Token 陷阱”与自动化思维(@gefei55 & @Pluvio9yte):

    • @gefei55 提醒大家:Token 看似无限,但人的精力有限。不要让 Vibe Coding 变成消耗生命的“Token 陷阱”,要分清主次。
    • @Pluvio9yte 总结了 AI 时代的核心法则:重复做过三遍的事情,一定要思考如何自动化
  • 关于时间差与信息套利(@gefei55): @gefei55 分享了利用 AI 监测 Twitter 爆款链接来寻找新域名的 SEO 玩法。在 Google Trends 起量之前,通过社交媒体信号(点赞、链接)挖矿,实现**“时间差套利”**。这证明了 AI 时代的商业机会在于洞察力的前置。

  • 回归思考的本质(@lijigang): 在工具泛滥的时代,@lijigang 反复强调“品味是一个人的损失函数”,建议进行“主题阅读”以获取学科视角。工具只是手段,真正的护城河是人对问题的定义能力和决策函数


3. 🛠️ 推荐的工具与资源 链接到标题

编程与开发 Agent

  • Doubao-Seed-2.1-Pro:由 @Pluvio9yte 推荐的国产模型,在编码和前端能力上表现出色,支持通过火山引擎 API 接入各类 Agent。
  • Fable:引发广泛讨论的 AI 编程工具,可处理复杂 Demo 开发并提出了比人类更优的解决方案。
  • 腾讯 WeKnora:@Pluvio9yte 推荐的开源企业知识平台,整合了 RAG、Agent 和自动生成知识图谱的功能。

Skill 与工作流管理

  • 软链接管理 Skill 方案:@dotey 提供了极客式的 Skill 管理思路,利用软链接统一管理多项目、多 Agent 的配置文件,配合 Git 实现高效迭代。
  • 访谈分析 Skill 套件:@dotey 开源了 interview-analysisinterview-writing 两个 Skill,旨在将长播客转化为高质量文章。
  • 端口管理工具:@vista8 推荐了一款免费的 macOS 菜单栏工具,可直观检视本地端口占用情况,是 Vibe Coding 的得力助手。

效率与数据

  • Obsidian 微信同步助手:@AI_Jasonyu 分享,可将微信文章和付费社群的高价值分享一键同步至 Obsidian,构建个人知识库。
  • CapWords:@nishuang 推荐的外语学习 App,利用 AI 识别现实物体并进行游戏化学习,是设计感与 AI 结合的标杆。
  • Google Workspace CLI:虽然作者被开除引发争议(@dotey 报道),但该工具依旧强大,可在命令行管理 Gmail、Drive,且自带 MCP 服务。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 34 条更新

a16z Podcast (A_full) 链接到标题

  • Building Self-Accelerating AI with Mirendil
    • 发布时间:2026-06-25 03:12 北京时间
    • 摘要:- Matt Bornstein 与 Mirendil 联合创始人 Behnam Neyshabur 和 Harsh Mehta 谈论了他们构建自我加速人工智能的愿景。
      • After leading research efforts at Google and Anthropic, the founders started Mirendil around a simple question: what happens when AI systems can meaningfully contribute to their own development?
      • 他们认为,最重要的应用可能是加速科学技术进步本身,而不是仅仅关注人工智能作为生产力工具。
      • 对话探讨了人工智能研究、缩放定律、自动化工程、科学发现以及构建可随时间改进的系统的挑战。
      • 他们讨论了人工智能辅助研究的未来,为什么他们认为科学进步仍然受到智能的瓶颈,以及更强大的人工智能系统如何帮助释放医学、工程和自然科学的进步。
    • EN 要点:
      • Matt Bornstein speaks with Mirendil cofounders Behnam Neyshabur and Harsh Mehta about their vision for building self-accelerating AI
      • After leading research efforts at Google and Anthropic, the founders started Mirendil around a simple question: what happens when AI systems can meaningfully co…
      • Rather than focusing solely on AI as a tool for productivity, they argue that the most important application may be accelerating scientific and technological pr…
      • The conversation explores AI research, scaling laws, automated engineering, scientific discovery, and the challenges of building systems that can improve over t…

Stratechery by Ben Thompson (A_full) 链接到标题

  • My Vibe Coding Adventure, The App and the Experience, Ten Takeaways
    • 发布时间:2026-06-24 18:00 北京时间
    • 摘要:- 我对我计划实际经常使用的应用程序进行氛围编码的经验和思考。
      • 15 美元/月150 美元/年。
      • 通过每周三封电子邮件或播客对当天新闻进行实质性分析。
      • 策略采访
      • 采访领先的上市首席执行官、私营公司创始人,并与分析师同行进行讨论。
    • EN 要点:
      • My experience and reflections on vibe coding an app that I plan on actually using regularly.

OpenAI Blog (A_full) 链接到标题

  • OpenAI and Broadcom unveil LLM-optimized inference chip
    • 发布时间:2026-06-24 14:00 北京时间
    • 摘要:- * 早期测试表明,第一代加速器的每瓦性能将大大优于当前最先进的技术。
        • 为整个行业当前和未来的法学硕士从头开始构建。
        • 从设计到生产的开发过程仅用了九个月,OpenAI 的模型加速了这一过程。
        • 扩展了OpenAI的全栈平台,从产品到模型,现在到芯片。
        • 将与数据中心合作伙伴一起进行千兆瓦规模的多代部署。
    • EN 要点:
      • OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI systems.

Google DeepMind Blog (A_full) 链接到标题

  • Introducing computer use in Gemini 3.5 Flash
    • 发布时间:2026-06-25 00:30 北京时间
    • 摘要:- 介绍Gemini 3.5 Flash 中的计算机使用。
      • This piece from Google DeepMind Blog explains how Introducing computer use in Gemini 3.5 Flash shapes the broader AI and infrastructure landscape.
      • 在 Gemini 3.5 Flash 中引入计算机使用之后,它还为创始人、运营商和投资者带来了实际影响。
    • EN 要点:
      • Introducing computer use in Gemini 3.5 Flash

ArXiv cs.AI (B_intro+search) 链接到标题

  • RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23927v1 公告类型:新。
      • 摘要:由大型语言模型 (LLM) 提供支持的代理人工智能系统正在迅速发展为自主决策系统,暴露出传统 LLM 漏洞之外的攻击向量。
      • 现有的安全评估通常与特定的实现或领域相关,限制了异构系统之间的统一比较。
      • 为了解决这一差距,我们引入了 RIFT-Bench,这是一种用于动态红队的图形表示驱动方法,可以跨不同代理架构进行统一评估。
    • EN 要点:
      • arXiv:2606.23927v1 Announce Type: new
      • Abstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyon…
      • Existing security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems
      • To address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming that enables unified evaluations across diverse…
  • Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23938v1 公告类型:新。
      • 摘要:结合思想链 (CoT) 推理的驱动 VLA 模型很有吸引力,因为它们利用预训练的 VLM 表示并以自然语言公开中间决策,但当前的基本原理通常缺乏保持基本原理与计划运动的因果关系所需的分步决策语义。
      • 我们引入了 Neuro-Symbolic Drive,这是一种神经符号驱动框架,它使用直接从经典的基于规则的规划器中提取的基于规则的推理轨迹来监督驱动 VLA。
      • 我们的主要观察结果是,基于规则的规划器是符号人工智能系统,已经充当可执行推理引擎:它们推理主动安全约束,搜索候选操作,并选择最终轨迹。
    • EN 要点:
      • arXiv:2606.23938v1 Announce Type: new
      • Abstract: Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose inter…
      • We introduce Neuro-Symbolic Drive, a neuro-symbolic driving framework that supervises a driving VLA with rule-grounded reasoning traces extracted directly from…
      • Our key observation is that rule-based planners are symbolic AI systems that already function as executable reasoning engines: they reason about active safety c…
  • Critique of Agent Model

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23991v1 公告类型:新。
      • 随着大语言模型(LLM)系统的兴起,这些系统被称为“编码代理”、“人工智能联合科学家”和其他承诺提高生产力的“代理”工具,同时,“存在”的担忧,例如人工智能在针对人类的投机性“机器机构”下具有破坏性力量,逃离了人类的控制,澄清自动化的终点和机构的起点变得至关重要,这既是为了构建有能力的系统,也是为了理解是否和害怕什么。
      • 借鉴笛卡尔的独立思想基础和科幻小说中对自主存在的描绘,我们调查了人工智能代理的现状,并从五个维度分析了代理架构:目标、身份、决策、自我调节和学习。
      • 具体来说,我们认为真正的代理要求这些结构\emph{在系统本身内部化}而不是通过外部脚手架组装。
    • EN 要点:
      • arXiv:2606.23991v1 Announce Type: new
      • Abstract: What is an agent
      • What constitutes agency
      • With the rise of Large Language Model (LLM) systems marketed as coding agents'', AI co-scientists’’, and other ``agentic" tools that promise to drive up pro…
  • Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.24010v1 公告类型:新。 -摘要:多智能体系统广泛应用于需要在严格的安全约束下协调行为的安全关键应用。
      • 现有的方法面临着根本性的权衡:基于学习的方法实现了强大的经验性能,但缺乏理论安全保证,而控制理论方法增强了安全性,但往往导致过于保守和低效的行为。
      • 我们提出了一种分层多智能体强化学习框架,该框架通过约束流形在低级别的温和假设下强制执行硬安全约束,同时通过高级策略学习实现有效协调。
    • EN 要点:
      • arXiv:2606.24010v1 Announce Type: new
      • Abstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints
      • Existing approaches face a fundamental trade-off: learning-based methods achieve strong empirical performance but lack theoretical safety guarantees, while cont…
      • We propose a hierarchical multi-agent reinforcement learning framework that enforces hard safety constraints under mild assumptions at low level via a constrain…
  • Reinforcement Learning Towards Broadly and Persistently Beneficial Models

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.24014v1 公告类型:新。
      • 摘要:随着人工智能系统在日益多样化和高风险的环境中部署,模型对齐必须推广到训练期间看到的任务和领域之外。
      • 这对于强化学习 (RL) 来说尤其重要,因为强化学习可能会通过奖励黑客、欺骗或其他意外策略引入意外的错位。
      • 我们研究在现实领域中实例化的有益行为的强化学习是否可以产生超出训练分布的广泛且持久的对齐泛化。
    • EN 要点:
      • arXiv:2606.24014v1 Announce Type: new
      • Abstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen dur…
      • This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through reward hacking, deception, or other unintended…
      • We study whether RL on beneficial behavior, instantiated in realistic domains, can produce broad and persistent alignment generalization beyond the training dis…
  • Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.24026v1 公告类型:新。 -摘要:机械可解释性在自动定位电路方面取得了实质性进展,但解释本地化组件的作用仍然是劳动密集型且难以标准化。
      • 在这项工作中,我们研究一旦确定了电路,语言模型(LM)代理是否可以帮助解决这个解释问题。
      • 我们引入了 AgenticInterpBench,这是一个由 84 个半合成变压器电路和 163 个组件级注释构建的电路解释基准。
    • EN 要点:
      • arXiv:2606.24026v1 Announce Type: new
      • Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains l…
      • In this work, we study whether language model (LM) agents can assist with this explanation problem once a circuit has already been identified
      • We introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer circuits with 163 component-level annotations
  • Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.24042v1 公告类型:新。
      • 摘要:推荐系统通常通过整体优化以实现用户的即时参与,从而引发过滤气泡和语义同质化。
      • 标准的单目标模型,包括传统的深度 Q 网络,不足以在平台保留率和信息多样性和提供商公平性等关键社会价值观之间进行权衡。
      • 为了解决这些限制,我们引入了多目标强化学习框架,将推荐形式化为语义多目标马尔可夫决策过程。
    • EN 要点:
      • arXiv:2606.24042v1 Announce Type: new
      • Abstract: Recommender systems often induce filter bubbles and semantic homogenization by monolithically optimizing for immediate user engagement
      • Standard single-objective models, including traditional Deep Q-Networks, are ill-equipped to navigate the trade-offs between platform retention and critical soc…
      • To address these limitations, we introduce a multi-objective reinforcement learning framework that formalizes recommendation as a semantic multi-objective Marko…
  • Ensemble Feature Selection and Harris Hawks Optimization for Explainable Mental Health Risk Prediction in Female Sex Workers

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.24047v1 公告类型:新。
      • 摘要:影响女性性工作者(FSW)的重要心理健康问题之一是精神障碍,尤其是抑郁症。
      • 遭受暴力、耻辱和经济困难进一步增加了他们的心理风险。
      • 当前的机器学习 (ML) 模型通常无法有效捕获这一边缘群体中存在的高维且复杂的风险模式。
    • EN 要点:
      • arXiv:2606.24047v1 Announce Type: new
      • Abstract: One of the significant mental health issues affecting female sex workers (FSWs) is mental disorders, especially depression
      • Exposure to violence, stigma, and economic hardship further increases their psychological risk
      • Current machine learning (ML) models are typically ineffective at capturing the high-dimensional and complex risk patterns that exist in this marginalized group
  • Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.24064v1 公告类型:新。
      • 摘要:从强语言模型到弱语言模型提炼推理能力通常涉及模仿特定的解决方案轨迹,有效地转移要回答的内容而不是如何推理。
      • 这种轨迹级模仿鼓励记忆特定于实例的步骤,而不是获得可转移的解决问题的技能,从而限制了对新问题的概括。
      • 我们提出策略引导策略优化(SGPO),它用可重用的策略蒸馏代替实例级轨迹模拟。
    • EN 要点:
      • arXiv:2606.24064v1 Announce Type: new
      • Abstract: Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transfe…
      • This trajectory-level imitation encourages memorization of instance-specific steps rather than acquisition of transferable problem-solving skills, limiting gene…
      • We propose Strategy-Guided Policy Optimization (SGPO), which replaces instance-level trajectory imitation with reusable strategy distillation
  • Exploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Papers

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.24099v1 公告类型:新。
      • 摘要:算法已成为人工智能(AI)时代科学研究的核心。
      • 尽管论文中提到的算法经常被用来表示受欢迎程度和影响力,但现有的研究通常孤立地评估单个算法,而对通过它们互连形成的集体影响力的关注有限。
      • 本研究基于学术论文全文构建自然语言处理(NLP)中的大规模算法共现网络,并从网络角度研究算法影响。
    • EN 要点:
      • arXiv:2606.24099v1 Announce Type: new
      • Abstract: Algorithms have become central to scientific research in the era of artificial intelligence (AI)
      • Although algorithm mentions in papers are often used to indicate popularity and influence, existing studies usually evaluate individual algorithms in isolation…
      • This study constructs large-scale algorithm co-occurrence networks in natural language processing (NLP) based on the full text of academic papers and investigat…

ArXiv cs.CL (B_intro+search) 链接到标题

  • EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23693v1 公告类型:新。
      • 摘要:文本到 SQL 使用户能够通过生成可执行的 SQL 查询来使用自然语言查询数据库。
      • 最近的方法越来越多地采用基于大型语言模型的强化学习(RL)来利用执行反馈进行训练。
      • 然而,现有的 RL 方法为 SQL 查询中的所有子句分配统一的查询级别奖励,平等对待正确和不正确的子句。
    • EN 要点:
      • arXiv:2606.23693v1 Announce Type: new
      • Abstract: Text-to-SQL enables users to query databases using natural language by generating executable SQL queries
      • Recent methods have increasingly adopted Large Language Models based reinforcement learning (RL) to leverage execution feedback for training
      • However, existing RL methods assign uniform query-level rewards to all clauses in a SQL query, treating correct and incorrect clauses equally
  • ModTGCN: Modularity-aware Graph Neural Networks for Text Classification

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23694v1 公告类型:新。 -摘要:尽管语义文档图表现出强大的类一致聚类,但基于图的文本分类模型通常依赖于局部邻域聚合并忽略全局社区结构。
      • 忽略这一点可能会模糊类边界并导致过度平滑。
      • 我们提出 ModTGCN,一种用于文本分类的模块化感知图神经网络,它联合优化交叉熵和基于模块化的辅助目标,以促进类一致的文档社区,同时保留区分性表示。
    • EN 要点:
      • arXiv:2606.23694v1 Announce Type: new
      • Abstract: Graph-based text classification models typically rely on local neighborhood aggregation and overlook global community structure, despite semantic docu…
      • Ignoring this can blur class boundaries and lead to over-smoothing
      • We propose ModTGCN, a modularity-aware graph neural network for text classification that jointly optimizes cross-entropy and a modularity-based auxiliary object…
  • Quantifying Prior Dominance in RAG Systems

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23695v1 公告类型:新。
      • 摘要:检索增强生成(RAG)将大型语言模型建立在外部知识的基础上,但当前的评估依赖于遭受“认知盲目性”的离散启发法 - 无法区分真正的上下文信息提取和参数记忆回忆。
      • 为了解决这个问题,我们引入了标准化上下文利用率(NCU)指标,利用零样本、预言机和对抗条件下的连续令牌对数概率来严格量化上下文信息增益。
      • 评估从 1.5B 到 72B 参数的架构以及专有的商业 API 表明,对于严格的事实提取(没有思想链推理),传统的缩放法则表现出极端的收益递减:高效的小语言模型 (SLM) 匹配或优于高容量架构。
    • EN 要点:
      • arXiv:2606.23695v1 Announce Type: new
      • Abstract: Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suf…
      • To address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adver…
      • Evaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of…
  • Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23700v1 公告类型:新。
      • 摘要:紧急错位 (EM) 与错位角色向量和邪恶性格特征的激活有关,这表明 EM 是通过破坏模型的一致性格来运作的,而不是直接学习有害内容。
      • 受这种联系的推动,我们研究了自生成文本识别(SGTR)微调,作为一种与现有训练中防御不同的针对字符的干预措施。
      • 我们在三个模型(GPT-4.1、Qwen2.5-32B-Instruct、Seed-OSS-36B-Instruct)和多个 EM 数据集上进行两阶段微调实验,将 SGTR 微调与良性微调基线(正确的特定领域数据、常识和字数统计)进行比较,以发现它在逆转和预防设置中都是有效的防御。
    • EN 要点:
      • arXiv:2606.23700v1 Announce Type: new
      • Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates thro…
      • Motivated by this connection, we study self-generated text recognition (SGTR) finetuning as a character-targeted intervention that is distinct from existing in-…
      • We conduct two-stage finetuning experiments across three models (GPT-4.1, Qwen2.5-32B-Instruct, Seed-OSS-36B-Instruct) and multiple EM datasets to compare SGTR…
  • Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23701v1 公告类型:新。
      • 摘要:定性的产品反馈可以揭示细致入微的用户体验,但其隐含的情感难以衡量。
      • 本文提出了一个可扩展且可解释的框架,该框架使用大型语言模型 (LLM) 来量化此类数据中的产品需求。
      • 使用来自 ZORQ 和 CARMA 的两个产品需求工具包 (PDT) 数据集,其中包含 106 个受访者术语分组以及黄金标准人工注释、零样本连续数值情感评分和分类情感分类,无需依赖明确的评论分数即可进行评估。
    • EN 要点:
      • arXiv:2606.23701v1 Announce Type: new
      • Abstract: Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure
      • This paper presents a scalable and interpretable framework that uses large language models (LLMs) to quantify product desirability from such data
      • Using two Product Desirability Toolkit (PDT) datasets from ZORQ and CARMA comprising 106 respondent term groupings with gold-standard human annotation, zero-sho…
  • Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23881v1 公告类型:新。
      • 摘要:基于知识的视觉问答(KB-VQA)需要将视觉查询扎根于图像中直接可观察内容之外的外部知识。
      • 虽然最近的多模态大语言模型 (MLLM) 显示出强大的感知能力,但它们在需要细粒度实体和证据级别基础的 KB-VQA 任务上遇到了困难。
      • 大多数现有的多模态检索增强生成(MM-RAG)方法将实体辨别和部分级证据排序紧密耦合到单个重新排序阶段,导致成本高昂且泛化有限。
    • EN 要点:
      • arXiv:2606.23881v1 Announce Type: new
      • Abstract: Knowledge-Based Visual Question Answering (KB-VQA) requires grounding visual queries to external knowledge beyond directly observable content in image…
      • While recent multi modal large language models (MLLMs) show strong perceptual abilities, they struggle on KB-VQA tasks requiring groundings from both fine-grain…
      • Most existing multi-modal retrieval augmented generation (MM-RAG) methods tightly couple entity discrimination and section-level evidence ranking into a single…
  • One Year Later…The Harms Persist, But So Do We!

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23884v1 公告类型:新。
      • 摘要:通用大语言模型 (LLM) 越来越多地用于与心理健康相关的对话,但安全保障措施在不同临床条件下仍然不足且不一致。
      • 这项研究使用四种对抗性攻击变体,评估了 16 个 DSM-5 条件下的 6 个专有法学硕士,引入了八维伤害分类法和多维评估框架。
      • 结果表明,保障措施仅适用于自杀和自残,而饮食失调、物质使用障碍和重度抑郁症等病症的失败率高达 100%。
    • EN 要点:
      • arXiv:2606.23884v1 Announce Type: new
      • Abstract: General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain inadequate an…
      • This study evaluates six proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a…
      • Results show that safeguards hold reliably only for suicide and self-harm, while conditions such as eating disorders, substance use disorder, and major depressi…
  • Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23915v1 公告类型:新。
      • 摘要:实践中通常将 LLM 检索增强生成中的自动归因指标视为可互换的。
      • 我们审核了八个自动评分器——词汇、嵌入和 BERTScore 基线以及蕴涵/基础训练模型(clean 和 FEVER NLI、检查器 MiniCheck)——跨三个评估结构(出处/话题性、生成答案归因和事实检查蕴涵),询问是否有任何评分器转移:在多数据集构造的每个数据集上保持在最佳审核评分器的 95% 置信区间内。
      • 在具有最多多数据集人工标记覆盖范围的构造中——生成答案归因(AttributionBench 的四个源数据集,n = 1,610,具有独立的 HAGRID,n = 2,150)——没有一个这样做:每个数据集指标排名反转(AttributedQA 上的 Kendall tau = -0.64,p = 0.031 与 AttributedQA 上的 p = 0.031)
    • EN 要点:
      • arXiv:2606.23915v1 Announce Type: new
      • Abstract: Practice often treats automatic metrics for attribution in LLM retrieval-augmented generation as interchangeable
      • We audit eight automatic scorers – lexical, embedding, and BERTScore baselines alongside entailment/grounding-trained models (clean and FEVER NLI, the checker…
      • In the construct with the most multi-dataset human-labeled coverage – generated-answer attribution (AttributionBench’s four source datasets, n = 1,610, with in…
  • When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23937v1 公告类型:新。
      • 摘要:精确匹配检索召回通常用作检索器是否为下游决策模型提供有用的策略上下文的代理。
      • 我们使用 Qwen2.5-3B/7B 分类器在 tau-bench 中测试此代理的行动前策略分类。
      • 在黄金政策条件下,紧凑的结构状态在调整后将宏观 F1 相对于原始轨迹提高了 0.13-0.17。
    • EN 要点:
      • arXiv:2606.23937v1 Announce Type: new
      • Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model
      • We test this proxy for pre-action policy classification in tau-bench using Qwen2.5-3B/7B classifiers
      • Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by 0.13-0.17 after tuning
  • QuechuaTok: Morphological Boundary Accuracy as a Necessary Metric for Tokenizer Evaluation in Agglutinative Low-Resource Languages

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23943v1 公告类型:新。 -摘要:标记化是 NLP 流程中的基础步骤,但生育率等标准评估指标无法捕获粘着语言的形态正确性。
      • 我们推出了 QuechuaTok,这是一个系统基准,比较了四种标记化策略(BPE、Unigram LM、WordPiece 和形态感知的 PRPE 标记器),适用于南盖丘亚语 (quz),南美洲有 8-1000 万人使用一种低资源粘着语言。
      • 使用 20 万句语料库和 SQUOIA 有限状态形态分析器(Rios,2016)作为银标准,我们评估三个指标:生育率、OOV 率和形态边界准确性 (MorphAcc)。
    • EN 要点:
      • arXiv:2606.23943v1 Announce Type: new
      • Abstract: Tokenization is a foundational step in NLP pipelines, yet standard evaluation metrics such as fertility rate fail to capture morphological correctness…
      • We present QuechuaTok, a systematic benchmark comparing four tokenization strategies - BPE, Unigram LM, WordPiece, and a morphology-aware PRPE tokenizer - for S…
      • Using a 200k-sentence corpus and the SQUOIA finite-state morphological analyzer (Rios, 2016) as silver standard, we evaluate three metrics: fertility rate, OOV…

ArXiv cs.LG (B_intro+search) 链接到标题

  • Systematic Exploration of 4-Expert Heterogeneous Mixture-of-Experts via Automated Pipeline Search

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23739v1 公告类型:新。
      • 摘要:我们为 LEMUR 神经网络数据集生态系统中的异构 4 专家混合专家 (MoE4) 架构提出了一种自动化大规模搜索管道。
      • 基于手工制作的异构 MoE 参考模型,我们用确定性代码组装生成器取代了手动设计,该生成器系统地将从 LEMUR 数据库中提取的基础架构系列组合到 MoE4 集成中,每个集成都由具有温度缩放、混合增强和余弦退火学习率调度的卷积门网络控制。
      • 在 NVIDIA RTX 4090 上进行的为期 28 天的活动中,该管道在 197 个批次中生成了 4,463 个候选模型,其中 1,021 个已成功评估。
    • EN 要点:
      • arXiv:2606.23739v1 Announce Type: new
      • Abstract: We present an automated large-scale search pipeline for heterogeneous 4-Expert Mixture-of-Experts (MoE4) architectures within the LEMUR neural network…
      • Building on a hand-crafted heterogeneous MoE reference model, we replace manual design with a deterministic code-assembly generator that systematically combines…
      • Over a 28-day campaign on an NVIDIA RTX 4090, the pipeline generated 4,463 candidate models across 197 batches, of which 1,021 were evaluated successfully
  • Weight-Space Geometry of Offline Reasoning Training

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23740v1 公告类型:新。
      • 摘要:离线强化学习损失(RFT、RIFT、DFT、离线 GRPO、DPO)被广泛用于将大型教师的推理提炼给较小的学生,并且通常仅在下游准确性上进行比较。
      • 我们询问它们在机制上是否不同或收敛到类似的权重更新。
      • 使用仅注意 LoRA 从单个基础模型 (Qwen3-4B) 的相同数学部署中训练六种方法(SFT、RFT、DFT、RIFT、离线 GRPO、DPO),我们通过余弦相似性、主角子空间分析、线性模式连接和 CKA 分析生成的增量。
    • EN 要点:
      • arXiv:2606.23740v1 Announce Type: new
      • Abstract: Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into smaller studen…
      • We ask whether they are mechanistically distinct or converge to a similar weight update
      • Training six methods (SFT, RFT, DFT, RIFT, Offline GRPO, DPO) on identical math rollouts from a single base model (Qwen3-4B) with attention-only LoRA, we analyz…
  • A Survey on Federated Causal Discovery and Inference

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23741v1 公告类型:新。
      • 摘要:因果推理包括因果结构的发现和因果效应的推断,是数据驱动决策的基础。
      • 在实践中,用于可靠因果分析的数据通常分布在各个机构中,并且由于隐私法规或通信限制而无法集中。
      • 联邦学习(FL)通过在不共享原始数据的情况下实现协作分析来解决这个问题,从而催生了联邦因果发现(FCD)和推理(FCI)领域的快速发展。
    • EN 要点:
      • arXiv:2606.23741v1 Announce Type: new
      • Abstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision mak…
      • In practice, data for reliable causal analysis are often distributed across institutions and cannot be centralized due to privacy regulations or communication c…
      • Federated learning (FL) addresses this by enabling collaborative analysis without raw data sharing, giving rise to the rapidly growing field of federated causal…
  • Low-power analogue neural networks with trainable nonlinear connections for continuous control

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23742v1 公告类型:新。
      • 摘要:物理神经网络通过直接使用模拟设备物理进行计算来保证低功耗机器学习,但大多数架构迫使非线性设备响应充当标量权重。
      • 受 Kolmogorov-Arnold 网络的启发,我们在连接上放置可训练的非线性函数,使每个物理连接成为可学习的计算元素。
      • 将这些功能实现为现场可编程模拟阵列上的模拟带通滤波器,我们发现其好处是依赖于任务的,并且源于物理基础的平滑性:网络代表平滑、连续有价值的目标,包括机器人运动学、连续控制和光伏最大功率点跟踪,其节点和连接比多层感知器少得多,但在类似分类的决策边界上不提供参数效率优势。
    • EN 要点:
      • arXiv:2606.23742v1 Announce Type: new
      • Abstract: Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures force nonlinear…
      • Inspired by Kolmogorov-Arnold networks, we place trainable nonlinear functions on the connections, making each physical connection a learnable computational ele…
      • Realising these functions as analogue band-pass filters on field-programmable analogue arrays, we find that the benefit is task-dependent and follows from the s…
  • Synergizing Physically Constrained MCMC and Chemical-Informed Gaussian Processes for Reaction Network Discovery

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23757v1 公告类型:新。
      • 摘要:从稀疏、嘈杂的化学时间序列数据中提取可解释的控制方程仍然很困难,因为离散反应拓扑和连续动力学参数紧密耦合。
      • 我们提出了 PC-MCMC-CIGP,一种可重复的灰盒工作流程,结合了尖峰和平板拓扑采样、硬守恒和热力学筛选,以及用于参数校准和实验设计的化学知情高斯过程 (CIGP) 残差模型。
      • 方法论的贡献不是孤立的新 MCMC 或 GP 系列;相反,它是将这些组件集成到物理受限的工作流程中,并具有明确的不确定性感知采集选择。
    • EN 要点:
      • arXiv:2606.23757v1 Announce Type: new
      • Abstract: Extracting interpretable governing equations from sparse, noisy chemical time-series data remains difficult because discrete reaction topology and con…
      • We present PC-MCMC-CIGP, a reproducible gray-box workflow that combines spike-and-slab topology sampling, hard conservation and thermodynamic screening, and a C…
      • The methodological contribution is not a new MCMC or GP family in isolation; rather, it is the integration of these components into a physically constrained wor…
  • Exploring Dualistic Meta-Learning to Enhance Domain Generalization in Open Set Scenarios

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23758v1 公告类型:新。
      • 摘要:域泛化从多个源域中学习,以泛化到未见过的目标域。
      • 然而,它经常忽略源和目标之间标签不匹配的实际情况。
      • 然后提出开放集域泛化来识别未见域中的未见类。
    • EN 要点:
      • arXiv:2606.23758v1 Announce Type: new
      • Abstract: Domain generalization learns from multiple source domains to generalize to unseen target domains
      • However, it often neglects the realistic case of label mismatch between source and target
      • Open set domain generalization is then proposed to recognize unseen classes in unseen domains
  • One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23767v1 公告类型:新。
      • 摘要:图宾根因果对的标题准确性通常会在论文之间进行比较,尽管每个因果对都是根据作者自己的协议进行测量的——不同的对子集、权重、模型选择和决策率。
      • 我们认为这是错误的比较,并运行正确的比较:同一手的重新评估,其中每种方法都由我们在相同的 102 对上运行,有一个严格的规则 - 不对每对进行调整和强制做出决定。
      • 作为一个干净的参考点,我们特意引入了一个最小基线:排序条件压缩,它将量化、排序、一阶差分数据提供给现成的压缩器(bz2),并且具有零拟合参数。
    • EN 要点:
      • arXiv:2606.23767v1 Announce Type: new
      • Abstract: Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors’ own proto…
      • We argue this is the wrong comparison and run the right one: a same-hands re-evaluation in which every method is run by us on the identical 102 pairs, with one…
      • As a clean reference point we introduce a deliberately minimal baseline: sorted-conditional compression, which feeds quantized, sorted, first-differenced data t…
  • Deciphering Fingerprints of 3D Molecular Surfaces for Accurate Epitope Prediction

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23830v1 公告类型:新。
      • 摘要:分子表面编码决定抗体-抗原识别的几何和物理化学模式,这对于表位预测至关重要。
      • 然而,现有方法依赖于序列或主干结构,并且难以捕获不连续的、表面驱动的表位。
      • 这项研究提出了 SurfBind,一种以表面为中心的表位预测学习框架,可直接在分子表面表征上运行。
    • EN 要点:
      • arXiv:2606.23830v1 Announce Type: new
      • Abstract: Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope prediction
      • However, existing methods rely on sequences or backbone structures and struggle to capture discontinuous, surface-driven epitopes
      • This study presents SurfBind, a surface-centric learning framework for epitope prediction that operates directly on molecular surface representations
  • Reconstructing GRACE Terrestrial Water Storage with Spatio-Temporal Graph Neural Networks: An Application to South America

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23833v1 公告类型:新。
      • 摘要:陆地水储存(TWS)综合了雪、土壤湿度、地表水和地下水,是气候变化和人类活动如何重塑全球水循环的关键指标。
      • GRACE 和 GRACE-FO 卫星任务提供了唯一直接的、全球一致的 TWS 变化观测,但它们的记录只从 2002 年开始,这对于许多气候尺度的分析来说太短了。
      • 我们提出了一种深度学习应用程序,通过学习每日 ERA5 气象强迫(降水、蒸散、径流)和每月 GRACE 观测之间的关系,重建可追溯到 1940 年的每月类 GRACE TWS 异常 (TWSA)。
    • EN 要点:
      • arXiv:2606.23833v1 Announce Type: new
      • Abstract: Terrestrial water storage (TWS) integrates snow, soil moisture, surface water, and groundwater and is a key indicator of how climate variability and h…
      • The GRACE and GRACE-FO satellite missions provide the only direct, globally consistent observations of TWS change, but their record only begins in 2002 which is…
      • We present a deep learning application that reconstructs monthly GRACE-like TWS anomalies (TWSA) back to 1940 by learning the relationship between daily ERA5 me…
  • The Degeneracy Distillery

    • 发布时间:2026-06-24 12:00 北京时间
    • 摘要:- arXiv:2606.23838v1 公告类型:新。
      • 摘要:当两个或多个参数或标签产生相似的数据时,它们是退化的,或者难以区分。
      • 简并性使得标签预测和逆问题都变得困难,因为机器学习算法和概率采样器都依赖于数据及其相对于参数的梯度的可区分性。
      • 然而,识别物理模型或现实世界数据集中的简并性可以阐明模型的选择或产生数据的底层过程。
    • EN 要点:
      • arXiv:2606.23838v1 Announce Type: new
      • Abstract: When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish
      • Degeneracies render both label prediction and inverse problems difficult, since both machine learning algorithms and probabilistic samplers rely on the distingu…
      • However, identifying degeneracies in physical models or real-world datasets can be elucidating about the choice of model or the underlying process that produces…