🤖 AI 速览

OpenAI 披露 ChatGPT 广告年化收入已达 10 亿美元,并开始向更多地区开放投放,说明 AI 产品的商业化正在从订阅和 API 走向更成熟的广告模式。与此同时,Meta 相关内容治理和解再度提醒,AI 平台的增长与监管会同时成为长期变量。
📋 文章元数据
发布时间
2026-09-01
类型
ai-daily
字数
5611
阅读时长
27 min

2026-09-01 AI日更 | ChatGPT 广告收入破十亿美元,AI 平台商业化与治理同步加速 链接到标题

OpenAI 披露 ChatGPT 广告年化收入已达 10 亿美元,并开始向更多地区开放投放,说明 AI 产品的商业化正在从订阅和 API 走向更成熟的广告模式。与此同时,Meta 相关内容治理和解再度提醒,AI 平台的增长与监管会同时成为长期变量。

📖 本期 Watch List 深度导读 链接到标题

今天最值得关注的有三条主线。第一条是 LLM 工程化继续下沉到基础设施层:从向量索引加速推理、Publisher-Adaptive 内容抽取,到更稳的实体消歧,几篇论文都在解决“模型能不能更快、更稳地吃进数据并吐出结果”。第二条是评测与对齐开始反思方法本身,Rasch 评测、Rubric-Guided RL、情绪上下文对决策偏置的研究,值得做模型评估和对齐的团队重点读。第三条则是推理与知识应用的边界扩展,跨语言多跳问答、数学例题内化改进、医疗 EHR 证据问答,都在强调“先找对证据,再做判断”。另外,Meta 监管与内容治理的长文也很适合管理层和产品负责人补读。

🌐 X 平台 AI 热点快讯 链接到标题

话题 1:Tim Cook Bows Out as Apple CEO After 15 Years of Growth 链接到标题

  • 分类:AI · News
  • 概况:热度时间:1 day ago,相关帖子数:68000
  • 摘要:Tim Cook Bows Out as Apple CEO After 15 Years of Growth: Tim Cook Bows Out as Apple CEO After 15 Years of Growth Today wraps Tim Cook’s 15-year run as Apple CEO, during which the company soared from $350 billion to over $4.6 trillion in market value, launched hits like Apple Watch, AirPods, and Vis…

话题 2:Google Unveils TimesFM-3 for Smarter Time Series Forecasting 链接到标题

  • 分类:AI · News
  • 概况:热度时间:4 hours ago,相关帖子数:702
  • 是什么事:Google 发布了 TimesFM-3,一款面向时间序列预测的模型,主打更强的泛化能力和更准确的未来趋势预测。
  • 为什么重要:这对 AI 领域重要,因为时间序列预测是供应链、金融、能源和运维等场景的核心能力,基础模型如果能降低建模门槛并提升跨场景效果,会扩大 AI 在业务决策中的实际价值。
  • 讨论概况:X 上的讨论主要集中在它相较前代和传统预测方法是否有实质提升、是否足够通用、在真实业务数据上的效果如何,以及模型是否开放、成本和部署复杂度是否可接受。

话题 3:Musk Recalls Doubts That Nearly Stopped SpaceX Launch 链接到标题

  • 分类:AI · Other
  • 概况:热度时间:8 hours ago,相关帖子数:3600
  • 摘要:Musk Recalls Doubts That Nearly Stopped SpaceX Launch:

话题 4:Meta Launches Paid Muse Code AI for Complex Coding Tasks 链接到标题

  • 分类:AI · News
  • 概况:热度时间:4 hours ago,相关帖子数:1900
  • 摘要:Meta Launches Paid Muse Code AI for Complex Coding Tasks: 🧠 GPT-5.6 Launches: My 12-Hour Test Daily · 2026-07-10 I want to put three things side by side today: the GPT-5.6 launch, Meta’s Muse Spark 1.1 API, and what I noticed after running Sol for about 12 hours. Meta priced the API at $1.25 per mi…

话题 5:OpenAI Codex Reaches 25 Million Active Users with Paid Quota Reset 链接到标题

  • 分类:AI · News
  • 概况:热度时间:2 days ago,相关帖子数:15000
  • 摘要:OpenAI Codex Reaches 25 Million Active Users with Paid Quota Reset:

话题 6:OpenClaw 2.0 Launches with Massive Update and 933 Contributors 链接到标题

  • 分类:AI · News
  • 概况:热度时间:19 hours ago,相关帖子数:7800
  • 摘要:OpenClaw 2.0 Launches with Massive Update and 933 Contributors:

话题 7:Lionel Messi Retires from Argentina National Team 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:8 hours ago,相关帖子数:908000
  • 摘要:Lionel Messi Retires from Argentina National Team: 🚨🇦🇷 BREAKING: Lionel Messi RETIRES from international football! The Argentina captain will no longer represent the national team.

话题 8:Tottenham Secure Adarabioyo Permanently and Mudryk on Loan from Chelsea 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:1 day ago,相关帖子数:127000
  • 是什么事:热刺被传已永久签下阿达拉比奥尤,并从切尔西租借穆德里克,转会消息在 X 上快速发酵。
  • 为什么重要:这类高热度体育转会话题会推动 AI 在舆情监测、体育内容生成、转会概率分析和球员价值评估中的应用需求。
  • 讨论概况:X 上主要在讨论转会真实性、交易条款和两名球员对球队阵容的实际影响,分歧集中在消息来源可信度以及这笔操作是否划算。

话题 9:Newcastle Land Fernandez-Pardo and Loan Out Woltemade on Deadline Day 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:9 hours ago,相关帖子数:47000
  • 是什么事:纽卡斯尔联在转会截止日签下费尔南德斯·帕尔多,并将沃尔特马德外租。
  • 为什么重要:该事件本身属于足球转会,与人工智能领域没有直接关联,可能是平台分类或标签误标。
  • 讨论概况:X上的讨论预计集中于两笔操作对纽卡斯尔阵容深度、球员发展和球队竞争力的影响,但在缺少代表性推文的情况下,无法确认具体分歧。

话题 10:Overwatch Revives LE SSERAFIM Crossover with New Hero Skins 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:7 hours ago,相关帖子数:38000
  • 是什么事:《守望先锋》重新推出与韩国女子组合LE SSERAFIM的联动活动,并上线新的英雄皮肤。
  • 为什么重要:该事件本身与人工智能领域没有直接关联,但其高热度可反映数字内容、虚拟角色商业化及游戏与流行文化融合的传播价值。
  • 讨论概况:X上的讨论焦点可能集中在新皮肤的设计质量、联动内容是否值得购买、活动重启的原因,以及玩家对重复联动和商业化运营的看法。

话题 11:Liverpool Reject £80m Gakpo Bid from City, Shift Focus to Fernandez 链接到标题

  • 分类:AI · Sports
  • 概况:热度时间:1 day ago,相关帖子数:146000
  • 摘要:Liverpool Reject £80m Gakpo Bid from City, Shift Focus to Fernandez:

话题 12:Kirkentine Meme Marks Charlie Kirk Anniversary 链接到标题

  • 分类:AI · Entertainment
  • 概况:热度时间:,相关帖子数:1700
  • 是什么事:X 平台上出现了围绕“Charlie Kirk 周年纪念”的 Kirkentine meme 话题,相关内容被大量转发和二创。
  • 为什么重要:这类话题反映了 AI 生成内容、梗图传播和平台推荐机制如何放大公共议题与娱乐化表达,对理解 AI 在内容分发和舆论塑形中的作用有参考价值。
  • 讨论概况:讨论主要集中在这类 meme 是单纯的网络玩笑、政治表达,还是对当事人及相关事件的再包装;分歧点在于有人认为是正常的二创传播,也有人认为在借热点进行立场化输出。

话题 13:Fence Kiss Meme Revives Dating Money Debates 链接到标题

  • 分类:AI · Entertainment
  • 概况:热度时间:,相关帖子数:355
  • 是什么事:“Fence Kiss”表情包在X平台走红,并重新引发了关于约会时费用应由谁承担的讨论。
  • 为什么重要:这反映了AI生成或传播的娱乐内容如何影响社会议题和公众舆论,也体现了AI与网络文化、情感关系及消费观念的交叉影响。
  • 讨论概况:讨论主要集中在约会费用是否应平摊、主动邀约者是否应承担费用,以及性别角色和经济平等观念在当代约会中的冲突。

话题 14:Rick and Morty Fan Comic Captures Tearful Grandpa-Grandson Moment 链接到标题

  • 分类:AI · Entertainment
  • 概况:热度时间:,相关帖子数:493
  • 是什么事:一部描绘《瑞克和莫蒂》中祖孙含泪时刻的粉丝漫画在X平台引发关注,相关讨论约493条。
  • 为什么重要:这反映了生成式AI正在降低同人漫画等娱乐内容的创作门槛,也引发对角色表达、创作者权益和内容真实性的关注。
  • 讨论概况:讨论焦点主要集中在漫画是否由AI生成、情感表达是否自然,以及AI同人创作应被视为艺术创作还是对原作风格和版权的挪用。

话题 15:Tesla Registers First Steering-Wheel-Free Cybercabs in Texas Ahead of Launch 链接到标题

  • 分类:AI · News
  • 概况:热度时间:23 hours ago,相关帖子数:10000
  • 摘要:Tesla Registers First Steering-Wheel-Free Cybercabs in Texas Ahead of Launch:

今日 X 上的 AI 舆情小结 链接到标题

今天 X 上的舆论主线是“AI 能力是否真正落地”与“AI 生成/放大的内容如何继续重塑平台热度”并行推进:一边是 Google TimesFM-3 这类时间序列基础模型,另一边则是转会、联动皮肤、meme 和同人漫画等高传播话题,显示 AI 讨论已明显外溢到内容分发和舆情场景。共识大致集中在两点:基础模型如果真能提升跨场景预测能力,会有实用价值;而 AI 相关内容的传播速度和覆盖面,已经在明显改变平台上的注意力分配。分歧则主要在“是不是货真价实的提升”以及“这些内容到底是创作、玩梗,还是立场输出/商业化包装”,尤其对 TimesFM-3 的真实业务效果、AI 同人是否越界、meme 是否政治化,争议都比较强。潜在风险在于,过度包装技术进展会放大预期落差,平台推荐机制会进一步放大误导性或对立性内容,而版权、真实性和内容归属问题也会持续成为摩擦点。

💡 大佬观点(Influencer Insights) 链接到标题

今日大佬观点暂缺,推荐阅读 Watch List 深度内容。

📚 附录:今日 Watch List 更新源列表 链接到标题

时间窗口:最近 3 天;覆盖 22 个源;共 32 条更新

Stratechery by Ben Thompson (A_full) 链接到标题

  • Meta Settles, A Framework For Regulating Content, The Rest of Big Tech
    • 发布时间:2026-08-31 18:00 北京时间
    • 摘要:【待翻译】- Meta’s settlement makes sense for all parties, but the entire sage highlights why any solution to regulating technology feels off.
      • $15 / month or $150 / year.
      • Substantial analysis of the news of the day delivered via three weekly emails or podcasts.
      • Stratechery Interviews.
      • Interviews with leading public CEOs, private company founders, and discussions with fellow analysts.
    • EN 要点:
      • Meta’s settlement makes sense for all parties, but the entire sage highlights why any solution to regulating technology feels off.

OpenAI Blog (A_full) 链接到标题

  • A milestone in expanding access to AI
    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- In less than 200 days after launch, ChatGPT Ads has reached $1 billion in annualized revenue run rate.
      • The platform is now used by tens of thousands of advertisers and continues to expand globally.
      • Starting later today, advertisers can purchase ChatGPT ads directly via Ads Manager across India, Europe, the Middle East, and North Africa.
      • Advertising is one pillar of OpenAI’s diversified business model, alongside consumer subscriptions, enterprise offerings, and usage-based APIs.
      • Together, these offerings give people, developers, and businesses choice in how they access OpenAI products, including an advertising-supported free tier that helps keep ChatGPT available to more than 1 billion weekly active users.
    • EN 要点:
      • ChatGPT Ads reaches $1 billion in annualized revenue run rate and expands globally, supporting broader access to AI through free and affordable options.

ArXiv cs.AI (B_intro+search) 链接到标题

  • Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27459v1 Announce Type: new.
      • Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge.
      • Its DeepQA system defeated the strongest human Jeopardy!
      • champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy.
    • EN 要点:
      • arXiv:2608.27459v1 Announce Type: new
      • Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge
      • Its DeepQA system defeated the strongest human Jeopardy
      • champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impos…
  • Rating the Raters: Rasch Measurement Theory for LLM Evaluation

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27463v1 Announce Type: new.
      • Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content.
      • Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters.
      • Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured.
    • EN 要点:
      • arXiv:2608.27463v1 Announce Type: new
      • Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content
      • Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters
      • Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured
  • Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27464v1 Announce Type: new.
      • Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking.
      • Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three types of expected utility: instrumental (will it help me act better?), hedonic (will it make me feel better?), and cognitive (will it improve my understanding?).
      • Each utility is estimated through a lens shaped by well-documented cognitive biases, including illusion of control, automation bias, unrealistic optimism, impact bias, overconfidence, and confirmation bias.
    • EN 要点:
      • arXiv:2608.27464v1 Announce Type: new
      • Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking
      • Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three ty…
      • ), hedonic (will it make me feel better
  • Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27471v1 Announce Type: new.
      • Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped.
      • Spotting a fallacious argument requires contextual knowledge beyond its pure surface text.
      • This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within the argumentative discourse.
    • EN 要点:
      • arXiv:2608.27471v1 Announce Type: new
      • Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political d…
      • Spotting a fallacious argument requires contextual knowledge beyond its pure surface text
      • This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within th…
  • LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27472v1 Announce Type: new.
      • Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge.
      • We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs).
      • In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging.
    • EN 要点:
      • arXiv:2608.27472v1 Announce Type: new
      • Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offe…
      • We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs)
      • In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging
  • Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27475v1 Announce Type: new.
      • Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it.
      • These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single trajectory.
      • We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonparametric, time-invariant coefficient fields.
    • EN 要点:
      • arXiv:2608.27475v1 Announce Type: new
      • Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it
      • These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single tra…
      • We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonp…
  • Class-Based Heuristic Selection for Solving the Flying Block Puzzle

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27476v1 Announce Type: new.
      • Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to exploit the structural constraints that govern constrained spatial domains, causing search performance to degrade catastrophically on harder instances.
      • We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearance-to-size constraints encountered in multi-agent path finding, autonomous vehicle navigation, and block relocation systems.
      • We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units are scarce, a formal kinematic taxonomy partitioning the state space into seven mutually exclusive classes with provably admissible heuristics based on vacancy ratio and goal-piece geometry, and a class-conditional tie-breaking mechanism that dynamically switches between depth-priority and vertical-distance ordering to overcome f-value plateaus.
    • EN 要点:
      • arXiv:2608.27476v1 Announce Type: new
      • Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to explo…
      • We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearan…
      • We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units…
  • Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27477v1 Announce Type: new.
      • Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks.
      • Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use.
      • We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios.
    • EN 要点:
      • arXiv:2608.27477v1 Announce Type: new
      • Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks
      • Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design…
      • We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios
  • Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27480v1 Announce Type: new.
      • Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected.
      • This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations in tea plantations.
      • Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, with geographic coordinates recorded for spatial tracking.
    • EN 要点:
      • arXiv:2608.27480v1 Announce Type: new
      • Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial d…
      • This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations…
      • Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, wi…
  • Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27482v1 Announce Type: new.
      • Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems.
      • The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through conditional aggregation tests on admissible knowledge contexts.
      • The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level.
    • EN 要点:
      • arXiv:2608.27482v1 Announce Type: new
      • Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems
      • The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through condition…
      • The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level

ArXiv cs.CL (B_intro+search) 链接到标题

  • Accelerating LLM Inference via Vector Index Based Output Embeddings

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27460v1 Announce Type: new.
      • Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies.
      • We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary projection with an HNSW-based vector index.
      • The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retrieved logits into a sparse full-vocabulary tensor.
    • EN 要点:
      • arXiv:2608.27460v1 Announce Type: new
      • Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with larg…
      • We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary…
      • The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retri…
  • SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27461v1 Announce Type: new.
      • Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts.
      • This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding.
      • To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal academic dialog benchmark.
    • EN 要点:
      • arXiv:2608.27461v1 Announce Type: new
      • Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts
      • This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding
      • To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal acad…
  • Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27462v1 Announce Type: new.
      • Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging.
      • While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples.
      • This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases.
    • EN 要点:
      • arXiv:2608.27462v1 Announce Type: new
      • Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and context…
      • While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples
      • This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases
  • The Effect of Emotional Context on Large Language Models’ Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27465v1 Announce Type: new.
      • Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user’s emotional state has become an important safety problem.
      • We test whether emotional expression increases a model’s endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence).
      • As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length.
    • EN 要点:
      • arXiv:2608.27465v1 Announce Type: new
      • Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice accordin…
      • We test whether emotional expression increases a model’s endorsement (encouragement to proceed) when a user, holding the same objective information, is overconf…
      • As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating t…
  • PACE: Publisher-Adaptive Content Extraction via Agentic Automation

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27466v1 Announce Type: new.
      • Abstract: Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability.
      • General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, images, and tables.
      • Direct LLM-based extraction offers greater flexibility, but incurs substantial cost and latency at scale, while manually engineered publisher-specific parsers can achieve high accuracy but require substantial human effort to build and maintain.
    • EN 要点:
      • arXiv:2608.27466v1 Announce Type: new
      • Abstract: Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and…
      • General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, imag…
      • Direct LLM-based extraction offers greater flexibility, but incurs substantial cost and latency at scale, while manually engineered publisher-specific parsers c…
  • UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27467v1 Announce Type: new.
      • Abstract: We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records.
      • We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment).
      • For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer.
    • EN 要点:
      • arXiv:2608.27467v1 Announce Type: new
      • Abstract: We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records
      • We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment)
      • For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the f…
  • Select, Don’t Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27470v1 Announce Type: new.
      • Abstract: Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs.
      • State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selecting the correct one given context.
      • Dual-encoder models optimize for both within a shared embedding space, forcing representations to balance high-recall retrieval with fine-grained selection, and they require trained retrievers, which are costly to maintain as knowledge graphs change.
    • EN 要点:
      • arXiv:2608.27470v1 Announce Type: new
      • Abstract: Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs
      • State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selec…
      • Dual-encoder models optimize for both within a shared embedding space, forcing representations to balance high-recall retrieval with fine-grained selection, and…
  • XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27481v1 Announce Type: new.
      • Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language.
      • This hides failures at language boundaries inside the reasoning chain.
      • We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence.
    • EN 要点:
      • arXiv:2608.27481v1 Announce Type: new
      • Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually…
      • This hides failures at language boundaries inside the reasoning chain
      • We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence
  • INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27501v1 Announce Type: new.
      • Abstract: Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctness, raising the question whether models truly internalize mathematical concepts or merely memorize solution patterns.
      • In human mathematics education, example-based reasoning such as constructing counterexamples to test theorem boundaries reflects deep conceptual understanding, but remains underdeveloped in current LLMs.
      • Enhancing this capability through preference optimization presents two key challenges: (1) the model’s limited example-based reasoning ability makes constructing effective preference pairs inherently difficult; and (2) capability acquisition is progressive, as the model must first learn to adopt this strategy before learning to apply it correctly.
    • EN 要点:
      • arXiv:2608.27501v1 Announce Type: new
      • Abstract: Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctne…
      • In human mathematics education, example-based reasoning such as constructing counterexamples to test theorem boundaries reflects deep conceptual understanding,…
      • Enhancing this capability through preference optimization presents two key challenges: (1) the model’s limited example-based reasoning ability makes constructin…
  • A Survey on Rubric-Guided Reinforcement Learning for Language Models

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27505v1 Announce Type: new.
      • Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences.
      • However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality.
      • Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of reward design, feedback generation, and policy optimization.
    • EN 要点:
      • arXiv:2608.27505v1 Announce Type: new
      • Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences
      • However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality
      • Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of re…

ArXiv cs.LG (B_intro+search) 链接到标题

  • Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27507v1 Announce Type: new.
      • Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in replicated copies of the same environment.
      • However, its pooled team-entropy score measures only collective exploration and cannot identify policies that contribute non-redundant coverage.
      • We introduce Marginal Coverage Credit for PGPSE (MCC-PGPSE), which combines leave-one-policy-out coverage with state-owner specialization to estimate policy-specific credit.
    • EN 要点:
      • arXiv:2608.27507v1 Announce Type: new
      • Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in repli…
      • However, its pooled team-entropy score measures only collective exploration and cannot identify policies that contribute non-redundant coverage
      • We introduce Marginal Coverage Credit for PGPSE (MCC-PGPSE), which combines leave-one-policy-out coverage with state-owner specialization to estimate policy-spe…
  • Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation–Deployment Gap

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27512v1 Announce Type: new.
      • Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models.
      • When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow creates a structural validation–deployment gap: because quantization is a many-to-one mapping over parameter space, source-precision certification does not guarantee behavioral equivalence in the deployed configuration.
      • We formalize this gap through Quantization Behavioral Equivalence Classes (QBECs) and prove that QBEC membership does not imply behavioral equivalence, providing a theoretical basis for quantization-triggered backdoor attacks.
    • EN 要点:
      • arXiv:2608.27512v1 Announce Type: new
      • Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models
      • When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow creates a structural…
      • We formalize this gap through Quantization Behavioral Equivalence Classes (QBECs) and prove that QBEC membership does not imply behavioral equivalence, providin…
  • DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27513v1 Announce Type: new.
      • Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length.
      • Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the KV cache in most layers with fixed-size recurrent states.
      • However, these recurrent states are commonly stored in FP32 and consume substantial GPU memory; their updates are memory-bandwidth bound and contribute significantly to decoding latency.
    • EN 要点:
      • arXiv:2608.27513v1 Announce Type: new
      • Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length
      • Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the KV cache in most layers with fixed-siz…
      • However, these recurrent states are commonly stored in FP32 and consume substantial GPU memory; their updates are memory-bandwidth bound and contribute signific…
  • A Deeper Analysis of Block-Sparse Featurizers

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27515v1 Announce Type: new.
      • Abstract: The recently introduced block-sparse featurizer (BSF; Fel et al., 2026) is similar to a sparse autoencoder (SAE), but its atomic unit is a small subspace (a block of directions) rather than a single direction.
      • It is designed for features that live on low-dimensional manifolds, which are especially frequent in vision.
      • This work studies the BSF’s strengths and weaknesses, finding how it still somewhat suffers from classic SAE failure modes, like feature splitting and composition.
    • EN 要点:
      • arXiv:2608.27515v1 Announce Type: new
      • Abstract: The recently introduced block-sparse featurizer (BSF; Fel et al., 2026) is similar to a sparse autoencoder (SAE), but its atomic unit is a small subsp…
      • It is designed for features that live on low-dimensional manifolds, which are especially frequent in vision
      • This work studies the BSF’s strengths and weaknesses, finding how it still somewhat suffers from classic SAE failure modes, like feature splitting and compositi…
  • When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27518v1 Announce Type: new.
      • Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by catastrophic forgetting and weight-disentanglement error.
      • In the literature, these difficulties are merely treated separately and mitigated through a variety of solutions, while the geometry induced by the base optimizer is treated as an implementation detail.
      • In this work, we show that the two difficulties are in fact two instances of the same phenomenon: a parameter update useful for one task shifts the model’s outputs on another.
    • EN 要点:
      • arXiv:2608.27518v1 Announce Type: new
      • Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by…
      • In the literature, these difficulties are merely treated separately and mitigated through a variety of solutions, while the geometry induced by the base optimiz…
      • In this work, we show that the two difficulties are in fact two instances of the same phenomenon: a parameter update useful for one task shifts the model’s outp…
  • Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27521v1 Announce Type: new.
      • Abstract: Many dynamical processes unfold on the sphere but the default scientific machine learning architectures are Euclidean.
      • Applying these architectures on a regular lat-lon grid causes problems: Cartesian convolutions become distorted at high latitude; 2D FFTs in Fourier neural operators incorrectly assume double periodicity; Cartesian positional encodings in ViTs distort spherical geodesic distances.
      • Recent work moves towards natively spherical primitives, including spherical convolutions (e.g., DeepSphere or DISCO), Spherical Fourier Neural Operators (SFNOs), and geodesic attention.
    • EN 要点:
      • arXiv:2608.27521v1 Announce Type: new
      • Abstract: Many dynamical processes unfold on the sphere but the default scientific machine learning architectures are Euclidean
      • Applying these architectures on a regular lat-lon grid causes problems: Cartesian convolutions become distorted at high latitude; 2D FFTs in Fourier neural oper…
      • Recent work moves towards natively spherical primitives, including spherical convolutions (e.g., DeepSphere or DISCO), Spherical Fourier Neural Operators (SFNOs…
  • Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27574v1 Announce Type: new.
      • Abstract: Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple groups or consists of multiple objects.
      • To date, a handful of multi-label graph learning methods exist, but none of them integrate training-time interpretation capability.
      • While post-hoc graph explainers have been developed, they do not explicitly model label-dependent evidence sharing in multi-label graph learners, especially when label pairs are weakly or negatively associated.
    • EN 要点:
      • arXiv:2608.27574v1 Announce Type: new
      • Abstract: Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple group…
      • To date, a handful of multi-label graph learning methods exist, but none of them integrate training-time interpretation capability
      • While post-hoc graph explainers have been developed, they do not explicitly model label-dependent evidence sharing in multi-label graph learners, especially whe…
  • Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27634v1 Announce Type: new.
      • Abstract: Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN imposes the same neighborhood cardinality throughout the feature space.
      • This assumption can be inadequate for data whose local geometry varies substantially across the underlying manifold.
      • We introduce Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification (CARSANN), a geometry-driven framework that adapts the spatial support of each neighborhood according to local geometric complexity.
    • EN 要点:
      • arXiv:2608.27634v1 Announce Type: new
      • Abstract: Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN imposes the same neighborhood cardinality thr…
      • This assumption can be inadequate for data whose local geometry varies substantially across the underlying manifold
      • We introduce Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification (CARSANN), a geometry-driven framework that adapts the spatial suppor…
  • More Data Cannot Break a Symmetry: Identifiability by Design

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27651v1 Announce Type: new.
      • Abstract: Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus geometry bounds what any such alignment can identify, before data exist.
      • The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling creates near-duplicates whose transposition is nearly free.
      • We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention.
    • EN 要点:
      • arXiv:2608.27651v1 Announce Type: new
      • Abstract: Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus…
      • The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling create…
      • We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention
  • Unsupervised Continual Learning with Growing Self-Organizing Maps and Synthetic Replay

    • 发布时间:2026-08-31 12:00 北京时间
    • 摘要:【待翻译】- arXiv:2608.27662v1 Announce Type: new.
      • Abstract: This work presents a generative continual learning framework based on growing self-organizing maps (GSOMs) that are augmented with learned distributional statistics as well as encoder-decoder models for class-incremental learning.
      • The proposed approach enables exemplar-free replay using distributional statistical memory, which eliminates the need to store raw data.
      • Each GSOM unit maintains its own mean, variance, and covariance estimates, which are subsequently used to generate synthetic samples for replay; in encoder-decoder configurations, these samples are then decoded back into the input space (via ancestral sampling) for subsequent training.
    • EN 要点:
      • arXiv:2608.27662v1 Announce Type: new
      • Abstract: This work presents a generative continual learning framework based on growing self-organizing maps (GSOMs) that are augmented with learned distributio…
      • The proposed approach enables exemplar-free replay using distributional statistical memory, which eliminates the need to store raw data
      • Each GSOM unit maintains its own mean, variance, and covariance estimates, which are subsequently used to generate synthetic samples for replay; in encoder-deco…