🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-26
- 类型
- ai-daily
- 字数
- 5674
- 阅读时长
- 27 min
2026-08-26 AI日更 | OpenAI 把 AI 竞争推向全栈:自研推理芯片、算力系统与可信智能体同步加速 链接到标题
今天的主线从模型能力转向系统能力。OpenAI 继续强化从数据中心、芯片到产品的全栈策略,Jalapeño 指向更低成本推理;同时,智能体治理、引用归因与运行时证据协议升温,可靠性评测也开始深入长尾语言和文化场景。
📖 本期 Watch List 深度导读 链接到标题
今天最值得跟进的主线有三条。第一条是 AI 基础设施全面走向“全栈化”:OpenAI 继续把算力、芯片与推理系统打通,Jalapeño 的早期结果也指向更高效率的推理架构,值得工程团队重点看。第二条是智能体进入治理与安全的硬问题:从 sycophancy、引用归因,到 agentic security 和运行时证据协议,说明“能做事”之后,可信、可审计开始成为核心门槛。第三条是模型可靠性正向长尾语言和文化场景下沉,今天几篇关于 Khmer、Nigerian Pidgin、Cyrillic 及视觉语言偏差的研究,提醒我们评测体系仍远未成熟。顺带一提,Netflix 探索把自己做成流媒体聚合入口,也值得关注,它反映的其实是平台型分发逻辑正在重排。
🌐 X 平台 AI 热点快讯 链接到标题
话题 1:SpaceXAI and Cursor Boost Grok Model Usage Limits Again 链接到标题
- 分类:AI · News
- 概况:热度时间:7 hours ago,相关帖子数:5400
- 是什么事:xAI 和 Cursor 再次上调 Grok 模型的使用额度,引发开发者和 AI 用户对可用性与成本的关注。
- 为什么重要:这反映出 AI 产品正在通过更高调用额度争夺开发者入口和工作流占用时长,也说明推理模型的算力供给、成本控制和商业化竞争正在加剧。
- 讨论概况:X 上的讨论主要围绕额度提升是否意味着 Grok 基础设施和模型服务能力真的增强,Cursor 用户是否会获得更稳定的编码体验,以及这是否会压缩 Claude、GPT 等模型在开发者工具中的份额;同时也有人质疑这种提升的成本可持续性和实际质量提升幅度。
话题 2:Shopify CEO Pushes for Claude Code to Adopt AGENTS.md Standard 链接到标题
- 分类:AI · News
- 概况:热度时间:9 hours ago,相关帖子数:2900
- 是什么事:Shopify CEO 公开呼吁 Claude Code 采用 AGENTS.md 标准,用统一文件描述代码库中 AI 编程代理的规则、上下文和协作约定。
- 为什么重要:这关系到 AI 编程工具的标准化和互操作性:如果不同代理能读取同一套项目约定,就能降低配置成本、减少误操作,并提升多工具、多代理在真实代码库中的协作稳定性。
- 讨论概况:X 上讨论焦点集中在 Claude Code 是否会带动 AGENTS.md 成为事实标准,以及统一规范能否改善编码代理对项目上下文的理解;分歧在于它究竟是必要的行业基础设施,还是又一个增加维护负担、可能造成配置碎片化的项目文件。
话题 3:OpenAI Engineer Surprises with Changed Appearance in Interview 链接到标题
- 分类:AI · News
- 概况:热度时间:2 days ago,相关帖子数:4600
- 是什么事:有 X 用户热议一位 OpenAI 工程师在采访中的外貌变化,引发对其身份、状态和背景的讨论。
- 为什么重要:这类话题之所以重要,在于它反映了公众对 OpenAI 及其核心员工的高度关注,容易影响公司形象、人才叙事以及外界对 AI 行业文化与工作压力的判断。
- 讨论概况:讨论焦点主要集中在外貌变化是否属实、原因是什么,以及这类关注是正常的人物观察还是对个人隐私的过度解读;也有人借题发挥,联想到 AI 公司高强度工作环境与舆论放大效应。
话题 4:OpenAI Unveils Jalapeño Chip Outpacing Nvidia in Speed and Efficiency 链接到标题
- 分类:AI · News
- 概况:热度时间:9 hours ago,相关帖子数:13000
- 是什么事:OpenAI 据称公布了一款名为 Jalapeño 的芯片,并宣称其在速度和能效上超过了 Nvidia 的同类产品。
- 为什么重要:如果这一说法成立,说明头部 AI 公司正在通过自研芯片降低对通用 GPU 的依赖,这会影响算力成本、供应链和未来模型部署方式。
- 讨论概况:X 上的讨论主要集中在三个点:性能对比是否有公开基准支撑、这是否意味着 OpenAI 正在推进更深层的硬件自研,以及它对 Nvidia 在 AI 算力市场中的地位会产生多大冲击。
今日 X 上的 AI 舆情小结 链接到标题
今天 X 上的主线是:AI 竞争正在从“谁的模型更强”转向“谁能更便宜、更稳定、更深地嵌入开发者工作流”,无论是 Grok 提高额度、Claude Code 被呼吁统一 AGENTS.md,还是 OpenAI 传出自研芯片,都在指向算力、接口标准和入口控制这三件事。共识大致是,开发者工具的核心不只是模型能力本身,还包括可用性、成本和项目上下文的协作效率;分歧则集中在这些动作到底是实质能力提升,还是营销式放量、标准化负担或未经验证的性能叙事。对 AGENTS.md 的看法也明显分裂:支持者把它看成减少误操作、提升多代理协作的基础设施,反对者担心它会变成额外维护成本并制造新的碎片化。潜在风险主要有三类:一是高额度和自研芯片背后的算力成本是否可持续,二是未经充分基准支撑的性能宣称可能误导市场判断,三是围绕个体工程师外貌和状态的放大讨论,容易把行业压力、隐私边界和舆论猎奇混在一起。
💡 大佬观点(Influencer Insights) 链接到标题
今日大佬观点暂缺,推荐阅读 Watch List 深度内容。
📚 附录:今日 Watch List 更新源列表 链接到标题
时间窗口:最近 3 天;覆盖 22 个源;共 35 条更新
Stratechery by Ben Thompson (A_full) 链接到标题
- Netflix to Sell Streaming Services?, Streamers as Aggregators, Revisiting Roku
- 发布时间:2026-08-25 18:00 北京时间
- 摘要:【待翻译】- Netflix is considering selling other streaming services, and I think it’s a good idea; it’s also a let-down for Netflix’s original goals and potential pivots.
- $15 / month or $150 / year.
- Substantial analysis of the news of the day delivered via three weekly emails or podcasts.
- Stratechery Interviews.
- Interviews with leading public CEOs, private company founders, and discussions with fellow analysts.
- EN 要点:
- Netflix is considering selling other streaming services, and I think it’s a good idea; it’s also a let-down for Netflix’s original goals and potential pivots.
OpenAI Blog (A_full) 链接到标题
The full stack behind abundant intelligence
- 发布时间:2026-08-25 15:05 北京时间
- 摘要:【待翻译】- Progress in AI compounds fastest when the entire system improves together.
- That is how I think about OpenAI’s compute strategy: one integrated system spanning data centers and chips, frontier models, our developer platform, consumer and enterprise products, and AI-native devices, with each layer strengthening the next.
- Better software makes hardware more productive.
- Hardware designed for our workloads improves speed and efficiency.
- More capable models unlock better products, which generate more demand, usage, and learning.
- EN 要点:
- OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower…
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
- 发布时间:2026-08-25 15:00 北京时间
- 摘要:【待翻译】- Since announcing Jalapeño, OpenAI’s first custom inference chip, we have been testing the chip and the system built around it.
- The results show a significant performance advance: Jalapeño can serve more AI work per unit of power while also returning responses more quickly.
- Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two.
- For customers, that can mean faster responses, more responsive agents, and more reliable access as demand grows.
- Our mission is to ensure that artificial general intelligence benefits all of humanity.
- EN 要点:
- Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern mod…
Disrupting a new covert influence campaign from Russia
- 发布时间:2026-08-25 08:00 北京时间
- 摘要:【待翻译】- OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.
- This piece from OpenAI Blog explains how Disrupting a new covert influence campaign from Russia shapes the broader AI and infrastructure landscape.
- It also surfaces practical implications for founders, operators, and investors following Disrupting a new covert influence campaign from Russia.
- EN 要点:
- OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.
Introducing the Admin plugin for ChatGPT Work and Codex
- 发布时间:2026-08-25 08:00 北京时间
- 摘要:【待翻译】- Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.
- This piece from OpenAI Blog explains how Introducing the Admin plugin for ChatGPT Work and Codex shapes the broader AI and infrastructure landscape.
- It also surfaces practical implications for founders, operators, and investors following Introducing the Admin plugin for ChatGPT Work and Codex.
- EN 要点:
- Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.
ArXiv cs.AI (B_intro+search) 链接到标题
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21362v1 Announce Type: new.
- Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request.
- Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions.
- We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position.
- EN 要点:
- arXiv:2608.21362v1 Announce Type: new
- Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request
- Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at…
- We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position
AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21363v1 Announce Type: new.
- Abstract: A protocol is presented for recording the governance decisions of automated AI runtimes.
- When a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can check offline, independent of the runtime that produced it.
- A record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by value, and declares both what its evidence covers and what it does not.
- EN 要点:
- arXiv:2608.21363v1 Announce Type: new
- Abstract: A protocol is presented for recording the governance decisions of automated AI runtimes
- When a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can ch…
- A record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by val…
Reviewing Model Collapse and Countermeasures
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21366v1 Announce Type: new.
- Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors.
- The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models.
- Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply.
- EN 要点:
- arXiv:2608.21366v1 Announce Type: new
- Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors
- The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models
- Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply
AI Learning and Conceptual Transfer in the Game of Hidden Rules
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21372v1 Announce Type: new.
- Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback, representation design, rule difficulty analysis, transfer learning, generalization, and pseudo-bot-assisted human learning analysis.
- The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of human learning data.
- arXiv:2608.21372v1 Announce Type: new Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules… The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of huma….
- EN 要点:
- arXiv:2608.21372v1 Announce Type: new
- Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules…
- The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of huma…
LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21374v1 Announce Type: new.
- Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics.
- We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria.
- From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%.
- EN 要点:
- arXiv:2608.21374v1 Announce Type: new
- Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspe…
- We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-…
- From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems…
SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21375v1 Announce Type: new.
- Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores.
- Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload size, token use, and latency, and under-fetching, which omits fields needed to answer the query.
- We present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and license policies as a schema graph.
- EN 要点:
- arXiv:2608.21375v1 Announce Type: new
- Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and grap…
- Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload…
- We present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and lice…
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21379v1 Announce Type: new.
- Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the working population - yet it is typically identified only retrospectively, after academic decline has already occurred.
- A contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without interpreting it.
- This paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI architecture to surface personalized insights and early burnout signals.
- EN 要点:
- arXiv:2608.21379v1 Announce Type: new
- Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the work…
- A contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without…
- This paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI arch…
There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21382v1 Announce Type: new.
- Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model’s answer is read from generated text or from per-option likelihoods.
- Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next.
- We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items.
- EN 要点:
- arXiv:2608.21382v1 Announce Type: new
- Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and wh…
- Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items tha…
- We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21393v1 Announce Type: new.
- Abstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepower sits somewhere else, and moving sensitive records between the two creates real headaches around latency, security, and regulatory exposure.
- IBM’s Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation.
- In this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip accelerator for lightweight classification tasks, and Red Hat OpenShift for container orchestration.
- EN 要点:
- arXiv:2608.21393v1 Announce Type: new
- Abstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsep…
- IBM’s Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation
- In this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip acceler…
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21408v1 Announce Type: new.
- Abstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts.
- Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal grammar, inconsistent sen-tence structures, and multiple variations in word spellings.
- This research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data.
- EN 要点:
- arXiv:2608.21408v1 Announce Type: new
- Abstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress…
- Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal gram…
- This research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data
ArXiv cs.CL (B_intro+search) 链接到标题
Distinguishing Revision and Delayed Elaboration in Incremental Narrative Interpretation
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21364v1 Announce Type: new.
- Abstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations must be updated accordingly.
- Incremental interpretation, therefore, depends not only on what is represented but also on how the representational state evolves under new evidence.
- We distinguish two structurally different update operators that arise in narrative interpretation: revision-driven update and delayed elaboration.
- EN 要点:
- arXiv:2608.21364v1 Announce Type: new
- Abstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations…
- Incremental interpretation, therefore, depends not only on what is represented but also on how the representational state evolves under new evidence
- We distinguish two structurally different update operators that arise in narrative interpretation: revision-driven update and delayed elaboration
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21365v1 Announce Type: new.
- Abstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in multilingual embedding models, and frequent mixed Khmer-English usage.
- This paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search.
- We construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplication, and document-length control.
- EN 要点:
- arXiv:2608.21365v1 Announce Type: new
- Abstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in…
- This paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search
- We construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplica…
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21369v1 Announce Type: new.
- Abstract: Nigerian Pidgin is one of Africa’s most widely spoken languages, yet remains severely underrepresented in language model evaluation.
- Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured.
- We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning.
- EN 要点:
- arXiv:2608.21369v1 Announce Type: new
- Abstract: Nigerian Pidgin is one of Africa’s most widely spoken languages, yet remains severely underrepresented in language model evaluation
- Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language under…
- We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning
On the Role of Citations in Preference Data
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21376v1 Announce Type: new.
- Abstract: Many NLP tasks require systems to provide attribution in their outputs–i.e.
- citations to grounding sources.
- Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs.
- EN 要点:
- arXiv:2608.21376v1 Announce Type: new
- Abstract: Many NLP tasks require systems to provide attribution in their outputs–i.e
- citations to grounding sources
- Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs
Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21377v1 Announce Type: new.
- Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings.
- This paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse?
- Across 4,800 veracity judgments (200 statements $\times$ 6 models $\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic systems (feedback loops, reconsideration checkpoints, and iterative refinement) systematically amplifies sycophantic behavior.
- EN 要点:
- arXiv:2608.21377v1 Announce Type: new
- Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied pr…
- This paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse
- Across 4,800 veracity judgments (200 statements $\times$ 6 models $\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic sys…
Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21384v1 Announce Type: new.
- Abstract: Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating disparities in cost and context capacity.
- We quantify this overhead across nine production tokenizers and five languages with standardized Cyrillic and Latin representations, covering 8.37 million word forms.
- On a corpus benchmark, Ukrainian shows 68-121% token overhead on modern tokenizers and 220% on the older cl100k, measured through full-text fertility on the BrUK and Brown corpora.
- EN 要点:
- arXiv:2608.21384v1 Announce Type: new
- Abstract: Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating dispa…
- We quantify this overhead across nine production tokenizers and five languages with standardized Cyrillic and Latin representations, covering 8.37 million word…
- On a corpus benchmark, Ukrainian shows 68-121% token overhead on modern tokenizers and 220% on the older cl100k, measured through full-text fertility on the BrU…
A Social Media Analysis of Discourse on the Israel–Palestine Conflict on Telegram
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21385v1 Announce Type: new.
- Abstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have not been systematically compared at scale.
- This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning May 2021 to June 2026 and covering multiple conflict escalations.
- It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference, and a fine-tuned BERTweet model), and a framing analysis, all evaluated against 736 manually annotated messages.
- EN 要点:
- arXiv:2608.21385v1 Announce Type: new
- Abstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose br…
- This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning…
- It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference…
Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21415v1 Announce Type: new.
- Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social groups.
- Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limited by their reliance on a single, stereotyped viewpoint and fail to account for the diversity of social perspectives.
- Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs multi-group counterfactual perspectives within the visual representation space and integrates them during decoding to promote equitable model behavior.
- EN 要点:
- arXiv:2608.21415v1 Announce Type: new
- Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from…
- Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limi…
- Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs m…
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21423v1 Announce Type: new.
- Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools.
- As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures.
- We systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pipelines.
- EN 要点:
- arXiv:2608.21423v1 Announce Type: new
- Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools
- As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures
- We systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pi…
CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21462v1 Announce Type: new.
- Abstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties.
- Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages.
- But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?
- EN 要点:
- arXiv:2608.21462v1 Announce Type: new
- Abstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alph…
- Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages
- But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do
ArXiv cs.LG (B_intro+search) 链接到标题
Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21386v1 Announce Type: new.
- Abstract: Given a task described by a few examples, how should a model be specialized to it?
- Four mechanisms are available – zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork – yet the operating regime of the last is rarely mapped.
- We run the identical four-way comparison across six tasks spanning regression, generation, language modeling, reinforcement learning, and clinical and genomic classification, holding the specialist, the context, and (where we can) the training budget fixed.
- EN 要点:
- arXiv:2608.21386v1 Announce Type: new
- Abstract: Given a task described by a few examples, how should a model be specialized to it
- Four mechanisms are available – zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork – yet the…
- We run the identical four-way comparison across six tasks spanning regression, generation, language modeling, reinforcement learning, and clinical and genomic c…
Runtime Action Interference for AI Control of AlphaStar in StarCraft II
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21398v1 Announce Type: new.
- Abstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppresses, or replaces its proposed actions.
- We contribute \emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering configured action patterns after inference.
- RAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op.
- EN 要点:
- arXiv:2608.21398v1 Announce Type: new
- Abstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppre…
- We contribute \emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering co…
- RAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op
Federated Ensemble Forecasting Under Supply-Chain Market Volatility
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21399v1 Announce Type: new.
- Abstract: Supply chain forecasting systems increasingly operate under market shocks, non-identically distributed regional demand, and limited willingness to centralize commercial data.
- This work proposes Federated Ensemble Forecasting with Negative-Correlation Learning (FEF NCL), a distributed method that trains specialized forecasting experts across client nodes while discouraging redundant model errors.
- The framework combines temporal feature encoders, client level drift scoring, reliability-weighted aggregation, and an explain ability layer that exposes the market and supplier variables most responsible for each forecast.
- EN 要点:
- arXiv:2608.21399v1 Announce Type: new
- Abstract: Supply chain forecasting systems increasingly operate under market shocks, non-identically distributed regional demand, and limited willingness to cen…
- This work proposes Federated Ensemble Forecasting with Negative-Correlation Learning (FEF NCL), a distributed method that trains specialized forecasting experts…
- The framework combines temporal feature encoders, client level drift scoring, reliability-weighted aggregation, and an explain ability layer that exposes the ma…
Class-Conditioned Gaussian Mixture Modeling for Imbalanced Time Series Quantification
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21473v1 Announce Type: new.
- Abstract: Quantification, estimating class prevalences in bags of unlabeled instances is vital in domains where aggregate statistics are more important than individual instance labels, such as biosignal monitoring, fall detection, and activity recognition.
- We investigate this issue in the challenging setting of imbalanced time series data and develop CC-GMNet-TS, a class-conditioned Gaussian mixture quantifier that combines a Transformer-based feature extractor with per-class latent mixtures.
- Unlike previous mixture-based quantifiers, which use a single Gaussian mixture shared by all classes, CC-GMNet-TS assigns each class its own compact mixture in a bounded latent space and scores segment embeddings against these class-specific components to create bag-level representations that emphasize rare but informative patterns.
- EN 要点:
- arXiv:2608.21473v1 Announce Type: new
- Abstract: Quantification, estimating class prevalences in bags of unlabeled instances is vital in domains where aggregate statistics are more important than ind…
- We investigate this issue in the challenging setting of imbalanced time series data and develop CC-GMNet-TS, a class-conditioned Gaussian mixture quantifier tha…
- Unlike previous mixture-based quantifiers, which use a single Gaussian mixture shared by all classes, CC-GMNet-TS assigns each class its own compact mixture in…
Congruence Decomposition with Neural Block Solvers for Large-Scale PCI Assignment
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21485v1 Announce Type: new.
- Abstract: Physical Cell Identity (PCI) assignment is essential for interference management in dense 5G networks.
- As cellular networks scale, PCI reuse becomes unavoidable, which may cause collisions, confusions, and multiple forms of modular interference.
- Jointly mitigating these effects gives rise to a large-scale, multi-objective combinatorial optimization problem that is difficult to solve efficiently at practical network scales.
- EN 要点:
- arXiv:2608.21485v1 Announce Type: new
- Abstract: Physical Cell Identity (PCI) assignment is essential for interference management in dense 5G networks
- As cellular networks scale, PCI reuse becomes unavoidable, which may cause collisions, confusions, and multiple forms of modular interference
- Jointly mitigating these effects gives rise to a large-scale, multi-objective combinatorial optimization problem that is difficult to solve efficiently at pract…
KAN-Robust-Bench: A Benchmark for Evaluating the Robustness of Kolmogorov-Arnold Networks
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21488v1 Announce Type: new.
- Abstract: While machine learning models have demonstrated strong performance in many domains, these models have shown profound vulnerabilities when they are exposed to adversarial threats.
- While adversarial attacks fall into various categories, the most prominent category in research studies is evasion.
- In evasion attacks, the adversary generates perturbed versions of samples, which might not be observable by human eyes.
- EN 要点:
- arXiv:2608.21488v1 Announce Type: new
- Abstract: While machine learning models have demonstrated strong performance in many domains, these models have shown profound vulnerabilities when they are exp…
- While adversarial attacks fall into various categories, the most prominent category in research studies is evasion
- In evasion attacks, the adversary generates perturbed versions of samples, which might not be observable by human eyes
The geometry of AI validation: Exact certification limits for iid best-of-N search
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21496v1 Announce Type: new.
- Abstract: AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output.
- Validation is therefore target-relative: evidence certifies deployment only in directions resolved by the interventions that produced it.
- We represent validation and deployment rules as kernels over a reliability surface.
- EN 要点:
- arXiv:2608.21496v1 Announce Type: new
- Abstract: AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output
- Validation is therefore target-relative: evidence certifies deployment only in directions resolved by the interventions that produced it
- We represent validation and deployment rules as kernels over a reliability surface
Selection of Heart Sound Segments for Synchronous Classification of Multi-channel Heart Sounds
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21499v1 Announce Type: new.
- Abstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main auscultation spots.
- Despite this, automatic heart sound analysis algorithms mostly classify patients using a single heart sound (single-channel), or, when using more than one (multi-channel), analyze each channel individually.
- To our knowledge, no prior work classifies patients through the synchronous analysis of multi-channel heart sounds, following the procedure used by physicians.
- EN 要点:
- arXiv:2608.21499v1 Announce Type: new
- Abstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main ausculta…
- Despite this, automatic heart sound analysis algorithms mostly classify patients using a single heart sound (single-channel), or, when using more than one (mult…
- To our knowledge, no prior work classifies patients through the synchronous analysis of multi-channel heart sounds, following the procedure used by physicians
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21504v1 Announce Type: new.
- Abstract: The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry.
- However, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to variations in problem formulation and chemical representation.
- As a result, reported performance may overestimate a model’s true ability to reason consistently across realistic settings.
- EN 要点:
- arXiv:2608.21504v1 Announce Type: new
- Abstract: The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry
- However, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to varia…
- As a result, reported performance may overestimate a model’s true ability to reason consistently across realistic settings
Multimodal Injury Risk and Performance Prediction in Tennis Using Weighted Ensemble Learning
- 发布时间:2026-08-25 12:00 北京时间
- 摘要:【待翻译】- arXiv:2608.21530v1 Announce Type: new.
- Abstract: Machine learning has had a positive impact on the sports industry, with one of its most promising applications being the prediction of athlete performance and injury risk.
- Recent advances have employed state-of-the-art models to improve prediction accuracy, yet progress remains limited by data availability and the reliance on subjective observations or expert assessments.
- To address these limitations, researchers in sports such as soccer, basketball, and wrestling have begun integrating heterogeneous data sources, such as wearable device readings, with traditional subjective assessments.
- EN 要点:
- arXiv:2608.21530v1 Announce Type: new
- Abstract: Machine learning has had a positive impact on the sports industry, with one of its most promising applications being the prediction of athlete perform…
- Recent advances have employed state-of-the-art models to improve prediction accuracy, yet progress remains limited by data availability and the reliance on subj…
- To address these limitations, researchers in sports such as soccer, basketball, and wrestling have begun integrating heterogeneous data sources, such as wearabl…