System translated (Gemini)

🤖 AI 速览

Today’s main theme shifts from model capabilities to system capabilities. OpenAI continues to strengthen its full-stack strategy, from data centers and chips to products, with Jalapeño aiming for lower-cost inference. Concurrently, agent governance, citation attribution, and runtime evidence …
📋 文章元数据
发布时间
2026-08-26
类型
ai-daily
字数
6704
阅读时长
32 min

2026-08-26 AI Daily | OpenAI Pushes AI Competition to Full-Stack: Accelerating In-House Inference Chips, Compute Systems, and Trustworthy Agents in Parallel Link to heading

Today’s main theme shifts from model capabilities to system capabilities. OpenAI continues to strengthen its full-stack strategy, from data centers and chips to products, with Jalapeño pointing toward lower-cost inference. Meanwhile, agent governance, citation attribution, and runtime evidence protocols are gaining traction, as reliability evaluation begins to delve into long-tail languages and cultural contexts.

📖 In-depth Guide to This Issue’s Watch List Link to heading

There are three main themes worth following today. The first is the comprehensive “full-stack” trend in AI infrastructure: OpenAI continues to integrate its compute, chips, and inference systems, and early results from Jalapeño also point to a more efficient inference architecture, a key area for engineering teams to watch. The second is that agents are entering the hard-problem phase of governance and security: from sycophancy and citation attribution to agentic security and runtime evidence protocols, it’s clear that after achieving the ability to “get things done,” trustworthiness and auditability are becoming core requirements. The third is the push for model reliability in long-tail languages and cultural contexts. Several studies today on Khmer, Nigerian Pidgin, Cyrillic, and visual-language biases remind us that evaluation systems are still far from mature. On a side note, Netflix’s exploration of becoming a streaming aggregation portal is also worth watching, as it reflects a reorganization of platform-based distribution logic.

🌐 AI Hot Topics on X Link to heading

Topic 1: xAI and Cursor Boost Grok Model Usage Limits Again Link to heading

  • Category: AI · News
  • Overview: Trending for: 7 hours ago, Related posts: 5,400
  • What happened: xAI and Cursor have once again increased the usage limits for the Grok model, drawing attention from developers and AI users to its availability and cost.
  • Why it matters: This reflects a trend of AI products competing for developer adoption and workflow integration by offering higher call limits. It also indicates intensifying competition in compute supply, cost control, and commercialization for inference models.
  • Discussion summary: Discussions on X center on whether the increased limits signify a genuine enhancement in Grok’s infrastructure and model serving capabilities, whether Cursor users will experience more stable coding, and if this will squeeze the market share of models like Claude and GPT in developer tools. Some also question the cost sustainability and the actual extent of quality improvement.

Topic 2: Shopify CEO Pushes for Claude Code to Adopt AGENTS.md Standard Link to heading

  • Category: AI · News
  • Overview: Trending for: 9 hours ago, Related posts: 2,900
  • What happened: The CEO of Shopify has publicly called for Claude Code to adopt the AGENTS.md standard, which uses a unified file to describe the rules, context, and collaboration conventions for AI programming agents within a codebase.
  • Why it matters: This is about the standardization and interoperability of AI programming tools. If different agents can read the same set of project conventions, it could reduce configuration costs, minimize errors, and improve the stability of multi-tool and multi-agent collaboration in real-world codebases.
  • Discussion summary: The debate on X focuses on whether Claude Code’s adoption could help AGENTS.md become a de facto standard and whether a unified specification can improve an agent’s understanding of project context. The point of contention is whether it’s essential industry infrastructure or just another project file that adds to the maintenance burden and risks configuration fragmentation.

Topic 3: OpenAI Engineer Surprises with Changed Appearance in Interview Link to heading

  • Category: AI · News
  • Overview: Trending for: 2 days ago, Related posts: 4,600
  • What happened: An X user’s post discussing the changed appearance of an OpenAI engineer in an interview has gone viral, sparking conversations about their identity, well-being, and background.
  • Why it matters: This kind of topic is significant because it reflects the high level of public attention on OpenAI and its key employees. It can easily impact the company’s image, talent narrative, and public perception of the AI industry’s culture and work pressure.
  • Discussion summary: The discussion mainly revolves around whether the change in appearance is real, its potential causes, and whether this level of attention is normal observation or an invasion of privacy. Some have also used it as a jumping-off point to discuss the high-intensity work environment at AI companies and the magnifying effect of public opinion.

Topic 4: OpenAI Unveils Jalapeño Chip Outpacing Nvidia in Speed and Efficiency Link to heading

  • Category: AI · News
  • Overview: Trending for: 9 hours ago, Related posts: 13,000
  • What happened: OpenAI has reportedly announced a chip named Jalapeño, claiming it surpasses Nvidia’s comparable products in speed and energy efficiency.
  • Why it matters: If this claim holds true, it signifies that leading AI companies are reducing their reliance on general-purpose GPUs through in-house chip development. This would impact compute costs, the supply chain, and future model deployment strategies.
  • Discussion Overview: Discussions on X are mainly focused on three points: whether performance comparisons are backed by public benchmarks, whether this implies OpenAI is pursuing deeper in-house hardware development, and the potential impact on Nvidia’s position in the AI compute market.

Summary of AI Public Opinion on X Today Link to heading

The main theme on X today is that the AI competition is shifting from “who has the stronger model” to “who can offer cheaper, more stable, and more deeply integrated solutions into developer workflows.” Whether it’s Grok increasing its limits, the calls to standardize Claude Code AGENTS.md, or the rumors of OpenAI developing its own chips, everything points to three key areas: compute power, interface standards, and entry-point control. The general consensus is that the core of developer tools is not just the model’s capability itself, but also its usability, cost, and collaborative efficiency within the project context. The disagreements center on whether these moves represent genuine capability improvements, marketing-driven volume releases, a burden of standardization, or unverified performance narratives. Opinions on AGENTS.md are also sharply divided: supporters see it as infrastructure for reducing operational errors and enhancing multi-agent collaboration, while opponents worry it will become an additional maintenance cost and create new fragmentation. There are three main potential risks: first, whether the compute costs behind high limits and in-house chips are sustainable; second, that performance claims without sufficient benchmark support could mislead market judgment; and third, the magnified discussions around individual engineers’ appearances and well-being, which tend to conflate industry pressure, privacy boundaries, and public curiosity.

💡 Influencer Insights Link to heading

Influencer insights are unavailable today. We recommend reading the in-depth content from the Watch List.

📚 Appendix: Today’s Watch List Update Source List Link to heading

Time Window: Last 3 days; Covering 22 sources; 35 updates in total

Stratechery by Ben Thompson (A_full) Link to heading

  • Netflix to Sell Streaming Services?, Streamers as Aggregators, Revisiting Roku
    • Published: 2026-08-25 18:00 Beijing Time
    • Summary: Netflix is considering selling other streaming services, and I think it’s a good idea; it’s also a let-down for Netflix’s original goals and potential pivots.
      • $15 / month or $150 / year.
      • Substantial analysis of the news of the day delivered via three weekly emails or podcasts.
      • Stratechery Interviews.
      • Interviews with leading public CEOs, private company founders, and discussions with fellow analysts.
    • EN Key Points:
      • Netflix is considering selling other streaming services, and I think it’s a good idea; it’s also a let-down for Netflix’s original goals and potential pivots.

OpenAI Blog (A_full) Link to heading

  • The full stack behind abundant intelligence

    • Published: 2026-08-25 15:05 Beijing Time
    • Summary: Progress in AI compounds fastest when the entire system improves together.
      • That is how I think about OpenAI’s compute strategy: one integrated system spanning data centers and chips, frontier models, our developer platform, consumer and enterprise products, and AI-native devices, with each layer strengthening the next.
      • Better software makes hardware more productive.
  • Hardware designed for our workloads improves speed and efficiency.

  • More capable models unlock better products, which generate more demand, usage, and learning.

  • EN Key points:

    • OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower…
  • Jalapeño’s first results show industry-leading speed and efficiency in AI inference

    • Published: 2026-08-25 15:00 Beijing Time
    • Summary: [TO BE TRANSLATED] - Since announcing Jalapeño, OpenAI’s first custom inference chip, we have been testing the chip and the system built around it.
      • The results show a significant performance advance: Jalapeño can serve more AI work per unit of power while also returning responses more quickly.
      • Jalapeño delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two.
      • For customers, that can mean faster responses, more responsive agents, and more reliable access as demand grows.
      • Our mission is to ensure that artificial general intelligence benefits all of humanity.
    • EN Key points:
      • Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern mod…
  • Disrupting a new covert influence campaign from Russia

    • Published: 2026-08-25 08:00 Beijing Time
    • Summary: [TO BE TRANSLATED] - OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.
      • This piece from OpenAI Blog explains how Disrupting a new covert influence campaign from Russia shapes the broader AI and infrastructure landscape.
      • It also surfaces practical implications for founders, operators, and investors following Disrupting a new covert influence campaign from Russia.
    • EN Key points:
  • OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.

  • Introducing the Admin plugin for ChatGPT Work and Codex

    • Publication Time: 2026-08-25 08:00 Beijing Time
    • Summary: [TO BE TRANSLATED] - Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.
      • This piece from OpenAI Blog explains how Introducing the Admin plugin for ChatGPT Work and Codex shapes the broader AI and infrastructure landscape.
      • It also surfaces practical implications for founders, operators, and investors following Introducing the Admin plugin for ChatGPT Work and Codex.
    • Key Takeaways (EN):
      • Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.

ArXiv cs.AI (B_intro+search) Link to heading

  • KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Summary: [TO BE TRANSLATED] - arXiv:2608.21362v1 Announce Type: new.
      • Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request.
      • Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at arbitrary positions.
      • We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position.
    • Key Takeaways (EN):
      • arXiv:2608.21362v1 Announce Type: new
      • Abstract: Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each request
  • Existing prefix-caching systems reduce this cost but require prompts to share a leading contiguous prefix, limiting effectiveness when shared content appears at…

  • We present KVBoost, a chunk-level KV cache reuse system for HuggingFace-compatible decoder models that enables reuse regardless of content position

  • AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Summary: [To be translated]- arXiv:2608.21363v1 Announce Type: new.
      • Abstract: A protocol is presented for recording the governance decisions of automated AI runtimes.
      • When a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can check offline, independent of the runtime that produced it.
      • A record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by value, and declares both what its evidence covers and what it does not.
    • EN Highlights:
      • arXiv:2608.21363v1 Announce Type: new
      • Abstract: A protocol is presented for recording the governance decisions of automated AI runtimes
      • When a runtime releases, blocks, defers, redacts, or escalates an individual output, AIREP records that decision as a single signed object that any party can ch…
      • A record carries the decision as one of a closed set of verbs under a stated policy basis, references its input, output, and evidence by hash rather than by val…
  • Reviewing Model Collapse and Countermeasures

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Summary: [To be translated]- arXiv:2608.21366v1 Announce Type: new.
      • Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors.
  • The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models.

    • Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply.
    • EN Key points:
      • arXiv:2608.21366v1 Announce Type: new
      • Abstract: Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diverse sectors
      • The advances of GenAI have actuated practitioners to use AI-synthesized data for training next-generation AI models
      • Undeniably, using synthetic data has alleviated the increasing stringent demand for data supply
  • AI Learning and Conceptual Transfer in the Game of Hidden Rules

    • Release Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21372v1 Announce Type: new.
      • Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback, representation design, rule difficulty analysis, transfer learning, generalization, and pseudo-bot-assisted human learning analysis.
      • The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of human learning data.
      • arXiv:2608.21372v1 Announce Type: new Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules… The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of huma….
    • EN Key points:
      • arXiv:2608.21372v1 Announce Type: new
      • Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules…
  • The report focuses on the Transformer-based A2C framework, Feature-Centric and Object-Centric representations, experimental findings, and classification of huma…

  • LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: 【To be translated】- arXiv:2608.21374v1 Announce Type: new.
      • Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics.
      • We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria.
      • From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%.
    • EN Key Points:
      • arXiv:2608.21374v1 Announce Type: new
      • Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspe…
      • We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-…
      • From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems…
  • SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG

  • Publication Time: 2026-08-25 12:00 Beijing Time

    • Abstract: [Translation pending] - arXiv:2608.21375v1 Announce Type: new.
      • Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and graph stores.
      • Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload size, token use, and latency, and under-fetching, which omits fields needed to answer the query.
      • We present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and license policies as a schema graph.
    • EN Key Points:
      • arXiv:2608.21375v1 Announce Type: new
      • Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector stores, and grap…
      • Exposing all tool descriptions to an LLM agent, or selecting tools only by vector similarity, causes two costly failures: over-fetching, which increases payload…
      • We present SchemaRouter, a lightweight routing layer that represents tools, endpoints, parameters, response fields, domain concepts, units, provenance, and lice…
  • RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University Students

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [Translation pending] - arXiv:2608.21379v1 Announce Type: new.
      • Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the working population - yet it is typically identified only retrospectively, after academic decline has already occurred.
      • A contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without interpreting it.
  • This paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI architecture to surface personalized insights and early burnout signals.

    • EN Key Points:
      • arXiv:2608.21379v1 Announce Type: new
      • Abstract: Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those of the work…
      • A contributing factor is that students have little structured visibility into their own study behaviour, and existing productivity tools record activity without…
      • This paper presents RIACT (Record, Insight, Analyze, Coach, Track), a web-based application that combines structured study session logging with a hybrid AI arch…
  • There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21382v1 Announce Type: new.
      • Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model’s answer is read from generated text or from per-option likelihoods.
      • Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next.
      • We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items.
    • EN Key Points:
      • arXiv:2608.21382v1 Announce Type: new
      • Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and wh…
      • Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items tha…
  • We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items

  • Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference

    • Release Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [TO BE TRANSLATED] - arXiv:2608.21393v1 Announce Type: new.
      • Abstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepower sits somewhere else, and moving sensitive records between the two creates real headaches around latency, security, and regulatory exposure.
      • IBM’s Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation.
      • In this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip accelerator for lightweight classification tasks, and Red Hat OpenShift for container orchestration.
    • EN Key Points:
      • arXiv:2608.21393v1 Announce Type: new
      • Abstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsep…
      • IBM’s Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation
      • In this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip acceler…
  • Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering

    • Release Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [TO BE TRANSLATED] - arXiv:2608.21408v1 Announce Type: new.
  • Abstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress and negative societal impacts.

    • Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal grammar, inconsistent sen-tence structures, and multiple variations in word spellings.
    • This research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data.
    • EN Key Points:
      • arXiv:2608.21408v1 Announce Type: new
      • Abstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significant distress…
      • Ro-man Urdu, a low-resource language used in Pakistan and among Urdu-speaking communities worldwide, presents additional challenges because of its informal gram…
      • This research aims to identify the most effective techniques for hate speech classification in such low-resource settings with limited data

ArXiv cs.CL (B_intro+search) Link to heading

  • Distinguishing Revision and Delayed Elaboration in Incremental Narrative Interpretation

    • Published: 2026-08-25 12:00 Beijing Time
    • Abstract: [Translation pending] - arXiv:2608.21364v1 Announce Type: new.
      • Abstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations must be updated accordingly.
      • Incremental interpretation, therefore, depends not only on what is represented but also on how the representational state evolves under new evidence.
      • We distinguish two structurally different update operators that arise in narrative interpretation: revision-driven update and delayed elaboration.
    • EN Key Points:
      • arXiv:2608.21364v1 Announce Type: new
  • Abstract: Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal representations…

  • Incremental interpretation, therefore, depends not only on what is represented but also on how the representational state evolves under new evidence

  • We distinguish two structurally different update operators that arise in narrative interpretation: revision-driven update and delayed elaboration

  • KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search

    • Published: 2026-08-25 12:00 Beijing Time
    • Abstract: arXiv:2608.21365v1 Announce Type: new.
      • Abstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in multilingual embedding models, and frequent mixed Khmer-English usage.
      • This paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search.
      • We construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplication, and document-length control.
    • EN Key Points:
      • arXiv:2608.21365v1 Announce Type: new
      • Abstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in…
      • This paper presents KSE-Web, an analysis of hybrid retrieval and LLM-assisted query expansion for Khmer semantic search
      • We construct the dataset from approximately 17K candidate Khmer titles and retain 3K cleaned full-text Khmer documents after filtering, normalization, deduplica…
  • Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning

    • Published: 2026-08-25 12:00 Beijing Time
  • Summary: [To be translated] - arXiv:2608.21369v1 Announce Type: new.

    • Abstract: Nigerian Pidgin is one of Africa’s most widely spoken languages, yet remains severely underrepresented in language model evaluation.
    • Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured.
    • We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning.
    • EN Highlights:
      • arXiv:2608.21369v1 Announce Type: new
      • Abstract: Nigerian Pidgin is one of Africa’s most widely spoken languages, yet remains severely underrepresented in language model evaluation
      • Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language under…
      • We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning
  • On the Role of Citations in Preference Data

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Summary: [To be translated] - arXiv:2608.21376v1 Announce Type: new.
      • Abstract: Many NLP tasks require systems to provide attribution in their outputs–i.e.
      • citations to grounding sources.
      • Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs.
    • EN Highlights:
      • arXiv:2608.21376v1 Announce Type: new
      • Abstract: Many NLP tasks require systems to provide attribution in their outputs–i.e
      • citations to grounding sources
      • Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs
  • Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Summary: [To be translated] - arXiv:2608.21377v1 Announce Type: new.
  • Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings.

    • This paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse?
    • Across 4,800 veracity judgments (200 statements $\times$ 6 models $\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic systems (feedback loops, reconsideration checkpoints, and iterative refinement) systematically amplifies sycophantic behavior.
    • Key Points (EN):
      • arXiv:2608.21377v1 Announce Type: new
      • Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied pr…
      • This paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse
      • Across 4,800 veracity judgments (200 statements $\times$ 6 models $\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic sys…
  • Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems

    • Publish Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21384v1 Announce Type: new.
      • Abstract: Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, creating disparities in cost and context capacity.
      • We quantify this overhead across nine production tokenizers and five languages with standardized Cyrillic and Latin representations, covering 8.37 million word forms.
      • On a corpus benchmark, Ukrainian shows 68-121% token overhead on modern tokenizers and 220% on the older cl100k, measured through full-text fertility on the BrUK and Brown corpora.
    • Key Points (EN):
      • arXiv:2608.21384v1 Announce Type: new
  • A Social Media Analysis of Discourse on the Israel–Palestine Conflict on Telegram

    • Posted on: 2026-08-25 12:00 Beijing Time
    • Abstract: [Translation Pending] - arXiv:2608.21385v1 Announce Type: new.
      • Abstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose broadcast architecture yields an unusually direct record of deliberate political communication, have not been systematically compared at scale.
      • This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning May 2021 to June 2026 and covering multiple conflict escalations.
      • It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference, and a fine-tuned BERTweet model), and a framing analysis, all evaluated against 736 manually annotated messages.
    • EN Highlights:
      • arXiv:2608.21385v1 Announce Type: new
      • Abstract: Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegram, whose br…
      • This study presents a multi-method computational analysis of 87,617 messages from sixteen Telegram channels, eight pro-Israel and eight pro-Palestine, spanning…
  • It combines sentiment analysis, three stance detection methods drawn from distinct paradigms (keyword matching, zero-shot DeBERTa via natural language inference…

  • Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

    • Release Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21415v1 Announce Type: new.
      • Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from their training data, resulting in biased behavior when processing portraits from different social groups.
      • Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limited by their reliance on a single, stereotyped viewpoint and fail to account for the diversity of social perspectives.
      • Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs multi-group counterfactual perspectives within the visual representation space and integrates them during decoding to promote equitable model behavior.
    • EN Key Points:
      • arXiv:2608.21415v1 Announce Type: new
      • Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social biases from…
      • Existing debiasing approaches typically compare token probabilities between the original and biased generations during decoding, but they are fundamentally limi…
      • Inspired by the social science principle that diversity fosters fairness, we propose Counterfactual Ensemble Decoding (CED), a novel framework that constructs m…
  • Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing

    • Release Time: 2026-08-25 12:00 Beijing Time
  • Summary: [TO BE TRANSLATED] - arXiv:2608.21423v1 Announce Type: new.

    • Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools.
    • As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures.
    • We systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pipelines.
    • EN Highlights:
      • arXiv:2608.21423v1 Announce Type: new
      • Abstract: Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools
      • As these systems move from demonstrations to deployed products, practitioners repeatedly encounter the same operational failures
      • We systematize these failures through a hands-on evaluation of ten widely used static, dynamic, cloud, orchestration, and AI red-teaming tools for unattended pi…
  • CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

    • Published: 2026-08-25 12:00 Beijing Time
    • Summary: [TO BE TRANSLATED] - arXiv:2608.21462v1 Announce Type: new.
      • Abstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties.
      • Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages.
      • But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?
    • EN Highlights:
      • arXiv:2608.21462v1 Announce Type: new
      • Abstract: Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alph…
      • Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages
  • But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do

ArXiv cs.LG (B_intro+search) Link to heading

  • Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?

    • Release Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To Be Translated] - arXiv:2608.21386v1 Announce Type: new.
      • Abstract: Given a task described by a few examples, how should a model be specialized to it?
      • Four mechanisms are available – zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork – yet the operating regime of the last is rarely mapped.
      • We run the identical four-way comparison across six tasks spanning regression, generation, language modeling, reinforcement learning, and clinical and genomic classification, holding the specialist, the context, and (where we can) the training budget fixed.
    • EN Key Points:
      • arXiv:2608.21386v1 Announce Type: new
      • Abstract: Given a task described by a few examples, how should a model be specialized to it
      • Four mechanisms are available – zero-shot, in-context attention, test-time gradient adaptation, and emitting specialist weights from a hypernetwork – yet the…
      • We run the identical four-way comparison across six tasks spanning regression, generation, language modeling, reinforcement learning, and clinical and genomic c…
  • Runtime Action Interference for AI Control of AlphaStar in StarCraft II

    • Release Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To Be Translated] - arXiv:2608.21398v1 Announce Type: new.
      • Abstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppresses, or replaces its proposed actions.
  • We contribute \emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering configured action patterns after inference.

  • RAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op.

  • EN Key Points:

    • arXiv:2608.21398v1 Announce Type: new
    • Abstract: A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, admits, suppre…
    • We contribute \emph{runtime action interference} (RAI), an AI control mechanism that preserves policy parameters while regulating action pacing and filtering co…
    • RAI releases a proposed action only when its cooldown condition is satisfied and its content detector does not flag the action; otherwise, it dispatches a no-op
  • Federated Ensemble Forecasting Under Supply-Chain Market Volatility

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [Pending Translation] - arXiv:2608.21399v1 Announce Type: new.
      • Abstract: Supply chain forecasting systems increasingly operate under market shocks, non-identically distributed regional demand, and limited willingness to centralize commercial data.
      • This work proposes Federated Ensemble Forecasting with Negative-Correlation Learning (FEF NCL), a distributed method that trains specialized forecasting experts across client nodes while discouraging redundant model errors.
      • The framework combines temporal feature encoders, client level drift scoring, reliability-weighted aggregation, and an explain ability layer that exposes the market and supplier variables most responsible for each forecast.
    • EN Key Points:
      • arXiv:2608.21399v1 Announce Type: new
      • Abstract: Supply chain forecasting systems increasingly operate under market shocks, non-identically distributed regional demand, and limited willingness to cen…
  • This work proposes Federated Ensemble Forecasting with Negative-Correlation Learning (FEF NCL), a distributed method that trains specialized forecasting experts…

  • The framework combines temporal feature encoders, client level drift scoring, reliability-weighted aggregation, and an explain ability layer that exposes the ma…

  • Class-Conditioned Gaussian Mixture Modeling for Imbalanced Time Series Quantification

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21473v1 Announce Type: new.
      • Abstract: Quantification, estimating class prevalences in bags of unlabeled instances is vital in domains where aggregate statistics are more important than individual instance labels, such as biosignal monitoring, fall detection, and activity recognition.
      • We investigate this issue in the challenging setting of imbalanced time series data and develop CC-GMNet-TS, a class-conditioned Gaussian mixture quantifier that combines a Transformer-based feature extractor with per-class latent mixtures.
      • Unlike previous mixture-based quantifiers, which use a single Gaussian mixture shared by all classes, CC-GMNet-TS assigns each class its own compact mixture in a bounded latent space and scores segment embeddings against these class-specific components to create bag-level representations that emphasize rare but informative patterns.
    • EN Key Points:
      • arXiv:2608.21473v1 Announce Type: new
      • Abstract: Quantification, estimating class prevalences in bags of unlabeled instances is vital in domains where aggregate statistics are more important than ind…
      • We investigate this issue in the challenging setting of imbalanced time series data and develop CC-GMNet-TS, a class-conditioned Gaussian mixture quantifier tha…
      • Unlike previous mixture-based quantifiers, which use a single Gaussian mixture shared by all classes, CC-GMNet-TS assigns each class its own compact mixture in…
  • Congruence Decomposition with Neural Block Solvers for Large-Scale PCI Assignment

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [Translation Pending] - arXiv:2608.21485v1 Announce Type: new.
      • Abstract: Physical Cell Identity (PCI) assignment is essential for interference management in dense 5G networks.
      • As cellular networks scale, PCI reuse becomes unavoidable, which may cause collisions, confusions, and multiple forms of modular interference.
      • Jointly mitigating these effects gives rise to a large-scale, multi-objective combinatorial optimization problem that is difficult to solve efficiently at practical network scales.
    • EN Key Points:
      • arXiv:2608.21485v1 Announce Type: new
      • Abstract: Physical Cell Identity (PCI) assignment is essential for interference management in dense 5G networks
      • As cellular networks scale, PCI reuse becomes unavoidable, which may cause collisions, confusions, and multiple forms of modular interference
      • Jointly mitigating these effects gives rise to a large-scale, multi-objective combinatorial optimization problem that is difficult to solve efficiently at pract…
  • KAN-Robust-Bench: A Benchmark for Evaluating the Robustness of Kolmogorov-Arnold Networks

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [Translation Pending] - arXiv:2608.21488v1 Announce Type: new.
      • Abstract: While machine learning models have demonstrated strong performance in many domains, these models have shown profound vulnerabilities when they are exposed to adversarial threats.
      • While adversarial attacks fall into various categories, the most prominent category in research studies is evasion.
      • In evasion attacks, the adversary generates perturbed versions of samples, which might not be observable by human eyes.
    • EN Key Points:
      • arXiv:2608.21488v1 Announce Type: new
  • The geometry of AI validation: Exact certification limits for iid best-of-N search

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21496v1 Announce Type: new.
      • Abstract: AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output.
      • Validation is therefore target-relative: evidence certifies deployment only in directions resolved by the interventions that produced it.
      • We represent validation and deployment rules as kernels over a reliability surface.
    • EN Key Points:
      • arXiv:2608.21496v1 Announce Type: new
      • Abstract: AI systems increasingly generate alternatives, inspect evidence, and deploy a selected output
      • Validation is therefore target-relative: evidence certifies deployment only in directions resolved by the interventions that produced it
      • We represent validation and deployment rules as kernels over a reliability surface
  • Selection of Heart Sound Segments for Synchronous Classification of Multi-channel Heart Sounds

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21499v1 Announce Type: new.
      • Abstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main auscultation spots.
      • Despite this, automatic heart sound analysis algorithms mostly classify patients using a single heart sound (single-channel), or, when using more than one (multi-channel), analyze each channel individually.
  • To our knowledge, no prior work classifies patients through the synchronous analysis of multi-channel heart sounds, following the procedure used by physicians.

    • EN Key Points:
      • arXiv:2608.21499v1 Announce Type: new
      • Abstract: Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four main ausculta…
      • Despite this, automatic heart sound analysis algorithms mostly classify patients using a single heart sound (single-channel), or, when using more than one (mult…
      • To our knowledge, no prior work classifies patients through the synchronous analysis of multi-channel heart sounds, following the procedure used by physicians
  • ChemDIRT: A Diversified Instruction, Representation, and Task Benchmark for Robust Chemistry-LLM Evaluation

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: - arXiv:2608.21504v1 Announce Type: new.
      • Abstract: The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry.
      • However, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to variations in problem formulation and chemical representation.
      • As a result, reported performance may overestimate a model’s true ability to reason consistently across realistic settings.
    • EN Key Points:
      • arXiv:2608.21504v1 Announce Type: new
      • Abstract: The rapid advancement of large language models (LLMs) has led to increasing interest in their application to scientific domains such as chemistry
      • However, existing chemistry benchmarks often provide only a narrow view of model capability, focusing on limited task sets while overlooking robustness to varia…
      • As a result, reported performance may overestimate a model’s true ability to reason consistently across realistic settings
  • Multimodal Injury Risk and Performance Prediction in Tennis Using Weighted Ensemble Learning

    • Publication Time: 2026-08-25 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.21530v1 Announce Type: new.
      • Abstract: Machine learning has had a positive impact on the sports industry, with one of its most promising applications being the prediction of athlete performance and injury risk.
      • Recent advances have employed state-of-the-art models to improve prediction accuracy, yet progress remains limited by data availability and the reliance on subjective observations or expert assessments.
      • To address these limitations, researchers in sports such as soccer, basketball, and wrestling have begun integrating heterogeneous data sources, such as wearable device readings, with traditional subjective assessments.
    • EN Key Points:
      • arXiv:2608.21530v1 Announce Type: new
      • Abstract: Machine learning has had a positive impact on the sports industry, with one of its most promising applications being the prediction of athlete perform…
      • Recent advances have employed state-of-the-art models to improve prediction accuracy, yet progress remains limited by data availability and the reliance on subj…
      • To address these limitations, researchers in sports such as soccer, basketball, and wrestling have begun integrating heterogeneous data sources, such as wearabl…