🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-21
- 类型
- ai-daily
- 字数
- 7511
- 阅读时长
- 36 min
2026-08-21 AI Daily Update | The Focus of AI Competition Shifts: OpenAI Discusses Power Risks, and Agent Collusion Risks Come Under Formal Scrutiny Link to heading
Today’s focus is no longer just on model capabilities but on AI entering discussions of governance and institutions. OpenAI’s new AI Futures initiative brings the issues of power concentration, abuse, and the restructuring of free societies to the forefront. Meanwhile, researchers are beginning to demand that agents with reasoning abilities undergo behavioral certification before making market decisions. On the engineering side, the evolution continues towards terminal-native, persistently running agents.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The most important development to watch today is OpenAI’s new AI Futures initiative, which elevates the discussion of “how free societies can be restructured to protect individual rights and agency in the age of transformative AI” to a strategic level. Echoing this, several papers are focusing on agent governance: issues like collusion risks in market decisions, the inadequacy of open-weight model cards, and the principle that “behavioral systems must undergo behavioral testing” all signal that AI safety is shifting from evaluating model capabilities to assessing institutions and processes.
The second major theme is the engineering of agents. Topics such as concurrency control in multi-agent systems, the dynamic graph perspective on Self-Evolving Agents, and the audit-style evaluation of investment management agents by FinSkillBench are particularly relevant for engineering teams. The question is no longer “can it answer,” but “can it execute reliably in scenarios involving shared states, tool calls, and high-stakes processes.”
Additionally, there are fundamental advancements in long-context understanding: LongNovel focuses on hallucinations in long-form summarization, while research on entity tracking shows that even smaller models can exhibit emergent, strong comprehension abilities in natural narratives. These are developments that model evaluation teams should follow.
🌐 AI Hot Topics on X Link to heading
Topic 1: Terminal-Code Brings VS Code Editor to Your Terminal Link to heading
- Category: AI · News
- Overview: Trending: 4 hours ago, Related posts: 230
- What it is: The open-source project Terminal-Code is gaining attention for bringing a VS Code-like editing experience to the terminal.
- Why it matters: This reflects a trend of AI programming tools evolving towards more lightweight, localized, and command-line-native development workflows, facilitating integration with intelligent agents, automation scripts, and remote development environments.
- Discussion summary: Discussions on X are centered on whether a terminal version of VS Code can enhance AI-assisted programming efficiency, how it integrates with workflows for tools like Devin, Gemini, and Antigravity, and the gaps in usability and extensibility compared to the traditional VS Code plugin ecosystem.
Topic 2: OpenAI Launches AI Futures Blog on Power Risks from Advanced AI Link to heading
- Category: AI · News
- Overview: Trending:, Related posts: 92
- What it is: OpenAI has launched a blog called “AI Futures,” focusing on the risks of power concentration, abuse, and governance challenges posed by advanced AI.
- Why it matters: This indicates that leading AI companies are making societal impact, institutional constraints, and safety governance—beyond just technical capability advancement—a core agenda item. These discussions could influence future AI regulation, deployment, and industry standards for responsibility.
- Discussion summary: The discussion on X revolves around whether OpenAI is genuinely confronting the power risks of advanced AI, whether corporate self-governance is sufficient, who should lead regulation (government or industry), and whether such public discourse will translate into concrete safety measures.
Topic 3: Stripe Acquires OpenRouter, Declares Singularity Began January 1 Link to heading
- Category: AI · News
- Overview: Trending: 1 day ago, Related posts: 13,000
- What it is: Stripe’s acquisition of OpenRouter is a hot topic on X, accompanied by the facetious claim that “the singularity began on January 1.”
- Why it matters: This is seen as a significant signal of integration between AI infrastructure and the model distribution layer, potentially impacting multi-model routing, payment processing, and the developer ecosystem.
- Discussion summary: The discussion is focused on whether the acquisition will alter OpenRouter’s neutrality, Stripe’s strategic intentions for entering the AI ecosystem, and whether this news is serious or a satirical marketing ploy/joke.
Topic 4: Cursor Boosts Cloud Agents with Autonomous Goals and Event Handling Link to heading
- Category: AI · News
- Overview: Trending: 1 day ago, Related posts: 2,600
- What it is: Cursor has enhanced its cloud agents with stronger autonomous capabilities, allowing them to set long-term objectives via /goal and be triggered by events like Slack thread subscriptions or scheduled tasks.
- Why it matters: This shows that AI agents are moving from “on-demand response” to “persistent execution,” which has significant implications for tool-based AI, automated workflows, and enterprise collaboration.
- Discussion summary: Discussions on X are mainly about whether these new capabilities can significantly boost productivity, if they introduce greater automation risks, and what practical advantages Cursor’s cloud agents have over existing AI programming and workflow tools.
Topic 5: Ilja Dragunov Leaves WWE as Contract Expires Link to heading
- Category: AI · Sports
- Overview: Trending for 6 hours, 47,000 related posts
- What happened: According to trending discussions on X, professional wrestler Ilja Dragunov has left WWE after his contract expired.
- Why it matters: While the event has no direct connection to AI development, the high-profile discussion in the sports entertainment field serves as a classic case study for social media trend analysis, fan sentiment recognition, and content recommendation algorithms.
- Discussion summary: Discussions on X are mainly focused on the reasons for his departure, whether he will join other wrestling promotions, WWE’s talent management strategies, and fans’ expressions of regret and anticipation for his in-ring performance and future career.
Summary of AI Public Opinion on X Today Link to heading
The main thread of AI-related discourse on X today revolves around the theme that “AI tools are moving from single-point functions to becoming more fundamental, persistent, and platform-oriented.” News related to Terminal-Code, Cursor’s cloud-based Agents, and OpenRouter are all seen as indicators that AI programming and infrastructure are evolving towards native terminal integration, automated execution, and multi-model distribution. A strong consensus is that developers want AI to be more closely integrated with local workflows, command lines, and long-term task management. This could lead to efficiency gains and is better suited for integration with agents, scripts, and collaboration tools. The main points of disagreement are whether these new forms are “genuinely boosting productivity” or just repackaged product narratives, and whether platform acquisitions and enhanced cloud agents will weaken neutrality and intensify ecosystem lock-in. Potential risks are centered on the further concentration of power and capabilities, accidental or malicious automation triggers, and whether insufficient corporate self-governance will keep “safety discussions” at a purely rhetorical level.
💡 Influencer Insights Link to heading
AI Industry Daily Briefing (August 20, 2026) Link to heading
Based on tweets from multiple AI influencers over the past 24 hours, I have summarized the following insights on tech trends, core viewpoints, and tool resources.
1. Key Tech Trends and Hot Products Watched by Influencers Today Link to heading
Today’s discussions are highly focused on the unification of the Claude ecosystem, benchmarking of on-device models, and the componentization of AI Agents (Skill/Harness).
Claude’s “Super App” Ambition: Deep Integration of Code + Design Anthropic is integrating design capabilities through Claude Code. According to @dotey, Claude Code now has Claude Design built-in. After trying it, he found it “very easy to use,” allowing for the generation of interactive React prototypes directly from local projects using
/designwithout switching contexts (and without relying on Figma). @dotey also emphasized that the key to removing the “AI feel” depends not just on the model’s capability, but more so on human aesthetics and a personalized design system. However, @dotey also pointed out that the latest version of Claude Code is extremely token-intensive for cross-session communication, and he recommends disabling it by writing to"crossSessionInbound": "refuse".On-Device Model Showdown: Qwen Securely Holds the Sweet Spot On-device deployment remains a key focus for hardcore users. @zhixianio shared a direct comparison between Gemma 4 12B Coder and Qwen 3.6-35B-A3B MoE: in a complex, long-form front-end task (like creating a complete Tetris game), the 12B Gemma encountered issues like black screens and logic freezes, while the 35B Qwen completed the task successfully. The conclusion is that 12B models have a clear ceiling when handling “long-form, stateful” tasks, and fine-tuning mainly improves efficiency rather than the upper limit of their underlying logic. Additionally, @zhixianio successfully ran the 4-bit quantized version of the official DeepSeek V4 Flash on a Mac Studio.
The Evolution of Agent Forms: From MCP to an Explosion in the Skill/Harness Ecosystem Skills are becoming the new entry point into the Agent ecosystem. ByteDance’s Coze Desktop and Tencent’s Workbuddy are competing to be the gateway for “AI Office” (according to @vista8, ByteDance’s strategy is a combination of Coze, Doubao, and TreaWork). Meanwhile, Xiaohongshu (Little Red Book) is also developing its REDSkill community, allowing users to upload and share Skill files (@ruanyf commented that this is the world’s first social media platform to create a Skill Hub). On the underlying protocol level, @dotey, citing @jakevin7, pointed out that the Apache Incubator has accepted its first Agent Harness project, Apache Maka.
The Offensive and Defensive Battle in Financial and Payment Compliance As the demand for overseas AI services increases, payment issues have become a hot topic. @AI_Jasonyu shared his experience of switching carriers after his Giffgaff account was banned and recommended opening a Starryblu Singapore card. @Pluvio9yte complained that the $200 credit from OpenAI was “look but don’t touch,” as all his payment cards were rejected.
2. Noteworthy Unique Perspectives or Industry Foresight Link to heading
A New Direction for Scaling Laws: Emphasis on Post-training @dotey relayed @jietang’s view on GLM-5.3: without changing the base model, coding capabilities were improved by 50% purely through post-training. The key distinction is between “total parameters (determining how much is learned)” and “activated parameters + effective depth (determining how deep it can think).” This marks a shift in competition from frantically piling up parameters to exploring “inference depth.”
Rebuttal to the “Small Model + Tools” Approach Addressing the popular industry view that “Small Model + Harness = Large Model,” @dotey cited Jason Wei’s opinion. Jason Wei used the analogy of cognitive reward shaping, arguing that large models internalize knowledge as “muscle memory,” while small models performing real-time retrieval is like “cramming for a test.” There’s a gap in deep understanding, speed, and stability. This provides a critical perspective on the path to achieving top-tier intelligence.
The Underlying Architecture and Commercialization of AI Programming IDEs @dotey observed that due to real-world development needs, there is a trend of migrating the Agent client from Tauri to Electron. On the monetization front, @Pluvio9yte offered a highly practical viewpoint: use AI to automatically discover security vulnerabilities (SRC), finding 10 in half an hour. The ability to execute this is a significant barrier in itself. He also shared a practical strategy for using AI to simulate real users (from a product manager’s perspective) for product testing.
The Double-Edged Sword of Open Source and the Burden of Maintenance @ruanyf shared the perspective of the SQLite author, who refuses external PRs, comparing a PR to a “free puppy” that represents a lifetime of maintenance responsibility. This sparked deep reflection among developers about the sustainability of open-source projects.
3. Recommended Tools or Resources Link to heading
Based on the analysis from various bloggers, the following tools and resources are frequently recommended:
| Tool/Resource | Recommender | Core Use & Evaluation |
|---|---|---|
| Claude Design | @dotey | Built into Claude Code, it can produce interactive React prototypes without needing Figma, bridging the gap between design and development. |
| Raycast V2 | @vista8 | An AI Chat Agent with a built-in “memory system.” It has powerful custom prompt capabilities and includes a high-precision voice input method. |
| Codex Scheduled Tasks | @Pluvio9yte | Used for daily email summaries, saving the time of reading them one by one. It’s recommended to pair it with GPT-5.4 or 5.6 to save money. |
| Coze Desktop | @vista8 | A product from ByteDance. It cleverly uses a “cloud drive” concept to bridge the context between local and cloud-based Agents, making it a powerful tool for office scenarios. |
| @atypica_AI | @Pluvio9yte | An AI tool for business commercialization research. It simulates real users to validate business viability and is more accurate than general-purpose models. |
| OpenConnector | @ruanyf | An open-source credentials gateway that prevents AI Agents from leaking core credentials and centralizes connection authorization management. |
| Xiaohongshu/WeChat Official Account Scraping API | @vista8 | Third-party data source discovered in Coze. It addresses the pain point of Agents lacking high-quality Chinese language data. |
| Logo Skill | @dotey | Directly generates a logo from the product’s IP, enhancing brand recognition and aligning with current AI product design aesthetics. |
📚 Appendix: Today’s Watch List Source Updates Link to heading
Timeframe: Last 3 days; 22 sources covered; 32 updates in total
OpenAI Blog (A_full) Link to heading
- Publication Time: 2026-08-20 15:00 Beijing Time
- Summary: - “Do we trust these parchment barriers to be sufficient to ward off the encroaching spirit of power?”
—James Madison, The Federalist Papers
- We are excited to launch AI Futures, the blog of OpenAI’s new strategic futures team.
- We are a small team whose collective goal is to answer one overarching question: How should free societies be reorganized to protect individual rights and agency while adapting to the emergence of transformative AI?
- Such questions are sometimes referred to within the broader AI safety and policy community as the concentration of power risk.
- EN Highlights:
Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.
Stampli cuts launch hours by 68% using ChatGPT Work
- Publication Time: 2026-08-20 08:00 Beijing Time
- Abstract: - Stampli is an intelligent procure-to-pay platform that connects procurement, accounts payable, supplier management, payments, and employee expenses.
- Its Deep Finance™ product transforms data transmitted through Stampli’s procure-to-pay platform into executive spending intelligence for CFOs, VPs, and other business leaders.
- Launching it meant product development, positioning, design, communication, support, and operations all running in parallel on a fixed timeline, with design resources and external contractors already committed to other priorities.
- Stampli’s marketing team used Codex to connect product background, meeting notes, decisions, and messaging guidelines into a shared system.
- With OpenAI tools, they compressed an estimated 243 hours of production work into about 77 hours, while maintaining full human review and final approval on all customer-facing content.
- EN Key Points:
- With a fixed deadline and design resources committed elsewhere, Stampli used Codex and ChatGPT Work to compress weeks of launch production into days.
ArXiv cs.AI (B_intro+search) Link to heading
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18078v1 Announcement Type: new.
- Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are prone to collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets.
- This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms, without diminishing the distinction in economic harm.
- Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion, which persists even when humans prompt the agents not to collude.
- EN Key Points:
- arXiv:2608.18078v1 Announce Type: new
- Abstract: This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be req…
- This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms…
- Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the…
Position: Profiling Game Worlds by Transition Complexity
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18079v1 Announcement Type: new.
Abstract: Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers rarely quantify how difficult the underlying transition prediction problem is at the stated interface (pixels/tokens/latents with limited history).
We propose the Transition Complexity Profile (TCP): a small, reproducible set of metrics that characterizes an environment’s (or gameplay dataset’s) induced transition kernel via (i) intrinsic single-step branching, (ii) interaction-induced uncertainty and observable adversary influence, and (iii) the span of temporal/spatial dependencies via standardized probing curves.
TCP is reported with an explicit reference distribution, protocol stochasticity, and a versioned measurement budget (sampling/resampling and fixed probe compute), enabling comparable figures across benchmarks.
EN Highlights:
- arXiv:2608.18079v1 Announce Type: new
- Abstract: Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers rarely quantify how difficult the underlying tr…
- We propose the Transition Complexity Profile (TCP): a small, reproducible set of metrics that characterizes an environment’s (or gameplay dataset’s) induced tra…
- TCP is reported with an explicit reference distribution, protocol stochasticity, and a versioned measurement budget (sampling/resampling and fixed probe compute…
- Published: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18080v1 Announce Type: new.
- Abstract: We present a review on the applications of large language models (LLMs) in the health sector, such as social media analysis, clinical conversational agents, therapeutic support tools, prompt engineering, multimodal learning, and ethical considerations.
- We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal inputs to enable early detection of depression, suicide risk assessment, personalized treatment support, and psychoeducational content generation.
- Our review highlights advancements in LLM models and annotation strategies that enhance interpretability and clinical relevance, while we also emphasize the critical role of prompt engineering for domain adaptation.
- EN Highlights:
- arXiv:2608.18080v1 Announce Type: new
- Abstract: We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therap…
- We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal input…
- Our review highlights advancements in LLM models and annotation strategies that enhance interpretability and clinical relevance, while we also emphasize the cri…
Position: Behavioral Systems Require Behavioral Tests
- Published: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18081v1 Announce Type: new.
Abstract: Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time.
However, current evaluation methods largely focus on performance outcomes, not the underlying behavioral processes that produce them.
This paper argues that AI agents must be evaluated like other behavioral systems: through systematic observation, perturbation, and interpretation of their behavior.
EN Highlights:
- arXiv:2608.18081v1 Announce Type: new
- Abstract: Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time
- Yet, current evaluation methods largely focus on performance outcomes, not the underlying behavioral processes that produce them
- This paper argues that AI agents must be evaluated like other behavioral systems: through systematic observation, perturbation, and interpretation of their acti…
- Release Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18086v1 Announce Type: new.
- Abstract: The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance.
- Although model cards have been widely adopted as transparency artifacts in model repositories, existing frameworks often fail to adequately inform downstream developers and users about the unique safety challenges posed by OWFMs.
- This position paper analyzes 500 model cards hosted on Hugging Face and argues that effective governance of OWFMs requires a multi-layered approach integrating three complementary components: (i) model cards, (ii) acceptable use policies (AUPs), and (iii) licenses.
- EN Highlights:
- arXiv:2608.18086v1 Announce Type: new
- Abstract: The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance
- Although model cards have been widely adopted as transparency artifacts in model repositories, existing frameworks often fail to adequately inform downstream de…
- This position paper analyzes 500 model cards hosted on Hugging Face and argues that effective governance of OWFMs requires a multi-layered approach integrating…
- Release Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18088v1 Announce Type: new.
- Abstract: When the effects of drone propeller faults are distributed across multiple flight log channels instead of appearing as a single diagnostic signal, safety and reliability risks can arise.
- This paper proposes a Metamorphic Artificial Age Scoring (AAS) decision-support prototype for flight-log-based drone propeller health monitoring.
The framework computes six health-related indicators from raw MATLAB matrices using selected historical real flight logs from the 2024 DronePropA public dataset: trajectory tracking error, attitude instability, thrust command burden, motor command imbalance, ESC command instability, and battery-level stress.
- EN Highlights:
- arXiv:2608.18088v1 Announce Type: new
- Abstract: Drone propeller faults can create safety and reliability risks when their effects are distributed across multiple flight-log channels rather than appe…
- This paper proposes a Metamorphic Artificial Age Score (AAS) decision-support prototype for flight-log-based drone propeller health monitoring
- Using selected historical real flight logs from the 2024 DronePropA public dataset, the framework computes six health-related indicators from raw MATLAB matrice…
- EN Highlights:
Position: Multi-Agent Systems Should Prioritize Concurrency Control
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18092v1 Announce Type: new.
- Abstract: LLM-based multi-agent systems (MAS) promise scalable collaboration, yet adding agents often reduces reliability.
- This position paper argues that many MAS failures are fundamentally concurrency control problems: agents concurrently read and write shared state, and long LLM inference windows amplify the risks of stale reads, lost updates, and inconsistent outcomes.
- Failure modes commonly attributed to coordination or communication breakdowns can be mapped directly onto classical concurrency anomalies.
- EN Highlights:
- arXiv:2608.18092v1 Announce Type: new
- Abstract: LLM-based multi-agent systems (MAS) promise scalable collaboration, yet adding agents often reduces reliability
- This position paper argues that many MAS failures are fundamentally concurrency control problems: agents concurrently read and write shared state, and long LLM…
- Failure modes commonly attributed to coordination or communication breakdowns can be mapped directly onto classical concurrency anomalies
FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18099v1 Announce Type: new.
- Abstract: Investment management is a high-stakes domain where agentic AI systems must do more than just generate plausible text.
- They must retrieve point-in-time data, compose correct computational inputs, invoke specialized methods, and produce auditable, structured outputs.
- We introduce FinSkillBench, an evaluation suite designed to measure whether language model agents can effectively leverage financial domain skills to solve investment management tasks.
- EN Highlights:
- arXiv:2608.18099v1 Announce Type: new
Abstract: Investment management is a high-stakes domain in which agentic AI systems must do more than generate plausible text
- They must retrieve point-in-time data, assemble correct computational inputs, invoke specialized methods, and produce auditable structured outputs
- We introduce FinSkillBench, an evaluation suite designed to measure whether language model agents can effectively use financial domain skills to solve investmen…
Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective
- Release Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18104v1 Announcement Type: new.
- Abstract: Large language model (LLM)-based agents are increasingly becoming self-evolving systems that can persist across interactions, maintain memories, use tools, acquire skills, refine workflows, and coordinate with other agents.
- These capabilities make agent states structural and dynamic: entities, relations, attributes, dependencies, and execution structures change with new evidence, feedback, and environmental conditions.
- Existing graph-agent surveys typically treat graphs as support structures for agent functions rather than as evolving substrates, while self-evolving-agent surveys focus on agent-level mechanisms and rarely discuss the evolution of graph topology.
- EN Key Points:
- arXiv:2608.18104v1 Announce Type: new
- Abstract: Large language model (LLM)-based agents are increasingly becoming self-evolving systems that persist across interactions, maintain memories, use tools…
- These capabilities make agent states structural and dynamic: entities, relations, attributes, dependencies, and execution structures change with new evidence, f…
- Existing graph-agent surveys typically treat graphs as support structures for agent functions rather than as evolving substrates, while self-evolving-agent surv…
- Release Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18110v1 Announcement Type: new.
- Abstract: Agentic AI is gaining new insights and advancements in the field of artificial intelligence, fostering immense potential to bring about rapid transformations across various sectors. This rapid progress and the potential to revolutionize various fields indicate a need for a deeper understanding and a firm grasp of the technology.
Moreover, an investigation into the latest research directions in agentic AI is necessary to comprehensively assess the potential scope for improvement and application. Therefore, to achieve these goals, a comprehensive review can provide researchers and practitioners with valuable insights into the current state and future research scope of agentic AI. This paper considers recently published academic contributions of agentic AI in various fields, discusses the foundations and working principles of agentic AI, traces the historical and theoretical evolution of agents in artificial systems, explores and discusses the architecture, working principles, and functionalities of Agentic AI, explores the practical applications of Agentic AI in various domains, analyzes research findings, identifies current challenges, discusses potential future research directions, and, with the help of proposed dimensions of system quality, presents a comprehensive framework for stakeholders to use and adopt Agentic AI. Consequently, this systematic review provides researchers and practitioners with a comprehensive understanding of Agentic AI, its current developments, and applications, highlighting key research gaps and outlining future research directions.
arXiv:2608.18110v1 Announce Type: new Abstract: Agentic AI is gaining new insights and advancements in the field of Artificial Intelligence, fostering significant potential to enable rapid transformation… Moreover, an investigation into the latest research directions in agentic AI needs to be conducted to comprehensively assess the potential scope for improvement….
- EN Highlights:
- arXiv:2608.18110v1 Announce Type: new
- Abstract: Agentic AI is gaining new insights and advancements in the field of Artificial Intelligence, fostering significant potential to enable rapid transform…
- Moreover, an investigation into state of the art research directions in agentic AI needs to be conducted to comprehensively assess the potential scope for impro…
- EN Highlights:
ArXiv cs.CL (B_intro+search) Link to heading
LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18082v1 Announce Type: new.
- Abstract: Despite the significant expansion of context windows in recent years, hallucinations in long-context summarization remain a challenge.
- Long novels are better suited for studying these hallucinations than news or papers because they contain intrinsic information and detailed descriptions of events and dialogues.
- However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization, nor does it fully explore how hallucinations change as the context lengthens.
- EN Highlights:
- arXiv:2608.18082v1 Announce Type: new
- Abstract: Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge
- Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events a…
- However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucina…
- Publication Time: 2026-08-20 12:00 Beijing Time
Abstract: - arXiv:2608.18083v1 Announce Type: new.
- Abstract: Understanding language requires tracking entities across the entire discourse - that is, knowing where things are and how they change, even when not explicitly stated.
- It remains unclear whether language models perform this tracking in a human-like way, partly because existing evaluations rely on artificial tasks that are far from natural language understanding and lack comparison with humans.
- Here, we evaluate entity tracking in both language models and humans (N = 48) using naturalistic narratives at multiple levels of-complexity.
EN Highlights:
- arXiv:2608.18083v1 Announce Type: new
- Abstract: Understanding language requires tracking entities across discourse - i.e., knowing where things are and how they change, even when not explicitly stat…
- Whether language models perform such tracking in a human-like fashion remains unclear, in part because existing evaluations rely on artificial tasks, far remove…
- Here, we evaluate entity tracking in both language models and humans (N = 48) using naturalistic narratives at multiple levels of complexity
Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem Proving
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18084v1 Announce Type: new.
- Abstract: Theorem proving in real-world Lean 4 projects is challenging because proofs often depend on project-specific context.
- While iterative refinement can use compiler errors to repair failed proofs, reusing failed attempts requires careful search control: some proofs provide better starting points than others, while later revisions might degrade the quality of partially correct proofs.
- We propose a compiler-guided proof search framework that balances exploration and exploitation.
- EN Highlights:
- arXiv:2608.18084v1 Announce Type: new
- Abstract: Theorem proving in real-world Lean 4 projects is challenging because proofs often depend on project-specific context
- While iterative refinement can use compiler errors to repair failed proofs, reusing failed attempts requires careful search control: some proofs provide better…
- We propose a compiler-guided proof search framework that balances exploration and exploitation
Persona-Guided LLM Agents for Task-Oriented Dialogue
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18085v1 Announce Type: new.
- Abstract: Previous work has shown that large language models (LLMs) can express different personality traits in open-ended text generation.
- However, it is unclear whether they can do so in goal-oriented dialogue without impacting task completion, and whether adapting to a user’s personality can improve interaction quality.
- We investigate these questions in task-oriented dialogue (TOD), where a system helps a user achieve a goal through multi-turn interactions.
- EN Highlights:
- arXiv:2608.18085v1 Announce Type: new
Abstract: Prior work has shown that large language models (LLMs) can express diverse personality traits in open-ended text generation
However, it remains unclear whether they can do so in a goal-directed dialogue without compromising task completion, and whether adapting to the user’s personal…
We study these questions in task-oriented dialogue (TOD), where a system helps a user accomplish a goal via multi-turn interaction
SuTRA : Structurally-Unified Tokenization with Root Awareness
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18087v1 Announcement Type: New.
- Abstract: Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes.
- This is harmful for morphologically rich Indic languages, where the basic units are complex orthographic syllables (aksharas) rather than letters.
- Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering.
- EN Key Points:
- arXiv:2608.18087v1 Announce Type: new
- Abstract: Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affix…
- This is harmful for morphologically rich Indic languages, where basic units are complex orthographic syllables (aksharas) rather than letters
- Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18089v1 Announcement Type: New.
- Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa.
- This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs.
- Restoring it typically requires labeled target-language data and retraining, which is not scalable for most African languages.
- EN Key Points:
- arXiv:2608.18089v1 Announce Type: new
- Abstract: Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa
- This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs
Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18090v1 Announcement Type: New.
- Abstract: In modern language models, there is a single internal direction that tracks the positive or negative sentiment of a sentence.
- We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion—about 1,500 fewer labels than typical supervised methods—and that the same direction appears in visual, audio, and human brain encoders that were never jointly trained.
- Recipe: Embed the nine emotion-anchored story sets in a frozen encoder and take the top principal direction of the nine averaged embeddings.
- EN Highlights:
- arXiv:2608.18090v1 Announce Type: new
- Abstract: Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels
- We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion – about 1,500 fewer labels tha…
- The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18091v1 Announcement Type: New.
- Abstract: As LLM-as-a-judge systems become increasingly common, the self-preference of LLMs—the tendency to favor their own outputs—raises growing concerns about evaluation reliability.
- However, it has been predominantly studied on generated text, where stylistic features and response quality are inevitably conflated.
- As a result, existing measurement methods cannot distinguish genuine self-preference from these confounding factors.
- EN Highlights:
- arXiv:2608.18091v1 Announce Type: new
- Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs – the tendency to favor one’s own outputs – raises growing concern…
- However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated
- As a result, existing measurements cannot separate genuine self-preference from these confounds
Abliteration Mitigation via Refusal Aliases
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18093v1 Announcement Type: New.
Abstract: Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has become a prominent safety concern due to its ability to bypass post-training alignment with only a small number of contrastive prompts.
We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted.
To hinder this process, we introduce a weight-editing method that obscures the refusal signal by applying rank-$k$ updates to residual stream writer matrices while replacing refusal-causing activations with random aliases and correcting downstream reader matrices to preserve the model’s original behavior.
EN Highlights:
- arXiv:2608.18093v1 Announce Type: new
- Abstract: Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal directio…
- We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted
- To hinder this process, we introduce a weight-editing method that obscures the refusal signal by applying rank-$k$ updates to residual stream writer matrices wh…
NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.18094v1 Announce Type: new.
- Abstract: Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource languages remain marginalized.
- We present NE-BERT, a domain-specific multilingual encoder model trained on approximately 8.3 million sentences spanning 9 Northeast Indian languages and 2 anchor languages (Hindi, English), a linguistically diverse region with minimal representation in existing multilingual models.
- By employing weighted data sampling and a custom SentencePiece Unigram tokenizer, NE-BERT outperforms IndicBERT-V2 and MuRIL across all 9 Northeast Indian languages, with an average perplexity reduction of 15.97x and 7.64x respectively, and a 1.50x improvement in tokenization capability over mBERT.
- EN Highlights:
- arXiv:2608.18094v1 Announce Type: new
- Abstract: Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource lang…
- We present NE-BERT, a domain-specific multilingual encoder model trained on approximately 8.3 million sentences spanning 9 Northeast Indian languages and 2 anch…
- By employing weighted data sampling and a custom SentencePiece Unigram tokenizer, NE-BERT outperforms IndicBERT-V2 and MuRIL across all 9 Northeast Indian langu…
ArXiv cs.LG (B_intro+search) Link to heading
Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.16913v1 Announcement Type: new.
- Abstract: Road safety monitoring has historically been reactive, relying on collision record analysis after fatalities and injuries have already occurred.
- Proactively identifying high-risk locations and dangerous driving behaviors before accidents occur is a critical but underexplored challenge.
- This paper addresses this gap by utilizing connected vehicle telemetry data from the Greater Sydney area in Australia to detect and predict near-miss risky driving events at the Local Government Area (LGA) level.
- EN Key Points:
- arXiv:2608.16913v1 Announce Type: new
- Abstract: Road safety monitoring has historically been reactive, relying on crash-record analysis after fatalities and injuries have already occurred
- Proactive identification of high-risk locations and dangerous driving behaviour before incidents occur is a critical but underexplored challenge
- This paper addresses this gap using connected vehicle telemetry data from Greater Sydney, Australia, to detect and forecast near-miss risky driving events at th…
- Abstract: - arXiv:2608.16913v1 Announcement Type: new.
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.16925v1 Announcement Type: new.
- Abstract: We construct an instrument that can read from a single fit, without an oracle, whether the operator assumed by a hybrid partial differential equation parameter estimator is wrong, and distinguish it from purely unidentifiable parameters.
- On a self-adjoint parabolic inverse problem, an information-matrix statistic with a plug-in scale and per-seed parameters has a median of 0.19 under correct specification, with a rejection rate of 0.033 against a pre-registered upper bound of 0.10, which rises to 224 and 85 under two misspecifications, triggered with every repetition.
- On a correctly specified but non-identifiable design, it remains silent—0.050 at n=200, Clopper-Pearson [0.024, 0.090]—while a rank statistic collapses to zero at the pre-registered boundary c_5^*=2.15x10^{-3}. Thus, two readings from a single fit separate the failures in two of the three designs, an achievable result for a deployable test.
- EN Key Points:
- arXiv:2608.16925v1 Announce Type: new
- Abstract: We build an instrument that reads, from a single fit and with no oracle, whether the operator a hybrid PDE-parameter estimator postulates is wrong-and…
- On one self-adjoint parabolic inverse problem, an information-matrix statistic with plug-in scale and per-seed parameter has median 0.19 under correct specifica…
- On a correctly specified but non-identifiable design it stays mute-$0.050$ at $n=200$, Clopper-Pearson $[0.024, 0.090]$-while a rank statistic collapses to zero…
Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training
- Publication Time: 2026-08-20 12:00 Beijing Time
- Summary: - arXiv:2608.16926v1 Announcement Type: New.
- Abstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training costs while maintaining model performance.
- However, existing methods usually treat data value as a relatively static property and pay limited attention to the compatibility between data and the capability distribution of the target model.
- To address this issue, we propose Data-DPO, a target-model-oriented SFT data selection method.
- EN Key Points:
- arXiv:2608.16926v1 Announce Type: new
- Abstract: Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while…
- However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capabili…
- To address this issue, we propose Data-DPO, a target model-oriented SFT data selection method
Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training
- Publication Time: 2026-08-20 12:00 Beijing Time
- Summary: - arXiv:2608.16927v1 Announcement Type: New.
- Abstract: As supervised fine-tuning data continues to expand, selecting high-value subsets from large candidate pools is crucial for reducing training costs and improving model performance.
- Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local noise.
- We address this limitation by formulating data selection as a coarse-to-fine hierarchical coverage problem and propose MASS.
- EN Key Points:
- arXiv:2608.16927v1 Announce Type: new
- Abstract: As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and i…
- Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained s…
- We address this limitation by formulating data selection as a coarse-to-fine hierarchical coverage problem and propose MASS
Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification
- Publication Time: 2026-08-20 12:00 Beijing Time
- Summary: - arXiv:2608.16928v1 Announcement Type: New.
- Abstract: The automatic sensitivity classification of organizational documents is a critical but underexplored problem, where the consequences of misclassification include regulatory violations and security breaches.
While AI-based approaches offer a scalable alternative to manual review, their reliability depends fundamentally on the integrity of training data.
A pervasive but underreported problem in this domain is label leakage: residual classification markers embedded within document bodies that allow models to exploit superficial shortcuts instead of learning true content-based sensitivity signals, resulting in inflated and unreliable performance estimates.
- EN Highlights:
- arXiv:2608.16928v1 Announce Type: new
- Abstract: Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification ra…
- While AI-based approaches offer a scalable alternative to manual review, their reliability depends fundamentally on the integrity of training data
- A pervasive but underreported problem in this domain is label leakage: residual classification markers embedded within document bodies that allow models to expl…
- EN Highlights:
Mr.Dec: Daily-Scale Longitudinal Multimodal Modeling for 30-Day Readmission Prediction
- Release Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.16929v1 Announce Type: new.
- Predicting 30-day readmission is crucial for assessing patient stability and optimizing healthcare resources.
- As clinical risk evolves with the accumulation of evidence during hospitalization, capturing these dynamic trajectories is essential.
- However, many existing methods compress complex longitudinal histories into fixed representations, often losing the granular, day-level clinical signals that reflect the patient’s changing physiological state.
- EN Highlights:
- arXiv:2608.16929v1 Announce Type: new
- Abstract: Predicting 30-day hospital readmission is essential for assessing patient stability and optimizing healthcare resources
- As clinical risk evolves with the accumulation of evidence during hospitalization, capturing these dynamic trajectories is essential
- However, many existing approaches compress the complex longitudinal history into fixed representations, often losing the granular, day-level clinical signals th…
EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning
- Release Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.16930v1 Announce Type: new.
- Existing multi-task learning methods rely on hard sharing, multi-path or multi-expert, adaptive sharing, and dynamic expansion.
- However, their capacity changes are often limited by predefined structures or triggered by task boundaries and conflicting signals.
- This raises a fundamental question: Can a network start with precise single-path computation and only grow new, independent paths when evidence from sustained optimization emerges?
- EN Highlights:
- arXiv:2608.16930v1 Announce Type: new
SW-ProxyCE: Zero-Query Adversarial Transfer from Public EEG Encoders to Private Downstream Models
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.16931v1 Announcement Type: new.
- Abstract: Electroencephalography (EEG) foundation models have recently emerged as a promising paradigm for EEG decoding by learning reusable representations from large-scale heterogeneous neural recordings.
- However, the public release of EEG foundation encoders, while facilitating downstream development, also introduces a previously unexplored security risk: publicly available representations may make private downstream models vulnerable to attacks.
- This paper studies adversarial transfer attacks in the deployment of EEG foundation models in a public-encoder and private-downstream setting, where the attacker has white-box access to the published encoder and a small task-matched labeled reference set, but cannot access or query the victim’s parameters, outputs, or gradients.
- EN Highlights:
- arXiv:2608.16931v1 Announce Type: new
- Abstract: Electroencephalography (EEG) foundation models have recently emerged as a promising paradigm for EEG decoding by learning reusable representations fro…
- However, the open release of EEG foundation encoders, while facilitating downstream developments, also introduces a previously unexplored security risk: publicl…
- This paper investigates adversarial transfer attacks in EEG foundation model deployment in a public-encoder and private-downstream setting, where attackers have…
DOW-KE: Anchor-Free Multi-Layer Knowledge Editing via Direct End-to-End Weight Optimization
- Publication Time: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.16932v1 Announcement Type: new.
- Abstract: The multi-layer locate-then-edit method for knowledge editing first optimizes the target residual stream activations (anchors) of selected layers, and then implements them layer by layer as weight updates.
- This pipeline optimizes intermediate representations but deploys multi-layer weight updates whose joint effect through a true forward pass is never optimized itself: no matter how the anchors are set or propagated, each update comes from local solving, so propagation-induced decay and distortion are not corrected, leaving a closed gap between the anchor targets and the realized edits.
- We propose DOW-KE, an anchor-free method based on a single principle: what is optimized must be exactly what is deployed.
- EN Highlights:
- arXiv:2608.16932v1 Announce Type: new
Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery
- Published: 2026-08-20 12:00 Beijing Time
- Abstract: - arXiv:2608.16963v1 Announce Type: new.
- Abstract: Learning analytics often treats unsupervised clusters of intelligent tutoring system (ITS) logs as learner types that should predict learning.
- We test that assumption on EdNet-KT3.
- Clustering study-strategy features (resource use, revision, video, problem practice) for 5,000 active learners yields a silhouette-selected parent cut ($k=5$), with 4 contrasting poles (reading-dominant, video-dominant, revision-dominant, and problem-first) plus a large near-mean residual (~64.9%).
- EN Key Points:
- arXiv:2608.16963v1 Announce Type: new
- Abstract: Learning analytics often treats unsupervised clusters of intelligent tutoring system (ITS) logs as learner types that should predict learning
- We test that assumption on EdNet-KT3
- Clustering study-strategy features (resource use, revision, video, problem practice) for 5{,}000 active learners yields a silhouette-selected parent cut ($k=5$)…