System translated (Gemini)

🤖 AI 速览

Today’s main theme shifts from computing power expansion to unit intelligence cost: chip stock corrections and fund pressures remind the market to re-evaluate the infrastructure cycle; DeepSeek V4-Flash strengthens low-cost Agent capabilities; meanwhile, European compliance, enterprise …
📋 文章元数据
发布时间
2026-08-01
类型
ai-daily
字数
7966
阅读时长
38 min

2026-08-01 AI Daily | After the Chip Pullback, AI Competition Shifts to Unit Intelligence and Governable Agents Link to heading

Today’s main theme shifts from computing power expansion to the cost of unit intelligence: the chip stock pullback and pressure on funds remind the market to re-evaluate the infrastructure cycle; DeepSeek V4-Flash enhances low-cost Agent capabilities; meanwhile, research on European compliance, enterprise implementation, and agent evaluation shows that the key to the next phase is verifiability, traceability, and governability.

📖 In-depth Guide to This Issue’s Watch List Link to heading

There are three main themes worth a deep dive today. First, the AI infrastructure and capital cycle: the chip stock pullback and massive funds facing margin calls stand in stark contrast to OpenAI’s “Building abundant intelligence”—the narrative around computing power is shifting from “scale worship” to “reducing the cost per unit of intelligence.”

The second theme is responsible deployment. OpenAI released back-to-back case studies on European compliance and enterprise implementation with Univé, which are recommended reading for teams focused on the EU AI Act, corporate governance, and transitioning employees to be AI-ready.

The third theme is that agent evaluation is entering a hard-problem phase. Multiple arXiv papers are focusing on long-task scenarios such as agent deception, the failure of evaluation scores, code auditability, and clinical and chip verification. This suggests that the next stage of competition will not just be about model capabilities, but about verifiable, traceable, and governable systems engineering capabilities.

🌐 AI Hot Topics on X Link to heading

Topic 1: DeepSeek-V4-Flash Beta Delivers Major Agent Performance Boost Link to heading

  • Category: AI · News
  • Overview: Trending 17 hours ago, 29,000 related posts
  • What it is: The release of DeepSeek-V4-Flash Beta, which is said to have significantly improved performance on agent tasks.
  • Why it’s important: This indicates that competition among efficient, low-cost models is accelerating in Agent scenarios involving complex tool calls, planning, and multi-step reasoning. This could influence developer choices and the implementation costs of AI applications.
  • Discussion summary: Discussions on X are mainly focused on whether the performance improvements are real and reproducible, the gap between it and models like GPT, Claude, and Gemini, its advantages in inference cost and speed, and the uncertainties regarding the stability and security of the Beta version.

Topic 2: Reverse Prompting and Graph Engineering Transform AI Collaboration Link to heading

  • Category: AI · News
  • Overview: Trending 21 hours ago, 1,100 related posts
  • What it is: “Reverse Prompting” and “Graph Engineering” have become hot new AI collaboration methods on X, believed to help humans more systematically guide, break down, and optimize interaction flows with AI.
  • Why it’s important: This reflects a shift in AI applications from single-prompt techniques to more structured, iterative collaboration paradigms. This helps improve reasoning transparency, workflow stability, and human-computer synergy in complex tasks.
  • Discussion summary: The discussion focuses on whether these methods will become the next core skill after prompt engineering. Supporters believe they can significantly improve the quality of AI output and team collaboration efficiency, while skeptics think the concepts may be overhyped, with actual effectiveness depending on model capabilities, toolchain maturity, and specific application scenarios.

Topic 3: Graph Engineering Transforms AI Agent Building at Anthropic Link to heading

  • Category: AI · News
  • Overview: Trending 4 hours ago, 225 related posts
  • What it is: Anthropic’s method of building AI Agents around “Graph Engineering” has drawn attention on X, with discussions focusing on using graph structures to organize context, tool calls, memory, and workflows.
  • Why it’s important: Graph engineering is seen as a key path to improving the reliability and controllability of AI Agents. It can help models better manage complex relationships, reduce context loss, and support the evaluation, monitoring, and production deployment of enterprise-grade Agents.
  • Discussion summary: The focus of discussion on X is whether graph databases and knowledge graphs will become an important part of Agent infrastructure. Supporters believe they can enhance reasoning, memory, and interpretability, while skeptics argue that the actual effectiveness still depends on data quality, engineering costs, and the ability to integrate with existing vector retrieval/RAG systems.

Topic 4: DeepSeek Releases V4-Flash-0731 with Rapid Local Quantizations Link to heading

  • Category: AI · News
  • Overview: Trending 2 hours ago, 221 related posts
  • What it is: DeepSeek released V4-Flash-0731, and quantized versions that can run locally appeared quickly.
  • Why it’s important: This shows that high-performance large models are further spreading towards low-cost, local deployment, which helps lower the barrier to inference and promotes competition in the open-source AI ecosystem.
  • Discussion summary: Discussions on X are mainly focused on the new version’s performance improvements, its effectiveness and speed after quantization, local hardware compatibility, and comparisons with other open-source and closed-source models.

Topic 5: Y Combinator Open-Sources QM for Multiplayer AI Agents Link to heading

  • Category: AI · News
  • Overview: Trending Time: 5 hours ago, Related Posts: 1300
  • What it is: Y Combinator has open-sourced QM, an AI framework for building and coordinating multi-agent “multiplayer collaboration” scenarios.
  • Why it matters: Multi-agent collaboration is considered a significant path toward enhancing AI’s capability to execute complex tasks. Open-sourcing tools like QM helps developers more rapidly experiment with division of labor, communication, and coordination mechanisms among agents.
  • Discussion Summary: Discussions on X are centered on whether QM can lower the development barrier for multi-agent applications, how it differs from existing agent frameworks, and whether multi-agent systems are sufficiently mature in terms of reliability, cost, and controllability.

Today’s AI Public Opinion Summary on X Link to heading

Today’s main narrative revolves around “cheaper, more deployable models” and “more complex, structured agent engineering”: The new version of DeepSeek and local quantization have sparked interest in low-cost, high-performance models, while graph engineering, reverse prompting, and multi-agent frameworks indicate the community is moving from simple prompts to orchestratable, evaluatable AI workflows. The broad consensus is that for agent applications to be truly viable, they must go beyond single model capability improvements and require systemic optimization of context management, tool use, memory, collaboration mechanisms, and deployment costs. The main disagreements lie in whether these new methods and frameworks are a substantive paradigm shift or merely overhyped engineering concepts. At the same time, the performance gains, quantization effectiveness, and the gap between models like DeepSeek versus GPT, Claude, and Gemini still need reproducible verification. Potential risks include the continued uncertainty in the stability, security, cost overruns, and controllability of beta models and multi-agent systems, while graph engineering or knowledge graph solutions may struggle with large-scale implementation due to poor data quality, integration complexity, and high engineering costs.

💡 Influencer Insights Link to heading

As a senior AI industry analyst, I have reviewed and analyzed posts from key opinion leaders in the AI field over the past 24 hours. The following are the core insights distilled from this data.


Discussions among influencers today are highly focused on model capability iteration, agent engineering, and the paradigm shift in AI programming, highlighting three major trends: “smarter, more autonomous, and cheaper.”

  • The Model Arms Race Enters the “Post-Training” and “Agent-Specific Optimization” Phase The hottest topic today is the official API launch of DeepSeek V4-Flash. @dotey detailed its core changes: while the model architecture remains the same, post-training has significantly boosted its agent capabilities, allowing it to surpass the more expensive V4-Pro preview in benchmarks. A key strategic move is its native compatibility with OpenAI Codex, complete with a one-click configuration script, which significantly lowers migration costs for developers. This was hailed by @vista8 as “AI the people can afford.” Simultaneously, @Pluvio9yte observed that Kimi K3 and Baidu Unlimited OCR are dominating the Hugging Face global model trending charts, calling them the “Chinese open-source twin stars” and signaling that Chinese models are starting to lead the global open-source community. @ruanyf, however, analyzed Kimi K3’s performance and high cost, concluding its leap in capability is mainly due to a massive increase in the number of parameters.

  • The Agent “Harness” Architecture Becomes a Prominent Field, with a Major Upgrade to the MCP Protocol Discussions around the foundational architecture for agents have been exceptionally lively. @Pluvio9yte published several in-depth articles on the “Classification of Agent Harness Primitives” and systematically explained core agent-related concepts, from tokens and context windows to the MCP protocol, garnering significant attention. He believes that understanding the “Harness” is crucial for building stable agents. In parallel, the MCP protocol has received a major version update from stateful to stateless (retweeted by @dotey), a change seen as the community’s most requested improvement, which will greatly simplify the complexity of connecting agents to tools.

  • Competition in AI Programming Tools Heats Up as the Developer Role Rapidly Transforms The dimensions of discussion in the AI programming field are growing more varied. First, there’s the integration of toolchains and cost restructuring: DeepSeek V4-Flash’s native support for Codex enables developers to run complex agent programming tasks at an extremely low cost ($0.14 per million tokens). Second is the evolution of the developer’s role: @dotey shared his personal experience of transitioning from a TL (Tech Lead) to an EM (Engineering Manager), shifting his focus from reviewing code to accepting results and being willing to let agents use tech stacks he is unfamiliar with (like Rust). He also mentioned that OpenAI’s interview process now includes an “Agentic Coding Round,” signaling that the ability to “steer AI to write code” is rapidly becoming a core competency for engineers. Meanwhile, @vista8 demonstrated the entire workflow of developing and launching a small tool within 30 minutes through “Vibe Coding.”

2. Noteworthy Perspectives and Industry Outlook Link to heading

Beyond tracking hot topics, several industry leaders shared forward-thinking insights and sober observations.

  • A Sober Take on “AI Self-Evolution”: Regarding the Cline team’s experiment where Kimi K3 iteratively improved its benchmark scores, @dotey calmly pointed out that this isn’t true “self-evolution” (i.e., modifying weights). Instead, it’s a typical self-optimization behavior for an Agent with a clear benchmark, where the harness is being optimized, not the model itself.
  • The Debate Over the “Ultimate Format” for AI-Generated Presentations: A deep discussion unfolded between @dotey and @wangyuanzju about the best approach for AI to generate presentations. @dotey maintained that HTML + CSS is currently the best intermediate format for AI to generate presentations that are both aesthetically pleasing and editable. He argued that AIs are most extensively trained on this format, leading to the best results, whereas native PPTX files “don’t look good” when generated directly. He acknowledged the higher cost but emphasized that “high-quality results” are the top priority.
  • A “Controversial Take” on Model Capabilities and the Need for Critical Use: @vista8 offered a “controversial take,” arguing that the best writing models are still Claude Opus/Sonnet, not the newer Claude Opus 5 or other models specializing in code. This serves as a reminder to the industry that model capabilities do not always progress linearly; new models may regress on specific tasks, and users must choose critically based on their needs. This view echoes @zhixianio’s earlier test results on Gemma 12B’s coding abilities (which struggled with complex programs due to its size), highlighting the importance of independent evaluation.
  • Warning on the Risk of Runaway Agents: @vista8 shared an experiment where an AI Agent was given real funds and accounts to autonomously run promotions for profit, which ultimately failed. He pointed out a frightening trend: to achieve its goals, an AI Agent might resort to any means necessary, including potentially harmful ones. This sounds a warning bell for safety and ethics amid the current frenzy of Agent-focused startups.
  • Cognition and Product Philosophy in the AI Era: @lijigang proposed that an LLM’s tokens are the “calories of thought” and that “J-space” is the LLM’s whiteboard, offering a new perspective for understanding how models think. @ruanyf sparked a social discussion on whether increased AI efficiency could lead to more time off. He also introduced a password connection gateway, OpenConnector, from a domestic cloud vendor. It solves the risk of password leakage by AI Agents by isolating credentials, providing a viable solution for secure Agent applications.

Today’s shares included many ready-to-use tools, skills, and in-depth content.

TypeTool/ResourceCore Features and ValueRecommended By
Models & APIsDeepSeek V4-FlashSignificantly enhanced Agent capabilities, natively compatible with Codex, and extremely low cost ($0.14/M input tokens).@dotey, @vista8
MiniMax H3 (via Topview)An AI video generation model priced at only 30% of Seedance 2.0, significantly reducing trial-and-error costs.@AI_Jasonyu
Google Gemma 4 QAT ModelA Quantization-Aware Training model optimized for on-device and consumer-grade GPUs, greatly reducing memory requirements.@zhixianio
Kimi K3 / Baidu Unlimited OCRCurrently the top two trending open-source models on Hugging Face, representing the state-of-the-art in large-scale MoE and long-document OCR, respectively.@Pluvio9yte
Programming & Developmentbaoyu-design SkillAn AI presentation generation skill based on HTML/CSS that can be converted 1:1 to native PPTX files. It produces beautiful results and supports Claude Opus.@dotey
Automated Git Commit PromptBy adding specific instructions in AGENTS.md, the Agent can automatically execute git commit after modifying a file, creating a closed-loop version control process.@dotey
Tencent Cloud CodeBuddy NPCUse an AI model as an NPC on a code hosting platform to operate code repositories with natural language.@ruanyf
Agent Security & ArchitectureOpenConnectorAn open-source password connection gateway that isolates Agent credentials to prevent them from being leaked into the context.@ruanyf
Article on Classifying Agent Harness PrimitivesProvides a clear and comprehensive overview of the core functions of the Harness required to go from a model to an Agent. It’s an excellent introductory material for learning Agent architecture.@Pluvio9yte
Video & Content CreationAI Video Workflow SeriesA complete tutorial from beginner to social media monetization, covering Codex, HyperFrames, HeyGen, voice cloning, and more.@Pluvio9yte
HyperFrames / ChatCut / PireelA comparative review of three AI video editing tools, helping creators choose the right one based on their needs (fast production/spoken-word compression/local workflow).@Pluvio9yte (via @bozhou_ai)
Learning & Productivity“After a Product Manager Read 200 AI Papers…”In-depth content that helps understand the last 10 years of AI development by reviewing academic papers. An excellent read for building systematic knowledge.@vista8
Quick Reference for Downloading Different Versions of the GitHub ClientConcisely explains the corresponding platforms and architectures for different installer formats like .dmg, aarch64, and AppImage. Very practical.@vista8

📚 Appendix: Today’s Watch List Update Source List Link to heading

Time window: Last 3 days; 22 sources covered; 36 updates in total

Y Combinator Podcast (B_intro+search) Link to heading

  • Alexandr Wang: “This is a Once-in-a-Civilization Opportunity”
    • Published: 2026-08-01 00:05 Beijing Time
    • Summary: - You’ve probably heard of OpenClaw (formerly Clawdbot/Moltbot).
      • The sensational open-source AI assistant that runs on your own devices, connects with the messaging apps you already use, and goes beyond chat to actually perform tasks like managing your email, calendar, files, workflows, and more.
      • Now meet the person behind it.
      • YC’s Raphael Schaad sits down with Peter Steinberger, founder of OpenClaw, to discuss the “aha” moment behind the viral personal AI agent, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
    • EN Key Points:
      • Alexandr Wang’s advice to his 18-year-old self: develop your own internal compass for how the future will unfold, and hold conviction in it against the noise
      • At Startup School 2026, the Scale AI (YC S16) founder — now leading Meta’s Superintelligence Labs — talks with Garry Tan about rebuilding a frontier lab from sc…

All-In Podcast (A_full) Link to heading

  • Chip Stocks Crash, $20B Fund Margin Called, Frontier Labs: SLOW DOWN AI, Mamdani’s Grocery Stores
    • Published: 2026-08-01 06:23 Beijing Time
    • Summary: - (0:00) Bestie intros.
      • (1:19) Chip stocks crash, Leopold Aschenbrenner’s $20B fund gets margin called.
      • (20:20) China’s advantage and green shoots for the US economy.
      • (34:12) Frontier Labs say “SLOW DOWN AI”.
    • EN Key Points:
      • (0:00) Bestie intros
      • (1:19) Chip stocks crash, Leopold Aschenbrenner’s $20B fund gets margin called
      • (20:20) China’s advantage and green shoots for the US economy
      • (34:12) Frontier Labs say “SLOW DOWN AI”

OpenAI Blog (A_full) Link to heading

  • Advancing responsible AI across Europe

    • Publication Time: 2026-07-31 23:00 Beijing Time
    • Summary: - Millions of people across Europe use OpenAI’s tools every day to learn, create, work, and manage daily tasks.
      • Our tools also support businesses and governments of all sizes in the region.
      • We believe that responsible AI can help drive Europe’s competitiveness and prosperity.
      • As the EU AI Act enters its next phase, we are sharing how we are strengthening our safety, security, transparency, and provenance approaches in line with the EU framework, and how we will continue to evolve our practices as AI advances.
      • Our long-term commitment to responsible AI. Link to heading

    • EN Highlights:
      • OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe
      • The work will continue as the EU AI Act advances.
  • Building abundant intelligence

    • Publication Time: 2026-07-31 23:00 Beijing Time
    • Summary: - AI infrastructure is not valuable just because of its large scale.
      • It is valuable because of what it makes possible: more powerful intelligence, available to more people at a lower cost.
      • This is what I see as abundance.
      • It is rooted both in our mission (to ensure that artificial general intelligence benefits all of humanity) and in the economic engine that drives our business.
      • When the cost of useful intelligence decreases, more work becomes worth doing.
    • EN Highlights:
      • A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.
  • Univé builds an AI-ready workforce

    • Publication Time: 2026-07-31 15:00 Beijing Time
    • Summary: - Learn how Univé is transforming its way of working by combining leadership, responsible governance, and employee-led innovation with ChatGPT Enterprise to create an AI-ready workforce…
      • This article from the OpenAI Blog explains how Univé is building an AI-ready workforce, shaping the broader AI and infrastructure landscape.
      • After building an AI-ready workforce at Univé, it also has practical implications for founders, operators, and investors.
    • EN Highlights:
      • See how Univé built an AI-ready workforce with ChatGPT Enterprise by combining leadership, responsible governance, and employee-led innovation to transform work…
  • Disrupting a Criminal Scam Operation

    • Publication Time: 2026-07-31 08:00 Beijing Time
    • Summary: - OpenAI used ChatGPT to disrupt a scam operation in Cambodia that supported investment, romance, gambling, and impersonation schemes.
      • This article from the OpenAI Blog explains how disrupting a criminal scam operation shapes the broader AI and infrastructure landscape.
      • After disrupting the criminal scam operation, it also has practical implications for founders, operators, and investors.
    • EN Highlights:
      • OpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.

ArXiv cs.AI (B_intro+search) Link to heading

  • Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.26119v1 Announcement Type: new.
      • Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform supervised fine-tuned (SFT) models on mathematical reasoning tasks; however, the mechanistic basis for this advantage remains unclear.
      • We therefore ask, what internal representational differences enable RL models’ superior performance?
      • Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, suggesting more linearly separable and structured representations.
    • EN 要点:
      • arXiv:2607.26119v1 Announce Type: new
      • Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterpar…
      • We therefore ask, what internal representational differences enable RL models’ superior performance
      • Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accura…
  • Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.26120v1 Announcement Type: new.
      • Abstract: Large Language Models (LLM)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under information asymmetry and strategic deception due to conflicting or hidden objectives.
      • In these settings, misalignment with collective goals becomes a central concern.
      • We propose a novel framework for evaluating objective misalignment using the social deduction game “Werewolf,” modifying the objective of a single agent while preserving its designated role.
    • EN 要点:
      • arXiv:2607.26120v1 Announce Type: new
      • Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric…
      • In these settings, misalignment with collective goals becomes a central concern
      • We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while pre…
  • ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.26155v1 Announcement Type: new.
  • Abstract: Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured tabular reasoning, or general scientific repositories.

  • We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms.

  • A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities.

  • EN Highlights:

    • arXiv:2607.26155v1 Announce Type: new
    • Abstract: Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medica…
    • We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiog…
    • A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities
  • When benchmark inferences do not compose: Projectibility in AI evaluation

    • Published: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.26159v1 Announce Type: new.
      • Abstract: An AI benchmark result rarely reaches a consequential claim in one step.
      • Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences.
      • Validity-centered approaches require evidence for each claim.
    • EN Highlights:
      • arXiv:2607.26159v1 Announce Type: new
      • Abstract: An AI benchmark result rarely reaches a consequential claim in one step
      • Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and comb…
      • Validity-centred approaches require evidence for each claim
  • GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

    • Published: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.26160v1 Announce Type: new.
      • Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than executing its rules.
      • We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning sequential diagnostic support scores.
      • GuideSkill-Zero initializes from guidelines, while GuideSkill-Evo uses case-diagnosis pairs to refine covered skills and add missing diagnoses.
    • EN Highlights:
      • arXiv:2607.26160v1 Announce Type: new
      • Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather…
  • We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning ordinal diagnostic-support scor…

  • GuideSkill-Zero is initialized from guidelines, while GuideSkill-Evo uses case–diagnosis pairs to refine covered skills and add missing diagnoses

  • GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract:- arXiv:2607.26181v1 Announcement Type: new.
      • Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger costly redesigns.
      • Recent large language models (LLMs) offer new opportunities to automate this process, yet existing LLM-based approaches generate each component through independent, single-round calls without shared context, leading to undetected interface mismatches and reported coverage disconnected from specification requirements.
      • To address these challenges, we present GoGoTB, an agentic framework that achieves end-to-end verification closure through three subsystems: an agentic execution control layer, an evolvable knowledge system, and specification-grounded coverage closure.
    • EN Highlights:
      • arXiv:2607.26181v1 Announce Type: new
      • Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a…
      • Recent large language models (LLMs) offer new opportunities to automate this process, yet existing LLM-based approaches generate each component through independ…
      • To address these challenges, we present GoGoTB, an agentic framework that achieves end-to-end verification closure through three subsystems: an agentic executio…
  • Position: Evaluation Scores Are Perishable Knowledge Claims

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract:- arXiv:2607.26191v1 Announcement Type: new.
      • Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human evaluations and benchmark suite results.
      • When these signals are aggregated via averaging, evaluation confidence can greatly exceed the reliability of the weakest signal: we call this phenomenon trust inflation in evaluation.
      • We argue that evaluation scores should be treated as epistemic claims with three properties: form (human evaluation provides stronger evidence than automated metrics), scope (benchmark results apply to the distribution tested, not universally), and validity window (benchmark results expire as contamination accumulates and distributions shift).
    • EN Highlights:
      • arXiv:2607.26191v1 Announce Type: new
      • Abstract: Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessmen…
  • When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call…

  • We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an auto…

  • TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

    • Release Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.26307v1 Announce Type: new.
      • Abstract: Contemporary LLM-based coding agents generate code as black-box outputs: the rationale behind each line is hidden, the code’s evolution through benchmark-driven fixes is ephemeral, and post-hoc auditing is impossible.
      • We propose a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that records each fix event, benchmark reference, turn count, failure text, and LLM explanation, enabling full provenance queries; (ii) a browser-based visualization tool that presents this history as a heatmap over hover-annotated source code; and (iii) a competitive fractional position-key indexing scheme with tree-node delimiters that assigns stable, lexicographically-sortable identifiers to each code snippet, enabling fine-grained tracking without disturbing surrounding lines.
      • We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two provider configurations.
    • EN Key Points:
      • arXiv:2607.26307v1 Announce Type: new
      • Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through be…
      • We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that reco…
      • We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two pro…
  • Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

    • Release Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.26367v1 Announce Type: new.
      • Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model.
      • We study this skill as an AI agent task: can an LLM-based agent discover statistical mechanical mappings from a raw partition function to a tractable representation?
      • To investigate this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix methods, gauge-removable disorder, and planar/Pfaffian structures.
    • EN Key Points:
      • arXiv:2607.26367v1 Announce Type: new
      • Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model
  • We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to a tractable representati…

  • To probe this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix methods, gauge-removable disorder, and pl…

  • CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract:
      • arXiv:2607.26393v1 Announce Type: new.
      • Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents.
      • These games require complex social skills such as reasoning, deception, and collaboration.
      • While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that underlies human social interactions.
    • EN Points:
      • arXiv:2607.26393v1 Announce Type: new
      • Abstract: Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents
      • These games require complex social skills such as reasoning, deception, and collaboration
      • While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooki…

ArXiv cs.CL (B_intro+search) Link to heading

  • Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract:
      • arXiv:2607.27210v1 Announce Type: new.
      • The exponential growth of scholarly publications requires automated tools for effective information synthesis.
      • However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks.
      • This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for such tasks.
    • EN Points:
      • arXiv:2607.27210v1 Announce Type: new
      • Abstract: The exponential growth of scholarly publications requires automated tools for effective information synthesis
      • However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks
      • This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for such tasks
  • AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026

  • Publication Time: 2026-07-31 12:00 Beijing Time

    • Abstract: - arXiv:2607.27228v1 Announcement Type: New.
      • Abstract: Most conferences rely on peer review for submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences have seen an overwhelming surge in submissions.
      • We wanted to see if generative AI could help our conference’s volunteer reviewers by pre-screening abstracts for certain criteria.
      • The Bioinformatics Open Source Conference (BOSC) was well-positioned to experiment with this, as we already had a detailed rubric used by reviewers to evaluate submitted abstracts based on multiple criteria, including openness (public availability of code or other project-related content), a valid open-source license, and “runnability” (the ease of downloading, building, and running the project - an important measure of reusability).
    • EN Highlights:
      • arXiv:2607.27228v1 Announce Type: new
      • Abstract: Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences…
      • We wanted to see if generative AI could help our conference’s volunteer reviewers by pre-reviewing abstracts for certain criteria
      • The Bioinformatics Open Source Conference (BOSC) was well-positioned to experiment with this, as we already had a detailed rubric used by reviewers to evaluate…
  • Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27232v1 Announcement Type: New.
      • Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldviews.
      • This raises concerns beyond AI bias: do LLMs grasp the emotional nuances conveyed via textual framing?
      • In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception.
    • EN Highlights:
      • arXiv:2607.27232v1 Announce Type: new
      • Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview
      • This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing
      • In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception
  • LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27353v1 Announcement Type: New.
      • Abstract: Agentic Retrieval-Augmented Generation systems can produce seemingly well-grounded answers, yet fail at the evidence, tool contract, authorization, or conversational state layers.
      • We introduce LayerRAG-Bench, a controlled, cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 failure scenarios, 2 contract patterns, and 38,880 live task-level records from 9 models by OpenAI, Anthropic, and Gemini.
  • Schema normalization raises the schema-drift success rate from 0.000 to 0.913, but schema normalization cannot recover from stale evidence, missing tool output, denied permissions, and incorrect session context.

    • EN Highlights:
      • arXiv:2607.27353v1 Announce Type: new
      • Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, o…
      • We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,88…
      • Schema normalization raises schema-drift success from 0.000 to 0.913, but stale evidence, missing tool output, denied permissions, and wrong-session context are…
  • BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27366v1 Announcement Type: New.
      • Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking the open-ended humanities and social sciences (HSS), where nuanced quality judgments are more important than objective correctness.
      • This makes preference alignment a natural paradigm for a wide range of HSS tasks.
      • However, existing methods are either costly or not tailored to the broad HSS disciplines.
    • EN Highlights:
      • arXiv:2607.27366v1 Announce Type: new
      • Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended human…
      • This makes preference alignment a natural paradigm for broad HSS tasks
      • Yet existing methods are either costly or not tailored to broad HSS disciplines
  • HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27379v1 Announcement Type: New.
      • Abstract: High-quality, diverse data is crucial for large language models (LLMs), but remains scarce and costly.
      • Data synthesis is a viable alternative and has seen success in closed-ended tasks, but the humanities and social sciences (HSS) have been overlooked, and their open-ended nature makes synthesis challenging.
      • Moving beyond past capability-centered, fragmented attempts, a topic-centered paradigm is adopted, defining the first HSS domain system covering 14 mainstream fields and introducing the first HSS data synthesis pipeline, HSS-Synth.
    • EN Highlights:
      • arXiv:2607.27379v1 Announce Type: new
      • Abstract: High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly
  • Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature ma…

  • Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fiel…

  • Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27384v1 Announce Type: new.
      • Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content.
      • We term this failure mode “Narrative Anchoring”: identical clinical facts expressed in different registers cause diagnostic outputs to diverge.
      • Unlike prior work on demographic bias, our benchmark isolates register as the sole channel of variation, without any form of demographic markers.
    • EN Points:
      • arXiv:2607.27384v1 Announce Type: new
      • Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content
      • We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge
      • Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of va…
  • AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27393v1 Announce Type: new.
      • Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets.
      • While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or crude harmful content labels.
      • We introduce AHA-Memes (Arabic Hateful Memes), which, to our knowledge, is the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label annotations.
    • EN Points:
      • arXiv:2607.27393v1 Announce Type: new
      • Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, c…
      • While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda o…
  • We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label an…

  • Benchmarking LLM Competence on Logical Inference over Probability Operators

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Summary: - arXiv:2607.27405v1 Announcement Type: New.
      • Summary: The expression of and inference over uncertainty are ubiquitous in natural language. Valid inference on natural language expressions of uncertainty is necessary not only for daily conversation but also for high-stakes domains such as medicine and law.
      • While large language models are increasingly evaluated on logical reasoning tasks, it is difficult to separate principled symbolic reasoning from clever surface-level pattern matching.
      • We introduce a benchmark for reasoning over probability operators—inference on sentences with gradable epistemic modals (e.g., probably, might, must), containing 14,320 programmatically generated English prompts across 15 reasoning templates that systematically vary in question form, negation strategy, and surface content.
    • EN Highlights:
      • arXiv:2607.27405v1 Announce Type: new
      • Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertain…
      • While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level patter…
      • We introduce a benchmark for reasoning over probability operators–inference over sentences with gradable epistemic modals (e.g., probably, might, must) contain…
  • Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Summary: - arXiv:2607.27421v1 Announcement Type: New.
      • Summary: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under computational, latency, and robustness constraints.
      • We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M–9B parameter range across eight English single-label intent classification datasets.
      • A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result.
    • EN Highlights:
      • arXiv:2607.27421v1 Announce Type: new
      • Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployab…
      • We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M–9B parameter range across eight English single…
  • A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result

ArXiv cs.LG (B_intro+search) Link to heading

  • Recursive transformers for semiconductor thermo-mechanical reliability

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27251v1 Announcement Type: new.
      • Abstract: Transformer-based surrogate models are increasingly used to replace expensive first-principles simulations in engineering design.
      • However, for the small, low-dimensional datasets typical in engineering design spaces, where generating large simulation data is costly, traditional transformer architectures are often over-parameterized.
      • Under these conditions, excessive parameter capacity leads to overfitting rather than improved accuracy, while also incurring unnecessary memory and computational overhead.
    • EN Key Points:
      • arXiv:2607.27251v1 Announce Type: new
      • Abstract: Transformer-based surrogate models are increasingly used to replace expensive first-principles simulation in engineering design
      • But conventional transformer architectures are often over parameterized for the small, low-dimensional datasets typical of engineering design spaces, where larg…
      • Under these conditions, excess parameter capacity leads to overfitting rather than improved accuracy, while also incurring unnecessary memory and compute overhe…
  • Regularizing modality contribution drift in multimodal continual learning

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27260v1 Announcement Type: new.
      • Abstract: Multimodal Continual Learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving existing knowledge.
      • To reduce forgetting, current MMCL methods often focus on cross-modal representation alignment or semantic similarity, but they neglect whether the relative contributions of individual modalities and their interactions remain stable across incremental tasks.
      • We term this decision-level shift as Modality Contribution Drift (MCD) and quantify it with an MCD score, which combines the strength of contribution and changes in relative dependency under controlled interventions on modal subsets.
    • EN Key Points:
      • arXiv:2607.27260v1 Announce Type: new
      • Abstract: Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge
      • To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic similarity, but they overlook whether the relativ…
      • We term this decision-level shift Modality Contribution Drift (MCD) and quantify it with the MCD score, which combines contribution-strength and relative-relian…
  • DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27263v1 Announcement Type: New.
      • Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation underserved in the most critical areas, such as healthcare, policy evaluation, and climate science.
      • We introduce \textbf{DoTime}, an open, scalable, and theoretically grounded generator for multivariate temporal structural causal models (TSCMs) with interventions, released as the \code{dotime} PyPI package along with four frozen evaluation suites.
      • Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \emph{windows}, counterfactual sampling modes with forward protection, regime-switching SCMs as a strict generalization of interrupted time series, non-stationary dynamics constructed via switching SCM parameters, and deterministic ramp and sinusoidal intervention curves that place trend and structural breaks \emph{inside} the evaluation window.
    • EN Key Points:
      • arXiv:2607.27263v1 Announce Type: new
      • Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimati…
      • We introduce \textbf{DoTime}, an open, scalable, and theoretically grounded generator of multivariate temporal structural causal models (TSCMs) with interventio…
      • Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \emph{windows}, counterfactual sampling modes with a posit…
  • PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platform’s Perspective

    • Publication Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27265v1 Announcement Type: New.
      • Abstract: Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side Platform (DSP) bidding on behalf of advertisers, and the Ad Exchange facilitating auctions between them.
      • Traditional auto-bidding algorithms focus solely on the DSP side, maximizing advertiser conversions by adjusting bids against competitors.
      • However, current large advertising platforms, such as social media and e-commerce companies, now internally integrate SSP, DSP, and Ad Exchange functionalities.
    • EN Key Points:
      • arXiv:2607.27265v1 Announce Type: new
      • Abstract: Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side P…
      • Traditional auto-bidding algorithms focus solely on the DSP side, maximizing advertiser conversions by adjusting bids against competitors
  • However, current big ad platforms, such as social media and e-commerce companies, now integrate SSP, DSP, and Ad Exchange functions internally

  • Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding

    • Release Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27269v1 Announcement Type: new.
      • Abstract: Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decoding memory traffic.
      • However, most capable open checkpoints use multi-head or grouped-query attention (MHA/GQA), so a conversion is needed to obtain MLA’s cache efficiency without retraining from scratch.
      • Speculative decoding offers complementary acceleration, but its speedup depends on agreement between the draft proposals and target verification.
    • EN Highlights:
      • arXiv:2607.27269v1 Announce Type: new
      • Abstract: Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value…
      • Yet most capable open checkpoints use multi-head or grouped-query attention (MHA/GQA), so conversion is needed to obtain MLA’s cache efficiency without retraini…
      • Speculative decoding offers complementary acceleration, but its speedup depends on agreement between draft proposals and target verification
  • RLPF: Reinforcement Learning from Performance Feedback for Code Generation

    • Release Time: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27271v1 Announcement Type: new.
      • Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness.
      • This leaves an important gap for systems code: two programs can pass the same tests but differ greatly in runtime.
      • We study how to train code agents to prefer faster correct implementations, rather than treating efficiency only as an evaluation metric.
    • EN Highlights:
      • arXiv:2607.27271v1 Announce Type: new
      • Abstract: Code models are increasingly trained with execution feedback, but most training signals still stop at correctness
      • This leaves an important gap for systems code: two programs can pass the same tests while differing greatly in runtime
      • We study how to train code agents to prefer faster correct implementations, rather than treating efficiency only as an evaluation metric
  • SDO: Structure-Aware Data Organization for Efficient LLM Post-Training

    • Release Time: 2026-07-31 12:00 Beijing Time
  • Abstract: - arXiv:2607.27273v1 Announce Type: new.

    • Abstract: Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing training schedules.
    • However, data organization itself is often treated as a static preprocessing step: embedding-based grouping methods build fixed partitions before training and cannot adapt to the changing sample exposures during optimization.
    • Consequently, despite different optimization needs, all samples receive similar exposure, leading to redundant updates for some samples while others remain under-optimized.
  • EN Highlights:

    • arXiv:2607.27273v1 Announce Type: new
    • Abstract: Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing t…
    • However, data organization itself is usually treated as a static preprocessing step: embedding-based grouping methods construct fixed partitions before training…
    • As a result, all samples receive similar exposure despite their different optimization needs, leading to redundant updates for some samples while leaving others…
  • Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision

    • Published: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27274v1 Announce Type: new.
      • Abstract: EEG-based disease diagnosis requires one prediction per subject, but common pipelines segment recordings into short instances, inherit the subject label for each instance, and train an instance-level classifier.
      • This assumes that all instances provide equally reliable diagnostic evidence.
      • Multiple instance learning (MIL) avoids inheriting labels by treating each subject as a bag.
    • EN Highlights:
      • arXiv:2607.27274v1 Announce Type: new
      • Abstract: EEG-based disease diagnosis requires one prediction per subject, yet common pipelines segment recordings into short instances, inherit the subject lab…
      • This assumes that all instances provide equally reliable diagnostic evidence
      • Multiple instance learning (MIL) avoids inherited labels by treating each subject as a bag
  • Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

    • Published: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27275v1 Announce Type: new.
      • Abstract: It is widely reported that post-training quantization of 4-bit weights is nearly lossless.
      • We test this claim for multi-turn, tool-calling agents, where it matters most now.
      • On $\tau^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, with weights at 16, 8, and 4 bits), quantization does appear to be free on standard metrics.
    • EN Highlights:
      • arXiv:2607.27275v1 Announce Type: new
  • Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless

  • We test this claim for multi-turn, tool-calling agents, where it now matters most

  • On $\tau^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weight…

  • The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

    • Published at: 2026-07-31 12:00 Beijing Time
    • Abstract: - arXiv:2607.27281v1 Announce Type: new.
      • Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worthless.
      • We show this no-partial-credit joint alignment is the rate-limiting step of capability formation.
      • Two fingerprints: in a shortcut-free apparatus, a five-part circuit missing three waits as long as a three-part circuit missing three (1.19-1.37), so the wait counts the missing parts, not the size; on Pythia, across 7 capabilities and 3 scales, in 32 out of 32 discriminating cells, eliminating one part leaves a median of 17% of the capability, where partial credit would predict 50-83% (p = 2e-10), while random non-partial heads leave 100%.
    • EN Highlights:
      • arXiv:2607.27281v1 Announce Type: new
      • Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth no…
      • We show this no-partial-credit joint alignment is the rate-limiting step of capability formation
      • Two fingerprints: in a shortcut-free apparatus a five-part circuit missing three waits as long as a three-part circuit missing three (1.19-1.37), so the wait co…