🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-15
- 类型
- ai-daily
- 字数
- 5090
- 阅读时长
- 24 min
2026-08-15 AI Daily Update | AI CapEx Continues to Accelerate: Anthropic Valuation Rumors and Nvidia’s $500 Billion Bet Link to heading
This issue’s focus shifts from model performance to the industry’s foundation: Anthropic’s valuation rumors, Meta’s AI narrative, and Nvidia’s massive investment continue to escalate the infrastructure race. Meanwhile, DeepSeek reinforces the commoditization trend with low-cost models and open-source Agent tools. On the research front, there is a growing emphasis on the governance and controllability of multi-LLM collaboration.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The most important developments to follow today fall into three categories. The first is AI capital expenditure and industry narratives: Stratechery continues to ask “How long will the CapEx train run?”, which, combined with Anthropic’s valuation, Meta’s AI declaration, Nvidia’s $500 billion bet, and Grok’s resurgence, outlines the financial and strategic framework for the next round of infrastructure competition. The second is the shift of agents from “functional” to “governable”: multi-LLM collaborative governance, low-cost social simulations, and world model benchmarks are pushing the research focus towards system-level behavior and controllability. The third involves the detailed battles of engineering implementation: papers on MoE route flipping, segmented prompt optimization, and latent memory read/write operations remind us that beyond model capabilities, stability and efficiency are equally decisive factors.
🌐 AI Hotspots on X Link to heading
Topic 1: Andrew Ng Shares AI Engineering Skills Map from 10,000 Job Postings Link to heading
- Category: AI · News
- Overview: Trending Time: 2 hours ago, Related Posts: 655
- What it is: Andrew Ng shared an AI engineering skills map compiled from 10,000 job postings, summarizing the abilities most valued by companies.
- Why it’s important: This reflects that talent demand in the AI field is shifting from “knowing how to use models” to “being able to implement AI in products and business operations.” It provides valuable reference for learning paths, hiring criteria, and curriculum design.
- Discussion summary: The main discussion on X revolves around the accuracy of this skills map, which abilities are most worth prioritizing, and whether it indicates a current trend in AI employment towards engineering and application rather than pure research.
Topic 2: DeepSeek Launches V4-Pro with Agent Upgrades and Low Costs Link to heading
- Category: AI · News
- Overview: Trending Time: 2 days ago, Related Posts: 34,000
- What it is: DeepSeek launched V4-Pro, featuring upgraded Agent capabilities, flexible inference tiers, compatibility with the OpenAI Responses API, and lower API costs with peak/off-peak pricing.
- Why it’s important: This indicates that high-performance models are becoming cheaper, easier to integrate, and more oriented towards agent workflows. It could intensify price competition among AI models and lower the barrier to entry for developers.
- Discussion summary: The discussion on X is focused on whether V4-Pro’s benchmark performance is genuine, its cost-effectiveness compared to models like Grok, Claude, and OpenAI, and whether the low prices and peak/off-peak billing can truly attract developers and production use cases.
Topic 3: DeepSeek Releases Open-Source Agent Harness and V4 Models with Dynamic Pricing Link to heading
- Category: AI · News
- Overview: Trending Time: 9 hours ago, Related Posts: 1,000
- What it is: DeepSeek released an open-source Agent Harness and V4 models, introducing a dynamic pricing mechanism.
- Why it’s important: This signifies that large model capabilities, agent frameworks, and inference pricing are becoming more open and commoditized, potentially reshaping AI model competition, inference costs, and ecosystem control.
- Discussion summary: The discussion on X centers on whether dynamic pricing will drive down industry inference prices, whether the open-source Agent Harness could become a new de facto standard, and if this will impact the profits of vendors who rely on model resale or integration.
Topic 4: DeepSeek Releases Modular Open-Source AI Agent Harness Link to heading
- Category: AI · News
- Overview: Trending Time: 6 hours ago, Related Posts: 493
- What it is: DeepSeek released a modular, open-source AI Agent Harness for more convenient building, testing, and composition of agent workflows.
- Why it’s important: Such tools lower the barrier to Agent development, promote standardization in agent orchestration, evaluation, and deployment within the open-source ecosystem, and may also influence developers’ choices regarding closed-source Agent platforms.
- Discussion summary: The discussion on X is focused on whether its modular design is genuinely practical, how it differs from existing Agent frameworks, its degree of openness and reproducibility, and whether it could become the new infrastructure for developers building production-grade agents.
Topic 5: Pelican on Bike Becomes Top AI Illustration Test Link to heading
- Category: AI · News
- Overview: Trending Time: , Related Posts: 49
- What it is: “A pelican riding a bicycle” unexpectedly became a popular prompt on X for testing the capabilities of AI image generators, with many users comparing the output from different models.
- Why it matters: This is significant because it shows how the public uses a simple, reproducible, and creative prompt to quickly assess an AI’s image comprehension, compositional stability, and detail generation capabilities. It also reflects that the competition among image models has shifted to a more intuitive user experience battleground.
- Discussion overview: The discussion on X primarily focused on which models could most accurately and naturally depict the absurd scene of a “pelican riding a bicycle,” as well as the differences in the generated results in terms of humor, realism, and controllability. Some also used this to discuss whether prompt-based tests can serve as a valid benchmark for model capabilities.
Topic 6: God as First Vibe Coder Sparks Developer Humor Link to heading
- Category: AI · Entertainment
- Overview: Trending: 10 hours ago, Related posts: 198
- What it is: A joking topic on X about “God being the first vibe coder” has sparked humorous banter among developers and AI enthusiasts.
- Why it matters: This type of meme reflects the changing programming culture in the AI era. Code generation, natural language development, and a “just state the goal, not the details” workflow are being widely discussed. It also shows the growing acceptance and imagination surrounding AI-assisted programming.
- Discussion overview: The discussion focuses on whether this is a humorous analogy for vibe coding or a satirical take on the evolution of programming methods. Some find it a great meme that captures AI programming trends, while others believe it oversimplifies the complexity of software engineering.
Topic 7: Game Developers Test AI to Build Engines in Months Link to heading
- Category: AI · Entertainment
- Overview: Trending: 18 hours ago, Related posts: 659
- What it is: Game developers are testing AI coding agents to build game engines and playable prototypes within months.
- Why it matters: This shows that AI is moving beyond content generation tools and into the software engineering infrastructure layer, potentially shortening game development cycles significantly and validating the real-world capabilities of long-context agents in complex engineering tasks.
- Discussion overview: The focus on X is on whether AI can genuinely reduce the development time for engines, gameplay, and asset pipelines from years to months. Supporters are optimistic about the efficiency gains for small teams, while skeptics argue that reliability, debugging, aesthetics, asset consistency, and AAA-level complexity remain major bottlenecks.
AI Public Opinion Summary on X Today Link to heading
The main thread of public opinion today is clear: the AI discussion is shifting from “whose model is stronger” to “who is cheaper, more accessible, and can be practically implemented in engineering and business.” Whether it’s Andrew Ng’s skill map, DeepSeek’s V4-Pro, or the open-source Agent Harness, the focus is on applications, engineering, and agent-based workflows. A strong consensus is emerging that enterprises value the ability to integrate AI into their products, data, and processes. Similarly, developers are more concerned with API costs, compatibility, and reproducible toolchains rather than pure research metrics. Disagreements center on two points: first, how “real” are these benchmarks, skill maps, and product claims? Second, will open-source and low prices truly change the ecosystem, or just shift the competition from model performance to delivery experience? Potential risks include price wars and dynamic pricing further squeezing industry profits, while the overly optimistic hype around Agents may obscure the hard challenges of debugging, stability, aesthetics, and reliability in complex engineering projects.
💡 Influencer Insights Link to heading
No influencer insights for today. Recommended reading from the Watch List deep dives.
📚 Appendix: Today’s Watch List Source Updates Link to heading
Time window: Last 3 days; 22 sources covered; 23 updates in total
All-In Podcast (A_full) Link to heading
- Anthropic’s $2T IPO, Zuck’s AI Manifesto, Nvidia’s $500B AI Bet, Grok’s Comeback
- Published: 2026-08-15 04:11 Beijing Time
- Summary: - (0:00) Gavin Baker joins the show.
- (2:36) Anthropic IPO report: $2T valuation, $100B+ run rate, October listing.
- (27:32) Zuck’s AI manifesto: What it means for Meta and frontier AI.
- (56:41) Summit speaker announcement.
- EN Key Points:
- (0:00) Gavin Baker joins the show
- (2:36) Anthropic IPO report: $2T valuation, $100B+ run rate, October listing
- (27:32) Zuck’s AI manifesto: What it means for Meta and frontier AI
- (56:41) All-In Summit Speaker Announcements
Stratechery by Ben Thompson (A_full) Link to heading
- 2026.33: The CapEx Train Keeps Rolling
- Posted: 2026-08-15 01:00 Beijing Time
- Abstract: - (Photo by Underwood Archives/Getty Images).
- Welcome back to This Week in Stratechery!
- As a reminder, every week on Friday, we send an overview of the content in the Stratechery bundle; highlighted links are free for everyone.
- Additionally, you have complete control over what we send to you.
- On that note, here are some of our favorites this week.
- EN Points:
- (Photo by Underwood Archives/Getty Images)
- Welcome back to This Week in Stratechery
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
- Additionally, you have complete control over what we send to you
Two Minute Papers (B_intro+search) Link to heading
- Claude AI Failed 650 Times…Then Beat The Human Record
- Posted: 2026-08-14 16:42 Beijing Time
- Abstract: - ❤️ Check out Weights & Biases and sign up for a free demo here:.
- 📝 The paper is available here:.
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
- Claude AI failed 650 times… then beat the human record.
- EN Points:
- ❤️ Check out Weights & Biases and sign up for a free demo here:
- 📝 The paper is available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…
ArXiv cs.AI (B_intro+search) Link to heading
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
- Posted: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11207v1 Announcement Type: New.
- When two LLM agents with opposing goal structures interact over multiple rounds, the lack of a shared objective function produces not competition, but collapse: the visitor surrenders, the site agent stops varying its approach, and the dialogue terminates without achieving the stated goals of either agent.
- This paper asks whether a control-theoretic governance layer can substitute for the missing objective function.
- The Experience Orchestrator (EO) addresses this problem in a simulated financial services environment, where a site agent steers a visitor towards contacting an advisor while the visitor maintains psychologically realistic resistance.
- EN Points:
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11210v1 Announce Type: new.
- Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter.
- Despite decades of methodological work, researchers almost always fall back on uniform priors.
- The main reason is that constructing informative priors from scientific literature is slow and requires both domain and statistical expertise.
- EN Highlights:
- arXiv:2608.11210v1 Announce Type: new
- Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter
- Despite decades of methodological work, researchers almost always fall back on uniform priors
- The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise
A Forced-Structure Reduction and Verifiable Bounds for Conway’s 99-Graph
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11211v1 Announce Type: new.
- Abstract: The Conway 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists.
- We report on a systematic, fully reproducible attack initiated by an autonomous AI research agent, scored against the track’s partial credit metric.
- Our verifiable contributions are: (1) an exhaustive proof that no circulant graph on $\mathbb{Z}/99$ satisfies more than $3366/4950=68.0%$ of the constraints (33 of 49 difference classes), with the same upper bound for other abelian groups of order 99; (2) a forced-structure reduction: $\lambda=1$ makes each neighborhood a perfect matching and $\mu=2$ bijects external vertices to unmatched neighbor-pairs, collapsing existence to a 12-regular graph on 84 vertices, encoded for CP-SAT and verified by recovering the unique $\mathrm{srg}(9,4,1,2)$; (3) a verified framework for prescribed-automorphism orbit existence (fixed-point-free and single-fixed-point actions, checked on $\mathrm{srg}(9,4,1,2)$ and the Paley graph $\mathrm{srg}(13,6,2,3)$), and (4) a best-verified artifact at $69.43%$, with evidence suggesting this is a robust frontier (fourteen distinct approaches, none surpassing it) entangled with the outstanding problem, as any provable bound below 4950 would be a proof of non-existence.
- EN Highlights:
- arXiv:2608.11211v1 Announce Type: new
Abstract: Conway’s 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists
We report a systematic, fully reproducible attack by an autonomous AI research agent, scored under the track’s partial-credit metric
Our verifiable contributions are: (1) an exhaustive proof that no circulant graph on $\mathbb{Z}/99$ satisfies more than $3366/4950=68.0%$ of the constraints (…
- Release Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11212v1 Announcement Type: New.
- Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-driven numerical disturbance (simulated 4-bit KV cache quantization read by a protected BF16 gate) pushes tokens across the decision boundary and flips the expert-triggered token.
- This paper proposes no new mitigation measures; it provides a causal apparatus, empirical findings, and detection limit results.
- A four-run apparatus prices the route-mediated fraction (RMF) of quantization damage, token-level attribution decomposes it by mechanism, and pre-registered probes carry the results across three architectures.
- EN Highlights:
- arXiv:2608.11212v1 Announce Type: new
- Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance – simulated 4-bit KV-cache quantization read…
- This paper proposes no new mitigation; it supplies a causal apparatus, empirical findings, and a detection-limit result
- A four-run apparatus prices the route-mediated fraction (RMF) of quantization damage, a token-level attribution decomposes it by mechanism, and pre-registered p…
Poor Man’s Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
- Release Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11215v1 Announcement Type: New.
- Abstract: Simulating societies of many large language model (LLM) agents is costly, but the questions posed by such simulations are often macroscopic: phase behavior, stylized facts, and the scaling with the number of agents $N$, rather than the cognition of any single agent.
- We transform a statistical physics observation into a method: replace each LLM agent with a low-parameter model fitted to a few hundred to a few thousand inexpensive queries, then run the society at arbitrary $N$ on a laptop.
- Whether this is effective is determined before the simulation runs and depends primarily on each agent’s perception.
- EN Highlights:
- arXiv:2608.11215v1 Announce Type: new
- Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phas…
We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queri…
Whether this works is decided before the simulation runs, chiefly by what each agent perceives
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
- Release Time: 2026-08-14 12:00 Beijing Time
- Abstract:- arXiv:2608.11216v1 Announce Type: new.
- Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single method can dominate in an environment.
- This makes it an ideal testbed for AI coding agents acting as autonomous researchers – a setting in which the improvement direction is not specified in advance, unlike the engineering specification tasks that dominate current agent benchmarks.
- We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided world-model starter under a fixed computational budget.
- EN Key Points:
- arXiv:2608.11216v1 Announce Type: new
- Abstract: World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dom…
- This makes it an ideal testbed for AI coding agents acting as autonomous researchers–a setting in which the improvement direction is not specified in advance,…
- We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided world-model starter under a fixed com…
MaSRead: Content-Addressed Reading of Replicated Latent Stores
- Release Time: 2026-08-14 12:00 Beijing Time
- Abstract:- arXiv:2608.11218v1 Announce Type: new.
- Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text.
- Merged by a conflict-free replicated data type, these fragments form a store that aggregates under any delivery order or duplication.
- However, later queries (unknown at encoding time) cannot reliably read the merged cache: co-located fragments interfere, making co-location unaddressable.
- EN Key Points:
- arXiv:2608.11218v1 Announce Type: new
- Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text
- Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication
Yet a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so colocation is not addressability
From Monolithic to Modular: Segment-level Automatic Prompt Optimization
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11219v1 Announcement Type: New.
- Abstract: Automatic Prompt Optimization (APO) typically rewrites prompts monolithically, which can improve one behavior while degrading others.
- We propose SAPO, a segment-level APO method that decomposes prompts into role, context, task, and output format, then applies targeted improvements based on the top 5 and bottom 5 examples.
- The optimization loop uses an LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation.
- EN Highlights:
- arXiv:2608.11219v1 Announce Type: new
- Abstract: Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others
- We present SAPO, a segment-level APO method that decomposes prompts into role, context, tasks, and output format, then applies targeted improvements based on to…
- The optimization loop uses one LLM with static meta-prompts and structured outputs for segmentation, weakness analysis, and candidate generation
LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11220v1 Announcement Type: New.
- Abstract: Currently, the creation of Process Flow Diagrams (PFDs) and their subsequent transformation into Piping and Instrumentation Diagrams (P&IDs) is predominantly performed manually.
- Applying artificial intelligence to the task could not only lead to process automation and time savings but also financial gains by exploring numerous topological options for the diagrams and reducing manual labor.
- This research presents P&ID Pilot - a practical end-to-end AI pipeline capable of handling the two stages of flowsheet development.
- EN Highlights:
- arXiv:2608.11220v1 Announce Type: new
- Abstract: Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) is predomina…
- Applying artificial intelligence in the task could potentially lead not only to process automation and time savings, but also to financial gains by exploring nu…
- This research presents P&ID Pilot - a practical end-to-end AI pipeline capable of handling flowsheet developing for both stages
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.11221v1 Announcement Type: new.
- Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artifacts suited to their specific areas of expertise.
- The behavior of these systems arises from the interaction between these artifacts and their operational environment.
- Simulation and co-simulation have become essential methods for analyzing CPS behavior. Through simulation activities, developers can explore system responses under changing conditions, including interactions with the environment.
- EN Highlights:
- arXiv:2608.11221v1 Announce Type: new
- Abstract: Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of expertise
- The behaviour of these systems emerges from the interaction between those artefacts and their operational environment
- Simulation and co-simulation have become essential approaches for analysing CPS behaviour and, through simulation campaigns, developers can explore system respo…
ArXiv cs.LG (B_intro+search) Link to heading
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12419v1 Announcement Type: new.
- Abstract: Large language models (LLMs) have achieved significant breakthroughs in various applications.
- However, their architectures remain inefficient in pre-training due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of local information within sequences; and (ii) Mixture of Experts (MoE) implicitly couples knowledge storage with computational paths, hindering flexible access to global knowledge outside the sequence.
- To overcome these limitations, we propose LoKiFormer, a novel LLM architecture that enhances the standard decoder with two dedicated modules: 1) Local Fusion Attention (LFA), which integrates convolution with attention to explicitly capture local patterns and allow attention to operate on more information-rich representations; and 2) Knowledge Memory Module (KMM), which introduces a parameterized key-value memory to explicitly store global knowledge in addressable slots, decoupling storage from computation and enabling direct knowledge retrieval.
- EN Highlights:
- arXiv:2608.12419v1 Announce Type: new
- Abstract: Large language models (LLMs) have achieved remarkable breakthroughs across various applications
- However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, le…
- To overcome these limitations, we propose LoKiFormer, a novel LLM architecture that augments the standard decoder with two dedicated modules: 1) Local Fusion At…
- Posted: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12422v1 Announcement Type: new.
- Abstract: Two free satellite signals carry real information about the risk of glacial lake outbursts in the Nepal Himalayas: radar interferometry detects the slow subsidence of moraine dams, and satellite weather data indicates the weeks when the lakes are under pressure.
- A companion feasibility study found that deformation indicates which lake is destabilizing and weather indicates when it is at risk, but did not propose a predictive model.
- To address this gap, we propose and evaluate models to predict which site is susceptible and when a trigger arrives.
- EN Key Points:
- arXiv:2608.12422v1 Announce Type: new
- Abstract: Two free satellite signals carry real information about glacial-lake outburst risk in the Nepal Himalaya: radar interferometry sees a moraine dam slow…
- A companion feasibility study found that deformation indicates which lake is destabilizing and weather indicates when it is at risk, but proposed no predictive…
- To address this gap, we propose and evaluate models that predict which site is susceptible and when a trigger arrives
MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
- Posted: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12435v1 Announcement Type: new.
- Abstract: The powerful long-context retrieval capability of Transformers is largely due to their token-level memory that grows with the context length.
- However, this flexibility results in quadratic computational complexity during training and causes the key-value cache to grow linearly during autoregressive inference.
- Recurrent alternatives offer efficient decoding by compressing the entire history into a fixed-size state, but they often underperform on recall-intensive tasks because earlier associations are typically overwritten by subsequent updates, retaining only the most recent context information.
- EN Key Points:
- arXiv:2608.12435v1 Announce Type: new
- Abstract: Transformers owe much of their strong long-context retrieval capability to a token-level memory that grows with context length
- This flexibility, however, incurs a quadratic computation complexity during training and a key–value cache that grows linearly during autoregressive inference
- Recurrent alternatives offer efficient decoding by compressing the entire history into a fixed-size state, but often underperform on recall-intensive tasks sinc…
- Posted: 2026-08-14 12:00 Beijing Time
Abstract: - arXiv:2608.12436v1 Announce Type: new.
- Abstract: Target tracking based on multi-AUV ad-hoc networks requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topologies, and uncertain ocean disturbances.
- While multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensional joint state-action modeling and noise-sensitive policy generation, leading to training instability and degraded tracking performance.
- To address these issues, we propose VGG-MADiffRL (a Value-Gradient-Guided Multi-Agent Diffusion RL algorithm) and MDCA (a diffusion-based hierarchical control architecture).
- EN Highlights:
- arXiv:2608.12436v1 Announce Type: new
- Abstract: Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets und…
- Although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensiona…
- To address these issues, we propose VGG-MADiffRL, a value-gradient-guided multi-agent diffusion RL algorithm, and MDCA, a diffusion
Unifying Generative Models with Path Integrals
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12438v1 Announce Type: new.
- Abstract: We formulate generative modeling as a path integral, in which flow-based, diffusion-based, variational, and adversarial models arise as different evaluation principles of a single master action.
- Its Martin-Siggia-Rose-Janssen-de~Dominicis (MSRJD) form separates free from interacting probability flows and opens them to diagrammatic perturbation theory.
- The expansion yields a one-loop correction to deterministic samplers at no stochastic-sampling cost, which we validate on solvable and nonlinear drifts, where it reduces the 53% tree-level error to 1.6%.
- EN Highlights:
- arXiv:2608.12438v1 Announce Type: new
- Abstract: We formulate generative modeling as a path integral in which flow-based, diffusion-based, variational, and adversarial models arise as different evalu…
- Its Martin-Siggia-Rose-Janssen-de~Dominicis (MSRJD) form separates free from interacting probability flows and opens them to diagrammatic perturbation theory
- The expansion yields a one-loop correction to deterministic samplers at no stochastic-sampling cost, which we validate on solvable and nonlinear drifts, where i…
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12441v1 Announce Type: new.
Abstract: Deep learning detectors for anomalies in dynamic graphs have reached high accuracy, but they remain opaque: when an edge is flagged, an analyst receives a score but no reason.
This opacity is untenable in the cooperative, regulated information systems where such detectors are deployed, where automated decisions must be auditable and trustworthy.
We address this gap with AddGraph, a foundational GCN+GRU framework for edge-level anomaly detection in dynamic graphs, which to our knowledge has never been equipped with any form of explainability.
EN Highlights:
- arXiv:2608.12441v1 Announce Type: new
- Abstract: Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the analyst rec…
- This opacity is untenable in the cooperative, regulated information systems where such detectors are deployed, where automated decisions must be auditable and t…
- We address this gap for AddGraph, the foundational GCN+GRU framework for edge-level anomaly detection in dynamic graphs, which to our knowledge has never been e…
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12446v1 Announcement Type: New.
- Abstract: Sleep stage classification is important for the diagnosis and management of sleep disorders, but most automatic staging studies evaluate models against a single reference hypnogram, despite known inter-scorer variability.
- This study investigates whether multi-scorer datasets can be used to construct more reliable reference labels from the collective behavior of multiple experts.
- We use the publicly available DOD-H and DOD-O datasets.
- EN Highlights:
- arXiv:2608.12446v1 Announce Type: new
- Abstract: Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models agains…
- This study investigates whether multi-scored datasets can be used to construct more reliable reference labels from the collective behavior of multiple experts
- We use the publicly available DOD-H and DOD-O datasets
Geometric and Behavioral Stratification in Transformer Residual Streams
- Publication Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12447v1 Announcement Type: New.
- Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream.
- – But what kind of directions do such bases choose?
- We study the prediction direction, the unembedding direction of the token the model is currently predicting, and find it acts as a content-defined privileged anchor.
- EN Highlights:
- arXiv:2608.12447v1 Announce Type: new
Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream
But what kind of direction does such a basis select
We investigate the prediction direction, the unembedding direction of the token a model currently predicts, and find that it functions as a content-defined priv…
Exemplar-based objective classification of gust-induced loads across multiple flight conditions
- Release Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12448v1 Announcement Type: new.
- Abstract: Is it possible to find an objective classification criterion to organize the complexity of gust-induced loads across multiple flight conditions?
- And is there a label that is as interpretable as one based on coarse parameters (e.g., flight attitude)?
- Our method encodes a large number of experimental observations through a machine-learned representation and applies a summarization process to select a minimal subset of highly important examples.
- EN Key Points:
- arXiv:2608.12448v1 Announce Type: new
- Abstract: Is it possible to find an objective classification criterion that organizes the complexity of gust-induced loads across many flight conditions
- And one that remains as interpretable as a labelling based on coarse parameters, such as the flight attitude
- Our approach encodes a large number of experimental observations through a machine-learned representation and applies a summarization procedure to select a mini…
- Release Time: 2026-08-14 12:00 Beijing Time
- Abstract: - arXiv:2608.12477v1 Announcement Type: new.
- Abstract: The development of clinical prediction models often proceeds as if the outcome of interest were cleanly observed for every patient.
- This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable.
- As a case study for this problem, we consider using a cohort of 2,497 patients for neurological prognostication after cardiac arrest, which includes 1,429 patients whose outcomes are indeterminate due to treatment decisions.
- EN Key Points:
- arXiv:2608.12477v1 Announce Type: new
- Abstract: Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient
- This assumption fails when treatment decisions make the clinically relevant outcome permanently unobservable
As a case study of this problem, we consider post-cardiac-arrest neurological prognostication using a cohort of 2,497 patients, including 1,429 patients whose o…