🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-11
- 类型
- ai-daily
- 字数
- 8598
- 阅读时长
- 41 min
2026-08-11 AI Daily | Models Go Local, Content Becomes Traceable Link to heading
Today’s signal is clear: the AI competition continues to shift from cloud capabilities toward local deployment, cost efficiency, and controllable governance. Meta is launching an open-source model for local hardware, strengthening the on-device ecosystem, while Anthropic is adding invisible watermarks to Claude’s text output to promote content traceability and compliance. At the same time, enterprise AI and Agent infrastructure are accelerating their entry into production scenarios.
📖 Deep Dive: This Issue’s Watch List Link to heading
There are three key threads to follow today: First, the case of OpenAI’s finance team and Model ML, which continues to apply “AI-native functions” to forecasting, reconciliation, and the last mile of PPT/Excel work—a must-read for operations, finance, and entrepreneurs. Second, OpenAI’s statements on Texas infrastructure and trusted licensing for frontier network models, indicating the industry’s focus is shifting from simply building large models to compute deployment, regulatory collaboration, and security governance. Third, the concentration of research on arXiv regarding MoE adaptation, cross-lingual understanding, personality evolution, and interpretability sends a clear signal: the next phase of competition is not just about more powerful models, but about capabilities that are more controllable, diagnosable, and implementable.
🌐 AI Hot Topics on X Link to heading
Topic 1: Meta Releases Muse Glimmer Open AI Model for Local Hardware Link to heading
- Category: AI · News
- Overview: Trending Time: 13 hours ago, Related Posts: 29,000
- What it is: Meta has released Muse Glimmer, an open-weights model designed to run on local hardware, and announced that the weights for Muse Spark 1.2 will be released soon.
- Why it matters: This indicates the large model competition is shifting from the cloud to local AI that can be deployed on consumer-grade devices. This shift impacts cost, privacy, latency, and ecosystem control, and will also influence the open-source model landscape and the development of agent forms.
- Discussion summary: Discussions on X are focused on whether it is truly suitable for single-GPU local execution, its competitive impact on Chinese and US open-source models like Qwen, and the strategic significance of Meta’s move to win over developers and promote a local agent ecosystem by opening up its weights. Some also question whether its promotional hype outweighs actual breakthroughs.
Topic 2: DeepSeek V4 Flash Shines in Agent Coding Tests with Pi Harness Link to heading
- Category: AI · Other
- Overview: Trending Time: , Related Posts: 176
- What it is: DeepSeek V4 Flash demonstrated outstanding performance in the Pi Harness agent programming tests, drawing attention from the AI community on X.
- Why it matters: This result suggests that lightweight or cost-effective models may be closing the capability gap with leading frontier models on code agent tasks, which has implications for developer tools, automated programming, and model cost competition.
- Discussion summary: The discussion is mainly centered on the reliability of the test benchmark, whether DeepSeek V4 Flash’s actual coding capabilities and cost advantages can be replicated in real-world projects, and its competitiveness compared to models from OpenAI, Anthropic, Google, and others.
Topic 3: Developers Share Mixed Views on AI Coding Tools Link to heading
- Category: AI · Other
- Overview: Trending Time: , Related Posts: 21
- What it is: A discussion on X surrounding Snowflake’s earnings report and AI product progress suggests that its performance indicates enterprise AI is moving from experimentation to production-level use. However, developers and investors have mixed views on the impact of AI coding tools, data platforms, and model costs.
- Why it matters: This shows that AI is driving actual spending growth in cloud data infrastructure, enterprise inference demand, and automated code migration. It also highlights that data governance, model routing, cost control, and platform control will become key battlegrounds for enterprise AI adoption.
- Discussion summary: The focus of the discussion is on whether Snowflake has evolved from a data warehouse to an enterprise AI “control plane”; whether AI programming tools will significantly reduce the cost of migrating legacy systems; who stands to benefit more among AWS, Arm, model vendors, and data platforms; and how enterprises can balance effectiveness, cost, and governance when choosing between frontier models and smaller ones.
Topic 4: AI Ends Zero-Marginal-Cost Era for SaaS Companies Link to heading
- Category: AI · News
- Overview: Trending Time: , Related Posts: 53
- What it is: The discussion on X about “AI ending the zero-marginal-cost era for SaaS” is gaining traction, with the view that AI products are more like selling services that replace human labor rather than traditional software licenses.
- Why it matters: This implies that the cost structure, pricing logic, and business models for AI applications may differ from traditional SaaS. Inference compute, customized delivery, and accountability for results will increase marginal costs, changing how startups are valued and how they compete.
- Discussion summary: Key points of discussion include whether AI products should be priced per seat, per usage, or based on results; whether enterprises will shift from buying general-purpose SaaS to cheaper, custom AI tools; and how open-source models, compute costs, and infrastructure players like Nvidia will affect the profit margins of AI software companies.
Topic 5: Debate Heats Over Vibe Coding and Programmer Skills Link to heading
- Category: AI · News
- Overview: Trending for: 2 days ago, Related posts: 6100
- What’s happening: A discussion is heating up on X about “Vibe Coding” (which primarily relies on AI to generate code from natural language intent), focusing on whether it weakens programmers’ fundamental skills.
- Why it matters: This concerns the positioning of AI programming tools in software development: are they auxiliary tools to boost productivity, or a paradigm shift that could alter developer training paths, code quality control, and the division of engineering responsibilities?
- Discussion summary: Supporters believe Vibe Coding can lower the barrier to entry for development, accelerating prototyping and daily tasks. Critics worry that over-reliance on AI will cause programmers to neglect core competencies like algorithms, architecture, debugging, and security, leading to unmaintainable code. The debate centers on balancing “efficiency gains” against “skill degradation and quality risks.”
Topic 6: Anthropic Adds Invisible Watermarks to Claude AI Text Worldwide Link to heading
- Category: AI · News
- Overview: Trending for: 3 hours ago, Related posts: 4600
- What’s happening: Anthropic is reportedly adding “invisible watermarks” to text generated by Claude worldwide to make it easier to identify AI-generated content.
- Why it matters: This is significant for the AI field as it relates to the traceability of generated content, platform governance, copyright, and abuse prevention. It could also influence the industry’s push for “AI identification” and content provenance standards.
- Discussion summary: Discussions on X primarily revolve around whether the watermarks can be truly effective, if they will affect text quality or editability, and the balance between such practices and transparency, privacy, and regulatory compliance. Some also question whether other models will follow suit.
Summary of AI Public Opinion on X Today Link to heading
Today’s main narrative focuses on AI’s shift from “showcasing and experimentation” to local deployment, enterprise production, and real commercialization. Meta’s open-sourcing of local models, the strong performance of DeepSeek’s lightweight model in programming tests, and Snowflake’s enterprise AI progress all point to cost, latency, privacy, and controllability becoming core competitive factors. The broad consensus is that AI capabilities are trickling down to cheaper, more flexible models and tools. Enterprises and developers will increasingly prioritize model routing, data governance, inference costs, and local agent ecosystems, rather than just chasing the largest models. The main points of divergence lie in the actual value of these advancements: whether benchmarks can represent real-world projects, if open-source weights are truly groundbreaking, whether AI programming is a productivity revolution or a path to skill degradation, and how AI software should be priced—by seat, usage, or outcome. Potential risks include a disconnect between model hype and practical capabilities, the erosion of SaaS profits by enterprise AI cost structures, code quality and security hazards from over-reliance on Vibe Coding, and new controversies arising from governance measures like invisible watermarks regarding transparency, privacy, and effectiveness.
💡 Influencer Insights Link to heading
AI Industry Influencer Daily: Tech Trends and Deep Dives Link to heading
1. Today’s Core Focus: The Explosion of Agent Infrastructure and Anxiety Over “De-AI-ing” Content Link to heading
Today’s influencers are focused on three key areas: the underlying standardization of Agent workflows, fine-grained control over AI content quality, and the diversified restructuring of computing power supply.
Agent Browsers and Runtime Environments: Infrastructure Becomes the New Battlefield Link to heading
Cloudflare’s launch of Kitesurf has garnered widespread attention. It’s seen as a browser engine specifically for Agents, designed to replace the expensive and cumbersome Chromium, providing a lighter and more scalable environment for automated tasks.
- Origin and Value: @Pluvio9yte retweeted that Kitesurf follows the Chrome DevTools Protocol, runs on Workers, and can replace the Agent’s expensive “eyes” (Chrome) with cheaper infrastructure, with the beta version being free. This is a huge boon for scenarios like bulk monitoring and scraping public pages.
- Agent Plugin Standardization: @Pluvio9yte also mentioned the Agent Plugins 1.0.0 specification released by a Google DeepMind engineer. It is a vendor-neutral packaging solution aimed at solving the reusability problem of Agent Skills and MCPs across different clients (like Claude Code, Cursor, etc.). @dotey also retweeted a discussion about the value of
harness, indicating a strong industry demand for a unified “middleware” layer.
“De-AI-ing” Becomes a Core Pain Point in Content Creation Link to heading
How to make AI-generated content shed its cold, mechanical feel has become a common topic, from code to text.
- De-AI-ifying Creative Writing: @Pluvio9yte proposed a proven and effective combined solution: using a “de-AI-ifying” skill to set constraints + having the AI learn from past writing styles. More critically, @Pluvio9yte believes the best “human touch” comes from conversational input, which involves first dictating using voice input tools and then having the large model polish it. This fundamentally changes the text’s logical structure and thought patterns.
- A Shift in Programming Mindset: @zhixianio praised the /wait-what skill, whose core capability is to “speak human” (using ASD-STE100 to simplify model output). This resonates with the need in programming scenarios to make AI output more precise and aligned with human comprehension.
- Philosophical Discussions: @lijigang raised a sharp point: “When you heavily use a model for conversation, its linguistic style will influence you. You might start speaking with a ‘Claude accent’… Ultimately, our brain’s neural networks are heavily influenced by context.” This suggests that “de-AI-ifying” is not just a content issue but also a risk of “linguistic assimilation” that can occur when humans interact with AI.
2. Noteworthy Unique Perspectives & Industry Outlook Link to heading
Anthropic’s Mathematical Breakthrough and the Cost of Transparency Link to heading
- @dotey reported in detail that Claude has made significant progress on a problem related to the Riemann Hypothesis: proving that at least 67.2% of non-trivial zeros lie on the critical line, far surpassing the academic community’s previous record of 41.6%. Notably, the research process was completed in 1.5 days by non-mathematicians coordinating about 60 sub-agents via Claude Code. @dotey emphasized that this demonstrates a leap in AI’s ability for original mathematical research, not just problem-solving.
- Meanwhile, @dotey also pointed out that Anthropic has begun embedding invisible watermarks and C2PA metadata in Claude’s output to comply with the EU AI Act. This means all text polished or generated by Claude will carry detectable markers, which could directly affect users who rely on Claude for content production, sparking a new round of discussions on content originality and traceability.
“Componentization” of Edge Models and Compute Supply Link to heading
- Pushing the Limits on the Edge: @Pluvio9yte discovered a project called Swiftlet that can run an 80B MoE model on a Mac with 4.3GB of RAM by streaming weights, and can even run a 35B model on an iPhone. @zhixianio also highly praised the full-duplex multimodal performance of the locally run MiniCPM-o 4.5 and believes this aligns with the “Model-Pak” (model cartridge) trend envisioned by @geekbb, signaling that large on-device models are rapidly approaching the threshold of usability.
- A Diversified Landscape for Compute Supply: @Pluvio9yte shared news that Anthropic signed a ten-billion-dollar compute deal with Volta, a cloud startup only a few months old. The essence of this purchase is the ability to “assemble power and GPUs on time,” leveraging sites from the Bitcoin miner Bitdeer. This signals a shift in compute supply from a monopoly by major cloud providers to a more flexible “assembly plant” model.
- The Controversy over the FDE Role: @dotey shared the view from Cursor’s head of talent that the Forward Deployed Engineer (FDE) is “the most sought-after position in tech,” but later shared another real-world interpretation of the FDE’s responsibilities, likening it to “hiring a hitman for the price of a kitchen knife” or simply “a placebo for AI-anxious bosses,” revealing a huge gap between the ideal and reality of the position.
3. Recommended Tools & Resources at a Glance Link to heading
Agents & Automation
- Kitesurf (recommended by @Pluvio9yte): A free, agent-specific browser engine from Cloudflare, serving as a lightweight alternative to Chromium.
- Herdr (recommended by @vista8): A persistent terminal tool that surpasses tmux, backed by YC, and suitable for geeks.
- Airtap (recommended by @AI_Jasonyu): Remotely control a real US-based phone in the cloud via iMessage, for stable account maintenance, registration, and daily use.
Development & Model Tools
- BaoCut (recommended and developed by @dotey): A video transcription, translation, and editing tool with separate GUI and CLI, supporting agent calls.
- OpenCodex / OpenClaw (recommended by @vista8): A terminal tool
ocx, that allows you to configure Codex to access various third-party models for easy management. - Reasonix (recommended by @Pluvio9yte): A programming framework officially recommended by DeepSeek that uses a prefix caching mechanism to optimize token costs in long sessions.
Gemma 4 12B Coder (tested by @zhixianio): A new option for local code generation, but practical tests show that for complex, long-term tasks, its 12B parameter size remains a hard ceiling, making it less stable than Qwen 35B MoE.
Design & Frontend Resources
- Baoyu-Design Skill (recommended by @dotey): A localized solution for maintaining consistency between UI prototypes and code, advocating for a “prototype first, function second” development process.
- UI Animation Component Libraries (compiled by @AI_Jasonyu): A collection including Motion Sites, React Bits, Uiverse, Anime.js, and more, specifically to address the pain point of AI-generated websites lacking design flair.
Content & Community
- RedSkill Community (discovered by @ruanyf): Xiaohongshu (Little Red Book) now supports uploading and distributing Skill files, attempting to become the “GitHub for Skills,” allowing programmers to reach a massive user base.
- Bento PPT (recommended by @vista8): An open-source HTML PPT generator, serving as an alternative to Nextslide, which was acquired by OpenAI.
📚 Appendix: Today’s Watch List Source Updates Link to heading
Timeframe: Last 3 days; Covers 22 sources; 41 updates in total.
Acquired.fm (A_full) Link to heading
- Disney: The Renaissance and the Empire
- Publication Time: 2026-08-10 12:41 Beijing Time
- Summary: - In 1984, The Walt Disney Company was worth more dead than alive.
- Disney Animation—the heart of Walt’s famous flywheel—had stagnated for years, bleeding talent while corporate raiders circled, salivating at the prospect of selling the film library to MGM and the parks to hotel operators.
- But what followed was the greatest turnaround in media history under the leadership of Michael Eisner and Frank Wells.
- Bringing the Disney Vault home through VHS and DVD.
- And the greatest media acquisition of all time—ESPN.
- EN Key Points:
- In 1984, the Walt Disney Company was worth more dead than alive
- Disney Animation — the heart of Walt’s famous flywheel — had stagnated for years, bleeding away talent while corporate raiders circled, salivating over offers t…
- But what followed instead was the greatest turnaround in media history under Michael Eisner and Frank Wells
- Beauty and the Beast
Y Combinator Podcast (B_intro+search) Link to heading
- Max Hodak: How Startups Build Speed
- Publication Time: 2026-08-11 05:41 Beijing Time
- Summary: - You’ve likely heard of OpenClaw (formerly Clawdbot/Moltbot).
- The sensational open-source AI assistant that runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually perform tasks like managing emails, calendars, files, and workflows.
- Now, meet the person behind it.
- YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to discuss the “aha” moment behind the viral personal AI agent, why local-first agents could replace many of today’s apps, and how personal agents are set to reshape the future of software.
- EN Key Points:
- Science is building a retinal implant that restores vision to people who have gone blind
- One patient has already used it to read a 300-page novel.Building a company like that requires a lot more than getting the technology right
- At Startup School 2026, Science CEO Max Hodak explains how the company buys things and hires people, and why those systems determine how fast it can move.He als…
Stratechery by Ben Thompson (A_full) Link to heading
- Apple Earnings, More on Amazon’s Earnings
- Publication Time: 2026-08-10 18:00 Beijing Time
- Summary: - Apple’s earnings (and stock) are not limited by memory, but by chip shortages; then, more on Amazon’s earnings and Andy Jassy’s market analysis.
- $15/month* or *$150/year.
- Substantive analysis of the day’s news via three weekly emails or a podcast.
- Strategy Interviews.
- Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
- EN Highlights:
- Apple’s earnings (and stock) are limited not by memory but rather chip shortages; then, more on Amazon’s earnings and Andy Jassy’s market analysis.
OpenAI Blog (A_full) Link to heading
What building an AI-native finance function taught me
- Publication Time: 2026-08-11 01:00 Beijing Time
- Summary: - OpenAI CFO Sarah Friar shares five lessons learned from building an AI-native finance function, from automated forecasting to stronger controls and AI ROI.
- This article on the OpenAI blog explains how the lessons from building an AI-native finance function can shape the broader AI and infrastructure landscape.
- It also reveals the practical implications for founders, operators, and investors.
- EN Highlights:
- OpenAI CFO Sarah Friar shares five lessons for building an AI-native finance function, from automated forecasting to stronger controls and AI ROI.
OpenAI’s letter to Governor Abbott on responsible AI infrastructure in Texas
- Publication Time: 2026-08-10 22:00 Beijing Time
- Summary: - OpenAI sent a letter to Texas Governor Greg Abbott outlining our commitment to the responsible development of AI infrastructure in Texas.
- We look forward to working with state and local leaders, utilities, and communities to ensure that AI infrastructure brings meaningful benefits to Texans.
- [Advancing responsible AI in Europe.
- [Building AI infrastructure with the Effingham County community. –[Advancing the next era of national science.
- EN Highlights:
- OpenAI sent Governor Greg Abbott a letter outlining its commitment to responsible AI infrastructure in Texas
- The letter supports reliable, transparent growth that benefits Texans.
Model ML completes finance work more efficiently with GPT-5.6 Sol
- Publication Time: 2026-08-10 20:00 Beijing Time
- Summary: - Before a financial analysis can stand up to scrutiny from clients or senior decision-makers, the team must complete the demanding last mile: reconciling evidence, building and formatting documents, checking every number, and linking every statement to its source.
- The completed PowerPoint presentation or Excel workbook must be editable and available for review.
After two successful exits, Model ML co-founders and brothers Arnie and Chaz Englander started investing through a private family office and, building software to help themselves like the builders they are, saw how much work needed to be done.
Model ML’s agents, which grew out of that software, help finance professionals with workflows from initial request to research, analysis, and a finished deck or workbook.
At the center, a core agent plans the work, selects the right tools, coordinates evidence, and runs calculations, routing each step to the model best suited for it, usually GPT-5.6 Sol.
- EN Highlights:
- Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
- EN Highlights:
Expanding Daybreak as the Cyber Defense Window Narrows
- Posted: 2026-08-10 18:00 Beijing Time
- Abstract: - Learn about GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model, available through Daybreak Red for authorized vulnerability research, exploit validation, and security testing.
- Learn about GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model, available through Daybreak Red for authorized vulnerability research, exploit validation, and security….
- As the cyber defense window narrows, Daybreak expands.
- EN Highlights:
- Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit validation, and security…
Putting frontier cyber models in more trusted hands
- Posted: 2026-08-10 18:00 Beijing Time
- Abstract: - Approved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, regulated cybersecurity services to customers.
- This post from the OpenAI blog explains how putting frontier cyber models in more trusted hands can shape the broader AI and infrastructure landscape.
- It also has practical implications for founders, operators, and investors after putting frontier cyber models in more trusted hands.
- EN Highlights:
- Approved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, governed cybersecurity services to customers.
Premium seats are coming to ChatGPT Business
- Posted: 2026-08-10 08:00 Beijing Time
- Abstract: - Premium seats are coming to ChatGPT Business.
- Sign up by August 20 to get $100 in workspace credits and unlock higher usage for your team’s most demanding work.
- Premium seats for ChatGPT Business will be available for sign-up by August 20 to receive $100 in workspace credits and unlock higher usage for your team’s most demanding work.
- EN Highlights:
- Premium seats are coming to ChatGPT Business
- Sign up by August 20 to get $100 in workspace credits and unlock higher usage for your team’s most demanding work.
How Zapier transformed core marketing processes with ChatGPT Work
- Posted: 2026-08-10 08:00 Beijing Time
- Abstract: - Zapier’s enterprise marketing team uses ChatGPT Work to reduce churn in its lead funnel, build marketing campaign assets, and automate reporting.
This article from the OpenAI blog explains how Zapier is transforming its core marketing processes with ChatGPT Work, shaping the broader AI and infrastructure landscape.
- It also provides practical implications for founders, operators, and investors by detailing how Zapier leverages ChatGPT Work to transform its core marketing processes.
- EN Key Points:
- The enterprise marketing team at Zapier uses ChatGPT Work to reduce the number of drop-offs in its lead funnel, build campaign assets, and automate reporting.
Virgin Atlantic sharpens customer journeys with ChatGPT Work
- Published: 2026-08-10 08:00 Beijing Time
- Summary: - Virgin Atlantic is using ChatGPT Work to accelerate research, product planning, and decision-making, helping teams connect signals across the entire customer journey.
- This article from the OpenAI blog explains how Virgin Atlantic is sharpening customer journeys with ChatGPT Work, shaping the broader AI and infrastructure landscape.
- It also provides practical implications for founders, operators, and investors following Virgin Atlantic’s use of ChatGPT Work to sharpen its customer journeys.
- EN Key Points:
- Virgin Atlantic is accelerating research, product planning, and decision-making with ChatGPT Work, helping teams connect signals across the customer journey.
ArXiv cs.AI (B_intro+search) Link to heading
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06394v1 Announce Type: new.
- Abstract: Multi-label node classification is an important yet challenging task in graph learning, where nodes exhibit multiple semantics simultaneously.
- Existing multi-label node classification methods can effectively model multiple labels, but they only consider in-domain scenarios where the model needs to be trained and tested within the same graph domain, resulting in limited cross-domain generalization capabilities.
- Recently, Graph Foundation Models (GFMs) have emerged as a promising paradigm for learning transferable graph representations across diverse graph domains and downstream tasks.
- EN Key Points:
- arXiv:2608.06394v1 Announce Type: new
- Abstract: Multi-label node classification is an important yet challenging task in graph learning, where nodes exhibit multiple semantics simultaneously
- Existing methods for multi-label node classification can effectively model multiple labels, while only considering in-domain scenarios where the model needs to…
- Recently, Graph Foundation Models (GFMs) have emerged as a promising paradigm for learning transferable graph representations across diverse graph domains and d…
EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06398v1 Announce Type: new.
Abstract: Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches.
However, existing byte-patch architectures still apply the same dense feed-forward computation to every patch.
This uniform computation cannot adapt model capacity to variations in patch semantics and granularity.
- EN Highlights:
- arXiv:2608.06398v1 Announce Type: new
- Abstract: Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized pat…
- However, existing byte-patch architectures still apply the same dense feed-forward computation to every patch
- This uniform computation cannot adapt model capacity to variations in patch semantics and granularity
- EN Highlights:
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06400v1 Announce Type: new.
- Abstract: Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging.
- Recent sparse Mixture-of-Experts (MoE) reward models seek to improve interpretability by routing prompts to specialized experts and characterizing experts through examples with high routing weights.
- However, routing weights only reveal which prompts an expert $\textit{receives}$, not how it $\textit{judges}$ responses, providing only a partial account of expert behavior.
- EN Highlights:
- arXiv:2608.06400v1 Announce Type: new
- Abstract: Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging
- Recent sparse Mixture-of-Experts (MoE) reward models seek to improve interpretability by routing prompts to specialized experts and characterizing experts throu…
- However, routing weights only reveal which prompts an expert $\textit{receives}$, not how it $\textit{judges}$ responses, providing only a partial account of ex…
Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06402v1 Announce Type: new.
- Abstract: Community detection is a fundamental task in graph analysis, aiming to identify cohesive groups of entities with similar behaviors or interests.
- Classical objective-driven methods struggle with complex graph structures, while deep learning approaches improve performance at the cost of interpretability and rely on labeled data and training.
- Large Language Models (LLMs), with their powerful reasoning capabilities and world knowledge, are promising for interpretable, label-free community detection.
- EN Highlights:
- arXiv:2608.06402v1 Announce Type: new
Abstract: Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or interests
Classic objective-driven methods struggle with complex graph structures, while deep-learning approaches improve performance at the expense of interpretability a…
Large language models (LLMs), with strong reasoning capabilities and world knowledge, are promising for interpretable, label-free community detection
ADIAS: Automated Design of Interactive Agentic Systems
- Release Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06410v1 Announce Type: new.
- Abstract: Automated agent design improves agent tools through iterative modification, evaluation, and feedback summarization.
- Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit.
- This causes inefficient repair targeting, slow consolidation of partial progress, and propagation of ineffective interventions across rounds.
- EN 要点:
- arXiv:2608.06410v1 Announce Type: new
- Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization
- Existing methods are largely candidate-centric: cross-round experience is organized around candidate agents, which leaves the repair progress implicit
- This causes inefficient repair targeting, slow consolidation of partial progress, and propagation of ineffective interventions across rounds
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin
- Release Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06411v1 Announce Type: new.
- Abstract: Multimodal large language models (MLLMs) achieve powerful performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing a large number of visual tokens.
- Visual token pruning can reduce this cost but requires accurate token importance estimation.
- Recent research shows that text-to-vision attention from intermediate language model layers can effectively guide visual token pruning, often using attention from pre-defined intermediate layers to select visual tokens to retain.
- EN 要点:
- arXiv:2608.06411v1 Announce Type: new
- Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost…
- Visual token pruning can reduce this cost, but requires accurate token importance estimates
Recent studies have demonstrated that text-to-vision attention from middle language model layers can effectively guide visual token pruning, typically using att…
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06474v1 Announcement Type: new.
- Abstract: Large language models increasingly generate complete websites from natural language descriptions, and reinforcement learning has become a core method for closing their remaining functionality gap.
- This training regime is bottlenecked by reward design.
- Hand-authored browser scripts are executable but costly to write for open-ended requirements, while VLM and GUI agent graders are scalable but may issue verdicts before observing decisive states.
- EN Key Points:
- arXiv:2608.06474v1 Announce Type: new
- Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central appr…
- This training regime is bottlenecked by reward design
- Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts befor…
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06501v1 Announcement Type: new.
- Abstract: The creativity of MLLMs is important in design, communication, education, and human-AI collaboration, yet remains difficult to evaluate because explicit goals and reward signals are scarce compared to accuracy-oriented tasks.
- Cross-concept understanding is a core cognitive capacity underlying receptive creativity.
- It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations.
- EN Key Points:
- arXiv:2608.06501v1 Announce Type: new
- Abstract: Creative capabilities of MLLMs matter in design, communication, education, and human–AI collaboration, yet remain difficult to evaluate because expli…
- Cross-concept understanding is a core cognitive capacity underlying receptive creativity
- It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations
KNOWPLAN: Knowledge-Driven AI Agents for Smart Degree Pathway Planning
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06530v1 Announcement Type: new.
- Abstract: Planning a degree from official university sources requires sequentially solving two problems.
- An institution’s curriculum must first be reconstructed from catalogs, department pages, JSON endpoints, and PDFs that do not share a schema, before student-specific pathways can be optimized under prerequisite logic and overlapping requirement constraints.
Coupling the two lets each failure mode hide the other, because a planner that drives its own crawling never learns facts its current plan does not need.
- EN Highlights:
- arXiv:2608.06530v1 Announce Type: new
- Abstract: Planning a degree from official university sources requires solving two problems in order
- The institution’s curriculum must first be reconstructed from catalogs, departmental pages, JSON endpoints, and PDFs that share no schema, and only then can a s…
- Coupling the two lets each failure mode hide the other, because a planner that drives its own crawling never learns facts its current plan does not need
- EN Highlights:
TaskSense: Focusing on What Matters in World Models
- Release Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06544v1 Announce Type: new.
- Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input.
- However, task-relevant content often occupies only a small fraction of the observation, while background clutter and distractors consume valuable representation capacity.
- This mismatch between visual reconstruction and control objectives biases latent representations to model task-irrelevant visual content, diluting the learning signal for control-relevant features and severely degrading downstream performance under visual distractions.
- EN Highlights:
- arXiv:2608.06544v1 Announce Type: new
- Abstract: World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preser…
- However, task-relevant content often occupies only a small fraction of the observation, while background clutter and distractors consume valuable representation…
- This mismatch between visual reconstruction and control objectives biases latent representations to model task-irrelevant visual content, diluting learning sign…
ArXiv cs.CL (B_intro+search) Link to heading
TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation
- Release Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06396v1 Announce Type: new.
- Abstract: Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relevant experts during downstream adaptation.
- However, current methods have two limitations: task experts are often identified based on aggregate routing statistics reflecting usage rather than association with successful task completion, and the activation of task experts as a signal for supervised allocation remains underexplored.
- We introduce Task-Expert-Aware Supervision (TEXAS), which combines correctness-conditioned task expert discovery with token-level supervised allocation.
- EN Highlights:
- arXiv:2608.06396v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying task-relev…
Yet current approaches have two limitations: task experts are typically identified from aggregate routing statistics that reflect usage rather than association…
We introduce Task-Expert-Aware Supervision (TEXAS), which combines correctness-conditioned task expert discovery with token-level supervision allocation
Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06409v1 Announcement Type: New.
- Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation.
- We introduce a generation-aligned diagnostic ladder that compares the emitted answer, option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token.
- Successive differences separate endpoint, decision-rule, and readout-coverage gaps.
- EN Key Points:
- arXiv:2608.06409v1 Announce Type: new
- Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures a…
- We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout…
- Successive differences separate endpoint, decision-rule, and readout-coverage gaps
NTDH: Complex Reasoning for Comprehensive Affective Analysis
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06425v1 Announcement Type: New.
- Abstract: Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring coordination of conflicting cues rather than a direct mapping to labels.
- Existing methods learn this mapping directly and do not explicitly model the coordination.
- We redefine the task as a complex reasoning problem that yields an output interface across heterogeneous label spaces and a trajectory that can be optimized for a verifiable reward; to our knowledge, this is the first such treatment involving sentiment and emotion.
- EN Key Points:
- arXiv:2608.06425v1 Announce Type: new
- Abstract: Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label out…
Existing methods learn this mapping directly and do not model the reconciliation explicitly
We recast the task as a complex-reasoning problem, which yields one output interface across heterogeneous label spaces and a trajectory over which a verifiable…
Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06429v1 Announcement Type: New.
- Abstract: Interpretability methods for large language models (LLMs) describe internal states but do not directly test whether that state is sufficient to produce observed behavior.
- In earlier work, we lesioned LLMs to produce error profiles in picture naming, a core task for assessing aphasia, and found that specific lesions produced errors similar to those of individual stroke survivors.
- Here we pose the inverse problem: given an error profile, can the lesion parameters that produced it be recovered, and what does this inverse problem reveal about the transformer’s computation?
- EN Highlights:
- arXiv:2608.06429v1 Announce Type: new
- Abstract: Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient t…
- In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific lesions produced er…
- Here we ask the inverse question: given an error profile, can the lesion parameters that produced it be recovered, and what does this inverse problem reveal abo…
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06485v1 Announcement Type: New.
- Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used for emotional support, social simulation, and role-playing, motivating the development of lifelong agents that maintain consistency in long-term interactions.
- A key component of this consistency is personality evolution: when agents experience life events in different environments, they should undergo reasonable, psychology-based changes.
- Although previous research has shown that the personality of LLMs may change under environmental disturbances, how these changes vary with different traits, events, personas, and models remains little understood.
- EN Highlights:
- arXiv:2608.06485v1 Announce Type: new
- Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the develop…
A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in diffe…
Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models rem…
ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06495v1 Announcement Type: New.
- Construction accident narratives contain rich causal information, but the evidence is often implicit, long-span, and distributed.
- We introduce ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA construction accident reports.
- The dataset uses a hierarchical schema for accident types, causal factors, sub-causal factors, and supporting evidence spans.
- EN Highlights:
- arXiv:2608.06495v1 Announce Type: new
- Abstract: Construction accident narratives contain rich causal information, but the evidence is often implicit, long-span, and distributed
- We introduce ConstructCIE, a manually annotated dataset for Causal Information Extraction from OSHA construction accident reports
- The dataset uses a hierarchical schema for accident types, causal factors, sub-causal factors, and supporting evidence spans
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06506v1 Announcement Type: New.
- Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages.
- Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant.
- We define the Cross-Lingual Comprehension Gap (CLCG) as the drop in response quality when the same content and question are presented in a target language rather than English.
- EN Highlights:
- arXiv:2608.06506v1 Announce Type: new
- Abstract: Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in othe…
- Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant
We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language…
GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization
- Posted: 2026-08-10 12:00 Beijing Time
- Summary: - arXiv:2608.06526v1 Announcement type: new.
- Summary: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into a privacy risk.
- Adversarial anonymization defends against this by rewriting text using a powerful language model that also plays the role of an attacker. However, it requires a powerful model at inference time, which results in sending private text to a third party—an exposure that anonymization is supposed to prevent.
- Recent work distills this behavior into a small, on-device model using supervised fine-tuning and Direct Preference Optimization (DPO), but DPO only imitates the teacher’s offline choices and never directly optimizes for the privacy-utility objectives we care about.
- EN Key Points:
- arXiv:2608.06526v1 Announce Type: new
- Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing int…
- Adversarial anonymization defends against this by rewriting a text with a capable language model that also plays the attacker, but it needs a powerful model at…
- Recent work distills this behavior into a small on-device model using supervised fine-tuning and direct preference optimization (DPO), but DPO only imitates the…
Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models
- Posted: 2026-08-10 12:00 Beijing Time
- Summary: - arXiv:2608.06529v1 Announcement type: new.
- Summary: Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs).
- Existing formulations build this blend via linear interpolation (LERP) in the raw embedding space, implicitly treating that space as Euclidean.
- We analyze the embedding space of MDLMs and find that the mask and predicted token embeddings maintain a near-constant angle (≈ 73^\circ) throughout the training process, while the embedding norms remain largely flat across vocabulary frequency ranks.
- EN Key Points:
- arXiv:2608.06529v1 Announce Type: new
- Abstract: Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs)
- Existing formulations build this blend with linear interpolation (LERP) in the raw embedding space, which implicitly treats that space as Euclidean
- We analyze the embedding space of MDLMs and find that the mask and predicted-token embeddings maintain a near-constant angle of (≈ 73^\circ) throughout tr…
Confidence Estimation for Financial Vision-Language Models in Chart and Document Understanding
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06532v1 Announcement Type: new.
- Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can change a decision, and the most authoritative-sounding answers are sometimes generated by models without reading the chart.
- Therefore, the operational question is trust, not accuracy: which answers can be acted on, and which can be escalated to a reviewer.
- We evaluate seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three financial visual question answering benchmarks, one of which is bilingual; each probe was trained only on natural images and applied to finance without adaptation, so the results measure out-of-distribution transfer.
- EN Key Points:
- arXiv:2608.06532v1 Announce Type: new
- Abstract: LVLMs are increasingly used to read financial charts, tables, and documents, where a single misread figure can move a decision and the most authoritat…
- The operational question is therefore trust, not accuracy: which answers can be acted on, and which escalated to a reviewer
- We evaluate seven confidence estimators, three inference-only and four trained internal probes, across five open-weight LVLMs and four conditions from three fin…
ArXiv cs.LG (B_intro+search) Link to heading
Latent Fact-Checking: Detecting Misinformation through Activation Engineering
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06417v1 Announcement Type: new.
- Abstract: The proliferation of misinformation online has driven the demand for scalable detection systems.
- While most existing methods rely on surface-level linguistic features or external knowledge retrieval, we treat truthfulness as a geometric property of a language model’s representation space.
- We introduce a misinformation detection framework grounded in activation engineering, which leverages the latent geometry of transformer models.
- EN Key Points:
- arXiv:2608.06417v1 Announce Type: new
- Abstract: The proliferation of misinformation online has driven demand for scalable detection systems
- While most existing approaches rely on surface-level linguistic features or external knowledge retrieval, we examine truthfulness as a geometric property of a l…
- We introduce a misinformation detection framework grounded in activation engineering, which leverages the latent geometry of transformer models
Risk-Aware Decision Policies for Agents Under Noisy Perception
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06420v1 Announcement Type: new.
Abstract: Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty, where misclassification can be costly or fatal.
- We present an Artificial Life predator-prey model of foraging under noisy perception, and compare agent performance when using various policies that take into account noisy predictions.
- Through controlled experiments under both symmetric and asymmetric perceptual noise, we show that as noise increases, blindly trusting perceptual labels leads to catastrophic failure, while uncertainty-aware policies significantly improve survival rates and reduce fatal errors.
- EN 要点:
- arXiv:2608.06420v1 Announce Type: new
- Abstract: Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can be costly or…
- We present an Artificial Life predator-prey model of foraging under noisy perception, and compare agent performance when using various policies that take into a…
- Through controlled experiments under both symmetric and asymmetric perceptual noise, we show that blindly trusting perceptual labels leads to catastrophic failu…
Sharding Prevents LLM Oversight Failures and Adversarial Exploitation
- Publication Date: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06422v1 Announce Type: new.
- Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements.
- When a call must return multiple verdicts, some decisions become weakly grounded in the evidence, even when the call receives the same token or tool budget as a set of separate calls.
- In expert-graded research replications, legal work, and clinical trial evaluations, agreement with experts declines as the number of verdicts per call increases.
- EN 要点:
- arXiv:2608.06422v1 Announce Type: new
- Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements
- When one call must return many verdicts, some decisions become weakly grounded in the evidence, even when that call receives the same token or tool budget as a…
- Across expert-graded research replications, legal work, and clinical-trial assessments, agreement with experts falls as the number of verdicts per call grows
Adversarial Causal Intervention Falsification
- Publication Date: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06427v1 Announce Type: new.
- Abstract: Generative models can reproduce observed distributions while encoding incorrect causal structures.
- We study a sequential game where a structural causal generator proposes observational and interventional distributions, and an adversarial experimentalist chooses interventions designed to maximally falsify the generator.
- Thus, the discriminator is not merely a real vs. synthetic classifier: it is indexed by interventions and tests whether the generator reproduces the corresponding post-intervention laws.
- EN 要点:
- arXiv:2608.06427v1 Announce Type: new
Abstract: Generative models can reproduce an observational distribution while encoding an incorrect causal structure
We study a sequential game in which a structural causal generator proposes observational and interventional distributions, while an adversarial experimentalist…
The discriminator is therefore not merely a real-versus-synthetic classifier: it is indexed by an intervention and tests whether the generator reproduces the co…
Fixed and Adaptive Topological DeepONets: Functional Measurements on Hausdorff Locally Convex Spaces
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06428v1 Announce Type: new.
- Abstract: Deep Operator Networks (DeepONets; arXiv:1910.03193) typically encode input functions through point values on a fixed discretization.
- Building on Ismailov’s Topological DeepONet framework (arXiv:2603.11972), we replace point samples with continuous linear functionals extracted from the continuous dual of a Hausdorff locally convex space $({V},{p_\alpha}_{\alpha\in A})$, whose topology is generated by a family of point-separating seminorms rather than a single norm, and develop fixed and adaptive functional measurement systems.
- Measurements are combined with the coefficient-space Two-Step procedure of Lee and Shin (arXiv:2309.01020), while training-only decoders and regularization stabilize the adaptive coordinates.
- EN 要点:
- arXiv:2608.06428v1 Announce Type: new
- Abstract: Deep Operator Networks (DeepONets; arXiv:1910.03193) typically encode an input function through point values on a fixed discretization
- Building on the Topological DeepONet framework of Ismailov (arXiv:2603.11972), we replace point samples by continuous linear functionals drawn from the continuo…
- Measurements are combined with the coefficient-space Two-Step procedure of Lee and Shin (arXiv:2309.01020), while a training-only decoder and regularization sta…
MiGHT-EHR: A Multi-task Graph Transformer for Heterogeneous Temporal Electronic Health Records
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06430v1 Announce Type: new.
- Abstract: Electronic Health Record (EHR) learning has garnered significant attention due to its potential to improve clinical predictions.
- However, effective learning remains challenging because EHRs encode heterogeneous, chronologically ordered clinical interactions.
- Specifically, EHRs contain: (i) heterogeneous clinical entities, including patients, visits, diagnoses, prescriptions, and procedures, along with their heterogeneous interactions, (ii) longitudinal patient trajectories across hospital visits, and (iii) statistical dependencies shared among related clinical prediction tasks.
- EN 要点:
- arXiv:2608.06430v1 Announce Type: new
Abstract: Learning from Electronic Health Records (EHRs) has gained significant attention due to its potential to improve clinical prediction
However, effective learning remains challenging because EHRs encode heterogeneous, temporally ordered clinical interactions
In particular, EHRs contain: (i) heterogeneous clinical entities, including patients, visits, diagnoses, prescriptions, and procedures, together with their hete…
SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06441v1 Announcement Type: New.
- Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges.
- We introduce SNI-GNN, a SmartNIC-assisted full-graph training system that reduces communication while preserving accuracy by predicting remote embeddings in-network.
- SNI-GNN deploys a lightweight linear-trend predictor on SmartNICs to refine cached historical embeddings, coupled with an importance-based boundary-node sampling strategy and an asynchronous DPU-GPU data pipeline with intermediate result reuse.
- EN Key Points:
- arXiv:2608.06441v1 Announce Type: new
- Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges
- We present SNI-GNN, a SmartNIC-assisted full-graph training system that reduces communication while preserving accuracy by predicting remote embeddings in-netwo…
- SNI-GNN deploys a lightweight linear-trend predictor on SmartNICs to refine cached historical embeddings, coupled with an importance-based boundary-node samplin…
ED-CSP: Crystal Structure Prediction from Electron Diffraction
- Publication Time: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06448v1 Announcement Type: New.
- Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem.
- Existing ED-based learning methods primarily predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve candidates from a limited structure library.
- Here, we introduce ED-CSP, a machine learning framework that can predict crystal structures based on chemical composition, atom counts, and multiple detector-plane ED point sets.
- EN Key Points:
- arXiv:2608.06448v1 Announce Type: new
- Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem
Existing ED-based learning methods mainly predict crystallographic labels, reconstruct structures from indexed reflections, or retrieve candidates from finite s…
Here, we introduce ED-CSP, a machine learning framework that predicts crystal structures from chemical composition, atom count, and multiple detector-plane ED s…
Beyond Attention: Signed Integrated Gradients Attribution in a BiomeGPT-Style Microbiome Transformer
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06486v1 Announce Type: new.
- Abstract: In a feature-tokenized transformer (arXiv:2106.11959) such as BiomeGPT (doi:10.64898/2026.01.05.697599), each input token is built by fusing a fixed identity with a sample-specific measurement: fixed species and variable abundance, T = S + A.
- To interpret downstream classification in such models, prior work inspects the attention weights of the special [CLS] token (arXiv:2106.11959, arXiv:1810.04805, BiomeGPT) to rank sample tokens by importance.
- These weights have two critical limitations: they are nonnegative, so they cannot separate disease-supporting from health-supporting evidence (arXiv:2201.12114), and they operate after token fusion, obscuring how input sources S and A each influence the output.
- EN Highlights:
- arXiv:2608.06486v1 Announce Type: new
- Abstract: In a feature-tokenized transformer (arXiv:2106.11959) such as BiomeGPT (doi:10.64898/2026.01.05.697599), each input token is built by fusing a fixed i…
- To interpret downstream classification in such models, prior work inspects the attention weights of the special [CLS] token (arXiv:2106.11959, arXiv:1810.04805,…
- These weights have two critical limitations: they are nonnegative, so they cannot separate disease-supporting from health-supporting evidence (arXiv:2201.12114)…
- Published: 2026-08-10 12:00 Beijing Time
- Abstract: - arXiv:2608.06503v1 Announce Type: new.
- Abstract: Recurrent context compression controls context growth for long-horizon agents, yet its behavioral effects remain poorly understood.
- In this preliminary empirical study, we show that compression can weaken the influence of recent interactions, increasing blocking actions, repetitive exploration, and in-run instability.
- Inspired by these observations, we introduce TRACE, a validator-guided framework that evaluates individual compression events via paired closed-loop continuations from the same environment state, and uses summary preferences to optimize natural language compression prompts while keeping all models frozen.
- EN Highlights:
- arXiv:2608.06503v1 Announce Type: new
Abstract: Recurrent context compression controls context growth in long-horizon agents, but its behavioral effects remain poorly understood
In this preliminary empirical study, we show that compression can weaken the influence of recent interactions, increasing blocked actions, repeated exploration,…
Motivated by these observations, we introduce TRACE, a verifier-guided framework that evaluates individual compaction events through paired closed-loop continua…