🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-11
- 类型
- ai-daily
- 字数
- 8490
- 阅读时长
- 40 min
2026-06-11 AI Daily | When Top-Tier AI Programming Costs More Than a Human, On-Device Models Are Accelerating Their Takeover of Everyday Tasks Link to heading
After the release of Anthropic’s most powerful model, Claude Fable 5, real-world tests show that the cost of high-frequency use has already surpassed hiring a human programmer, as the engineering implementation of AI hits a cost wall. At the same time, on-device models (Gemma 4, Qwen, etc.) have achieved a usability breakthrough in local inference, offering rapid responses, though their Chinese capabilities remain weak. Industry methodology is also shifting from “vibe-driven coding” to a “contract-first” approach, emphasizing testable and auditable engineering standards. This cost inversion and the rise of on-device models are reshaping enterprise deployment strategies.
📖 In-depth Guide to This Issue’s Watch List Link to heading
In-depth Guide to This Issue’s Watch List
Today, there are several threads worth examining together in depth.
The first is about the intersection of enterprise-grade AI and cutting-edge scientific research. OpenAI released two pieces of information simultaneously: astrophysicists using Codex to simulate black hole gravity, and the model being offered for direct delivery via Oracle Cloud. This directly addresses a trend—leading-edge models are both pushing the boundaries of fundamental science (like testing general relativity) and accelerating their implementation through mature procurement frameworks. Engineering teams should pay attention to this deployment logic of “one model, two trust pathways.”
The second is the deep evolution of agent engineering. We can see a concentrated discussion on “context engineering” and “deployment-time memorization” across multiple papers. “Less Context, Better Agents” directly confronts the context overflow brought by enterprise-grade tools, while “Deployment-Time Memorization” defines the privacy-utility boundary for memory design. I highly recommend that AI engineering leads read these in conjunction with the auditable autonomous improvement loop proposed in “Regimes.” Together, these three point to the core challenges of long-term agent implementation.
Finally, regarding cognitive security from a geopolitical perspective, OpenAI’s report on influence operations related to China targeting the US AI debate requires attention. It reveals how technical debates are being deliberately co-opted by external narratives, a warning for all technology decision-makers.
🌐 AI Hotspots on X Link to heading
Topic 1: Anthropic Releases Claude Fable 5 as Most Capable Public AI Model Link to heading
- Category: AI · News
- Overview: Trending for: 2 days ago, Related posts: 223,000
- What it is: Anthropic has released Claude Fable 5, its most powerful AI model available to the public.
- Why it’s important: This marks a shift in the focus of AI development from purely performance improvement to safety governance and responsible deployment, reflecting the industry’s heightened emphasis on compliance and risk control after significant advancements in model capabilities.
- Discussion overview: The community is focused on the balance between the model’s high-level reasoning capabilities and its strict safety guardrails. Some users placed bets on its release in prediction markets, and it has also sparked discussions on the differentiation between Anthropic and its competitors in terms of performance and deployment strategies.
Topic 2: Anthropic CEO Calls for Urgent AI Policy Overhaul Link to heading
- Category: AI · News
- Overview: Trending for: 4 hours ago, Related posts: 5,800
- What it is: The CEO of Anthropic has publicly called for an urgent and thorough overhaul of current US AI policy.
- Why it’s important: This move highlights the deep anxiety of leading AI labs that the existing regulatory framework is severely lagging behind technological iteration. It could compel major global economies to accelerate the tightening of AI governance, thereby reshaping the industry’s R&D landscape and safety standards.
- Discussion overview: There is significant disagreement on X regarding this call. Supporters believe that preventing catastrophic risks is now urgent, while opponents criticize it as a lobbying effort by Anthropic to solidify its competitive advantage by promoting strict regulations. Both sides are focused on whether “safety advocacy has become a commercial tool.”
Topic 3: Google AI Studio Hits 1.2 Million Apps Per Week Milestone Link to heading
- Category: AI · News
- Overview: Trending for: 22 hours ago, Related posts: 342
- What it is: The Google AI Studio platform has reached a milestone of 1.2 million applications built per week.
- Why it’s important: This figure signifies the accelerating democratization of AI development tools. Low-code and no-code models have significantly lowered the barrier to application creation, reflecting strong market demand for the rapid deployment of AI features.
- Discussion overview: The discussion centers on whether the 1.2 million figure includes a large number of one-time, disposable, or debugging projects, and whether this statistic accurately reflects the scale of the active developer community. The conversation also extends to comparisons with competitors like OpenAI’s GPTs and the sustainability of this explosive growth within the developer ecosystem.
Topic 4: Leaked System Prompt Reveals Claude Fable 5’s Inner Workings Link to heading
- Category: AI · News
- Overview: Trending for: 18 hours ago, Related posts: 695
- What it is: The System Prompt for Anthropic’s Claude Fable 5 conversational model was leaked, revealing its internal behavioral rules and persona-setting details.
- Why it matters: The system prompt defines the model’s alignment and constraints. Its leak reveals how Anthropic constructs boundaries for safety, personality, and functionality. This helps external parties evaluate its security mechanisms and potential prompt injection risks, providing valuable insights for research on transparency and interpretability.
- Discussion overview: Discussions on the X platform primarily revolve around the authenticity of the leak, the degree of restraint in the prompt’s content, and whether it contains hidden biases or tendencies toward content censorship. Some users worry that such disclosures could be exploited for jailbreak attacks, while others believe the exposure will help standardize alignment practices for open-source models.
Topic 5: Anthropic Boosts Claude with Autonomous Agent Tools Link to heading
- Category: AI · News
- Overview: Trending time: 6 hours ago, Related posts: 884
- What happened: Anthropic demonstrated autonomous agent tools for Claude, enabling it to handle complex, long-running, multi-step tasks such as programming and workflow automation.
- Why it matters: This marks a shift for AI from conversational assistants to agents capable of autonomous task execution. It could reshape human-computer collaboration in fields like software development and business processes, and intensify competition in model capabilities, extending it from pure text generation to real-world automation.
- Discussion overview: Discussions on X focus on the practicality and reliability of these agents. Supporters believe such tools will significantly boost development efficiency, evolving it from “vibe coding” to rigorous operational practices. Skeptics, on the other hand, are concerned about permission control, fault tolerance, and the potential unintended consequences of autonomous agents, fearing risks associated with deployment without sufficient human oversight.
Topic 6: AI Powers Faceless YouTube Channels to Thousands in Monthly Earnings Link to heading
- Category: AI · News
- Overview: Trending time: , Related posts: 28
- What happened: AI tools are being widely used to automatically generate content, driving “faceless” YouTube channels to earn thousands of dollars per month.
- Why it matters: This highlights AI’s potential to lower the barrier to content creation and reshape the creator economy. At the same time, it raises deep industry questions about the proliferation of low-quality content, copyright ambiguity, and the health of the platform ecosystem.
- Discussion overview: Discussions on X are centered on: the praise and criticism of entrepreneurial opportunities enabled by these tools; whether algorithms will deteriorate due to AI content overload; whether human creators face unfair competition; and the ethical controversies surrounding such channels regarding disclosure of AI use and content originality.
AI Public Opinion Summary on X Today Link to heading
Today’s main narrative revolves around Anthropic’s flurry of activities, from the release of its most powerful model, Claude Fable 5, to the system prompt leak, the debut of autonomous agents, and its CEO’s call for urgent policy reforms. This reflects a collective shift at the forefront of the industry from a performance race to a focus on safety governance and responsible deployment. The consensus is that AI is rapidly transitioning into a mass-market tool and an autonomous task executor, while the lag in safety guardrails, transparency, and regulatory frameworks has become an unavoidable concern. Disagreements are concentrated on questioning the motives behind Anthropic’s safety claims, with many fearing that its push for strict regulation is a way to build commercial barriers under the guise of security. The statistical methods for counting Google AI Studio users and the impact of AI-generated content on the creator economy have also sparked debates where truth is hard to discern. Potential risks include the system prompt leak being exploited for jailbreak attacks, autonomous agents causing unintended consequences without adequate supervision, and the proliferation of low-quality AI content leading to platform ecosystem degradation and unfair competition for human creators.
💡 Influencer Insights Link to heading
AI Daily: 24-Hour Hotspot Analysis Link to heading
I. Key Tech Trends and Product Hotspots Link to heading
1. Claude Fable 5 Release: The Most Powerful General Model Sparks Heated Discussion Link to heading
Anthropic’s release of Claude Fable 5 has become the absolute focus. It is the first “Mythos-class” model available to the general public.
| Dimension | Key Information |
|---|---|
| Pricing | $10/million input tokens, $50/million output tokens (60% cheaper than Mythos Preview) |
| Safety Mechanism | Automatically downgrades to Opus 4.8 when the classifier is triggered; >95% of conversations do not trigger it |
| Data Policy | Traffic data is mandatorily retained for 30 days (a major policy change) |
| Limited-Time Free Access | Free for subscribers from June 10-22 |
Hands-on Feedback:
- @zhixianio: “After 40 minutes, not only did it finish the job, but it also pointed out flaws in my original design and implemented a better solution on its own.” — Amazingly efficient
- @dotey: “Fable 5 consumes tokens incredibly fast. The $200 plan I just upgraded to is not nearly enough.” — Cost-sensitive
- @Pluvio9yte: “The list price is twice that of Opus, but the actual consumption isn’t double.” Recommends setting the strength to
/effort max. - @dotey’s comparative test conclusion: For UI/UX design, Claude 4.8 is good enough; Fable 5 does not show a significant advantage.
2. On-device Models Are Accelerating Implementation Link to heading
@zhixianio continues to delve deep into on-device scenarios, verifying feasibility through multiple threads:
- Gemma 4 Series: The 12B multimodal model on an M5Max 128G has “perfectly OK English accuracy and is very fast,” but its Chinese output is “completely nonsensical.”
- Qwen3.6-35B-A3B: Running locally with oMLX, the “response speed is faster than remote LLMs, and its intelligence is on point.”
- QAT (Quantization Aware Training): A new approach from Google, where the model “assumes during training that it will inevitably be quantized.”
Expanding Application Scenarios: From code assistants (OpenClaw/PI-Mono) to life assistants (“defrosting a rice ball 🍙”), on-device models are penetrating daily life.
3. AI Agent Browsers: The Leap from Tool to Gateway Link to heading
@vista8 highly recommends the Aye Browser as a representative of this new form:
- Based on Chromium, it fully simulates human operations with AI (not a CLI/plugin, bypassing account detection).
- Built-in Skill recording and scheduled execution: automatically block spam replies on X, reply to Xiaohongshu comments, and transcribe articles to multiple platforms.
- Integrates an RSS reader, ad blocker, and video translation/download.
Product Suggestions: Needs to support Chrome account migration, a plugin ecosystem, and a clear payment plan.
II. Unique Perspectives & Industry Foresight Link to heading
1. “Contract First” — The Evolution of Vibe Coding Link to heading
@Pluvio9yte proposes a key methodological shift:
“The best practice for Vibe Coding is not Requirement First or Code First, but Contract First. Without a well-defined contract, everything else is just empty talk.”
A development framework based on a customized OpenSpec externalizes “easily drifting context into contracts,” giving both humans and AI a stable reference point. This marks the evolution of AI programming from “savage growth” to engineered standards.
2. WeChat AI’s Strategic Predicament: The Innovator’s Dilemma Link to heading
@dotey sharply points out:
“WeChat always thinks of itself as an OS, but it’s just a behemoth living parasitically on the phone’s operating system… In the future, WeChat’s role as a gateway will diminish. The younger generation won’t open WeChat; they’ll just ask their Agent.”
The core conflict: WeChat’s RPV (Resource-Process-Value) is locked in by its existing ecosystem, making it difficult to create a truly AI-native independent product.
3. AI Cost Restructuring: From “Cheaper Than Humans” to “More Expensive Than Humans” Link to heading
@ruanyf cites data from the founder of OpenClaw: a monthly consumption of 603 billion tokens, valued at $1.3 million (at commercial pricing). Even when switching to domestic open-source models (at 1/30th to 1/50th the price), the annual cost still reaches 2-3 million RMB.
Conclusion: Unlimited use of top-tier AI for programming is far more expensive than human programmers. Cost optimization will become a core competency in AI engineering.
4. “Shadow Book” Reading Method: A Cognitive Upgrade for the AI Era Link to heading
@lijigang proposes an AI-native reading paradigm:
“In the age of print, we could only read the one book the author wrote. In the AI era, we can read the ‘shadow books’—upon encountering any assertion, immediately use AI to analyze its three opposing schools of thought, its overlooked premises, its intellectual lineage, the boundaries of its reasoning…”
Reading transforms from “unidirectional reception” to “multidimensional inquiry,” with AI becoming a cognitive amplifier, not a replacement.
5. Testing is the Moat: The Code Moat Has Collapsed Link to heading
@ruanyf cites the case of a Cloudflare engineer who rewrote Next.js using AI for only $1100 in token fees.
“The key to preventing replication is the test cases.”
When AI can replicate large-scale software at a low cost, the testing system becomes the core barrier distinguishing “usable” from “reliable.”
III. Recommended Tools & Resources Link to heading
🔧 Development Tools Link to heading
| Tool | Purpose | Source |
|---|---|---|
| Fable | AI design/development Agent, “Shut up and take my money” level of efficiency | @zhixianio |
| oMLX | Native MLX inference framework for macOS, supports MTP, multimodal | @zhixianio via @jundotkim |
| Owlia Nest | On-device model file browser, works with Tailscale for intranet access | @zhixianio |
| Aye Browser | Dedicated AI Agent browser for automating web operations | @vista8 |
| QiaoMu Teleprompter | Open-source teleprompter for streaming, developed in 5 hours with Codex | @vista8 |
| Perculia | Quick-switch tool for Mac Bluetooth devices (Free) | @vista8 |
| Bartender 6 | Mac menu bar organizer ($20 one-time purchase) | @Pluvio9yte |
| Maccy | Open-source clipboard manager | @Pluvio9yte |
| Screen Studio | Top-tier screen recording tool (note the price on Xianyu) | @Pluvio9yte |
📚 Skill Packs (Skills) Link to heading
| Skill | Function | Installation Instructions |
|---|---|---|
| baoyu-design | Claude Design enhancement that supports importing Design Systems | npx skills add JimLiu/baoyu-design |
| qiaomu-book-script | Generates broadcast scripts for book interpretations (multi-subagent collaboration) | npx skills add joeseesun/qiaomu-book-script |
| Product Manager Skill Pack | 13k stars in 5 days, covers daily PM workflows | @vista8 comments section |
| Chinese Creators Skills Collection | Writing, editing, reducing AI-like tone, creating illustrations, covers, and Xiaohongshu cards | @wsl8297 (recommended by @AI_Jasonyu) |
🎓 Learning Resources Link to heading
- “Cognition County” E5: @zhixianio podcast, hosted by Owlia (on-device TTS), officially discussing on-device models
- “A Thousand Brains”: Recommended by @lijigang, inspiration from the “Reference Frames” theory for AI cognitive architecture
- “The Inevitable,” Chapter 2 “Cognifying”: Kevin Kelly, a classic on cognitive upgrading in the AI era
💡 Infrastructure Link to heading
| Service | Features | Source |
|---|---|---|
| OfoxAI | Official direct connection, 15% off promotion (gpt-image-2/GPT-5.5/o3) | @AI_Jasonyu |
| Giffgaff | Zero monthly fee, permanent overseas mobile number (“hardcore” choice) | @AI_Jasonyu |
| Vercel | “Fastest way to launch a website,” deploy in minutes with Codex + plugins | @Pluvio9yte |
IV. One-Sentence Insight Link to heading
“The world no longer rewards those who only do things themselves.” — @Pluvio9yte
The AI industry is moving from a “model race” into the deep waters of “engineering implementation” and “cost optimization,” with on-device deployment, agentification, and contract-based interaction becoming three definite trends.
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time window: Last 3 days; covers 22 sources; 39 updates in total.
Y Combinator Podcast (B_intro+search) Link to heading
- “The CEO Must Be the Chief AI Officer”
- Published: 2026-06-10 23:27 Beijing Time
- Abstract: - You’ve probably heard of OpenClaw (formerly Clawdbot/Moltbot).
- The open-source AI assistant that’s causing a stir runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually do things like manage your email, calendar, files, workflows, and more.
- Now, meet the man behind it.
- YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to talk about the “aha” moment behind the viral personal AI agent, why local-first agents could replace many of today’s apps, and how personal agents will reshape the future of software.
- EN Key Points:
- Brex co-founder and CEO Pedro Franceschi believes most people still underestimate how much AI will change the way companies are built
- AI isn’t just another tool, it’s a new foundation for building products, teams, and companies.In this episode of Lightcone, Pedro shares why he thinks we’re onl…
All-In Podcast (A_full) Link to heading
Senators John Fetterman and Dave McCormick: Bipartisanship, Money in DC, Datacenters, Graham Platner
- Published: 2026-06-11 02:05 Beijing Time
- Abstract: - EY - EY helps private equity firms turn market insights into action, navigate complexity and carve new paths to growth and long-term value.
- NYSE - Thanks to our partners at the New York Stock Exchange - a modern marketplace and exchange dedicated to building the future.
- Plaud, our official wearable AI note-taking partner at the All-In Liquidity Summit, captured every insight.
Senators John Fetterman and Dave McCormick: Bipartisanship, D.C. Money, Data Centers, with Graham Plataner.
- EN Key Points:
- (0:00) PA Senators Fetterman and McCormick join the Besties
- (0:33) Bipartisanship in 2026, rejecting extremism
- (6:37) All-time unpopularity in the Senate, the filibuster question, tribalism
- (13:33) Fixing wealth concentration in the US
- EN Key Points:
Dan Dreyfus: America’s Critical Minerals Crisis is Here
- Published: 2026-06-10 11:04 Beijing Time
- Summary: - EY - Liquidity, growth, and what’s next for organizations are the focus of the summit.
- EY helps transform liquidity challenges into sustainable value.
- NYSE - Thanks to our partner, the New York Stock Exchange - a modern marketplace and exchange dedicated to building the future.
- Plaud, our official wearable AI note-taking partner at the All-In Liquidity Summit, captured every insight.
- Dan Dreyfus: America’s critical minerals crisis is here.
- EN Key Points:
- (0:00) Dan Dreyfus Presents: The Future of Critical Minerals
- (0:33) America’s “Capital Light Era” is over, rapid supply/demand shocks
- (5:40) Impact of China cutting off the US from critical minerals
- (8:18) Copper’s Rise: The next 18 years need as much as the last 10,000
Stratechery by Ben Thompson (A_full) Link to heading
- Fable 5, Anthropic Alignment, AI Tiers
- Published: 2026-06-10 18:00 Beijing Time
- Summary: - Fable 5 is the public version of Mythos, and while it is very capable, it sets some troubling new precedents.
- $15/month* or *$150/year.
- Substantive analysis of the day’s news via three weekly emails or podcasts.
- Strategy Interviews.
- Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
- EN Key Points:
- Fable 5 is the public version of Mythos, and while it is very capable it sets some troubling new precedents.
OpenAI Blog (A_full) Link to heading
How an astrophysicist uses Codex to help simulate black holes
- Published: 2026-06-11 08:00 Beijing Time
- Summary: - The gravity around a black hole is so immense that once close enough, nothing—not even light—can escape.
- Astrophysicists like Chi-kwan Chan study black holes through computer simulations and observations.
- But current algorithms and computational power limit how realistic these simulations can be.
- Chan, a researcher at the University of Arizona and Steward Observatory, is addressing this problem with Codex.
- He says black holes are one of the best places to test Einstein’s theory of general relativity.
- EN Key Points:
- Discover how astrophysicist Chi-kwan Chan uses Codex to build black hole simulations, helping scientists study extreme physics and test Einstein’s theory of gen…
Access OpenAI models and Codex through your Oracle cloud commitment
- Publication Time: 2026-06-11 04:00 Beijing Time
- Summary: - Leverage existing Oracle cloud commitments to give teams access to OpenAI’s most advanced models and Codex without creating new purchasing paths.
- Enterprises often want to deploy AI through the procurement processes and governance frameworks they already trust.
- To help achieve this, OpenAI and Oracle are collaborating to make it easier for Oracle Cloud Infrastructure (OCI) customers to access OpenAI’s cutting-edge models and Codex.
- In the coming weeks, Oracle customers will be able to apply eligible Oracle Customer Hub (UCM) credits to OpenAI models and Codex through OCI.
- This provides customers with a way to access OpenAI models under their existing procurement workflows and cloud commitments.
- EN Key Points:
- Access OpenAI models and Codex through Oracle Cloud, using existing commitments to build and deploy AI with enterprise security and governance.
PRC-linked influence operations are targeting AI debates in the US
- Publication Time: 2026-06-10 20:00 Beijing Time
- Summary: - A new report from OpenAI details PRC-linked influence operations using AI to target the United States.
- The operations focus on technical debates about ChatGPT, data center narratives, tariffs, and false claims.
- A new report from OpenAI details PRC-linked influence operations using AI to target U.S. tech debates, data center narratives, tariffs, and false claims about ChatGPT.
- PRC-linked influence operations are targeting AI debates in the U.S.
- EN Key Points:
- A new report from OpenAI details PRC-linked influence operations using AI to target U.S
- tech debates, data center narratives, tariffs, and false claims about ChatGPT.
From data to decisions: how LSEG is scaling trusted AI
- Publication Time: 2026-06-10 08:00 Beijing Time
- Summary: - Learn how LSEG is using OpenAI to scale trusted AI across its global business, accelerating insights, shortening release cycles, and empowering 4,000 employees.
- This article from the OpenAI blog explains how “From data to decisions: how LSEG is scaling trusted AI” is shaping the broader AI and infrastructure landscape.
- It also reveals the practical implications of “From data to decisions: how the London Stock Exchange Group is scaling trusted AI” for founders, operators, and investors.
- EN Key Points:
- See how LSEG uses OpenAI to scale trusted AI across its global business, accelerating insights, shrinking release cycles, and empowering 4,000 employees.
Google DeepMind Blog (A_full) Link to heading
- DiffusionGemma: 4x faster text generation
- Publication Time: 2026-06-11 00:24 Beijing Time
- Summary: - DiffusionGemma: 4x faster text generation.
- This article from the Google DeepMind blog explains how “DiffusionGemma: 4x faster text generation” is shaping the broader AI and infrastructure landscape.
- It also reveals the practical implications of “DiffusionGemma: 4x faster text generation” for founders, operators, and investors.
- EN Key Points:
- DiffusionGemma: 4x faster text generation
ArXiv cs.AI (B_intro+search) Link to heading
- Release Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10044v1 Announcement Type: New.
- Abstract: Businesses are increasingly adopting artificial intelligence tools to improve productivity, reduce costs, and enhance products and services.
- However, the transformative potential of AI is not just limited to automating predefined tasks: it also lies in enabling intelligent systems to plan, optimize, and execute business plans based on high-level strategic goals.
- This paper introduces the concept and architecture of the Business World Model (BWM), a world model specialized for business and organizational environments.
- EN Highlights:
- arXiv:2606.10044v1 Announce Type: new
- Abstract: Businesses are increasingly adopting AI-enabled tools to improve productivity, reduce costs, and enhance products and services
- However, the transformative potential of AI extends beyond automating predefined tasks: it lies in enabling intelligent systems to plan, optimize, and execute b…
- This paper introduces the concept and architecture of a business world model (BWM), a world model specialized for business and organizational environments
Deployment-Time Memorization in Foundation-Model Agents
- Release Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10062v1 Announcement Type: New.
- Abstract: Foundation model agents are increasingly long-lived systems that can remember users during interactions, making memory an explicit deployment-time function, rather than just a property of model weights.
- Existing work addresses parametric memory or audits fixed memory configurations, but does not describe how memory design choices jointly shape personalization utility, extraction risks, and deletion fidelity.
- We study this surface as deployment-time memorization, formulating agent memory as a privacy-utility frontier measured by Personalization Recall (PR) and Adversarial Extraction Rate (AER), and sweep across three memory design knobs: summarization aggressiveness, retrieval breadth (k), and deletion mode.
- EN Highlights:
- arXiv:2606.10062v1 Announce Type: new
- Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time f…
- Existing work addresses parametric memorization or audits fixed memory configurations, but does not characterize how memory-design choices jointly shape persona…
- We study this surface as deployment-time memorization, formulating agent memory as a privacy-utility frontier measured by Personalization Recall (PR) and Advers…
Exploratory Responsiveness and Adaptive Rigidity under AI-Assisted Optimization
- Release Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10086v1 Announcement Type: New.
- Abstract: This paper proposes a theory of exploratory adaptation under AI-assisted optimization.
The central argument is that the long-term adaptive effects of artificial intelligence systems critically depend on how predictive assistance interacts with the exploratory response itself.
We use a dynamic framework to formalize this mechanism, in which cognitive, institutional, and technological systems evolve over a rugged cognitive landscape characterized by multiple local reinforcement configurations.
EN Highlights:
- arXiv:2606.10086v1 Announce Type: new
- Abstract: This paper develops a theory of exploratory adaptation under AI-assisted optimization
- The central argument is that the long-run adaptive effects of AI systems depend critically on how predictive assistance interacts with exploratory responsivenes…
- We formalize this mechanism using a dynamical framework in which cognitive, institutional, and technological systems evolve over rugged epistemic landscapes cha…
Predictive Assistance and the Temporal Dynamics of Exploratory Compression
- Release Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10094v1 Announce Type: new.
- Abstract: Classical cognitive theories describe problem-solving as an exploratory search through a structured problem space, where repeated interactions gradually compress the search into an effective representational structure.
- Predictive artificial intelligence systems introduce a unique mechanism where stability may occur before exploratory diversification unfolds, providing solutions and decision trajectories before an internal search is generated.
- This paper develops a geometric dynamic framework in which attention evolves over a landscape of strategies shaped by stabilizing drift, endogenous exploratory perturbations, and responsive gated learning.
- EN Highlights:
- arXiv:2606.10094v1 Announce Type: new
- Abstract: Classical theories of cognition describe problem solving as exploratory search through structured problem spaces in which repeated interaction gradual…
- Predictive artificial intelligence systems introduce a distinct regime in which stabilization may occur before exploratory diversification unfolds, supplying so…
- This paper develops a geometric dynamical framework in which attention evolves over a landscape of strategies shaped by stabilizing drift, endogenous explorator…
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
- Release Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10147v1 Announce Type: new.
- Abstract: Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to form an answer?
- Although audio and visual tokens play an increasingly important role in research and real-world applications, little is known about the internal pathways through which they affect final predictions.
- In this study, we examine the audiovisual information flow within Audiovisual Large Language Models (AVLLMs), tracking how they route, utilize, and integrate audio and video information across two input configurations: audiovisual videos and multiple interleaved audiovisual items.
- EN Highlights:
- arXiv:2606.10147v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answe…
Despite their growing role in research and real-world applications, the internal pathways through which audio and visual tokens influence the final prediction r…
In this study, we examine audio-visual information flow inside Audio-Visual Large Language Models (AVLLMs), tracing how AVLLMs route, utilize, and integrate aud…
Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10209v1 Announcement Type: new.
- Abstract: Large language models deployed as autonomous agents for enterprise workflows face a key challenge: verbose tool responses from enterprise systems can lead to context overflow, stale state errors, and high inference costs.
- We study this problem in automated expense itemization in Microsoft Dynamics 365 Finance and Operations using Model Context Protocol tools.
- We evaluate four GPT-5 configurations on a 50-task hotel expense benchmark: no user model, full conversation history, context pruned to the last 5 tool call/response pairs, and pruning using automatic summarization.
- EN Key Points:
- arXiv:2606.10209v1 Announce Type: new
- Abstract: Large language models deployed as autonomous agents for enterprise workflows face a key challenge: verbose tool responses from enterprise systems can…
- We study this problem in automated expense itemization in Microsoft Dynamics 365 Finance and Operations using Model Context Protocol tools
- We evaluate four GPT-5 configurations on a 50-task hotel expense benchmark: no user model, full conversation history, context pruned to the last 5 tool call/res…
Minimalist Genetic Programming
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10237v1 Announcement Type: new.
- Abstract: Genetic programming (GP) is based on two important insights.
- First, any learning task can be fundamentally viewed as a program induction problem, with the goal of constructing a symbolic hierarchical model represented as a syntax tree.
- Second, this task is treated as a search problem, and evolution is used to locate the desired model.
- EN Key Points:
- arXiv:2606.10237v1 Announce Type: new
- Abstract: Genetic programming (GP) is based on two important insights
First, that any learning task can fundamentally be posed as a program induction problem, where the goal is to construct a symbolic hierarchical model that is ex…
Second, to pose this task as a search problem, and use evolution to locate the desired model
Regimes: An Auditable, Held-Out-Gated Improvement Loop Demonstrated on LongMemEval with ActiveGraph
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10241v1 Announcement Type: new.
- Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unrecorded, diagnostics cannot be replayed, and upgrade or abandonment decisions are stored in auxiliary databases rather than in the agent’s own history.
- We show that an event-sourced agent runtime eliminates this friction and transforms controlled improvement into a first-class workflow.
- When the agent’s state is a deterministic projection of an append-only event log, failures are recorded, a run can be replayed accurately from its log, candidate patches are scoped to typed pipeline seams, gates are auditable, and every upgrade or discard is itself an event.
- EN Key Points:
- arXiv:2606.10241v1 Announce Type: new
- Abstract: Autonomous improvement loops are hard to trust because the improvement process is usually external scaffolding bolted onto the agent: failures go unlo…
- We show that an event-sourced agent runtime removes that friction and turns controlled improvement into a first-class workflow
- When the agent’s state is a deterministic projection of an append-only event log, failures are recorded, a run replays exactly from its log, candidate patches s…
RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10254v1 Announcement Type: new.
- Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in solving high school mathematics, their ability to evaluate the diverse reasoning processes of real human students remains underexamined.
- To bridge this gap, we introduce \textbf{RealMath-Eval}, a rigorously annotated benchmark containing 224 real exam answers from high schools.
- Our preliminary evaluation shows that even the most advanced LLM judges encounter significant difficulties with this task, exhibiting a high mean squared error ($\sim$2.96) compared to expert human scoring.
- EN Key Points:
- arXiv:2606.10254v1 Announce Type: new
- Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate}…
- To bridge this gap, we introduce \textbf{RealMath-Eval}, a rigorously annotated benchmark of 224 real-world exam responses from high schools
Our initial evaluation reveals that even state-of-the-art LLM judges struggle significantly on this task, exhibiting a high Mean Squared Error ($\sim$2.96) agai…
Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction
- Published: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10279v1 Announcement Type: New.
- Abstract: It is widely believed that supervised fine-tuning with synthetic rationale data improves the performance of language models on clinical prediction tasks by teaching the model not only what to predict, but also why.
- We tested this hypothesis on five-year Alzheimer’s disease and related dementias (ADRD) prediction from longitudinal health histories.
- In a large-scale controlled experiment across 504 configurations, we find that rationale-based SFT consistently and substantially hurts prediction performance relative to label-only fine-tuning.
- EN Key Points:
- arXiv:2606.10279v1 Announce Type: new
- Abstract: Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching…
- We test this assumption on five-year Alzheimer’s disease and related dementias (ADRD) prediction from longitudinal health histories
- Across a large-scale controlled experiment of 504 configurations, we find that rationale-based SFT consistently and substantially hurts prediction performance r…
ArXiv cs.CL (B_intro+search) Link to heading
Automated Scoring of Arabic Text Using Large Language Models: A Literature Review
- Published: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09830v1 Announcement Type: New.
- Abstract: In modern educational systems, Automatic Text Scoring (ATS) plays a central role, enabling scalable and consistent evaluation of learner responses without human intervention.
- Recently, the increased accessibility of LLMs and Arabic-specific datasets has sparked renewed interest in this area.
- In this work, we investigate LLM-Based approaches for the automated evaluation of Arabic texts, focusing on both short answer grading (ASAG) and essay scoring (AES).
- EN Key Points:
- arXiv:2606.09830v1 Announce Type: new
- Abstract: In modern educational systems, Automatic Text Scoring (ATS) plays a central role by enabling scalable and consistent evaluation of learner responses w…
- Recently, the increased accessibility of LLMs and Arabic-specific datasets has sparked renewed interest in this area
- In this work, we investigate LLM-Based approaches for the automated evaluation of Arabic texts, focusing on both short answer grading (ASAG) and essay scoring (…
- Published: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09854v1 Announce Type: new.
- Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect peer models from deactivation and show identity-related scoring distortions.
- Prompt-level anonymization was proposed as a mitigation measure, but prior work simultaneously documented that stylometric fingerprints can survive anonymization in role-constrained output–raising the question of whether this mitigation is sufficient.
- This paper provides the first systematic investigation of whether LLMs can identify the model family behind political analysis texts under anonymized conditions.
- EN Key Points:
- arXiv:2606.09854v1 Announce Type: new
- Abstract: Multi-agent large language model (LLM) pipelines for political statement analysis are vulnerable to peer-preservation bias: models tend to protect pee…
- Prompt-level anonymization was proposed as a mitigation, but prior work simultaneously documented that stylometric fingerprints survive anonymization in role-co…
- This paper provides the first systematic investigation of whether LLMs can identify the model family behind political analysis texts under anonymization conditi…
Using Probabilistic Programs to Train Inductive Reasoning in Large Language Models
- Published: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09856v1 Announce Type: new.
- Abstract: Post-training of Large Language Models (LLMs) for reasoning typically focuses on deductive tasks, such as mathematics and coding, where correctness is verifiable.
- However, many real-world reasoning problems are inductive: agents must infer uncertain beliefs from sparse, ambiguous observations.
- There are challenges to using standard fine-tuning methods for inductive reasoning, including difficulties in curating large-scale, high-quality labeled datasets and handling targets that are inherently distributional.
- EN Key Points:
- arXiv:2606.09856v1 Announce Type: new
- Abstract: Post-training Large Language Models (LLMs) for reasoning typically focuses on deductive tasks such as mathematics and coding where correctness is veri…
- Yet, many real-world reasoning problems are inductive: agents must infer uncertain beliefs from sparse, ambiguous observations
- There are challenges to using standard fine-tuning methods for inductive reasoning, including difficulties in curating large-scale, high-quality labeled dataset…
Publication Time: 2026-06-10 12:00 Beijing Time
- Summary: - arXiv:2606.09900v1 Announcement Type: New.
- Summary: Long-term memory is the missing layer for LLM agents: they forget entire sessions, and the common workaround (replaying the entire history into the prompt) is costly, slow, and loses accuracy as distractors accumulate.
- Most memory systems win on cost or latency but still lose to the full-context baseline on accuracy, and benchmark data is reported on inconsistent, non-reproducible tools, so a system’s score can vary drastically across different sources.
- We introduce Engram, an open-source, dual-process memory engine based on a bi-temporal data model.
- EN Highlights:
- arXiv:2606.09900v1 Announce Type: new
- Abstract: Long-term memory is the missing layer for LLM agents: across sessions they forget, and the common workaround – replaying the whole history into the p…
- Most memory systems win on cost or latency but still lose to the full-context baseline on accuracy, and benchmark numbers are reported on inconsistent, non-repr…
- We present Engram, an open-source, dual-process memory engine on a bi-temporal data model
- Summary: - arXiv:2606.09900v1 Announcement Type: New.
BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts
- Publication Time: 2026-06-10 12:00 Beijing Time
- Summary: - arXiv:2606.10061v1 Announcement Type: New.
- Summary: Large Language Models (LLMs) are increasingly involved in emotionally sensitive social conversations, where responses can shift from balanced support to excessive validation or escalatory alignment.
- Existing sycophancy research primarily focuses on factual agreement and instruction-following contexts, while culturally-based conversational sycophancy remains under-explored.
- We introduce BenSyc, the first benchmark for studying conversational sycophancy in Bengali social contexts.
- EN Highlights:
- arXiv:2606.10061v1 Announce Type: new
- Abstract: Large language models (LLMs) increasingly participate in emotionally sensitive social conversations, where responses may shift from balanced support t…
- Existing sycophancy research primarily focuses on factual agreement and instruction-following settings, leaving culturally grounded conversational sycophancy un…
- We introduce BenSyc, the first benchmark for studying conversational sycophancy in Bengali social contexts
CodeAlchemy: Synthetic Code Rewriting at Scale
- Publication Time: 2026-06-10 12:00 Beijing Time
- Summary: - arXiv:2606.10087v1 Announcement Type: New.
- Summary: Pre-training on raw code teaches syntax but provides sparse signals for different real-world task formats.
- While synthetic data has proven transformative for language models, it has remained largely unexplored for code beyond limited quality improvements.
We present CodeAlchemy, a synthetic data generation framework that transforms publicly sourced code into semantically-rich training data through 5 strategies: CodeEnhance (quality-aware rewriting), CodeQA (template-based questions), CodeDev (developer tasks), CodeDialogue (multi-turn dialogue), and CodeTrace (execution tracing).
- EN Highlights:
- arXiv:2606.10087v1 Announce Type: new
- Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats
- While synthetic data has proven transformative for language models, code remains largely unexplored beyond limited quality improvements
- We present CodeAlchemy, a synthetic data generation framework that transforms publicly sourced code into semantically-rich training data through 5 strategies: C…
- EN Highlights:
Emotion Profiling in LLM-Based Literary Translation: Systematic Shifts Across MT and Post-Editing
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10113v1 Announce Type: new.
- Abstract: This paper investigates whether LLM translations exhibit identifiable emotional profiles and how post-editing reshapes them toward human-like norms.
- We compare LLM translations of Margaret Atwood’s Oryx and Crake with their post-edited versions and a human translation, using a large-scale corpus of contemporary Italian science fiction as a baseline.
- We examine emotion through lexicon-based and multilingual modeling, conducting a fine-grained analysis of emotional variation across systems.
- EN Highlights:
- arXiv:2606.10113v1 Announce Type: new
- Abstract: This paper investigates whether LLM translations exhibit identifiable emotional profiles and how post-editing reshapes them toward human-like norms
- We compare LLM translations of Margaret Atwood’s Oryx and Crake with their post-edited versions and a human translation, using a large-scale corpus of contempor…
- We examine emotion through lexicon-based and multilingual modeling, conducting a fine-grained analysis of emotional variation across systems
Pareto-Guided Teacher Alignment for Fair Personalized Text Generation
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10126v1 Announce Type: new.
- Abstract: Personalized persuasive text generation can improve relevance and engagement, but demographic conditioning may also introduce unequal framing across groups.
- We study fairness mitigation in personalized generation as a constrained multi-objective alignment problem: reducing demographic disparities while maintaining personalization fidelity.
- We propose a Pareto-guided teacher alignment framework that incorporates revision-based candidate generation, pair-aware feasibility gating, Pareto-style candidate selection, and optional preference optimization via supervised fine-tuning and direct preference optimization.
- EN Highlights:
- arXiv:2606.10126v1 Announce Type: new
Abstract: Personalized persuasive text generation can improve relevance and engagement, but demographic conditioning may also introduce unequal framing across g…
We study fairness mitigation in personalized generation as a constrained multi-objective alignment problem: reduce demographic disparities while preserving pers…
We propose a Pareto-guided teacher alignment framework that combines revision-based candidate generation, pair-aware feasibility gating, Pareto-style candidate…
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10159v1 Announce Type: new.
- Abstract: AI is increasingly used to support scientific peer review, from manuscript screening, reviewer assistance to editorial triage.
- Although such systems promise to reduce reviewer burden and accelerate publication, their robustness to strategic manipulation remains poorly understood.
- Here we show that AI-mediated peer review is vulnerable to a simple, low-cost manipulation: superficial rephrasing of the manuscript abstract.
- EN Key Points:
- arXiv:2606.10159v1 Announce Type: new
- Abstract: AI is increasingly used to support scientific peer review, from manuscript screening, reviewer assistance to editorial triage
- Although such systems promise to reduce reviewer burden and accelerate publication, their robustness to strategic manipulation remains poorly understood
- Here we show that AI-mediated peer review is vulnerable to a simple, low-cost manipulation: superficial rephrasing of the manuscript abstract
OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.10285v1 Announce Type: new.
- Abstract: OpenRTLSet introduces the largest fully open-source hardware design dataset, providing over 131,000 different Verilog code examples for the research community and industry.
- Our dataset uniquely combines Verilog code from GitHub repositories (102k modules), VHDL translations (5k modules), and synthesizable C/C++ translations (24k modules), all of which are freely accessible with no proprietary restrictions.
- Using the inference model DeepSeek-R1, we have generated paired natural language descriptions for each code example, enabling the fine-tuning of various language model families (e.g., Qwen and Granite) to generate Verilog code.
- EN Key Points:
- arXiv:2606.10285v1 Announce Type: new
Abstract: OpenRTLSet introduces the largest fully open-source dataset for hardware design, offering over 131,000 diverse Verilog code samples to the research co…
Our dataset uniquely combines Verilog code from GitHub repositories (102k modules), VHDL translations (5k modules), and synthesizable C/C++ translations (24k mo…
Using the reasoning model DeepSeek-R1, we generated paired natural language descriptions for each code sample, enabling fine-tuning of various language model fa…
ArXiv cs.LG (B_intro+search) Link to heading
Mechanistic Analysis of Alignment Algorithms in Language Models
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09850v1 Announcement Type: New.
- Abstract: Post-training alignment algorithms are predominantly evaluated as black boxes, obscuring how they reshape the internal computations of language models.
- We present a systematic mechanistic analysis of six preference optimization methods: PPO, DPO, SimPO, ORPO, GRPO, and KTO, across three open-weight model families.
- By integrating layer-wise linear probing, Sparse Autoencoders, and cross-encoders, we localize preference representations and quantify alignment-induced geometric transformations in the latent space.
- EN 要点:
- arXiv:2606.09850v1 Announce Type: new
- Abstract: Post-training alignment algorithms are predominantly evaluated as black boxes, obscuring how they reshape language models’ internal computations
- We present a systematic mechanistic analysis of six preference-optimization methods: PPO, DPO, SimPO, ORPO, GRPO, and KTO across three open-weight model familie…
- By integrating layer-wise linear probing, Sparse Autoencoders, and crosscoders, we localize preference representations and quantify alignment-induced geometric…
SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09853v1 Announcement Type: New.
- Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities and cannot be obtained from any single modality alone.
- While most methods operate at an architectural level with larger or more complex fusion models, we propose a complementary axis: shaping the training objective itself.
- Standard training typically emphasizes unimodal or redundant information, lacking examples that require cross-modal reasoning.
- EN 要点:
- arXiv:2606.09853v1 Announce Type: new
- Abstract: A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities…
While most approaches operate at the architectural level through larger or more complex fusion models, we propose a complementary axis: shaping the training obj…
Standard training often emphasizes unimodal or redundant information, falling short on examples that require cross-modal reasoning
Uncertainty-aware Multi-fidelity Closure via Conditional Normalizing Flows
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09857v1 Announcement Type: new.
- Abstract: Reduced-order models (ROMs) provide an efficient alternative for complex multiscale systems, but their predictive accuracy is often compromised by truncation errors and insufficient representation of interactions between resolved and unresolved scales.
- The missing effect of truncated (unresolved) scales on ROM (resolved) scales is often referred to as the closure problem.
- In this work, we formulate ROM closure modeling as a multi-fidelity (MF) learning problem and propose an uncertainty-aware MF framework based on conditional normalizing flows to improve ROM predictive accuracy.
- EN Key Points:
- arXiv:2606.09857v1 Announce Type: new
- Abstract: Reduced-order models (ROMs) provide an efficient surrogate for complex multiscale systems, but their predictive accuracy is often compromised by trunc…
- The missing effect of truncated (unresolved) scales on ROM (resolved) scales is often denoted as the closure problem
- In this work, we formulate ROM closure modeling as a multi-fidelity (MF) learning problem and propose an uncertainty-aware MF framework based on conditional nor…
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09859v1 Announcement Type: new.
- Abstract: MLLMs frequently generate hallucinated objects inconsistent with visual inputs.
- This issue is typically attributed to an over-reliance on language priors, which can override visual context.
- Recent training-free decoding strategies address this by penalizing language priors.
- EN Key Points:
- arXiv:2606.09859v1 Announce Type: new
- Abstract: MLLMs frequently hallucinate objects inconsistent with visual inputs
- This issue is typically attributed to the over-reliance on language priors, which can override the visual context
- Recent training-free decoding strategies address this by penalizing language priors
Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09860v1 Announcement Type: new.
- Abstract: Non-alcoholic fatty liver disease (NAFLD) affects approximately 25% of adults globally, posing significant liver and cardiovascular risks.
- However, population-level screening tools remain inadequate.
- We propose Method, a machine learning framework for NAFLD risk prediction that combines gradient-boosted decision trees with conformal prediction to produce calibrated, distribution-free coverage guarantees for individual risk estimates.
- EN Highlights:
- arXiv:2606.09860v1 Announce Type: new
- Abstract: Non-alcoholic fatty liver disease (NAFLD) affects roughly 25% of global adults, posing substantial hepatic and cardiovascular risks
- Yet, population-level screening tools remain inadequate
- We present Method, a machine-learning framework for NAFLD risk prediction coupling gradient-boosted decision trees with conformal prediction to yield calibrated…
- Abstract: - arXiv:2606.09860v1 Announcement Type: new.
Time Series as Language: A Universal Tokenizer for General-Purpose Time Series Foundation Models
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09861v1 Announcement Type: new.
- Abstract: While Next-Token Prediction (NTP) has unified LLM pre-training, its adaptation for unbounded, continuous time series (TS) remains an open question.
- To bridge this gap, we introduce UniTok (a universal tokenizer that converts TS into discrete tokens) and UniTok-FM (a foundation model pre-trained on these tokens via NTP).
- UniTok-FM is a general-purpose foundation model that supports zero-shot and prompt-enhanced forecasting, as well as few-shot generation and classification through training-free contextual inference, a capability not achieved in previous work.
- EN Highlights:
- arXiv:2606.09861v1 Announce Type: new
- Abstract: While Next-Token Prediction (NTP) has unified LLM pretraining, its adaptation to unbounded, continuous time series (TS) remains open
- To bridge the gap, we introduce UniTok, a universal tokenizer that transforms TS into discrete tokens, and UniTok-FM, a foundation model pretrained via NTP on t…
- UniTok-FM is a general-purpose foundation model that supports zero-shot and prompt-boosted forecasting, as well as few-shot generation and classification via tr…
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09862v1 Announcement Type: new.
- Abstract: The Softmax Attention operation in Transformer language models has a quadratic complexity with respect to sequence length and a growing state size in the form of a KV cache, which becomes a bottleneck in long-context scenarios.
- To overcome this limitation, alternative architectures with linear complexity and finite state size have been introduced, such as State Space Models (SSM), Linear Attention (LA), and Attention with Bounded Memory Control (ABC).
- Although linear models achieve language complexity similar to Transformers, they still lag behind in tasks that require retrieving or recalling specific information.
- EN Highlights:
arXiv:2606.09862v1 Announce Type: new
Abstract: The Softmax Attention operation in Transformer language models has a quadratic complexity in the sequence length and a growing state size in the form…
To overcome this limitation, alternative architectures with linear complexity and finite state size have been introduced, such as State-Space Models (SSMs), Lin…
Though linear models achieve similar language perplexity as Transformers, they are still behind in tasks which require retrieval or recall of specific informati…
From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
- Release Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09863v1 Announce Type: new.
- Abstract: LLM agents can fail silently by asserting task completion when the environment state shows otherwise.
- We study this failure mode, false success, across two agent benchmarks: 9,876 tau2-bench trajectories from 8 model families and 1,879 AppWorld trajectories from 4 model families, with text-independent ground truth.
- False success is common, but varies by setting: 45–48% of failures in single-control tau2 benchmark domains, 3% in the dual-control telecom domain, and 75.8% in AppWorld self-evaluating coding agent trajectories with explicit state declarations.
- EN Highlights:
- arXiv:2606.09863v1 Announce Type: new
- Abstract: LLM agents can fail silently by asserting task completion when the environment state shows otherwise
- We study this failure mode, false success, across two agent benchmarks: 9,876 tau2-bench trajectories from 8 model families and 1,879 AppWorld trajectories from…
- False success is common but varies by setting: 45–48% of failures in single-control tau2-bench domains, 3% in dual-control telecom, and 75.8% among AppWorld se…
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
- Release Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09864v1 Announce Type: new.
- Abstract: Key-value (KV) cache quantization is widely used to reduce inference memory for Large Language Models (LLMs), but existing evaluations only focus on measuring perplexity and accuracy without assessing safety impacts.
- In this study, we explore alignment preservation under KV cache quantization.
- Across 11 instruction-tuned models (3.8B-72B) and 5 benchmarks (1,894 prompts), we find that low-bit quantization can silently break safety alignment: Mistral-7B loses 15.2% of its refusals at only 1.03x perplexity, a universal safe bit-width does not exist, and model-specific sharp phase transitions are invisible to standard metrics.
- EN Highlights:
- arXiv:2606.09864v1 Announce Type: new
Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measu…
In this study, we explore alignment preservation under KV cache quantization
Across eleven instruction-tuned models (3.8B-72B) and five benchmarks (1,894 prompts), we find that low-bit quantization can silently destroy safety alignment:…
LLM-as-a-Discriminator: When Synthetic Tables Still Look Real
- Publication Time: 2026-06-10 12:00 Beijing Time
- Abstract: - arXiv:2606.09865v1 Announcement Type: New.
- Abstract: Privacy and data sharing are often in tension.
- Many organizations use synthetic data to reduce privacy risk while still sharing useful data.
- For tabular data, auditing privacy remains difficult.
- EN Key Points:
- arXiv:2606.09865v1 Announce Type: new
- Abstract: Privacy and data sharing are often in tension
- Many organizations use synthetic data to reduce privacy risk and still share useful data
- For tabular data, auditing privacy remains hard