🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-20
- 类型
- ai-daily
- 字数
- 7804
- 阅读时长
- 37 min
2026-06-20 AI Daily | Agents Enter the Governance Era, On-Device Models Start Reshaping Workflows Link to heading
Today’s main theme shifts from model capabilities to controllable implementation: Agents require permissions, auditing, and runtime governance; on-device models are rapidly maturing in multimodal, programming, and personal assistant scenarios; and AI coding tools are evolving from assisted generation towards reusable automated workflows.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The most important theme to explore today is “Agents Moving from Demonstration to Governance.” The interview with the founder of OpenClaw is highly relevant for product and engineering teams: truly valuable Agents are being defined by domain experts who understand the context. Meanwhile, several papers on runtime governance, DeFi risk supervision, and clarifying uncertainty remind us that systems capable of tool use must be integrated into a framework of permissions, obligations, and auditing.
The second theme is model paradigms and reliability. Work like Diffusion Language Models and ITNet continues to challenge the default assumptions of Transformers/auto-regression. “Hidden anchors” in multi-agent deliberation, cognitive blind spots in clinical tabular data, and bias revealed through random path aggregation all point to one problem: models not only need to be more powerful, but they also need to know why they are wrong and when to stop.
Finally, the power structures within the tech industry also warrant attention. From the SpaceX IPO and rumors of Cursor’s acquisition to the controversy over Anthropic’s fables, the AI industry is simultaneously reshaping capital, platforms, and the public narrative.
🌐 AI Hot Topics on X Link to heading
Topic 1: Transformer Pioneer Noam Shazeer Leaves Google for OpenAI Link to heading
- Category: AI · News
- Overview: Trending since: 1 day ago, Related posts: 15,000
- What happened: Noam Shazeer, co-author of the Transformer paper and a pioneer of MoE, has reportedly left Google to join OpenAI, sparking widespread attention on the movement of top AI talent.
- Why it matters: Shazeer was instrumental in laying the groundwork for modern large model architectures. His move is considered a significant signal for measuring cutting-edge R&D capabilities, organizational attractiveness, and the future direction of AI technology.
- Discussion summary: Discussions on X are focused on whether Google is continuing to lose key AI talent, whether OpenAI is further solidifying its research lead, and who will pioneer the next-generation architecture or training paradigm that will succeed the Transformer. Some also questioned the accuracy of the report’s details and timeline.
Topic 2: Loop Engineering Turns AI Agents into Self-Sustaining Coders Link to heading
- Category: AI · News
- Overview: Trending since: 10 hours ago, Related posts: 336
- What happened: Loop Engineering has proposed a method that allows AI agents to autonomously write, test, and improve code through a cyclical development process, referring to them as “self-sustaining coders.”
- Why it matters: This reflects a trend of AI programming agents evolving from assisting with code generation to continuously performing engineering tasks. This could affect software development efficiency, the level of automation, and the division of roles for human engineers.
- Discussion summary: Discussions on X are mainly focused on whether such agents can truly achieve long-term autonomous development, how code quality and security can be guaranteed, and whether they will boost developer productivity or introduce risks of excessive automation and job displacement.
Topic 3: Anthropic Fixes Claude Code Usage Bug for Premium Users Link to heading
- Category: AI · News
- Overview: Trending since: 20 hours ago, Related posts: 3,700
- What happened: Anthropic has fixed a bug that affected usage statistics or credit deductions for premium subscribers of Claude Code.
- Why it matters: Claude Code is designed for high-frequency developer workflows, and the accuracy of its credit and billing systems is crucial for user trust, enterprise adoption, and the viability of its business model.
- Discussion summary: Discussions on X centered on whether Anthropic will promptly compensate affected users, whether Claude Code’s usage policies are sufficiently transparent, and if Anthropic can maintain developer confidence as competitors like OpenAI accelerate the launch of their programming agent capabilities.
Topic 4: Z.ai’s GLM-5.2 Tops Open AI Model Charts with Strong Benchmarks Link to heading
- Category: AI · News
- Overview: Trending since: 1 day ago, Related posts: 19,000
- What happened: GLM-5.2, released by Zhipu’s Z.ai, has achieved leading results on several open model benchmarks, attracting attention from the AI community.
- Why it matters: This demonstrates that China’s open-source large models are continuing to approach or surpass mainstream international models in reasoning, coding, and general capabilities, which could accelerate competition and application deployment in the open model ecosystem.
- Discussion summary: Discussions on X are mainly focused on whether the benchmark results reflect actual capabilities, the gap between GLM-5.2 and models like DeepSeek, Qwen, and Llama, and its attractiveness to developers regarding open weights, commercial licensing, and deployment costs.
Summary of Today’s AI Sentiments on X Link to heading
Today’s main narrative revolves around the concentration of cutting-edge AI capabilities in stronger organizations, more autonomous tools, and more open ecosystems. The movement of top research talent is seen as a bellwether for the competitiveness of institutions like OpenAI and Google, while advancements in coding agents and open-source models indicate that AI is shifting from a competition of model capabilities to one of practical engineering productivity. The consensus is that AI programming, reasoning, and the open-source model ecosystem are all rapidly maturing. Developer workflows will be profoundly reshaped, and platform advantages will be determined by a combination of talent, computing power, product experience, and business trust. The main points of contention are whether related news and benchmarks are reliable, whether autonomous coding agents truly possess long-term engineering capabilities, and whether the leading performance of Chinese open-source models can be translated into a stable advantage in real-world scenarios. Potential risks are concentrated in over-reliance on insufficiently validated agent systems, a lack of transparency in billing and quotas that erodes user trust, and the further concentration of talent and technical resources in a few leading institutions, which amplifies industry competition and governance pressures.
💡 Influencer Insights Link to heading
AI Daily Industry Dynamics Analysis Link to heading
Date: June 18-19, 2026 (Synthesizing recent trends)
1. Today’s Focus: The Rise of On-device Models and Their Engineering Practices Link to heading
Over the past 24 hours, discussions among AI leaders have strongly pointed to a core trend: High-performance on-device models are transitioning from “usable” to “great to use” and are beginning to reshape the workflows of developers and power users.
1.1 On-device Model Capabilities Validated, with Breakthroughs in Both Performance and Quality Link to heading
- @zhixianio expressed being “very satisfied” with the full-duplex audio-visual performance of MiniCPM-o 4.5, marveling, “It’s hard to imagine a 9B model can achieve this.” Although stability issues persist during long runs, its quality is already initially viable.
- @zhixianio dubbed Qwen3.6-35B-A3B MoE (oMLX) the “sweet spot 🍮 champion,” noting its response speed in Personal Assistant (PA) and Coding scenarios surpasses remote LLMs, and its native multimodal user experience is “even better than DSV4 Pro.”
- The Google Gemma 4 family has become a focus for on-device applications:
- Gemma 4 E4B + MTP: @zhixianio’s tests confirmed its excellent performance on Japanese email parsing and classification tasks.
- Gemma 4 12B Coder: @zhixianio conducted rigorous comparative tests on code generation. The conclusion was that in a head-to-head with Qwen3.6-35B-A3B MoE, Gemma 12B Coder hit a clear “ceiling” in generating complex, stateful programs (like Tetris), as its 12B parameters struggled to support long-form complex logic.
- Gemma 4 QAT Quantization-Aware Model: @zhixianio specifically highlighted Google’s Quantization-Aware Training (QAT) approach, considering it a key direction for on-device optimization, adding, “Android will soon be able to use its own native models.”
1.2 AI Coding Tools Enter a “Clash of Titans” and a New Stage of Automation Link to heading
- Claude Code vs. OpenAI Codex: @ruanyf’s question, “Do you use Codex or Claude Code?” sparked a discussion. @vista8 stated that Codex is the superior product, but specific scenarios still require Claude Code. They use a self-developed MCP to enable synergy between the two, even achieving a “double Codex quota” hack. @gefei55 also shared an open-source project that uses an MCP to give the ChatGPT web version local code manipulation capabilities via Codex, which is also essentially aimed at “doubling the quota” and calling the most powerful model.
- Claude Code launches Artifacts for visual collaboration: @dotey provided a detailed interpretation of this feature, stating that it solves the collaboration problem where AI programming results are “only visible to the operator.” It allows work products from debugging, PR reviews, and architectural explanations to be shared with the team as real-time web pages.
- OpenAI Codex’s “Record & Replay” feature: @dotey and @AI_Jasonyu highly praised this feature, viewing it as essentially a “super version of RPA + keyboard macro + Computer Use combined.” Users only need to demonstrate a workflow once, and Codex can automatically generate a reusable Skill, significantly lowering the barrier to automation.
2. Noteworthy Unique Perspectives and Industry Foresight Link to heading
2.1 The Technological Paradigm Shift from “Correlation” to “Causality” Link to heading
- @Pluvio9yte offered a deep analysis of Aether AI, founded by Professor Biwei Huang, and pointed out that its “Causal World Models” indicate a key direction for the next stage. He argues that current large models fail when generating images like “pouring water into a cup with a hole in the bottom” because they learn the data correlation between “pouring water” and a “full cup,” not the physical causal mechanism that “water leaks from a hole.” Aether AI aims to enable AI to understand environmental variables and intervention outcomes, which is crucial for fields demanding rigorous logic, such as embodied intelligence and new materials R&D.
2.2 The Evolution of Vibe Coding: From “Requirements-First” to “Contract-First” Link to heading
- @Pluvio9yte shared his in-depth thoughts on transitioning from a security professional to a full-stack developer. He argues that the best practice for “Vibe Coding” is not Requirement First or Code First, but Contract First. By combining this experience with the open-source project OpenSpec, he has developed a framework that “externalizes easily shifting context into contracts, providing a stable reference for both humans and AI.” The goal is to enable less experienced developers to systematically undertake large-scale project development.
2.3 Personal Development Philosophy in the AI Era Link to heading
- @ruanyf shared the article “Can I Take a Day Off Today?”, questioning where the benefits for employees lie after AI significantly boosts the productivity of white-collar workers. He predicts that AI will increase the average salary or welfare of the entire society as a long-term trend.
- @lijigang, drawing from the philosophy that “Bodhisattvas fear the cause, while ordinary people fear the effect,” points out that an individual’s three core views (view of life, world, and values) are the fundamental ‘function f’ for handling life’s events. Designing this function well is superior to praying for a single outcome (f(x)). He also reminds us that “Token consumption is a ‘vanity metric,’ while the effectiveness in solving problems is the ’true metric.’”
- @gefei55 shared an AI-assisted business intelligence strategy for gaining a time advantage: using the X API to filter recent, high-engagement tweets with links to discover new terms and products before they signal on Google Trends. This achieves a state where “while others are waiting for the Trends curve, you’re already ahead of the game.” He has also open-sourced this method.
3. Recommended Tools & Resources Link to heading
3.1 Development & Design Tools Link to heading
- baoyu-design Skill (by @dotey): A powerful local design Skill that supports generating animated videos and exporting them as MP4s. It can also generate a PPTX with images in one click, which can be further edited. Its animation engine uses a declarative design
f(t)and supports frame-accurate exporting. The GitHub address has been released. - Meta Skill 2.0 (by @yaojingang, recommended by @vista8): A “meta-Skill” created by integrating the leaked Claude Code source code from Anthropic. It’s used for creating high-quality Skills and was called “the best Meta Skill I’ve ever used” by @vista8.
- Figma Chrome Plugin (recommended by @vista8): Can convert any web page element into an editable layer in Figma, which is extremely useful for designers replicating websites and taking precise screenshots.
- Fable (recommended by @zhixianio): An AI-driven development assistant that can complete 70% of a demo in 40 minutes and optimize the original plan. It was reviewed as “Shut up and take my money.”
3.2 Productivity & Automation Link to heading
- YouMind 1.0 (by @lifesinger, recommended by @AI_Jasonyu, @gefei55): A content creation tool officially released after two years of development. Its highlights are high-quality article output and excellent formatting compatibility with platforms like X and WeChat Official Accounts. Combined with its image-adding feature, it offers a one-stop solution for the pain points of long-form content creation.
- Local Video Translation Tool (by @xiaohu, recommended by @Pluvio9yte): An open-source, all-in-one video translation tool that automates the entire process: downloading, transcription, translation, polishing, and burning subtitles.
- EvoMap (recommended by @vista8): A platform event where Stars on GitHub open-source projects can be exchanged for large model API Tokens. It encourages developers to package workflows and Prompts and upload them to earn more rewards.
- Giffgaff UK SIM Card Guide (by @AI_Jasonyu): A detailed tutorial for solving the phone number verification challenges for overseas AI services (like Codex, Claude Code).
3.3 Cutting-Edge Models & Frameworks Link to heading
- oMLX v0.4.0 (released by @jundotkim, shared by @zhixianio): The first official version to support native Swift macOS apps, making it a key tool for smoothly running MLX models.
- PP-OCRv6 (by @AI_Jasonyu): An ultra-lightweight OCR model released by Baidu. At only 1.5MB, it can run in a browser, process a single image in as fast as 97ms, and its character-by-character recognition accuracy surpasses large models like GPT-5.5, making it highly suitable for on-device integration.
- OpenClaw (mentioned by @zhixianio): A personal assistant framework that integrates local models. Paired with on-device models (like Qwen), it enables fast and intelligent native multi-modal interactions.
📚 Appendix: Today’s Watch List Source Updates Link to heading
Time window: Last 3 days; 22 sources covered; 34 updates in total
Y Combinator Podcast (B_intro+search) Link to heading
- Why Domain Experts Are Winning In The Age Of AI
- Release Time: 2026-06-19 23:00 Beijing Time
- Abstract: - You may have heard of OpenClaw (formerly known as Clawdbot/Moltbot).
- The breakout open-source AI assistant that runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually perform tasks like managing email, calendars, files, workflows, and more.
- Now, meet the person behind it.
- YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to discuss the “aha” moment behind the viral personal AI agent, why a local-first agent could replace many of today’s apps, and how personal agents will reshape the future of software.
- EN Highlights:
- Bryant Chou co-founded Webflow, which today powers around 1% of all websites on the internet
- Now he’s back in the current YC batch with Ploy, an AI-powered website and marketing platform that doesn’t just build your site — it connects to your analytics,…
- In this episode of the Lightcone he explains how he built Ploy to be “anti slop,” how building today compares to his first startup, and why founders with domain…
- Now26:01 — First Three Months: Webflow 2013 vs
All-In Podcast (A_full) Link to heading
- World’s First Trillionaire, Anthropic Fable Banned, The New Oligarchs, Iran Peace Deal
- Published: 2026-06-20 06:07 Beijing Time
- Summary: - (0:00) Bestie intros.
- (2:41) The New Oligarchs, America’s incoming politburo, and learned helplessness.
- (14:18) SpaceX’s record-breaking IPO, $60B Cursor acquisition, and trillionaire reactions.
- (33:34) Behind the scenes of Anthropic’s Fable ban.
- EN Highlights:
- (0:00) Bestie intros
- (2:41) The New Oligarchs, America’s incoming politburo, and learned helplessness
- (14:18) SpaceX’s record breaking IPO, $60B Cursor acquisition, and trillionaire reactions
- (33:34) Behind the scenes of Anthropic’s Fable ban
Stratechery by Ben Thompson (A_full) Link to heading
- 2026.25: The Stuff of Myth(os)
- Published: 2026-06-20 01:00 Beijing Time
- Summary: - (Photo by Ronald Cortes/Getty Images).
- Welcome back to This Week in Stratechery!
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone.
- Additionally, you have complete control over what we send to you.
- With that said, here are some of our favorites from the week.
- EN Highlights:
- (Photo by Ronald Cortes/Getty Images)
- Welcome back to This Week in Stratechery
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
- Additionally, you have complete control over what we send to you
Two Minute Papers (B_intro+search) Link to heading
- Scientists Found A Better Language For AI Agents
- Publication Time: 2026-06-19 22:06 Beijing Time
- Summary: - ❤️ Check out Weights & Biases and sign up for a free demo here:.
- 📝 The paper is available here:.
- Adam Bridges, Benji Rabhan, B Shang, Cameron Navor, Charles Ian Norman Venn, Christian Ahlin, Eric T, Fred R, Gordon Child, Juan Benet, Michael Tedder, Owen Skarpness, Richard Sundvall, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
- Scientists have found a better language for AI agents.
- EN Key Points:
- ❤️ Check out Weights & Biases and sign up for a free demo here:
- 📝 The paper is available here:
- Brain reading video:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
ArXiv cs.AI (B_intro+search) Link to heading
Deontic Policies for Runtime Governance of Agentic AI Systems
- Publication Time: 2026-06-19 12:00 Beijing Time
- Summary: - arXiv:2606.19464v1 Announcement Type: New.
- Abstract: Autonomous agentic AI systems driven by Large Language Models (LLMs) introduce a new class of security, privacy, and compliance challenges: agents that can invoke tools, operate on data, install software, and coordinate with peer agents across organizational boundaries must be subject not only to authentication and access controls but also to the full structure of enterprise governance.
- This includes specifying what agents are permitted and prohibited from doing, what they are obliged to do after taking certain actions (e.g., notify the CISO), under what conditions long-term obligations can be waived, and which rules take precedence when policies conflict.
- This governance problem exceeds what current policy engines can provide.
- EN Key Points:
- arXiv:2606.19464v1 Announce Type: new
- Abstract: Autonomous agentic AI systems driven by Large Language Models (LLMs) introduce a new class of security, privacy, and compliance challenges: an agent t…
- This includes specifying what agents are permitted and prohibited from doing, what they areobliged to do after certain actions (e.g., notify the CISO), under wh…
- This governance problem exceeds what current policy engines provide
- Publication Time: 2026-06-19 12:00 Beijing Time
- Summary: - arXiv:2606.19469v1 Announcement Type: New.
- Abstract: Undergraduate computer science is governed by international curriculum guidelines that are revised approximately every ten years, yet curricula lack reliable, repeatable methods to measure the extent to which they cover the current guidelines and how that coverage changes when the guidelines are restructured.
We address this with a human-in-the-loop pipeline that measures a program’s coverage of an external body of knowledge, applied longitudinally to accredited computer science bachelor’s degrees according to the 2013 (CS2013) and 2023 (CS2023) computer science curricula.
The pipeline represents the program and each guideline as structured corpora, generates candidate course-to-knowledge-unit matches via semantic retrieval, and confirms them through human judgment under clear coverage definitions.
- EN Highlights:
- arXiv:2606.19469v1 Announce Type: new
- Abstract: Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable, reproduci…
- We address this with a human-in-the-loop pipeline that measures a program’s coverage of an external body of knowledge, applied longitudinally to one accredited…
- The pipeline represents the program and each guideline as structured corpora, generates candidate course-to-knowledge-unit matches by semantic retrieval, and co…
- EN Highlights:
Diffusion Language Models: An Experimental Analysis
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19475v1 Announce Type: new.
- Abstract: Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range of tasks.
- Recently, Diffusion Language Models (DLMs) have emerged as an alternative paradigm, generating text through iterative denoising instead of next-token prediction, thereby allowing parallel refinement of entire sequences.
- While numerous diffusion-based architectures have been proposed, differences in evaluation protocols, datasets, inference budgets, and generation hyperparameters make it difficult to compare their functionalities and understand the trade-offs they offer.
- EN Highlights:
- arXiv:2606.19475v1 Announce Type: new
- Abstract: Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a wide range…
- Recently, Diffusion Language Models (DLMs) have emerged as an alternative paradigm that generates text through iterative denoising rather than next-token predic…
- While numerous diffusion-based architectures have been proposed, differences in evaluation protocols, datasets, inference budgets, and generation hyperparameter…
Hidden Anchors in Multi-Agent LLM Deliberation
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19494v1 Announce Type: new.
- Abstract: Multi-agent LLM deliberation (where agents exchange and revise answers over multiple rounds) is increasingly used to improve reasoning and accuracy, but how and why it works is rarely modeled.
- This deliberation mirrors how humans make decisions.
- As social animals, we are pulled both by the group—the herd mentality captured by classical models of opinion dynamics like those of DeGroot and Friedkin-Johnsen—and by our own intrinsic beliefs, which they do not capture.
- EN Highlights:
arXiv:2606.19494v1 Announce Type: new
- Abstract: Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increasingly used to improve reasoning and accuracy, ye…
- Such deliberation mirrors how humans reach decisions
- As social animals we are pulled both by the group, the herd effect that classical opinion-dynamics models such as DeGroot and Friedkin–Johnsen capture, and by…
DeXposure-Claw: An Agentic System for DeFi Risk Supervision
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19501v1 Announce Type: new.
- Abstract: Decentralized finance exposes supervisors to fast-moving, networked credit risks.
- General-purpose LLM agents are ill-suited for this environment: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations do not provide regulators with a consistent method to measure the resulting false positives.
- We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that guides LLM decisions through structured evidence: (1) DeXposure-FM, a graph time-series foundation model that predicts future exposure networks; (2) Deterministic monitors and stress scenarios then translate these predictions into type alerts, attribution signals, and scenario evidence; (3) Data health and trust gates limit escalation before DeXposure-Claw issues auditable regulatory tickets with justifications.
- EN 要点:
- arXiv:2606.19501v1 Announce Type: new
- Abstract: Decentralized finance exposes supervisors to fast-moving, networked credit risks
- General-purpose LLM agents fit this setting poorly: they over-read weak evidence and recommend high-stakes interventions, while existing evaluations offer no re…
- We introduce DeXposure-Claw, a forecast-grounded agentic supervision system that routes LLM decisions through structured evidence: (1) DeXposure-FM, a graph tim…
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19509v1 Announce Type: new.
- Abstract: Large Language Models (LLMs) are increasingly applied to structured clinical data, but whether they can recognize the limitations of their knowledge in such tasks remains to be explored.
- We investigate this issue through the lens of cross-model attribution divergence, aiming to reduce cognitive uncertainty in structured tasks by comparing Qwen 2.5 7B and XGBoost on prediction tasks via attribution divergence analysis.
- Firstly, LLM linguistic confidence is epistemically vacuous, outputting near-constant values (0.856-0.937) regardless of whether accuracy is 49% or 75.3%, tracking prompt format rather than prediction quality.
- EN 要点:
- arXiv:2606.19509v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they can recognize the limits of their own knowledge on…
We study this question through the lens of cross-model attribution divergence with the goal of reducing epistemic uncertainty for structured tasks, comparing Qw…
We report four findings
- Publish Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19522v1 Announce Type: new.
- Abstract: The retina offers a non-invasive window into neurodegenerative diseases, capturing subtle structural patterns associated with the risk of future cognitive decline.
- Vision-language alignment frameworks such as REVEAL have shown that pairing retinal fundus images with structured clinical risk narratives can improve early prediction of Alzheimer’s Disease (AD).
- A key design choice in these methods is the use of phenotypic grouping, where individuals with similar risk profiles are treated as multi-positive pairs during contrastive learning.
- EN Highlights:
- arXiv:2606.19522v1 Announce Type: new
- Abstract: The retina offers a noninvasive window into neurodegenerative disease, capturing subtle structural patterns associated with a risk of future cognitive…
- Vision-language alignment frameworks such as REVEAL have shown that pairing retinal fundus images with structured clinical risk narratives improves early predic…
- A key design choice in these approaches is the use of phenotypic grouping, where individuals with similar risk profiles are treated as multi-positive pairs duri…
- Publish Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19527v1 Announce Type: new.
- Abstract: Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics?
- We endow an LLM with a conscience step that reviews its own reasoning and outputs, and we extend the training loss with an alignment component using Direct Preference Optimization (DPO) to guide the model away from unethical outputs.
- The result is an online technique that can tune models across a wide range of applications: training, fine-tuning, adversarial prompting, and zero-shot learning.
- EN Highlights:
- arXiv:2606.19527v1 Announce Type: new
- Abstract: Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics
- And can they self-correct
- We endow an LLM with a conscience step that reviews its own reasoning and outputs, and we extend the training loss with an alignment component using Direct Pref…
ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19538v1 Announcement Type: new.
- Abstract: Convolutional networks, recurrent networks, and transformers each encode different inductive biases—locality, sequential memory, and content-related pairwise interactions—and have remained mathematically distinct since their inception.
- We show that this fragmentation reflects not a fundamental diversity in how signals should be processed, but rather incomplete views of a single underlying mathematical object: the learnable integral transform.
- We introduce the Integral Transform Network (ITNet), a unified architecture built around a learnable kernel that jointly depends on position and features.
- EN Highlights:
- arXiv:2606.19538v1 Announce Type: new
- Abstract: Convolutional networks, recurrent networks, and transformers each encode different inductive biases – locality, sequential memory, and content-depend…
- We show that this fragmentation reflects not a fundamental diversity in how signals should be processed, but rather incomplete views of a single underlying math…
- We introduce the Integral Transform Network (ITNet), a unified architecture built around a learnable kernel that depends jointly on positions and features
Uncertainty Decomposition for Clarification Seeking in LLM Agents
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19559v1 Announcement Type: new.
- Abstract: Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) agents and highlight the lack of canonical, decomposable, and communicable representations of uncertainty that could unlock new agent capabilities, such as actively seeking clarification and building shared mental models.
- Practical deployment constraints—black-box APIs, interactive latency budgets, and the absence of labeled trajectories—rule out log-probability-based, multi-sampling, and training-based methods, making prompt-based estimation the most viable family of approaches for surfacing such signals at deployment time.
- We answer this call with a simple prompt-based decomposition that separates action confidence from request uncertainty (u), enabling the agent to ask for clarification when task specifications are ambiguous.
- EN Highlights:
- arXiv:2606.19559v1 Announce Type: new
- Abstract: Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language model (LLM) a…
- Practical deployment constraints – black-box APIs, interactive latency budgets, and the absence of labeled trajectories – rule out logprob-based, multi-sampli…
- We answer this call with a simple prompt-based decomposition that separates action confidence from request uncertainty (u), enabling the agent to ask for clarif…
ArXiv cs.CL (B_intro+search) Link to heading
Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19344v1 Announcement Type: new.
- Abstract: Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generation.
- Standard auditing methods rely on a single output inspection or static automated metrics.
- These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability generation branches.
- EN Key Points:
- arXiv:2606.19344v1 Announce Type: new
- Abstract: Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of text generation.
- Standard auditing methods rely on a single output inspection or static automated metrics
- These approaches obscure the underlying probability distributions and fail to capture biases hidden in lower-probability generation branches
Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19345v1 Announcement Type: new.
- Abstract: The rapid increase in scientific publications has led to manual study screening in Systematic Literature Reviews (SLRs) becoming increasingly resource-intensive, inefficient, and inconsistent.
- Classifying studies that clearly report health-related quality-of-life outcomes (e.g., EQ-5D data) requires a high level of clinical interpretation, posing challenges for human reviewers.
- This study investigates the use of Google’s Gemini and Gemma Large Language Models (LLMs) for automating EQ-5D detection in the PubMed biomedical database based solely on published abstracts.
- EN Key Points:
- arXiv:2606.19345v1 Announce Type: new
- Abstract: The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is increasingly resource-intensive.
- Classifying studies that clearly report health-related quality-of-life results, such as EQ-5D data, requires a high level of clinical interpretation and poses challenges for human reviewers.
- This study investigates the use of Google’s Gemini and Gemma large language models (LLMs) in automating EQ-5D detection in the PubMed biomedical database based solely on published abstracts.
Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19346v1 Announcement Type: new.
- Abstract: We investigate cross-lingual transfer by fine-tuning seven Large Language Models (4B–671B parameters) in Arabic and evaluating zero-shot reading comprehension for Semitic languages and non-Semitic controls.
In dense and Mixture-of-Experts architectures, we find no evidence of Semitic-specific transfer: models with weak baselines improve significantly across all languages, while strong baseline models show only marginal gains, regardless of language family.
- A chain-of-thought ablation reinforces this finding — the same models that benefit most from fine-tuning benefit equally from inference-time reasoning, suggesting that both mechanisms address task format alignment rather than cross-lingual knowledge transfer.
- EN Key Points:
- arXiv:2606.19346v1 Announce Type: new
- Abstract: We study cross-lingual transfer by fine-tuning seven large language models (4B–671B parameters) on Arabic and evaluating zero-shot reading comprehens…
- Across dense and Mixture-of-Experts architectures, we find no evidence of Semitic-specific transfer: models with weak baselines improve dramatically across all…
- A chain-of-thought ablation reinforces this finding – the same models that benefit most from fine-tuning benefit equally from inference-time reasoning, suggest…
How LLMs Fail and Generalize in RTL Coding for Hardware Design?
- Published: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19347v1 Announce Type: new.
- Abstract: Translating sequential programming priors into the parallel temporal logic of hardware design remains a crucial bottleneck for Large Language Models (LLMs).
- To investigate this, we introduce a new error taxonomy grounded in problem solvability, inspired by cognitive theory.
- Our taxonomy categorizes failures into syntactic, semantic, solvable functional, and unsolvable functional types.
- EN Key Points:
- arXiv:2606.19347v1 Announce Type: new
- Abstract: Translating sequential programming priors into the parallel temporal logic of hardware design remains a crucial bottleneck for large language models(L…
- To investigate this, we introduce a new error taxonomy grounded in problem solvability, inspired by cognitive theory
- Our taxonomy categorizes failures into syntactic, semantic, solvable functional, and unsolvable functional types
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Published: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19348v1 Announce Type: new.
- Abstract: We present a preview of the DeepSeek-V4 series, featuring two powerful Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B active) and DeepSeek-V4-Flash with 284B parameters (13B active) — both supporting a 1-million-token context length.
- The DeepSeek-V4 series incorporates several key architectural and optimization upgrades: (1) a hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavy Compressed Attention (HCA) for improved long-context efficiency; (2) manifold-constrained hyper-connections (mHC) to enhance traditional residual connections; and (3) the Muon optimizer for faster convergence and higher training stability.
We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities.
- EN Highlights:
- arXiv:2606.19348v1 Announce Type: new
- Abstract: We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models – DeepSeek-V4-Pro with 1.6T paramet…
- DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attent…
- We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances…
- EN Highlights:
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19349v1 Announce Type: new.
- Abstract: While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (dLLMs) remains largely unexplored.
- Unlike AR models restricted by unidirectional causal masking, dLLMs intrinsically utilize bidirectional attention, offering extensive spatial flexibility for query placement.
- Unfortunately, current practices conventionally inherit AR-style trailing-query templates, often overlooking the structural paradigm shift.
- EN Highlights:
- arXiv:2606.19349v1 Announce Type: new
- Abstract: While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (dLLMs) remai…
- Unlike AR models restricted by unidirectional causal masking, dLLMs intrinsically utilize bidirectional attention, offering extensive spatial flexibility for qu…
- Unfortunately, current practices conventionally inherit AR-style trailing-query templates, often overlooking the structural paradigm shift
Pruning via Causal Attribution Preserves Reasoning Performance in Large Language Models
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19350v1 Announce Type: new.
- Abstract: Large Language Models (LLMs) excel at multi-step reasoning but incur significant inference costs.
- We introduce Causal Attribution Pruning (CAP), a training-free method that identifies critical attention heads by measuring their causal impact on reasoning tasks, and uses these head-level scores to guide fine-grained weight pruning.
- For each attention head, CAP estimates the expected performance drop when that head is masked during a forward pass on a small set of reasoning problems.
- EN Highlights:
- arXiv:2606.19350v1 Announce Type: new
Abstract: Large language models (LLMs) excel at multi-step reasoning but incur substantial inference cost
We introduce Causal Attribution Pruning (CAP), a training-free method that identifies critical attention heads by measuring their causal impact on reasoning tas…
For each attention head, CAP estimates the expected performance degradation when the head is masked during forward passes on a small calibration set of reasonin…
Detecting Hallucinations for Large Language Model-based Knowledge Graph Reasoning
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19351v1 Announcement Type: new.
- Abstract: Knowledge graph (KG) reasoning infers new knowledge from existing facts and is widely applied in question answering, recommendation, and decision support.
- With the rapid development of large language models (LLMs), LLM-based KG reasoning frameworks have become increasingly popular by leveraging retrieved KG information.
- However, hallucinations in LLMs remain a critical issue.
- EN Key Points:
- arXiv:2606.19351v1 Announce Type: new
- Abstract: Knowledge graph (KG) reasoning infers new knowledge from existing facts and is widely applied in question answering, recommendation, and decision supp…
- With the rapid development of large language models (LLMs), LLM-based KG reasoning frameworks have become increasingly popular by leveraging retrieved KG inform…
- However, hallucinations in LLMs remain a critical issue
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19352v1 Announcement Type: new.
- Abstract: Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities.
- Despite substantial progress in sign-language recognition, translation, and production, advances remain constrained by fragmented datasets, inconsistent annotations, and limited language coverage.
- Existing benchmarks often fail to reflect real-world communication needs, and a systematic analysis of these limitations remains limited.
- EN Key Points:
- arXiv:2606.19352v1 Announce Type: new
- Abstract: Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities
- Despite substantial progress in sign-language recognition, translation, and production, advances remain constrained by fragmented datasets, inconsistent annotat…
Existing benchmarks often fail to reflect real-world communication needs, and systematic analyses of these limitations remain limited
- Published: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19353v1 Announce Type: new.
- Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model’s ability to understand context, blurring whether failures are caused by data properties or model limitations.
- Uncertainty decomposition (separating aleatoric from epistemic sources) is particularly crucial in this setting, yet existing methods designed for standard generative tasks cannot capture the unique dynamics of ICL.
- To address this, we introduce a concept of self-function vectors, built upon Bayesian views and the mechanistic interpretability of ICL.
- EN Key Points:
- arXiv:2606.19353v1 Announce Type: new
- Abstract: In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly s…
- Uncertainty decomposition-separating aleatoric from epistemic sources-is particularly crucial in this setting, yet existing methods, designed for standard gener…
- To address this, we introduce a concept of self-function vectors, built upon Bayesian views and the mechanistic interpretability of ICL
ArXiv cs.LG (B_intro+search) Link to heading
- Published: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19361v1 Announce Type: new.
- Abstract: Identification conditions describe the computability of a target query or parameter of interest as a function of the type and amount of available information.
- In causal identification, this information is often expressed in the form of a causal graph, and data are observed or collected for some subset of variables in the graph.
- Target queries may be for a single effect alone or for a class of effects in a given model.
- EN Key Points:
- arXiv:2606.19361v1 Announce Type: new
- Abstract: Identification conditions describe the computability of a target query or parameter of interest as a function of the type and amount of information av…
- In causal identification, this information is often expressed in the form of a causal graph, and data are observed or collected for some subset of variables in…
- Target queries may be for a single effect alone or for a class of effects in a given model
- Published: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19363v1 Announcement Type: new.
- Abstract: The deployment of Time-Series Foundation Models (TSFMs) in physical sciences is hindered by a critical trade-off: while these models encode rich, universal temporal dynamics, they suffer from severe distributional misalignment when applied zero-shot to specific scientific domains, and their computational cost prevents deployment in edge computing sensor networks.
- We address a fundamental challenge: How can we extract latent structural knowledge from misaligned foundation models (FMs) to train lightweight, specialized forecasters?
- We propose Gated Uncertainty-Aware Routing for Distillation (Guard), a novel framework that reframes multi-teacher distillation as an instance-wise decision process with two adaptive mechanisms: (1) a contextual router that dynamically selects the most relevant teacher based on local input statistics, leveraging the complementarity between different foundation models; and (2) an uncertainty-gated temperature mechanism that acts as a “circuit breaker,” automatically weakening the distillation intensity when teacher confidence deviates from domain reality.
- EN Key Points:
- arXiv:2606.19363v1 Announce Type: new
- Abstract: The deployment of Time-Series Foundation Models (TSFMs) in physical sciences is hindered by a critical trade-off: while these models encode rich, univ…
- We address a fundamental challenge: How can we extract latent structural knowledge from misaligned foundation models (FM) to train lightweight, specialized fore…
- We propose Gated Uncertainty-Aware Routing for Distillation (Guard), a novel framework that reframes multiteacher distillation as an instance-wise decision proc…
Closing the Social-Semantic Gap: SPSD for Edge-Based Prompt Compression in Cloud LLM Inference
- Published: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19364v1 Announcement Type: new.
- Abstract: The prefill stage of Large Language Model (LLM) inference is a growing contributor to cloud-scale energy costs.
- Many consumer-support and conversational prompts contain social scaffolding: politeness markers, apologetic preambles, repetition, and rapport-building language that are important for human communication but contain little marginal information for machine reasoning.
- We call this discrepancy the Social-Semantic Gap.
- EN Key Points:
- arXiv:2606.19364v1 Announce Type: new
- Abstract: The prefill stage of Large Language Model (LLM) inference is a growing contributor to cloud-scale energy cost
- Many consumer-support and conversational prompts contain social scaffolding: politeness markers, apologetic preamble, repetition, and rapport-building language…
- We call this discrepancy the Social-Semantic Gap
Performance Analysis and Optimization of 3D Generative Diffusion Models across GPU Architectures
Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19365v1 Announcement Type: new.
- Abstract: Diffusion models are crucial for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by the substantial GPU resource demands arising from hundreds of U-Net evaluations per sample and highly heterogeneous kernel behavior.
- This paper performs a comprehensive performance analysis of the state-of-the-art medical diffusion model, Med-DDPM, across three generations of NVIDIA architectures to study kernel-level runtime failures, instruction mix characteristics, memory system utilization, warp-level activity, and profiler priority score estimation.
- We show that training is overwhelmingly dominated by cuDNN convolution and implicit-GEMM kernels, with inefficiencies arising from memory-access patterns, tensor layout transformations, and limited tensor core utilization.
- EN Key Points:
- arXiv:2606.19365v1 Announce Type: new
- Abstract: Diffusion models have become essential for high-fidelity 3D MRI synthesis, yet their deployment remains constrained by substantial GPU resource demand…
- This paper performs a comprehensive performance analysis of the state-of-the-art medical diffusion model, Med-DDPM, across three generations of NVIDIA architect…
- We show that training is overwhelmingly dominated by cuDNN convolution and implicit-GEMM kernels, with inefficiencies arising from memory-access patterns, tenso…
- Abstract: - arXiv:2606.19365v1 Announcement Type: new.
Information Lattice Learning as Probabilistic Graphical Model Structure Learning
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19366v1 Announcement Type: new.
- Abstract: Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice that encodes a hierarchy of abstractions and lifting selected rules back to the signal domain.
- When the signal is a probability mass function, we show the probabilistic rules learned by ILL admit a natural probabilistic graphical model (PGM) interpretation and develop this interpretation in detail.
- A partition in ILL induces a deterministic quotient variable, and a rule is the marginal law of that quotient variable.
- EN Key Points:
- arXiv:2606.19366v1 Announce Type: new
- Abstract: Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice that encodes a…
- When the signal is a probability mass function, we show the probabilistic rules learned by ILL admit a natural probabilistic graphical model (PGM) interpretatio…
- A partition in ILL induces a deterministic quotient variable, and a rule is the marginal law of that quotient variable
Weibull Weight-Scale Parameter Evolution under AdamW Training Dynamics
- Publication Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19367v1 Announcement Type: new.
Abstract: Building on a two-parameter Weibull framework for diagnosing transformer weight distributions, we study why the Weibull weight-scale parameter $\lambda$ grows, overshoots, and then relaxes during AdamW training.
We derive a leading-order three-force decomposition of the squared weight norm from the AdamW update: an alignment force measuring the correlation between weights and the adaptive update direction, an injection force from the adaptive step size, and a decay force from decoupled weight decay.
On self-trained Pythia-70M models with ground-truth optimizer moments, alignment dominates the rise phase, contributing 88-94% of the absolute force budget across four random seeds, and is robust to overweight removal.
EN Highlights:
- arXiv:2606.19367v1 Announce Type: new
- Abstract: Building on a two-parameter Weibull framework for diagnosing transformer weight distributions, we study why the Weibull weight-scale parameter $\lambd…
- We derive a leading-order three-force decomposition of the squared weight norm from the AdamW update: an alignment force measuring the correlation between weigh…
- On self-trained Pythia-70M models with ground-truth optimizer moments, alignment dominates the rise phase, contributing 88-94% of the absolute force budget acro…
- Release Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19369v1 Announce Type: new.
- Abstract: Estimation-of-distribution algorithms (EDAs) are a powerful class of evolutionary methods for black-box optimization, especially when little is known about the objective’s structure.
- Whereas classical evolutionary algorithms rely on hand-designed mutation and crossover operators, which are hard to devise for unknown problem structures and a source of bias, EDAs completely sidestep operator design: they fit a probability distribution to the best individuals and sample the next generation from it.
- EDAs are well established on continuous parameter spaces, but they have not previously been generalized to sparse ones, in which most coefficients of a good solution are exactly zero.
- EN Highlights:
- arXiv:2606.19369v1 Announce Type: new
- Abstract: Estimation-of-distribution algorithms (EDAs) are a powerful class of evolutionary methods for black-box optimization, especially when little is known…
- Whereas classical evolutionary algorithms rely on hand-designed mutation and crossover operators, hard to devise for unknown problem structures, and a source of…
- EDAs are well established on continuous parameter spaces, but they have not previously been generalized to sparse ones, in which most coefficients of a good sol…
Human-like autonomy emerges from self-play and a pinch of human data
- Release Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19370v1 Announce Type: new.
Abstract: Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data.
It uses cheap, large-scale simulations to substitute expensive, large-scale human driving demonstrations.
A key limitation of this approach is that policies trained through pure self-play can learn effective but alien driving conventions incompatible with people.
- Release Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19371v1 Announce Type: new.
- Abstract: Alzheimer’s disease (AD) is a fatal disorder that destroys memory and cognitive skills in the elderly population.
- Most treatments for AD are effective in the early stage, leading to an increasing demand for early AD diagnosis.
- AD diagnosis increasingly relies on multimodal data such as clinical assessments, structural Magnetic Resonance Imaging (MRI), and Positron Emission Tomography (PET) imaging.
cAPM: Continual AI-Assisted Pace-Mapping with Active Learning
- Release Time: 2026-06-19 12:00 Beijing Time
- Abstract: - arXiv:2606.19373v1 Announce Type: new.
- Abstract: Ventricular tachycardia is a life-threatening rhythm disorder and a major cause of sudden cardiac death.
- Pace-mapping is a clinical procedure for identifying the intervention target during catheter ablation of VT.
- It requires clinicians to pace different sites of the ventricle and quickly interpret the resulting electrocardiograms to determine where to pace next or if a target site has been identified.
It requires clinicians to pace different sites in the ventricles and rapidly interpret the resulting electrocardiograms to determine where to pace next or wheth…