System translated (Gemini)

🤖 AI 速览

Today’s core signals come from two ends: GPT-5 Pro is entering the immunology experiment judgment phase, indicating that AI for Science is transitioning from data assistance to scientific reasoning. Concurrently, OpenAI is promoting advanced AI sharing standards, leading to more …
📋 文章元数据
发布时间
2026-06-24
类型
ai-daily
字数
7494
阅读时长
36 min

2026-06-24 AI Daily | GPT-5 Cracks Immunology Puzzle, AI Evaluation and On-Device Models Heat Up Link to heading

Today’s key signals come from two ends: GPT-5 Pro is entering the experimental judgment phase in immunology, indicating that AI for Science is shifting from data assistance to scientific reasoning; OpenAI is promoting standards for sharing advanced AI, making industry governance and capability assessment more institutionalized. Meanwhile, on-device models, AI programming tools, and enterprise collaboration assistants continue to accelerate their adoption, with cost, controllability, and real-world workflow value becoming the focus of competition.

📖 In-depth Guide to This Issue’s Watch List Link to heading

Today, three areas deserve the most attention: First, GPT-5 helping immunologists solve a three-year-old puzzle is a prime example of “AI for Science” transitioning from auxiliary search to experimental judgment, which is recommended reading for research and medical teams. Second, OpenAI’s push for standards for sharing advanced AI, along with papers like VeriBound and FirstPass, indicates the industry is incorporating model capability assessment, peer review, and process-reward models into a more serious governance framework. Third, research on on-device RAG compression, language steering, post-training differences in multi-agent systems, and emotional TTS reminds us that competition among models isn’t just about parameter scale, but also about controllability, cost, and interaction details. Other topics to watch include memory chips and Chinese models, as AI infrastructure and geopolitical supply chains remain long-term variables.

🌐 AI Hot Topics on X Link to heading

Topic 1: Z.ai’s GLM-5.2 Tops Open AI Leaderboards, Matches Closed Models at Lower Cost Link to heading

  • Category: AI · News
  • Overview: Trending for: 1 day ago, Related posts: 11,000
  • What it is: Z.ai’s released GLM-5.2 has ranked high on several public AI benchmark leaderboards and is said to approach or match the performance of some closed-source models at a lower cost.
  • Why it matters: This indicates that open-source or open-weight models continue to close the gap with cutting-edge closed-source models in capability and cost-efficiency, potentially accelerating enterprise adoption, model commoditization, and a shift in the global AI competitive landscape.
  • Discussion summary: Discussions on X focus on whether GLM-5.2’s leaderboard scores reflect real-world application capabilities, the sustainability of its low-cost advantage, and the competitive pressure open models place on closed-source vendors like OpenAI and Anthropic. Some also express skepticism that benchmarks can be optimized for, calling for more independent evaluations.

Topic 2: Anthropic Launches Claude Tag for Slack Workspaces Link to heading

  • Category: AI · News
  • Overview: Trending for: 6 hours ago, Related posts: 7,100
  • What it is: Anthropic has launched the Claude Tag feature for Slack workspaces, allowing teams to invoke Claude directly within Slack to assist with conversations and workflows.
  • Why it matters: This signals a deeper embedding of AI assistants into enterprise collaboration scenarios, moving from standalone chatbots to deeply integrated tools within daily office platforms. This could affect corporate knowledge management, automated collaboration, and AI assistant adoption rates.
  • Discussion summary: Discussions on X are primarily focused on its potential to enhance team efficiency, its competitive standing against Slack’s native AI and other office assistants, and concerns regarding enterprise data privacy, permission controls, and the risk of misuse.

Topic 3: OpenAI Engineer Draws Late-Night Codex Projects from Devs Worldwide Link to heading

  • Category: AI · News
  • Overview: Trending for: 14 hours ago, Related posts: 418
  • What it is: An OpenAI engineer initiated a late-night call on X, inviting developers from around the world to showcase their projects built with Codex, sparking hundreds of replies.
  • Why it matters: This reflects the transition of AI programming tools from the demo phase to real-world development scenarios, where the developer ecosystem and practical use cases are becoming crucial indicators of a model’s capability and product value.
  • Discussion summary: The discussion centers on whether Codex genuinely enhances development efficiency, which projects best showcase its capabilities, and whether AI programming tools will serve as developer assistants or eventually replace some basic coding jobs.

Today’s AI Public Opinion Summary on X Link to heading

Today’s main narrative is the shift of AI from a race for model capabilities to real-world implementation: GLM-5.2 represents open models catching up to the cutting edge of closed-source models in performance and cost; Claude Tag exemplifies assistants becoming deeply integrated into enterprise workflows; and the Codex discussion indicates that AI programming tools are now entering the practical development ecosystem. The consensus is that the value of AI increasingly depends on its ability to be embedded into workflows at a lower cost, enhance productivity, and establish sustainable use cases for both developers and enterprises. Points of disagreement primarily revolve around whether public benchmarks are equivalent to real-world capabilities, whether the low-cost advantage can be sustained long-term, and whether AI assistants will ultimately augment human work or replace certain jobs. Potential risks include gamed benchmarks leading to skewed expectations, ambiguous boundaries concerning enterprise data privacy and permissions, and the misuse, liability questions, and security vulnerabilities associated with over-reliance on AI tools.

💡 Influencer Insights Link to heading

Okay, based on the compilation of tweets from the last 24 hours, here is the analysis report on AI industry trends.


Core Trend: The fierce competition in on-device models and AI programming tools has reached a white-hot stage, with cost and efficiency becoming the core metrics.

  • On-device Model Capabilities Mature, on the Eve of an Explosion: Multiple industry leaders have expressed amazement at the capabilities of small models running locally.

    • @zhixianio tested the full-duplex audio and video effects of MiniCPM-o 4.5 locally and found the quality satisfactory, “hard to imagine” for a 9B model. He also shared his experience of forcing himself to use a local model (Qwen3.6-35B-A3B) for programming and daily tasks, with the result that “the speed and intelligence were online, and the multimodal experience was even better than remote models.” He also conducted an in-depth review of the potential of Google’s Gemma 4 series (including the 12B multimodal and E4B models) on-device.
    • His proposed “Model-Pak” (large model cartridge) concept has also sparked interest in the future distribution formats for on-device models.
  • AI Programming Tools Become a “Red Ocean” Battlefield, with Codex Gaining Momentum:

    • New Model Releases and Comparisons: The release of Google Gemma 4 12B Coder sparked heated discussions. @zhixianio conducted a deep comparative review and concluded that while it performs on par with 35B MoE models on simple tasks, the 12B size has a clear capability ceiling for complex, stateful program generation. His “sweet spot” remains Qwen 35B.
    • OpenAI Codex Ecosystem and Capability Expansion: Bloggers like @vista8 and @AI_Jasonyu have been intensively discussing new Codex features, such as Record & Replay (seen as a super RPA), an embedded browser for unlimited canvas image generation (the Cowart project), and updates to the Codex CLI. @dotey also mentioned an open-source project that gives ChatGPT local code manipulation capabilities similar to Codex.
    • Choice Between Tools: @ruanyf launched a poll on “Codex vs Claude Code,” reflecting developers’ indecision between the two.
  • Inference Costs for AI Generation Become a Prominent Focus:

    • @ruanyf shared data about an OpenAI employee consuming an estimated $1.3 million worth of tokens in a month, sparking a debate on whether AI is more expensive than human programmers. @vista8 and @Pluvio9yte also mentioned that advanced models like Codex/Fable consume tokens extremely quickly, pushing the community to look for more economical APIs or reset solutions.

2. Noteworthy Unique Perspectives or Industry Foresight Link to heading

  • From “Correlation” to “Causality”: The Next Key Variable in AI Development: @Pluvio9yte provided an in-depth review of Aether AI, founded by Professor Biwei Huang, and its “Causal Large Model” philosophy. He pointed out the limitations of current large models based on correlation prediction, such as their inability to understand that “pouring water into a cup with a hole will cause it to leak.” He believes that introducing causal reasoning is key to solving complex problems like embodied intelligence and scientific discovery, representing an important exploratory direction in technological evolution. Viewpoint source: @huang_biwei

  • Rethinking and Reshaping the Value of Labor with AI: @ruanyf shared a Hacker News discussion on “Can we have time off after AI improves efficiency?” raising a profound question: If AI only increases productivity without improving employee welfare, what is its significance for individuals? This touches on the core contradiction in the transformation of production relations in the AI era.

  • “Testing is the New Moat” as the Barrier of Code Itself Disappears: @ruanyf observed an engineer who replicated Next.js using AI at a cost of $1,100, pointing out that the moat of code no longer exists. The real barrier for software in the future will be comprehensive test cases and domain know-how, rather than the code itself. This is a judgment with profound implications for the software industry.

  • Technical Innovation in Attention Mechanisms Exclusive to On-device Models: @vista8 provided a deep analysis of the technical principles behind Baidu’s Unlimited OCR model. Its “sliding attention window” mechanism, inspired by how humans copy text, successfully solves the KV cache expansion problem when processing long documents. This is an excellent example of engineering and technical innovation enhancing a model’s application effectiveness.

  • The Dilemma of AI Innovation Within Enterprises: @dotey reported in detail on the incident where a Google engineer was fired for developing the popular Google Workspace CLI. This case profoundly reveals the potential conflict between the AI agent strategy within large companies and the interests of existing management, as well as the suppression of innovation by bureaucracy.

  • Combining Social Media Traffic Methodologies with AI: @gefei55 shared an overseas SEO strategy that “beats the clock” by using an AI agent to monitor high-engagement tweets with links on Twitter, allowing for the discovery of new terms and projects before they are reflected in Google Trends. @vista8 even used Codex to develop a Skill that simulates the headline style of the well-known media outlet “Xinzhiyuan,” demonstrating AI’s ability to deconstruct specific content production patterns.


  • AI Programming and Development Tools

    • Gemma 4 12B Coder: The latest local code generation model released by Google (@HuggingModels, review @zhixianio).
    • Cowart: A Codex + infinite canvas tool plugin that is now open-source. It supports a more intuitive way to annotate and modify images (@zhongerxin, recommended by @vista8).
    • DevSpace: An open-source MCP server that gives the ChatGPT web interface local code manipulation capabilities similar to Codex (@gefei55).
    • getdesign.md: A collection of DESIGN.md design system files from real brands like Linear, Vercel, and Notion. Providing these to AI programming tools can significantly improve the professionalism and consistency of the generated UI (@Pluvio9yte).
    • Tencent Cloud EdgeOne Makers: A newly released one-stop Agent hosting platform that solves common pain points encountered when deploying local demos, such as concurrency, storage, and security. It offers a free quota (@AI_Jasonyu).
  • AI Productivity and Automation

    • Fable-5 (Usage Tip): Use the /effort command to switch to max intensity, which can unlock more powerful capabilities (@Pluvio9yte).
    • Claude Tag: A new Slack feature released by Anthropic that allows Claude to reside in channels as a colleague. It can be assigned tasks via @-mentions and supports multi-person collaboration, continuous learning, and proactive notifications (@claudeai, analysis by @dotey).
    • WeChat Public Account to PPT Tool: An open-source project currently under development by @vista8 that can convert WeChat Public Account articles into editable PPTX files with a single click.
    • YouMind: A content creation tool recommended by @AI_Jasonyu and @gefei55, particularly skilled at integrating with the formatting of various publishing platforms (like X Article).
  • Open-Source Models and Knowledge Bases

    • Baidu Unlimited OCR: An open-source long-document OCR model that uses a sliding attention mechanism. It performs exceptionally well on extra-long PDFs and is only 1.5MB in size (@Pluvio9yte, @vista8).
    • WeKnora: An enterprise-level knowledge platform quietly open-sourced by Tencent. It integrates RAG Q&A, a ReAct Agent, and self-maintaining Wiki/knowledge graph functionalities (@Pluvio9yte).
    • “Deep Agents in Action”: The third book on Agent development, open-sourced by @zhanghaili0610 (@vista8).

📚 Appendix: Today’s Watch List Source Updates Link to heading

Timeframe: Last 3 days; 22 sources covered; 35 updates in total

All-In Podcast (A_full) Link to heading

  • GameStop CEO Ryan Cohen’s $56B Plan to Take Over eBay
    • Release Time: 2026-06-24 04:52 Beijing Time
    • Abstract: - AppLovin Ads — AppLovin’s AI advertising platform reaches over 1 billion daily active users in the mobile gaming sector.
      • Full-screen video ads with a median watch time of 35 seconds.
      • Advertisers spend hundreds of thousands of dollars daily for profit, and advertiser access is still in closed beta.
      • Nasdaq - Nasdaq is positioned at the nexus of technology and capital markets, providing world-class platforms and services to global capital markets and beyond with unparalleled technology, insights, and market expertise.
      • GameStop CEO Ryan Cohen’s $56 billion plan to take over eBay.
    • EN Highlights:
      • (0:00) David Friedberg intros GameStop CEO Ryan Cohen
  • (1:56) Building and selling Chewy for $3.35B, how to compete with Amazon in e-commerce
  • (11:58) Post-Chewy life, activist investing, the road to GameStop CEO, expanding into collectibles
  • (26:39) Why he wants to buy eBay for $56B: Massive potential, poor execution (slow growth, rising expenses, seller relationship failure)

Stratechery by Ben Thompson (A_full) Link to heading

  • Memory Chips and China, Microsoft and Chinese Models
    • Published: 2026-06-23 18:00 Beijing time
    • Summary: - The three major memory manufacturers may regret opening the door to Chinese memory manufacturers; meanwhile, Microsoft is highly incentivized to use Chinese models.
      • $15/month* or *$150/year.
      • Substantive analysis of the day’s news through three weekly emails or podcasts.
      • Strategy Interviews.
      • Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
    • EN Key Points:
      • The big three memory makers may come to regret opening up the door to Chinese memory makers; Microsoft, meanwhile, is very incentivized to use Chinese models.

OpenAI Blog (A_full) Link to heading

  • How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery

    • Published: 2026-06-24 01:00 Beijing time
    • Summary: - Physician and immunologist Derya Unutmaz has been interested in artificial intelligence for years.
      • But his “aha” moment came in late 2025 when GPT-5 Pro helped him and his lab revisit a three-year-old puzzle centered on a special type of immune cell that helps the body fight cancer and other diseases.
      • The mystery focused on a fundamental but important question in immunology: How does glucose influence the development and specialization of T cells?
      • T cells are immune cells that help the body fight viruses, kill cancer cells, respond to certain bacteria and parasites, and distinguish between healthy cells and threats.
      • As they mature, they take on different jobs, including roles that can lead to cancer, autoimmune diseases, and infections.
    • EN Key Points:
      • GPT-5 Pro helped solve a 3-year-old immunology mystery, offering insights into T cell behavior
      • The breakthrough could support cancer and autoimmune research.
  • Helping build shared standards for advanced AI

    • Published: 2026-06-23 21:00 Beijing time
    • Summary: - Increasingly capable models can enhance cybersecurity, accelerate scientific discovery, and expand access to expertise.
      • However, they can also pose security risks if their capabilities are misunderstood, safeguards are inadequate, or governments lack the necessary information to respond.
      • To realize the benefits safely and confidently, society will need institutions with the technical and governance capabilities to evaluate, safeguard, and govern increasingly capable systems.
      • Appia will develop open, modular specifications designed to translate international standards and established frameworks into practical assessment criteria across the entire AI value chain.
      • Its work can help develop a critical missing layer of trust, through which third parties can check for compliance with standards, generating clearer and more reusable evidence as different organizations develop models, infrastructure, and applications.
    • EN Key Points:
  • OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the Appia Foundation.

  • How Omio is building the future of conversational travel

    • Published: 2026-06-23 08:00 Beijing Time
    • Summary: - Discover how Omio uses OpenAI to enhance conversational travel experiences, accelerate product development, and transform into an AI-native company.
      • This article from the OpenAI blog explains how Omio is building the future of conversational travel, shaping the broader AI and infrastructure landscape.
      • Following up on how Omio is building the future of conversational travel, it also offers practical implications for founders, operators, and investors.
    • EN Key Points:
      • Discover how Omio uses OpenAI to power conversational travel experiences, accelerate product development, and transform into an AI-native company.

ArXiv cs.AI (B_intro+search) Link to heading

  • On the Identifiability of User Adaptation in Co-Adaptive Neural Interfaces

    • Published: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20569v1 Announcement Type: new.
      • Abstract: We analyze identifiability in co-adaptive human-machine systems.
      • We show that closed-loop encoder estimates do not uniquely identify user adaptation, but instead reflect properties of the joint system.
      • We discuss the implications for interpreting behavioral adaptation and propose conditions for identification.
    • EN Key Points:
      • arXiv:2606.20569v1 Announce Type: new
      • Abstract: We analyze identifiability in co-adaptive human-machine systems
      • We show that closed-loop encoder estimates do not uniquely identify user adaptation, but instead reflect properties of the joint system
      • We discuss implications for interpreting behavioral adaptation and propose conditions for identification.
  • Beyond Fixed Budgets: Characterizing the Inelasticity and Limitations of Tree-of-Thought Reasoning Strategies

    • Published: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20599v1 Announcement Type: new.
      • Abstract: Tree-of-Thought (ToT) search has emerged as a promising direction for enhancing the reasoning capabilities of large language models, but deploying these methods in practice raises a question that has received little systematic attention: How do different search strategies perform across varying computational budgets, model sizes, and problem difficulties?
      • In this work, we evaluate two representative ToT methods; DPTS (a method based on Monte Carlo Tree Search) and SSDP (a method based on semantic deduplication) across two mathematical reasoning benchmarks (Math500 and GSM8K), two model scales (Llama-3B and Llama-8B), and four token budgets (3k–10k).
      • Our analysis reveals that these two methods exhibit limitations in opposite directions.
    • EN Key Points:
      • arXiv:2606.20599v1 Announce Type: new
  • Abstract: Tree of Thought (ToT) search has become a promising direction for improving the reasoning capabilities of large language models, but deploying these m…

  • In this work, we evaluate two representative ToT methods; DPTS, a Monte Carlo tree search based approach, and SSDP, a semantic deduplication based approach, acr…

  • Our analysis reveals that the two methods exhibit limitations that pull in opposite directions

  • The New Associationism: Lessons from Deep Learning

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20600v1 Announce Type: new.
      • Abstract: What can the success of modern AI tell us about how humans learn?
      • This paper argues that taking AI seriously as a model of human learning supports a modest but genuine associationism.
      • The core finding is that supervised learning (learning driven by evaluative feedback) is the basis for a surprisingly wide range of contemporary AI systems, from large language models to game-playing agents, with the main difference being how much work is required to generate the relevant feedback signal.
    • EN Key Points:
      • arXiv:2606.20600v1 Announce Type: new
      • Abstract: What can the success of modern AI tell us about how humans learn
      • This paper argues that taking AI seriously as a model of human learning supports a modest but genuine associationism
      • The central finding is that supervised learning – learning driven by evaluative feedback – underlies a surprisingly wide range of contemporary AI systems, fro…
  • Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20615v1 Announce Type: new.
      • Abstract: AI agents now participate as first-class team members across the entire software development lifecycle, yet no specification language exists to express the human-agent responsibility boundaries, approval gates, and governance constraints required for this collaboration.
      • Existing approaches encode processes in agent prompts (which can drift), target adjacent domains (workflow management, business processes), or only address fragments (access control, approval gates).
      • We propose a domain-specific language for specifying AI-SDLC processes as protocols, with a formal syntax, well-formedness conditions, operational semantics, and enforced invariants.
    • EN Key Points:
      • arXiv:2606.20615v1 Announce Type: new
      • Abstract: AI agents now participate as first-class team members across the software development lifecycle, yet no specification language exists for expressing t…
      • Existing approaches encode process in agent prompts (subject to drift), target adjacent domains (workflow management, business processes), or address only fragm…
  • We propose a domain-specific language for specifying AI-SDLC processes as protocols, with formal syntax, well-formedness conditions, operational semantics, and…

  • PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent Debate

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20621v1 Announcement Type: New.
      • Abstract: Multi-agent debate improves the reliability of large language models (LLMs) through iterative peer critiques.
      • However, fixed topologies often introduce persistent positional biases, amplify unreliable agents, and lead to high sensitivity to role assignments.
      • We introduce \textit{Permutation-Equivariant Adaptive Routing Multi-Agent Debate (PEAR)}, an inference-time protocol that can dynamically reconfigure communication roles and sparse topologies in successive debate rounds.
    • EN Key Points:
      • arXiv:2606.20621v1 Announce Type: new
      • Abstract: Multi-agent debate improves the reliability of large language models (LLMs) through iterative peer critiques
      • However, fixed topologies often introduce persistent positional biases, amplify unreliable agents, and cause high sensitivity to role assignments
      • We introduce \textit{Permutation-Equivariant Adaptive Routing Multi-Agent Debate (PEAR)}, an inference-time protocol that dynamically reconfigures communication…
  • Darwin Mobile Agent: A Roadmap for Self-Evolution

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20622v1 Announcement Type: New.
      • Abstract: The goal of artificial intelligence is to create agents capable of general, adaptive behavior in open-ended environments.
      • Guided by the “Bitter Lesson,” we argue that the most effective path to achieving this goal is to systematically eliminate human priors and allow intelligence to emerge naturally through interaction with a “large world” that is several orders of magnitude more complex than the agent itself.
      • We propose the mobile Graphical User Interface (GUI) as a practical proxy for such a world and introduce the Darwin Mobile Agent, an open-source infrastructure intended to serve as the foundation for autonomous reinforcement learning in this domain.
    • EN Key Points:
      • arXiv:2606.20622v1 Announce Type: new
      • Abstract: The goal of artificial intelligence is to create agents capable of general, adaptive behaviour in open-ended environments
      • Guided by the “Bitter Lesson”, we argue that the most effective path toward this goal is to systematically remove human priors and allow intelligence to natural…
      • We propose the mobile Graphical User Interface (GUI) as a practical proxy for such a world and introduce Darwin Mobile Agent, an open-source infrastructure desi…
  • Path-dependent program induction under resource constraints explains human sequence learning

    • Published: 2026-06-23 12:00 Beijing Time
    • Summary: - arXiv:2606.20623v1 Announce Type: new.
      • Abstract: How do people build abstract, reusable knowledge from sequential experience under bounded cognitive resources?
      • To answer this question, we integrate rate-distortion theory with recent advances in program induction to describe how prior knowledge shapes which future structures are low-cost to encode and easy to discover.
      • We formalize this in a Hierarchical Adaptor Grammar (HAG) with distinct local (within-task) and global (across-task) libraries, governed jointly by constraints of memory and computation.
    • EN Highlights:
      • arXiv:2606.20623v1 Announce Type: new
      • Abstract: How do people build abstract, reusable knowledge from sequential experience under bounded cognitive resources
      • To answer this question, we integrate rate-distortion theory with recent advances in program induction to describe how prior knowledge shapes which future struc…
      • We formalize this in a hierarchical Adaptor Grammar (HAG) with distinct local (within-task) and global (across-task) libraries, governed jointly by constraints…
  • In LLM Reasoning, there is Irrationality on top of Value Misalignment

    • Published: 2026-06-23 12:00 Beijing Time
    • Summary: - arXiv:2606.20624v1 Announce Type: new.
      • Abstract: Significant progress has been made in aligning LLMs with target value functions.
      • We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximize the aligned value in reasoning.
      • We mathematically formalize this gap as rational value risk: the utility discrepancy between a model’s deployed reasoning strategy and its rational counterpart, defined as the response that maximizes expected utility in the steepest direction.
    • EN Highlights:
      • arXiv:2606.20624v1 Announce Type: new
      • Abstract: Significant progress has been made in aligning LLMs with target value functions
      • We argue that, even when an LLM has been well aligned in (post-)training, it may still fail to maximise the aligned value in reasoning
      • We mathematically formalise this gap as rational value risk: the utility discrepancy between a model’s deployed reasoning strategy and its rational counterpart,…
  • AlphaMemo: Structured Search-Process Memory for Self-Evolving Alpha Mining Agents

    • Published: 2026-06-23 12:00 Beijing Time
    • Summary: - arXiv:2606.20625v1 Announce Type: new.
      • Abstract: LLM agents show promise for alpha mining by combining financial priors, symbolic reasoning, executable factor generation, and feedback-driven refinement.
      • However, they face challenges with combinatorial search spaces, noisy and non-stationary feedback, redundant discoveries, and the risk of overfitting from naively reusing past successes.
  • To address these challenges, we propose AlphaMemo, a self-evolving alpha mining agent with a structured search process memory.

    • EN Highlights:
      • arXiv:2606.20625v1 Announce Type: new
      • Abstract: LLM agents are promising for alpha mining via combining financial priors, symbolic reasoning, executable factor generation, and feedback-driven refine…
      • Yet, they face a combinatorial search space, noisy non-stationary feedback, redundant discoveries, and overfitting risks from naively reusing past successes
      • To address these challenges, we propose AlphaMemo, a self-evolving alpha mining agent with Structured Search-Process Memory
  • Latent Goal Prediction from Language for Model-Based Planning

    • Publish Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20627v1 Announce Type: new.
      • Abstract: Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals.
      • Visual targets provide precise local gradients but poor distant guidance, while language, although flexible, is limited by noisy cross-modal alignment or reliance on large generative models, making it unsuitable for the high-sampling nature of model-based planning.
      • To address these challenges, we introduce Latent Goal Prediction from Language (LAGO), a framework that predicts a sequence of intermediate goal states based on language instructions and action conditions, all within the same latent space.
    • EN Highlights:
      • arXiv:2606.20627v1 Announce Type: new
      • Abstract: Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals
      • Visual targets provide precise local gradients but poor distant guidance, while language is flexible yet limited by noisy cross-modal alignment or dependence on…
      • To address these challenges, we introduce Latent Goal Prediction from Language (LAGO), a framework that predicts both sequences of intermediate goal states from…

ArXiv cs.CL (B_intro+search) Link to heading

  • Less is More: Lightweight Prompt Compression for Question Answering Applications on Edge Devices

    • Publish Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20571v1 Announce Type: new.
      • Abstract: In agent-driven Question Answering (QA) applications, Retrieval-Augmented Generation (RAG) is often introduced to improve the response accuracy of Large Language Models (LLMs) by providing additional context.
      • Due to the inherent noise in retrieval results and the coarse granularity of document-level retrieval, the retrieved context often contains a large amount of redundant information.
      • In this setup, the agent’s prompt, consisting of the user query and the associated retrieved context, leads to unnecessary computational overhead during LLM inference.
    • EN Highlights:
      • arXiv:2606.20571v1 Announce Type: new
  • Abstract: In agent-driven question answering (QA) applications, retrieval-augmented generation (RAG) is commonly introduced to enhance the response accuracy of…

  • Due to the inherent noise in retrieval results and the coarse granularity of document-level retrieval, the retrieved context often contains substantial redundan…

  • In this setting, the agent prompt, consisting of the user query and the associated retrieved context, leads to unnecessary computational overhead during LLM inf…

  • Investigating Linguistic Steering: An Analysis of Adjectival Effects Across Large Language Model Architectures

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20572v1 Announcement Type: New.
      • Abstract: Achieving reliable control of Large Language Models (LLMs) requires a precise, scalable understanding of how they interpret linguistic cues.
      • We introduce a rigorous framework using Shapley values to quantify the steering effect of individual adjectives on model performance, moving beyond anecdotal heur…
      • Applying this method to 100 adjectives across a diverse suite of models (including o3, gpt-4o-mini, phi-3, llama-3-70b, and deepseek-r1) on the MMLU benchmark,…
    • EN Key Points:
      • arXiv:2606.20572v1 Announce Type: new
      • Abstract: Achieving reliable control of Large Language Models (LLMs) requires a precise, scalable understanding of how they interpret linguistic cues
      • We introduce a rigorous framework using Shapley values to quantify the steering effect of individual adjectives on model performance, moving beyond anecdotal he…
      • Applying this method to 100 adjectives across a diverse suite of models (including o3, gpt-4o-mini, phi-3, llama-3-70b, and deepseek-r1) on the MMLU benchmark,…
  • Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20632v1 Announcement Type: New.
      • Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other’s outputs, or coordinate as agents.
      • Their value depends on models producing observably different conversational behaviors when given the same input.
      • Prior offline research suggested mapping a behavioral diversity model for each family, as LLMs prefer outputs from their own family when evaluating each other separately.
    • EN Key Points:
      • arXiv:2606.20632v1 Announce Type: new
      • Abstract: Multi-LLM systems use multiple language models to deliberate, judge each other’s outputs, or coordinate as agents
  • Their value depends on the models producing measurably different conversational behaviors when given the same input

  • Prior offline studies recommend drawing one model per family for behavioral diversity, because LLMs prefer outputs from their own family when rating one another…

  • EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

    • Posted: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20650v1 Announce Type: new.
      • Abstract: Instruction-based controllable speech synthesis enables users to specify emotions through natural language.
      • However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained intensity.
      • We propose EmoInstruct-TTS, a dual-path instruction-guided framework for emotional speech synthesis.
    • EN Highlights:
      • arXiv:2606.20650v1 Announce Type: new
      • Abstract: Instruction-based controllable speech synthesis enables users to specify emotions through natural language
      • However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained intensity
      • We propose EmoInstruct-TTS, a dual-path instruction-guided framework for emotional speech synthesis
  • Specific Domain Ontology Construction Using Large Language Models

    • Posted: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20691v1 Announce Type: new.
      • Abstract: Ontologies are useful structures for organizing and maintaining information that can be understood by both humans and systems.
      • However, since their manual creation is a laborious task, many specific domains lack reference ontologies.
      • The outstanding ability of Large Language Models (LLMs) to understand natural language has prompted their application in various fields, including ontology development.
    • EN Highlights:
      • arXiv:2606.20691v1 Announce Type: new
      • Abstract: Ontologies are useful structures to organize and maintain information that can be understood both by humans and systems
      • However, since their manual crafting is a laborious task, many specific domains lack reference ontologies
      • The outstanding ability for understanding natural language demonstrated by the Large Language Models (LLMs) has motivated their application to aid on a variety…
  • MindAlign: Decoding Inner Speech from fMRI Signals via Multimodal Embedding Alignment under Limited Data

    • Posted: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20696v1 Announce Type: new.
  • Abstract: Decoding inner speech from non-invasive brain signals remains a fundamental challenge due to the absence of overt linguistic output, limited training data, and large inter-subject variability.

    • Existing brain-to-text approaches often rely on task-specific decoder fine-tuning, which restricts scalability and complicates adaptation to new participants.
    • We propose MindAlign, a decoupled two-stage brain-to-language framework that can generate open-ended text from fMRI signals without modifying the underlying language model.
    • EN Highlights:
      • arXiv:2606.20696v1 Announce Type: new
      • Abstract: Decoding inner speech from non-invasive brain signals remains a fundamental challenge due to the absence of overt linguistic output, limited training…
      • Existing brain-to-text approaches often rely on task-specific decoder fine-tuning, which restricts scalability and complicates adaptation to new participants
      • We propose MindAlign, a decoupled two-stage brain-to-language framework that enables open-ended text generation from fMRI signals without modifying the underlyi…
  • VeriBound: PAC-Bayesian Generalization Bounds for Process Reward Models Trained with Formal Verification Tools

    • Release Date: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20740v1 Announce Type: new.
      • Abstract: Process Reward Models (PRMs) provide step-level verification for Large Language Model (LLM) reasoning, but their training data acquisition remains a bottleneck: manual annotation is costly, and Monte Carlo rollout estimates are noisy.
      • A recent approach, FOVER, trains PRMs on step-level error labels automatically annotated by formal verification tools such as Z3 and Isabelle, and empirically observes cross-task generalization from symbolic tasks to different reasoning benchmarks.
      • However, this generalization phenomenon lacks any theoretical explanation, and no formal bounds exist on the generalization error, sample complexity, convergence speed, or downstream Best-of-K performance of such PRMs.
    • EN Highlights:
      • arXiv:2606.20740v1 Announce Type: new
      • Abstract: Process Reward Models (PRMs) provide step-level verification for Large Language Model (LLM) reasoning, yet their training data acquisition remains a b…
      • A recent approach, FOVER, trains PRMs on step-level error labels automatically annotated by formal verification tools such as Z3 and Isabelle, and empirically o…
      • However, this generalization phenomenon lacks any theoretical explanation, and no formal bounds exist on the generalization error, sample complexity, convergenc…
  • From Sentiment to Actionable Insights: A Data-Driven Public Sentiment Analysis of Advanced Air Mobility

    • Release Date: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20751v1 Announce Type: new.
  • Abstract: Advanced Air Mobility (AAM) is an emerging low-altitude air transportation system whose successful deployment depends not only on technological advancements but also on public acceptance.

  • This acceptance will drive government support, regulations, noise standards, and willingness to fly, which in turn will improve the overall commercial viability of AAM.

  • Therefore, understanding public sentiment towards AAM is crucial for identifying its societal barriers and informing its adoption strategies.

    • EN points:
      • arXiv:2606.20751v1 Announce Type: new
      • Abstract: Advanced Air Mobility (AAM) is an emerging low-altitude air transportation system whose successful deployment depends not only on technological advanc…
      • This acceptance will drive government support, regulations, noise standards, and willingness to fly, and in turn the overall commercial viability of AAM
      • Understanding public sentiment toward AAM is therefore essential for identifying its societal barriers and informing its adoption strategies
  • FirstPass: Grounding AI Scientific Judgment in Multi-Round Editorial Outcomes

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20769v1 Announce Type: new.
      • Abstract: AI systems for peer review fail on three fronts: they are trained exclusively on computer science and machine learning venues, neglect the iterative dialogue essential for validating science, and are evaluated based on stylistic mimicry rather than genuine editorial judgment.
      • We introduce FirstPass, a dataset and fine-tuned model that addresses all three issues.
      • Curating 3,668 complete multi-round peer-review dialogues from Nature Communications across five scientific domains (biology, chemistry, neuroscience, physics, and earth sciences), we leverage mandatory transparent peer review (commencing November 2022) and verify 100% content integrity through automated audits.
    • EN points:
      • arXiv:2606.20769v1 Announce Type: new
      • Abstract: AI systems for peer review fail on three fronts: they train on Computer Science and Machine Learning venues alone, ignore the iterative dialogue that…
      • We introduce FirstPass, a dataset and fine-tuned model that addresses all three
      • Curating 3,668 complete multi-round peer-review dialogues from Nature Communications across five scientific domains (biology, chemistry, neuroscience, physics,…
  • Beyond ‘One Language, One Script’: Quantifying Orthographic Bias in Multilingual VLMs with PuMVR

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20770v1 Announce Type: new.
      • Abstract: Current Visual Language Models (VLMs) are renowned for their multilingual capabilities, yet they operate under a flawed assumption: one language corresponds to one writing system.
      • This overlooks billions of users of multi-script languages such as Punjabi, Serbian, Hindi-Urdu, and Kurdish, for whom model capabilities may be adversely affected by orthographic bias.
  • We introduce PuMVR (Punjabi Multimodal Visual Reasoning), the first benchmark designed to quantify script-dependent bias through 375 culturally-grounded image reasoning tasks in the three active scripts of Punjabi (Gurmukhi, Shahmukhi, Roman).

    • EN Key Points:
      • arXiv:2606.20770v1 Announce Type: new
      • Abstract: Current Vision-Language Models (VLMs) are celebrated for their multilingual capabilities, yet they operate under a flawed assumption: that one languag…
      • This overlooks billions of users of multi-script languages like Punjabi, Serbian, Hindi-Urdu, Kurdish, among many others, for whom a model’s capability may be f…
      • We introduce PuMVR (Punjabi Multimodal Visual Reasoning), the first benchmark designed to quantify script-dependent bias through 375 culturally grounded image-r…

ArXiv cs.LG (B_intro+search) Link to heading

  • Towards CSI-Native Foundation Models: A Channel-Adaptive Roadmap for 6G

    • Release Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20670v1 Announce Type: new.
      • Abstract: Wireless foundation models offer a path toward reusable channel state information (CSI) intelligence for sixth-generation (6G) systems.
      • However, existing generic-backbone adaptation and CSI pretraining methods often treat CSI as task tensors rather than propagation-conditioned channel responses, thus failing to capture the inherent time-frequency-space geometry of the wireless environment.
      • This paper presents a channel-adaptive roadmap toward CSI-native foundation models, proposing a unified framework that aligns pretraining, positional modeling, and attention control with three channel requirements: scale-aware heterogeneous exposure, physical time-frequency-antenna coordinates, and correlation-constrained token interactions.
    • EN Key Points:
      • arXiv:2606.20670v1 Announce Type: new
      • Abstract: Wireless foundation models offer a path toward reusable channel state information (CSI) intelligence for sixth-generation (6G) systems
      • However, existing generic-backbone adaptation and CSI pretraining methods often treat CSI as task tensors rather than propagation-conditioned channel responses,…
      • This paper presents a channel-adaptive roadmap toward CSI-native foundation models, proposing a unified framework that aligns pretraining, positional modeling,…
  • NeuroShield: A Device-Agnostic Foundation Model for EEG Authentication

    • Release Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20673v1 Announce Type: new.
      • Abstract: A core challenge in EEG authentication is that models are often tied to the acquisition settings they were trained on.
      • Specifically, variations in headset hardware, channel layouts, and signal durations create heterogeneous recordings that existing models cannot handle, causing each new headset or dataset to be treated as a separate model development problem.
      • This fragmentation limits multi-dataset learning, hinders knowledge transfer, and reduces model reusability.
    • EN Key Points:
  • arXiv:2606.20673v1 Announce Type: new

  • Abstract: A central challenge in EEG authentication is that models are typically tied to the acquisition settings in which they are trained

  • In particular, variations in headset hardware, channel layout, and signal duration create heterogeneous recordings that existing models are not designed to hand…

  • This fragmentation limits multi-dataset learning, hinders knowledge transfer, and reduces model reusability

  • Massive Activations Are Architecturally Robust: A Controlled Scratch/Commitment Residual Stream Test

    • Release Date: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20743v1 Announce Type: new.
      • Abstract: Trained Transformers reliably develop massive activations, i.e., a small number of hidden dimensions whose magnitude is far above the median and which is concentrated on the sequence start token.
      • Whether these outliers are a removable artifact of the residual stream’s overloaded read and write role, or instead a functional necessity, is actively debated.
      • We test the artifact hypothesis directly with an architectural intervention.
    • EN Key Points:
      • arXiv:2606.20743v1 Announce Type: new
      • Abstract: Trained transformers reliably develop massive activations, a small number of hidden dimensions whose magnitude is far above the median and which conce…
      • Whether these outliers are a removable artifact of the residual stream’s overloaded read and write role, or instead a functional necessity, is actively debated
      • We test the artifact hypothesis directly, with an architectural intervention
  • CIExplainer++: Generating Causal and Interpretable Explanations for Graph Neural Networks

    • Release Date: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20747v1 Announce Type: new.
      • Abstract: Explainable Artificial Intelligence aims to make black-box models more trustworthy by presenting, in a human-understandable manner, the elements that lead to the model’s output.
      • This involves (i) identifying constituent parts and connections that have a truly causal effect on the output, and (ii) translating such structures into interpretable representations.
      • For the former, we introduce CIExplainer, a novel perturbation-based approach grounded in causal inference for explaining Graph Neural Networks (GNNs).
    • EN Key Points:
      • arXiv:2606.20747v1 Announce Type: new
      • Abstract: Explainable Artificial Intelligence aims to make black-box models more trustworthy by presenting, in a human-understandable manner, the elements that…
  • This involves both (i) identifying components and connections with genuine causal influence on outputs and (ii) translating such structures into an interpretabl…

  • For the former, we introduce CIExplainer, a novel perturbation-based method grounded in causal inference for explaining Graph Neural Networks (GNNs)

  • Evidential Fusion Network for Multimodal Survival Prediction under Missing Modalities

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract:- arXiv:2606.20757v1 Announcement Type: new.
      • Abstract: Recent multimodal survival prediction models have demonstrated strong predictive performance by leveraging complementary information across modalities.
      • However, such models generally assume data completeness and exhibit limited robustness toward missing modalities, which are frequently encountered in real-world clinical settings.
      • We propose the Evidential Missing Modality Survival Fusion (EMMS) model for multimodal survival prediction under missing modalities.
    • EN Key Points:
      • arXiv:2606.20757v1 Announce Type: new
      • Abstract: Recent multimodal survival prediction models have demonstrated strong predictive performance by leveraging complementary information across modalities
      • However, such models generally assume data completeness and exhibit limited robustness toward missing modalities, which are frequently encountered in real-world…
      • We propose the Evidential Missing Modality Survival Fusion (EMMS) model for multimodal survival prediction under missing modalities
  • ELADO: Elliptic PDE Assessment Datasets for Operator Learning

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract:- arXiv:2606.20771v1 Announcement Type: new.
      • Abstract: We introduce ELADO (Elliptic PDE Assessment Datasets for Operator Learning), a systematic benchmark suite constructed to show and quantify failure modes of neural operator architectures when learning solution operators for elliptic partial differential equations.
      • While the benchmarks of existing datasets focus on average-case performance, the ELADO datasets are constructed to highlight challenges that arise naturally in elliptic PDE problems.
      • In particular, we construct several datasets built around the Poisson and Helmholtz equations, each with non-constant coefficients.
    • EN Key Points:
      • arXiv:2606.20771v1 Announce Type: new
      • Abstract: We introduce ELADO (Elliptic PDE Assessment Datasets for Operator Learning), a systematic benchmark suite constructed to show and quantify failure mod…
      • While the benchmarks of existing datasets focus on average case performance, the ELADO datasets are constructed to highlight challenges that arise naturally in…
  • In particular, we construct several datasets built around Poisson’s equation and the Helmholtz equation, each with non-constant coefficients

  • B[FM]$^2$: Brain Foundation Model via Flow Matching with SplitUNet

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20812v1 Announce Type: new.
      • Abstract: EEG foundation models can learn generalizable representations from large-scale EEG corpora, enabling single-backbone transfer across diverse clinical and brain-computer interface tasks.
      • Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervision.
      • Recognizing that this discretization disrupts continuous brain rhythms and obscures fine-grained temporal dynamics, we propose B[FM]$^2$ (Brain Foundation Model via Flow Matching), whose inductive bias aligns with the data by pre-training directly on the raw signal using continuous-time flow matching, without the need for patches, tokenization, or masking.
    • EN Key Points:
      • arXiv:2606.20812v1 Announce Type: new
      • Abstract: EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical…
      • Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervis…
      • Recognizing that this discretization fragments continuous brain rhythms and obscures fine-grained temporal dynamics, we present B[FM]$^2$(Brain Foundation Model…
  • CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20820v1 Announce Type: new.
      • Abstract: Can we trust evaluation scores to reflect an LLM’s true real-world performance?
      • Certifiable evaluation answers this question by providing guarantees for LLM evaluation.
      • Specifically, existing methods sequentially curate evaluation samples and continuously update a confidence interval (CI) that covers the true performance with a high probability (e.g., 95%), until certain conditions are met, such as the CI width reaching a target precision.
    • EN Key Points:
      • arXiv:2606.20820v1 Announce Type: new
      • Abstract: Can we trust evaluation scores to capture an LLM’s true real-world performance
      • Certifiable evaluation answers this question by providing guarantee for LLM evaluation
      • In particular, existing methods sequentially curate evaluation samples and keep updating confidence intervals (CIs) that cover the true performance with high pr…
  • Evolutionary Discovery of Developmental Reward Schedules in Deep Reinforcement Learning

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract: - arXiv:2606.20858v1 Announce Type: new.
  • Abstract: The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and remains fixed throughout training, while the progression of motivational priorities is largely unexplored.

  • In this work, we propose an evolutionary framework for discovering developmental reward schedules, in which three distinct biologically inspired motivational components—agency, novelty, and reactivity—are combined through time-varying weights that dynamically change during training.

  • Evaluated on two sparse-reward MiniGrid tasks: DoorKey-6x6 and KeyCorridorS3R1, our framework compares the generalizability of four evolutionary algorithms: CMA-ES, xNES, DE, and L-SHADE with an extrinsic motivation baseline (our primary comparison point) and three other hand-designed methods.

    • EN Highlights:
      • arXiv:2606.20858v1 Announce Type: new
      • Abstract: The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the…
      • In this work, we propose an evolutionary framework for discovering developmental reward schedules, in which three distinct biologically inspired motivational co…
      • Evaluated on two sparse-reward MiniGrid tasks: DoorKey-6x6 and KeyCorridorS3R1, our framework compares the generalizability of four evolutionary algorithms: CMA…
  • Machine Learning Classification of Cryopathy Syndromes: A Comprehensive Comparative Study

    • Publication Time: 2026-06-23 12:00 Beijing Time
    • Abstract:- arXiv:2606.20874v1 Announce Type: new.
      • Abstract: Cryopathy syndromes are difficult to classify because laboratory patterns often overlap across diagnostic categories, and some diagnoses are rare.
      • This makes routine interpretation of cryoglobulin-related tests challenging and increases dependence on expert judgment.
      • The aim of this study was to develop and compare machine learning approaches for automated classification of cryopathy syndromes from laboratory data and to identify practical strategies for clinical decision support.
    • EN Highlights:
      • arXiv:2606.20874v1 Announce Type: new
      • Abstract: Cryopathy syndromes are difficult to classify because laboratory patterns often overlap across diagnostic categories, while some diagnoses are rare
      • This makes routine interpretation of cryoglobulin-related tests challenging and increases dependence on expert judgment
      • The aim of this study was to develop and compare machine learning approaches for automated classification of cryopathy syndromes from laboratory data and to ide…