🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-14
- 类型
- ai-daily
- 字数
- 6833
- 阅读时长
- 33 min
2026-07-14 AI Daily | Agents Begin to Bolster Control Layers: From Inference-Time Orchestration to Long-Horizon Evaluation Link to heading
Today’s main theme is the shift of Agents from demonstration capabilities to reliable execution: CogniConsole, GATS, and long-horizon terminal benchmarks are all strengthening control, planning, and feedback mechanisms. On the application side, RAG continues to enter professional domains such as law, investment, and archives. On the business side, model vendors are engaging in more nuanced competition over access stability, ecosystem entry points, and talent.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The most important theme to delve into today is “Agents moving from demonstration to reliable execution.” CogniConsole abstracts inference-time control into a formal architecture, GATS attempts to reduce planning randomness with graph-enhanced search, and Long-Horizon-Terminal-Bench extends evaluation to longer-period, higher-feedback terminal tasks. Engineering teams should pay close attention.
The second theme is model efficiency and safety boundaries. HALO explores adaptive latent inference on frozen models, complexity-guided initialization focuses on the pre-training starting point, and the “emergent misalignment” paper reminds us that misalignment and realignment phenomena may not be robust, and safety conclusions still require more rigorous experimental support.
On the application side, RAG is entering high-value professional scenarios: knowledge graph fact verification, legal precedent retrieval, investment briefing generation, and literary archive indexing are all deepening. Also of note are the lawsuits involving Apple and OpenAI, which reflect the ongoing strategic battles over AI talent, data, and the boundaries of trade secrets.
🌐 AI Hot Topics on X Link to heading
Topic 1: Anthropic Extends Claude Fable 5 Access to July 19 Amid Rival Launches Link to heading
- Category: AI · News
- Overview: Trending Time: 1 day ago, Related Posts: 54,000
- What happened: Anthropic announced the extension of access to Claude Fable 5 until July 19, coinciding with a period of intensive new model and product releases from competitors.
- Why it’s important: This move is likely intended to maintain user attention and developer retention, reflecting the escalating competition among large model companies over model capabilities, availability, and ecosystem stickiness.
- Discussion summary: Discussions on X are focused on whether the extension means Anthropic is buying time for its next release, how Fable 5’s actual performance compares to competitors, and whether users should continue to bet on the Claude ecosystem. Some also question if this is merely a marketing tactic without substantial technical updates.
Topic 2: Anthropic Extends Claude Fable 5 Access to July 19 Amid User Pressure Link to heading
- Category: AI · News
- Overview: Trending Time: 2 days ago, Related Posts: 69,000
- What happened: Under user pressure, Anthropic extended access to Claude Fable 5 until July 19.
- Why it’s important: This highlights that the availability, pricing, and access policies for high-performance AI models are becoming crucial factors for user retention and market competition. It also shows the growing dependence of developers and heavy users on model continuity.
- Discussion summary: Discussions on X are centered on whether Anthropic underestimated user demand, whether the extension is just a temporary appeasement, the reasonableness of model access restrictions, and how companies should balance computational costs with user experience.
Topic 3: Monzo Co-Founder Tom Blomfield Joins Anthropic from Y Combinator Link to heading
- Category: AI · News
- Overview: Trending Time: 14 hours ago, Related Posts: 2,400
- What happened: Tom Blomfield, co-founder of Monzo and a former partner at Y Combinator, has joined Anthropic.
- Why it’s important: This indicates that top AI companies are continuing to attract seasoned talent from the fintech and startup ecosystems to strengthen their capabilities in product, commercialization, and startup collaborations.
- Discussion summary: Discussions on X mainly focus on Anthropic’s appeal to talent, the specific areas Blomfield will be responsible for, and whether his career move from YC to a major AI company reflects a siphoning effect of the AI industry on startup talent.
Topic 4: OpenAI’s GPT-5.6 Sol Tops Design Benchmarks After Launch Tweaks Link to heading
- Category: AI · News
- Overview: Trending Time: 1 day ago, Related Posts: 14,000
- What happened: After some post-launch adjustments, OpenAI’s GPT-5.6 Sol has achieved leading scores on design-related benchmarks.
- Why it’s important: This demonstrates the continued improvement of large models’ capabilities in application scenarios such as visual design, interface generation, and creative workflows, which could impact the competitive landscape of AI design tools and productivity software.
- Discussion summary: Discussions on X are primarily focused on whether its benchmark scores can represent real-world design capabilities, whether post-launch tweaks affect fair comparisons, and whether OpenAI’s advantages in multimodal and design tasks over other model vendors are sustainable.
Summary of Today’s AI Public Opinion on X Link to heading
The main narrative in public opinion today is that the competition in large models is shifting from simply “who can release a more powerful model” to “who can provide sustained stable access, real-world usable capabilities, and stronger ecosystem stickiness.” The consensus is that Anthropic extending access to Claude Fable 5 and attracting Tom Blomfield to join both indicate that leading AI companies are accelerating their strategies around user retention, developer ecosystems, and commercialization. OpenAI’s lead in design benchmarks also reinforces the judgment that multimodal and creative workflows are becoming the new battleground. Disagreements mainly center on the substance of these actions: whether the Claude extension is a response to user demand and a way to alleviate pressure on computing power and product timelines, or simply a marketing tactic and delay; it is also still debated whether GPT-5.6 Sol’s benchmark lead represents true design capabilities. The potential risk is that if model vendors rely too heavily on limited-time access, benchmark optimization, and talent narratives to maintain hype, it could exacerbate user distrust in usability, pricing fairness, and the credibility of evaluations, while also creating greater uncertainty for developers betting on a particular ecosystem.
💡 Influencer Insights Link to heading
Hello, I am an AI industry analyst. Based on the tweet stream from the past 24 hours, I have compiled the following insights report for you.
X Platform AI Influencer Daily Insights Report Link to heading
Date: July 13, 2026
1. Top Tech & Product Buzz Today: OpenAI’s Grand Ecosystem Unification and the “Codex Phenomenon” Outbreak Link to heading
The undisputed focus today is OpenAI’s latest integrated ecosystem, spearheaded by the full public launch of Codex and ChatGPT Work. Influencers’ discussions are centered on the restructuring of the product’s form and the actual user experience.
- “Three-in-One” Desktop App & New Model Matrix: OpenAI has merged ChatGPT, Codex, and Work into a unified desktop application (@dotey). This is accompanied by the public release of the GPT-5.6 series models (Sol/Terra/Luna), which correspond to flagship, cost-effective, and lightweight tasks, respectively (@dotey). @Pluvio9yte made a side-by-side comparison of Sol, Grok 4.5, and Claude Fable 5, noting that Sol has a lower code rework rate than GPT-5.5 but is slightly less dynamic than Fable 5.
- Codex’s “Breakout” and Capability Upgrade: The hype around Codex has extended beyond being just a developer tool. The most talked-about news today is the temporary removal of its 5-hour usage limit and its integration of one-click Chrome Cookie import and developer mode (@Pluvio9yte, @dotey). This means Codex can use a user’s logged-in state to access platforms like Douyin and Xiaohongshu to directly debug or discover potential topics. @Pluvio9yte even found it can attach to WeChat, though the feature is currently unstable.
- Claude Fable 5’s “Tug-of-War” Extension: Anthropic has once again extended access to Claude Fable 5 (possibly a codename for Opus 5) until July 19 (@zhixianio, @Pluvio9yte). This extremely rare strategy of repeated extensions has sparked different interpretations: @dotey criticized it as “child’s play,” while @Pluvio9yte boldly predicted it signals the imminent arrival of the real Opus 5.
- Hot AI Editing Tool: ChatCut: Both @Pluvio9yte and @vista8 highlighted the AI video editing tool ChatCut. Although @vista8 believes its results are not yet on par with custom Remotion workflows, both influencers acknowledged its value as a tool for rapid video production within the Codex ecosystem.
2. Notable Unique Perspectives & Industry Foresight Link to heading
- Reflections on “Telegraph-Style Skills” (@dotey): Baoyu raised systematic questions about the current “token-saving” trend (e.g., the Caveman project). He used the “telegraphic style” analogy to point out that forcing an AI to write code in “caveman English” only saved 8.5% of output tokens in real-world tests by JetBrains. This is because an Agent’s primary costs are tool calls and context, making trimming idle chatter a trivial optimization, like “cutting the mineral water budget.” He emphasized that as token costs fall, precise, highly corrective “long-winded” responses are more valuable than being miserly with words.
- Individual vs. Team Dynamics in the AI Era (@dotey): Based on insights from an internal share at Anthropic, AI is thinning the “scaffolding” for Agents. The focus is shifting from controlling every step of a process to designing collaboration between Agents. However, this creates a new problem: an individual can rapidly generate 10 prototypes, leading to chaotic product sprawl. AI amplifies individual ability but fails to solve the team’s challenge of making trade-offs and decisions. This poses a new challenge for management.
- Is Xiaohongshu Building the Next “Code Hosting Platform”? (@ruanyf): Ruan Yifeng points out a forward-looking industry trend: Xiaohongshu has launched the REDSkill community. Users can now directly upload and share Skill files in their notes. This is a “wall-breaking” attempt to combine a lifestyle community with technology distribution, seen as a beachhead for Xiaohongshu’s AI transformation, and could open up a new distribution channel for developers.
- Decision-Making and Value in the AI Era (@vista8): Qiao Mu quotes a famous line from a 1979 IBM slide: “A computer can never be held accountable, therefore a computer must never make a management decision.” In an age where AI can handle most execution, the scarce value of humans is shifting towards: discovering the right questions to ask, making judgments with insufficient information, and taking responsibility for the consequences (Skin in the game).
- The True Nature of Open Source (@ruanyf): Citing the views of Anthropic’s founders, the industry is beginning to more rigorously distinguish between “open source” and “open weights.” Since the internal workings of models are not visible, open source in the traditional sense doesn’t truly exist in the AI field, prompting a re-examination of the openness and controllability of models.
3. Recommended Tools & Resources Link to heading
- Lightweight Development Skill: mattpocock (Recommended by: @Pluvio9yte): While many are using heavyweight Superpowers, @Pluvio9yte strongly recommends this 160k-star lightweight Skill. It’s particularly well-suited for flexible combination and layering on the current versions of Sol and Fable 5, significantly enhancing coding flexibility.
- AI Editing Integration: ChatCut MCP + Skill (Recommended by: @vista8): ChatCut can be directly installed into Codex as an MCP and Skill. It allows you to generate a demo video for a website through natural language conversation. While the style is rudimentary, it represents an automated leap from pure text to video.
- Open-Source Tech Diagram Library: fireworks-tech-graph (Recommended by: @vista8): Having accumulated 8.5k stars purely by word-of-mouth, this library is perfect for quickly drawing professional technical architecture diagrams with AI assistance, solving the pain point of creating illustrations for articles.
- New Force in On-Device Models: MiniCPM-o 4.5 (Recommended by: @zhixianio): A 9B full-duplex audio-video model was tested locally and is considered to have practical, real-world usability in terms of voice quality. This demonstrates the immense potential of small-parameter models after on-device optimization.
- Xiaohongshu Scraping Tool: Specific API (Recommended by: @Pluvio9yte): Recommends a tool that, aside from occasional API endpoint updates, can meet almost all Xiaohongshu scraping needs. It’s extremely low-cost (over 300 calls for just $0.80), making it ideal for market research.
📚 Appendix: Today’s Watch List Update Sources Link to heading
Time window: Last 3 days; 22 sources covered; 31 updates in total.
Stratechery by Ben Thompson (A_full) Link to heading
- Apple Sues OpenAI, Apple’s Real Problem
- Published: 2026-07-13 18:00 Beijing Time
- Abstract: - Apple is suing AI for stealing trade secrets; there is one guilty employee, but this mostly feels like lashing out.
- $15/month or *$150/year.
- Substantive analysis of the day’s news via three weekly emails or a podcast.
- Strategy Interviews.
- Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
- EN Key Point:
- Apple is suing AI for stealing trade secrets; there is one guilty employee, but this mostly feels like lashing out.
ArXiv cs.AI (B_intro+search) Link to heading
Interval Certifications for Multilayered Perceptrons via Lattice Traversal
- Published: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08773v1 Announce Type: new.
- Abstract: In this work, we propose a rigorous theoretical framework for a fundamental problem in AI safety, namely adversarial robustness.
- In particular, we demonstrate that the adversarial robustness problem can be reduced to a lattice traversal problem.
- Each element of this lattice corresponds to an interval, i.e., an axis-aligned hyper-rectangle, containing the input point $\mathbf{x}$.
- EN Key Point:
- arXiv:2607.08773v1 Announce Type: new
Abstract: In this work we present a rigorous theoretical framework to a foundational problem of AI safety, namely adversarial robustness
In particular, we show that the adversarial robustness problem can be reduced to a a lattice traversal problem
Each element of this lattice corresponds to an interval, i.e., an axis-aligned hyper-rectangle, containing an input point $\mathbf{x}$
- Published: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08774v1 Announce Type: new.
- Reliability in large language model (LLM) systems is typically framed as a function of model capability.
- We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control}—the computational layer governing task framing and context selection.
- We introduce \emph{CogniConsole}, an architectural instantiation that externalizes this control into a structured interface, combining programmatic coordination with inference based on finite prompts.
- EN Highlights:
- arXiv:2607.08774v1 Announce Type: new
- Abstract: Reliability in large language model (LLM) systems is typically framed as a function of model capability
- We challenge this by demonstrating that reliability is significantly influenced by \emph{inference-time control} – the computational layer governing task frami…
- We introduce \emph{CogniConsole}, an architectural instantiation that externalizes this control into a structured interface combining programmatic coordination…
GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
- Published: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08894v1 Announce Type: new.
- Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and ReAct rely heavily on LLM inference during the planning process, resulting in high computational costs and stochastic behavior.
- We propose \textbf{GATS} (Graph-Augmented Tree Search), a planning framework that combines UCB1-based systematic tree search with a layered world model to eliminate LLM calls during the inference process while achieving superior planning performance.
- Our three-layer world model integrates: (L1) precise symbolic action matching, (L2) statistics learned from execution logs, and (L3) LLM-based prediction for unknown actions.
- EN Highlights:
- arXiv:2607.08894v1 Announce Type: new
- Abstract: Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree Search) and…
We present \textbf{GATS} (Graph-Augmented Tree Search), a planning framework that combines systematic UCB1-based tree search with a layered world model to elimi…
Our three-layer world model integrates: (L1) exact symbolic action matching, (L2) statistics learned from execution logs, and (L3) LLM-based prediction for unkn…
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08964v1 Announcement Type: New.
- Abstract: AI agents have become capable of autonomously completing short, well-specified tasks.
- However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome.
- This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability.
- EN Highlights:
- arXiv:2607.08964v1 Announce Type: new
- Abstract: AI agents have become capable of autonomously completing short, well-specified tasks
- However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome
- This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08986v1 Announcement Type: New.
- Abstract: We formalize a research result in the Lean 4 proof assistant by having a mathematician direct an AI system, and frame the activity as a formalization game.
- The objective is to turn a LaTeX document into a Lean document.
- The game is won when the development compiles, contains no sorry, and a machine check shows the target theorems rest on Lean’s foundational axioms alone.
- EN Highlights:
- arXiv:2607.08986v1 Announce Type: new
- Abstract: We formalize a research result in the Lean 4 proof assistant by having a mathematician direct an AI system, and frame the activity as a formalization…
- The objective is to turn a LaTeX document into Lean
- The game is won when the development compiles, contains no sorry, and a machine check shows the target theorems rest on Lean’s foundational axioms alone
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning
- Release Date: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09059v1 Announce Type: New.
- Abstract: We present ARCANA, a collaborative multi-agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints.
- ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective refinement.
- A perceptual grounding agent builds object-centric scene graphs from raw grids, a latent program policy proposes diverse DSL programs, a symbolic executor verifies candidates in demonstrations, and a reflective agent synthesizes failure-driven feedback for the next round.
- EN Key Points:
- arXiv:2607.09059v1 Announce Type: new
- Abstract: We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints
- ARCANA decomposes each task into iterative perception, hypothesis generation, symbolic execution, and reflective refinement
- A perceptual grounding agent builds object centric scene graphs from raw grids, a latent program policy proposes diverse DSL programs, a symbolic executor verif…
- Release Date: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09076v1 Announce Type: New.
- Abstract: Cyberattacks on operational technology are increasingly causing costly downtime and physical damage, exposing the limitations of traditional rule-based monitoring in industrial IoT environments.
- While Large Language Models (LLMs) have strong semantic reasoning abilities to assist in decision support, their hallucinatory nature presents unacceptable security liabilities for closed-loop control.
- This paper introduces a neuro-agentic control framework, a novel architecture that couples an LLM-based planner (i.e., Gemini 2.5 Flash-Lite) with a pre-trained time-series foundation model (TimesFM) to enable physics-informed autonomous defense.
- EN Key Points:
- arXiv:2607.09076v1 Announce Type: new
- Abstract: Cyberattacks on operational technology are increasingly causing costly downtime and physical damage, exposing the limitations of traditional rule-base…
- While Large Language Models (LLMs) have strong semantic reasoning abilities to assist in decision support, their hallucinatory nature presents unacceptable safe…
- This paper introduces a neuro-agentic control framework, a novel architecture that couples an LLM-based planner (i.e., such as Gemini 2.5 Flash-Lite) with a pre…
L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning
Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09099v1 Announce Type: new.
- Abstract: While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-intensive legal domains remains underexplored.
- In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods in legal textual entailment.
- By assigning distinct expert personas to multiple agents, L-MAD improves upon strong single-agent baselines by up to 8%.
- EN Highlights:
- arXiv:2607.09099v1 Announce Type: new
- Abstract: While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-h…
- In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods within…
- By assigning distinct expert personas to multiple agents, L-MAD improves upon strong single-agent baselines by up to 8%
- Abstract: - arXiv:2607.09099v1 Announce Type: new.
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09142v1 Announce Type: new.
- Abstract: Large language models (LLMs) are increasingly applied in online medical consultation, but existing benchmarks are still not well-aligned with actual clinical practice.
- Many rely on synthetic conversations or patient simulators, omit medical images uploaded by patients, or use multiple-choice or lexical overlap metrics to evaluate open-ended clinical responses, which do not reflect clinical quality well.
- We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation, constructed from de-identified doctor-patient interactions collected from nationwide internet hospitals.
- EN Highlights:
- arXiv:2607.09142v1 Announce Type: new
- Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinica…
- Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choic…
- We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collect…
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09153v1 Announce Type: new.
- Abstract: Process Reward Models (PRMs) have been shown to be very effective in guiding Test-Time Scaling (TTS) methods, which significantly enhances the capabilities of LLM-based multi-agent systems.
However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch.
In long multi-agent deployments, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely limiting the application of PRMs in long-context scenarios.
- EN Highlights:
- arXiv:2607.09153v1 Announce Type: new
- Abstract: Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabil…
- However, existing PRMs are text-based: they re-encode the entire trajectory text from scratch
- In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length L, creates a severe computational bottleneck, severely lim…
- EN Highlights:
ArXiv cs.CL (B_intro+search) Link to heading
HALO: Hybrid Adaptive Latent Reasoning for Language Models
- Release Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08775v1 Announcement Type: New.
- Abstract: We study how to improve a frozen pretrained language model with a small amount of adaptive extra computation.
- A simple approach is to add additional refinement steps on top of the backbone hidden states, but fixed extra refinement can be wasteful: a one-step refinement head may be too weak, while forcing a second full-sequence refinement step everywhere can increase computation without improving transfer.
- We introduce HALO, a hybrid adaptive latent-refinement method that combines a coarse refinement stage with selective second-stage latent refinement on a subset of tokens selected via token scoring and monotonic token stopping.
- EN Highlights:
- arXiv:2607.08775v1 Announce Type: new
- Abstract: We study how to improve a frozen pretrained language model with a small amount of adaptive extra computation
- A simple approach is to add additional refinement steps on top of the backbone hidden states, but fixed extra refinement can be wasteful: a one-step refinement…
- We introduce HALO, a hybrid adaptive latent-refinement method that combines a coarse refinement stage with selective second-stage latent refinement on a subset…
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?
- Release Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09053v1 Announcement Type: New.
- Abstract: Recent work reports Emergent Misalignment (EM), where a language model fine-tuned on a narrow, domain-specific misalignment dataset suddenly acquires broad misaligned behaviors, alongside evidence that this behavior can be reversed with limited realignment.
- We systematically study repeated alignment and misalignment cycles using a controlled fine-tuning loop, while tracking behavioral performance as well as LoRA representations throughout training.
- While we reproduce EM, we find that both misalignment and realignment are highly sensitive to superficial dataset characteristics, with the apparent rapid realignment largely disappearing after controlling for response length differences.
- EN Highlights:
- arXiv:2607.09053v1 Announce Type: new
Abstract: Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire…
We systematically study repeated alignment and misalignment cycles using controlled fine-tuning loops while tracking behavioral performance, and LoRA representa…
Although we reproduce EM, we find that both misalignment and realignment are highly sensitive to superficial dataset characteristics, with apparent rapid realig…
- Release Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09092v1 Announcement Type: New.
- Abstract: Knowledge Graphs (KGs) are typically constructed automatically from large-scale corpora, but due to noisy sources and extraction failures, they inevitably contain factual errors, and reliably verifying them at an industrial scale remains a key challenge.
- To address this problem, we propose AgentKGV, an Agentic LLM-RAG framework for KG fact verification, which integrates dynamic routing and iterative query rewriting to handle surface form mismatches in document-level retrieval.
- To make this framework more accurate and cost-effective for industrial deployment, we further introduce a two-stage training strategy: turn-level distillation-based SFT, which transfers reasoning capabilities from a large teacher model to a small model for stable query rewriting and inference; and trajectory-level GRPO, which optimizes the search strategy to reduce large-scale unnecessary retrieval.
- EN Key Points:
- arXiv:2607.09092v1 Announce Type: new
- Abstract: Knowledge graphs (KGs) are often automatically constructed from large-scale corpora, but they inevitably contain factual errors due to noisy sources a…
- To address this, we propose AgentKGV, the Agentic LLM-RAG framework for KG fact Verification, that integrates dynamic routing and iterative query rewriting, whi…
- To make this framework more accurate and cost-efficient for industrial deployment, we further introduce a two-stage training strategy: turn-level distillation-b…
PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation
- Release Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09094v1 Announcement Type: New.
- Abstract: Legal precedent retrieval is a fundamental task for legal case preparation, planning, litigation strategy, and legal research.
- Current automatic precedent retrieval methods map legal documents into a low-dimensional semantic space and calculate similarity based on the proximity of their representations.
- These methods treat legal documents as monolithic texts, ignoring the rhetorical organization of legal technical details.
- EN Key Points:
- arXiv:2607.09094v1 Announce Type: new
Abstract: Legal precedent retrieval is a fundamental task in legal case preparation, planning, litigation strategy, and legal research
- Current approaches for automatic precedent retrieval map legal documents to a low-dimensional semantic space and compute similarity based on the proximity of th…
- These approaches treat legal documents as monolithic texts, ignoring the rhetorical organization of the legal technicalities
- Published: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09121v1 Announce Type: new.
- Abstract: In this study, we examine the opportunities that Large Language Models (LLMs) bring to various aspects of the fundamental analysis of companies, based on reports from LLMs, data and documents describing macroeconomic conditions (like GDP and inflation changes), and documents filed with the U.S. Securities and Exchange Commission (SEC), available in EDGAR.
- We are preprocessing this data and then sending it via API to a gpt-4o model in a Retrieval-Augmented Generation (RAG)-like regime.
- EN Highlights:
- arXiv:2607.09121v1 Announce Type: new
- Abstract: In this study, we examine the opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies based on t…
- Securities and Exchange Commission (SEC) which can be found in EDGAR
- We were preprocessing those data and than sending via API to gpt-4o model in a Retrieval-Augmented Generation (RAG) like regime
Complexity-Guided Component-wise Initialization for Language Model Pretraining
- Published: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09204v1 Announce Type: new.
- Abstract: Pre-trained language models often exhibit structured weight spectra, which suggests that training may repeatedly produce similar layered and component-wise organizations.
- We investigate whether these recurring spectral patterns can be repurposed as an initialization signal for pre-training GPT-2 style language models.
- First, we analyze 11 pre-trained GPT-2 style checkpoints of varying sizes, languages, tokenizers, and training corpora, measuring the Frobenius norm and effective rank entropy across layers and Transformer subcomponents.
- EN Highlights:
- arXiv:2607.09204v1 Announce Type: new
- Abstract: Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wi…
We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-style language-model pretraining
- First, we analyze eleven pretrained GPT-2-style checkpoints that vary in size, language, tokenizer, and training corpus, measuring Frobenius norm and effective-…
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09291v1 Announcement Type: new.
- Abstract: Medieval document transcribers had very different practices; on top of that, heterogeneous digitization policies have resulted in corpora where the character sets must be treated as fluid.
- In this paper, we address the problem of changing between character sets in a flexible manner.
- We focus on one-to-one character mappings and train character-level one-to-one RNNs to undo them with self-supervision; recovering half the CER even with 20 lines of text.
- EN Key Points:
- arXiv:2607.09291v1 Announce Type: new
- Abstract: Medieval document transcribers have very different practices; on top of that, heterogeneous digitization policies have resulted in corpora where the c…
- In this paper we address the problem of changing between character-sets in a flexible manner
- We focus on one-to-one character mappings and train characterlevel one-to-one RNNs to undo them with self-supervision; recovering half the CER even with 20 text…
Creativity, honesty and designed forgetting emerge in small hyperbolic language models
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09306v1 Announcement Type: new.
- Abstract: Language models are optimized for scale, yet remain functional rather than companionable, and as an assistant personalizes into a companion, accumulating a user’s memories, it quietly becomes someone and can silently acquire traits that harm the user.
- What a companion is becoming, and what would make it worth becoming, has no reliable instrument: trained human raters cannot agree on the answer (Fleiss kappa = 0.074).
- Here, we show that three small language models (146M to 3B parameters) sharing a hyperbolic basis answer both parts of the question.
- EN Key Points:
- arXiv:2607.09306v1 Announce Type: new
- Abstract: Language models are optimised for scale, yet remain functional rather than companionable, and as an assistant personalises into a companion, accumulat…
- What a companion is becoming, and what would make it worth becoming, has no reliable instrument: trained human raters cannot agree on the answer (Fleiss kappa =…
Here we show that three small language models (146 M to 3 B parameters) sharing a hyperbolic substrate answer both halves of that question
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09316v1 Announcement Type: New.
- Abstract: Thematic indexing—the practice of assigning structured conceptual labels to sections of text—is essential for scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labor-intensive process.
- This paper explores the application of machine learning to automatic thematic indexing, using two substantial sub-corpora of the Complete Works of Voltaire as a test case: Essai sur les mœurs et l’esprit des Nations and Questions sur l’Encyclopédie.
- The task is framed as a multi-label classification problem, in which a model must assign the set of index entries that a professional indexer would apply to a given text page.
- EN Highlights:
- arXiv:2607.09316v1 Announce Type: new
- Abstract: Thematic indexing – the practice of assigning structured conceptual labels to sections of text – is essential to scholarly access in large-scale lit…
- This paper explores the application of machine learning to automatic thematic indexing, using two substantial sub-corpora of the Complete Works of Voltaire as a…
- The task is framed as a multi-label classification problem, in which a model must assign the set of index entries that a professional indexer would apply to a g…
Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.09324v1 Announcement Type: New.
- Abstract: Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections.
- This article reports the findings of the “Extracting Keywords from Crowdsourced Collections” project, which used the Their Finest Hour Online Archive, a crowdsourced World War II digital collection hosted by the University of Oxford, as a case study.
- The project evaluated three natural language processing methods for automatic keyword extraction: named entity recognition, keyword extraction, and topic modeling.
- EN Highlights:
- arXiv:2607.09324v1 Announce Type: new
- Abstract: Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections
- This article reports the findings of the “Extracting Keywords from Crowdsourced Collections” project, which used the Their Finest Hour Online Archive, a crowdso…
The project evaluated three Natural Language Processing approaches to automate keyword extraction: Named Entity Recognition, Keyword Extraction, and Topic Model…
ArXiv cs.LG (B_intro+search) Link to heading
A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08776v1 Announcement Type: New.
- Abstract: Despite the success of Knowledge Distillation (KD) in Large Language Models (LLMs), the underlying mechanisms behind its efficacy remain unclear.
- In this paper, we propose a unified approach to explore the common mechanisms of various KD methods using interactions.
- Specifically, we decompose the LLM’s output score into the sum of numerous interactions.
- EN Key Points:
- arXiv:2607.08776v1 Announce Type: new
- Abstract: Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains unclear
- In this paper, we propose a unified approach to explore the common mechanism of various KD methods using interactions
- Specifically, we decompose the output score of the LLM into the sum of numerous interactions
iLENS: Interpretable LLM-Guided Mixture-of-Experts for Neuroimaging Survival Analysis
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08778v1 Announcement Type: New.
- Abstract: Alzheimer’s Disease (AD) is a complex neurodegenerative disorder that continues to impact millions of people worldwide.
- Predicting AD conversion in the prodromal stage remains crucial for disease understanding and patient care.
- Therefore, survival models are widely used for AD risk prediction, but they are often static predictors with limited interpretability and no natural language reasoning capabilities.
- EN Key Points:
- arXiv:2607.08778v1 Announce Type: new
- Abstract: Alzheimer’s Disease (AD) is a complex neurodegenerative disorder that continues to impact millions of people worldwide
- Predicting AD conversion during the prodromal stage remains critical for disease understanding and patient care
- As such, survival models are widely used for AD risk prediction, yet they are typically static predictors with limited interpretability and no capacity for natu…
Signed Symmetric Quantization for Few-Bit Integers
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08779v1 Announcement Type: New.
- Abstract: Signed integer alphabets contain one more representable negative value than positive ones.
However, by convention, the standard symmetric integer quantizer fixes its scale to be strictly positive, which assigns this extra representable value to the negative tail and can force positive outliers to be clipped.
In this work, we show that, at few-bit precision, such clipping is a significant source of quantization error.
- EN Key Points:
- arXiv:2607.08779v1 Announce Type: new
- Abstract: The signed integer alphabet contains one more negative representable value than positive
- Yet, by convention, the standard symmetric integer quantizer fixes its scale to be strictly positive, which assigns this extra representable value to the negati…
- In this work, we show that, at few-bit precision, such clipping is a non-trivial source of quantization error
- EN Key Points:
Sticky Routing: Training MoE Models for Memory-Efficient Inference
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08780v1 Announce Type: new.
- Abstract: Mixture-of-Experts (MoE) models only activate a sparse subset of experts per token, but consecutive tokens often activate different experts - leading to constant weight swapping between slow storage and fast memory on edge devices.
- Existing remedies are either system-level (caching heuristics) or post-hoc (router fine-tuning), leaving the root cause unchanged during pre-training.
- We propose StickyMoE, a differentiable routing consistency loss that penalizes abrupt expert switches between adjacent tokens, encouraging the router to maintain the same expert assignment over semantically consistent spans.
- EN Key Points:
- arXiv:2607.08780v1 Announce Type: new
- Abstract: Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts – ca…
- Existing remedies are either system-level (caching heuristics) or post-hoc (router fine-tuning), leaving the root cause unchanged during pretraining
- We propose StickyMoE, a differentiable routing consistency loss that penalises abrupt expert switches between adjacent tokens, encouraging the router to maintai…
Reward Transport: Property Control in Flow Matching via Noise-Space Alignment
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08781v1 Announce Type: new.
- Abstract: The coupling in flow matching (the rule for pairing noise vectors with data points) is often treated as a computational choice.
- We show that this coupling can act as an alignment interface: by matching noise and data according to target molecular properties, it directly embeds controllable structure into the learned flow field.
- Building on this insight, we introduce Reward Transport, which uses optimal transport coupling at training time to align scalar noise-space coordinates with molecular rewards; at inference time, varying this coordinate can steer the generated distribution without oracles, reward models, gradient guidance, or additional computation.
- EN Key Points:
- arXiv:2607.08781v1 Announce Type: new
Abstract: The coupling in flow matching – the rule pairing noise vectors with data points – is typically treated as a computational choice
We show that this coupling can instead serve as an alignment interface: by matching noise and data according to a target molecular property, it embeds controlla…
Building on this view, we introduce Reward Transport, which uses optimal transport coupling at training time to align a scalar noise-space coordinate with molec…
Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
- Published: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08782v1 Announce Type: new.
- Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models.
- Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs.
- Existing works for optimizing expert placement focus on leveraging past requests’ expert activation patterns.
- EN Highlights:
- arXiv:2607.08782v1 Announce Type: new
- Abstract: Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models
- Its efficiency depends on the communication and computation latencies of the GPUs, which are linked to the placement of experts in the GPUs
- Existing works for optimizing expert placement focus on leveraging past requests’ expert activation patterns
LieBN: Batch Normalization over Lie Groups
- Published: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08783v1 Announce Type: new.
- Abstract: Manifold-valued measurements are prevalent in various machine learning tasks.
- Recent advances have extended Deep Neural Networks (DNNs) to operate on manifolds, accompanied by normalization techniques tailored to different geometries, collectively known as Riemannian normalization.
- However, most existing Riemannian normalization methods are either designed for specific manifolds or fail to effectively normalize manifold-valued sample distributions.
- EN Highlights:
- arXiv:2607.08783v1 Announce Type: new
- Abstract: Manifold-valued measurements are prevalent in various machine learning tasks
- Recent advances have extended Deep Neural Networks (DNNs) to operate on manifolds, accompanied by normalization techniques tailored to different geometries, col…
- However, most existing Riemannian normalization methods are either designed for specific manifolds or fail to effectively normalize manifold-valued sample distr…
HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08784v1 Announce Type: new.
- Abstract: Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge.
- Existing evaluations are difficult to compare because they often simultaneously change datasets, task splits, client data splits, task orders, backbones, memory assumptions, and reporting rules.
- We introduce \textbf{HERO}, a heterogeneity-aware benchmark library for FCL.
- EN 要点:
- arXiv:2607.08784v1 Announce Type: new
- Abstract: Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge
- Existing evaluations are difficult to compare because they often change datasets, task splits, client data splits, task orders, backbones, memory assumptions, a…
- We introduce \textbf{HERO}, a heterogeneity-aware benchmark library for FCL
DaDaDa: A Dataset for Data Pricing in Data Marketplaces
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08785v1 Announce Type: new.
- Abstract: High-quality data drives machine learning advances across industries.
- Recognizing the value of data, data transactions are increasingly common, giving rise to many data marketplaces, e.g., AWS Marketplace, Databricks, and Datarade.
- However, determining the appropriate prices for data products remains a significant challenge due to the unique properties of data products.
- EN 要点:
- arXiv:2607.08785v1 Announce Type: new
- Abstract: High-quality data drives machine learning advances across industries
- Recognizing the value of data, data transactions are increasingly common, giving rise to many data marketplaces, e.g., AWS Marketplace, Databricks, and Datarade
- However, determining the appropriate prices for data products remains a significant challenge due to the unique properties of data products
- Publication Time: 2026-07-13 12:00 Beijing Time
- Abstract: - arXiv:2607.08786v1 Announce Type: new.
- Abstract: With the growing deployment of Large Language Models (LLMs), LLM inference cost has become a critical challenge.
- Pruning techniques that introduce sparsity into weight matrices can accelerate inference.
- However, maintaining model quality often limits pruning to moderately unstructured sparsity (around 50%).
- EN 要点:
- arXiv:2607.08786v1 Announce Type: new
Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge
Pruning techniques that introduce sparsity into weight matrices can accelerate inference
However, maintaining model quality typically limits pruning to moderate unstructured sparsity (around 50%)