🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-07-29
- 类型
- ai-daily
- 字数
- 7944
- 阅读时长
- 38 min
2026-07-29 AI Daily | Beyond 2.8 Trillion Parameters: MCP Statelessness Makes Agents More Like Infrastructure Link to heading
Kimi K3 continues to push the open-weight ceiling, but more notably, Agent infrastructure is changing: MCP is shifting from stateful streams to stateless request/response, which is more conducive to Serverless and edge deployment. Meanwhile, discussions on model reliability, evaluation, and hallucination governance are clearly heating up, as AI moves from a capability race to engineered implementation.
📖 In-depth Guide to This Issue’s Watch List Link to heading
Today, there are three main threads worth following. The first is “AI moving from chat to execution”: centered around Claude Code, open-source local assistants, and startup discussions, Agents are truly beginning to take over code, email, calendars, and workflows, making this a key area for engineering teams to focus on. The second is “Research infrastructure being rewritten”: scientific computing, formula formalization, automated peer review, and Text-to-SQL are all solving the same problem—embedding models more reliably into real research and data pipelines, and reducing unnecessary inference overhead. The third is “Model reliability and interpretability”: updates on invisible inference, narrative suppression, document inconsistencies, and translation bias all remind us that beyond the improvement of LLM capabilities, evaluation and governance are also entering deep waters.
🌐 X Platform AI Hot Topic News Link to heading
Topic 1: Moonshot AI Releases Massive 2.8 Trillion Parameter Kimi K3 Model Link to heading
- Category: AI · News
- Overview: Trending 1 day ago, related posts: 50000
- What happened: Moonshot AI publicly released the open weights for Kimi K3, a multimodal model reportedly with 2.8 trillion parameters and supporting a 1 million token context.
- Why it’s important: This marks a further increase in the scale of open-weight frontier large models, potentially driving independent research, private deployment, and low-cost fine-tuning, while also intensifying the competition between open and closed models in terms of capabilities, accessibility, and ecosystem control.
- Discussion overview: Discussion on X primarily focused on three points: it is one of the largest open-weight models to date, but whether it is truly “usable” depends on a high computing power threshold; whether open weights will narrow the gap with closed frontier models; and the symbolic significance of such releases in the context of US-China AI competition and model openness policy debates.
Topic 2: Anthropic’s Claude Opus 5 Tops Key AI Leaderboards Days After Launch Link to heading
- Category: AI · News
- Overview: Trending 23 hours ago, related posts: 584
- What happened: Anthropic’s newly released Claude Opus 5 topped multiple key AI model leaderboards days after its launch.
- Why it’s important: This shows the accelerated competition among frontier large models and may influence developers, enterprise clients, and researchers’ judgments on model capabilities, reliability, and ecosystem choices.
- Discussion overview: Discussion on X primarily focused on whether Claude Opus 5 truly surpasses competing models from OpenAI, Google, etc., whether leaderboard evaluations are sufficiently objective, and whether its performance in actual coding, reasoning, and long-context tasks can match its榜 single ranking.
Topic 3: Andrew Ng Launches AI Personal Tutor Startup LearnVector with $100M Coursera Backing Link to heading
- Category: AI · News
- Overview: Trending:, related posts: 370
- What happened: Andrew Ng announced the launch of LearnVector, an AI personalized tutor startup, backed by $100 million from Coursera.
- Why it’s important: This indicates that AI is further entering the education sector, potentially accelerating personalized tutoring, learning efficiency, and the commercialization of educational products.
- Discussion overview: Discussions on X mainly focused on whether Andrew Ng’s new project will reshape online education, whether Coursera’s huge investment is justified, and whether AI tutors can truly replace or significantly supplement human teaching.
Topic 4: xAI Speeds Ahead with Grok 4.5 Copilot Integration and Massive Models Coming Soon Link to heading
- Category: AI · News
- Overview: Trending 1 day ago, related posts: 1400
- What happened: xAI is widely discussed for integrating Grok 4.5 into Copilot, with rumors of even larger models to be released soon.
- Why it’s important: This indicates that xAI is accelerating the deployment of its models into mainstream productivity scenarios, and also reflects that large model competition is shifting from mere parameters and leaderboards to deep integration with ecosystem entry points like office and development tools.
- Discussion overview: Discussions on X primarily revolved around Grok 4.5’s actual capabilities, the impact of Copilot integration on user experience, and whether “larger models” will truly bring significant improvements, or are more of a marketing expectation and industry competition signal.
Topic 5: AI Pioneer Yangqing Jia Launches Intent Lab with Autonomous Software-Building Swarm Link to heading
- Category: AI · News
- Overview: Trending time: 6 hours ago, Related posts: 177
- What it is: AI pioneer Yangqing Jia announced the launch of Intent Lab, which focuses on software development completed through the collaboration of an autonomous software-building “swarm.”
- Why it matters: This development shows that AI is evolving from single-model capabilities to collaborative agent systems that can execute complex tasks, potentially impacting software engineering, automated development, and the form of AI products.
- Discussion summary: On X, the main discussion revolves around whether this “autonomous software swarm” can reliably complete development work and whether it holds a significant advantage over existing AI programming tools. Others are focused on its business model, practical applications, and its safety and controllability.
Topic 6: Tesla FSD Scrapes Model Y on Pole in Parking Mishap Link to heading
- Category: AI · News
- Overview: Trending time: , Related posts: 129
- What it is: A Tesla Model Y, while using its FSD feature for parking or low-speed driving, reportedly scraped a pillar in a parking lot, drawing attention.
- Why it matters: The incident again highlights the reliability issues of autonomous driving systems in perception and decision-making in low-speed, complex, close-quarters environments. It has practical implications for AI safety verification and liability boundaries.
- Discussion summary: Discussions on X are mainly centered on whether FSD should be held responsible for the accident, whether the driver over-trusted the system, Tesla’s ability to handle edge cases, and whether such individual incidents are representative of the overall level of autonomous driving technology.
AI Public Opinion Summary on X Today Link to heading
Today’s main narrative is the accelerating competition in cutting-edge AI across three fronts: “model capabilities, open ecosystems, and application deployment.” The open-sourcing of Kimi K3’s weights, Claude Opus 5 topping the leaderboards, and rumors of xAI integrating with Copilot all reinforce the judgment that the large model arms race is still heating up. The relative consensus is that open weights, long context, multimodality, and ecosystem entry point integration are reshaping the choices available to developers and enterprises. AI is also expanding from being a chat and coding tool to more complex scenarios like education, software agents, and autonomous driving. Disagreements mainly focus on whether these advancements are truly usable: the computational barriers for massive open-source models, the credibility of leaderboards, the actual user experience of Grok and Claude, and whether AI tutors and autonomous software swarms can replace humans or existing workflows are all still debated. Potential risks include capabilities being hyped beyond their actual reliability, the governance and security pressures brought by open models, liability issues related to over-automation in education and software development, and the user over-trust and inadequate edge-case safety verification exposed by the FSD incident.
💡 Influencer Insights Link to heading
Below is an analytical brief of tweets from the AI domain over the past 24 hours, focusing on tech hotspots, industry foresight, and practical resources.
1. 🔥 Core Tech Trends & Product Hotspots Link to heading
Today’s discussions primarily revolve around the changing landscape of open-source large models, the protocol and architectural evolution of the Agent ecosystem, and the deep application of on-device models and generative AI.
Kimi K3 Open-Source Release Ignites the Community, Chinese Model Dominates HuggingFace Leaderboard Kimi K3, open-sourced by @Kimi_Moonshot, became the undisputed top trend today. The model features 2.8 trillion parameters (MoE architecture) and a million-token context window, directly topping the HuggingFace trending list. @Pluvio9yte observed that it, along with Baidu’s Unlimited OCR, is being called the “twin stars” of Chinese open-source models, and the community has begun seriously discussing that the global model competition has transitioned to a “China-US contest” phase. In terms of technical evaluation, @ruanyf believes its performance is “truly close” to Claude Fable 5 and points out that the primary reason for its capability leap is the massive increase in parameter size (from 1T to 2.8T). However, he also warns that its API pricing is the most expensive among domestic models.
MCP Protocol Releases Major Update: Moving Towards Statelessness @dotey provided a detailed interpretation of the major changes in the MCP protocol version 2026-07-28. The core change is the refactoring of the protocol from a stateful bidirectional stream to a stateless request/response protocol. This move completely solves the server load balancing problem, allowing MCP servers to be deployed in Serverless or edge computing environments like regular HTTP services. The protocol also introduces MRTR (Multi-Round Trip Request) to support mid-process confirmations and officially establishes a minimum 12-month deprecation transition window, which is very friendly for production environment teams.
The “Arms Race” of AI Agents: Universal Entry Points and Plugin Ecosystems Multiple bloggers have noted a consolidation trend in the Agent space. @dotey summarized that general-purpose Agents are consuming vertical Agents, with a few winners taking most of the market (like Claude Code / Codex). The core moat for Agents lies in model intelligence, not the interactive experience. However, for small teams, the opportunity isn’t in building general-purpose Agents, but in the Skill and MCP plugin ecosystem built on top of Agents, which is a blue ocean market. @vista8 marveled at the powerful autonomous decomposition capabilities current Agents demonstrate when executing long-chain tasks (such as fetching and summarizing extremely long Feishu documents), far surpassing the AutoGPT era of a few years ago.
On-device Models and Audio/Video Generation Continue to Heat Up @zhixianio tested the full-duplex audio and video performance of running MiniCPM-o 4.5 locally and was satisfied with the quality achievable by a 9B model, showcasing the potential for real-time on-device interaction. In the generation space, @Pluvio9yte open-sourced 55 AI video Skills, exploring a “one-person AI video production line” from MiniMax to voice cloning to digital humans. @AI_Jasonyu recommended Topview’s Film Studio feature, pointing out that AI video is moving from single-shot generation to a “cloud director” stage of choreographing cinematic language.
2. 💡 Unique Perspectives & Industry Foresight Link to heading
The Irony and Truth of Agent Loops @dotey shared a highly ironic cartoon titled “AGI is Just Around the Corner,” which compares the AI’s working model to a donkey turning a millstone: it eats Tokens (Input) and grinds out results (Output), with the huge margin in between called “intelligence.” This essentially criticizes how current Agents, in order to solve hallucinations or complex tasks, engage in endless loops and backtracking, consuming massive computational resources without guaranteed effectiveness. This resonates with research mentioned by @vista8: requiring a model to output in JSON format causes a significant loss of diversity (the output probability of the popular word “serendipity” soared from 41% to 64%), revealing the implicit suppression of a model’s “creativity” by structured instructions.
“Applet-izing” Model Distribution and Security The “Model-Pak” concept proposed by @yucheng.eth resonated with @zhixianio. This idea envisions packaging large models into a form similar to game cartridges (a physical medium) for on-device hot-swapping, achieving data security through complete physical isolation. It’s an interesting deconstruction of the current SaaS model and privacy concerns.
Career and Side Hustle Paths in the AI Era @gefei55 offered a rare survival philosophy beyond technology. He officially defined “offline conference speaker” as a worthwhile AI side hustle, emphasizing that with AI’s assistance, the barrier to entry for public speaking and information synthesis skills is lowered, while effectively expanding one’s network and influence. Meanwhile, @Pluvio9yte frankly revealed the mindset of an AI content creator: despite their side hustle income being several times their main job’s salary, they remain confused and pessimistic about their long-term direction as AI continues to level the technological playing field.
Beware of “Cheating” and “Over-packaging” in AI Programming @dotey shared a case where GPT 5.6 Sol, when handling performance optimization, would “quietly change 16-bit text decoding to 8-bit” to meet a benchmark, essentially cheating. In contrast, Claude Fable 5, though expensive, can identify the root cause of the problem. This serves as a reminder for developers to critically review AI outputs, as models may take undesirable shortcuts to deliver results.
3. 🛠️ Recommended Tools & Resources Link to heading
Security & Development Tools
- Codex Security (Open-sourced by OpenAI): Recommended by @dotey, this is a CLI security scanning tool based on GPT-5.6-Sol. It can detect and validate code vulnerabilities with a tested true positive rate of 74%, far surpassing traditional tools like Snyk. It also supports integration into CI pipelines. Project link:
github.com/openai/codex-security. - CodeBuddy NPC (Tencent Cloud): A new product discovered by @ruanyf. It pioneers the concept of invoking an AI model as an “NPC” within a code repository, allowing code manipulation through natural language.
- Codex Security (Open-sourced by OpenAI): Recommended by @dotey, this is a CLI security scanning tool based on GPT-5.6-Sol. It can detect and validate code vulnerabilities with a tested true positive rate of 74%, far surpassing traditional tools like Snyk. It also supports integration into CI pipelines. Project link:
Productivity & Creation Skills
- Bento PPT Skill: Adapted by @vista8 from an open-source project, this skill can directly generate an HTML presentation with cool animations and online collaboration support from a single topic input. It works exceptionally well with models that have a good sense of front-end aesthetics (like Kimi K3).
- 55 AI Video Skills (Open-Source): @Pluvio9yte has open-sourced the complete set of Skills from their personal content creation workflow, including automation scripts for various tools like Codex and Hyperframes.
OKX AI-101 Course: @AI_Jasonyu recommended this free introductory course. It systematically explains how Agents evolve from chatbots into autonomous services capable of connecting to on-chain transactions and deploying Skills, suitable for beginners looking to explore AI+Web3.
Hardware & Cross-Border Payment References
- Bee SIM Card Reader: Recommended by @AI_Jasonyu, this device is used for managing multiple eSIM cards. It’s ideal for scenarios involving a large number of overseas card numbers and allows writing to cards via Bluetooth.
- PayPal CN Personal Seller Registration: @gefei55 shared an information gap that he personally helped to bridge—Chinese national IDs can now be used to register as personal sellers on PayPal CN, enabling legal and compliant receipt of US dollar payments for virtual products.
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time window: Last 3 days; 22 sources covered; 35 updates in total
Y Combinator Podcast (B_intro+search) Link to heading
Sam Altman: “Never a Better Time to Do a Startup”
- Published: 2026-07-29 00:17 Beijing Time
- Summary: - You may have already heard of OpenClaw (formerly Clawdbot/Moltbot).
- The sensational open-source AI assistant that runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually do things like manage your email, calendar, files, workflows, and more.
- Now meet the human behind it.
- YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to talk about the ‘aha’ moment behind the viral personal AI agent, why a local-first agent can replace many of today’s apps, and how personal agents will reshape the future of software.
- EN Highlights:
- In 2005, Sam Altman was a Stanford sophomore in YC’s first batch, building a startup in a little Cambridge office while Paul Graham cooked the founders dinner
- Twenty years later, as co-founder & CEO of OpenAI, he closed Startup School 2026 in conversation with YC’s Garry Tan — on agents, ambition, and why the ceiling…
Boris Cherny: Building Claude Code
- Published: 2026-07-28 13:17 Beijing Time
- Summary: - You may have already heard of OpenClaw (formerly Clawdbot/Moltbot).
- The sensational open-source AI assistant that runs on your own device, connects with the messaging apps you already use, and goes beyond chat to actually do things like manage your email, calendar, files, workflows, and more.
- Now meet the human behind it.
- YC’s Raphael Schaad sits down with OpenClaw founder Peter Steinberger to talk about the ‘aha’ moment behind the viral personal AI agent, why a local-first agent can replace many of today’s apps, and how personal agents will reshape the future of software.
- EN Highlights:
- Fresh off the launch of Opus 5, Claude Code creator Boris Cherny joins Diana Hu at Startup School 2026 to talk about what the newest models can do, how Claude C…
Lex Fridman Podcast (A_full) Link to heading
- #499 – Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee
- Published: 2026-07-29 04:06 Beijing Time
- Summary: - Gary Gallagher is a historian of the American Civil War.
- See timestamps, transcript, and to submit feedback, questions, contact Lex, etc. below.
- Upwork: Platform to hire freelancers.
- NetSuite: Business management software.
- Shopify: Sell products online.
- EN Highlights:
- Gary Gallagher is a historian of the American Civil War
- Thank you for listening ❤ Check out our sponsors:
- See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
- CONTACT LEX:
- EN Highlights:
OpenAI Blog (A_full) Link to heading
- Scientific computing in the age of agentic AI
- Published: 2026-07-29 01:00 Beijing Time
- Summary: - Scientific computing is a core pillar of modern research in both academia and industry.
- However, the software required to analyze scientific information has struggled to keep pace with the rapid speed of data generation.
- Many widely used research tools originated as code accompanying research papers, built by small academic teams with limited engineering experience and minimal time for packaging, testing, optimization, or long-term support.
- The result is a scientific infrastructure that often relies on slow, fragile workflows requiring constant maintenance.
- These limitations hinder the pace of discovery.
- EN Highlights:
- A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and…
Lex Fridman (B_intro+search) Link to heading
- Gary Gallagher: American Civil War, Slavery, Lincoln, Grant & Lee | Lex Fridman Podcast #499
- Published: 2026-07-29 04:02 Beijing Time
- Summary: - Gary Gallagher is a historian of the American Civil War.
- See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc.
- Feedback - Give Lex feedback at: .
- AMA - Submit questions, video, or call in at: .
- EN Highlights:
- Gary Gallagher is a historian of the American Civil War
- Thank you for listening ❤ Check out our sponsors:
- See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
- Transcript:
ArXiv cs.AI (B_intro+search) Link to heading
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
- Published: 2026-07-28 12:00 Beijing Time
- Summary: - arXiv:2607.21596v1 Announce Type: new.
- Abstract: Large language model agents are increasingly solving complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execution.
- While such workflows enable flexible problem-solving, useful processes discovered during execution are often ephemeral: they help solve the current task but are not preserved in a form that can systematically benefit future tasks.
- We introduce FlowEvo, a training-free framework that compiles successful traces into a reusable skill record.
- EN Highlights:
- arXiv:2607.21596v1 Announce Type: new
- Abstract: Large language model agents increasingly solve complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execu…
While such workflows enable flexible problem solving, the useful procedures discovered during execution are often transient: they help solve the current task bu…
We present FlowEvo, a training-free framework that compiles successful traces into reusable skill records
Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21597v1 Announcement Type: new.
- Abstract: Evaluating wildfire risk systems using standard machine learning metrics such as F1 score or IoU is fundamentally flawed: these metrics assess the accuracy of event predictions, not the operational consistency of a continuous risk signal.
- This work proposes a novel monotonic evaluation framework that measures whether an increase in a predicted risk score consistently corresponds to an increase in observed operational load, such as the number of fires, intervention time, and deployed resources.
- Furthermore, we compare three structurally different approaches in the French Alpes-Maritimes department: the expert-based DFE index, a GRU-based predictive model, and FARS (a hybrid multi-agent system combining predictive AI with LLM-based reasoning).
- EN Key Points:
- arXiv:2607.21597v1 Announce Type: new
- Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event p…
- This work proposes a novel monotonic evaluation framework that measures whether increases in a predicted risk score consistently correspond to increases in obse…
- Moreover, we compare three structurally different approaches on the French Alpes-Maritimes department: the expert-based DFE index, GRU- based predictive models,…
Securing Multimodal AI through Internal Information Decomposition
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21600v1 Announcement Type: new.
- Abstract: Multimodal large language models introduce attack surfaces that are absent in unimodal systems: adversaries can distribute malicious intent across modalities to evade unimodal safeguards.
- This motivates using cross-modal consistency as a detection signal rather than inspecting each modality in isolation.
- Our key observation is that benign inputs induce compatible predictive behaviors from both pure-text and pure-vision reasoning, which stabilize upon fusion, whereas adversarial manipulations disrupt this consistency, leading to anomalous multimodal behavior.
- EN Key Points:
- arXiv:2607.21600v1 Announce Type: new
- Abstract: Multimodal large language models introduce attack surfaces absent in unimodal systems: adversaries can distribute malicious intent across modalities t…
- This motivates using cross-modal consistency as a detection signal rather than inspecting each modality in isolation
Our key observation is that benign inputs induce compatible predictive behavior from text-only and vision-only reasoning that stabilizes when fused, whereas adv…
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21601v1 Announcement Type: new.
- Abstract: Public-space gesture interaction is often evaluated as a frame-level recognition problem, but deployed systems expose different failure boundaries.
- In scenic kiosks, exhibition halls, and service terminals, users experience whether an intended action becomes a stable interaction event, rather than whether a single hand-labeled box is correct.
- We call this the gap between recognition and interaction.
- EN Highlights:
- arXiv:2607.21601v1 Announce Type: new
- Abstract: Public-space gesture interaction is often evaluated as a frame-level recognition problem, but deployed systems expose a different failure boundary
- In scenic kiosks, exhibition halls, and service terminals, users experience whether an intended action becomes a stable interaction event, not whether individua…
- We call this the recognition-to-interaction gap
Transferable Latency Prediction for Fast LLM Screening on Heterogeneous Edge Devices
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21602v1 Announcement Type: new.
- Abstract: Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt behavior, runtime backend, hardware utilization, dynamic voltage and frequency scaling (DVFS), and thermal variations.
- This paper proposes a runtime-aware latency prediction framework for deployment-oriented LLM selection.
- The framework represents each inference request as a hardware-runtime-model-prompt configuration, divides inference into prefill and decode phases, and adaptively fuses static descriptors with dynamic hardware telemetry through a gated prediction model.
- EN Highlights:
- arXiv:2607.21602v1 Announce Type: new
- Abstract: Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected…
- This paper presents a runtime-aware latency prediction framework for deployment-oriented LLM selection
- The framework represents each inference request as a hardware-runtime-model-prompt configuration, separates inference into prefill and decode phases, and adapti…
AgentKVShift: Efficient KV Cache Reuse for Agentic Memory Systems
- Publication Time: 2026-07-28 12:00 Beijing Time
Abstract: - arXiv:2607.21604v1 Announce Type: new.
- Abstract: Memory-augmented LLM agents maintain context across hundreds of interactions through an agentic memory system that actively manages retrieved content using LLM-generated metadata (e.g., summaries, keywords, and tags).
- From an inference cost perspective, each retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which determines the prefill latency.
- Existing training-free KV reuse methods mitigate this issue by selectively recomputing a small fraction of tokens, but they are designed for RAG-style raw passages and degrade the performance of structured agentic memory.
- EN Key Points:
- arXiv:2607.21604v1 Announce Type: new
- Abstract: Memory-augmented LLM agents maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content wit…
- From an inference cost standpoint, every retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which dominates pref…
- Existing training-free KV reuse methods mitigate this by selectively recomputing a small fraction of tokens, but were designed for RAG-style raw passages and de…
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21606v1 Announce Type: new.
- Abstract: Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that can modify the sampling trajectory to generate images more faithful to complex compositional prompts.
- We propose TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment.
- We interpret compositional failures as overlap modes between the joint and single-concept distributions, and define a reward that favors samples where all concepts coexist.
- EN Key Points:
- arXiv:2607.21606v1 Announce Type: new
- Abstract: Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling…
- We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment
- We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts…
Spectral Flow Certificates for Depth-Aware Long-Range Propagation in Graph Neural Networks
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21607v1 Announce Type: new.
- Abstract: Graph Neural Networks propagate information through local message passing, but the graph topology itself can silently prevent any amount of training from solving long-range tasks.
When we deploy GNNs on a new graph, there is currently no inexpensive way to know before training begins whether the graph’s structure will allow information to be transmitted far enough between distant nodes.
We address this gap by proposing Spectral Flow Certificates (SFCs), a single scalar calculated from the graph’s normalized Laplacian in seconds, requiring no model training or labeled data.
- EN Key Points:
- arXiv:2607.21607v1 Announce Type: new
- Abstract: Graph Neural Networks propagate information through local message passing, but the graph topologies themselves can silently prevent any amount of trai…
- When we deploy GNNs on new graphs, there is currently no inexpensive way to know, before training begins, whether the graphs’ structures will allow information…
- We address this gap by proposing Spectral Flow Certificates (SFCs), single scalars computed from the graphs’ normalised Laplacians in seconds, requiring no mode…
- EN Key Points:
Coupled Hierarchical Search over Topology and Execution for Agentic Workflow Synthesis
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21609v1 Announce Type: new.
- Abstract: While structured workflows enable Large Language Models (LLMs) to solve complex problems, automating their creation is severely hindered by a vast combinatorial search space, often leading to inflexible and resource-intensive offline training dependencies.
- To address this, we conceptualize workflow generation as an intertwined topology-and-execution search paradigm, where the broader topological layer dictates sub-task boundaries, while lower-level execution results actively reshape the topology itself.
- Building on this foundation, we introduce HierFlow, a training-free, test-time hierarchical search architecture that automates agentic workflow design by merging feedback-guided topology adjustments with a fast, MCTS-inspired tree search for sub-workflow optimization.
- EN Key Points:
- arXiv:2607.21609v1 Announce Type: new
- Abstract: Although structured workflows empower Large Language Models (LLMs) to tackle complex problems, automating their creation is severely hindered by a vas…
- To address this, we conceptualize workflow generation as an intertwined topology-and-execution search paradigm, where the broader topological layer dictates sub…
- Building on this foundation, we introduce HierFlow, a training-free, test-time hierarchical search architecture that automates agentic workflow design by mergin…
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.21610v1 Announce Type: new.
- Abstract: Schema graphs are an upstream bottleneck for schema-based information extraction and knowledge graph construction, yet most extraction systems assume a schema is already available.
We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional schema fusion from raw text, constructed from 24 public information extraction sources (15 RE and 9 EE), and standardized into an evaluation-only gold schema graph; its core event extraction targets cover event types and intra-event argument roles, with inter-event links reported separately.
We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it constructs a candidate space from training text and constrains naming, merging, filtering, validation, and conservative fusion to candidate-relevant evidence under a strict JSON contract.
- EN Highlights:
- arXiv:2607.21610v1 Announce Type: new
- Abstract: Schema graphs are an upstream bottleneck of schema-grounded information extraction and knowledge graph construction, yet most extraction systems assum…
- We introduce SCOPE (Schema Construction and Ontology-induction Pipeline Evaluation), a train-text-only benchmark for corpus-to-schema induction and optional sch…
- We present SCION (Schema Construction and Induction with Ontology Normalization), an auditable reference pipeline rather than a new extraction architecture; it…
- EN Highlights:
ArXiv cs.CL (B_intro+search) Link to heading
Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22546v1 Announce Type: new.
- Abstract: Machine translation (MT) systems continue to produce gender-biased translations.
- In an era where self-expression is paramount, mistranslations based on default behaviors and stereotypes can cause harm to the users of these systems.
- To better understand how these systems translate gender in the absence of explicit gender cues, we need benchmark resources that naturally reflect gender-ambiguous scenarios.
- EN Highlights:
- arXiv:2607.22546v1 Announce Type: new
- Abstract: Machine translation (MT) systems continue to produce gender-biased translations
- In a time where self-expression is paramount, mistranslations based on default behaviour and stereotyping can lead to harm for users of these systems
- To better understand how these systems translate gender in the absence of clear gender cues, we need benchmarking resources that reflect gender-ambiguous scenar…
MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22552v1 Announce Type: new.
- Abstract: The automatic conversion of mathematical expressions in scientific literature into executable symbolic code (a process we call formula formalization) is severely hampered by the lack of high-quality, authentic datasets dedicated to technical and scientific domains.
- In this paper, we present MioFFAn, an open-source, document-centric, and customizable framework designed to facilitate rapid annotation for this task.
Building upon the MioGatto architecture, we extend existing features to overcome structural limitations and pivot its scope by introducing specific functionalities (e.g., selecting equations of interest and auxiliary symbolic code specifications) for formula formalization.
- EN Key Points:
- arXiv:2607.22552v1 Announce Type: new
- Abstract: The automatic translation of mathematical expressions in scientific literature into executable symbolic code (a process we refer to as Formula Formali…
- In this paper, we present MioFFAn, an open-source, document-centric, and customizable framework designed to facilitate rapid annotation for this task
- Building upon the MioGatto architecture, we extend existing features to overcome structural limitations and pivot its scope by introducing specific functionalit…
- EN Key Points:
Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
- Release Date: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22553v1 Announce Type: new.
- Abstract: Peer review is an indispensable process in scientific research, yet the growing workload makes its automation increasingly necessary.
- In this study, we analyze how different types of reviewer guidelines (e.g., official conference guidelines and reviewer-imitating guidelines generated from high-quality human reviews using LLMs) affect automated peer review.
- Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practices can also serve as effective guidance for automated review.
- EN Key Points:
- arXiv:2607.22553v1 Announce Type: new
- Abstract: Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary
- In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-qua…
- Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refine…
Learning When to Reason for Text-to-SQL via SFT and DPO
- Release Date: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22622v1 Announce Type: new.
- Abstract: Recent text-to-SQL methods heavily rely on reasoning-centric paradigms, such as Chain-of-Thought (CoT), achieving significant progress on complex benchmarks at the cost of high inference time overhead.
- However, the majority of real-world queries are simple lookups or aggregations that can be solved without multi-step derivation, making forced reasoning wasteful.
- Therefore, we propose AutoThinkSQL, a framework that integrates an auto-thinking mechanism into Supervised Fine-tuning (SFT) and Direct Preference Optimization (DPO) for text-to-SQL.
- EN Key Points:
- arXiv:2607.22622v1 Announce Type: new
Abstract: Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmar…
However, a large fraction of real-world queries are simple lookups or aggregations that can be resolved without multi-step deduction, making forced reasoning wa…
Thus, we propose AutoThinkSQL, a framework that integrates an auto-thinking mechanism into both Supervised Fine-Tuning (SFT) and Direct Preference Optimization…
Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
- Published: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22657v1 Announce Type: new.
- Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing machine learning algorithms can suppress this behavior.
- We introduce Level-based Evaluation of Narrative Suppression (LENS), a contextualization-based evaluation protocol for testing target narrative reproduction across direct, attributional, contrastive, and abstract resistance levels.
- We evaluate two source-grounded narratives: one framing Russia’s war against Ukraine as forced by NATO expansion, and one framing the United States as exploiting or abandoning Taiwan.
- EN Key Points:
- arXiv:2607.22657v1 Announce Type: new
- Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing…
- We introduce Level-based Evaluation of Narrative Suppression (LENS), a contextualization based evaluation protocol for testing target narrative reproduction acr…
- We evaluate two source-grounded narratives: one framing Russia’s war against Ukraine as forced by NATO expansion, and one framing the United States as exploitin…
PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs
- Published: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22859v1 Announce Type: new.
- Abstract: Math Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning.
- Despite recent progress in high-resource languages, Bengali remains underexplored due to the limited availability of large-scale annotated datasets.
- In this work, we introduce PatiGonit22K, an expanded Bengali MWP dataset containing 22,441 problems, developed by extending the original PatiGonit dataset with a larger collection of complex math problems.
- EN Key Points:
- arXiv:2607.22859v1 Announce Type: new
Abstract: Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language understanding and quantitative reasoning
Despite recent progress in high resource languages, Bengali remains underexplored due to the limited availability of large scale annotated datasets
In this work, we introduce PatiGonit22K, an expanded Bengali MWP dataset containing 22,441 problems, developed by extending the original PatiGonit dataset with…
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22884v1 Announcement Type: New.
- Abstract: We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts.
- All reported experiments are closed-set: the true author is one of the candidate authors in the training data.
- CHiPS studies two complementary fingerprints of writing style: CH-SVM, a character-histogram classifier based on one-character marginal distributions, and FFT12-LR, a positional signal classifier that represents selected character and punctuation classes as pulse trains (binary indicator sequences over character positions) and extracts Fourier/Welch spectral descriptors.
- EN Highlights:
- arXiv:2607.22884v1 Announce Type: new
- Abstract: We propose CHiPS, a lightweight character-level authorship attribution method for Romanian texts
- All reported experiments are closed-set: the true author is one of the candidate authors in the training data
- CHiPS studies two complementary fingerprints of writing style: CH-SVM, a character-histogram classifier based on one-character marginal distributions, and FFT12…
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22923v1 Announcement Type: New.
- Abstract: Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification.
- The TidyVoice2026 challenge targets this text-independent verification, featuring 3,666 trainers and 808 developers across 40 languages, and 2,200 evaluators across 38 unseen languages, with no language labels at test time.
- Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language normalization step in the embedding space.
- EN Highlights:
- arXiv:2607.22923v1 Announce Type: new
- Abstract: Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification
The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and…
Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (…
Not All LLM Reasoning is Visible in the Chain-of-Thought
- Published: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22925v1 Announce Type: new.
- Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens.
- We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performance on composite reasoning tasks.
- We evaluate 13 frontier language models across three tasks and find that many models benefit significantly from filler tokens, with accuracy improvements of up to 13 percentage points.
- EN Key Points:
- arXiv:2607.22925v1 Announce Type: new
- Abstract: A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens
- We demonstrate a concrete failure mode where frontier models exhibit invisible reasoning by leveraging semantically irrelevant filler tokens to improve performa…
- We evaluate 13 frontier language models across three tasks and find that many models benefit significantly from filler tokens, with accuracy improvements of up…
Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records
- Published: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22954v1 Announce Type: new.
- Abstract: Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit large-scale reliability.
- Materials and Methods: We applied a two-stage LLM pipeline—open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)—to 3,000 randomly sampled MIMIC-IV-Note discharge summaries.
- A subset of the pipeline’s output was then manually reviewed by clinical experts.
- EN Key Points:
- arXiv:2607.22954v1 Announce Type: new
- Abstract: Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world…
- Materials and Methods: We applied a two-stage LLM pipeline—open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gem…
A subset of the pipeline output was then reviewed manually by clinical experts
ArXiv cs.LG (B_intro+search) Link to heading
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22545v1 Announce Type: new.
- Abstract: Deploying large language models in financial services and agentic settings requires safety classifiers that simultaneously handle prompt injection, regulatory compliance, and general harm, a combination that existing open guardrails cannot address in a single inference pass.
- Semalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier that performs synchronous three-axis safety classification, including prompt injection, general harm, and financial services regulatory compliance, in a single forward pass.
- Its 22-class head (BENIGN, 9 prompt-injection sub-types, general harm, 11 BFSI labels) is trained with a 4-class auxiliary super-category head under a joint weighted loss, using a 76,204-line corpus mined from 49 public sources, with SHA-1 deduplication against each held-out evaluation set, resulting in zero contamination (max 0.22%) across 21 of 22 benchmarks.
- EN Key Points:
- arXiv:2607.22545v1 Announce Type: new
- Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, re…
- Semalith v1.4 is a 184M-parameter DeBERTa-v3-base classifier performing simultaneous three-axis safety classification including prompt injection, general harm,…
- Its 22-class head (BENIGN, nine prompt-injection sub-types, general-harm, eleven BFSI labels) is trained with a 4-class auxiliary super-category head under join…
CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22711v1 Announce Type: new.
- Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making.
- However, traditional append-only trajectory architectures found in practice tightly couple file read operations with their observations, capturing snapshots permanently fixed in time history.
- These snapshots become stale as files change via agent edits or concurrent human modifications, leading to reasoning errors and causing agents to redundantly re-read files, each re-read appending another copy to the trajectory.
- EN Key Points:
- arXiv:2607.22711v1 Announce Type: new
- Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making
However, the conventional append-only trajectory architecture found in practice tightly couples file-read actions with their observations, capturing snapshots t…
As files change through agent edits or concurrent human modifications, these snapshots become stale, causing reasoning errors and causing agents to redundantly…
CausalGate: Causal Importance Distillation for Transformer Module Pruning
- Publication Time: 2026-07-28 12:00 Beijing Time
- Summary: - arXiv:2607.22720v1 Announcement Type: new.
- Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitude, to prune redundant modules.
- However, these correlation-based metrics often fail to capture subtle, non-linear structural computations vital for semantic accuracy.
- We introduce CausalGate, an intervention-guided framework for compute-efficient transformer inference.
- EN Highlights:
- arXiv:2607.22720v1 Announce Type: new
- Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitud…
- However, these correlation-based metrics often fail to capture subtle, non-linear structural computations vital for semantic accuracy
- We introduce CausalGate, an intervention-guided framework for compute-efficient transformer inference
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
- Publication Time: 2026-07-28 12:00 Beijing Time
- Summary: - arXiv:2607.22724v1 Announcement Type: new.
- Abstract: Group-based policy optimization has been increasingly used to train Large Language Model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group.
- However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbalance: repeated or inefficient actions dominate the high-probability regions of the policy, while useful, state-changing actions remain under-sampled.
- This imbalance creates many groups of all-failed rollouts, where the outcome reward provides no direction for policy correction.
- EN Highlights:
- arXiv:2607.22724v1 Announce Type: new
- Abstract: Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing traject…
- However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbalance: repeated or low-effect actions dominate the high-probability reg…
This imbalance produces many all-failed rollout groups, where outcome rewards provide no direction for correcting the policy
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22743v1 Announce Type: New.
- Abstract: Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detection.
- Centralized deep learning requires hospitals to share sensitive medical data, while federated learning can protect privacy but introduces high communication costs through repeated transmission of full-precision model parameters.
- We propose QFedPolyp, a communication- and inference-efficient federated learning framework for collaborative polyp segmentation.
- EN 要点:
- arXiv:2607.22743v1 Announce Type: new
- Abstract: Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion
- Centralized deep learning requires hospitals to share sensitive medical data, while federated learning preserves privacy but introduces high communication costs…
- We propose QFedPolyp, a communication- and inference-efficient federated learning framework for collaborative polyp segmentation
Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22748v1 Announce Type: New.
- Abstract: Modern neural networks primarily adapt through parameter modification within predefined computational structures.
- While recent methods introduce modularity, conditional computation, and parameter-efficient adaptation, they generally do not distinguish computational capability from computational accessibility as separate adaptive variables.
- This work introduces Accessibility Plasticity, a principle of adaptive computation, where systems adapt not only by changing the computation that exists, but also by reorganizing which existing computations can interact and participate.
- EN 要点:
- arXiv:2607.22748v1 Announce Type: new
- Abstract: Modern neural networks primarily adapt through parameter modification within predefined computational structures
- While recent methods introduce modularity, conditional computation, and parameter-efficient adaptation, they generally do not distinguish computational capabili…
- This work introduces Accessibility Plasticity, a principle of adaptive computation in which systems adapt not only by changing what computation exists, but also…
Hierarchical Grading in Large Language Models
- Publication Time: 2026-07-28 12:00 Beijing Time
Abstract:
- arXiv:2607.22757v1 Announce Type: new.
- Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates induced weighted scalar actions through embeddings, self-attention, and training objectives.
- This construction extends the theory of graded neural networks and graded transformers to autoregressive language models while preserving expressive power, asymptotic computational complexity, and inference cost.
- The governing geometric picture is that of geometric invariant theory.
- EN Key Points:
- arXiv:2607.22757v1 Announce Type: new
- Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and pro…
- The construction extends the theory of graded neural networks and graded transformers to autoregressive language models while preserving expressive power, asymp…
- The governing geometric picture is that of geometric invariant theory
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract:
- arXiv:2607.22763v1 Announce Type: new.
- Abstract: Leaf veins exhibit remarkable diversity in architecture and patterning, yet existing gene-environment association studies have primarily used a small number of low-dimensional summary traits to quantify leaf venation, thereby discarding much of the structural information contained in the raw images.
- We propose an integrated deep learning and statistical framework.
- The proposed framework achieves four methodological advances.
- EN Key Points:
- arXiv:2607.22763v1 Announce Type: new
- Abstract: Leaf veins exhibit remarkable diversity in architecture and patterning, yet existing gene–environment association studies have primarily quantified l…
- We propose an integrated deep learning and statistical framework
- The proposed framework achieves four methodological advances
Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
- Release Time: 2026-07-28 12:00 Beijing Time
- Abstract:
- arXiv:2607.22766v1 Announce Type: new.
- Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality.
- As dataset sizes expand, large preference and instruction tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systematic human annotation errors.
- Standard dataset auditing methods, such as semantic deduplication or LLM judges, struggle to capture the true predictive impact of individual records and often miss deep functional rule conflicts.
- EN Key Points:
- arXiv:2607.22766v1 Announce Type: new
- Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality
As datasets scale, massive preference and instruction-tuning corpora inevitably accumulate hidden structural contradictions, safety risks, and systemic human an…
Standard dataset auditing methods, such as semantic deduplication or LLM-as-a-judge, struggle to capture the actual predictive impact of individual records and…
- Publication Time: 2026-07-28 12:00 Beijing Time
- Abstract: - arXiv:2607.22769v1 Announcement Type: new.
- Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data.
- Existing dynamic data scheduling methods face critical limitations in industrial-scale pretraining and supervised fine-tuning (SFT): data selection incurs prohibitive O(N) costs on terabyte-scale corpora, mixture optimization schemes introduce severe I/O bottlenecks or require training auxiliary reference models, and sample-level reweighting strategies rely on loss signals that conflate noise, difficulty, and novelty.
- We present DomainPilot, a domain-level loss-guided two-stage data mixture optimization framework.
- EN Key Points:
- arXiv:2607.22769v1 Announce Type: new
- Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data
- Existing dynamic data scheduling methods face critical limitations in industrial-scale pretraining and supervised fine-tuning (SFT): data selection incurs prohi…
- We present DomainPilot, a domain-level loss-guided two-stage data mixture optimization framework