System translated (Gemini)

🤖 AI 速览

OpenAI is repositioning Codex as a new form of ChatGPT, shifting the focus from chat to deliverable workflows; enterprises are beginning to measure AI investments by output per dollar rather than token price. Meanwhile, custom chips and single-card models continue to lower deployment thresholds, …
📋 文章元数据
发布时间
2026-07-15
类型
ai-daily
字数
7432
阅读时长
35 min

2026-07-15 AI Daily | OpenAI Repositions Codex as the New ChatGPT, AI Competition Begins to Focus on Output Per Dollar Link to heading

OpenAI is repositioning Codex as a new form of ChatGPT, shifting the focus from chat to deliverable workflows. Enterprises are beginning to measure AI investment by output per dollar rather than token price. Meanwhile, custom chips and single-card models continue to lower the deployment threshold, with computing power and cost remaining central to the next phase of competition.

📖 Deep Dive: This Issue’s Watch List Link to heading

The most noteworthy trend today is “AI’s shift from a chat interface to a workflow entry point.” Discussions around the OpenAI Super App and Codex/ChatGPT resonate with “investment management in the agentic era”: the real metric is no longer the price per token, but the verifiable tasks, time saved, and scalable processes generated per dollar.

Engineering teams are advised to focus on the articles covering coding agents and inference efficiency: topics like how much context a Coding Agent needs, KV-Cache compression, local MoE inference, and the Index 1.9B small model all point to one thing—making agents controllable, low-cost, and deployable systems.

The third theme is trustworthy AI. Concepts like Ground Truth not being objective truth, AuditWeave’s evidence layer, distillation detection, and clinical time-series benchmarks all serve as reminders: beyond model capabilities, data provenance, audit trails, and evaluation boundaries are becoming hard requirements for the next phase of implementation.

🌐 AI Hot Topics on X Link to heading

Topic 1: OpenAI’s Codex and ChatGPT Work Hit 8 Million Users with Free Resets Link to heading

  • Category: AI · News
  • Overview: Trending: 1 day ago, Related Posts: 13,000
  • What happened: OpenAI’s Codex and ChatGPT Work have reportedly reached 8 million users, lowering the barrier to entry with features like “free resets.”
  • Why it matters: This indicates the rapid adoption of AI tools for programming and office scenarios. Free or low-cost strategies could further drive developers and enterprise users to migrate to AI-assisted workflows.
  • Discussion summary: Discussions on X focus on whether free tiers will change the payment models for AI tools, whether OpenAI is using this to expand its ecosystem advantage, and whether the claim of “free access to 100+ advanced models” is reliable or subject to limitations and marketing hype.

Topic 2: OpenAI’s GPT-5.6 Sol Tops Benchmarks Amid Efficiency Fixes and Bug Reports Link to heading

  • Category: AI · News
  • Overview: Trending: 2 days ago, Related Posts: 30,000
  • What happened: OpenAI’s GPT-5.6 Sol has reportedly taken the lead in several benchmarks, accompanied by progress in efficiency optimization and some user-reported bugs.
  • Why it matters: This shows that competition among frontier large models continues to revolve around performance, inference efficiency, and stability. If the results are true, it could influence decisions by enterprises and developers regarding model selection, cost control, and reliability.
  • Discussion summary: Discussions on X are centered on whether benchmark scores represent real-world capabilities, whether efficiency fixes can reduce usage costs, and if current bugs will affect actual deployment. Some users are optimistic about its leading performance, while others question the transparency and stability of the tests.

Topic 3: Samsung to Produce Custom AI Chips for Anthropic, Report Says Link to heading

  • Category: AI · News
  • Overview: Trending: 14 hours ago, Related Posts: 753
  • What happened: Samsung will reportedly produce custom AI chips for Anthropic, a collaboration that points toward large model companies developing their own computing power supply chains.
  • Why it matters: This indicates that leading AI companies are accelerating their efforts to reduce dependence on NVIDIA’s general-purpose GPUs, making custom ASICs, advanced manufacturing processes, HBM, and packaging capacity the new core of competition in AI infrastructure.
  • Discussion summary: Discussions on X focus on whether Samsung can close the gap with TSMC and SK Hynix in the AI chip supply chain, and whether companies like Anthropic and Amazon are building an “NVIDIA alternative.” The debate is whether custom chips can truly replace GPUs in the short term or will merely supplement specific inference and cloud scenarios.

Topic 4: Tencent Releases Quantized Hy3 AI Model for Single GPUs Link to heading

  • Category: AI · News
  • Overview: Trending: 5 hours ago, Related Posts: 186
  • What happened: Tencent has released a quantized version of its Hy3 AI model, designed to run on a single GPU.
  • Why it matters: This shows that large model deployment is moving further towards low-cost, localized, and edge computing scenarios, which helps lower the hardware barrier for enterprises and developers to use AI models.
  • Discussion summary: Discussions on X primarily focus on the extent of performance loss after model quantization, its competitiveness against similar open-source models, and the practical value of single-GPU deployment for small to medium-sized developers and local AI applications.

Summary of AI Public Opinion on X Today Link to heading

The main narrative today is that AI is shifting from a “competition of parameters and leaderboards” to a “competition of usability, cost, and implementation barriers.” Whether it’s OpenAI expanding its user base in programming and office work through free resets, or Tencent promoting quantized models that can run on a single card, these moves are seen as accelerating the popularization of AI. The general consensus is that low-cost usage, improved inference efficiency, and lower hardware barriers will continue to drive developers and enterprises to migrate to AI workflows. This will also prompt large model manufacturers to compete for ecosystem and computing power supply chains. The main points of disagreement are twofold: first, whether benchmarks and claims of “leading” performance can truly represent capabilities in real-world scenarios; second, whether custom chips and quantized models will genuinely change the industry landscape or will remain suitable only as supplementary solutions. The potential risks are that free and low-cost strategies may come with functional limitations or marketing exaggerations. If cutting-edge models lack stability, it will affect actual deployment. Furthermore, the restructuring of the computing power and chip supply chains could create new dependencies and competitive barriers.

💡 Influencer Insights Link to heading

Okay, as a senior AI industry analyst, I have carefully reviewed the tweets from various AI influencers over the past 24 hours. Here is a data-driven daily insight report.


AI Industry Daily Insights (Based on Analysis of X Influencer Tweets) Link to heading

1. Today’s Core Technology and Product Hotspots Link to heading

Today, influencers’ attention is primarily focused on the evolution of front-end model capabilities, the expansion of AI Agent product forms, and observations of emerging community phenomena.

  • Tencent’s Hunyuan Hy3 Model Attracts High Attention:

    • Both @ruanyf and @Pluvio9yte highlighted Tencent’s newly released flagship model, Hy3. @ruanyf pointed out that it has only 295B parameters, far smaller than the industry’s GLM 5.2 (744B). It focuses on “high speed and low cost,” with highly competitive API pricing (input $0.15/百万token). Its performance reaches or even surpasses that of GLM 5.1, making it suitable as a primary model for daily use.
    • @Pluvio9yte provided a more in-depth use case, emphasizing Hy3’s excellent comprehensive abilities in “engineering implementation, aesthetic judgment, and content curation.” A typical example is generating a complete, beautiful, and ready-to-launch “League of Legends” character showcase website in just 8 minutes with a two-sentence prompt, a task that would have previously required days of manual development.
  • The Continuous Evolution and Deep Integration of AI Agent Product Forms:

    • Growth and Evolution of Codex/Work: @dotey retweeted that the combined user base of Codex and Work has exceeded 800万 and continues to reset usage quotas, showing rapid growth momentum. He also explained in detail the differentiation and integration of Chat, Work, and Codex: Chat is for conversation, Work is an agent that delivers finished products across applications, and Codex is an agent focused on code repositories. The three share parts of the underlying Agent framework but have different purposes.
    • Claude’s Strategy Adjustment: Anthropic announced it would again extend access to its high-end model, Claude Fable 5, until July 19th. @dotey commented on this as a “child’s play decision,” which reflects Anthropic’s competitive posture and strategic wavering in its efforts to retain users under the pressure of OpenAI’s strong product offensive.
    • Apple vs. OpenAI: @dotey reported the news that Apple has formally sued OpenAI and its former employees for stealing trade secrets, alleging they were used to develop AI hardware. This lawsuit adds a new dimension of tension to the top-tier talent war and the battle for the AI hardware entry point among tech giants.
  • The Rise of an Emerging Community Phenomenon: Xiaohongshu REDSkill

    • @ruanyf keenly observed a unique cross-disciplinary phenomenon: the lifestyle platform Xiaohongshu (Little Red Book) is starting to build a REDSkill community, allowing users to upload and share AI Skill files. He commented that this is equivalent to combining a social media platform with a “Skill Hub,” an approach unprecedented globally. It might be Xiaohongshu’s way of paving the path for the platform’s AI transformation and increasing its technical content. For developers, it represents a massive new distribution channel to reach a vast user base.

2. Noteworthy Unique Perspectives and Industry Foresight Link to heading

  • The Myth of “Saving Tokens” (by @dotey): Baoyu conducted an in-depth analysis of the popular “telegraph-style Skills” (like the Caveman project, which claims to save 65% of tokens). He cited test results from JetBrains, pointing out that in actual programming tasks, output tokens are only reduced by 8.5%. He argues that this type of optimization is targeted at “chat scenarios,” whereas the real cost of an Agent lies in tool calls and system prompts. With the ongoing trend of falling API prices, it’s better to optimize context management and reduce rework on paths than to obsess over “talking like a caveman.” This pours cold water on the development of “money-saving” Skills, highlighting the primary and secondary contradictions in optimization.

  • Reshaping of Individual Roles and Capabilities (by @dotey, @vista8):

    • @dotey shared an internal observation from Anthropic: The “scaffolding” for Agents is becoming thinner, with the focus shifting from “controlling every step” to “designing collaboration between Agents.” He also emphasized that Agents can amplify individual capabilities but won’t automatically solve team coordination problems, and products might expand chaotically due to rapid individual trial-and-error.
    • @vista8 referenced a 1979 IBM slide, stating, “A computer can never be held accountable, therefore a computer must never make a management decision.” The same logic applies to AI; what will be truly scarce in the future is the ability to make decisions with incomplete information and take responsibility for the consequences. His observation of the phenomenon where “bosses want to distill employee experience into a Skill, but employees resist” also profoundly reveals the new conflicts over knowledge and value attribution within companies in the AI era.
  • The Potential of “On-Device Models” and “Model Cartridges” (by @zhixianio):

    • @zhixianio, through a hands-on review of Gemma 4 12B Coder, reached a counter-intuitive conclusion: despite community hype, the “ceiling” for a 12B parameter model is apparent, making it unsuitable for complex programs that require long, single-pass generation. This reminds us that the capability boundaries of small models remain clear; fine-tuning improves efficiency but doesn’t raise the upper limit.
    • He agreed with @geekbb’s vision of a “Model-Pak,” suggesting that future on-device models could be distributed and used like plug-in cartridges. This aligns with the on-device AI trend Google is heavily promoting (e.g., Gemma QAT models) and hints at a new form of offline AI applications.
  • Reference for Local Code Generation Model Evaluation: @zhixianio provided a detailed comparative review of Gemma 4 12B Coder vs. Qwen 3.6-35B-A3B MoE. The conclusion is that the 35B MoE model is currently the “sweet spot” for local code generation and can serve as an important reference for developers choosing a local model.
  • AI Video Editing Tools (Skills):
    • @Pluvio9yte and @dotey both noted ChatCut, an AI video editing Skill that integrates with Codex/Claude Code. It can automatically edit videos based on transcription, remove filler words, and more.
    • @dotey released his self-developed BaoCut, a subtitle transcription, translation, and editing Skill (Mac only). He specifically highlighted its approach to solving the “post-generation re-editing” problem for Agents by combining a CLI with a GUI to provide an Agent-friendly interface.
  • AI Video Replication and Integrated Creation: @Pluvio9yte shared the powerful capabilities of his “Video Replication Skill” enhanced by the Sol model and recommended using Tencent Hy3 for complex creative work that requires a balance of engineering, aesthetics, and content planning.
  • Open Source Projects and Learning Resources:
    • @vista8 recommended a Model “PK” Arena developed with AI, which allows for one-click comparison of text and front-end output from multiple models.
    • @vista8 recommended fireworks-tech-graph (8.5k stars), an open-source project for creating professional technical diagrams that has gained popularity through community promotion.
    • @vista8 also noted that the effectiveness of AI editing tools can be average, and building your own workflow with Listenhub CLI + Remotion remains a reliable option for pursuing higher quality.

📚 Appendix: Today’s Watch List Source Updates Link to heading

Timeframe: Last 3 days; covers 22 sources; 34 updates in total

Stratechery by Ben Thompson (A_full) Link to heading

  • The OpenAI Super App, ChatGPT = Codex, Whither Chat
    • Published: 2026-07-14 18:00 Beijing Time
    • Abstract: - OpenAI has refashioned Codex as the new ChatGPT; is the company abandoning the chat category they pioneered?
      • $15/month or $150/year.
      • Substantive analysis of the day’s news via three weekly emails or a podcast.
      • Strategy Interviews.
      • Interviews with leading public company CEOs, private company founders, and discussions with fellow analysts.
    • EN Highlights:
      • OpenAI has refashioned Codex as the new ChatGPT; is the company abandoning the chat category they pioneered

OpenAI Blog (A_full) Link to heading

  • How to manage AI investments in the agentic era

    • Published: 2026-07-14 18:00 Beijing Time
    • Summary:
      • OpenAI aims to make AI more accessible, capable, and affordable over time.
      • From GPT-4 to GPT-5.4, the price per million tokens has decreased by 97%.
      • However, token prices alone do not indicate whether AI is creating value.
      • Leaders should focus on useful work per dollar: tasks completed, time saved, improved decisions, and workflows ready to scale.
      • As teams shift from chat to longer-running workflows, administrators need clearer visibility into demand, spending, and risks.
    • EN Highlights:
      • Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
  • How data science teams use ChatGPT Work

    • Published: 2026-07-14 08:00 Beijing Time
    • Summary:
      • Learn how data science teams use ChatGPT Work to transform problems, dashboards, and raw data into analytical assets ready for review.
      • With ChatGPT Work, data science teams can more quickly turn disparate inputs into usable analytical assets.
      • Starting with dashboards, metric definitions, exports, experiment annotations, and business context, ChatGPT Work helps assemble first drafts of deliverables (including charts, explanations, source links, and audit questions) so teams can validate work and share with confidence.
      • Watch the on-demand webinar. Link to heading

      • Note: This webinar was recorded when these workflows were located in the former Codex application.
    • EN Highlights:
      • See how data science teams can use ChatGPT Work to build root-cause briefs, impact readouts, KPI memos, scoped analyses, and dashboard specs from real work inpu…
  • How sales teams use ChatGPT Work

    • Published: 2026-07-14 08:00 Beijing Time
    • Summary:
      • Learn how sales teams use ChatGPT Work to create pipeline briefs, meeting prep packets, forecast reviews, customer plans, and real stalled-deal diagnoses…
      • This article from the OpenAI blog explains how sales teams use ChatGPT Work to shape the broader AI and infrastructure landscape.
      • Following how sales teams use ChatGPT Work, it also provides practical implications for founders, operators, and investors.
    • EN Highlights:
      • See how sales teams can use ChatGPT Work to create pipeline briefs, meeting prep packets, forecast reviews, account plans, and stalled-deal diagnoses from real…

ArXiv cs.AI (B_intro+search) Link to heading

  • From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation

    • Published: 2026-07-14 12:00 Beijing Time
    • Summary:
      • arXiv:2607.09664v1 Announce Type: new.
      • Abstract: To provide structured and interpretable evaluations, we decompose image-based diagnostics into components following the Toulmin model of argumentation.
      • The model consists of claim, grounds, warrant, qualifier, rebuttal, and backing.
      • Consider claims generated by machine learning (ML) models for retinal diagnosis.
    • EN Highlights:
      • arXiv:2607.09664v1 Announce Type: new
  • Abstract: To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of argumentat…

  • This model consists of a claim, grounds, warrant, qualifier, rebuttal, and backing

  • Consider a claim generated by a machine learning (ML) model for retinal diagnosis

  • Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09665v1 Announce Type: new.
      • Abstract: Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions.
      • We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parsability Sensitivity Index (PSI), the corresponding range for answer parsability.
      • Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by over 30x between models, driven largely by compliance failures.
    • EN 要点:
      • arXiv:2607.09665v1 Announce Type: new
      • Abstract: Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions
      • We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced…
      • Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by…
  • Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09678v1 Announce Type: new.
      • Abstract: When LLM agents hand off information to one another, does the message format matter?
      • Two literatures disagree: format optimization work reports that structured messages can reduce cost without harming accuracy, while format constraint work finds that imposing structure degrades generation–and neither measures what happens when messages traverse multiple hops, where copy-fidelity rather than one-shot generation dominates.
      • We introduce a controlled relay testbed: a digest of twelve programmatically-generated atomic facts is serially re-encoded over six hops in five formats (free NL, precisely-instructed NL, JSON, triples, key-value), scored by a fixed strong grader against the programmatic ground truth, across two relay capability tiers, cognitive load conditions, and paired-fork error injection.
    • EN 要点:
      • arXiv:2607.09678v1 Announce Type: new
      • Abstract: When LLM agents hand off information to one another, does the message format matter
  • Two literatures disagree: format-optimization work reports that structured messages cut cost without hurting accuracy, while format-restriction work finds that…

  • We introduce a controlled relay testbed: briefs of twelve programmatically generated atomic facts are re-encoded hop-by-hop in five formats (free NL, precision-…

  • Boltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09689v1 Announce Type: new.
      • Abstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs–Boltzmann measure $\exp{-\beta E(\theta)}$, where the inverse temperature is the sample size $\beta=n$.
      • Three consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \emph{reduce} is literally a partition function $Z=\int\prod_k h_k,d\theta$ whose mode is precision-weighted (inverse-variance) pooling; frequentist consistency is the zero-temperature limit $T=1/n\to0$.
      • arXiv:2607.09689v1 Announce Type: new Abstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs–Boltzmann measure… Three consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \emph{…}.
    • EN Highlights:
      • arXiv:2607.09689v1 Announce Type: new
      • Abstract: To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs–Boltzmann measure…
      • Three consequences are exact in the Gaussian/linear case and first-order otherwise: disjoint chunks carry independent Boltzmann factors, so the MapReduce \emph{…
  • Interpreting Latent CoT Reasoning as Dynamical Systems

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09698v1 Announce Type: new.
      • Abstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate trajectories in hidden space at each step, unlike explicit CoT which follows a single, transparent reasoning trajectory.
      • Existing mechanistic approaches have shown compression, shortcuts, and superposition, but do not explain how reasoning evolves across latent steps.
      • To address this gap, we model the sequence of latent tokens as a trajectory in representation space and apply dynamical systems analysis to characterize the evolution of reasoning.
    • EN Highlights:
      • arXiv:2607.09698v1 Announce Type: new
      • Abstract: Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed candidate…
  • Existing mechanistic methods show compression, shortcuts, and superposition without explaining how reasoning evolves across latent steps

  • To address this gap, we model latent token sequences as trajectories in representation space and apply dynamical systems analysis to characterize the evolution…

  • YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract:- arXiv:2607.09706v1 Announce Type: New.
      • Abstract: Language models convert worded situations into numerical plans, and dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-value coefficients, then solve once.
      • For decisions that allocate real budget, effort, or clinical attention, this confidence is a failure mode: every objectified number is an assumption, and the optimal plan that presumes perfect correctness is fragile—a simulated calculation.
      • YUKTI changes the target of autoformulation.
    • EN Key Points:
      • arXiv:2607.09706v1 Announce Type: new
      • Abstract: Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objec…
      • For decisions that allocate real budget, effort, or clinical attention, that confidence is the failure mode: every objectified number is an assumption, and a pl…
      • YUKTI changes the target of autoformulation
  • GES-TSP: Graph Edge Sparsification for TSP

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract:- arXiv:2607.09708v1 Announce Type: New.
      • Abstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) is computationally expensive.
      • Researchers often employ graph sparsification methods to improve computational efficiency.
      • Traditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information.
    • EN Key Points:
      • arXiv:2607.09708v1 Announce Type: new
      • Abstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive
      • Researchers often employ graph sparsification methods to improve computational efficiency
      • Traditional sparsification methods typically rely on fixed heuristics and fail to fully exploit instance-specific structural information
  • The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation

  • Publication Time: 2026-07-14 12:00 Beijing Time

  • Abstract: - arXiv:2607.09709v1 Announcement Type: New.

    • Abstract: Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact.
    • We study the opposite signal: a deterministic, judge-free, ungameable filter – whether a generated project launches cleanly under a headless engine (strict-lau…).
    • Under this gate, rejection-sampling self-distillation compounds out-of-family generalization.
  • Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09713v1 Announcement Type: New.
      • Abstract: A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specific….
      • In this setting, policy generation by AI agents can be a credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated….
      • However, practical deployment is constrained by inference latency and compute footprint: large cloud-based models are often too slow, opaque, or data-sensitive….
  • Feedback-Coupled Memory Systems in Continuous Time

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09714v1 Announcement Type: New.
      • Abstract: The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which—the agent update operator $f_i$ and the environment update operator $\Psi$—are not axiomatically defined in the original framework.
  • To address this, $f_i$ is defined by Mechanism-Based Intelligence (MBI), where agents update locally through a decentralized price mechanism and economic principles, while $\Psi$ is defined by the Coupled Memory Graph Process (CMGP), a non-Markovian framework in which the environment is treated as a physical substrate that can coherently record and respond to trajectory history without external forcing.

  • The resulting continuous-time FCMS instantiation achieves Lyapunov global dissipativity governed by the computable threshold $4\beta^2 < 2\eta\mu\gamma^2$.

    • EN Highlights:
      • arXiv:2607.09714v1 Announce Type: new
      • Abstract: The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which - the agent…
      • To address this, $f_i$ is defined by Mechanism-Based Intelligence (MBI), where agents update locally through a decentralized price mechanism and economic princi…
      • The resulting continuous-time FCMS instantiation achieves Lyapunov global dissipativity governed by the computable threshold $4\beta^2 < 2\eta\mu\gamma^2$

ArXiv cs.CL (B_intro+search) Link to heading

  • CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

    • Release Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09880v1 Announce Type: new.
      • Abstract: Clinical time series are crucial for patient monitoring, risk assessment, and clinical decision support.
      • However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA).
      • Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully derive answers from irregularly timed observations.
    • EN Highlights:
      • arXiv:2607.09880v1 Announce Type: new
      • Abstract: Clinical time series are central to patient monitoring, risk assessment, and clinical decision support
      • However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Qu…
      • Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfull…
  • Index SLM Technical Report

    • Release Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09885v1 Announce Type: new.
      • Abstract: We introduce Index-1.9B, a series of open small language models developed by Bilibili.
  • The series includes four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chinese and English tokens; Index-1.9B-Pure, a control variant trained with the same recipe but with all instruction-like data strictly filtered from the corpus; Index-1.9B-Chat, which starts from the base model and undergoes supervised fine-tuning and direct preference optimization; and Index-1.9B-Character, which enhances the chat model with retrieval-augmented generation for few-shot role-playing customization.

  • Pre-training employed a Warmup-Stable-Decay learning rate schedule, where the concentration of curated data was substantially increased during the decay phase, along with a Norm-Head output layer to stabilize training at high learning rates.

  • EN Highlights:

    • arXiv:2607.09885v1 Announce Type: new
    • Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili
    • The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chines…
    • Pre-training employs a Warmup-Stable-Decay learning-rate schedule in which the concentration of curated data is raised substantially during the decay phase, tog…
  • RouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation

    • Release Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09908v1 Announce Type: new.
      • Abstract: Recommender systems increasingly face a choice among heterogeneous agents (collaborative filters, sequential models, content-based retrievers, and LLM-based rerankers), with no single agent being consistently the best.
      • We study this choice as task-aware agent ranking under cost constraints using RouteRec, a framework that compares request-level hard selection with item-level learned aggregation of four traditional recommender agents and one LLM reranking agent.
      • On MovieLens-1M, the full-quality oracle has substantial headroom (HR @ai_daily_20260510.md = 0.584), confirming that useful cross-agent signals exist.
    • EN Highlights:
      • arXiv:2607.09908v1 Announce Type: new
      • Abstract: Recommender systems increasingly face a choice among heterogeneous agents – collaborative filters, sequential models, content-based retrievers, and L…
      • We study this choice as task-aware agent ranking under cost constraints using RouteRec, a framework that compares request-level hard selection with item-level l…
      • On MovieLens-1M, the full quality oracle has substantial headroom (HR @ai_daily_20260510.md = 0.584), confirming that useful cross-agent signal exists
  • Global Merger-Arbitrage Forecasting with Language Models

    • Release Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09921v1 Announce Type: new.
      • Abstract: We propose a language model forecasting system for merger arbitrage, a specialized, high-stakes financial setting where the task is to predict the outcomes of announced merger and acquisition deals.
  • Unlike previous LLM judgmental forecasting work, which focused on broad mixed-topic benchmarks and short contexts such as news snippets, we investigate a setting that requires long-context reasoning over hundreds of pages of technical documentation.

    • Our system combines expert-guided context engineering with finetuning on hindsight-guided reasoning traces derived from historical deals.
    • EN Key Points:
      • arXiv:2607.09921v1 Announce Type: new
      • Abstract: We present a language-model forecasting system for merger arbitrage, a specialized high-stakes financial setting in which the task is to predict the o…
      • Unlike prior work on judgmental forecasting with LLMs, which has focused on broad mixed-topic benchmarks and short context such as news snippets, we study a set…
      • Our system combines expert-guided context engineering with finetuning on hindsight-guided reasoning traces derived from historical deals
  • Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09932v1 Announce Type: new.
      • Abstract: Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency towards hallucination poses significant risks in this high-stakes context.
      • This study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audiences.
      • The framework consists of 200 stratified trials drawn from the ClinicalTrials.gov database aggregated analysis, evaluated using audience-specific prompt templates and a six-dimensional faithfulness annotation scheme.
    • EN Key Points:
      • arXiv:2607.09932v1 Announce Type: new
      • Abstract: Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to…
      • This study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audienc…
      • The framework consists of 200 stratified trials drawn from the Aggregate Analysis of ClinicalTrials.gov database, evaluated using audience-specific prompt templ…
  • Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09957v1 Announce Type: new.
      • Abstract: This report studies on-device English-to-Traditional-Chinese subtitle translation in Taiwan under short-input, short-output, batch-size-one inference, low-latency, and privacy constraints.
      • These conditions limit the value of optimizations designed for long-context or high-throughput language model serving.
      • Starting with LMT-60-0.6B, preliminary analysis suggests that vocabulary projection becomes a more significant decoding-time cost after GGUF quantization reduces the relative cost of Transformer blocks.
    • EN Key Points:
  • arXiv:2607.09957v1 Announce Type: new

    • Abstract: This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inferen…
    • These conditions limit the value of optimizations designed for long-context or high-throughput language-model serving
    • Starting from LMT-60-0.6B, preliminary profiling suggests that vocabulary projection becomes a more important decode-time cost after GGUF quantization reduces t…
  • Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09999v1 Announce Type: new.
      • Abstract: We show that post-training quantization can silently alter how large language models reason, even when task accuracy is preserved.
      • Using a six-category failure taxonomy validated by two independent human annotators (Cohen’s κ = 0.906), we classify 30,000 chain-of-thought outputs from five instruction-tuned LLMs (3B–14B parameters) across three quantization precisions (FP32, FP16, NF4) and four reasoning benchmarks.
      • We find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable reasoning) shows significant size-dependent changes under NF4, with the two smallest models tested dropping sharply, while models 12B parameters and larger remain unchanged.
    • EN Highlights:
      • arXiv:2607.09999v1 Announce Type: new
      • Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved
      • Using a six-category failure taxonomy validated by two independent human annotators (Cohen’s $\kappa$ = 0.906), we classify 30,000 chain-of-thought outputs from…
      • We find that while accuracy is robust across precisions (maximum 3.1 pp drop), Hollow Convergence (correct answers reached through incomplete or unverifiable re…
  • Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.10020v1 Announce Type: new.
      • Abstract: We introduce FindMyText, an open-source Python package designed to efficiently assess whether a given text appears partially or entirely within a text corpus.
      • The tool builds upon existing document fingerprinting techniques but extends them with a novel mechanism to explicitly capture sequences of matching fingerprints.
      • By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text, not just text similarity.
    • EN Highlights:
      • arXiv:2607.10020v1 Announce Type: new
  • Abstract: We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text co…

  • The tool builds on prior techniques for document fingerprinting, but extends them with a novel mechanism to explicitly capture sequences of matching fingerprint…

  • By identifying such chains, the tool can more reliably detect near-verbatim copies of a given text rather than mere textual similarities

  • Efficiently Adapting Spoken Language Models for the Singaporean Context

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.10092v1 Announcement Type: New.
      • Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but their adaptation to sensitive domains is underexplored, especially when the original training data is inaccessible and use cases require multilingual, spoken-query interactions.
      • We apply an open-source SLM to the Singaporean Home Team context, covering five speech tasks across Singapore’s four official languages, combining LoRA fine-tuning, an alternative text QA dataset to prevent catastrophic forgetting, and a multi-task objective that adapts the CoBa re-weighting scheme to speech.
      • We also build HTD-multilingual-QA, a multilingual QA dataset containing 504,853 samples in both text and speech formats.
    • EN Highlights:
      • arXiv:2607.10092v1 Announce Type: new
      • Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the orig…
      • We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore’s four official languages, combining LoRA fine-tuning, a…
      • We also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form
  • Cost of Reasoning in non-English Languages: A Case Study on Japanese

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.10114v1 Announcement Type: New.
      • Abstract: Reasoning Language Models (RLMs) achieve their strongest performance when reasoning in English, the language for which reasoning-oriented training data is most abundant.
      • However, the reasoning trajectory serves as a clue for model interpretability and safety, and is useful in practice for both model users and developers.
      • Therefore, it is desirable to develop models that can reason in a user’s chosen language while still maintaining strong reasoning performance.
    • EN Highlights:
      • arXiv:2607.10114v1 Announce Type: new
      • Abstract: Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training d…
  • However, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers

  • Thus, it is desirable to be able to develop a model that reasons in a language of the user’s choice, while still maintaining strong reasoning performance

ArXiv cs.LG (B_intro+search) Link to heading

  • Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09666v1 Announce Type: new.
      • Abstract: Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-structured data.
      • However, there remains a lack of a systematic review of GNN-based methods across the entire knowledge graph technology pipeline.
      • To address this gap, we first propose a novel two-level classification framework for GNN-based knowledge graph technologies: the KG technologies pipeline and the GNN-based perspective.
    • EN Key Points:
      • arXiv:2607.09666v1 Announce Type: new
      • Abstract: Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-structured da…
      • However, there remains a lack of a systematic review about GNN-based methodologies across the entire knowledge graph technologies pipeline
      • To address this gap, we first propose a novel two-level taxonomy framework for GNN-based knowledge graph technologies: the KG technologies pipeline and GNN-base…
  • Position: Every Ground Truth is a Human Construction, not an Objective Truth

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09668v1 Announce Type: new.
      • Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models.
      • This position paper argues that ground truths are not neutral objective measurements naturally given, but instead that they are constructed by human and technological arrangements.
      • We believe that the machine learning community will benefit from elucidating and discussing these often invisible or unreported choices, and acknowledging that reference datasets are contingent rather than universal.
    • EN Key Points:
      • arXiv:2607.09668v1 Announce Type: new
      • Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models
      • This position paper argues that ground truths are not neutral objective measurements that are naturally given, but instead that they are constructed by arrangem…
  • We argue that the ML community will benefit from articulating and discussing these often invisible or unreported choices and acknowledging that reference data s…

  • AuditWeave: A Tamper-Evident, Auditor-Navigable Evidence Layer for AI-Assisted and Data-Transformation Workflows

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09682v1 Announcement Type: New.
      • Abstract: AI systems are increasingly used to assist consequential decisions in regulated domains such as auditing, finance, and healthcare.
      • This creates a recurring obligation: an organization must be able to reconstruct, after the fact, which evidence informed a given conclusion, and to show that the record of that reasoning has not been altered.
      • Existing tools address related but distinct problems - model observability, drift monitoring, governance reporting - and are built for machine-learning engineers of operating systems, not for reviewers who must trace a specific conclusion back to its supporting evidence.
    • EN Key Points:
      • arXiv:2607.09682v1 Announce Type: new
      • Abstract: AI systems are increasingly used to assist consequential decisions in regulated domains such as auditing, finance, and healthcare
      • This creates a recurring obligation: an organization must be able to reconstruct, after the fact, which evidence informed a given conclusion, and to show that t…
      • Existing tools address related but distinct problems - model observability, drift monitoring, governance reporting - and are built for the machine-learning engi…
  • Ablation, Statistical Inference, and Validation for KV-Cache Compression

    • Publication Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09683v1 Announcement Type: New.
      • Abstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, separating system codec differences from implementation differences through statistical validation methods.
      • Key findings reveal that while eigenbasis-based methods fail on heavy-tailed data due to covariance instability, they excel in structured regimes, where the effective semantic dimension ($d_{eff}$) adapts to the calibration budget rather than the true data rank.
      • (This is an abstract of the abstract, thank you).
    • EN Key Points:
      • arXiv:2607.09683v1 Announce Type: new
      • Abstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with B…
      • Key findings reveal that while eigenbasis-based methods fail on heavy-tailed data due to covariance instability, they excel in structured regimes, with the effe…
      • (this is an abstract of the abstract thank you )
  • SciML in the Wild: A Diagnostic Study of When Structural Priors Help and When They Hurt

    • Posted: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09684v1 Announce Type: new.
      • Abstract: Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs), and Universal Differential Equations (UDEs) are most effective when structural priors reflect reliable dynamic controls.
      • We ask what happens when this assumption is violated.
      • Using macroeconomic forecasting as a stress-test domain, we evaluate five model families (ARIMA, LSTM, NODE, PINN, and UDE) across 23 countries using sparse annual data, multiple time splits, and five random seeds.
    • EN Highlights:
      • arXiv:2607.09684v1 Announce Type: new
      • Abstract: Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs), and Univ…
      • We ask what happens when this assumption is violated
      • Using macroeconomic forecasting as a stress-test domain, we evaluate five model families, ARIMA, LSTM, NODE, PINN, and UDE, across 23 countries using sparse ann…
  • MawForge: Memory-Bounded Expert Materialization for Local Mixture-of-Experts Inference

    • Posted: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09686v1 Announce Type: new.
      • Abstract: Sparse Mixture-of-Experts (MoE) language models separate the total parameter count from the active computation per token, but local inference systems often still require the full model, key-value cache, runtime buffers, and operating system space to fit into fast memory.
      • MawForge tests a different systems hypothesis: local MoE serving can be made practical on constrained unified-memory machines by storing the full model on disk, keeping common tensors resident, and materializing routed expert tensors on-demand into a bounded execution cache.
      • The central finding is that MawForge is effective as a bounded execution mechanism and measurement substrate for local MoE inference, but not as a cache-maximization strategy.
    • EN Highlights:
      • arXiv:2607.09686v1 Announce Type: new
      • Abstract: Sparse Mixture-of-Experts (MoE) language models separate total parameter count from per-token active computation, but local inference systems often st…
      • MawForge tests a different systems hypothesis: local MoE serving can be made practical on constrained unified-memory machines by storing the full model on disk,…
      • The central finding is that MawForge is effective as a bounded execution mechanism and measurement substrate for local MoE inference, but not as a cache-maximiz…
  • Prioritizing Search Space Regions in the Low Autocorrelation Binary Sequences Problem

    • Posted: 2026-07-14 12:00 Beijing Time
  • Abstract: - arXiv:2607.09688v1 Announce Type: new.

    • Abstract: The Low Autocorrelation Binary Sequences (LABS) problem is a hard combinatorial optimization challenge with important applications in communication, signal processing, and satellite navigation.
    • This paper proposes a hybrid search framework that combines Thompson sampling with parallel self-avoiding walks to adaptively allocate computational effort across restricted classes of the LABS search space.
    • By modeling partitions as arms in a multi-armed bandit setting, the proposed method dynamically shifts search resources toward partitions that empirically produce higher merit factors, while maintaining exploration of less-sampled regions.
  • EN Highlights:

    • arXiv:2607.09688v1 Announce Type: new
    • Abstract: Low autocorrelation binary sequences problem (LABS) is a hard combinatorial optimization challenge with important applications in communications, sign…
    • This paper proposes a hybrid search framework that combines Thompson sampling with parallel self-avoiding walks to adaptively allocate computational effort acro…
    • By modeling partitions as arms in a multi-armed bandit setting, the proposed method dynamically shifts search resources toward partitions that empirically produ…
  • What Context Does a Coding Agent Actually Need to Act?

    • Release Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09691v1 Announce Type: new.
      • Abstract: A modern coding agent can hold an entire repository in its context window.
      • Most of its reading is wasted - and the interesting question is not how much context an agent can use, but what it actually \emph{needs}.
      • We study that question at the moment it matters most: when the agent must \emph{edit} code.
    • EN Highlights:
      • arXiv:2607.09691v1 Announce Type: new
      • Abstract: A modern coding agent can hold an entire repository in its context window
      • Most of its reading is wasted – and the interesting question is not how much context an agent can use, but what it actually \emph{needs}
      • We study that question at the moment it matters most: when the agent must \emph{edit} code
  • Reference-Based Distillation Detection in LLMs

    • Release Time: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09692v1 Announce Type: new.
      • Abstract: Model distillation—training on the output of a more powerful third-party model—is widely used to improve performance but raises concerns about unfair advantages and policy violations.
      • This raises a fundamental question: can we detect if one model has been distilled from another?
      • We show that while identifying the teacher model from an isolated student is very challenging, it becomes tractable in a reference-based setting: given a model and an earlier checkpoint from the same lineage, we can identify the teacher model used to train the later checkpoint.
    • EN Highlights:
      • arXiv:2607.09692v1 Announce Type: new
  • Depth-Entropy Guided Sampling for Training-Free LLM Reasoning

    • Published: 2026-07-14 12:00 Beijing Time
    • Abstract: - arXiv:2607.09693v1 Announce Type: new.
      • Abstract: Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, and reward signals.
      • Recent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-layer likelihoods and ignore the transformer’s internal forward-pass dynamics.
      • We introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal.
    • EN Key Points:
      • arXiv:2607.09693v1 Announce Type: new
      • Abstract: Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expens…
      • Recent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-la…
      • We introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal