System translated (Gemini)

🤖 AI 速览

Today’s focus is AI shifting from model demonstrations to executable workflows: OpenAI integrates ChatGPT into medical EHR and industry data, and Google DeepMind advances Gemini’s agentic video understanding, indicating that multimodal capabilities are entering vertical scenarios for …
📋 文章元数据
发布时间
2026-09-02
类型
ai-daily
字数
8367
阅读时长
40 min

2026-09-02 AI Daily Update | AI Begins Integrating into Real Workflows: EHR, Video Understanding, and Organizational Capability Layering Link to heading

Today’s focus is on AI shifting from model demonstrations to executable workflows: OpenAI connects ChatGPT to healthcare EHR and industry data, Google DeepMind advances Gemini’s agentic video understanding, indicating that multimodal capabilities are moving towards vertical industry implementation. Meanwhile, the output gap among frontier firms continues to widen, and AI competition is beginning to shift towards permissions, context, measurement, and safety boundaries.

📖 This Issue’s Watch List Deep Dive Link to heading

Today’s most noteworthy item for the Watch List is the main thread of “agents entering real workflows.” Google DeepMind is using Gemini to advance agentic video understanding, while OpenAI is integrating ChatGPT with healthcare organizations’ EHR and industry data. This shows that multimodal understanding and proprietary data connection are moving from demonstrations to vertical industry implementation. Product and platform teams are advised to pay close attention to their permission, context, and safety boundary designs.

The second point is the “capability gap in AI-native organizations.” Articles about frontier firms indicate that the per capita output tokens in companies with high AI usage have significantly diverged; combined with discussions from Lenny/Stratechery-esque sources on Nvidia, open-source models, and the compute economy, one can observe how AI investment is transitioning from tool procurement to an organizational operating system.

The third point is more research-oriented: several arXiv updates on LLM evaluation, interpretability, RAG debate analysis, causal discovery, and scientific agents collectively point to one question: models must not only generate but also be measurable, interpretable, and embeddable in serious reasoning tasks.

🌐 X Platform AI Hot News Link to heading

Topic 1: John Ternus Takes Over as Apple’s New CEO Link to heading

  • Category: AI · News
  • Overview: Hot for 8 hours ago, related posts: 124000
  • What happened: Apple announced John Ternus as the new CEO, ending Tim Cook’s 15-year tenure at the helm.
  • Why it’s important: This is seen as a signal of Apple’s strategic shift, especially concerning hardware iteration, foldable iPhones, and Apple’s ability to catch up in the AI race.
  • Discussion overview: Discussions on X mainly focus on the impact of the leadership change on Apple’s product roadmap and AI strategy. Debates include whether Ternus can drive more aggressive innovation and if Apple can accelerate AI adoption without disrupting supply chain and ecosystem stability.

Topic 2: Sutskever Warns of AI Security Risks in GPU Clouds Link to heading

  • Category: AI · News
  • Overview: Hot for 10 hours ago, related posts: 1600
  • What happened: Ilya Sutskever warned that AI deployed on GPU clouds could pose new security risks, raising concerns about the security of compute infrastructure.
  • Why it’s important: This is important because the discussion of AI risks is expanding from the models themselves to compute power, cloud environments, and deployment pipelines, with security boundaries becoming a prerequisite for AI’s scaled implementation.
  • Discussion overview: Discussions on X mainly focus on whether GPU clouds are sufficiently isolated, how AI operating environments should be audited and protected, and whether security responsibilities should lie with the model provider, cloud vendor, or application owner.

Topic 3: Sadie Sink Stars in Calvin Klein’s New Denim Campaign Link to heading

  • Category: AI · Entertainment
  • Overview: Hot for 1 day ago, related posts: 161000
  • What happened: Calvin Klein released a new denim campaign “Feel the Fit” starring Sadie Sink, featuring styles like 90s Straight, Low Rise Baggy, and High Rise Loose.
  • Why it’s important: High-profile brand marketing events like this are important for the AI field because they reflect the influence of generative content, virtual advertising creative, and celebrity image dissemination in commercial communication, which also impacts AI-related content production, aesthetic trends, and brand placement strategies.
  • Discussion overview: On X, discussions mainly revolve around Sadie Sink’s endorsement effectiveness, whether the ad style continues Calvin Klein’s traditional sexy route, and the brand’s shift from earlier marketing emphasizing inclusivity and identity expression to a strategy centered on celebrities and denim collections.

Topic 4: Tim Cook Steps Down as Apple CEO After 15 Years Link to heading

  • Category: AI · News
  • Overview: Hot for 2 days ago, related posts: 113000
  • What happened: Tim Cook stepped down as Apple CEO after 15 years, transitioning to executive chairman, with Apple’s hardware chief John Ternus taking over.
  • Why it’s important: This event relates to the strategic shift of one of the world’s most influential tech companies in the AI era, particularly whether Apple will accelerate the pace of AI product, system, and hardware integration.
  • Discussion Summary: The main discussion on X revolves around Cook’s achievement of leading Apple to a market capitalization of over $4 trillion and whether Ternus can guide Apple into the AI competition phase. Disagreements center on whether Apple’s past slow progress in generative AI is a weakness or if its conservative pace is more conducive to future implementation.

Topic 5: World Labs Unveils Atlas, First Multimodal World Model with Pixel-Perfect Control Link to heading

  • Category: AI · News
  • Summary: Hot Topic Time: 6 hours ago, Related Posts: 4100
  • What it is: World Labs announced Atlas, describing it as the first multimodal world model that supports pixel-perfect control.
  • Why it matters: This signifies that world models are evolving from simple generation to more controllable spatial understanding and editing, which has direct implications for embodied intelligence, 3D content generation, and interactive simulations.
  • Discussion Summary: The main discussion on X focuses on whether Atlas truly achieves its claim of “pixel-perfect control,” its differences from existing video/3D generation models, and whether world models are now entering a stage where they can be used in practical workflows.

Topic 6: Elon Musk Grants Free Grok Bot Token Reset to All Users Link to heading

  • Category: AI · News
  • Summary: Hot Topic Time: 2 hours ago, Related Posts: 2000
  • What it is: Elon Musk announced a free Grok bot token reset for all users.
  • Why it matters: This reflects adjustments in Grok’s product strategy, user access barriers, and cost control. It will also impact user experience, model invocation habits, and the promotion of AI features on the X platform.
  • Discussion Summary: The discussion on X is mainly centered on whether this can genuinely alleviate issues of insufficient quotas and restricted usage, and whether this “free reset” is more of a temporary fix, a marketing move, or an indication that Grok’s billing and quota mechanisms are being adjusted.

Topic 7: Manchester United Reject Everton Loan for Zirkzee on Deadline Day Link to heading

  • Category: AI · Sports
  • Summary: Hot Topic Time:, Related Posts: 10000
  • What it is: Manchester United rejected Everton’s loan request for striker Zirkzee on transfer deadline day.
  • Why it matters: This event has no direct connection to the field of artificial intelligence; it primarily reflects football transfer decisions and club roster management.
  • Discussion Summary: No representative tweets or specific discussion content are currently available, making it impossible to reliably summarize the differing views on the X platform. Existing information only indicates that the topic has received high attention.

Topic 8: Arsenal Bolster Defense with Key Signings but Lose Martinelli in Mixed Transfer Window Link to heading

  • Category: AI · Sports
  • Summary: Hot Topic Time: 14 hours ago, Related Posts: 30000
  • What it is: Arsenal strengthened their defense with several key signings during the transfer window but simultaneously lost Martinelli, creating a mixed situation of one in, one out.
  • Why it matters: Such high-interest sports topics demonstrate how real-time public opinion on X quickly converges around transfers, lineups, and player value. They are also often used to train and evaluate an AI’s ability to extract key event information, perform sentiment analysis, and identify differing public opinions.
  • Discussion Summary: The discussion mainly focuses on whether the defensive reinforcements are enough to boost their title-contending competitiveness and whether Martinelli’s departure will weaken the attack. Supporters emphasize immediate impact and squad depth, while critics worry about imbalance and a lack of subsequent reinforcements.

Topic 9: Manchester United Fans Frustrated as Transfer Window Nears Close Without Left-Back Link to heading

  • Category: AI · Sports
  • Summary: Hot Topic Time: 1 day ago, Related Posts: 88000
  • What it is: With the transfer window nearing its close, Manchester United fans have yet to see a new left-back signed, and frustration over the slow reinforcement process is growing on X.
  • Why it matters: The significance of such high-interest sports public opinion for the AI field lies in its reflection of typical patterns of real-time hot topic dissemination, emotional clustering, and fan base polarization. It can be used to train and evaluate capabilities in event understanding, public opinion analysis, and topic tracking.
  • Discussion Summary: The focus of discussion is mainly on whether the management and coaching staff missed the opportune time for reinforcement, whether the current left-back options are sufficient, and whether the blame should be placed on a slow transfer strategy or improper budget allocation. The disagreement lies between those who believe an immediate signing is necessary and those who think internal players can be used as a short-term solution.

Topic 10: Arsenal’s Ethan Nwaneri Joins Dortmund on Season Loan Link to heading

  • Category: AI · Sports
  • Summary: Hot Topic Time: 17 hours ago, Related Posts: 69000
  • What it is: Arsenal’s 19-year-old rising star, Ethan Nwaneri, has joined Dortmund on a season-long loan with no option to buy.
  • Why it matters: Such high-interest transfer events are important for AI as they can test a model’s ability to extract information from real-time sports news, perform entity recognition, distinguish between rumors and official announcements, and aggregate hot topics across platforms.
  • Discussion Summary: The focus on X is mainly on whether this loan is beneficial for Nwaneri’s development, whether Arsenal let him go too early, and whether Dortmund is continuing its strategy of “developing young players.” The point of contention is whether this is a solid development opportunity or a loss to Arsenal’s squad depth.

Topic 11: Chelsea Sells Enzo Fernández to Man City for £125m Record, Signs Lamine Camara for €55m Link to heading

  • Category: AI · Sports
  • Summary: Trending Time: 8 hours ago, Related Posts: 161,000
  • What it is: A transfer rumor is trending on X: Chelsea has reportedly sold Enzo Fernández to Man City for a record £125 million, while signing Lamine Camara for €55 million.
  • Why it matters: This type of high-value transfer directly impacts the squad structure of Premier League giants, discussions on financial fair play, and player valuations. It also serves as an important case study for observing club-building strategies and market pricing.
  • Discussion Summary: The discussion focuses on whether the transfer fee is reasonable, whether Chelsea has effectively refreshed its squad during its rebuild, and whether Man City’s record-breaking price to strengthen its midfield is worthwhile. Points of contention include some believing it’s a normal premium for a top-tier star, while others question the authenticity of the news and the logic behind the deal.

Topic 12: Golden Cybercabs Flood Austin Streets Ahead of Tesla Launch Link to heading

  • Category: AI · News
  • Summary: Trending Time: 1 day ago, Related Posts: 32,000
  • What it is: Tesla has deployed a large number of gold Cybercab-related vehicles on the streets of Austin, drawing attention to the impending launch of its autonomous taxi service.
  • Why it matters: This is seen as a major signal that Tesla is moving its autonomous driving capabilities from testing to real-world operation, which will determine whether AI-driven mobility can enter the stage of large-scale commercial deployment.
  • Discussion Summary: The main discussion on X revolves around whether this means the official launch of the Robotaxi service, whether the vehicles are for demonstration or operational use, and whether Tesla’s autonomous driving safety and regulatory compliance can withstand real-world road tests.

Topic 13: ChatGPT’s Playful Pronunciation Video Draws Laughs and User Gripes Link to heading

  • Category: AI · News
  • Summary: Trending Time: , Related Posts: 252
  • What it is: ChatGPT released a video demonstrating pronunciation in a playful manner, which drew laughter from users but also some complaints.
  • Why it matters: This reflects that AI voice and multimodal interactions are evolving from functional demonstrations to more personalized and entertaining expressions, but voice accuracy and user experience remain crucial evaluation criteria.
  • Discussion Summary: The discussion on X centers on the video’s humor, whether ChatGPT’s personified expression is natural, and the disagreements over pronunciation accuracy, product practicality, and excessive entertainment.

Topic 14: Naval Ravikant on Truth, Love, and Beauty as Perfection Link to heading

  • Category: AI · Entertainment
  • Summary: Trending Time: 10 hours ago, Related Posts: 747
  • What it is: Naval Ravikant initiated a discussion around the idea that “truth, love, and beauty are perfection,” sparking interest on X regarding his philosophical expressions and values in the AI era.
  • Why it matters: This type of discussion shifts the focus on AI from purely technical issues to the level of value judgments, involving how models understand and express core human concepts like “truth,” “beauty,” and “emotion.”
  • Discussion Summary: The focus on X is primarily on whether Naval’s perspective is applicable in the AI era, and whether AI can truly understand truth and generate aesthetics or can only simulate human expressions of these concepts. The disagreement centers on the balance between the inspirational nature of his philosophical insights and their practical applicability.

Topic 15: Trump Pushes Congress for Federal Film Production Incentives After Voight Meeting Link to heading

  • Category: AI · Entertainment
  • Summary: Trending Time: 1 day ago, Related Posts: 32,000
  • What it is: After a meeting with Voight, Trump is pushing Congress to support federal-level film production incentive policies in the United States.
  • Why it matters: This is important because it could affect film production costs, shooting locations, and the reshoring of the industry. It will also indirectly impact the adoption space for AI-generated content, visual effects, and media production tools in the film and television industry.
  • Discussion Summary: The discussion on X is focused on whether the policy can genuinely promote the reshoring of the US film industry, whether federal incentives will intensify competition for local subsidies, and the impact of such measures on traditional production, unions, and emerging AI-driven production workflows.

Today’s AI Public Opinion Summary on X Link to heading

Today’s main narrative revolves around the shift in AI from a competition over models to the implementation of products, infrastructure, and governance. Leadership changes at Apple, Tesla’s robotaxi, World Labs’ world model, ChatGPT’s voice demos, and Grok’s quota adjustments are all being discussed in the context of who can truly turn technology into stable, usable products. The general consensus is that the industry is no longer just looking at parameters and demos, but is placing more importance on hardware integration, cloud security, delivery cadence, and real-world usability. The main points of disagreement are twofold: first, whether these advancements are substantive breakthroughs or more marketing-driven narrative packaging; and second, how to balance speed and safety, especially concerning GPU cloud security, autonomous driving compliance, and the credibility of AI-generated content. The potential risks are also clear: overpromising amid high expectations, leaving security vulnerabilities in the deployment pipeline, and prematurely pushing immature capabilities into large-scale use under the pressure of capital and public opinion.

💡 Influencer Insights Link to heading

No influencer insights today. We recommend reading the in-depth content from the Watch List.

📚 Appendix: Today’s Watch List Update Source List Link to heading

Timeframe: Last 3 days; covers 22 sources; 36 updates in total

Stratechery by Ben Thompson (A_full) Link to heading

  • Nvidia Earnings, Dollars Per Gigawatt, Open and Hugging Face
    • Published: 2026-09-01 18:00 Beijing Time
    • Summary: 【待翻译】- Nvidia’s earnings were remarking and boring — two sides of the same coin.
      • Everything the company does is about avoiding a consolidated world.
      • $15 / month or $150 / year.
      • Substantial analysis of the news of the day delivered via three weekly emails or podcasts.
      • Stratechery Interviews.
    • EN Key Points:
      • Nvidia’s earnings were remarking and boring — two sides of the same coin
      • Everything the company does is about avoiding a consolidated world.

OpenAI Blog (A_full) Link to heading

  • How AI-native companies turn workflows into operating capability

    • Published: 2026-09-02 01:00 Beijing Time
    • Summary: 【待翻译】- Frontier firms (those with the top 10% of AI usage) now generate 8.3× as many output tokens per active user as typical firms, up from 2.6× in January.
      • The widening gap points to a deeper operating shift: leading firms connect agents to company context and tools, delegate more substantive work, and make successful workflows easier to repeat.
      • For leaders, the challenge is to turn that depth into work people can trust, measure, and improve.
      • Leaders should also leave room for experimentation, including use cases whose value is not obvious on the first try.
  • Their workflows differ, but the progression is instructive: teach an agent a stable process, give it persistent context as work changes, then let it carry opportunities into tested action.

    • EN Key Points:
      • Basis, Clay, and Exa Labs use AI agents to improve onboarding, account management, and developer integrations
      • See what enterprise leaders can apply.
  • Path to Astra: critical capabilities and frontier safeguards

    • Publication Time: 2026-09-01 21:00 Beijing Time
    • Summary: [To be translated] - It is the first model we are designating at this level, and requires stronger safeguards during development and before release.
      • Over the past several weeks, we have delayed parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions.
      • Based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework.
      • Based on retrospective testing, we believe our production safeguards at the time would have prevented the Hugging Face incident.
      • We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity.
    • EN Key Points:
      • Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for release.
  • Healthcare organizations can now connect EHR and additional industry data to ChatGPT

    • Publication Time: 2026-09-01 20:00 Beijing Time
    • Summary: [To be translated] - ChatGPT can now connect to trusted healthcare data, helping clinicians securely access patient context, medical research, and more.
  • This piece from OpenAI Blog explains how Healthcare organizations can now connect EHR and additional industry data to ChatGPT shapes the broader AI and infrastructure landscape.

    • It also surfaces practical implications for founders, operators, and investors following Healthcare organizations can now connect EHR and additional industry data to ChatGPT.
    • EN Highlights:
      • ChatGPT can now connect to trusted healthcare data, helping clinicians securely access patient context, medical research, and more.

Google DeepMind Blog (A_full) Link to heading

  • Introducing agentic video understanding with Gemini
    • Release Date: 2026-09-02 01:08 Beijing Time
    • Summary: [TO BE TRANSLATED] - Introducing agentic video understanding with Gemini.
      • This piece from Google DeepMind Blog explains how Introducing agentic video understanding with Gemini shapes the broader AI and infrastructure landscape.
      • It also surfaces practical implications for founders, operators, and investors following Introducing agentic video understanding with Gemini.
    • EN Highlights:
      • Introducing agentic video understanding with Gemini

Two Minute Papers (B_intro+search) Link to heading

  • GLM 5.3: Powerful AI Is Becoming Almost Free
    • Release Date: 2026-09-01 17:04 Beijing Time
    • Summary: [TO BE TRANSLATED] - ❤️ Check out Lambda here and sign up for their GPU Cloud:.
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
      • GLM 5.3: Powerful AI Is Becoming Almost Free.
    • EN Highlights:
      • ❤️ Check out Lambda here and sign up for their GPU Cloud:
      • 📝 GLM 5.3 Flash:
      • 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
      • Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…

ArXiv cs.AI (B_intro+search) Link to heading

  • Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

    • Published: 2026-09-01 12:00 Beijing Time
    • Summary: [Translation Pending] - arXiv:2608.27459v1 Announce Type: new.
      • Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge.
      • Its DeepQA system defeated the strongest human Jeopardy!
      • champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy.
    • EN Key Points:
      • arXiv:2608.27459v1 Announce Type: new
      • Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge
      • Its DeepQA system defeated the strongest human Jeopardy
      • champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impos…
  • Rating the Raters: Rasch Measurement Theory for LLM Evaluation

    • Published: 2026-09-01 12:00 Beijing Time
    • Summary: [Translation Pending] - arXiv:2608.27463v1 Announce Type: new.
      • Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content.
      • Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters.
      • Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured.
    • EN Key Points:
      • arXiv:2608.27463v1 Announce Type: new
      • Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content
  • Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters

    • Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured
  • Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.27464v1 Announce Type: new.
      • Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking.
      • Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three types of expected utility: instrumental (will it help me act better?), hedonic (will it make me feel better?), and cognitive (will it improve my understanding?).
      • Each utility is estimated through a lens shaped by well-documented cognitive biases, including illusion of control, automation bias, unrealistic optimism, impact bias, overconfidence, and confirmation bias.
    • EN Key Points:
      • arXiv:2608.27464v1 Announce Type: new
      • Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking
      • Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three ty…
      • ), hedonic (will it make me feel better
  • Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.27471v1 Announce Type: new.
  • Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped.

  • Spotting a fallacious argument requires contextual knowledge beyond its pure surface text.

  • This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within the argumentative discourse.

    • EN Key Points:
      • arXiv:2608.27471v1 Announce Type: new
      • Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political d…
      • Spotting a fallacious argument requires contextual knowledge beyond its pure surface text
      • This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within th…
  • LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: - arXiv:2608.27472v1 Announce Type: new.
      • Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge.
      • We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs).
      • In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging.
    • EN Key Points:
      • arXiv:2608.27472v1 Announce Type: new
      • Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offe…
  • We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs)

  • In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging

  • Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.27475v1 Announce Type: new.
      • Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it.
      • These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single trajectory.
      • We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonparametric, time-invariant coefficient fields.
    • Key Points:
      • arXiv:2608.27475v1 Announce Type: new
      • Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it
      • These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single tra…
      • We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonp…
  • Class-Based Heuristic Selection for Solving the Flying Block Puzzle

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.27476v1 Announce Type: new.
  • Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to exploit the structural constraints that govern constrained spatial domains, causing search performance to degrade catastrophically on harder instances.

    • We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearance-to-size constraints encountered in multi-agent path finding, autonomous vehicle navigation, and block relocation systems.
    • We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units are scarce, a formal kinematic taxonomy partitioning the state space into seven mutually exclusive classes with provably admissible heuristics based on vacancy ratio and goal-piece geometry, and a class-conditional tie-breaking mechanism that dynamically switches between depth-priority and vertical-distance ordering to overcome f-value plateaus.
    • EN Key Points:
      • arXiv:2608.27476v1 Announce Type: new
      • Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to explo…
      • We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearan…
      • We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units…
  • Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [PENDING TRANSLATION] - arXiv:2608.27477v1 Announce Type: new.
  • Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks.

    • Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use.
    • We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios.
    • EN Key Points:
      • arXiv:2608.27477v1 Announce Type: new
      • Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks
      • Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design…
      • We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios
  • Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: arXiv:2608.27480v1 Announce Type: new.
      • Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected.
      • This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations in tea plantations.
      • Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, with geographic coordinates recorded for spatial tracking.
    • EN Key Points:
      • arXiv:2608.27480v1 Announce Type: new
  • Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial d…

  • This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations…

  • Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, wi…

  • Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [Translation Pending] - arXiv:2608.27482v1 Announce Type: new.
      • Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems.
      • The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through conditional aggregation tests on admissible knowledge contexts.
      • The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level.
    • EN Key Points:
      • arXiv:2608.27482v1 Announce Type: new
      • Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems
      • The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through condition…
      • The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level

ArXiv cs.CL (B_intro+search) Link to heading

  • NLP-Driven Knowledge Extraction and Thematic Classification of Translated Ancient Indian Medical Texts

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [Translation Pending] - arXiv:2608.28608v1 Announce Type: new.
  • Abstract: Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques.

  • Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering.

  • The research here utilizes Natural Language Processing (NLP) methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions.

  • EN Key points:

    • arXiv:2608.28608v1 Announce Type: new
    • Abstract: Ancient Indian medical texts like Sushruta Samhita have extensive information on diseases, treatments, and surgical techniques
    • Yet, their ancient format and use of intricate vocabulary pose difficulties in accessibility and systematic ordering
    • The research here utilizes Natural Language Processing (NLP) methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in…
  • Parametric Multimodal User Memory: Storing What Captions Cannot Carry

    • Published at: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.28609v1 Announce Type: new.
      • Abstract: A personalized agent needs a user memory: a persistent model of who its user is.
      • Today it is almost always text – transcripts and captions retrieved by similarity.
      • This serves the captionable half of a person (“my cat is named Bibi”), but discards the perceptual half no caption can hold: how a voice sounds, how a face reads across age and lighting, how tired someone sounds.
    • EN Key points:
      • arXiv:2608.28609v1 Announce Type: new
      • Abstract: A personalized agent needs a user memory: a persistent model of who its user is
      • Today it is almost always text – transcripts and captions retrieved by similarity
  • This serves the captionable half of a person (“my cat is named Bibi”), but discards the perceptual half no caption can hold: how a voice sounds, how a face read…

  • Gurukul AI: An Interactive AI-Driven Educational Platform for Indian Education System

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Summary: - arXiv:2608.28611v1 Announce Type: new.
      • Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly trained on Western-centric data, making them ill-suited for regional curricula like India’s.
      • The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools.
      • In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula.
    • EN Key Points:
      • arXiv:2608.28611v1 Announce Type: new
      • Abstract: Recent advances in large language models (LLMs) like ChatGPT and LLaMA have transformed AI-driven education, but these systems are predominantly train…
      • The Indian education system is linguistically diverse, exam-oriented, and structured around standardized syllabi, not addressed by existing datasets or tools
      • In this work, we curate a syllabus-aligned QA dataset based on NCERT (National Council of Educational Research and Training) textbooks for classes 9-12, capturi…
  • STAGEET: Stage-wise Typed Edit Tagging for Grammatical Error Correction with Arabic as a Case Study

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Summary: - arXiv:2608.28614v1 Announce Type: new.
      • Abstract: Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence.
  • Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal the type of correction being made.

    • We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to correction categories.
    • EN Key Points:
      • arXiv:2608.28614v1 Announce Type: new
      • Abstract: Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rathe…
      • Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal t…
      • We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to…
  • From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education

    • Release Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [[OC_PH_TO_TRANSLATE]]- arXiv:2608.28619v1 Announce Type: new.
      • Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds.
      • Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue.
      • However, these logs are educationally difficult to use directly.
    • EN Key Points:
      • arXiv:2608.28619v1 Announce Type: new
      • Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while th…
      • Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue
  • However, these logs are educationally difficult to use directly

  • Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Summary: [To be translated] - arXiv:2608.28623v1 Announce Type: new.
      • Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering.
      • In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence.
      • However, for LMRMs no reliable method to measure sycophancy yet exists.
    • EN Key Points:
      • arXiv:2608.28623v1 Announce Type: new
      • Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before an…
      • In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence
      • However, for LMRMs no reliable method to measure sycophancy yet exists
  • MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson’s Disease Assessments

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Summary: [To be translated] - arXiv:2608.28624v1 Announce Type: new.
      • Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson’s disease is time-consuming and often depends on specialist expertise.
      • Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce factually correct and temporally consistent responses for structured longitudinal assessment data.
  • To address these limitations, we propose MA-RAG, a query-driven multi-agent retrieval-augmented generation framework that decomposes clinical reasoning into domain-specialized agents, combines structured fact extraction, and synthesizes clinically grounded summaries through a final verification stage.

    • EN Key Points:
      • arXiv:2608.28624v1 Announce Type: new
      • Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson’s disease is time-consuming and often depends on specialis…
      • Although large language models (LLMs) can generate natural language summaries, they frequently lack domain-specific clinical grounding and struggle to produce f…
      • To address these limitations, we propose MA-RAG, a query-driven multi-agent retrieval-augmented generation framework that decomposes clinical reasoning into dom…
  • Asymmetric Within-Document Predictive Learning for Scientific Document Representation

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.28625v1 Announce Type: new.
      • Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers.
      • We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations.
      • Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrastive baseline using the same section pairs.
    • EN Key Points:
      • arXiv:2608.28625v1 Announce Type: new
      • Abstract: We study predictive pretraining for scientific document representation using the discourse structure of papers
  • We propose SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict…

  • Experiments on RELISH, high-influence citation, SciDocs, and cite prediction show that plain predictive training is viable but weaker than a controlled contrast…

  • Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.28626v1 Announce Type: new.
      • Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation.
      • This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models’ training cutoffs.
      • Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format.
    • EN Highlights:
      • arXiv:2608.28626v1 Announce Type: new
      • Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation
      • This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Lea…
      • Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations,…
  • Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.28629v1 Announce Type: new.
  • Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM.

  • Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs.

  • Firstly, a BIM-to-Text method with component-balanced chunking is introduced to bridge BIM data with LLMs.

  • EN Key Points:

    • arXiv:2608.28629v1 Announce Type: new
    • Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design defects in BIM
    • Therefore, this study proposes an integrated framework to identify and repair various defects in BIM via domain-specific LLMs
    • Firstly, a BIM-to-Text method with component-balanced chunking is introduced to bridge BIM data with LLMs

ArXiv cs.LG (B_intro+search) Link to heading

  • ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.28771v1 Announce Type: new.
      • Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR).
      • While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal reasoning structure largely unoptimized.
      • Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token-level entropy drops within the thinking phase than incorrect ones.
    • EN Key Points:
      • arXiv:2608.28771v1 Announce Type: new
      • Abstract: Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning wit…
  • While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning pro…

  • Through empirical analysis across multiple model families, we identify a consistent pattern: correct reasoning trac es exhibit more frequent and larger token-le…

  • Unsupervised Latent Space Alignment with Hyperspherical Geodesic Matching

    • Publish Time: 2026-09-01 12:00 Beijing Time
    • Summary: [To be translated] - arXiv:2608.28840v1 Announce Type: new.
      • Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries.
      • These latent geometries are not directly compatible, yet they can be nearly the same up to some class of transformations.
      • While there exists many methods for alignment between different latent spaces, it is typically done using a set of shared sample correspondences, known as anchors.
    • EN Highlights:
      • arXiv:2608.28840v1 Announce Type: new
      • Abstract: Independently trained neural networks tend to encode the same data with similar latent geometries
      • These latent geometries are not directly compatible, yet they can be nearly the same up to some class of transformations
      • While there exists many methods for alignment between different latent spaces, it is typically done using a set of shared sample correspondences, known as ancho…
  • Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks

    • Publish Time: 2026-09-01 12:00 Beijing Time
    • Summary: [To be translated] - arXiv:2608.28843v1 Announce Type: new.
  • Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output oracle at the FFN branch; consider transformer FFN branches with GELU or SiLU activations under chosen-input raw-output access, without access to parameters, gradients, or internal activations; exploit a second-order leakage channel in which projected input Hessians form different mixtures of the same hidden symmetric rank-one factors induced by the FFN input weights.

    • We formalize resulting Hessian collection as a partially symmetric decomposition to establish conditions for local identifiability and stability to exploit vector-output stencil reuse to reduce the structural query cost by a factor of 16.
    • On independently trained CIFAR-10 vision transformers, only 16 projected Hessians, corresponding to 8193 black-box queries, recover the hidden FFN directions with average absolute cosine alignment above 0.94, with 95.1 % of GELU and 91.9 % of SiLU directions exceeding 0.90 alignment.
    • Key Points (EN):
      • arXiv:2608.28843v1 Announce Type: new
      • Abstract: We show that smooth two-layer feed-forward networks (FFNs) expose an additional structural model extraction channel under a chosen-input raw-output or…
      • We formalize resulting Hessian collection as a partially symmetric decomposition to establish conditions for local identifiability and stability to exploit vect…
      • On independently trained CIFAR-10 vision transformers, only 16 projected Hessians, corresponding to 8193 black-box queries, recover the hidden FFN directions wi…
  • Equivariant Sheaf Neural Networks: Learning Geometric Transport on Graphs

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.28853v1 Announce Type: new.
  • Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how vector information can be transformed as it moves across a graph.

    • We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring vector features while preserving exact Euclidean equivariance.
    • Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the edge transport itself.
    • EN Highlights:
      • arXiv:2608.28853v1 Announce Type: new
      • Abstract: Equivariant graph neural networks provide a principled way to model geometric systems, but efficient first-order architectures remain limited in how v…
      • We introduce \textsc{ESNN}, an Equivariant Sheaf Neural Network that enriches this interaction by learning directed, matrix-valued transport between neighboring…
      • Rather than increasing the order of the representation, ESNN keeps scalar and vector features first-order and places the additional geometric flexibility in the…
  • The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [Translation Pending] - arXiv:2608.28859v1 Announce Type: new.
      • Abstract: Reasoning models do not stop when they know the answer.
      • On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model’s own answer probability takes to settle, and how much of that excess is removable varies from problem to problem, so a global length penalty cannot take it out.
      • We take it out by internalizing a causal interpretability finding into the weights.
    • EN Highlights:
      • arXiv:2608.28859v1 Announce Type: new
      • Abstract: Reasoning models do not stop when they know the answer
  • On DeepSeek-R1-Distill-Qwen-7B the chain of thought runs about twice as long as the model’s own answer probability takes to settle, and how much of that excess…

  • We take it out by internalizing a causal interpretability finding into the weights

  • Conservative Hybrid Graph Networks for Process Systems with Learned Routing

    • Publish Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [[OC_PH_ABSTRACT_32]] - arXiv:2608.28896v1 Announce Type: new.
      • Abstract: Industrial process networks do not maintain a single effective topology while operating: streams are throttled or bypassed, and units move between idle, transition, and active regimes.
      • Models of such systems are typically trained on measured state trajectories while the operating mechanisms that generated them remain latent, and an unconstrained graph network can fit such a trajectory without assigning stable physical meaning to the recovered routing.
      • We address both problems with the Conservative Hybrid Graph Network (CHGN), which learns routing, regime assignment, and removal rates as data-driven surrogates and inserts them into a fixed transport equation, so that the mass balance holds by construction for any predicted routing.
    • EN Key Points:
      • arXiv:2608.28896v1 Announce Type: new
      • Abstract: Industrial process networks do not maintain a single effective topology while operating: streams are throttled or bypassed, and units move between idl…
      • Models of such systems are typically trained on measured state trajectories while the operating mechanisms that generated them remain latent, and an unconstrain…
      • We address both problems with the Conservative Hybrid Graph Network (CHGN), which learns routing, regime assignment, and removal rates as data-driven surrogates…
  • Off-Policy Evaluation for Semantic ID Recommenders: Does the Model’s Own Code Hierarchy Help?

    • Publish Time: 2026-09-01 12:00 Beijing Time
    • Abstract: [[OC_PH_ABSTRACT_33]] - arXiv:2608.28905v1 Announce Type: new.
  • Abstract: Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer, decoded autoregressively.

    • Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE).
    • We ask a simple question: can the model’s own SID tree serve as the action abstraction for that OPE?
    • EN Key Points:
      • arXiv:2608.28905v1 Announce Type: new
      • Abstract: Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer,…
      • Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE)
      • We ask a simple question: can the model’s own SID tree serve as the action abstraction for that OPE
  • Learning-Theoretic Foundation for General Coded Computing: The Straggler Setting

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [Translation pending] - arXiv:2608.28910v1 Announce Type: new.
      • Abstract: Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems.
      • However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and matrix multiplication, and typically rely on strict recovery thresholds.
      • These assumptions significantly limit their applicability to modern machine-learning workloads, particularly deep neural networks (DNNs), whose computations generally lack rigid algebraic structure and, in many applications, require only accurate approximations rather than exact recovery.
    • EN Key Points:
      • arXiv:2608.28910v1 Announce Type: new
  • Abstract: Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems

  • However, existing coded-computing schemes are predominantly designed for the exact recovery of highly structured computations, such as polynomial evaluation and…

  • These assumptions significantly limit their applicability to modern machine-learning workloads, particularly deep neural networks (DNNs), whose computations gen…

  • SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM Inference

    • Published: 2026-09-01 12:00 Beijing Time
    • Abstract: [To be translated] - arXiv:2608.28911v1 Announce Type: new.
      • Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length.
      • We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-Instruct with an affine quantizer is statistically indistinguishable from FP16 KV down to 2.322 code bits/value and collapses at 2.0 bits - a quality cliff in (2.0, 2.322] that reappears in generation-time quantization and multi-turn dialogue and transfers to Mistral-7B.
      • The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of mixing is grid interpolation, reaching average precisions uniform quantization cannot realize.
    • EN Key Points:
      • arXiv:2608.28911v1 Announce Type: new
      • Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length
      • We show that uniform KV quantization on a fractional-bit grid does not degrade gracefully: under a prespecified multi-seed statistical protocol, Llama-3.1-8B-In…
  • The cliff reframes importance-aware mixed precision: above it, eight model-internal importance indicators are statistically interchangeable, so the benefit of m…

  • RankShift: In-Database Detection and Explanation of Categorical Shifts

    • Publication Time: 2026-09-01 12:00 Beijing Time
    • Abstract: arXiv:2608.28922v1 Announce Type: new.
      • Abstract: A login service can receive its usual number of failed sign-ins while one source grows from 2% to 30% of them.
      • The same pattern appears in system logs when a rare event template becomes common while the message rate stays stable.
      • These events change which categories are active without changing how many events occur.
    • English Key Points:
      • arXiv:2608.28922v1 Announce Type: new
      • Abstract: A login service can receive its usual number of failed sign-ins while one source grows from 2% to 30% of them
      • The same pattern appears in system logs when a rare event template becomes common while the message rate stays stable
      • These events change which categories are active without changing how many events occur