🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-09-01
- 类型
- ai-daily
- 字数
- 7029
- 阅读时长
- 33 min
2026-09-01 AI Daily | ChatGPT Ad Revenue Exceeds $1 Billion, AI Platform Commercialization and Governance Accelerate Simultaneously Link to heading
OpenAI disclosed that ChatGPT’s annualized advertising revenue has reached $1 billion and has begun to open up advertising placements to more regions, indicating that the commercialization of AI products is moving from subscriptions and APIs to a more mature advertising model. At the same time, Meta’s content governance settlement once again reminds us that the growth and regulation of AI platforms will both be long-term variables.
📖 In-depth Guide to This Issue’s Watch List Link to heading
Today, three main threads are most noteworthy. The first is the continued deepening of LLM engineering into the infrastructure layer: from vector indexing to accelerate inference, Publisher-Adaptive content extraction, to more stable entity disambiguation, several papers are addressing “whether models can ingest data faster and more stably and output results.” The second is that evaluation and alignment are beginning to reflect on the methods themselves, with Rasch evaluation, Rubric-Guided RL, and research on emotional context and decision bias, which are worth a focused read for teams doing model evaluation and alignment. The third is the expansion of the boundaries of reasoning and knowledge application, with cross-language multi-hop question answering, improved internalization of mathematical examples, and medical EHR evidence-based question answering, all emphasizing “finding the right evidence first, then making a judgment.” Additionally, Meta’s long article on regulation and content governance is also highly recommended for management and product leads.
🌐 X Platform AI Hot News Link to heading
Topic 1: Tim Cook Bows Out as Apple CEO After 15 Years of Growth Link to heading
- Category: AI · News
- Overview: Trending time: 1 day ago, Related posts: 68000
- Summary: Tim Cook Bows Out as Apple CEO After 15 Years of Growth: Tim Cook Bows Out as Apple CEO After 15 Years of Growth Today wraps Tim Cook’s 15-year run as Apple CEO, during which the company soared from $350 billion to over $4.6 trillion in market value, launched hits like Apple Watch, AirPods, and Vis…
Topic 2: Google Unveils TimesFM-3 for Smarter Time Series Forecasting Link to heading
- Category: AI · News
- Overview: Trending time: 4 hours ago, Related posts: 702
- What happened: Google released TimesFM-3, a model for time series forecasting, featuring stronger generalization capabilities and more accurate future trend predictions.
- Why it’s important: This is important for the AI field because time series forecasting is a core capability in scenarios such as supply chain, finance, energy, and operations. If foundation models can lower the modeling barrier and improve cross-scenario effectiveness, it will expand the practical value of AI in business decision-making.
- Discussion overview: Discussions on X mainly focused on whether it has substantial improvements over its predecessors and traditional forecasting methods, whether it is general enough, its effectiveness on real business data, and whether the model is open, and whether its cost and deployment complexity are acceptable.
Topic 3: Musk Recalls Doubts That Nearly Stopped SpaceX Launch Link to heading
- Category: AI · Other
- Overview: Trending time: 8 hours ago, Related posts: 3600
- Summary: Musk Recalls Doubts That Nearly Stopped SpaceX Launch:
Topic 4: Meta Launches Paid Muse Code AI for Complex Coding Tasks Link to heading
- Category: AI · News
- Overview: Trending time: 4 hours ago, Related posts: 1900
- Summary: Meta Launches Paid Muse Code AI for Complex Coding Tasks: 🧠 GPT-5.6 Launches: My 12-Hour Test Daily · 2026-07-10 I want to put three things side by side today: the GPT-5.6 launch, Meta’s Muse Spark 1.1 API, and what I noticed after running Sol for about 12 hours. Meta priced the API at $1.25 per mi…
Topic 5: OpenAI Codex Reaches 25 Million Active Users with Paid Quota Reset Link to heading
- Category: AI · News
- Overview: Trending since: 2 days ago, Related posts: 15000
- Summary: OpenAI Codex Reaches 25 Million Active Users with Paid Quota Reset:
Topic 6: OpenClaw 2.0 Launches with Massive Update and 933 Contributors Link to heading
- Category: AI · News
- Overview: Trending since: 19 hours ago, Related posts: 7800
- Summary: OpenClaw 2.0 Launches with Massive Update and 933 Contributors:
Topic 7: Lionel Messi Retires from Argentina National Team Link to heading
- Category: AI · Sports
- Overview: Trending since: 8 hours ago, Related posts: 908000
- Summary: Lionel Messi Retires from Argentina National Team: 🚨🇦🇷 BREAKING: Lionel Messi RETIRES from international football! The Argentina captain will no longer represent the national team.
Topic 8: Tottenham Secure Adarabioyo Permanently and Mudryk on Loan from Chelsea Link to heading
- Category: AI · Sports
- Overview: Trending since: 1 day ago, Related posts: 127000
- What it is: Tottenham is rumored to have permanently signed Adarabioyo and loaned Mudryk from Chelsea, with the transfer news rapidly gaining traction on X.
- Why it’s important: High-traction sports transfer topics like this drive the demand for AI applications in public opinion monitoring, sports content generation, transfer probability analysis, and player valuation assessment.
- Discussion summary: Discussions on X primarily revolve around the authenticity of the transfer, the terms of the deal, and the actual impact of the two players on the team’s lineup. Disagreements are centered on the credibility of the sources and whether the move is cost-effective.
Topic 9: Newcastle Land Fernandez-Pardo and Loan Out Woltemade on Deadline Day Link to heading
- Category: AI · Sports
- Overview: Trending since: 9 hours ago, Related posts: 47000
- What it is: Newcastle United signed Fernandez-Pardo and loaned out Woltemade on transfer deadline day.
- Why it’s important: This event is a football transfer and has no direct connection to the field of artificial intelligence. It might be a misclassification or incorrect tagging by the platform.
- Discussion summary: Discussions on X are expected to focus on the impact of these two moves on Newcastle’s squad depth, player development, and team competitiveness. However, without representative tweets, the specific points of disagreement cannot be confirmed.
Topic 10: Overwatch Revives LE SSERAFIM Crossover with New Hero Skins Link to heading
- Category: AI · Sports
- Overview: Trending since: 7 hours ago, Related posts: 38000
- What it is: Overwatch has re-launched its collaboration event with the K-pop girl group LE SSERAFIM, releasing new hero skins.
- Why it’s important: While this event has no direct connection to artificial intelligence, its high popularity reflects the communication value of digital content, virtual character commercialization, and the fusion of gaming with pop culture.
- Discussion summary: The focus of discussion on X is likely on the design quality of the new skins, whether the collaboration content is worth purchasing, the reasons for restarting the event, and players’ opinions on repeated collaborations and commercialization strategies.
Topic 11: Liverpool Reject £80m Gakpo Bid from City, Shift Focus to Fernandez Link to heading
- Category: AI · Sports
- Overview: Trending since: 1 day ago, Related posts: 146000
- Summary: Liverpool Reject £80m Gakpo Bid from City, Shift Focus to Fernandez:
Topic 12: Kirkentine Meme Marks Charlie Kirk Anniversary Link to heading
- Category: AI · Entertainment
- Overview: Trending since:, Related posts: 1700
- What it is: On the X platform, a “Kirkentine meme” topic has emerged centered around the “Charlie Kirk anniversary,” with related content being widely shared and remixed.
- Why it’s important: Topics like this reflect how AI-generated content, meme propagation, and platform recommendation algorithms can amplify public issues and entertaining expressions. This provides valuable insight into understanding the role of AI in content distribution and public opinion shaping. Discussion Overview: The discussion primarily focuses on whether these memes are purely online jokes, political expressions, or a re-packaging of the individuals and related events involved; the point of contention is whether some consider it normal derivative content propagation, while others believe it’s using hot topics for opinionated output.
Topic 13:Fence Kiss Meme Revives Dating Money Debates Link to heading
- Category: AI · Entertainment
- Overview: Trending Time:, Related Posts: 355
- What Happened: The “Fence Kiss” meme went viral on the X platform, reigniting discussions about who should bear the costs during dating.
- Why It Matters: This reflects how AI-generated or propagated entertainment content influences social issues and public opinion, and also demonstrates the intersecting impact of AI with internet culture, emotional relationships, and consumer attitudes.
- Discussion Overview: The discussion primarily centers on whether dating expenses should be split, whether the initiator of the date should cover the costs, and the conflicts between gender roles and economic equality in contemporary dating.
Topic 14:Rick and Morty Fan Comic Captures Tearful Grandpa-Grandson Moment Link to heading
- Category: AI · Entertainment
- Overview: Trending Time:, Related Posts: 493
- What Happened: A fan comic depicting a tearful moment between grandpa and grandson in “Rick and Morty” garnered attention on the X platform, with approximately 493 related discussions.
- Why It Matters: This reflects how generative AI is lowering the barrier to creation for entertainment content like fan comics, and also raises concerns about character expression, creator rights, and content authenticity.
- Discussion Overview: The discussion focuses primarily on whether the comic was AI-generated, whether the emotional expression is natural, and whether AI fan creations should be considered artistic creation or an appropriation of the original work’s style and copyright.
Topic 15:Tesla Registers First Steering-Wheel-Free Cybercabs in Texas Ahead of Launch Link to heading
- Category: AI · News
- Overview: Trending Time: 23 hours ago, Related Posts: 10000
- Summary: Tesla Registers First Steering-Wheel-Free Cybercabs in Texas Ahead of Launch:
Today’s AI Public Opinion Summary on X Link to heading
Today’s main public opinion trend on X is the parallel progression of “whether AI capabilities are truly implemented” and “how AI-generated/amplified content continues to reshape platform virality”: on one side are time-series foundational models like Google TimesFM-3, and on the other are highly viral topics such as transfers, collaborative skins, memes, and fan comics, indicating that AI discussions have clearly spilled over into content distribution and public sentiment scenarios. Consensus generally converges on two points: foundational models, if they can truly enhance cross-scenario prediction capabilities, will have practical value; and the speed and reach of AI-related content dissemination are already significantly altering attention allocation on platforms. Disagreements mainly revolve around “whether it’s a genuine improvement” and “whether this content is creation, meme-playing, or opinionated output/commercial packaging,” with particularly strong debate regarding the actual business effects of TimesFM-3, whether AI fan creations cross boundaries, and whether memes are politicized. Potential risks include over-packaging technological advancements amplifying expectation gaps, platform recommendation mechanisms further amplifying misleading or divisive content, and copyright, authenticity, and content attribution issues continuously becoming points of friction.
💡 Influencer Insights Link to heading
Today’s influencer insights are temporarily unavailable. Recommended reading: Watch List in-depth content.
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time Window: Last 3 days; Covering 22 sources; Total 32 updates
Stratechery by Ben Thompson (A_full) Link to heading
- Meta Settles, A Framework For Regulating Content, The Rest of Big Tech
- Published: 2026-08-31 18:00 Beijing Time
- Summary: [To be translated] - Meta’s settlement makes sense for all parties, but the entire sage highlights why any solution to regulating technology feels off.
- $15 / month or $150 / year.
- Substantial analysis of the news of the day delivered via three weekly emails or podcasts.
- Stratechery Interviews.
- Interviews with leading public CEOs, private company founders, and discussions with fellow analysts.
- EN Key Points:
- Meta’s settlement makes sense for all parties, but the entire sage highlights why any solution to regulating technology feels off.
- EN Key Points:
OpenAI Blog (A_full) Link to heading
- A milestone in expanding access to AI
- Published: 2026-08-31 12:00 Beijing Time
- Summary: [To be translated] - In less than 200 days after launch, ChatGPT Ads has reached $1 billion in annualized revenue run rate.
- The platform is now used by tens of thousands of advertisers and continues to expand globally.
- Starting later today, advertisers can purchase ChatGPT ads directly via Ads Manager across India, Europe, the Middle East, and North Africa.
- Advertising is one pillar of OpenAI’s diversified business model, alongside consumer subscriptions, enterprise offerings, and usage-based APIs.
- Together, these offerings give people, developers, and businesses choice in how they access OpenAI products, including an advertising-supported free tier that helps keep ChatGPT available to more than 1 billion weekly active users.
- EN Key Points:
- ChatGPT Ads reaches $1 billion in annualized revenue run rate and expands globally, supporting broader access to AI through free and affordable options.
ArXiv cs.AI (B_intro+search) Link to heading
Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
- Published: 2026-08-31 12:00 Beijing Time
- Summary: [To be translated] - arXiv:2608.27459v1 Announce Type: new.
- Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge.
- Its DeepQA system defeated the strongest human Jeopardy!
- champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy.
- EN Key Points:
- arXiv:2608.27459v1 Announce Type: new
Abstract: In 2011, IBM’s Watson was something like a sealed capsule of its era’s queryable knowledge
Its DeepQA system defeated the strongest human Jeopardy
champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impos…
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27463v1 Announce Type: new.
- Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content.
- Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters.
- Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured.
- EN Key Points:
- arXiv:2608.27463v1 Announce Type: new
- Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’ outputs, and raters of human-generated content
- Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters
- Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured
Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27464v1 Announce Type: new.
- Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking.
Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three types of expected utility: instrumental (will it help me act better?), hedonic (will it make me feel better?), and cognitive (will it improve my understanding?).
- Each utility is estimated through a lens shaped by well-documented cognitive biases, including illusion of control, automation bias, unrealistic optimism, impact bias, overconfidence, and confirmation bias.
- EN Key Points:
- arXiv:2608.27464v1 Announce Type: new
- Abstract: This position paper argues that human-centered explainable AI (HCXAI) should incorporate insights from the psychology of information seeking
- Drawing on Sharot and Sunstein’s framework of information-seeking motives, we propose that people evaluate whether to engage with explanations based on three ty…
- ), hedonic (will it make me feel better
Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2608.27471v1 Announce Type: new.
- Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political debates, where public opinion is shaped.
- Spotting a fallacious argument requires contextual knowledge beyond its pure surface text.
- This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within the argumentative discourse.
- EN Key Points:
- arXiv:2608.27471v1 Announce Type: new
- Abstract: Fallacies are arguments that employ invalid reasoning, making their automatic detection critical in sensitive contexts such as high-stakes political d…
- Spotting a fallacious argument requires contextual knowledge beyond its pure surface text
This entails world knowledge pertaining to the subject matter under discussion, as well as knowledge of the relationships that exist between arguments within th…
LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27472v1 Announce Type: new.
- Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge.
- We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs).
- In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging.
- EN Key Points:
- arXiv:2608.27472v1 Announce Type: new
- Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offe…
- We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs)
- In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighted averaging
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27475v1 Announce Type: new.
- Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it.
- These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single trajectory.
We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonparametric, time-invariant coefficient fields.
- EN Key points:
- arXiv:2608.27475v1 Announce Type: new
- Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it
- These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single tra…
- We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE structure together with nonp…
- EN Key points:
Class-Based Heuristic Selection for Solving the Flying Block Puzzle
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2608.27476v1 Announce Type: new.
- Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to exploit the structural constraints that govern constrained spatial domains, causing search performance to degrade catastrophically on harder instances.
- We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearance-to-size constraints encountered in multi-agent path finding, autonomous vehicle navigation, and block relocation systems.
We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units are scarce, a formal kinematic taxonomy partitioning the state space into seven mutually exclusive classes with provably admissible heuristics based on vacancy ratio and goal-piece geometry, and a class-conditional tie-breaking mechanism that dynamically switches between depth-priority and vertical-distance ordering to overcome f-value plateaus.
- EN Key Points:
- arXiv:2608.27476v1 Announce Type: new
- Abstract: Heuristic search underlies planning in autonomous systems ranging from warehouse logistics to robotic navigation, yet generic heuristics fail to explo…
- We study this problem through the two-column Flying Block Puzzle, a rigorously NP-complete spatial planning microworld whose bottleneck geometry mirrors clearan…
- We introduce the Class-Based Heuristic A* (CBHA*) algorithm, which integrates a General Move Constraint to capture minimum displacement costs when vacant units…
- EN Key Points:
Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2608.27477v1 Announce Type: new.
- Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks.
- Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design do not yet fully capture the diversity and complexity of realistic mobile use.
- We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios.
- EN Key Points:
- arXiv:2608.27477v1 Announce Type: new
- Abstract: Graphical user interfaces have emerged as an important environment for evaluating autonomous AI agents on multimodal interactive tasks
Existing benchmarks such as AndroidWorld and MobileWorld provide strong foundations for mobile agent evaluation, but their application coverage and task design…
We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27480v1 Announce Type: new.
- Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial damage when infestations remain undetected.
- This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations in tea plantations.
- Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, with geographic coordinates recorded for spatial tracking.
- EN Key Points:
- arXiv:2608.27480v1 Announce Type: new
- Abstract: Tea plantations are vulnerable to Postelectrotermes militaris, commonly known as the Upcountry Live Wood Termite (ULWT), which can cause substantial d…
- This study proposes an IoT-enabled acoustic monitoring framework integrated with deep learning for early detection and severity assessment of ULWT infestations…
- Research Method: Audio signals were captured non-invasively from tea trunks using a high-sensitivity microphone connected to a Raspberry Pi-based IoT device, wi…
Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27482v1 Announce Type: new.
Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems.
- The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through conditional aggregation tests on admissible knowledge contexts.
- The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level.
- EN Key Points:
- arXiv:2608.27482v1 Announce Type: new
- Abstract: We study context localization for generalized level-based evaluation in knowledge-based systems
- The framework models situations where a structured nonnegative score, defined on facts, rules, cases, criteria or evidence units, is evaluated through condition…
- The generalized level measure maximizes a monotone set function over all contexts whose aggregated support reaches a prescribed level
ArXiv cs.CL (B_intro+search) Link to heading
Accelerating LLM Inference via Vector Index Based Output Embeddings
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27460v1 Announce Type: new.
- Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with large multilingual vocabularies.
- We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary projection with an HNSW-based vector index.
- The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retrieved logits into a sparse full-vocabulary tensor.
- EN Key Points:
- arXiv:2608.27460v1 Announce Type: new
Abstract: Large output embedding matrices create a significant memory bandwidth bottleneck during autoregressive decoding, especially for compact LLMs with larg…
- We reformulate the output projection followed by top-k token selection as a maximum inner product search over token embeddings and replace the dense vocabulary…
- The resulting output head retrieves only a small candidate set of high-scoring tokens and can be integrated into existing decoding pipelines by scattering retri…
- Release Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Pending Translation] - arXiv:2608.27461v1 Announce Type: new.
- Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts.
- This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding.
- To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal academic dialog benchmark.
- EN Highlights:
- arXiv:2608.27461v1 Announce Type: new
- Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts
- This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding
- To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal acad…
Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech
- Release Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Pending Translation] - arXiv:2608.27462v1 Announce Type: new.
Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging.
- While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples.
- This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases.
- EN Highlights:
- arXiv:2608.27462v1 Announce Type: new
- Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and context…
- While existing PLM- or LLM-based methods perform well, they typically apply a single reasoning process to all samples
- This overlooks fine-grained linguistic nuances and causes unnecessary computation for simpler cases
- Posted: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27465v1 Announce Type: new.
- Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user’s emotional state has become an important safety problem.
- We test whether emotional expression increases a model’s endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence).
- As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length.
- EN Highlights:
- arXiv:2608.27465v1 Announce Type: new
Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice accordin…
- We test whether emotional expression increases a model’s endorsement (encouragement to proceed) when a user, holding the same objective information, is overconf…
- As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating t…
PACE: Publisher-Adaptive Content Extraction via Agentic Automation
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Pending Translation] - arXiv:2608.27466v1 Announce Type: new.
- Abstract: Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability.
- General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, images, and tables.
- Direct LLM-based extraction offers greater flexibility, but incurs substantial cost and latency at scale, while manually engineered publisher-specific parsers can achieve high accuracy but require substantial human effort to build and maintain.
- EN Key Points:
- arXiv:2608.27466v1 Announce Type: new
- Abstract: Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and…
- General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, imag…
- Direct LLM-based extraction offers greater flexibility, but incurs substantial cost and latency at scale, while manually engineered publisher-specific parsers c…
UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
Published: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27467v1 Announce Type: new.
- Abstract: We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records.
- We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment).
- For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the full evidence set, exploiting the asymmetry between judging relevance in the abstract versus relative to a generated answer.
- EN Highlights:
- arXiv:2608.27467v1 Announce Type: new
- Abstract: We describe the UIC-AIHealth4All system for ArchEHR-QA 2026, a shared task on grounded question answering from electronic health records
- We participated in Subtasks 2 (evidence identification), 3 (answer generation), and 4 (answer-evidence alignment)
- For Subtasks 2 and 3, we propose an answer-first pipeline in which the model generates candidate answers citing specific note sentences before classifying the f…
- Abstract: [Translation Pending] - arXiv:2608.27467v1 Announce Type: new.
Select, Don’t Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27470v1 Announce Type: new.
- Abstract: Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs.
- State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selecting the correct one given context.
- Dual-encoder models optimize for both within a shared embedding space, forcing representations to balance high-recall retrieval with fine-grained selection, and they require trained retrievers, which are costly to maintain as knowledge graphs change.
- EN Highlights:
- arXiv:2608.27470v1 Announce Type: new
Abstract: Entity Disambiguation (ED) is a key task for constructing and using knowledge graphs
- State-of-the-art neural approaches commonly model ED as a single task, although it consists of two distinct subproblems: retrieving candidate entities and selec…
- Dual-encoder models optimize for both within a shared embedding space, forcing representations to balance high-recall retrieval with fine-grained selection, and…
XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2608.27481v1 Announce Type: new.
- Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually translate an entire example into one language.
- This hides failures at language boundaries inside the reasoning chain.
- We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence.
- EN Key Points:
- arXiv:2608.27481v1 Announce Type: new
- Abstract: Knowledge-intensive multi-hop question answering requires systems to select evidence and compose dependent facts, yet multilingual benchmarks usually…
- This hides failures at language boundaries inside the reasoning chain
- We introduce XHotpotQA, a controlled benchmark for cross-lingual knowledge composition over mixed-language evidence
INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2608.27501v1 Announce Type: new.
- Abstract: Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctness, raising the question whether models truly internalize mathematical concepts or merely memorize solution patterns.
In human mathematics education, example-based reasoning such as constructing counterexamples to test theorem boundaries reflects deep conceptual understanding, but remains underdeveloped in current LLMs.
- Enhancing this capability through preference optimization presents two key challenges: (1) the model’s limited example-based reasoning ability makes constructing effective preference pairs inherently difficult; and (2) capability acquisition is progressive, as the model must first learn to adopt this strategy before learning to apply it correctly.
- EN 要点:
- arXiv:2608.27501v1 Announce Type: new
- Abstract: Mathematical reasoning has seen rapid progress in large language models (LLMs), yet existing methods optimize predominantly for final-answer correctne…
- In human mathematics education, example-based reasoning such as constructing counterexamples to test theorem boundaries reflects deep conceptual understanding,…
- Enhancing this capability through preference optimization presents two key challenges: (1) the model’s limited example-based reasoning ability makes constructin…
A Survey on Rubric-Guided Reinforcement Learning for Language Models
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated]- arXiv:2608.27505v1 Announce Type: new.
- Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences.
- However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality.
- Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of reward design, feedback generation, and policy optimization.
- EN 要点:
- arXiv:2608.27505v1 Announce Type: new
Abstract: Reinforcement learning from human feedback (RLHF) has become the dominant paradigm for aligning large language models (LLMs) with human preferences
However, traditional RLHF relies on scalar reward signals that lack interpretability and fail to capture the multifaceted nature of response quality
Rubric-guided reinforcement learning addresses these limitations by introducing structured, interpretable evaluation criteria, or rubrics, as the backbone of re…
ArXiv cs.LG (B_intro+search) Link to heading
Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization
- Release Time: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2608.27507v1 Announce Type: new.
- Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in replicated copies of the same environment.
- However, its pooled team-entropy score measures only collective exploration and cannot identify policies that contribute non-redundant coverage.
- We introduce Marginal Coverage Credit for PGPSE (MCC-PGPSE), which combines leave-one-policy-out coverage with state-owner specialization to estimate policy-specific credit.
- EN Key Points:
- arXiv:2608.27507v1 Announce Type: new
- Abstract: Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in repli…
- However, its pooled team-entropy score measures only collective exploration and cannot identify policies that contribute non-redundant coverage
- We introduce Marginal Coverage Credit for PGPSE (MCC-PGPSE), which combines leave-one-policy-out coverage with state-owner specialization to estimate policy-spe…
- Release Time: 2026-08-31 12:00 Beijing Time
Summary: - arXiv:2608.27512v1 Announce Type: new.
- Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models.
- When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow creates a structural validation–deployment gap: because quantization is a many-to-one mapping over parameter space, source-precision certification does not guarantee behavioral equivalence in the deployed configuration.
- We formalize this gap through Quantization Behavioral Equivalence Classes (QBECs) and prove that QBEC membership does not imply behavioral equivalence, providing a theoretical basis for quantization-triggered backdoor attacks.
EN Key Points:
- arXiv:2608.27512v1 Announce Type: new
- Abstract: Post-training quantization is often treated as a semantically neutral optimization for edge deployment of Large Language Models
- When a full-precision source checkpoint is evaluated and quantization is applied downstream without equivalent re-evaluation, this workflow creates a structural…
- We formalize this gap through Quantization Behavioral Equivalence Classes (QBECs) and prove that QBEC membership does not imply behavioral equivalence, providin…
DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
- Publish Time: 2026-08-31 12:00 Beijing Time
- Summary: - arXiv:2608.27513v1 Announce Type: new.
- Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length.
- Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the KV cache in most layers with fixed-size recurrent states.
However, these recurrent states are commonly stored in FP32 and consume substantial GPU memory; their updates are memory-bandwidth bound and contribute significantly to decoding latency.
- EN Key Points:
- arXiv:2608.27513v1 Announce Type: new
- Abstract: Softmax attention stores key and value vectors for every preceding token, causing inference memory to grow with sequence length
- Recent language models incorporating Gated DeltaNet (GDN) or Kimi Delta Attention (KDA) reduce this cost by replacing the KV cache in most layers with fixed-siz…
- However, these recurrent states are commonly stored in FP32 and consume substantial GPU memory; their updates are memory-bandwidth bound and contribute signific…
- EN Key Points:
A Deeper Analysis of Block-Sparse Featurizers
- Published At: 2026-08-31 12:00 Beijing Time
- Abstract: [To be translated] - arXiv:2608.27515v1 Announce Type: new.
- Abstract: The recently introduced block-sparse featurizer (BSF; Fel et al., 2026) is similar to a sparse autoencoder (SAE), but its atomic unit is a small subspace (a block of directions) rather than a single direction.
- It is designed for features that live on low-dimensional manifolds, which are especially frequent in vision.
- This work studies the BSF’s strengths and weaknesses, finding how it still somewhat suffers from classic SAE failure modes, like feature splitting and composition.
- EN Key Points:
- arXiv:2608.27515v1 Announce Type: new
- Abstract: The recently introduced block-sparse featurizer (BSF; Fel et al., 2026) is similar to a sparse autoencoder (SAE), but its atomic unit is a small subsp…
- It is designed for features that live on low-dimensional manifolds, which are especially frequent in vision
- This work studies the BSF’s strengths and weaknesses, finding how it still somewhat suffers from classic SAE failure modes, like feature splitting and compositi…
When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
- Publication Time: 2026-08-31 12:00 Beijing Time
- Summary: [Pending Translation] - arXiv:2608.27518v1 Announce Type: new.
- Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by catastrophic forgetting and weight-disentanglement error.
- In the literature, these difficulties are merely treated separately and mitigated through a variety of solutions, while the geometry induced by the base optimizer is treated as an implementation detail.
- In this work, we show that the two difficulties are in fact two instances of the same phenomenon: a parameter update useful for one task shifts the model’s outputs on another.
- EN Highlights:
- arXiv:2608.27518v1 Announce Type: new
- Abstract: Continual learning (CL) and model merging (MM) both aim to obtain a single model that performs well across multiple tasks, challenged respectively by…
- In the literature, these difficulties are merely treated separately and mitigated through a variety of solutions, while the geometry induced by the base optimiz…
- In this work, we show that the two difficulties are in fact two instances of the same phenomenon: a parameter update useful for one task shifts the model’s outp…
Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics
- Publication Time: 2026-08-31 12:00 Beijing Time
- Summary: [Pending Translation] - arXiv:2608.27521v1 Announce Type: new.
- Abstract: Many dynamical processes unfold on the sphere but the default scientific machine learning architectures are Euclidean.
- Applying these architectures on a regular lat-lon grid causes problems: Cartesian convolutions become distorted at high latitude; 2D FFTs in Fourier neural operators incorrectly assume double periodicity; Cartesian positional encodings in ViTs distort spherical geodesic distances.
Recent work moves towards natively spherical primitives, including spherical convolutions (e.g., DeepSphere or DISCO), Spherical Fourier Neural Operators (SFNOs), and geodesic attention.
- EN Key points:
- arXiv:2608.27521v1 Announce Type: new
- Abstract: Many dynamical processes unfold on the sphere but the default scientific machine learning architectures are Euclidean
- Applying these architectures on a regular lat-lon grid causes problems: Cartesian convolutions become distorted at high latitude; 2D FFTs in Fourier neural oper…
- Recent work moves towards natively spherical primitives, including spherical convolutions (e.g., DeepSphere or DISCO), Spherical Fourier Neural Operators (SFNOs…
- EN Key points:
Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution
- Publication Time: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation pending] - arXiv:2608.27574v1 Announce Type: new.
- Abstract: Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple groups or consists of multiple objects.
- To date, a handful of multi-label graph learning methods exist, but none of them integrate training-time interpretation capability.
- While post-hoc graph explainers have been developed, they do not explicitly model label-dependent evidence sharing in multi-label graph learners, especially when label pairs are weakly or negatively associated.
- EN Key points:
- arXiv:2608.27574v1 Announce Type: new
- Abstract: Multi-label graph learning intends to capture the intrinsic complexity of real-world applications, where one sample is often related to multiple group…
- To date, a handful of multi-label graph learning methods exist, but none of them integrate training-time interpretation capability
- While post-hoc graph explainers have been developed, they do not explicitly model label-dependent evidence sharing in multi-label graph learners, especially whe…
Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27634v1 Announce Type: new.
- Abstract: Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN imposes the same neighborhood cardinality throughout the feature space.
- This assumption can be inadequate for data whose local geometry varies substantially across the underlying manifold.
- We introduce Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification (CARSANN), a geometry-driven framework that adapts the spatial support of each neighborhood according to local geometric complexity.
- EN Key Points:
- arXiv:2608.27634v1 Announce Type: new
- Abstract: Nearest neighbor classification relies fundamentally on how locality is defined, yet conventional $k$-NN imposes the same neighborhood cardinality thr…
- This assumption can be inadequate for data whose local geometry varies substantially across the underlying manifold
- We introduce Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification (CARSANN), a geometry-driven framework that adapts the spatial suppor…
More Data Cannot Break a Symmetry: Identifiability by Design
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: [Translation Pending] - arXiv:2608.27651v1 Announce Type: new.
- Abstract: Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus geometry bounds what any such alignment can identify, before data exist.
- The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling creates near-duplicates whose transposition is nearly free.
- We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention.
- EN Key Points:
arXiv:2608.27651v1 Announce Type: new
- Abstract: Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus…
- The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling create…
- We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention
Unsupervised Continual Learning with Growing Self-Organizing Maps and Synthetic Replay
- Published: 2026-08-31 12:00 Beijing Time
- Abstract: arXiv:2608.27662v1 Announce Type: new.
- Abstract: This work presents a generative continual learning framework based on growing self-organizing maps (GSOMs) that are augmented with learned distributional statistics as well as encoder-decoder models for class-incremental learning.
- The proposed approach enables exemplar-free replay using distributional statistical memory, which eliminates the need to store raw data.
- Each GSOM unit maintains its own mean, variance, and covariance estimates, which are subsequently used to generate synthetic samples for replay; in encoder-decoder configurations, these samples are then decoded back into the input space (via ancestral sampling) for subsequent training.
- EN Key Points:
- arXiv:2608.27662v1 Announce Type: new
- Abstract: This work presents a generative continual learning framework based on growing self-organizing maps (GSOMs) that are augmented with learned distributio…
- The proposed approach enables exemplar-free replay using distributional statistical memory, which eliminates the need to store raw data
- Each GSOM unit maintains its own mean, variance, and covariance estimates, which are subsequently used to generate synthetic samples for replay; in encoder-deco…