🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-06-27
- 类型
- ai-daily
- 字数
- 7391
- 阅读时长
- 35 min
2026-06-27 AI Daily | GPT-5.6 Limited Preview, Frontier Models Enter Tiered Access Era Link to heading
Today’s main thread focuses on the changing release mechanisms for frontier models: GPT-5.6 is being previewed in tiers as Sol, Terra, and Luna, with the safety stack and access cadence becoming key points of focus. Meanwhile, the reproducibility of evaluations, knowledge boundaries, and the side effects of alignment are being re-examined. AI for Science and applications in high-risk scenarios continue to advance, as the industry shifts from purely pursuing capabilities to prioritizing reliability, constraints, and verifiable implementation.
📖 This Issue’s Watch List In-Depth Link to heading
There are three main themes to watch today: First, large model capabilities and evaluation governance. The preview of GPT-5.6 Sol shows that frontier models continue to advance in stratified layers of performance, cost, and safety stacks. At the same time, the reproducibility of LLM-as-Judge, Know2Guess knowledge boundary evaluations, and a paper on “helpfulness training weakening value retention” remind teams not to just look at leaderboards, but to also prioritize evaluation stability and alignment side effects.
The second theme is AI for Science penetrating deeper into physical, biological, and chemical systems: physics-guided CNNs, reinforcement learning in chemical reaction networks, and KG-TRACE for antimicrobial resistance prediction are all worth noting for their “neural network + domain constraints” engineering paradigm.
The third theme is reliable AI in socio-technical systems: financial anti-money laundering, media bias detection, algorithmic fairness, and remote sensing for flood identification demonstrate the higher demands for explainability, contextual modeling, and addressing structural bias as AI moves from general capabilities to high-risk scenarios.
🌐 AI Hot Topics on X Link to heading
Topic 1: OpenAI Investigates Codex Usage Limits Draining Too Quickly Link to heading
- Category: AI · News
- Overview: Trending for: 12 hours ago, Related posts: 1500
- What it is: OpenAI is investigating reports from some users that their Codex usage limits are being consumed too quickly.
- Why it matters: Codex is a key product form for AI programming assistants. Anomalies in usage metering can directly impact developer trust in the reliability, cost transparency, and productivity value of AI tools.
- Discussion summary: Discussions on X are focused on whether there is a bug in the limit calculation, whether subscriber benefits are affected, and whether OpenAI should provide clearer usage details and compensation. Some also question the sustainability of the cost model for high-intensity AI programming tools.
Topic 2: Meme Captures Semi-Bull Eye-Roll on AI Hype Link to heading
- Category: AI · Entertainment
- Overview: Trending for: 6 hours ago, Related posts: 56
- What it is: A meme poking fun at the AI boom from the perspective of a “semi-bullish eye-roll” is circulating on X, reflecting a sense of fatigue among some users regarding the overheated AI narrative.
- Why it matters: This indicates that beyond the continuous capital investment and product launches, the sentiment of the public and practitioners regarding technological value, commercialization pace, and bubble risks is becoming more complex.
- Discussion summary: The discussion centers on whether AI still possesses long-term disruptive potential or has been over-marketed. Supporters believe short-term noise doesn’t affect the long-term trend, while skeptics argue there is a gap between current valuations, hype, and actual user experience.
Topic 3: OpenAI Launches Limited GPT-5.6 Preview After U.S. Government Request Link to heading
- Category: AI · News
- Overview: Trending for: 17 hours ago, Related posts: 59000
- What it is: OpenAI has reportedly launched a limited preview deployment of GPT-5.6 following a U.S. government request. Access may initially be directed towards specific federal-related users or channels.
- Why it matters: If true, this indicates that the release of frontier AI models is increasingly influenced by government security, regulatory, and strategic needs. It could also change the cadence of model evaluation, open access, and commercial releases.
- Discussion summary: Discussions on X focus on whether the government should get priority access to new models, and whether this “federal gatekeeping” helps with safety testing or exacerbates concerns about a lack of transparency, concentration of power, and the exclusion of ordinary users.
Topic 4: OSWorld 2.0 Reveals AI Agents’ Limits on Hour-Long Tasks Link to heading
- Category: AI · News
- Overview: Trending for: 7 hours ago, Related posts: 245
- What it is: The OSWorld 2.0 benchmark shows that current AI agents still face significant capability bottlenecks in complex computer operation tasks that last for about an hour.
- Why it matters: This suggests that although AI agents have made rapid progress in short tasks and single-step automation, they have not yet reached a reliable and practical level in long-term planning, error recovery, interface understanding, and task persistence. This has significant implications for the real-world deployment of general agents.
- Discussion summary: Discussions on X are mainly focused on whether this benchmark more accurately reflects the upper limits of agent capabilities, and whether the failures of existing models stem from insufficient reasoning ability, unstable tool use, or overly complex evaluation tasks. Some also believe this indicates that industry expectations for the commercialization of “autonomous agents” need to be tempered.
Topic 5: Trump Administration Requests Staggered GPT-5.6 Release from OpenAI Link to heading
- Category: AI · News
- Overview: Trending time: 2 days ago, Related posts: 39,000
- What it is: According to multiple posts and reports, the U.S. Trump administration has requested that OpenAI delay and release GPT-5.6 in stages, initially making it available only to a select few government-vetted partners.
- Why it matters: This indicates that frontier AI models are being treated as dual-use technologies with cybersecurity and national security risks. The government may become more deeply involved in model release schedules, customer access, and security assessments, impacting the commercialization of AI and its open ecosystem.
- Discussion Overview: Discussions on X are focused on whether government review will become a de facto AI release license, whether closed-source frontier models will lose developer trust due to regulatory and access uncertainties, and whether low-cost open-source or open-weight models like DeepSeek and Qwen will see greater adoption as a result. The disagreement lies in whether this is necessary security governance or excessive intervention that weakens innovation and market competition.
Topic 6: ByteDance Unveils Seedance 2.5 with 30-Second 4K Video Generation Link to heading
- Category: AI · News
- Overview: Trending time: 5 hours ago, Related posts: 1,800
- What it is: ByteDance unveiled Seedance 2.5 at FORCE 2026, featuring support for up to 30 seconds of native video generation, up to 50 multi-modal reference inputs, local editing, 3D pre-visualization, and 4K video capabilities.
- Why it matters: This indicates that AI video generation is moving from short clip demonstrations to an industrialized production workflow with longer durations, higher consistency, and greater controllability, potentially accelerating adoption in advertising, film pre-visualization, and content production.
- Discussion Overview: Discussions on X are primarily focused on whether Seedance 2.5 can surpass competitors like Sora and Veo in character consistency, camera control, and commercial viability. There is also interest in its enterprise testing, global release schedule, copyright commercialization platform, and the impact of rapid AI video iteration on creators and copyright.
AI Public Opinion Summary on X Today Link to heading
The main theme in today’s public opinion is that as AI continues to iterate rapidly, trust, governance, and practical applicability are facing stricter scrutiny. The consensus is that frontier models and AI video still possess strong technological momentum. However, controversies over Codex quotas, bottlenecks in agents’ long-task capabilities, and the proliferation of AI-hype memes all indicate that users are growing more sensitive to cost transparency, real-world productivity, and excessive marketing. The primary point of disagreement is whether government intervention in the phased release of GPT-5.6 is a necessary security measure or if it will lead to a concentration of power, opaque market access, and diminished developer trust. The potential risk is that if the release schedules and usage rules for closed-source models continue to be opaque, developers may shift towards more open or lower-cost alternatives. At the same time, the rapid advancement of AI video capabilities will amplify issues surrounding copyright, creators’ rights, and content authenticity.
💡 Influencer Insights Link to heading
Based on intelligence from AI influencer tweets over the past 24 hours (with a focus on June 24-26, 2026), the following is a senior industry analysis report.
AI Industry Dynamics Daily: Model Regulation, On-Device Surge, and Agent Ecosystem Restructuring Link to heading
1. Key Technology Trends and Product Hotspots Watched by Influencers Today Link to heading
Today’s discussion centers on “Power Shift” and “Architecture Restructuring” and is extremely information-dense.
🔥 Core Events: The “Regulated Release” of GPT-5.6 and Anthropic’s Trade Accusations Link to heading
This is the most significant and intensely discussed event today, marking the official entry of AI industry competition into deep geopolitical waters.
- Tiered Release of GPT-5.6: @dotey provided a detailed analysis of the three versions of GPT-5.6 released by OpenAI (the flagship Sol, the everyday Terra, and the economy Luna). The main highlight is not the model’s capabilities, but the “gatekept approval” release mechanism required by the U.S. government. He stressed that this sets an unprecedented precedent, dramatically widening the gap between the company’s internal capabilities and those available to the public. Additionally, Sol’s Ultra mode, which uses multiple parallel sub-agents to handle complex tasks, points towards an architectural direction of “AI self-management.”
- Anthropic Pressures Alibaba: @dotey reported that Anthropic sent a letter to the White House accusing Alibaba of launching a massive distillation attack on Claude using 25,000 fake accounts (28.8 million interactions), aiming to steal core coding and reasoning capabilities to train the Qwen model. This move comes at an awkward time, as Anthropic’s own Fable 5 was globally recalled by the Department of Commerce due to a jailbreak vulnerability, showcasing a highly defensive business posture.
💻 The Agent Operating System Race: Codex Faces Restructuring, and Domestic Cloud Vendors Enter the Fray Link to heading
The Agent race is starting to evolve from tools to operating systems, while the barrier to entry for infrastructure is being significantly lowered.
- Codex Evolves into “Agent OS”: @dotey and @turingbook agree that Codex is becoming the operating system of the AI era, with all OpenAI staff having switched to using Codex. However, @Pluvio9yte reported a significant reduction in Token consumption, indicating it’s no longer an unlimited spree.
- Tencent Cloud EdgeOne Makers Released: Several bloggers (@AI_Jasonyu, @vista8) strongly promote this product, addressing pain points in Agent development and deployment (sandboxing, memory, concurrency). The intent is clear: to let developers focus on business logic while the platform hosts the infrastructure, which may impact the existing application models of Claude Code or Codex.
- Low-Code Alternative: @Pluvio9yte open-sourced video production Skills and tested the Volcano Coding Plan, which costs only 9.9 yuan/month, demonstrating that the Agent ecosystem is being rebuilt from top to bottom by domestic cloud vendors.
📱 On-Device Model Boom and Cost Reduction Link to heading
The dramatic contrast between rising hardware prices (Apple’s price hike theory shared by @zhixianio) and improved on-device performance.
- MiniCPM-o Potential: @zhixianio tested a local on-device full-duplex model and was amazed by the performance of the 9B model, indicating the impending popularization of on-device audio and video interaction.
- Gemma 4 Hands-on Test: @zhixianio extensively tested Google’s on-device model system (E4B, 12B Coder), concluding that the 12B scale still has a code ceiling, but its Quantization Aware Training (QAT) approach offers a new direction for on-device optimization.
- Content Replication: @Pluvio9yte open-sourced a pipeline for replicating hyperframes video styles, emphasizing that repetitive tasks like video editing should be fully automated.
2. Noteworthy Unique Perspectives or Industry Outlook Link to heading
- “Token Trap” and Energy Management: @gefei55 proposed that while Tokens are now infinite, human energy is limited. “How to avoid getting lost in the Token trap where anything can be done” has become a core ability humans need to master in the AI era.
- AI Programming’s “Hallucination” and True Value: @gefei believes that Vibe Coding no longer requires looking at code, but @ruanyf mentioned a surge in GitHub code commits (14x year-over-year), which may lead to an oversupply of AI-generated code. As @nishuang once pointed out, test cases are the new moat.
- IP Lock-in and “Fable 5” Blunder: @dotey observed that after Anthropic’s Fable 5 was forcibly removed by the U.S. Department of Commerce, co-founder Tom Brown replaced the “difficult to communicate with” Amodei to negotiate, and the model is expected to return to subscription. This shows that in the face of regulation, the internal technological idealism of AI companies must yield to political realities.
- Domestic Substitution of Toolchains and Contrast: @gefei55 pointed out that developing niche tool websites with a $150/month subscription fee is readily paid by European and American users; while domestically, Volcano Engine directly reduced Coding Agent to 9.9 yuan. The divergence of the two business ecosystems is becoming increasingly apparent.
- “Flavor” Adaptation Between Models: @lijigang suggested that heavy use of a certain model can lead to speaking with a “Claude flavor,” pointing out that human neural networks are highly “context-sensitive,” which hints at the risk of thought-shaping when choosing models.
3. Recommended Tools or Resources Link to heading
High-frequency tools shared by experts in practice today:
- Tencent Cloud EdgeOne Makers (Recommenders: @AI_Jasonyu, @vista8):
- Positioning: Agent deployment and operation platform. Solves the persistent problems of local execution and online crashes.
- Benefits: Free 500,000 Tokens for beta testing, extremely friendly to individual developers.
- PPT Master / Multi-Agent Collaboration Methods (Recommender: @dotey):
- Recommended Skill collaboration solutions, such as pipelines from interview analysis to article generation; and taught the technique of “merging multiple drafts” to prevent AI from missing details.
- Codex Orange Book / deobfuscate-javascript (Recommenders: @AI_Jasonyu, @dotey):
- If Codex feels like a black box, you can learn from @bozhou_ai’s open-source Orange Book tutorial, or use @dotey’s decompiler project to study the mechanisms of closed-source Agents.
- Doubao Seed 2.1 Pro Integration with Claude Code (Recommender: @Pluvio9yte):
- Provided a detailed tutorial on Volcano Engine API key integration with CC Switch model mapping, which is currently a popular path for low-cost access to high-performance models.
- VoxCPM2 (Voice) (Recommender: @AI_Jasonyu):
- Hailed as the “king of open-source speech,” it can generate audio from natural language descriptions and has garnered 22.9k GitHub Stars.
- Giffgaff Overseas Resources (Recommended by: @AI_Jasonyu):
- Provides highly practical tutorials and purchasing strategies for overseas SIM cards, an essential resource for registering and paying for international AI applications.
📚 Appendix: Today’s Watch List Update Source List Link to heading
Time window: Last 3 days; 22 sources covered; 32 updates in total.
Stratechery by Ben Thompson (A_full) Link to heading
- 2026.26: Summer Vibes
- Publication Time: 2026-06-27 01:00 Beijing Time
- Summary: - Welcome back to This Week in Stratechery!
- As a reminder, every Friday, we send out an overview of the content in the Stratechery bundle; highlighted links are free for everyone.
- Additionally, you have complete control over the content we send you.
- With that, here are some of our favorites from this week.
- A Vibe Coding Adventure. Being an analyst in the age of AI is exciting, especially because the questions seem so important.
- EN Key Points:
- Welcome back to This Week in Stratechery
- As a reminder, each week, every Friday, we’re sending out this overview of content in the Stratechery bundle; highlighted links are free for everyone
- Additionally, you have complete control over what we send to you
- If you don’t want to receive This Week in Stratechery emails (there is no podcast), please uncheck the box in your delivery settings
OpenAI Blog (A_full) Link to heading
- Previewing GPT-5.6 Sol: a next-generation model
- Publication Time: 2026-06-26 18:00 Beijing Time
- Summary: - We are beginning a limited preview of the GPT-5.6 series: Sol, our flagship model; Terra, a balanced model for everyday tasks; and Luna, a fast and affordable model.
- Terra offers performance competitive with GPT-5.5 at half the price, while Luna provides powerful capabilities at the lowest cost.
- GPT-5.6 Sol is being launched with our most robust security stack to date.
- We have strengthened protections against high-risk activities, sensitive network requests, and repeated abuse. We have also spent weeks identifying vulnerabilities, stress-testing our systems, and hardening them against real-world attacks.
- We believe in broad access and plan to make GPT-5.6 Sol, Terra, and Luna generally available in the coming weeks.
- EN Key Points:
- OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most advanced safety stac…
ArXiv cs.AI (B_intro+search) Link to heading
Detecting and Controlling Sycophancy with Cascading Linear Features
- Publication Time: 2026-06-26 12:00 Beijing Time
- Summary: - arXiv:2606.26155v1 Announcement Type: New.
- Abstract: Explaining and controlling model behavior via activation steering methods requires numerous pairs of contrastive samples that clearly exhibit the desired or undesired behavior.
- These data pairs determine the degree to which an interpretability framework can reliably detect the model features causing the behavior, and consequently, the ability to steer the model toward or away from such behavior.
- In this work, we present an iterative data generation pipeline that isolates the cascading linear features responsible for the behavior.
- EN Key Points:
arXiv:2606.26155v1 Announce Type: new
Abstract: Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly exhibit desir…
These data pairs determine the degree to which interpretability frameworks can reliably detect model features responsible for a behavior, and therefore the abil…
In this work, we present an iterative data generation pipeline that isolates cascading linear features responsible for a behavior
Life After Benchmark Saturation: A Case Study of CORE-Bench
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract:
- arXiv:2606.26158v1 Announce Type: new.
- When a benchmark’s accuracy saturates, it is often retired and replaced with a more challenging version.
- We show that this approach privileges accuracy and misses the opportunity to study six other key dimensions of agent performance: construct validity issues, such as shortcuts, out-of-distribution generalization, efficiency, reliability, the relative importance of the model versus scaffolding, and the boost from human-agent collaboration.
- We use CORE-Bench Hard (a benchmark for the computational reproducibility of scientific code) as a case study to demonstrate that even after accuracy saturation, measuring agents along these dimensions can yield meaningful insights into agent performance.
- EN Key Points:
- arXiv:2606.26158v1 Announce Type: new
- Abstract: When a benchmark’s accuracy saturates, it is often retired and replaced with a more challenging version
- We show that this approach privileges accuracy and misses the opportunity to study six other key dimensions of agent performance: construct validity issues such…
- We use CORE-Bench Hard, a benchmark for computational reproducibility of scientific code, as a case study to demonstrate that measuring agents along these dimen…
Refusal Lives Downstream of Persona in Chat Models
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract:
- arXiv:2606.26161v1 Announce Type: new.
- In instruction-tuned chat models, linear directions in activation space for refusal and persona traits have been identified, but the two have been studied as separate mechanisms.
- We show their interaction: a subservient persona leads to refusal.
- In Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, we extract the subservient model persona direction and the refusal direction and intervene on both.
- EN Key Points:
- arXiv:2606.26161v1 Announce Type: new
- Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been…
We show they interact: a compliant persona gates refusal
In Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, we extract a compliant model-persona direction and a refusal direction and intervene on both
AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26173v1 Announcement Type: New.
- Abstract: Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and proofs.
- Most current applications focus on static coding benchmarks.
- We extend this paradigm to algorithmic trading.
- EN Highlights:
- arXiv:2606.26173v1 Announce Type: new
- Abstract: Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and proofs
- Most current applications focus on static coding benchmarks
- We extend this paradigm to algorithmic trading
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26203v1 Announcement Type: New.
- Abstract: As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined.
- We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multilayer network analysis to study sociotechnical power structures at scale.
- We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led).
- EN Highlights:
- arXiv:2606.26203v1 Announce Type: new
- Abstract: As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined
- We introduce an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and mul…
- We validate it on two contrasting standards for agent interoperability: ERC-8004 (permissionless, on-chain) and Google A2A (corporate-led)
Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26205v1 Announcement Type: New.
- Abstract: Patients are increasingly seeking medication information online, but safety knowledge for psychiatric drugs is divided between authoritative but abstract regulatory adverse event records and experience-proximate but unverified patient narratives.
Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly contextualized information can amplify fear, nocebo responses, and non-adherence.
- Here, we develop a provenance-aware, knowledge-graph-based multi-agent framework unifying 466,525 Reddit posts, 60,782 WebMD reviews, and twenty years of U.S. history.
- EN Key Points:
- arXiv:2606.26205v1 Announce Type: new
- Abstract: Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records…
- Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly contextualised information can amplify fear, n…
- Here we develop a provenance-aware, knowledge-graph-based multi-agent framework unifying 466,525 Reddit posts, 60,782 WebMD reviews, and twenty years of U.S
Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System
- Release Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26267v1 Announcement Type: New.
- Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess.
- However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the granular quality of gameplay.
- Nevertheless, incorporating move-by-move information into rating adjustments presents a significant challenge given the substantial noise and the vastness of the game state space.
- EN Key Points:
- arXiv:2606.26267v1 Announce Type: new
- Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess
- However, they inherently suffer from response lag due to their exclusive reliance on match outcomes, neglecting the granular quality of gameplay
- Nevertheless, incorporating move-by-move information into rating adjustments presents a significant challenge given the substantial noise and the vastness of th…
- Release Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26298v1 Announcement Type: New.
- Abstract: Autonomous AI agents may begin to execute consequential, irreversible actions, such as clinical prescription and production software deployment.
- This paper observes that human institutions govern powerful autonomous actors not by monitoring their reasoning but by requiring independently-attested evidence when consequential actions are taken.
- We formalize this institutional pattern as a computational governance model for AI agent systems.
- EN Key Points:
- arXiv:2606.26298v1 Announce Type: new
Abstract: Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployment
- This paper observes that human institutions have governed powerful autonomous actors not by monitoring their reasoning but by requiring independently attested e…
- We formalise this institutional pattern as a computational governance model for AI agent systems
COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26299v1 Announcement Type: New.
- Abstract: While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict geometric constraints and subjective visual aesthetics remains a challenge.
- This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic design on the equations of flat-foldability.
- We introduce COrigami, an end-to-end AI-driven pipeline that assists the design cycle by generating crease patterns from natural language.
- EN Highlights:
- arXiv:2606.26299v1 Announce Type: new
- Abstract: While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict…
- This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic de…
- We present COrigami, an end-to-end AI-driven pipeline that assists the design cycle by generating crease patterns from natural language
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26300v1 Announcement Type: New.
- Abstract: A classical intuition holds that verifying a solution is easier than producing one.
- For today’s coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabilities and engineering harnesses grow more sophisticated, generating complex candidate solutions is no longer the hard part – reliably verifying them has become the harder problem.
- Every verifier we can build is only a proxy for human intent, not the intent itself.
- EN Highlights:
- arXiv:2606.26300v1 Announce Type: new
- Abstract: A classical intuition holds that verifying a solution is easier than producing one
- For today’s coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabilities and engineering harnesses grow more so…
Every verifier we can build is only a proxy for human intent, never the intent itself
ArXiv cs.CL (B_intro+search) Link to heading
HierBias: Context-Conditioned Hierarchical Media Bias Detection with Multi-Task Type Classification
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26100v1 Announcement Type: New.
- Abstract: Media bias detection is a critical task for ensuring fair and balanced information dissemination, yet existing sentence-level methods classify each sentence independently, ignoring the inter-sentential contextual signals that human annotators naturally utilize.
- We propose \textbf{HierBias}, a hierarchical context-conditioned media bias detector that formally models document context in its bias predictions.
- We introduce the \emph{context-conditioned bias probability} and theoretically prove that leveraging document context strictly reduces the Bayesian error of sentence-level classification when the inter-sentential mutual information is non-zero.
- EN Highlights:
- arXiv:2606.26100v1 Announce Type: new
- Abstract: Media bias detection is a critical task for ensuring fair and balanced information dissemination, yet existing sentence-level approaches classify each…
- We present \textbf{HierBias}, a hierarchical context-conditioned media bias detector that formally models document context in bias prediction
- We introduce the \emph{context-conditioned bias probability} and prove theoretically that leveraging document context strictly reduces the Bayes error of senten…
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26101v1 Announcement Type: New.
- Abstract: Reliable evaluation of large language models should separate supported answers from unsupported guesses, without conflating them with data contamination, prompt idiosyncrasies, or general refusal behaviors.
- We propose a contamination-aware, multi-zone benchmark to measure the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels.
- The benchmark includes 1,200 items across five domains, explicit abstention expectations, contamination risk metadata, and dual parsing using both an official strict parser and a standardized robustness parser.
- EN Highlights:
- arXiv:2606.26101v1 Announce Type: new
- Abstract: Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contami…
- We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build…
The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict…
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26102v1 Announcement Type: new.
- Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may unintentionally degrade values instilled during pre-training.
- We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, using SFT (helpfulness via Dolly-15k vs. helpfulness via Dolly-15k)
- and GRPO (helpfulness via RLHFlow vs. GRPO) with coding via Magicoder-110K.
- EN 要点:
- arXiv:2606.26102v1 Announce Type: new
- Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these process…
- We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on…
- coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs
Investigating LLM’s Problem Solving Capability – a Study on Statics Questions
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26103v1 Announcement Type: new.
- Abstract: Large Language Models (LLMs) have rapidly influenced many aspects of society, particularly education, due to their ability to complete assignments and exams across a wide range of subjects.
- Although previous studies have examined the educational impact of LLMs, most existing work relies on public or open question datasets and lacks topic-specific analysis.
- In engineering education, especially within mechanical engineering, systematic studies of LLM performance on specific problem types remain limited.
- EN 要点:
- arXiv:2606.26103v1 Announce Type: new
- Abstract: Large Language Models (LLMs) have rapidly influenced many aspects of society, particularly education, due to their demonstrated ability to complete as…
- Although prior studies have examined the educational impact of LLMs, much of the existing work relies on public or open problem datasets and lacks topic-specifi…
- In engineering education, especially within mechanical engineering, systematic investigations of LLM performance on specific problem types remain limited
Assert, don’t describe: Linguistic features that shift LLM reasoning about animal welfare
- Publication Time: 2026-06-26 12:00 Beijing Time
- Summary: - arXiv:2606.26104v1 Announcement Type: New.
- Abstract: Animal welfare advocates write a large volume of articles, and this material is increasingly used to train language models that millions of people then ask about animal welfare.
- Using vocabulary-matched stance-contrast probes on a held-out animal welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B’s preference for pro-animal welfare reasoning when used as fine-tuning data.
- Eight of the ten features produce statistically significant changes.
- EN Highlights:
- arXiv:2606.26104v1 Announce Type: new
- Abstract: Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about ani…
- Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B’s pr…
- Eight of the ten features produce statistically significant shifts
Context Recycling for Long-Horizon LLM Inference
- Publication Time: 2026-06-26 12:00 Beijing Time
- Summary: - arXiv:2606.26105v1 Announcement Type: New.
- Abstract: Large Language Models (LLMs) exhibit strong capabilities in short-context reasoning, but their performance degrades over long conversational horizons due to context window limitations and inefficient token usage.
- We introduce ContextForge, a context recycling system that maintains task-relevant information across turns by combining structured query generation, external memory retrieval, and controlled synthesis.
- The system enables efficient reuse of prior computations without relying on full context replay, reducing token overhead while preserving answer quality.
- EN Highlights:
- arXiv:2606.26105v1 Announce Type: new
- Abstract: Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational horizons due t…
- We introduce ContextForge, a system for context recycling that maintains task-relevant information across turns by combining structured query generation, extern…
- The system enables efficient reuse of prior computation without relying on full context replay, reducing token overhead while preserving answer quality
- Publication Time: 2026-06-26 12:00 Beijing Time
- Summary: - arXiv:2606.26106v1 Announcement Type: New.
- Abstract: Large Language Models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and distress.
While previous safety research has focused on preventing explicit harms such as toxic or policy-violating content, less attention has been paid to conversational behaviors that might unintentionally escalate conflict.
- In this paper, we investigate whether LLMs can be guided toward more de-escalating dialogue behavior through lightweight prompt-level constraints derived from Nonviolent Communication (NVC).
- EN Highlights:
- arXiv:2606.26106v1 Announce Type: new
- Abstract: Large language models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and distress
- While prior safety research has focused on preventing explicit harms such as toxic or policy-violating content, less attention has been paid to conversational b…
- In this paper, we investigate whether LLMs can be guided toward more de-escalating dialogue behavior through lightweight prompt-level constraints derived from N…
- Published: 2026-06-26 12:00 Beijing Time
- Abstract:- arXiv:2606.26107v1 Announce Type: new.
- Abstract: Sign language communication systems that integrate emotional expression remain underexplored, particularly for low-resource languages.
- This pilot study presents NEST-V1 (Nepali Emotion and Speech Transformer - Version 1), a proof-of-concept multimodal framework that demonstrates the feasibility of generating emotion-conditioned Nepali Sign Language avatars from spoken input.
- As a preliminary investigation, we focus on four common Nepali words (“thank you”, “hello”, “house”, “me”) across three emotional states (happy, neutral, sad) to validate our core technical approach.
- EN Highlights:
- arXiv:2606.26107v1 Announce Type: new
- Abstract: Sign language communication systems, that integrate emotional expression remain underexplored, particularly for low-resource languages
- This pilot study presents NEST-V1 (Nepali Emotion and Speech Transformer - Version 1), a proof-of-concept multimodal framework that demonstrates the feasibility…
- As a preliminary investigation, we focus on four common Nepali words (“thank you”, “hello”, “house”, “me”) across three emotional states (happy, neutral, sad) t…
Where Larger Models Excel: The Primacy of Constraint-Guided Reasoning
- Published: 2026-06-26 12:00 Beijing Time
- Abstract:- arXiv:2606.26108v1 Announce Type: new.
- Abstract: Larger language models consistently outperform smaller ones on reasoning benchmarks, yet the reasoning differences behind this gap remain underexplored.
- Across benchmarks in mathematics, physics, chemistry, and programming, we observe a robust performance gap: averaged over datasets, Qwen3-32B outperforms Qwen3-8B by 6.43%, and GPT-OSS-120B outperforms GPT-OSS-20B by 7.38%.
To study the reasoning differences behind these gains, we developed AdvCluster, an automated framework that identifies questions where the larger model shows a stable advantage, extracts fine-grained descriptions of advantages from paired reasoning trajectories produced by the larger and smaller models, organizes them through semantic clustering, and quantitatively evaluates and selects them under the guidance of a reviewer model.
- EN Highlights:
- arXiv:2606.26108v1 Announce Type: new
- Abstract: Larger language models consistently outperform smaller ones on reasoning benchmarks, yet the reasoning differences underlying this gap remain underexp…
- Across benchmarks in mathematics, physics, chemistry, and programming, we observe stable performance gaps: averaged over datasets, Qwen3-32B outperforms Qwen3-8…
- To study the reasoning differences behind these gains, we develop AdvCluster, an automated framework that identifies questions where the larger model shows a st…
- EN Highlights:
- Published: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26112v1 Announcement Type: New.
- Abstract: Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora.
- We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical databases can serve as an effective foundation for conversational AI development.
- Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs and fine-tunes a 12B-parameter language model using resource-efficient LoRA with 4-bit quantization.
- EN Highlights:
- arXiv:2606.26112v1 Announce Type: new
- Abstract: Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training cor…
- We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical data…
- Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs, fine-tunes a 12B-parameter language model using resource-efficient LoR…
ArXiv cs.LG (B_intro+search) Link to heading
- Published: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26128v1 Announcement Type: New.
- Abstract: The spatio-temporal evolution of many physical, chemical, and biological systems is described by non-linear partial differential equations (PDEs).
Recently, deep neural network-based surrogate models have gained increasing interest as efficient alternatives to computationally expensive traditional numerical solvers.
In this work, we propose an attention-based, physics-guided convolutional neural network as a surrogate model to learn the microstructural evolution of such systems.
- EN Highlights:
- arXiv:2606.26128v1 Announce Type: new
- Abstract: The spatiotemporal evolution of many physical, chemical, and biological systems is described by nonlinear partial differential equations (PDEs)
- Recently, deep neural network-based surrogate models have gained increasing interest as efficient alternatives to computationally expensive traditional numerica…
- In this work, we propose an attention-based, physics-guided convolutional neural network as a surrogate model to learn the microstructural evolution of such sys…
- EN Highlights:
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26164v1 Announce Type: new.
- Abstract: Finding all modes of a multimodal black-box function is a fundamental challenge in optimization, Bayesian inference, and scientific computing.
- Existing approaches – basin-hopping, CMA-ES, multistart gradient descent – operate sequentially and cannot exploit the massive parallelism of modern GPU hardware.
- We introduce \chisao{} (\textbf{C}onvergence-\textbf{H}alt-\textbf{I}nvert-\textbf{S}tick-\textbf{A}nd-\textbf{O}scillate), a GPU-native population optimizer that operates on an entire batch of samples simultaneously and leverages intentional convergence-anticonvergence oscillation cycles to escape local traps while freezing confirmed modes.
- EN Highlights:
- arXiv:2606.26164v1 Announce Type: new
- Abstract: Finding all modes of a multimodal black-box function is a fundamental challenge in optimization, Bayesian inference, and scientific computing
- Existing approaches – basin-hopping, CMA-ES, multistart gradient descent – operate sequentially and cannot exploit the massive parallelism of modern GPU hardw…
- We introduce \chisao{} (\textbf{C}onvergence-\textbf{H}alt-\textbf{I}nvert-\textbf{S}tick-\textbf{A}nd-\textbf{O}scillate), a GPU-native population optimizer th…
- Publication Time: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26168v1 Announce Type: new.
- Abstract: Living systems use noisy and incomplete sensory signals to navigate their environment.
In unicellular algae, phototaxis is often modeled as a mechanistic run-and-tumble process driven by stimulus-response rules.
- However, such descriptions overlook how organisms actively sample their environment to reduce sensory ambiguity.
- EN Highlights:
- arXiv:2606.26168v1 Announce Type: new
- Abstract: Living systems navigate environments using noisy and incomplete sensory signals
- In unicellular algae, phototaxis is often modeled as a mechanistic run–tumble process driven by stimulus–response rules
- However, such descriptions overlook how organisms actively sample their environment to reduce sensory ambiguity
- Posted: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26169v1 Announce Type: new.
- Abstract: Neural Architecture Search (NAS) has emerged as a pivotal technique in optimizing the design of Generative Adversarial Networks (GANs), automating the search for effective architectures while addressing the challenges inherent in manual design.
- This paper provides a comprehensive review of NAS methods applied to GANs, categorizing and comparing various approaches based on criteria such as search strategies, evaluation metrics, and performance outcomes.
- The review highlights the benefits of NAS in improving GAN performance, stability, and efficiency, while also identifying limitations and areas for future research.
- EN Highlights:
- arXiv:2606.26169v1 Announce Type: new
- Abstract: Neural Architecture Search (NAS) has emerged as a pivotal technique in optimizing the design of Generative Adversarial Networks (GANs), automating the…
- This paper provides a comprehensive review of NAS methods applied to GANs, categorizing and comparing various approaches based on criteria such as search strate…
- The review highlights the benefits of NAS in improving GAN performance, stability, and efficiency, while also identifying limitations and areas for future resea…
- Posted: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26179v1 Announce Type: new.
- Abstract: While WGS-based AMR prediction has achieved high accuracy, existing models lack mechanisms for grounding neural attributions in established biological pathways.
- We propose KG-TRACE, a novel neuro-symbolic framework that integrates the WHO mutation Knowledge Graph (KG) as a structured biological constraint for neural genomic models.
- Unlike existing methods that learn statistical patterns in isolation, KG-TRACE fuses genomic features and RotatE-based KG embeddings through a learned cognitive trust gate, dynamically weighting neural evidence against symbolic biological knowledge.
- EN Highlights:
- arXiv:2606.26179v1 Announce Type: new
Abstract: While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established biological pat…
We present KG-TRACE, a novel neuro-symbolic framework that integrates the WHO mutation knowledge graph (KG) as a structured biological constraint on a neural ge…
Unlike existing methods that learn statistical patterns in isolation, KG-TRACE fuses genomic features and RotatE-based KG embeddings through a learned epistemic…
- Published: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26185v1 Announce Type: new.
- Abstract: LLM-as-judge (“grader”) components have now become standard in evaluation tools, including safety evaluations where pass/fail verdicts can affect downstream deployment decisions.
- A common assumption is that setting the grader’s sampling temperature to 0 makes the grading deterministic.
- We test this assumption against a real-world safety evaluation codebase (Japan AISI’s open-source aisev) and show that it fails on two levels.
- EN Highlights:
- arXiv:2606.26185v1 Announce Type: new
- Abstract: LLM-as-judge (“grader”) components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstrea…
- A widespread assumption is that setting the grader’s sampling temperature to 0 makes grading deterministic
- We test this assumption against a real safety-evaluation codebase (Japan AISI’s open-source aisev) and show it fails on two levels
Clue-Guided Money Laundering Group Discovery
- Published: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26189v1 Announce Type: new.
- Abstract: Money Laundering Group Discovery (MLGD) aims to identify hidden criminal groups and recover their complete structures in large-scale financial networks.
- Existing graph anomaly detection methods mainly produce node-level risk alerts, while global group discovery methods passively search the entire network for suspicious groups.
- Neither approach aligns with real-world Anti-Money Laundering (AML) investigations, where analysts typically start with concrete clues and gradually expand the investigation’s scope to track down the responsible parties.
- EN Highlights:
- arXiv:2606.26189v1 Announce Type: new
- Abstract: Money Laundering Group Discovery (MLGD) aims to identify hidden criminal groups and recover their complete structures in large-scale financial network…
Existing graph anomaly detection methods mainly produce node-level risk alerts, while global group discovery methods passively search for suspicious groups over…
- Both are mismatched with real Anti-money-laundering (AML) investigations, where analysts usually start from a concrete clue and gradually expand the investigati…
Federated Hash Projected Latent Factor Learning
- Published: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26192v1 Announce Type: new.
- Abstract: Hash Learning (HL) is an effective representation learning method that maps real-valued data into compact binary representations.
- Traditional HL methods often require users to upload their personal data to a central server, which is incompatible with increasingly stringent data security regulations.
- Federated Learning (FL) offers a decentralized paradigm for learning a globally optimal model without centralizing private data.
- EN Highlights:
- arXiv:2606.26192v1 Announce Type: new
- Abstract: Hash Learning (HL) is an efficient representation learning approach that maps real-valued data into compact binary representations
- Traditional HL methods typically require users to upload personal data to a central server, which is incompatible with increasingly stringent data security regu…
- Federated Learning (FL) provides a decentralized paradigm for learning globally optimal models without centralizing private data
Statistical and Structural Approaches to Algorithmic Fairness
- Published: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26200v1 Announce Type: new.
- Abstract: Modern machine learning systems have moved beyond their origins as isolated predictive structures, evolving into complex socio-technical architectures that actively regulate human opportunities.
- As algorithms increasingly determine access to economic and social opportunities, it has become widely recognized that these systems are deeply embedded with the structural inequalities and biases of their environment.
- There is a growing recognition that models optimized for predictive accuracy can systematically disadvantage marginalized groups, and the field of algorithmic fairness has emerged in response to this.
- EN Highlights:
- arXiv:2606.26200v1 Announce Type: new
- Abstract: Modern machine learning systems have outgrown their origins as isolated predictive constructs, evolving into complex socio-technical architectures tha…
- As algorithms increasingly determine access to economic and social opportunities, it has become widely recognized that these systems are deeply embedded with th…
The field of algorithmic fairness emerged in response to the growing recognition that models optimized for predictive accuracy can systematically disadvantage m…
- Published: 2026-06-26 12:00 Beijing Time
- Abstract: - arXiv:2606.26204v1 Announcement Type: new.
- Abstract: Floods frequently impact regions around the world.
- Rapid and accurate flood detection is crucial for emergency response and timely mitigation of human and economic loss.
- The expanding availability of satellite data and advances in artificial intelligence have enhanced monitoring of environmental hazards, but many flood events remain difficult to detect due to cloud cover obscuring optical satellite imagery.
- EN Key Points:
- arXiv:2606.26204v1 Announce Type: new
- Abstract: Floods frequently impact regions around the world
- Rapid and accurate flood detection is crucial for emergency response and timely mitigation of human and economic loss
- The expanding availability of satellite data and advances in artificial intelligence have enhanced monitoring of environmental hazards, but many flood events re…