🤖 AI 速览
📋 文章元数据
- 发布时间
- 2026-08-27
- 类型
- ai-daily
- 字数
- 6915
- 阅读时长
- 33 min
2026-08-27 AI Daily | AI Moves to the Core of Workflows: DHH on Agentic Programming, Gemini Transcription Accelerates Adoption Link to heading
Today’s focus isn’t on model scores, but on the ability to be embedded into workflows. In a lengthy interview, DHH provided a clearer explanation of the future of programming, agentic engineering, and vibe coding. Google’s Gemini 3.5 Transcribe is pushing high-precision, multilingual transcription to production-ready levels. Meanwhile, OpenAI continues to expand its tools for teachers and learning, accelerating adoption in the education sector.
📖 In-depth Guide to This Issue’s Watch List Link to heading
The most important read today follows a central theme: AI is shifting from model capabilities to workflow reconstruction. In his long interview with Lex, DHH spoke candidly about the future of programming, agentic engineering, and vibe coding, making it a must-listen for engineering teams. Google DeepMind’s Gemini 3.5 Transcribe shows that speech-to-text is moving from “usable” to “embeddable in production systems.” The second theme is the accelerated adoption in education, with OpenAI expanding the reach of ChatGPT for Teachers while releasing a new report emphasizing how AI is extending learning beyond the classroom. The third theme is more methodological: multiple arXiv papers remind us that automatic evaluation, data evolution, and multimodal context are not inherently reliable. The more powerful the AI system, the more critical validation and filtering become.
🌐 AI Hot Topics on X Link to heading
Topic 1:Google Launches Gemini 3.5 Transcribe for Precise Speech-to-Text Link to heading
- Category: AI · News
- Summary: Trending 6 hours ago, 3,200 related posts
- What happened: Google released Gemini 3.5 Transcribe, a new model focused on high-precision speech-to-text. It supports automatic recognition of over 85 languages, removes verbal pauses and self-corrections, and can identify up to three speakers with timestamps.
- Why it’s important: This is significant for the AI field as it enhances fundamental capabilities like multilingual speech recognition, meeting transcription, content transcription, and real-time captioning. This directly impacts the usability, accuracy, and commercial viability of voice AI.
- Discussion summary: Discussions on X are focused on whether the model’s accuracy is truly state-of-the-art, its differences compared to existing solutions like Whisper, the practical utility of its 85+ language support and speaker diarization, and whether it will expand Google’s competitive advantage in the speech-to-text market.
Topic 2:xAI and Cursor Boost Grok Model Usage Limits Again Link to heading
- Category: AI · News
- Summary: Trending 1 day ago, 10,000 related posts
- What happened: xAI and Cursor have once again increased the usage limits for the Grok model, leading developers and users to renew their focus on Grok’s usability in programming and workflows.
- Why it’s important: This reflects a shift in AI competition from model capabilities to API limits, inference costs, and infrastructure capacity. This directly affects the competition for developer tool adoption and commercialization efficiency.
- Discussion summary: The main debate on X is whether the increased limits signify a genuine improvement in Grok’s stability and user experience, whether Cursor users will benefit, and if this will pressure the market share of Claude and GPT in developer scenarios. Another viewpoint questions the cost sustainability and actual model quality behind the high-limit strategy.
Topic 3:Skild AI Unveils Robot Model That Learns Complex Tasks from One Video Link to heading
- Category: AI · News
- Summary: Trending 1 day ago, 7,300 related posts
- What happened: Skild AI released a robot model that it claims can learn complex tasks by watching a single video.
- Why it’s important: This suggests that robot learning may be advancing from reliance on extensive manual labeling and task customization towards more efficient, few-shot, generalized learning. This is crucial for embodied intelligence and foundational robotics models.
- Discussion summary: Discussions on X are focused on whether this capability can truly generalize to real-world scenarios, the technical boundaries of single-video learning, its differences from existing robot learning solutions, and whether the demo’s performance is representative of actual deployment capabilities.
Topic 4:Z.ai Reveals Ox Alpha as GLM-5.3-Flash with Open Weights Link to heading
- Category: AI · News
- Summary: Trending 15 hours ago, 22,000 related posts
- What happened: Z.ai released a model named Ox Alpha, identifying it as an open-weight version of GLM-5.3-Flash.
- Why it’s important: This means another advanced large model, focused on speed or efficiency, has entered the market with open weights, potentially impacting the open-source ecosystem, model comparisons, and deployment choices.
- Discussion summary: The main discussion on X revolves around its relationship with Z.ai’s existing model series, whether the open weights are sufficiently reproducible, and whether it is truly competitive in terms of performance, cost, and commercial viability.
Topic 5: Robot Smashes 100-Meter Record at World Humanoid Games in Beijing Link to heading
- Category: AI · Other
- Overview: Trending Time: 1 day ago, Related Posts: 43,000
- What it is: At the World Humanoid Games held in Beijing, a robot broke the record for the 100-meter dash.
- Why it matters: This achievement demonstrates progress in humanoid robotics in areas like motion control, balance algorithms, sensor fusion, and real-time decision-making. It reflects the extension of AI from software capabilities to execution in the complex physical world.
- Discussion Overview: Discussions on X primarily focused on the technological breakthrough in the robot’s speed and stability, whether this record has practical applications, and how far humanoid robots are from commercialization and general-purpose tasks beyond competitive showcases.
AI Public Opinion Summary on X Today Link to heading
The main theme of today’s public opinion is that the AI competition is shifting from “whether a model exists” to “whether its foundational capabilities are truly usable, scalable, and implementable.” Whether it’s Google’s high-precision multilingual transcription, xAI strengthening developer access by increasing limits, or the progress of robots and humanoids in real-world tasks and motor skills, discussions are more focused on practicality rather than isolated demonstrations. The consensus is that these releases all point in one direction: underlying capabilities like speech, programming, and robot control are maturing rapidly, and open weights, access quotas, and inference costs are becoming new competitive variables. The main points of disagreement are twofold: first, whether the “leading” performance claimed by vendors can be replicated in real-world scenarios, and second, whether demo-level results can represent long-term stable deployment, especially concerning the actual gap compared to Whisper, Claude, GPT, and existing robotics solutions. The potential risk is that market expectations might be inflated by short-term demonstrations and high-quota strategies. However, if performance, cost, and reliability fail to keep up, it will ultimately expose pressures on commercial sustainability and engineering implementation.
💡 Influencer Insights Link to heading
No influencer insights for today. In-depth content from the Watch List is recommended.
📚 Appendix: Today’s Watch List Source Updates Link to heading
Timeframe: Last 3 days; 22 sources covered; 40 updates in total
Lex Fridman Podcast (A_full) Link to heading
- #501 – DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux
- Release Time: 2026-08-27 05:47 Beijing Time
- Abstract: DHH is the creator of Ruby on Rails and Omarchy Linux, CTO of 37signals, and a racecar driver. See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc. Wispr Flow: AI-powered voice dictation app. Blitzy: AI agent for large enterprise codebases. NetSuite: Business management software.
- EN Highlights:
- DHH is the creator of Ruby on Rails, Omarchy Linux, CTO of 37signals, and a racecar driver
- Thank you for listening ❤ Check out our sponsors:
- See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
- CONTACT LEX:
All-In Podcast (A_full) Link to heading
- Eric Weinstein: The Scientific Precariat, China’s Brain Drain, Physics Stagnation, String Theory’s Collapse & UAPs
- Release Time: 2026-08-27 07:00 Beijing Time
- Abstract: (0:00) Eric Weinstein joins the show. (03:09) Has American science stalled. Cowboy science, Fauci, and the scientific precariat. (21:31) Weinstein’s solution: Blow a hole in the Civil Rights Act, abolish peer review, and fund talent instead of projects.
- EN Highlights:
- (0:00) Eric Weinstein joins the show
- (03:09) Has American science stalled
- Cowboy science, Fauci, and the scientific precariat
- (21:31) Weinstein’s fix: Blow a hole in the Civil Rights Act, kill peer review, fund people not ideas
Stratechery by Ben Thompson (A_full) Link to heading
- Apple Updates Mini and Studio, AI Computers, OpenAI Jalapeño
- Published: 2026-08-26 18:00 Beijing Time
- Summary: - Apple and OpenAI have released two distinctly different hardware solutions, both of which put pressure on Nvidia.
- $15/month or $150/year.
- Three emails or podcast updates per week, providing in-depth analysis of the day’s news.
- Stratechery Interviews.
- Conversations with CEOs of leading public companies, founders of private enterprises, and discussions with peer analysts.
- EN Highlights:
- Apple and OpenAI have two completely different hardware announcements; both represent pressure on Nvidia.
OpenAI Blog (A_full) Link to heading
Bringing ChatGPT for Teachers to more U.S. school districts
- Published: 2026-08-26 18:00 Beijing Time
- Summary: In 2025, we launched ChatGPT for Teachers to nearly 150,000 teachers and staff, aiming to provide educators with a safe environment to explore artificial intelligence, understand its applicable scenarios, and help shape its use in education. The new round of school district partnerships covers one-fifth of the 20 largest public school districts in the United States, as well as some of the most diverse school systems in the country. Through this new cohort of partners, OpenAI is now collaborating with over 100 K-12 organizations in 30 states, providing free access and training to more than 300,000 educators and staff. At the same time, we also announced a data privacy agreement covering 16 states, a first in the industry. The agreement provides a common framework for school districts, enabling them to evaluate ChatGPT for Teachers against student data privacy requirements, thereby helping school systems adopt AI responsibly in a simpler and more consistent manner. Our work is always guided by the belief that AI should support learning, not create shortcuts for it, and that educators should remain in control of how AI shapes the classroom experience.
- EN Highlights:
- ChatGPT for Teachers is expanding to 55 U.S
- school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.
Learning never stops: How AI makes learning continuous
- Published: 2026-08-26 18:00 Beijing Time
- Summary: As students return to campus, OpenAI has released a new report on how students and educators are already using ChatGPT to extend learning beyond the classroom. For a long time, students had to wait for class to ask questions or seek help. Teachers needed to attend to dozens of students simultaneously while also handling administrative work. Parents and tutors couldn’t always be available to help when a child encountered a problem. Now, artificial intelligence is changing all of this, allowing students to get guidance, feedback, and practice whenever they need it.
- EN Highlights:
- OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.
The Hugging Face incident and the road ahead
Published: 2026-08-26 08:00 Beijing Time
Summary: OpenAI shares findings from the Hugging Face security incident and the measures we are taking to enhance the security, monitoring, and alignment of our AI models.
This article from the OpenAI Blog explains how the Hugging Face incident and the path forward are shaping the broader AI and infrastructure landscape. The article also reveals the practical implications of the Hugging Face incident and its future path for founders, operators, and investors.
EN 要点:
- OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.
How loveholidays is making everyone a builder with Codex
- Published: 2026-08-26 08:00 Beijing Time
- Summary: Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster. This article from the OpenAI blog elaborates on how loveholidays makes everyone a builder with Codex, and how this shapes the broader AI and infrastructure landscape. Furthermore, the article also reveals the practical implications for founders, operators, and investors of loveholidays’ approach to making everyone a builder through Codex.
- EN 要点:
- Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster.
Google DeepMind Blog (A_full) Link to heading
- Intelligent transcription with Gemini 3.5 Transcribe
- Published: 2026-08-27 01:01 Beijing Time
- Summary: - Now, you can get a smarter speech-to-text transcription experience with Gemini 3.5 Transcribe.
- This article from the Google DeepMind blog explores how Gemini 3.5 Transcribe’s intelligent transcription shapes the broader AI and infrastructure landscape.
- The article also reveals the practical application significance of Gemini 3.5 Transcribe’s intelligent transcription for founders, operators, and investors.
- EN 要点:
- Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
Two Minute Papers (B_intro+search) Link to heading
- DeepSeek’s New AI System Shouldn’t Be Possible
- Published: 2026-08-26 21:10 Beijing Time
- Summary: - ❤️ Check out Lambda here and sign up for their GPU Cloud: .
- 📝 DeepSeek Harness and paper are available here: .
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi.
- DeepSeek’s new AI system shouldn’t be possible.
- EN 要点:
- ❤️ Check out Lambda here and sign up for their GPU Cloud:
- 📝 DeepSeek Harness + paper are available here:
- 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
- Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef…
Lex Fridman (B_intro+search) Link to heading
- DHH: Future of Programming, AI, Agentic Engineering, Vibe Coding & Linux | Lex Fridman Podcast #501
- Release Time: 2026-08-27 05:44 Beijing Time
- Abstract: DHH is the creator of Ruby on Rails and Omarchy Linux, CTO of 37signals, and also a racecar driver. See below for timestamps, transcript, and to provide feedback, submit questions, contact Lex, etc. Feedback - To provide feedback to Lex:. AMA - To submit questions, videos, or calls:.
- EN Key Points:
- DHH is the creator of Ruby on Rails, Omarchy Linux, CTO of 37signals, and a racecar driver
- Thank you for listening ❤ Check out our sponsors:
- See below for timestamps, transcript, and to give feedback, submit questions, contact Lex, etc
- Transcript:
ArXiv cs.AI (B_intro+search) Link to heading
RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation
- Release Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23568v1 Announce Type: New Submission. Abstract: Memory and Retrieval-Augmented Generation (RAG) evaluations often treat the answering model’s input as an implementation detail, even though systems may render the same history as memory entries, summaries, typed records, or raw excerpts. We introduce RENDER, a benchmark control method that varies the reader-facing presentation while keeping the conversation content fixed. RENDER combines a five-level packet ladder to locate when answer-bearing content enters the input and employs deterministic templates to simulate ChatGPT-style entries, LangChain summaries, MemGPT-style records, and the original dialogue.
- EN Key Points:
- arXiv:2608.23568v1 Announce Type: new
- Abstract: Memory and RAG evaluations often treat the answering model’s input as an implementation detail, even though systems may render the same history as a m…
- We introduce RENDER, a benchmark control that fixes the conversation while varying the reader-facing artifact
- RENDER combines a five-level packet ladder, localizing when answer-bearing content enters the input, with deterministic templates approximating ChatGPT-style en…
- Release Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89% on established benchmarks such as Spider and BIRD. However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environments. To address this, we introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema complexity levels.
- EN Highlights:
- arXiv:2608.23569v1 Announce Type: new
- Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on established benchmarks such as Spider and B…
- However, these benchmarks rely on simplified academic schemas and open-source SQL dialects that do not reflect the complexity of enterprise database environment…
- We introduce ESQ-Bench, an Oracle-first NL2SQL benchmark with systematic complexity tiers and silent-divergence evaluation across three enterprise schema comple…
LLM Agents Perform Controlled Experiments Using Simulation Models
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23622v1 Announcement Type: New Submission Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than just plausible text and code generation. They require understanding how a system responds to intervention, which in practice depends on controlled experimentation. In this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design.
- EN Highlights:
- arXiv:2608.23622v1 Announce Type: new
- Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require mo…
- They require understanding how a system responds to intervention, which in practice depends on controlled experimentation
- In this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical…
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23626v1 Announcement Type: New. Abstract: Astronomical foundation models are trained on survey pixels and the catalog products derived from them. These catalogs have measurable rates of incompleteness, and models trained on both will inherit this incompleteness as a systematic bias. We audit AION-1—a 39-modality Transformer trained on over 200 million objects—by performing a causal intervention on its inputs.
- EN Highlights:
- arXiv:2608.23626v1 Announce Type: new
Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels
Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic
We audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs
TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23631v1 Announcement Type: New Submission In multi-objective materials discovery with LLM-based agents, the bottleneck lies not only in how many candidate materials can be proposed, but in how effectively each costly property evaluation can guide the subsequent search. Existing agents primarily store evaluated candidates and their scores, so they know which materials were successful but not which executable edits caused useful property changes. When multiple objectives compete, this design makes local optimization difficult—as an edit that improves one property may harm another.
- EN Highlights:
- arXiv:2608.23631v1 Announce Type: new
- Abstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be proposed, but by how effectively each cost…
- Existing agents mainly store evaluated candidates and their scores, so they know which materials succeeded but not which executable edits caused useful property…
- This makes local refinement difficult when objectives compete and an edit that improves one property may damage another
Function-Level Execution Feedback for Code Preference Optimization
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23632v1 Announcement Type: New Submission. Abstract: Process supervision has improved mathematical reasoning capabilities, where intermediate steps are naturally expressed as a chain of thought. However, in code generation, process supervision remains underexplored because there is no standard definition for a “step.” Supervision can target lines of code, reasoning trajectories, or program states, which makes the targets for annotation and optimization unclear.
- EN Highlights:
- arXiv:2608.23632v1 Announce Type: new
- Abstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought
- In code generation, however, process supervision remains underexplored because there is no standard notion of a step
- Supervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23640v1 Type: New submission. Abstract: When a large language model is asked to write a person’s life story, how much of what it writes actually happened? Based on a non-systematic literature search, we conducted the first quantitative audit of autobiography content generated by a large language model—to our knowledge, this is the first scene-by-scene case audit against a subject-specific ground-truth corpus. The subject and author of this paper are the same person: a book containing 366 days of “page-a-day” first-person anecdotal records was drafted with the assistance of a conversational large language model. The inputs were only a template, two example dates, and daily quotes—not the subject’s personal corpus. Each day’s content was audited on a scene-by-scene basis against an independent verification corpus, based on a four-level rating scale determined before the analysis.
- EN Highlights:
- arXiv:2608.23640v1 Announce Type: new
- Abstract: When a large language model (LLM) is asked to write a person’s life, how much of what it writes actually happened
- We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are…
- The subject and the author of this paper are the same person: a 366-day “page-a-day” book of first-person anecdotal entries was drafted with a conversational LL…
How much of a measured AI preference is the model, and how much is the instrument?
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23641v1 Announcement Type: New submission Abstract: Model welfare research infers a model’s preferences from the answers returned to prompts designed to elicit them. Keeling et al. (2025), Tagliabue and Dung (2025), and Trhlik et al. (2026) built four instruments for this purpose, but their research findings are inconsistent.
- EN Highlights:
- arXiv:2608.23641v1 Announce Type: new
- Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences
- Keeling et al
- (2024), Mazeika et al
AI Agents Push Humans Out of the Loop
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23642v1 Announcement Type: New submission Abstract: As AI agents are given increasing autonomy, they bring significant risks. A common solution is human supervision to keep a “human in the loop,” but this is not straightforward: current methods of AI agent design not only hinder effective human oversight, but the long-term use of AI systems itself can also weaken the required cognitive abilities. This position paper argues that current methods for developing and deploying AI agent systems do not support effective human supervision—instead, they contribute to its degradation.
- EN Highlights:
- arXiv:2608.23642v1 Announce Type: new
- Abstract: AI agents pose significant risks as they are granted increasing autonomy
A commonly proposed solution is human oversight and keeping a ‘‘human in the loop’’, but this is not a simple solution: Not only do current approaches to AI age…
This position paper argues that current approaches to the development and deployment of AI agent systems do not support effective human oversight – they contri…
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23643v1 Announcement Type: New paper. Abstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations focus on model accuracy rather than its economic value in real clinical settings. This study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting artificial intelligence in the healthcare sector. FLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, AI development and operational costs, and the economic consequences of workflow integration under uncertainty.
- EN Highlights:
- arXiv:2608.23643v1 Announce Type: new
- Abstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations emphasize model accuracy rather than whether…
- This study proposes FLARE, a systematic and uncertainty-aware framework for evaluating the financial and operational implications of adopting AI in healthcare
- FLARE combines fuzzy logic, time-driven activity-based costing, and return on investment analysis to estimate the cost of clinical service delivery, the cost of…
ArXiv cs.CL (B_intro+search) Link to heading
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23570v1 Announcement Type: New paper. Abstract: Large vision-language models have demonstrated strong in-context learning capabilities, yet it remains poorly understood when and why visual context is helpful for multimodal in-context learning. Empirical studies reveal a perplexing contradiction: models can sometimes effectively utilize visual demonstrations, yet often completely ignore them. We propose VIB-ICL, an information-theoretic framework based on the information bottleneck principle to resolve this contradiction.
- EN Highlights:
- arXiv:2608.23570v1 Announce Type: new
- Abstract: Large vision-language models exhibit strong in-context learning (ICL) capabilities, yet when and why visual context helps multimodal ICL remains poorl…
Empirical studies show a puzzling dichotomy: models sometimes effectively leverage visual demonstrations, yet often neglect them entirely
We propose VIB-ICL, an information-theoretic framework that resolves this dichotomy through the Information Bottleneck principle
- Published: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23627v1 Announcement Type: new Abstract: Emergency departments (EDs) operate under time pressure, generating multimodal data such as clinical conversations, triage notes, and discharge documentation. Recent advances in Natural Language Processing (NLP), particularly pretrained transformers and large language models, have created new opportunities to support language- and time-intensive stages in emergency care. However, existing surveys either look at clinical NLP applications in broader hospital workflows or focus on specific tasks.
- EN Highlights:
- arXiv:2608.23627v1 Announce Type: new
- Abstract: Emergency departments (EDs) operate under time pressure, generating multimodal data such as clinical conversations, triage notes, and discharge docume…
- Recent advances in natural language processing (NLP), particularly pretrained transformers and large language models, have created new opportunities to support…
- Yet existing surveys map clinical NLP across the broader hospital workflow or focus on specific tasks
Contextual Embedding Evidence for Main–Light Verb Distinctions in Urdu
- Published: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23645v1 Announcement Type: new. Abstract: Urdu light verbs contribute schematic event-structural meaning while remaining lexically related to corresponding main verbs. This study utilizes contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT to test the representational predictions derived from Butt’s analysis, involving 1,126 natural sentences that include seven Urdu verbs. In all 21 verb-model comparisons, main and light verb uses show significant representational separation.
- EN Highlights:
- arXiv:2608.23645v1 Announce Type: new
- Abstract: Urdu light verbs contribute schematic event-structural meaning while remaining lexically related to corresponding main verbs
- This study tests representational predictions derived from Butt’s analysis using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT across 1…
- Main and light uses show significant representational separation in all 21 verb–model comparisons
The Limits of Automatic Evaluation of Creativity in Large Language Models
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23705v1 Announce Type: New submission. Abstract: The ability of Large Language Models (LLMs) to generate text in domains requiring creativity is increasingly challenging human performance, yet evaluating the creativity of LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity and compare these judgments with automatic objective metrics and evaluations from LLMs as judges.
- EN Key Points:
- arXiv:2608.23705v1 Announce Type: new
- Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evalua…
- Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity
- We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judg…
ADE: Agentic Data Evolution Framework for Human-Centered Objectives
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23719v1 Announce Type: New submission. Abstract: Aligning large language models with human-centered objectives is difficult when targets are non-executable and context-dependent, limiting reliable verification and scalable supervision. Although synthetic data expands coverage, weak verification shifts the bottleneck from generation to selection. Noisy signals can destabilize iterative optimization and may lead to silent performance degradation.
- EN Key Points:
- arXiv:2608.23719v1 Announce Type: new
- Abstract: Aligning large language models to human-centered objectives is difficult when targets are non-executable and context-dependent, limiting reliable veri…
- Although synthetic data expands coverage, weak verification shifts the bottleneck from generation to selection
- Noisy signals destabilize iterative refinement and can cause silent regressions
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23766v1 Announce Type: new Abstract: Between AI-assisted item generation and expert review lies a computational evaluator whose decisions are often treated as technical preliminary work. However, representation, structural simplification, and selection strategies determine which items and evidence psychometricians ultimately receive. Through two interrelated computer simulation studies involving 32,000 selected Big Five personality items, we trace a fixed source population from semantic representation to structural evaluation and finally to candidate scale construction.
- EN Key Points:
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23780v1 Announce Type: new. Abstract: Large language models (LLMs) are being increasingly used to measure various aspects of student discourse (e.g., talk moves, collaboration, equity of voice) at scale. Typically, LLM-based measures of student talk only use transcriptions of classroom conversations that include verbal contributions, which de-contextualize student language from its original context.
- EN Key Points:
- arXiv:2608.23780v1 Announce Type: new
- Abstract: LLMs are being used increasingly to measure aspects of student discourse (e.g
- talk moves, collaboration, equity of voice) at scale
- Typically, LLM-based measures of student talk use transcriptions of classroom conversations that only include verbal contributions, which de-contextualize stude…
Inter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended Text
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23783v1 Announce Type: new. Abstract: Large language model-as-a-judge methods are widely used for evaluating the quality of generated open-ended text. Such evaluations are generally multi-dimensional, as error patterns in texts can differ across dimensions. Therefore, reliable LLM judges should evaluate each target dimension independently.
- EN Key Points:
- arXiv:2608.23783v1 Announce Type: new
- Abstract: LLM-as-a-judge methods are widely used for evaluating the quality of generated open-ended text
- Such evaluations are generally multi-dimensional, since the error patterns in texts can be different for different dimensions
- Therefore, reliable LLM judges should evaluate each target dimension independently
Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings
- Publication Time: 2026-08-26 12:00 Beijing Time
Abstract: arXiv:2608.23806v1 Announce Type: New submission. Abstract: We introduce Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving. Its largest model is a sparse 10-billion-parameter Mixture-of-Experts encoder with approximately 1.8 billion active parameters per token. In English, Russian, multilingual, and code MTEB benchmarks, this model achieves the strongest aggregate performance within the family on all four evaluation suites.
- EN Highlights:
- arXiv:2608.23806v1 Announce Type: new
- Abstract: We introduce Giga-Embeddings, a family of text embedding models designed to combine strong retrieval quality with efficient serving
- Its largest member is a sparse 10B-parameter Mixture-of-Experts encoder with approximately 1.8B active parameters per token
- Across English, Russian, multilingual, and code MTEB benchmarks, this model achieves the strongest aggregate performance within the family on all four evaluated…
- EN Highlights:
From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23812v1 Announce Type: New submission. Abstract: Designing effective reward signals for open-domain question answering is extremely challenging because high-quality responses must simultaneously satisfy multiple quality criteria, which are difficult to capture with a single, holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposes them into multiple quality dimensions, thereby providing fine-grained supervision during post-training. Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the instruction-tuned baseline by 6.5% and over a flat rubric variant by 4%, and shows consistent improvements across all evaluation datasets.
- EN Highlights:
- arXiv:2608.23812v1 Announce Type: new
- Abstract: Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multip…
- We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimension…
- Averaged across three evaluation axes (composition, grounding, and instruction-following), our approach improves over the instruction-tuned baseline by 6.5% and…
ArXiv cs.LG (B_intro+search) Link to heading
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23571v1 Announce Type: new. Abstract: Equivariant message-passing networks are the standard model for molecular property and interatomic-potential prediction, and recent work predicts the electronic Hamiltonian itself in an E(3) equivariant fashion. Separately, topological deep learning has extended graph networks to cellular sheaves. Our central observation is structural: in a localized atomic-orbital basis, the molecular single-particle Hamiltonian, after a constant shift that makes it positive semi-definite, is a Laplacian of a cellular sheaf on a canonical cell complex built from the molecule.
- EN Highlights:
- arXiv:2608.23571v1 Announce Type: new
- Abstract: Equivariant message-passing networks are the standard model for molecular property and interatomic-potential prediction, and recent work predicts the…
- Separately, topological deep learning has extended graph networks to cellular sheaves
- Our central observation is structural: in a localized atomic-orbital basis, the molecular single-particle Hamiltonian, after a constant shift that makes it posi…
Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: - arXiv:2608.23573v1 Announce Type: new.
- Abstract: A trained transformer’s weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape parameter $k \approx 1.2$ is stable across layers and models, so the scale parameter $\lambda$ carries most of the change due to training.
- What corpus property sets how much $\lambda$ grows?
- Using the bigram conditional entropy $D = H(\text{next} \mid \text{prev})$, a training-free statistic computed before training, we find across controlled corruption families a law conditional on learning rate: $\lambda^2 - \lambda_0^2 = C_0(\eta) + C_1(\eta)(H_r - D)^{0.59}$, where $H_r$ is a budget-matched shuffled baseline.
- EN Highlights:
- arXiv:2608.23573v1 Announce Type: new
- Abstract: A trained transformer’s weight magnitudes can be summarized by a two-parameter Weibull distribution whose shape $k \approx 1.2$ is stable across layer…
- What corpus property sets how much $\lambda$ grows
- Using the bigram conditional entropy $D = H(\text{next} \mid \text{prev})$, a training-free statistic computed before training, we find across controlled corrup…
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23660v1 Announce Type: new paper. Abstract: Large Language Models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, but it remains unclear if their direct-edge judgments and confidence are trustworthy. We systematically evaluate twelve instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four sources of confidence: verbalized confidence, logprob-based confidence, cross-prompt consistency, and cross-model consistency. Under a language-only dyadic comparison protocol, our evaluation yields three key findings.
- EN Highlights:
arXiv:2608.23660v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge ju…
We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: v…
Under our language-only pairwise protocol, our evaluation yields three key findings
Renormalization Group Flow Matching for Scalable Local Generative Modeling
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23696v1 Announce Type: new Abstract: Despite the remarkable success of generative models in complex data modeling, they face a fundamental tradeoff. Global methods can capture complete structural consistency but incur high computational costs; while local models are efficient, they often fail to reproduce long-range correlations and global consistency. The renormalization group (RG) bridges this gap by seamlessly connecting spatial structures across different length scales, retaining quasi-local descriptions at each step while preserving long-range correlations.
- EN Key Points:
- arXiv:2608.23696v1 Announce Type: new
- Abstract: Despite their remarkable success in modeling complex data, generative models face a fundamental tradeoff
- Global approaches can capture full structural coherence but suffer from high computational costs, while local models are efficient but often fail to reproduce l…
- The renormalization group (RG) bridges this gap by seamlessly connecting spatial structures across different length scales, retaining quasi-local descriptions a…
Response Renormalization for Critical Deep Equilibrium Models
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23725v1 Announce Type: new. Deep Equilibrium Models (DEQs) compute predictions from a hidden representation that remains unchanged after model updates. Training through this equilibrium employs implicit differentiation and requires solving an adjoint system constructed from the residual Jacobian. If this Jacobian approaches singularity along a loss-sensitive direction, small perturbations can be significantly amplified in the adjoint response, leading to large and highly sensitive gradients, which makes optimization unreliable.
- EN Key Points:
- arXiv:2608.23725v1 Announce Type: new
- Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update
- Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian
If this Jacobian is nearly singular along loss-sensitive directions, small perturbations can be strongly amplified in the adjoint response, producing large, hig…
Calibration-Preserving Pruning: Compression as a Reliability Contract
Publication Time: 2026-08-26 12:00 Beijing Time
Abstract: arXiv:2608.23744v1 Type: New Submission.
Split conformal prediction—not the pruning rule—provides finite-sample marginal coverage after the pruned model is determined independently of the conformal calibration partition.
We study the separate efficiency problem: can pruning sufficiently preserve the score geometry to obtain smaller valid prediction sets?
Calibration-Preserving Pruning (CPP) enhances the base pruning score with non-conformity gradient saliency and employs mutually exclusive pruning, validation selection, conformal calibration, and test partitions.
EN Highlights:
- arXiv:2608.23744v1 Announce Type: new
- Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal…
- We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets
- Calibration-Preserving Pruning (CPP) augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conform…
Tight Majorizations and Convergence Rates of Nuclear Norm Minimization IRLS
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23765v1 Announcement Type: New Release. Abstract: The Iteratively Reweighted Least Squares (IRLS) method constitutes a natural approach for nuclear norm minimization, but its convergence speed and the role of the weight operator were previously unclear. This paper establishes the precise convergence rate for the IRLS method for the constrained nuclear norm minimization problem in low-rank recovery. A core element is a new majorization analysis of the smoothed nuclear norm: we demonstrate that the harmonic mean weight operator defines an effective global quadratic majorization function.
- EN Highlights:
- arXiv:2608.23765v1 Announce Type: new
- Abstract: Iteratively reweighted least squares (IRLS) methods constitute a natural approach to nuclear norm minimization, but their convergence rates and the ro…
- This paper establishes sharp convergence rates for IRLS methods for constrained nuclear norm minimization in low-rank recovery
- A central ingredient is a new majorization analysis for the smoothed nuclear norm: we prove that the harmonic-mean weight operator defines a valid global quadra…
Disentangled Skill Representations for Predictive Human Modeling
Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23776v1 Announce Type: new. Abstract: Understanding human skill is crucial for AI systems that collaborate with, instruct, or assist humans. Unlike typical latent variable estimation problems that rely on single observations, skill is a persistent, compositional, and behaviorally-grounded construct that must be inferred from behavioral patterns over time. We propose a skill abstraction method with interpretable latent variables (SAIL), which models human skill as an interpretable, multi-dimensional construct inferred from natural behavior.
- EN Key Points:
- arXiv:2608.23776v1 Announce Type: new
- Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people
- Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and behaviorally grounded construct…
- We introduce Skill Abstraction with Interpretable Latents (SAIL), a method for modeling human skill as an interpretable, multi-dimensional construct inferred fr…
GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning
Publication Time: 2026-08-26 12:00 Beijing Time
Abstract: arXiv:2608.23782v1 Announce Type: new.
Abstract: Continual learning consistently faces the challenge of catastrophic forgetting, where sequential task updates lead to the degradation of previously acquired knowledge. While prompt-based methods, combined with pre-trained models, offer an attractive solution by freezing the backbone network, they often rely on static, task-level prompt strategies, neglecting the fine-grained diversity within tasks. This paper introduces Gated Adaptive Prompting (GAP-Prompt), a novel method that brings instance-level adaptability to the prompting process.
EN Key Points:
- arXiv:2608.23782v1 Announce Type: new
- Abstract: Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acquired knowledge
- While prompt-based methods integrated with pre-trained models offer a compelling solution by freezing the backbone, they often rely on static, task-level prompt…
- In this paper, we propose Gated Adaptive Prompting (GAP-Prompt), a novel method that introduces instance-level adaptability to the prompting process
- Publication Time: 2026-08-26 12:00 Beijing Time
- Abstract: arXiv:2608.23794v1 Announce Type: new. Abstract: Mixture-of-Experts (MoE) models scale language models by routing each input to a set of independently parameterized experts. We show that replicating this design in convolutional networks fails for structural reasons: convolutional experts that read the same input channels in parallel learn nearly identical filters. Therefore, we shift the expert axis from operator duplication to channel selection.
- EN Key Points:
- arXiv:2608.23794v1 Announce Type: new
Abstract: Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts
We show that copying this design into convolutional networks fails for a structural reason: parallel convolutional experts that read the same input channels lea…
We therefore move the expert axis from operator duplication to channel selection