System translated (Gemini)

🤖 AI 速览

Today’s focus is on model internal mechanisms and controllability: research is beginning to interpret MoE routing as frequency patterns similar to Huffman coding, transforming black-box routing into an analyzable subject; meanwhile, expert-aware contrastive decoding is being used to mitigate …
📋 文章元数据
发布时间
2026-07-26
类型
ai-daily
字数
3958
阅读时长
19 min

2026-07-26 AI Daily | MoE Routing Becomes Interpretable, Hallucination Governance Shifts to Expert-Level Decoding Link to heading

Today’s focus is on the internal mechanisms and controllability of models: research is beginning to interpret MoE routing as a frequency pattern similar to Huffman coding, turning black-box routing into an analyzable subject. Meanwhile, expert-aware contrastive decoding is being used to mitigate hallucinations, indicating that governance methods are moving from prompt engineering to structural-level optimization.

📖 In-depth Guide to This Issue’s Watch List Link to heading

There are three main themes worth exploring in detail today: First, MoE and the internal mechanisms of models, with papers explaining routing as Huffman coding and the use of expert-aware contrastive decoding to reduce hallucinations. This shows that “black-box routing” is becoming an analyzable and intervenable subject. Second, controllable workflows in industry scenarios, where areas like adverse medical event detection, material science literature analysis, and paragraph-level argument mining are emphasizing retrieval, human-computer collaboration, and multi-agent task decomposition—highly beneficial for engineering teams. Third, output quality and diversity, with discussions on literary evaluation, viewpoint de-homogenization, and the boundaries of natural language, reminding us to look beyond generation capabilities to norms, disagreements, and verifiability.

🌐 AI Hot Topics on X Link to heading

Topic 1: Veteran Engineers Skip Reviews on Reliable AI-Generated Code Link to heading

  • Category: AI · News
  • Overview: Trending time:, Related posts: 264
  • What happened: Reports indicate that some veteran engineers are skipping code review processes when they deem AI-generated code to be sufficiently reliable.
  • Why it matters: This reflects AI’s deep integration into the software development lifecycle, directly impacting code quality, engineering efficiency, and accountability boundaries. It also sparks debate on whether auto-generated code can replace manual oversight.
  • Discussion summary: Discussions on X are mainly focused on two points: one side argues it significantly boosts development efficiency, while the other worries that over-reliance on AI could introduce security and maintenance risks, which are amplified without manual review.

Topic 2: DeepSeek Pauses Fundraising After CEO Comments Leak Link to heading

  • Category: AI · News
  • Overview: Trending time: 8 hours ago, Related posts: 1100
  • What happened: DeepSeek has reportedly paused its planned fundraising process after comments from its CEO were leaked.
  • Why it matters: This highlights the sensitivity of AI companies regarding rapid expansion, capital operations, and internal communication. It could also affect external perception of their governance, strategy, and fundraising capabilities.
  • Discussion summary: Discussions on X are centered on the specific content of the leaked comments, whether it will impact DeepSeek’s valuation and fundraising prospects, and if it could disrupt its momentum and market confidence in the competitive AI landscape.

Topic 3: Anthropic Launches Claude Opus 5 at Half the Cost of Fable 5 Link to heading

  • Category: AI · News
  • Overview: Trending time: 2 days ago, Related posts: 18000
  • What happened: Anthropic has released Claude Opus 5, offering near state-of-the-art performance at a price roughly half that of competing products.
  • Why it matters: This indicates that the large model competition is shifting from purely chasing leaderboard rankings to focusing on cost-effectiveness and inference tunability. This could further drive down the cost of using AI and change how products are designed at the agent and application layers.
  • Discussion summary: The main discussion on X revolves around whether this launch is a “leaderboard takeover” or an “escalation of the price war,” and whether the model’s built-in “effort” adjustment means the future is not about choosing models, but about allocating inference intensity based on the task.

Topic 4: Tech Leaders Warn Against AI Open-Weight Restrictions Link to heading

  • Category: AI · News
  • Overview: Trending time: 1 day ago, Related posts: 197000
  • What happened: Several tech leaders have issued warnings against imposing excessive restrictions on open-weight AI models, arguing such policies could impact model releases and usage.
  • Why it matters: This issue concerns the balance between open innovation, reproducibility, competitive landscape, and security governance within the AI ecosystem. It will also affect model distribution methods and industry standards.
  • Discussion summary: Discussions on X are centered on whether open-weight models increase the risk of misuse versus whether restrictions would stifle innovation. The debate focuses on the trade-offs between security regulation, the open-source spirit, industrial competition, and national AI leadership.

Topic 5: Anthropic Engineers Shift AI from Prompts to Graphs and Lean Context Link to heading

  • Category: AI · News
  • Overview: Trending time: 15 hours ago, Related posts: 1100
  • What happened: Anthropic engineers proposed and are discussing a shift in AI workflows from being “long-prompt driven” to using “graph-structured orchestration + lean context.”
  • Why it matters: This reflects a shift in AI system design from single-prompt optimization to more stable, controllable, and scalable architectures. This could impact the efficiency and cost of agents, tool use, and enterprise-level applications.
  • Discussion summary: Discussions on X are focused on whether this approach can significantly improve inference stability and reduce context redundancy and hallucinations. At the same time, some question whether graph structures might increase engineering complexity and reduce the flexibility of prompts.

Topic 6: Developer Builds Rocket League Clone in One Claude Opus 5 Prompt Link to heading

  • Category: AI · Entertainment
  • Overview: Trending for: 19 hours ago, Related posts: 265
  • What it is: A developer claims to have generated a Rocket League-style clone game with just a single prompt to Claude Opus 5.
  • Why it matters: This demonstrates the capability of large models in complex code generation, rapid prototyping, and automated creation, reflecting how AI is further lowering the barrier to entry for game development.
  • Discussion summary: Discussions on X are mainly focused on whether this was truly achieved with a “single prompt,” the quality and playability of the generated code, whether significant manual fixes were applied, and if such demos represent actual capabilities or are just marketing hype.

Topic 7: Terminator 2 Meme Revives AI Data Center Fears Link to heading

  • Category: AI · News
  • Overview: Trending for: , Related posts: 528
  • What it is: A meme related to Terminator 2 has resurfaced on X, sparking renewed discussion about the expansion of AI data centers and their potential risks.
  • Why it matters: This reflects public concern over the rapid expansion of AI infrastructure, the concentration of computing power, and energy consumption. It also influences external perceptions of the sustainability and security of the AI industry.
  • Discussion summary: The discussion is mainly split into two camps: one side believes that large-scale data centers are a necessary foundation for AI development, while the other is concerned about high power consumption, environmental pressure, centralization risks, and science-fiction-like fears of “runaway AI.”

Summary of AI Public Opinion on X Today Link to heading

The main narrative on X today can be summarized as a comprehensive shift in focus for AI: from “model capability showcases” to parallel discussions on “engineering implementation, cost competition, and governance risks.” There is a general consensus that AI can now significantly enhance development efficiency, prototype generation, and workflow automation. The consensus is that the emergence of low-cost, high-performance models like Claude Opus 5, along with new paradigms such as graph-structured orchestration, indicates that the industry’s competitive focus is shifting from simply topping leaderboards to competing on cost-effectiveness, stability, and controllability. The point of contention centers on “whether AI can be trusted.” Some argue for skipping partial manual reviews and promoting open-sourcing of weights to foster innovation, while others worry this will amplify security vulnerabilities, maintenance costs, and misuse risks. The potential risks fall into three main categories: first, security vulnerabilities from the over-automation of code and agent systems; second, the impact on market confidence from leaked corporate governance and funding news; and third, the pressures of energy consumption, centralization, and long-term sustainability from data center expansion.

💡 Influencer Insights Link to heading

Alright, as a senior AI industry analyst, I have conducted a deep dive into the thought fragments of AI influencers over the past 24 hours. Here are the core insights distilled from a summary of their posts.


Today’s discussions were highly focused on the democratization of model capabilities, the evolution of AI Agent interaction paradigms, and the practical competition of on-device models.

🔥 Price Wars and Expanded Access for Model Capabilities Link to heading

  • Claude Opus 5 Launch: Unquestionably the top headline today. Anthropic released the precisely positioned Opus 5. @dotey provided a detailed breakdown of its pricing strategy: offering near-frontier intelligence at half the price of $Fable 5$ ($5/$25 per 1M tokens for input/output). Bloggers generally see this as a precise “surgical strike” by Anthropic on the high-end model market, targeting complex task processing with high cost-effectiveness. Meanwhile, both @zhixianio and @AI_Jasonyu confirmed that Fable 5 access has been extended to Max and Team plans.
  • AI Coding Tools Get Another Boost: In response to the Claude Opus 5 launch, both @dotey and @Pluvio9yte predict that Codex and Claude Code will likely reset their usage quotas to compete for developer users. This reflects the intense competition among leading AI coding tools, which frequently offer “perks” to maintain user loyalty.
  • Real-time Voice and Multimodal Interaction Become Standard for Agents: @dotey reported that the ChatGPT desktop client (formerly the Codex App) has fully launched voice control. Based on the GPT-Live full-duplex architecture, it allows users to give spoken commands while managing multiple background Agents simultaneously. Combined with screen perception (Appshots), this marks a new era for AI assistants, evolving from “conversational” to “context-aware” collaboration. Meanwhile, @Pluvio9yte announced that Codex will also introduce a real-time voice mode, suggesting that voice interaction is set to become the primary interface for Agents.

🧠 “Pragmatic” Validation of On-Device Models Link to heading

  • Hands-on Test of Small-Parameter Model Capability Boundaries: @zhixianio conducted a “brutal” comparison between Gemma 4 12B Coder and their daily-use Qwen 3.6 35B MoE using real-world programming tasks. The conclusion is that 12B models have a physical ceiling when it comes to generating complex, long-form, stateful programs in a single pass, and this is unrelated to fine-tuning. This provides extremely valuable real-world data for the current debate on the “sweet spot” of popular small models.
  • Quantization Technology Becomes Key for On-Device Implementation: Google released the Gemma 4 QAT (Quantization-Aware Training) model. @zhixianio pointed out that this method, which specializes in adapting to quantization capabilities during the training phase, is a crucial step towards unlocking high-performance local inference for edge devices and consumer GPUs, heralding that system-level AI will soon be available on Android devices.

🎨 The “Motion Graphics Arms Race” in AI Video Generation Tools Link to heading

  • Upgrades in Productized Video Tools: @Pluvio9yte and @AI_Jasonyu both promoted the update to HeyGen, which now integrates the HTML motion graphics framework HyperFrames. This represents a shift in AI video generation from simple digital human narration to the automatic generation of dynamic graphics with a more cinematic and commercial aesthetic.
  • Open-Sourcing of a Community Motion Graphics Library: The MotionForge motion graphics Skill, open-sourced by @QingQ77 and discovered by @Pluvio9yte, provides 106 shot cards and 161 dynamic samples. It enables Claude Code / Codex to directly replicate motion effects on par with promotional videos from major companies, marking a significant attempt at automated video creation by Agents.

2. Noteworthy Unique Perspectives or Industry Foresight Link to heading

  • The Debate on “Should AI Make Us Stop Reading Code?”: This was the most in-depth debate of the day. @dotey shared the extreme viewpoint of software engineering guru Bob Martin: “I don’t read code written by AI.” The core idea is to control quality through extreme constraints like unit tests and Gherkin tests, because “human reading speed is far slower than AI generation speed.” @dotey added a perfect supplement: while the cost of AI refactoring is low, the software engineering discipline of “writing tests before refactoring” has become even more critical in the AI era. Good test coverage is a prerequisite for efficient AI self-correction.
  • The “Role Reversal” Strategy in AI Collaboration: @dotey raised profound questions about a popular Agent collaboration model. The current popular /advisor model of “small model designs, large model acts as a consultant” has a paradox: if the small model’s design is poor or overconfident, even a powerful consultant is useless. He proposed a more performance-optimal combination: “large model for design + small model for execution + large model for acceptance testing.” This optimizes the cost-effective allocation of AI resources.
  • A New Version of Geographic Arbitrage: “Currency Arbitrage”: @Pluvio9yte keenly spotted a case where someone took advantage of the Bolivian currency collapse to purchase a Codex 20x subscription at an extremely low price (approximately 845 RMB). This reflects the significant arbitrage opportunities that exist between global AI service pricing and regional economic fluctuations.
  • The Potential of Lightweight HTML Presentations: The open-source HTML PPT project Bento, showcased by @vista8, has all its documents in plain-text JSON, inherently designed to be modified and iterated upon by AI. This suggests that the paradigm for future content presentations is shifting towards “AI-native formats,” where static files will gradually be replaced by interactive web pages that can be programmed in real-time.
  • OpenCodex: Recommended by @Pluvio9yte. This is an open-source project that allows Codex to connect with other large models like Kimi, Grok, and GLM, breaking the single-provider lock-in of GPT and offering a multi-model switching workflow.
  • Topview MCP: Strongly recommended by @AI_Jasonyu. A “full-stack marketing MCP” designed for cross-border e-commerce, it has built-in data signals from platforms like Amazon, YouTube, and TikTok. It provides an end-to-end solution from product selection data analysis to generating images and videos, representing a significant efficiency boost for e-commerce professionals.
  • Bento HTML PPT: Discovered by @vista8. This is a plain-text JSON-driven HTML presentation framework with cool animations, perfectly suited for programmatic generation and modification by AI Agents.
  • HeyGen (Limited-Time Offer): The company is officially giving away a one-month Creator membership with the discount code 5GESQHRN. Those in need of digital human video production should get on board promptly.
  • Overseas Infrastructure: The Saily US eSIM tutorial series by @AI_Jasonyu (starting as low as $0.98) has become today’s featured low-cost solution for obtaining an overseas phone number. Both he and @gefei55 have repeatedly emphasized that the barrier to obtaining various overseas resources (phone numbers, payment methods, accounts) is rising. The core consensus is that “the earlier you get them, the more stable and cheaper they are.”

📚 Appendix: Today’s Watch List Source Update Link to heading

Time window: Last 3 days; Covering 22 sources; 10 updates total

ArXiv cs.CL (B_intro+search) Link to heading

  • What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20425v1 Announcement Type: new.
      • What makes writing “good” remains a persistent question in literary studies and computational linguistics.
      • We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality.
      • In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model’s implicit theories of quality from its reasoning traces.
    • EN Key Points:
      • arXiv:2607.20425v1 Announce Type: new
      • Abstract: What makes writing “good” remains a persistent question in literary studies and computational linguistics
      • We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality
      • In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model’s im…
  • Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs’Hallucinations

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20426v1 Announcement Type: new.
      • Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either struggle to alter models’ internal knowledge or have poor cross-domain generalization.
      • Contrastive decoding mitigates hallucinations by using layer-wise differences in LLMs.
      • However, prior studies have only explored transformer-based models (e.g., GPT), ignoring other effective frameworks like mixture-of-experts (MoE) models.
    • EN Key Points:
      • arXiv:2607.20426v1 Announce Type: new
      • Abstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models’internal knowledge or h…
      • Contrastive decoding mitigates hallucinations by using layer-wise differences in LLMs
      • However, prior studies only explore transformer-based models (e.g., GPT), ignoring other effective frameworks like mixture-of-experts (MoE) models
  • Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20427v1 Announcement Type: new.
      • The Mixture-of-Experts architecture has revolutionized scaling, but the underlying logic of its routing remains a black box.
      • In this paper, we uncover a fundamental governing principle: MoE routing is not merely selection, but an embodiment of Huffman coding.
  • We introduce the Frequency-Diversity Law, revealing that state-of-the-art models, such as Phi-3.5-MoE and Gemma-4-27B-A4B, spontaneously act as information-theoretic engines.

    • EN Highlights:
      • arXiv:2607.20427v1 Announce Type: new
      • Abstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box
      • In this paper, we uncover a fundamental governing principle: MoE routing is not merely selection, but a manifestation of Huffman Coding
      • We introduce the Frequency-Diversity Law, revealing that state-of-the-art models, such as Phi-3.5-MoE and Gemma-4-27B-A4B, spontaneously act as information-theo…
  • Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20428v1 Announcement Type: new.
      • Abstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-related adverse events (cirAE) from clinical records.
      • Compared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement measured by Cohen’s kappa (kappa = 0.82 vs 0.50), and reduced the average review time by approximately half.
      • This framework pilots how LLMs can be applied to identify immune-related toxicities across organ systems and, more broadly, enable accurate, scalable, and transparent adverse event data extraction.
    • EN Highlights:
      • arXiv:2607.20428v1 Announce Type: new
      • Abstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-…
      • Compared with unassisted manual review, the LLM-assisted workflow improved accuracy (F1 = 0.88 vs 0.77), inter-rater agreement measured by Cohen’s kappa (kappa…
      • This framework pilots how LLMs can be applied to identify immune-related toxicities across organ systems and, more broadly, enable accurate, scalable, and trans…
  • More Is Not More: What Matters for Diversity in LLM Opinions?

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20429v1 Announcement Type: new.
      • Abstract: Large language models are increasingly used to simulate diverse human perspectives in open-ended tasks such as synthesizing surveys, modeling focus groups, and forecasting public opinion.
      • However, the outputs of LLMs exhibit systematic opinion homogenization.
      • Practitioners have explored various interventions to increase diversity, but the picture remains fragmented: different methods are evaluated individually with incomparable metrics, and in practice they are often deployed and scaled concurrently, making it difficult to attribute gains to specific components.
    • EN Highlights:
      • arXiv:2607.20429v1 Announce Type: new
  • Abstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, an…

  • However, LLM outputs exhibit systematic opinion homogenization

  • Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with…

  • LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20430v1 Announcement Type: New.
      • Abstract: We introduce LLM-INSTRUCT, the winning system for the UZH shared task at ArgMining 2026, which involves paragraph-level argument mining in UN and UNESCO resolutions.
      • The task requires paragraph-type classification, subset prediction of 141 official tags, and directed relation prediction under a strict JSON schema setting using only open-weight models of up to 8B parameters.
      • We frame the task as constrained structured prediction.
    • EN Key Points:
      • arXiv:2607.20430v1 Announce Type: new
      • Abstract: We present LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining in UN and UNESCO resolutions
      • The task requires paragraph-type classification, prediction of a subset of 141 official tags, and directed relation prediction under a strict JSON schema settin…
      • We frame the task as constrained structured prediction
  • Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20431v1 Announcement Type: New.
      • Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, but traditional retrieval-augmented generation pipelines struggle to coordinate heterogeneous tasks in a single retrieve-then-generate architecture.
      • Here, we introduce AlphaAgent, a skill-driven agent framework that separates retrieval-based question-answering from paper report generation through explicit skill contracts.
      • A dedicated retrieval skill rewrites user requests into material-specific search intents, queries a curated index of over 300,000 papers in the Metallurgy and Metallurgical Engineering categories from journal citation reports, and reformulates queries when initial evidence is insufficient.
    • EN Key Points:
      • arXiv:2607.20431v1 Announce Type: new
      • Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet co…
  • Here we present AlphaAgent, a skill-driven agent framework that decouples retrieval-based question answering from paper-level report generation through explicit…

  • A dedicated retrieval skill rewrites user requests into material-specific search intents, queries a curated index of more than 300,000 papers from the Journal C…

  • Position: Natural Language Should Not Fully Replace Formal Languages

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20432v1 Announce Type: new.
      • Abstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal languages, such as programming languages for software design.
      • In this position paper, we argue that this perspective overlooks fundamental linguistic properties of natural language, specifically that it is optimized for informality in open-ended contexts.
      • We introduce a formal framework centered on “task specificity,” defining it as the information-theoretic reduction of uncertainty in an output space (e.g., all possible images) according to the user’s specific requirements.
    • EN Key Points:
      • arXiv:2607.20432v1 Announce Type: new
      • Abstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal langua…
      • In this position paper, we argue that this perspective overlooks fundamental linguistic properties of natural language, specifically that it is optimized for un…
      • We introduce a formal framework centered on task specificity, defining it as the information-theoretic reduction of uncertainty in an output space – such as…
  • Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

    • Publication Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20433v1 Announce Type: new.
      • Abstract: While language models remain frozen at their training state, the world evolves continuously.
      • Knowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and programmatic reasoning collapse, while encyclopedic recall remains intact.
      • We trace this asymmetric degradation to a distributional mismatch.
    • EN Key Points:
      • arXiv:2607.20433v1 Announce Type: new
      • Abstract: While language models remain frozen at their training state, the world evolves continuously
      • Knowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and…
      • We trace this asymmetric degradation to a distributional mismatch
  • Break Through the Compression Bottleneck: From Theory to Practice

    • Release Time: 2026-07-25 12:00 Beijing Time
    • Abstract: - arXiv:2607.20434v1 Announce Type: New.
      • Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.
      • Existing compression methods suffer from bottleneck issues: when the compression ratio is increased, performance degrades significantly.
      • Low-rank decomposition and quantization are two prominent compression methods that have been proven to significantly reduce the computational and memory requirements of large language models (LLMs) while maintaining model accuracy.
    • EN Key Points:
      • arXiv:2607.20434v1 Announce Type: new
      • Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead
      • Existing compression methods suffer from bottleneck issues: when the compression ratio is increased, performance degrades significantly
      • Low-rank decomposition and quantization are two prominent compression methods that have been proven to significantly reduce the computational and memory require…