Login
Sign Up
Woofun AI reports that Changxin Storage, labeled the "first domestic storage stock," officially landed on the ChiNext board, igniting the market with an astonishing surge of 500%. This event underscores a stark divergence in capital valuation between AI-driven storage infrastructure and decentralized storage within the Web3 domain, a phenomenon analyzed by Jacob Zhao of IOSG. While AI storage continues to be crazily revalued by capital in the current wave of technological narrative, decentralized storage has fallen into a long period of silence and loss. The fundamental answer to this disparity lies in the complete divergence of underlying value functions: one serves the immediate maximization of computing power, while the other defends the long-term integrity of data.
The core conflict defining this market split is the distinction between hot data efficiency and cold data trustworthiness. The revaluation of storage in the AI era is essentially a carnival about "hot data efficiency," serving the ultimate maximization of computing power utilization and commercial monetization. In contrast, decentralized storage adheres to the value proposition of "cold data trustworthiness," defending data fairness, anti-censorship, and the long-term memory of human civilization. The former operates as an efficiency system for hot data, whereas the latter functions as a trust system for cold data. The current capital market undoubtedly firmly stands on the side of "efficiency," but human civilization ultimately still needs an immutable memory foundation. The long-term value of trustworthy cold storage has never disappeared; it has merely been dormant in the dark side of the cycle, waiting to be repriced by the times.
In the traditional IT era, storage was viewed as a "capacity business" where CIOs focused on unit capacity costs, hard drive reliability, disaster recovery plans, and archiving strategies, with equipment update cycles lasting 3 to 5 years. Storage was seen as an accessory that followed server procurement.
However, this round of storage boom is not a traditional cycle recovery but a revaluation of data mobility capabilities driven by AI. In the era of large models, storage logic has transformed from "capacity first" to "efficiency supreme," focusing on extreme metrics such as GPU feeding rates, Checkpoint writes, and extremely low latency for RAG. This marks a leap in storage value from "the final resting place of data" to "the high-speed passage for data entering computation."
The evolution of resource bottlenecks in AI infrastructure is essentially a battle to complete the "barrel effect." The true utilization rate of computing power is not a linear addition of single assets but a harsh multiplicative effect: true utilization rate = GPU × HBM × DRAM × SSD × network × file system. Any shortcoming in one link will lead to the collapse of overall computing power utilization. In the AI era, storage has transformed for the first time from a "cost center" to an "efficiency engine." This is the fundamental logic behind the repricing of storage, shifting the focus from mere capacity retention to the active acceleration of computational throughput.
AI storage is by no means a mere accumulation of single hardware but a complex system of tightly coupled, layered scheduling. To clearly dissect its value flow, the architecture is divided into four core levels. The Compute Proximal Memory Layer, dominated by HBM and supplemented by DRAM and CXL memory pooling technology, directly interfaces with GPU/CPU packaging or buses, aiming to break the "memory wall." The High-Speed Persistent Storage Layer relies on enterprise-grade SSDs to handle high-frequency Checkpoint writes and massive training set loading.
The Low-Cost Large Capacity Storage Layer, composed of HDDs and cold storage, provides an irreplaceable TCO advantage for historical logs and compliance backups. Finally, the AI Storage Systems and Data Software layer includes high-performance parallel file systems and vector databases, ensuring that what AI truly consumes is data availability efficiently organized by the software stack. As an ecological extension, decentralized storage anchors public dataset notarization and AI training data provenance, establishing its unique niche as a "trustworthy cold layer."
High Bandwidth Memory (HBM) represents the "Bandwidth Organ" closest to computing power in the AI storage chain. Its core mission is not to store data but to continuously 'feed' data to computing power at extremely high bandwidth. The core architecture of HBM is "3D DRAM stacking + 2.5D advanced packaging," achieved through TSV vertical stacking and CoWoS heterogeneous integration, resulting in a generational leap in bandwidth. The industrial barriers involve DRAM manufacturing processes, TSV, ultra-thin stacking, packaging, heat dissipation, testing, and customer certification. Currently, only SK Hynix, Samsung, and Micron can stably mass-produce, establishing a triple moat of top-tier DRAM manufacturing processes, packaging capabilities, and NVIDIA/AMD customer certifications. Changxin Storage (CXMT) emerges as a core variable in China's DRAM domestic substitution, challenging this oligopoly.
While HBM addresses extreme bandwidth near the GPU, DRAM solidifies the server system memory foundation, and CXL attempts to break physical boundaries. DRAM primarily carries CPU-side cache, data preprocessing, and intermediate state storage. The global DRAM market is highly concentrated among SK Hynix, Samsung, and Micron. CXL (Compute Express Link) is a next-generation cache coherence interconnect protocol aimed at breaking the limitations of traditional DIMM slots and server memory resource islands, promoting the evolution of memory architecture towards expansion, pooling, and sharing. Core companies driving this transition include Astera Labs and Lanqi Technology, positioning CXL as a high mid-to-long-term architectural value asset as it moves from platform support to large-scale deployment.
Enterprise-grade SSDs serve as the data hub built from NAND chips, SSD controllers, and NVMe/PCIe pathways. These components represent independent links in the industrial chain. NAND chips, determined by Samsung, SK Hynix (Solidigm), Micron, Kioxia, Western Digital, and Yangtze Memory Technologies, define storage density and unit cost. SSD controllers, managed by Phison, Silicon Motion, Marvell, and Maxio, determine performance release, data error correction, and QoS stability. The NVMe/PCIe data pathway layer, represented by Broadcom, Marvell, and Astera Labs, determines transmission efficiency. Combined with GPUDirect Storage technology, this layer reduces CPU memory bounce buffers and CPU involvement, significantly alleviating I/O bottlenecks across the entire lifecycle of training data loading, Checkpoint writing, and RAG retrieval.
Woofun AI data shows that the low-cost foundation of AI data lakes relies on HDDs and cold storage, with representative companies including Seagate, Toshiba, and Western Digital. AI will not eliminate HDDs; rather, the demand for multimodal large models for video and image data, along with the exponential growth of enterprise compliance logs, is surging. In the AI storage architecture, SSDs manage high throughput while HDDs manage low costs and long-term preservation.
The AI Storage Software Stack acts as the scheduling hub, transforming underlying hardware into knowledge assets. High-Performance Storage Systems from VAST Data, WEKA, and Pure Storage solve the 'data hunger' of GPU clusters. Object Storage from AWS S3 builds a capacity foundation for unstructured data. Vector Databases from Pinecone and Milvus handle semantic indexing. The RAG Data Layer, exemplified by Databricks, ensures enterprise data is safely and traceably invoked by large models through data slicing, cleaning, and citation provenance.
Decentralized storage projects face a dilemma rooted in the mismatch between productization, retrieval experience, and Token incentives. Filecoin constructs a verifiable economic system through PoRep and PoSt, shifting towards AI data provenance and public dataset hosting. It must package as an S3-compatible API to upgrade from a "cheap storage market" to "verifiable computing infrastructure." Arweave, through Blockweave and SPoRA mechanisms, positions itself as a foundation for human public memory, preserving human rights records and war crime evidence. The value of Arweave lies not in speed but in its capacity to carry civilizational memory across cycles.
However, both face challenges: Supply-Demand Incentive Mismatch rewards "I can store" rather than "I need to store"; Lack of Enterprise Service Capability means enterprises purchase "peace of mind" from AWS rather than experimental infrastructure; Retrieval Experience Shortcomings make node dispersion difficult for AI hot data workflows; and Insufficient Privacy Compliance creates conflict between the right to delete and permanent immutability. Other projects like Storj and Sia have weaker Web3 narrative influence compared to Filecoin and Arweave.
BNB Greenfield and Walrus are tied to specific public chain ecosystems like BNB or SUI. Celestia and EigenDA belong to the data availability (DA) layer, serving Rollup transaction confirmations rather than long-term archiving. Projects like 0G attempt to integrate storage, data availability, computation, and AI agent settlement into AI-native modular infrastructure, but their real demand and commercialization closed loops still need verification.
The future outlook suggests a repricing of trust in the age of data monopoly. AI Data Provenance, combining cryptographic proof to construct "data lineage proof," addresses regulatory and auditing pressures. Public Datasets and Civilizational Archives anchor censored archives and cultural heritage, building irreplaceable memories. Trustworthy Archiving and Compliance Notarization achieve self-proof through Hash notarization. The integration of ZK, TEE, and DID technologies resolves privacy tensions, upgrading from a single "storage protocol" to "trustworthy data infrastructure."
Invisible Product Routes provide S3-compatible APIs and fiat billing, allowing users to purchase "trustworthy archiving" services directly. As the AI era amplifies data monopolies and copyright disputes, decentralized storage may welcome a repricing of value as a "trustworthy cold layer." Memories that cannot be easily erased by platforms or any single power may transform from romantic idealism into necessary infrastructure, marking a critical shift in how humanity preserves its civilizational archives.