When Wang Jian, the founder of AliCloud, took the stage at the 2026 World AI Conference and declared that the next paradigm of artificial intelligence would shift from text-and-code models to multi-modal scientific data, the room of venture capitalists and engineers blinked. They expected another hot take on scaling laws or agent frameworks. Instead, they got a foundational pivot: AI must become infrastructure, like mathematics, and the raw material is not just tokens of language but the complex, heterogeneous data of physics, biology, and climate. The auditor blinked; the market didn’t. But for those of us watching on-chain liquidity flows and regulatory corridors, the signal was unmistakable: the decoupling of value from compute is about to begin, and it will route through tokenized data markets.
The context here matters more than the soundbite. Wang Jian’s speech was not a random visionary rant. It came from a man who built one of the world’s largest cloud platforms and has spent years watching the AI industry burn capital on bigger models while ignoring the data supply chain. His core thesis—that the next leap forward requires a universal technical architecture capable of ingesting scientific data as naturally as it handles text—is a direct challenge to the current GPU-and-parameter arms race. For crypto, this has profound implications. If AI is to become an infrastructure layer akin to mathematics, then the data that feeds it must become a programmable asset. And programmable assets are what crypto does best.
Enter the technical frontier: tokenizing scientific data. The challenge is not trivial. Current tokenization methods—Byte Pair Encoding (BPE), WordPiece—are designed for discrete text sequences. Scientific data is inherently non-discrete: protein folding structures, radar images, gravitational wave outputs. These data types require lossless representation and provenance tracking. My audit experience with ERC-20 smart contracts in 2017 taught me that security and liquidity are intertwined. Here, the security of data integrity and the liquidity of data access are the same. Blockchain’s immutable hash chains can anchor scientific datasets, while zero-knowledge proofs and decentralized storage (IPFS, Arweave) enable verifiable computation without exposing proprietary information. This is not a pipe dream. Over the past six months, I have tracked three protocols attempting to tokenize climate model outputs and genomic sequences. Their TVL is negligible—under $50 million—but their architecture is sound.
The core insight is that the liquidity premium will shift from compute tokens (like GPU-backed coins) to data tokens (datasets as NFTs or fungible pools). Consider the macro picture: global R&D spending exceeds $2 trillion annually. A fraction of that data, once tokenized, could be repurposed for AI training, creating a secondary market for “data as a service.” DeFi’s yield farming was a tax on ignorance; the tokenization of scientific data is a tax on institutional inefficiency. The academic and corporate silos that guard their data will be forced to unbundle when they see the arbitrage. And the infrastructure to do it? Layer-2 sequencers, which are already centralized in practice, can be repurposed as data ordering layers for scientific feeds. The irony is delicious: the same centralized sequencers I have criticized for years might become the backbone of AI’s data pipeline before they ever become “decentralized.”
Now the contrarian angle—because liquidity doesn’t care about architecture; it flows to the most efficient node. The consensus narrative today is that AI model size is the ultimate moat. Bigger models, better performance, higher market cap. Wang Jian’s speech directly contradicts that. He argues that the bottleneck is not compute but high-quality, structured scientific data. If he is right, then the true moat is data ownership and tokenization, not model parameters. This flips the investment thesis for crypto. Instead of betting on GPU networks (Render, Akash) or model marketplaces (Bittensor subnetworks), the smart capital should be looking at data provenance protocols and decentralized science (DeSci) platforms. Furthermore, the idea of a “universal architecture” for all scientific modalities is ambitious but fragile. In crypto, we have learned that monolithic chains fail; modularity wins. The same will apply to scientific data. We will see specialized application chains (app-chains) for genomics, another for climate, another for particle physics. The universal layer will be the settlement and staking layer, not the data processing layer.
Let’s ground this in practical signals. Over the short term (0–6 months), watch for top-tier journals like Nature or Science to publish papers on scientific data tokenization methods. If we see a breakthrough in lossless serialization of protein or radar data for transformer models, the market will react. I also track the funding rounds of startups claiming to bridge blockchain and scientific data—currently dormant, but likely to heat up. In the mid-term (6–18 months), compare the performance of GPT-5 or Llama 4 on scientific benchmarks versus specialized models. If the generalists close the gap, the “universal architecture” thesis gains credibility. If they don’t, the app-chain thesis wins. Either way, the tokenized data infrastructure will be needed.
The market is currently sideways, and chop is for positioning. The sideways price action in Bitcoin and Ethereum has lulled most traders into waiting for a macro catalyst. But the real catalyst is structural: AI’s pivot to scientific data will create new asset classes before the next halving cycle. My experience in the 2022 Terra collapse taught me that algorithmic dependencies are fragile; data dependencies are far stickier. The next bear market will not be about stablecoin depegs but about data oracle latency. Chainlink’s current solution—centralized nodes feeding decentralized contracts—is itself a joke when applied to multi-modal scientific streams. Expect new oracle designs to emerge, specifically focused on high-frequency, high-integrity scientific data.
Takeaway: The next crypto cycle will be defined not by DeFi leverage or NFT profiling, but by the tokenization of scientific data. Position your portfolio not around GPU scarcity but around data access rights. The infrastructure is being built now, away from the headlines. If the auditor blinks again, the market will have already moved. The question is: where is your data staked?


