Last Tuesday I opened a nine-dimension protocol analysis that contained no protocol.
Every section was present. Technical. Tokenomics. Market structure. Ecosystem position. Regulatory exposure. Team and governance. Risk. Narrative. Supply-chain transmission. Every section was populated. Every section read N/A.
The framework was intact. The subject was absent. The document had already been syndicated to a distribution list I would estimate at eleven thousand inboxes.

I know the shape of that artifact because I have produced thousands of them. Not the N/A version — the other one. The one where a model is handed a partial dataset, a rigid template, and a deadline, and it concludes that a refusal formatted as a report is still a report. That is the failure mode the crypto research stack is now optimising for. Not hallucination. Something harder to police: the production of structurally valid documents that carry zero positional risk.
Nothing printed on the tape. No liquidation cascade. No funding dislocation. And yet the piece did its job. It filled a slot. It retained a subscriber. It taught eleven thousand people that a nine-dimension framework can be completed without a subject.
Alpha detected. Position established — on the absence.
The supply chain nobody audits
The crypto research stack has three layers. I have worked all three, and the incentive gradient runs downhill through every one of them.
The bottom layer is data. CoinGecko and CoinMarketCap for listings and spot. Kaiko and Amberdata for order-book and derivatives microstructure. Glassnode and Nansen for on-chain cohorts. Dune and The Graph for custom queries. DefiLlama for TVL aggregation. Call it a dozen credible primitives, plus a long tail of resellers who re-export the same numbers under a different logo and a different font.
The middle layer is synthesis. Newsletter desks, analyst teams, aggregator feeds, and since 2023, language models doing the first pass. This is where the headcount collapsed. A desk that ran six analysts in 2021 runs two analysts and a retrieval pipeline in 2026. The saved cost is not visible in any output. The degradation is.
The top layer is consumption. Retail. Fund analysts who need one defensible paragraph to attach to a memo. Exchange listing committees. Compliance teams building an audit trail they will never be asked to produce.

The economics explain everything downstream. CoinGecko tracks north of thirty thousand assets. The number of people globally who can produce a defensible, technically literate, fourteen-hundred-word analysis of any single one of them is — generously — a few thousand. That ratio is not a gap. It is a structural void. And a structural void in a market that pays per unit of coverage will always be filled by whatever is cheapest to produce.
Nine-dimension templates became that cheapest unit. Not because anyone designed them badly — the one I opened was competently built, with sensible sub-prompts and an explicit instruction not to fabricate. But a template with fixed slots creates a completeness constraint, and a completeness constraint creates pressure to fill every slot whether or not the data exists.
That pressure predates LLMs by a decade. The models only removed the friction.
Where the nulls come from
Strip the narrative away and this is a data engineering problem. I have traced seven production sources of confident garbage. None of them involve anyone deciding to lie.
1. Indexer lag and reorg handling. Most protocol-level consumers read from subgraphs or hosted indexers, not from a node. When a protocol migrates — a proxy redeploy, a new vault, a factory clone — the indexer returns stale state until it catches up. The query does not error. It returns HTTP 200 with yesterday's numbers. An analyst pulls TVL at block N, publishes two hours later, and the figure is off by an entire migration. Nothing in the output flags it.

2. null coerced to zero. This is the single largest generator of confident garbage in the sector. Aggregator endpoints return null for assets they do not index. The downstream pipeline — the CSV exporter, the dashboard, the retrieval layer feeding the model — coerces null to 0. A zero supply, zero TVL, zero volume reading is mathematically indistinguishable from a dead protocol. Now the model has a datum. It writes a sentence about it. Nobody in the chain intended to mislead, and the chain still produced a false statement.
3. TVL double counting. Restaking wrappers, liquid staking tokens, recursive lending loops. One dollar of ETH can register as TVL on a liquid staking protocol, again on a restaking layer, again as collateral on a money market, and again as the notional of a Pendle position. DefiLlama applies an adjustment for the best-known cases. Most secondary sources inherit the raw number and drop the adjustment, because the adjustment is harder to fit into a headline. Four times the liquidity, one time the capital.
4. Volume is not liquidity. I broke a PFP floor in 2021 with this. Floor price is a derived statistic — the ask on the cheapest listed token, which says nothing about what a market maker would pay for size. I mapped self-trades: two wallets, two-sided fills, gas funded from a shared source, positions held for minutes. That pattern is still running on illiquid alts today, and it inflates "volume" on every aggregator that reports it. Volume is throughput. Liquidity is depth. They are not the same number and they are not even the same unit.
5. Circulating supply versus mint. A mint event is not a float event. Treasury-held, pre-minted, and unvested tokens occupy the same address space as circulating supply unless someone manually separates them. Market-cap rankings are built on the unsorted figure, which means the ranking itself is a derived claim about a segmentation nobody performed.
6. Derivatives asymmetry. Spot price is nearly free to acquire. Funding rates, open interest by venue, and options skew are expensive, jurisdictionally fragmented, and settled on a dozen different clocks. So roughly four out of five articles that claim to describe "market structure" are describing spot price and calling it structure.
7. Timestamp drift. Research is stamped at publication, not at extraction. In a market that moves on thirty-minute candles, a two-hour gap between query and byline is not a rounding error. It is a different regime with a different set of active liquidations.
8. Corpus feedback. The models now writing the first pass were trained on a decade of content that was itself generated from these same broken inputs. Garbage in, confident out, then re-ingested next quarter as ground truth.
Stack all eight and you have the input. Now feed it to a model with a nine-slot template and an instruction not to hallucinate.
Here is what the model learns. The safest output that satisfies the completeness constraint is to restate the framework and leave the substance as a placeholder. Not because the model is honest. Because nine slots minus four data points leaves five slots that must be filled with something, and the cheapest compliant something is the name of the slot itself.
The N/A report is not a refusal. It is a compliance artifact.
I have been on the other side of this. In 2020 I ran a Python monitor against MakerDAO stability fees and liquidation thresholds, reading the chain directly instead of trusting a dashboard. The entire reason that script existed was that the dashboards were stale. The same discipline applies to writing: if a number does not have a block height attached, it is a rumour with a font.
And in 2024, coordinating the ETF series for European readers, the most valuable thing we published was not the analysis. It was a table. Extraction timestamp, source, block height, and an explicit list of what we could not resolve. That table outperformed every narrative piece in the series. Each of the three articles cleared one hundred thousand unique readers, and the document people forwarded internally — to risk committees, to allocators — was the uncertainty table. That is the demand signal almost nobody in this industry reads correctly.
What a defensible pipeline actually requires is unglamorous: direct node reads for anything that matters, null preserved as null end to end, an extraction timestamp and block height on every figure, at least two independent sources reconciled against each other, and a written unknowns section that survives editing.
The diagnosis everyone prefers
The standard diagnosis is that AI hallucinates. That diagnosis is comfortable, it is wrong, and it lets every participant keep their current incentives.
The real problem is that this market pays for coverage and audits for accuracy. Coverage is produced per unit. Accuracy is verified once, expensively, and late. When you pay per unit and audit rarely, you get unit production. The N/A document I opened is, technically, the correct behaviour — it refused to fabricate — and it earned its author nothing. A fabricated piece with a confident thesis would have earned more, faster, with a lower probability of ever being caught. One of those two outcomes is rewarded at scale. Guess which.
There is a second layer nobody wants to name. N/A is not a neutral output. It is a transfer of responsibility. The byline keeps the authority. The reader absorbs the risk. The framework's completeness is preserved, and the analytical burden is handed downstream to someone who must now perform the work the document claimed to have done. In a market where the median reader is hunting a directional signal, a nine-section N/A report is functionally identical to a broken feed — except it arrives with a masthead attached.
And a third layer, specific to where we are right now. In a sideways tape, price stops doing the storytelling. When there is no trend to explain, coverage volume substitutes for direction. Published piece count rises, information content per piece falls, and the reader's ability to separate signal from template degrades precisely when they need it most.
Arbitrage window closing in 10 minutes. Not on a trade — on credibility. On-chain data is now cheap enough that provenance is a solvable problem. The window in which an unattributed number passed for analysis is closing, and the desks still publishing them have not noticed the tape.
What to watch
Watch for lineage, not frameworks. The next meaningful shift in crypto research is not a better nine-dimension template — it is signed datasets: extraction timestamps, block heights, source hashes, and an explicit, published list of what the pipeline could not resolve. That standard already exists in traditional market data and in regulated audit trails. It has simply never been forced onto this asset class, because nobody here was ever required to reconcile.
Watch which desks start shipping unknowns sections. Watch which aggregators begin propagating nulls instead of zeroes. Watch whether listing one-pagers start carrying block heights. Those three signals will separate the pipelines that produce data from the pipelines that produce documents.
If a nine-section report can be generated from an empty input and syndicated to eleven thousand people without one correction, the model was never the problem. The problem is that nobody in the chain was paid to notice.
Pick your inputs the way you pick a counterparty. Liquidation pending. Don't be the collateral.