I've been saying this for years, but the market finally has proof: AI isn't a tool anymore. It's a utility. And utilities don't just change your code — they change your schedule.
A 10-person startup in China just adjusted their entire team's working hours so they could afford to keep using AI coding tools. They subscribed to four different AI services — MiniMax, GLM, DeepSeek, and Volcano Engine — and discovered the token bills were eating them alive. So they did what factories did fifty years ago with electricity prices: they moved their work to off-peak hours.
One day off during the week, one day on the weekend. Lunch pushed to 2 PM. And why? Because DeepSeek charges 2x for peak-hour API calls, and Zhipu gives 50% off after hours. The human was rescheduled around the machine.
This is the story the industry doesn't want you to read. But I've been inside these systems — I audited DeFi protocols in 2020, I've stress-tested bonding curves against flash loan attacks, and I've seen firsthand what happens when infrastructure costs start making operational decisions for a company. This is that moment.
The Peak-Pricing Mechanism Nobody Designed
Let me break down what actually happened, because the mechanics matter more than the drama.
DeepSeek's pricing is straightforward. During peak hours — weekday business hours, 9 AM to 6 PM — you pay 2x. The weekend is entirely off-peak, which is where the startup moved their work. Zhipu's GLM followed with a simpler structure: 50% off all off-peak calls. These are GPU pricing models that look exactly like electricity's peak-valley pricing. And they work because AI inference costs are time-sensitive.
Think about this in terms of the actual hardware. A single GPU cluster running an AI inference workload has a marginal cost curve that's close to zero during off-peak hours. At 2 AM, that A100 or H100 is sitting idle, burning power for nothing. But at 2 PM on a Tuesday, that same GPU is contested — every developer in China is calling the API simultaneously. So the provider charges more because the opportunity cost is higher.
This is peak-load pricing, and it's not new. Cloud providers did it. Electricity grids did it. Airlines did it. But this is the first time it's being applied to AI code generation at scale — and the first time it's directly affecting how software teams structure their workdays.
And here's the thing: the user response is rational. It's exactly what the pricing model intends. If your startup has a modest token spend per month and you're burning through 30-50% more by working during peak hours, the math says move your hours. It's not about feeling exploited. It's about survival.
What Your Developer's New Schedule Really Means
The deeper story isn't about the pricing structure. It's about the fact that a 10-person team is now running four different AI coding services simultaneously. That tells you something critical about the current state of AI infrastructure: it's no longer optional, and it's no longer cheap.
According to IDC 2024 data, over 40% of software developers in China already use AI-assisted coding tools daily. When the penetration rate hits 40%, you're no longer looking at a niche productivity tool — you're looking at an infrastructure layer. And infrastructure has a way of restructuring your organization around its costs.
I've seen this pattern before. In the early days of cloud computing, startups would run entire server fleets on AWS and then discover their bill was bigger than their rent. The ones that survived weren't necessarily the best product teams — they were the ones that understood how to optimize their cloud spend. The same thing is happening now with token spend, except the stakes are higher because the cost is more volatile.
The token bill is now the incremental cost on top of a subscription. It's not the $20/month subscription fee that's the problem. It's the per-token usage. And if you're a small team doing serious code generation — building an entire product with AI assistance — you're hitting tens of millions of tokens a month. The cost is real.
What's more telling is the fact that this team didn't cut back on AI usage. They didn't say, "Let's use fewer AI calls." They said, "Let's move our work hours so we can afford to use the same amount of AI." That's a profound signal. It means AI coding has become non-negotiable infrastructure — it's the water supply, not the coffee machine. You don't drink less water because the price goes up. You just figure out how to buy it cheaper.
The Peak-Valley Economics of Inference
This is where the crypto background kicks in. For years, we've been talking about "time value" in decentralized networks — about stake, about bonding curves, about the opportunity cost of locking capital. Now, AI infrastructure is showing the same principle, but in the physical world.
DeepSeek's 2x peak-hour premium is based on the real cost of GPU time. It's not a marketing gimmick. When your inference cluster runs at 80% capacity during the day and 15% at night, the marginal cost of serving a request during peak hours is genuinely higher. You can't scale instantaneously. You can't spin up more H100s in a second. So you ration demand through price.
And the numbers are real. According to industry estimates, the average AI inference cluster utilization is somewhere between 30% and 50% overall, but peak hours see over 80% while off-peak hours drop to single digits. If you can shift 10-20 percentage points of that utilization off-peak, you're effectively adding capacity without adding hardware. That's the economics of the peak pricing.
The question is: how far will this go? We're seeing the early days of a trend where AI providers get sophisticated about pricing. DeepSeek has already made waves by publishing its training costs — roughly $5.57 million for the V3 model — which is absurdly low compared to GPT-4's estimated $100 million. That kind of cost efficiency gives them pricing power. They can afford to be aggressive.
But here's the catch: price discrimination requires user discrimination. And users are fighting back in the only way they can — by moving their working hours.
The Hidden War: Labor vs. The Machine's Schedule
Now let's get to the part that nobody in the industry is talking about, because it's uncomfortable.
This startup didn't just "adjust" their schedule. They restructured their team's life around a pricing model. And that's a labor issue, not just a tech issue.
A Chinese startup with 10 people just decided that their employees will work five days a week but on a shifted schedule — one weekday off, one weekend day on, lunch pushed to 2 PM. That's a significant change to the human experience of work. It's not a small operational tweak. It's a fundamental shift in work-life rhythm.
And why? Because the company wants to save on token costs. They're saving maybe 30-50% of their AI spend, depending on how much they moved. But what's the human cost? What does it mean to force engineers to work at night and on weekends because the GPU cluster is idle?
This is a labor issue that touches on something I've been thinking about since the 2021 NFT bubble, when we were all excited about "decentralized identity" and "the future of work." The future of work turned out to be much more literal: adapting your work schedule to the availability of the hardware that runs the AI.
And here's the part that bothers me: the company is externalizing the cost of AI infrastructure onto its employees. They're taking the cost of the AI's peak-hour premium and shifting it onto the workers' sleep schedule. That's not a theoretical concept — that's a real shift in the power structure between capital and labor.
But I'm also a pragmatist. The counterpoint is equally real: this is what every technology does when it becomes infrastructure. Factories in the 1920s moved their production schedules to take advantage of lower electricity rates. Data centers do this today. It's called "demand response" and it's a standard practice in the energy industry.
If AI is truly becoming like electricity — and it is — then demand response is a natural extension. The question is: are we comfortable with that? We're not talking about machines anymore. We're talking about human workers.
The fact that a V2EX post about this went viral says a lot. A developer posted about this schedule shift and the comment section was split — some called it brilliant cost optimization, others were shocked that human beings would adjust their lives to accommodate a GPU's idle time. That reaction is the real story. We're all just starting to realize what "AI as infrastructure" really means.
The Competitive Landscape Shift
Let me zoom out for a second, because this isn't just about one startup.
This single phenomenon tells us the Chinese AI coding market has entered a new phase. It's no longer about who has the better model. It's about who has the better cost structure.
DeepSeek is leading the charge. They've used a combination of open-source releases and aggressive pricing to destabilize the market. By setting a 2x peak price, they're signaling that they have the compute capacity and the cost efficiency to handle demand. They're also setting a price anchor for the entire market — and by making their models open-source, they're ensuring that if the API gets too expensive, users can just self-host. That's a pretty wild play.
Zhipu's response is just as interesting. They're the "old guard" of Chinese AI, a spinoff of the Tsinghua lab. They responded to DeepSeek's peak pricing by offering a 50% off-peak discount — matching the effective off-peak price but with a different structure. That's a defensive move.
MiniMax and Volcano Engine are also in the game. But the competitive dynamics are shifting from "who has the smartest model" to "who can give you the most tokens for the least money without making you lose your mind."
Here's the thing that matters for the entire industry: this is what happens when a market matures. It's not about "AI is coming to take your jobs" — that's the old narrative. The new narrative is: "AI is here, and now we have to figure out how to pay for it."
The Untold Opportunity: The Cost Optimization Layer
So what does this mean for investors and builders?
There's a huge opportunity in this chaos. We've seen this pattern before. When cloud computing became a cost burden, a whole category of "cloud cost management" tools emerged — CloudHealth, Densify, all these companies. The same thing is going to happen with AI.
I'm predicting that by 2025-2026, we'll see a new wave of startups focused on "AI cost optimization" — tools that help you schedule your AI calls to off-peak hours, route between different models based on pricing, and monitor your token spend in real-time. The API layer is becoming commoditized, and the layer above it — the cost-optimization layer — is where the value is being created.
But there's an even bigger implication for the broader software industry. This "off-peak programming" phenomenon is going to hit the outsourcing market. If you're a software outsourcing company in China, and you're taking on a project that uses AI coding, the cost of tokens is now a line item in your project budget. You might have to charge your clients more, or you might have to work at night to keep your margins. Either way, the cost of AI coding is changing the economics of software development itself.
The "token cost" is becoming a hidden tax on all software development. And like all taxes, it affects the smallest players the most. A 10-person startup can't negotiate a special rate with DeepSeek. They can't afford a dedicated compute cluster. All they can do is adjust their work hours.
The Risk That No One Is Pricing In
The other thing I want to flag — the elephant in the room — is the risk of labor arbitrage. The startup in this story is saving 30-50% of its token spend by shifting hours. But they're also shifting the cost to their employees. And there's a real chance this becomes a trend, and it becomes a problem.
We're not going to see a labor movement for AI tokens. But we will see labor standards get defined. And we'll see some startups get a competitive advantage because they're willing to make their teams work at 2 AM while others are not.
That's not a story about AI. That's a story about how technology changes work. And it's the oldest story in the world.
The Takeaway: The True Cost of Convenience
So here's where I land on this. The phenomenon of a 10-person startup scheduling their work around GPU idle time is not a bug. It's a feature of an infrastructure shift. We're at the moment where AI is moving from being a magic tool to being a utility — with all the pricing, load-balancing, and schedule-arbitrage that comes with utilities.
If you're a developer or a founder, you should stop thinking of AI tokens as a fun budget line item. They're now a core cost component of your business. And if you're not already thinking about how to optimize that cost — whether through scheduling, routing, or architecture — you're leaving money on the table.
But more importantly: we should be thinking about what it means for people. As we optimize our businesses for the new AI cost structure, we need to make sure that we're not optimizing people out of their lives. The schedule shift is the first test. Let's watch how the industry handles it.
I've been in this space for over a decade, and I've seen the bull runs and the bear markets. This is not a market cycle. This is a fundamental shift in how software is built and paid for. The teams that figure out the cost game will win the next era. The teams that don't — well, they'll be working the night shift.
We didn't build decentralized ledgers because they were easy. We built them because they were necessary. And it looks like AI infrastructure is heading down the same path — just with better pricing.