Oh look, we need to be more efficient. Welcome to the new reality.
Caricature Nano Banana 2, Prompt Der Promptologe
Update from August 3, 2026: Figures on model prices are particularly fleeting in this story. While this article was being written, OpenAI cut its prices for GPT-5.6 Luna by 80 percent — a great example of the pace described below. All price information has been updated accordingly and dated.
Just a few months ago, burning as many tokens as possible was a status symbol in American companies. There were internal leaderboards where teams were celebrated if they had spent a particularly large amount on AI queries. This was called “tokenmaxxing,” a play on Silicon Valley slang, where you “maxx” everything that can be “maxxed.” Those who spent a lot were supposedly proving: We’re at the forefront.
The german version:
This logic has flipped within a matter of weeks. As the Wall Street Journal described it in late July, “Tokenmaxxing” has turned into “Thrift-Maxxing” — companies are now competing to see who can save the most skillfully. The shift has come so quickly that it’s surprising even people whose job is to advise companies on precisely this issue.
The Lamborghini for Getting Milk
Mike Saeks is one such person. The former investment banker was hired two months ago by Cursor, the coding startup that’s currently being sold to Elon Musk’s SpaceX for $60 billion. His role as “Field Chief Technology Officer”: helping companies measure and optimize their AI spending. Cursor refers to this internally as “tokenomics” — a play on words that’s now so commonplace that you almost forget just how new the mindset behind it actually is.
Saeks sums it up in a way worth keeping in mind: “It’s like driving a Lamborghini to the grocery store to pick up milk when it was designed to be raced around a track.” Using the most expensive, most powerful model for every query makes about as much sense as using a race car for your weekly grocery shopping.
Cursor itself has calculated just how expensive the difference can actually be. Building a web browser from scratch — using only OpenAI’s GPT-5.5 — cost a good $10,000. The same task using a combination of Cursor’s own Composer model and Anthropic’s Opus 4.8: $1,339. A savings of 87 percent, with comparable results. Saeks’ conclusion is almost tossed off in passing, but it describes the new reality quite accurately: “The best model for a task used to change every few months. Now it feels like it’s happening multiple times a week.”
From a race to squander to a race to save
Marty Kausas, CEO of the AI customer support platform Pylon, puts it more bluntly: “There’s zero loyalty that I’m seeing. It really feels like a bloodbath right now.” His company alone received about $1.6 million in free tokens from one provider in 2026, plus $65,000 from a second and $10,000 from a third. AI labs are fighting “like crazy” with discounts to retain their existing customers — a behavior that seems remarkable in an industry that was just recently talking about scarcity and waiting lists.
What happened? In short: Companies have stopped treating AI as a strategic matter of faith and are starting to treat it like any other commodity — they compare, they negotiate, they mix and match. And suddenly, names are appearing on the shopping lists of major American companies that were still considered exotic niche players just two years ago: DeepSeek, Moonshot, Zhipu, Alibaba’s Qwen.
What the Numbers Show
The best way to gauge just how significant this shift really is is through OpenRouter, the industry’s largest model-agnostic brokerage platform. At the end of 2024, Chinese open-weight models accounted for less than 1.2 percent of the token volume there. By the spring of 2026, their share exceeded 50 percent on some days; by June, it had settled at around 44 to 46 percent among the ten largest providers — DeepSeek alone now accounts for about 16 percent, ahead of Google, Anthropic, and OpenAI taken individually. The exact figure varies depending on the reporting date and measurement method; however, the order of magnitude and trend are consistent across multiple sources.
Analysts at Citi arrive at a similar picture using a different metric: The share of open-source models among the tokens processed on OpenRouter jumped from 34 percent in January 2026 to 65 percent in June.
And the cost difference driving this shift is real — and not insignificant. In May 2026, the benchmarking firm Artificial Analysis ran each major model through the same ten test tasks and tallied the costs. The result at the time: Anthropic’s top model came in at $4,811, OpenAI’s at $3,357—DeepSeek, by contrast, at $1,071, Moonshot’s Kimi at $948, and Zhipu’s GLM at just $544. Claude was thus nearly nine times as expensive as the cheapest Chinese alternative, for practically the same task.
But here’s the thing: This snapshot is already outdated—and by the U.S. labs themselves. On July 30, 2026, in the middle of researching this article, OpenAI lowered the API prices for its fastest model, GPT-5.6 Luna, by 80 percent (from $1.00/$6.00 to $0.20/$1.20 per million input/output tokens), and for the mid-range model, Terra, by 20 percent (new: $2.00/$12.00). The most expensive model, Sol, remained at its original price of $5.00/$30.00. Reason: Efficiency gains in the model, server software, and agent infrastructure — according to OpenAI, the Sol model itself achieved part of this by autonomously rewriting production code, thereby reducing server costs by 20 percent. This places Luna in a price range previously reserved for Chinese open-weight models: On the so-called “Agents’ Last Exam” benchmark, Luna is expected to outperform Anthropic’s flagship model, Fable 5, at approximately 99 percent lower cost per task, according to OpenAI. Replit’s AI chief, Michele Catasta, puts it this way: “GPT-5.6 Luna is the closest we’ve come to intelligence too cheap to meter. I’ve never seen a model this affordable be this powerful.”
The pricing landscape as of early August 2026
Official API prices, dollars per million tokens, input/output:
Anthropic’s Fable 5 is at 10/50, while Claude Opus 5 — available since July 24 and, according to Anthropic, priced at half that of Fable 5 — is at 5/25.
OpenAI’s Sol is at 5/30, Terra at 2/12, and Luna at 0.20/1.20.
Google’s Gemini 3.1 Pro is at 2/12 — coincidentally almost identical to OpenAI’s Terra.
Among the Chinese open-weight models, the picture is more mixed than the “China is cheap” narrative suggests:
DeepSeek V4 Flash, at 0.14/0.28, remains by far the cheapest model in the comparison
Zhipu’s GLM-5.2 comes in at 1.40/4.40
but Moonshots’ Kimi K3, at 3.00/15.00 per million tokens, costs more than Google’s Gemini or OpenAI’s Terra.
So the “China discount” isn’t a law of nature; rather, it’s negotiable on a model-by-model basis.
Try it out: If you want to experiment with volume and model mix yourself to see how much costs shift depending on the proportion and model selection.
The calculator uses the current prices from early August.
It gets interesting when you look at how companies actually build this mix—because “we’re going with the cheap model now” is rarely the whole story.
At Telnyx, an infrastructure provider for real-time AI agents, this happened via a detour the company hadn’t chosen. There, 1,000 agents were running on a top Anthropic model plus an open-source operating system, at a subscription price of $200 per employee per month. Then Anthropic prohibited third-party systems from using its services on a subscription basis—a violation of its terms of service, as the company explained. Switching to a pay-per-use model would have cost around $100,000 a day, according to CEO David Casem: “We did the math and it was going to be like 100 grand per day.” So Telnyx switched providers. Today, a family of models from the Chinese provider Z.AI powers the company’s 1,400 agents, at a cost of about $100 per agent per day. Anthropic’s most powerful model, internally known as “Fable,” continues to orchestrate the work like a conductor, while OpenAI’s “Sol” reviews the results produced by the open models. Casem’s conclusion sounds almost conciliatory: “It’s not like we don’t use OpenAI or Anthropic models—we still do. They just don’t do everything anymore.”
At Harvey, an AI startup for law firms, the same principle works in reverse: The company has trained its own GLM-5.2—the model from the Chinese provider Zhipu—and equipped it with a tool that automatically calls on Anthropic’s Fable 5 for particularly difficult cases. “We work with all of them,” says President Gabe Pereyra, “we’re figuring out a bunch of solutions like this to maintain or improve performance and achieve much better cost efficiency.”
This exact pattern—a low-cost model as the default, an expensive one only on demand—has been given a name. Databricks CEO Ali Ghodsi calls it the “Advisor model”: A low-cost open-source model handles the bulk of the work; when it encounters a task it can’t solve, it’s given a tool that allows it to ask a Frontier model from OpenAI or Anthropic for help. “You can curb costs really well this way,” says Ghodsi. His company operates an AI gateway through which thousands of enterprise customers route their model requests every day—so Ghodsi has a pretty direct view of the trend before it hits the press.
At Hex, a provider of AI data analytics, the shift happened remarkably quickly: CEO Barry McCardel reports that within two weeks, about half of the customers had incorporated Kimi, Moonshot’s model, into their workflow. His reasoning is indicative of the uncertainty currently prevailing in the industry: “We are continually assessing how much we want to commit to any one lab given how dynamic a moment this is. Any day, the labs can drop a model that hits the frontier.”
And even where the transition has been underway for some time, it’s now being held up as a model: Zoom has been using a finely tuned, open Llama model from Meta for three years, combined with services from Anthropic and OpenAI. CTO Xuedong Huang describes the principle using a Chinese proverb in which three ordinary people together achieve the wisdom of a genius—his “secret sauce” for companies.
The Mixers: Eleven companies in focus—including Coinbase, Microsoft, and Cursor—can be explored in detail here.
Why China, of All Places
There is a structural reason why Chinese models, of all things, are taking on this role as the affordable workhorse. The best-known U.S. models are closed-source: They run exclusively on their manufacturers’ infrastructure, trained with the most expensive chips available, in a power grid that can barely keep up—costs that are inevitably passed on to customers. Chinese labs, on the other hand, operate under U.S. export restrictions on high-performance chips and have had to aggressively optimize their systems to train and efficiently run competitive models with less computing power. Necessity became a strategy—and the models are also mostly “open-weight”: their training parameters can be downloaded and run on users’ own servers. No one can cut off access by decree.
This is no longer just a theoretical risk. The Trump administration temporarily prohibited Anthropic from selling its most powerful model—known internally as “Mythos” and publicly available as the scaled-down “Fable 5”—to foreign entities. The ban was not lifted until early July 2026, subject to new security requirements. For many in the industry, this was a wake-up call: A U.S. Frontier model is politically revocable. A model hosted on one’s own server is not.
Just how quickly the gap between the systems is narrowing is, naturally, a matter of debate—the U.S. standardization institute CAISI estimated the lag of Chinese models at around eight months in May 2026, an estimate that should be interpreted with due caution. But even Anthropic admits in its own policy paper that its models are now only “several months ahead” of the Chinese ones—and warns that Beijing is gaining ground “in global adoption on cost.”
Whether this is a security problem or an opportunity is already the subject of heated debate in Europe. A pros-and-cons analysis in the Handelsblatt sums up both sides: One position warns that Chinese open-source models could contain backdoors that activate automatically under certain circumstances—after all, every company in China is legally required to cooperate with the Communist Party. The opposing view counters this: Unlike with Anthropic or OpenAI, the most powerful Chinese models can be downloaded for free and run locally, under the user’s full control. Even if Beijing were to impose an export ban one day, the models would simply continue to run on European customers’ servers. An “Anthropic moment,” such as the one the U.S. has just experienced, is structurally impossible with open models.
Meanwhile, an unusual coalition is forming on the U.S. side: At the end of July, Nvidia, Microsoft, and Palantir jointly signed an open letter in support of open models and urged their own government to exercise restraint regarding potential restrictions—a sign that the battle lines in this debate do not run along the expected lines of national interests.
The Top Players’ Backlash
For Anthropic and OpenAI, this is more than just an unpleasant footnote. Both companies are preparing billion-dollar IPOs whose valuations—each well over $800 billion—are based on the assumption that corporate customers will permanently pay a premium for the best model available. It is precisely this premium that is eroding the fastest in the very market segment that is crucial to the IPO narrative.
Both labs are responding with a two-pronged strategy. First: more affordable models—and here, Anthropic and OpenAI are now engaged in a full-blown race on a weekly basis. On July 24, Anthropic released Claude Opus 5, a model designed to approach the capabilities of its in-house flagship model, Fable 5, but at half the cost—intended as the new standard for everyday office tasks. Six days later, OpenAI followed suit, taking a much more aggressive approach: prices for GPT-5.6 Luna dropped by 80 percent, and for Terra by 20 percent (details and current pricing table in the previous section). Second, and likely more important in the long term: building customer loyalty through the software layer rather than the model itself. Both companies are aggressively developing tools, coding agents, and deep integrations into enterprise workflows—the moat is no longer the model itself, but rather the implementation surrounding it. Ashwin Gopinath, a former MIT assistant professor and now CEO of the startup Sentra, warns in this context against tools that become too deeply embedded in business processes: Models can be replaced, but a company’s “operational memory” can hardly be—a “Trojan horse,” as he calls it.
Whether this strategy will pay off remains to be seen. But one thing is certain: The era in which “the best model” was automatically also “the only sensible model” came to an end within a matter of months amid the hypersonic pace of AI development. It is being replaced by something more down-to-earth—the question of which tool is actually worth the money for which task.
A sober assessment to conclude
It would be too simplistic to turn all of this into a straightforward narrative about the end of the AI hype. The same companies that are now counting every penny when purchasing models continue to invest hundreds of billions elsewhere in data centers and chips—the cost-cutting is happening at the application level, not the infrastructure level. And the figures on the market shift itself—and this is part of due diligence—are largely third-party platform and analyst data, not audited financial statements—snapshots that could shift again in a few weeks, just as they have done multiple times in the past few weeks.
But what can be said with certainty is this: For the first time since the start of the generative AI boom, companies are treating AI providers just like any other supplier—they’re comparing options, they’re negotiating, and they’re willing to switch. It is precisely this level-headedness—this return to ordinary market dynamics—that may be the real takeaway of the summer of 2026: not that AI has gotten worse, but that customers have stopped acting like fans and started acting like buyers.
And something else is quietly shifting along with this: The story can no longer be neatly framed as “expensive America versus cheap China.” With Luna, OpenAI has pushed its own closed model into a price range that was previously reserved for Chinese open-weight providers—while Kimi K3 from China is, at the same time, more expensive than Google’s Gemini. The real divide no longer runs between two countries, but between two questions that every company must now ask anew for each individual task: How much intelligence does this one step really require—and who is currently offering it at the lowest price? As July 30 demonstrated, the answer to this question can literally change over the weekend.
Der Promptologe, July 26, 2026
Disclosure: This post was created with the help of Anthropic Claude.
Translation: DeepL
My insights from the wonderful world of AI development
List of Sources
WSJ: “Corporate America Has Suddenly Decided to Stop Blowing Money on AI”, 24.7.2026 — https://www.wsj.com/business/china-us-ai-model-costs-53a12e96
WSJ: “Meet the Companies Shelling Out for Top AI Models” — https://www.wsj.com/cio-journal/meet-the-companies-shelling-out-for-top-ai-models-e1fe3375
WSJ: “AI’s Wider Availability Is Good for China, Not Great for OpenAI and Anthropic” — https://www.wsj.com/tech/ai/cheaper-ai-commodity-openai-anthropic-0111da73
WSJ: “What to Know About the Chinese AI Models Rattling U.S. Stocks” — https://www.wsj.com/tech/ai/what-to-know-about-the-chinese-ai-models-rattling-u-s-stocks-1a80a479
CNBC: “Cheap AI could derail OpenAI and Anthropic’s IPOs”, 20.5.2026 — https://www.cnbc.com/2026/05/20/cheap-ai-could-derail-openai-and-anthropics-ipos.html
CNBC: “Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge”, 7.7.2026 — https://www.cnbc.com/2026/07/07/chinese-ai-models-costs-us-openai-anthropic.html
Bloomberg: “Anthropic Unveils More Cost-Efficient Model for Everyday Tasks” — https://www.bloomberg.com/news/articles/2026-07-24/anthropic-unveils-more-cost-efficient-model-for-everyday-tasks
Bloomberg: “China’s Powerful New Moonshot AI Model Closes Gap With US Rivals” — https://www.bloomberg.com/news/articles/2026-07-17/china-s-powerful-new-moonshot-ai-model-closes-gap-with-us-rivals
Bloomberg: “Microsoft Is Replacing OpenAI Image Models in PowerPoint, Bing” — https://www.bloomberg.com/news/articles/2026-07-23/microsoft-replacing-openai-image-ai-models-in-powerpoint-bing
Bloomberg: “Washington’s Foreign Ban on Anthropic’s Top Models May Backfire” — https://www.bloomberg.com/news/newsletters/2026-06-26/white-house-s-ban-on-anthropic-ai-access-may-boost-china-s-open-source-models
Bloomberg Opinion: “China Follows Up DeepSeek Moment by Surprising Silicon Valley with Kimi” — https://www.bloomberg.com/opinion/newsletters/2026-07-17/china-follows-up-deepseek-moment-by-surprising-silicon-valley-with-kimi
manager magazin: “OpenAI vs Anthropic: Wie Anthropic die ChatGPT-Firma OpenAI bei KI-Modellen abhängt” — https://www.manager-magazin.de/unternehmen/tech/openai-vs-anthropic-wie-anthropic-die-chatgpt-firma-openai-bei-ki-modellen-abhaengt-a-b239b174-9959-42f0-a0c6-3a6ce4f58a5a
Handelsblatt: “Kommentar: KI aus China birgt ein unkalkulierbares Sicherheitsrisiko” — https://www.handelsblatt.com/meinung/kommentare/kommentar-ki-aus-china-birgt-ein-unkalkulierbares-sicherheitsrisiko/100241883.html
Handelsblatt: “Kommentar: KI aus China kann ein Segen für Europa sein” — https://www.handelsblatt.com/meinung/kommentare/kommentar-ki-aus-china-kann-ein-segen-fuer-europa-sein/100242119.html
Handelsblatt KI-Briefing: “Murati kannte Geheimnisse von OpenAI – und wählte Chinas Weg” — https://www.handelsblatt.com/technik/ki/ki-briefing/ki-briefing-murati-kannte-geheimnisse-von-openai-und-waehlte-chinas-weg/100240739.html
Handelsblatt: “Moonshot: China-Start-up fordert mit KI Anthropic und OpenAI heraus” — https://www.handelsblatt.com/technik/ki/moonshot-china-start-up-fordert-mit-ki-anthropic-und-openai-heraus/100241002.html
Tomasz Tunguz: “The Unsustainable Subsidy” — https://tomtunguz.com/ai-model-inflation/
metacheles.de: “Kreislauf-Invest: Wie Big Techs KI-Umsaetze sich im Kreis drehen!” (Sekundärquelle mit umfangreicher eigener Quellenliste, s. Artikel) — https://www.metacheles.de/kreislauf-invest-wie-big-techs-ki-umsaetze-sich-im-kreis-drehen/
FourWeekMBA: “OpenRouter Usage Data Shows US AI Models Fell From ~70% to ~30% of Token Volume” — https://fourweekmba.com/ai-openrouter-us-models-token-share-deepseek-volume-revenue-spl/
TechTimes: “Chinese AI Models Lead OpenRouter Traffic” — https://www.techtimes.com/articles/317352/20260529/chinese-ai-models-lead-openrouter-traffic-coding-gains-come-china-data-risk.htm
Yahoo Finance: “Chinese AI Models Now Capture Up to 46% of US Enterprise Token Usage” — https://finance.yahoo.com/technology/ai/articles/chinese-ai-models-now-capture-020440715.html
Yahoo Finance: “Citi flags open-source moment in AI as model access tightens” — https://finance.yahoo.com/technology/ai/articles/citi-flags-open-source-moment-135800457.html
Investing.com: “Open-source AI models benefit from restrictions on frontier systems, Citi says” — https://www.investing.com/news/stock-market-news/opensource-ai-models-benefit-from-restrictions-on-frontier-systems-citi-says-4764190
digitalapplied.com: “The AI Agent Build & Run Cost Index 2026” — https://www.digitalapplied.com/blog/ai-agent-build-run-cost-index-2026
New Sources as of August 3, 2026:
OpenAI: “Advancing the price-performance frontier with GPT-5.6”, 30.7.2026 https://openai.com/de-DE/index/advancing-the-price-performance-frontier-with-gpt-5-6/
VentureBeat: “AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost” — https://venturebeat.com/technology/ai-price-wars-openai-cuts-gpt-5-6-luna-prices-by-80-as-model-competition-shifts-toward-cost
Axios: “OpenAI discounts GPT-5.6 Luna and Terra”, 30.7.2026 — https://www.axios.com/2026/07/30/openai-cuts-prices-gpt-terra-luna5
Forbes: “OpenAI Cuts GPT-5.6 Pricing Up To 80%, As AI Costs Come Under Scrutiny”, 31.7.2026 — https://www.forbes.com/sites/rachelwells/2026/07/31/openai-cuts-gpt-56-pricing-up-to-80-as-ai-costs-come-under-scrutiny/
qz.com / technology.org: “Anthropic launched Claude Opus 5 at half the price of its most powerful AI model”, 24.7.2026
BenchLM.ai: OpenAI/Anthropic/DeepSeek/Google API-Pricing-Übersichten (August 2026) — Sekundärquelle zum Preisabgleich, keine Primärquelle
Raindrop-Sammlung “KI-Newsletter”: Bookmark 1806200869 (”Inkling-Small, GPT-5.6 price cuts, Gemini Robotics 2”, TLDR-Newsletter 31.7.2026) und 1806134348 (”Tesla SpaceX merger, OpenAI price cuts, Stripe’s AI knowledge platform”, 31.7.2026) — Gmail-Links, Inhalt nicht direkt einsehbar, aber als Beleg, dass die Preissenkung in der Fachpresse breit aufgegriffen wurde
Kimi K3, GLM-5.2, DeepSeek V4, Gemini 3.1 Pro, Claude Fable 5/Opus 5, GPT-5.6 Sol: jeweils per Web-Suche gegen aggregierte Pricing-Tracker (pricepertoken.com, openrouter.ai, aipricing.guru u.a.) verifiziert — Sekundärquellen







