Kimi Sends Stock Prices Plummeting and Shakes Up the Charts
The AI industry has gotten into the habit of equating quality with benchmark performance. That’s convenient but inaccurate. More relevant: Quality for what purpose, at what cost, under whose control?
The Chinese Kimi K3 model is causing a stir and sending stock prices plummeting: It’s time to compare the most important AI models.
Graphic: Nano Banana 2, Prompt: The Promptologist
On July 16, 2026, Beijing-based Moonshot AI releases a language model called Kimi K3. Two hours after the announcement, chip stocks begin to fall. The Philadelphia Semiconductor Index loses a double-digit percentage of its annual gains within a few days. Nvidia drops just under two percent on the day of the announcement. $392 billion are wiped within hours from expected valuations of OpenAI and Anthropic.
Headlines read: “DeepSeek Moment Number Two.”
The german version:
The pattern is familiar: A Chinese model is released, benchmark tables circulate, the markets react, and Western commentators debate the state of the arms race. The answer then depends on which table you consult.
The CAISI Institute, an agency of the U.S. National Institute of Standards and Technology (NIST), published its own assessment in April 2026: According to this, leading Chinese models lag eight months behind the U.S. frontier. The Stanford AI Index, published a month earlier, reaches a different conclusion: The leading U.S. model is 2.7 percent ahead of the Chinese one. An eight-month lag on one hand, 2.7 percent on the other—the same question, two conflicting reports.
Not two opinions. Two different benchmark sets.
Anyone who understands why this is the case also understands why the standard comparison—benchmark table on the left, Chinese model on the right—measures the wrong thing. And why the more interesting questions lie elsewhere: in price, in ownership, and in the right to shut down a model.
What Benchmarks Measure—and What They Don’t
Depending on the source, Kimi K3 is currently either the best coding model in the world or about six months behind the U.S. Frontier. Both are true—because both measure different things.
The rankings from reputable testing providers show a wider range than the headlines suggest. Vals AI places Kimi K3 between Claude Fable 5 and GPT-5.6 Sol; Artificial Analysis certifies that the model is comparable to top Western models on complex tasks; the Arena ranking places it at number one in the field of web programming. At the same time, the Epoch Capabilities Index places Kimi K3 exactly on the Chinese trend line—not an outlier, but the expected continuation of a trend that has been consistent for two years. Ryan Greenblatt, an AI security researcher, estimates the pre-training gap to the current U.S. frontier at eight months.
Kimi K3 roughly corresponds to the performance level of American closed models from late 2025. That’s impressive. Moonshot AI admits this itself: “While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT-5.6 Sol.”
The Maximum-Effort Problem
Benchmark tables are not neutral measurements. They are marketing material that looks like science—if you don’t look too closely.
Moonshot AI evaluated Kimi K3 on all standard tests using “maximum effort”. That means more computation time per task and more generated intermediate tokens than other labs typically use. The result: numbers that are consistent in themselves but skew comparisons with competitors who measured using standard settings. AI researcher Gavin Leech estimates that about 25 percent of Kimi K3’s benchmark gains are due to targeted “benchmaxxing”—the systematic optimization of a model for the specific tasks in the test suites, rather than for general performance.
Leech attributes another 15 percent of the improvements to a different phenomenon: distillation from Claude’s training or output data. In tests, Kimi K3 responds to the question “Who developed you?” with “I am Claude” in about 10 percent of cases. Formatting patterns, style of argumentation, and sentence structure — in random samples, these closely resemble the patterns of Anthropic models. This is legally permissible. For benchmark comparisons, however, it poses a problem: Kimi K3 inherits Claude’s strengths — and a specific weakness. For security reasons, Claude refuses to provide certain answers in the cybersecurity domain. Distilled models inherit this restraint without the ethical consideration behind it — and therefore perform worse than expected in security tests.
Ethan Mollick (Wharton School, University of Pennsylvania) refers to the third structural problem as “saturated benchmarks”: The standard tests—HLE, GPQA Diamond, SWE-bench—are now so well-known and so frequently included in training data that top models consistently achieve scores of 90 percent and higher. Differences of just one percent then determine whether a model ranks first or fifth, even though such differences would be barely measurable in practice. “Public benchmarks are decreasingly useful,” states OpenAI researcher Dean Ball.
Taking this into account, what remains is this: Kimi K3 is a very good open-weight model — strong in coding, web programming, and agentic tasks, but weaker at complex, multi-step tasks without a technical focus. Independent long-term tests are still needed for a complete assessment.
What matters most to you in an AI model? Use the sliders to weight performance, price, openness, response speed, and German language proficiency according to your own priorities—the ranking changes in real time
Programming Claude / Prompt The Promptologist
The Other Ranking
The benchmark table isn’t wrong. It just doesn’t measure what’s relevant for most decisions.
Anyone who uses a language model not for testing but for work — for document summarization, code generation, translations, or customer service automation — pays per processed token. And here, the ranking looks fundamentally different.
Claude Fable 5, Anthropic’s current flagship, costs $50 per million output tokens. While Kimi K3 currently tops the Chinese models in benchmark tests, DeepSeek V4 Pro pursues a different strategy: nearly comparable performance at a fraction of the cost — $0.87. The difference: a factor of 57.
This is no exception. It’s the norm.
OpenAI’s GPT-5.6 Sol costs $30 per million output tokens, while Gemini 3.1 Pro costs $12. On the Chinese side, not every model follows the same strategy:
Kimi K3 positions itself as the most powerful model and costs $15 per million output tokens.
DeepSeek, on the other hand, consistently focuses on price — V4 Pro costs $0.87
GLM 5.2 at $4.
The most affordable Frontier-class model in this comparison is DeepSeek V4 Flash at just $0.28 per million output tokens—a factor of 179 compared to Claude Fable 5.
What this price difference means in practice can be seen in specific cases. Coinbase, the U.S. crypto exchange, has cut its AI spending in half and now runs around 1,200 automated agents on Chinese models — specifically GLM and Kimi. Shopify is said to have switched its automated product catalog processing to Qwen, achieving a 75-fold reduction in costs per processed item. Airbnb CEO Brian Chesky publicly described Qwen as “very good… fast and cheap.” On the OpenRouter platform, which is considered a neutral marketplace for model access, DeepSeek is now the largest single provider, accounting for 17.6 percent of all processed tokens—ahead of Alibaba’s Qwen at 13.9 percent and Anthropic’s Claude at 13.3 percent.
The Trend Reversal
The price history of language models is not a simple downward curve. From 2023 to 2024, the costs of U.S. flagship models fell dramatically. When it launched in March 2023, GPT-4 still cost $60 per million output tokens. With GPT-4o — released in May 2024 — that price dropped to $10: an 83 percent decline in 14 months.
Then a trend reversal began, one that is clearly visible in the data but has received little attention in the public eye. Claude Fable 5 costs $50 today—five times as much as GPT-4o did back then. GPT-5.6 Sol is priced at $30. This is no mistake. It is the result of a structural shift.
U.S. providers have become entangled in an infrastructure race that exceeds their free cash flow. Microsoft, Google, and Amazon are investing hundreds of billions in data centers and chip capacity. These expenses must be recouped—through the prices that enterprise customers pay. At the same time, Chinese providers are strategically subsidizing their inference costs. Market observers consider it unlikely that DeepSeek is breaking even with an output price of $0.87 per million tokens. Prices are driven by market share, not margins.
The result is a market divided into two segments: U.S. models optimize for performance and profitability. Chinese models optimize for market penetration.
What exactly does your use case cost? Select a scenario and a monthly document volume — the calculator shows what you would actually pay with each of the models.
Programming Claude / Prompt The Promptologist
Open Weights: The Ownership Question
Kimi K3 is not an open-source model. It has been positioned as an open-weight model — a distinction that is often confused but is significant for practical decisions.
Open Weight means: The model’s trained parameters are made publicly available. Kimi did so as announced on July 27. Anyone who wishes to do so can download them, run them locally, customize them, and integrate them into their own systems. What is not publicly available is the source code of the training framework, the complete training data, and the exact post-processing procedures. In the case of Kimi K3, the license is “Modified MIT” — a permissive license with restrictions that Moonshot has not yet fully disclosed in detail. DeepSeek V4 Pro and GLM 5.2 are already available as Open Weight models.
For companies and public institutions, Open Weight offers an option that simply doesn’t exist with closed models: self-hosting. Anyone who runs a model on their own infrastructure isn’t dependent on American cloud providers. No API that can be shut down. No terms of service that can be changed unilaterally.
This is not an abstract consideration. In July 2026, parts of the Trump administration considered placing Moonshot AI on the Entity List — a measure that would effectively prohibit U.S. companies from doing business with the Chinese provider. Another precedent: The Trump administration temporarily forced Anthropic to withdraw a model. For companies that treat AI as strategic infrastructure, the question of who can shut down the model in case of doubt is not a technical one — it is a political one.
For European companies, a third dimension comes into play. Anyone who sends personal data to a closed model operated from the U.S. faces compliance issues that open-weight models running on their own European infrastructure generally do not raise. The GDPR compliance of U.S. cloud services has been a persistent source of uncertainty since the Schrems II ruling and its follow-ups. Self-hosted open-weight models do not fully resolve this — but they shift the point of control to where European law applies.
How have costs per million tokens evolved since 2023? The logarithmic timeline shows the dramatic price decline through 2024 — and the current trend reversal among U.S. flagship models
Programming Claude / Prompt The Promptologist
Six Models Profiled
Three U.S. providers, three Chinese. All currently available, all accessible via API, all with their own characteristics. What they can do, what they cost—and for whom they’re suitable in which situations.
Claude Fable 5 — Anthropic, U.S. — Closed-source
$10 input / $50 output per million tokens
Anthropic’s current flagship model positions itself on quality, not price. Strengths: in-depth analysis, nuanced handling of ambiguous issues, long context up to one million tokens, coding, security. Moonshot AI itself cites Fable 5 as the model that even surpasses Kimi K3. Recommendation: for high-quality individual tasks where response quality matters most and the volume is manageable. Clearly the most expensive provider in this comparison in terms of price—not suitable if millions of documents need to be processed monthly.
GPT-5.6 Sol — OpenAI, USA — Closed-source
$5 input / $30 output per million tokens
OpenAI’s current standard model strikes a comparatively good balance between price and performance in the U.S. market. Broad developer ecosystem, numerous third-party integrations. Long-context tier starting at 272,000 tokens ($10 input / $45 output). The most obvious choice for teams with existing OpenAI infrastructure. Price becomes a critical factor when high volumes are planned—in that case, a cost comparison with the Chinese alternatives is recommended.
Gemini 3.1 Pro — Google, USA — Closed-source
$2 input / $12 output per million tokens
The most affordable U.S. frontier option in this comparison. Unique feature: seamless integration with Google Workspace, Google Cloud, and the entire Google ecosystem. Long-context tier starting at 200,000 tokens ($4 / $18). Those already working within the Google infrastructure still pay a significant premium compared to DeepSeek—but in return, they get direct integration with existing productivity tools without additional middleware.
Kimi K3 — Moonshot AI, China — Open Weight
(since July 27, 2026)
$3 input / $15 output per million tokens
In the same price segment as GPT-5.6 Sol, but as an open-weight model with the option to run it on your own infrastructure. Key strengths: front-end development, 3D rendering, agentic coding, parallel task processing (K3 Swarm Max variant). Weakness: complex, multi-step non-coding tasks; insufficient independent test data for long-term evaluation. The “maximum effort” problem with benchmark figures must be taken into account.
DeepSeek V4 Pro — DeepSeek, China — Open Weight
$0.44 input / $0.87 output per million tokens
The price leader in the high-performance segment. 57 times cheaper than Fable 5 in terms of output; the largest single provider on OpenRouter by token volume. Suitable for high-volume applications: document processing, automated summaries, batch translations, customer service agents. Open Weight, self-hostable. Consider this: DeepSeek is a Chinese company, and the regulatory landscape in the EU and the U.S. is evolving.
GLM 5.2 — Zhipu AI, China — Open Weight
$1 input / $4 output per million tokens
The middle ground between price and quality on the Chinese side. Coinbase runs some of its 1,200 agents on GLM. Benchmarks place it below Kimi K3 and DeepSeek V4 Pro, but well above the budget segment. Particularly suitable: as a basis for custom fine-tuning, where the favorable licensing terms have a direct impact. Context: 205,000 tokens — the smallest context window in the comparison, but sufficient for most applications.
When were the most important models released—and how large was the gap between the strongest U.S. model and the strongest Chinese model? The timeline shows how this gap has evolved.
Claude Programming / Prompt: The Prompt Log
What Remains
Benchmark tables aren’t worthless. When read with caution, they are a useful starting point. What they don’t do: provide answers to the questions that matter in practice.
Anyone who wants to process 50 million documents a month doesn’t ask about the GPQA Diamond Score. Anyone who wants to know whether they can shut down a model if the political situation changes doesn’t ask about the SWE-bench ranking. Anyone checking whether a solution can be built to be GDPR-compliant doesn’t ask about the Arena leaderboard.
The AI industry has gotten into the habit of equating quality with benchmark rankings. That’s convenient but inaccurate. The more relevant question is: Quality for what, at what cost, and under whose control?
The Kimi-K3 moment wasn’t a shock. It was further confirmation of a trend that has been emerging in the data for the past year and a half: High-performance Chinese models are arriving faster than expected, cost significantly less than U.S. alternatives, and a growing proportion of them are available for free use as open-source software. At the same time, U.S. providers are raising prices under capex pressure.
This is neither a disaster nor a surrender. It is a market undergoing structural change—and one in which the old benchmarks no longer provide a sufficient explanation.
All interactive graphics in the overview:
Please note:
Prices and all other information in this article: As of July 2026.
API prices from Chinese providers fluctuate monthly. U.S. providers change their prices less frequently, but tiered pricing and caching conditions are subject to change without notice: Be sure to check the providers’ pricing pages directly before making any usage decisions!
Information regarding Shopify (75× cost reduction) is based solely on indirect sources; there is no official confirmation.
Calculator programming: Claude, prompt by The Promptologist
The Promptologist, July 27, 2026
Disclosure: This post was created with the help of Anthropic Claude.
Translation: DeepL










