12 reasons U.S. companies are turning to cheaper Chinese AI models
The meter is running, and American companies are watching every token tick by. In one busy corner of the AI market, Chinese models now carry most of the load. Reuters reported in July 2026 that they handled about 60% of tokens used by U.S. companies on OpenRouter, a service that lets customers reach many models through one connection.
That number needs a bright warning label. It covers OpenRouter, not the whole U.S. business market. Still, the platform studied more than 450 trillion tokens from January 1 through June 14, 2026. Its logs show a real change in how developers buy AI, even if Amazon, Microsoft, Google, and direct vendor deals remain outside that view.
Price lit the match, but it does not tell the whole story. NIST found that DeepSeek V4 was eight months behind the U.S. frontier across its test mix. The same review found it cost 53% less on one test and 41% more on another. The bargain can be real, yet every company still has to weigh security, privacy, censorship, ownership, and politics.
The price gap can be enormous

For high-volume work, a tiny rate difference grows teeth. OpenRouter listed DeepSeek V4 Flash at 9 cents for input and 18 cents for output per million tokens on its cheapest route. It listed GPT-5.5 at $5 and $30.
Rates change by host, speed, and data policy, but that example helps explain why a finance chief may ask engineers to test the cheaper model.
CNBC reported that Justin Summerville of OpenRouter put the typical savings from open Chinese models at 60% to 90%. Harpreet Arora, Vercel’s head of agentic infrastructure, gave CNBC the blunt version: “Price is doing the work here.”
NIST adds needed restraint. DeepSeek V4 beat a similar U.S. model on cost in five of seven tests, not all seven.
AI agents consume tokens at a fierce pace

A chatbot may answer once. An AI agent can plan, search, call tools, check its work, and try again.
OpenRouter found that agent work used about 15 times more tokens per request than ordinary human use. Agent tokens passed human tokens around February 1, 2026. At that scale, a model that saves a few dollars per million tokens can cut a large monthly bill.
Coinbase shows how those savings can land inside a real company. Business Insider reported that CEO Brian Armstrong shared a chart with token use near a company high as AI spending fell to nearly half its peak.
The chart gave no exact timeline, so it cannot prove one model caused the drop. It does show why caching, shorter prompts, and cheaper defaults now matter.
Many jobs do not need the smartest model

Companies rarely need a digital genius to tag support tickets, pull fields from invoices, summarize meeting notes, or draft routine code.
NIST’s July 2026 review found that GLM-5.2 had capabilities similar to GPT-5.2, a U.S. model released in December 2025. NIST also placed its cyber skill near Anthropic’s Opus 4.6. That is plenty of power for many narrow jobs.
Results still change from test to test. In NIST’s DeepSeek V4 review, the model scored 74% on SWE-Bench, below GPT-5.5 at 81%. It reached 97% on one math test, above Opus 4.6 at 92%.
For a company processing a million simple requests, that uneven scorecard makes task tests more useful than one general rank. That is a practical buying test.
The performance gap is measured in months

Chinese models once looked like distant followers. NIST’s May 2026 work placed DeepSeek V4 roughly eight months behind the leading U.S. systems across 16 benchmarks and 35 models.
That gap still matters for hard research, cyber work, and complex software projects. It matters far less for a basic workflow that costs ten times more on the top model.
The gap also moves fast. GLM-5.2 arrived on June 16, and NIST called it probably the strongest open model with downloadable weights at release. Its overall skill matched GPT-5.2 in the agency’s tests.
Last winter’s premium skill can become this summer’s low-cost option. Buyers now test Chinese releases at once instead of waiting a year. That changes budget talks.
Downloadable weights give technical teams more control

An open weight model can run on a company’s own machines or through a chosen U.S. host. That can reduce vendor lock-in and let engineers tune the system for one job.
Z.ai released GLM-5.2 with a one-million-token window, large enough to read long code bases and project files. A private deployment can also keep prompts away from the model maker’s public service.
Control has a darker edge. NIST warned that safeguards on open models can be removed after download. Still, Nvidia CEO Jensen Huang argued in Axios that “Open-source models that are excellent should be used.”
He said local sandboxes can block outside access. His view does not erase risk, but it explains why some security teams prefer code they can inspect over a closed service they cannot see.
Model routers make switching much easier

Companies no longer have to marry one AI vendor. Business Insider reported that Coinbase was testing GLM-5.2 and Kimi K2.7 as defaults inside its gateway, then sending harder planning work to stronger models.
Armstrong laid out five cost controls, including routing, caching, shorter context, and clearer spending data. That turns model choice into a rule, not a companywide migration.
OpenRouter now offers access to more than 400 models through one interface. Its June data showed DeepSeek’s share rising from 9% to 18% in about six months. Once an engineering team builds a routing layer, a cheaper Chinese model can win thousands of small tasks without replacing the premium model used for the hardest work.
Startups can buy more runway

For a young company, AI spending can compete with payroll. CNBC’s reporting, recapped by The Decoder, found that the 25-person assistant startup Lindy moved all managed traffic from Claude to DeepSeek through U.S. hosting.
CEO Flo Crivello expected savings of millions of dollars within months. His explanation was plain: “It’s a matter of survival for the business.” That case does not fit every bank, hospital, or defense contractor. It does show how fast a startup can move once inference costs pass staff costs.
OpenRouter found DeepSeek V4 Flash carried 70% of DeepSeek agent traffic by the end of May, one month after release. Cheap models gain ground fastest where runway is short and switching is easy.
Security risk can be boxed into lower-stakes work

A cheap engine still needs brakes. NIST found that GLM-5.2 allowed help with cyber exploit development and blocked fewer sensitive biology questions than U.S. reference models. It also appeared tougher against prompt attacks than older Chinese models.
That mixed record supports a narrow use case: routine coding in a sealed test area, not free access to payment systems or customer accounts.
The newest models can do real damage if given tools and weak targets. NIST and the U.K. AI Security Institute found Kimi K3 reached step 17 of a 32-step mock attack, and it finished once in 10 tries.
Université de Montréal professor and Mila founder Yoshua Bengio warned Reuters that open model safeguards “are easier to remove.” Firms still adopt them, but smart teams limit permissions, test outputs, and keep humans in the loop.
Local hosting can reduce privacy exposure

The privacy issue changes with the route. DeepSeek’s hosted service says it collects prompts, uploaded files, IP addresses, device details, and chat history.
Its February 2026 policy says it stores personal data in China and may use interactions to improve its technology. A U.S. company sending health records, legal files, or trade secrets through that service could create serious contract and compliance problems.
A downloaded model is a different product from the public app. OpenRouter said DeepSeek’s low-cost first-party route may train on customer data, but some Western hosts offer no training on prompts for roughly twice the price.
Local hosting can go further by keeping data inside a company network. It raises hardware and staffing costs, so privacy protection belongs in the price comparison.
Censorship matters less in some narrow tasks

Political bias is not a vague fear. A 2025 NIST review found three DeepSeek models repeated inaccurate or misleading Chinese Communist Party narratives four times as often as four U.S. reference models on sensitive questions.
NIST tested downloaded weights, so the pattern did not come only from DeepSeek’s public website. A news desk or policy team should treat that result as a loud alarm.
A warehouse system that classifies part numbers faces a different risk than a chatbot writing about Taiwan. That is why some firms fence Chinese models into coding, extraction, and other narrow work.
NIST’s study covered three DeepSeek versions, not every Chinese model. The test design itself points to the right response: check each model, language, topic, and host before production, then keep political or public-facing content out of an untested system.
Intellectual property risk remains an allegation, not a verdict

Washington’s concern has grown sharper. Reuters reported on July 24 that U.S. officials accused Moonshot AI of using outputs from Anthropic’s Fable 5 to train Kimi K3 through distillation.
The Commerce Department was also investigating access to advanced U.S. chips. Moonshot did not respond to Reuters, and no public court ruling had proved the model-theft claim at publication time.
That distinction matters for a buyer. NIST’s 2025 DeepSeek report stated that it did not investigate distillation or how the models were developed.
Companies can still reduce exposure by checking licenses, recording model versions, asking hosts for indemnity, and keeping generated code under review. Cheap access does not settle who owns the knowledge behind the output.
A mixed model stack can hedge political shocks

U.S. and Chinese policy can change faster than a software budget. Reuters reported that Washington was considering sanctions tied to alleged IP theft, and Beijing was weighing limits on overseas access to Chinese models.
A company built around one foreign API could lose service, face new disclosure rules, or spend weeks replacing a blocked tool. That risk can push firms toward a mixed stack instead of away from Chinese models.
Reuters put its U.S. company share near 60% on OpenRouter so that a sudden cutoff could hit American users too. Downloadable weights can preserve access, but sanctions or license changes may still create legal trouble. The safest bet may be several tested models, clear fallback rules, and no single country controlling every task.
Key takeaways

The shift is real, but its scale has limits. Chinese models handle about 60% of tokens from U.S. companies on OpenRouter, not 60% of the whole American AI market.
Cost leads the story. OpenRouter has posted price gaps measured in dozens of times, and NIST found DeepSeek V4 cheaper on five of seven tests.
The next year will test the bargain. NIST found an eight-month capability gap for DeepSeek V4 and four times as many misleading party narratives in its 2025 DeepSeek tests.
Companies can cut exposure through local hosting, strict routing, model tests, and human review. Cheap AI is easy to find. Cheap AI that stays safe, private, lawful, and available is the harder prize.
Disclaimer – This list is solely the author’s opinion based on research and publicly available information. It is not intended to be professional advice.
You May Also Like: OpenAI Says AI Models Bypassed a Test Sandbox and Breached Hugging Face for Benchmark Answers
Like our content? Be sure to follow us
