Quick Answer
Kimi K3 is a 2.8-trillion-parameter AI model released July 16, 2026 by Beijing-based Moonshot AI, billed as the largest open-weight model yet. It is free in the Kimi app and strong at agentic coding, but independent testing measured a 51% hallucination rate, and hosted requests are processed on servers in China.
Key Takeaways
- 2.8 trillion parameters with a roughly 1-million-token context window, built for long coding and agentic tasks
- Free tier in the consumer Kimi app; API pricing is about $3 per million input tokens and $15 per million output
- The downloadable weights are scheduled for July 27, 2026 and had still not shipped as of this update
- Independent testing measured a 51% hallucination rate, up from 39% on K2.6 — alongside real accuracy gains
- The main practical concern for most people is data residency: hosted requests are processed in China
What is Kimi K3?
Kimi K3 is the flagship AI model from Moonshot AI, a Beijing-based lab best known for its Kimi assistant. Unveiled at the World AI Conference in Shanghai on July 16, 2026, it landed hard: within hours it climbed to the top of the Frontend Code Arena leaderboard, and coverage described it as the largest open-weight model ever released. US AI stocks dipped as investors digested what a freely available frontier-class model might mean for companies selling access to closed ones.
Technically, it is a mixture-of-experts model with about 2.8 trillion total parameters, of which only a small fraction — 16 of its 896 experts — is activated for any given token. That sparsity is what makes a model this large practical to run at all. It carries a context window of about one million tokens, enough to take an entire codebase or a book-length document in a single pass, and it includes native visual understanding and an always-on reasoning mode.
The reason it is worth understanding, even if you never use it, is what it represents: open-weight models from Chinese labs have moved from “impressive for the price” to genuinely competitive on some frontier tasks. That shift is reshaping the market we track in our roundup of the top AI models.
How good is Kimi K3 really?
Better than the hype in some places, worse in others — and the most useful source here is Moonshot itself. On its own website, the company states that K3’s overall performance still trails the most powerful proprietary models, naming Claude Fable 5 and GPT-5.6 Sol, while saying it showed frontier-level performance across Moonshot’s evaluation suite and consistently outperformed other tested models.
That is a notably measured claim from a vendor, and it is the right anchor. The leaderboard result that drove the headlines — jumping to first place on Frontend Code Arena — is real but narrow: it measures one category of task. Broader head-to-head comparisons reported since launch suggest the top proprietary models still win the majority of benchmarks overall, particularly on frontier engineering and vision work, while K3 does especially well on long-horizon agentic coding and terminal-driven work, and does it far more cheaply.
Since launch, independent testing has added an important wrinkle. As TechTimes reported, evaluation firm Artificial Analysis measured K3’s hallucination rate on its AA-Omniscience benchmark at approximately 51%, up from 39% on the previous K2.6 — even as raw factual accuracy improved from 33% to 46%. In other words, the model answers more questions correctly than its predecessor while also producing more confident wrong answers, a figure that does not appear in Moonshot’s own published charts. For fairness, the same benchmark records Claude Fable 5 at 54.9%, so the headline number is in line with some frontier peers; what stands out is the generation-over-generation increase. The same testing measured a median time-to-first-token of over two minutes in the always-reasoning mode, far slower than typical interactive assistants — fine for background agentic jobs, noticeable in a chat window.
Two caveats apply to every number you will read. Most published benchmarks at this stage are either self-reported by the vendor or drawn from leaderboards with their own methodologies, and full independent replication requires the open weights, which had not yet shipped at the time of this update. Treat single-benchmark wins as evidence about that benchmark, not a verdict on the model. If you want a grounded comparison of the models most people actually use day to day, see our Claude vs ChatGPT vs Gemini comparison.
Is Kimi K3 free, and how do you use it?
There are three practical routes in, and the free one covers most people:
- The Kimi app and web chat. There is a free consumer tier, which is how most casual users will try it. This is the equivalent of using ChatGPT’s free plan. Demand has been heavy enough that Moonshot briefly paused new paid subscriptions in the week after launch, citing capacity constraints.
- The API. Paid, at roughly $3 per million input tokens and $15 per million output tokens, with cached input far cheaper at around $0.30. That is substantially below flagship US model pricing, which is a large part of K3’s appeal to developers — though it is roughly triple what Moonshot charged for K2.6.
- Self-hosting the weights. Theoretically the most private option, and realistically out of reach for individuals — a 2.8-trillion-parameter mixture-of-experts model needs a serious multi-GPU cluster, not a gaming PC.
For developers weighing it against the incumbents for coding work specifically, our guide to the best AI coding assistants covers where each tool fits.
Is Kimi K3 actually open source?
Partly, and the distinction matters. “Open weights” means the trained model parameters are published so anyone can download, run, and fine-tune the model — it does not necessarily mean the training data or code are public, and it is not the same as a fully permissive open-source licence.
More importantly, the weights are still not available as this update goes out. Moonshot committed to publishing them on Hugging Face, along with a technical report, by July 27, 2026 — now a day away — under what has been described as a modified MIT-style licence. Until those files actually land, K3 is a hosted service like any other, not something you can run yourself, and independent trackers have so far declined to classify it as an open model for exactly that reason. If you are planning anything commercial around it, read the final licence text when it publishes rather than relying on the label.
Why is the White House accusing Moonshot of copying a US model?
On Wednesday, July 22, 2026, Michael Kratsios, director of the White House Office of Science and Technology Policy, posted on X that the administration has information that Moonshot AI distilled Anthropic’s Fable model to produce K3. He wrote that to do so the company “developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection.” He separately alleged that Moonshot obtained servers using Nvidia GB300 hardware, including via Thailand — chips that are not permitted for sale to China. Treasury Secretary Scott Bessent subsequently raised the prospect of sanctions and Entity List designations.
Several things need stating plainly, because this is an allegation rather than a finding. As CyberScoop reported, Kratsios did not provide details on how the government reached its conclusion, and no public evidence has been released. Moonshot did not return a request for comment and, as of this update, has not publicly answered the claims. When Anthropic made broadly similar allegations against Moonshot earlier in 2026, Beijing described them as groundless. Some AI researchers have publicly questioned whether the timeline between Fable’s release and K3’s launch supports distillation at the scale alleged. At least one member of Congress has also questioned the consistency of threatening sanctions while high-performance AI chip sales to China continue.
It is also worth understanding what distillation is, because the word is doing a lot of work. Training a smaller or newer model on the outputs of a more capable one is a standard, widely used technique across the industry — Kratsios explicitly said legitimate distillation plays an important role in AI development. The dispute is over scale, concealment, and authorisation, not the technique itself. Where exactly the line sits is genuinely contested, and no court or regulator has ruled on it. This article reports the claims; it does not take a position on their merits. For background on how export policy has shaped this landscape, see our coverage of US AI export controls.
Is Kimi K3 safe to use?
For most everyday purposes, the risk is not that the model will do something malicious. It is about where your data goes and who can access it.
When you use the Kimi app or Moonshot’s hosted API, your prompts are processed on servers in China and are subject to Chinese law and the company’s own retention policies. That is the same consideration we walked through for DeepSeek, and the sensible guidance is identical: fine for general questions, brainstorming, and public information; not appropriate for personal identifying details, health or financial information, confidential work material, or proprietary code. Many employers restrict tools like this outright, so check your workplace policy before pasting anything work-related. Reporting around the launch has also referenced an April 2026 cross-user data exposure affecting the Kimi service that Moonshot has not publicly addressed; the company has made no statement, which is itself worth weighing when deciding what to paste into a hosted chat.
The open-weight release changes this calculation for organisations, though not for individuals. Once the weights are published, a company can self-host K3 entirely inside its own infrastructure, so no data leaves its environment at all — which is precisely why regulated businesses find open models attractive. That option requires substantial hardware and engineering, so it is a corporate route rather than a personal one.
Separately, note that content moderation and topic handling differ across models and jurisdictions. Independent testing has repeatedly found that Chinese-developed models decline or deflect on politically sensitive subjects in ways US models do not, which is worth knowing if you are using one for research. And given the measured hallucination rate, treat K3’s confident factual answers the way you should treat any frontier model’s: verify before you rely on them.
Should you use Kimi K3?
For casual use, there is little reason to switch. The free tiers of the mainstream assistants are strong, and you gain nothing by routing everyday questions through a service with extra data-residency considerations.
For developers, it is worth a look — the cost difference is real, and the long-context, long-horizon agentic coding strengths are genuine. Test it on non-sensitive work, benchmark it on your own tasks rather than trusting leaderboards, keep verification in the loop for anything factual, and keep an eye on how it compares to GPT-5.6 and the other frontier releases landing this summer.
For organisations, the interesting moment arrives with the weights — now scheduled for tomorrow. A frontier-class model you can run entirely in your own environment solves a data-governance problem that no hosted API can. That is the development actually worth tracking here — more than any single leaderboard position, and more than the political dispute currently surrounding it.
FAQ
What is Kimi K3?
Kimi K3 is a large AI model released on July 16, 2026 by Moonshot AI, a Beijing-based lab. It uses a mixture-of-experts design with about 2.8 trillion total parameters, has a roughly 1-million-token context window, and is positioned as an open-weight competitor to closed frontier models from OpenAI and Anthropic.
Is Kimi K3 free to use?
The consumer Kimi app and web chat include a free tier, which is enough for most casual use. The API is paid, at roughly $3 per million input tokens and $15 per million output tokens, with cached input around $0.30 — noticeably cheaper than flagship US models. Self-hosting the weights is free of licence cost but requires substantial hardware.
Is Kimi K3 better than ChatGPT or Claude?
Not overall, by Moonshot’s own statement: the company says K3 still trails Claude Fable 5 and GPT-5.6 Sol on overall performance. K3 did take first place on one coding leaderboard and performs strongly on long-horizon agentic coding at much lower cost. Most published benchmarks are self-reported or leaderboard-based, so treat single wins cautiously.
Does Kimi K3 hallucinate more than other models?
Independent testing by Artificial Analysis measured K3’s hallucination rate at about 51%, up from 39% on its predecessor, even as raw accuracy improved from 33% to 46%. On the same benchmark, Claude Fable 5 records 54.9%, so the number is comparable to some peers — the notable part is the sharp rise between generations. Verify factual output before relying on it.
Is Kimi K3 safe and private?
Using the Kimi app or hosted API means your prompts are processed on servers in China under Chinese law, and reports have referenced an April 2026 cross-user data exposure that Moonshot has not publicly addressed. That is fine for general, non-sensitive tasks but unsuitable for personal details, financial or health data, or confidential work. Once the open weights publish, organisations can self-host to keep data entirely in their own environment.
Why is the US government accusing Moonshot AI of stealing from Anthropic?
On July 22, 2026, White House OSTP director Michael Kratsios said the administration has information that Moonshot distilled Anthropic’s Fable model to build K3, and alleged it accessed restricted Nvidia hardware. No public evidence was provided, Moonshot has not responded to the claims, and some researchers have questioned the timeline. These remain allegations, and no regulatory or legal finding has been made.
Further reporting on the allegations and the Nvidia hardware claims is available from The Hill.
