The Kimi K3 flagship model just became Moonshot AI’s only current flagship, as the company retired its older Kimi K2 and K2.5 lines in favor of one unified release. Moonshot built K3 as a 2.8-trillion-parameter open-weight model with a full 1-million-token context window, and it’s positioned as one of the strongest open models available today — strong enough that it briefly topped the Frontend Code Arena leaderboard ahead of Claude and GPT-class models on that specific benchmark. This guide covers what K3 actually offers and how to start using it, whether you want a quick chat or a full coding setup.
What Makes Kimi K3 Different
Because K3 runs on a new architecture called Kimi Delta Attention, paired with a technique Moonshot calls Attention Residuals, it handles very long sequences more efficiently than earlier Kimi models. As a result, Moonshot reports roughly 2.5x the scaling efficiency of K2 — meaning it converts the same compute into noticeably more usable capability. The model also includes native visual understanding, so it can read screenshots and images directly rather than needing a separate vision model bolted on.
| Spec | Kimi K3 |
|---|---|
| Parameters | 2.8 trillion (Mixture-of-Experts) |
| Context window | 1,048,576 tokens |
| Architecture | Kimi Delta Attention + Attention Residuals |
| Vision support | Yes, native — images and video |
| Reasoning mode | Always-on thinking (adjustable effort: low/high/max) |
| License | Open weights, Kimi K3 License (revenue-based terms for large companies) |
What You’ll Need
- For casual use: nothing but a browser and a free account at kimi.com
- For API access: a Moonshot account with at least a $1 top-up (this unlocks the flagship tier)
- For coding workflows: Python 3.9+ if you’re calling the API directly, or a tool like Kimi Code / OpenCode if you want an agent-style setup
Step 1: Try Kimi K3 in the Browser
Step 1Go to kimi.com and start a new chat — no account is strictly required to try a basic conversation. This is the fastest way to get a feel for K3’s reasoning and long-context handling before committing to API setup.
Step 2: Get API Access
Step 2Create an account on platform.kimi.ai, then top up at least $1 to unlock the flagship K3 model — this top-up amount also determines your rate limits, so a larger initial top-up gives you more concurrency and throughput from the start.
Step 3: Make Your First API Call
Step 3Install the OpenAI Python SDK, since K3’s API is OpenAI-compatible. Then point the base URL to Moonshot’s endpoint instead of OpenAI’s, and reference the model as “kimi-k3” in your request.
pip install --upgrade 'openai>=1.0'
from openai import OpenAI
client = OpenAI(
api_key="YOUR_MOONSHOT_API_KEY",
base_url="https://api.moonshot.ai/v1",
)
completion = client.chat.completions.create(
model="kimi-k3",
messages=[{"role": "user", "content": "Explain this codebase in one paragraph."}]
)
Step 4: Adjust Reasoning Effort
Step 4K3 always thinks before answering — there’s no way to fully disable it. If a task is dragging on longer than needed, set the reasoning_effort parameter to “low” instead of the default “max” to speed up simpler requests.
Step 5: Use It for Long-Context Coding
Step 5K3’s 1-million-token window means you can paste an entire mid-sized codebase into a single request. Context caching happens automatically too, so repeated requests against the same large document or codebase don’t re-process the full context each time, which keeps costs down on multi-turn sessions.
Step 6: Connect It to a Coding Agent (Optional)
Step 6Because K3 exposes an OpenAI-compatible API, you can plug it into agent tools like OpenCode instead of only using it through raw API calls. Moonshot also offers its own open-source coding tool, Kimi Code, which defaults to K3 but can work with other model providers too.
The Kimi K3 Workflow at a Glance
Frequently Asked Questions
Is the Kimi K3 flagship model free? Chatting with it at kimi.com is free to try. API access requires a minimum $1 top-up, after which usage is billed at $3 per million input tokens and $15 per million output tokens.
Can I run Kimi K3 on my own hardware? Technically yes, since the weights are open — but at 2.8 trillion parameters, realistic self-hosting requires datacenter-grade GPU hardware, not a consumer machine.
How does Kimi K3 compare to Claude or GPT? Moonshot’s own benchmarks place K3 just behind the very top proprietary models on overall performance, while it led specific benchmarks like front-end coding arena votes on launch day. Results vary by task, so it’s worth testing K3 directly against your own use case rather than trusting a single leaderboard.