Kimi K3: Model Introduction & Practical Guide
Kimi K3 is Moonshot AI's flagship model and, at 2.8 trillion parameters, the first open-source model at the three-trillion scale. Built on the KDA (Kimi Delta Attention) hybrid linear attention mechanism and Attention Residuals, it pairs a 1M-token context window with native vision - and it is designed for long-horizon coding, knowledge work, and agentic execution rather than conversation.
Here's the short version: K3 is the model that closed the gap between open-weight and closed frontier systems on agentic work. The architecture is genuinely novel - hybrid linear attention with attention residuals, sparse MoE routing across 896 experts (16 active), and a Stable LatentMoE framework that lifts scaling efficiency roughly 2.5x over K2. The capability mix is distinctive: it sustains long engineering tasks with minimal supervision, coordinates terminal tools, uses screenshots and visual feedback to refine game/frontend/CAD work, and ships full weights for self-hosting. The trade-offs: it is a premium tier on Moonshot's platform (paid access from a ¥10 minimum), the complete technical report is still rolling out, and at this scale self-hosting requires serious infrastructure.
This guide covers model overview, core features, technical specifications, capability comparison, core advantages, recommended use cases, example prompts, and selection recommendations.
Quick Facts
| Attribute | Value |
|---|---|
| Model name | Kimi K3 |
| Developer | Moonshot AI (Dark Side of the Moon) |
| Category | Flagship MoE large language model (open source) |
| Parameters | 2.8T total (sparse MoE: 896 experts, 16 activated) |
| Architecture | KDA (Kimi Delta Attention) hybrid linear attention + Attention Residuals + Stable LatentMoE |
| Context window | 1M tokens |
| Modalities | Text + native vision (images, screenshots, video frames) |
| Reasoning control | reasoning_effort: low / high / max (default max); thinking mode always on |
| Key strengths | Long-horizon coding, knowledge work, visual reasoning (game/frontend/CAD) |
| Weights | Full open weights released (GitHub / Hugging Face) |
| API | kimi-k3 via api.moonshot.cn / platform.kimi.com |
| Access terms | Flagship tier - requires account recharge (min ¥10); new-user vouchers not applicable |
Table of Contents
- Model Overview
- Core Features
- Technical Specifications
- Capability Comparison
- Core Advantages
- Recommended Use Cases
- Example Prompts
- Selection Recommendations
- FAQ
- Sources & Further Reading
1. Model Overview
Moonshot's K-series has spent the past year repeatedly setting the open-source scale ceiling - by the company's own count, Kimi models held that ceiling in 9 of the 12 months from July 2025 to July 2026. K3 is the next step in that campaign: 2.8 trillion parameters, trained with architectural innovations that specifically target the two constraints that make very long, very deep models hard to train and serve.
The first innovation is KDA (Kimi Delta Attention), a hybrid linear attention mechanism designed to keep information flowing smoothly across very long sequences - the enabler for the 1M-token context. The second is Attention Residuals (AttnRes) - it has delivered on both: K3 sustains long-horizon engineering tasks with minimal human supervision, handles large codebases, coordinates terminal toolchains, and combines software engineering with visual reasoning. The screenshot-driven workflows (game development, frontend work, CAD) point at something new in the open ecosystem: a model that uses visual feedback the way a developer actually does.
The launch surface matters too. K3 powers Moonshot's Kimi product (with Swarm agent clusters and Goal-mode parallel execution for knowledge work), an API with dynamic tool loading and 1M-context caching, and - critically - published weights. For enterprises weighing open versus closed, K3 is currently the largest open option in its capability class.
2. Core Features
Long-horizon coding. Sustains long engineering tasks in large codebases with minimal supervision: understanding, planning, implementation, and terminal tool orchestration.
Visual + engineering fusion. Uses screenshots and visual feedback to improve game development, frontend, and CAD work - iterating on what the output looks like, not just what the code says.
End-to-end knowledge work. Strong performance on public benchmarks and internal evaluations derived from real user-agent workflows: research, analysis, and deliverable generation.
Agent cluster execution. The Kimi product layer supports Swarm agent clusters and Goal-mode parallel task execution - multiple K3 agents working in parallel on decomposed goals.
1M-token context with automatic caching. Long-document and repo-scale work without chunking, with server-side caching to control cost on repeated context.
Dynamic tool loading. Tools and tool_choice are dynamically loaded rather than fixed up front - a practical detail for agents with large tool libraries.
Configurable reasoning effort. reasoning_effort in three levels (low / high / max, default max); thinking mode is always on and reasoning streams separately from the answer.
3. Technical Specifications
| Specification | Detail |
|---|---|
| Total parameters | 2.8T (sparse MoE) |
| Expert routing | 896 experts, 16 activated per token (Stable LatentMoE) |
| Attention | KDA (Kimi Delta Attention) hybrid linear attention + Attention Residuals |
| Scaling efficiency | ~2.5x vs K2 |
| Context | 1M tokens + automatic caching |
| Modalities | Text, native vision (image/video input) |
| Reasoning | Always-on thinking; reasoning_effort low/high/max |
| Streaming | Separate reasoning_content and content streams |
| Weights | Full open weights published |
| API | kimi-k3 (OpenAI-SDK compatible at api.moonshot.cn/v1) |
Access notes. K3 is a flagship-tier model on Moonshot's open platform: API access unlocks after a minimum ¥10 recharge, and cumulative recharge determines account tiers with corresponding concurrency/RPM/TPM limits. The 15-yuan new-user voucher does not apply to K3 - a detail worth knowing before planning an evaluation budget.
4. Capability Comparison
| Dimension | Kimi K3 | Qwen3.8-Max | DeepSeek-V4 Pro |
|---|---|---|---|
| Parameters | 2.8T MoE | 2.4T MoE | 671B MoE (37B active) |
| Context | 1M | 1M | 1M-class |
| Open weights | Yes (first 3T-class open model) | Planned with formal release | Open (MIT) |
| Positioning | Agentic coding + knowledge work | Multi-agent productivity | Reasoning + efficiency |
| Reasoning control | reasoning_effort (3 levels) | Thinking toggle | Two-mode design |
| Ecosystem | Kimi app + Swarm clusters, API | Qoder / Token Plan | DeepSeek API + self-host |
Where K3 wins: open weights at unmatched scale, screenshot-driven engineering workflows, and agent-cluster product features (Swarm/Goal mode) that go beyond single-agent usage. Where to watch: throughput and cost at 2.8T scale depend heavily on Moonshot's inference stack; self-hosting the full model is a data-center project, not a workstation one. For most teams, the API is the realistic path, with open weights as the strategic hedge.
5. Core Advantages
- The largest open weights in the industry. 2.8T parameters published - the strongest available hedge against vendor lock-in at the frontier.
- 2.5x scaling efficiency. The KDA + AttnRes + LatentMoE stack converts compute into capability far more efficiently than K2 - the reason a 3T-class model is trainable and servable at all.
- Visual feedback loops. Screenshot-driven iteration for games, frontends, and CAD is a differentiated capability in open models.
- Long-horizon reliability. Multi-hour engineering tasks with terminal tools and minimal supervision, validated on real internal workflows.
- Dynamic tool loading + 1M context caching. Engineering details that make production agents cheaper and more robust in practice.
- A product ecosystem around the model. Swarm clusters, Goal mode, and the Kimi app give the model immediate application surfaces.
6. Recommended Use Cases
- Long-horizon coding agents: multi-hour engineering tasks, large-codebase comprehension, repo-scale refactors, terminal-driven development.
- Game, frontend, and CAD development with visual iteration: generate, screenshot, critique, refine.
- Knowledge-work deliverables: research synthesis, reports, and presentations generated end to end, with parallel Swarm execution for large scopes.
- Long-document analysis: 1M-token processing of contracts, filings, and technical corpora with cache-assisted repeat queries.
- Self-hosted frontier deployments: organizations that need a top-tier model on their own infrastructure, accepting the scaling cost.
- Tool-heavy agent platforms: dynamic tool loading suits agents with large, shifting tool libraries.
7. Example Prompts
1. Long-horizon coding task
2. Visual frontend iteration
3. Swarm-style knowledge work
4. Million-token contract review
5. Mining game prototype from one prompt
8. Selection Recommendations
Choose Kimi K3 if:
- You want frontier-class agentic capability with open weights as a strategic option.
- Your work is long-horizon coding or engineering with visual iteration.
- You need 1M-token context with caching for production document workflows.
- You want parallel agent execution (Swarm/Goal) as a product feature.
- You are building tool-heavy agents that benefit from dynamic tool loading.
Consider alternatives if:
- You need the absolute cheapest high-volume inference - mid-size models (Qwen3.8-Flash, MiniMax M3) undercut it on price.
- You need multi-agent concurrency plans with predictable flat subscriptions (Qwen3.8-Max Token Plan).
- You require the complete technical report and fully settled evaluation methodology before adoption - parts of K3's documentation are still rolling out.
- Your evaluation budget is minimal: K3 requires a paid account, and new-user credits don't apply.
Sources & Further Reading
- Kimi K3 quickstart - Kimi API docs (official)
- Kimi K3 - Kimi product site
- Kimi API platform
- Moonshot AI official site
- Kimi K3 overview - Kimi K3 guide (Chinese)
Access terms, pricing, and rate limits are set by Moonshot AI and change over time; benchmark figures should be validated against your own workloads. Running the full open weights requires substantial infrastructure - plan self-hosting costs carefully.



