Wan 2.7 Video: Alibaba's Full-Modal Video Creation and Editing Model
Wan 2.7 Video (catalog slug wan-2-7-video) is Alibaba's April 2026 video creation model: text, image, video, and audio inputs in one system, with generation and editing treated as the same workflow - the "video that can be edited like a document" release.
Here's the short version: Wan2.7-Video moved the Wan family "from acting to directing." It takes full-modal input - text, images, video, audio - and handles both creation and post-hoc editing: change a character's face, swap a role, rewrite a plot beat, switch styles, adjust details, and control timing. It also does continuation (extending a 2-second clip to 15 seconds), first/last-frame control, and action imitation. Alibaba shipped it on the Tongyi Wan app, the Qwen app, and Alibaba Cloud Model Studio, with 720P generation priced around ¥0.6 per second on the Qwen platform. This is the model that made "AI video is a one-shot generator" feel outdated.
This guide covers model overview, core features, technical specifications, capability comparison, core advantages, recommended use cases, example prompts, selection recommendations, a workflow, and my verdict.
Quick Facts
| Attribute | Value |
|---|---|
| Catalog slug | wan-2-7-video |
| Developer | Alibaba (Tongyi Lab / Wan) |
| Release | April 3, 2026 |
| Inputs | Text, image, video, audio |
| Capabilities | Generation + editing + continuation |
| Editing | Structure, plot, details, timing |
| Continuation | 2 s to 15 s |
| Control | First/last frame, action imitation |
| Access | Tongyi Wan, Qwen app, Model Studio |
| Pricing (720P) | ~¥0.6 / second (Qwen platform) |
Table of Contents
- Model Overview
- Core Features
- Technical Specifications
- Capability Comparison
- Core Advantages
- Recommended Use Cases
- Example Prompts
- Selection Recommendations
- Workflow: Create, Then Edit Like a Document
- The Bottom Line
- FAQ
- Sources & Further Reading
1. Model Overview
On April 3, 2026, Alibaba's Tongyi Lab released Wan2.7-Video with a clear thesis: video generation had become good at creating clips but bad at fixing them. The release reframed the problem - "make video content editable like a document" - and shipped a system that accepts text, image, video, and audio input, then lets you edit structure, plot direction, local details, and temporal changes.
The practical capabilities follow from that framing: face and character swaps, role changes, style switches, dialogue adjustments with natural lip sync, action imitation, and continuation from short clips (2 seconds to 15 seconds) with first/last-frame control. It is not one more one-shot generator; it is a creation-and-editing system.
Alibaba made it available immediately across its consumer surfaces (Tongyi Wan, Qwen app) and through Model Studio for developers, with the Qwen platform listing 720P video generation at roughly ¥0.6 per second. The "editing" framing is the product, not the marketing.
2. Core Features
Full-modal input. Text, image, video, and audio all feed the same system.
Post-hoc editing. Change faces, characters, plot beats, styles, and details after generation.
Continuation. Extend a 2-second clip up to 15 seconds with coherent motion.
First/last-frame control. Anchor the start and end of generated or extended footage.
Action imitation. Copy a performance's motion into new footage.
Dialogue and lip sync. Adjust lines while keeping natural mouth alignment.
Multi-surface access. Tongyi Wan, Qwen app, and Model Studio API.
3. Technical Specifications
| Specification | Detail |
|---|---|
| Model | Wan2.7-Video |
| Release | April 3, 2026 |
| Input modalities | Text, image, video, audio |
| Generation | T2V / I2V / reference-based |
| Editing | Structure, plot, details, timing |
| Continuation | 2 s -> up to 15 s |
| Controls | First/last frame, action imitation |
| Resolution | 720P class (platform-dependent) |
| Access | Tongyi Wan, Qwen app, Model Studio |
| Pricing | ~¥0.6 / s (720P, Qwen platform) |
Note: exact duration, resolution, and pricing vary by surface (app vs API vs region). The Model Studio model IDs carry snapshot dates (e.g.,
wan2.7-i2v-2026-04-25) - pin the version you test.
4. Capability Comparison
| Capability | Wan 2.7 Video | Wan 2.6 (class) | Hosted flagships (Veo/Sora class) |
|---|---|---|---|
| Editing after generation | Core feature | Limited | Limited |
| Full-modal input | Text/image/video/audio | Text/image/audio | Text/image |
| Continuation | 2 s -> 15 s | Yes | Yes |
| Character/face swap | Yes | Some | No |
| Action imitation | Yes | No | No |
| Lip-synced dialogue edits | Yes | Partial | Yes |
| Open weights | Partial lineage | Open family | No |
Where Wan 2.7 wins: editing depth - no other flagship treats video as a document. Where it loses: vendor benchmark drama aside, hosted Western flagships still lead on some raw generation aesthetics; this is a workflow-led release.
5. Core Advantages
- Edit after generate. The model's reason to exist - fix what generation got wrong.
- Full-modal input. Text, image, video, and audio in one system.
- Continuation and control. 2 to 15 seconds with first/last-frame anchoring.
- Action imitation. Reuse performances across scenes and characters.
- Alibaba surfaces. App, API, and cloud access with clear pricing.
- Document-like workflow. Structure, plot, and timing edits map to editorial thinking.
6. Recommended Use Cases
- Short dramas: generate, then fix plot, characters, and dialogue.
- Advertising: swap faces/characters and restyle campaign footage.
- Content repair: correcting details in already-generated video.
- Storyboarding: continuation and action imitation for previs.
- Localization: adjusting dialogue with natural lip sync.
7. Example Prompts
1. Character swap
2. Plot edit
3. Continuation
4. Action imitation
Prompting guidance: for edits, separate the target change from the invariants ("keep lighting and motion unchanged"). For continuation, describe the action arc so the extension has somewhere to go.
8. Selection Recommendations
Choose Wan 2.7 Video if:
- Your workflow needs editing after generation - swaps, plot fixes, style changes.
- Full-modal input and continuation are requirements.
- You are on Alibaba surfaces (Tongyi Wan, Qwen app, Model Studio).
- "Video as a document" matches how your team thinks.
Choose alternatives if:
- You need open weights for everything (the full 2.7 hosted system is API-first).
- Your benchmark priority is pure generation aesthetics without editing needs.
- You require a Western-cloud vendor relationship.
9. Workflow: Create, Then Edit Like a Document
Practical notes:
- Generate first, then edit - that is the model's intended loop.
- Pin the Model Studio snapshot you test; behavior can shift between releases.
- Name invariants in every edit prompt; the editing depth is a feature only if you use it precisely.
10. The Bottom Line
Verdict: The release that made "AI video is a one-shot generator" sound outdated. Wan 2.7 Video's contribution is workflow, not just quality: full-modal input, post-generation editing, continuation, first/last-frame control, and action imitation, all in one system with Alibaba's apps and API behind it. It is API-first rather than fully open, and raw-generation benchmark debates continue - but for teams that need to fix, extend, and direct video rather than just spawn it, this is the most complete package in the Chinese ecosystem.
Sources & Further Reading
- Wan2.7-Video release coverage - ZhidX
- Wan2.7-Video on Qwen app - C114
- Wan2.7-T2V model page - Qwen AI platform
- Text-to-video guide (Wan 2.5-2.7) - Alibaba Cloud Model Studio
- Tongyi Wan official
Information reflects Alibaba's official release materials and platform documentation as of August 2026. Pricing and limits are vendor-published; verify current values on your platform.

