1. Native Audio-Video Generation
Wan 2.5 generates video and synchronized audio together, combining visuals, dialogue, music, ambience, and sound effects in one process to reduce separate audio generation and post-production work.
Wan 2.5, an Alibaba Wan video model, is now available on Lanta AI. Turn text, images, and audio into cinematic videos with synchronized sound, natural motion, and precise lip sync. Effortlessly bring characters and scenes to life for talking videos, music performances, product promotion, or social media content.
Amusement Park Popcorn Selfie
Puppy Weightlifting
Poppy Garden Portrait
Underwater Dolphin Selfie
Streetwear on a Giant Soda Can
Woman Riding a Giant Mushroom
Floating Pastry Hot Air Balloon
Arctic Fox in Snow
Wan 2.5 goes beyond basic video generation with native audio, audio-driven performance, automatic sound, precise lip sync, and better visual consistency, giving you more control from prompt to final video.
Wan 2.5 generates video and synchronized audio together, combining visuals, dialogue, music, ambience, and sound effects in one process to reduce separate audio generation and post-production work.
Wan 2.5 can use uploaded voice, dialogue, music, or other audio to drive character performance, helping movements and lip sync follow the sound for more natural talking, singing, and music videos.
Wan 2.5 synchronizes character lip movements with speech and singing, making dialogue and performances feel more natural without requiring a separate lip-sync tool or additional editing.
Wan 2.5 can automatically create voices, ambient sound, music, and sound effects that match the scene, helping creators produce more complete videos without manually finding and layering audio tracks.
Wan 2.5 lets you describe dialogue, speaking style, music, ambience, and sound effects directly in the prompt, giving you more control over both how a scene looks and how it sounds.
Wan 2.5 follows detailed instructions for actions, camera movement, dialogue, sound, and composition more accurately, helping creators get closer to the intended result with fewer generations and prompt revisions.
Wan 2.5 animates images while better preserving character appearance, composition, lighting, and visual identity, making it easier to add natural motion without losing important details from the original image.
Wan 2.5 generates synchronized videos up to 1080P, 10 seconds, and 30fps, providing smoother, longer, higher-quality clips that are more practical for ads, social media, and character content.
| Capability | Wan 2.2 | Wan 2.5 | Wan 2.7 | Wan 3.0 |
|---|---|---|---|---|
| Model Positioning | Open video foundation | Audio-visual generation | Controlled cinematic generation | All-in-One multimodal creation |
| Text → Video | ✅ | ✅ | ✅ | ✅ |
| Image → Video | ✅ | ✅ | ✅ | ✅ |
| First Frame | ✅ | ✅ | ✅ | ✅ |
| First + Last Frame | ✅ Separate model | — | ✅ | ✅ |
| Reference → Video | Dedicated model / later capability | — | ✅ | ✅ Unified |
| Document → Video | ❌ | ❌ | ❌ | ✅ |
| Video Editing | Basic generation control | Improved prompt control | Strong reference & video control | Unified multimodal control |
| Native Audio | ❌ Core T2V / I2V are silent | ✅ | ✅ | ✅ Enabled by default |
| Video Duration | 5s | 10s | 15s | 30s |
| Resolution | 1080P | 1080P | 1080P | 1080P |
| Frame Rate | 30fps | 30fps | 30fps | 30fps |
| Key Innovation | MoE + cinematic quality | Native Audio | Multishot + R2V + control | All-in-One + Omni-Reference + 30s + Documents |
/w1440.webp)
With Wan 2.5 image-to-video generation, users can turn a single character image into a speaking or singing video with mouth movements synchronized to the audio. This makes it easier to create AI presenters, singing characters, virtual avatars, and dialogue scenes without using a separate lip-sync tool or manually matching speech in post-production.
Generate Wan 2.5 Video/w1440.webp)
Wan 2.5 performs particularly well when a scene has one clear subject, a defined composition, and one or two primary movements, helping faces, poses, and scene structure remain more stable from frame to frame. This makes it well suited for product shots, character close-ups, conversations, portrait animation, and short cinematic moments where consistency matters more than complex multi-scene action.
Animate with Wan 2.5/w1440.webp)
Wan 2.5 image-to-video generation is designed to preserve the subject’s appearance, framing, lighting, perspective, and overall composition while adding movement. This helps creators animate portraits, anime characters, mascots, products, or artwork without unnecessarily changing the visual identity of the original image, making the result more usable for branded and character-focused videos.
Create Video with Audio/w1440.webp)
Wan 2.5 can automatically generate voices, background music, environmental ambience, and sound effects based on what happens in the scene. A talking character can receive matching speech, while streets, weather, or other environments can receive contextual sound, reducing the need to search for separate audio assets and manually build a soundtrack after the video is generated.
Create a 1080p Video/w1440.webp)
Wan 2.5 supports video generation up to 10 seconds, 1080P, and 30fps, with synchronized audio available for both text-to-video and image-to-video workflows. The longer duration and high-resolution output give creators more room for complete actions, dialogue, product moments, and social media scenes without having to assemble several very short generations.
Choose Your FormatWan 2.5 combines flexible text and image generation with synchronized audio, stronger motion control, and high-resolution output for complete short-form video creation.
Create videos from a written prompt or animate a still image while controlling the subject, action, visual style, and camera movement you want.
Generate dialogue, ambient sound, music, and sound effects together with the visuals so key sounds follow what happens on screen more naturally.
Use voice, dialogue, songs, or music to guide lip movement and character performance, making talking, singing, and music-driven videos easier to create.
Choose a 5- or 10-second duration for quick concepts, character performances, product moments, ads, and social media content without stitching multiple clips together.
Generate videos in 480P, 720P, or up to 1080P depending on whether you are testing an idea or preparing a sharper final video for publishing.
Create square 1:1, vertical 9:16, widescreen 16:9, and other supported formats for TikTok, Reels, Shorts, ads, presentations, and social platforms.
Create a Wan 2.5 AI video on Lanta AI in three simple steps. Add your prompt or image, choose your video settings, then generate a synchronized short video ready to preview and use.
Write a prompt describing the subject, action, camera movement, mood, dialogue, or sound you want, or upload a source image to turn it into a video.
Choose Wan 2.5, then set your preferred 5- or 10-second duration, aspect ratio, resolution, and other available generation settings.
Click Generate and let Wan 2.5 create your video. Preview the completed result, then download it for social media, advertising, character content, or your next project.
Turn text, images, or audio into synchronized videos with natural motion, lip sync, and built-in sound. Start creating with Wan 2.5 on Lanta AI today.