Generative Footage

Kling AI Motion Brush and Scene Control Features

Clarify which Kling motion feature solves your creative problem.

Staff Writer · · 11 min read
Cover illustration for “Kling AI Motion Brush and Scene Control Features”
Model Comparisons · September 20, 2026 · 11 min read · 2,501 words

Kling AI runs four separate systems under the same rough label of "motion control," and creators mix them up constantly. That mix-up costs real money: wasted credits, generations that miss the brief entirely, retries that add up fast. This piece breaks down what each system actually does, where the lines between them sit, and which one to reach for depending on the shot in front of you.

The core confusion: four systems that all get called "motion control"

Kuaishou launched Kling AI in June 2024. By early 2026 it had 60 million creators and 30,000 enterprise clients running through it, and the revenue curve backs up that adoption story: $100 million in annual recurring revenue about 10 months in (March 2025), then $240 million ARR by December 2025, at the 19-month mark. That's a fast climb for a generative video product, and the version roadmap moved just as fast: 1.6 in December 2024, 2.0 in April 2025, 2.1 that May, and 3.0 arriving with cinematic features like multi-shot storyboarding and 4K at 60fps.

Under the hood, Kling runs on a diffusion-based transformer (DiT) paired with a self-developed 3D VAE. That VAE compresses video across both space and time, which is what makes precise motion control workable instead of a guessing game.

Four separate systems live inside Kling, and they get talked about as if they're one thing: Motion Brush, Motion Control, the Camera Movement panel, and Multi-Shot, plus prose camera direction written straight into a text prompt.

Lay them side by side and the differences get clear fast. Motion Brush moves painted regions of an image. Motion Control moves a character's body and face. The Camera Movement panel moves the camera. Multi-Shot moves the camera across a whole sequence of shots. Each one takes different inputs too: Motion Brush wants brush strokes plus a trajectory, Motion Control wants a character image and a reference clip, the camera panel wants a slider value, Multi-Shot wants a written shot list. Element caps differ as well. Motion Brush allows up to 6 motion elements, Motion Control locks to one primary subject, and Multi-Shot caps out at 6 shots.

Some newsletters have claimed Motion Brush got renamed to Motion Control. Some newsletters have claimed Motion Brush got renamed to Motion Control, but it did not. Both are live, both are documented separately on kling.ai, and they do genuinely different jobs. Picking the wrong one doesn't just waste a generation, it burns credits on output that never had a shot at matching the brief. The next three sections walk through each system in the order a creator would actually run into them, starting simple and building toward full scene control.

Motion Brush: how region-based animation works

Think of Motion Brush as a mask sitting on top of Kling's image-to-video pipeline. Most image-to-video tools animate the entire frame at once. Motion Brush lets a creator paint specific regions and decide what moves inside them, while everything outside the brushed area stays locked in place.

Kling's documentation lays out two ways to select an area. Automatic selection uses smart segmentation to grab an object's outline and generate a motion path fast. Manual brushing gives finer control, with a brush size that adjusts up to 50 pixels, useful for anything from a broad sweep to a single strand of hair. Up to 6 elements can carry independent trajectories in one generation, so a scene with a bird, a flag, and a candle flame can have three different motions running at the same time.

The feature that gets underused here is Static Brush. It locks pixels in place, and its main job is stopping camera drift that a motion trajectory can trigger elsewhere in frame without anyone asking for it. Kling's own tip: drop a static brush along the bottom of the image to kill that drift before it starts. Static Brush can cover more than one area at once.

Both the direction and the length of the curve drawn on screen change the output. The direction and length of the curve guide where the element moves within the frame. The model treats the trajectory as a direct instruction rather than a loose suggestion.

Prompts have to match the brush work. Kling's documentation calls for "element + motion" phrasing, with "a puppy running on the road" as the sample. A brush stroke with no matching prompt language, or a prompt that contradicts the stroke, produces movement nobody can predict. One more precision tip straight from the docs: brush only the part of the object that should move, an animal's head instead of its whole body, for tighter results.

Draw a trajectory on something that can't physically move, a building, say, and the model reads that line as a camera movement path instead. Static Brush is the fix. One brushing rule rounds this out: one motion brush should cover a single element rather than several unrelated objects.

What Motion Brush can't do matters just as much as what it can. It doesn't pull motion from a reference clip, it doesn't touch the camera directly, and it has nothing to do with multi-shot sequencing. Those jobs live in separate systems.

Motion Control: applying movement from a reference video to a character image

Motion Control works on a different premise. Feeding it a character image and a reference clip causes Kling to read the motion in that clip, map it onto the proportions of the character, and generate a new video of that character performing the same movement.

One hard limit defines the whole feature: it's image-to-video only. Text-to-video prompts don't support Motion Control, so a character image is non-negotiable, every time. A second limit sits on the subject side. Only one primary subject gets tracked. If a frame has multiple people in it, the model picks whoever takes up the most screen space; if two people occupy roughly equal portions of the frame, the model picks no one.

In practice this covers walking cycles, dance sequences, hand gestures, head turns, and synced facial expressions, all lifted from the reference footage rather than typed into a prompt. Reference clips run 3 to 30 seconds, and the feature works on both Kling 2.6 and 3.0.

Prompt strategy flips here compared to Motion Brush. Best practice holds that the prompt should describe the setting and context, not the movement. The reference clip already handles the choreography. Orientation choice matters too, with different modes suited to different shot types and output lengths.

Lifting motion from one clip and applying it to a completely different subject, with no comparable native tool on rival platforms, made dance-transfer and performance-replication content a natural showcase for the feature. Camera replication from a reference clip exists elsewhere. Transferring motion at the subject level, onto a specific character, is something Kling does that the rest of the current top tier doesn't match.

Motion Control handles one character in one shot. Once a project needs several shots stitched into a single sequence, Kling 3.0's scene tools take over.

Kling 2.6 vs. 3.0 for Motion Control: when each version is the right call

Version confusion is one of the most common complaints developer forums have raised since Kling 3.0 launched in May 2026, and the confusion is fair: both versions run Motion Control, but the gap between them changes what's actually achievable.

Character reference support is the biggest jump. Kling 2.6 accepts a single reference image. The 3.0 web interface accepts up to 7, though the API still processes one image per request no matter which version is running. Output length grows too, from 10 seconds max on 2.6 to 15 seconds on 3.0. Audio tells its own story: 2.6 has no audio sync at all, while 3.0 adds native lip-sync across five languages, Chinese, English, Japanese, Korean, and Spanish. Face consistency is handled differently at the architecture level too: 2.6 anchors identity to a single frame, 3.0 builds a multi-frame anchor instead.

That multi-frame anchor solves a real problem. On 2.6, a reference clip with heavy head rotation tends to break face consistency, because the model only has one angle of the subject to work from. Give it multiple angles, as 3.0 does, and it builds a fuller identity map that holds up even when the motion gets complicated.

The decision isn't close once you lay it out. Reach for 2.6 with a single frontal-facing character, a generation under 10 seconds, and no audio requirement: it's capable and it costs less, so there's no reason to pay for 3.0's overhead. Reach for 3.0 the moment the reference involves face turns, complicated body motion, audio sync, or a clip longer than 10 seconds, where the multi-reference handling earns the extra spend. That's where the multi-reference handling actually earns the extra spend.

On pricing, pay-as-you-go rates for 3.0 run $0.071 per second on Standard and $0.095 per second on Professional, no minimum spend required. 2.6 Pro runs around $0.07 per second, with 3.0 climbing to roughly $0.20 per second on the higher end. Those numbers only tell half the story, though. A creator hitting a 70% first-take success rate is effectively paying a much higher real cost per usable clip, since every failed take still burns money. Sharper reference images and tighter image-to-video setup cut that regeneration tax down more than any other single fix. One more practical note from community testing: even though the documented range for reference clips is 3 to 30 seconds, the sweet spot for actual Motion Control quality is 2 to 5 seconds.

Kling 3.0 scene-level features: Multi-Shot, camera control, and native audio

Multi-Shot, sometimes called AI Director mode, is the headline addition in Kling's VIDEO 3.0 release. It runs in two modes. Automatic lets the model plan its own shot transitions. Custom Multi-Shot hands that control to the creator, shot by shot. One gate matters here: the Multi-Shot toggle has to be switched on before Custom Multi-Shot even becomes available. If the toggle is left off, the output comes back as a single shot no matter how the prompt is written.

A generation can chain up to 6 connected shots, each running 3 to 15 seconds, with the full sequence capped at 15 seconds total. Inside Custom mode, each shot gets its own settings: duration down to the second, framing (wide, medium, close-up), camera movement (static, dolly, orbit, tracking), narrative content like dialogue or action beats, and its own character references. The whole plan executes in a single generation pass, not six separate ones stitched together afterward.

Separate from Multi-Shot is the 6-axis Camera Movement panel, a UI tool rather than a prompt-based one. It covers six basic moves, pan, tilt, zoom, roll, track, and pedestal, plus four preset Master Shots: move left and zoom in, move right and zoom in, move forward and zoom up, move down and zoom out. Each axis takes an intensity value between -10 and 10. Kling's camera control guide adds prompt vocabulary that works alongside the panel and recommends sticking to one main camera move per shot instead of stacking several on top of each other. The guide also draws a clear line between motion vocabulary and framing vocabulary, keeping terms like ultra-wide, bokeh, telephoto, and low-angle in a separate category from movement terms.

Native 4K output arrived April 27, 2026, with the 4K API going live a few days earlier on April 23. The resolution is confirmed at true 4K, generated natively rather than pushed through an upscale pass after the fact.

Audio got a real upgrade too. VIDEO 3.0 produces native multilingual audio and lip-sync across five languages, English, Chinese, Japanese, Korean, and Spanish, with American, British, and Indian accent variants available. It supports multi-person dialogue with lip-sync and voice attribution per speaker. A Multi-Dialogue feature lets a creator assign specific lines to specific characters in the prompt, and the model matches each line to the right face, solving the speech-confusion problem that occurs constantly in busy, multi-person frames. Text rendering also received attention in the VIDEO 3.0 release.

Character consistency across shots: Elements 3.0 and facial anchoring

Faces changing between frames is one of the most-reported complaints from developers and creators since Kling 3.0 shipped. Elements 3.0 is the system built to fix it.

It works by taking a 3 to 8 second reference video and locking the character's appearance and voice from it. Voice binding keeps visual identity and audio identity moving together across different scenes, so a character doesn't just look the same from shot to shot, they sound the same too. Feeding the system multiple angles improves how well it understands the character in three dimensions, and reference images can be combined with a voice clip so both traits lock in together.

VIDEO 3.0's Motion Control also picked up a facial consistency upgrade on its own, holding stable features and smooth expressions through complex, multi-angle, long-duration motion. That's a meaningful shift: it pushes Motion Control from a novelty trick into something closer to actual cinematic performance capture.

The multi-reference approach is doing the real work here. Giving the model several angles of a face lets it build a richer identity map, one that resolves most of the head-rotation face-breaking problems that showed up constantly on 2.6. One catch for anyone building automated pipelines: even on 3.0, the API still accepts only one character image per request. Multi-reference support exists in the web interface only, so teams running things programmatically need to plan around that gap. For anyone building a multi-shot story, the strongest setup right now combines Elements 3.0 reference videos with Custom Multi-Shot storyboarding, keeping one locked character identity running across all 6 shots in a single generation pass.

Choosing the right system for the job: a decision map for common use cases

Reach for Motion Brush when the starting point is a static image and only certain regions of it, not the whole frame, need to move. It's also the right call when the background has to stay completely locked (Static Brush handles that directly), and when a scene needs several independent elements moving in different directions at once, up to 6 of them. Product shots, nature scenes with subtle movement, and portraits with one selectively animated detail all fall into this category.

Reach for Motion Control when there's a specific performance, a walk, a dance, a gesture, that needs to transfer onto a character who never actually performed it. Reach for the Camera Movement panel when the shot is simple and the only thing that needs direction is the camera itself.

Reach for Multi-Shot, paired with Elements 3.0, when the project is a sequence: several shots, one consistent character, dialogue that has to land on the right face every time. The mistake creators keep making is defaulting to whichever tool sounds most powerful rather than matching the tool to the actual job in front of them. Matching the tool to the actual job, rather than defaulting to whichever sounds most powerful, is what separates a clean first-take generation from a stack of wasted credits.

Sources

  1. Kling AI Motion Control Guide: 2.6 vs 3.0, Brush & Free Access
  2. How to Animate Image Parts | Kling AI
  3. kling.ai
  4. kling.ai
  5. kling.ai
  6. kling.ai
  7. Kling Camera Control and Motion Control: What's Documented (2026)
  8. kling.ai

More in Model Comparisons