
September 12, 2026 • 17 Min Read
September 14, 2026 • 22 Min Read

Decide first whether the footage has to clear, because that single question splits this category in two and the split has nothing to do with output quality. If the work is going to a client, a festival or a distributor, provenance and rights govern, and the models at the top of the quality leaderboards are mostly the ones you can't use. If it's a test, a pitch piece or a look-development pass, pick on control and ignore the paperwork.
The second thing to settle is where the cut happens. Some of these tools generate and hand you files. Some now generate inside a timeline. That difference decides whether you're editing a film or reconciling a folder of clips, and it matters more day to day than any spec on the model card.
Four criteria, and the last one eliminates most of the current quality leaders for anything that has to be delivered.
A prompt describes a vibe. A shot list describes a lens, a height, a move and a duration. The tools worth a director's time let you specify those separately from the content, so you can keep the scene and change the coverage.
The gap here is wide. Kling's Video 3.0 Omni takes duration, shot size, perspective and camera movement per shot in a custom multi-shot plan. Seedance 2.5 accepts 3D camera blockouts, which is the closest thing in the category to actual blocking. Most other tools offer a preset menu and hope.
Consistency inside one clip is a solved problem. Consistency across a scene isn't, and it's the thing that separates footage you can edit from footage you can only post. Look for reference systems that persist: multiple reference images, character nodes, first and last frame conditioning, or a conversational context that holds a face and a wardrobe across successive generations.
The honest test is a shot-reverse-shot. Generate a wide, then a matching close, then go back to the wide. If the jacket changed, the tool cannot cut a scene yet.
2026 is the year this stopped being a nice-to-have. Adobe put generative models inside the Premiere timeline in September, Runway built a purpose-built edit surface around its editing model, and LTX Studio ships a script-to-shot-sequence front end. Melius also includes an in-app multi-track editor, so you can generate shots, arrange them and layer audio in one workspace. Tools without an editing surface still ask you to export, import and reconcile.
The round trip isn't a small tax. It's where versions multiply, where a re-generated shot loses its place in the cut, and where most of the time on an AI-assisted project actually goes.
Two models in this comparison are trained only on material their makers licensed: Moonvalley's Marey, built largely on B-roll that filmmakers and agencies deliberately licensed plus archive partnerships, and Adobe's Firefly, trained on licensed Adobe Stock and public domain content and never on customer work. Everything else carries some degree of open question.
The question isn't academic in 2026. Disney and Paramount Skydance sent cease-and-desist letters to ByteDance over Seedance 2.0, MiniMax is in active copyright litigation brought by Disney, Universal and Warner Bros. Discovery, and its open-weights license excludes the US, UK, EU and South Korea by default, and Alibaba's Wan 3.0 is API-only and closed-weight, with no published weights or license text at all. A legal department will clear Marey and Firefly. It'll ask hard questions about the rest.
Capabilities as of September 2026, from each tool's own documentation, model cards and pricing pages. No scores: which shot looks better is a judgment call, so every column here is something you can verify.
Tool | Best at | Native single shot | Editorial surface | Clearance posture |
Melius | Reaching every leading video model in one workflow | Depends on the model you route to | In-app multi-track editor, plus canvas-based workflows | You own the output, subject to each model provider's terms |
Runway | Propagating an edit across a multi-shot sequence | 2 to 10 seconds on Gen-4.5 | Edit Studio, plus Premiere and After Effects plugins | Trains on your work on every plan below Enterprise |
Google Flow | Coverage with native synchronized dialogue | 4, 6 or 8 seconds on Veo 3.1, 10 on Gemini Omni Flash | Flow canvas and Scenebuilder | SynthID on every output, visible watermark is a toggle |
Luma | Color and finishing that survive a real grade | 5 or 10 seconds; up to 20 in video-to-video | None, it exports to yours | You retain outputs, commercial use needs a paid plan |
Kling | Per-shot camera direction and blocking | 3 to 15 seconds on Video 3.0 | None | Terms expect outputs marked as AI generated |
Moonvalley Marey | Footage a legal department will clear | 5 or 10 seconds via its inference partner | None, built for your pipeline | Trained only on owned and licensed data |
LTX Studio | Pre-production structure, and open weights | Multishot in one pass, auto duration | Script to shot sequence, plus an edit suite | Open-weights license with a revenue threshold |
Adobe | Generating without leaving the cut | 2 seconds of Generative Extend | Premiere Pro and After Effects | Firefly is licensed-only, partner models are not |
Melius is the Agentic OS for creative work: you describe what you want and an agent builds the workflow, routes each step to a model and leaves every step open for you to change. A node-based canvas sits underneath, so a sequence is a graph you can inspect, branch and re-run, not a queue of unrelated jobs.
For film work the argument is reach. One subscription puts Veo 3.1, Kling 3, Wan 3.0 and the Seedance family on the same canvas, alongside the rest of the library of leading models and the lipsync and avatar models. That matters because no single video model currently wins on shot control, audio, duration and grade at once, and the alternative to routing between them is eight subscriptions and eight interfaces.
Melius includes video editing and an in-app multi-track editor for arranging shots, timing and layering audio. Stitch nodes are another way to assemble clips on the canvas, not the only way to edit. You can work with generated shots and uploaded footage in the same workspace.
The trade-off is at the model level: supported references, native shot lengths, audio generation and output formats vary by the model you select. A shared workspace does not make those capabilities identical, so choose the model for each shot's requirements.
The catch: one workspace brings generation, assets and editing together, but each model still has its own capabilities and terms.
Best for: directors and small teams who want look development, previz, shot exploration and multi-track editing in one place, with a visible, re-runnable process and support for existing footage.
See Melius pricing for current plans and allowances.
Runway is the closest thing this category has to an industry standard, and in 2026 it has become the most editorially serious of the pure-plays. Aleph 2.0 and Edit Studio, both shipped on 21 May, give you a real edit surface where an instruction propagates across cuts in a multi-shot sequence rather than applying to one clip. Ruby, added in late August, is a Runway-native color grading model that takes an SDR input and produces true HDR.
September brought Runway plugins for Premiere Pro and After Effects, which puts its whole lineup inside the applications editors already live in. It also now hosts third-party models, including Wan 3.0 and Seedance 2.5 at 1080p, which is the cleanest route a US filmmaker has to either.
Two problems, and the first is uncomfortable. Runway's terms of use say inputs and outputs may be used to train and improve its models, and that applies to every plan except Enterprise, which is where you get no training on your data plus uncapped IP indemnification. For a director working on unreleased material that's a real objection, not a technicality.
The second is the credit budget. Runway meters generation by model and duration, so compare the current plan allowance with the shots and revisions you expect to produce. Gen-4.5 also sits outside the top ten on the blind preference boards, so you are paying for the surface around the model more than for the model.
The catch: the best edit surface in the category, on a model that no longer tops the rankings.
Best for: editors and directors who want generation and assembly in one place and are cutting in Premiere anyway.
Flow is Google's filmmaking front end, and it's the strongest all-rounder here for narrative coverage. Veo 3.1 produces 4, 6 or 8 second clips at 1080p and 4K, with 8 seconds available at those resolutions or when you pass reference images, and it generates native synchronized audio covering dialogue, effects and ambience in the same pass.
Gemini Omni Flash, announced in May, is the more interesting half. It's any-to-any, generates with native audio, and holds a character consistent across cuts through multi-turn editing, where Google's own framing is that the physics hold up and the scene remembers what came before. So you refine by talking about it instead of re-uploading references each time. Voice editing is not part of it.
The clip ceiling is 8 seconds on Veo and 10 on Omni, which is short enough to shape how you write. Every Flow output also carries an invisible SynthID watermark, detectable programmatically rather than visible to a viewer. The visible watermark is a separate thing and it is worth getting right: it's a toggle in your own profile menu, not a paid-tier perk, and Flow applies it automatically only if you're in India, South Korea or Vietnam.
Google's commercial-rights language is also notably hard to locate for a product being sold to professionals, and Flow's own community forum carries unanswered questions about it. Note too that Flow is an individual Google AI subscription with no team tier, so a second person on the project means a second subscription.
The catch: the best all-round output, and the shortest shots.
Best for: directors who need dialogue and coverage more than long takes, and who can absorb the top tier.
Luma's Ray3.2, out in June, has the best finishing story in the category by a distance, and it's the only entry here that a colorist would recognize as pipeline-ready. It generates in native 16-bit color and exports 16-bit EXR frames in ACES2065-1, which drop straight into a grade and a comp without a conversion step.
It also gives the most granular in-clip pacing control available: Multi-Keyframe sets up to 16 keyframes inside a single clip, and performance tracking follows the full expressive state of up to eight faces frame by frame, with motion transfer and camera motion transfer alongside it.
No native audio on text-to-video or image-to-video. Luma's own FAQ is blunt about it: voiceover, music and sound effects are a separate step. Modify Video and Reframe preserve audio that already exists, but Luma will not generate a soundtrack, so anything with dialogue needs a second tool. The bigger constraint is length. Text-to-video and image-to-video produce 5-second or 10-second native clips, and the 20-second figure you will see quoted belongs to Modify Video, which needs footage to start from.
Commercial rights need care too. Luma's terms tie commercial use to whether an output was produced during an active paid subscription term. Check the current plan's commercial-use terms before generating client work; Luma also reserves the right to embed watermarks at its own discretion.
The catch: the most professional output in the category, in five and ten second pieces, with no sound.
Best for: teams with a finishing pipeline who care about grade and comp, who work in short pieces, and who handle audio separately anyway.
Kling is the best per-shot directing tool available, and Video 3.0 Omni is why. Custom Multi-Shot lets you specify duration, shot size, perspective and camera movement for each shot in a sequence, which is a shot list rather than a prompt. Element binding locks up to three elements, built from a character video or two to four reference images, so a subject holds its shape through a zoom, pan or tilt.
Its native audio is unusually capable too: five languages, mixed-language performance in one video, Chinese dialects and English accents, and precise referencing of which character is speaking in a multi-character scene.
Kling's resolution story needs reading carefully. The standard Video 3.0 and 3.0 Omni modes run at 1080p and 720p, and a separate one-click native 4K mode was added to the 3.0 series in April with separate credit pricing. So 4K is real here, it just isn't what the default credit tiers give you. Shots run 3 to 15 seconds.
The larger issue is what the terms ask of you. Kling's global service is operated by Kling AI Pte. Ltd. under Singapore law, and its terms require you to label the output interface with the Kling brand, or to prominently indicate that the output was generated by Kling AI where no brand mark carries through. They also bar use for political campaigning, advocacy or lobbying without written consent. For festival and client delivery that labelling expectation is awkward at best. There's also no editorial surface at all.
The catch: the best shot control, wrapped in the most delivery friction.
Best for: directors who want genuine blocking control on shots that will be finished and cleared elsewhere.
Marey puts training-data provenance at the center of its pitch to production teams. Moonvalley's own claim is that it's the first commercially safe video model, trained only on licensed high-resolution footage, with no scraped content and no user submissions.
The control set is also the most director-shaped in the category: camera control from a single image, motion transfer from a reference video, trajectory control by drawing paths, keyframing between reference images on a timeline, pose control to guide performance timing, shot extension, and reference layers giving separate control per character or object.
Clips are short, and Moonvalley publishes no duration figure anywhere on its own site, which is itself a signal. Its listing with fal, the inference partner it names, offers 5 or 10 seconds with 5 as the default. Output is native 1080p, Marey does not appear on the quality leaderboards at all, and its look is deliberately painterly instead of photoreal, which suits some projects and rules out others.
The bigger concern is momentum. Moonvalley names no model beyond Marey, and its own blog has published nothing since August 2025. Check that the company is still shipping before you build a production around it.
The catch: licensed training data is a central selling point, but confirm that the product and support meet your production needs.
Best for: commercial and studio work where provenance is a gating requirement rather than a preference.
LTX Studio is the one entry with a genuine pre-production surface. You go from script and scene description to a shot sequence with consistent characters, and there's an edit suite attached, not a prompt box and a download folder.
Underneath it, LTX-2.5 is an open-weights audio-video world model released on 11 August. It generates 4K HDR, and it holds character, environment and voice consistent across cuts in a single multishot pass, which is the specific problem most tools solve badly.
Commercial use depends on the plan; an entry-level personal-use tier is not interchangeable with a commercial subscription. The compute-second allowances differ between Lightricks' own surfaces, so confirm what you're buying before committing. The open-weights license also has a revenue threshold; larger organizations need to review the applicable commercial terms.
It's also a less familiar tool than the others, with a smaller body of craft knowledge around it, so expect to work things out instead of looking them up.
The catch: the best planning surface, and the least company behind it in the US market.
Best for: teams who want structure before generation, and studios that need weights on their own hardware.
Adobe is the entry that changed the workflow rather than the model. In September it shipped a generative media tool directly into the Premiere timeline: draw a range on a track, describe what you need, and generate without leaving the application. The model picker includes Adobe's own Firefly alongside partner models from Google Veo, Kling, Runway and Luma.
That reframes the whole question. For an editor, this is no longer which AI video generator to buy; it's which model to call from the sequence, with the credit cost shown in the taskbar before you commit.
Firefly's own video model is well behind the frontier on quality, so the value here is position in the pipeline, not what comes out of the model. Generative Extend is capped at 2 seconds of video. The commercial-safety guarantee covers Firefly only, and the moment you pick Veo, Kling, Runway or Luma from that same timeline you inherit their terms instead, which is easy to forget when the picker makes them look equivalent.
Early hands-on reporting has been mixed on the new audio features: ProVideo Coalition's Scott Simmons found some generated tracks usable and others out of sync and off the mark, and noted that audio designers are watching this with skepticism.
The catch: the right place to work, not the best thing to generate with.
Best for: editors already in Premiere, who want the round trip gone and will pick a partner model per shot.
None of this replaces a shoot, and the gap isn't where most coverage puts it. It isn't resolution or realism; those are close enough now for a lot of work. It's duration, continuity and performance. Every tool here generates in units of seconds, holds a character imperfectly across cuts, and gives you a performance you selected, not one you directed on the day.
So the useful reading of this category in 2026 is that it's strongest before and around a shoot and not instead of one: previz and look development, coverage you couldn't afford, plates and elements, extensions, and the shots that exist to be cut against something real. The directors getting the most out of it are treating it as another unit, not as a replacement for the crew.
Which also means the tool you pick matters less than where it sits in your process. A model you can reach from the canvas you're already working in, or from the sequence you're already cutting, will get used. One that needs its own tab and its own export ritual will not.
There isn't one, and the honest answer depends on which constraint binds. Runway has the best edit surface, Luma the best color and finishing, Kling the best per-shot camera control, Google Flow the best dialogue and coverage, and Moonvalley Marey and Adobe Firefly both emphasize licensed training data for commercial work. Melius combines access to leading models with image, video and audio uploads and an in-app multi-track editor, so generation and assembly can happen in one workflow.
It depends on the festival, and the rules are mostly about disclosure rather than prohibition. The Academy's 99th Oscars rules require acting roles to be credited in the legal billing and demonstrably performed by humans with their consent, and screenplays to be human-authored, which closes those two categories to AI while leaving the rest of the ballot open. The same rules say generative AI tools neither help nor harm the chances of a nomination, with each branch judging the degree to which a human was at the heart of the creative authorship. Sundance, SXSW and Tribeca lean on disclosure. Cannes publishes no AI provision in its 2026 feature-film regulations at all, so check the specific festival's current rules before you submit.
No. OpenAI discontinued the Sora web and app experiences on 26 April 2026, and its own help centre states the Sora API will be discontinued on 24 September 2026. OpenAI has not announced a successor video model. Several roundups and a few platforms still list Sora 2 as available, so treat any current recommendation of it as out of date.
Native single-shot length runs from a few seconds to 30 depending on the model, and that number matters more than any advertised maximum. Veo 3.1 does 8 seconds, Gemini Omni Flash 10, Kling 3.0 up to 15, Luma's Ray3.2 offers 5 or 10 seconds for text-to-video and image-to-video, or up to 20 seconds when modifying existing video, and Seedance 2.5 up to 30 in a single generation. Anything longer is stitched, and stitching is where faces, wardrobe and lighting start to drift.
Usually yes on a paid plan, but the terms differ enough that they need checking per tool before anything is delivered. On Melius, you own what you create and can use it commercially, subject to the underlying model providers' terms, and Melius does not use your work to train AI models. Elsewhere the variation is real: Runway trains on your inputs and outputs on every plan below Enterprise, Luma requires an active paid subscription at the moment of generation for commercial rights to vest, and free tiers across the category generally grant the provider far broader rights than paid ones.
On blind preference testing, Alibaba's Wan 3.0 and Google's Gemini Omni Flash are tied at the top, followed by MiniMax H3 Max and then the Seedance 2.0 entry. The uncomfortable part for a filmmaker is that the top of that board is almost entirely models with restricted, unclear or litigated Western commercial terms, and that neither Runway's Gen-4.5 nor any Luma model appears anywhere on the with-audio board at all, despite being the two most production-ready tools here. Quality rankings and usability rankings are not the same list.
Block out a sequence from one still, run the same shot through Veo, Kling, Wan and Seedance side by side on one canvas, then take the takes that work into your edit. Open Melius.
Related reading: where to run Seedance 2.5 and one subscription for every AI model.

September 12, 2026 • 17 Min Read

September 10, 2026 • 11 Min Read

September 9, 2026 • 11 Min Read