How to Edit Videos With Claude: Three Real Routes (and When Each Wins)
Yes, you can edit video with Claude — three real routes: MCP servers that drive your NLE, Claude Code rendering, and conversational editors, with honest costs.

The question "can Claude edit videos?" has an annoying answer: yes, but Anthropic ships nothing called a video editor. There is no timeline in Claude, no export button, no "edit my footage" feature. Anthropic has never marketed video editing as a Claude capability — its closest product, the Cowork desktop agent released in January 2026, is a general-purpose file-and-workspace assistant, not a creative tool.
Everything people mean when they say "editing video with Claude" is something you assemble around the model. That assembling has gotten dramatically easier in 2026, which is why the search results are suddenly full of it. But the results quietly mix three very different setups that suit three very different people, and picking the wrong one wastes a weekend.
This guide maps the three routes as they actually exist in late September 2026: what you build, what each is genuinely good at, what it costs you in setup and maintenance, and which one fits your footage and your skills. If what you actually want is generating a video from nothing — motion graphics, explainers, animated stats — that's a related but different question covered in our Claude video generator explainer; this article is about getting a finished edit out of footage you already shot.
The Honest One-Paragraph Answer
Claude can drive real video editing, but only as an orchestrator. It plans the edit, writes the instructions or code, and operates other software — your NLE through an MCP server, a rendering framework like Remotion or HyperFrames, or a hosted editing agent. What it never does is open your footage natively and cut it inside a Claude-branded timeline. So the real decision is not "Claude yes or no." It's which of the three plumbing styles below matches how you work, because they differ more than the marketing suggests.
Route 1: Claude Drives Your NLE Through MCP
The most literal version of "Claude edits my videos" is the Model Context Protocol route. MCP is the open standard, released by Anthropic in late 2024, that lets Claude connect to external tools the way a plugin connects to a DAW. Install an MCP server for your editor, point Claude Desktop or Claude Code at it, and the model can read your project, move clips, set markers, and trigger renders — while the actual cutting still happens inside real editing software.
The ecosystem as of late September 2026:
| Server | Stars (Sep 2026) | What Claude gets control of |
|---|---|---|
| samuelgursky/davinci-resolve-mcp | ~3.2k | DaVinci Resolve timelines, bins, renders, via Resolve's Python API |
| hetpatel-11/Adobe_Premiere_Pro_MCP | ~629 | Premiere Pro projects: sequences, clips, exports |
| KyaniteLabs/kinocut | ~177 | Guardrailed FFmpeg and HyperFrames operations, repurposing tools |
Around these sit a long tail of small community servers built on MoviePy and FFmpeg that expose splice, trim, caption, and transcode operations to any MCP client.
What this route is genuinely great at: batch operations on real footage inside professional tools. If you have sixty talking-head clips that all need the same rough pass — silence trims, marker placement at topic changes, timeline assembly from a shot list — Claude executing an NLE's scripting API does this with a precision that browser tools can't match, and the result stays a normal Resolve or Premiere project your team can open. You also keep everything: the tools are free or already-owned, the pipeline is inspectable, and nothing uploads to someone else's cloud. For developers and technical teams, this is the control-maximal option, and no honest comparison should pretend otherwise.
The costs are equally real. You are now a systems integrator. Each server has its own setup ritual — Resolve scripting requires enabling external scripting and, for full API access, a Studio license; Premiere's servers need the right extension host running; FFmpeg-based servers need the binary on your path. When an NLE update changes its API, your pipeline breaks until the server maintainer catches up. And the model can only be as careful as the server's guardrails allow: an MCP session that can move clips can also mis-move them, so versioned project files stop being optional. If the phrase "debug a JSON-RPC handshake" doesn't sound like a reasonable Saturday, route 1 will fight you.
Route 2: Claude Code Renders the Video as Code
The second route abandons the timeline metaphor entirely. You open Claude Code, describe a video, and the model writes a program that renders one — typically Remotion, the React-based video framework, or HyperFrames, HeyGen's open-source framework whose tagline is the whole idea in six words: "Write HTML. Render video. Built for agents." In effect this is vibe coding applied to video — the same you-direct, it-assembles shift traced in the rise of vibe video editing — except the thing being assembled is a renderable scene graph rather than an edit.
This is the route behind most of the viral "Claude made this video" posts you have seen. The model writes component code, the framework renders deterministic frames, and a prompt-to-MP4 loop emerges: change a line, re-render, diff the output. Add a voiceover track from a TTS service like ElevenLabs and you have a full explainer pipeline with no camera involved. Our companion article on making videos with Claude walks through why this exploded after the Opus 5.5 release and what the outputs realistically look like.
The critical thing to understand — and where most "Claude edits my video" hype goes wrong — is that this route does not edit footage. Remotion and HyperFrames compose graphics: animated text, charts, UI mockups, scenes built from code. There is no meaningful way to hand this pipeline forty minutes of shaky phone video from your trip and get a paced vlog back. If your source material is a camera roll, route 2 is the wrong tool no matter how impressive its demo reel is. If your source material is an idea — a product explainer, a data story, a title sequence, a kinetic typography short — it is hard to beat, and the iteration speed is addictive.
Cost-wise, expect to be comfortable reading code even if you don't write it. The frameworks are free and open source; the bills are your time on the prompt-render-fix loop and compute for rendering longer pieces.
Route 3: A Conversational Editor That Already Works
The third route is what most people asking the question actually want: upload footage, describe the edit, get a cut — no assembly. Hosted conversational editors package the same agent pattern behind a product. You upload clips, type something like "cut the dead air, tighten the pacing, add captions, make three Shorts from the best moments," review the draft, and refine it in follow-up messages.
Sparki, the product we build, works this way — upload raw footage, describe the edit in plain language, review and revise across rounds, export. We'll be explicit about our bias: we're one option in this route, not a neutral judge of it. The honest framing of the category's trade-off: you give up the deep control of route 1 (you can't rewire the pipeline when it does something you don't like) and the creative rawness of route 2 (you're working within a product's editing vocabulary, not writing your own). What you get back is the part routes 1 and 2 make you build yourself: the model, the tools, the file handling, and the render farm already wired together, with zero setup between "I have footage" and "there's a first cut."
The pattern itself is going mainstream independent of us: at its Made on YouTube event on September 23, 2026, YouTube announced a conversational editing tool that lets creators make edits in natural language — a strong signal that "describe the edit, review the result" is becoming a default interface rather than a novelty. The broader argument for why briefing beats timeline-operation lives in our piece on chat-to-edit commentary editing, and the paradigm-level view of how agentic editors compare to timeline, text-driven, and template tools is in our AI video editor paradigms report.
Which Route Fits You
The deciding factors are your footage and your tolerance for plumbing, in that order.
| Your situation | Best fit | Why |
|---|---|---|
| Real footage, batch or recurring edits, you write code | Route 1 | Reusable pipeline, NLE-grade output, full control |
| Real footage, one-off edits, no coding appetite | Route 3 | Setup time would exceed editing time otherwise |
| No footage — explainers, motion graphics, data stories | Route 2 | Code rendering is the native medium for graphics |
| Mixed: footage plus animated segments | Routes 1 or 3 for footage + Route 2 for graphics | No single route covers both well yet |
| Team with an existing Resolve/Premiere workflow | Route 1 | Meets the work where it already lives |
Two footnotes to that table. First, the boundary that matters most: which decisions you hand to Claude at all. The mechanical first pass — trims, captions, silence removal, draft assembly — hands over well; pacing taste and story judgment come back worse if you stop reviewing them, a boundary we draw carefully in what AI video editors can and can't automate. Second, route 1 and route 3 converge over time — kinocut already mixes MCP with repurposing presets — so re-check the landscape before assuming this table is still current.
Your First Session on Each Route
Route 1, a realistic first run: install the MCP server for your NLE (Resolve's needs external scripting enabled in preferences), connect it in Claude Desktop, and start with read-only tasks — "list this project's bins," "find clips longer than 30 seconds," "add markers at each silence over two seconds." Only let it mutate the project once you've copied the project file. Your first win should be a batch marker or rough-assembly pass, not a finished video.
Route 2, a realistic first run: scaffold a Remotion project (or clone HyperFrames' examples), and ask Claude Code for a ten-second animated title card — one composition, two text elements, one easing change. Re-render, adjust the prompt, repeat. Scale up to a thirty-second explainer only after you've felt the prompt-render loop at small scale, and read the license terms for any TTS voice and music you pipe in.
Route 3, a realistic first run: upload one real clip — not your best one — and describe an edit you could verify: "remove the pauses, keep the part about pricing, add burned-in captions." Judge the draft on whether the cuts land where you'd have cut, then push one refinement round. That single loop tells you more about whether conversational editing fits you than any feature list. You can run it with Sparki, whose long-to-short and copy-style features cover the two most common follow-up needs — cutting Shorts from a long source and matching the structure of a reference edit.
https://sparki.io/blog/how-to-edit-videos-with-claude
