Pexo
Pexo/Blog/AI Video News & Trends/How to Make Videos With Claude Code Using Opus 5: A Step-by-Step Guide

How to Make Videos With Claude Code Using Opus 5: A Step-by-Step Guide

Liora Adler avatarLiora Adler
·Last updated Jul 25, 2026
How to Make Videos With Claude Code Using Opus 5: A Step-by-Step Guide
Summary

A step-by-step guide to making videos with Claude Code on Claude Opus 5 (released July 24, 2026, the new default on Claude Max) using the Pexo skill: install the skill, add your API key, and describe the video in plain language. Opus 5 reads the request and orchestrates; Pexo auto-routes each shot across Seedance 2.0, Kling 3.0, Veo 3.1, and Sora 2 and returns a finished, scored MP4 — no editing, no model selection. Covers what you need, the install and prompt steps, why the model version does not change the workflow, troubleshooting, and an 11-question FAQ.

To make videos with Claude Code using Claude Opus 5, you install the Pexo skill, add your API key, and describe the video you want in plain language — the agent then generates a finished result with no editing software, no prompt engineering, and no model selection. Pexo provides a skill you install into Claude Code; Opus 5 is the reasoning model that reads your request, calls the skill, and monitors the job, while Pexo auto-routes each shot across models like Seedance 2.0, Kling 3.0, Veo 3.1, and Sora 2 and returns a finished, multi-shot MP4. Claude Opus 5 shipped on July 24, 2026 as the new default model on Claude Max and the strongest model on Claude Pro, but the important point for this workflow is that the mechanics do not depend on the model version: install the skill, describe the video, review, and export works the same on Opus 5 as it did on Opus 4.8 or Fable 5. This guide walks the full path end to end.

What Claude Opus 5 Changes (and What It Does Not)

Claude Opus 5 is Anthropic's newest Opus-generation model, released July 24, 2026 with a 1,000,000-token context window and API pricing of $5 per million input tokens and $25 per million output tokens. It becomes the default model on Claude Max and the strongest option on Claude Pro, and it adds a toggle that lets you trade cost against capability per task. For video, a stronger orchestrating model means the agent parses a messy, multi-part brief more reliably, keeps a long shot list coherent, and reasons about revisions better. What it does not change is the actual generation: Opus 5 does not itself render pixels. The video comes from the Pexo skill, which routes to dedicated video models. So the workflow is model-agnostic — you get the same finished video whether Claude Code is running Opus 5, Opus 4.8, or Fable 5.

FactClaude Opus 5
ReleasedJuly 24, 2026
Default onClaude Max (strongest on Claude Pro)
Context window1,000,000 tokens (default and max)
API pricing$5 / M input, $25 / M output tokens
PositioningComes close to Fable 5 at roughly half the price; Fable 5 remains the frontier flagship
Role in this workflowOrchestrates the request and calls the skill — it does not generate the video itself

What You Need

Two things, and you can be running in a few minutes:

RequirementWhat it is
Claude Code on Opus 5Anthropic's terminal coding agent, with Claude Opus 5 selected (the default on Claude Max). The same steps work on any current Claude model, and in Codex, Cursor, and OpenClaw.
The Pexo skill + an API keyPexo provides a skill you install into Claude Code; it adds the video-generation capability the model does not have on its own. Sign in at pexo.ai and copy your API key.

Claude Code does not generate video on its own, and neither does Opus 5 — the skill is what adds the capability. Claude Code is the most stable host for it because Pexo ships a native SKILL.md there. Once the skill is installed and your key is set, everything else happens inside the conversation. For the different ways an agent can make video, see can Claude Code make videos.

Step 1: Install the Pexo Skill and Add Your API Key

Pexo provides a skill you install into Claude Code — it is not built into the model. Sign in at pexo.ai, copy your API key, install the skill from the open-source repo, then confirm the agent can see it:

# Inside Claude Code (running Opus 5), list installed skills
> /skills
# You should see "pexo" in the list

The Pexo skill ships its helper scripts plus a SKILL.md that tells the agent how to behave — including a rule to pass your request through faithfully rather than rewriting it. Set your API key as the skill instructs (an environment variable), then run the included diagnostic to confirm the config, dependencies, and key are all valid before you start. The skills repo is github.com/pexoai/pexo-skills. Never assume Claude Code or Opus 5 has Pexo "built in" — you install it once, and the same package also runs in Codex, Cursor, and OpenClaw.

Step 2: Describe the Video You Want

This is the whole interface: tell Opus 5 what you want in plain English. You do not pick a model or write a technical prompt. Useful things to specify:

  • What it's about — "a product video for these wireless headphones"
  • Length and shots — "15 seconds, three shots"
  • Mood — "cinematic and premium," "fast-paced for TikTok"
  • Music — "ambient electronic," or let the agent choose
  • Aspect ratio — "9:16 for Reels," "16:9 for YouTube"

A complete first request looks like this:

> Make a 15-second cinematic product video for these wireless headphones —
  three shots, a slow orbit on the first, premium feel, with ambient music. 9:16.

That is enough. Opus 5 reads the brief, decomposes it into a plan, and hands the job to the Pexo skill.

Step 3: Let the Agent Generate

Once you send the request, Opus 5 dispatches it to Pexo, which runs the full production: it writes a shot script, auto-selects the best model for each shot from 10+ options (a product close-up might route to Kling 3.0, a motion scene to Seedance 2.0, a cinematic wide to Veo 3.1), generates each shot, adds transitions, composes an original score, and mixes a three-layer soundtrack of voiceover, music, and Foley sound effects.

A 15-second, three-shot video takes roughly 8–10 minutes end to end — far faster than choosing models, writing per-model prompts, and assembling clips by hand. While it runs, Opus 5 polls Pexo for progress and reports back; you do not need to do anything until the finished file returns. The larger 1M-token context of Opus 5 helps here because the agent can hold the full brief, the shot list, and the progress log in one conversation without losing track.

Step 4: Review and Iterate

When the video comes back, you review it and ask for changes the same way you described it — in plain language. There is no timeline to edit:

> Make the second shot slower, and swap the music for something more upbeat.

Opus 5 passes the revision through the skill and returns an updated cut. Because the request stays conversational, you iterate by talking, not by re-rendering anything yourself. A stronger reasoning model tends to interpret vague revision notes ("make it feel more energetic") more usefully, but the mechanics are identical on any model.

Step 5: Export and Use It

The finished video comes back as a standard MP4, mastered and ready to post. Because the production is multi-format aware, you can ask for the same content in different aspect ratios — 9:16 for TikTok and Reels, 16:9 for YouTube, 1:1 for feed — without regenerating from scratch. Download it and publish.

Tips for Better Results

A few habits produce noticeably better videos from the same skill, regardless of which Claude model is running:

  • Describe the mood, not the model. "Premium and cinematic" guides Pexo's routing better than naming a model — the routing layer translates intent into the right model for each shot.
  • Name the platform. Saying "for TikTok" or "for a YouTube pre-roll" sets the aspect ratio, length, and pacing conventions automatically.
  • Give 2–4 reference images for products. When accuracy matters — a specific product, logo, or packaging — hand over a few clear photos at 1080p or higher.
  • Be explicit about shot count and rhythm. "Three shots, quick cuts" versus "one slow continuous move" produces very different edits.
  • Generate variants in the same conversation. Opus 5's long context keeps the brand look consistent across "now a punchier 9-second version" requests in one session.

Other Ways to Start: Image, URL, Script, Audio

Text is only one input. Pexo accepts five, so Opus 5 can start from whatever you already have:

InputHow you startExample
TextDescribe the video"a cinematic ad for my coffee brand"
ImageHand over product photosturn studio shots into a moving product video — see the image-to-video guide
URLPaste a product pagethe agent extracts images, copy, and price into an ad
ScriptProvide a written scriptthe agent segments it into scenes
AudioSupply a track or voiceoverthe agent generates visuals to match

Troubleshooting

SymptomLikely causeFix
Agent says it can't make videoPexo skill not installed, or model has no video capabilityInstall the skill; run /skills to confirm "pexo" appears
"Invalid API key" or auth errorKey not set or mistypedRe-copy the key from pexo.ai and set the environment variable the skill expects
Agent rewrites your briefSkill config not loadedConfirm SKILL.md is present; the diagnostic verifies the config is valid
Job seems stuckGeneration still runningA 15s/3-shot video takes 8–10 min; let the agent keep polling
Wrong aspect ratioRatio not specifiedState "9:16" or "16:9" in the request; ask for multiple ratios at once

Which "Model" Does the Work?

There are two model layers here, and it helps to keep them straight. Claude Opus 5 is the reasoning model inside Claude Code — it reads your brief, calls the Pexo skill, and manages the conversation. The video models — Seedance 2.0, Kling 3.0, Veo 3.1, Sora 2, Runway Gen-4.5, and more — are what Pexo routes to per shot to actually render footage. You never pick either one manually. This is why the tutorial stays valid across Claude releases: upgrading the reasoning model from Opus 4.8 to Opus 5 (or to Fable 5, which Anthropic still recommends for the most advanced autonomous work) improves orchestration, while the video quality comes from Pexo's routing layer, which reshuffles its own model roster every few weeks.

Resources

ResourceURLWhat it is
Pexopexo.aiThe video skill used in this guide
Pexo Skills (GitHub)github.com/pexoai/pexo-skillsOpen-source skills for coding agents
Best video skills for Claude Codepexo.ai/blogThe full ranking of options

Frequently Asked Questions (FAQ)

How do I make a video with Claude Code using Opus 5?

Install the Pexo skill into Claude Code, add your API key, then describe the video you want in plain language while Claude Opus 5 is the selected model. Opus 5 reads the request and calls the skill; Pexo writes a shot script, auto-selects video models, generates the shots, adds transitions and a three-layer soundtrack, and returns a finished MP4 — typically in 8–10 minutes for a 15-second, three-shot video. You do not pick a model or edit a timeline; you describe the result and review it.

Does Claude Opus 5 have Pexo built in?

No. Pexo provides a skill you install into Claude Code — it is not built into Opus 5 or into Claude Code itself. Opus 5 is the reasoning model that orchestrates the request; the Pexo skill adds the actual video-generation capability. You install it once from github.com/pexoai/pexo-skills and set your API key. The same skill also runs in OpenAI Codex, Cursor, and OpenClaw, with Claude Code being the most stable host because Pexo ships a native SKILL.md there.

What is Claude Opus 5 and when was it released?

Claude Opus 5 is Anthropic's newest Opus-generation model, released July 24, 2026. It became the default model on Claude Max and the strongest option on Claude Pro, ships with a 1,000,000-token context window, and is priced at $5 per million input tokens and $25 per million output tokens. Anthropic says it comes close to Fable 5's capability at roughly half the price, while Fable 5 remains the frontier flagship recommended for the most advanced autonomous projects.

Do I need Opus 5 specifically to make videos in Claude Code?

No. The workflow is model-agnostic. Because the video is generated by the Pexo skill rather than by the reasoning model, you get the same finished result on Opus 5, Opus 4.8, or Fable 5. A stronger model like Opus 5 parses complex briefs and revisions more reliably and holds a long shot list in its 1M-token context, but the install-describe-review-export steps are identical on any current Claude model.

Which models does Pexo use to generate the video?

Pexo auto-selects per shot from 10+ video models including Seedance 2.0, Kling 3.0, Veo 3.1, Sora 2, Runway Gen-4.5, and MiniMax/Hailuo — routing a product close-up to one model and a cinematic wide to another. You never name a model; the routing layer picks the best fit for each shot. This is separate from Claude Opus 5, which is the reasoning model doing the orchestration, not the video rendering.

Do I need to know how to edit video or write prompts?

No. You describe the outcome in plain English — mood, length, shots, music — and the agent handles model selection, prompting, and assembly internally. There is no timeline editor and no per-model prompt syntax to learn. You iterate by asking for changes in words, like "make the first shot slower," and Opus 5 passes the revision through the Pexo skill.

How long does it take to make a video this way?

A 15-second, three-shot video with auto model selection and a mixed three-layer score takes about 8–10 minutes end to end. A single raw clip from one model returns faster (1–3 minutes) but is unassembled. The time scales with length, shot count, and whether you want a finished cut versus a single clip — it does not depend on which Claude model is orchestrating.

Can I make a video from product photos or a URL instead of text?

Yes. Pexo accepts five input types: text, image, product URL, script, and audio. You can hand over product photos, paste a Shopify or Amazon URL (the agent extracts images and copy), provide a script, or supply an audio track. For a photo-based walkthrough, see the image-to-video guide.

Does this work in Codex, Cursor, and OpenClaw too?

Yes. Because Pexo ships as an installable skill built on the open Agent Skills standard, the same package runs in Claude Code, OpenAI Codex, Cursor, and OpenClaw. The install location differs slightly per agent, but the workflow — install the skill, describe, generate, review, export — is identical. Claude Code is the most stable host because Pexo provides a native SKILL.md there.

How much does it cost to make videos with Claude Code and Opus 5?

The skill itself is free to install; generation runs on Pexo credits, and new accounts include a starting allowance to try it. Separately, Claude Opus 5 usage runs on your Claude plan (it is the default on Claude Max) or on API pricing of $5/$25 per million input/output tokens. So your total cost has two parts: your Claude plan for the reasoning model, and Pexo credits for the video generation.

Can I make vertical videos for TikTok, Reels, or Shorts?

Yes. When you describe the video, specify the aspect ratio — 9:16 for TikTok, Instagram Reels, and YouTube Shorts, 16:9 for standard YouTube, or 1:1 for feed posts — and Pexo exports in that format. You can ask for the same video in several ratios at once, so one request can produce both a vertical social cut and a widescreen version without regenerating from scratch.

Pexo Recommend