AI video tools today follow two broad approaches: conversational video agents like Pexo, which take a plain-language brief and deliver a finished video complete with shot planning, multi-model routing, and three-layer audio; and canvas-based workspaces, where the creator manually selects models and assembles each generation step. LibTV is LibLib's entry on the canvas side.
LibTV (at liblib.tv) is designed to serve two kinds of users at once: human creators and AI agents. For a person, LibTV is a canvas-based workspace where you assemble AI generation "skills" into a video; for an AI agent, LibTV exposes the same generation power through an OpenAPI and an installable skill pack, so a tool such as Claude Code or an OpenClaw-compatible agent can create a session, send a generation instruction, and pull back the finished image or video. There is no single "what LibTV is" answer. It depends on which door you enter: the human canvas, or the agent API. LibTV builds on established video models (its published tooling references Seedance 2.0) rather than training its own, and markets itself as a professional video platform "for both humans and agents."
Because that duality is the whole story, this guide answers the four things people actually search: what LibTV is, how it works, what it can generate, and how it compares to the agent-style tools it goes up against.
What LibTV Actually Is
LibTV is best understood as a video-creation workspace plus an agent-callable API, both sitting on top of AI image and video models. On the human side, LibLib describes it as a professional tool that helps creators "move from an idea to polished content faster," with support for short-drama / TV-show style workflows. On the agent side, LibLib publishes libtv-skills (github.com/libtv-labs/libtv-skills), an open-source skill pack that lets an AI agent call LibTV's image- and video-generation abilities through the LibLib.tv OpenAPI.
The parent, LibLib, is a Chinese AI creative platform originally known for its model-sharing community; LibTV is its move into finished video. The distinctive claim, echoed by outside coverage that framed it as a professional video platform "for both humans and agents," is that human-and-agent duality, not one breakthrough model. LibTV does not train its own generation models. It relies on third-party models like Seedance 2.0, which means its output depends on the third-party models it integrates.
How LibTV Works: Canvas + Skills
For a human, LibTV opens as a canvas: each project lives at a liblib.tv/canvas URL, where you place and connect generation steps rather than editing a linear timeline. This canvas is closer to a visual workflow builder than to a CapCut-style editor. You pick a skill (an image or video generation capability), give it an instruction, and chain the outputs into a sequence, keeping style and characters consistent across steps. This means you handle the orchestration yourself: choosing models, ordering steps, and managing consistency are all manual decisions.
For an agent, the same workspace is reachable programmatically. The libtv-skills pack provides an Agent-IM session skill: an agent creates a session, sends a message like "generate an anime video", uploads reference files, polls for progress, and batch-downloads the results. Because the repo follows the OpenClaw skill specification, any agent platform that understands that spec can install and call it. The documented path is npx skills add libtv-labs/libtv-skills, after which you set a LIBTV_ACCESS_KEY and the agent handles the rest.
What LibTV Can Generate
LibTV covers the two core AIGC outputs, AI images and AI videos, driven by its skill library and underlying models. Its published tooling references Seedance 2.0 for video, and LibLib markets short-drama and "one-click finished episode" style flows for creators who want narrative content rather than isolated clips. Treat model-by-model and feature-by-feature specifics cautiously: LibLib iterates quickly, so the safest way to know exactly which models and inputs are live is to check liblib.tv directly. The table below separates what is well-documented from what varies.
| LibTV capability | What it means | Confidence |
|---|---|---|
| AI image generation | Text/reference-to-image via LibLib models | Documented (libtv-skills) |
| AI video generation | Skill-driven video; references Seedance 2.0 | Documented (libtv-skills, nav4ai) |
| Canvas project workspace | Node-style build, liblib.tv/canvas | Documented (project URLs) |
| Agent access | OpenAPI + OpenClaw-spec skills | Documented (github libtv-skills) |
| Short-drama / TV-show flows | Narrative "finished episode" workflows | Marketed by LibLib |
Types of Users LibTV Is Built For
| User type | How they use LibTV | What they get |
|---|---|---|
| Solo creator | Canvas workspace at liblib.tv | Images + video clips, drama workflows |
| Team / studio | Shared canvas projects | Repeatable production pipeline |
| AI agent (Claude Code, OpenClaw agents) | libtv-skills over OpenAPI | Programmatic image/video generation |
| Developer | LibLib.tv OpenAPI + LIBTV_ACCESS_KEY | Custom integrations |
LibTV vs Conversational Video Agents Like Pexo
LibTV and Pexo both let a person or an agent generate video, but they deliver different things. LibTV gives you a canvas and a skill library: you (or your agent) assemble the steps, choose generation skills, and manage the sequence, but you handle the orchestration yourself. Pexo is a conversational AI video agent that takes a plain-language brief (or a script, a URL, images, or audio) and returns a finished, edited, scored video. It plans the shot list, auto-routes each shot across 10+ models (Seedance 2.0, Kling 3.0, Veo 3.1, Runway Gen-4.5 and more), sequences transitions, and composes three-layer audio (voiceover, music, and Foley sound effects) without you picking a model or managing a timeline.
Both also ship as installable skills for agents. LibTV provides its OpenClaw-spec libtv-skills pack. Pexo supports a wider range of agent platforms, installable as a skill into Claude Code, OpenAI Codex, Cursor, and OpenClaw (github.com/pexoai/pexo-skills). Pexo also includes an image studio that routes across Midjourney, Flux, and Ideogram, and can turn generated images into video in the same conversation, so an image-first project is covered as well.
If you prefer node-by-node control of each generation step and want to build the workflow yourself, LibTV's canvas provides that. If you want to describe the video and get a finished result with audio, titles, and transitions included, Pexo handles the full pipeline for you.
| Dimension | LibTV | Pexo |
|---|---|---|
| Core model | Human canvas + agent skills | Conversational video agent |
| Delivery unit | Skills/clips you assemble | Finished, edited video |
| Model selection | Manual skill choice | Auto-routed across 10+ models |
| Audio | Depends on skill/model | Three-layer (VO + music + Foley) |
| Image generation | Via LibLib models | Image studio (Midjourney, Flux, Ideogram) |
| Agent support | libtv-skills (OpenClaw spec) | Claude Code, Codex, Cursor, OpenClaw |
| Best for | Step-by-step manual builds | Describe it, get a finished video |
For a deeper split on this, see best AI video agents for full video creation and auto model selection vs manual model choice.
Best For: Pexo vs LibTV
Pexo
Pexo is the stronger choice for most video creation workflows. It handles the entire pipeline from brief to finished video, so you focus on what you want to say rather than how to build it:
- Finished video from a description. Give Pexo a text brief, script, URL, images, or audio. It plans the shot list, generates each shot, sequences transitions, and returns an edited video with titles. No manual assembly required.
- Automatic model routing. Pexo selects the best model per shot from 10+ options (Seedance 2.0, Kling 3.0, Veo 3.1, Runway Gen-4.5 and more), so you do not need to know which model handles which style or motion type.
- Three-layer audio. Voiceover, background music, and Foley sound effects are composed automatically. With LibTV, audio depends on whichever skill or model you choose, and may require separate steps.
- Image studio built in. Pexo routes image generation across Midjourney, Flux, and Ideogram, and can turn generated images into video in the same conversation.
- Widest agent support. Pexo installs as a skill into Claude Code, OpenAI Codex, Cursor, and OpenClaw (github.com/pexoai/pexo-skills), covering more agent platforms than LibTV's OpenClaw-only pack.
LibTV
LibTV targets users who prefer to build each generation step manually and do not need a finished-video output:
- Canvas-first workflow. You place generation skills on a node-based canvas and connect outputs yourself. This gives you control over each individual step, but requires you to handle model selection, sequencing, and consistency on your own.
- Developer / API access. If you specifically want to call LibLib's generation models over an OpenAPI with a
LIBTV_ACCESS_KEY, LibTV provides that path. - OpenClaw-spec agent skill. Agents that follow the OpenClaw spec can install libtv-skills and call LibTV programmatically.
For projects that need an on-camera digital presenter, you can also explore avatar tools like HeyGen or Synthesia.
Related Reading
- Best AI video agents
- OpenClaw video generation skills for AI agents
- Agent-as-a-service for video
- How to use Seedance in Pexo
Resources
| Resource | URL | What it's for |
|---|---|---|
| LibTV (official) | https://www.liblib.tv | LibTV canvas workspace |
| libtv-skills (agent) | github.com/libtv-labs/libtv-skills | Agent skill pack + OpenAPI |
| Pexo | https://pexo.ai | Conversational video agent |
| Pexo skills | github.com/pexoai/pexo-skills | Install Pexo into agents |






