← Variant Studio

Short-form video, rendered on your own GPU

A photo, a ticker, and something to say. It writes the narration, speaks it, animates a presenter and composites a 9:16 clip.

Install

$ pip install video-gen-local
$ vgl install       # GPU environment + models, ~1 hour
$ vgl serve         # studio on localhost:7860
COMMAND WHAT LANDS ON YOUR MACHINE pip install video-gen-local the CLI the studio renderers text-to-speech captions vgl install the part pip cannot do measure hardware venv · 3.12 venv · 3.10 ~110 GB weights verify each stage vgl serve studio on localhost:7860 Nothing here runs in a cloud. The two Python versions are not a preference — the avatar model needs 3.10 and everything else needs 3.12.
Three commands, and the middle one does the heavy lifting. About an hour on a fresh machine, nearly all of it downloading.

That is the whole install — no clone, no checkout. The setup scripts ship inside the package, so vgl install runs from wherever pip put it.

Why two steps and not one

pip install gives you the CLI, the studio, the renderers, text-to-speech and captions. It cannot give you the avatar model, for two reasons that are not going away:

And the weights are about 110 GB of model downloads, which are not Python packages at all — pip was never going to fetch those. vgl install does both jobs, in the order that matters (torch before the requirements that would overwrite it, numpy re-pinned last).

Prefer to work from source? git clone then pip install -e . gives you the same three commands.

What you need

ComponentMinimumComfortable
GPU24 GB VRAM, CUDA 12.x80 GB (H100 / A100)
Disk150 GB250 GB
RAM32 GB64 GB
OSLinux, ffmpegPython 3.10 and 3.12

A GPU is required. There is no hosted fallback and no CPU path. If nvidia-smi does not print your card, nothing here works.

Runs the same on your own machine, on Nebius, or on AWS. The only real difference between providers is whether stopping the machine wipes its disk — on a VM cloud it does not, on RunPod or Vast it does, and vgl install simply re-provisions.

Keys

Every account is your own. There are no shared credentials in the source, and the studio refuses a job whose template needs a key you have not set rather than failing partway through a render.

VariableNeeded for
POLYGON_API_KEYPrices, company name, market cap
FAL_KEYThe fight card only — it generates its artwork
HEDGEFUND_API_KEYThe events feed — catalyst, showdown and recent-event read it
FMP_API_KEYearnings-call only — call transcripts. A plan that includes the transcript endpoints (they sit above FMP's Starter tier)
SEC_USER_AGENTrecent-event's quote. The SEC requires a declaring User-Agent with a contact address on EDGAR requests
$ cp .env.example .env # required keys are at the top, uncommented

Make one

Open the studio, pick talking-head, drop in a front-facing photo, type a ticker and a sentence of notes, press Generate. About fourteen minutes at 480p.

The studio has no authentication. Reach it over an SSH tunnel rather than exposing the port: ssh -fN -L 7860:localhost:7860 user@host

Templates

Five layouts. They differ in what they put on screen and what they need from you — not in quality.

A presenter on an orange field beneath a rotating card graphic.
catalyst A ~10s alert. Presenter on a solid field under a looping graphic. ticker · notes or events feed
Bull and bear presenters stacked on a black field either side of a divider.
showdown A three-hander debate. Narrator, bull and bear, each with their own face and voice. insights feed · three faces
A bull fighter in green rim light with a claim panel beneath.
showdown-no-lipsync A fight card. Generated art per round, claims spoken over it, no lip-sync. insights feed · two faces · FAL_KEY
not yet rendered
on this machine
recent-event A filing, explained. Headline card, presenter over a chart on a forest-green field, animated points, and a quote copied verbatim from the SEC filing with the executive who said it. Writes itself; a blank ticker covers the newest confident catalyst in the feed. nothing required · optional ticker · HEDGEFUND_API_KEY
not yet rendered
on this machine
earnings-call The latest earnings call, on a violet field. The narration and on-screen points are written from management's prepared remarks in one pass, and the quote is a sentence lifted from the transcript with its named speaker — never composed. A blank ticker covers the newest call on a US listing. nothing required · optional ticker · FMP_API_KEY
not yet rendered
on this machine
social-clip Presenter over a live stock chart, with an optional X post card above it. ticker · one face · optional post URL
not yet rendered
on this machine
talking-head Presenter only, over a flat background. The simplest thing that works. one face · notes

These are frames from real renders on this machine, not mockups. The two without one have simply never been run here — the studio shows every template regardless.

How it works

YOU SUPPLY FROM VARIANT, OR YOUR OWN RUNS ON YOUR GPU PRODUCES COMPOSITE Variant API key pasted in the browser ticker ADBE the research bull case and bear case, written for the ticker or your own notes your avatar the face image display name · channel tier band or upload any photo Polygon the one outbound call Qwen3-14B writes the narration Kokoro-82M speaks it LongCat-Video-Avatar the face, lip-synced to the voice, background out chart drawn locally narration.wav presenter chart.png name, channel and tier band become the top bar ffmpeg presenter over the chart Whisper captions burnt in top bar from your avatar music bed closing scene final.mp4 · 9:16 PUBLISH BACK — OPTIONAL upload to your CDN Cloudinary · S3 · Blob POST the public URL Variant fetches the file live on your channel title · badge · alerts
runs on your GPU Variant supply it yourself instead leaves your machine
Variant is a shortcut, not a requirement. Connect a key and it supplies the face, the branding and the research, then takes the finished video back. Skip it and you upload a photo, write your own notes, and keep the file. Everything between is a model on your own hardware either way.

Variant integration

If you publish to Variant, the studio can render as one of your avatars and post the finished video back. Your token is the identity — paste it once and everything else is read from it.

The key is the identity, so the server holds none. Each person supplies their own in the browser and sees only their own avatars — two people can share one studio without either being able to post as the other. Clearing the field disconnects.

ONE RENDER ONE PUBLIC URL EVERY DESTINATION THAT ACCEPTS ONE final.mp4 rendered once your CDN Cloudinary · S3 · Blob V Variant wired today · title and badge X post with the video attached YouTube Shorts, 9:16 already Instagram Reels Twitch clips
The render is the expensive part; the posting is not. Once the file sits at a public URL, every destination is one API call against the same asset — which is how one GPU-hour becomes five audiences instead of one. Variant is the one wired today; the rest are what that URL makes straightforward, not features that already exist.

Publishing needs the CDN step. Variant fetches the file from a public URL, so a render that was never uploaded has nothing to post.

If something breaks

SymptomFix
libcublas.so.12 missingTorch drifted to CUDA 13. vgl install venvs
Python.h missingapt-get install python3.12-dev
Render dies at the last segmentVRAM. vgl install detect re-measures and caps clip length
A template is greyed outIt needs a key you have not set. The message names it
Slower than yesterdayOn a container cloud you were rescheduled onto a different host

vgl doctor --deep checks every stage and tells you which one is unhappy.