How Clipper works, and how to run it
The detailed guide: every feature, the command line, running it on a NAS with PM2, and how the code is laid out. For an overview, go back to the README; for running it for other people, see ADMIN.md.
Run on the NAS (venv + PM2)
# keep pip's cache off the system partition
export PIP_CACHE_DIR=/path/on/data/volume/pip-cache
./setup.sh # creates ./venv, installs pinned requirements
./venv/bin/clipper create-user YOUR_NAME # prompts for a password (min 8 chars)
# data lives outside the code folder, ideally on the data volume
export CLIPPER_ROOT=/path/on/data/volume/clipper-data
export CLIPPER_INBOX_DIR=/path/on/data/volume/clipper-inbox
pm2 start ecosystem.config.js && pm2 save
Then open http://<nas-address>:5050. PM2 runs two processes:
clipper-web (the UI) and clipper-worker (runs queued stages, one at a time).
Logs: pm2 logs clipper-worker. Restarting the worker is safe: an interrupted item goes
back in the queue.
Requires ffmpeg/ffprobe on the PATH. Settings are environment variables; see .env.example.
Generic mode: from a talking video to suggested clips
- ingest: copies the file in and reads its length, size, and audio.
- transcribe: speech-to-text with word timestamps (faster-whisper, run as a separate
process so a crash can't take the worker down). Models download once into
<CLIPPER_ROOT>/models. - analyze: sends the timestamped transcript to a language model, which proposes up to 8 self-contained 20-60 second clips. Clipper never trusts those times: each pick is checked (inside the video, right length, no overlaps) and snapped to real word boundaries, so clips don't start mid-word. One retry with feedback if the answer is unusable.
Tick "Continue automatically" on the new-job page (or clipper new FILE --auto) to run all three.
The job page shows the transcript, the suggested clips, and a Preview button for each.
Needs an API key for the model that picks the clips. OpenAI is the default; Gemini and Anthropic also work. Keep keys in a file only you can read (outside the code folder and git):
Paste this one line on its own; with more lines in the same paste, read would swallow the next one.
It refuses to save anything that does not look like a key:
read -rsp "OpenAI API key: " K; echo; if [ "${#K}" -ge 20 ] && [ "${K:0:3}" = "sk-" ]; then printf 'OPENAI_API_KEY=%s\n' "$K" >> "$CLIPPER_ROOT/clipper.env" && echo "saved (${#K} characters)"; else echo "that does not look like an OpenAI key; nothing saved"; fi; unset K
Then, as a separate paste:
chmod 600 "$CLIPPER_ROOT/clipper.env" && pm2 restart clipper-worker
To use another provider, add CLIPPER_LLM_PROVIDER=gemini (or anthropic) and its key
(GEMINI_API_KEY / ANTHROPIC_API_KEY) to the same file. CLIPPER_MODEL=... overrides the default
model (gpt-6-astra, gemini-3.8-flash, claude-sonnet-5-5). Only the transcript text is sent.
Video with no speech (silent screen recordings, music only) stops at analyze with a clear message.
- Music mode: the file must be a video (a still album cover with the track on it is fine). Transcription is
tuned for singing, and your typed lyrics are lined up with what is heard, so captions use your exact words; the
job page shows how much matched. language: in the details helps when a song opens with a long instrumental.
Clips are then chosen from the timed lyrics (whole lines, 15-35 s, hook_hints always honoured; instrumentals use
the hints and the loudest stretches) and rendered like any other video. The end card points at the song's stream link.
Every rendered clip keeps a moment after its last word (CLIPPER_TAIL, CLIPPER_MUSIC_TAIL; never into the next
word) and fades out over it, so a held note is not cut off.
- Post text: approve clips, then "Write post text" (an extra step, never automatic) writes a caption and hashtags for each
approved clip that has none, keeping your own hashtags first. Each clip has a box to edit it and a Copy button.
Text you typed is never overwritten.
- Number of clips: a detail on each job (New job form, or Edit details), 1 to 20, applied the next time clips are chosen. Blank uses
CLIPPER_MAX_CLIPS (talk, default 8) or CLIPPER_MUSIC_MAX_CLIPS (songs, default 4). 20 is a hard limit on everything.
- Names and look: on screen, jobs are Projects, New job is Create clips, Queue is Activity and generic mode is Spoken word, under a
forest-green menu on the left (across the top on a phone). Stored names, links and URLs are unchanged. The palette's contrast is checked by
tests/test_shell.py.
- Batch actions: "Approve/Reject all N undecided" (never touches decided clips), and tick-boxes on the clips with Approve, Reject, Undecide,
Set text and Clear text for the ticked ones (shortcuts to tick all, approved, undecided, none).
- Find more clips: a form under the clips asks for up to N further clips (never past 20 in all) from parts the current ones do not
cover. Existing clips, approvals, text and renders are untouched; new ones go at the end, so numbers do not change.
- Bundles: "Download all approved clips" (and a Bundle link per clip) gives a zip with a folder per clip: its video, post.txt,
details.txt and cover.jpg, plus contents.txt. Only clips whose render is up to date; the rest are listed as left out.
- Unknown lyrics: tick "I don't know the lyrics: listen for them" and the words are taken from what is heard in the track (one line
per phrase, marked as heard, not typed). Edit details pre-fills the lyrics box with them to correct; saving re-times at once.
- Renders: each render has its own timestamped name (clip-03-20261006-111742.mp4), downloads carry it, and the page lists a
clip's earlier renders. The newest CLIPPER_KEEP_RENDERS (default 5) per clip are kept; older ones are deleted.
- On the video: one choice per video for the caption band: lyrics/captions, hashtags, both (hashtags fade in where no caption is
showing, the default for songs), or nothing. The hashtags are those in the clip's post text, else your own; at most five. The
preview shows them too, and a change marks renders outdated.
Each approved clip can also carry a line of your own text (50 characters), shown above the hashtags.
- Edit details: the Details card has an Edit link (the same form as New job, filled in). Saving never re-runs anything and says
what to do next. For a song, edited lyrics are re-timed at once from the saved recogniser output (no new transcription); clips
keep their times and captions you have not edited are refreshed. Language and instrumental changes need transcribe again.
Using it
- New job: pick a file from the inbox folder (drop files there over SMB), upload one, or give a URL. The details form is generated from the mode. Music mode will not create a job until artist, title, rights, a link, and lyrics (or instrumental) are filled in.
- If a
.yamlfile sits beside an inbox file and the details form is left empty, it is used. - From a URL: YouTube and most video sites, via yt-dlp used as a library (
pip install -r requirements.txt). Downloads one video at up to 1080p and keeps its title for the clip-picking prompt. YouTube needs node 20+ on the PATH (orCLIPPER_JS_RUNTIME).clipper check-url [VIDEO_URL]shows what works on this machine. When downloads start failing, update withpip install -U "yt-dlp[default]"and restart the worker. If YouTube refuses a file (error 403) Clipper retries other ways automatically;clipper check-url LINK --testshows which work here. - Job page: status, source info, progress while a stage runs, and buttons for the stages that can run next.
CLI equivalents: clipper template music, clipper new FILE --mode music,
clipper enqueue JOB ingest, clipper queue, clipper worker, clipper status,
clipper import-fs (import old job.json folders).
Layout
src/clipper/
models.py Job, SourceInfo, Segment, QueueItem (pure data, no I/O)
config.py Settings from env vars (the only machine-specific part)
store/ JobStore + JobQueue protocols; sqlite (default) and fs backends
modes/ Mode protocol + generic, music (metadata rules per use case)
stages/ Stage protocol + ingest, transcribe, analyze (cut/render come next)
transcript.py transcript model, phrase grouping, word-boundary snapping
asr.py speech-to-text child process and its parent-side runner
llm.py OpenAI / Gemini / Anthropic client (stdlib only, injectable)
pipeline.py create_job, run_stage (no CLI/web dependency)
worker.py claim queued stages and run them (embeddable)
users.py web logins (same SQLite file)
web/ Flask app: app.py routes, forms.py (schema -> form), templates/, static/
cli.py argparse client
Reuse and spin-offs
- Other apps:
from clipper.pipeline import create_jobandfrom clipper.context import Context. Nothing in the core imports the CLI or Flask. - New use case: write a
Mode(pydantic schema + template) andregister_mode()it. The web form for it is generated automatically (modes with unsupported field shapes get a YAML box instead). - New step: write a
Stage(name, accepts, produces, run) andregister_stage()it. Callctx.report(percent, "note")inside it to drive the progress bar. - Another database: implement
JobStoreandJobQueue,register_backend("name", factory), setCLIPPER_STORE. SQLite is the reference implementation.
Notes
- The web server is Flask's built-in threaded server: fine for a handful of people. Accounts are by invitation and failed logins are slowed down (see ADMIN.md). Before exposing it to the internet, put HTTPS in front of it (a reverse proxy or a Cloudflare tunnel); for heavier use, run it under a production server such as gunicorn.
- Moving to another machine: copy the folder and the data directory, edit the environment
variables, run
./setup.sh. - Docker files are kept as an alternative (
docker compose up -d worker) but the NAS setup above is the primary path. - Tests:
pip install -e ".[dev]" && pytest.