REST API reference
Two API surfaces: the plugin controller (proxied behind Jellyfin auth at /Upscaler/*) and the AI service (direct, :5000).
Auth
- Plugin controller: Jellyfin's standard
X-Emby-Tokenaccess-token header. Elevated endpoints additionally require admin role (RequiresElevation). - AI service: shared-secret
X-Api-Tokenheader whenAPI_TOKENenv var is set on the container. Plain HTTP with no token if unset (homelab-friendly default).
๐ Securing the AI service (API token)
The AI service on :5000 does the GPU work and has no login of its own. If it's reachable beyond a fully trusted LAN, protect it with the X-Api-Token header. There are three distinct pieces - keep them apart and the rest is easy:
| Piece | Where you set it | Role |
|---|---|---|
API_TOKEN (env) | Docker container environment variable | Master / bootstrap secret. Read once at container startup. Always valid and cannot be revoked from the UI - so you can never lock yourself out. |
| AI Service API Token | Plugin โ Settings tab (field) | What the plugin itself sends on every call. Must equal a value the service accepts (the env token, or a managed one). |
| Managed tokens | Plugin โ API Tokens tab | Named, optionally-expiring tokens for other clients (scripts, integrations) that hit :5000 directly. Generated in the UI, stored hashed, individually revocable. |
A request is accepted when its X-Api-Token matches either the env API_TOKEN or any active (non-expired) managed token - both compared in constant time (hmac.compare_digest).
Turn auth on - 3 steps
- Generate a strong secret locally:
openssl rand -base64 32(โ256 bit). Don't reuse it anywhere else. - Set it as
API_TOKENon the container (compose / run below) and restart the container. - Paste the identical value into Plugin โ Settings โ AI Service API Token and save. โ ๏ธ If the two differ, every call returns
401/403- the #1 setup mistake. Then hit Test Connection: green means the tokens match.
API_TOKEN=disable and leave the plugin field empty - the service runs open (the homelab default). You can still create managed tokens, but they are not enforced until you switch disable off.Generating named tokens in Jellyfin (API Tokens tab)
Once auth is on, the plugin's API Tokens tab (Jellyfin admin only) mints extra tokens without ever touching the env var:
- Type a name (e.g. Living-room TV), pick an expiry (30 / 90 / 365 days, or never), click Create token.
- The value is shown once - copy it immediately. Only a SHA-256 hash is stored; it can never be displayed again.
- Revoke any token individually; the others and the env bootstrap stay valid. Expired tokens stop working automatically (no sweep needed).
API_TOKEN is the bootstrap that lets the plugin talk to the service; once that handshake works, the admin-gated API Tokens tab can issue the rest. Managed tokens live in a dedicated /app/config volume (not the cache) so they survive container recreates.docker-compose
services:
ai-upscaler:
image: kuscheltier/jellyfin-ai-upscaler:docker7 # or :docker7-intel / :docker7-amd / :docker7-cpu (CPU)
environment:
- API_TOKEN=change-me-to-a-long-random-string # or: disable
ports:
- "5000:5000"
volumes:
- upscaler-models:/app/models
- upscaler-config:/app/config # persists managed API tokens across recreates
restart: unless-stopped
docker run
docker run -d --name jellyfin-ai-upscaler \
-e API_TOKEN=change-me-to-a-long-random-string \
-p 5000:5000 \
kuscheltier/jellyfin-ai-upscaler:docker7
Confirm it's active
Call GET /status on the service - "api_token_configured": true and "auth_enabled": true mean the token is enforced. The Operator Console at :5000 shows the same on its API tab. Then click Test Connection in the plugin: a green result means the plugin's token matches the container's.
Plugin controller endpoints (/Upscaler/*)
| Method | Path | Purpose |
|---|---|---|
GET | /Upscaler/health/detailed | Proxies the AI service health check; returns status, circuit-breaker state and GPU posture. (There is no bare /Upscaler/health — it was documented for a while and never existed.) |
GET | /Upscaler/libraries | Enumerates Jellyfin virtual folders (ID, name, collectionType, paths). Used by the Select-Libraries chip picker. |
GET | /Upscaler/models | Model catalog proxied from service, merged with per-user preference marks. |
POST | /Upscaler/models/{id}/load | Warms the model in GPU memory via service. Admin only. |
POST | /Upscaler/models/{id}/download | Pulls ONNX weights. Admin only. |
GET | /Upscaler/jobs | Active jobs list. Admin only - previously leaked paths to any authenticated user, fixed in v1.6.1.11. |
POST | /Upscaler/upscale-frame | Base64-encoded frame in, upscaled frame out. Uses currently-loaded model. |
GET | /Upscaler/benchmark-frame?width=&height= | Single-frame benchmark pass; returns fps + ms. GET, not POST — documented wrongly until v1.8.3.27. Needs a loaded model, otherwise 400 with detail: "No model loaded". |
GET | /Upscaler/filter-preview/frame/{itemId} | Extracts a real frame from a Jellyfin library item, applies the selected preset, returns before/after PNGs. |
POST | /Upscaler/filter-config | Persists the 6 advanced filter params (gamma/sharpness/temp/vignette/grain/denoise). |
GET | /Upscaler/recommend-model | Content-based pick for one video: model, scale, reason, signals, and a filter suggestion (never applied automatically). Used by Auto-Mode. |
GET | /Upscaler/hardware-benchmark | Local hardware benchmark: CPU cores, GPU/provider detection, platform. Replaces /Upscaler/recommendations, which remains as a deprecated alias with an identical payload. |
GET | /Upscaler/tokens | List managed AI-service tokens (proxied). Admin only. |
POST | /Upscaler/tokens | Create a managed token (name, optional expiresDays). Returns the value once. Admin only. |
DELETE | /Upscaler/tokens/{id} | Revoke a managed token. Admin only. |
AI service endpoints (:5000)
http://<your-host>:5000/docs (Swagger UI), /redoc (ReDoc) or /openapi.json right in the browser. The Operator Console's API tab links the same three. The table below is a hand-kept overview; the live schema is the source of truth.System
| Method | Path | Purpose |
|---|---|---|
GET | /health | Basic liveness - returns 200 once the FastAPI app is up. |
GET | /status | Detailed: service version, provider, api_token_configured, auth_enabled. |
GET | /features | Feature capability matrix (face-restore, rife, tensorrt, etc). |
GET | /metrics | Prometheus-compatible - per-model latency, queue depth, VRAM, request counts. |
GET | /openapi.json | Full OpenAPI 3.1 schema. Feed it to any generator. |
GET | /logs-stream | Server-Sent Events stream of service logs. Powers the in-UI Console tab. |
GET | /auth/tokens | List managed tokens (hash + secret never returned; each has a derived expired flag). |
POST | /auth/tokens | Create a hashed token (name, optional expires_days); returns the plaintext once. |
DELETE | /auth/tokens/{id} | Revoke a managed token by id. |
Hardware
GET | /hardware | Providers with version strings, ONNX runtime version, build flags. |
GET | /gpus | Per-GPU: name, VRAM total/free, driver, compute capability. |
POST | /benchmark/run | Triggers a full benchmark sweep across all loaded models; async, poll /benchmark/status. |
Models
GET | /models | Full catalog with per-model state. |
GET | /models/{id} | Single model details - scale, input shape, provider compat, SHA. |
POST | /models/{id}/download | Fetch ONNX weights. Streams progress via /models/{id}/progress. |
POST | /models/{id}/load | Create inference session, preallocate tensors, optional TensorRT conversion. |
POST | /models/{id}/unload | Dispose session, free VRAM. |
Inference
POST | /upscale-frame | Single-frame inference. Body: { image, model, scale }. |
POST | /upscale-batch | Multi-frame batched inference; better throughput than per-frame. |
POST | /face-restore/detect | Returns face bounding boxes (Haar cascade). |
POST | /face-restore/restore | Applies GFPGAN or CodeFormer to a face crop. |
POST | /face-restore/pipeline | End-to-end: detect โ upscale โ feathered paste-back. |
POST | /models/load-detector | Loads an ONNX object detector alongside the upscaler. Form: model_name, input_size. |
POST | /detect-mask | Frame in, frame out, with detected objects covered. Query: classes, mode, confidence, pad. |
Object masking
Since v1.8.3.24 this runs during playback. The player already captured
frames for server-side upscaling; with masking enabled it sends them to
Upscaler/detect-mask instead and draws the covered frame back onto the overlay.
It replaces upscaling on that stream โ two full inference passes per frame do not
keep up with playback. Enable it on the Settings tab, or in the player menu under
Auto → Cover objects.
POST | /Upscaler/detect-mask | Plugin proxy used by the player. Masking parameters come from the plugin config, not the query string. |
POST | /Upscaler/object-mask/load-model | Loads the configured detector into the AI service. Admin only. |
Covers things you do not want on screen — the original request was to stop a dog
reacting to dogs and cats on TV. Detection runs here rather than in ffmpeg, because
jellyfin-ffmpeg is built without any DNN backend, so dnn_detect cannot run
there at all; and accepting arbitrary -vf would expose ffmpeg's file-reading
filters (movie=, subtitles=) to every authenticated user.
No detection model ships with the service. Every catalog entry carries a
verified sha256 pin, and inventing one for an unhashed model would break the guarantee the
importer exists to provide. Import your own — for example
tiny-yolov3-11.onnx from the ONNX model zoo — through the normal upload
path, which pins and verifies it.
Both detector families work and the service picks between them by reading the loaded model's own inputs and outputs, not by guessing: single-head exports (YOLOv5/v7/v8/v9, one tensor of raw anchors) and the NMS-head ONNX YOLOv3 exports (two inputs, three outputs, boxes already in source coordinates). A model it does not recognise is rejected at load time, because a misread tensor does not raise — it paints boxes over the wrong part of the picture.
TOKEN=your-api-token
BASE=http://localhost:5000
# 1. load a detector you imported earlier
curl -X POST "$BASE/models/load-detector" -H "X-Api-Token: $TOKEN" \
-F model_name=tiny-yolov3 -F input_size=416
# 2. cover every animal in a frame with a solid box
curl -X POST "$BASE/detect-mask?classes=animals&mode=box&pad=12" \
-H "X-Api-Token: $TOKEN" --data-binary @frame.jpg -o masked.jpg
# blur instead of a filled box, and only dogs
curl -X POST "$BASE/detect-mask?classes=dog&mode=blur" \
-H "X-Api-Token: $TOKEN" --data-binary @frame.jpg -o masked.jpg
The response carries X-Detections with the number of regions covered.
classes takes COCO names or the animals group;
pad grows each box, which matters because a detector's box hugs the animal and
the ears sticking out set a dog off just as well.
Example: load a model + upscale a frame
TOKEN=your-api-token
BASE=http://localhost:5000
# 1. Load model
curl -X POST -H "X-Api-Token: $TOKEN" $BASE/models/realesrgan-x4plus/load
# 2. Upscale a PNG (base64 in, base64 out)
IMG=$(base64 -w 0 input.png)
curl -X POST -H "X-Api-Token: $TOKEN" -H "Content-Type: application/json" \
-d "{\"image\":\"$IMG\",\"model\":\"realesrgan-x4plus\",\"scale\":4}" \
$BASE/upscale-frame | jq -r .image | base64 -d > output.png
Rate limiting & concurrency
The service has a single GPU serialiser - concurrent /upscale-frame requests queue behind one active inference. The MaxConcurrentStreams plugin config gates how many requests the plugin will have in-flight at once; tune this based on VRAM headroom.
ID validation regex
Model IDs and filter preset IDs are validated against ^[a-zA-Z0-9_-]+(?:\.[a-zA-Z0-9_-]+)*$ - dot-separated segments are allowed (for rife-v4.9 etc) but traversal patterns (.., leading/trailing dots) are blocked.