Two channels, chosen by one rule: if losing a message is a problem, it rides TCP; if only the freshest sample matters, it rides UDP.
| Worker | Receives | Default ports |
|---|---|---|
| Script | session + events (TCP) | TCP 9000 · web viewer 9001 |
| Lighting | session + events (TCP) and zone/pose telemetry (UDP = TCP port + 1) | TCP 9100 · UDP 9101 · web viewer 9102 |
In the app's Script / Lighting settings you enter host:port plus optional
?args. The query string is forwarded verbatim to the worker in the session — it is the
single config surface for both the worker and the app.
192.168.1.30:9000?key=YOURKEY
192.168.1.30:9100?key=YOURKEY&pose=both&baseline=0.064
| Arg | Read by | Meaning |
|---|---|---|
| key | worker | Shared secret. Demo workers reject sessions (and web viewers) without the right key. |
| pose | app | off (default, zero pose work) · 1 single-eye 2D poses · both both eyes, so the worker can triangulate 3D depth from disparity. |
| video | worker | Added by the app, not typed in the URL: the playing video's file name without extension, percent-encoded (decode it!). Only present when the user enables Send video name in that worker's settings — never assume it; fall back to sidecars. Stem only, never a path. |
| anything else | worker | Forwarded untouched — calibration, mode flags, whatever your worker wants (e.g. baseline, proj). |
Everything is little-endian and tightly packed (no alignment padding, no type tags). Reference codecs — reader and writer — ship in Python, Java, and Swift and are verified byte-identical.
| Type | Size | Encoding |
|---|---|---|
| byte / bool | 1 | unsigned · 0/1 |
| int16 / int32 / int64 | 2 / 4 / 8 | two's-complement LE |
| float / double | 4 / 8 | IEEE-754 LE |
| short ASCII | 1 + N | u8 length (N < 128), then N bytes |
| long ASCII | 4 + N | int32 length, then N bytes |
| short Unicode | 1 + 2N | u8 length in UTF-16 code units (N < 128), then UTF-16LE |
| long Unicode | 4 + 2N | int32 length in code units, then UTF-16LE |
longASCII args // endpoint URL query (+ "&video=<stem>" if the user opted in), e.g. "key=YOURKEY&pose=both"
int32 fileCount // sidecar files: names starting with the video's stem + "."
-- per file --
shortASCII ext // everything after "stem." — "srt", "json", or qualified: "roll.funscript"
int32 contentLength
bytes content
int32 action // see table below
double position // seconds into the video
double length // total duration, seconds
| action | Name | Sent when |
|---|---|---|
| 0 | start | a video finishes loading |
| 1 | play | playback resumes |
| 2 | pause | user pauses, or the player closes |
| 3 | skip | ±15 s skip buttons |
| 4 | scrub | scrubber drag released (final position) |
| 5 | end | clip reaches its end. The player then loops: it seeks back to 0 and keeps playing, without a further event — so treat end as "restarting", not "stopped" |
| 6 | heartbeat | about once a second by default, also while paused. Carries the true playhead, so use it to correct drift in your own clock — and as a watchdog. Change the cadence with ?hb=<seconds> in the endpoint URL, clamped 0.25–30 (a logger that wants quiet asks for ?hb=30). |
| 7 | scrubBegin | scrubber drag starts (position = playhead at drag start) |
| 8 | scrubMove | ~10 Hz live position while the drag is in progress |
Workers must ignore (but still reply to) action codes they don't recognize — new ones may be added.
int32 code // your status / command code (403 = unauthorized in the demos)
int32 dataLength
bytes data // free-form payload back to the app
The reply channel is how a worker talks back. Replies are never matched up with what the app sent, so a worker may push a frame at any moment — which is what makes the control commands below possible on the same connection, with no listening port on the headset.
Send the same frame with one of these codes. The player performs the change and emits the matching playback event, which comes back to every worker — so you learn the result the same way you learn everything else, and never have to track a playhead the player didn't confirm.
Reply codes below 100 are yours; 100 and up are reserved for commands. Demo workers ack an event by echoing its action code, which is safe because action ids stay under 100 — that is a promise, not a coincidence. Don't ack with a code ≥ 100 or the player will read your acknowledgement as an instruction.
| code | Command | Payload |
|---|---|---|
| 100 | play | — |
| 101 | pause | — |
| 102 | play/pause toggle | — |
| 103 | skip | double seconds, signed (−10 back, +10 forward) |
| 104 | seek | double absolute position in seconds |
start and play mean playing, pause and end
mean paused. Asking for a state the player is already in is harmless: it re-sends the event,
which is how a worker that has drifted gets corrected.The demo script_handler.py shows the whole loop — its web page has
play/pause and ±10 s buttons, and the state it displays is driven purely by the events coming back.
int32 seq // drop stale/reordered datagrams
double timestamp
u8 cols, rows // zone grid, currently 8×4
-- per zone (row-major, cols×rows) --
u8 r, g, b // average color, LINEAR RGB (convert for your lights)
u8 motion // color change vs previous sample
u8 coverage // fraction of zone NOT keyed out as green (0–255)
-- models --
int32 modelCount // detected people
-- per model --
u8 eye // 0 = left/primary, 1 = right (pose=both only)
u8 jointCount
u8 id · int16 x · int16 y · int16 confidence // coords/conf ×10000, per-eye normalized, bottom-left origin
u8 r, g, b // frame color AT the landmark (gamma RGB)
u8 regionCount
u8 id · int16 x · int16 y · u8 r, g, b // calculated clothing sample points
| 0–4 | nose · leftEye · rightEye · leftEar · rightEar |
| 5 | neck |
| 6–11 | shoulders · elbows · wrists (L, R) |
| 12–17 | hips · knees · ankles (L, R) |
| 18 | root |
| 0 | hat |
| 1 | hair |
| 2 | shirt |
| 3 | pants |
| 4 · 5 | left shoe · right shoe |
?pose=both the same person arrives once per eye.
Pair models across eyes (same vertical position), then depth ∝ horizontal disparity
xL − xR. Pass your rig's calibration through the args
(e.g. baseline=0.064&proj=equirect) to get metric depth.
Two reference workers, with the Python and Java codecs, are packaged as a download. Each worker opens its data port(s) plus a live auto-refreshing web viewer.
Download reference workers (38 KB) Python 3.9+, no third-party packages.
unzip vrvideoplayer-workers.zip && cd vrvideoplayer-workers
python3 script_handler.py # TCP 9000 · viewer http://localhost:9001/?key=YOURKEY
python3 lighting_handler.py # TCP 9100 · UDP 9101 · viewer http://localhost:9102/?key=YOURKEY
Enter the endpoints in the app (use your machine's LAN IP, not localhost):
Script 192.168.1.30:9000?key=YOURKEY
Lighting 192.168.1.30:9100?key=YOURKEY&pose=both
The viewers show the connection, forwarded args, received files with previews, the live event log — and for lighting, the colored zone grid and detected poses with clothing/skin swatches. They're the fastest way to confirm your own worker's parsing against known-good output.
key= gate is a casual-access check, not security — traffic is
plaintext. Set your own value in APPKEY at the top of each handler script (it ships as changeme) and use
the same value in the app's endpoint URL. For a real deployment add TLS and your own token.The app deliberately stays "dumb" — it grabs and streams; workers interpret. That makes the worker side a playground. Some ideas, roughly ordered by ambition.
Parse a .srt sidecar, follow position from events + heartbeats, render captions or lyrics on a TV, tablet, or projector near the viewer.
Relay play/pause/scrub between several headsets watching the same file — everyone's worker mirrors the leader's timeline.
A .json cue sidecar maps timestamps to fans, scent diffusers, heat lamps, or misters — the worker fires GPIO/smart-plug cues as the timeline crosses each mark.
Timeline-cued rumble and impact tracks (bHaptics, buttkicker) from a haptics sidecar, tightly synced by the heartbeat position.
Log every start/pause/skip/scrub with positions: watch-through rates, drop-off points, most-rewatched moments — per file, across sessions.
Accessibility: play recorded descriptions (a sidecar audio file per scene) on external speakers at the right timestamps, ducking during dialogue.
For live venues: the sidecar is a full show file — the worker drives DMX fixtures, motors and practical props in lockstep with the video, and pause/scrub rehearses any cue instantly.
Use the reply channel: the worker inspects a choices sidecar and replies with commands — the seed of interactive, choose-your-path VR films.
The classic: map the 8×4 zone grid to Hue / WLED / LIFX around the room — left zones drive left lights — so the room glows with the scene. Colors arrive linear; convert to your lights' space.
Per-zone motion spikes on action — drive intensity chases and strobes from it for concert or music-video footage.
The What To Wear headline act: shirt / pants / shoes / hat colors per person, live — match against a product catalog and surface "get this outfit" suggestions for whoever is on screen.
Translate zones + motion into sACN/Art-Net universes: wash fixtures follow scene color, movers react to motion, all synced to the video on the headset.
Retarget the 19 landmarks onto a rigged avatar (VMC protocol, Live2D, Unity) — the on-screen performer drives a character in real time.
Compare the instructor's pose stream to the user's (from a webcam running the same landmark model) and score sync, rep counts, and form.
Film-production tool: log clothing-region colors per performer per take; flag when a shirt color drifts between setups.
With pose=both, triangulate each performer's distance from disparity and render a top-down blocking map — where everyone stands over time.
Slow-averaged scene palette drives calm ambient washes for spas, focus rooms, or sleep-wind-down content.
Feed zone motion into OBS/vMix scene switching for a companion 2D broadcast of what the headset viewer is watching.
binary_reader.py (or the Java/Swift codec),
accept the session, and read the stream. The demo handlers are working scaffolds — fork one.