A working, offline assistant for the problem statement's expected solution. Two fixed payload cameras watch a crew member perform a six-step sample-transfer protocol at the rack. Models we trained on our own synthetic dataset track the crew member and the hardware in 3D relative to the rack. The assistant recognises each step, says what to do next, raises a voice alert when a step is skipped or done out of order, writes a timestamped step log, and records and streams the annotated video. Everything below was produced by the system itself on held-out runs it had never seen.
The research behind it (downlinking 3D crew state from real ISS footage, orientation, compression) is a separate report: bas.triixd.tech.
| Expected | What the system does | Evidence |
|---|
Each camera frame goes through one detector that we trained: it returns the crew member's 20 body keypoints, the two samples, and the box with its lid open or closed. Because both cameras are calibrated to the rack once, matching keypoints and samples are triangulated straight into the rack's frame. There is no floor and no "up": positions are relative to the rack face. A small causal network labels each frame with the current action from the hands, the samples and the lid, and a protocol checker turns those labels into confirmed steps.
The recogniser labels a frame once it has seen the next two seconds, so a step is confirmed about three seconds after it starts. The research report measures this trade-off: without the look-ahead, ordering mistakes are confused with skipped samples (bas.triixd.tech, E11). The checker reports a step out of sequence as soon as it is confirmed, names the steps it skipped, flags steps that are not in the protocol, and at the end lists anything never done.
The dataset comes from our own Blender generator. Each run places a mannequin crew member, floating at a random orientation to the rack (full 360° roll, random pitch and drift), in front of a rack with a hinged box, two samples and two pads. It performs the protocol, correctly in 70% of runs and with one injected mistake in the rest: samples swapped in order, a sample skipped, a sample placed on the wrong pad, or the box left open. Body size, clothing and skin colours, lighting, and the two cameras' positions and focal lengths are randomised per run. Every frame comes with exact labels: 3D and 2D positions of 20 body joints and both samples, the lid angle, which hand holds which sample (hand–object contact), and the current step.
Perception. One YOLO11 pose model, started from public COCO weights and fine-tuned only on our renders. It finds five classes (crew, red sample, yellow sample, box closed, box open) and the crew member's 20 keypoints. Training adds random image rotations up to 180° so the detector has no preferred "up" either.
Steps. A causal temporal convolutional network with 84 thousand parameters and 12 seconds of history. It was trained on the generator's labels with perception-like corruption: position noise, the lid seen only as open or closed, and positions that freeze when detection drops out.
The complete system (video in, alerts out) was run on held-out runs, rendered at five frames per second from both cameras, that the models never saw. Each stage is scored against the generator's ground truth.
| Measure | Result |
|---|
One command runs the assistant on two camera feeds (camera indices, video files, stream URLs or folders of frames) with a rack calibration file. It needs no network: the models, the speech synthesiser and the console are all local. We ran it on a laptop with an RTX 4060.
python har/run.py --cam0 0 --cam1 1 --calib rack_calibration.json --out sessions/today \
--voice --gui 8080 --stream udp://192.168.1.20:5000
The console opens at http://127.0.0.1:8080. Each session folder holds video.mp4 (annotated, also streamed), steps.jsonl (the step log), frames.jsonl (per-frame state) and voice/ (every spoken message). Training from scratch: generate.py (Blender) → har/yolo_data.py → har/train_perception.py, and har/train_steps.py.
A step-log entry: