Implementation · BAS on-board activity recognition

BAS HAR Assistant

A working, offline assistant for the problem statement's expected solution. Two fixed payload cameras watch a crew member perform a six-step sample-transfer protocol at the rack. Models we trained on our own synthetic dataset track the crew member and the hardware in 3D relative to the rack. The assistant recognises each step, says what to do next, raises a voice alert when a step is skipped or done out of order, writes a timestamped step log, and records and streams the annotated video. Everything below was produced by the system itself on held-out runs it had never seen.

The research behind it (downlinking 3D crew state from real ISS footage, orientation, compression) is a separate report: bas.triixd.tech.

Figure 1. The operator console, replaying sessions exactly as the system recorded them: the annotated camera video it stored and streamed, the voice messages it spoke (press play; voice can be muted), its step log, the protocol checklist and the crew in the rack's frame. Choose a session at the top right; each is a held-out run with the mistake noted in its title. Run live, the same console reads from the system instead of from files.

1The expected solution, delivered

Table 1. Each item of the problem statement, what the system does for it, and the evidence.
ExpectedWhat the system doesEvidence

2How it works

Each camera frame goes through one detector that we trained: it returns the crew member's 20 body keypoints, the two samples, and the box with its lid open or closed. Because both cameras are calibrated to the rack once, matching keypoints and samples are triangulated straight into the rack's frame. There is no floor and no "up": positions are relative to the rack face. A small causal network labels each frame with the current action from the hands, the samples and the lid, and a protocol checker turns those labels into confirmed steps.

Payload camera 05 frames/s Payload camera 15 frames/s Detector (trained)keypoints · samples · lid Triangulation3D in the rack frame Step recogniser (trained)action per frame Protocol checkernext · skipped · order Voice alerts Step log (JSONL) Console (GUI) MP4 on disk+ UDP stream to IP annotated frames
Figure 2. The on-board pipeline. Teal boxes are models we trained. Nothing leaves the station except the optional video stream, and the system runs without a network.

The recogniser labels a frame once it has seen the next two seconds, so a step is confirmed about three seconds after it starts. The research report measures this trade-off: without the look-ahead, ordering mistakes are confused with skipped samples (bas.triixd.tech, E11). The checker reports a step out of sequence as soon as it is confirmed, names the steps it skipped, flags steps that are not in the protocol, and at the end lists anything never done.

3Dataset generation

The dataset comes from our own Blender generator. Each run places a mannequin crew member, floating at a random orientation to the rack (full 360° roll, random pitch and drift), in front of a rack with a hinged box, two samples and two pads. It performs the protocol, correctly in 70% of runs and with one injected mistake in the rest: samples swapped in order, a sample skipped, a sample placed on the wrong pad, or the box left open. Body size, clothing and skin colours, lighting, and the two cameras' positions and focal lengths are randomised per run. Every frame comes with exact labels: 3D and 2D positions of 20 body joints and both samples, the lid angle, which hand holds which sample (hand–object contact), and the current step.

Figure 3. Held-out validation images with the trained detector's output: crew keypoints and boxes for the crew, the samples and the box (open or closed).

4Trained models

Perception. One YOLO11 pose model, started from public COCO weights and fine-tuned only on our renders. It finds five classes (crew, red sample, yellow sample, box closed, box open) and the crew member's 20 keypoints. Training adds random image rotations up to 180° so the detector has no preferred "up" either.

Steps. A causal temporal convolutional network with 84 thousand parameters and 12 seconds of history. It was trained on the generator's labels with perception-like corruption: position noise, the lid seen only as open or closed, and positions that freeze when detection drops out.

keypoints mAP50-95boxes mAP50-95
Figure 4. Validation accuracy of the perception model over training (runs held out from training).

5Results on held-out runs

The complete system (video in, alerts out) was run on held-out runs, rendered at five frames per second from both cameras, that the models never saw. Each stage is scored against the generator's ground truth.

Table 2. The full system on held-out runs.
MeasureResult

6Running it offline

One command runs the assistant on two camera feeds (camera indices, video files, stream URLs or folders of frames) with a rack calibration file. It needs no network: the models, the speech synthesiser and the console are all local. We ran it on a laptop with an RTX 4060.

python har/run.py --cam0 0 --cam1 1 --calib rack_calibration.json --out sessions/today \
    --voice --gui 8080 --stream udp://192.168.1.20:5000

The console opens at http://127.0.0.1:8080. Each session folder holds video.mp4 (annotated, also streamed), steps.jsonl (the step log), frames.jsonl (per-frame state) and voice/ (every spoken message). Training from scratch: generate.py (Blender) → har/yolo_data.py → har/train_perception.py, and har/train_steps.py.

A step-log entry:


  

7Limitations