Logo
Orbis Health

Client

Orbis Health

Duration

2 Weeks

Year

2026

#AI / Machine Learning#Computer Vision#Healthcare

Orbis Health

AI Posture Detection from a Ceiling Camera

Orbis Health moved from having no usable training data at all to a lightweight, real-time posture detection system that reliably distinguishes standing, sitting, and lying down from directly overhead, without wires or wearable sensors on the patient.

Challenges

Almost No Data to Learn From

No Labeled Training Data

Only one video of one person in one room was available. No manual labels existed to train a standard classifier

Extreme Data Scarcity

A single video source meant any model trained directly on it would overfit to one person and one bed setup

Unusual Camera Geometry

Straight-down ceiling view breaks standard pose-estimation and off-the-shelf posture models built for eye-level photos

Class Imbalance

Real patient behavior skewed the data heavily toward “lying down,” under-representing standing and sitting

Poor Generalization Risk

Early models confused “lying down” with “sitting” when tested on a different person in the same room

Edge Deployment Constraints

The final system had to run in real time on a low-power Jetson Orin device, ruling out large, slow models

Solution

A Three-Part Pipeline Built for Data Scarcity

The team designed a pipeline that automatically labeled the raw video, trained a lightweight classifier on top of an existing detector, and deliberately sourced a second dataset that matched the camera's viewing angle rather than simply adding more data.

1

Automatic Data Labeling

YOLO + MetaCLIP 2 + OpenCLIP Ensemble

Detects and crops the person in each frame, then uses a weighted, agreement-checked ensemble of zero-shot vision-language models to label posture, discarding any frame that isn't confidently and consistently classified

2

Lightweight Posture Classifier

Frozen YOLO Backbone + 2-Layer Classifier Head

Reuses YOLO's internal learned features instead of retraining from scratch; only a small trainable head is fit to the three posture classes, avoiding overfitting on a small dataset

3

Targeted Data Augmentation

Flip, Color Jitter, Crop, Blur, Occlusion

Squeezes additional variety out of a small training set, while deliberately excluding rotation, since the camera angle itself carries the signal that separates lying from sitting

4

Angle-Matched Domain Generalization

FallDataset (Corner-Mounted, Downward-Angled Cameras)

Adds real variety across 5 people and 5 rooms filmed from a similar overhead-style angle, after a mismatched public dataset made results worse instead of better

5

Real-Time Edge Inference

Jetson Orin Deployment

Runs the full detect–extract–classify pipeline in ~30ms per frame (~27–28 FPS), fast enough for live clinical monitoring on-device

Transformation

Before vs. After

Before
After

No labels, no way to train a classifier

Automated ensemble labeling pipeline (610 clean, verified images)

Single-person, single-room training data

Domain-generalized dataset via angle-matched FallDataset

Confidently wrong zero-shot models used blindly

Multi-model agreement checks reject unreliable labels

Fine-tuning the full backbone hurt performance

Frozen backbone + small trainable head, tuned for small data

Eye-level datasets added, accuracy dropped sharply

Camera-angle-matched dataset added, accuracy jumped

Vision-language models too slow for live use

Lightweight classifier runs in real time on Jetson Orin

Results

The Numbers Speak

0%

Final accuracy after domain generalization

0

Clean, auto-labeled training images

~30ms

Per-frame inference time (~27–28 FPS)

0

Posture classes reliably distinguished

Technology

Stack at a Glance

Person Detection & Feature Extraction

YOLO (frozen backbone)

Zero-Shot Labeling

MetaCLIP 2 + OpenCLIP Ensemble

Posture Classification

Custom 2-Layer Trainable Head

Domain Generalization

FallDataset (angle-matched transfer data)

Data Augmentation

Flip, Color Jitter, Crop, Blur, Random Erasing

Deployment Target

NVIDIA Jetson Orin (on-device, real-time inference)

Outcome

What Our Client Gained

Our client moved from having no usable training data at all to a lightweight, real-time posture detection system that reliably distinguishes standing, sitting, and lying down from directly overhead, without wires or wearable sensors on the patient. By solving the data problem with careful automatic labeling, targeted augmentation, and angle-matched domain generalization instead of simply chasing more data, the team delivered a model accurate and fast enough to run live on low-power edge hardware, ready for real clinical deployment.

Contact

AI-Accelerated Engineering with Real-World Impact.

We help businesses move faster, work smarter, and scale with confidence. Whether you're looking to automate operations, build AI-powered products, or modernise your data infrastructure, we'll map the fastest path to value.

Response timeWithin 24 hours
EngagementsFrom 4-week sprints to ongoing
First stepFree discovery call

© Technovate Global, All rights reserved.

Technovate Global