Computer Vision

Multimodal OR Video Analytics Pipeline

A research-grade multimodal video analytics pipeline for operating rooms - fusing full-body pose, hand state, gaze direction, posture classification, and multi-person tracking with a LLaVA v1.5-7B vision-language model fine-tuned via QLoRA to recognise surgical staff roles from multi-camera OR footage.

Pakistan2025

Key Outcome

Research-grade OR understanding

Automated multi-camera surgical scene analysis - roles, gaze, posture, and activities from video alone.

The Problem

Surgical workflow monitoring in operating rooms still relies on manual observation, paper checklists, and retrospective video review - slow, inconsistent, and unable to scale across hundreds of procedures for quality assurance or trainee feedback.

Our Solution

Built an end-to-end multimodal pipeline that processes OR video feeds through specialised perception modules - body pose, hand state, gaze, posture, and multi-person tracking - and fuses all signals to recognise surgical staff roles using a LLaVA v1.5-7B model fine-tuned with QLoRA. Output is structured scene graphs and activity timelines suitable for clinical audit, trainee performance dashboards, and surgical workflow research.

Key Features

Full-body pose estimation providing the spatial foundation for all higher-level understanding

Hand pose and state detection: open, closed, holding instrument, or in contact with surface

Head orientation and gaze estimation showing where each OR staff member is looking

Posture classification: actively operating, standing and observing, leaning in, stepping back

Multi-person tracking with persistent identity and trajectory analysis across frames

LLaVA v1.5-7B fine-tuned with QLoRA for surgical staff role recognition from visual context

Activity understanding module recognising higher-level events: instrument handoffs, incision, timeout

Scene graph generation: entities (people, instruments, zones) and their relationships in a queryable format

Technology Stack

PythonLLaVAQLoRAYOLOv8OpenCVPyTorchMediaPipeMedical AIComputer Vision

Project Visuals

Multimodal OR Video Analytics Pipeline screenshot 1
Multimodal OR Video Analytics Pipeline screenshot 2
Multimodal OR Video Analytics Pipeline screenshot 3
Multimodal OR Video Analytics Pipeline screenshot 4

Interested in a similar project?

Tell us your requirements and we will respond within one business day.