Autonomous Assistive Robot
Recipient-directed assistance through active perception and bounded motion
An assistive robot designed to reach a selected caregiver. A tracked LEGO EV3 platform and Jetson perception stack coordinate enrolled identity, camera viewpoints, local route assessment, and short movements. The workflow searches a prepared indoor room and reacquires the recipient between approach segments, making physical action depend on current visual and motor evidence.
Contents
Identity
Face Recognition and Appearance Continuity
YOLOX-s detects person candidates, YuNet locates faces, and InsightFace embeddings compare observations with enrolled profiles. Repeated matches provide temporal confirmation rather than treating one frame as sufficient evidence. TensorRT inference keeps the person and face pipeline on the Jetson, close to the camera.
Face-anchored clothing memory supports continuity when the recipient turns away or becomes partly obscured. Colour and texture descriptors compare body crops, including ordered strip signatures for partial views. Identity evidence and ambiguity checks qualify those appearance matches before the mission uses them.
| Configuration | Identity retained | Evaluable pairs |
|---|---|---|
| Baseline tracker | 220 | 297 |
| Appearance continuity | 289 | 297 |
Both configurations use the same recorded frame–detection pairs and a single identity anchor. The eight remaining pairs precede that anchor; this comparison measures continuity on this recording.
Camera
Active Perception and Camera Coordination
A motorized webcam handles both recipient observation and floor inspection. Useful-view calibration discovers suitable camera positions within the current motor reference, rather than assuming an encoder count always represents the same optical direction. Camera leases and head–track interlocks coordinate those transitions; movement invalidates old position evidence before the recipient is reacquired.
Approach
Local Route Assessment and Bounded Motion
SegFormer floor segmentation proposes local image corridors, while Gemini reviews their visible hazards. A fresh scene check follows the cloud response, and missing or unusable route evidence withholds movement. Monocular depth informs the approach policy without being treated as an exact physical measurement.
Short movement primitives limit reliance on any single observation. A retained supervised trial combined a −39.82° turn, 6.39 cm of encoder-estimated travel, and facial reacquisition; the result captures one checked approach segment.
Execution
State-Bound Distributed Inference
The Jetson coordinates the mission through ROS 2 while a companion Mac runs depth, segmentation, and cloud adapters. Asynchronous requests keep robot feedback processing active during heavier inference. Returned assessments are checked against their source images and applicable camera, identity, and motion bindings before use. Serialized EV3 transactions and bounded recovery coordinate actuator commands, while an EV3-local watchdog stops the tracks when drive-command refreshes expire or the client disconnects.
Assistance
Recipient-Bound Request Delivery
Typed text or an edited speech transcript is reviewed before becoming a recipient-bound request. Generated audio can be previewed through the robot speaker. Delivery playback requires current recipient identity, validated measured distance, and stopped track and camera feedback. Profile and message revisions keep approval attached to the intended person and wording; incomplete arrival evidence leaves playback withheld.
Technology
Perception
- Jetson Orin Nano
- TensorRT
- OpenCV
Coordination
- ROS 2
- Gemini
Actuation and storage
- LEGO EV3
- SQLite