Autonomous Assistive Robot

Recipient-directed assistance through active perception and bounded motion

An assistive robot designed to reach a selected caregiver. A tracked LEGO EV3 platform and Jetson perception stack coordinate enrolled identity, camera viewpoints, local route assessment, and short movements. The workflow searches a prepared indoor room and reacquires the recipient between approach segments, making physical action depend on current visual and motor evidence.

Contents

Identity

Face Recognition and Appearance Continuity

YOLOX-s detects person candidates, YuNet locates faces, and InsightFace embeddings compare observations with enrolled profiles. Repeated matches provide temporal confirmation rather than treating one frame as sufficient evidence. TensorRT inference keeps the person and face pipeline on the Jetson, close to the camera.

Face-anchored clothing memory supports continuity when the recipient turns away or becomes partly obscured. Colour and texture descriptors compare body crops, including ordered strip signatures for partial views. Identity evidence and ambiguity checks qualify those appearance matches before the mission uses them.

Identity continuity on one retained recording
ConfigurationIdentity retainedEvaluable pairs
Baseline tracker220297
Appearance continuity289297

Both configurations use the same recorded frame–detection pairs and a single identity anchor. The eight remaining pairs precede that anchor; this comparison measures continuity on this recording.

Camera

Active Perception and Camera Coordination

A motorized webcam handles both recipient observation and floor inspection. Useful-view calibration discovers suitable camera positions within the current motor reference, rather than assuming an encoder count always represents the same optical direction. Camera leases and head–track interlocks coordinate those transitions; movement invalidates old position evidence before the recipient is reacquired.

Perception and movement with one shared camera The robot confirms a selected person, switches to a floor view for route assessment, makes one short encoder-monitored move, stops, and returns to a person view to reacquire identity. Fresh observations are required before the next movement segment. Perception and movement One shared camera changes view as the mission progresses. Person view Floor view Person view Confirm identity Inspect the route Move one segment Reacquire identity Face and appearancecontinuity Visible corridorsand hazard assessment Short, monitoredmovement, then stop Fresh recipientobservations Recheck current observations before the next segment Conceptual sequence; the route is assessed from current camera observations.
Perception and movement with one shared camera A sequence of person observation, floor inspection, short monitored movement and stopped person reacquisition. The camera changes view, and the cycle returns to fresh identity observations before another segment. Perception and movement One camera changes view through the movement cycle. Confirm identity Inspect the route Move one segment Reacquire identity Person viewFace and appearancecontinuity Floor viewVisible corridorsand hazard assessment Short, monitoredmovement, then stop Person viewFresh recipientobservations Recheck current observationsbefore the next segment.
The shared camera alternates between person observation and floor inspection. Each bounded movement ends with a stop and fresh recipient observations before the next segment.

Approach

Local Route Assessment and Bounded Motion

SegFormer floor segmentation proposes local image corridors, while Gemini reviews their visible hazards. A fresh scene check follows the cloud response, and missing or unusable route evidence withholds movement. Monocular depth informs the approach policy without being treated as an exact physical measurement.

Short movement primitives limit reliance on any single observation. A retained supervised trial combined a −39.82° turn, 6.39 cm of encoder-estimated travel, and facial reacquisition; the result captures one checked approach segment.

Execution

State-Bound Distributed Inference

The Jetson coordinates the mission through ROS 2 while a companion Mac runs depth, segmentation, and cloud adapters. Asynchronous requests keep robot feedback processing active during heavier inference. Returned assessments are checked against their source images and applicable camera, identity, and motion bindings before use. Serialized EV3 transactions and bounded recovery coordinate actuator commands, while an EV3-local watchdog stops the tracks when drive-command refreshes expire or the client disconnects.

Assistance

Recipient-Bound Request Delivery

Typed text or an edited speech transcript is reviewed before becoming a recipient-bound request. Generated audio can be previewed through the robot speaker. Delivery playback requires current recipient identity, validated measured distance, and stopped track and camera feedback. Profile and message revisions keep approval attached to the intended person and wording; incomplete arrival evidence leaves playback withheld.

Technology

Perception

  • Jetson Orin Nano
  • TensorRT
  • OpenCV

Coordination

  • ROS 2
  • Gemini

Actuation and storage

  • LEGO EV3
  • SQLite