01 / CORE TECHNOLOGY
Integrated AI architecture
Voice & visual input
Speech, expression and gesture cues
Perception & dialogue
Context and permitted preferences
Response planning
Coordinate voice, expression and gesture
Check & execute
Check motion limits in local control
The presentation brings LLMs, Vision AI, Speech, Emotion AI and Robot AI into a shared platform. Stage 1 focuses on turning user input into appropriate responses expressed consistently through voice, facial animation and movement.
The following is an explanatory development design based on that direction. It does not represent completed implementation or validated performance, and may change after pilots.
| Layer | Input and processing | Output / boundary |
|---|---|---|
| Perception | Extract speech and interaction cues from microphones, camera and touch | Text, object or gesture candidates and confidence. |
| Understanding | Use dialogue context and permitted preferences | Response text and bounded expressive intent. |
| Expression | Coordinate voice, animation and feasible motion | Approved motion requests and audio output. |
| Local control | Monitor actuator state, sensors and motion limits | Actuation, stop behavior and error state. |
| Operations | Manage device health, versions and consent settings | Updates, diagnostics and support. |
02 / CORE TECHNOLOGY
Conversation, speech & memory
Speech-to-text converts utterances into text. The dialogue engine uses the question and conversation context to formulate a response, which text-to-speech turns into audio. Clarification flows are needed for noise, distance and overlapping speakers.
Personalized memory is intended to use preferences and context the user permits. Development separates short-term context from longer-term preferences and defines storage scope, retention and deletion controls. Questions requiring external information depend on the availability of connected services.
Validation: transcription error rate, response latency distribution, context continuity, interruption handling and network failure feedback. Model providers and service configurations will be selected during integration.
03 / CORE TECHNOLOGY
Vision AI · visual perception
Visual perception extracts interaction cues such as people, objects and hand gestures from camera input. Stage 1 considers cues relevant to engagement, while spatial awareness connects to later navigation development.
The proposed pipeline processes images, detects objects, extracts features and classifies results, with confidence and timestamps attached. Under occlusion or changing light, uncertain results should lead to clarification or use of other sensors.
Validation: detection across lighting, distance and viewing angles, false positives and misses, frame processing time and behavior when the camera is blocked. Illustrative metrics in the presentation are not product guarantees.
04 / CORE TECHNOLOGY
Emotion AI · estimation & expression
Emotion AI considers facial cues, vocal prosody, spoken content and conversational context when choosing a response. A facial expression alone cannot establish a person’s actual feelings or health, so classifications are treated as uncertain cues.
Responses combine wording, vocal style, on-screen expressions and head movement. Development considers how to adjust when the user reacts differently and how users can reduce unwanted interventions.
Validation: conflicting cues, synchronization across voice and movement, user evaluations and variation across ages and settings. Emotion estimates are not intended for medical judgments.
05 / CORE TECHNOLOGY
ROS 2 · motion control
The platform direction uses ROS 2 to connect sensing, perception, behavior and actuation as modules. Defined interfaces and module states help keep changes to dialogue logic from directly affecting motion safety.
Natural-language responses should not directly become motor commands. The proposed behavior layer selects an approved gesture, and the control layer checks joint limits, speed and sensor state before execution. Lost connections or stale commands require defined stop or safe idle behavior.
Validation: motion repeatability, limit enforcement, sensor and communication faults, and restart recovery. Stage 2 adds localization and route planning; Stage 3 adds grasping and arm planning.
06 / CORE TECHNOLOGY
Hardware & embedded systems
The Stage 1 concept integrates facial and chest displays, camera and microphones, servo actuators, wheels and a battery. Development must account for balance, cable routing, noise, heat and maintenance access as well as component placement.
The embedded layer is intended to handle actuation requests, sensor collection and power monitoring. Battery protection, low-voltage behavior, actuator overload and permitted motion during charging need testing under realistic conditions.
Validation: temperature and current under load, extended operation, power loss and recovery, mechanical interference and pinch hazards. Certification and production conditions will be determined after prototype validation.
07 / CORE TECHNOLOGY
Edge AI, cloud & OTA
Local processing is intended for latency-sensitive sensor responses and device state management. Cloud services support expanded dialogue, device management and updates. The design must distinguish functions that remain available offline from those that require connectivity.
The operations platform direction includes enrollment, access roles, software versions, battery and network state. OTA development considers package integrity, staged rollout, failure detection and rollback, with updates scheduled for suitable power and safe idle conditions.
Validation: connectivity loss, power interruption during updates, version compatibility and recovery. Fleet operation and subscription management are later service extensions.
08 / CORE TECHNOLOGY
Privacy, safety & validation
Development separates information needed for personalization from diagnostic logs, defining purpose and consent scope early. Considerations include understandable microphone and camera status, access controls and tools to review or delete stored information.
Testing progresses from individual modules to integrated scenarios involving speech, displays and actuators. Failure cases include recognition errors, network delay, sensor faults and low battery, with defined feedback and stopping behavior.
| Phase | Validation outputs |
|---|---|
| Module testing | Input/output definitions, error handling and test results. |
| Prototype integration | Interaction traces, load testing and power behavior. |
| Field pilots | Usability observations, reproducible issues and improvement priorities. |
| Release preparation | Final specifications, applicable certification review and operating procedures. |







