KERKorea Emotional Robot Inc.KO

KER / CORE TECHNOLOGY

Integrated development across perception, conversation, expression and control

Core technology

Lumira robot and application scene

01 / CORE TECHNOLOGY

Integrated AI architecture

01

Voice & visual input

Speech, expression and gesture cues

02

Perception & dialogue

Context and permitted preferences

03

Response planning

Coordinate voice, expression and gesture

04

Check & execute

Check motion limits in local control

The presentation brings LLMs, Vision AI, Speech, Emotion AI and Robot AI into a shared platform. Stage 1 focuses on turning user input into appropriate responses expressed consistently through voice, facial animation and movement.

The following is an explanatory development design based on that direction. It does not represent completed implementation or validated performance, and may change after pilots.

LayerInput and processingOutput / boundary
PerceptionExtract speech and interaction cues from microphones, camera and touchText, object or gesture candidates and confidence.
UnderstandingUse dialogue context and permitted preferencesResponse text and bounded expressive intent.
ExpressionCoordinate voice, animation and feasible motionApproved motion requests and audio output.
Local controlMonitor actuator state, sensors and motion limitsActuation, stop behavior and error state.
OperationsManage device health, versions and consent settingsUpdates, diagnostics and support.
View detailed information +
View detailed information
Complete original figure. Click to view at full size. Metrics and schedules are planned or illustrative; original Korean labels are retained.

The input layer will associate timestamps with speech, images, touch and device status. The decision layer connects the current dialogue state with the request, while the output layer coordinates speech, display and movement to reduce conflicting responses.

Defined inputs, outputs and timeout behavior make it easier to replace a model or sensor without redesigning the whole system. Integration tests will trace delays through the layers and check that uncertain results are not treated as established facts downstream.

02 / CORE TECHNOLOGY

Conversation, speech & memory

Speech-to-text converts utterances into text. The dialogue engine uses the question and conversation context to formulate a response, which text-to-speech turns into audio. Clarification flows are needed for noise, distance and overlapping speakers.

Personalized memory is intended to use preferences and context the user permits. Development separates short-term context from longer-term preferences and defines storage scope, retention and deletion controls. Questions requiring external information depend on the availability of connected services.

Validation: transcription error rate, response latency distribution, context continuity, interruption handling and network failure feedback. Model providers and service configurations will be selected during integration.

View detailed information +
View detailed information
Complete original figure. Click to view at full size. Metrics and schedules are planned or illustrative; original Korean labels are retained.

Audio development will examine interference from the robot’s own speaker and background noise. Voice activity detection, switching back to listening when the user interrupts, and concise clarification after recognition failure will be designed together. Latency will be measured separately for recognition, reasoning and synthesis.

Memory design distinguishes preferences worth retaining from information needed only within the current conversation. Users should be able to correct or remove inaccurate stored information. Repeated questions and missing context will also be considered when evaluating response quality.

03 / CORE TECHNOLOGY

Vision AI · visual perception

Visual perception extracts interaction cues such as people, objects and hand gestures from camera input. Stage 1 considers cues relevant to engagement, while spatial awareness connects to later navigation development.

The proposed pipeline processes images, detects objects, extracts features and classifies results, with confidence and timestamps attached. Under occlusion or changing light, uncertain results should lead to clarification or use of other sensors.

Validation: detection across lighting, distance and viewing angles, false positives and misses, frame processing time and behavior when the camera is blocked. Illustrative metrics in the presentation are not product guarantees.

View detailed information +
View detailed information
Complete original figure. Click to view at full size. Metrics and schedules are planned or illustrative; original Korean labels are retained.

Visual outputs will include position, timestamp and confidence so that stale observations do not drive current behavior. Development will examine how to select a conversation partner when several people are visible and how to respond when that person leaves the frame.

Camera position and field of view interact with the face display, lighting and installation height. Trials will distinguish distance, backlighting and partial occlusion. Processing speed and temperature will inform model size and input resolution for the deployed hardware.

04 / CORE TECHNOLOGY

Emotion AI · estimation & expression

Emotion AI considers facial cues, vocal prosody, spoken content and conversational context when choosing a response. A facial expression alone cannot establish a person’s actual feelings or health, so classifications are treated as uncertain cues.

Responses combine wording, vocal style, on-screen expressions and head movement. Development considers how to adjust when the user reacts differently and how users can reduce unwanted interventions.

Validation: conflicting cues, synchronization across voice and movement, user evaluations and variation across ages and settings. Emotion estimates are not intended for medical judgments.

View detailed information +
View detailed information
Complete original figure. Click to view at full size. Metrics and schedules are planned or illustrative; original Korean labels are retained.

When speech, facial and language cues disagree, the design will consider their confidence and context instead of immediately forcing a single label. Responses can clarify intent or use a low-intensity reaction. An inferred emotion should not override what a person says about their own feelings.

The expression library will connect facial and motion patterns for listening, greeting, positive engagement and guidance. Evaluation will consider repetition fatigue and differences in interpretation across age groups and cultures, with adjustable motion and vocal intensity as a design direction.

05 / CORE TECHNOLOGY

ROS 2 · motion control

The platform direction uses ROS 2 to connect sensing, perception, behavior and actuation as modules. Defined interfaces and module states help keep changes to dialogue logic from directly affecting motion safety.

Natural-language responses should not directly become motor commands. The proposed behavior layer selects an approved gesture, and the control layer checks joint limits, speed and sensor state before execution. Lost connections or stale commands require defined stop or safe idle behavior.

Validation: motion repeatability, limit enforcement, sensor and communication faults, and restart recovery. Stage 2 adds localization and route planning; Stage 3 adds grasping and arm planning.

View detailed information +
View detailed information
Complete original figure. Click to view at full size. Metrics and schedules are planned or illustrative; original Korean labels are retained.

Gesture commands will include timing, validity and cancellation conditions, with a defined waiting pose after interruption. Paths and speeds will be coordinated across the head, arms and waist. Repeated commands caused by communication errors should not trigger unintended continued motion.

Tests will progress from individual joint limits to sequences and extended repetition. Actuators with usable current, temperature or position feedback will be considered. Fault and command-timeout responses will be designed to operate independently of the conversational AI connection.

06 / CORE TECHNOLOGY

Hardware & embedded systems

The Stage 1 concept integrates facial and chest displays, camera and microphones, servo actuators, wheels and a battery. Development must account for balance, cable routing, noise, heat and maintenance access as well as component placement.

The embedded layer is intended to handle actuation requests, sensor collection and power monitoring. Battery protection, low-voltage behavior, actuator overload and permitted motion during charging need testing under realistic conditions.

Validation: temperature and current under load, extended operation, power loss and recovery, mechanical interference and pinch hazards. Certification and production conditions will be determined after prototype validation.

View detailed information +
View detailed information
Complete original figure. Click to view at full size. Metrics and schedules are planned or illustrative; original Korean labels are retained.

Commercial computing boards and sensor modules will be used to validate key functions before deciding which custom electronics are needed. Interface definitions across cameras, microphones, displays and motor control will support replacement and diagnosis as space, wiring and power requirements become clearer.

Mechanical and electrical design will be evaluated together. Repeated joint motion can stress cables, while cooling affects appearance and noise. Connector identification, assembly order and access to service parts will be checked during prototyping to make later builds more repeatable.

07 / CORE TECHNOLOGY

Edge AI, cloud & OTA

Local processing is intended for latency-sensitive sensor responses and device state management. Cloud services support expanded dialogue, device management and updates. The design must distinguish functions that remain available offline from those that require connectivity.

The operations platform direction includes enrollment, access roles, software versions, battery and network state. OTA development considers package integrity, staged rollout, failure detection and rollback, with updates scheduled for suitable power and safe idle conditions.

Validation: connectivity loss, power interruption during updates, version compatibility and recovery. Fleet operation and subscription management are later service extensions.

View detailed information +
View detailed information
Complete original figure. Click to view at full size. Metrics and schedules are planned or illustrative; original Korean labels are retained.

Time-sensitive responses will be kept local where practical, with cloud services supporting external knowledge and expanded conversation. Response policies will reflect connectivity, and requests will be associated with results to reduce duplicate execution after reconnection.

Operations will track device versions and fault histories to identify the effects of a change. Updates will be considered for staged rollout after limited validation, with recovery tested under storage shortage or power interruption. Cloud usage will be managed alongside user experience and operating cost.

08 / CORE TECHNOLOGY

Privacy, safety & validation

Development separates information needed for personalization from diagnostic logs, defining purpose and consent scope early. Considerations include understandable microphone and camera status, access controls and tools to review or delete stored information.

Testing progresses from individual modules to integrated scenarios involving speech, displays and actuators. Failure cases include recognition errors, network delay, sensor faults and low battery, with defined feedback and stopping behavior.

PhaseValidation outputs
Module testingInput/output definitions, error handling and test results.
Prototype integrationInteraction traces, load testing and power behavior.
Field pilotsUsability observations, reproducible issues and improvement priorities.
Release preparationFinal specifications, applicable certification review and operating procedures.

Each function will have defined normal behavior and acceptable failure states. Reproducible tests will combine audio, display and motor activity. Defects will be tied to triggering conditions and software revisions, followed by checks of affected functions after a correction.

Information needs will be defined before deciding access and retention. Pilot participants should understand the purpose of use and the information involved. Clear camera and microphone status remains a design goal. Product reliability will be developed through these decisions and repeated verification.

LET’S TALK

Let’s talk about
life with robots.

Write an inquiry ↗
Contact ↗