
Over the past two decades, the core thread of human-machine interaction has been enabling machines to better comprehend human "commands".
From keyboards and mice to touchscreens, and then to voice assistants and large language models, interaction efficiency has kept improving, yet the underlying logic remains unchanged: humans first translate their intentions into machine-interpretable instructions, which the machine then executes. Even today’s AI capable of understanding natural language mostly only reaches the stage of “hearing what you say”, rather than “grasping what you are doing, your motivations behind it, and what you intend to do next”.
Embodied intelligence is breaking down this barrier. As robots are deployed in factories, households and service scenarios, the real challenge is no longer mere reasoning, but how to perform stable, delicate and continuous motions in the physical world. Meanwhile, neurotechnology is expanding its reach beyond medical rehabilitation and wearable interaction to machine training, motion comprehension and physical control.
The two originally separate technological paths are converging: embodied intelligence demands data that closely mirrors human control mechanisms, while neurotechnology requires a sufficiently large application landscape. The intersection of the two may spawn the next-generation human-machine interaction paradigm—machines will no longer wait for explicit instructions. Instead, they interpret human intentions from electroencephalography (EEG), electromyograph
01
Shift from "Command Input" to "Intention Decoding"
Traditional interactive devices only capture the outcomes of human movements. A mouse records displacement; cameras capture hand gestures; motion capture systems reconstruct joint trajectories. Such data can answer what a person has done, yet struggle to reveal the underlying motives, exerted force magnitudes, and real-time fine adjustments throughout the movement.
Neural signals precisely fill this gap.
Surface electromyography (sEMG) lies midway along the pathway through which motor commands travel from the brain to muscles. It captures neuromuscular activity preceding movement onset, as well as force exertion and fine motor control signals. Electroencephalography (EEG), by contrast, provides access to higher-level mental states such as motor imagery, attention, surprise, and error perception. Accordingly, the input end of human-machine interaction is shifting forward from overt physical movements to covert internal intentions.
The non-invasive brain-computer interface (BCI) system demonstrated by a research team from Carnegie Mellon University serves as a typical example of this shift. Twenty-one test subjects wore 128-channel EEG caps and controlled individual fingers of a robotic hand in real time via motor imagery. The online classification accuracy hit 80.56% for two-finger tasks and 60.61% for three-finger tasks.

What matters far more than the numerical results themselves is the natural mapping that enables the robotic hand to move whichever finger the user imagines moving. Users no longer need to learn counterintuitive control codes; instead, the system adapts to human motor intent. Online fine-tuning mitigates EEG drift, while smoothing mechanisms suppress label jumping, transforming control from discrete one-shot classification into continuous, stable interaction.
This indicates that future human-machine interfaces will neither necessarily rely on novel displays nor more complex gesture languages. Instead, they are far more likely to be intention-layer interfaces: the machine starts to interpret user motor plans the moment they form; the system captures error-related brain states as soon as the user detects operational mistakes; and the robot initiates coordinated assistance in advance before the user completes the full movement.
02
What Embodied Intelligence Truly Lacks Is Not Merely Video Data, But the Full Control Process
The leaps of large models stem from internet-scale data, yet physical artificial intelligence lacks a ready-made equivalent of the Internet. Data required by robots can only be generated through human interactions with the physical world.
First-person videos, teleoperation and motion capture can record movement trajectories, yet they frequently fail to capture occluded hand postures, muscle force exertion, and the iterative motion correction process.
Take the paper EgoEMG: A Multimodal Egocentric Dataset with Bilateral EMG and Vision for Hand Pose Estimation released by the team led by Associate Professor Jianjiang Feng (tenured-track) from the Department of Automation, Tsinghua University as an example. The EgoEMG dataset unveils another viable pathway for neurotechnology to integrate with embodied intelligence. This dataset synchronously collects bilateral EMG signals, IMU measurements, egocentric RGB frames, external RGB-D footage and optical motion capture data from 41 subjects, covering 60 gesture categories with a total duration exceeding 10 hours, and reconstructs hand kinematics with 22 degrees of freedom (DoF) joint angles. It aligns three layers of information onto a unified timeline: what the user sees, how muscles generate force, and the ultimate kinematic movement of the hand.

The EMGFormer proposed by the team achieves a 22% performance improvement over previous baselines in cross-subject generalization tasks. Moreover, the fusion of EMG and vision outperforms vision-only approaches under challenging conditions including occlusion, motion blur and depth ambiguity.
From an industrial perspective, this reshapes the fundamental unit of robotic training data. In the past, a single training sample typically consisted of paired video footage and corresponding motions; in the future, each sample may encompass scene context, human intent, muscle activation signals, motion trajectories, contact feedback and task outcomes. Robots will no longer merely mimic superficial movements, but approximate the complete human pipeline spanning perception, decision-making and motor execution.
This indicates that the data competition for embodied intelligence is shifting from sheer data volume to causal density. For a given segment of motion video, if it simultaneously incorporates neural intent, force variation and error feedback, its value to the model lies not merely in an extra data modality, but in an additional causal chain that explains why each movement occurs.
03
Closed Loops Between Neural Intent and Robotic Tactile Sensing Are Forging a New Paradigm
Merely decoding human intention is insufficient to achieve complete interaction.
Robots need to perceive whether their grasp is stable, whether objects slip, whether the contacted surface is soft or rigid, and the physical implication behind human commands such as "grasp more gently". Accordingly, neurotechnology addresses the question of what humans intend to do, while tactile sensing reveals what the robot physically achieves in reality.
Existing multimodal tactile systems generally follow a four-stage pipeline. First, deformations, forces and vibrations are converted into digital signals. Next, visual, tactile and linguistic information are encoded separately. Then joint representations are generated via attention mechanisms or contrastive learning. Finally, the system outputs recognition results, textual descriptions or control commands.
Tactile research is also evolving from simple object recognition toward material discrimination, grasp success prediction, cross-modal generation, and continuous manipulation guided by linguistic instructions. For instance, when a user issues the command “gently grasp that soft object”, the system needs to jointly comprehend linguistic semantics, visual geometric features and real-time contact forces, rather than mechanically mapping a sentence to a fixed gripper force value.
This outlines the closed-loop architecture of next-generation interaction: EEG or EMG signals deliver human intentions and physiological states, robot vision interprets scene context, tactile sensing captures contact outcomes, and embodied models handle motion planning and control. Feedback signals are then relayed back to the human operator or leveraged for iterative motion adjustments.
Interaction is no longer a one-way pipeline of "human issuing commands and machines executing them"; instead, it evolves into shared control where both sides continuously predict, calibrate and adapt to one another.
This paradigm delivers direct practical value across rehabilitation prosthetics, industrial human-robot collaboration, smart eyewear and service robots. For prosthetic users, neural interfaces enable intuitive natural control, while tactile feedback helps them judge grasp stability. In industrial collaborative scenarios, robots integrate workers’ EMG signals and brain states to proactively detect takeover intentions or operational hazards. For AI smart glasses, neural wristbands serve as voice-free implicit input devices. As for household robots, subtle human movements, force tendencies and even perceived operational errors can all be converted into collaborative control signals.
04
Industrial Opportunities Will Shift From Hardware Sales to Closed-Loop Mastery
The most transformative shift brought by this integration is not merely an additional category of wearable hardware, but a complete rearrangement of the industrial value chain.
At one end lie data ingress terminals. Neural wristbands, EEG headbands, tactile skins and egocentric devices continuously capture human manipulation data generated in real-world scenarios.
The forward-looking technical roadmap centers on building the data infrastructure for Physical AI via non-invasive neural interfaces. Hardware undertakes signal collection, while AI decoding models translate raw EMG signals into hand postures, motion intentions, force trends and fine-grained micro-control commands, which in turn support robotic training and human-machine interaction. Consequently, its commercial value stems not only from hardware sales, but more significantly from long-term data services and proprietary model capabilities.
At the other end sit closed-loop models. Companies with genuine competitive moats in the future will not merely own a single EEG cap, EMG wristband or electronic skin, but master the full-stack pipeline covering signal acquisition, decoding, multimodal fusion, motion control and sensory feedback.

The underlying reason lies in significant individual variations in neural signals, inconsistent form factors of tactile sensors, and spatiotemporal misalignment across cross-modal data. Standalone hardware products are highly replaceable; the core proprietary assets consist of continuously accumulated human biometric data, adaptive algorithms and mature closed-loop application scenarios.
Naturally, this development path is still constrained by practical challenges: EEG suffers from limited spatial resolution, EMG models struggle with cross-subject generalization, tactile datasets remain far smaller in scale compared to vision and language corpora, there is a lack of unified standards for diverse sensors, and real-time systems must simultaneously balance power consumption, latency, durability and safety requirements.
More importantly, neural data contains highly sensitive human physiological information, which necessitates well-defined mechanisms for data authorization, usage scope delineation and privacy governance in the future. Otherwise, the more seamless the interaction, the deeper the potential exposure of personal data.
Nevertheless, the general technological roadmap has become increasingly clear.
In traditional human-machine interaction, humans learn the machine’s operational logic. The next phase of the intelligent era will see machines learning humanity’s neural language. Embodied intelligence delivers large-scale application scenarios for neural technology, while neural technology endows embodied systems with the capability to interpret human intent, force dynamics and physiological states. Once the two technologies form a closed loop integrated with tactile sensing, computer vision and world models, human-machine relations will evolve from humans merely using tools toward human and robots collaboratively accomplishing tasks.
At that stage, the quality of interaction experience will no longer be determined by the number of physical buttons, screen size or voice recognition accuracy. Instead, the core metric lies in whether machines can interpret human intent at the right moment, execute appropriate motions within real-world environments, and iteratively adjust behaviors in collaboration with humans via continuous sensory feedback.
The integration of neural technology and embodied intelligence does not merely reshape a new interaction entry point; it fundamentally redefines how humans and machines mutually comprehend each other and jointly act upon the physical world.









