Why physical AI’s next frontier is the robot’s periphery, not its brain
A humanoid hand is lifting a glass of water. In one fingertip, a tactile array registers the first micro-slip: a shear pattern that means the glass has started to move against the skin. Before it moves much further, grip force across several fingers has to rise, in proportion, together. The model that chose to pick up the glass is not involved. It is milliseconds away, in a different timing domain, thinking about what to do next.
That correction is a physical AI achievement, and it does not happen in the robot’s brain. It happens in its body, in a chain of sensing, communication and actuation that runs from the fingertip to the joint and back, and that must complete on time, every time.

For a decade, “edge AI” meant moving inference out of the data centre and into the edge device. A robot is not one edge device. It is a distributed real-time system, and the new edge is the robot’s own body.
A hand is a distributed real-time system
Consider what a dexterous hand actually contains. The Shadow Dexterous Hand, one of the most capable commercial hands, has 20 actuated degrees of freedom across 24 joints and 129 sensors: absolute position at each joint, force at each actuator, tactile sensing at the fingertips, plus temperature, current and voltage. Each motor node runs its own controller, and the sensor data travels over EtherCAT.
That is not one computer. It is dozens of control loops, running at different rates, physically distributed across a structure the size of a human hand, all of which must agree on what is happening and when.
| Timing domain | Where it runs | Typical rate |
|---|---|---|
| Motor current loop (field-oriented control) | At each actuator | Tens of kHz |
| Joint position and force loops | At or near each joint | 1–4 kHz |
| Tactile and proprioceptive sampling | Fingertips, joints, palm | Around 1 kHz per sensor |
| Grasp and reflex coordination | Hand or wrist controller | Hundreds of Hz to 1 kHz |
| Policy and planning | Central compute | Tens of Hz |
The rates span three orders of magnitude, and every domain feeds the one above it. The policy at the top of the table only works if the loops beneath it deliver their results on schedule. The earlier argument in this series, that Physical AI needs two kinds of compute, described that split as two layers. In a real robot, the reflex layer is not a layer. It is a fabric, spread across the body.
Determinism composes, or it doesn’t
Follow the slip correction through the hand. It is a chain of hops, and each hop has its own worst case:
- The tactile array at the fingertip is sampled and timestamped.
- A local node filters the shear signal and decides that a slip event has started.
- The event crosses a link to the hand controller.
- The controller computes a coordinated grip-force increase for several fingers.
- Commands cross links back out to the finger actuators.
- Each actuator’s force and current loops apply the new set-point.
The end-to-end bound on that path is, roughly, the sum of the worst case at every hop, plus the error between the clocks that timestamp each step. If every hop is bounded, the path is bounded, and the bound can be computed at design time. That is what it means for determinism to compose.
The converse is the uncomfortable part. A single hop without a bound makes the whole path unbounded. It might be a node whose interrupt latency depends on what else it is doing, a shared bus, or a link whose arbitration varies with load. Its average may be excellent; it is still the hop that sets the guarantee, because the guarantee is a worst-case property.
Time coherence matters as much as latency. A grasp controller fusing tactile, joint-position and force data needs those samples to describe the same instant. As an earlier post in this series put it, misaligned inputs do not produce slightly worse outputs; they produce well-formed answers to the wrong question. Across a distributed body, alignment is only as good as the shared notion of time between nodes.
This is the same asymmetry the XCORE GenSoC substrate argument made for a single chip, now applied to the whole robot. A deterministic fabric can carry best-effort traffic on top of it. No amount of software above a non-deterministic hop can recover a bound that the hop never provided. The weakest hop decides what the whole robot can promise.
The fragmentation tax
Today’s robots meet these constraints largely through careful integration, joint by joint. A typical actuator node pairs a motor-control microcontroller with a separate fieldbus interface. Sensor aggregation may sit on an FPGA or a further microcontroller, with signal processing on another part again. Each device is individually well understood, and each brings its own timing model, toolchain and interrupt structure.
That approach works, and it has shipped a great deal of excellent machinery. But it carries three costs that grow with the number of joints.
- No end-to-end analysis. Each part’s timing can be characterised in isolation. The path through several of them is usually validated empirically – by measurement – not computed. Measurement tells you what happened in the test; it does not bound what will happen across a fleet.
- Integration effort per node type. Every distinct actuator, sensor or interface tends to become its own small engineering project. A humanoid has many distinct node types, not one repeated many times.
- Rigidity. When the timing behaviour of a node is fixed by its silicon, adding a sensor, a protocol or a new control law can mean a new board rather, not just new firmware.
In industrial arms, with six or seven joints and a fixed task, that tax is affordable. In a humanoid, with 20 actuated degrees of freedom in each hand alone and a task that changes every few seconds, it compounds.
Won’t robots simply centralise, like cars?
The strongest objection to this argument comes from automotive. Vehicles are consolidating from many separate control units to a few central computers and zonal controllers. Robots, the argument goes, will follow, and the periphery will shrink to dumb drivers wired to a powerful brain.
The automotive precedent actually points the other way. Zonal architectures centralise decision-making, but they keep controllers physically close to the sensors and actuators they serve. Wiring, latency and electrical noise all forbid running every high-rate loop back to a central computer. What changes is the character of the edge node: fewer of them, each serving a region rather than a single function, each running more varied work.
Translated to a robot, a zone might be a hand, a forearm or a leg. That node must close several motor loops, aggregate dozens of sensors, keep time with the rest of the body and talk to central compute, all with bounded behaviour. Consolidation does not remove the deterministic edge. It raises the bar for it, from a fixed-function part per joint to a programmable, many-loop, time-aware controller per zone.
Determinism that scales
The XCORE® architecture was designed around the property this argument requires: a timing model that holds as the system grows, rather than one that holds only inside a single core.
On a single tile, each hardware thread is guaranteed a known share of the pipeline. Threads execute against tightly coupled single-cycle SRAM with no caches, and I/O is handled through hardware ports that respond to pin events and timed operations with cycle-level precision. The XS3 architecture makes each thread’s timing independent of what its neighbours are doing. That is the per-node bound.
The composition comes from how threads communicate. They use hardware channels, and a channel behaves the same way whether its two ends sit in the same core, on different tiles of one chip, or on different chips. Between tiles and devices, channels run over the xCONNECT switch and links, so multiple devices can be joined into a single system programmed with one model. The xcore.ai technical overview describes this, starting with two devices side by side on a board being programmed as one system.
For a robot, this matters more than any single benchmark. A fingertip aggregator, a hand controller and a wrist node can be separate devices placed where the wiring wants them, yet share one communication model with no shared bus and no locks between them. Crossing a link adds latency, but bounded latency, so the timing analysis for one node extends to the path across several.
Two honest limits apply. First, the guarantee covers the XCORE fabric; once traffic leaves it over Ethernet, determinism depends on the network, which is why time-sensitive networking matters (next section). Second, XCORE is not the robot’s brain. Policy, perception and planning belong on throughput-optimised compute. The point is integration: a bounded body that a probabilistic brain can rely on.
A track record at the edge of the body
The robotic edge needs three things done deterministically: sensing, actuation and communication. XMOS was shipping each of them on the same architecture, long before “physical AI” had a name.
Communication. XMOS built an Ethernet AVB endpoint providing time-synchronised, low-latency streaming over IEEE 802 networks, including a gPTP time server and media clock recovery. This was released as the industry’s first open-source AVB implementation, and continues in the AVB/TSN library today. AVB is the direct predecessor of time-sensitive networking (TSN). Shared time across nodes and reserved bandwidth are exactly what a distributed body needs.
Actuation. XMOS developed a multi-axis motor control platform for BLDC and PMSM motors, with dual-axis field-oriented control and inner loops up to 142 kHz. Partner Synapticon’s SOMANET motion cores combine xCore and Arm processors for deterministic real-time motion control in their ground-breaking servo drives built for robotic joints. The result is several motor loops on one device, with timing that can be analysed rather than tuned, already deployed in robot and cobot drives.
Sensing. XMOS’s sensing heritage began with sound, but it does not end there. Its voice processors capture far-field audio from multi-microphone arrays: Pollen Robotics’ Reachy Mini uses an array built on the XVF3800 for spatial hearing, and XCORE has been used for spatial audio recording with large-scale microphone arrays. XCORE-VISION brings the same architecture to imaging, with an 8 MP camera, capture and image signal processing, and on-device models such as YOLOv8 and MobileNetV2. Beyond audio and vision, interfaces such as I²C, SPI and PWM are implemented in software, so accelerometers, gyroscopes, encoders, force and tactile sensors connect directly to the ports that serve them.
The architecture is what turns this breadth into an advantage. Each sensor interface is a software task on its own hardware thread, not a fixed peripheral block, so one device can be configured for whatever mix of sensors a node needs. That can be an array of one kind, such as a microphone array, or a heterogeneous array: tactile pads, joint encoders, an IMU and a camera on a single wrist node. Each stream is captured on its own deterministic schedule against a common time base, so alignment across sensor types is a correct by construction. That is exactly the discipline tactile and proprioceptive fusion require.
The common thread is not a product category. It is a property: delivering the right result at the right instant, across many concurrent streams, on silicon whose timing can be reasoned about. Audio processing taught the company that lateness is wrongness, a robot’s body applies the same lesson with higher stakes.
Where the platform will be decided
Most of the attention in physical AI goes to the brain: the models, the accelerators, the central compute module. That is where one socket per robot is fiercely contested.
The body is a different market. It has many nodes per robot, not one, and its number of timing domains grows with every degree of freedom added. It is where today’s designs are most fragmented, and where the hardest guarantees a robot makes are actually kept. It is also where no deterministic platform yet spans sensing, actuation and communication end to end.
The companies that come to lead robotics will need a credible answer for that edge: a substrate on which the slip at the fingertip reaches the finger on time, every time, and can be shown to do so before the robot ships. Building that answer from first principles takes years. XMOS has spent two decades building it, one deterministic stream at a time.
Physical AI begins at the edge of the body. That is where the next frontier is.

Mark Lippett
CEO, XMOS
Sources
- XMOS, Physical AI Needs Two Kinds of Compute
- XMOS, Physical AI: First timing, then AI
- XMOS, The GenSoC Execution Substrate
- XMOS, The XMOS XS3 Architecture
- XMOS, XU316-1024 xcore.ai datasheet
- XMOS, xcore.ai technical overview
- XMOS, AVB Endpoint Design Guide
- Embedded.com, AVB software reference design goes open source
- XMOS, lib_tsn on GitHub
- Electronics Weekly, XMOS offers specialist motor control platform (March 2012)
- XMOS, sw_motor_control on GitHub
- Synapticon, SOMANET Motion Core
- Pollen Robotics / Hugging Face, Building Reachy Mini’s media stack
- Robots Guide, Shadow Hand; Clearpath, Shadow Dexterous Hand



