Heterogeneous Compute at the Edge

by Kevin Cloutier, September 14, 2026

Before you Begin

This research builds upon the previous work with the SO-101 robot and Vision Language Action. If you haven't read that, please do so before continuing as this post assumes a solid foundation in those systems: Vision Language Action: Model Development & Validation

Go... read... come back!

At the end of Vision Language Action: Model Development & Validation, I laid out an overview of what I wanted to achieve. At the time, I wasn't entirely sure where the research would take me, or what obstacles I would encounter along the way.

We left off having trained a model, evaluated it, and even ran it on a development machine. That first phase was fundamentally an R&D effort: part ideation, part technical landscaping, and part feasibility assessment to answer a simple question: is this even possible?

Now that we have proven a Vision Language Action model can control a robot, and it can control it well, we have a much better understanding of how the technology works. The questions become considerably more interesting: How do we turn this into a real-world system? What does integration look like? How do we deploy multiple AI capabilities together, across the compute resources available at the edge?

This work quickly became a story about why compute architecture matters for advanced systems. Software and software frameworks are important, they always will be, and I discuss several throughout this article. But the larger takeaway is that heterogeneous compute, exemplified by Intel's tiled SoC design, enables a fundamental shift in how one architects intelligent machines.

Traditionally, robotic systems have separated the “mind” from the “body”: control and execution on one side, computational intelligence on another. Heterogeneous architectures allow us to move beyond that separation, distributing workloads across multiple specialized processing resources and bringing intelligence, perception, decision-making, and control together within a more tightly integrated system using physically shared memory. The days of shuttling data from one processor to another are over.

Continue reading to learn about how a shift in computer engineering architectures is laying the groundwork for more intelligent systems and machines.

The System

I assume if you are reading this website you have a fairly strong engineering background, probably in computer science, electrical engineering, computer engineering, etc. If so, then you have likely come across the Tower of Hanoi puzzle and I bet you can recall when it all "just clicked". For this work I needed a problem that was simple enough to communicate, but complex enough to prove a heterogenous compute system without forcing a square peg into a round hole (see what I did there, peg, Tower of Hanoi?).

For the uninitiated, the classic Tower of Hanoi puzzle consists of three pegs and a stack of game pieces arranged by size: smallest on top, largest on bottom. The objective is to move the entire stack from one peg to another, following three simple rules:

  1. Only one game piece can be moved at a time.
  2. A larger game piece can never be placed on top of a smaller game piece.
  3. All game pieces must be moved from the source peg to the destination peg, using an auxiliary peg.

Despite these simple rules, the puzzle produces a sequence of dependent decisions and a myriad of possible solutions where the minimal amount of moves can be expressed as the function f = { 2n-1 }, where n is the number of game pieces. Therefore a tower with 3 game pieces has a minimum set of moves of 7 where (23 - 1) == 7. There is a killer explanation over at Khan Academy: Tower of Hanoi Explained, which is where I sourced the image below.

article image
The Tower of Hanoi, source: Khan Academy

The Tower of Hanoi is useful here because the rules of the puzzle don't require a machine-learning model, they can be represented deterministically in software. Once the current state of the board is known, conventional game logic can determine the next legal move. Additionally, I can easily decouple the game state; next-move generation; and controlling the robotic arm; thus allowing me to make this an extremely interesting example of a system of systems. Who would have thought? Well, I did!

Physical Limitations

As described above and shown in the image from Khan Academy, the classic Tower of Hanoi puzzle utilizes discs that are placed on 3 pegs. It resembles a child's toy with its simplicity, but if you read Vision Language Action: Model Development & Validation then you know that this research uses a less-than-precise robotic arm, the SO-101. As we have seen, this robot is a fantastic learning platform but it's dexterity is, well, it's just awful. Which in turn means grasping the traditional discs is really out of the question.

I needed a way to physically solve the puzzle without having to grab discs from a peg. After days of riffing on ideas it dawned on me that I could combine the peg and game piece, making manipulation orders of magnitude easier. See the images below:

article image
The game piece top (left) and bottom (right). Note I have built a peg into the top of the piece and a recessed area for the game piece to be placed on another game piece.

In this design, the robot can pick a game piece by grasping the "peg" rather than a disc threaded onto a peg. The SO-101 then moves the piece to the correct location and drops it onto the peg extending from the piece below it. Incorporating the peg into the game piece in this manner retained the mechanics of the game, while allowing for less-than-accurate robotic manipulators to be utilized (and thus reducing the overall cost of this line of research). A short video will make this clear:

A short video depicting the game pieces and the first move of the Tower of Hanoi puzzle.

Heterogeneous Architecture

Now that we understand the Tower of Hanoi problem, we can decompose the system into a set of individual workloads and begin mapping those workloads onto the compute resources available in the target system: the Intel Core Ultra X7 358H. This processor is an interesting platform for this experiment because it brings several fundamentally different compute engines together in a single package: 16 CPU cores consisting of 4 Performance-cores, 8 Efficient-cores, and 4 Low Power Efficient-cores; an integrated Intel Arc B390 GPU with 12 Xe-cores; and a dedicated NPU capable of up to 50 INT8 TOPS.

This is important because these processors are not interchangeable. The CPU is well suited to general-purpose computation, orchestration, control, and the many sequential and latency-sensitive tasks that make up a robotic system. The GPU provides substantial parallel compute for workloads such as vision and large neural network inference, while the NPU is purpose-built for efficient AI workloads that can run independently of the CPU and GPU. Intel explicitly positions the integrated GPU and NPU as resources for running edge AI alongside the work already being performed by the CPU.

The architectural challenge, then, isn't simply to make the models run. It is to understand the workloads that make up the complete system and assign each one to the compute resource where it makes the most sense. The Tower of Hanoi is a useful vehicle for exploring exactly that problem: how to turn a collection of AI models, computer vision, control logic, and software services into a coordinated system that makes effective use of heterogeneous compute.

Central Processing Unit

The CPU is the general-purpose processor and acts as the system coordinator. Unlike the GPU and NPU, the CPU is designed to execute arbitrary sequences of instructions, including branches, loops, system calls, I/O operations, and application logic. This makes it particularly well suited to tasks that require decisions, coordination, and interaction with the rest of the system.

Integrated Graphics Processing Unit

The integrated graphics processing unit (iGPU) is a highly parallel processor designed to perform computations across large amounts of data simultaneously. Unlike the CPU, which is optimized for general-purpose sequential and control-oriented workloads, the GPU contains many parallel execution resources that can perform the same or similar operations on many pieces of data at the same time.

A standard discrete GPU (dGPU) is typically installed on a PCIe expansion card, containing its own dedicated high-bandwidth memory. For example, an Intel Arc Pro B70 contains its own GPU and 32 GB of DDR6 memory. The dGPU therefore has a dedicated memory subsystem that is physically separate from the system memory used by the CPU, this requires moving data from one system to the other.

By contrast, an integrated GPU (iGPU) is incorporated into the processor, or system-on-chip, and does not normally have a separate pool of dedicated graphics memory. The iGPU instead shares the system's main memory with the CPU and other components of the processor. This avoids the need for a separate GPU memory subsystem and allows the iGPU and CPU to operate on data within the same system-memory environment.

On the other hand, a discrete GPU can provide substantially greater memory bandwidth and dedicated memory capacity, making it well suited to very large, compute-intensive models.

Neural Processing Unit

The Neural Processing Unit (NPU) is a specialized accelerator designed specifically for neural-network workloads. An NPU sacrifices much of the general-purpose flexibility of a CPU, and some of the programmability of a GPU, in exchange for highly efficient execution of the mathematical operations commonly found in neural networks. This allows neural-network inference to be performed with significantly greater power efficiency than would typically be possible using a general-purpose processor.

Mapping the Problem to the Processors

With the characteristics of each processor established, we can now return to the Tower of Hanoi and examine the individual problems that make up the system. The complete operation can be thought of as a pipeline:

  1. A set of cameras observes the board
  2. The system determines the current state
  3. Game logic determines the next legal move
  4. The system describes the state to the user
  5. A neural-network policy translates the next move into robot actions
  6. The robot executes the movement.

Although these stages are part of a single application, they do not perform the same type of computation. Some require general-purpose program execution and decision-making. Others involve large amounts of mathematical computation that can be performed in parallel. Still others are specifically neural-network workloads.

1. A set of cameras observes the board

A camera itself is not a processor workload in the conventional sense. It produces a stream of image data that must be captured and made available to the rest of the system. The CPU can manage the camera interface and the movement of image data through the application. Once an image has been captured, however, the computational work required to interpret it is better suited to an accelerator.

2. The system determines the current state

Determining the state of the physical board is a classic computer-vision problem, but it does not require a neural network. The system uses conventional computer-vision techniques to analyze the camera image and determine the position of the game pieces. This workload can be processed on the CPU. Unlike similar systems, I have made a conscious decision to not utilize a vision language model here as our board has a known physical structure and the pieces occupy known positions, allowing the state to be determined deterministically from the camera image.

3. Game logic determines the next legal move

Once the state of the board is known, determining the next move is fundamentally different from recognizing the board. The rules of the Tower of Hanoi are deterministic. Given the current state, the application can calculate which move should occur next using conventional program logic. This involves comparisons, conditional branches, data structures, and state management and thus is exactly the type of workload for which the CPU is designed.

4. The system describes the state to the user

A small language model can provide a natural-language interface to the system, allowing the system to describe what it sees and communicate the state of the game to the user. This is a neural-network workload and is therefore well suited to execution on the NPU.

There is a key consideration here, the frequency and timing requirements of this inference. The SLM is not part of the robot's real-time control loop. There is no need to run it every 200 milliseconds, or even have it's inference chain come below the Doherty Threshold. The system can invoke the SLM approximately before each game move, when an updated natural-language description is useful to the user. This makes the NPU a good fit, as its dedicated neural-network hardware can perform inference efficiently (~50ms throughput) without occupying the GPU. Therefore, the GPU remains 100% available for a VLA workload, which has much more demanding timing requirements.

5. The neural-network policy translates the move into robot actions

The game logic can produce a discrete instruction such as: Move the game piece from A3 to C1. The robot, however, cannot execute that instruction directly. It requires a continuous sequence of joint positions and gripper movements that describe how to physically perform the action. This is the role of the Vision-Language-Action (VLA) policy.

Like the SLM, the VLA is a neural network and therefore contains substantial amounts of tensor and matrix computation. However, the characteristics of this particular workload make the integrated GPU a better target than the NPU. The GPU provides a large number of programmable parallel execution resources and is capable of efficiently executing the tensor operations required by the vision model.

6. The robot executes the movement

Finally, the CPU coordinates communication with the robotic arm and sends the generated actions to the robot. The actual physical movement is performed by the robot's motors and controllers, not by any of the processors in the Core Ultra system. The CPU manages this interaction, including timing, communication, and monitoring of the robot's execution.

A Deeper Look

Let's dive one level deeper into each major component.

article image
The system architecture depicting the CPU, iGPU, and NPU

Vision Language Action

The VLA is the component that determines how the robot moves. Recall our system design from above. I am not asking the VLA to solve the Tower of Hanoi, understand the rules, or even determine the state of the board. Those problems are being handled elsewhere. The VLA has a much simpler job: given a natural-language instruction describing a single legal move, determine the robot moves required to execute it. For example, the system may determine that the next move is: Move the game piece from A3 to C1.

That instruction becomes the task presented to the VLA. It receives the camera observations and the instruction, and generates the sequence of robot actions required to locate the piece, grasp it, move it to the destination, and release it.

Similar to the Vision Language Action: Model Development & Validation research, I trained the SmolVLA policy specifically for this type of pick-and-place operation using demonstrations with the SO-101. The training data focused on the physical manipulation task rather than the game itself. In other words, I am teaching the model how to perform the movement, not which movement it should choose. This keeps the policy focused on a relatively constrained problem and allows the rest of the system to handle the higher-level reasoning.

At runtime, the trained policy will run on the integrated GPU. The CPU provides the observations (camera images and servo states) along with a natural language prompt as input. In turn, the VLA generates its action sequence and the CPU feeds those actions to the robot. Once the movement is complete, the system can observe the board again and determine whether the expected state transition actually occurred.

Why "Move the game piece from A3 to C1"?

Several times I have used the example "Move the game piece from A3 to C1". This is no accident as there is a bit of intentional design hiding in that sentence. The A3 and C1 aren't arbitrary names; they describe specific physical locations on the board that were originally represented by pegs. The three locations in this implementation are named A, B, and C. The available positions at each location are numbered from the bottom up: 1 is the bottom position, 2 is the middle position, and 3 is the top position. That gives us a simple coordinate system for describing the nine possible locations on the board: A1, A2, A3, B1, B2, B3, C1, C2, and C3.

So when the system says "Move the game piece from A3 to C1", it isn't identifying a particular game piece. It is describing a transformation between two physical locations: take whatever game piece is currently at position 3 on peg A and move it to position 1 on peg C. Again, the VLA does not know, or care about, the rules of the Tower of Hanoi puzzle. I don't want the model learning that a particular color or particular game piece belongs at a particular location. The pieces are interchangeable from the perspective of the manipulation policy. What matters is where the piece is and where it needs to go.

A benefit of this design is it significantly reduces the training space. Instead of training separate behaviors based on the identity, color, or configuration of individual pieces, the policy learns the physical task of moving a piece from one location to another. The same learned behavior can therefore be applied regardless of which piece happens to occupy that location.

Small Language Model

The Small Language Model has a very different job in this system. Its job is to describe the board in a conversational way to the end user. That sounds simple, but it provides a useful interface between the physical world and the rest of the system. The SLM can describe the current arrangement of the game pieces, identify what is on the board, answer questions about the pieces, and provide a natural-language description of the current scene. This capability allows the system to be a bit more interactive where one can questions about what the camera is seeing such as:

  • Which pieces are on peg A?
  • What is currently at position A3?
  • Describe the current state of the board.

The SLM generates the answers from the input and essentially provides a natural-language interpretation of the physical environment.

Computer Vision and Game Logic

This is where conventional software comes back into the picture. While the SLM is responsible for describing the board and providing a natural-language interface to the physical environment, a conventional computer vision algorithm can determine the actual game state. A cv system can consume camera images and determine where the game pieces are located on the board. It can then maps their locations into the coordinate system used by the rest of the application: A1, A2, A3, B1, B2, B3, C1, C2, and C3.

One challenge with this is we can't simply look at a single frame and assume that it represents the current state of the board. The robot may still be moving a piece and the camera may capture the board while something is in motion. Therefore it is necessary to add a debouncing routine that waits for the board to reach a steady state before accepting the detected positions as the current game state.

Once the board is stable, the detected piece locations can be mapped into an array representing the current board state. The game logic then compares that state against the known set of valid moves and determines what should happen next. The result is a simple source and destination instruction, such as Move the game piece from A3 to C1. This instruction is passed to the VLA, which generates the robot actions required to execute the move.

Developing with Heterogeneous Compute

At this point, let's fast forward a bit... I have described the architecture above, and we already discussed training SmolVLA in Vision Language Action: Model Development & Validation. So, let's assume we have a Vision Language Action policy (SmolVLA) trained on the NVIDIA CUDA stack and that we have selected the underlying Small Language Model (Qwen 3 4B IT) for scene descriptions. This is where the computer engineering fun truly begins...

Targeting the Integrated GPU with a Vision Language Action Model

I have a trained VLA model that works great, but there is just one problem: I didn't train it on the Intel SoC where I want to run it. The model was trained on a NVIDIA DGX Spark inside of the CUDA ecosystem. That's a great environment for model development and training, and it's the environment most AI development is currently executed, but it isn't the target environment for my system. The goal of this experiment is to run the complete Physical AI system at the edge on an Intel Core Ultra 3 platform, using the CPU, integrated GPU, and NPU as described above. Therefore, we needed to figure out how to deploy this work on a different hardware stack.

This is where OpenVINO becomes important. OpenVINO is a free, open-source software toolkit by Intel used to optimize and deploy deep learning and AI models on multiple stacks. This framework allows one to take a model developed and trained in the NVIDIA ecosystem and transform it into something that can be deployed efficiently on Intel hardware, including the heterogenous computer of the Core Ultra 3 SoC. The key is the model itself doesn't change its job, it is still the same VLA policy, but the way that model is represented and executed has to change to match the target hardware.

If you are familiar with embedded systems, this is cross-compilation: where we develop the model is different from where we where we deploy it. In reality, the hardware used to train policies and the hardware used to run AI inference have very different requirements. I don't want my deployed system to require a physically large, hot, and costly, discrete GPU containing multiple memory systems simply because that is where the model was trained. I want to train the policy where it is efficient to do so, then deploy where the application actually needs to run... at the edge, 100% locally.

Surveying what I actually need at runtime made it clear that the entire LeRobot framework was unnecessary, and I certainly did not need any of the training code. At runtime, the VLA has one job: take the camera observations, the language instruction, and the current robot state, then generate the next chunk of robot actions. The challenge is to identify exactly where that operation occurs within the much larger LeRobot and SmolVLA implementation so it can be isolated, allowing OpenVINO to export the model to the target hardware.

To do this, I started at the LeRobot policy level and traced the inference path down through the SmolVLA implementation rather than trying to guess which part of the framework was required. At this point in time, this isn't as simple as clicking a few buttons, it requires some digging and understanding of the execution paths. This exploration led to the VLAFlowMatching model in modeling_smolvla.py, where I found the sample_actions() method. Unlike the surrounding policy code, this method contains the actual action-generation process. It takes the images, language tokens, masks, robot state, and noise, then performs the prefix embedding and VLA processing, runs the flow-matching denoising loop, and finally returns the action chunk used to control the robot.

That was the boundary I was looking for. The code above sample_actions() is concerned with preparing inputs and managing the policy, while the code inside sample_actions() performs the computation that actually turns those inputs into robot actions. I therefore did not need to reproduce the LeRobot policy framework; I only needed to make this existing inference operation available to the OpenVINO conversion process.

But there was a complication. sample_actions() is already part of the loaded SmolVLA model, but it is not the model's standard forward() entry point that OpenVINO requires. PyTorch modules are normally invoked through forward(), and the PyTorch-to-OpenVINO conversion process uses that module interface to determine the computation it needs to convert. Simply giving the converter the existing model would therefore expose its normal forward() path, not the specific sample_actions() path I had identified as the runtime operation I wanted to deploy.

The solution was to create a small adapter class. OpenVINO requires an object derived from torch.nn.Module, so I extended that class and implemented forward() to simply call the existing model's sample_actions() method. The adapter does not create another SmolVLA model or contain another set of weights. It simply holds a reference to the model loaded from the checkpoint and redirects the module's standard forward() call to the existing action-generation method. A snippet from the class I wrote can be seen below:

    
class SmolVLAInferenceWrapper(torch.nn.Module):

    def __init__(self, policy):
        super().__init__()
        self.model = policy.model

    def forward(
        self,
        images,
        img_masks,
        lang_tokens,
        lang_masks,
        state,
        noise,
    ):
        return self.model.sample_actions(
            images,
            img_masks,
            lang_tokens,
            lang_masks,
            state,
            noise=noise,
        )
    

The forward() method is important because OpenVINO is not converting the Python class itself; it uses the method as the entry point to trace the computation that needs to be converted. When the OpenVINO converter invokes forward(), the call is now passed directly into SmolVLA's existing sample_actions() implementation. OpenVINO can then follow the operations performed by that method, including the model's embedding, VLA processing, and flow-matching denoising operations, and represent them as an OpenVINO computation graph. Once converted, that graph can be serialized as an OpenVINO model and executed without the original LeRobot policy machinery. The forward() method therefore isn't part of the VLA inference algorithm; it provides the entry point OpenVINO needs to identify and capture the inference computation I want to deploy.

With the inference boundary isolated, the next step was to convert that PyTorch computation into a form that could run natively through OpenVINO. This is where the ov.convert_model() function comes in:

    
ov_model = ov.convert_model(
    model,
    example_input=example_inputs,
)
    

Creating Example Input

At this point one obvious question is "how did he define the example inputs to provide the converter?". I did not need real camera images or a real robot state to perform the conversion. OpenVINO only needs representative tensors that match the interface of the computation. The starting point for that is the signature of sample_actions() itself. Its arguments define the six inputs required by the inference path: images, image masks, language tokens, language masks, robot state, and the initial noise used by the flow-matching process.

I then traced each of those inputs back through the SmolVLA implementation to determine their expected shapes and data types. The model configuration provided values such as the action chunk size, while the policy and VLA implementation established the image resolution, number of camera inputs, language sequence length, and padded state and action dimensions. This gave me the tensor interface that the exported model needed to accept. With those dimensions established, I just created representative PyTorch tensors for each input

Somewhat surprisingly, these values do not need to represent a real robot observation. Their purpose is to just establish a valid example of the model's input interface. Random values are sufficient for tensors such as the image and noise inputs because the converter is interested in the computation performed by the model, not the semantic meaning of the particular data used during conversion. The masks and token tensors still need the correct data types and dimensions because those properties determine how the model's operations are constructed.

Lastly, the tensors are assembled in exactly the same order as the arguments to the original forward() function.

Here, model is the small adapter class I created above, not another copy of SmolVLA. OpenVINO invokes its forward() method using the supplied example inputs. That call immediately enters the existing sample_actions() implementation, allowing OpenVINO to capture the computation that produces the action chunk. The example inputs provide the converter with representative tensors from which it can determine the inputs, shapes, data types, and operations involved in the computation.

OpenVINO then constructs its own representation of that computation as an OpenVINO model graph. The trained SmolVLA parameters are carried into this representation along with the operations that use them. What began as a PyTorch model executing Python and PyTorch operations is therefore transformed into an OpenVINO graph that describes the neural-network computation independently of the original LeRobot policy runtime that can be executed on Intel hardware.

Once the conversion is complete, the resulting OpenVINO model can be serialized to disk:

    
ov.save_model(
    ov_model,
    "smolvla.xml",
)
    

The above produces the OpenVINO model files, with the XML file describing the graph and the associated binary file containing the model parameters. At this point, the original LeRobot policy object is no longer required to execute the converted inference graph. The Intel deployment runtime can load the OpenVINO model directly and provide it with the same inputs: camera observations, language tokens, masks, robot state, and noise—to produce the action chunk.

In effect, the conversion process takes the inference computation I identified in the SmolVLA source code, captures it through the adapter's forward() interface, and turns it into a self-contained OpenVINO model that can be deployed on the target hardware.

What about SmolVLA on the NPU?

With the iGPU working, I then started targeting the NPU. Unfortunately, the same OpenVINO graph could not be compiled for the NPU without additional changes, so the deployment process became considerably more interesting. Thankfully, the compiler errors told me where the problems were, and inspecting the generated graph told me what was actually happening at those points in the computation. Add in some Vibe Coding, and I worked through those issues one at a time until I had the VLA running on the NPU as well as the iGPU. The underlying application code I wrote was designed for this very scenario. Change a few parameters in a config file, and voilà, the model moves to a different piece of hardware (e.g. to the Intel NPU instead of the iGPU, NVIDIA CUDA, etc.)

Interestingly to me, getting the VLA running on the NPU didn't make it the better choice. When I compared the inference performance, the NPU was taking roughly six times as long (1.5s) to generate an action chunk with this particular model as the iGPU (250ms). That's just not responsive enough for this use case, and though I could go down the route of compression, etc. I truly had a different model targeting the NPU anyway. So, while the NPU could run the model, my workload wasn't a good fit for that particular hardware device.

Targeting the NPU with a Small Language Action Model

As described above, the Small Language Model is not determining the state of the Tower of Hanoi game. It is not deciding which move is legal, or controlling the robot. Those responsibilities remain with the deterministic application logic and computer vision system running on the CPU (see further below for info on the CPU workloads). The SLM receives an already-established board state and turns that structured information into a sentence that a person watching the demonstration can understand.

A language model is useful for generating flexible, natural language, but it is not the right component to serve as the authoritative source of a deterministic physical game's state.

Another OpenVINO Inference Graph

Qwen 3 4B was selected for the SLM at somewhat random. I was searching for a model that was lightweight and previously known to work on an Intel NPU, hoping to leverage the engineering work of others for this workload! After a short internet search I settled on Qwen 3. Unfortunately, once I got into the weeds I found this model, though it may work on some NPUs, mine was not one of them. From what I could tell, the combination of my device hardware and the OpenVINO framework version I had standardized on earlier ion this project, may not work well with this model.

With so much work completed already, I was not about to start updating / upgrading libraries. So, I had one of two choices, find another model, or follow a similar approach to exporting as I described with the VLA. I opted for the latter as I felt I would have more control over the process and output. I'm happy I did.

Exporting Qwen 3 4B

Similar to SmolVLA, I had downloaded the Qwen model via the Hugging Face ecosystem. This was a PyTorch model supporting CUDA and though very convenient during development, it was not the representation I needed for edge deployment on the Intel SoC. Now the goal was once again to remove the dependency on the original Python model execution path and produce one that could be loaded directly by OpenVINO for assignment to the NPU.

The first step was to identify the actual computation that needed to be exported. The full transformer model includes functionality for model loading, generation, caching, configuration, tokenization, and other framework-level operations. Most of that is not part of the neural-network computation that I need on the NPU. For this research, the core operation I needed was only the model's forward pass: given token IDs, an attention mask, and position IDs, calculate the output logits. This led to another adapter class:

    
class Qwen3InferenceWrapper(torch.nn.Module):

    def __init__(self, model):
        super().__init__()
        self.model = model

    def forward(
        self,
        input_ids,
        attention_mask,
        position_ids,
    ):
        outputs = self.model(
            input_ids=input_ids,
            attention_mask=attention_mask,
            position_ids=position_ids,
            use_cache=False,
        )

        return outputs.logits
    

As was the case with the VLA export, the wrapper doesn't create a new language model and does not contain a second set of weights. It holds a reference to the original Qwen3 model and exposes the specific forward-pass interface that I want to convert. By defining the forward() method explicitly, I could control the inputs presented to the exporter and ensure that the exported graph represented the actual inference operation required by the deployed application. The text-generation loop itself remains outside this exported forward graph. The system application running on the CPU is responsible for preparing the inputs, invoking the model, selecting the next token, and repeating the process as needed.

Once the inference wrapper was defined, I again used OpenVINO's conversion functionality to transform the PyTorch computation into an OpenVINO model representation. Nothing new here, OpenVINO invokes the model's forward() method using the supplied example inputs, follows the operations performed by the underlying Qwen3 model, and constructs an OpenVINO computation graph representing that operation.

    
ov_model = ov.convert_model(
    model,
    example_input=example_inputs,
)
    

For the initial export, I selected a fixed sequence length of 512 tokens. A sequence length is the number of token positions the model is configured to process in one inference operation, this includes the system prompt, user prompt, everything. 512 provided a large capacity while I was establishing the export and deployment process, avoiding the risk of making the graph too restrictive before I fully understood what I was dealing with.

My export utilized a static input length so that OpenVINO knows the exact tensor dimensions it needs to compile and optimize for the target hardware. That gives the compiler a much more constrained computation graph than one supporting arbitrary sequence lengths, often making the model smaller. In the end, the export resulted in an XML file describing the computation graph and an associated binary file containing the model parameters. At this point, the model exporting from the PyTorch representation into an OpenVINO intermediate representation was complete.

Sequence Length Testing

With the model exported, the next question was whether 512 tokens was actually practical for this application. I would have preferred to keep the 512-token sequence length because it provided substantial headroom for the model's context. However, the measured inference latency on the NPU was too high for the workload, exceeding 110 ms per inference with a throughout of ~4.5 tokens/s. Even though this processing is not part of the control system, that felt too long. Well under that Doherty Threshold, but we could do better. So I went about testing multiple models:

Sequence Average latency Throughput
64 ~30.5 ms ~32.75 tokens/s
128 ~51.2 ms ~19.53 tokens/s
256 ~95.7 ms ~10.45 tokens/s
512 >110 ms ~4.5 tokens/s

At first glance, the 64-token model was the obvious winner. With an average inference time of only ~30.5 ms, it was more than three times faster than the original 512-token configuration. However, the 64-token sequence was simply too short for the application. The system prompt and user prompt together left insufficient room for the model to generate the response I wanted.

I therefore moved to the 128-token configuration. This provided enough capacity after I reduced the system prompt to its absolute minimum. Even with an extremely compact prompt, I was able to generate useful, natural-language responses describing the board. At approximately 51.2 ms per inference, the 128-token configuration provided a substantial improvement over the original 512-token model while still giving the application enough context to produce useful output.

For this use case that was the balance I was looking for. The SLM is not part of the robot's real-time control loop, so I did not need to optimize for the absolute lowest possible latency. I needed enough context for the prompts and response, while keeping the inference time low enough that the generated description felt responsive to someone watching the demonstration.

Onto the final device....

Targeting the CPU with the Control System and Orchestration

The CPU is responsible for coordinating the entire system. While the integrated GPU and NPU execute the neural-network workloads, the CPU remains the central orchestration layer that connects those workloads to the physical robot and maintains the state of the demonstration. This includes configuration, initialization, hardware discovery, computer vision, game-state management, prompt generation, robot control, event handling, logging, and thread coordination.

I kept these responsibilities on the CPU because they are primarily control-flow and application-level operations, not large numerical workloads. This is where a general-purpose central processing unit excels. These processes require flexibility, access to the operating system, interaction with the robot and cameras, and the ability to coordinate several activities that operate at different rates. Moving these tasks to an accelerator would provide little benefit and make the system more complicated.

The Control Loop Worker Thread

The control loop is the most time-sensitive part of the CPU application. It owns the action queue and continuously removes the next available action from that queue for execution by the SO-101 robot. The queue contains decoded joint commands produced by the VLA, allowing the robot to continue moving while the next inference is running in a separate thread.

The VLA Worker Thread

The VLA thread waits until the control loop signals that more actions are required. Depending on which hardware is targeted for the VLA (e.g. dGPU, iGPU, NPU), the rate of inference can vary widely. Therefore, a configurable replan_threshold argument allows different hardware to compensate for different latencies.

When notification arrives to run the inference, images are saved of the current state and the VLA is called. All the while, the control loop is eating through the remaining actions and moving the robot. The goal is to set a replan_threshold large enough to ensure the queue is never starved, but also as minimal as possible. This is very hardware- and workload-specific, so tinkering is necessary.

The actual hard part is when the results return from the VLA, they must be interwoven back into the queue, knowing full well that the items at the start of the latest results may have already been sent to the robot while the inference was happening. This needs to be accounted for, and the CPU is again the right place for that work.

Computer Vision and Authoritative Board State Worker Thread

Computer vision also runs as part of the CPU orchestration layer. Each VLA observation includes the camera images required by the model, but the world-facing image is also passed to a deterministic board state manager. That component identifies the pieces and their positions on the pegs, applies the relevant state-processing logic, and maintains the application's authoritative representation of the board.

Why the CPU Matters in This Architecture

Although the GPU and NPU receive the attention because they execute the neural-network models, the CPU is what makes the complete system function as an application that can drive physical hardware. It connects the cameras, computer vision, game-state representation, language prompts, inference backends, action queue, robot driver, user controls, and logging into one coordinated system.

This is the practical meaning of heterogeneous compute in this research. The accelerators are not competing to run the entire application. Each is assigned the workload for which it is best suited: the integrated GPU performs VLA inference, the NPU performs the lightweight language-generation workload, and the CPU manages the deterministic logic, concurrency, device interaction, and real-time orchestration that ties everything together.

Live System Visualization

At this point, the models are running and the different pieces of the system distributed across the Intel hardware, but there was still a problem... most of what was happening was invisible. The CPU was orchestrating the system, the NPU was interpreting the scene with natural language, the iGPU was generating robot actions, and the SO-101 was physically executing those actions, but unless you were looking at the source code or the logs, one couldn't really see how all of those pieces were working together. This system was doing a very poor job of communicating to the world the advancements supported by heterogeneous compute!

For demonstration purposes, I wanted everyone to see the system, to visualize what was going on under the hood. To that end, I built a small web-based interface that provides a live view of the system while the robot is running. I wanted this interface to provoke questions during demonstrations and, because things always go wrong in demos, provide a chance for me to explain what happened using the system's actual output while showing how the robot responded. The simple UI seen below did the trick.

article image
The user interface I built for demonstration at AI Infra Summit 2026

What's Next?

So... where does this leave us?

I started this experiment with a fairly simple question: can I take a Vision Language Action model that was developed and trained on a NVIDIA CUDA system and put it to work as part of a Physical AI deployment at the edge targeting Intel Core Ultra 3?

The answer is a resounding yes.

If you made it this far, you may be surprised to read that I'm not particularly interested in building a robot that can solve a common puzzle (e.g. the Tower of Hanoi). I'm interested in understanding what it takes to build a Physical AI system that can see, reason, and act on an edge device, completely at the edge.

Now the real fun begins, bringing this architecture to the market.

Onward...

Notes on VLA Training Data

I'm sure you have heard the mantra "you need clean training data for imitation learning" over and over again. This is true, and it is extremely important, but sometimes there is more at play than just clean data.

Take the third move in one of the minimal solutions for the Tower of Hanoi, C3 to B2. This picks the smallest game piece from the farthest position on the board relative to the camera, and moves it to the middle position where it is placed on top of the green game piece already there. I went to great lengths to record this movement without interfering with the world camera. I even used two hands on the robot arm to steady it during the demonstrations, and after reviewing the data and associated videos, all looked good.

After training, however, the policy missed this move more than 90% of the time. It almost never successfully grasped the game piece. My first thought was simple: I need more data. So I threw out the original dataset (in case it was bad) for that one move and recorded 100 new episodes, twice as many as I originally recorded. I then retrained the policy, tested it again, and... zero improvement. I think it was actually worse, but that's subjective as I didn't properly record the stats during evaluation (just pen and paper in a notebook).

I went back and watched the demonstration videos carefully. Nothing obvious jumped out at me. The demonstrations looked clean, the robot was stable, and the movement looked consistent. I started wondering if the position of the world camera was making depth estimation difficult at that particular angle, especially since this was the smallest piece and the one farthest from the camera. Then I had another thought: what if the robot orientation I trained for this move was simply too difficult for the policy?

It was a bit counterintuitive, but I threw out the 100 new episodes and recorded another 50. This time, I approached the game piece from what I hoped was an easier-to-manipulate angle, even though that meant slightly occluding the world camera during the grasp. And guess what? It worked. The policy successfully performed the movement more than 95% of the time.

So, in the end, some of imitation learning really isn't an exact science, or at least it doesn't feel like one from this side of the mathematics. Sometimes the answer isn't more data. Sometimes the answer is changing what you are asking the model to learn.

Below is a quick video walking through this, and a few other tips that may help.

A short video depicting some training issues and how to overcome them.

Follow me on the socials!