What goes into running a fine-tuned π0.5 policy on a Franka robot

Draft — intro goes here. One or two paragraphs on what this post is about, why running a fine-tuned π0.5 policy1 on a real Franka arm is harder than the paper makes it look, and who should read it.

1. Setup

This section describes our robot setup. The layout is shown in Figure 1.

Robot setupA Franka arm is connected to its control PC. The Franka Control Interface (FCI) links the control PC to a NUC running polymetis. A GPU machine reads the robot state from the NUC through polymetis and images from two RealSense D435 cameras and one D405 camera, then sends the observation, language command and robot state to a server running pi0.5. The server denoises a random action trajectory into a smooth action chunk and sends it back; it flows through the NUC and FCI to the arm. D435 D435 D405 FRANKA ARM images polymetis obs · prompt · state action chunk LANGUAGE COMMAND › "move the cube to the right" FCI Control PC FRANKA CONTROL NUC RUNS POLYMETIS GPU machine CAMERAS · STATE Server machine RUNS π0.5 DENOISING STEP 0/10 x z grip
  1. 01Sense
  2. 02Request
  3. 03Infer
  4. 04Respond
  5. 05Act

Figure 1. The robot setup and one control cycle. Robot state and images from two RealSense D435 cameras and one D405 reach the GPU machine, which queries the π0.5 server. The server denoises random noise into a smooth action chunk, which is streamed back to the arm.

The Franka arm is connected to its control PC. The Franka Control Interface (FCI)2 sits between the control PC and a NUC, which talks to the robot through polymetis3. The workstation also has three Intel RealSense cameras (two D435s and one D405), which are connected to a separate GPU machine. The GPU machine collects the camera observations and the robot state, which it gets from the NUC over polymetis.

At each inference step, the GPU machine sends the current observation, the language command and the robot state to a server machine running π0.5. The server replies with an action chunk, which the GPU machine sends back through polymetis and the FCI to the arm.

Placeholder: task and evaluation protocol.

2. Data collection

To collect demonstrations, we teleoperate the arm with a Meta Quest 2 VR headset, using the DROID4 teleoperation stack and its Oculus controller5. Figure 2 shows how it works.

Teleoperation with a Meta Quest 2An operator holds a Meta Quest 2 controller. The headset sits on a stand in front of them and tracks the controller. While the grip button is held, the controller's displacement from where the grip was pressed is applied to the Franka's end-effector, relative to where the robot was at that moment, so the end-effector copies the hand's motion. The trigger closes the gripper. OPERATOR FRANKA ARM META QUEST 2 same Δp pee = pee,0 + Δp Δp Δp
  1. 01Hold grip
  2. 02Move
  3. 03Trigger
  4. 04Release

Figure 2. Teleoperation with a Meta Quest 2, following DROID's VR controller. While the grip is held, the controller's displacement Δp since the grip was pressed is added to the end-effector pose the robot had at that moment.

The headset sits in front of the operator and tracks the hand-held controller. Holding the grip button enables movement and records two origins: the controller pose and the robot’s end-effector pose at that instant. From then on, the controller’s displacement from its origin is added to the robot’s origin pose, so the end-effector copies the motion of the operator’s hand. Orientation works the same way. The trigger opens and closes the gripper. Releasing the grip stops the robot following, and pressing it again sets new origins, so the operator can reposition their hand without moving the robot. The A and B buttons mark an episode as a success or a failure.

One decision that matters in demo collection is the frequency you record at. In practice it depends on many things: your hardware, your task, your inference server, your network delays, and so on. The first place it hits you is dataset size.

A standard VLA collection setup like ours has three RealSense cameras: two D435s and a D405. Out of the box, over USB 3, librealsense6 streams the D435’s color camera at 640×480 and the D405’s at 848×480, both at 30 fps. Each recorded step stores one frame from every camera, so the dataset grows linearly with the recording frequency, the length of each demo and the number of demos. For reference, DROID records at 15 Hz7. Figure 3 lets you try different settings.

D435 × 2

D405 × 1

One second of camera frames and the frames that get recordedCAMERA RECORDED 0 s 1 s

Dataset size

Native resolution
Resized to 224×224

Figure 3. How the recording frequency drives dataset size. Sizes are for uncompressed RGB frames from all three cameras; robot state and actions add only a few hundred bytes per step. Video encoding shrinks the stored size a lot, but every frame is still decoded to this size when you train. π0.5 resizes each image to 224×224.

Placeholder: tasks, number of demonstrations, and episode length.

3. Deployment on the Franka

One thing that matters at deployment is matching the control frequency to the frequency of the data. Most of our demos are recorded at somewhere between 8 and 15 Hz, mainly because we don’t have much storage.

The first question is how fast you can send a command to the robot at all. This is where the NUC from Figure 1 comes in. The Franka expects a new command every millisecond: the FCI runs a 1 kHz control loop, and the round trip plus the controller’s own computation has to fit inside that 1 ms, or the robot stops with a communication error. A stock Linux kernel can’t promise that, so the loop runs on a NUC running Ubuntu with a real-time (PREEMPT_RT) kernel8. The GPU machine never talks to the arm directly. It sends targets to polymetis on the NUC, and the NUC keeps up the 1 kHz loop on its own.

In our setup, one env.step(action) from the GPU machine takes about 3.4 ms on average, so the fastest we could command the robot is roughly 290 Hz. Recording at that rate isn’t realistic. With the three cameras at their defaults (Figure 3), a single one-minute demo would be about 54 GB, and the cameras only produce 30 new frames a second anyway, so most of those frames would be repeats. At 8 Hz the same minute is about 1.5 GB.

So we slow the loop down on purpose. To run at a target frequency, say 8 Hz to match the data, each step spends its ~3.4 ms sending the action and then sleeps for the rest of the period:

period = 1 / 8                  # s, the data frequency
while True:
    start = time.perf_counter()
    env.step(action)            # ~3.4 ms
    time.sleep(max(0.0, period - (time.perf_counter() - start)))

At 8 Hz that is about 3.4 ms of work and 121.6 ms of sleep per step.

Here is why this matters. The policy has no notion of seconds. It learns that the next waypoint is one step away, and in our demos one step was 125 ms. An action chunk is a sequence of waypoints meant to be played one step apart. If you play it faster than the data, the robot does the same motion in less time. If you play it slower, the motion drags. Either way the robot drifts away from the pace it was trained on. Figure 4 lets you try this.

Deployment frequency versus data frequencyLeft: during demo collection the arm moves through a chunk of eight waypoints, one every 125 ms, so the chunk takes one second. Right: at deployment the same chunk is played at the frequency set by the knob. At 8 Hz the two arms move together. Faster, the arm finishes the same motion early and runs ahead of the demo pace; slower, it falls behind. Demo collection 8 HZ · ONE STEP EVERY 125 MS Deployment
Deployment 8 Hz

Figure 4. The policy learns a chunk as a sequence of waypoints, one step apart. In our demos a step is 125 ms (8 Hz). Drag the knob to play the same chunk at a different deployment frequency: the path is the same, but the time it takes is not. On the right, the dashed ring is where the arm would be at the demo pace, and the gold dashed line is how far it has drifted from it.

Placeholder: inference, latency, and everything else between the checkpoint and the robot moving.

4. Takeaways & limitations

Placeholder.

5. What’s next

Placeholder.

  1. github.com/Physical-Intelligence/openpi ↩

  2. Franka Control Interface documentation ↩

  3. Polymetis, a PyTorch-based real-time robot control framework. ↩

  4. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset ↩

  5. droid/controllers/oculus_controller.py ↩

  6. Default stream profiles in librealsense, src/ds/d400/d400-factory.cpp (rs435_device and rs405_device). The D435i defaults to 1280×720 color instead. ↩

  7. The DROID dataset ↩

  8. PREEMPT_RT patches the Linux kernel so that almost all of it can be interrupted. Interrupt handlers run as ordinary threads and most kernel locks can be preempted, so a high-priority real-time thread gets the CPU within a small, bounded delay instead of waiting behind the kernel or other processes. Franka requires the controller to “run with real-time priority under a PREEMPT_RT kernel” (Setting up the real-time kernel), and “round-trip time + control loop execution must remain below 1 ms” (Troubleshooting). ↩