Draft — intro goes here. One or two paragraphs on what this post is about, why running a fine-tuned π0.5 policy1 on a real Franka arm is harder than the paper makes it look, and who should read it.
1. Setup
This section describes our robot setup. The layout is shown in Figure 1.
The Franka arm is connected to its control PC. The Franka Control Interface (FCI)2 sits between the control PC and a NUC, which talks to the robot through polymetis3. The workstation also has three Intel RealSense cameras (two D435s and one D405), which are connected to a separate GPU machine. The GPU machine collects the camera observations and the robot state, which it gets from the NUC over polymetis.
At each inference step, the GPU machine sends the current observation, the language command and the robot state to a server machine running π0.5. The server replies with an action chunk, which the GPU machine sends back through polymetis and the FCI to the arm.
Placeholder: task and evaluation protocol.
2. Data collection
To collect demonstrations, we teleoperate the arm with a Meta Quest 2 VR headset, using the DROID4 teleoperation stack and its Oculus controller5. Figure 2 shows how it works.
The headset sits in front of the operator and tracks the hand-held controller. Holding the grip button enables movement and records two origins: the controller pose and the robot’s end-effector pose at that instant. From then on, the controller’s displacement from its origin is added to the robot’s origin pose, so the end-effector copies the motion of the operator’s hand. Orientation works the same way. The trigger opens and closes the gripper. Releasing the grip stops the robot following, and pressing it again sets new origins, so the operator can reposition their hand without moving the robot. The A and B buttons mark an episode as a success or a failure.
One decision that matters in demo collection is the frequency you record at. In practice it depends on many things: your hardware, your task, your inference server, your network delays, and so on. The first place it hits you is dataset size.
A standard VLA collection setup like ours has three RealSense cameras: two D435s and a D405. Out of the box, over USB 3, librealsense6 streams the D435’s color camera at 640×480 and the D405’s at 848×480, both at 30 fps. Each recorded step stores one frame from every camera, so the dataset grows linearly with the recording frequency, the length of each demo and the number of demos. For reference, DROID records at 15 Hz7. Figure 3 lets you try different settings.
D435 × 2
D405 × 1
Dataset size
Placeholder: tasks, number of demonstrations, and episode length.
3. Deployment on the Franka
One thing that matters at deployment is matching the control frequency to the frequency of the data. Most of our demos are recorded at somewhere between 8 and 15 Hz, mainly because we don’t have much storage.
The first question is how fast you can send a command to the robot at all. This is where the NUC from Figure 1 comes in. The Franka expects a new command every millisecond: the FCI runs a 1 kHz control loop, and the round trip plus the controller’s own computation has to fit inside that 1 ms, or the robot stops with a communication error. A stock Linux kernel can’t promise that, so the loop runs on a NUC running Ubuntu with a real-time (PREEMPT_RT) kernel8. The GPU machine never talks to the arm directly. It sends targets to polymetis on the NUC, and the NUC keeps up the 1 kHz loop on its own.
In our setup, one env.step(action) from the GPU machine takes about 3.4 ms on average, so the fastest we could command the robot is roughly 290 Hz. Recording at that rate isn’t realistic. With the three cameras at their defaults (Figure 3), a single one-minute demo would be about 54 GB, and the cameras only produce 30 new frames a second anyway, so most of those frames would be repeats. At 8 Hz the same minute is about 1.5 GB.
So we slow the loop down on purpose. To run at a target frequency, say 8 Hz to match the data, each step spends its ~3.4 ms sending the action and then sleeps for the rest of the period:
period = 1 / 8 # s, the data frequency
while True:
start = time.perf_counter()
env.step(action) # ~3.4 ms
time.sleep(max(0.0, period - (time.perf_counter() - start)))
At 8 Hz that is about 3.4 ms of work and 121.6 ms of sleep per step.
Here is why this matters. The policy has no notion of seconds. It learns that the next waypoint is one step away, and in our demos one step was 125 ms. An action chunk is a sequence of waypoints meant to be played one step apart. If you play it faster than the data, the robot does the same motion in less time. If you play it slower, the motion drags. Either way the robot drifts away from the pace it was trained on. Figure 4 lets you try this.
Placeholder: inference, latency, and everything else between the checkpoint and the robot moving.
4. Takeaways & limitations
Placeholder.
5. What’s next
Placeholder.
Polymetis, a PyTorch-based real-time robot control framework. ↩
DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset ↩
Default stream profiles in librealsense,
src/ds/d400/d400-factory.cpp(rs435_deviceandrs405_device). The D435i defaults to 1280×720 color instead. ↩PREEMPT_RTpatches the Linux kernel so that almost all of it can be interrupted. Interrupt handlers run as ordinary threads and most kernel locks can be preempted, so a high-priority real-time thread gets the CPU within a small, bounded delay instead of waiting behind the kernel or other processes. Franka requires the controller to “run with real-time priority under aPREEMPT_RTkernel” (Setting up the real-time kernel), and “round-trip time + control loop execution must remain below 1 ms” (Troubleshooting). ↩