Raspberry Pi 5 + Hailo-8L  ·  27 FPS  ·  everything onboard

Seeing to Drive on a Single Board

An attention-free student of Depth Anything V2 that runs on a 13-TOPS integer NPU with int8 equal to float, and the onboard obstacle avoidance, global planning and SLAM it makes possible on a small rover.

Itai M.  ·  DANIEL: Depth Anything Navigation In Exteriors, Learned  ·  2026

27.1 FPSdepth at 288×512 on the Hailo-8L, batch 1, measured on the Pi
r = 0.99993int8 on the chip vs the same network in float, 500 frames
0.269AbsRel on held-out metric ground truth; next best on the chip 0.354
8.2 / 10scenes reached with no contact, with the real Pi + Hailo in the real-time loop

Foundation depth does not reach a small NPU through better quantisation.

The Hailo-8L runs attention in int8, and diffuse attention at a few hundred tokens falls below one quantisation step: our build of Depth Anything V2 and the vendor's both erase obstacles. We distil the 335 M-parameter teacher into a 24.5 M-parameter network with no attention at all, range obstacles from where they meet the floor, and run mapping, hybrid-A* planning and SLAM on the Pi itself.

Demos

The real board, in the loop and on real video.

4K source · every frame on the Hailo-8L

Outdoors: the deployed student, frame by frame on the NPU

A low camera crossing a field toward rocks and posts. Every frame went through the student v5 HEF on the Raspberry Pi 5's Hailo-8L (34.6 ms per frame on the NPU) and was read back; the camera half is the 4K original. Frames from this video (every 10th) were in the student's training set, so it is a demonstration, not a held-out result. Download the 4K version.

Real Pi 5 + Hailo-8L in the loop

The deployed stack drives on its own SLAM pose

Isaac Sim streams the car's camera to the real Pi; the Pi computes depth on the Hailo, maps, plans and replies with steering and throttle, applied only when they arrive. The six scenes one recorded pass completed with no contact, each uncut; that pass scored 6 clean and 7 reached of 10, the five-seed mean is 8.2.

Real robot camera

Depth, point cloud and obstacle scan from one camera

Real robot video with depth computed on the Hailo-8L (student v3 HEF, every frame read back from the chip), the metric point cloud scaled by the floor, and the ground-contact scan with the chosen arc.

All ten scenes

A full real-board pass, failures included

Every scene of the same pass, labelled with its outcome on the car's real footprint: chairs, a cone slalom, a barrel gate, people, mixed clutter and a crate wall, with timeouts where they happened.

Real robot camera frame Depth computed on the Hailo-8L for the same frame
cameraHailo-8L depth
Drag

Camera ↔ depth from the chip

Held-out real frames; depth from the deployed student v5, read back from the Hailo-8L.

Method

Three ideas, each measured on the chip.

01

Why transformer depth breaks on the Hailo-8L

At the chip's 449 tokens, a diffuse softmax weight is about 1/449 ≈ 0.0022, below one uniform int8 step of 1/255 ≈ 0.0039. Measured in float, 91 % of attention weights fall below that step and carry 17–33 % of each query's attention mass. Re-calibration and quantisation fine-tuning do not fix it; removing attention does.

Per-block share of attention weights below one int8 step and the attention mass they carry
02

An attention-free student

A ResNet-34 encoder and a U-Net decoder of convolution, batch norm, ReLU and bilinear upsampling: every operation quantises cleanly. 24.5 M parameters (DAv2-S: 24.8 M), distilled from DAv2-Large with a scale- and shift-invariant loss, then fine-tuned on rendered ground truth. On the chip: 35.1 ms, int8 = float at r 0.99993.

ResNet-34 encoder U-Net decoder
03

Metric space without metric depth

Relative depth has no scale. But where an obstacle meets the floor, the image row alone gives its distance from the camera height. The network only has to say where each column's floor ends; the columns become a 69°, 4 m laser scan, and the floor profile becomes a ruler for a metric point cloud.

Z = fy · h / (vc − cy)

Ground-contact scan hits on a held-out frame and in bird's-eye view against the true prop footprints
  1. cameraRGB frame30 Hz
  2. Hailo-8L NPUStudent depth35.0 ms, int8
  3. Pi CPUGround-contact scan + cloudin parallel with the NPU
  4. Pi CPUHit-count map + SLAMground VO, encoder, gyro
  5. Pi workerHybrid A* + safety cascade110 ms median
  6. Pi CPUPure pursuit → PWMESC brake gate

Results

Every depth network we could run on the Hailo-8L.

400 held-out simulated frames with the renderer's metric depth and 100 held-out real robot frames. Every Hailo model runs on the chip and is read back; per-frame least-squares scale and shift in disparity. Obstacle pixels are those that land on a prop's collider, which removes the renderer's floor speckle.

Accuracy against measured on-chip speed

Hover or focus a point. Dashed lines: float GPU models that cannot run on the Pi at this rate.

ModelRuns onAbsRel allAbsRel obst.δ1r (real)Speed
Depth on real robot frames computed on the Hailo-8L by every model, against the float teacher
Real robot frames, depth as computed on the Hailo-8L (except the GPU teacher). The int8 transformer builds lose the scene; the student keeps the teacher's structure at 27 FPS. Dark areas lie outside a model's input crop.

Closing the loop: a real-time error budget.

Ten frozen scenes in a Gaussian-splat reconstruction of a real yard, eleven obstacle types with honest colliders, at least five seeds per arm, contacts scored on the car's real 0.557 × 0.294 m footprint. Clean goals of 10:

Oracle: exact props, lockstep
9.8
Real time, exact obstacles
9.2
Hailo depth, simulator pose
7.9 ± 1.1
Hailo depth, own SLAM pose
7.7 ± 0.7
Own pose, real Pi + Hailo
8.2 ± 0.4
Camera-only odometry
2.6

Execution costs about 0.6, perception 1.3, self-localisation 0.2 clean goals (0.9 goals reached). Ten seeds for the deployed arms, five for the others; the own-pose arms use a simulated wheel encoder and gyro.

Clean success per scene for the main configurations
Per scene: the misses concentrate on the cone slalom, the gate, the chair cluster, the crate wall and the person-and-cones scene.
Every hybrid-A* plan in one episode with the SLAM estimate against the truth
Onboard SLAM and planning on the robot's own pose: every plan the Pi made in one slalom episode. The car reached the goal despite 0.47 m peak drift.
A real robot frame lifted to a metric point cloud on the Pi
One real frame to metric 3D on the Pi: chip depth, scale from the floor profile, flying pixels removed.
Per-frame stage timing on the Pi and hybrid-A* search times
The budget on the Pi: NPU 35 ms overlapped with 47–49 ms of CPU; 16.5 Hz on 720p video with SLAM, 27 Hz without; 28 h of real-board runs without throttling.

Honestly

What is not solved yet.

Resources

Everything to reproduce the numbers.

Citation

BibTeX

@misc{itaim2026seeing,
  title  = {Seeing to Drive on a Single Board: Attention-Free Distilled Depth on a
            Raspberry Pi 5 + Hailo-8L for Onboard Obstacle Avoidance, Planning and SLAM},
  author = {M., Itai},
  year   = {2026},
  url    = {https://itaim18.github.io/seeing-to-drive/}
}

Built on Depth Anything V2 and Depth Anything 3, the Hailo Model Zoo, NVIDIA Isaac Sim / Isaac Lab and its neural-reconstruction rendering. Weights and data CC-BY-NC-4.0 (the teacher's licence).