Fully label-free
Neither RSU nor ego training uses human annotations.
Scaling autonomous driving to a new city almost always means repeating the same expensive loop: collect ego-side data, annotate it, retrain. At the same time, a different kind of sensor is quietly becoming common in cities鈥攔oadside units (RSUs), fixed in place, watching traffic from a stable vantage point with a wide field of view. We wanted to know whether the second thing could replace the first.
Concretely: can this infrastructure become a network of distributed local teachers? The property that makes this possible is stationarity. An RSU sees the same intersection, from the same angle, over and over. Averaged over time, what's persistent is background, and what moves is traffic鈥攁 distinction that emerges from repetition alone, with no human label in sight.
Neither RSU nor ego training uses human annotations.
Local RSU predictions become ego-centric pseudo-labels.
The trained ego detector runs without RSUs or V2X links.
We kept each stage of the pipeline deliberately simple. The goal was not to push any single module as far as it could go, but to make the whole idea of infrastructure-taught perception concrete enough to measure, and simple enough to extend.
Persistence-point scores separate stable background from transient objects. DBSCAN and tracking refine proposals into pseudo-labels for each RSU.
82.8% average local RSU APNearby predictions are transformed into ego coordinates, filtered, and fused with distance-weighted NMS as the vehicle drives.
89.5% car pseudo-label recallAggregated RSU predictions supervise a single ego detector trained offline across many locations and viewpoints.
Standalone at deploymentAn RSU never sees a human label, so the persistence-point score has to do real work: over many observations of the same fixed scene, a point that keeps reappearing in roughly the same place is almost certainly background, and a point that doesn't is almost certainly part of something that moved. That asymmetry turns out to be sharp rather than fuzzy鈥攊n practice, static points collapse toward a score near 1.0 by orders of magnitude more than dynamic ones do, which is exactly the separation the pipeline needs to turn raw, unlabeled streams into usable pseudo-labels.
If a single RSU's detector generalized well to other locations, there would be little reason to run this pipeline at city scale鈥攐ne well-placed RSU could teach the whole town. It doesn't. Training a detector at one RSU and evaluating it at every other RSU's viewpoint shows strong performance only along the diagonal, each RSU's own trained view, and a sharp drop everywhere off it.
Studying this pipeline requires a dataset most existing V2X benchmarks don't provide: dense RSU coverage across a whole geo-fenced area, with RSU and ego views of the same traffic both available and accurately posed. Most collaborative-perception datasets place a handful of roadside units at an isolated intersection or corridor, not a city's worth of them. CIVET is a controlled multi-agent dataset built on CARLA and V2Xverse, spanning four towns with different traffic compositions, object scales, road layouts, and RSU鈥揺go spatial relationships.
On the primary urban town, the fully label-free CenterPoint pipeline reaches 82.3% AP for vehicles鈥攖welve points below a detector trained on full ego ground truth, but built entirely without a human label.
Before asking whether this scales, we needed to know whether it works at all. On a single urban town, 12 RSUs each learn their own local detector from nothing but repeated observation, broadcast pseudo-labels to the ego vehicle, and together supervise one standalone detector鈥攂uilt without a single human label.
| Training supervision | Car | Ped. | Cyc. | Avg. |
|---|---|---|---|---|
| 12 unsup. RSUs | 82.3 | 79.3 | 68.5 | 76.7 |
| Ego ground-truth | 94.4 | 91.8 | 96.6 | 94.3 |
Feasibility on one town isn't the same as geographic scalability. We combine four towns鈥攖wo contrasting rural and urban environments, plus two more added specifically to test scale鈥攊nto a single unified deployment of 48 RSUs, and train one ego detector across all of them at once. Car AP holds at 82.7%, essentially unchanged from the single-town result, even though the combined test set now spans larger vehicles and sparser traffic in one town alongside denser pedestrian and cyclist traffic in another. More RSUs and more diversity did not erode what the pipeline had learned.
| Training supervision | Car | Ped. | Cyc. | Avg. |
|---|---|---|---|---|
| 48 unsup. RSUs | 82.7 | 81.0 | 78.3 | 80.7 |
| Ego ground-truth | 91.0 | 90.6 | 95.3 | 92.3 |
Infrastructure supervision and ego-centric unsupervised learning draw on different signals鈥攐ne from a fixed RSU watching the same scene repeatedly, the other from an ego vehicle with its own viewpoint鈥攕o we'd expect the two to complement rather than duplicate each other. We find that combining our RSU pseudo-labels with Oyster's ego-centric unsupervised labels improves dynamic-object AP over either source alone.
| Annotation source | Dyn. AP |
|---|---|
| MODEST | 18.0 |
| Oyster | 39.8 |
| Ours | 62.0 |
| Oyster + Ours | 64.9 |
@inproceedings{xu2026civet,
title={When the City Teaches the Car:
Label-Free 3D Perception from Infrastructure},
author={Xu, Zhen and Yoo, Jinsu and Bautista, Cristian and
Huang, Zanming and Pan, Tai-Yu and Liu, Zhenzhen and
Luo, Katie Z and Campbell, Mark and Hariharan, Bharath and
Chao, Wei-Lun},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026}
}