VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps

Sunesh Praveen Raja Sundarasami1,2,*, Taehyoung Kim1,*,†, Johannes Scherer1, Tomaž Cotič1,3, Sivasubiramaniam Subbiah1,4, Andreas Greiner1,5, Paul Spannaus1, Sebastian Houben2
1Fraunhofer IVI 2Hochschule Bonn-Rhein-Sieg 3University of Bologna 4FAU 5THI
*Equal contribution †Corresponding author
Preprint
VoxelFix overview showing observations mapped into a completed semantic voxel map, followed by graph-based geometric and semantic correction.
Map-only post-hoc semantic correction. A standard mapping pipeline aggregates semantic predictions into a completed voxel map. VoxelFix begins after construction and uses only fixed voxel geometry and semantic labels to correct accumulated errors. Source images, depth maps, camera poses, and intermediate reconstruction state are not required.

Abstract

Semantic 3D maps are increasingly constructed automatically for aerial robotics by integrating learned semantic predictions into 3D representations. While this avoids costly manual 3D annotation, errors in the perception and mapping pipeline can persist in the resulting map, reducing its reliability for downstream autonomous tasks. Existing 3D semantic map refinement methods either rely on the original observations, treat occupancy as part of the prediction problem, or apply non-learned local regularization to completed maps.

Instead, we study post-hoc semantic correction, asking whether semantic accuracy can be recovered directly from the completed map while keeping its geometry and occupancy fixed. We introduce VoxelFix, a graph-based model that corrects voxel labels based on local geometry and neighboring semantic information. To obtain training pairs, we corrupt contiguous regions of annotated OccuFly maps according to class confusions observed in upstream maps.

We evaluate VoxelFix on completed OccuFly maps generated from predictions of four independently trained 2D segmentation models. VoxelFix consistently improves mIoU by 4.23-5.00 percentage points, with gains broadly distributed across the evaluated semantic classes and particularly strong improvements for tree, roof, and wall. Results on an independently reconstructed out-of-distribution aerial scene further suggest that the learned correction can transfer beyond the environments seen during training.

Method Overview

VoxelFix reasons over complementary geometric and semantic neighborhoods without altering map geometry or occupancy.

VoxelFix architecture with a fixed k-nearest-neighbor graph, geometric and semantic GATv2 branches, learned fusion, and correction and error-detection heads.
VoxelFix architecture. Each occupied voxel becomes a graph node described by local geometry, a frozen volumetric embedding, its current label, and class-conditioned agreement features. Parallel GATv2 branches process geometric and semantic edge information, then a voxel-wise gate fuses both views for relabeling. Dashed paths indicate training-only corruption, prototype updates, and auxiliary error-detection supervision.

Quantitative and Qualitative Results

In-Distribution Evaluation

Baseline comparison table. VoxelFix reaches 30.66 mIoU, gains 4.86 points, reaches 68.34 percent accuracy and 25.36 percent error correction rate, with 3.55 percent damage rate.
Map-level correction on the fixed OccuFly test scene. With SegFormer-MiT-B3 upstream predictions, VoxelFix reaches 30.66 mIoU, a 4.86-point gain over the uncorrected Radix map and 2.14 points over MinkUNet. It corrects 25.36% of initially erroneous voxels while changing 3.55% of initially correct labels incorrectly. Learned models report mean and standard deviation over eight cross-validation splits.

Robustness table showing VoxelFix gains from 4.23 to 5.00 mIoU points across SegFormer, two UPerNet variants, and Mask2Former.
Robustness across upstream models. On the same fixed test scene, VoxelFix improves completed maps from SegFormer, UPerNet with SwinV2-T, UPerNet with ConvNeXtV2-T, and Mask2Former. Gains range from 4.23 to 5.00 mIoU points despite different initial map quality and error characteristics.

Three in-distribution scenes comparing Radix input, KNN, CRF, geometry heuristic, MinkUNet, VoxelFix, and ground truth voxel maps.
Qualitative correction on the fixed OccuFly test scene. VoxelFix more consistently repairs spatially coherent roof, wall, tree, and grass errors while preserving surrounding labels than local smoothing, geometry-only refinement, and learned volumetric correction.

Out-of-Distribution Evaluation

Out-of-distribution comparison table. VoxelFix improves the reconstructed scene from 30.95 to 32.66 mIoU, a gain of 1.71 points.
Correction on an independently reconstructed aerial scene. The scene is excluded from training and model selection. Using Mask2Former predictions, VoxelFix raises mIoU from 30.95 to 32.66, a 1.71-point gain, with a 15.65% error correction rate and a 1.03% damage rate.

Three out-of-distribution scenes comparing Radix input, KNN, CRF, geometry heuristic, MinkUNet, VoxelFix, and ground truth voxel maps.
Qualitative OOD correction. VoxelFix repairs several coherent errors on unseen scene geometry while largely preserving correct labels. Ambiguous structures remain challenging when the completed map does not contain enough evidence for the correct class.

Interactive Voxel Viewer

Inspect the results above in 3D. Select a scene and upstream model, then compare the input, VoxelFix output, and ground truth with synchronized cameras.

Preparing hosted examples...
A
Noisy inputCompleted Radix map
Drag to orbit / Scroll to zoom
C
Ground truthEvaluation reference
Geometry is matched by 3D coordinate
Semantic taxonomy

Available classes and colors

Counts show input / output / ground truth. Select a class to hide it.

BibTeX

@misc{sundarasami2026voxelfixposthocsemanticcorrection,
  title = {VoxelFix: Post-Hoc Semantic Correction of Completed 3D Voxel Maps},
  author = {Sunesh Praveen Raja Sundarasami and Taehyoung Kim and Johannes Scherer and Tomaž Cotič and Sivasubiramaniam Subbiah and Andreas Greiner and Paul Spannaus and Sebastian Houben},
  year = {2026},
  eprint = {2609.05114},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url = {https://arxiv.org/abs/2609.05114}
}