Overview
RoofVIP v1.0.3.0 extends the RoofVIP Benchmark Dataset into a multimodal, fixed-scene representation for building-roof analysis and cross-modal benchmarking. It is derived from geospatial products provided by the Bavarian Surveying Administration (Bayerische Vermessungsverwaltung) through Bavarian Open Geodata and combines very-high-resolution RGB imagery, elevation data, LiDAR point clouds, and manually annotated roof geometry.
The release covers approximately 18 km² of the Munich metropolitan area. Whereas the original RoofVIP benchmark is organized as object-centred image-vector pairs, v1.0.3.0 reformats the data into fixed 256 × 256 pixel scenes. At the native ground sampling distance of 0.2 m, each scene covers approximately 51.2 × 51.2 m on the ground.
A scene can contain multiple buildings, surrounding background objects, and buildings intersected by the scene boundary. Partial buildings are retained. This scene-based organization provides a consistent spatial unit for training and evaluating image-, elevation-, point-cloud-, and multimodal learning methods.
10,245 scenes · 8,195 train · 1,022 validation · 1,028 test · DOI 10.5281/zenodo.22015175
Data modalities
Each scene may contain the following co-registered data modalities:
Very High Resolution RGB imagery (VHR RGB)
Derived from the Digitales Orthophoto RGB 20 cm (DOP20 RGB) provided by the Bavarian Surveying Administration.
Digital Surface Model (DSM)
Derived from the Digitales Oberflächenmodell 20 cm (DOM20).
Digital Terrain Model (DTM)
Derived from the Digitales Geländemodell 1 m (DGM1).
LiDAR point cloud
Point-cloud data covering the same geographic scenes are included to support multimodal and cross-modal roof-reconstruction experiments.
Normalized Digital Surface Model (nDSM)
The nDSM is a derived elevation representation calculated directly from the supplied DSM and DTM:
nDSM = DSM − DTM
Because the nDSM is generated from the supplied DSM and DTM products, it has no separate original-data download source.
Vector roof annotations
Building roof structures are manually labelled following a Level of Detail (LoD) 2.0 representation. Each individual roof plane is represented as a polygon in a common projected coordinate reference system. Polygon vertices and boundaries can therefore be used as point, edge, polygon, or graph targets for roof-reconstruction tasks.
The vector annotations were initially created in ESRI Shapefile (.shp) format and additionally converted into NumPy polygon representations (.npy) for machine-learning applications.
Together, the RGB, DSM, DTM, nDSM, LiDAR point-cloud, and vector-label products provide a common spatial basis for evaluating unimodal and multimodal building-roof reconstruction approaches.
Dataset processing
The release was generated from the original Bavarian Open Geodata products and the RoofVIP roof annotations through the following workflow:
- The geographic domain was partitioned into spatially independent regions before scene generation.
- Original orthophotos, elevation products, and point-cloud data were spatially matched to the RoofVIP geographic domain.
- The data were reformatted into fixed 256 × 256 pixel scenes.
- Roof-plane polygons were associated with corresponding scenes while preserving buildings intersected by scene boundaries.
- DSM, DTM, RGB, and LiDAR data were spatially cropped and co-registered to the corresponding scene extent.
- The nDSM was derived from the DSM and DTM.
- Manually annotated roof-plane polygons were retained as the reference vector labels and converted into machine-learning-compatible representations where required.
- Files were organized using consistent scene identifiers so that raster, point-cloud, derived-elevation, and vector products can be directly associated across modalities.
Within each geographic partition, scenes are generated with 25% forward and side overlap, corresponding to a stride of 75% of the scene size. Scenes containing no annotated building geometry are excluded from the benchmark.
No additional geometric modification is applied to the source DSM, DTM, or LiDAR measurements beyond the spatial processing required for scene extraction, alignment, and dataset organization.
Training, validation, and test split
RoofVIP v1.0.3.0 provides predefined training, validation, and test subsets using an approximately 80–10–10 split:
| Subset | Scenes |
|---|---|
| Training | 8,195 |
| Validation | 1,022 |
| Test | 1,028 |
| Total | 10,245 |
The split follows the structural-complexity-based strategy used in RoofVIP. Descriptors representing building roof geometry and graph topology characterize roof-structure complexity. These descriptors are reduced using Principal Component Analysis (PCA), producing a PCA-based representation used to balance the structural-complexity distribution across the training, validation, and test subsets.
Importantly, the geographic area is partitioned before tiling. This prevents overlapping regions, spatially adjacent duplicate content, or the same individual building from appearing in more than one subset. Scene generation is then performed independently within each partition. This procedure reduces spatial leakage while retaining comparable distributions of roof-structure complexity across the three subsets.
The predefined split should therefore be retained when reporting benchmark results to facilitate reproducibility and direct comparison between methods.
Relationship to the original RoofVIP benchmark
The original RoofVIP benchmark consists of manually annotated two-dimensional roof-plane polygons and very-high-resolution RGB orthophotos organized primarily around individual building objects. Version 1.0.3.0 retains the same underlying geographic domain and roof annotations while introducing a fixed-scene representation and additional co-registered modalities.
The principal extensions are:
- conversion from object-centred samples to fixed 256 × 256 pixel scenes;
- inclusion of DSM and DTM elevation data;
- derivation of corresponding nDSM data;
- inclusion of LiDAR point-cloud data for multimodal experiments;
- preservation of manually annotated LoD 2.0 roof-plane vector geometry;
- standardized correspondence among all modalities using common scene identifiers; and
- a predefined, geographically separated and PCA-complexity-balanced 80–10–10 train/validation/test split.
These additions allow RoofVIP to support not only RGB-based roof reconstruction, but also systematic comparison of image, elevation, point-cloud, and multimodal approaches under a common dataset configuration.
Access and persistent identifier
- Zenodo record: https://zenodo.org/records/22015175
- DOI: 10.5281/zenodo.22015175
- Archive:
RoofVIP_1.0.3.0_FixTileScene_Size256.zip
Licensing
The original Bavarian Open Geodata products used to construct this dataset are distributed under the Creative Commons Attribution 4.0 International License (CC BY 4.0): creativecommons.org/licenses/by/4.0/.
The image tiles, elevation crops, and point-cloud subsets contained in this release remain derived from their respective original Bavarian Open Geodata products and remain subject to the applicable CC BY 4.0 attribution requirements.
The manually produced roof annotations, converted polygon representations, derived nDSM data, scene organization, dataset curation, and associated benchmark structure are distributed as part of this dataset under CC BY 4.0.
Users must provide appropriate attribution to:
- the Bavarian Surveying Administration (Bayerische Vermessungsverwaltung) as provider of the underlying geospatial data; and
- the authors of the RoofVIP dataset for the roof annotations, processing, dataset organization, and benchmark preparation.
Appropriate credit to the original data providers must be preserved. No endorsement by the original licensors is implied.
When using or redistributing the dataset, preserve the original source and licensing information and cite the corresponding RoofVIP publication where applicable.
Citation
For the original RoofVIP scientific publication, please cite:
Amrullah, C., Panangian, D., Mutreja, G., Abdelhedi, Y., and Bittner, K.: RoofVIP Benchmark Dataset: 2D Roof Planar Polygons and Very High-Resolution Digital Orthophotos Pairs for Building Roof Reconstruction, ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sci., XI-2-2026, 207–216, 2026. https://doi.org/10.5194/isprs-annals-XI-2-2026-207-2026
For reproducibility, also identify the exact dataset release as RoofVIP v1.0.3.0 and cite its Zenodo DOI: 10.5281/zenodo.22015175.