JEOS RP ISSN03 | Page 493

486
J. Eur. Opt. Society-Rapid Publ. 22, 49( 2026)
Fig. 4. Overview of the synthetic UAV-swarm dataset generation pipeline.
within the bounded observation volume W. The software supports two complementary placement modes. In structured mode, UAV instances are arranged according to a regular three-dimensional grid, with configurable spacing along each axis( controlled by XCount, YCount, ZCount and a LoadFactor parameter), enabling precise control over formation geometry. In randomized mode, positions are sampled stochastically within W, governed by an explicit random seed that ensures full reproducibility. A global density factor regulates spatial sparsity, preventing unrealistic overlaps while allowing the generation of both sparse and dense swarm configurations. In our experiments, all results are obtained on datasets generated using eight different UAV 3D models, as illustrated in Figure 5.
Camera placement and scene orientation are parameterized independently. The virtual camera is constrained to face the center of the observation volume, guaranteeing that all UAV instances remain within the field of view. Camera translation along the three spatial axes( CameraX, CameraY, CameraZ) and global rotation of the scene( PlaneRX, PlaneRY, PlaneRZ) are specified using start, end, and step values, enabling automated traversal of the parameter space and efficient batch generation of images covering diverse viewpoints. Notably, the CameraZ parameter represents a relative scale factor rather than an absolute distance; the actual camera-to-swarm distance is automatically computed based on the spatial extent of the loaded models. A real-time preview mechanism allows users to inspect parameter effects before committing to large-scale generation, reducing wasted computation and improving dataset quality.
The rendering subsystem supports multiple output resolutions, including full-HD( 1920 1080), 2 K( 2560 1440), and 4K( 4096 2160), as well as both lossy( JPEG, compressed to 85 % quality) and PNG image formats. Scene appearance diversity is introduced through seven selectable skybox environments provided by Unity’ s rendering pipeline, including Daytime, Sunset, Beautiful Pasture, Snowy Bridge, Small Harbour, Tellsplatte, and Country Road, representing a variety of sky textures, weather conditions, and lighting scenarios( see Fig. 6). During batch generation, all parameters are locked to ensure consistency across the produced images.
For each rendered frame, ground-truth bounding boxes are generated automatically using the projection-based annotation procedure described previously. The software outputs both the rendered images and their corresponding annotation files in standard formats compatible with common deep learning frameworks, enabling direct integration into detector training pipelines.
By combining explicit parameterization, deterministic generation, and automated annotation, the proposed software tool provides a flexible and efficient foundation for synthesizing large-scale UAV swarm datasets. The following section presents a quantitative analysis of the dataset produced using this system.
3.4
Dataset analysis and characteristics
This section presents a quantitative analysis of the proposed synthetic UAV swarm dataset, focusing on dataset scale, target size distribution, and swarm density characteristics relevant to long-range multi-UAV detection.
The dataset consists of 7000 RGB images rendered at a resolution of 1920 1080, containing a total of 31,542 annotated UAV instances. Each image includes between 1 and 15 UAVs, with an average of 4.5 instances per image, emphasizing multi-target swarm scenarios rather than isolated UAV detection. Overall statistics are summarized in Table 2.
Following common practice, we characterize p ffiffiffiffiffiffi target scale by the square root of bounding box area, wh, wherew and h denote the bounding box width and height in pixels. Based on this measure, UAV instances are categorized into small, medium, and large targets. As shown in Table 2, the dataset is strongly biased p ffiffiffiffiffiffi toward small objects, with 67.3 % of instances having wh < 32 pixels. This distribution closely reflects real-world long-range UAV observation scenarios and poses significant challenges for generic object detectors.