J. Eur. Opt. Society-Rapid Publ. 22, 49( 2026) 485
Fig. 2. The arrows indicate roll, pitch and yaw rotations about the body-fixed axes, corresponding to the attitude parameters(/ i, h i, w i) in the 6-DOF state representation.
target scale in the image is governed by the depth component z i: smallerz i yields larger bounding boxes, while larger z i produces small-scale targets.
The imaging sensor is modeled as an ideal pinhole camera with intrinsic matrix K 2 R 3 3 and extrinsic parameters [ R c | t c ], where R c 2 SO( 3) is the rotation matrix and
t c 2 R 3 is the translation vector. A 3D point
v ¼ ðX; Y; ZÞin world coordinates is projected onto the image plane as:
2 3 2 3
X u 6 7
Y k4 v 5 ¼ KR ½ c; jt c Š6 7 ð8Þ 4 Z 5
1
1
where( u, v) are pixel coordinates and k is a scale factor. In our implementation, the camera is positioned at a fixed location within a virtual outdoor scene constructed in Unity, configured to match a standard full-HD resolution of 1920 1080 pixels.
Ground-truth bounding boxes are generated automatically by projecting each UAV’ s 3D bounding volume onto the image plane. Let B i R 3 denote the axis-aligned bounding box of UAV i in world coordinates, with eight corner vertices V i ={ v 1, v 2,..., v 8 }. Each vertex is projected to obtain its 2D image coordinates:
ðu k; v k Þ ¼ pðv k; K; R c; t c Þ; k ¼ 1;...; 8 ð9Þ
where p() denotes the perspective projection function. The 2D bounding box b i =( x min, y min, x max, y max) is then computed as:
b i ¼ ðmin k u k; min k v k; max k u k; max k v k Þ: ð10Þ
This projection-based annotation guarantees pixel-level accuracy and perfect consistency across the dataset, completely eliminating the cost and potential errors associated with manual labeling. Combined with the trajectory-based generation scheme, the proposed pipeline enables efficient synthesis of large-scale, densely annotated UAV swarm datasets with controlled variation in pose, formation and imaging conditions.
Fig. 3. The arrows indicate translations along the three orthogonal axes of the world coordinate frame, corresponding to the position components( x i, y i, z i) in the 6-DOF state representation.
3.3
Software implementation
The data generation pipeline described above is implemented as a standalone software tool named 3D Model Shots Generator, built on the Unity engine. This tool integrates all components of the proposed methodology, including 6-DOF state initialization, trajectory interpolation, camera configuration, rendering, and automatic annotation, into a unified and user-friendly framework.
The software adopts a parameter-driven architecture in which scene configuration, object placement, camera settings, and rendering control are explicitly decoupled and exposed to the user, as illustrated in Figure 4. Thisdesign philosophy enables systematic exploration of the data generation space while maintaining strict reproducibility. Given identical 3D models and parameter settings, the system deterministically produces the same swarm configurations and rendered outputs, which is essential for controlled experimental analysis and fair benchmarking across different detection algorithms.
During initialization, one or more 3D UAV models in common formats are loaded and instantiated multiple times