Abstract
Underwater object detection remains challenging due to severe visual degradation caused by light absorption, scattering, and depth-dependent color attenuation, as well as the limited availability of large-scale and diverse underwater datasets. Recent advances in large-scale text-to-image generative models offer a scalable alternative for synthesizing underwater imagery across varied environmental conditions. This paper investigates whether depth-aware text-to-image generation can produce synthetic underwater scenes with perceptual characteristics suitable for data augmentation in underwater vision. We generate underwater images across shallow (0–6 ft), mid-depth (6–30 ft), and deep-water (30+ ft) ranges using modern text-to-image models, and systematically compare them with real underwater images. A comprehensive, reference-free evaluation is conducted using both underwater-specific perceptual quality metrics and complementary generic image analytics to assess color fidelity, structural clarity, contrast, and perceptual naturalness. The results reveal consistent depth-dependent visual trends shared by real and synthetic images, along with distinct model-specific behaviors across depth ranges. In particular, different generative models exhibit clear trade-offs between perceptual enhancement, visual stability, and adherence to natural image statistics. While improved perceptual quality does not necessarily guarantee better detection performance, our findings suggest that depth-aware synthetic imagery can effectively complement real datasets by enriching underrepresented depth conditions and visual characteristics, providing practical insights for data augmentation in underwater object detection.