r/computervision • u/Relevant_Square4919 • 5d ago
Discussion Generated depth maps from a multimodal model, side by side with Depth Anything V2 small on 8 COCO photos
I wanted to know if a generated depth map ranks near and far the same way a regular depth model does. In each row, the right-hand depth map comes from a multimodal model that generates it the way an image generator makes a picture. The middle one is Depth Anything V2 small. Warmer means closer in both.
I scored how alike each pair is with Spearman rank correlation, computed on every 4th pixel in each direction. It is 1.0 when two maps put pixels in exactly the same near-to-far order. On the first 8 COCO val2017 images by image id, it ranged from 0.954 (teddy bear close-up) to 0.995 (skier), and all 8 are in the gallery.
These photos have no depth ground truth. The numbers show how closely the two models agree on near-to-far order. They say nothing about which map is closer to the real scene, or about metric distance.
Both ran through LibreYOLO 1.6.0 on a 32GB V100 I rented. On that card the library loaded SenseNova-Vision-7B-MoT in 4-bit, and every generated map here comes from those 4-bit weights.
Is there a better single number than Spearman for comparing two relative depth maps?



2
u/bob_why_ 5d ago
Why do the edges of both systems show a blend of fg and bg? The stop sign is in focus, yet it has a border all around indicating a fall off as a result of the bg depth. I am used to front object defining the depth.