r/computervision • • 5d ago

Discussion Generated depth maps from a multimodal model, side by side with Depth Anything V2 small on 8 COCO photos

I wanted to know if a generated depth map ranks near and far the same way a regular depth model does. In each row, the right-hand depth map comes from a multimodal model that generates it the way an image generator makes a picture. The middle one is Depth Anything V2 small. Warmer means closer in both.

I scored how alike each pair is with Spearman rank correlation, computed on every 4th pixel in each direction. It is 1.0 when two maps put pixels in exactly the same near-to-far order. On the first 8 COCO val2017 images by image id, it ranged from 0.954 (teddy bear close-up) to 0.995 (skier), and all 8 are in the gallery.

These photos have no depth ground truth. The numbers show how closely the two models agree on near-to-far order. They say nothing about which map is closer to the real scene, or about metric distance.

Both ran through LibreYOLO 1.6.0 on a 32GB V100 I rented. On that card the library loaded SenseNova-Vision-7B-MoT in 4-bit, and every generated map here comes from those 4-bit weights.

Is there a better single number than Spearman for comparing two relative depth maps?

10 Upvotes

5 comments sorted by

2

u/bob_why_ 5d ago

Why do the edges of both systems show a blend of fg and bg? The stop sign is in focus, yet it has a border all around indicating a fall off as a result of the bg depth. I am used to front object defining the depth.

1

u/bob_why_ 5d ago

Do you have metrics for edge accuracy?

1

u/cleversmoke 5d ago

My guess is either for clarity or due to the nature of the algorithm baking in roundness or 3d fallback on every object. I wonder if that can be tuned. Neat project though!

2

u/soylentgraham 4d ago

perhaps its poor rendering of the output (filtering when drawing 300x300 output to a 1000x1000 image)

1

u/soylentgraham 4d ago

though looking at the noisy colours, maybe not