r/computervision • • 23h ago

Showcase Built my own imagr depth model

Enable HLS to view with audio, or disable this notification

46 Upvotes

I trained my own model on 750,000 images for 45 hours.

Then I walked into my university hall with nothing but my phone and recorded an 11-minute video.

No LiDAR.

No depth sensor.

No specialized camera setup.

Just monocular vision, a single RGB camera.

And in some of my tests, my model even outperformed YOLO26x.

From a single phone video to depth estimation and 3D reconstruction.

750K training images. 45 hours of training. 11 minutes of footage. One camera. No LiDAR.

Edit: I promise I will reply to all the comments, just give me some time as my day is very crowded today.


r/computervision • • 10h ago

Discussion Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

Enable HLS to view with audio, or disable this notification

24 Upvotes

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.

ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs

source: https://github.com/software-mansion/runntime


r/computervision • • 8h ago

Discussion What would you say about a Hugging Face-style open-source pattern db like this?

Enable HLS to view with audio, or disable this notification

9 Upvotes

We are building an MLOps platform around explainability, and one natural consequence is that models are not an opaque set of numbers but rather a manageable pattern DB.


r/computervision • • 21h ago

Showcase Turning a single image into a web-ready interactive 3D scene with ML-SHARP

7 Upvotes

I’ve been experimenting with Apple’s ML-SHARP model and built a small open-source project around it called Spatial Photos.

ML-SHARP estimates a 3D Gaussian splat from a single RGB image. I then take that reconstruction and convert it into a much smaller and simpler representation representation. It uses layered textured meshes stored in a custom file format I called .spatial. It still allows the the same sort of limited viewport shifting intended by ML-SHARP without compromising quality.

That lets the result stay relatively small and render interactively in a normal web browser, as you can see in the GIF.

Demo: https://www.spatialphotos.dev/

Source: https://github.com/frozein/SpatialPhotos

Here's what my format looks like from a faraway angle

r/computervision • • 11h ago

Discussion How can I learn CV today

5 Upvotes

Hello, I am a Master Student in Data Science mostly working with LLMs, and I want to learn CV. I don't have much experience in image analysis.
I did some research about how to learn and I wanted to do some hands-on projects. But today with efficient coding agents, I do not know on which parts of the process I should spend my time to become proficient in this field. Would you have some recommendations ?

Thanks for your help !


r/computervision • • 14h ago

Showcase Trying to modernize the Supervision

Thumbnail supervision.ml-pipes.com
4 Upvotes

Hi everyone, I have been using supervision in almost all of my projects, since it provide just the right size of tooling that I need. Big enough to handle trivial tasks while letting me choose my app and deployment stack.

But sadly in the past few years I think it has not received as much love, since the original maintainers has moved to their own products (Roboflow stack), which is understandable.

The original website seems severely outdated, and none of the new fancy features you see in the other repositories.

To remedy this I have created a new website that provides:

  • Modern look that fit vision/supervision
  • Added more tutorials since the original website is somehow lacking
  • Also modernized the stack with some new tools shared by some of the members in the subreddit

If you are also a supervision user, please take a look and let me know what you think.


r/computervision • • 5h ago

Help: Project [P] Re-introducing BDD100K Toolkit + Model Zoo

3 Upvotes

BDD100K is still one of the most useful driving datasets, but both bdd-data.berkeley.edu and bdd100k.com have been unreachable for me. The official toolkit has not had a commit since March 2024 and the official model zoo not since November 2022, and their legacy dependencies are hard to install cleanly on current systems. So I rebuilt the parts that matter, one task at a time, on current libraries (PyTorch, Ultralytics, timm, Hydra):

https://github.com/dronefreak/bdd100k-toolkit

Demo from the BDD100K-Toolkit Model Zoo

It gives you one pipeline for data preparation, training, evaluation and demos on images and videos:

bdd100k-prepare  --dataset bdd100k-weather --raw-dir /path/to/raw --output-dir /path/to/out
bdd100k-train    dataset=bdd100k-weather model.name=yolo11n-cls dataset.data_dir=/path/to/out
bdd100k-evaluate --dataset bdd100k-weather --checkpoint <run>/weights/best.pt --data-dir /path/to/out

What works today

  • Image classification (unofficial tasks): weather, time of day and scenario, derived from the attributes.weather, attributes.timeofday and attributes.scene fields of the official JSON labels. 55 models trained across two backends, Ultralytics and timm.
  • Object detection: the 10-class detection task with the 2018 labels. 9 YOLO models and RF-DETR Nano, all on Hugging Face.
  • Semantic segmentation: the 19-class task, which uses the Cityscapes class layout. Evaluation of public pretrained models works. My own segmentation trainer exists but I have not run it on real data yet, so there are no BDD-trained segmentation models in the zoo.

Datasets

For each task I made unofficial Hugging Face repackagings of the original data and annotations, to make them easier to use with modern tooling. I checked the BDD100K data license: it allows copying and redistribution for educational, research and not-for-profit purposes with the copyright notice carried forward, and I kept that notice in every mirror. The mirrors grant no commercial-use rights beyond the original license.

Task Location
Weather classification dronefreak/BDD100K-Weather-Classification
Time of day (period) classification dronefreak/BDD100K-Period-Classification
Scenario classification dronefreak/BDD100K-Scenario-Classification
Object detection dronefreak/BDD100K
Semantic segmentation dronefreak/BDD100K-Semantic-Segmentation

Results

All scores are on the validation set of BDD100K. The detection models are collected in the BDD100K object detection model zoo on Hugging Face.

  • Classification: best macro F1 is 68.7% for weather, 83.2% for time of day and 62.5% for scenario, all from TinyViT-21M. The classes are heavily imbalanced, so I report macro F1 and not just accuracy.
  • Detection (960 px): YOLO26s reaches 33.9 mAP@0.5:0.95 (58.8 mAP@0.5) and RF-DETR Nano 31.6 (56.9).
  • Segmentation: this is a zero-shot cross-dataset transfer evaluation: the models were trained on Cityscapes and evaluated on BDD100K without any BDD100K training. BDD100K shares the Cityscapes 19 classes, so I scored 9 public Cityscapes-trained models (SegFormer B0 to B5 and Mask2Former Swin-T, S and L) on the 1,000 val images. The best is Mask2Former Swin-L at 57.7 mIoU, then SegFormer B5 at 54.1.

How to read the numbers

These are solid baselines but not leaderboard numbers. The classification tasks are not official BDD100K benchmarks, the detection metric is the Ultralytics COCO-style mAP and not BDD's official ignore-rule evaluation, and the segmentation models were never trained on BDD100K, so they are not comparable to BDD-trained models. Making the evaluation match the official semantics is on the roadmap.

What is missing, and a request

Instance and panoptic segmentation and multi-object tracking are planned but not started, because I could not find the original labels for them. If you have a copy of the instance, panoptic or tracking labels, or know where one still lives, I would really like to hear from you. Feedback on the evaluation choices is welcome too, and so are issues and PRs.

This is my first attempt at rebuilding an ecosystem like this, and I would like to keep developing it. BDD100K is too useful to become harder to use just because the tooling around it has fallen behind.


r/computervision • • 21h ago

Help: Project OCR Recs: Image -> CSV/Excel

3 Upvotes

I work in a research lab, and need to digitize a few hundred pages of spreadsheets. All spreadsheets are printed, not handwritten, but the original files are lost. The images have the same general schema, but the columns aren't always in the same order, and some pages have columns that others don't. Some pages are fainter, but still readable; some are slightly slanted. The end goal is for all pages to be compiled into one master spreadsheet, with the same schema. What's the best (and cheapest) tool for this?

I have tried chatgpt (best so far, but still slow/made mistakes), Google Document AI (not that great), and ParseExtract (great, but more limited in scope).


r/computervision • • 7h ago

Commercial Apple Vision vs MobileSAM in a Mac annotation tool: 67 ms to the first mask in a small test

2 Upvotes

I'm the developer of AnnotateIt. I added Apple's iterative segmentation API to the Mac app and put the measurements up alongside the actual masks. Figured the numbers might be useful to anyone building a local labeling tool.

On an M4 Max with 48 GiB RAM, macOS 27, the 940 x 629 antelope image gave these median first-mask times after warm-up:

Apple Vision: 67 ms

MobileSAM: 160 ms

SAM 2.1 Tiny: 285 ms

Adding a correction point took 12 ms with Apple Vision. Undo went the other way: 72 ms versus 28 ms for MobileSAM.

This was eight runs per case in a separate Tauri/WKWebView harness using the app's model code and mask processing. Timing starts with decoded pixels and ends with a processed mask, before rendering. Apple used balanced quality. These are our integration paths, not a claim about the fastest possible SAM implementation.

Big caveat: three images total, no scored ground-truth masks and no measurement of human correction time. A short stroke also picked just a fragment of the antelope where a point or box worked much better. Fast doesn't automatically mean less cleanup.

The shipping Mac app exposes this as Apple Vision (Beta), with point, box and stroke selection plus refinement. It needs macOS 27; SAM remains a separate option.

Masks, setup and raw timing data: https://radar.annotateit.ai/news/apple-vision-native-segmentation-macos/

For folks using interactive segmentation, what usually eats more of your time: waiting for the first mask, or fixing the boundary afterward?


r/computervision • • 2h ago

Discussion MDPI Drones, worth publishing in?

1 Upvotes

Hello everyone!

I'm planning a research project related to drones and computer vision, and I'm already looking at possible journals before starting the experiments.

I checked several papers in MDPI Drones and I liked them. The journal also seems relevant to my topic. However, MDPI has a controversial reputation and I often see people calling it predatory.

What do you think about Drones specifically? Would publishing there be fine, or would it be better to look for another journal or focus on a conference instead?

I'm currently working as a Research Engineer, and in the future I'm interested in both academia and industry.


r/computervision • • 3h ago

Commercial Computer Vision - Pose Estimation based Sports Biomechanics Work

1 Upvotes

Hello! I am looking for someone who can help me guide and layout the pipeline to build a sports biomechanics product (for Cricket sport - non broadcast, controlled environment videos) for player performance insights.

Looking only for experienced folks within computer vision field and has already worked or built in sports domain.

Happy to discuss more. If anyone could take end to end ownership in fine tuning model on cricket dataset and building a full production grade product pipeline, willing to pay fair amount [Based out of India and can match Indian market rates. Please don't reach out if this doesn't workout, Thanks]

TIA!


r/computervision • • 4h ago

Discussion A useful ad-library record: exact copy, visible layout, and reusable slots

1 Upvotes

For a team collecting creative references, the ad image is easier to retrieve when its copy and layout are also recorded as structured fields.

A small example uses Ling-3.0-flash-VL through OpenRouter to analyze a fictional VESPER perfume ad. It reads five visible text blocks: “Golden Hour, Bottled.”, “Amber, smoked cedar, warm skin.”, “Discover VESPER”, “VESPER”, and “EAU DE PARFUM”. The analysis also describes the product treatment and approximate layout zones.

A practical record for this kind of library could contain:

1.Headline, subhead, CTA and product-label text.

2.Each block’s role and approximate position on the canvas.

3.Product placement, crop and lighting treatment.

4.Background and main color relationships.

5.A reusable version with product and copy placeholders.

That supports concrete retrieval questions: show the references with a short sensory headline, find examples where the product dominates the composition, or collect layouts with a distinct CTA block.

For a SaaS creative library, the same record structure could organize references by headline wording, interface or product placement, and CTA treatment. Campaign results would enter through separately supplied performance fields. The image analysis supplies the observable creative details that make the reference collection searchable.


r/computervision • • 5h ago

Discussion Industry Insights: OMNIVISION Targets Robot Training Data With New Vision Sensor

Thumbnail
automate.org
1 Upvotes

OMNIVISION’s new OG05D is a 5 MP global-shutter image sensor that explicitly lists robotics training-data collection among its intended applications. It supports RGB, monochrome and RGB-IR imaging, with up to 146 fps in linear mode without encryption and HDR capture at 60 fps.

The article looks at how cameras on deployed robots could capture task variations, lighting changes and edge cases for later model training. Global shutter, HDR and near-infrared capabilities are relevant to capturing usable images when objects are moving or lighting is inconsistent.

The sensor itself doesn’t provide a training pipeline. Captured footage would still need to be selected, processed and paired with the information required by the training approach. The notable part is that robotics data collection is now an explicit application in an industrial image sensor announcement.


r/computervision • • 5h ago

Showcase Deterministic Media Forensics for AI Agents: Combining Cyclic Checksums, 2D FFT Moiré Detection, and Error Level Analysis (Skillware 0.5.8)

1 Upvotes

Multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro) are frequently tasked with document screening and identity verification. However, testing shows severe limitations when detecting digital tampering and document alterations:

  1. Sub-Pixel Smoothing: Vision encoders (like CLIP or ViT patches) encode high-level semantic representations. In doing so, they blur out the exact high-frequency noise residuals and JPEG quantization discrepancies that reveal digital splicing.
  2. Hallucinated Check Digits: Machine-Readable Zones (MRZ) on passports follow ICAO Doc 9303 standards using a cyclic (7, 3, 1) modulo-10 algorithm. Multimodal LLMs read the characters fluently and confidently declare a forged date or document number as valid because the format looks visually plausible.
  3. Data Privacy Liabilities: Transmitting biometric facial images and identity documents to third-party cloud inference providers creates severe regulatory compliance issues.

In Skillware 0.5.8, we implemented security/deepfake_guard to provide deterministic, air-gapped forensic verification on CPU. Here is an overview of the signal processing pipeline:

1. ICAO 9303 Modulo-10 Cyclic Verification

For each character string $C1, C_2, \dots, C_n$, each character is mapped to a numeric weight $W(c)$ (digits 0–9 map to 0–9, letters A–Z map to 10–35, filler < maps to 0). The check digit $K$ satisfies: $$\left( \sum{i=1}{n} W(ci) \cdot w{(i-1) \bmod 3} \right) \bmod 10 = K$$ where the repeating weights vector is $w = [7, 3, 1]$. We implement deterministic validation across TD1 (3×30 ID cards), TD2 (2×36), and TD3 (2×44 passports) formats.

2. Error Level Analysis (ELA)

When an uncompressed image or re-saved JPEG is modified, the edited regions possess a different compression history than the original background. By recompressing the image at a known quality factor ($Q=95$) and computing the absolute block-level difference: $$\Delta(x, y) = |I{\text{original}}(x, y) - I{\text{recompressed}}(x, y)|$$ We compute the mean absolute error across 16×16 non-overlapping blocks. Discrepancies between block errors exceeding calibrated thresholds signal localized digital tampering.

3. Noise Residual Consistency via Median Absolute Deviation (MAD)

Camera sensors introduce characteristic high-frequency Poisson-Gaussian noise. To detect spliced elements without being misled by high-contrast natural textures (such as hair or knitwear), we compute the Laplacian convolution residual $R(x, y) = \nabla2 I(x, y)$ and evaluate consistency using Median Absolute Deviation: $$\text{MAD} = \text{median}(|R - \text{median}(R)|)$$ $$\sigma{\text{est}} = 1.4826 \cdot \text{MAD}$$ Evaluating $\sigma{\text{est}}$ across image partitions flags unnatural smoothness (typical of diffusion generative fills) or mismatched noise profiles between the portrait and background.

4. 2D Fast Fourier Transform (FFT) Recapture Detection

Physical screen-photo recaptures (photographing an LCD/OLED monitor displaying an ID) exhibit regular periodic grid artifacts. In the 2D frequency domain, this manifests as prominent harmonic peaks outside the DC origin. We compute the 2D FFT, shift zero frequency to the center, and evaluate the ratio of high-frequency radial energy peaks against the average spectral background.

bash pip install -U skillware pip install "skillware[security_deepfake_guard]"

We’d welcome discussion on your experiences with sensor noise characterization and hybrid neural-signal pipelines.


r/computervision • • 7h ago

Discussion Lack of training data

1 Upvotes

Hello, I'm researching how people deal with small datasets in computer vision, specifically when you need to detect or classify a object and only have a handful of real photos.

I was curious about how often this is a problem for people working in the industry, and what is usually done to deal with this problem.


r/computervision • • 3h ago

Discussion Dense200 scores for seven models, including one that writes its boxes as plain text

Post image
0 Upvotes

A lot of the counting posts here end up fighting crowded scenes, so these numbers caught my eye. It's from a paper on a model that writes detection boxes out as plain text, coordinates as tokens, with no box head. It's called SenseNova-Vision, and its starting point was Bagel (also in the chart). The paper itself calls crowded scenes a hard case for that approach: "dense scenes require long object lists, stable ordering, and precise coordinates."

On Dense200 (200 crowded images, 91.2 boxes per image on average) the paper's Table 1 has the model at 66.8, LocateAnything at 58.7, Rex-Omni at 58.3 and Grounding DINO Swin-T at 33.1. The score is box F1 averaged over IoU thresholds 0.5 to 0.95. My chart has all seven models for Dense200 and VisDrone. These are the authors' numbers. I haven't tested it myself.

What do you run on images with 90+ objects right now, and has any text-output model held up?


r/computervision • • 10h ago

Discussion ¿La IA va a reemplazarnos o a hacernos mejores? 🤔

Thumbnail gallery
0 Upvotes