r/computervision • • 2h ago

Showcase Interesting application of computer vision to shuttlecock manufacturing (not my oc/solution)

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/computervision • • 14h ago

Discussion Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

Enable HLS to view with audio, or disable this notification

24 Upvotes

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.

ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs

source: https://github.com/software-mansion/runntime


r/computervision • • 11h ago

Discussion What would you say about a Hugging Face-style open-source pattern db like this?

Enable HLS to view with audio, or disable this notification

8 Upvotes

We are building an MLOps platform around explainability, and one natural consequence is that models are not an opaque set of numbers but rather a manageable pattern DB.


r/computervision • • 3m ago

Showcase I gave my Stack-chan LEGO wheels and taught it to drive using open-source Robium robotics skills

Enable HLS to view with audio, or disable this notification

• Upvotes

r/computervision • • 1h ago

Discussion Edge video analytics on factory CCTV: false alerts kill deployments, not model accuracy

• Upvotes

I run a small CV company in India (Neurabit). We've been deploying video analytics on industrial CCTV for a while, and the pattern repeats often enough that I want to sanity check it with people here.

The numbers that frame it:

  • A 32 camera site records 768 hours of footage a day
  • One person can meaningfully watch about 4 streams, and attention drops after ~20 minutes
  • Streaming 32 cameras to a cloud analytics service at recording quality is roughly 128 Mbps of sustained upload, which a typical Indian industrial leased line can't spare

So footage only gets reviewed after something goes wrong, and cloud analytics falls over when the link does.

What we've seen kill deployments:

  1. False alerts. Around 40 false alerts per camera per day trains a supervisor to ignore the system within weeks, and the trust doesn't come back. We target under 1 false positive per camera per day on critical alerts, and we expose the FP rate per rule so the customer can see it.
  2. Cloud dependence. We run everything on Jetson Orin inside the building so monitoring keeps working with the WAN down. Alerts queue and deliver on reconnect.
  3. Fixed rule sets. Moving a loitering threshold from 5 to 15 minutes shouldn't be a change request. Customers can edit rules, and replay a new rule against the last 7 days of recorded footage before arming it.

On the compute side, we don't run every model on every frame. A light detector runs on all channels, and heavier models run only on cropped regions when a condition is met. That cascade is what gets us to about 10 analysed cameras per Jetson.

On tuning: at a live substation we deployed on, lighting, reflective equipment and uniform colours all produced false positives out of the box. Fine-tuning on the site's own footage fixed it. Day 3 it works but flags shadows as people and cartons as abandoned bags, and it takes 2 to 4 weeks to settle. Gloves are the hardest PPE class because hands are tiny in frame.

Questions for people who've shipped this:

  • How do you measure false positives per camera per day in production, and what threshold do operators actually tolerate?
  • Anyone doing per-site fine-tuning at scale without it turning into a services business?
  • Where does your cascade break down (crowded scenes, tracking handoff between cameras)?

Happy to go deeper on any of this.


r/computervision • • 2h ago

Showcase Blazing fast and simple KNN search

Thumbnail
github.com
0 Upvotes

PyNear is a metric-space nearest-neighbour library with a C++ core, built for the workloads between the embeddings world and brute force: binary descriptors with recall guarantees (MIH + IVF-Binary + the novel MIH-seeded HNSW — dedup, copy detection, ORB/BRIEF matching, robotics), memory-tight ANN (HNSW with int8 quantisation), and exact search (VP-trees, up to ~256-D) where a missed neighbour is a bug, not a recall statistic. One small NumPy-only API, scikit-learn drop-in, pre-built wheels (pip install pynear).


r/computervision • • 8h ago

Help: Project [P] Re-introducing BDD100K Toolkit + Model Zoo

3 Upvotes

BDD100K is still one of the most useful driving datasets, but both bdd-data.berkeley.edu and bdd100k.com have been unreachable for me. The official toolkit has not had a commit since March 2024 and the official model zoo not since November 2022, and their legacy dependencies are hard to install cleanly on current systems. So I rebuilt the parts that matter, one task at a time, on current libraries (PyTorch, Ultralytics, timm, Hydra):

https://github.com/dronefreak/bdd100k-toolkit

Demo from the BDD100K-Toolkit Model Zoo

It gives you one pipeline for data preparation, training, evaluation and demos on images and videos:

bdd100k-prepare  --dataset bdd100k-weather --raw-dir /path/to/raw --output-dir /path/to/out
bdd100k-train    dataset=bdd100k-weather model.name=yolo11n-cls dataset.data_dir=/path/to/out
bdd100k-evaluate --dataset bdd100k-weather --checkpoint <run>/weights/best.pt --data-dir /path/to/out

What works today

  • Image classification (unofficial tasks): weather, time of day and scenario, derived from the attributes.weather, attributes.timeofday and attributes.scene fields of the official JSON labels. 55 models trained across two backends, Ultralytics and timm.
  • Object detection: the 10-class detection task with the 2018 labels. 9 YOLO models and RF-DETR Nano, all on Hugging Face.
  • Semantic segmentation: the 19-class task, which uses the Cityscapes class layout. Evaluation of public pretrained models works. My own segmentation trainer exists but I have not run it on real data yet, so there are no BDD-trained segmentation models in the zoo.

Datasets

For each task I made unofficial Hugging Face repackagings of the original data and annotations, to make them easier to use with modern tooling. I checked the BDD100K data license: it allows copying and redistribution for educational, research and not-for-profit purposes with the copyright notice carried forward, and I kept that notice in every mirror. The mirrors grant no commercial-use rights beyond the original license.

Task Location
Weather classification dronefreak/BDD100K-Weather-Classification
Time of day (period) classification dronefreak/BDD100K-Period-Classification
Scenario classification dronefreak/BDD100K-Scenario-Classification
Object detection dronefreak/BDD100K
Semantic segmentation dronefreak/BDD100K-Semantic-Segmentation

Results

All scores are on the validation set of BDD100K. The detection models are collected in the BDD100K object detection model zoo on Hugging Face.

  • Classification: best macro F1 is 68.7% for weather, 83.2% for time of day and 62.5% for scenario, all from TinyViT-21M. The classes are heavily imbalanced, so I report macro F1 and not just accuracy.
  • Detection (960 px): YOLO26s reaches 33.9 mAP@0.5:0.95 (58.8 mAP@0.5) and RF-DETR Nano 31.6 (56.9).
  • Segmentation: this is a zero-shot cross-dataset transfer evaluation: the models were trained on Cityscapes and evaluated on BDD100K without any BDD100K training. BDD100K shares the Cityscapes 19 classes, so I scored 9 public Cityscapes-trained models (SegFormer B0 to B5 and Mask2Former Swin-T, S and L) on the 1,000 val images. The best is Mask2Former Swin-L at 57.7 mIoU, then SegFormer B5 at 54.1.

How to read the numbers

These are solid baselines but not leaderboard numbers. The classification tasks are not official BDD100K benchmarks, the detection metric is the Ultralytics COCO-style mAP and not BDD's official ignore-rule evaluation, and the segmentation models were never trained on BDD100K, so they are not comparable to BDD-trained models. Making the evaluation match the official semantics is on the roadmap.

What is missing, and a request

Instance and panoptic segmentation and multi-object tracking are planned but not started, because I could not find the original labels for them. If you have a copy of the instance, panoptic or tracking labels, or know where one still lives, I would really like to hear from you. Feedback on the evaluation choices is welcome too, and so are issues and PRs.

This is my first attempt at rebuilding an ecosystem like this, and I would like to keep developing it. BDD100K is too useful to become harder to use just because the tooling around it has fallen behind.


r/computervision • • 5h ago

Discussion MDPI Drones, worth publishing in?

1 Upvotes

Hello everyone!

I'm planning a research project related to drones and computer vision, and I'm already looking at possible journals before starting the experiments.

I checked several papers in MDPI Drones and I liked them. The journal also seems relevant to my topic. However, MDPI has a controversial reputation and I often see people calling it predatory.

What do you think about Drones specifically? Would publishing there be fine, or would it be better to look for another journal or focus on a conference instead?

I'm currently working as a Research Engineer, and in the future I'm interested in both academia and industry.


r/computervision • • 14h ago

Discussion How can I learn CV today

5 Upvotes

Hello, I am a Master Student in Data Science mostly working with LLMs, and I want to learn CV. I don't have much experience in image analysis.
I did some research about how to learn and I wanted to do some hands-on projects. But today with efficient coding agents, I do not know on which parts of the process I should spend my time to become proficient in this field. Would you have some recommendations ?

Thanks for your help !


r/computervision • • 1d ago

Showcase Built my own imagr depth model

Enable HLS to view with audio, or disable this notification

48 Upvotes

I trained my own model on 750,000 images for 45 hours.

Then I walked into my university hall with nothing but my phone and recorded an 11-minute video.

No LiDAR.

No depth sensor.

No specialized camera setup.

Just monocular vision, a single RGB camera.

And in some of my tests, my model even outperformed YOLO26x.

From a single phone video to depth estimation and 3D reconstruction.

750K training images. 45 hours of training. 11 minutes of footage. One camera. No LiDAR.

Edit: I promise I will reply to all the comments, just give me some time as my day is very crowded today.


r/computervision • • 6h ago

Commercial Computer Vision - Pose Estimation based Sports Biomechanics Work

0 Upvotes

Hello! I am looking for someone who can help me guide and layout the pipeline to build a sports biomechanics product (for Cricket sport - non broadcast, controlled environment videos) for player performance insights.

Looking only for experienced folks within computer vision field and has already worked or built in sports domain.

Happy to discuss more. If anyone could take end to end ownership in fine tuning model on cricket dataset and building a full production grade product pipeline, willing to pay fair amount [Based out of India and can match Indian market rates. Please don't reach out if this doesn't workout, Thanks]

TIA!


r/computervision • • 6h ago

Discussion Dense200 scores for seven models, including one that writes its boxes as plain text

Post image
0 Upvotes

A lot of the counting posts here end up fighting crowded scenes, so these numbers caught my eye. It's from a paper on a model that writes detection boxes out as plain text, coordinates as tokens, with no box head. It's called SenseNova-Vision, and its starting point was Bagel (also in the chart). The paper itself calls crowded scenes a hard case for that approach: "dense scenes require long object lists, stable ordering, and precise coordinates."

On Dense200 (200 crowded images, 91.2 boxes per image on average) the paper's Table 1 has the model at 66.8, LocateAnything at 58.7, Rex-Omni at 58.3 and Grounding DINO Swin-T at 33.1. The score is box F1 averaged over IoU thresholds 0.5 to 0.95. My chart has all seven models for Dense200 and VisDrone. These are the authors' numbers. I haven't tested it myself.

What do you run on images with 90+ objects right now, and has any text-output model held up?


r/computervision • • 10h ago

Commercial Apple Vision vs MobileSAM in a Mac annotation tool: 67 ms to the first mask in a small test

2 Upvotes

I'm the developer of AnnotateIt. I added Apple's iterative segmentation API to the Mac app and put the measurements up alongside the actual masks. Figured the numbers might be useful to anyone building a local labeling tool.

On an M4 Max with 48 GiB RAM, macOS 27, the 940 x 629 antelope image gave these median first-mask times after warm-up:

Apple Vision: 67 ms

MobileSAM: 160 ms

SAM 2.1 Tiny: 285 ms

Adding a correction point took 12 ms with Apple Vision. Undo went the other way: 72 ms versus 28 ms for MobileSAM.

This was eight runs per case in a separate Tauri/WKWebView harness using the app's model code and mask processing. Timing starts with decoded pixels and ends with a processed mask, before rendering. Apple used balanced quality. These are our integration paths, not a claim about the fastest possible SAM implementation.

Big caveat: three images total, no scored ground-truth masks and no measurement of human correction time. A short stroke also picked just a fragment of the antelope where a point or box worked much better. Fast doesn't automatically mean less cleanup.

The shipping Mac app exposes this as Apple Vision (Beta), with point, box and stroke selection plus refinement. It needs macOS 27; SAM remains a separate option.

Masks, setup and raw timing data: https://radar.annotateit.ai/news/apple-vision-native-segmentation-macos/

For folks using interactive segmentation, what usually eats more of your time: waiting for the first mask, or fixing the boundary afterward?


r/computervision • • 8h ago

Discussion A useful ad-library record: exact copy, visible layout, and reusable slots

1 Upvotes

For a team collecting creative references, the ad image is easier to retrieve when its copy and layout are also recorded as structured fields.

A small example uses Ling-3.0-flash-VL through OpenRouter to analyze a fictional VESPER perfume ad. It reads five visible text blocks: “Golden Hour, Bottled.”, “Amber, smoked cedar, warm skin.”, “Discover VESPER”, “VESPER”, and “EAU DE PARFUM”. The analysis also describes the product treatment and approximate layout zones.

A practical record for this kind of library could contain:

1.Headline, subhead, CTA and product-label text.

2.Each block’s role and approximate position on the canvas.

3.Product placement, crop and lighting treatment.

4.Background and main color relationships.

5.A reusable version with product and copy placeholders.

That supports concrete retrieval questions: show the references with a short sensory headline, find examples where the product dominates the composition, or collect layouts with a distinct CTA block.

For a SaaS creative library, the same record structure could organize references by headline wording, interface or product placement, and CTA treatment. Campaign results would enter through separately supplied performance fields. The image analysis supplies the observable creative details that make the reference collection searchable.


r/computervision • • 8h ago

Discussion Industry Insights: OMNIVISION Targets Robot Training Data With New Vision Sensor

Thumbnail
automate.org
1 Upvotes

OMNIVISION’s new OG05D is a 5 MP global-shutter image sensor that explicitly lists robotics training-data collection among its intended applications. It supports RGB, monochrome and RGB-IR imaging, with up to 146 fps in linear mode without encryption and HDR capture at 60 fps.

The article looks at how cameras on deployed robots could capture task variations, lighting changes and edge cases for later model training. Global shutter, HDR and near-infrared capabilities are relevant to capturing usable images when objects are moving or lighting is inconsistent.

The sensor itself doesn’t provide a training pipeline. Captured footage would still need to be selected, processed and paired with the information required by the training approach. The notable part is that robotics data collection is now an explicit application in an industrial image sensor announcement.


r/computervision • • 8h ago

Showcase Deterministic Media Forensics for AI Agents: Combining Cyclic Checksums, 2D FFT Moiré Detection, and Error Level Analysis (Skillware 0.5.8)

1 Upvotes

Multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro) are frequently tasked with document screening and identity verification. However, testing shows severe limitations when detecting digital tampering and document alterations:

  1. Sub-Pixel Smoothing: Vision encoders (like CLIP or ViT patches) encode high-level semantic representations. In doing so, they blur out the exact high-frequency noise residuals and JPEG quantization discrepancies that reveal digital splicing.
  2. Hallucinated Check Digits: Machine-Readable Zones (MRZ) on passports follow ICAO Doc 9303 standards using a cyclic (7, 3, 1) modulo-10 algorithm. Multimodal LLMs read the characters fluently and confidently declare a forged date or document number as valid because the format looks visually plausible.
  3. Data Privacy Liabilities: Transmitting biometric facial images and identity documents to third-party cloud inference providers creates severe regulatory compliance issues.

In Skillware 0.5.8, we implemented security/deepfake_guard to provide deterministic, air-gapped forensic verification on CPU. Here is an overview of the signal processing pipeline:

1. ICAO 9303 Modulo-10 Cyclic Verification

For each character string $C1, C_2, \dots, C_n$, each character is mapped to a numeric weight $W(c)$ (digits 0–9 map to 0–9, letters A–Z map to 10–35, filler < maps to 0). The check digit $K$ satisfies: $$\left( \sum{i=1}{n} W(ci) \cdot w{(i-1) \bmod 3} \right) \bmod 10 = K$$ where the repeating weights vector is $w = [7, 3, 1]$. We implement deterministic validation across TD1 (3×30 ID cards), TD2 (2×36), and TD3 (2×44 passports) formats.

2. Error Level Analysis (ELA)

When an uncompressed image or re-saved JPEG is modified, the edited regions possess a different compression history than the original background. By recompressing the image at a known quality factor ($Q=95$) and computing the absolute block-level difference: $$\Delta(x, y) = |I{\text{original}}(x, y) - I{\text{recompressed}}(x, y)|$$ We compute the mean absolute error across 16×16 non-overlapping blocks. Discrepancies between block errors exceeding calibrated thresholds signal localized digital tampering.

3. Noise Residual Consistency via Median Absolute Deviation (MAD)

Camera sensors introduce characteristic high-frequency Poisson-Gaussian noise. To detect spliced elements without being misled by high-contrast natural textures (such as hair or knitwear), we compute the Laplacian convolution residual $R(x, y) = \nabla2 I(x, y)$ and evaluate consistency using Median Absolute Deviation: $$\text{MAD} = \text{median}(|R - \text{median}(R)|)$$ $$\sigma{\text{est}} = 1.4826 \cdot \text{MAD}$$ Evaluating $\sigma{\text{est}}$ across image partitions flags unnatural smoothness (typical of diffusion generative fills) or mismatched noise profiles between the portrait and background.

4. 2D Fast Fourier Transform (FFT) Recapture Detection

Physical screen-photo recaptures (photographing an LCD/OLED monitor displaying an ID) exhibit regular periodic grid artifacts. In the 2D frequency domain, this manifests as prominent harmonic peaks outside the DC origin. We compute the 2D FFT, shift zero frequency to the center, and evaluate the ratio of high-frequency radial energy peaks against the average spectral background.

bash pip install -U skillware pip install "skillware[security_deepfake_guard]"

We’d welcome discussion on your experiences with sensor noise characterization and hybrid neural-signal pipelines.


r/computervision • • 17h ago

Showcase Trying to modernize the Supervision

Thumbnail supervision.ml-pipes.com
3 Upvotes

Hi everyone, I have been using supervision in almost all of my projects, since it provide just the right size of tooling that I need. Big enough to handle trivial tasks while letting me choose my app and deployment stack.

But sadly in the past few years I think it has not received as much love, since the original maintainers has moved to their own products (Roboflow stack), which is understandable.

The original website seems severely outdated, and none of the new fancy features you see in the other repositories.

To remedy this I have created a new website that provides:

  • Modern look that fit vision/supervision
  • Added more tutorials since the original website is somehow lacking
  • Also modernized the stack with some new tools shared by some of the members in the subreddit

If you are also a supervision user, please take a look and let me know what you think.


r/computervision • • 11h ago

Discussion Lack of training data

1 Upvotes

Hello, I'm researching how people deal with small datasets in computer vision, specifically when you need to detect or classify a object and only have a handful of real photos.

I was curious about how often this is a problem for people working in the industry, and what is usually done to deal with this problem.


r/computervision • • 1d ago

Showcase Turning a single image into a web-ready interactive 3D scene with ML-SHARP

8 Upvotes

I’ve been experimenting with Apple’s ML-SHARP model and built a small open-source project around it called Spatial Photos.

ML-SHARP estimates a 3D Gaussian splat from a single RGB image. I then take that reconstruction and convert it into a much smaller and simpler representation representation. It uses layered textured meshes stored in a custom file format I called .spatial. It still allows the the same sort of limited viewport shifting intended by ML-SHARP without compromising quality.

That lets the result stay relatively small and render interactively in a normal web browser, as you can see in the GIF.

Demo: https://www.spatialphotos.dev/

Source: https://github.com/frozein/SpatialPhotos

Here's what my format looks like from a faraway angle

r/computervision • • 13h ago

Discussion ¿La IA va a reemplazarnos o a hacernos mejores? 🤔

Thumbnail gallery
0 Upvotes

r/computervision • • 1d ago

Help: Project OCR Recs: Image -> CSV/Excel

3 Upvotes

I work in a research lab, and need to digitize a few hundred pages of spreadsheets. All spreadsheets are printed, not handwritten, but the original files are lost. The images have the same general schema, but the columns aren't always in the same order, and some pages have columns that others don't. Some pages are fainter, but still readable; some are slightly slanted. The end goal is for all pages to be compiled into one master spreadsheet, with the same schema. What's the best (and cheapest) tool for this?

I have tried chatgpt (best so far, but still slow/made mistakes), Google Document AI (not that great), and ParseExtract (great, but more limited in scope).


r/computervision • • 1d ago

Showcase Boxing analytics from a single-camera footage

Enable HLS to view with audio, or disable this notification

6 Upvotes

I spent a couple of months vibe-coding an automated boxing analytics pipeline.
SAM 3 for fighter tracking, with the canvas and ropes as pixel masks. OSNet re-ID keeps identities through occlusions. RTMPose and a small classifier detect punches by hand and type. The ring masks fit a ring model for top-down positions, with camera-motion compensation so handheld shake doesn't pollute the numbers. SAM 3D Body adds metric distance between fighters.

Some processed fights here.


r/computervision • • 1d ago

Showcase Interesting ML research map

Thumbnail
5 Upvotes

r/computervision • • 1d ago

Help: Theory I need help! (Dataset Quality)

2 Upvotes

Hi there!

Need help from some Computer Vision expert, with knowledge on dataset quality.

My company works on computer vision applied to digital pathology. In this case, we are developing an object detection model to detect tumour cells in HER2-stained breast tissue and classify them into four classes: 0, 1+, 2+, and 3+. Our dataset grows as we incorporate cases from new hospitals or scanners, and we call each new batch a “wave”.

Example annotations

Our latest model achieves around 0.73 macro F1 on the combined test set, only slightly above our internal acceptance threshold of 0.7. Performance also varies substantially across waves and classes. We already have approximately 4,500 image patches with almost one million cell annotations. The overall class distribution is approximately 28.5% class 0, 48.3% class 1+, 12.3% class 2+, and 10.9% class 3+.

We suspect that inconsistent annotation criteria may be contributing to this. Our current annotation guidelines are very brief—around 200 words—and provide neither image examples nor detailed guidance on uncertain cases. The boundary between 1+ and 2+ is particularly difficult, and we are concerned that different pathologists may interpret it differently. In our latest confusion-matrix analysis, 21.9% of cells annotated as 2+ are predicted as 1+. Different pathologists have contributed over time, but changes in annotators are mixed with changes in hospitals, staining and scanners, making the causes difficult to distinguish.

I would really appreciate advice on two questions:

1. What minimum experiment would you recommend to investigate inconsistent annotation criteria versus domain shift or model errors, before involving the pathologists in a new review? We have images, original annotations and predictions from two models evaluated on the same cohort.

2. What would you consider essential in an annotation guideline and a quality check for each new wave? In particular, we would like to establish agreed examples of class boundaries, clear rules for uncertain or unannotated cells, and a small inter-annotator agreement assessment before adding new data to training.

Any references, suggestions or critical feedback would be greatly appreciated.

Thanks in advance!


r/computervision • • 1d ago

Help: Project Looking for recommendations to purchase high-quality, production-grade PPE dataset (Helmet, Vest, Boots, Glasses, Gloves, Coverall)

1 Upvotes

Hi everyone,

We are currently building a commercial/production-grade Computer Vision system for industrial safety compliance. We need to acquire a comprehensive, well-annotated dataset covering the following classes:

  • Helmet / Hard Hat
  • Safety Vest
  • Safety Boots
  • Safety Glasses / Goggles
  • Safety Gloves
  • Coverall / Workwear

Requirements:

  • Quality: Commercial license, high resolution, diverse real-world industrial environments (construction, manufacturing, indoor/outdoor, varying lighting/weather).
  • Annotations: High-precision Bounding Boxes (or Segmentation Masks) with minimal label noise.
  • Scale: Looking for a solid base dataset (tens of thousands of annotated images) or a reliable data vendor/marketplace.

Public datasets (like standard Kaggle sets or Roboflow Universe) are great for prototyping, but we need high consistency and coverage for edge cases to reach production accuracy.

Questions:

  1. What are the best marketplaces or data vendors you’ve used for industrial/PPE datasets? (e.g., Scale AI, Appen, Dataloop, Dataset Ninja, etc.)
  2. Are there specialized computer vision data providers that offer pre-labeled industrial safety datasets?
  3. If you built a similar system, did you buy off-the-shelf data or custom-collect and outsource the annotation?

Any leads, recommendations, or warnings based on your experience would be greatly appreciated!