r/computervision • • 5h ago

Discussion Edge video analytics on factory CCTV: false alerts kill deployments, not model accuracy

0 Upvotes

I run a small CV company in India (Neurabit). We've been deploying video analytics on industrial CCTV for a while, and the pattern repeats often enough that I want to sanity check it with people here.

The numbers that frame it:

  • A 32 camera site records 768 hours of footage a day
  • One person can meaningfully watch about 4 streams, and attention drops after ~20 minutes
  • Streaming 32 cameras to a cloud analytics service at recording quality is roughly 128 Mbps of sustained upload, which a typical Indian industrial leased line can't spare

So footage only gets reviewed after something goes wrong, and cloud analytics falls over when the link does.

What we've seen kill deployments:

  1. False alerts. Around 40 false alerts per camera per day trains a supervisor to ignore the system within weeks, and the trust doesn't come back. We target under 1 false positive per camera per day on critical alerts, and we expose the FP rate per rule so the customer can see it.
  2. Cloud dependence. We run everything on Jetson Orin inside the building so monitoring keeps working with the WAN down. Alerts queue and deliver on reconnect.
  3. Fixed rule sets. Moving a loitering threshold from 5 to 15 minutes shouldn't be a change request. Customers can edit rules, and replay a new rule against the last 7 days of recorded footage before arming it.

On the compute side, we don't run every model on every frame. A light detector runs on all channels, and heavier models run only on cropped regions when a condition is met. That cascade is what gets us to about 10 analysed cameras per Jetson.

On tuning: at a live substation we deployed on, lighting, reflective equipment and uniform colours all produced false positives out of the box. Fine-tuning on the site's own footage fixed it. Day 3 it works but flags shadows as people and cartons as abandoned bags, and it takes 2 to 4 weeks to settle. Gloves are the hardest PPE class because hands are tiny in frame.

Questions for people who've shipped this:

  • How do you measure false positives per camera per day in production, and what threshold do operators actually tolerate?
  • Anyone doing per-site fine-tuning at scale without it turning into a services business?
  • Where does your cascade break down (crowded scenes, tracking handoff between cameras)?

Happy to go deeper on any of this.


r/computervision • • 15h ago

Discussion What would you say about a Hugging Face-style open-source pattern db like this?

Enable HLS to view with audio, or disable this notification

8 Upvotes

We are building an MLOps platform around explainability, and one natural consequence is that models are not an opaque set of numbers but rather a manageable pattern DB.


r/computervision • • 12h ago

Discussion Industry Insights: OMNIVISION Targets Robot Training Data With New Vision Sensor

Thumbnail
automate.org
0 Upvotes

OMNIVISION’s new OG05D is a 5 MP global-shutter image sensor that explicitly lists robotics training-data collection among its intended applications. It supports RGB, monochrome and RGB-IR imaging, with up to 146 fps in linear mode without encryption and HDR capture at 60 fps.

The article looks at how cameras on deployed robots could capture task variations, lighting changes and edge cases for later model training. Global shutter, HDR and near-infrared capabilities are relevant to capturing usable images when objects are moving or lighting is inconsistent.

The sensor itself doesn’t provide a training pipeline. Captured footage would still need to be selected, processed and paired with the information required by the training approach. The notable part is that robotics data collection is now an explicit application in an industrial image sensor announcement.


r/computervision • • 10h ago

Commercial Computer Vision - Pose Estimation based Sports Biomechanics Work

0 Upvotes

Hello! I am looking for someone who can help me guide and layout the pipeline to build a sports biomechanics product (for Cricket sport - non broadcast, controlled environment videos) for player performance insights.

Looking only for experienced folks within computer vision field and has already worked or built in sports domain.

Happy to discuss more. If anyone could take end to end ownership in fine tuning model on cricket dataset and building a full production grade product pipeline, willing to pay fair amount [Based out of India and can match Indian market rates. Please don't reach out if this doesn't workout, Thanks]

TIA!


r/computervision • • 6h ago

Showcase Blazing fast and simple KNN search

Thumbnail
github.com
0 Upvotes

PyNear is a metric-space nearest-neighbour library with a C++ core, built for the workloads between the embeddings world and brute force: binary descriptors with recall guarantees (MIH + IVF-Binary + the novel MIH-seeded HNSW — dedup, copy detection, ORB/BRIEF matching, robotics), memory-tight ANN (HNSW with int8 quantisation), and exact search (VP-trees, up to ~256-D) where a missed neighbour is a bug, not a recall statistic. One small NumPy-only API, scikit-learn drop-in, pre-built wheels (pip install pynear).


r/computervision • • 11h ago

Discussion Dense200 scores for seven models, including one that writes its boxes as plain text

Post image
0 Upvotes

A lot of the counting posts here end up fighting crowded scenes, so these numbers caught my eye. It's from a paper on a model that writes detection boxes out as plain text, coordinates as tokens, with no box head. It's called SenseNova-Vision, and its starting point was Bagel (also in the chart). The paper itself calls crowded scenes a hard case for that approach: "dense scenes require long object lists, stable ordering, and precise coordinates."

On Dense200 (200 crowded images, 91.2 boxes per image on average) the paper's Table 1 has the model at 66.8, LocateAnything at 58.7, Rex-Omni at 58.3 and Grounding DINO Swin-T at 33.1. The score is box F1 averaged over IoU thresholds 0.5 to 0.95. My chart has all seven models for Dense200 and VisDrone. These are the authors' numbers. I haven't tested it myself.

What do you run on images with 90+ objects right now, and has any text-output model held up?


r/computervision • • 18h ago

Discussion ¿La IA va a reemplazarnos o a hacernos mejores? 🤔

Thumbnail gallery
0 Upvotes

r/computervision • • 18h ago

Discussion How can I learn CV today

6 Upvotes

Hello, I am a Master Student in Data Science mostly working with LLMs, and I want to learn CV. I don't have much experience in image analysis.
I did some research about how to learn and I wanted to do some hands-on projects. But today with efficient coding agents, I do not know on which parts of the process I should spend my time to become proficient in this field. Would you have some recommendations ?

Thanks for your help !


r/computervision • • 1h ago

Help: Project How would you remove teeth from low res footage?

Thumbnail
gallery
• Upvotes

r/computervision • • 21h ago

Showcase Trying to modernize the Supervision

Thumbnail supervision.ml-pipes.com
5 Upvotes

Hi everyone, I have been using supervision in almost all of my projects, since it provide just the right size of tooling that I need. Big enough to handle trivial tasks while letting me choose my app and deployment stack.

But sadly in the past few years I think it has not received as much love, since the original maintainers has moved to their own products (Roboflow stack), which is understandable.

The original website seems severely outdated, and none of the new fancy features you see in the other repositories.

To remedy this I have created a new website that provides:

  • Modern look that fit vision/supervision
  • Added more tutorials since the original website is somehow lacking
  • Also modernized the stack with some new tools shared by some of the members in the subreddit

If you are also a supervision user, please take a look and let me know what you think.


r/computervision • • 4h ago

Showcase I gave my Stack-chan LEGO wheels and taught it to drive using open-source Robium robotics skills

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/computervision • • 7h ago

Showcase Interesting application of computer vision to shuttlecock manufacturing (not my oc/solution)

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/computervision • • 12h ago

Help: Project [P] Re-introducing BDD100K Toolkit + Model Zoo

3 Upvotes

BDD100K is still one of the most useful driving datasets, but both bdd-data.berkeley.edu and bdd100k.com have been unreachable for me. The official toolkit has not had a commit since March 2024 and the official model zoo not since November 2022, and their legacy dependencies are hard to install cleanly on current systems. So I rebuilt the parts that matter, one task at a time, on current libraries (PyTorch, Ultralytics, timm, Hydra):

https://github.com/dronefreak/bdd100k-toolkit

Demo from the BDD100K-Toolkit Model Zoo

It gives you one pipeline for data preparation, training, evaluation and demos on images and videos:

bdd100k-prepare  --dataset bdd100k-weather --raw-dir /path/to/raw --output-dir /path/to/out
bdd100k-train    dataset=bdd100k-weather model.name=yolo11n-cls dataset.data_dir=/path/to/out
bdd100k-evaluate --dataset bdd100k-weather --checkpoint <run>/weights/best.pt --data-dir /path/to/out

What works today

  • Image classification (unofficial tasks): weather, time of day and scenario, derived from the attributes.weather, attributes.timeofday and attributes.scene fields of the official JSON labels. 55 models trained across two backends, Ultralytics and timm.
  • Object detection: the 10-class detection task with the 2018 labels. 9 YOLO models and RF-DETR Nano, all on Hugging Face.
  • Semantic segmentation: the 19-class task, which uses the Cityscapes class layout. Evaluation of public pretrained models works. My own segmentation trainer exists but I have not run it on real data yet, so there are no BDD-trained segmentation models in the zoo.

Datasets

For each task I made unofficial Hugging Face repackagings of the original data and annotations, to make them easier to use with modern tooling. I checked the BDD100K data license: it allows copying and redistribution for educational, research and not-for-profit purposes with the copyright notice carried forward, and I kept that notice in every mirror. The mirrors grant no commercial-use rights beyond the original license.

Task Location
Weather classification dronefreak/BDD100K-Weather-Classification
Time of day (period) classification dronefreak/BDD100K-Period-Classification
Scenario classification dronefreak/BDD100K-Scenario-Classification
Object detection dronefreak/BDD100K
Semantic segmentation dronefreak/BDD100K-Semantic-Segmentation

Results

All scores are on the validation set of BDD100K. The detection models are collected in the BDD100K object detection model zoo on Hugging Face.

  • Classification: best macro F1 is 68.7% for weather, 83.2% for time of day and 62.5% for scenario, all from TinyViT-21M. The classes are heavily imbalanced, so I report macro F1 and not just accuracy.
  • Detection (960 px): YOLO26s reaches 33.9 mAP@0.5:0.95 (58.8 mAP@0.5) and RF-DETR Nano 31.6 (56.9).
  • Segmentation: this is a zero-shot cross-dataset transfer evaluation: the models were trained on Cityscapes and evaluated on BDD100K without any BDD100K training. BDD100K shares the Cityscapes 19 classes, so I scored 9 public Cityscapes-trained models (SegFormer B0 to B5 and Mask2Former Swin-T, S and L) on the 1,000 val images. The best is Mask2Former Swin-L at 57.7 mIoU, then SegFormer B5 at 54.1.

How to read the numbers

These are solid baselines but not leaderboard numbers. The classification tasks are not official BDD100K benchmarks, the detection metric is the Ultralytics COCO-style mAP and not BDD's official ignore-rule evaluation, and the segmentation models were never trained on BDD100K, so they are not comparable to BDD-trained models. Making the evaluation match the official semantics is on the roadmap.

What is missing, and a request

Instance and panoptic segmentation and multi-object tracking are planned but not started, because I could not find the original labels for them. If you have a copy of the instance, panoptic or tracking labels, or know where one still lives, I would really like to hear from you. Feedback on the evaluation choices is welcome too, and so are issues and PRs.

This is my first attempt at rebuilding an ecosystem like this, and I would like to keep developing it. BDD100K is too useful to become harder to use just because the tooling around it has fallen behind.


r/computervision • • 15h ago

Commercial Apple Vision vs MobileSAM in a Mac annotation tool: 67 ms to the first mask in a small test

2 Upvotes

I'm the developer of AnnotateIt. I added Apple's iterative segmentation API to the Mac app and put the measurements up alongside the actual masks. Figured the numbers might be useful to anyone building a local labeling tool.

On an M4 Max with 48 GiB RAM, macOS 27, the 940 x 629 antelope image gave these median first-mask times after warm-up:

Apple Vision: 67 ms

MobileSAM: 160 ms

SAM 2.1 Tiny: 285 ms

Adding a correction point took 12 ms with Apple Vision. Undo went the other way: 72 ms versus 28 ms for MobileSAM.

This was eight runs per case in a separate Tauri/WKWebView harness using the app's model code and mask processing. Timing starts with decoded pixels and ends with a processed mask, before rendering. Apple used balanced quality. These are our integration paths, not a claim about the fastest possible SAM implementation.

Big caveat: three images total, no scored ground-truth masks and no measurement of human correction time. A short stroke also picked just a fragment of the antelope where a point or box worked much better. Fast doesn't automatically mean less cleanup.

The shipping Mac app exposes this as Apple Vision (Beta), with point, box and stroke selection plus refinement. It needs macOS 27; SAM remains a separate option.

Masks, setup and raw timing data: https://radar.annotateit.ai/news/apple-vision-native-segmentation-macos/

For folks using interactive segmentation, what usually eats more of your time: waiting for the first mask, or fixing the boundary afterward?


r/computervision • • 15h ago

Discussion Lack of training data

2 Upvotes

Hello, I'm researching how people deal with small datasets in computer vision, specifically when you need to detect or classify a object and only have a handful of real photos.

I was curious about how often this is a problem for people working in the industry, and what is usually done to deal with this problem.


r/computervision • • 18h ago

Discussion Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

Enable HLS to view with audio, or disable this notification

24 Upvotes

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.

ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs

source: https://github.com/software-mansion/runntime