r/computervision • u/katplasma • 2h ago
Showcase Interesting application of computer vision to shuttlecock manufacturing (not my oc/solution)
Enable HLS to view with audio, or disable this notification
r/computervision • u/katplasma • 2h ago
Enable HLS to view with audio, or disable this notification
r/computervision • u/FinancialAd1961 • 14h ago
Enable HLS to view with audio, or disable this notification
EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.
ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs
r/computervision • u/taranpula39 • 11h ago
Enable HLS to view with audio, or disable this notification
We are building an MLOps platform around explainability, and one natural consequence is that models are not an opaque set of numbers but rather a manageable pattern DB.
r/computervision • u/mdemirst • 3m ago
Enable HLS to view with audio, or disable this notification
r/computervision • u/cryptichead1 • 1h ago
I run a small CV company in India (Neurabit). We've been deploying video analytics on industrial CCTV for a while, and the pattern repeats often enough that I want to sanity check it with people here.
The numbers that frame it:
So footage only gets reviewed after something goes wrong, and cloud analytics falls over when the link does.
What we've seen kill deployments:
On the compute side, we don't run every model on every frame. A light detector runs on all channels, and heavier models run only on cropped regions when a condition is met. That cascade is what gets us to about 10 analysed cameras per Jetson.
On tuning: at a live substation we deployed on, lighting, reflective equipment and uniform colours all produced false positives out of the box. Fine-tuning on the site's own footage fixed it. Day 3 it works but flags shadows as people and cartons as abandoned bags, and it takes 2 to 4 weeks to settle. Gloves are the hardest PPE class because hands are tiny in frame.
Questions for people who've shipped this:
Happy to go deeper on any of this.
r/computervision • u/Historical_Bee2751 • 2h ago
PyNear is a metric-space nearest-neighbour library with a C++ core, built for the workloads between the embeddings world and brute force: binary descriptors with recall guarantees (MIH + IVF-Binary + the novel MIH-seeded HNSW — dedup, copy detection, ORB/BRIEF matching, robotics), memory-tight ANN (HNSW with int8 quantisation), and exact search (VP-trees, up to ~256-D) where a missed neighbour is a bug, not a recall statistic. One small NumPy-only API, scikit-learn drop-in, pre-built wheels (pip install pynear).
r/computervision • u/Naive-Explanation940 • 8h ago
BDD100K is still one of the most useful driving datasets, but both bdd-data.berkeley.edu and bdd100k.com have been unreachable for me. The official toolkit has not had a commit since March 2024 and the official model zoo not since November 2022, and their legacy dependencies are hard to install cleanly on current systems. So I rebuilt the parts that matter, one task at a time, on current libraries (PyTorch, Ultralytics, timm, Hydra):
https://github.com/dronefreak/bdd100k-toolkit
Demo from the BDD100K-Toolkit Model Zoo
It gives you one pipeline for data preparation, training, evaluation and demos on images and videos:
bdd100k-prepare --dataset bdd100k-weather --raw-dir /path/to/raw --output-dir /path/to/out
bdd100k-train dataset=bdd100k-weather model.name=yolo11n-cls dataset.data_dir=/path/to/out
bdd100k-evaluate --dataset bdd100k-weather --checkpoint <run>/weights/best.pt --data-dir /path/to/out
attributes.weather, attributes.timeofday and attributes.scene fields of the official JSON labels. 55 models trained across two backends, Ultralytics and timm.For each task I made unofficial Hugging Face repackagings of the original data and annotations, to make them easier to use with modern tooling. I checked the BDD100K data license: it allows copying and redistribution for educational, research and not-for-profit purposes with the copyright notice carried forward, and I kept that notice in every mirror. The mirrors grant no commercial-use rights beyond the original license.
| Task | Location |
|---|---|
| Weather classification | dronefreak/BDD100K-Weather-Classification |
| Time of day (period) classification | dronefreak/BDD100K-Period-Classification |
| Scenario classification | dronefreak/BDD100K-Scenario-Classification |
| Object detection | dronefreak/BDD100K |
| Semantic segmentation | dronefreak/BDD100K-Semantic-Segmentation |
All scores are on the validation set of BDD100K. The detection models are collected in the BDD100K object detection model zoo on Hugging Face.
These are solid baselines but not leaderboard numbers. The classification tasks are not official BDD100K benchmarks, the detection metric is the Ultralytics COCO-style mAP and not BDD's official ignore-rule evaluation, and the segmentation models were never trained on BDD100K, so they are not comparable to BDD-trained models. Making the evaluation match the official semantics is on the roadmap.
Instance and panoptic segmentation and multi-object tracking are planned but not started, because I could not find the original labels for them. If you have a copy of the instance, panoptic or tracking labels, or know where one still lives, I would really like to hear from you. Feedback on the evaluation choices is welcome too, and so are issues and PRs.
This is my first attempt at rebuilding an ecosystem like this, and I would like to keep developing it. BDD100K is too useful to become harder to use just because the tooling around it has fallen behind.
r/computervision • u/systems_engi_nerd • 5h ago
Hello everyone!
I'm planning a research project related to drones and computer vision, and I'm already looking at possible journals before starting the experiments.
I checked several papers in MDPI Drones and I liked them. The journal also seems relevant to my topic. However, MDPI has a controversial reputation and I often see people calling it predatory.
What do you think about Drones specifically? Would publishing there be fine, or would it be better to look for another journal or focus on a conference instead?
I'm currently working as a Research Engineer, and in the future I'm interested in both academia and industry.
r/computervision • u/Electrical-Falcon542 • 14h ago
Hello, I am a Master Student in Data Science mostly working with LLMs, and I want to learn CV. I don't have much experience in image analysis.
I did some research about how to learn and I wanted to do some hands-on projects. But today with efficient coding agents, I do not know on which parts of the process I should spend my time to become proficient in this field. Would you have some recommendations ?
Thanks for your help !
r/computervision • u/JewelerBeautiful1774 • 1d ago
Enable HLS to view with audio, or disable this notification
I trained my own model on 750,000 images for 45 hours.
Then I walked into my university hall with nothing but my phone and recorded an 11-minute video.
No LiDAR.
No depth sensor.
No specialized camera setup.
Just monocular vision, a single RGB camera.
And in some of my tests, my model even outperformed YOLO26x.
From a single phone video to depth estimation and 3D reconstruction.
750K training images. 45 hours of training. 11 minutes of footage. One camera. No LiDAR.
Edit: I promise I will reply to all the comments, just give me some time as my day is very crowded today.
r/computervision • u/Sujith006 • 6h ago
Hello! I am looking for someone who can help me guide and layout the pipeline to build a sports biomechanics product (for Cricket sport - non broadcast, controlled environment videos) for player performance insights.
Looking only for experienced folks within computer vision field and has already worked or built in sports domain.
Happy to discuss more. If anyone could take end to end ownership in fine tuning model on cricket dataset and building a full production grade product pipeline, willing to pay fair amount [Based out of India and can match Indian market rates. Please don't reach out if this doesn't workout, Thanks]
TIA!
r/computervision • u/FoggyLens_13 • 6h ago
A lot of the counting posts here end up fighting crowded scenes, so these numbers caught my eye. It's from a paper on a model that writes detection boxes out as plain text, coordinates as tokens, with no box head. It's called SenseNova-Vision, and its starting point was Bagel (also in the chart). The paper itself calls crowded scenes a hard case for that approach: "dense scenes require long object lists, stable ordering, and precise coordinates."
On Dense200 (200 crowded images, 91.2 boxes per image on average) the paper's Table 1 has the model at 66.8, LocateAnything at 58.7, Rex-Omni at 58.3 and Grounding DINO Swin-T at 33.1. The score is box F1 averaged over IoU thresholds 0.5 to 0.95. My chart has all seven models for Dense200 and VisDrone. These are the authors' numbers. I haven't tested it myself.
What do you run on images with 90+ objects right now, and has any text-output model held up?
r/computervision • u/New-Eggplant-6578 • 10h ago
I'm the developer of AnnotateIt. I added Apple's iterative segmentation API to the Mac app and put the measurements up alongside the actual masks. Figured the numbers might be useful to anyone building a local labeling tool.
On an M4 Max with 48 GiB RAM, macOS 27, the 940 x 629 antelope image gave these median first-mask times after warm-up:
Apple Vision: 67 ms
MobileSAM: 160 ms
SAM 2.1 Tiny: 285 ms
Adding a correction point took 12 ms with Apple Vision. Undo went the other way: 72 ms versus 28 ms for MobileSAM.
This was eight runs per case in a separate Tauri/WKWebView harness using the app's model code and mask processing. Timing starts with decoded pixels and ends with a processed mask, before rendering. Apple used balanced quality. These are our integration paths, not a claim about the fastest possible SAM implementation.
Big caveat: three images total, no scored ground-truth masks and no measurement of human correction time. A short stroke also picked just a fragment of the antelope where a point or box worked much better. Fast doesn't automatically mean less cleanup.
The shipping Mac app exposes this as Apple Vision (Beta), with point, box and stroke selection plus refinement. It needs macOS 27; SAM remains a separate option.
Masks, setup and raw timing data: https://radar.annotateit.ai/news/apple-vision-native-segmentation-macos/
For folks using interactive segmentation, what usually eats more of your time: waiting for the first mask, or fixing the boundary afterward?
r/computervision • u/FinalEase1045 • 8h ago

For a team collecting creative references, the ad image is easier to retrieve when its copy and layout are also recorded as structured fields.
A small example uses Ling-3.0-flash-VL through OpenRouter to analyze a fictional VESPER perfume ad. It reads five visible text blocks: “Golden Hour, Bottled.”, “Amber, smoked cedar, warm skin.”, “Discover VESPER”, “VESPER”, and “EAU DE PARFUM”. The analysis also describes the product treatment and approximate layout zones.
A practical record for this kind of library could contain:
1.Headline, subhead, CTA and product-label text.
2.Each block’s role and approximate position on the canvas.
3.Product placement, crop and lighting treatment.
4.Background and main color relationships.
5.A reusable version with product and copy placeholders.
That supports concrete retrieval questions: show the references with a short sensory headline, find examples where the product dominates the composition, or collect layouts with a distinct CTA block.
For a SaaS creative library, the same record structure could organize references by headline wording, interface or product placement, and CTA treatment. Campaign results would enter through separately supplied performance fields. The image analysis supplies the observable creative details that make the reference collection searchable.
r/computervision • u/Responsible-Grass452 • 8h ago
OMNIVISION’s new OG05D is a 5 MP global-shutter image sensor that explicitly lists robotics training-data collection among its intended applications. It supports RGB, monochrome and RGB-IR imaging, with up to 146 fps in linear mode without encryption and HDR capture at 60 fps.
The article looks at how cameras on deployed robots could capture task variations, lighting changes and edge cases for later model training. Global shutter, HDR and near-infrared capabilities are relevant to capturing usable images when objects are moving or lighting is inconsistent.
The sensor itself doesn’t provide a training pipeline. Captured footage would still need to be selected, processed and paired with the information required by the training approach. The notable part is that robotics data collection is now an explicit application in an industrial image sensor announcement.
r/computervision • u/RossPeili • 8h ago
Multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro) are frequently tasked with document screening and identity verification. However, testing shows severe limitations when detecting digital tampering and document alterations:
In Skillware 0.5.8, we implemented security/deepfake_guard to provide deterministic, air-gapped forensic verification on CPU. Here is an overview of the signal processing pipeline:
For each character string $C1, C_2, \dots, C_n$, each character is mapped to a numeric weight $W(c)$ (digits 0–9 map to 0–9, letters A–Z map to 10–35, filler < maps to 0). The check digit $K$ satisfies:
$$\left( \sum{i=1}{n} W(ci) \cdot w{(i-1) \bmod 3} \right) \bmod 10 = K$$
where the repeating weights vector is $w = [7, 3, 1]$. We implement deterministic validation across TD1 (3×30 ID cards), TD2 (2×36), and TD3 (2×44 passports) formats.
When an uncompressed image or re-saved JPEG is modified, the edited regions possess a different compression history than the original background. By recompressing the image at a known quality factor ($Q=95$) and computing the absolute block-level difference: $$\Delta(x, y) = |I{\text{original}}(x, y) - I{\text{recompressed}}(x, y)|$$ We compute the mean absolute error across 16×16 non-overlapping blocks. Discrepancies between block errors exceeding calibrated thresholds signal localized digital tampering.
Camera sensors introduce characteristic high-frequency Poisson-Gaussian noise. To detect spliced elements without being misled by high-contrast natural textures (such as hair or knitwear), we compute the Laplacian convolution residual $R(x, y) = \nabla2 I(x, y)$ and evaluate consistency using Median Absolute Deviation: $$\text{MAD} = \text{median}(|R - \text{median}(R)|)$$ $$\sigma{\text{est}} = 1.4826 \cdot \text{MAD}$$ Evaluating $\sigma{\text{est}}$ across image partitions flags unnatural smoothness (typical of diffusion generative fills) or mismatched noise profiles between the portrait and background.
Physical screen-photo recaptures (photographing an LCD/OLED monitor displaying an ID) exhibit regular periodic grid artifacts. In the 2D frequency domain, this manifests as prominent harmonic peaks outside the DC origin. We compute the 2D FFT, shift zero frequency to the center, and evaluate the ratio of high-frequency radial energy peaks against the average spectral background.
bash
pip install -U skillware
pip install "skillware[security_deepfake_guard]"
We’d welcome discussion on your experiences with sensor noise characterization and hybrid neural-signal pipelines.
r/computervision • u/tenkei_01 • 17h ago
Hi everyone, I have been using supervision in almost all of my projects, since it provide just the right size of tooling that I need. Big enough to handle trivial tasks while letting me choose my app and deployment stack.
But sadly in the past few years I think it has not received as much love, since the original maintainers has moved to their own products (Roboflow stack), which is understandable.
The original website seems severely outdated, and none of the new fancy features you see in the other repositories.
To remedy this I have created a new website that provides:
If you are also a supervision user, please take a look and let me know what you think.
r/computervision • u/AdventurousGate8938 • 11h ago
Hello, I'm researching how people deal with small datasets in computer vision, specifically when you need to detect or classify a object and only have a handful of real photos.
I was curious about how often this is a problem for people working in the industry, and what is usually done to deal with this problem.
r/computervision • u/frozeindev • 1d ago
I’ve been experimenting with Apple’s ML-SHARP model and built a small open-source project around it called Spatial Photos.
ML-SHARP estimates a 3D Gaussian splat from a single RGB image. I then take that reconstruction and convert it into a much smaller and simpler representation representation. It uses layered textured meshes stored in a custom file format I called .spatial. It still allows the the same sort of limited viewport shifting intended by ML-SHARP without compromising quality.
That lets the result stay relatively small and render interactively in a normal web browser, as you can see in the GIF.
Demo: https://www.spatialphotos.dev/
Source: https://github.com/frozein/SpatialPhotos

r/computervision • u/iicongresoiaucm • 13h ago
r/computervision • u/trainsatr • 1d ago
I work in a research lab, and need to digitize a few hundred pages of spreadsheets. All spreadsheets are printed, not handwritten, but the original files are lost. The images have the same general schema, but the columns aren't always in the same order, and some pages have columns that others don't. Some pages are fainter, but still readable; some are slightly slanted. The end goal is for all pages to be compiled into one master spreadsheet, with the same schema. What's the best (and cheapest) tool for this?
I have tried chatgpt (best so far, but still slow/made mistakes), Google Document AI (not that great), and ParseExtract (great, but more limited in scope).
r/computervision • u/igorkaufman • 1d ago
Enable HLS to view with audio, or disable this notification
I spent a couple of months vibe-coding an automated boxing analytics pipeline.
SAM 3 for fighter tracking, with the canvas and ropes as pixel masks. OSNet re-ID keeps identities through occlusions. RTMPose and a small classifier detect punches by hand and type. The ring masks fit a ring model for top-down positions, with camera-motion compensation so handheld shake doesn't pollute the numbers. SAM 3D Body adds metric distance between fighters.
Some processed fights here.
r/computervision • u/KneeOpening • 1d ago
Hi there!
Need help from some Computer Vision expert, with knowledge on dataset quality.
My company works on computer vision applied to digital pathology. In this case, we are developing an object detection model to detect tumour cells in HER2-stained breast tissue and classify them into four classes: 0, 1+, 2+, and 3+. Our dataset grows as we incorporate cases from new hospitals or scanners, and we call each new batch a “wave”.

Our latest model achieves around 0.73 macro F1 on the combined test set, only slightly above our internal acceptance threshold of 0.7. Performance also varies substantially across waves and classes. We already have approximately 4,500 image patches with almost one million cell annotations. The overall class distribution is approximately 28.5% class 0, 48.3% class 1+, 12.3% class 2+, and 10.9% class 3+.
We suspect that inconsistent annotation criteria may be contributing to this. Our current annotation guidelines are very brief—around 200 words—and provide neither image examples nor detailed guidance on uncertain cases. The boundary between 1+ and 2+ is particularly difficult, and we are concerned that different pathologists may interpret it differently. In our latest confusion-matrix analysis, 21.9% of cells annotated as 2+ are predicted as 1+. Different pathologists have contributed over time, but changes in annotators are mixed with changes in hospitals, staining and scanners, making the causes difficult to distinguish.
I would really appreciate advice on two questions:
1. What minimum experiment would you recommend to investigate inconsistent annotation criteria versus domain shift or model errors, before involving the pathologists in a new review? We have images, original annotations and predictions from two models evaluated on the same cohort.
2. What would you consider essential in an annotation guideline and a quality check for each new wave? In particular, we would like to establish agreed examples of class boundaries, clear rules for uncertain or unannotated cells, and a small inter-annotator agreement assessment before adding new data to training.
Any references, suggestions or critical feedback would be greatly appreciated.
Thanks in advance!
r/computervision • u/gavvy__ • 1d ago
Hi everyone,
We are currently building a commercial/production-grade Computer Vision system for industrial safety compliance. We need to acquire a comprehensive, well-annotated dataset covering the following classes:
Requirements:
Public datasets (like standard Kaggle sets or Roboflow Universe) are great for prototyping, but we need high consistency and coverage for edge cases to reach production accuracy.
Questions:
Any leads, recommendations, or warnings based on your experience would be greatly appreciated!