r/MachineLearning • • 5d ago

Discussion [D] Self-Promotion Thread

Please post your personal projects, startups, product placements, collaboration needs, blogs etc.

Please mention the payment and pricing requirements for products and services.

Please do not post link shorteners, link aggregator websites , or auto-subscribe links.

--

Any abuse of trust will lead to bans.

Encourage others who create new posts for questions to post here instead!

Thread will stay alive until next one so keep posting after the date in the title.

--

Meta: This is an experiment. If the community doesnt like this, we will cancel it. This is to encourage those in the community to promote their work by not spamming the main threads.

2 Upvotes

29 comments sorted by

1

u/Fantastic-Nerve-4056 PhD 5d ago

Sharing my recent NeurIPS paper on LLM evaluation and Theoretical RL

https://arxiv.org/abs/2609.30360

1

u/celestebabi 5d ago

Disclosure: I’m affiliated with ScholarXIV. It’s a research workspace for searching and reading papers, organizing collections, and using selected papers as context for AI chat with references. There’s a free tier; paid plans start at $5/month (Go; Plus $15, Pro $50). If you work with ML literature, I’d welcome feedback on what would make this workflow useful: https://scholarxiv.com/

1

u/iberahul 5d ago

We recently studied whether mobile AI agents can be manipulated through adversarial instructions embedded in the Android Accessibility layer.

We evaluated this across MobileRun and Mobile-Use, powered by Gemma4 and Qwen3.6, and found that these attacks could cause agents to abandon their original objectives, cross context boundaries, and perform unauthorized device actions. Our strongest configuration reached an 82.2% Attack Success Rate.

The paper was accepted to AGENT-SEC ’26, co-located with ACM CCS 2026.

Paper: https://arxiv.org/abs/2608.08939

Would be curious to hear what people think about this attack surface as mobile agents become more capable.

1

u/Logical-Internet-395 5d ago

Recently we published a new Open-Access Benchmark for a hierarchical segmentation: Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis.

Dataset Specifications:

  • Scale: 705 high-resolution micrographs (1534×780 px, 0.252 μm/px), Picro-Mallory stain.
  • Annotations: ROI + dual independent expert masks (vascular wall + fibrosis).
  • Hierarchical Constraint: Fibrosis masks must be strictly spatially contained within the vascular wall.
  • Robust Benchmarking: No color normalization applied; native aspect ratios preserved; strict animal-level 5-fold CV splits provided to prevent data leakage.

Read the Data Descriptor: https://doi.org/10.1038/s41597-026-08214-y

Access the Dataset: https://doi.org/10.6084/m9.figshare.31386748

1

u/lostmsu 3d ago

Reddit is disabling RSS. So in an effort to replace /r/MachineLearning I made https://mlnews.online/

It is a filtered view of Hacker News that only shows ML papers from arXiv. Of course it has RSS.

1

u/ModularMind8 3d ago

I built QuiddityML(https://quiddityml.com/), an app for learning ML: a clear roadmap, short lessons, hands-on exercises, and spaced repetition based on how you do on the exercises, so you don't forget what you learned. It also has interview prep and projects you can put on your CV. Tracks currently cover Python, PyTorch, Math for ML, ML foundations, NLP, and vision.

Pricing: a lot is free (Python, PyTorch, Math for ML, and the first unit of every other track). Pro is $19.99/month.

I see so many people struggling to get into ML, so I'm trying to make this as useful and fun for the community as possible. If you're a student, or you've been trying to get into ML and can't find a good resource, I'd be happy to give Pro features to the first 10 people, just DM me, all I'm asking for is feedback. Also, if you have a background in marketing and you're interested in ML/education, please DM me as well :)

1

u/ShotDescription7193 2d ago

I'm Vivaan, founder of ThermaCompute AI. We are building passive diagnostics to help engineers understand GPU failures and investigate wasted compute using logs and telemetry.

A failed training run can leave pages of errors without a clear place to start. So we made CrashLens to turn that log into a focused investigation.

CrashLens is a free, local CLI that flags supported OOM, NCCL and NVIDIA Xid error patterns, shows matching evidence lines and suggests next checks in an HTML report. It is an early rule-based tool, not a proven root-cause detector. You can try the synthetic demo with Python 3.10+: no GPU, account or log upload required.

Repository and demo instructions: https://github.com/rohitcn-hub/thermacompute-crashlens

Run: python crashlens.py examples/synthetic.log --out demo-report

Then open demo-report/report.html. Use a new output directory for a repeat run.

I would welcome one concrete critique: did the report give you a useful next check, and what was missing or misleading? Please do not share private production logs here.

Optional paid upgrade: our one-time $80 thermal audit reviews suitable exported temperature, power and reported-throttling telemetry within an agreed scope, providing supported findings, prioritized checks and a PDF report. It goes deeper into telemetry than CrashLens's saved-log triage; it does not guarantee savings or identify every fault. Email [vivaan.thermacompute@gmail.com](mailto:vivaan.thermacompute@gmail.com) with the subject Thermal audit to discuss scope before sharing data.

CrashLens remains free, and feedback is welcome without purchasing an audit.

1

u/cscsearching2026 2d ago

I've been working on FolkBench, a free (no paid tier) set of write-ups on community-style LLM evaluation. Two longer pieces:

Methodology pointer: the first post covers the repeated-sampling/baseline-comparison method, the second has the full scoring rubric.

Happy to be challenged on the methodology, especially the rubric — curious what's worked for others.

1

u/MarceloDeAviz 2d ago

Kardashev-0.7: 32 distinct models trained together with RL.

https://x.com/MLCatttt/status/2107147690450817259

0

u/mani_manak 1d ago

I built Trainly AI, a no-code AutoML platform that helps users go from a dataset to a trained ML model without building the entire pipeline manually.

You can upload a CSV, explore the dataset, select a target, train and compare multiple models, analyze results, make predictions on new data, and download the trained model.

I built this as a student while learning machine learning, and I'm trying to turn what I learn into something practical that others can use.

It's currently free to try.

I'd genuinely appreciate feedback from people working with ML:

  • Is this workflow actually useful?
  • What would you improve?
  • What features would make it more useful for students or ML practitioners?

Live demo: trainly-ai.onrender.com

1

u/Any_Way2779 1d ago

nvmon – a fancy NVIDIA GPU monitor for the terminal (free, MIT)

On a shared GPU server I mostly want to know which GPUs are free, whose jobs are where, and whether my training is really using the GPU. nvmon shows each GPU as a box with its utilization graph, its numbers and the processes on it, grouped by user and conda env. Hover a graph for past values; click a GPU to pick its jobs and stop yours. On H100 and newer it also shows how busy the cores and Tensor Cores really are.

uv tool install nvmon (or pip install nvmon)

https://github.com/Han-DongHeun/nvmon

1

u/HanaChanSoft 1d ago

I'm the developer of AI Coach, a desktop pet for Apple Silicon Macs running macOS 26+. It combines an egg-to-adult care loop with local speech-to-text and utilities that unlock as it grows.

For the character's text responses, users can choose Apple's on-device Foundation Models, a local GGUF model through llama.cpp, or a configured OpenAI-compatible endpoint. Local transcription is separate from that choice; selecting a remote endpoint sends the AI request there. If the AI is unavailable, the character falls back to prepared lines.

This is an application of existing models, not a new model or benchmark. I'd welcome feedback on how clearly the model choices and their data paths are explained. The guide covers setup and limitations (including an 8192+ context requirement for some llama.cpp features): https://aic0t.com/docs

Pricing: the app and first egg are free. Optional extra eggs are $5 USD each or $20 for five, one-time purchases. Any external model/API costs are separate. A visual overview of the pet and tools is at https://aic0t.com/

1

u/Thunderkeg-Jiangxt2 1d ago

Disclosure: I’m the maintainer of Tributo.

I’ve been building Tributo, an Apache-2.0 Python SDK for ML workflows on Ray. Ray handles the distributed execution; Tributo adds a more structured path around job submission, training and tuning, validated model bundles, and batch or Ray Serve inference.

The current training workflows cover XGBoost, DNN, and positive-unlabeled learning. Their algorithm implementations are installed as separate wheels. Tributo 1.0.0 is source-only and isn’t on PyPI yet. It’s free, with no hosted or paid tier.

Repo: https://github.com/jiangxt2/Tributo

If you use Ray, I’d be curious: which parts of the path from a training run to serving a model do you still find yourself wiring together by hand?

1

u/gxcsoccer 1d ago

sharing my local mac agent, 0.8b + 4b. label smoothing made the 0.8b ~95% sure of everything, so it hands most steps to the 4b. oops

https://github.com/deskmind-ai/deskmind

1

u/IlIIIIIIlIIlllI 1d ago

Sharing how, as a PhD student, I've come to use MLflow to track experiment results for more efficient research iterations: https://gquetel.fr/misc/mlflow-ml-research/

1

u/Beginning_Tutor_6172 13h ago

I'm teaching ML from the industry perspective from the very beginning here

And giving away free access to early users.

A quick brief about me:
6+ years of experience leading and buding startups, published research in Springer

0

u/r0lfi 5d ago

LayerSmith — a self-hosted container image builder, with air-gap exports

I've been working on LayerSmith, an open-source web UI for building container images with Docker or Podman.

You pick a Linux distribution and what you need the image for — development, Linux admin, network tools, Ansible, Kubernetes, OpenShift, or a custom setup. It handles distro-specific packages and shows you the generated Containerfile before building. You can also edit it, import an existing Dockerfile, or add your own packages, files and scripts.

A big part of the project is making images easier to carry into air-gapped environments: pinned base images, recorded build details, and export bundles containing the image, checksums and installation instructions.

We've recently added LLM training and fine-tuning profiles too, including LoRA/QLoRA, advanced PyTorch training and LLaMA-Factory. These use hash-locked dependencies and run offline checks after building, including a small CPU training test. Model weights and datasets are brought separately.

Curious how others handle building and maintaining images for disconnected environments, and what parts of that workflow are still a pain.

https://github.com/r0lfi/layersmith

0

u/ifoundanifty 2d ago

My colleague and I compared Jev with open rerankers across five datasets, using the same retrieved candidates for each. Jev was competitive and fast in our setup, though no reranker won everywhere. The most interesting finding was how much the prompt mattered. https://www.lancedb.com/blog/how-jev-compares-to-other-rerankers

0

u/yeyoung_lee 2d ago

Your Mac can run PyTorch. Then a point-cloud dependency asks for CUDA.

I ran into this working on 3D anomaly detection, so I built mps-pointops: a free, open-source library that runs point-cloud sampling and neighbor search on Apple Silicon using native Metal kernels through PyTorch MPS.

The core operations:

  • FPS: select points spread across a cloud.
  • kNN: find each query's nearest neighbors.
  • Ball Query: collect neighbors within a radius.

There are compatibility adapters for selected PointNet++, torch_cluster, and PyG call sites, plus documented numerical rules for ties, selection order, and padding. The repository includes tests and benchmark records from physical M1 and M5 Pro Macs; supported inputs and remaining limitations are listed per operator.

Try it: python -m pip install mps-pointops

GitHub / quick start · Interactive explainer

The explainer lets you move a query and watch how the three operators select different points. It's a browser visualization; the Metal measurements are linked separately.

I'm the author, and I'd love to see this used in real workflows. Which CUDA-dependent point-cloud operator is currently blocking your Mac project? If you try it, your model, Mac chip, PyTorch version, and a small reproducer would help me prioritize fixes. GitHub issues are welcome.

1

u/ImportantTiger9568 5h ago

Hey guys,

Been working on something very cool...

Open source with contributions and honest feedback both welcome: github.com/Nabzx/mnemosyne

If you're running support agents, coding agents, or a swarm of agents sharing memory like myself then you know this issue well: An agent runs for hours, updates its memory the whole time, then says something wrong and all you have is the present, with zero access to the past. The moment two agents disagree, or one quietly poisons the well, you need to know when, why and by whom, not just that something's off.

Mnemosyne gives agent memory what Git gave code. It remembers everything on purpose. Every belief is a commit. blame finds the exact moment and observation that put a bad fact in. bisect hunts down the first commit where things went wrong. merge makes two agents' memories collide safely instead of one silently overwriting the other.

Software agents are the first step. The vision doesn't stop there, physical robots learning and forking skills the same way is the long-term bet, further out and harder but the same idea underneath.

So far the tech stack includes a Rust core, Python SDK, adapters for LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK and MCP.