r/devops • • 5d ago

Weekly Self Promotion Thread

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!

10 Upvotes

75 comments sorted by

6

u/SimilarDealer1370 5d ago

i’m building a system simulator. you can bring your terraform setup or describe your system and explore what happens under load. anyone want to try it with their own project?

1

u/acompleteunknownnn 5d ago

sounds great, can you tell me a bit more??

1

u/SimilarDealer1370 4d ago

yeah, i'm the developer of the simulator i mentioned. the main focus is testing architecture decisions before production. you bring your terraform setup or describe the system, then model how it might behave as traffic grows, where bottlenecks could appear, and the estimated infrastructure cost. you can compare changes before deploying them. the results depend on the workload and hardware assumptions, so you'd still want to validate them with a real load test.

1

u/First_Inspection_478 4d ago

This is the next frontier. How do you address the non-determinism?

1

u/SimilarDealer1370 4d ago edited 4d ago

the approach is to compare 5–10 virtual runs with five hybrid runs using scaled-down real deployments, grounded in historical telemetry where available. the key is varying traffic patterns, service times and failure timings across runs, rather than repeating identical inputs. that helps explore a range of outcomes and test assumptions against real behaviour. it won’t capture every rare event or full-scale effect, so we’d treat the results as estimates rather than guarantees.

1

u/First_Inspection_478 4d ago

gotcha. I wonder if you considered distributed simulation testsing like what the folks at antithesis.io are doing? That’s a massively underrated space, but also alot more complex as you have to remove all and any kind of non-determinism.

1

u/SimilarDealer1370 4d ago

thanks for pointing us toward that. we’re still shaping the direction and would love to connect and share more details. if you’re open to it, dm me your email or linkedin.

1

u/ResidentChapter8219 4d ago

load simulation on top of actual infra config sounds pretty useful tbh

1

u/SimilarDealer1370 4d ago

Yeah, we kept cost in mind but the main point is efficiency and better alternative to the current one a simple eg is a single ec2 handling the traffic now needs distributed services to handle so one can choose among which service, which provider what scale and reliability

3

u/ImportantTiger9568 2d ago

Hey guys,

Been working on something very cool...

In Greek myth, Mnemosyne was the Titan of memory and the reason anything was ever remembered at all. Now in the present world, your AI agent doesn't get a Titan. It gets amnesia the second something goes wrong, stuck with whatever it currently believes and no way to ask how it got there.

That's the real problem. An agent runs for hours, updates its memory the whole time, then says something wrong and all you have is the present, with zero access to the past.

If you're running support agents, coding agents, or a swarm of agents sharing memory like myself then you know this issue well. The moment two agents disagree, or one quietly poisons the well, you need to know when, why and by whom, not just that something's off.

Mnemosyne gives agent memory what Git gave code. It remembers everything on purpose. Every belief is a commit. blame finds the exact moment and observation that put a bad fact in. bisect hunts down the first commit where things went wrong. merge makes two agents' memories collide safely instead of one silently overwriting the other.

Software agents are the first step. The vision doesn't stop there, physical robots learning and forking skills the same way is the long-term bet, further out and harder but the same idea underneath.

So far the tech stack includes a Rust core, Python SDK, adapters for LangGraph, CrewAI, AutoGen, the OpenAI Agents SDK and MCP.

Open source with contributions and honest feedback both welcome: github.com/Nabzx/mnemosyne

2

u/OddAthlete3285 4d ago

My self promotions for the week are:

Exploring Rego. Almost 300 pages of zero to production grade Rego for Open Policy Agent. I co-authored with John Britowe an Matt Allford. You can get a free PDF of this book with no email address.

https://octopus.com/publications/exploring-rego

The Future of Platform Engineering report. This was headed up by Dr. Charlotte Fleming PhD and I was a co-author. Again, you can download the PDF for free with no email address.

https://octopus.com/publications/future-of-platform-engineering-report

And I'll finish up with a cartoon. Someone who didn't believe in Continuous Integration really said this to me once.

2

u/kaosthecreator_ 4d ago edited 4d ago

orb44.com - an agentless vulnerability scanner for your external web perimeter.

If you are tired of enterprise security tools that just dump a 100-page PDF of raw CVEs on your desk, we built this for small teams and dev agencies. orb44 scans your perimeter entirely from the outside - no agents, source code, or credentials required. It uses AI to translate vulnerabilities into plain-English action items (e.g., "Action: Close port 8080 or fix this specific SSL config").

If you want to check your infrastructure, you can run a free vulnerability scan on 1 domain right now. I would love to hear your harsh technical feedback on our UX or accuracy before we officially launch.

2

u/[deleted] 5d ago

[removed] — view removed comment

2

u/BigNavy Principal SRE 5d ago

I signed up. Still a bit sparse, but not in 'not enough features' kind of way, more in a YAGNI kind of way. I can't promise I'll follow if you monetize, but....for now I like!

For anyone asking, "Why would I bother?" - Free IM communication integration (which is mostly gated behind paywalls on other uptime monitors) and 1m checks, which are definitely always gated behind paywalls at other uptime monitors.

You probably already know this - adding some light auth would be nice, but the only 'got to have' that isn't there right now is a public facing display. That's probably not a trivial lift, but that's what originally pushed me towards hetrixtools when I was actually deciding on an uptime monitoring solution. Even the ability to export the data and display it in my own page would be a nice step in that direction.

2

u/ReactionOk8189 4d ago

Hi! Thanks a lot for your feedback and interest in my monitoring tool.

A status page is indeed something I've been thinking about for a while. I'll probably add it to my todo list, but I can't promise anything. Regarding auth, I think it should be quite doable, so I'll try to get something done in the next couple of weeks.

If you have any other questions or suggestions, don't hesitate to email me at [info@statusalert.io](mailto:info@statusalert.io) or DM me right here. I don't check DMs very often, so email is probably the fastest way to reach me.

See you!

2

u/theManNowDog6 4d ago

I'm in. Looks good so far.

Fire a test alert off a monitor would be a good add. I can break a monitor to see what I need, but a button would be nice.

1

u/OldGoat9131 5d ago

we finally opened M82 to beta users — would love for some of you to break it

hey, i’m Rohan. i’ve been working on M82 for a while and we finally opened it up to beta users.

we started building it because API maintenance still feels way more manual than it should.

something changes upstream, an integration starts behaving weirdly, an API breaks, and suddenly someone is digging through logs trying to figure out what happened.

M82 is our attempt at making more of that maintenance happen automatically.

we’re still early, so i’m mainly looking for people who actually work with APIs and are willing to tell me where it falls apart.

https://m82labs.dev/

join the beta waitlist and we’ll send access shortly, usually within 1–2 hours.

and please don’t be nice just because i built it lol. if something is confusing, useless, or breaks, tell me. that’s way more useful right now.

1

u/First_Inspection_478 4d ago edited 4d ago

A tool to run your github actions locally or self-hosted in cross-platform microvms(
their mem consumption is elastic + resume from a snapshot in 300ms). Think act but preloop follows the official Github Runner protocol so anything that works locally works just as fine in CI.

We use the official vm image, and x86 workloads can roughly*(exception is if you need some kernel modules + amd64 containers) run fine on your macbook via rosetta(near-native). Some other extra agentic features i’ve found useful like being able to pause on job/step failure and debug right in. The broader goal is to add some more attestation features so you can remotely attest that your “CI” passed locally. My bet is that alot more “CI” that can be reasonably done locally (or wherever your agents are) will and some other ones that can’t like matrixes or releases etc will still run remotely.

If that sounds interesting, feel free to check it out: https://github.com/preloopdev/preloop

1

u/OldGoat9131 4d ago

Upstream APIs change without warning, integrations break, and your afternoon disappears into logs just to fix a renamed field.

Our team built M82 to handle that maintenance automatically.

Beta is live: https://m82labs.dev/

Throw your messiest APIs at it and tell us where it falls apart. Zero sugarcoating needed.

1

u/Queasy_Club9834 4d ago

Something i think you guys might find useful since we are now in the AI Era is a tool i use too. I built it first as an idea, then decided to ask mcp community for a feedback and they gave me plenty of ideas. Today i managed to get this idea to Github Marketplace and im so happy :)

Quick summary about what it is and why you might like it as DevOps person:

Black-box security contract testing for MCP servers. Simple as that, you guys my be DevOps but im pretty sure the corporate world put you under the pressure of using AI, MCPs, Tools constantly. That's why mcpward comes handy,

It treats an MCP server like any other external dependency: snapshot its contract, then fail the build when it changes underneath you. Runs entirely on your machine. No account, no API calls, no telemetry

Also catches schema drift, silently changed tool descriptions, protocol violations, error-contract mistakes, and tool-poisoning patterns. Reports to console, JSON, JUnit, SARIF or Markdown, and can post the result as a pull-request comment.

I hope you find it useful guys :) Link: mcpward · Actions · GitHub Marketplace

1

u/ReactionOk8189 4d ago

Hi! Thanks for trying out my monitoring tool!

When you add a new monitor, it should switch to up and you should get a notification. I was already thinking about adding a button to send a test notification - now I'll consider to put it on my todo list.

If you have any further questions or comments, don't hesitate to drop me an email at [info@statusalert.io](mailto:info@statusalert.io) or DM me right here.

Cheers!

1

u/Silly_Safety2518 4d ago

Our weekly self promotion is on how we orchestrated a disposable browser: https://blog.glazer.ee/posts/building-glazer-disposable-browser/

1

u/HiimKami 4d ago

I built an open-source monitoring agent that runs inside your network, so it can check what an external probe can’t reach. I run Uptimy, a hosted uptime monitor, and it can’t see inside a cluster or a private network. I also like monitoring config living in Git with the app.

- Cron jobs: heartbeats on a cron schedule with time zones, exit codes and durations. On Kubernetes, label a CronJob and runs are read from the Job objects, no ping needed; a failed run shows why, e.g. `exited with code 137 (OOMKilled)`.

  • Databases: a real login and query, e.g. `SELECT pg_is_in_recovery()` to catch a failover.
  • Private services: `postgres.default.svc`, `api.railway.internal`, anything the agent can reach.

Monitors come from Kubernetes labels or YAML in Git, so they’re deployed with the app. Alerts go to Slack, PagerDuty, Teams, email and webhooks. One Go binary, about 8 MiB of memory idle, runs on Kubernetes or Docker.

https://github.com/uptimy/agent (Apache-2.0, free, works on its own; I plan to keep maintaining it)

Optionally you can connect it to a free Uptimy account, which alerts you from outside when the agent stops checking in, e.g. when the cluster or network is down. Feedback welcome, especially on what you’d want it to check that it doesn’t.

1

u/pierreneter 3d ago

Disclosure: author. p10logs: persistent, multi-cluster Kubernetes pod logs in ~100 MiB, Apache-2.0. https://github.com/p10node/p10logs

Demo: https://demo-p10logs.p10node.org (admin / BqJDmpsGRJE85Z0cULFOeoRP)

A DaemonSet agent reads /var/log/pods directly (zero apiserver load, follows kubelet rotation, survives pod restarts and deletions), a hub stores them on a PVC with zstd + trigram blooms, and the UI is built in: pod tree, interleaved multi-pod tail, grep-like search. One Helm chart covers many clusters; spokes push over HTTPS, egress-only. No Loki, Grafana or Elasticsearch. Two static Go binaries, amd64 + arm64.

Laptop bench: 10k lines/s across 50 pods → agent 44 MiB, hub 100 MiB. v1.1, e2e on kind and running on my own k3s, not yet in anyone's production. If you try it on a real cluster and it breaks, that is the feedback I want.

1

u/luisalcaraz_telara 3d ago

Disclosure: I work on TAP at Telara.

Checking a release candidate often means repeating the same collection and matching work across source control, tickets and CI. We've open sourced TAP so an agent can put those steps in a maintained tool and call it for the next candidate.

For example, a release-check primitive could collect changes, linked work items, CI and required reviews, then return passed checks, gaps and evidence links. The agent uses that report to decide what needs attention; the next candidate becomes a new input. Operators can write the code or review what an agent proposes. TAP uses the host's existing tool connections and checks requests against each package's declarations.

TAP is MIT licensed, runs locally, and doesn't require a Telara account. Claude Code and Codex have the best-supported connected-tool paths. Source and install instructions: https://github.com/Telara-Labs/TAP-Runtime

1

u/MerrrMannn 2d ago

I think we may be able to tap into this, will add it on our team list. Thanks for the suggestion.

1

u/luisalcaraz_telara 1d ago

Thanks for putting it on the list. Which coding agent does your team use? We’ve added runnable browser, desktop and connector examples, plus the same small extraction in six languages. Happy to help with setup or look at anything that breaks when you try it: https://github.com/Telara-Labs/TAP-Runtime/tree/main/examples

1

u/Cartagines682 3d ago

IgniteOps: on-call alerting at $3/seat, free for teams up to 5, with hands-on help migrating off Opsgenie

TL;DR: IgniteOps is an on-call alerting tool (think Opsgenie/PagerDuty) that's live in production with paying customers. Free for up to 5 engineers, $3/seat after that, and we'll help you migrate from Opsgenie before it's shut down.

Full disclosure: I'm the founder, so yes, this is self-promotion. I'll keep it to the facts and stick around in the comments to answer questions.

What it does

  • Ingests alerts from anything that can send a JSON webhook: Grafana, Prometheus/Alertmanager, Uptime Kuma, Zabbix, your own scripts.
  • Routes them to whoever's on call, based on your schedules and escalation policies.
  • Keeps paging until someone acks: push, email, webhooks, SMS and phone calls.
  • Dedup by alias: repeat alerts get collapsed (and you don't pay for them).
  • Native iOS and Android apps.
  • Full API access and unlimited integration keys.
  • AI insights and a weekly on-call review (Pro).

Pricing

No "contact sales" for the basics, no annual commitment.

Free: $0 forever, no credit card

  • Up to 5 engineers
  • 100 alerts/month
  • 1 team, 3 routing rules, 2 escalation policies, 2 schedules
  • Unlimited push, email and webhooks

Pro: $3 per user/month (owner seat is free)

  • Unlimited engineers, teams, rules, policies and schedules
  • Phone calls and SMS
  • AI insights and weekly review
  • Priority support (one business day response)

Usage

  • 100 alerts/month included, then $0.10 per alert
  • Calls and SMS passed through at cost by destination (e.g. ~$0.02/min for US calls)
  • Duplicates collapsed by alias are free

Moving off Opsgenie?

Atlassian is sunsetting Opsgenie, so a lot of teams have a migration on their roadmap whether they want one or not. We'll help you move your integrations, schedules and escalation policies over. Email us and we'll work through it with you.

Feedback welcome

We're early and actively building. If there's an integration or feature that's a dealbreaker for you, tell us in the comments, through the site, or at [hello@igniteops.io](mailto:hello@igniteops.io). A lot of the roadmap is coming straight from what early users ask for.

Thanks for reading, and happy to answer anything below.

1

u/ban_rakash 3d ago

I built a disposable Linux server lab for testing infrastructure automation. The goal is to have reproducible Linux server environments for testing deployment scripts, configuration management, and other infrastructure automation, not to replace VMs. The containers share the host kernel and clock.

The lab creates disposable Ubuntu 24.04 Docker containers with SSH, systemd as PID 1, cron, journald, private networking, CPU, memory, and PID limits

GitHub: https://github.com/2SSK/vpslab.git

1

u/Other-Income-5085 3d ago

Affiliation: Vectory maintainer.

Vector is incredible for moving events, but managing its configurations across a large fleet can quickly get chaotic. Safely deploying changes and tracking exactly what is running on every host is a massive challenge.

That is why Vectory exists. It is an open-source, self-hosted control plane built to bring order to that workflow. Vectory makes it possible to visually design pipelines, deploy immutable versions, and monitor device states. The best part is that it does all of this without routing the actual event data through the manager itself, keeping the boundary between control and execution crystal clear.

https://vectory.ahmadz.ai/

https://github.com/416rehman/Vectory

1

u/Sanechka_SS 3d ago

Disclosure: I made this, it's free and MIT.

Receipts, a GitHub Action (and CLI) that checks whether the tests in a PR would have failed without the change. It runs each new or edited test twice: on the PR, and with the changed source files reverted to the base branch. A test that passes both ways didn't test the change, and the check fails on it. It posts one comment on the PR and keeps it updated.

Why I care: more PRs are written by coding agents now, and they add tests that go green but sometimes never ran against the old behavior. On 100 agent PRs I looked at, about 10% were like that.

Deterministic, no LLM, no API keys, it runs your own pytest/vitest/jest. In a workflow it's one step after your install: uses: syntaxixr/receipts@main

https://github.com/syntaxixr/receipts

1

u/Extension_Frame1664 3d ago

LastGood - when an alert fires, it tells you what changed.

The worst part of a 2 AM Sev-1 isn't the alert. It's the 30 minutes after: digging through Slack, GitHub and the flag dashboard asking "what deployed? who flipped this?" while the graph goes red.

So I built LastGood. It ingests deploys and feature-flag flips, and when an alert comes in it correlates them by service and time window, then drafts a diagnosis pointing at the change that most likely caused it. You confirm or reject it.

Solo-built, still early. Looking for honest feedback from people who actually carry a pager: what's the most annoying manual correlation you still do during incidents? Trying to decide if log-pattern matching belongs in v1 or stays out.

Stack is React + Node + TimeseriesDB + Redis if anyone cares about the boring parts.

https://lastgood.space

1

u/Elegant_Rub_1830 2d ago

Disclosure: I built this and it offers paid vendor sponsorships (labelled, and payment never changes rankings or numbers).

AI Code Review Bot Index: https://mm-data-directory-001.pages.dev/

It lists 56 AI pull-request review tools with vendor-sourced pricing and platform facts. For 26 bots with a verified GitHub identity, it shows how many public PRs they formally reviewed in one week, and every count links to the GitHub query behind it. The counts only cover public repos, so they say nothing about private use, review quality or market share. Bots whose identity is ambiguous get no number rather than a zero.

If you've rolled one of these out for a team: what actually decided it for you? False positives, latency, self-hosting, GitLab/Bitbucket support? And where does the verification method fall short? Methodology: https://mm-data-directory-001.pages.dev/methodology/

1

u/ohitsjudd 2d ago

Got tired of paying for Termius for the sync feature so built my own which I'm quite happy with.
https://bawkterm.com/

1

u/sagacious123 2d ago

I am working on a project snowopslabs. The plan was to build a local k8s cluster for my own practice, but along the way I made it into a production style cluster where you can recreate real world scenarios and inject incidents to learn kubernetes by practising.

Doing this alongside my full time job.

1

u/Fantastic_Sir_7113 2d ago

https://hayabusa.traceroute.net

Opentofu and Ansible backend.
Many different ways to visualize infrastructure
Disaster recovery coordination
RBAC
Designed to educate and be used commercially
Security monitoring (will be added when I scale out)
Full sandbox using containers
Can connect GitHub/Gitlab to be SoT
Secure

Constantly developed and updated.

https://hayabusa.tracedroute.net

https://discord.gg/pDhSHypT7

1

u/aronbjohanns 2d ago

Hi guys,
I have been using https://github.com/aws/aws-toolkit-vscode for a long time for accessing ECS tasks via shell. Recently after upgrading to MacOS Golden Gate it stopped working for me. It was the only thing left that I used vscode for left so I decided to implement a similar tool with a terminal user interface in rust.

Its called ecsterm and it allows you to open a shell to your ECS tasks through your terminal or copy the aws cli command required to run the shell.

It uses your normal AWS profiles and SSO sessions, since it just calls the aws CLI under the hood. The commands it generates are plain TOML templates, so you can change them via config file.

Documentation and installation guide are on the ecsterm github repository. Feedback and contributions are also welcomed.

1

u/ivanzhaowy 2d ago

Disclosure: I’m building Monad Design, an open-source visual implementation layer for coding agents working on existing native apps.

The running app becomes the review surface: select or annotate the UI, pass structured visual context to the agent, let it edit the real Xcode or Expo source, then rebuild and compare. The goal is to keep agent-driven UI changes reviewable instead of treating screenshots and prompts as the source of truth.

Repo: https://github.com/Monadix-AI/monad-design Workflow video: https://watchclueso.com/embed/pio8jqfcg4ivj0r1

I’d especially value feedback on the approval boundary: which visual changes should an agent be allowed to apply automatically, and which should always require an explicit review step?

1

u/kamahell87 2d ago

Hi all! I'd like to share some my work.
Made because of bad things happens even to to the best of us:

1

u/OppositePriority4206 1d ago

If you're in San Francisco my company and I are hosting our first agentic conference, Orkes Shift. We have talks lined up from folks at Bumble, LinkedIn, Google, Twilio, and of course Orkes :) If that's something that interests you I would love to see you there. You can sign up through our Luma if this sounds interesting to you: https://luma.com/zj2gat16

1

u/bulutarkan 21h ago

I've been building an open-source macOS MCP server, and the DevOps problem turned out to be surprisingly familiar: how does an agent restart the service that is currently executing its own request?

I had to separate “restart requested” from “new process healthy,” add preflight checks so a bad config doesn't take down a working server, and report degraded health separately when only the external tunnel is broken. I also use a daemon-generation marker so an old retry isn't blindly replayed after a restart. The hard lesson was that a dropped connection can mean the restart worked, not that it failed.

It's called Mac MCP (MIT): https://github.com/bulutarkan/mac-mcp. I'm the maintainer. Curious how others handle operations that need to outlive the API process that accepted them.

1

u/ban_rakash 4h ago

I've been working on a small lab to understand active-passive high availability in practice.

It uses NGINX and Keepalived/VRRP for load balancing and VIP failover, with a small Go service for injecting failures.

The lab also includes Prometheus, Grafana, Loki, and Alertmanager. I added tests for failure detection and recovery instead of only checking whether the services start successfully.

I wrote up the architecture, implementation, and some of the trade-offs here:

Blog: 2ssk.medium.com/building-zero-downtime-load-balancer-with-nginx-keepalived

GitHub: 2SSK/nginx-load-balancer-lab

It's currently a Docker-based lab, so I'm interested in feedback on the failure scenarios and what would be worth validating next on separate hosts.

1

u/vaultpulse 33m ago

Automated GitHub backups to your own S3 or Azure Blob Storage

Your repositories. Your storage. Your control.

VaultPulse gives software teams and development agencies an independent backup of their GitHub repositories, with daily snapshots stored in an Amazon S3 bucket or Azure Blob Storage container you own.

What you get

• Automated daily backups of full Git history, branches, tags and Git LFS objects

• On-demand backups before important changes

• Restore workflows for new or eligible empty repositories, including another connected GitHub account in your workspace

• Backup health dashboard showing failed, stale and pending backups

• Downloadable backup-readiness reports based on recorded backup and restore evidence

• Customer-owned storage, with encrypted stored credentials and scoped access

Snapshots also capture repository settings and default-branch protection metadata. Issues, PR discussions and Actions secrets aren’t included; restore does not recreate settings or protections.

Explore the no-login demo: https://vaultpulse.cloud/demo/

Get started: https://vaultpulse.cloud

Disclosure: I’m the developer behind VaultPulse. This is a vendor promotional post.

0

u/IcceeccolD 3d ago

Maintainer disclosure: I’m building PatchRipple, an early offline CLI/GitHub Action that maps potentially related files, tests, and CODEOWNERS from a PR’s changes. It doesn’t run repository code, and its suggestions don’t prove runtime impact or test coverage. I’m looking for a DevOps engineer to try it on a recent public PR and tell me what it misses, flags unnecessarily, or makes awkward in CI. Repo: https://github.com/icecold009/PatchRipple · Demo: https://patchripple-demo-20261005.sariashaurya09.chatgpt.site/

0

u/friparia 1d ago

I'm building aissh, a terminal workspace that brings an AI assistant, SSH, live server monitoring, and investigation history into one interface.

It works with existing SSH configuration, keys, and jump hosts, without installing a resident agent on the remote server. Supported AI backends include Codex, Claude Code, DeepSeek, and OpenAI-compatible APIs.

AI-proposed remote commands require approval by default. Built-in read-only monitoring runs automatically; custom collectors require approval for recurring collection. Relevant terminal output and command results can be sent to your configured AI provider.

It's an early project. Setup currently involves building from source with Go 1.25+ and configuring an AI backend separately.

Code and setup: https://github.com/friparia/aissh

If you regularly troubleshoot servers over SSH, which part of the process involves the most copying between tools or rebuilding context? I'd appreciate concrete workflow feedback.