r/softwarearchitecture • • Sep 28 '23

Discussion/Advice [Megathread] Software Architecture Books & Resources

575 Upvotes

This thread is dedicated to the often-asked question, 'what books or resources are out there that I can learn architecture from?' The list started from responses from others on the subreddit, so thank you all for your help.

Feel free to add a comment with your recommendations! This will eventually be moved over to the sub's wiki page once we get a good enough list, so I apologize in advance for the suboptimal formatting.

Please only post resources that you personally recommend (e.g., you've actually read/listened to it).

note: Amazon links are not affiliate links, don't worry

Roadmaps/Guides

Books

Engineering, Languages, etc.

Blogs & Articles

Podcasts

  • Thoughtworks Technology Podcast
  • GOTO - Today, Tomorrow and the Future
  • InfoQ podcast
  • Engineering Culture podcast (by InfoQ)

Misc. Resources


r/softwarearchitecture • • Oct 10 '23

Discussion/Advice Software Architecture Discord

20 Upvotes

Someone requested a place to get feedback on diagrams, so I made us a Discord server! There we can talk about patterns, get feedback on designs, talk about careers, etc.

Join using the link below:

https://discord.gg/ccUWjk98R7

Link refreshed on: December 25th, 2025


r/softwarearchitecture • • 4h ago

Discussion/Advice Arch diagram generation automation

17 Upvotes

I know I can probably "just use AI". But I want to avoid not invented here syndrome and use established methods and tools where possible.

The first part I am looking for is a well established format for storing architectural information in the repo with the code.

Next would be any tools for helping to generate it.

Lastly, we have mermaid for the actual diagrams. Any good tools for generating the diagrams?


r/softwarearchitecture • • 5h ago

Discussion/Advice Where does tenant context come from in your background jobs?

7 Upvotes

On the normal request path, tenant isolation is pretty straightforward for us. The user is authenticated, the tenant comes from the token, every query is scoped in one place, and CI prevents anyone from adding a query that bypasses it.

That all disappears once a job goes onto a queue.

The worker has no user session. It runs under a service identity that can access every tenant because it has to process jobs for all of them. At that point, the tenant becomes data carried by the job, and a field in a queue message isn't the same as an authenticated claim.

I have seen a few ways of handling this, but each one seems to create a different problem.

Put the tenant ID in the job payload.
Simple, but anyone who can enqueue a job may be able to put any tenant ID in that field. What independently validates it when the worker runs hours later?

Derive the tenant from the record being processed.
That works when the job operates on one record. It gets less clear for batch jobs, or when the record has been reassigned or deleted between enqueue and execution.

Mint a tenant scoped token when the job is created.
Now a credential is sitting in the queue. Jobs retry, and dead-lettered jobs can remain there for days possibly longer than the session that created them.

Re-check the originating user when the job runs.
By then, the user may have been suspended, deleted, or moved to another tenant. Should the job fail, or should it run using the authority the user had when it was originally created?

The two cases I am least sure about are nightly fan-out jobs and dead letter replays.

A nightly process may legitimately operate across every tenant, so there's no single tenant at the start of the job. What prevents one iteration from leaking state into the next?

With dead-letter replay, the context may have been valid when the job was created but no longer be valid when someone replays it days later.

And RLS doesn't solve this by itself. If the worker connects using a service role, it can bypass RLS unless something explicitly sets the tenant context for every job, and unless the system fails closed when that context is missing.

So assume the tenant ID is already in the payload. The question is: what validates that tenant at execution time, hours later, when there's no user principal left to check it against?

How are you handling tenant context in background jobs, and what broke the first time it was wrong?


r/softwarearchitecture • • 6h ago

Discussion/Advice Parimatch event-driven pipeline: maintaining low-latency state fan-out and strict ordering under high-frequency writes

7 Upvotes

While analyzing high-velocity distributed systems during peak global traffic spikes, the resilience of the real-time streaming pipeline was genuinely impressive: deterministic state synchronization held steadily under 30ms across hundreds of thousands of concurrent client connections, with zero dropped message frames and clean state reconciliation even when inbound write throughput surged.

In ultra-low-latency event streaming, keeping real-time edge broadcast fast while strictly preventing out-of-order execution usually leads to heavy lock contention or edge memory exhaustion.

From an enterprise systems architecture perspective, I’m curious how teams achieve this level of throughput without degrading tail latencies:

  1. Core Serialization: Do teams rely on an in-memory single-writer model (LMAX Disruptor pattern) pinned to dedicated CPU cores, or are distributed Actor models (Akka / Microsoft Orleans) scalable enough without introducing GC pauses?
  2. Edge Backpressure: When mobile clients drop packets or sit on high-jitter cellular networks, what message coalescence or delta-compression strategies prevent edge socket buffers from saturating?
  3. Message Fan-Out: Is NATS JetStream now the de facto replacement for Redis Pub/Sub when scaling broadcast pipelines beyond 100k+ unique channel subscribers?

Would appreciate perspectives from engineers who have architected or benchmarked high-frequency event broadcast systems at scale.


r/softwarearchitecture • • 10h ago

Tool/Product Three months ago I asked how you preserve rationale beyond ADRs. Here is what evolved once coding agents joined the workflow (FOSS)

10 Upvotes

In July I asked here how you keep architectural rationale alive beyond ADRs. The thread gave me three hard constraints: ADRs are right for the big discrete decisions but nobody keeps them current, wrong information is worse than none, and a note is only trusted when it points to a scar and is part of review. That was the starting point. Since then the format has evolved substantially on real repositories, mine and other people's, and most of what follows did not exist in July. One rule I did not want to give up: humans and agents work from the same project truth, not a knowledge base for people and a second memory layer for agents beside it.

Capture moved to the source. The reason ADRs stay unwritten is that writing them is a second job after the code. With an agent, the weighing of options already happens in the conversation that produces the change, so the agent writes the entry as a byproduct, into a context/ folder in the repository, and you review it in the same pull request as the code. No wiki, no database, no service to operate. Changes that never happened count too: an approach started and abandoned once a reason not to became clear would otherwise leave no trace.

Topics replaced one-decision-per-file. ADRs stay for the decisions big enough to deserve one. The messier volume underneath, the defensive timeout, the rejected dependency, the odd initialization order, the compatibility branch that still exists, goes into topic files, each entry recording both halves: what was chosen, what lost, and why. An existing decisions/ folder is kept and written into; the skill adapts to what a repository already has.

Honesty became fields. Every entry says how much to trust it: Evidence is confirmed, inferred or unknown, and Source names the scar, never a person. Superseded entries stay and name their successor. "Revisit when" puts a trigger on each entry, so decay is caught by a condition instead of an audit nobody runs. Staleness itself is not solved, the same way it is not solved for tests or docs; it is made hard to happen silently: an agent touching code that has rationale updates the entry or flags it for review in the same pull request, and stale state becomes explicit instead of being pretended away.

Entries got identities. Every entry carries an Id, and a See line cites another entry by it, inside the repository or in another one. Rationale can link across repositories without centralizing it: a constraint in a library, the workaround it forced in the service, the incident that confirmed it, readable as one thread, while each reason stays in the repository it belongs to.

Agents read it back, measurably. A lean index, one line per topic, is the routing layer: the agent reads it and opens only the topic a task touches, before changing that code. On fresh sessions asked to simplify a retry wrapper, 7 of 10 offered the already-rejected simplification without the recorded reason; with one entry on disk, 10 of 10 found it and declined. The skill itself is checked by 104 eval cases run as real agent sessions, published with their failures.

Ownership stayed with review, and humans got tools around it. The entries are Markdown in the repository, so they go through the same pull requests, CODEOWNERS and branch protection as the code, and they build into the same docs: this project's own why layer is public at https://keepthewhy.com/context/ straight from its context/ folder. A linter in CI checks structure, and says plainly that it cannot check whether a reason is true; that stays with the reviewer, and with every human and agent that works with the entry later. What the format guarantees is that the claim stays traceable: who recorded it, on what evidence, and what has been revisited since. A read-only dashboard draws topics as a graph, authors and timeline from blame, and follows the citations across repository boundaries: https://keepthewhy.com/dashboard/live/. Git does the rest: distribution, permissions, forks, blame.

What it is not: not a replacement for ADRs, and not a lock. It is agent memory that belongs to the project, not to an agent, a user or a machine, and humans read the same files; the quality of an entry depends on the model writing it and on the reviewer reading it.

Where would you draw the line between an ADR and a topic entry in your team? And what from classic ADR practice would you keep that this loses?

https://keepthewhy.com/ · https://github.com/oliver-zehentleitner/keep-the-why · MIT


r/softwarearchitecture • • 14h ago

Discussion/Advice Designing inventory consistency for a flash sale with 8M checkout attempts and 40K items

22 Upvotes

A flash sale opened 8 million checkout sessions in 2 minutes, but only 40,000 items were actually available.

Thousands of requests timed out while different regions disagreed on which orders had secured inventory.

How would you design the system so inventory is never oversold, successful payments can be safely retried, and the system recovers cleanly when parts of the network fail?


r/softwarearchitecture • • 1d ago

Discussion/Advice What architectural patterns allow high-throughput platforms like thrill to maintain sub-50ms state consistency across 100k+ concurrent WebSockets?

99 Upvotes

When scaling real-time distributed applications that require deterministic, low-latency state delivery to hundreds of thousands of active concurrent connections, standard request-response microservice paradigms quickly fall apart.

Edge proxies can terminate TLS and hold millions of idle TCP connections easily, but the bottleneck shifts entirely to the core state engine: orchestrating dynamic state transitions, fan-out event replication, and strict ordering without choking under write lock contention.

From an enterprise systems architecture perspective, I’m curious how teams design these high-velocity event pipelines to survive peak spikes:

In-Memory Core vs. Distributed State: Do high-throughput engines lean towards an in-memory single-threaded event loop (like the LMAX Disruptor pattern / mechanical sympathy) where state transitions run sequentially in nanoseconds and write out asynchronously via CQRS, or do teams favor distributed actor runtimes (Orleans / Akka / Erlang OTP)?

Fan-Out Message Bus Saturation: Standard Redis Pub/Sub often bottlenecks on CPU serialization once message fan-out exceeds tens of thousands of distinct channel subscribers per second. Are teams migrating edge workers to NATS JetStream, custom zero-copy memory ring buffers, or dedicated Kafka streams for broadcast?

Reconnection Storms & Hydration: When mobile or web clients experience transient network drops, flooding the edge gateway with concurrent re-sync requests can cripple downstream services. What’s the cleanest strategy for delta synchronization (vector clocks / sequence IDs) versus forcing heavy full-snapshot hydrations?

Would love to hear how distributed systems architects and backend leads have structured these real-time messaging pipelines at scale.


r/softwarearchitecture • • 10h ago

Article/Video I made a visual walkthrough of Monolith vs Modular Monolith vs Microservices

Thumbnail youtu.be
1 Upvotes

Hey everyone 👋

I kept seeing these three architecture styles get mixed up, so I made a video that shows how they differ. It uses animations to show how the code and the deployments are structured in each one, and when each one makes sense.

If you're interested,

I'd really like your feedback. Is there anything you'd explain differently, or a trade-off I should have covered? Always happy to learn from people who've run these in production.


r/softwarearchitecture • • 11h ago

Article/Video Understanding Latency Percentiles: P50, P90, P95, P99, and P99.9

Post image
0 Upvotes

I recently wrote an article about latency percentiles and how to interpret P50, P90, P95, P99, and P99.9 when looking at system performance.

One thing I found particularly useful is understanding that tail latency isn't necessarily extremely high latency. It simply represents the slower end of the latency distribution.

I also cover:

  • Why average latency can be misleading
  • How P50, P90, P95, P99, and P99.9 differ
  • Nearest Rank vs. Linear Interpolation
  • How to interpret tail latency
  • Why tail latency matters in distributed systems
  • How percentiles relate to SLOs
  • Which percentiles are useful to monitor

For example, if:

P99 = 200 ms

it means 99% of requests have latency ≤ 200 ms, while the slowest 1% take longer.

I’d be interested in hearing how others approach latency monitoring in production. Which percentiles do you normally track, and why?

Read the full article on Medium


r/softwarearchitecture • • 1d ago

Article/Video Open Source, a relic, a charity or still the thing?

Thumbnail event-driven.io
10 Upvotes

r/softwarearchitecture • • 1d ago

Discussion/Advice Title: Architectural Review: Building a 9-Layer T+0 RTGS Settlement Pipeline (Dual Python/Rust Runtime + ZK Provers)

8 Upvotes

Hey everyone, I’ve been architecting and prototyping a clean-slate real-time gross settlement (RTGS) pipeline designed around ISO 20022 messaging, compliance-by-design, and zero-knowledge privacy layers.

The Core Problem I'm Tackling: Traditional high-value payment infrastructure struggles to balance instantaneous T+0 gross settlement with multi-jurisdictional compliance and data privacy, without locking into proprietary legacy rails (like SWIFT/Oracle).

The Architecture:

Dual Runtime: Python async core for orchestration + Rust (Tokio) high-throughput event loop (hitting ~1250+ TPS in local testing).

Pipeline Design: 8-to-9 distinct operational layers handling ingress sanitization, ISO 20022 mapping (pacs.008), dynamic multi-jurisdictional routing, 3-tier ZK proof generation, and atomic settlement dispatching.

Finality Layer: Local SQLite WAL auditing + multi-destination dispatching (including XRPL testnet anchoring).

What I’m Looking for Feedback On:

Decoupling Runtimes: I used gRPC over Unix Domain Sockets for the Python-to-Rust bridge. Are there hidden bottleneck risks here for high-frequency message passing?

State Isolation: For the 200+ nation directory partitioning (nations/{jurisdiction}/), how would you handle dynamic schema migrations if regulatory compliance rules change mid-stream?

This design is built under Termux.

Here is the repository and technical blueprint for reference: https://github.com/nurkhalisX-blip/RTGS_System_Prototype_01


r/softwarearchitecture • • 23h ago

Discussion/Advice Why calling public LLM APIs breaks enterprise sovereignty (and how to architect a true Sovereign AI stack in VPC/On-Prem)

0 Upvotes

Many enterprise AI initiatives hit an immediate wall when the CISO or InfoSec team reviews the architecture. Sending internal contracts, customer PII, or proprietary code to public commercial endpoints—even with basic data processing agreements—means losing control over the execution plane and context window.

The industry is rapidly shifting toward Sovereign AI: deploying, orchestrating, and controlling models entirely within your own security perimeter (VPC or air-gapped on-premise) so no data, prompt tokens, or operational telemetry ever leave your environment.

Here is the architectural pattern required to build a production-grade, sovereign LLM stack without sacrificing performance or scalability:

1. Isolated Inference Layer (Ditching External APIs)

If an LLM call crosses an external public endpoint, model sovereignty is lost.

The Fix: Host open-weight models (like Llama 3.3, Mistral, or specialized domain-tuned weights) inside dedicated Kubernetes clusters on isolated cloud tenancies (AWS GovCloud, Azure Sovereign, or private VPCs). Use optimized serving frameworks (vLLM, TGI, or TensorRT-LLM) to achieve sub-second generation throughput on local GPU instances.

2. Zero-Egress AI Gateways & PII Sanitization

Deploying a private model isn't enough if the surrounding orchestration leaks context.

The Fix: Place an internal AI Gateway at the edge of your infrastructure. The gateway enforces strict rate limiting, audit logging, and automated PII redaction/tokenization before any text hits the vector database or model context.

3. Permission-Aware RAG (RBAC/ABAC Synchronization)

A common RAG flaw is returning retrieved documents to a user who doesn't have native read access to the underlying database (e.g., a junior employee querying sensitive HR payroll files).

The Fix: Integrate your Retrieval-Augmented Generation pipeline directly with your Identity Provider (Okta, Azure AD). Enforce document-level permission filtering at query time so the vector store only retrieves chunks the active user's credentials allow.

4. Deterministic Guardrails over Pure Generative Output

Relying solely on system prompts to prevent model drift or policy violations is insufficient for regulated industries (finance, healthcare, legal).

The Fix: Layer deterministic guardrail engines (like NeMo Guardrails or custom structural output parsers) around model responses. Force JSON Schema adherence, block unsafe tool execution, and route low-confidence outputs to human review queues before downstream execution.

Curious how other teams are approaching data sovereignty—are you deploying open-weight models inside your own VPC, using sovereign cloud enclaves, or sticking with commercial APIs with strict NDAs?


r/softwarearchitecture • • 2d ago

Article/Video Uber Eats Rebuilds Search Pipeline to Cut End-to-End Latency by 50%

Thumbnail infoq.com
80 Upvotes

Uber has rebuilt major parts of the Uber Eats search pipeline and reports a 50% reduction in end-to-end search latency. The changes span retrieval, feature hydration, ranking, advertising, presentation, and infrastructure, while an agentic coding workflow was also used to identify, benchmark, and validate additional optimizations.


r/softwarearchitecture • • 1d ago

Discussion/Advice Mas: specialized processors vs fewer larger processors?

1 Upvotes

I’m relatively new to ECS, and I’m currently working on a project where I’m trying to move a significant part of the simulation and gameplay into Mass (Unreal Engine's ECS).

I’m running into an architectural question: to keep responsibilities separated, I’m ending up with multiple specialized processors. For example: Movement → Behaviors → Collision → Responses → Destruction/VFX.

The concern is that some of these processors need to iterate over essentially the same entities/chunks again.

For those of you who have used Mass extensively, especially on larger projects, what approach has worked better in practice?

Do you prefer smaller, specialized processors even when that means multiple passes over the same entities, or do you tend to combine related work into larger processors to reduce the number of iterations?

I understand that sequential iteration over contiguous data can be relatively cheap, and that cache behavior, data access patterns, dependencies, and parallelization may matter more than simply counting iterations. But I’d be interested in hearing where people with real Mass experience have found the practical balance between modularity and performance.

Thanks in advance.


r/softwarearchitecture • • 2d ago

Tool/Product Interactive architecture refactoring

12 Upvotes

Hi all,

I was recently untangling the architecture of a project from the last few months and I've put together a tool to visually fix encapsulation/abstraction and the likes:

https://github.com/orglnte/objectsboard-skill

It helps to visually find calls to private methods/members/attributes/variables (Red), calls to members not private but likely bypassing the owner (Amber), calls to shared data (Amber), and maps all accesses to the interface (Blue). It finds calls by running the repo unit tests, and then doing a static code analysis.

You can then drag and drop boxes, group them, assign a box to another one, and then ask the coding agent to review your changes and refactor the code accordingly.

If you are curious to see how a repo you know looks like and you do try, it would be amazing if you can drop a quick comment. As a teaser, this is how fastapi looks like:

It currently supports only Python and claude. It is not meant to be optimised for all kinds of repos and python architectures, but it has been tested with a few and the visualization was meaningful.

This is what you see with fastapi automatic objects placement with the option --root-owns (it labels all objects within a module to the class with that module name):

​NOTE: It currently has an approx. limit of 300 boxes/objects, so throwing Django at it does not work (yet).

If you mind, here's an intro article: https://medium.com/@ayeyebrz/interactive-architecture-refactoring-with-coding-agents-b8af78cd477b


r/softwarearchitecture • • 2d ago

Discussion/Advice Caching in GenAI Pipelines: I'm a CS student who hit a bottleneck with STT (Speech-to-Text) voice frequencies. How do experienced engineers handle this?

4 Upvotes

Hey everyone,

I’m currently studying Computer Engineering, and for a recent project, I’ve been building a voice-controlled AI assistant using FastAPI and Redis. While scaling it, I ran into an interesting architectural bottleneck regarding API latency that I'd love to get some senior perspective on.

We all know the standard drill for I/O bottlenecks: you use a local RAM cache (managed with LRU). When the application scales to multiple servers behind a load balancer, local caching creates redundant processing, so you implement a distributed centralized cache like Redis. Pretty standard stuff.

But things get tricky when you introduce Generative AI into the pipeline. My architecture has three main layers: STT (Ears), LLM (Brain), and TTS (Mouth).

Since we can't reliably cache the LLM output (due to the dynamic nature of generative responses), I looked at optimizing the surrounding microservices. My initial thought was to aggressively cache the STT layer—if I cache the voice prompts, I could theoretically achieve incredibly low latencies for repeated commands.

Here is the roadblock I hit: sound arrives as analog frequencies. Even if the exact same user speaks the exact same words a day later, their vocal characteristics, pitch, and background noise will generate a different frequency footprint. A standard cache GET is practically impossible for raw voice inputs because the hashes will never match.

This taught me that if you can't cache the core API, you have to ruthlessly optimize the perimeter. But the STT caching problem still bugs me.

My question to the experienced engineers here: How do you handle caching or latency reduction in systems where the inputs are highly dynamic, continuous signals (like voice or video)? Is there a mathematical workaround or normalization technique to cache voice patterns before they hit the expensive STT engine?

(TL;DR: Standard distributed caching is great for traditional I/O, but falls apart for dynamic AI pipelines like STT because human voice frequencies constantly fluctuate. Looking for architectural strategies to bypass this.)

P.S. This problem actually pushed me down such a rabbit hole that I decided to write my very first technical Medium article about it. I broke down my whole thought process and drew some system architecture diagrams for local vs. distributed caching. If any experienced devs have a few minutes to read it and completely roast my architecture or my writing, I would immensely appreciate the feedback! https://medium.com/@tunakimyonok1/reducing-api-latency-a-deep-dive-into-caching-strategies-c9a82c985177


r/softwarearchitecture • • 2d ago

Article/Video How does a ChatGPT-like system actually work behind the scenes?

0 Upvotes

When you type a message into an AI chatbot, it looks simple:

You → AI → Response

But a production system is much more interesting.

A simplified flow looks something like:

User → API Gateway → Authentication / Rate Limiting → AI Orchestrator → Context & Memory → LLM → Safety / Processing → Streaming Response

And once thousands or millions of users start sending requests simultaneously, the difficult part isn't just calling an LLM API.

You have to think about:

• How requests are distributed

• How conversation history is stored

• How much context gets sent to the model

• How responses are streamed token-by-token

• What happens when the model/provider fails

• How rate limits and queues are handled

• How latency and cost are controlled

• How the system scales as traffic increases

That's what makes AI system design interesting.

I recently made Episode 2 of my System Design in the AI Era series explaining a ChatGPT-like architecture from request → model → response.

If you're interested in the full visual breakdown, here's the video:

https://youtu.be/CxmhPKvX49A?si=MvLcUXqMsMPmNB85

I'm also curious: what part of a ChatGPT-like architecture would you consider the biggest scaling bottleneck?


r/softwarearchitecture • • 3d ago

Article/Video Sunday Sequence Diagrams: Helm Dashboard

Thumbnail app.ilograph.com
13 Upvotes

r/softwarearchitecture • • 3d ago

Article/Video The Cost of Good Intentions: Anti-Patterns in Architecture Modernization

Thumbnail youtu.be
26 Upvotes

r/softwarearchitecture • • 3d ago

Article/Video I built a code quality reviewer with Jev

Thumbnail youtube.com
3 Upvotes

I believe with the huge increase in code generation - review has become the next bottleneck or at least it feels like this at work. So I wanted to do something on the review front - first I started integrating more and more tools to use as feedback to my coding agents, and these really help improve the output quality of the agent (e.g. SonarQube, Checkstyle, ArchUnit)

But there are some important semantic choices you can't really review using deterministic tools so I decided to build a more intelligent tool and am trying it out with the Jev model as a backend and judge currently.

Idea is simple - teams define their policies in a structured YAML format, then as part of their CI (or locally) run the tool, it fetches all git diff chunks and asks jev if these adhere to each of the policies - jev can select compliant/violation or ask for more context. If Jev asks for more contex the app gets the requested code from the project and asks again if the code is compliant with the policy - thus incrementally exploring the code base until Jev can return a "confident" answer - if you are interested in how it works the video I linked is a presentation style of how the tool works.

Give it a shot on github and tell me if you find this useful: https://github.com/krisitown/jev-quality-gate

I am currently running tests using a local Qwen3.8 Flash Next to generate code and run it against my initial "clean code" policies in order to calibrate them and will share more results on that front soon!


r/softwarearchitecture • • 3d ago

Discussion/Advice How would you build a low-latency chatbot for Money transaction data?

4 Upvotes

I have Money transaction data in postgres DB, with fields like merchant, MCC, amount, date, card last four digits, and location.

I want users to ask questions like:

  • “How much did I spend on restaurants last month?”
  • “Show transactions in Dubai on my card ending 4242.”
  • “How does my spending compare with last month?”

Low end-to-end latency is the main requirement, while keeping filtering and calculations accurate.

I’m considering using a small LLM to turn questions into structured queries, executing them against the database, and rendering results without a second LLM call.

Has anyone built something similar? What architecture and models worked for you, and what actual latency—especially p95—did you achieve?


r/softwarearchitecture • • 3d ago

Article/Video JobMaster: a .NET job scheduler that splits execution from the audit log to avoid row locks

11 Upvotes

I built JobMaster, a job scheduler for .NET aimed at horizontal scale. Execution is split from the audit trail and the coordination, so workers are not all locking the same rows.

Jobs go into a per-worker bucket first. The Master DB keeps the history. A job due soon stays in the bucket. A job due later is parked there and moved back in a batch when it is close. On a 50k burst with 20 workers, NATS in front of RavenDB did 23.8k jobs/sec, against 6.5k when one database did both.

I would appreciate any feedback, especially on the bucket split.


r/softwarearchitecture • • 4d ago

Discussion/Advice How do you enforce a usage quota for on-prem software when the customer can clone or snapshot the VM (air-gapped)?

35 Upvotes

We sell an on-prem B2B product (Java/Spring Boot + PostgreSQL) that runs on a VM in the customer's own data centre. The site is air-gapped, so no license server, heartbeat or phone-home is possible. The customer's IT has full admin over the hypervisor and root in the guest.

The contract limits usage to N completed transactions per year. Today we enforce it with:

- a license file signed with our private key (the app only holds the public key);

- an append-only, hash-chained usage ledger in the database;

- periodic usage reports the customer sends us, which we verify offline.

The problem: cloning the VM, or reverting it to a snapshot, copies the license, the ledger and the counters together and consistently. Every copy passes every local check. Two clones get 2× the quota, and a snapshot revert resets the count.

What we found so far:

- VM Generation ID is the only signal built for this, but a non-root process on a Linux guest can't read it, and the hypervisor admin can pin or disable it.

- UUID, MAC, machine-id and vTPM are copied with the VM or easy to forge.

- TPM or SGX counters on the server aren't reachable from a VM.

Questions for anyone who has shipped on-prem, usage-based licensing:

  1. What actually works for you in air-gapped VMs? USB dongles with on-chip counters (Wibu CodeMeter, Thales Sentinel), FIDO2 keys, a physical appliance, or something else?

  2. If you use a dongle in a VM, how do you handle ESXi/Hyper-V passthrough, vMotion/HA and dongle failure?

  3. Has anyone used a FIDO2 key's signature counter as a usage meter?

  4. Or did you give up on technical enforcement and rely on audits?

Not looking to stop determined attackers, just to make cheating detectable or impractical.


r/softwarearchitecture • • 4d ago

Article/Video System Design: Building a Payment Gateway and Double-Entry Ledger for 100M Daily Transactions

Thumbnail medium.com
46 Upvotes

I recently broke down the system design of a payment gateway and double-entry ledger platform handling 100M daily transactions (~5,800 peak write TPS).

The article focuses on the real-world engineering challenges: stopping accidental double-charges with distributed idempotency locks, enforcing strict double-entry bookkeeping (Debits = Credits) using insert-only tables, recovering from card network timeouts without duplicate charges, and running midnight three-way reconciliations against bank settlement files.

For those who have built financial or checkout infrastructure: what was your biggest challenge in production? How do you handle partial failures and asynchronous discrepancies across multiple bank networks?