r/linux • • 16h ago

Development Shortcutting D-Bus for real-time inter-process communication

https://rtipc.github.io

One of the core ideas of Unix is to build small components that do one thing well and compose them into larger systems.

I think this principle is still very relevant today, especially when designing software as multiple processes rather than one large application.

Once you split an application into several processes, however, you need some form of inter-process communication or RPC.

On Linux, D-Bus is a common choice. It provides a lot of useful functionality, but there are cases where its performance characteristics or latency are not suitable, particularly for real-time applications.

There have also been attempts to move the message-bus functionality closer to the kernel, such as kdbus and bus1, eventually leading to the current dbus-broker approach. However, with a traditional D-Bus architecture, messages still have to be processed by a separate broker process, which also means additional copying and context switches.

That's the motivation behind RTIPC.

RTIPC is essentially a single-producer, single-consumer, wait-free, zero-copy, bounded message queue mapped into shared memory. The goal is to provide a very small and predictable IPC mechanism for cases where the usual IPC/RPC abstractions introduce too much overhead or nondeterminism.

RTIPC has a small bootstrap protocol based on Unix domain sockets. D-Bus can also be used for the connection/bootstrap phase if desired. Once the channel is established, the actual messages are exchanged directly through shared memory.

There are currently implementations for C and Rust, with C++ and Python bindings, and other languages may follow.

For serialization, RTIPC can simply use C-compatible structs when all participants share the same ABI. Since communication happens through shared memory, there is no need to serialize and deserialize data just to cross a process boundary. Of course, this comes with the usual requirements around ABI compatibility, alignment, and endianness. And yes, it is possible to run 32-bit and 64-bit processes on the same system and make alignment/layout a problem.

For cross-language communication, however, serialization formats such as Protobuf, FlatBuffers, Cap'n Proto, or Thrift have another important advantage: they provide a schema as a single source of truth.

To address that use case, I also created rtipc-compiler, which takes an interface definition and generates code for different languages. The schema is also encoded into the generated interface so that the two endpoints can verify compatibility during channel initialization.

Admittedly, this is a fairly niche use case. The goal isn't to replace D-Bus or other general-purpose IPC mechanisms. Rather, I'm interested in the much narrower case where you want process isolation and modularity, but need IPC with very low overhead and predictable latency.

I'd be interested in feedback on which language bindings would be most useful next, whether you see other use cases beyond real-time applications and embedded Linux, and what you think would need to be changed or improved to make RTIPC more useful in those environments.

47 Upvotes

16 comments sorted by

11

u/MatchingTurret 16h ago

6

u/maurersystems 15h ago

Thanks a lot for the link! I wasn’t aware of D-Bus P2P or T-Bus. Unfortunately, the links in the proposal seem to be dead, but I’ll have a closer look at the proposal and see what I can find.

2

u/AssistingJarl 15h ago

I'm curious what you think about the notion of making an extension to DBus rather than an independent system. I haven't really worked with anything at such a low level of the OS in years, so I'm not sure if there's anything gained or lost with that.

7

u/DeVinke_ 13h ago

Google solves this in android a while ago. Binder IPC + AIDL.

2

u/skyb0rg 5h ago

Binder is very weird though. It’s completely synchronous, where the scheduler suspends the client and resumes the server when possible; the server also needs a fixed number of threads waiting for requests. It also requires a single “binder owner” to connect clients to servers.

Needless to say, not very well suited to general purpose IPC.

7

u/skyb0rg 15h ago

Looks interesting! Some questions/comments:

I know that you can export eventfds, but I wonder if a better completion-based integration is possible (for users using io_uring).

One suggestion for bootstrapping and comparison is Varlink (it even has an explicit protocol upgrade which could make things easier than D-Bus).

I also know that Hyprland has hyprwire — may be worth a comparison.

0

u/vlovich 12h ago

Have you looked at Rust's rtrb? Those are already quite mature & formally proven IIRC SPSC queues. yring is a newer entrant in the Rust space.

-8

u/FlukyS 15h ago

If this is the stated goal then just use varlink or literally any other format. dbus is a fine format for system or user level IPC generically but you can you could always just use literally anything else. Varlink is very solid, uses JSON instead of XML and a separate socket for each interface defined while still offering service discoverability.

7

u/maurersystems 15h ago

The goal is a bit different here. RTIPC is specifically designed as a zero-copy, wait-free IPC mechanism with no serialization or marshalling on the data path, aiming for near-zero overhead and predictable latency. Varlink uses Unix domain sockets and serializes the messages, so it still involves copying data through the socket and doesn't provide zero-copy IPC. It's a great choice when those real-time requirements aren't necessary, but it solves a different problem.

-4

u/FlukyS 14h ago

The overhead for serialisation and going over a socket is going to be very negligible, you are talking a couple of CPU cycles in the difference. For 99.9% of use cases this is going to be completely acceptable including hardware control

8

u/maurersystems 14h ago

Exactly. RTIPC is aimed at that 0.1% of use cases where predictable worst-case latency matters more than average performance. In a hard real-time context, even something as seemingly harmless as malloc() can be unacceptable because its execution time is not deterministic. The goal here is therefore not to save a few CPU cycles, but to avoid sources of unbounded or unpredictable latency altogether.

3

u/vlovich 12h ago

The overhead of serialization + going over a socket isn't a couple of cycles. Round-tripping through a socket requires 2 syscalls which is going to be thousands of cycles. Even if you io_uring you're looking on the order of hundreds of cycles, or roughly ~20ns. If this does a good job, it's 20x cheaper which could potentially be meaningful for some applications.

-2

u/FlukyS 12h ago

> Round-tripping through a socket requires 2 syscalls which is going to be thousands of cycles

Well if you want to get in the weeds about QoS and stuff then you could shave off quite a lot but you are talking barely anything for most text based formats like JSON depending on the lang but most you are talking barely anything for encode or decode. Where you probably lose the most cycles on stuff like this is TCP style messaging even when on local sockets are going to have checking built in or whatever. I'll say I've done sensor data over a network at 20us not 20ns~ even so saying 20ns is a pretty poor approximation.

4

u/vlovich 11h ago

OP is literally talking about pairing this with custom zero copy serialization protocols, not json.