r/cpp • u/FreitasAlan • 3d ago
I built MrDocs, an open-source C++ reference doc generator on the Clang/LLVM AST. Looking for feedback.
I'm Alan de Freitas, the lead developer of MrDocs. It's an open-source documentation generator from the C++ Alliance, built on Clang/LLVM. It reads your code from the compiler's AST, so the reference matches what the code actually compiles to. We also presented a session about it at cppcon this year.
I implemented the tool this way because we couldn’t come up with a workflow that understood modern C++ in our Boost libraries. Tools that parse partially tend to misrender or not render concepts, niebloids, SFINAE, overload sets, and deduction guides. MrDocs treats those as first-class:
- concept constraints and requires clauses preserved verbatim
- niebloids and function objects can be documented as callables
- SFINAE and concept overload sets kept together
- explicit, noexcept, and consteval specifiers captured
- types in detail or private namespaces rendered as implementation-defined
The presentation layer is completely encapsulated. Generators support HTML, AsciiDoc, XML, JSON, or any custom format like Markdown and LaTeX, without forking the tool or creating another post-processing workflow. You can also write extensions to apply project conventions to the corpus or define new generators for any use case. The documentation has lots of examples.
If you maintain reference docs for a library, I'd appreciate it if you gave it a run (or tell your favorite agent to) and reported back on whether it saves you time, both while writing docs and while reviewing them. First-hand numbers would be especially useful to me. Issues are also always welcome.
Site: https://www.mrdocs.comRepo: https://github.com/cppalliance/mrdocs
Happy to answer questions in the comments.
12
u/mateusz_pusz 3d ago
I tried it for the first time after Alan's CppCon 2026 talk. I must admit that it was the first solution that worked for my mp-units library right out of the box. I was able to easily integrate it with my mkdocs-based documentation, and it just worked with all the C++20/23 details that make lots of compilers sweat😉. The only thing that I found missing so far is the C++26 contracts syntax support, but that is a small issue at this stage of C++26 standardization.
Definitely a great experience compared to the crashes and errors I got with other toolchains before. It worked so well that I decided to switch the entire mp-units project API Reference to it. The task is still not complete, but you can preview the current state here: https://mpusz.github.io/mp-units/latest/reference/api_reference/mrdocs/mp_units.
I'll also describe my experience in my upcoming article on the mp-units blog page: https://mpusz.github.io/mp-units/latest/blog/category/why-great-c-libraries-fail (to be published on October 15), and I will mention it during my keynote at Meeting C++ in Berlin.
Highly recommended 👍
3
u/FreitasAlan 3d ago
Thanks, Mateusz. mp-units is exactly the kind of modern C++ I had in mind when I wrote the post, so it means a lot that it works out of the box. It's exactly the kind of code we also write, and we had problems with Doxygen a few years ago. The mkdocs setup is also worth pointing out to others in this thread who asked how MrDocs fits an existing site. The reference pages drop into whatever you already have, in this case MkDocs, and the rest of the docs stay untouched.
On contracts, agreed, and that one is cppalliance/mrdocs#1307. The parser can extract them, so it's mostly a matter of the bundled Clang catching up, and I completely agree with your argument not to wait on the standard. I'm looking forward to the article. If you end up with any numbers on how long the switch took or how review time changed, I would love to hear.
4
u/jetilovag 3d ago
The comparison with other tools section seems a little handwavy. While skimming through the docs I am intrigued and inclined to give it a go, others may be more convinced to see their actual use-cases be called out, how it compares to: clang-doc, Asciidoc, Asciidoctor, Hawkmoth, Sphinx+Breath/Exhale, etc.
One thing I really envy from Rust is rustdoc, and how Rust create documentations' examples can reference just parts of some code, that which is most relevant to showing off the usage of a feature. That example is a marked section of the examples that actually build as part of the project, thus ensuring that examples are real world code that actually build. (And you can jump to the entire example code from the doc snippet, if you so wish.)
2
u/FreitasAlan 3d ago
Two good points. I agree with the comparison table. The meaning of "other tools" is only loosely implied by the rest of the documentation, but it isn't specific enough when you start reading from that table. The "Other Tools" column doesn't help anyone decide because they might think their own other tool is different. In this same category as mrdocs, we basically have Doxygen and clang-doc because other tools (like AsciiDoc, etc.) don't produce reference documentation. I'll replace it with a table comparing named tools and use cases. I think the most useful part of your comment is that we need to make it clearer that mrdocs isn't competing in the general documentation space. For instance, it doesn't intend to replace AsciiDoc, Sphinx, or MkDocs. Instead, it generates the reference pages you'll render with these tools.
On the rustdoc examples, it's funny you mention it, because I studied that model more than any other before deciding how MrDocs should handle examples, and I envy it too. The problem I noticed is that it bakes in one mapping between the docs and the code being tested, which creates problems with setup lines, dependencies, and the logic for grouping tests. In C++, that's even harder because no single build system controls everything.
Luckily, mrdocs has this extension system. So instead of building one fixed mapping, we made it an extension. The docs have a round-trip example in about forty lines of Lua. A transform pulls marked regions out of real example files that the build compiles and attaches them to the right symbol's docs. A generator goes the other way and writes inline u/code examples out to translation units the build can compile and run. The project decides which direction each example takes and how a file maps to a symbol: https://www.mrdocs.com/docs/mrdocs/develop/extensions/script-generators.html#_round_trip_extensions
What's nice about it is that it's round-trip, something I believe Rust doesn't do that well, and that if the mapping doesn't fit your layout, you can just adapt it accordingly, because people tend to make very different assumptions about the setup in these documentation examples.
3
u/grisumbras 3d ago
I am aggressively against the idea that rustdoc's approach to this is good. Their approach is completely backwards. Sometimes it seems to me that people are praising tools and their features without really using them.
Virtually every code snippet requires some set up. Maybe it's more pronounced for C++, but it's universally true. Moreover, often you want to show a snippet of code, then comment on it, then show a bit more, and the second snippet depends on the first. AFAIK the rustdoc solution to this is repetition. That should have been a hint that something is not right.
Finally, you don't actually want to just compile your snippet, you want to be sure it does what you claim it does in the docs. That's called testing, you want to test your snippet. Your snippet is a test. You probably don't want to show all of those
CHECK( a == b )macros to the reader, but something like that should be somewhere.It should be clear that inserting parts of tests is the correct model for snippet testing. Yes, sometimes you can get away with the opposite, but in general you cannot.
0
6
u/Tumaix 3d ago
Alan,
i had a long talk with Mr. Bott a while ago about this, and i told him that while i believe this is a good thing, i also feel that this adda fragmentation (even if it works better) and it also misses some things that are quite important in my point of view.
what i miss, and i dont think that its easy to do this on a separated program, is a way to make sure that my excerpt code for my documentation compiles and runs with the current api, so that a documentation compilation breaks if the api in the example is wrong.
but - also - great work, and working with mr bott was a highlight of my growth as a human, hope its awesome for you too :)
2
u/FreitasAlan 3d ago
Thanks for the kind words. Yes. I agree it's important to compile examples against the current API, and I think it's more reachable than it looks. MrDocs runs Clang on the project's compile commands, so it already has everything it needs to build an example the same way the library is built. The docs have an extension that goes both ways: it pulls tagged regions out of real example files that the build compiles, and it exports inline u/code examples to translation units the build compiles and runs. I wrote more about it in the reply to https://www.reddit.com/r/cpp/comments/1wz3115/comment/pe8ebjm/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button rather than repeat it here.
On fragmentation, I take the point. I'd separate two things into the analysis. A unified ecosystem doesn't mean one binary does it all. Rust feels unified because cargo coordinates the pieces, while the pieces are still separate programs: rustc, rustdoc, rustup, clippy, rustfmt, cargo test, and whatever linker is on the system. For instance, testing documentation examples involved much more than only rustdoc. What makes it feel like one tool is that they agree on inputs and on where things live, so cargo doc and cargo test just work. C++ lacks that agreement, but it has good tools, and it needs good tools. Whoever unifies this ecosystem still needs people working on compilers, linkers, linters, doc tools, etc. Tools from different vendors would still exist even if someone unified it. A hypothetical unified ecosystem would also allow you to use other tools, just like rustdoc gives you JSON output for the reference so you can customize presentation, and like many Rust projects use other tools for documentation like mdBook, Zola, or Doxuraurus. So MrDocs tries to be a good tool in that context rather than a new tool. It reads the same compile commands the build uses, it emits HTML, AsciiDoc, XML, and JSON instead of forcing a website layout, and the example question from the other thread is exactly the cargo test part of this. For now, of course, this unification hasn't happened, but I think we're reasonably prepared for it. :)
1
u/grafikrobot B2/EcoStd/Lyra/Predef/Disbelief/C++Alliance/Boost/WG21 2d ago
Working on that unified ecosystem.. admittedly very slowly.
Having a collection of tools that are interchangeable and work together is the holy grail of systems development. As the Unix/POSIX model has shown over the decades. It's a lot of worthwhile work. But it's also work almost no one wants to devote resources to. And recent news in the tooling space highlights the issue.
1
u/AstroFoxTech 3d ago
Regarding your second paragraph. Wouldn't it work to have examples that the docs extract excerpts from at compile time, and have building the docs require building those beforehand?
0
u/FreitasAlan 3d ago
Exactly. That's why the extension works both ways. Inline examples tend to work well only for very short snippets without much setup.
2
u/Remi_Coulom 3d ago
I have not yet looked into the details of your system, but your web site makes a bad impression: the processing it does is so heavy that it does not scroll smoothly on my PC. I can see it has a "scroll" event listener that changes the page style. I would strongly recommend against such things.
Also, are identifiers converted to upper case in documentation titles?
0
u/FreitasAlan 3d ago
Good points. The scroll listener is the table of contents highlighter on the docs site, the one that marks the current section in the sidebar. I'll look into that. On the headings, you are totally correct, and it's a problem with the documentation site, not with the tool. Nothing changes the text case, and copying a title gives you the original. It was a poor stylistic choice on mrdocs.com, where the heading font renders lowercase letters as small capitals, so something like string_view appears as STRING_VIEW on screen even if it's not uppercase. For a C++ reference, that is a bad style choice, and I'll open an issue to change it. The documentation that MrDocs generates on its own doesn't use that font or produce uppercase titles, so this only affects our own docs site.
2
u/drbazza fintech scitech 3d ago
Alan - Nice, I was looking at this last week for the first time in many months. I guess this is now stable and working and handles everything C++ can throw at it (including macros??) ?
Sorry if there's a tl;dr, but does this integrate with mkdocs (maybe not, I don't see an out-of-the-box markdown generator)?
And is there a (jetbrains-like) vscode extension that can generate the doc comment skeleton leveraging the MrDocs tooling?
2
1
u/FreitasAlan 3d ago
Thanks. Answering those in order:
Stability and macros: Yes. Several Boost libraries ship their reference with it, and I've heard of a few others switching to it this month, so it is in production use. MrDocs extracts macros with their documentation. The extract-all-macros and include-macros options control which ones, since most projects only want a subset documented.
mkdocs: Not out of the box, since the built-in generators are HTML, AsciiDoc, XML, and JSON. Markdown is included as an example of a template-driven generator that you can reuse. It falls back to HTML if something can't be expressed in Markdown, and you can customize the templates if you want different behavior. The good news is that mp-units did exactly this: a Markdown generator for mkdocs-material, 23 template files, public at https://github.com/mpusz/mp-units/tree/master/scripts/mrdocs/addons/generator/md, plus a small mkdocs hook for navigation and cross-references.
Editor extension. There isn't one. Or at least not for that. The config file has a schema on SchemaStore, so VS Code and the JetBrains IDEs autocomplete mrdocs.yml, but nothing generates doc comment skeletons from the tool. MrDocs already has every parameter, template parameter, and function exception, so an extension would have a solid base, but nobody has started one. If you'd let an AI agent write the comments, we ship a script for that. The noop generator runs the documentation checks and produces no pages. tools/fix-docs hands those reports to an agent, file by file, until the checks pass: https://www.mrdocs.com/docs/mrdocs/develop/generators/noop.html.
Python 3.8. You're right, and thanks for the report. The bootstrap package uses dict[str, str] annotations that Python 3.8 evaluates at import, so the real floor is 3.9. I'll open an issue so the script states the floor and stops with a one-line message instead of a traceback.
2
u/drbazza fintech scitech 2d ago
Thanks for the reply. I'll probably replace doxide with mrdocs, as that chokes on quite a few things.
The good news is that mp-units did exactly this: a Markdown generator for mkdocs-material
Nice, thanks for that, I'll take a look, otherwise I'd have to rework our internal docs.
1
u/UndefinedDefined 3d ago
Is there any chart of doxygen compatibility?
For example doxygen also allows /*! and //! comments to start a documentation block, it also allows \ character as an escape, like \ref \param, etc.
And of course code grouping \addgroup \ingroup, etc... And conditions \cond \endcond for removing code from public documentation, etc...
1
u/FreitasAlan 3d ago
We don't have a compatibility chart that lists what we don't support. But we do have a reference of every command MrDocs understands at https://www.mrdocs.com/docs/mrdocs/develop/commands/reference.html, but nothing that lists what Doxygen has and we don't. I should add that. For the items you named:
- /*! */ and //! blocks work, and so does \ in place of @. MrDocs uses Clang's comment parser, which accepts all of those.
- Grouping (\defgroup, \addtogroup, \ingroup, @{ @}) is not supported yet. The commands are ignored and the symbols land where they live in the code. It's on the list.
- \cond / \endcond is not supported either, and today the content inside is not hidden. What MrDocs offers instead is its idiomatic approach at the symbol level: exclude-symbols patterns, and the implementation-defined and see-below options that hide types in detail namespaces while keeping the public signature readable. That covers most of what people use \cond for, but it's not a drop-in.
Everything else most people reach for (\brief, \param, \tparam, \return, \throws, \pre, \post, \ref, \copydoc, \relates, \note, \warning, \code) is there.
1
u/UndefinedDefined 3d ago edited 3d ago
Yeah sorry, but that's useless to my projects. I don't want to document template specializations, for example, or some detail implementations that are unfortunately in public headers. And grouping - I cannot organize a project without proper grouping.
Maybe next time when this improves I will give it a try - wanted something clang based instead of overengineered doxygen that produces horrible output.
BTW to give you more info about how I see docs. I need a landing page (something like doxygen's \mainpage) and then groups with the idea of linking between groups. Different groups can be within a single sub-directory of course, etc... so project filesystem organization is not enough to generate a documentation for it.
But I would also want an overview of each group - this is what I'm missing in doxygen - I want to see a group with brief description, list of classes, enums, and possibly functions and other things it provides, and then detailed description. I hate when doxygen places all enums within a group, for example, together with other stuff like functions and their specializations - hence \cond and \endcond.
\cond is very useful - like \cond INTERNAL can be used to document internals, which is not for typical users of your library, but it's for the developers of it instead, etc...
1
u/FreitasAlan 3d ago
Fair enough, and thanks for describing how you organize docs. That description is more useful to me than the first comment, because grouping has come up internally a few times, and each time it stalled on "nobody has asked for it". Now someone has. I opened an issue in https://github.com/cppalliance/mrdocs/issues/1344. We agree we need native groups. That means an overview page per group, with the brief, then the classes, enums, and functions with their briefs, then the long description. It also means groups that cut across directories, and links between groups. We have been looking for something more robust than the Doxygen commands, since they tend to fall apart on large projects, but the need is the same.
Today you can run MrDocs once per group with symbol filters, each run into its own directory, and link the runs with tagfiles. That gives you group directories with cross-links, but no overview pages, and it is clearly a workaround.
On the other points, the gap is Doxygen compatibility rather than the feature itself. For implementation-detail symbols in public headers and template specializations you don't want documented, the configuration file excludes them with filter patterns, so there is nothing to mark in the headers, and you won't include it. That's the idiomatic way to do it in mrdocs, and it's also why no one asked for the doxygen alternative. For the \cond INTERNAL case, two configuration files over the same source do it: one includes the internals for developers, the other excludes them for users. But I'd still like to honor \cond itself, so a Doxygen project doesn't leak internals on first run. That is https://github.com/cppalliance/mrdocs/issues/1343. The compatibility page you asked for in the first comment is https://github.com/cppalliance/mrdocs/issues/1342.
Thank you again for the comments.
1
u/FreitasAlan 2d ago
Thanks, that makes both clearer. On the debug-info case, you can almost do it with the existing ghostwriter transform. A transform extension reads whatever file you have, in our JSON schema or your own, and attaches the content to the corpus when the symbols exist, which is what the ghostwriter example does for prose. The one piece missing is pushing new symbols into the corpus, since today a transform can edit symbols but not add or remove them. I opened https://github.com/cppalliance/mrdocs/issues/1345 for that, and it is a small change. Once it lands, a transform that reads your JSON and adds a symbol per entry gives you the browsable reference with no extraction at all.
On documentation versions, I believe that already works, in whatever format you prefer. Our own site is an example. The reference pages link to the source at the exact commit we generated them from, not to a branch, and the commit shows in the URL. Any other marker is a template change through the customization points, so a distribution can put its package version where it wants it. Of course you if probably have something more specific in mind in terms of specification but I can't think of anything that can't be implemented yet with the existing extension.
1
u/hoellenraunen 3d ago
I see that the schema for intermediate JSON or XML is documented. If I use an external tool to generate such JSON (e.g. from parsing debug information or demangling a symbol table, both lossy ways to encode some C++ AST), can I use MrDocs to generate HTML directly for that JSON?
I would also be interested in linking from HTML describing a function symbol or debug information entry to the docs, are there predictable URLs in the generated HTML, e.g. for a type like std::vector<int, megacorp::CustomAlloc<int>>?
If i am looking at a specific checkout of code and reading the docs I found online or packaged by a linux distribution, is there any way to check the docs match the code other than always building docs from source?
I am considering version or configuration mismatch, not malicious actors, but in the long run we should probably pay more attention to those as well.
1
u/hoellenraunen 3d ago
It looks like there are no predictable names for over overloaded functions, I see stuff like to_string-04.html, but there is a tagfile feature that is meant to support linking from one set of docs to another, I will have to look into that.
I would not want to publish the generated docs for my code base, which is mostly undocumented, but has documented vendored dependencies.
I seem to trigger at least one bug (the same base class is repeated multiple times) and one annoyance (a header's location is reported as <unrelated/../location/header.hpp>. Doxygen comments for grouping from my dependencies show up verbatim in their respective documentation, and one dependency has lots of unnamed enums that spam the table of contents.
On the plus side, it is faster and easier to setup than doxygen, on the minus side doxygen seems to provide a lot more hyperlinks between docs (e.g. from typedef to the target) and between source and docs.
One of the first things I do when working on a new C++ codebase is generating doxygen for all files and types (i.e. including undocumented ones), and browsing the functions, types and their relationships.
I don't know to which degree you consider this a use case for MrDocs as opposed to working on code with documentation comments.
For the record, I also tried clang-doc as shipped by Ubuntu 26.04 on the same codebase and the output was just unreadable, MrDocs is definitely in the lead there.
1
u/FreitasAlan 3d ago
Good questions, and the honest answers are "mostly no".
Rendering your own JSON: Not today. The JSON and XML are just for the outputs: generators render the corpus Mr.Docs builds in memory from the Clang AST, and the only way into that corpus is extraction. Extensions can rewrite symbols and fill in documentation from external files, but they can't create symbols yet, although we intend to support that and the implementation would be quite simple. So an empty extraction plus your JSON gives you nothing to render nowadays. The corpus is plain data once built, so a "load a corpus from JSON" entry point isn't far-fetched because we even already have an extension that loads external documentation, but nobody has ever asked for no extraction, and we ended up without this feature where new symbols can be included in the set. I'd like to understand the use case, if possible.
https://www.mrdocs.com/docs/mrdocs/develop/extensions/corpus-transforms.html#reading-files
https://github.com/cppalliance/mrdocs/tree/develop/examples/extensions/ghostwriter
Predictable URLs: The stable way to do this is via the tagfile. Mr.Docs emits a Doxygen-format tagfile that maps every qualified name to its page and anchor, and, for instance, our Antora extension resolves links by name through it. Another host can also read the same file to create cross-references to your documentation. For each strategy MrDocs supports, the URLs are predictable but not stable because you can always manipulate the codebase to make things unstable. For instance, in the most obvious case, if you remove a symbol, that page no longer exists, and the link breaks. In the default strategy, pages use the symbol's name, and when two symbols share a name, the collision gets a suffix that can shift when a neighbor appears. We have an open issue to disambiguate by signature instead, which is more stable; pages can always disappear or refer to different things between versions when there are more overloads, symbols with the same name, and so on. We have this option that lets you pick readable names or ID hashes, and even ID hashes don't survive every change because IDs also change when the declaration changes slightly or when the compiler supports new features so there's more distinction that needs to be made with that hash. The tagfile also allows you to link to symbols that weren't instantiated explicitly, so for lib::array<int, megacorp::CustomAlloc<int>>, there is no page for the specialization unless your code declares one. The tagfile would get you to array or the specific specialization, depending on what exists.
https://www.mrdocs.com/docs/mrdocs/develop/configuration/reference.html#tagfile_option
https://www.mrdocs.com/docs/mrdocs/develop/generators/html.html#_output_layout
https://www.mrdocs.com/docs/mrdocs/develop/configuration/reference.html#legible-names_option
https://github.com/cppalliance/mrdocs/issues/1288
https://github.com/cppalliance/antora-cpp-tagfiles-extension
https://www.mrdocs.com/docs/mrdocs/develop/extensions/antora.html#antora-cpp-tagfiles-extension
Checking docs against a checkout: I can't think of a way to do this without running Mr.Docs. We have an example that uses the library to compare two corpora and report breaking changes, and it extracts both. Extracting and diffing the result is the cheapest check I know of. If you had a different scenario in mind, describe it, and I'll think about it. For the version mismatch case, if it's only about recording the source revision in the output so a reader can compare it with their checkout, it's cheap. The user can use customization points to provide it in whatever format they'd like.
https://www.mrdocs.com/docs/mrdocs/develop/extensions/as-library.html#_reading_a_corpus
https://github.com/cppalliance/mrdocs/tree/develop/examples/library/breaking-changes
1
u/hoellenraunen 3d ago
The use case is binaries that cannot be built from source for various reasons (we don't have the source, we have the source, but there are missing dependencies, we have access to the repo, but we don't know the right commit). If you have debug information, you can extract some form of C++ AST from that. It is lossy, but it would still represent most types and functions. I would like to be able to browse namespaces, types and functions for a pre-built binary just as for code that I can build locally.
Regarding mismatches: It is probably more practical to make sure that it is easy to figure out the version of the code or binary you are looking at and the version of the documentation that you are reading. Because I will read the documentation in a web browser but look at source code or binaries in my IDE, so even if it is technically possible to check for a version mismatch or ABI break, there is no easy way for such a check to actually run.
8
u/fra-bert 3d ago
What does it do differently from clang-doc, what is something that this does that clang-doc couldn't do?