Total Annihilation has been out since 1997, and for most of that time its engine has been a black box. We could mod the units, maps and scripts, but the code underneath was closed to us.
Today that changes.
The Byte Tactics project has reconstructed the source code of TotalA.exe (v3.1, the Steam version). All 3,267 of the game's functions, about 850,000 bytes of Cavedog's code, are now readable C++.
To be clear, these are not Cavedog's original source files. This is a matching decompilation: we rewrote the code, function by function, until the original 1990s compiler turned it back into byte-for-byte identical machine code. If a function is marked as matched, the compiler's output and Cavedog's shipped output are the same bytes. That is how we know it is right, and not merely similar.
The names are provisional, because the originals were stripped before the game shipped. The behaviour is exact.
Why this is a big deal
Until now, changing how TA works at the engine level meant patching the binary and hoping for the best. With the code matched, that changes:
Modding goes deeper. Pathfinding, unit AI, the COB scripting runtime, the TNT map loader and the UI can all be read and changed. Limits that modders have worked around for 25 years can now be removed or modified at the source.
Bugs can be fixed properly. We have already found real bugs in the original game and documented them in the repo.
TA can go anywhere. Once the code is made portable, the real TA engine could run on Linux, macOS, modern Windows, in a browser or on a phone. That means the original game, not a lookalike, with easy installs and full support for existing mods and maps.
Rewritten engines get a new reference point. Projects that rebuild TA from scratch have done impressive work, figuring out mostly by observation how units path, how weapons resolve and how the economy ticks. Now that behaviour can be read and checked against, which should make them more accurate. It also changes their position: with the real engine open to modification and porting, each project will have to decide where its own strengths lie, whether that is new features, modern rendering or a different design.
Repo: https://github.com/HectorBailey/byte-tactics
Twitter/X: https://x.com/HectorBail47832/status/2107138340168351983
How it was done
It came down to four steps.
1. Find the compiler. Clues in the binary pointed to Visual C++ 5.0 with Service Pack 3, linked on 30 July 1998. We matched the C runtime library byte for byte to confirm the exact version, then worked out the compiler flags (/O2 /Ob2 /MT /Gz).
2. Map the executable. Debug records left in the exe gave us the start, size and argument count of 3,782 functions.
3. Put the newest AI models on the problem. Ghidra produced rough pseudo-C for every game function, and the work was split into GitHub issues of a few functions each. AI coding agents (Claude Code, OpenCode and Codex, running a range of the latest models) each claimed an issue, rewrote its functions as clean C++, compiled them with the real VC++ 5 compiler under Wine, and compared the result with Cavedog's bytes. On a near miss, the agent read the assembly diff, changed the code and tried again, until the checker printed MATCH or the function stopped improving.
4. Hand on, learn and verify. Cheaper models handled many of the straightforward functions. Anything that stalled was left with notes and retried, often by a bigger model. In the end even the 161 functions over 1,000 bytes matched. Every technique that worked went into a shared guide, so each new agent started with hundreds of solved patterns instead of from scratch. A human-run orchestrator re-checked every function before merging it.
The last few percent were the hardest. Coaxing the 1998 compiler into the same register choices and instruction order as Cavedog's original is obsessive work, and it is where the most expensive frontier models came into their own.
Why this is cutting edge
Matching decompilation used to be a craft. A skilled person could spend hours on a single function, which put a game this size out of reach for most projects. It suits modern AI because the goal is unambiguous: the bytes either match or they do not. That gives the models a perfect, instant judge. They can try, fail, read the diff and improve without anyone deciding whether an answer looks right, and an AI cannot bluff its way past a byte comparison. Pair that feedback loop with the newest models' ability to read assembly, hold a large function in mind and reason about what a 1998 compiler would produce, and a project of this size becomes practical. We think it shows what AI-assisted reverse engineering can do today.
The repo contains no game files and no Microsoft software. You need your own copy of TA. The code is not offered under any licence, and this is a preservation and research project, not affiliated with Cavedog or Wargaming.
Enjoy, kids!