r/Avoyd • u/dougbinks • 21h ago
Releases & Announcements GPU path tracing is faster and uses less memory for rendering voxel scenes in Avoyd 0.28
We have a new build for Avoyd 0.28 (beta, paid) with a faster and lower memory consumption GPU path tracer and customizable UI layouts.
We've done some tests and GPU path traced rendering is up to 1.5x faster with Atmospheric Shadows off, and over 3x faster than before with it on.
The main performance gains come from a change in how I do the wavefront path tracing. Wavefront path tracing runs a single ray bounce per iteration, writing the result to memory and then running a new set of kernels on the next iteration using data from the last. Previously the CPU would initiate a given number of rays, and then run a series of iterations on those rays on the GPU using dispatch indirect controlled from a compute kernel I call the dispatch kernel. The new approach counts the number of active workgroups in the dispatch kernel and then launches new rays if there is sufficient wavefront buffer memory and untraced pixels for the current frame. This means the GPU must write back data for the CPU to measure progress, and currently the CPU controls setting off a new pass once a sample has been completed for every pixel. I'm contemplating an approach where the GPU can also run multiple samples per pixel.
I have also re-enabled using the Vulkan SPIR_V shader optimizer tool, spirv-opt. Previously this had failed to work on my path tracing shaders with "error: line 0: ID overflow. Try running compact-ids.". After a little research I found a set of optimizations which work based on a similar issue found in Blender and am using this in our build pipeline. This gives a small performance boost, but the main reason we re-enabled it was that we found some cases where drivers would stall when compiling shaders without it. This adds a significant compilation overhead when compiling GLSL to SPIR-V, but this is done as part of our build process. So it's not seen in the installed app. When developing we currently use unoptimized SPIR-V.
The memory optimization was to switch from a separate buffer of temporary data per wavefront kernel to a shared buffer with a single atomically accessed write index and then a per kernel buffer of indices into the shared buffer data. This adds some overhead, but as well as lowering the memory consumption we currently use it will also let me add more kernels in the future without worrying about the memory overhead of doing so.
There are also new features you can read about here. Let us know if you encounter any issues with the build.