The compute graph¶
A klartraum::ComputeGraph is a directed acyclic graph of
klartraum::ComputeGraphElement nodes. Each element produces an
output (a buffer, a tensor, an image) and can take the outputs of other
elements as inputs.
auto blur = std::make_shared<BlurOp>(); blur->setInput(renderpass);
auto noise = std::make_shared<NoiseOp>(); noise->setInput(blur);
auto add = std::make_shared<AddOp>(); add->setInput(blur, 0); add->setInput(noise, 1);
klartraum::ComputeGraph graph(vulkanContext, numberPaths);
graph.compileFrom(add);
graph.submitAndWait(vulkanContext.getGraphicsQueue(), 0);
Compiling¶
compileFrom(element) starts at the given output element and walks its
inputs. It then:
Orders the elements with Kahn’s topological-sort algorithm, so every element runs after all of its inputs.
Sets up each element (
_setup()): pipelines, descriptor sets and buffers are created.Records one command buffer per element and path (
_record()).Connects the command buffers with semaphores, so each element waits for the elements it depends on.
Paths¶
A graph is compiled for a fixed number of paths. Every path has its own command buffers and, where needed, its own output buffers. In a windowed application there is one path per swapchain image, so several frames can be processed at the same time without sharing buffers that are still in use.
Elements whose data is identical for all paths (a loaded model, for example) keep a single buffer; elements whose output changes per frame keep one per path.
Submitting¶
submitTo(queue, pathId)runs the host-side updates of the path, submits its command buffers and returns a semaphore that is signalled when the path has finished. The frame loop uses this before presenting.submitAndWait(queue, pathId)submits and blocks until the GPU is done, which is convenient for tests and one-off computations.
Element types¶
Element |
Purpose |
|---|---|
A GPU buffer of a fixed type |
|
A tensor with separate dimension and data buffers |
|
A uniform buffer updated from the CPU, e.g. the camera |
|
A compute shader with inputs, outputs and push constants |
|
A compute shader mapping one buffer to another |
|
Copies a buffer |
|
Classic rasterization into an image |
Profiling¶
Before compiling, enableProfiling() adds GPU timestamp queries around every
element, and enablePerformanceProfiling() adds hardware performance counters
where the driver supports VK_KHR_performance_query (with pipeline statistics
as a fallback). getProfilingResults() returns the mean time per element.