Hi, I’ve been experimenting with using the AsyncGPUReadbackRequest to fetch data from a compute shader. It’s great when a gpu takes too long to finish a task, but no matter how quick and simple the shader is I can’t seem to be able to fetch data and request more during the same frame; at best I can only get a result every other frame. Is this an expected side effect of async readback or is there some trick to it (that doesn’t involve multiple instances of the buffers and shader)?
This is deliberate. GPU pipelines are quite long - it’s why they’re so fast. But if you read back immediately without a delay it will stall the entire pipeline. It will stop what it is doing concurrently and perform your request. This would be a major performance loss.
It’s supposed to be delayed.
You can fetch the results without a delay but not using this API. The whole point of this API was to avoid the performance loss that comes with doing that.
Just to add onto the performance side of this and how important it is,
I have a second camera rendering the scene with custom material to a render texture for pixel perfect screen selections.
Switching to AyncGPUReadback took ~2-3ms off my frame time - this is huge.
I just cache the last frame camera ray and use that to avoid any frame delay issues.
It’s pretty wonderful.
To test what it runs like fetching the results immediately try Unity - Scripting API: Experimental.Rendering.AsyncGPUReadbackRequest.WaitForCompletion
I understand the need for some delay, I was just wondering if there was some timing trick to, for example dispatch on frame 1, then read from the first job on frame 2 and dispatch the second job, read from the second job on frame 3 and dispatch the third job etc.
If, on a fast system, the shader can finish a job well within a frame it seems odd to wait until the frame after the next frame to read the data back, but I don’t know much about what’s going on under the hood so I’ll work around it for now. I might also experiment with switching between a pair of buffers and seeing if that helps
It’s not really that fine grained, and will vary wildly between GPUs anyway. It’s ideal for temporal things like cloud rendering systems or getting water / terrain heights - where a few frames really doesn’t make much of a difference. You can also do things like store it and interpolate/extrapolate on the CPU side. That’s the nature of this particular beast.
You could even process a ton of AI or vegetation spawning for collider pooling on the cpu and so on, I can think of millions of reasons it’s still totally useful even like this.
Whats your use case?
I’m just using it to experiment with generating planets, the idea being that a compute shader would generate normal & splat maps and vertex heights, then the textures would be copied to the material and the vertex heights would be read back and fed into the mesh and collider of each terrain tile. But even with almost ‘blank’ data being generated it seems impossible to receive data from the frame before (or confirmation that data is ready to be copied, in the case of the textures) and dispatch a new job every frame, only every other frame.
Up to 30 tiles generated per second (with Vsync on) is plenty as it is, it just bugs me that even in the best case scenario an entire frame between each update seems to go to waste. Utilizing that extra time would allow the camera to move around the planet much faster