First point, this thread should help
model.Execute schedules the work to GPU/Burst. There might have some layers that needs to run on the CPU (shape calculations…) too.
Could very well be that the cumulative of those two cost 12ms…
To aleviate this, I would try two things:
- fix your input shape of your model as much as possible to make our shape inference more able to bake things down.
- schedule N layers at a time
Run a model | Sentis | 1.2.0-exp.2 - avoid blocking read from tensor (
MakeReadable/CompletePendingTransactions)
Read output from a model asynchronously | Sentis | 1.2.0-exp.2
If you think the 12ms is excessive, do file a bug report