r/computervision
DO NOT TRUST “convenient” inference APIs provided by famous repos!
- upvotes
- 14
- comments
- 2
Post
Highlighted: the lines this signal was extracted from
I had a hunch: those nice one-line inference APIs like model(["input1", "input2", ...]) are screwing the preprocessing, leaving GPU wasting a surprising amount of time waiting. So I tested it! To check it, I built explicit pipelines around the very same models and optimized the preprocessing. I wanted to see how much latency could change before touching the model itself. On an NVIDIA L4, with the same models and eight-image batches, moving the input work into an explicit pipeline cut end-to-end latency by: Workload Native API Explicit route YOLO detection 65.95 ms 32.37 ms PaddleOCR text detection 374.25 ms 196.90 ms Hugging Face ViT classification 1786.22 ms 593.35 ms That is 51%, 47%, and 67% lower latency respectively. The model did not get faster. I just stopped letting opaque convenience code decide when the model should run. This is not “framework X is bad,” and it is not a universal result. It is one hardware setup and a few common inference paths. But it is a useful reminder that the friendly API boundary is also a performance decision.
Also quoted as evidence
On an NVIDIA L4, with the same models and eight-image batches, moving the input work into an explicit pipeline cut end-to-end latency by: [...] PaddleOCR text detection 374.25 ms 196.90 ms
From the post