When an ASR pipeline is pushed to production, the interesting question is not only how fast it runs, but how much throughput you can extract from each GPU before latency starts to break. In the setup described here, that tradeoff was the main lever for reducing inference cost by 75% using NVIDIA...
Source: [Dev.to](https://dev.to/noah_taro_1e3297725e3fcbf/cutting-asr-inference-cost-with-nvidia-mps-on-amazon-ec2-13an)