Request and token performance
Queue time, inference latency, time to first token, inter-token latency, token throughput, and error rate describe different parts of the request lifecycle. A single latency number cannot explain the entire experience.
GPU and replica efficiency
GPU utilization, memory utilization, replica count, and endpoint health help teams understand how capacity is being used. Idle replicas and memory pressure require different operational responses.
Performance in economic context
Cost by model, endpoint, cluster, and environment connects infrastructure decisions to workload economics. Consider throughput, latency, availability, and cost together when iterating on runtime configuration.
Based on InferGrid’s company and product context.