Engineering teams often prioritize raw throughput when tuning LLM serving, but high request rates can mask performance degradation that hurts the end-user experience. Focusing on goodput metrics ensures that completed tasks meet critical latency targets, providing a superior measure of real-world system effectiveness.