vLLM Benchmarks Five Speculative Decoding Methods on AMD Instinct GPUs, Reporting Gains Up to 2.87x
AMD and Embedded LLM tested five speculative-decoding drafting methods in vLLM on AMD Instinct GPUs, with some configurations reaching up to 2.87x throughput over the non-speculative baseline.