FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

## Computer Science > Artificial Intelligence

## Computer Science > Artificial Intelligence ## Title:FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving > Abstract:Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one…

Читать полностью →

Источник: ArXiv cs.AI

Подключаюсь к источникам…

30 главных источников
о мире ИИ

Автоматический перевод, курирование и красивая подача главных статей об искусственном интеллекте.

0
статей
0
источников
9
разделов