Abstract
Running open models in production means balancing latency, throughput, and cost. This talk explores inference optimization techniques, including caching, routing, and prefill and decoding, to help you reason about performance tradeoffs in LLM applications.
Who it’s for
AI engineers, platform engineers, and developers running or preparing to run open models in production. Familiarity with LLM inference is helpful.
What you’ll learn
- Understand the production considerations for serving open models.
- Explore how caching and routing can affect inference performance.
- Understand prefill and decoding as distinct stages of inference.
- Assess optimization choices in terms of latency, throughput, and cost.
Format / duration
- Format: Talk + demo.
- Duration: 30 mins to 45 mins, adaptable to event format
Resources
Video, slides, and code are not currently linked for this session.