⭐️ Engineering open models for production

Type
Topics

Abstract

Running open models in production means balancing latency, throughput, and cost. This talk explores inference optimization techniques, including caching, routing, and prefill and decoding, to help you reason about performance tradeoffs in LLM applications.

Who it’s for

AI engineers, platform engineers, and developers running or preparing to run open models in production. Familiarity with LLM inference is helpful.

What you’ll learn

  • Understand the production considerations for serving open models.
  • Explore how caching and routing can affect inference performance.
  • Understand prefill and decoding as distinct stages of inference.
  • Assess optimization choices in terms of latency, throughput, and cost.

Format / duration

  • Format: Talk + demo.
  • Duration: 30 mins to 45 mins, adaptable to event format

Resources

Video, slides, and code are not currently linked for this session.

Delivered at (4 events)