LLM
LLM Inference Infrastructure Provisioning Course, Part 4: Selecting an Inference Engine
Hello! In the previous installments of this course, we covered in detail how to estimate the request volume an LLM service needs to handle and how to calculate the memory a model consumes during inference. This time we dig into the fourth of the seven steps: selecting an inference engine.