Deep Dive into Inference Optimization for LLMs with Philip Kiely


Episode Artwork
1.0x
0% played 00:00 00:00
Nov 05 2024 64 mins   5

Today we have Philip Kiely from Baseten on the show. Baseten is a Series B startup focused on providing infrastructure for AI workloads.


We go deep on Inference Optimization. We cover choosing a model, discuss the hype around Compound AI, choosing an Inference Engine, Optimization Techniques like Quantization and Speculative Decoding all the way down to your GPU choice.