Supercharging GenAI: Ray, Kubernetes, and TPUs for Lightning-Fast Inference
KubeCon
Nov 15, 2024
📍 Salt Lake City, UT
Abstract
Talk at KubeCon 2024 on supercharging GenAI inference.
Key points:
- Distributed inference with Ray for scaling LLM serving
- Kubernetes orchestration for ML workloads
- TPU acceleration for production AI inference
- Performance benchmarks and cost optimization
- Real-world deployment architecture
Takeaway: Combining Ray, Kubernetes, and TPUs provides a scalable, cost-effective infrastructure for production GenAI inference at enterprise scale.

Authors
Staff Agentic AI Researcher