AI 日报hiw3c.com

生产环境中的 LLM 推理路由:从引擎信号到策略

BestBlogs·AI 高分精选 www.bestblogs.dev 网页快照

Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

OpenAI inference engineers Lu Zhang and Qianru Lao recount how IRB evolved from a feedback-loop load balancer into a control-plane and data-plane design whose global optimizer assigns routing weights to minimize end-to-end latency under capacity, health, and cache constraints, with penalties, capped retries, and load shedding for production stability.