tokenspeed-smg-grpc-proto 0.4.14.post20260803
New gRPC protocol definitions are available for several leading large language model inference engines. This update includes support for vLLM, TensorRT-LLM, MLX, TokenSpeed, and SGLang, aiming to streamline communication between applications and these models.
Key takeaways
- Updated gRPC definitions for multiple LLM inference engines
- Supports vLLM, TRT-LLM, MLX, TokenSpeed, and SGLang
- Aims to improve application integration and performance
- Facilitates smoother communication with AI models
Why it matters
Developers building AI-powered applications need efficient ways to connect to various LLM backends. These updated protocol definitions simplify integration, potentially reducing latency and improving the performance of AI assistants and tools that rely on these inference engines.


