otf-llm 3.2.0

Source: Pypi.org· team.gtlabs@gmail.com, team.gtlabs@gmail.com· August 15, 2026

High-performance hybrid LLM inference engine with Fused OpenAI Triton INT4 GEMM kernels, 98.16% logit parity, Zero-RAM streaming quantizer, and 3-Tier MoE offloading.

This story was reported by Pypi.org. Read the full original article:
Read on Pypi.org

More in Products & Launches

View all