otf-llm 4.0.0

Source: Pypi.org· team.gtlabs@gmail.com, team.gtlabs@gmail.com· August 15, 2026

High-performance hybrid LLM inference engine with Adaptive Non-Uniform 2-Bit Quantization (Lloyd-Max + Fused Triton INT2 GEMM), 98.2% logit parity, Zero-RAM quantizer, and 3-Tier MoE offloading.

This story was reported by Pypi.org. Read the full original article:
Read on Pypi.org

More in Products & Launches

View all