fermion-research 0.1.9

Source: Pypi.org· July 28, 2026

Five-value sub-2-bit LLMs: chat with a ~2 GB 8B container at native-runtime speed, load it as a Transformers model, or serve it on an OpenAI-compatible endpoint

This story was reported by Pypi.org. Read the full original article:
Read on Pypi.org

More in AI Research

View all