1 project found for "llm-serving"
High-throughput, memory-efficient inference and serving engine for large language models, built for running LLMs in production at scale rather than on a single local machine.