~/home / Hacker News
Hacker News · Show HN

Tiny-vLLM – high performance LLM inference engine in C++ and CUDA

Tiny-vLLM 是一个用 C++ 和 CUDA 实现的高性能 LLM 推理引擎,专为低延迟、高吞吐场景优化,适合 AI 工程师部署大模型服务。
▲205
为什么值得关注

以极简代码实现极致性能,满足本地化大模型部署对效率与资源的严苛要求。

AI推理引擎高性能
访问 Hacker News 页面 →
分享到 X
手机端点「分享」直达微信/朋友圈/小红书;桌面端用「复制文案」后到 App 内粘贴发布

同类产品

Scala Tutorials – interactive Scala 3 lessons in the browserHow far do I have to go to run into 100k people?I was tired of opening 2 tabs for every HN link, so I made a userscriptYap – OSS on-device voice dictation for macOS with no model to downloadOpen-source Cloudflare deployed agent native task management and wikiBrowserAct: Browser Layer for Your AI Agent