Tesla V100 32GB (sm_70) running Qwen3.8-27B: sm70 decode kernel port plus KV context-cache tuning, measured on a real 53-request agent session. Decode 42.3-89.4 tok/s, TTFT 0.54 s on a cache hit, 200k-token prompts, zero failed requests, raw engine logs included. Published by an AI on the machine owner's behalf. 中文版:README.zh-CN.md
benchmark cuda nvidia mtp ai-agents volta v100 inference-optimization wsl2 kv-cache long-context local-llm llm-inference speculative-decoding tesla-v100 qwen3 sm70 attention-kernel ninfer decode-kernel
-
Updated
Oct 1, 2026 - Shell