Luce PFlash v1.0
Luce PFlash is a C++/CUDA speculative prefill system that achieves ~10.4x faster Time-To-First-Token on long-context inference. A small Qwen3-0.6B drafter scores token importance so the heavy 27B target only prefills the spans that matter — cutting 128K token prefill from 257s to 24.8s on a single RTX 3090.
~10.4x faster TTFT on 128K context vs standard llama.cpp
Small 0.6B drafter scores token importance to skip irrelevant spans
Single daemon compress command integrates into existing dflash stack
PricingFree / Open Source
Version1.0