Prefix-Aware Attention for LLM Decoding
Python 47 3
An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
Python 39 4
C++ 45 58
MOSAIC: Unlocking Over 30× Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming