Skip to content

Pinned Loading

  1. PAT PAT Public

    Prefix-Aware Attention for LLM Decoding

    Python 47 3

  2. RAGPulse RAGPulse Public

    An Open-Source RAG Workload Trace to Optimize RAG Serving Systems

    Python 39 4

  3. flash-linear-attention-npu flash-linear-attention-npu Public

    C++ 45 58

Repositories

Showing 4 of 4 repositories
  • flashserve/flash-linear-attention-npu's past year of commit activity
    C++ 45 58 85 113 Updated Sep 21, 2026
  • RAGPulse Public

    An Open-Source RAG Workload Trace to Optimize RAG Serving Systems

    flashserve/RAGPulse's past year of commit activity
    Python 39 MIT 4 0 0 Updated Sep 1, 2026
  • PAT Public

    Prefix-Aware Attention for LLM Decoding

    flashserve/PAT's past year of commit activity
    Python 47 MIT 3 0 0 Updated May 26, 2026
  • Mosaic Public

    MOSAIC: Unlocking Over 30× Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming

    flashserve/Mosaic's past year of commit activity
    Python 6 Apache-2.0 0 0 0 Updated May 23, 2026