Documentation History

This page records documentation changes by release.

For runtime, API, and compatibility changes, see GitHub Releases.

The current docs site history starts at 0.2.2.

0.2.7

  • Updated the installation and compatibility guidance for the MUSA SDK 5.2.0 build baseline.

  • Added FlashAttention forward support for chunked large head dimensions, including Gemma 512-512 attention.

  • Added FlashMLA metadata support for arbitrary batch sizes, including the additional multi-batch cases covered by the release tests.

  • Added KDA state-pool support.

  • Added FlashInfer Norm API reference.

  • Extended the DeepGEMM wrapper with cuBLASLt-style GEMM, M-grouped and K-grouped FP8 interfaces, scale-layout helpers, and disable_ue8m0_cast=True compatibility. The wrapper page documents its MUSA-specific behavior and limits.

  • Fixed hangs in affected MUTLASS S4FP8 decode scenarios.

  • Extended the SageAttention wrapper with dense FP8/FP16 compatibility names and a varlen entry point. The wrapper page documents which arguments are signature-compatible only and which MATE paths execute.

  • Added MSA notes for indexer behavior and short-query metadata handling.

  • Added the Glossary page for core MATE terms.

  • Documented the MATE_DRY_RUN developer and diagnostic option in the environment-variable reference.

0.2.6

0.2.5

  • Added Kimi Delta Attention (KDA) decode coverage to Python APIs and KDA.

  • Expanded Attention for FMHA only_qv and MLA head-ratio coverage.

  • Added MSA direct APIs and the MSA/fmha_sm100 wrapper for MiniMax Sparse Attention (MSA) workflows on MUSA.

  • Updated Installing MATE for the MUSA simple package index.

  • Added MUBIN artifact management commands to Command Line Interface.

  • Refined FlashMLA/DenseMLA support for multiple head-ratio configurations.

0.2.4

0.2.3

0.2.2