Wrappers

MATE uses the packages in the wrappers/ directory as a compatibility layer to run CUDA software on MUSA. They preserve familiar package names and high-level APIs while routing execution to MATE operators and kernels. This enables existing integrations to migrate to Moore Threads platforms with minimal code changes.

Wrappers are the default integration path when your framework already targets a supported CUDA-oriented Python package. For delivered packages, install the matching wrapper first. pip installs the matching mate dependency automatically from the same wheel source. Build from wrappers/ only when you are developing a wrapper locally.

How the wrappers work

Each wrapper keeps the upstream-facing Python package surface stable while routing supported execution paths to MATE-backed implementations on MUSA.

Key mechanisms

  • API mapping: Maps upstream-style calls to MATE operator paths.

  • Namespace preservation: Preserves expected package names and import paths.

  • Kernel routing: Runs calls on MATE-optimized operators and MUSA kernels.

  • Distribution identification: Uses the PEP 440 local version suffix +musa so installed MUSA wrappers can be distinguished from native implementations with python -m pip show <package>.

Why use wrappers

  • Lower migration overhead: Minimizes code changes and avoids separate hardware-specific code paths.

  • Faster integration: Accelerates deployment of common tools and libraries on Moore Threads GPUs.

Wrapper support at a glance

Select a wrapper package to open its documentation page.

MSA (MiniMax Sparse Attention) workloads map to the fmha_sm100 wrapper. Use it when your project already targets the fmha_sm100 package surface and you want the same import path on MUSA.

KDA (Kimi Delta Attention) workloads map to the flash_kda wrapper. Use it when your project already targets the flash_kda package surface and you want the same import path on MUSA.

Wrapper package

Import path

Best fit

Current scope

flash_attn_3

flash_attn_interface

FlashAttention-3 style APIs

Dense FMHA, varlen FMHA, KV-cache attention, scheduler metadata

sageattention

sageattention

SageAttention style APIs

Dense SageAttention-compatible path

flash_mla

flash_mla

FlashMLA style APIs

MLA metadata, decode, sparse prefill

fmha_sm100

fmha_sm100

MSA style APIs

MSA planning, sparse prefill/decode, sparse top-k selection

flash_kda

flash_kda

FlashKDA / KDA style APIs

KDA forward, workspace-size compatibility helper

deep-gemm

deep_gemm

DeepGEMM style APIs

Grouped GEMM, dense GEMM, prenorm GEMM, MQA logits