Skip to content

mistralrs bench

Run performance benchmarks for base or LoRA model generation

mistralrs bench [OPTIONS] [COMMAND]
OptionDefaultDescription
-m, --model-id <MODEL_ID>HuggingFace model ID or local path to model directory
-t, --tokenizer <TOKENIZER>Path to local tokenizer.json file
-a, --arch <ARCH>Model architecture (auto-detected if not specified)
--dtype <DTYPE>autoModel data type
--format <FORMAT>Model format: plain (safetensors), gguf, or ggml Auto-detected if not specified Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE>Quantized model filename(s) for GGUF/GGML (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID>Model ID for tokenizer when using quantized format
--gqa <GQA>1GQA value for GGML models
--enable-lorafalseEnable dynamic LoRA without preloading an adapter. Supports text models. Qwen3.5/3.6 MoE requires automatic model selection; vision-tower adapters are unsupported
--lora <ALIAS=SOURCE|JSON>Preload a language-model LoRA adapter as ALIAS=SOURCE. Remote adapters use revision main. May be repeated. Qwen3.5/3.6 MoE conditional-generation models require auto model selection; vision-tower adapters are unsupported
--lora-max-adapters <LORA_MAX_ADAPTERS>16Maximum loaded LoRA aliases and, independently, resident adapter generations
--lora-max-rank <LORA_MAX_RANK>256Maximum rank accepted for a LoRA adapter
--lora-max-bytes <BYTES>8589934592Maximum memory used by loaded adapters
--legacy-lora <SOURCE>Legacy LoRA adapter source for a raw GGUF or GGML model
--legacy-lora-order <LEGACY_LORA_ORDER>Ordering JSON file for a legacy raw GGUF or GGML LoRA adapter
--xlora <XLORA>X-LoRA adapter model ID
--xlora-order <XLORA_ORDER>X-LoRA ordering JSON file
--tgt-non-granular-index <TGT_NON_GRANULAR_INDEX>Target non-granular index for X-LoRA
--quant <QUANT>Quantization front-door: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) This prefers prebuilt UQFF from mistralrs-community/<model>-UQFF, so use --isq if you do not want to switch to a prebuilt UQFF
--isq <IN_SITU_QUANT>In-situ quantization: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) and quantizes the selected model in-place (in-situ)
--from-uqff <FROM_UQFF>UQFF file(s) to load from. Accepts numeric shorthands (2, 3, 4, 5, 6, 8) to auto-detect the appropriate UQFF file (e.g., --from-uqff 8 finds q8_0-0.uqff or afq8-0.uqff). Also accepts ISQ type names (e.g., q4k, afq8). Shards are auto-discovered: specifying the first shard (e.g., q4k-0.uqff) automatically finds q4k-1.uqff, etc. Use semicolons to separate different quantizations
--isq-organization <ISQ_ORGANIZATION>ISQ organization strategy: default or moqe
--imatrix <IMATRIX>imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE>Calibration file for imatrix generation
--cpufalseForce CPU-only execution
-n, --device-layers <DEVICE_LAYERS>Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY>Topology YAML file for device mapping
--hf-cache <HF_CACHE>Custom HuggingFace cache directory
--max-seq-len <MAX_SEQ_LEN>4096Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE>1Max batch size for automatic device mapping
--paged-attn <MODE>autoPagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN>Allocate KV cache for this context length. If not specified, defaults to using 90% of available VRAM
--pa-memory-mb <MEMORY_MB>GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION>GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE>Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE>autoKV cache quantization type
--max-edge <MAX_EDGE>Maximum edge length for image resizing (aspect ratio preserved)
--max-num-images <MAX_NUM_IMAGES>Maximum number of images per request
--max-image-length <MAX_IMAGE_LENGTH>Maximum image dimension for device mapping
--no-kv-cachefalseDisable KV cache entirely
--matformer-config-path <MATFORMER_CONFIG_PATH>Path to a MatFormer config (CSV/JSON describing available slices). See model card
--matformer-slice-name <MATFORMER_SLICE_NAME>MatFormer slice to load (must match a slice name in the config file)
--mtp-model <MTP_MODEL>MTP assistant model id or path
--mtp-n-predict <MTP_N_PREDICT>Number of MTP draft tokens to propose per target step
--adapter <ADAPTER>LoRA adapter alias to benchmark. Omit to benchmark the base model
--prompt-len <PROMPT_LEN>512Input lengths used to measure time to first token. Zero skips TTFT. Accepts comma-separated values for sweeps
--gen-len <GEN_LEN>128Output tokens per decode request. Values below 2 skip decode metrics
--depth <DEPTH>4Input context lengths used to measure decode TPOT. Accepts comma-separated values for sweeps
--iterations <ITERATIONS>3Number of benchmark iterations
--warmup <WARMUP>1Number of warmup runs per benchmark case (discarded)

Auto-detect model type (recommended)

mistralrs bench auto [OPTIONS] --model-id <MODEL_ID>
OptionDefaultDescription
-m, --model-id <MODEL_ID>requiredHuggingFace model ID or local path to model directory
-t, --tokenizer <TOKENIZER>Path to local tokenizer.json file
-a, --arch <ARCH>Model architecture (auto-detected if not specified)
--dtype <DTYPE>autoModel data type
--format <FORMAT>Model format: plain (safetensors), gguf, or ggml Auto-detected if not specified Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE>Quantized model filename(s) for GGUF/GGML (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID>Model ID for tokenizer when using quantized format
--gqa <GQA>1GQA value for GGML models
--enable-lorafalseEnable dynamic LoRA without preloading an adapter. Supports text models. Qwen3.5/3.6 MoE requires automatic model selection; vision-tower adapters are unsupported
--lora <ALIAS=SOURCE|JSON>Preload a language-model LoRA adapter as ALIAS=SOURCE. Remote adapters use revision main. May be repeated. Qwen3.5/3.6 MoE conditional-generation models require auto model selection; vision-tower adapters are unsupported
--lora-max-adapters <LORA_MAX_ADAPTERS>16Maximum loaded LoRA aliases and, independently, resident adapter generations
--lora-max-rank <LORA_MAX_RANK>256Maximum rank accepted for a LoRA adapter
--lora-max-bytes <BYTES>8589934592Maximum memory used by loaded adapters
--legacy-lora <SOURCE>Legacy LoRA adapter source for a raw GGUF or GGML model
--legacy-lora-order <LEGACY_LORA_ORDER>Ordering JSON file for a legacy raw GGUF or GGML LoRA adapter
--xlora <XLORA>X-LoRA adapter model ID
--xlora-order <XLORA_ORDER>X-LoRA ordering JSON file
--tgt-non-granular-index <TGT_NON_GRANULAR_INDEX>Target non-granular index for X-LoRA
--quant <QUANT>Quantization front-door: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) This prefers prebuilt UQFF from mistralrs-community/<model>-UQFF, so use --isq if you do not want to switch to a prebuilt UQFF
--isq <IN_SITU_QUANT>In-situ quantization: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) and quantizes the selected model in-place (in-situ)
--from-uqff <FROM_UQFF>UQFF file(s) to load from. Accepts numeric shorthands (2, 3, 4, 5, 6, 8) to auto-detect the appropriate UQFF file (e.g., --from-uqff 8 finds q8_0-0.uqff or afq8-0.uqff). Also accepts ISQ type names (e.g., q4k, afq8). Shards are auto-discovered: specifying the first shard (e.g., q4k-0.uqff) automatically finds q4k-1.uqff, etc. Use semicolons to separate different quantizations
--isq-organization <ISQ_ORGANIZATION>ISQ organization strategy: default or moqe
--imatrix <IMATRIX>imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE>Calibration file for imatrix generation
--cpufalseForce CPU-only execution
-n, --device-layers <DEVICE_LAYERS>Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY>Topology YAML file for device mapping
--hf-cache <HF_CACHE>Custom HuggingFace cache directory
--max-seq-len <MAX_SEQ_LEN>4096Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE>1Max batch size for automatic device mapping
--paged-attn <MODE>autoPagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN>Allocate KV cache for this context length. If not specified, defaults to using 90% of available VRAM
--pa-memory-mb <MEMORY_MB>GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION>GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE>Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE>autoKV cache quantization type
--max-edge <MAX_EDGE>Maximum edge length for image resizing (aspect ratio preserved)
--max-num-images <MAX_NUM_IMAGES>Maximum number of images per request
--max-image-length <MAX_IMAGE_LENGTH>Maximum image dimension for device mapping

Text generation model with explicit configuration

mistralrs bench text [OPTIONS] --model-id <MODEL_ID>
OptionDefaultDescription
-m, --model-id <MODEL_ID>requiredHuggingFace model ID or local path to model directory
-t, --tokenizer <TOKENIZER>Path to local tokenizer.json file
-a, --arch <ARCH>Model architecture (auto-detected if not specified)
--dtype <DTYPE>autoModel data type
--format <FORMAT>Model format: plain (safetensors), gguf, or ggml Auto-detected if not specified Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE>Quantized model filename(s) for GGUF/GGML (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID>Model ID for tokenizer when using quantized format
--gqa <GQA>1GQA value for GGML models
--enable-lorafalseEnable dynamic LoRA without preloading an adapter. Supports text models. Qwen3.5/3.6 MoE requires automatic model selection; vision-tower adapters are unsupported
--lora <ALIAS=SOURCE|JSON>Preload a language-model LoRA adapter as ALIAS=SOURCE. Remote adapters use revision main. May be repeated. Qwen3.5/3.6 MoE conditional-generation models require auto model selection; vision-tower adapters are unsupported
--lora-max-adapters <LORA_MAX_ADAPTERS>16Maximum loaded LoRA aliases and, independently, resident adapter generations
--lora-max-rank <LORA_MAX_RANK>256Maximum rank accepted for a LoRA adapter
--lora-max-bytes <BYTES>8589934592Maximum memory used by loaded adapters
--legacy-lora <SOURCE>Legacy LoRA adapter source for a raw GGUF or GGML model
--legacy-lora-order <LEGACY_LORA_ORDER>Ordering JSON file for a legacy raw GGUF or GGML LoRA adapter
--xlora <XLORA>X-LoRA adapter model ID
--xlora-order <XLORA_ORDER>X-LoRA ordering JSON file
--tgt-non-granular-index <TGT_NON_GRANULAR_INDEX>Target non-granular index for X-LoRA
--quant <QUANT>Quantization front-door: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) This prefers prebuilt UQFF from mistralrs-community/<model>-UQFF, so use --isq if you do not want to switch to a prebuilt UQFF
--isq <IN_SITU_QUANT>In-situ quantization: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) and quantizes the selected model in-place (in-situ)
--from-uqff <FROM_UQFF>UQFF file(s) to load from. Accepts numeric shorthands (2, 3, 4, 5, 6, 8) to auto-detect the appropriate UQFF file (e.g., --from-uqff 8 finds q8_0-0.uqff or afq8-0.uqff). Also accepts ISQ type names (e.g., q4k, afq8). Shards are auto-discovered: specifying the first shard (e.g., q4k-0.uqff) automatically finds q4k-1.uqff, etc. Use semicolons to separate different quantizations
--isq-organization <ISQ_ORGANIZATION>ISQ organization strategy: default or moqe
--imatrix <IMATRIX>imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE>Calibration file for imatrix generation
--cpufalseForce CPU-only execution
-n, --device-layers <DEVICE_LAYERS>Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY>Topology YAML file for device mapping
--hf-cache <HF_CACHE>Custom HuggingFace cache directory
--max-seq-len <MAX_SEQ_LEN>4096Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE>1Max batch size for automatic device mapping
--paged-attn <MODE>autoPagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN>Allocate KV cache for this context length. If not specified, defaults to using 90% of available VRAM
--pa-memory-mb <MEMORY_MB>GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION>GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE>Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE>autoKV cache quantization type

Multimodal model

mistralrs bench multimodal [OPTIONS] --model-id <MODEL_ID>
OptionDefaultDescription
-m, --model-id <MODEL_ID>requiredHuggingFace model ID or local path to model directory
-t, --tokenizer <TOKENIZER>Path to local tokenizer.json file
-a, --arch <ARCH>Model architecture (auto-detected if not specified)
--dtype <DTYPE>autoModel data type
--format <FORMAT>Model format: plain (safetensors), gguf, or ggml Auto-detected if not specified Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE>Quantized model filename(s) for GGUF/GGML (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID>Model ID for tokenizer when using quantized format
--gqa <GQA>1GQA value for GGML models
--quant <QUANT>Quantization front-door: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) This prefers prebuilt UQFF from mistralrs-community/<model>-UQFF, so use --isq if you do not want to switch to a prebuilt UQFF
--isq <IN_SITU_QUANT>In-situ quantization: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) and quantizes the selected model in-place (in-situ)
--from-uqff <FROM_UQFF>UQFF file(s) to load from. Accepts numeric shorthands (2, 3, 4, 5, 6, 8) to auto-detect the appropriate UQFF file (e.g., --from-uqff 8 finds q8_0-0.uqff or afq8-0.uqff). Also accepts ISQ type names (e.g., q4k, afq8). Shards are auto-discovered: specifying the first shard (e.g., q4k-0.uqff) automatically finds q4k-1.uqff, etc. Use semicolons to separate different quantizations
--isq-organization <ISQ_ORGANIZATION>ISQ organization strategy: default or moqe
--imatrix <IMATRIX>imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE>Calibration file for imatrix generation
--cpufalseForce CPU-only execution
-n, --device-layers <DEVICE_LAYERS>Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY>Topology YAML file for device mapping
--hf-cache <HF_CACHE>Custom HuggingFace cache directory
--max-seq-len <MAX_SEQ_LEN>4096Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE>1Max batch size for automatic device mapping
--paged-attn <MODE>autoPagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN>Allocate KV cache for this context length. If not specified, defaults to using 90% of available VRAM
--pa-memory-mb <MEMORY_MB>GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION>GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE>Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE>autoKV cache quantization type
--max-edge <MAX_EDGE>Maximum edge length for image resizing (aspect ratio preserved)
--max-num-images <MAX_NUM_IMAGES>Maximum number of images per request
--max-image-length <MAX_IMAGE_LENGTH>Maximum image dimension for device mapping

Image generation model (diffusion)

mistralrs bench diffusion [OPTIONS] --model-id <MODEL_ID>
OptionDefaultDescription
-m, --model-id <MODEL_ID>requiredHuggingFace model ID or local path to model directory
-t, --tokenizer <TOKENIZER>Path to local tokenizer.json file
-a, --arch <ARCH>Model architecture (auto-detected if not specified)
--dtype <DTYPE>autoModel data type
--cpufalseForce CPU-only execution
-n, --device-layers <DEVICE_LAYERS>Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY>Topology YAML file for device mapping
--hf-cache <HF_CACHE>Custom HuggingFace cache directory
--max-seq-len <MAX_SEQ_LEN>4096Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE>1Max batch size for automatic device mapping

Speech synthesis model

mistralrs bench speech [OPTIONS] --model-id <MODEL_ID>
OptionDefaultDescription
-m, --model-id <MODEL_ID>requiredHuggingFace model ID or local path to model directory
-t, --tokenizer <TOKENIZER>Path to local tokenizer.json file
-a, --arch <ARCH>Model architecture (auto-detected if not specified)
--dtype <DTYPE>autoModel data type
--cpufalseForce CPU-only execution
-n, --device-layers <DEVICE_LAYERS>Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY>Topology YAML file for device mapping
--hf-cache <HF_CACHE>Custom HuggingFace cache directory
--max-seq-len <MAX_SEQ_LEN>4096Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE>1Max batch size for automatic device mapping

Embedding model

mistralrs bench embedding [OPTIONS] --model-id <MODEL_ID>
OptionDefaultDescription
-m, --model-id <MODEL_ID>requiredHuggingFace model ID or local path to model directory
-t, --tokenizer <TOKENIZER>Path to local tokenizer.json file
-a, --arch <ARCH>Model architecture (auto-detected if not specified)
--dtype <DTYPE>autoModel data type
--format <FORMAT>Model format: plain (safetensors), gguf, or ggml Auto-detected if not specified Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE>Quantized model filename(s) for GGUF/GGML (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID>Model ID for tokenizer when using quantized format
--gqa <GQA>1GQA value for GGML models
--quant <QUANT>Quantization front-door: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) This prefers prebuilt UQFF from mistralrs-community/<model>-UQFF, so use --isq if you do not want to switch to a prebuilt UQFF
--isq <IN_SITU_QUANT>In-situ quantization: accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.) and quantizes the selected model in-place (in-situ)
--from-uqff <FROM_UQFF>UQFF file(s) to load from. Accepts numeric shorthands (2, 3, 4, 5, 6, 8) to auto-detect the appropriate UQFF file (e.g., --from-uqff 8 finds q8_0-0.uqff or afq8-0.uqff). Also accepts ISQ type names (e.g., q4k, afq8). Shards are auto-discovered: specifying the first shard (e.g., q4k-0.uqff) automatically finds q4k-1.uqff, etc. Use semicolons to separate different quantizations
--isq-organization <ISQ_ORGANIZATION>ISQ organization strategy: default or moqe
--imatrix <IMATRIX>imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE>Calibration file for imatrix generation
--cpufalseForce CPU-only execution
-n, --device-layers <DEVICE_LAYERS>Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY>Topology YAML file for device mapping
--hf-cache <HF_CACHE>Custom HuggingFace cache directory
--max-seq-len <MAX_SEQ_LEN>4096Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE>1Max batch size for automatic device mapping
--paged-attn <MODE>autoPagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN>Allocate KV cache for this context length. If not specified, defaults to using 90% of available VRAM
--pa-memory-mb <MEMORY_MB>GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION>GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE>Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE>autoKV cache quantization type