diff --git a/README.md b/README.md index 50107d7..f450e1d 100644 --- a/README.md +++ b/README.md @@ -126,6 +126,7 @@ curl -X POST "http://localhost:8001/v1/chat/completions" \ - 报错 `Model architectures ['Qwen3_5MoeForConditionalGeneration'] are not supported for now` 或 `The Transformers implementation ... is not compatible with vLLM` 时,说明当前 vLLM 栈与该模型架构不兼容,需切换到兼容模型或改用其他推理后端。 - 报错 `StrictDataclassClassValidationError` 且包含 `validate_rope` / `unsupported operand type(s) for -=: 'set' and 'list'` 时,移除 Dockerfile 中对 `transformers --upgrade --pre` 的强制升级,使用镜像内置依赖重建。 +- 报错 `moe_wna16 quantization is currently not supported in rocm` 时,将该模型的 `quantization` 改回 `gptq`。 - 报错 `model config (gptq) does not match quantization argument (gptq_marlin)` 时,将该模型配置改为 `dtype=float16` 且 `quantization=gptq`。 - 报错 `RPC call to sample_tokens timed out` 或出现 `GPU core dump` 时,先下调模型配置为更稳参数:`ctx=32768`、`max_num_seqs=4`、`max_tokens=2048`、`gpu_util=0.90`,并开启 `enforce_eager=true`。 - 若模型目录存在但仍加载失败,检查挂载路径是否为 `/opt/model:/opt/model:ro`,并确认容器内可见模型文件。 diff --git a/config.json b/config.json index f3c6e72..73505cf 100644 --- a/config.json +++ b/config.json @@ -72,7 +72,7 @@ "Qwen3.5-35B-A3B-GPTQ-Int4": { "local_path": "Qwen3.5-35B-A3B-GPTQ-Int4", "dtype": "float16", - "quantization": "moe_wna16", + "quantization": "gptq", "ctx": "32768", "trust_remote": true, "valid_tp": [1, 2],