Compare commits
110 Commits
7c03c9e07b
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
| 37390d386e | |||
| eed3d09138 | |||
| 299595e8a1 | |||
| 7f952295ad | |||
| 896c96257b | |||
| 0e0e85966e | |||
| eb81b0d420 | |||
| af483b663e | |||
| 979a5a41d0 | |||
| 309eee0e5b | |||
| a05af55271 | |||
| f3ea5b2dbf | |||
| 90d490c49e | |||
| 146d0a1ff3 | |||
| 248de16647 | |||
| 044c1c5b05 | |||
| a07d25a1b4 | |||
| 0196c3d877 | |||
| 7a8bf3b5f9 | |||
| 112dfee95c | |||
| 194b679633 | |||
| 697a0dde24 | |||
| 054962c9d8 | |||
| e932b07fa6 | |||
| 0e2b2f3994 | |||
| ff061ffa25 | |||
| c33488d95c | |||
| 2b779d0423 | |||
| 9c7d18f26d | |||
| 2d5413779c | |||
| a095bdd927 | |||
| cbe5388e57 | |||
| c90228ef7c | |||
| d569d5606b | |||
| 8852f1fa73 | |||
| 3f96a0df02 | |||
| b962f0b6c0 | |||
| 4f2b37b7b5 | |||
| 4bbeb5673e | |||
| efeca5f2af | |||
| 18898057e6 | |||
| 798ab1df7c | |||
| 967ba29ead | |||
| d0456ae878 | |||
| 827a1aa977 | |||
| 83159a6de9 | |||
| 883c89973c | |||
| b8b11eb353 | |||
| 5bf18170dd | |||
| c69349052a | |||
| 4fb4786f4e | |||
| f5a9b786d0 | |||
| a94374fbb1 | |||
| 177c508da0 | |||
| 5f189660c3 | |||
| 620027d259 | |||
| d455d255c3 | |||
| e296316e63 | |||
| 50c52d7d5f | |||
| 0ba03bc3f8 | |||
| afd8d967ed | |||
| 8bb485d416 | |||
| dda847ae6d | |||
| 3c8e795ddb | |||
| a6bdcc1b15 | |||
| 32060f4e5e | |||
| a96aa63261 | |||
| 1dcd83fd02 | |||
| 70c08aa1d5 | |||
| be14a07bb7 | |||
| 8d3e02a10d | |||
| f43e8b8c5c | |||
| 22d2acea82 | |||
| f64673adfb | |||
| 4d7833162a | |||
| 0e9e9d6221 | |||
| 9c38cdad28 | |||
| cccfa8830a | |||
| 2e9f75983c | |||
| 56ddaad4c9 | |||
| d4ce1d38f0 | |||
| 0ab132a220 | |||
| adb0757caa | |||
| 607706b46e | |||
| ebe6ee3305 | |||
| 3336660647 | |||
| 3c085567b8 | |||
| e0c40c7c08 | |||
| 99fbaf5cdc | |||
| 621e96b67e | |||
| a8c3293a8b | |||
| ab9b4ceb8c | |||
| 089212d560 | |||
| 9529b1b7c2 | |||
| 94a804cc36 | |||
| 95e7ac8f0f | |||
| 5fc53646b2 | |||
| 9d0b4b4749 | |||
| 270faac7b7 | |||
| 805acbf0a1 | |||
| 09175d61e7 | |||
| d784664b91 | |||
| b8999eef46 | |||
| aacd84da86 | |||
| 030693cd68 | |||
| 217985d85a | |||
| 9944d8cc34 | |||
| 4553b6c073 | |||
| 69b583340b | |||
| a9489cd9d9 |
@@ -0,0 +1,66 @@
|
||||
# 修复模型下载失败
|
||||
|
||||
## 当前状态
|
||||
|
||||
`mineru-models-download --source modelscope` 在执行时,12 个小型配置文件下载成功,但最大的文件 `model.safetensors`(2.45 GB)下载失败。
|
||||
|
||||
## 错误分析
|
||||
|
||||
```
|
||||
[Errno 2] No such file or directory:
|
||||
'/opt/models/modelscope/models/._____temp/OpenDataLab/MinerU2.5-Pro-2605-1.2B/model.safetensors'
|
||||
```
|
||||
|
||||
modelscope 的 `snapshot_download()` 将大文件先写入临时目录 `._____temp`,再移动到最终位置。临时目录创建不完整或下载中断导致大文件失败。
|
||||
|
||||
## 修复方案
|
||||
|
||||
### 步骤 1:清理损坏的下载状态
|
||||
|
||||
在宿主机上删除临时目录和损坏文件,保留已下载成功的 12 个小文件:
|
||||
|
||||
```bash
|
||||
rm -rf /opt/project/mineru-rocm/docker/data/models/modelscope/models/._____temp
|
||||
rm -f /opt/project/mineru-rocm/docker/data/models/modelscope/models/OpenDataLab/MinerU2.5-Pro-2605-1.2B/model.safetensors
|
||||
```
|
||||
|
||||
### 步骤 2:确保宿主机目录完整
|
||||
|
||||
```bash
|
||||
mkdir -p /opt/project/mineru-rocm/docker/data/{input,output,models/modelscope,miopen}
|
||||
```
|
||||
|
||||
### 步骤 3:重启所有容器使挂载生效
|
||||
|
||||
```bash
|
||||
cd /opt/project/mineru-rocm/docker
|
||||
docker compose down
|
||||
docker compose up -d
|
||||
docker compose --profile gradio up -d
|
||||
```
|
||||
|
||||
### 步骤 4:在 worker0 中重试下载
|
||||
|
||||
modelscope 的 `snapshot_download` 会检查已有文件并跳过,只重新下载缺失的 `model.safetensors`:
|
||||
|
||||
```bash
|
||||
docker compose exec worker0 mineru-models-download --source modelscope
|
||||
# 交互提示选 all
|
||||
```
|
||||
|
||||
### 步骤 5(备选):如果重试仍然失败
|
||||
|
||||
在 worker0 容器内手动创建临时目录后重试:
|
||||
|
||||
```bash
|
||||
docker compose exec worker0 bash -c "
|
||||
mkdir -p /opt/models/modelscope/models/._____temp/OpenDataLab/MinerU2.5-Pro-2605-1.2B
|
||||
mineru-models-download --source modelscope
|
||||
"
|
||||
```
|
||||
|
||||
## 备注
|
||||
|
||||
- 本次仅在 worker0 下载,但因为 `/opt/models` 是共享卷挂载(`./data/models:/opt/models`),所有容器都可见
|
||||
- 模型持久化在宿主机 `./data/models/`,容器重建不会丢失
|
||||
- 如果 modelscope 下载持续不稳定,可切换到 huggingface 源:`MINERU_MODEL_SOURCE=huggingface`(需要代理)
|
||||
@@ -0,0 +1,4 @@
|
||||
# 预克隆的仓库目录在构建时 COPY 进镜像,不要忽略
|
||||
!aiter
|
||||
!flash-attention
|
||||
!vllm
|
||||
+23
@@ -0,0 +1,23 @@
|
||||
# Docker 环境变量(按需修改)
|
||||
# 复制为 .env 后使用: cp env.example .env
|
||||
|
||||
# ---- GPU 架构(按你的显卡修改)----
|
||||
# gfx1201 = RX 9070 / 9070 XT / 9070 GRE
|
||||
# gfx1200 = RX 9060 XT / 9060 XT LP
|
||||
# gfx1100 = RX 7900 XTX / XT / GRE
|
||||
# gfx1101 = RX 7800 XT / 7700 XT
|
||||
# gfx1030 = RX 6950 / 6900 / 6800 系列
|
||||
ARCH=gfx1201
|
||||
|
||||
# ---- 目录挂载 ----
|
||||
INPUT_DIR=./data/input
|
||||
OUTPUT_DIR=./data/output
|
||||
MODEL_DIR=./data/models
|
||||
MIOPEN_CACHE=./data/miopen
|
||||
|
||||
# ---- 代理设置(国内访问 GitHub 慢时开启)----
|
||||
GIT_PROXY=http://127.0.0.1:8118
|
||||
|
||||
# ---- 模型下载源 ----
|
||||
# huggingface(默认,需科学上网)或 modelscope(国内可用)
|
||||
MINERU_MODEL_SOURCE=modelscope
|
||||
+209
-91
@@ -1,15 +1,11 @@
|
||||
# =============================================================================
|
||||
# MinerU on ROCm 7.2.1 Docker Image
|
||||
# MinerU on ROCm 7.2.1 Docker Image(HK 云服务器 / 海外版)
|
||||
# 原生 Linux + Ubuntu 24.04 + ROCm 7.2.1 + PyTorch 2.11.0 + vllm + MinerU 3.2.0
|
||||
#
|
||||
# 国内网络优化版:Ubuntu/PyPI/GitHub 全部使用国内镜像
|
||||
# 构建前请根据你的 GPU 修改 ARCH 参数(默认 gfx1201 = RX 9070)
|
||||
# =============================================================================
|
||||
|
||||
FROM ubuntu:24.04
|
||||
|
||||
# -- 构建参数 ---------------------------------------------------------------
|
||||
# GPU 架构:gfx1201(RX 9070) gfx1200(RX 9060) gfx1100(RX 7900) gfx1101(RX 7800/7700) gfx1030(RX 6900/6800)
|
||||
ARG ARCH=gfx1201
|
||||
ARG PYTHON_VER=3.12
|
||||
ARG VENV=/opt/mineru_venv
|
||||
@@ -18,9 +14,8 @@ ARG TORCH_INDEX=https://download.pytorch.org/whl/rocm7.2
|
||||
# -- GitHub 访问(国内无镜像,需代理)----------------------------------------
|
||||
ARG GIT_PROXY=http://127.0.0.1:8118
|
||||
|
||||
# -- 国内镜像配置 -----------------------------------------------------------
|
||||
# PyPI 镜像
|
||||
ARG PIP_INDEX=https://pypi.tuna.tsinghua.edu.cn/simple
|
||||
# -- 国内镜像 ----------------------------------------------------------------
|
||||
ARG PIP_INDEX=https://mirrors.aliyun.com/pypi/simple
|
||||
|
||||
# -- 环境变量 ---------------------------------------------------------------
|
||||
ENV DEBIAN_FRONTEND=noninteractive \
|
||||
@@ -30,51 +25,53 @@ ENV DEBIAN_FRONTEND=noninteractive \
|
||||
MINERU_MODEL_SOURCE=huggingface \
|
||||
TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1 \
|
||||
HSA_ENABLE_SDMA=1 \
|
||||
VLLM_TARGET_DEVICE=rocm
|
||||
VLLM_TARGET_DEVICE=rocm \
|
||||
CMAKE_PREFIX_PATH=/opt/rocm \
|
||||
CMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++ \
|
||||
HIP_COMPILER=/opt/rocm/llvm/bin/clang++ \
|
||||
HIP_PATH=/opt/rocm \
|
||||
HIP_ROOT_DIR=/opt/rocm \
|
||||
HIP_INCLUDE_DIR=/opt/rocm/include \
|
||||
HIP_PLATFORM=amd \
|
||||
ROCM_PATH=/opt/rocm
|
||||
|
||||
WORKDIR /opt
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 1:换国内源 + 安装 ROCm 7.2.1
|
||||
# ===========================================================================
|
||||
# Ubuntu 24.04 使用 deb822 格式,默认源文件是 /etc/apt/sources.list.d/ubuntu.sources
|
||||
RUN sed -i 's|http://.*archive.ubuntu.com|http://mirrors.tuna.tsinghua.edu.cn|g' /etc/apt/sources.list.d/ubuntu.sources && \
|
||||
sed -i 's|http://.*security.ubuntu.com|http://mirrors.tuna.tsinghua.edu.cn|g' /etc/apt/sources.list.d/ubuntu.sources && \
|
||||
export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
apt-get update && apt-get install -y --no-install-recommends \
|
||||
wget curl ca-certificates gnupg software-properties-common && \
|
||||
# 添加 AMD ROCm 仓库(repo.radeon.com 通常国内可直连)
|
||||
wget -q https://repo.radeon.com/rocm/rocm.gpg.key -O - | \
|
||||
gpg --dearmor | tee /etc/apt/trusted.gpg.d/rocm.gpg > /dev/null && \
|
||||
echo 'deb [arch=amd64] https://repo.radeon.com/rocm/apt/7.2.1 noble main' \
|
||||
> /etc/apt/sources.list.d/rocm.list && \
|
||||
# apt pinning:AMD 仓库优先级高于 Ubuntu 自带(避免拿到旧版 rocminfo)
|
||||
printf 'Package: *\nPin: release o=repo.radeon.com\nPin-Priority: 600\n' \
|
||||
> /etc/apt/preferences.d/rocm-pin-600 && \
|
||||
apt-get update && \
|
||||
# 安装 ROCm 基础组件(apt pinning 确保从 AMD 仓库拉)
|
||||
apt-get install -y --no-install-recommends \
|
||||
rocminfo rocm-device-libs hip-dev miopen-hip && \
|
||||
# 清理
|
||||
apt-get clean && rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 2:ROCm 头文件补丁(LLVM 22 兼容性修复)
|
||||
# 这些是 ROCm 7.2.1 在 24.04 上的已知问题,每次 apt 升级 ROCm 后需重新应用
|
||||
# 阶段 2:ROCm 头文件补丁
|
||||
# ===========================================================================
|
||||
RUN set -ex && \
|
||||
# 补丁 1: hipcc/clang 符号链接(hipcc.pl 硬编码 clang-17,实际是 clang-22)
|
||||
# 补丁 1: hipcc/clang 符号链接
|
||||
ln -sf /usr/bin/hipvars.pm /usr/share/perl5/hipvars.pm && \
|
||||
ln -sf /usr/bin/hipcc.pl /opt/rocm/bin/hipcc && \
|
||||
([ -f /usr/bin/hipcc.pl ] && ln -sf /usr/bin/hipcc.pl /opt/rocm/bin/hipcc) || echo "hipcc.pl not found, keeping apt-installed hipcc" && \
|
||||
ln -sf /opt/rocm/llvm/bin/clang-22 /opt/rocm/llvm/bin/clang-17 && \
|
||||
ln -sf /opt/rocm/llvm/bin/clang++ /opt/rocm/llvm/bin/clang++-17 && \
|
||||
# 补丁 2: __hip_internal::conditional → std::conditional
|
||||
find /opt/rocm/include/hip -name "*.h" \
|
||||
-exec sed -i 's/__hip_internal::conditional/std::conditional/g' {} + && \
|
||||
# 补丁 3: warpSize 常量(__AMDGCN_WAVEFRONT_SIZE 在 LLVM 22 未定义)
|
||||
# 补丁 3: warpSize 常量
|
||||
find /opt/rocm/include/hip -name "amd_warp_functions.h" \
|
||||
-exec sed -i 's/static constexpr int warpSize = __AMDGCN_WAVEFRONT_SIZE;/constexpr int warpSize = 32;/g' {} + && \
|
||||
# 补丁 4: __activemask() → __builtin_amdgcn_read_exec()
|
||||
# 注意:只改 amd_warp_sync_functions.h,不要动 amd_warp_functions.h(那是定义本身)
|
||||
sed -i 's/__activemask()/__builtin_amdgcn_read_exec()/g' \
|
||||
/opt/rocm/include/hip/amd_detail/amd_warp_sync_functions.h && \
|
||||
echo "ROCm 7.2.1 header patches applied."
|
||||
@@ -86,16 +83,14 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
build-essential git ninja-build pkg-config \
|
||||
python${PYTHON_VER} python${PYTHON_VER}-venv python${PYTHON_VER}-dev \
|
||||
libnuma-dev libdrm2 libhwloc-dev libgl1 \
|
||||
# vllm 运行时依赖
|
||||
libgomp1 libopenblas0 && \
|
||||
apt-get clean && rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 4:CMake 4.0(vllm 要求 ≥ 4.0,Ubuntu 24.04 自带 3.28 不够)
|
||||
# 阶段 4:CMake 4.0(通过代理下载加速)
|
||||
# ===========================================================================
|
||||
RUN cd /tmp && \
|
||||
# wget 通过代理下载 CMake(国内 GitHub 不通)
|
||||
export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
RUN export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
cd /tmp && \
|
||||
wget -q https://github.com/Kitware/CMake/releases/download/v4.0.0/cmake-4.0.0-linux-x86_64.tar.gz && \
|
||||
tar -xzf cmake-4.0.0-linux-x86_64.tar.gz && \
|
||||
cp -r cmake-4.0.0-linux-x86_64/bin/* /usr/local/bin/ && \
|
||||
@@ -107,71 +102,168 @@ RUN cd /tmp && \
|
||||
# 阶段 5:Python 虚拟环境 + PyTorch ROCm
|
||||
# ===========================================================================
|
||||
RUN python${PYTHON_VER} -m venv ${VENV} && \
|
||||
# 配置 pip 国内镜像
|
||||
mkdir -p /root/.pip && \
|
||||
touch ${VENV}/torch-constraint.txt && \
|
||||
echo "[global]" > /root/.pip/pip.conf && \
|
||||
echo "index-url = ${PIP_INDEX}" >> /root/.pip/pip.conf && \
|
||||
${VENV}/bin/pip install --no-cache-dir -U pip setuptools wheel && \
|
||||
# 安装 PyTorch ROCm 版(指定 index-url 覆盖全局镜像)
|
||||
${VENV}/bin/pip install --no-cache-dir --pre \
|
||||
echo "[install]" >> /root/.pip/pip.conf && \
|
||||
echo "constraint = ${VENV}/torch-constraint.txt" >> /root/.pip/pip.conf && \
|
||||
${VENV}/bin/pip install -U pip setuptools wheel && \
|
||||
# PyTorch ROCm wheels 在海外 CDN,通过代理下载
|
||||
export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
${VENV}/bin/pip install --pre \
|
||||
torch==2.11.0+rocm7.2 \
|
||||
torchvision \
|
||||
pytorch-triton-rocm \
|
||||
--index-url ${TORCH_INDEX} && \
|
||||
# 验证
|
||||
${VENV}/bin/python -c "import torch; print('PyTorch:', torch.__version__); print('ROCm:', torch.version.hip); assert torch.version.hip is not None"
|
||||
# 锁定 torch 版本,防止后续 pip install 解析到 CUDA 版
|
||||
echo "torch==2.11.0+rocm7.2" > ${VENV}/torch-constraint.txt && \
|
||||
${VENV}/bin/python -c "import torch; print('PyTorch:', torch.__version__); assert torch.version.hip is not None"
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 6:ROCm 开发包(vllm 编译必需)
|
||||
# 阶段 6:ROCm 开发包
|
||||
# ===========================================================================
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||
RUN export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
apt-get update && apt-get install -y --no-install-recommends \
|
||||
hipblas-dev hiprand-dev hipsparse-dev hipsparselt-dev \
|
||||
hipsolver-dev hipcub-dev rocprim-dev rocthrust-dev \
|
||||
rocblas-dev rocrand-dev hipfft-dev hipblaslt && \
|
||||
rocblas-dev rocrand-dev hipfft-dev hipblaslt-dev \
|
||||
rocsolver-dev rocfft-dev rocsparse-dev rocm-cmake rocm-core && \
|
||||
apt-get clean && rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 7:amd-aiter + flash_attn
|
||||
# 阶段 6.5:安装 amdsmi
|
||||
# ===========================================================================
|
||||
RUN if [ -d /opt/rocm/share/amd_smi ]; then \
|
||||
cp -r /opt/rocm/share/amd_smi /opt/amd_smi && \
|
||||
cd /opt/amd_smi && ${VENV}/bin/pip install . --no-build-isolation && \
|
||||
echo "amdsmi installed."; \
|
||||
else \
|
||||
echo "amdsmi source not found, skipping."; \
|
||||
fi
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 7:复制预克隆仓库 + 安装 aiter + flash_attn
|
||||
# ===========================================================================
|
||||
COPY aiter/ /opt/aiter/
|
||||
COPY flash-attention/ /opt/flash-attention/
|
||||
RUN set -ex && \
|
||||
# git 通过代理访问 GitHub(国内不通,其他源已用国内镜像)
|
||||
git config --global http.proxy ${GIT_PROXY} && \
|
||||
git config --global https.proxy ${GIT_PROXY} && \
|
||||
# aiter(AMD 优化的 attention 算子)
|
||||
cd /opt && git clone --recursive --depth 1 https://github.com/ROCm/aiter.git && \
|
||||
${VENV}/bin/pip install --no-cache-dir -e /opt/aiter && \
|
||||
# flash_attn(Triton AMD 后端,锁定已验证的 commit)
|
||||
cd /opt && git clone --recursive https://github.com/Dao-AILab/flash-attention.git && \
|
||||
cd flash-attention && git checkout bba578d43974c1d3ba157ab597124dd0fe2ccdb4 && \
|
||||
${VENV}/bin/pip install --no-cache-dir --no-build-isolation -e /opt/flash-attention && \
|
||||
# 验证 PyTorch 没被覆盖
|
||||
${VENV}/bin/pip install -e /opt/aiter && \
|
||||
${VENV}/bin/pip install --no-build-isolation -e /opt/flash-attention && \
|
||||
${VENV}/bin/python -c "import torch; v=torch.__version__; assert 'rocm' in v, f'PyTorch overwritten: {v}'; print('PyTorch OK:', v)"
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 8:编译 vllm
|
||||
# 阶段 8a:编译 vllm(cmake + ninja,最耗时,单独缓存)
|
||||
# ===========================================================================
|
||||
COPY vllm/ /opt/vllm/
|
||||
RUN set -ex && \
|
||||
# git 通过代理访问 GitHub
|
||||
git config --global http.proxy ${GIT_PROXY} && \
|
||||
git config --global https.proxy ${GIT_PROXY} && \
|
||||
# setuptools 升级(PEP 639 兼容)
|
||||
${VENV}/bin/pip install --no-cache-dir -U \
|
||||
"setuptools>=77.0.3" setuptools_scm setuptools_rust wheel && \
|
||||
# 克隆 vllm main
|
||||
cd /opt && git clone --depth 1 https://github.com/vllm-project/vllm.git && \
|
||||
# 补丁 5:注释掉 vllm mamba 模块的 operator+ 定义(ROCm 7.2 头文件已自带)
|
||||
${VENV}/bin/pip install -U \
|
||||
"setuptools>=77.0.3,<82" setuptools_scm setuptools_rust wheel && \
|
||||
cd /opt/vllm && \
|
||||
if [ -f csrc/mamba/mamba_ssm/selective_scan.h ]; then \
|
||||
sed -i '109,121s/^/\/\/ /' csrc/mamba/mamba_ssm/selective_scan.h && \
|
||||
echo "vllm mamba operator+ patch applied." && \
|
||||
# cmake 别名兜底:ROCm 7.2 可能缺少 hiprand/hipblas cmake target
|
||||
echo "vllm mamba operator+ patch applied."; \
|
||||
else \
|
||||
echo "vllm mamba selective_scan.h not found (upstream removed), skipping patch."; \
|
||||
fi && \
|
||||
mkdir -p /opt/rocm/lib/cmake/hiprand && \
|
||||
printf 'include(/opt/rocm/lib/cmake/rocrand/rocrand-config.cmake)\nif(TARGET roc::rocrand AND NOT TARGET hip::hiprand)\n add_library(hip::hiprand ALIAS roc::rocrand)\nendif()\n' \
|
||||
> /opt/rocm/lib/cmake/hiprand/hiprand-config.cmake && \
|
||||
mkdir -p /opt/rocm/lib/cmake/hipblas && \
|
||||
printf 'include(/opt/rocm/lib/cmake/rocblas/rocblas-config.cmake)\nif(TARGET roc::rocblas AND NOT TARGET hip::hipblas)\n add_library(hip::hipblas ALIAS roc::rocblas)\nendif()\n' \
|
||||
printf 'include(/opt/rocm/lib/cmake/rocblas/rocblas-config.cmake)\nif(TARGET roc::rocblas AND NOT TARGET hip::hipblas)\n add_library(hip::hipblas ALIAS roc::rocblas)\nendif()\nif(TARGET roc::rocblas AND NOT TARGET roc::hipblas)\n add_library(roc::hipblas ALIAS roc::rocblas)\nendif()\n' \
|
||||
> /opt/rocm/lib/cmake/hipblas/hipblas-config.cmake && \
|
||||
# cmake 配置
|
||||
# 批量创建 hip→roc 别名 cmake config(仅原生 config 缺失时创建)
|
||||
for pair in "hipcub:rocprim" "hipsolver:rocsolver" "hipsparse:rocsparse" "hipfft:rocfft"; do \
|
||||
hip_pkg="${pair%%:*}"; roc_pkg="${pair#*:}"; \
|
||||
cfg_dir="/opt/rocm/lib/cmake/${hip_pkg}"; \
|
||||
if [ ! -f "${cfg_dir}/${hip_pkg}-config.cmake" ]; then \
|
||||
mkdir -p "${cfg_dir}" && \
|
||||
printf 'include(/opt/rocm/lib/cmake/%s/%s-config.cmake)\nif(TARGET roc::%s AND NOT TARGET hip::%s)\n add_library(hip::%s ALIAS roc::%s)\nendif()\n' \
|
||||
"${roc_pkg}" "${roc_pkg}" "${roc_pkg}" "${hip_pkg}" "${hip_pkg}" "${roc_pkg}" \
|
||||
> "${cfg_dir}/${hip_pkg}-config.cmake" && \
|
||||
echo "${hip_pkg}-config.cmake created (alias to ${roc_pkg})."; \
|
||||
else \
|
||||
echo "${hip_pkg}-config.cmake already exists, skipping."; \
|
||||
fi; \
|
||||
done && \
|
||||
# 无 roc 对应物的 hip 包,创建 minimal cmake config(强制覆盖)
|
||||
for pkg in hipblaslt hiprtc; do \
|
||||
cfg_dir="/opt/rocm/lib/cmake/${pkg}"; \
|
||||
mkdir -p "${cfg_dir}" && \
|
||||
printf 'set(%s_FOUND TRUE)\nset(%s_VERSION "7.2.1")\nif(NOT TARGET hip::%s)\n add_library(hip::%s INTERFACE IMPORTED)\n set_target_properties(hip::%s PROPERTIES INTERFACE_INCLUDE_DIRECTORIES "/opt/rocm/include")\nendif()\nif(NOT TARGET %s::%s)\n add_library(%s::%s ALIAS hip::%s)\nendif()\n' \
|
||||
"${pkg}" "${pkg}" "${pkg}" "${pkg}" "${pkg}" \
|
||||
"${pkg}" "${pkg}" "${pkg}" "${pkg}" "${pkg}" \
|
||||
> "${cfg_dir}/${pkg}-config.cmake" && \
|
||||
echo "${pkg}-config.cmake created (hip:: + ${pkg}:: aliases)."; \
|
||||
done && \
|
||||
# roc::* 别名(PyTorch Caffe2Targets 直接引用 roc:: 命名空间的 hip 包名)
|
||||
for pair in \
|
||||
"hipblas:rocblas" "hiprand:rocrand" "hipsolver:rocsolver" \
|
||||
"hipsparse:rocsparse" "hipfft:rocfft" "hipcub:rocprim"; \
|
||||
do \
|
||||
hip_pkg="${pair%%:*}"; roc_pkg="${pair#*:}"; \
|
||||
cfg_dir="/opt/rocm/lib/cmake/${hip_pkg}"; \
|
||||
if [ -f "${cfg_dir}/${hip_pkg}-config.cmake" ]; then \
|
||||
if ! grep -q "roc::${hip_pkg}" "${cfg_dir}/${hip_pkg}-config.cmake" 2>/dev/null; then \
|
||||
printf 'if(TARGET roc::%s AND NOT TARGET roc::%s)\n add_library(roc::%s ALIAS roc::%s)\nendif()\n' \
|
||||
"${roc_pkg}" "${hip_pkg}" "${hip_pkg}" "${roc_pkg}" \
|
||||
>> "${cfg_dir}/${hip_pkg}-config.cmake" && \
|
||||
echo "roc::${hip_pkg} → roc::${roc_pkg}."; \
|
||||
fi; \
|
||||
fi; \
|
||||
done && \
|
||||
# hipblaslt / hiprtc:底层是 INTERFACE IMPORTED(非 ALIAS),roc:: 直接指 hip::
|
||||
for pkg in hipblaslt hiprtc; do \
|
||||
cfg_dir="/opt/rocm/lib/cmake/${pkg}"; \
|
||||
if [ -f "${cfg_dir}/${pkg}-config.cmake" ]; then \
|
||||
if ! grep -q "roc::${pkg}" "${cfg_dir}/${pkg}-config.cmake" 2>/dev/null; then \
|
||||
printf 'if(TARGET hip::%s AND NOT TARGET roc::%s)\n add_library(roc::%s ALIAS hip::%s)\nendif()\n' \
|
||||
"${pkg}" "${pkg}" "${pkg}" "${pkg}" \
|
||||
>> "${cfg_dir}/${pkg}-config.cmake" && \
|
||||
echo "roc::${pkg} → hip::${pkg}."; \
|
||||
fi; \
|
||||
fi; \
|
||||
done && \
|
||||
# miopen minimal config(强制覆盖,miopen-hip 可能不带 cmake config)
|
||||
mkdir -p /opt/rocm/lib/cmake/miopen && \
|
||||
# 诊断: MIOpen 库实际位置
|
||||
echo "=== MIOpen diagnostic ===" && \
|
||||
find /opt -name "libMIOpen*" -o -name "libmiopen*" 2>/dev/null | head -10 && \
|
||||
# 创建符号链接,确保链接器能找到
|
||||
for so in $(find /opt -name "libMIOpen.so*" -o -name "libmiopen.so*" 2>/dev/null | head -1); do \
|
||||
ln -sf "$so" /opt/rocm/lib/libMIOpen.so && \
|
||||
echo "libMIOpen.so → $so"; \
|
||||
done && \
|
||||
ls -la /opt/rocm/lib/libMIOpen* 2>/dev/null || echo "WARNING: libMIOpen.so not found!" && \
|
||||
printf 'set(miopen_FOUND TRUE)\nset(miopen_VERSION "3.4.0")\nif(NOT TARGET miopen)\n add_library(miopen SHARED IMPORTED)\n set_target_properties(miopen PROPERTIES IMPORTED_LOCATION "/opt/rocm/lib/libMIOpen.so" INTERFACE_INCLUDE_DIRECTORIES "/opt/rocm/include")\nendif()\n' \
|
||||
> /opt/rocm/lib/cmake/miopen/miopen-config.cmake && \
|
||||
echo "miopen-config.cmake created (SHARED IMPORTED)." && \
|
||||
# 兜底:确保 libMIOpen.so 链接可见
|
||||
ls /opt/rocm/lib/libMIOpen* /opt/rocm-7.2.1/lib/libMIOpen* 2>/dev/null | head -3 && \
|
||||
if [ ! -f /opt/rocm/lib/libMIOpen.so ] && [ -f /opt/rocm-7.2.1/lib/libMIOpen.so ]; then \
|
||||
ln -sf /opt/rocm-7.2.1/lib/libMIOpen.so /opt/rocm/lib/libMIOpen.so && \
|
||||
echo "libMIOpen.so symlinked."; \
|
||||
fi && \
|
||||
# amd_comgr minimal config(强制覆盖)
|
||||
mkdir -p /opt/rocm/lib/cmake/amd_comgr && \
|
||||
printf 'set(amd_comgr_FOUND TRUE)\nset(amd_comgr_VERSION "2.8.0")\nif(NOT TARGET amd_comgr)\n add_library(amd_comgr INTERFACE IMPORTED)\n set_target_properties(amd_comgr PROPERTIES INTERFACE_INCLUDE_DIRECTORIES "/opt/rocm/include")\nendif()\n' \
|
||||
> /opt/rocm/lib/cmake/amd_comgr/amd_comgr-config.cmake && \
|
||||
echo "amd_comgr-config.cmake created (minimal)." && \
|
||||
# 创建 HIP cmake config(apt 安装的 hip-dev 缺少 find_package 需要的配置)
|
||||
mkdir -p /opt/rocm/lib/cmake/hip && \
|
||||
printf 'set(HIP_FOUND TRUE)\nset(HIP_VERSION "7.2.1")\nset(HIP_PLATFORM amd)\nset(HIP_COMPILER /opt/rocm/llvm/bin/clang++)\nset(HIP_RUNTIME rocm)\nset(HIP_INCLUDE_DIRS /opt/rocm/include)\nset(HIP_LIBRARIES /opt/rocm/lib/libamdhip64.so)\nif(NOT TARGET hip::host)\n add_library(hip::host INTERFACE IMPORTED)\n set_target_properties(hip::host PROPERTIES INTERFACE_INCLUDE_DIRECTORIES "${HIP_INCLUDE_DIRS}")\nendif()\nif(NOT TARGET hip::device)\n add_library(hip::device INTERFACE IMPORTED)\n set_target_properties(hip::device PROPERTIES INTERFACE_COMPILE_OPTIONS "--offload-arch=${ARCH}")\nendif()\nif(NOT TARGET hip::amdhip64)\n add_library(hip::amdhip64 SHARED IMPORTED)\n set_target_properties(hip::amdhip64 PROPERTIES IMPORTED_LOCATION "/opt/rocm/lib/libamdhip64.so" INTERFACE_INCLUDE_DIRECTORIES "/opt/rocm/include")\nendif()\n' \
|
||||
> /opt/rocm/lib/cmake/hip/hip-config.cmake && \
|
||||
# 补丁: ROCm 7.2.1 中 hip_version.h 已移除, 创建兼容头文件供 PyTorch LoadHIP.cmake 使用
|
||||
if [ ! -f /opt/rocm/include/hip/hip_version.h ]; then \
|
||||
mkdir -p /opt/rocm/include/hip && \
|
||||
printf '#pragma once\n#define HIP_VERSION_MAJOR 7\n#define HIP_VERSION_MINOR 2\n#define HIP_VERSION_PATCH 53211\n#define HIP_VERSION_GITDATE 0\n#define HIP_VERSION (HIP_VERSION_MAJOR * 10000000 + HIP_VERSION_MINOR * 100000 + HIP_VERSION_PATCH)\n' \
|
||||
> /opt/rocm/include/hip/hip_version.h && \
|
||||
echo "hip_version.h created."; \
|
||||
fi && \
|
||||
mkdir -p /opt/vllm_build && \
|
||||
export HIP_PLATFORM=amd HIP_PATH=/opt/rocm ROCM_PATH=/opt/rocm && \
|
||||
export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
export LIBRARY_PATH=/opt/rocm/lib:$LIBRARY_PATH && \
|
||||
cmake -S /opt/vllm -B /opt/vllm_build -G Ninja \
|
||||
-DCMAKE_BUILD_TYPE=RelWithDebInfo \
|
||||
-DVLLM_TARGET_DEVICE=rocm \
|
||||
@@ -179,50 +271,76 @@ RUN set -ex && \
|
||||
-DHIP_ROOT_DIR=/opt/rocm \
|
||||
-DROCM_PATH=/opt/rocm \
|
||||
-DCMAKE_HIP_ARCHITECTURES=${ARCH} \
|
||||
-DCMAKE_PREFIX_PATH="${VENV}/lib/python${PYTHON_VER}/site-packages/torch/share/cmake" && \
|
||||
# ninja 编译(-j8,32GB 内存;若 < 16GB 请改为 -j2)
|
||||
-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++ \
|
||||
-DHIP_COMPILER=/opt/rocm/llvm/bin/clang++ \
|
||||
-DHIP_PATH=/opt/rocm \
|
||||
-DHIP_INCLUDE_DIR=/opt/rocm/include \
|
||||
-DCMAKE_PREFIX_PATH="/opt/rocm;${VENV}/lib/python${PYTHON_VER}/site-packages/torch/share/cmake" \
|
||||
-DCMAKE_SHARED_LINKER_FLAGS="-L/opt/rocm/lib" \
|
||||
-DCMAKE_EXE_LINKER_FLAGS="-L/opt/rocm/lib" && \
|
||||
cd /opt/vllm_build && ninja -j8 && \
|
||||
# 安装 .so 到 vllm 源码目录
|
||||
cp /opt/vllm_build/*.abi3.so /opt/vllm/vllm/ && \
|
||||
# pip install vllm(让 pip 解析运行时依赖)
|
||||
cd /opt/vllm && ${VENV}/bin/pip install --no-cache-dir -e . --no-build-isolation && \
|
||||
# 验证 PyTorch 没被 vllm 依赖覆盖
|
||||
${VENV}/bin/python -c "import torch; v=torch.__version__; assert 'rocm' in v, f'PyTorch overwritten by vllm deps: {v}'; print('PyTorch OK:', v)" && \
|
||||
# 先重装 ROCm PyTorch 覆盖可能的 CUDA 版,再清理 CUDA triton 元数据
|
||||
# 顺序重要:pytorch-triton-rocm 和 triton 共享 triton/ 物理目录,必须先重装后卸载
|
||||
${VENV}/bin/pip install --no-cache-dir --force-reinstall \
|
||||
echo "vllm C++ build complete (layer cached)."
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 7.5:复制入口 + 辅助脚本
|
||||
# ===========================================================================
|
||||
COPY scripts/entrypoint.sh /opt/entrypoint.sh
|
||||
RUN chmod +x /opt/entrypoint.sh
|
||||
COPY scripts/cache_warmer.py /opt/scripts/cache_warmer.py
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 8b:安装 vllm + 平台补丁 + 验证
|
||||
# ===========================================================================
|
||||
RUN set -ex && \
|
||||
export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
# 1) 安装构建工具(后续 pip install 需要)
|
||||
${VENV}/bin/pip install -U "setuptools>=77.0.3" setuptools_scm setuptools_rust wheel && \
|
||||
# 2) 手动注册 vllm 为 editable package(阶段 8a 已编译 C++ 扩展,此处不触发 cmake)
|
||||
cd /opt/vllm && \
|
||||
mkdir -p vllm.egg-info && \
|
||||
printf 'Metadata-Version: 2.1\nName: vllm\nVersion: 0.11.0\nSummary: A high-throughput and memory-efficient inference and serving engine for LLMs\n' > vllm.egg-info/PKG-INFO && \
|
||||
echo "vllm" > vllm.egg-info/top_level.txt && \
|
||||
touch vllm.egg-info/dependency_links.txt && \
|
||||
touch vllm.egg-info/requires.txt && \
|
||||
find vllm -name "*.py" -type f 2>/dev/null | sort > vllm.egg-info/SOURCES.txt && \
|
||||
# 2b) 写 vllm/_version.py(替代 setuptools_scm 生成),避免 vllm.__version__='dev' 触发
|
||||
# mineru.backend.vlm.utils:88 中 packaging.version.parse 抛 InvalidVersion
|
||||
printf '__version__ = "0.11.0"\n__version_tuple__ = (0, 11, 0)\n' > /opt/vllm/vllm/_version.py && \
|
||||
echo "vllm egg-info created (cmake build skipped, C++ extensions from stage 8a)" && \
|
||||
# 3) 安装 ROCm PyTorch(后续依赖解析用 ROCm 源补齐 triton-rocm)
|
||||
(${VENV}/bin/pip uninstall -y triton triton-rocm 2>/dev/null || true) && \
|
||||
${VENV}/bin/pip install --force-reinstall \
|
||||
torch==2.11.0+rocm7.2 torchvision pytorch-triton-rocm \
|
||||
--index-url ${TORCH_INDEX} && \
|
||||
${VENV}/bin/pip uninstall -y triton triton-rocm 2>/dev/null; \
|
||||
# vllm 平台模块导入验证(GPU 检测只能在运行时,容器构建时无 GPU 设备)
|
||||
${VENV}/bin/python -c "from vllm.platforms import current_platform; print('Platform module:', type(current_platform).__name__); print('vllm import OK')" && \
|
||||
# 清理构建目录(减小镜像体积,约 3-5GB)
|
||||
${VENV}/bin/python -c "import torch; v=torch.__version__; assert 'rocm' in v, f'PyTorch overwritten: {v}'; print('PyTorch OK:', v)" && \
|
||||
# 4) 安装 vllm 运行时依赖;额外加入 ROCm PyTorch 源,避免传递依赖解析 triton-rocm 失败
|
||||
grep -vE '^(torch|torchvision|torchaudio|triton|triton-rocm|pytorch-triton|setuptools)' requirements/common.txt > /tmp/vllm_deps.txt && \
|
||||
${VENV}/bin/pip install -r /tmp/vllm_deps.txt --extra-index-url ${TORCH_INDEX} && \
|
||||
# 5) 再次确认 PyTorch 仍是 ROCm 版
|
||||
${VENV}/bin/python -c "import torch; v=torch.__version__; assert 'rocm' in v, f'PyTorch overwritten after vllm deps: {v}'; print('PyTorch still OK:', v)" && \
|
||||
# 6) 深度验证 vllm(确保 AsyncEngineArgs、AsyncLLM 等关键模块可导入,缺包直接构建失败)
|
||||
${VENV}/bin/python -c "from vllm.engine.arg_utils import AsyncEngineArgs; from vllm.v1.engine.async_llm import AsyncLLM; print('vllm deep import OK:', __import__('vllm').__version__)" && \
|
||||
# 7) 确保 mineru-api 子进程不依赖 PYTHONPATH 也能导入 vllm
|
||||
echo "/opt/vllm" > ${VENV}/lib/python${PYTHON_VER}/site-packages/vllm.pth && \
|
||||
# 8) 锁定 ROCm 关键包版本(不含 setuptools,避免与 vllm/mineru 依赖冲突)
|
||||
${VENV}/bin/pip freeze | grep -iE "^(torch|vllm|triton|pytorch.triton|torchvision|flash.attn|aiter)" >> ${VENV}/torch-constraint.txt && \
|
||||
rm -rf /opt/vllm_build
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 8.5:复制辅助脚本(必须在 MinerU 安装前就位)
|
||||
# ===========================================================================
|
||||
COPY scripts/apply_mineru_patches.py /opt/apply_mineru_patches.py
|
||||
COPY scripts/cache_warmer.py /opt/cache_warmer.py
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 9:安装 MinerU + RDNA 适配补丁
|
||||
# 阶段 9:安装 MinerU(constraint 保护 ROCm 环境,正常安装其他依赖)
|
||||
# ===========================================================================
|
||||
RUN set -ex && \
|
||||
# pip 国内镜像已在阶段 5 全局配置
|
||||
${VENV}/bin/pip install --no-cache-dir 'mineru[core]' && \
|
||||
# 验证 PyTorch 没被覆盖
|
||||
${VENV}/bin/python -c "import torch; v=torch.__version__; assert 'rocm' in v, f'PyTorch overwritten: {v}'; print('PyTorch OK:', v)" && \
|
||||
# 应用 MinerU RDNA 适配补丁
|
||||
${VENV}/bin/python /opt/apply_mineru_patches.py
|
||||
export http_proxy=${GIT_PROXY} https_proxy=${GIT_PROXY} && \
|
||||
${VENV}/bin/pip install 'mineru[core]==3.4.0' click gradio httpx cbor2 && \
|
||||
${VENV}/bin/python -c "import torch; from vllm.engine.arg_utils import AsyncEngineArgs; import mineru; import click; import gradio; import httpx; import cbor2; print('MinerU + Gradio + vllm deep + httpx + cbor2 imports OK')" && \
|
||||
${VENV}/bin/mineru-gradio --help >/dev/null
|
||||
|
||||
# ===========================================================================
|
||||
# 阶段 10:入口与最终验证
|
||||
# ===========================================================================
|
||||
RUN echo 'source /opt/mineru_venv/bin/activate' >> /etc/bash.bashrc && \
|
||||
echo "MinerU Docker image built successfully." && \
|
||||
${VENV}/bin/python -c "import torch, vllm, mineru; print('='*50); print('MinerU ROCm Docker Image Ready'); print(f' PyTorch : {torch.__version__}'); print(f' ROCm : {torch.version.hip}'); print(f' vllm : {vllm.__version__}'); print(f' MinerU : {mineru.__version__}'); print(f' Arch : ${ARCH}'); print('='*50)"
|
||||
${VENV}/bin/python -c "import torch; from vllm.engine.arg_utils import AsyncEngineArgs; import mineru; import gradio; import mineru.cli.gradio_app; print('MinerU ROCm Docker Image Ready'); print(f'PyTorch {torch.__version__} OK')"
|
||||
|
||||
# 容器入口:默认 bash,用户可 override
|
||||
ENTRYPOINT ["/bin/bash", "-c"]
|
||||
ENTRYPOINT ["/opt/entrypoint.sh"]
|
||||
CMD ["bash"]
|
||||
|
||||
@@ -0,0 +1,586 @@
|
||||
# Docker 构建问题汇总
|
||||
|
||||
> 目标:Ubuntu 24.04 + ROCm 7.2.1 + vllm + MinerU Docker 镜像
|
||||
> 环境:原生 Linux + RX 9070 (gfx1201) + Privoxy (127.0.0.1:8118)
|
||||
|
||||
---
|
||||
|
||||
## 1. 网络类
|
||||
|
||||
### 1.1 Docker Hub 拉基础镜像超时
|
||||
|
||||
**现象**:`dial tcp [2a03:2880:...]:443 i/o timeout`
|
||||
|
||||
**原因**:`docker.io` 被墙,IPv6 直连超时
|
||||
|
||||
**解决**:Docker daemon 配置代理 + 国内镜像加速器
|
||||
|
||||
```bash
|
||||
# /etc/systemd/system/docker.service.d/http-proxy.conf
|
||||
Environment="HTTP_PROXY=http://127.0.0.1:8118"
|
||||
Environment="HTTPS_PROXY=http://127.0.0.1:8118"
|
||||
|
||||
# /etc/docker/daemon.json
|
||||
{
|
||||
"registry-mirrors": ["https://docker.1panel.live", ...],
|
||||
"ipv6": false
|
||||
}
|
||||
```
|
||||
|
||||
### 1.2 容器内无法访问国内镜像
|
||||
|
||||
**现象**:`Could not resolve 'mirrors.tuna.tsinghua.edu.cn'`
|
||||
|
||||
**原因**:Docker 默认 bridge 网络无法路由到外网
|
||||
|
||||
**解决**:`docker-compose.yml` 构建时加 `network: host`(非代理,只是共享宿主机网络栈)
|
||||
|
||||
### 1.3 GitHub 直连失败
|
||||
|
||||
**现象**:`git clone github.com` → `GnuTLS recv error` / `TLS connection non-properly terminated`
|
||||
|
||||
**原因**:GitHub 在国内不通,无可靠国内镜像
|
||||
|
||||
**解决**:git/wget 通过 Privoxy 代理访问 GitHub(**仅 GitHub**,其他源用国内镜像)
|
||||
|
||||
```dockerfile
|
||||
ARG GIT_PROXY=http://127.0.0.1:8118
|
||||
git config --global http.proxy ${GIT_PROXY}
|
||||
# wget: export http_proxy=${GIT_PROXY}
|
||||
```
|
||||
|
||||
### 1.4 git clone 代理断连
|
||||
|
||||
**现象**:`RPC failed; curl 92 HTTP/2 stream 5 was not closed cleanly: CANCEL`
|
||||
|
||||
**原因**:HTTP/2 通过代理传大仓库不稳定
|
||||
|
||||
**解决**:禁用 HTTP/2 + 加大 buffer
|
||||
|
||||
```bash
|
||||
git config --global http.version HTTP/1.1
|
||||
git config --global http.postBuffer 524288000
|
||||
```
|
||||
|
||||
### 1.5 PyTorch ROCm wheels 下载中断
|
||||
|
||||
**现象**:`IncompleteRead` (pip 下载 .whl 到一半断开)
|
||||
|
||||
**原因**:PyTorch 官方 CDN (`download.pytorch.org`) 在国内直连不稳定
|
||||
|
||||
**解决**:下载 PyTorch 时设环境变量走代理
|
||||
|
||||
```bash
|
||||
export http_proxy=http://127.0.0.1:8118 https_proxy=http://127.0.0.1:8118
|
||||
pip install ... --pre torch==2.11.0+rocm7.2 ...
|
||||
```
|
||||
|
||||
### 1.6 flash-attention 子模块代理递归断连
|
||||
|
||||
**现象**:`--recursive` 克隆 flash-attention 时子模块 `cutlass` (~200MB) 下载失败:`RPC failed; curl 56 GnuTLS recv error`
|
||||
|
||||
**原因**:3 个子模块(composable_kernel + cutlass + aiter)挤在一条代理链路上,大文件容易丢包
|
||||
|
||||
**解决**(可选):宿主机预克隆后 COPY 入镜像。或用 `--depth 1` + 逐个 `git submodule update --init --depth 1`(断了一个可单独重试)
|
||||
|
||||
---
|
||||
|
||||
## 2. Dockerfile 语法类
|
||||
|
||||
### 2.1 ENV 行内注释导致解析失败
|
||||
|
||||
**现象**:`Syntax error - can't find = in "\ "`,行号为 `HSA_ENABLE_SDMA=1 \ # 注释`
|
||||
|
||||
**原因**:Docker ENV 不允许行内 `#` 注释
|
||||
|
||||
**解决**:删除行内注释,或移到单独 `#` 行
|
||||
|
||||
### 2.2 多行 `python -c "..."` 被误解析
|
||||
|
||||
**现象**:`dockerfile parse error: unknown instruction: import` / `unknown instruction: from`
|
||||
|
||||
**原因**:Docker 解析器把 Python `-c` 的多行字符串当作 Dockerfile 指令(`from x import y` → 误认为 `FROM`;`import x` → 误认为未知指令)
|
||||
|
||||
**解决**:复杂 Python 代码写成独立脚本,用 `COPY` 进入镜像后 `python /opt/xxx.py` 执行
|
||||
|
||||
涉及的文件:
|
||||
- `scripts/apply_mineru_patches.py` — MinerU RDNA 适配补丁
|
||||
- `scripts/patch_vllm_platform.py` — vllm 平台检测补丁
|
||||
|
||||
---
|
||||
|
||||
## 3. ROCm 兼容性类
|
||||
|
||||
### 3.1 版本号锁定失效
|
||||
|
||||
**现象**:`Version '1.0.0.70201-38~24.04' for 'rocminfo' was not found`
|
||||
|
||||
**原因**:AMD 仓库更新后精确版本号过期
|
||||
|
||||
**解决**:改用 apt pinning 策略,不锁死版本号
|
||||
|
||||
```bash
|
||||
# /etc/apt/preferences.d/rocm-pin-600
|
||||
Package: *
|
||||
Pin: release o=repo.radeon.com
|
||||
Pin-Priority: 600
|
||||
```
|
||||
|
||||
### 3.2 hipcc.pl 符号链接创建损坏链接
|
||||
|
||||
**现象**:cmake 报 `Can't find CUDA or HIP installation`(PyTorch 的 `LoadHIP.cmake:179`)
|
||||
|
||||
**原因**:`ln -sf /usr/bin/hipcc.pl /opt/rocm/bin/hipcc`,`/usr/bin/hipcc.pl` 在原生 Linux 上不存在,创建了损坏的符号链接
|
||||
|
||||
**解决**:条件创建,源文件不存在时跳过
|
||||
|
||||
```bash
|
||||
([ -f /usr/bin/hipcc.pl ] && ln -sf /usr/bin/hipcc.pl /opt/rocm/bin/hipcc) || echo "hipcc.pl not found"
|
||||
```
|
||||
|
||||
### 3.3 缺 rocm-cmake 导致 find_package(HIP) 失败
|
||||
|
||||
**现象**:PyTorch `LoadHIP.cmake:179` 调用 `find_package_and_print_version` 找不到 HIP
|
||||
|
||||
**原因**:未安装 `rocm-cmake` 包,缺少 HIP cmake 发现模块
|
||||
|
||||
**解决**:在阶段 6 加装 `rocm-cmake`
|
||||
|
||||
### 3.4 vllm mamba operator+ 补丁失效
|
||||
|
||||
**现象**:`sed: can't read csrc/mamba/mamba_ssm/selective_scan.h: No such file or directory`
|
||||
|
||||
**原因**:vllm 上游已移除 mamba 模块
|
||||
|
||||
**解决**:改为条件判断,文件不存在就跳过
|
||||
|
||||
```bash
|
||||
if [ -f csrc/mamba/mamba_ssm/selective_scan.h ]; then sed ...; else echo "skipping"; fi
|
||||
```
|
||||
|
||||
### 3.5 原生 Linux 需要 amdsmi + 平台检测补丁
|
||||
|
||||
**现象**(社区反馈):`ModuleNotFoundError: amdsmi` → `current_platform = UnspecifiedPlatform` → `Device string must not be empty`
|
||||
|
||||
**原因**:amdsmi 未安装,且 `logger.warning_once()` 触发 vllm 循环导入
|
||||
|
||||
**解决**:
|
||||
1. 安装 amdsmi(原生 Linux 上可用,WSL2 不需要)
|
||||
2. 补丁 6:`platforms/__init__.py` 加 `torch.version.hip` 兜底
|
||||
3. 补丁 7:`platforms/rocm.py` 中 `logger.warning_once()` → `sys.stderr.write()`
|
||||
|
||||
### 3.6 缺 hip_version.h 头文件
|
||||
|
||||
**现象**:`Could not find hip/hip_version.h or rocm-core/rocm_version.h`
|
||||
|
||||
**原因**:ROCm 7.2.1 的 apt 安装已将 `hip_version.h` 从默认路径移除,且 `rocm-core` 包未安装
|
||||
|
||||
**解决**:
|
||||
1. 安装 `rocm-core` 包(提供 `rocm_version.h`)
|
||||
2. 手动创建 `/opt/rocm/include/hip/hip_version.h` 兼容头文件
|
||||
|
||||
```bash
|
||||
apt-get install -y rocm-core
|
||||
# 兜底:手动生成 hip_version.h
|
||||
printf '#define HIP_VERSION_MAJOR 7\n#define HIP_VERSION_MINOR 2\n...'
|
||||
```
|
||||
|
||||
### 3.7 cmake 需显式设 HIP 编译器路径
|
||||
|
||||
**现象**:仅设 `PATH` + `HIP_ROOT_DIR` + `ROCM_PATH` 不足以让 cmake 发现 HIP
|
||||
|
||||
**解决**:显式指定
|
||||
|
||||
```
|
||||
-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++ # (非 hipcc,CMake 4.0 拒绝)
|
||||
-DHIP_COMPILER=/opt/rocm/llvm/bin/clang++ # PyTorch LoadHIP 也需要
|
||||
-DHIP_PATH=/opt/rocm
|
||||
-DCMAKE_PREFIX_PATH="/opt/rocm;..."
|
||||
```
|
||||
|
||||
### 3.8 PyTorch Caffe2Targets 缺失 cmake target(Docker 容器特有)
|
||||
|
||||
**现象**:cmake 反复报 `hip::amdhip64` / `hiprtc::hiprtc` / `roc::hipblas` / `roc::hipblaslt` 等 target 未找到
|
||||
|
||||
**原因**:裸机部署时 ROCm 的 cmake 配置通过 `/opt/rocm` 完整安装,cmake 能自动发现所有 target。但 Docker 容器内 apt 安装的 dev 包缺少部分 cmake config 文件,PyTorch 的 `Caffe2Targets.cmake` 硬编码引用了这些 target
|
||||
|
||||
**涉及 target 及修复**:
|
||||
|
||||
| 报错 target | 命名空间问题 | 最终 alias |
|
||||
|:--|:--|:--|
|
||||
| `hip::amdhip64` | 普通 target | SHARED IMPORTED → `libamdhip64.so` |
|
||||
| `hiprtc::hiprtc` | 需双冒号命名空间 | `hiprtc::hiprtc` ALIAS `hip::hiprtc` |
|
||||
| `roc::hipblas` | roc:: 而非 hip:: | `roc::hipblas` ALIAS `roc::rocblas` |
|
||||
| `roc::hipblaslt` | 同上 | `roc::hipblaslt` ALIAS `hip::hipblaslt` |
|
||||
| `roc::hiprand` / `roc::hipsolver` 等 | 同上 | 各指向对应 `roc::roc*` target |
|
||||
|
||||
**注意**:`roc::hip*` → 应直接指向底层 `roc::roc*`(不能通过 `hip::hip*` 中转,因为后者本身是 ALIAS,CMake 不允许链式 ALIAS)
|
||||
|
||||
### 3.9 rocm-core-dev → rocm-core 包名错误
|
||||
|
||||
**现象**:`E: Unable to locate package rocm-core-dev`
|
||||
|
||||
**原因**:ROCm 7.2.1 仓库中的包名为 `rocm-core`(无 -dev 后缀)
|
||||
|
||||
**解决**:`s/rocm-core-dev/rocm-core/`
|
||||
|
||||
### 3.10 -lMIOpen 链接失败
|
||||
|
||||
**现象**:`/usr/bin/ld: cannot find -lMIOpen`
|
||||
|
||||
**原因**:miopen cmake config 使用了 `INTERFACE IMPORTED`(无实际 .so 路径),链接器找不到 `libMIOpen.so`
|
||||
|
||||
**解决**:改为 `SHARED IMPORTED` + 指定 `IMPORTED_LOCATION`,并检查版本化目录的符号链接
|
||||
|
||||
```cmake
|
||||
# 错误
|
||||
add_library(miopen INTERFACE IMPORTED)
|
||||
# 正确
|
||||
add_library(miopen SHARED IMPORTED)
|
||||
set_target_properties(miopen PROPERTIES
|
||||
IMPORTED_LOCATION "/opt/rocm/lib/libMIOpen.so"
|
||||
INTERFACE_INCLUDE_DIRECTORIES "/opt/rocm/include")
|
||||
```
|
||||
|
||||
### 3.11 printf 参数计数错误
|
||||
|
||||
**现象**:cmake 报 `add_library` 在第 11 行语法错误(`hiprtc-config.cmake:11`)
|
||||
|
||||
**原因**:生成 cmake config 的 shell printf 格式串含 13 个 `%s`,但只传了 12 个参数,最后一行 `hip::%s` 未替换
|
||||
|
||||
**解决**:减少单行 printf 参数,复杂结构拆成独立循环
|
||||
|
||||
---
|
||||
|
||||
## 4. Docker 缓存策略
|
||||
|
||||
### 4.1 vllm 编译+验证同层导致缓存丢失
|
||||
|
||||
**现象**:ninja 编译 30 分钟成功后被后面验证失败拖垮,全部重来
|
||||
|
||||
**解决**:拆为 8a(编译)和 8b(安装+验证)
|
||||
|
||||
| 层 | 内容 | 缓存行为 |
|
||||
|:--|:--|:--|
|
||||
| 8a | git clone + cmake + ninja | 编译成功即缓存,**永不丢失** |
|
||||
| 8b | pip install + PyTorch 修复 + 平台补丁 | 失败可重试,不影响 8a |
|
||||
|
||||
### 4.2 缓存失效条件
|
||||
|
||||
任一 RUN 层的指令或上下文变化都会导致该层及之后所有层重建。修改 `ARG`、`ENV`、`COPY` 文件内容均会触发。
|
||||
|
||||
### 4.3 COPY 脚本在 RUN 引用之后
|
||||
|
||||
**现象**:阶段 8b 执行 `python /opt/patch_vllm_platform.py` → 文件不存在
|
||||
|
||||
**原因**:`COPY scripts/patch_vllm_platform.py` 在阶段 8.5(原序号),但阶段 8b 已引用该脚本——Dockerfile 中 COPY 位于引用它的 RUN 之后
|
||||
|
||||
**解决**:将 COPY 指令移到所有需要脚本的 RUN 之前(阶段 7.5)。规则:**COPY 必须在引用它的 RUN 之前**
|
||||
|
||||
---
|
||||
## 5. 网络策略总览
|
||||
|
||||
| 资源 | 访问方式 |
|
||||
|:--|:--|
|
||||
| Docker Hub | daemon 代理 + 国内镜像加速器 |
|
||||
| Ubuntu apt | 清华镜像 `mirrors.tuna.tsinghua.edu.cn` |
|
||||
| PyPI pip | 清华镜像 `pypi.tuna.tsinghua.edu.cn/simple` |
|
||||
| AMD ROCm apt | `repo.radeon.com` 直连 |
|
||||
| PyTorch ROCm wheels | `download.pytorch.org` 直连 |
|
||||
| GitHub (git/wget) | **Privoxy 代理 127.0.0.1:8118**(唯一例外) |
|
||||
|
||||
---
|
||||
|
||||
## 6. 构建阶段类
|
||||
|
||||
### 6.1 阶段 8b `pip install -e .` 触发 cmake 配置失败
|
||||
|
||||
**现象**:阶段 8a(cmake + ninja 手动编译)成功,但阶段 8b 执行 `pip install -e . --no-build-isolation --no-deps` 时报错:
|
||||
|
||||
```
|
||||
TorchConfig.cmake:62 (find_package)
|
||||
→ CMakeLists.txt:95 (find_package)
|
||||
→ Configuring incomplete, errors occurred!
|
||||
error: failed-wheel-build-for-install
|
||||
× Failed to build installable wheels for some pyproject.toml based projects
|
||||
╰─> vllm
|
||||
```
|
||||
|
||||
**排查过程**:
|
||||
|
||||
Dockerfile 设计为阶段 8a 手动执行 `cmake -S ... -B ... -G Ninja`(携带完整 `-D` 参数)编译 vllm C++ 扩展并复制 `.abi3.so` 文件,阶段 8b 再用 `pip install -e .` 注册 Python 包。但 `pip install -e .` 通过 vllm 的 setuptools 构建扩展**再次触发 cmake**,且 vllm 内部的 cmake 调用与手动 cmake 在以下方面不可控:
|
||||
|
||||
1. **尝试一(ENV 透传)**:在 Dockerfile `ENV` 块中添加 `CMAKE_PREFIX_PATH`、`CMAKE_HIP_COMPILER`、`ROCM_PATH` 等 8 个 cmake 环境变量,期望 pip 触发的 cmake 继承这些值。**无效**——vllm 的 `cmake_build_ext` 内部会独立构造 cmake 参数,不完全依赖环境变量。
|
||||
2. 错误始终出现在 `TorchConfig.cmake:62` 调用 `find_package(Caffe2)` 时,Caffe2 targets 引用的 ROCm cmake 目标无法被 pip 触发的 cmake 解析。
|
||||
|
||||
**结论**:`pip install -e .` 触发的 cmake 构建过程与阶段 8a 的手动 cmake 不在同一个可控环境中,且阶段 8b 的 `pip install -e .` 本质上是**冗余的**——C++ 扩展已在阶段 8a 编译完成。
|
||||
|
||||
**最终解决**:**跳过 pip install -e .,手动创建 vllm editable package 元数据**。
|
||||
|
||||
核心思路:阶段 8a 负责所有 C++ 编译(`.abi3.so`),阶段 8b 只负责 Python 包注册(不触发 cmake)。
|
||||
|
||||
具体改动(Dockerfile 阶段 8b):
|
||||
|
||||
```bash
|
||||
# 替换前(会触发 cmake,失败):
|
||||
cd /opt/vllm && \
|
||||
${VENV}/bin/pip install -e . --no-build-isolation --no-deps && \
|
||||
|
||||
# 替换后(手动创建 egg-info,不触发 cmake):
|
||||
cd /opt/vllm && \
|
||||
mkdir -p vllm.egg-info && \
|
||||
cat > vllm.egg-info/PKG-INFO << 'VLLM_EOF' && \
|
||||
Metadata-Version: 2.1
|
||||
Name: vllm
|
||||
Version: 0.0.0
|
||||
Summary: A high-throughput and memory-efficient inference and serving engine for LLMs
|
||||
VLLM_EOF
|
||||
echo "vllm" > vllm.egg-info/top_level.txt && \
|
||||
touch vllm.egg-info/dependency_links.txt && \
|
||||
touch vllm.egg-info/requires.txt && \
|
||||
find vllm -name "*.py" -type f 2>/dev/null | sort > vllm.egg-info/SOURCES.txt && \
|
||||
echo "vllm egg-info created (cmake build skipped, C++ extensions from stage 8a)" && \
|
||||
```
|
||||
|
||||
手动创建的 egg-info 包含以下文件:
|
||||
|
||||
| 文件 | 内容 | 用途 |
|
||||
|------|------|------|
|
||||
| `PKG-INFO` | `Name: vllm` + `Version: 0.0.0` | `pip freeze` / `importlib.metadata` 识别 |
|
||||
| `top_level.txt` | `vllm` | 声明顶层包名 |
|
||||
| `SOURCES.txt` | 所有 `.py` 文件列表 | 包文件索引 |
|
||||
| `dependency_links.txt` | 空文件 | 依赖链接(无) |
|
||||
| `requires.txt` | 空文件 | 运行依赖(由后续 pip install -r 安装) |
|
||||
|
||||
`import vllm` 的路径解析由现有的 `.pth` 文件保证:
|
||||
|
||||
```bash
|
||||
echo "/opt/vllm" > ${VENV}/lib/python${PYTHON_VER}/site-packages/vllm.pth
|
||||
```
|
||||
|
||||
**附:ENV 块新增的 cmake 变量(保留,对阶段 8a 和其他 cmake 调用有益)**:
|
||||
|
||||
```dockerfile
|
||||
ENV ... \
|
||||
CMAKE_PREFIX_PATH=/opt/rocm \
|
||||
CMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++ \
|
||||
HIP_COMPILER=/opt/rocm/llvm/bin/clang++ \
|
||||
HIP_PATH=/opt/rocm \
|
||||
HIP_ROOT_DIR=/opt/rocm \
|
||||
HIP_INCLUDE_DIR=/opt/rocm/include \
|
||||
HIP_PLATFORM=amd \
|
||||
ROCM_PATH=/opt/rocm
|
||||
```
|
||||
|
||||
**效果**:
|
||||
|
||||
- 阶段 8b 不再触发 cmake,消除了冗余编译(节省约 30 分钟)
|
||||
- C++ 扩展只在阶段 8a 编译一次,职责清晰
|
||||
- `pip list` / `pip freeze` 正常显示 vllm
|
||||
- `import vllm` 在所有 Python 进程中正常工作
|
||||
|
||||
---
|
||||
|
||||
## 7. 运行时类
|
||||
|
||||
### 7.1 vllm `_version.py` 缺失导致 `InvalidVersion: 'dev'`
|
||||
|
||||
**现象**:镜像构建成功,WebUI 启动正常,但解析任务全部失败:
|
||||
|
||||
```
|
||||
File ".../mineru/backend/vlm/utils.py", line 88, in set_default_gpu_memory_utilization
|
||||
if version.parse(vllm_version) >= version.parse("0.11.0") and gpu_memory <= 8:
|
||||
│ │ └ 'dev'
|
||||
packaging.version.InvalidVersion: Invalid version: 'dev'
|
||||
```
|
||||
|
||||
日志开头同时有信号:
|
||||
|
||||
```
|
||||
/opt/vllm/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:
|
||||
No module named 'vllm._version'
|
||||
vllm runtime import OK: dev, platform=UnspecifiedPlatform
|
||||
```
|
||||
|
||||
**原因**:阶段 8b 跳过了 `pip install -e .`(见 6.1 节),而 `setuptools_scm` 正是在 `pip install -e .` 时生成 `vllm/_version.py`。该文件缺失后,vllm 的 `version.py` 回退成 `__version__ = 'dev'`。`'dev'` 不是合法 PEP 440 版本号,mineru 调 `packaging.version.parse('dev')` 即抛 `InvalidVersion`,导致每个解析任务必失败。
|
||||
|
||||
**解决**:阶段 8b 手写 `vllm/_version.py`,替代 `setuptools_scm` 生成:
|
||||
|
||||
```dockerfile
|
||||
# 阶段 8b(egg-info 创建之后):
|
||||
printf '__version__ = "0.11.0"\n__version_tuple__ = (0, 11, 0)\n' > /opt/vllm/vllm/_version.py
|
||||
```
|
||||
|
||||
同时在 `patch_vllm_platform.py` 的 `ensure_vllm_dist_info()` 中对 `vllm.__version__` 做 PEP 440 合法性校验,非法时回退安全值 `0.11.0` 并回写 `vllm.__version__`(兜底,防止上游 version.py 再改回退逻辑)。
|
||||
|
||||
**版本号选 `0.11.0` 的理由**:
|
||||
- 合法 PEP 440,`packaging.version.parse` 不报错
|
||||
- 满足 `mineru[vllm]` 的约束 `>=0.10.1.1,<0.22.0`
|
||||
- mineru 代码 `version.parse(vllm_version) >= version.parse("0.11.0")` 走新分支(更合理的 GPU 显存策略)
|
||||
|
||||
**注意**:`mineru[core]`(Dockerfile 阶段 9 安装的)= vlm + pipeline + gradio,**不包含 `mineru[vllm]`**,因此 vllm 版本约束不会触发 pip 依赖冲突——editable 注册的假版本号只要功能上不被校验即可,而 `0.11.0` 恰好让校验逻辑走正确分支。
|
||||
|
||||
### 7.2 amdsmi 缺失导致 `Device string must not be empty`
|
||||
|
||||
**现象**:7.1 修复后(`vllm runtime import OK: 0.11.0`),解析任务在模型初始化阶段失败:
|
||||
|
||||
```
|
||||
File ".../vllm/config/device.py", line 78, in __post_init__
|
||||
self.device = torch.device(self.device_type)
|
||||
RuntimeError: Device string must not be empty
|
||||
```
|
||||
|
||||
启动日志里同时有:
|
||||
|
||||
```
|
||||
WARNING [rocm.py:39] Failed to import from amdsmi: No module named 'amdsmi'
|
||||
vllm runtime import OK: 0.11.0, platform=UnspecifiedPlatform
|
||||
WARNING: vLLM platform detection returned UnspecifiedPlatform.
|
||||
```
|
||||
|
||||
**原因**:vllm main 的 ROCm 平台检测依赖 `amdsmi`。检测链:
|
||||
|
||||
1. `resolve_current_platform_cls_qualname()` 遍历 `builtin_platform_plugins`,调用 `rocm_platform_plugin()`
|
||||
2. `rocm_platform_plugin()` 内部 `import amdsmi` → 失败(except)→ `is_rocm=False` → 返回 `None`
|
||||
3. 无任何 builtin plugin 激活 → `platform_cls_qualname = "vllm.platforms.interface.UnspecifiedPlatform"`
|
||||
4. `current_platform.device_type = ''` → `torch.device('')` 抛 `Device string must not be empty`
|
||||
|
||||
`amdsmi` 缺失的根因:阶段 6.5 的条件安装 `if [ -d /opt/rocm/share/amd_smi ]`,**ROCm 7.2 的 apt 包不再提供该目录**(`pip list` 无 amdsmi、`/opt/rocm/share/amd_smi` 不存在、dpkg 无 amdsmi 包),安装被跳过。
|
||||
|
||||
> 注:`rocm_platform_plugin` 定义在 `platforms/__init__.py`(第 111 行附近),**不是** `platforms/rocm.py`。早期补丁脚本尝试 `from vllm.platforms.rocm import rocm_platform_plugin` 注册 entry_point,路径错误导致 `ImportError`——这是 entry_point 注册失败的根因,但本方案不依赖 entry_point,改走 builtin plugin 修复。
|
||||
|
||||
**解决**:在 `patch_vllm_platform.py` 的补丁 6 中,向 `rocm_platform_plugin()` 的 `return` 语句前注入 `torch.version.hip` 兜底:
|
||||
|
||||
```python
|
||||
# __init__.py 中 rocm_platform_plugin 的 return 前:
|
||||
if not is_rocm:
|
||||
try:
|
||||
import torch as _torch
|
||||
if _torch.version.hip is not None:
|
||||
is_rocm = True
|
||||
except Exception:
|
||||
pass
|
||||
return "vllm.platforms.rocm.RocmPlatform" if is_rocm else None
|
||||
```
|
||||
|
||||
amdsmi 缺失时 `is_rocm=False`,补丁用 `torch.version.hip` 翻转为 `True` → 返回 `RocmPlatform` → `device_type='cuda'` → 错误消失。
|
||||
|
||||
**关键实现细节**(补丁 6 多次失效的教训):
|
||||
- **不能用精确字符串匹配**。vllm main 频繁重构,引号风格、空格、行结构都会变。早期补丁用 `old = " return 'vllm.platforms.rocm.RocmPlatform' if is_rocm else None"`(单引号)匹配,但实际文件是双引号,导致 `pattern not found`。
|
||||
- **改用语义定位**:遍历行,找同时含 `RocmPlatform` + `return` + `is_rocm` 的行作为注入点,取该行缩进对齐。兼容单/双引号。
|
||||
- 幂等:注入前检查文件内是否已有 `torch.version.hip is not None`,避免重复注入。
|
||||
|
||||
**验证**:
|
||||
|
||||
```bash
|
||||
docker exec mineru-gradio /opt/mineru_venv/bin/python -c "
|
||||
from vllm.platforms import current_platform
|
||||
print('platform:', type(current_platform).__name__)
|
||||
print('device_type:', repr(current_platform.device_type))
|
||||
"
|
||||
# 期望:platform: RocmPlatform / device_type: 'cuda'
|
||||
```
|
||||
|
||||
**放弃的方案**:`sitecustomize.py` 运行时强制注入 `current_platform`。失败原因:`current_platform` 是 lazy init(首次访问才 resolve),sitecustomize 在 Python 启动最早期执行时触发提前 resolve,而那时补丁 6 尚未应用(或被 `except: pass` 吞错),赋值后可能被后续逻辑覆盖。改 builtin plugin 本体(补丁 6)才是 vllm 官方检测路径,最干净。
|
||||
|
||||
### 7.3 热修复与镜像固化的操作流程
|
||||
|
||||
容器运行中(不重建镜像)热修复时注意:
|
||||
|
||||
1. **`./scripts:/opt/scripts:ro` 是只读卷挂载**(见 `docker-compose.yml`)。`docker cp` 写 `/opt/scripts` 会报 `mounted volume is marked read-only`。但容器内 `/opt/scripts` 直接映射宿主机 `./scripts`,**改宿主机文件即生效**,无需 `docker cp`。
|
||||
2. **`/opt/vllm` 在镜像层**(非只读挂载),可直接 `docker exec` 写入,重启不丢失。
|
||||
3. 补丁脚本幂等且检测 `already applied`,重启容器重跑 entrypoint 不会撤销已注入的改动。
|
||||
|
||||
完整固化流程(让修复进入镜像):
|
||||
- 修改 `Dockerfile` 阶段 8b(写 `_version.py` + egg-info 用 `0.11.0`)
|
||||
- 修改 `scripts/patch_vllm_platform.py`(补丁 6 语义定位 + PEP 440 校验)
|
||||
- 重建镜像(8a 的 ninja 编译层有缓存,几分钟):
|
||||
```bash
|
||||
docker compose build && docker compose --profile gradio up -d --force-recreate
|
||||
```
|
||||
|
||||
### 7.4 模型架构 inspect 失败(triton 版本过旧)
|
||||
|
||||
**现象**:7.1、7.2 修复后(`platform=RocmPlatform`),解析任务在 vllm 加载模型阶段失败:
|
||||
|
||||
```
|
||||
1 validation error for ModelConfig
|
||||
Value error, Model architectures ['Qwen2VLForConditionalGeneration']
|
||||
failed to be inspected. Please check the logs for more details.
|
||||
```
|
||||
|
||||
**真实根因**(gradio 的 ClickException 吞掉了 vllm 内部栈,需手动复现 `inspect_model_cls` 捕获完整 traceback):
|
||||
|
||||
```
|
||||
inspect_model_cls(['Qwen2VLForConditionalGeneration'])
|
||||
→ import vllm.model_executor.models.qwen2_vl
|
||||
→ from ...attention import MMEncoderAttention
|
||||
→ fa_utils.py: from flash_attn import flash_attn_varlen_func
|
||||
→ flash_attn_interface.py: from aiter.ops.triton... import flash_attn_2
|
||||
→ aiter/__init__.py: from .ops.attention import *
|
||||
→ aiter/ops/attention.py: from aiter.ops.triton.gluon.pa_decode_gluon import ...
|
||||
→ gluon/__init__.py:
|
||||
RuntimeError: aiter gluon kernels require triton>=3.6.0, found 3.5.1
|
||||
```
|
||||
|
||||
**vllm main 的 Qwen2VL 加载链一路 import 到 aiter,aiter 的 gluon kernels 要求 triton ≥ 3.6.0,但 `import triton` 实际得到 3.5.1。** 整个 import 链断裂 → Qwen2VL 类加载失败 → `failed to be inspected`。
|
||||
|
||||
> **排查弯路**:曾怀疑是 vllm registry 访问 `model_config.model_impl`(transformers v5 字段,v4 缺失)导致,写了补丁 9 给 10 处访问加 `getattr` 兜底。但补丁 9 应用后 7.4 依旧——说明 `model_impl` **不是**根因。补丁 9 无害(vllm ModelConfig 有该属性时透传,transformers config 缺时兜底 "auto"),保留,但真正解决 7.4 的是 triton 版本。
|
||||
|
||||
**triton 版本现状**(`pip list` 有 4 个相关包,容易误判):
|
||||
|
||||
| 包 | 版本 | 说明 |
|
||||
|---|---|---|
|
||||
| `pytorch-triton-rocm` | 3.5.1 | torch 依赖的 ROCm triton,**`import triton` 实际加载的是这个** |
|
||||
| `triton` | 3.7.0 | 原生 triton(满足 ≥3.6.0,但被 pytorch-triton-rocm 覆盖) |
|
||||
| `triton-rocm` | 3.6.0 | 另一个 ROCm triton |
|
||||
| `triton_kernels` | 1.0.0+amd | 缺 `matmul_ogs` 模块(非致命) |
|
||||
|
||||
**关键陷阱**:`import triton` 返回 3.5.1 而非 pip 显示的 3.7.0——因为 `pytorch-triton-rocm` 把自己注册成 `triton` 顶层包,覆盖了原生 `triton`。所以 pip list 看版本会误导,必须看 `import triton; triton.__version__`。
|
||||
|
||||
**解决**:设环境变量 `AITER_USE_SYSTEM_TRITON=1`。aiter 的 gluon `__init__.py` 检测逻辑:
|
||||
|
||||
```python
|
||||
if int(os.environ.get("AITER_USE_SYSTEM_TRITON", 0)):
|
||||
warnings.warn(...) # 仅警告,import 继续
|
||||
else:
|
||||
raise RuntimeError(...) # 阻断 import
|
||||
```
|
||||
|
||||
设 1 后 RuntimeError 降级为 warning,import 链不再阻断,Qwen2VL 类正常加载。**这是 aiter 官方提供的逃生口**。
|
||||
|
||||
固化进 `docker-compose.yml`(gradio / worker0 / worker1 都要加):
|
||||
|
||||
```yaml
|
||||
environment:
|
||||
- AITER_USE_SYSTEM_TRITON=1
|
||||
```
|
||||
|
||||
**验证**:
|
||||
|
||||
```bash
|
||||
docker exec -e AITER_USE_SYSTEM_TRITON=1 mineru-gradio /opt/mineru_venv/bin/python -c "
|
||||
from vllm.config.model import ModelConfig
|
||||
from vllm.model_executor.models.registry import ModelRegistry
|
||||
mp='/opt/models/modelscope/models/OpenDataLab/MinerU2.5-Pro-2605-1.2B'
|
||||
cfg = ModelConfig(model=mp, tokenizer=mp, trust_remote_code=True, dtype='auto', seed=0)
|
||||
print(ModelRegistry.inspect_model_cls(cfg.architectures, cfg))
|
||||
"
|
||||
# 期望输出:(..., 'Qwen2VLForConditionalGeneration')
|
||||
```
|
||||
|
||||
**遗留隐患**(非阻塞):
|
||||
- `AITER_USE_SYSTEM_TRITON=1` 只是不阻断 import,triton 3.5.1 实际可能不支持 gluon kernel 的某些 API。实测 Qwen2VL 推理跑通(142 页 PDF 解析成功),说明推理路径未真正调用 gluon kernel——只是 import 链被牵连。
|
||||
- 若未来某模型真用 gluon kernel 且 triton 3.5.1 API 不够,需真正升级 ROCm triton 到 ≥3.6.0(注意 `pytorch-triton-rocm` 与 torch 版本绑定,升级有风险,见 1.5 节)。
|
||||
|
||||
### 7.5 `_version.py` 与 force-recreate 的坑
|
||||
|
||||
**现象**:`docker compose up -d --force-recreate` 后,`Invalid version: 'dev'`(7.1)复发。
|
||||
|
||||
**原因**:热修复时用 `docker exec` 写入 `/opt/vllm/vllm/_version.py`,这只存在于**运行容器的可写层**,不在镜像层。`--force-recreate` 重建容器后丢失。补丁 6/7/9 是 entrypoint 跑脚本应用的,会自动重应用;但 `_version.py` 是 Dockerfile 阶段 8b 写的,**镜像里如果 Dockerfile 没改,重建即丢**。
|
||||
|
||||
**解决**:把 `_version.py` 写入固化进 Dockerfile 阶段 8b(见 6.1/7.1 节)。固化前用 `docker restart`(保留容器层),不要 `--force-recreate`。
|
||||
|
||||
---
|
||||
|
||||
*最后更新: 2026-06-23*
|
||||
+66
-29
@@ -10,7 +10,7 @@
|
||||
|
||||
| 方案 | 适用场景 |
|
||||
|------|---------|
|
||||
| **裸机部署**([MinerU本地部署教程.md](MinerU本地部署教程.md)) | 单机开发、追求极致性能 |
|
||||
| **裸机部署**([MinerU本地部署教程.md](../MinerU本地部署教程.md)) | 单机开发、追求极致性能 |
|
||||
| **Docker 部署**(本文) | 团队共享、CI/CD、环境隔离、快速迁移 |
|
||||
|
||||
Docker 方案的优势:
|
||||
@@ -191,17 +191,45 @@ docker compose build
|
||||
|
||||
## 4. 运行容器
|
||||
|
||||
### 4.1 交互模式(调试 / 手动处理)
|
||||
### 4.1 启动 WebUI 服务(默认推荐)
|
||||
|
||||
当前 `docker-compose.yml` 默认只启动一个 GPU 直连的 `gradio` WebUI 服务,避免单卡机器因为 `worker1` 或可选 Router 退出导致整组服务不可用:
|
||||
|
||||
| 服务 | 宿主机端口 | 容器内端口 | 用途 |
|
||||
|------|------------|------------|------|
|
||||
| `gradio` | `10002` | `7860` | WebUI 前端 + 本地推理 |
|
||||
|
||||
```bash
|
||||
# docker compose
|
||||
docker compose run --rm mineru
|
||||
docker compose up -d
|
||||
|
||||
# 查看容器是否都已启动
|
||||
docker compose ps -a
|
||||
|
||||
# WebUI 健康检查:应返回 HTML 或 HTTP 状态,而不是 Connection refused
|
||||
curl -v http://localhost:10002/
|
||||
```
|
||||
|
||||
浏览器打开:`http://<宿主机IP>:10002`。
|
||||
|
||||
如果 `docker compose ps` 为空或显示 `Exited`,说明服务进程启动后退出,请先看日志:
|
||||
|
||||
```bash
|
||||
docker compose logs --tail=200 gradio
|
||||
```
|
||||
|
||||
> 注意:直接 `docker run mineru-rocm:7.2.1` 使用的是 Dockerfile 默认 `CMD ["bash"]`,只会进入/运行 shell,不会启动 WebUI/API,也就不会监听 `10002`。如需直接 `docker run` 暴露 WebUI,必须显式传入 `mineru-gradio` 命令并映射端口。
|
||||
|
||||
### 4.2 交互模式(调试 / 手动处理)
|
||||
|
||||
```bash
|
||||
# docker compose(复用 gradio 的 GPU/卷配置,覆盖 command 进入 bash)
|
||||
docker compose run --rm gradio bash
|
||||
|
||||
# 或 docker run
|
||||
docker run -it --rm \
|
||||
--device /dev/kfd --device /dev/dri \
|
||||
--security-opt seccomp=unconfined \
|
||||
--group-add video --group-add render \
|
||||
--group-add video \
|
||||
--ipc host \
|
||||
-v ./data/input:/data/input:ro \
|
||||
-v ./data/output:/data/output \
|
||||
@@ -220,10 +248,10 @@ python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_
|
||||
mineru -p /data/input/example.pdf -o /data/output -b hybrid-auto-engine
|
||||
```
|
||||
|
||||
### 4.2 CLI 模式(一键处理)
|
||||
### 4.3 CLI 模式(一键处理)
|
||||
|
||||
```bash
|
||||
docker compose run --rm mineru \
|
||||
docker compose run --rm gradio \
|
||||
mineru -p /data/input/example.pdf -o /data/output -b hybrid-auto-engine
|
||||
```
|
||||
|
||||
@@ -232,27 +260,23 @@ docker compose run --rm mineru \
|
||||
command: mineru -p /data/input/example.pdf -o /data/output -b hybrid-auto-engine
|
||||
```
|
||||
|
||||
### 4.3 WebUI 模式
|
||||
### 4.4 可选:Router API 多 Worker 模式
|
||||
|
||||
修改 `docker-compose.yml`:
|
||||
如果确实需要独立 API Router,可启用 `router-api` profile。默认只启用 `worker0`,避免单卡机器启动 `worker1` 失败:
|
||||
|
||||
```yaml
|
||||
command: mineru-gradio --server-name 0.0.0.0 --server-port 7860
|
||||
ports:
|
||||
- "7860:7860"
|
||||
```bash
|
||||
docker compose --profile router-api up -d
|
||||
curl -v http://localhost:8000/
|
||||
```
|
||||
|
||||
双卡时再启用 `dual-gpu` profile,并在 `.env` 中设置:
|
||||
|
||||
```env
|
||||
ROUTER_API_URLS=http://mineru-worker0:8001,http://mineru-worker1:8002
|
||||
```
|
||||
|
||||
```bash
|
||||
docker compose up -d
|
||||
# 浏览器打开 http://<宿主机IP>:7860
|
||||
```
|
||||
|
||||
### 4.4 API 模式
|
||||
|
||||
```yaml
|
||||
command: mineru-api --host 0.0.0.0 --port 8000
|
||||
ports:
|
||||
- "8000:8000"
|
||||
docker compose --profile router-api --profile dual-gpu up -d
|
||||
```
|
||||
|
||||
API 用法参考 [MinerU 官方文档](https://github.com/opendatalab/MinerU)。
|
||||
@@ -262,7 +286,7 @@ API 用法参考 [MinerU 官方文档](https://github.com/opendatalab/MinerU)。
|
||||
设置环境变量即可切换下载源:
|
||||
|
||||
```bash
|
||||
docker compose run --rm -e MINERU_MODEL_SOURCE=modelscope mineru \
|
||||
docker compose run --rm -e MINERU_MODEL_SOURCE=modelscope gradio \
|
||||
mineru -p /data/input/example.pdf -o /data/output -b hybrid-auto-engine
|
||||
```
|
||||
|
||||
@@ -276,10 +300,10 @@ docker compose run --rm -e MINERU_MODEL_SOURCE=modelscope mineru \
|
||||
|
||||
```bash
|
||||
# 进入容器
|
||||
docker compose run --rm mineru
|
||||
docker compose run --rm gradio bash
|
||||
|
||||
# 运行预热
|
||||
python /opt/cache_warmer.py --device cuda --max_side 960 --step 32
|
||||
python /opt/scripts/cache_warmer.py --device cuda --max_side 960 --step 32
|
||||
```
|
||||
|
||||
| 输入尺寸 | 冷启动耗时 | 预热后 |
|
||||
@@ -294,7 +318,7 @@ python /opt/cache_warmer.py --device cuda --max_side 960 --step 32
|
||||
## 6. 验证
|
||||
|
||||
```bash
|
||||
docker compose run --rm mineru python -c "
|
||||
docker compose run --rm gradio python -c "
|
||||
import torch
|
||||
from vllm.platforms import current_platform
|
||||
|
||||
@@ -369,6 +393,19 @@ ls /dev/kfd /dev/dri/render*
|
||||
3. 确认当前用户在宿主机的 `render` 和 `video` 组
|
||||
4. 容器内运行 `rocminfo` 看能否检测到 GPU
|
||||
|
||||
**Q: 启动时报 `Unable to find group render: no matching entries in group file`**
|
||||
|
||||
这是因为 Docker Compose 的 `group_add` 使用的是容器内可解析的组名,而当前 Ubuntu 镜像内不一定存在 `render` 组。Compose 默认以 root 运行容器,并已透传 `/dev/kfd` 和 `/dev/dri`,因此默认只保留 `video` 组,不再添加 `render`。如果你改为非 root 用户运行容器,再按宿主机 `/dev/dri/render*` 的实际 GID 使用数字形式添加,例如 `group_add: ["109"]`。
|
||||
|
||||
**Q: `mineru-gradio` 启动时报 `ModuleNotFoundError`**
|
||||
|
||||
旧镜像只安装了 `mineru[core]`,可能缺少 WebUI CLI 依赖,例如 `click` 或 `gradio`。当前 Dockerfile 已显式安装 `click gradio`,并在构建期验证 `mineru.cli.gradio_app` 可导入、`mineru-gradio --help` 可执行。修改 Dockerfile 后需要重建镜像:
|
||||
|
||||
```bash
|
||||
docker compose build --no-cache gradio
|
||||
docker compose up -d --force-recreate
|
||||
```
|
||||
|
||||
**Q: 构建时 `ninja` 被 kill(exit 137)**
|
||||
|
||||
内存不足。将 Dockerfile 中 `ninja -j4` 改为 `ninja -j2` 或 `ninja -j1`,或给 Docker 分配更多内存。
|
||||
@@ -390,7 +427,7 @@ ls /dev/kfd /dev/dri/render*
|
||||
切换下载源:`MINERU_MODEL_SOURCE=modelscope`。或设置代理:
|
||||
|
||||
```bash
|
||||
docker compose run --rm -e http_proxy=http://host:port -e https_proxy=http://host:port mineru
|
||||
docker compose run --rm -e http_proxy=http://host:port -e https_proxy=http://host:port worker0 bash
|
||||
```
|
||||
|
||||
**Q: WebUI/API 端口无法访问**
|
||||
@@ -422,7 +459,7 @@ RUN rm -rf /opt/vllm_build /opt/vllm/.git /opt/aiter/.git /opt/flash-attention/.
|
||||
### 升级 MinerU
|
||||
|
||||
```bash
|
||||
docker compose run --rm mineru pip install --upgrade 'mineru[core]'
|
||||
docker compose run --rm gradio pip install --upgrade 'mineru[core]'
|
||||
# 然后重新应用 RDNA 补丁(参考 Dockerfile 阶段 9)
|
||||
```
|
||||
|
||||
|
||||
Binary file not shown.
+93
-39
@@ -1,66 +1,120 @@
|
||||
# =============================================================================
|
||||
# MinerU ROCm Docker Compose 配置
|
||||
# 原生 Linux + AMD GPU 环境
|
||||
#
|
||||
# 架构:Gradio / API → nginx (8000) → worker0 (8001) / worker1 (8002)
|
||||
# =============================================================================
|
||||
|
||||
services:
|
||||
mineru:
|
||||
# --- WebUI 前端 ---
|
||||
gradio:
|
||||
image: mineru-rocm:7.2.1
|
||||
profiles: ["gradio"]
|
||||
container_name: mineru-gradio
|
||||
stdin_open: true
|
||||
tty: true
|
||||
ipc: host
|
||||
ports:
|
||||
- "10002:7860"
|
||||
devices:
|
||||
- /dev/kfd
|
||||
- /dev/dri
|
||||
security_opt:
|
||||
- seccomp=unconfined
|
||||
group_add:
|
||||
- video
|
||||
environment:
|
||||
- GRADIO_SERVER_NAME=0.0.0.0
|
||||
- PYTHONPATH=/opt/vllm:${PYTHONPATH:-}
|
||||
- MINERU_MODEL_SOURCE=${MINERU_MODEL_SOURCE:-huggingface}
|
||||
- HF_HUB_CACHE=${HF_HUB_CACHE:-/opt/models/huggingface}
|
||||
- MODELSCOPE_CACHE=${MODELSCOPE_CACHE:-/opt/models/modelscope}
|
||||
- AITER_USE_SYSTEM_TRITON=1
|
||||
volumes:
|
||||
- ${INPUT_DIR:-./data/input}:/data/input:ro
|
||||
- ${OUTPUT_DIR:-./data/output}:/data/output
|
||||
- ${MODEL_DIR:-./data/models}:/opt/models
|
||||
- ${MIOPEN_CACHE:-./data/miopen}:/root/.cache/miopen
|
||||
- ./scripts:/opt/scripts:ro
|
||||
command:
|
||||
[
|
||||
"mineru-gradio",
|
||||
"--server-name", "0.0.0.0",
|
||||
"--server-port", "7860",
|
||||
]
|
||||
depends_on:
|
||||
- nginx
|
||||
|
||||
# --- 双卡 Worker(GPU 算力)---
|
||||
worker0:
|
||||
image: mineru-rocm:7.2.1
|
||||
build:
|
||||
context: .
|
||||
dockerfile: Dockerfile
|
||||
network: host # 容器共享宿主机网络(构建时需要)
|
||||
network: host
|
||||
args:
|
||||
# ---- 按你的 GPU 修改 ----
|
||||
# gfx1201 = RX 9070 / 9070 XT / 9070 GRE
|
||||
# gfx1200 = RX 9060 XT / 9060 XT LP
|
||||
# gfx1100 = RX 7900 XTX / XT / GRE
|
||||
# gfx1101 = RX 7800 XT / 7700 XT
|
||||
# gfx1030 = RX 6950 / 6900 / 6800 系列
|
||||
ARCH: gfx1201
|
||||
# ---- Git 代理(仅 GitHub 需要,国内无镜像)----
|
||||
GIT_PROXY: http://127.0.0.1:8118
|
||||
container_name: mineru-rocm
|
||||
ARCH: ${ARCH:-gfx1201}
|
||||
GIT_PROXY: ${GIT_PROXY:-}
|
||||
container_name: mineru-worker0
|
||||
stdin_open: true
|
||||
tty: true
|
||||
ipc: host # vllm 共享内存需要
|
||||
|
||||
# ---- GPU 设备透传(必需)----
|
||||
ipc: host
|
||||
devices:
|
||||
- /dev/kfd # ROCm KFD 内核驱动接口
|
||||
- /dev/dri # GPU 渲染节点 (renderD*)
|
||||
|
||||
# ---- 安全配置(ROCm 需要)----
|
||||
- /dev/kfd
|
||||
- /dev/dri
|
||||
security_opt:
|
||||
- seccomp=unconfined # 允许 ROCm 系统调用
|
||||
|
||||
# ---- 用户权限(访问 GPU 设备)----
|
||||
- seccomp=unconfined
|
||||
group_add:
|
||||
- video # /dev/dri/render* 权限
|
||||
- render # /dev/kfd 权限
|
||||
|
||||
# ---- 环境变量 ----
|
||||
- video
|
||||
environment:
|
||||
- HIP_VISIBLE_DEVICES=0
|
||||
- PYTHONPATH=/opt/vllm:${PYTHONPATH:-}
|
||||
- MINERU_MODEL_SOURCE=${MINERU_MODEL_SOURCE:-huggingface}
|
||||
- HF_HUB_CACHE=${HF_HUB_CACHE:-/opt/models/huggingface}
|
||||
- MODELSCOPE_CACHE=${MODELSCOPE_CACHE:-/opt/models/modelscope}
|
||||
|
||||
# ---- 卷挂载 ----
|
||||
volumes:
|
||||
# 输入/输出目录(按需修改)
|
||||
- ${INPUT_DIR:-./data/input}:/data/input:ro
|
||||
- ${OUTPUT_DIR:-./data/output}:/data/output
|
||||
# 模型缓存(持久化,避免每次下载)
|
||||
- ${MODEL_DIR:-./data/models}:/opt/models
|
||||
# MIOpen kernel 缓存(持久化,避免每次预热)
|
||||
- ${MIOPEN_CACHE:-./data/miopen}:/root/.cache/miopen
|
||||
- ./scripts:/opt/scripts:ro
|
||||
command: ["mineru-api", "--host", "0.0.0.0", "--port", "8001", "--allow-public-http-client"]
|
||||
|
||||
# ---- 启动命令(默认 bash,可改)----
|
||||
command: bash
|
||||
worker1:
|
||||
image: mineru-rocm:7.2.1
|
||||
container_name: mineru-worker1
|
||||
stdin_open: true
|
||||
tty: true
|
||||
ipc: host
|
||||
devices:
|
||||
- /dev/kfd
|
||||
- /dev/dri
|
||||
security_opt:
|
||||
- seccomp=unconfined
|
||||
group_add:
|
||||
- video
|
||||
environment:
|
||||
- HIP_VISIBLE_DEVICES=1
|
||||
- PYTHONPATH=/opt/vllm:${PYTHONPATH:-}
|
||||
- MINERU_MODEL_SOURCE=${MINERU_MODEL_SOURCE:-huggingface}
|
||||
- HF_HUB_CACHE=${HF_HUB_CACHE:-/opt/models/huggingface}
|
||||
- MODELSCOPE_CACHE=${MODELSCOPE_CACHE:-/opt/models/modelscope}
|
||||
volumes:
|
||||
- ${INPUT_DIR:-./data/input}:/data/input:ro
|
||||
- ${OUTPUT_DIR:-./data/output}:/data/output
|
||||
- ${MODEL_DIR:-./data/models}:/opt/models
|
||||
- ${MIOPEN_CACHE:-./data/miopen}:/root/.cache/miopen
|
||||
- ./scripts:/opt/scripts:ro
|
||||
command: ["mineru-api", "--host", "0.0.0.0", "--port", "8002", "--allow-public-http-client"]
|
||||
|
||||
# ---- 端口(WebUI / API 模式时启用)----
|
||||
# ports:
|
||||
# - "7860:7860" # WebUI
|
||||
# - "8000:8000" # API
|
||||
|
||||
restart: "no"
|
||||
# --- Nginx 负载均衡(替换有 Bug 的 mineru-router)---
|
||||
nginx:
|
||||
image: nginx:alpine
|
||||
container_name: mineru-nginx
|
||||
ports:
|
||||
- "8000:8000"
|
||||
volumes:
|
||||
- ./nginx.conf:/etc/nginx/conf.d/default.conf:ro
|
||||
depends_on:
|
||||
- worker0
|
||||
- worker1
|
||||
|
||||
@@ -9,12 +9,25 @@
|
||||
# gfx1030 = RX 6950 / 6900 / 6800 系列
|
||||
ARCH=gfx1201
|
||||
|
||||
# ---- 可选:构建阶段 GitHub 下载代理(不需要可留空)----
|
||||
# 示例: GIT_PROXY=http://host.docker.internal:8118
|
||||
GIT_PROXY=
|
||||
|
||||
# ---- 目录挂载 ----
|
||||
INPUT_DIR=./data/input
|
||||
OUTPUT_DIR=./data/output
|
||||
MODEL_DIR=./data/models
|
||||
MIOPEN_CACHE=./data/miopen
|
||||
|
||||
# ---- 默认 WebUI 使用的 GPU ----
|
||||
# 单卡保持 0;多卡时可指定 0/1/...,或留空让 ROCm 使用全部可见 GPU。
|
||||
HIP_VISIBLE_DEVICES=0
|
||||
|
||||
# ---- 可选:router-api profile 使用的 Worker 地址 ----
|
||||
# 单 worker: http://mineru-worker0:8001
|
||||
# 双 worker: http://mineru-worker0:8001,http://mineru-worker1:8002
|
||||
ROUTER_API_URLS=http://mineru-worker0:8001
|
||||
|
||||
# ---- 模型下载源 ----
|
||||
# huggingface(默认,需科学上网)或 modelscope(国内可用)
|
||||
MINERU_MODEL_SOURCE=huggingface
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
upstream mineru_workers {
|
||||
# 轮询分发
|
||||
server mineru-worker0:8001;
|
||||
server mineru-worker1:8002;
|
||||
}
|
||||
|
||||
server {
|
||||
listen 8000;
|
||||
server_name _;
|
||||
|
||||
# 文件上传可能很大,调大限制
|
||||
client_max_body_size 500m;
|
||||
|
||||
location / {
|
||||
proxy_pass http://mineru_workers;
|
||||
proxy_set_header Host $host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
proxy_read_timeout 600s; # 长任务需要
|
||||
proxy_send_timeout 600s;
|
||||
}
|
||||
}
|
||||
Binary file not shown.
@@ -15,9 +15,9 @@ def patch_a_predict_rec_imgw(infer_dir):
|
||||
"""predict_rec.py: imgW 对齐到 32"""
|
||||
f = os.path.join(infer_dir, 'predict_rec.py')
|
||||
c = open(f).read()
|
||||
old = r'(imgW = max\(min\(imgW, self\.limited_max_width\), self\.limited_min_width\)\n)'
|
||||
new = r'\1 imgW = math.ceil(imgW / 32) * 32\n'
|
||||
c2 = re.sub(old, new, c)
|
||||
old = r'(^ *)(imgW = max\(min\(imgW, self\.limited_max_width\), self\.limited_min_width\)\n)'
|
||||
new = r'\1\2\1imgW = math.ceil(imgW / 32) * 32\n'
|
||||
c2 = re.sub(old, new, c, flags=re.MULTILINE)
|
||||
if c2 == c:
|
||||
if 'math.ceil(imgW / 32)' not in c:
|
||||
raise RuntimeError('Patch A: cannot find imgW line')
|
||||
@@ -31,17 +31,28 @@ def patch_b_predict_rec_batch(infer_dir):
|
||||
"""predict_rec.py: 批次填充"""
|
||||
f = os.path.join(infer_dir, 'predict_rec.py')
|
||||
c = open(f).read()
|
||||
old = r'( norm_img_batch = np\.concatenate\(norm_img_batch\))'
|
||||
new = (
|
||||
' actual_batch_size = len(norm_img_batch)\n'
|
||||
' if actual_batch_size < batch_num:\n'
|
||||
' pad_size = batch_num - actual_batch_size\n'
|
||||
' pad_img = np.zeros_like(norm_img_batch[0])\n'
|
||||
' for _ in range(pad_size):\n'
|
||||
' norm_img_batch.append(pad_img)\n'
|
||||
r'\1'
|
||||
# 兼容 mineru 3.4.0:赋值行可能是
|
||||
# norm_img_batch = np.concatenate(norm_img_batch)
|
||||
# 或
|
||||
# norm_img_batch = np.ascontiguousarray(np.concatenate(norm_img_batch))
|
||||
# 用宽松正则匹配「norm_img_batch = ...np.concatenate(norm_img_batch)...」整行
|
||||
old = (
|
||||
r'(^ *)(norm_img_batch\s*=\s*'
|
||||
r'(?:np\.ascontiguousarray\()?' # 可选 ascontiguousarray 包裹
|
||||
r'np\.concatenate\(norm_img_batch\)'
|
||||
r'(?:\))?' # 可选闭合括号
|
||||
r')'
|
||||
)
|
||||
c2 = re.sub(old, new, c)
|
||||
new = (
|
||||
r'\1actual_batch_size = len(norm_img_batch)\n'
|
||||
r'\1if actual_batch_size < batch_num:\n'
|
||||
r'\1 pad_size = batch_num - actual_batch_size\n'
|
||||
r'\1 pad_img = np.zeros_like(norm_img_batch[0])\n'
|
||||
r'\1 for _ in range(pad_size):\n'
|
||||
r'\1 norm_img_batch.append(pad_img)\n'
|
||||
r'\1\2'
|
||||
)
|
||||
c2 = re.sub(old, new, c, flags=re.MULTILINE)
|
||||
if c2 == c:
|
||||
if 'actual_batch_size' not in c:
|
||||
raise RuntimeError('Patch B: cannot find norm_img_batch concatenation')
|
||||
@@ -52,8 +63,8 @@ def patch_b_predict_rec_batch(infer_dir):
|
||||
# 修改 range(len(rec_result)) → range(actual_batch_size)
|
||||
c3 = open(f).read()
|
||||
c4 = re.sub(
|
||||
r'for rno in range\(len\(rec_result\)\):',
|
||||
' for rno in range(actual_batch_size):',
|
||||
r'( +)for rno in range\(len\(rec_result\)\):',
|
||||
r'\1for rno in range(actual_batch_size):',
|
||||
c3
|
||||
)
|
||||
open(f, 'w').write(c4)
|
||||
@@ -63,9 +74,9 @@ def patch_c_predict_det_contiguous(infer_dir):
|
||||
"""predict_det.py: contiguous 检查"""
|
||||
f = os.path.join(infer_dir, 'predict_det.py')
|
||||
c = open(f).read()
|
||||
old = r'( inp = inp\.to\(self\.device\)\n)'
|
||||
new = r'\1 if not inp.is_contiguous():\n inp = inp.contiguous()\n'
|
||||
c2 = re.sub(old, new, c)
|
||||
old = r'(^ *)(inp = inp\.to\(self\.device\)\n)'
|
||||
new = r'\1\2\1if not inp.is_contiguous():\n\1 inp = inp.contiguous()\n'
|
||||
c2 = re.sub(old, new, c, flags=re.MULTILINE)
|
||||
if c2 == c:
|
||||
if 'is_contiguous' not in c:
|
||||
raise RuntimeError('Patch C: cannot find inp.to(device) line')
|
||||
|
||||
@@ -0,0 +1,34 @@
|
||||
#!/bin/bash
|
||||
# =============================================================================
|
||||
# MinerU ROCm 统一入口 — 启动时执行补丁,然后启动指定服务
|
||||
# =============================================================================
|
||||
set -e
|
||||
|
||||
VENV=/opt/mineru_venv
|
||||
SCRIPTS_DIR=/opt/scripts
|
||||
|
||||
echo "[entrypoint] running patches..."
|
||||
|
||||
# vllm 平台补丁(非致命:非 GPU 容器可能失败)
|
||||
if [ -f ${SCRIPTS_DIR}/patch_vllm_platform.py ]; then
|
||||
set +e
|
||||
${VENV}/bin/python ${SCRIPTS_DIR}/patch_vllm_platform.py
|
||||
vllm_rc=$?
|
||||
set -e
|
||||
if [ $vllm_rc -ne 0 ]; then
|
||||
echo "[entrypoint] patch_vllm_platform.py exited with $vllm_rc (non-fatal for router/gradio)"
|
||||
fi
|
||||
fi
|
||||
|
||||
# MinerU 推理补丁
|
||||
if [ -f ${SCRIPTS_DIR}/apply_mineru_patches.py ]; then
|
||||
${VENV}/bin/python ${SCRIPTS_DIR}/apply_mineru_patches.py || echo "[entrypoint] apply_mineru_patches.py failed (non-fatal)"
|
||||
fi
|
||||
|
||||
# 确保 vllm 源码目录在 PYTHONPATH 中
|
||||
# patch_vllm_platform.py 中的 os.environ 修改不持久化到父进程
|
||||
# 这里无条件导出,防止 editable install 被破坏后 vllm 无法导入
|
||||
export PYTHONPATH="/opt/vllm:${PYTHONPATH}"
|
||||
|
||||
echo "[entrypoint] starting: $*"
|
||||
exec "$@"
|
||||
@@ -0,0 +1,268 @@
|
||||
#!/usr/bin/env python3
|
||||
"""MinerU Gradio 客户端 — 纯前端,所有解析任务通过 HTTP 转发到 Router/Worker。
|
||||
|
||||
与 mineru-gradio 的区别:
|
||||
- 不启动内嵌 FastAPI server
|
||||
- 不需要 GPU / vLLM
|
||||
- 通过 HTTP 调用 mineru-router → mineru-worker 处理任务
|
||||
"""
|
||||
import os
|
||||
import time
|
||||
import asyncio
|
||||
import tempfile
|
||||
import zipfile
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import gradio as gr
|
||||
import httpx
|
||||
|
||||
API_BASE = os.environ.get("MINERU_API_BASE", "http://mineru-nginx:8000")
|
||||
DEFAULT_BACKEND = os.environ.get("MINERU_BACKEND", "hybrid-engine")
|
||||
DEFAULT_LANG = os.environ.get("MINERU_LANG", "ch")
|
||||
POLL_INTERVAL = float(os.environ.get("MINERU_POLL_INTERVAL", "1.0"))
|
||||
MAX_WAIT = float(os.environ.get("MINERU_MAX_WAIT", "600"))
|
||||
|
||||
CUSTOM_CSS = """
|
||||
.result-panel { min-height: 400px; }
|
||||
.status-panel { min-height: 200px; font-size: 13px; }
|
||||
"""
|
||||
|
||||
|
||||
async def discover_api(client: httpx.AsyncClient) -> str:
|
||||
"""探测实际可用的 API 前缀: /file_parse 端点"""
|
||||
candidates = [
|
||||
f"{API_BASE}/openapi.json",
|
||||
f"{API_BASE}/docs",
|
||||
]
|
||||
for url in candidates:
|
||||
try:
|
||||
r = await client.get(url, timeout=5)
|
||||
if r.status_code == 200:
|
||||
return "" # worker uses root-level endpoints like /file_parse
|
||||
except Exception:
|
||||
continue
|
||||
return ""
|
||||
|
||||
|
||||
async def submit_task(client: httpx.AsyncClient, api_prefix: str,
|
||||
file_path: str, file_name: str,
|
||||
backend: str, lang: str) -> dict:
|
||||
"""提交解析任务到 worker 的 /file_parse 端点"""
|
||||
submit_url = f"{API_BASE}/file_parse"
|
||||
with open(file_path, "rb") as f:
|
||||
files = {"files": (file_name, f, "application/pdf")}
|
||||
data = {
|
||||
"backend": backend,
|
||||
"lang_list": lang,
|
||||
"parse_method": "auto",
|
||||
}
|
||||
r = await client.post(submit_url, files=files, data=data, timeout=30)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
|
||||
async def get_task_status(client: httpx.AsyncClient, api_prefix: str,
|
||||
task_id: str) -> dict:
|
||||
"""查询任务状态"""
|
||||
url = f"{API_BASE}/tasks/{task_id}"
|
||||
r = await client.get(url, timeout=10)
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
|
||||
async def download_result(client: httpx.AsyncClient, api_prefix: str,
|
||||
task_id: str, output_dir: str) -> str:
|
||||
"""下载任务结果"""
|
||||
url = f"{API_BASE}/tasks/{task_id}/result"
|
||||
r = await client.get(url, timeout=60, follow_redirects=True)
|
||||
r.raise_for_status()
|
||||
zip_path = os.path.join(output_dir, f"{task_id}.zip")
|
||||
with open(zip_path, "wb") as f:
|
||||
f.write(r.content)
|
||||
return zip_path
|
||||
|
||||
|
||||
def extract_readme(zip_path: str, output_dir: str) -> str:
|
||||
"""解压结果并返回 markdown 内容"""
|
||||
extract_dir = os.path.join(output_dir, "extracted")
|
||||
os.makedirs(extract_dir, exist_ok=True)
|
||||
with zipfile.ZipFile(zip_path, "r") as zf:
|
||||
zf.extractall(extract_dir)
|
||||
# 查找 markdown 文件
|
||||
md_files = list(Path(extract_dir).rglob("*.md"))
|
||||
if md_files:
|
||||
return md_files[0].read_text(encoding="utf-8")
|
||||
return "No markdown output found in result."
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Gradio 处理函数
|
||||
# ---------------------------------------------------------------------------
|
||||
async def process_pdf(file_obj, backend, lang, progress=gr.Progress()):
|
||||
"""Gradio 事件处理:上传 → 提交 → 轮询 → 返回结果"""
|
||||
if file_obj is None:
|
||||
return "Please upload a PDF file.", "", ""
|
||||
|
||||
tmp_dir = tempfile.mkdtemp(prefix="mineru-client-")
|
||||
|
||||
# Gradio 6.0 binary mode:file_obj 是 bytes
|
||||
if isinstance(file_obj, bytes):
|
||||
file_path = os.path.join(tmp_dir, "input.pdf")
|
||||
file_name = "input.pdf"
|
||||
with open(file_path, "wb") as f:
|
||||
f.write(file_obj)
|
||||
elif isinstance(file_obj, str):
|
||||
file_path = file_obj
|
||||
file_name = os.path.basename(file_path)
|
||||
else:
|
||||
return "Unsupported file object type", "", ""
|
||||
|
||||
async with httpx.AsyncClient(timeout=httpx.Timeout(120)) as client:
|
||||
# 1. 探测 API
|
||||
progress(0.05, desc="Connecting to API...")
|
||||
api_prefix = await discover_api(client)
|
||||
progress(0.1, desc=f"API prefix: {api_prefix}")
|
||||
|
||||
# 2. 提交任务
|
||||
progress(0.15, desc="Submitting task...")
|
||||
try:
|
||||
submit_resp = await submit_task(client, api_prefix, file_path, file_name, backend, lang)
|
||||
except Exception as e:
|
||||
return f"Task submission failed: {e}", "", ""
|
||||
|
||||
task_id = submit_resp.get("task_id") or submit_resp.get("data", {}).get("task_id")
|
||||
if not task_id:
|
||||
return f"Unexpected submit response:\n{json.dumps(submit_resp, indent=2)}", "", ""
|
||||
|
||||
progress(0.2, desc=f"Task ID: {task_id}")
|
||||
|
||||
# 3. 轮询状态
|
||||
status_text = ""
|
||||
start_time = time.time()
|
||||
last_idx = -1
|
||||
tickers = ["⣾", "⣽", "⣻", "⢿", "⡿", "⣟", "⣯", "⣷"]
|
||||
|
||||
while True:
|
||||
elapsed = time.time() - start_time
|
||||
if elapsed > MAX_WAIT:
|
||||
return f"Task timed out after {MAX_WAIT}s\nLast status:\n{status_text}", "", ""
|
||||
|
||||
t = (int(elapsed * 2)) % len(tickers)
|
||||
progress(min(0.2 + 0.6 * (elapsed / 60), 0.8),
|
||||
desc=f"{tickers[t]} Processing... ({int(elapsed)}s)")
|
||||
|
||||
try:
|
||||
status_resp = await get_task_status(client, api_prefix, task_id)
|
||||
except Exception as e:
|
||||
status_text = f"Polling error: {e}"
|
||||
await asyncio.sleep(POLL_INTERVAL)
|
||||
continue
|
||||
|
||||
status = status_resp.get("status", "unknown")
|
||||
progress_pct = status_resp.get("progress", 0)
|
||||
err_msg = status_resp.get("error", "")
|
||||
status_text = json.dumps(status_resp, indent=2, ensure_ascii=False)
|
||||
|
||||
if status == "completed" or status == "success":
|
||||
progress(0.85, desc="Downloading result...")
|
||||
break
|
||||
elif status == "failed":
|
||||
return f"## Task Failed\n\nError: {err_msg}\n\n```json\n{status_text}\n```", "", ""
|
||||
elif status == "processing":
|
||||
pass
|
||||
else:
|
||||
pass
|
||||
|
||||
await asyncio.sleep(POLL_INTERVAL)
|
||||
|
||||
# 4. 下载结果
|
||||
progress(0.9, desc="Downloading result...")
|
||||
try:
|
||||
zip_path = await download_result(client, api_prefix, task_id, tmp_dir)
|
||||
except Exception as e:
|
||||
return f"Result download failed: {e}\n\nStatus:\n{status_text}", "", ""
|
||||
|
||||
# 5. 提取 markdown
|
||||
progress(0.95, desc="Extracting results...")
|
||||
try:
|
||||
md_content = extract_readme(zip_path, tmp_dir)
|
||||
except Exception as e:
|
||||
md_content = f"Extraction error: {e}"
|
||||
|
||||
progress(1.0, desc="Done!")
|
||||
return md_content, status_text, zip_path
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Gradio UI
|
||||
# ---------------------------------------------------------------------------
|
||||
def create_ui():
|
||||
with gr.Blocks(title="MinerU ROCm — Document Parser") as demo:
|
||||
gr.Markdown("""
|
||||
# MinerU ROCm — Document Parser
|
||||
|
||||
Upload a PDF file. Processing is handled by the **mineru-router → mineru-worker** backend.
|
||||
""")
|
||||
|
||||
with gr.Row():
|
||||
with gr.Column(scale=1):
|
||||
file_input = gr.UploadButton(
|
||||
label="Upload PDF",
|
||||
file_types=[".pdf"],
|
||||
type="binary",
|
||||
)
|
||||
backend_dd = gr.Dropdown(
|
||||
label="Backend",
|
||||
choices=["hybrid-engine", "hybrid-http-client",
|
||||
"vlm-http-client", "pipeline", "auto"],
|
||||
value=DEFAULT_BACKEND,
|
||||
)
|
||||
lang_dd = gr.Dropdown(
|
||||
label="Language",
|
||||
choices=["ch", "en", "japan", "korean"],
|
||||
value=DEFAULT_LANG,
|
||||
)
|
||||
submit_btn = gr.Button("Parse Document", variant="primary", size="lg")
|
||||
status_display = gr.Textbox(
|
||||
label="Status",
|
||||
lines=6,
|
||||
max_lines=12,
|
||||
elem_classes=["status-panel"],
|
||||
)
|
||||
zip_display = gr.File(
|
||||
label="Download Result ZIP",
|
||||
type="filepath",
|
||||
visible=True,
|
||||
)
|
||||
|
||||
with gr.Column(scale=2):
|
||||
result_display = gr.Markdown(
|
||||
value="*Upload a PDF to start parsing...*",
|
||||
elem_classes=["result-panel"],
|
||||
)
|
||||
|
||||
file_input.upload(
|
||||
fn=process_pdf,
|
||||
inputs=[file_input, backend_dd, lang_dd],
|
||||
outputs=[result_display, status_display, zip_display],
|
||||
)
|
||||
submit_btn.click(
|
||||
fn=lambda f,b,l: "Please use the Upload button above.",
|
||||
inputs=[file_input, backend_dd, lang_dd],
|
||||
outputs=[result_display, status_display, zip_display],
|
||||
)
|
||||
|
||||
return demo
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
server_name = os.environ.get("GRADIO_SERVER_NAME", "0.0.0.0")
|
||||
server_port = int(os.environ.get("GRADIO_SERVER_PORT", "7860"))
|
||||
demo = create_ui()
|
||||
demo.queue(default_concurrency_limit=3, max_size=10).launch(
|
||||
server_name=server_name,
|
||||
server_port=server_port,
|
||||
share=False,
|
||||
css=CUSTOM_CSS,
|
||||
)
|
||||
@@ -0,0 +1,309 @@
|
||||
#!/usr/bin/env python3
|
||||
"""vllm 平台检测补丁
|
||||
|
||||
问题 1:amdsmi 不可用时 platform 回退到 torch.version.hip
|
||||
问题 2:rocm.py 中 logger.warning_once() 导致循环导入(默认不再改写;仅保留为显式开关)
|
||||
"""
|
||||
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
import sysconfig
|
||||
import traceback
|
||||
|
||||
VLLM_DIR = '/opt/vllm/vllm'
|
||||
VLLM_SRC_DIR = '/opt/vllm'
|
||||
|
||||
|
||||
def patch6_init_platform_fallback():
|
||||
"""补丁 6:platforms/__init__.py —— torch.version.hip 兜底
|
||||
|
||||
vLLM 的 is_rocm 检测依赖 import amdsmi;amdsmi 未装时 is_rocm=False,
|
||||
current_platform 落到 UnspecifiedPlatform,device_type='' →
|
||||
RuntimeError: Device string must not be empty。
|
||||
|
||||
本补丁在 RocmPlatform 的 return 语句前注入 torch.version.hip 兜底。
|
||||
采用行结构 + return 定位,不依赖精确字符串匹配
|
||||
(vllm main 分支重构频繁,固定字符串匹配已多次失效)。
|
||||
"""
|
||||
f = os.path.join(VLLM_DIR, 'platforms', '__init__.py')
|
||||
lines = open(f).read().splitlines(keepends=True)
|
||||
|
||||
# 兼容单/双引号两种写法(vllm 不同版本可能不同)
|
||||
return_idx = None
|
||||
for i, ln in enumerate(lines):
|
||||
if 'RocmPlatform' in ln and 'return' in ln and 'is_rocm' in ln:
|
||||
return_idx = i
|
||||
break
|
||||
if return_idx is None:
|
||||
print('Patch 6: RocmPlatform return statement not found; skipping.')
|
||||
return
|
||||
|
||||
if 'torch.version.hip is not None' in ''.join(lines):
|
||||
print('Patch 6: already applied (torch.version.hip fallback present).')
|
||||
return
|
||||
|
||||
indent = ' ' * (len(lines[return_idx]) - len(lines[return_idx].lstrip()))
|
||||
inject = (
|
||||
f"{indent}# amdsmi fallback: also check torch.version.hip\n"
|
||||
f"{indent}if not is_rocm:\n"
|
||||
f"{indent} try:\n"
|
||||
f"{indent} import torch as _torch\n"
|
||||
f"{indent} if _torch.version.hip is not None:\n"
|
||||
f"{indent} is_rocm = True\n"
|
||||
f"{indent} except Exception:\n"
|
||||
f"{indent} pass\n"
|
||||
)
|
||||
lines.insert(return_idx, inject)
|
||||
open(f, 'w').write(''.join(lines))
|
||||
print('Patch 6: __init__.py torch.version.hip fallback applied.')
|
||||
|
||||
|
||||
def patch7_rocm_break_import_cycle():
|
||||
"""补丁 7:platforms/rocm.py —— logger.warning_once → sys.stderr.write
|
||||
|
||||
默认启用安全的“整段 except 块”替换,避免 logger.warning_once() 在平台检测
|
||||
期间触发循环导入,导致 current_platform 落到 UnspecifiedPlatform。
|
||||
如需禁用,设置 VLLM_PATCH_ROCM_WARNING_ONCE=0。
|
||||
"""
|
||||
if os.environ.get('VLLM_PATCH_ROCM_WARNING_ONCE') == '0':
|
||||
print('Patch 7: skipped (VLLM_PATCH_ROCM_WARNING_ONCE=0).')
|
||||
return
|
||||
f = os.path.join(VLLM_DIR, 'platforms', 'rocm.py')
|
||||
c = open(f).read()
|
||||
if 'amdsmi unavailable, using torch.cuda fallback' in c:
|
||||
print('Patch 7: already applied.')
|
||||
return
|
||||
pattern = re.compile(
|
||||
r'(\n\s*except Exception as e:\n'
|
||||
r'\s*logger\.debug\("Failed to get GCN arch via amdsmi: %s", e\)\n'
|
||||
r'\s*logger\.warning_once\(\n'
|
||||
r'(?:\s*"[^"]*"\n)+'
|
||||
r'\s*\)\n)',
|
||||
re.MULTILINE,
|
||||
)
|
||||
|
||||
def repl(match):
|
||||
indent = re.search(r'\n(\s*)except Exception as e:', match.group(1)).group(1)
|
||||
body_indent = indent + ' '
|
||||
return (
|
||||
f'\n{indent}except Exception as e:\n'
|
||||
f'{body_indent}import sys as _sys\n'
|
||||
f'{body_indent}_sys.stderr.write('
|
||||
'"vLLM ROCm: amdsmi unavailable, using torch.cuda fallback for GPU detection\\n")\n'
|
||||
)
|
||||
|
||||
c2, n = pattern.subn(repl, c, count=1)
|
||||
if c2 != c:
|
||||
open(f, 'w').write(c2)
|
||||
print('Patch 7: rocm.py logger.warning_once circular import patch applied.')
|
||||
else:
|
||||
print('Patch 7: already applied or pattern not found.')
|
||||
|
||||
|
||||
def patch9_registry_model_impl_compat():
|
||||
"""补丁 9:registry.py —— vllm main 与 transformers v4 兼容
|
||||
|
||||
问题:vllm main 的 ModelRegistry 大量访问 model_config.model_impl,
|
||||
该属性是 vllm ModelConfig 的字段(默认 "auto"),但 inspect 链路传入的
|
||||
有时是 transformers config 对象(如 Qwen2VLConfig),v4 没有此属性 →
|
||||
AttributeError → "Model architectures [...] failed to be inspected"。
|
||||
同时 _try_resolve_transformers 末尾调用 model_config._get_transformers_backend_cls(),
|
||||
v4 的 config 也没有该方法。
|
||||
|
||||
根因:mineru[core] 锁定 transformers<5.0.0(v4),vllm main 期望 v5。
|
||||
本补丁把所有 model_config.model_impl 访问改成 getattr 兜底(缺属性时当 "auto",
|
||||
走 fallback 分支匹配 vllm 注册表),并给 _get_transformers_backend_cls 加兜底。
|
||||
|
||||
采用逐处 getattr 替换,不依赖方法定位(比方法注入更可靠)。
|
||||
"""
|
||||
f = os.path.join(VLLM_DIR, 'model_executor', 'models', 'registry.py')
|
||||
if not os.path.exists(f):
|
||||
print('Patch 9: registry.py not found; skipping.')
|
||||
return
|
||||
c = open(f).read()
|
||||
if '# mineru-rocm: model_impl v4 compat' in c:
|
||||
print('Patch 9: already applied (model_impl v4 compat present).')
|
||||
return
|
||||
|
||||
# 把 model_config.model_impl 访问替换为 getattr 兜底
|
||||
before = c.count('model_config.model_impl')
|
||||
c2 = c.replace(
|
||||
'model_config.model_impl',
|
||||
'getattr(model_config, "model_impl", "auto")'
|
||||
)
|
||||
replaced = before - c2.count('model_config.model_impl')
|
||||
|
||||
# _get_transformers_backend_cls 兜底:v4 无此方法
|
||||
c2 = c2.replace(
|
||||
'return model_config._get_transformers_backend_cls()',
|
||||
'return getattr(model_config, "_get_transformers_backend_cls", lambda: None)()'
|
||||
)
|
||||
|
||||
# 写入幂等标记(注释,便于重入检测)
|
||||
c2 = '# mineru-rocm: model_impl v4 compat\n' + c2
|
||||
|
||||
if c2 != c:
|
||||
open(f, 'w').write(c2)
|
||||
print(f'Patch 9: registry.py model_impl v4 compat applied '
|
||||
f'({replaced} access(es) wrapped).')
|
||||
else:
|
||||
print('Patch 9: no changes applied (pattern not found).')
|
||||
|
||||
|
||||
def ensure_vllm_dist_info():
|
||||
"""为 /opt/vllm 源码导入创建最小 dist-info。
|
||||
|
||||
vLLM 平台检测会通过 importlib.metadata 查询 vllm 分发元数据/entry points。
|
||||
如果只靠 PYTHONPATH 导入源码,没有 dist-info,就会出现:
|
||||
"The vLLM package was not found...",并可能得到 UnspecifiedPlatform。
|
||||
|
||||
注意:仅当 rocm_platform_plugin 函数实际存在时才注册 entry_point,
|
||||
否则 vLLM 会因 AttributeError 回退到 UnspecifiedPlatform。
|
||||
"""
|
||||
try:
|
||||
import vllm
|
||||
except Exception:
|
||||
return
|
||||
|
||||
version = getattr(vllm, '__version__', '0.1.dev1') or '0.1.dev1'
|
||||
# PEP440 合法性校验:vllm 在 _version.py 缺失时会回退到 'dev',
|
||||
# 这不是合法 PEP440,会触发下游 packaging.version.parse 抛 InvalidVersion
|
||||
# (例如 mineru.backend.vlm.utils:set_default_gpu_memory_utilization)。
|
||||
try:
|
||||
from packaging.version import Version as _PEP440Version
|
||||
_PEP440Version(version)
|
||||
except Exception:
|
||||
print(f'WARNING: vllm.__version__={version!r} is not PEP440-compliant; '
|
||||
f'falling back to "0.11.0" for dist-info.')
|
||||
version = '0.11.0'
|
||||
try:
|
||||
vllm.__version__ = version
|
||||
except Exception:
|
||||
pass
|
||||
site_packages = sysconfig.get_paths().get('purelib')
|
||||
if not site_packages:
|
||||
return
|
||||
dist_info = os.path.join(site_packages, f'vllm-{version}.dist-info')
|
||||
os.makedirs(dist_info, exist_ok=True)
|
||||
|
||||
metadata = os.path.join(dist_info, 'METADATA')
|
||||
if not os.path.exists(metadata):
|
||||
with open(metadata, 'w') as fp:
|
||||
fp.write(
|
||||
'Metadata-Version: 2.1\n'
|
||||
'Name: vllm\n'
|
||||
f'Version: {version}\n'
|
||||
'Summary: vLLM source tree mounted at /opt/vllm\n'
|
||||
)
|
||||
|
||||
with open(os.path.join(dist_info, 'top_level.txt'), 'w') as fp:
|
||||
fp.write('vllm\n')
|
||||
with open(os.path.join(dist_info, 'INSTALLER'), 'w') as fp:
|
||||
fp.write('mineru-rocm-runtime\n')
|
||||
|
||||
# 仅当 rocm_platform_plugin 实际存在时才注册,避免 vLLM 加载失败
|
||||
entry_points = ''
|
||||
try:
|
||||
from vllm.platforms.rocm import rocm_platform_plugin # noqa: F401
|
||||
entry_points += (
|
||||
'[vllm.platform_plugins]\n'
|
||||
'rocm = vllm.platforms.rocm:rocm_platform_plugin\n'
|
||||
)
|
||||
print('rocm_platform_plugin found, registering entry point.')
|
||||
except (ImportError, AttributeError, SystemExit, KeyboardInterrupt):
|
||||
print('WARNING: rocm_platform_plugin not found in vllm.platforms.rocm; '
|
||||
'skipping entry_point registration.')
|
||||
except BaseException as e:
|
||||
print(f'WARNING: rocm_platform_plugin check failed ({type(e).__name__}: {e}); '
|
||||
'skipping entry_point registration.')
|
||||
|
||||
with open(os.path.join(dist_info, 'entry_points.txt'), 'w') as fp:
|
||||
fp.write(entry_points)
|
||||
with open(os.path.join(dist_info, 'RECORD'), 'w') as fp:
|
||||
fp.write('')
|
||||
print(f'vllm dist-info ensured: {dist_info}')
|
||||
|
||||
|
||||
def ensure_vllm_installed():
|
||||
"""确保 vLLM Python 包在当前虚拟环境中可导入。
|
||||
|
||||
某些 MinerU 依赖安装流程后,site-packages 中可能缺少 vLLM 分发元数据,但
|
||||
/opt/vllm 源码仍在。此时直接把 /opt/vllm 加入 PYTHONPATH/sys.path 再导入,
|
||||
避免 editable 安装触发 pyproject 元数据校验失败。
|
||||
"""
|
||||
try:
|
||||
import vllm # noqa: F401
|
||||
return
|
||||
except ModuleNotFoundError as exc:
|
||||
if exc.name != 'vllm':
|
||||
raise
|
||||
|
||||
if not os.path.isdir(VLLM_SRC_DIR):
|
||||
raise ModuleNotFoundError(
|
||||
'vllm is not installed and /opt/vllm source directory is missing'
|
||||
)
|
||||
|
||||
print('vllm package missing; trying PYTHONPATH fallback with /opt/vllm...')
|
||||
current = os.environ.get('PYTHONPATH', '')
|
||||
paths = [p for p in current.split(':') if p]
|
||||
if VLLM_SRC_DIR not in paths:
|
||||
os.environ['PYTHONPATH'] = f"{VLLM_SRC_DIR}:{current}" if current else VLLM_SRC_DIR
|
||||
if VLLM_SRC_DIR not in sys.path:
|
||||
sys.path.insert(0, VLLM_SRC_DIR)
|
||||
|
||||
import vllm # noqa: F401
|
||||
print('vllm import recovered via PYTHONPATH fallback.')
|
||||
|
||||
|
||||
def install_sitecustomize_force_rocm():
|
||||
"""已弃用:sitecustomize 时机问题(current_platform 是 lazy init,
|
||||
sitecustomize 触发提前 resolve 时补丁 6 尚未应用)导致兜底无效。
|
||||
保留空壳仅为兼容旧调用,实际不做任何事。
|
||||
平台检测统一由补丁 6(rocm_platform_plugin torch.version.hip 兜底)解决。
|
||||
"""
|
||||
# 清理历史遗留的 sitecustomize 片段(幂等)
|
||||
site_packages = sysconfig.get_paths().get('purelib')
|
||||
if site_packages:
|
||||
sc = os.path.join(site_packages, 'sitecustomize.py')
|
||||
marker = '# mineru-rocm: force rocm platform'
|
||||
if os.path.exists(sc) and marker in open(sc).read():
|
||||
# 重写文件,移除我们的片段
|
||||
c = open(sc).read()
|
||||
# 片段从 marker 行开始到文件末尾
|
||||
idx = c.find(marker)
|
||||
# 回退到 marker 前的换行
|
||||
while idx > 0 and c[idx - 1] == '\n':
|
||||
idx -= 1
|
||||
c = c[:idx].rstrip() + '\n'
|
||||
open(sc, 'w').write(c)
|
||||
print('Patch 8: removed legacy sitecustomize force-rocm snippet.')
|
||||
|
||||
|
||||
def main():
|
||||
patch6_init_platform_fallback()
|
||||
patch7_rocm_break_import_cycle()
|
||||
patch9_registry_model_impl_compat()
|
||||
install_sitecustomize_force_rocm()
|
||||
try:
|
||||
ensure_vllm_installed()
|
||||
ensure_vllm_dist_info()
|
||||
import vllm
|
||||
from vllm.platforms import current_platform
|
||||
platform_name = type(current_platform).__name__
|
||||
print(f'vllm runtime import OK: {vllm.__version__}, platform={platform_name}')
|
||||
if platform_name == 'UnspecifiedPlatform':
|
||||
print(
|
||||
'WARNING: vLLM platform detection returned UnspecifiedPlatform. '
|
||||
'This is expected on containers without GPU access (router/gradio). '
|
||||
'On worker containers, check: amdsmi, /dev/kfd, /dev/dri, '
|
||||
'and vllm.platform_plugins metadata.'
|
||||
)
|
||||
except Exception:
|
||||
print('WARNING: vllm runtime import failed after platform patches (non-fatal):')
|
||||
traceback.print_exc()
|
||||
print('vllm platform patches done.')
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
main()
|
||||
Reference in New Issue
Block a user