Creating model: ('PP-DocLayoutV3', None, None)
Model files already exist. Using cached files. To redownload, please delete the directory manually: `/home/paddleocr/.paddlex/official_models/PP-DocLayoutV3`.
python: /work/Paddle/third_party/eigen3/unsupported/Eigen/CXX11/src/Tensor/TensorExecutor.h:612: static void Eigen::internal::TensorExecutor<const Eigen::TensorAssignOp<Eigen::TensorMap<Eigen::Tensor<float, 2, 1, int>, 16>, const Eigen::TensorSlicingOp<const Eigen::DSizes<int, 2>, const Eigen::DSizes<int, 2>, const Eigen::TensorMap<Eigen::Tensor<const float, 2, 1, int>, 16>>>, Eigen::GpuDevice>::run(const Expression &, const GpuDevice &) [Expression = const Eigen::TensorAssignOp<Eigen::TensorMap<Eigen::Tensor<float, 2, 1, int>, 16>, const Eigen::TensorSlicingOp<const Eigen::DSizes<int, 2>, const Eigen::DSizes<int, 2>, const Eigen::TensorMap<Eigen::Tensor<const float, 2, 1, int>, 16>>>, Device = Eigen::GpuDevice, Vectorizable = false, Tiling = Eigen::internal::On]: Assertion `hipGetLastError() == hipSuccess' failed.
--------------------------------------
C++ Traceback (most recent call last):
--------------------------------------
0 paddle::framework::ThreadPoolTempl<paddle::framework::StlThreadEnvironment>::WorkerLoop(int)
1 paddle::framework::PirInterpreter::RunInstructionBaseAsync(unsigned long)
2 paddle::framework::PirInterpreter::RunInstructionBase(paddle::framework::InstructionBase*)
3 paddle::framework::PhiKernelInstruction::Run()
4 phi::KernelImpl<void (*)(phi::GPUContext const&, phi::DenseTensor const&, std::vector<long, std::allocator<long> > const&, paddle::experimental::IntArrayBase<phi::DenseTensor> const&, paddle::experimental::IntArrayBase<phi::DenseTensor> const&, std::vector<long, std::allocator<long> > const&, std::vector<long, std::allocator<long> > const&, phi::DenseTensor*), &(void phi::SliceKernel<float, phi::GPUContext>(phi::GPUContext const&, phi::DenseTensor const&, std::vector<long, std::allocator<long> > const&, paddle::experimental::IntArrayBase<phi::DenseTensor> const&, paddle::experimental::IntArrayBase<phi::DenseTensor> const&, std::vector<long, std::allocator<long> > const&, std::vector<long, std::allocator<long> > const&, phi::DenseTensor*))>::Compute(phi::KernelContext*)
5 void phi::SliceCompute<float, phi::GPUContext, 2ul>(phi::GPUContext const&, phi::DenseTensor const&, std::vector<long, std::allocator<long> > const&, std::vector<long, std::allocator<long> > const&, std::vector<long, std::allocator<long> > const&, std::vector<long, std::allocator<long> > const&, std::vector<long, std::allocator<long> > const&, phi::DenseTensor*)
6 phi::funcs::EigenSlice<Eigen::GpuDevice, float, 2>::Eval(Eigen::GpuDevice const&, Eigen::TensorMap<Eigen::Tensor<float, 2, 1, int>, 16, Eigen::MakePointer>, Eigen::TensorMap<Eigen::Tensor<float const, 2, 1, int>, 16, Eigen::MakePointer> const&, Eigen::DSizes<int, 2> const&, Eigen::DSizes<int, 2> const&)
----------------------
Error Message Summary:
----------------------
FatalError: `Process abort signal` is detected by the operating system.
[TimeInfo: *** Aborted at 1785858992 (unix time) try "date -d @1785858992" if you are using GNU date ***]
[SignalInfo: *** SIGABRT (@0x2f) received by PID 47 (TID 0x7749759ff6c0) from PID 47 ***]
+------------------------------------------------------------------------------+
| AMD-SMI 26.5.0+2b22ab01 |
| OS kernel Version: 6.17.13-2-pve |
| ROCm Version: 7.14.0 |
| VBIOS Version: 023.008.000.068.000001 |
| Platform: Linux Baremetal |
|-------------------------------------+----------------------------------------|
| BDF GPU-Name | Mem-Uti Temp UEC Power-Usage |
| GPU HIP-ID OAM-ID Partition-Mode | GFX-Uti Fan Mem-Usage |
|=====================================+========================================|
| 0000:03:00.0 ...Radeon AI PRO R9700 | 0 % 32 °C 0 42/300 W |
| 0 0 N/A N/A | 17 % 24.71 57/32624 MB |
+-------------------------------------+----------------------------------------+
+------------------------------------------------------------------------------+
| Processes: |
| GPU PID Process Name GTT_MEM VRAM_MEM MEM_USAGE CU % SDMA |
|==============================================================================|
| No running processes found |
+------------------------------------------------------------------------------+
======================================== ROCm System Management Interface ========================================
================================================== Concise Info ==================================================
Device Node IDs Temp Power Partitions SCLK MCLK Fan Perf PwrCap VRAM% GPU%
(DID, GUID) (Edge) (Avg) (Mem, Compute, ID)
==================================================================================================================
0 2 0x7551, 51350 N/A N/A N/A, N/A, 0 N/A N/A 0% unknown N/A 0% 0%
==================================================================================================================
============================================== End of ROCm SMI Log ===============================================
🔎 Search before asking
🐛 Bug (问题描述)
Two containers created as per PaddleOCR-VL-AMD_GPU Service Deployment.
Output of paddleocr-vl:
Output of paddleocr-genai-vllm-server:
(EngineCore_DP0 pid=59) INFO 08-04 15:46:48 [core.py:96] Initializing a V1 LLM engine (v0.14.0rc2.dev331+ga698e8e7a) with config: model='/home/paddleocr/.paddlex/official_models/PaddleOCR-VL-1.6', speculative_config=None, tokenizer='/home/paddleocr/.paddlex/official_models/PaddleOCR-VL-1.6', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=16384, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=True, quantization=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=PaddleOCR-VL-1.6-0.9B, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none', '+sparse_attn_indexer'], 'splitting_ops': ['vllm::unified_attention', 'vllm::unified_attention_with_output', 'vllm::unified_mla_attention', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::gdn_attention_core', 'vllm::kda_attention', 'vllm::sparse_attn_indexer', 'vllm::rocm_aiter_sparse_attn_indexer'], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [131072], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 96, 104, 112, 120, 128, 136, 144, 152, 160, 168, 176, 184, 192, 200, 208, 216, 224, 232, 240, 248, 256, 272, 288, 304, 320, 336, 352, 368, 384, 400, 416, 432, 448, 464, 480, 496, 512], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': True, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 512, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': True}, 'local_cache_dir': None} (EngineCore_DP0 pid=59) /usr/local/lib/python3.12/dist-packages/vllm/v1/executor/uniproc_executor.py:52: UserWarning: Failed to get the IP address, using 0.0.0.0 by default.The value can be set by the environment variable VLLM_HOST_IP or HOST_IP. (EngineCore_DP0 pid=59) distributed_init_method = get_distributed_init_method(get_ip(), get_open_port()) (EngineCore_DP0 pid=59) INFO 08-04 15:46:52 [parallel_state.py:1212] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://0.0.0.0:38993 backend=nccl [W804 15:46:52.570822714 socket.cpp:209] [c10d] The hostname of the client socket cannot be retrieved. err=-3 (EngineCore_DP0 pid=59) INFO 08-04 15:46:52 [parallel_state.py:1423] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A (EngineCore_DP0 pid=59) Using a slow image processor as `use_fast` is unset and a slow processor was saved with this model. `use_fast=True` will be the default behavior in v4.52, even if the model was saved with a slow processor. This will result in minor differences in outputs. You'll still be able to use a slow processor with `use_fast=False`. (EngineCore_DP0 pid=59) INFO 08-04 15:46:57 [gpu_model_runner.py:4021] Starting to load model /home/paddleocr/.paddlex/official_models/PaddleOCR-VL-1.6... (EngineCore_DP0 pid=59) INFO 08-04 15:46:58 [rocm.py:178] Set FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE to enable Flash Attention Triton backend on RDNA. (EngineCore_DP0 pid=59) INFO 08-04 15:46:58 [rocm.py:396] Using Torch SDPA backend for ViT model. (EngineCore_DP0 pid=59) INFO 08-04 15:46:58 [mm_encoder_attention.py:77] Using AttentionBackendEnum.TORCH_SDPA for MMEncoderAttention. (EngineCore_DP0 pid=59) WARNING 08-04 15:46:58 [activation.py:577] [ROCm] PyTorch's native GELU with tanh approximation is unstable. Falling back to GELU(approximate='none'). (EngineCore_DP0 pid=59) INFO 08-04 15:46:58 [rocm.py:338] Using Triton Attention backend. (EngineCore_DP0 pid=59) WARNING 08-04 15:46:58 [compilation.py:1047] Op 'sparse_attn_indexer' not present in model, enabling with '+sparse_attn_indexer' has no effect Loading safetensors checkpoint shards: 100% 1/1 [00:01<00:00, 1.30s/it] (EngineCore_DP0 pid=59) INFO 08-04 15:47:00 [default_loader.py:291] Loading weights took 1.51 seconds (EngineCore_DP0 pid=59) INFO 08-04 15:47:00 [gpu_model_runner.py:4118] Model loading took 1.97 GiB memory and 1.951697 seconds (EngineCore_DP0 pid=59) INFO 08-04 15:47:01 [gpu_model_runner.py:4946] Encoder cache will be initialized with a budget of 131072 tokens, and profiled with 104 image items of the maximum feature size. (EngineCore_DP0 pid=59) /usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/conv.py:124: UserWarning: Failed validator: GCN_ARCH_NAME (Triggered internally at /app/pytorch/aten/src/ATen/hip/tunable/Tunable.cpp:364.) (EngineCore_DP0 pid=59) x = F.linear( (EngineCore_DP0 pid=59) /usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/conv.py:124: UserWarning: failed to open file '/app/afo_tune_device_0_full.csv' for writing; your tuning results will not be saved (Triggered internally at /app/pytorch/aten/src/ATen/hip/tunable/Tunable.cpp:643.) (EngineCore_DP0 pid=59) x = F.linear( (EngineCore_DP0 pid=59) INFO 08-04 15:47:27 [backends.py:805] Using cache directory: /home/paddleocr/.cache/vllm/torch_compile_cache/f5684c14f6/rank_0_0/backbone for vLLM's torch.compile (EngineCore_DP0 pid=59) INFO 08-04 15:47:27 [backends.py:865] Dynamo bytecode transform time: 3.17 s (EngineCore_DP0 pid=59) INFO 08-04 15:47:31 [backends.py:267] Directly load the compiled graph(s) for compile range (1, 131072) from the cache, took 1.798 s (EngineCore_DP0 pid=59) INFO 08-04 15:47:31 [monitor.py:34] torch.compile takes 4.97 s in total (EngineCore_DP0 pid=59) INFO 08-04 15:47:33 [gpu_worker.py:356] Available KV cache memory: 9.34 GiB (EngineCore_DP0 pid=59) INFO 08-04 15:47:33 [kv_cache_utils.py:1307] GPU KV cache size: 544,080 tokens (EngineCore_DP0 pid=59) INFO 08-04 15:47:33 [kv_cache_utils.py:1312] Maximum concurrency for 16,384 tokens per request: 33.21x Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 10% 5/51 [00:00<00:01, 42.36Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 20% 10/51 [00:00<00:00, 42.6Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 29% 15/51 [00:00<00:00, 43.0Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 39% 20/51 [00:00<00:00, 42.8Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 49% 25/51 [00:00<00:00, 42.7Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 59% 30/51 [00:00<00:00, 42.7Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 69% 35/51 [00:00<00:00, 42.9Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 78% 40/51 [00:00<00:00, 42.9Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 88% 45/51 [00:01<00:00, 42.0Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 98% 50/51 [00:01<00:00, 42.5Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100% 51/51 [00:01<00:00, 42.68it/s] Capturing CUDA graphs (decode, FULL): 100% 35/35 [00:01<00:00, 34.01it/s] (EngineCore_DP0 pid=59) INFO 08-04 15:47:36 [gpu_model_runner.py:5051] Graph capturing finished in 3 secs, took 0.25 GiB (EngineCore_DP0 pid=59) INFO 08-04 15:47:36 [core.py:272] init engine (profile, create kv cache, warmup model) took 35.30 seconds (EngineCore_DP0 pid=59) INFO 08-04 15:47:36 [vllm.py:624] Asynchronous scheduling is enabled. (APIServer pid=28) INFO 08-04 15:47:36 [api_server.py:665] Supported tasks: ['generate'] (APIServer pid=28) INFO 08-04 15:47:36 [serving.py:177] Warming up chat template processing... (APIServer pid=28) INFO 08-04 15:47:36 [hf.py:308] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this. (APIServer pid=28) INFO 08-04 15:47:36 [serving.py:212] Chat template warmup completed in 16.4ms (APIServer pid=28) INFO 08-04 15:47:36 [api_server.py:946] Starting vLLM API server 0 on http://[::]:8080 (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:38] Available routes are: (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /docs, Methods: HEAD, GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /redoc, Methods: HEAD, GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /tokenize, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /detokenize, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /inference/v1/generate, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /pause, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /resume, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /is_paused, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /metrics, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /health, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/chat/completions, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/responses, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/audio/transcriptions, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/audio/translations, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/completions, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/completions/render, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/messages, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/models, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /load, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /version, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /ping, Methods: GET (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /ping, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /invocations, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /classify, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/embeddings, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /score, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/score, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /rerank, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v1/rerank, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /v2/rerank, Methods: POST (APIServer pid=28) INFO 08-04 15:47:36 [launcher.py:46] Route: /pooling, Methods: POST (APIServer pid=28) INFO: Started server process [28] (APIServer pid=28) INFO: Waiting for application startup. (APIServer pid=28) INFO: Application startup complete.🏃♂️ Environment (运行环境)
Server Hardware:
Container versions:
amd-smi:
rocm-smi:
🌰 Minimal Reproducible Example (最小可复现问题的Demo)
Pull both docker containers as per https://www.paddleocr.ai/latest/en/version3.x/pipeline_usage/PaddleOCR-VL-AMD-GPU.html#4-service-deployment.
Expected result:
Actual result:
Assertion `hipGetLastError() == hipSuccess' failed