Files
amd-r9700-vllm-toolboxes/benchmarks/benchmark_results_amd-r9700/cpatonn_Qwen3-Next-80B-A3B-Instruct-AWQ-4bit_tp2_server.log
T
2026-03-26 20:50:00 +08:00

1129 lines
133 KiB
Plaintext

/opt/venv/lib64/python3.13/site-packages/torch/library.py:357: UserWarning: Warning only once for all operators, other operators may also be overridden.
Overriding a previously registered kernel for the same operator and the same dispatch key
operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926
dispatch key: ADInplaceOrView
previous kernel: no debug info
new kernel: registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926 (Triggered internally at /__w/TheRock/TheRock/external-builds/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
self.m.impl(
(APIServer pid=22068) INFO 12-11 20:20:35 [api_server.py:1351] vLLM API server version 0.11.2.dev690+g67475a6e8.d20251209
(APIServer pid=22068) INFO 12-11 20:20:35 [utils.py:253] non-default args: {'model_tag': 'cpatonn/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit', 'host': '127.0.0.1', 'model': 'cpatonn/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit', 'trust_remote_code': True, 'max_model_len': 16384, 'tensor_parallel_size': 2, 'gpu_memory_utilization': 0.98, 'max_num_seqs': 32}
(APIServer pid=22068) The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.
(APIServer pid=22068) INFO 12-11 20:20:38 [model.py:629] Resolved architecture: Qwen3NextForCausalLM
(APIServer pid=22068) INFO 12-11 20:20:38 [model.py:1755] Using max model len 16384
(APIServer pid=22068) INFO 12-11 20:20:38 [scheduler.py:228] Chunked prefill is enabled with max_num_batched_tokens=2048.
(APIServer pid=22068) INFO 12-11 20:20:38 [config.py:310] Disabling cascade attention since it is not supported for hybrid models.
(APIServer pid=22068) INFO 12-11 20:20:38 [config.py:437] Setting attention block size to 544 tokens to ensure that attention page size is >= mamba page size.
(APIServer pid=22068) INFO 12-11 20:20:38 [config.py:461] Padding mamba page size by 1.49% to ensure that mamba page size and attention page size are exactly equal.
/opt/venv/lib64/python3.13/site-packages/torch/library.py:357: UserWarning: Warning only once for all operators, other operators may also be overridden.
Overriding a previously registered kernel for the same operator and the same dispatch key
operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926
dispatch key: ADInplaceOrView
previous kernel: no debug info
new kernel: registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926 (Triggered internally at /__w/TheRock/TheRock/external-builds/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
self.m.impl(
(EngineCore_DP0 pid=22230) INFO 12-11 20:20:42 [core.py:93] Initializing a V1 LLM engine (v0.11.2.dev690+g67475a6e8.d20251209) with config: model='cpatonn/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit', speculative_config=None, tokenizer='cpatonn/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=16384, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=True, quantization=compressed-tensors, enforce_eager=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False), seed=0, served_model_name=cpatonn/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit, enable_prefix_caching=False, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.VLLM_COMPILE: 3>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'splitting_ops': ['vllm::unified_attention', 'vllm::unified_attention_with_output', 'vllm::unified_mla_attention', 'vllm::unified_mla_attention_with_output', 'vllm::mamba_mixer2', 'vllm::mamba_mixer', 'vllm::short_conv', 'vllm::linear_attention', 'vllm::plamo2_mamba_mixer', 'vllm::gdn_attention_core', 'vllm::kda_attention', 'vllm::sparse_attn_indexer'], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.FULL_AND_PIECEWISE: (2, 1)>, 'cudagraph_num_of_warmups': 1, 'cudagraph_capture_sizes': [1, 2, 4, 8, 16, 24, 32, 40, 48, 56, 64], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': True, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': 64, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False}, 'local_cache_dir': None}
(EngineCore_DP0 pid=22230) WARNING 12-11 20:20:42 [multiproc_executor.py:880] Reducing Torch parallelism from 24 threads to 1 to avoid unnecessary CPU contention. Set OMP_NUM_THREADS in the external environment to tune this value as needed.
/opt/venv/lib64/python3.13/site-packages/torch/library.py:357: UserWarning: Warning only once for all operators, other operators may also be overridden.
Overriding a previously registered kernel for the same operator and the same dispatch key
operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926
dispatch key: ADInplaceOrView
previous kernel: no debug info
new kernel: registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926 (Triggered internally at /__w/TheRock/TheRock/external-builds/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
self.m.impl(
/opt/venv/lib64/python3.13/site-packages/torch/library.py:357: UserWarning: Warning only once for all operators, other operators may also be overridden.
Overriding a previously registered kernel for the same operator and the same dispatch key
operator: flash_attn::_flash_attn_backward(Tensor dout, Tensor q, Tensor k, Tensor v, Tensor out, Tensor softmax_lse, Tensor(a6!)? dq, Tensor(a7!)? dk, Tensor(a8!)? dv, float dropout_p, float softmax_scale, bool causal, SymInt window_size_left, SymInt window_size_right, float softcap, Tensor? alibi_slopes, bool deterministic, Tensor? rng_state=None) -> Tensor
registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926
dispatch key: ADInplaceOrView
previous kernel: no debug info
new kernel: registered at /opt/venv/lib64/python3.13/site-packages/torch/_library/custom_ops.py:926 (Triggered internally at /__w/TheRock/TheRock/external-builds/pytorch/pytorch/aten/src/ATen/core/dispatch/OperatorEntry.cpp:208.)
self.m.impl(
INFO 12-11 20:20:45 [parallel_state.py:1203] world_size=2 rank=0 local_rank=0 distributed_init_method=tcp://127.0.0.1:53731 backend=nccl
INFO 12-11 20:20:45 [parallel_state.py:1203] world_size=2 rank=1 local_rank=1 distributed_init_method=tcp://127.0.0.1:53731 backend=nccl
INFO 12-11 20:20:45 [pynccl.py:111] vLLM is using nccl==2.27.3
INFO 12-11 20:20:46 [parallel_state.py:1411] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank 0
INFO 12-11 20:20:46 [parallel_state.py:1411] rank 1 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 1, EP rank 1
(Worker_TP0 pid=22312) INFO 12-11 20:20:46 [gpu_model_runner.py:3544] Starting to load model cpatonn/Qwen3-Next-80B-A3B-Instruct-AWQ-4bit...
(Worker_TP1 pid=22313) WARNING 12-11 20:20:47 [compressed_tensors.py:742] Acceleration for non-quantized schemes is not supported by Compressed Tensors. Falling back to UnquantizedLinearMethod
(Worker_TP1 pid=22313) INFO 12-11 20:20:47 [layer.py:379] Enabled separate cuda stream for MoE shared_experts
(Worker_TP1 pid=22313) INFO 12-11 20:20:47 [compressed_tensors_moe.py:189] Using CompressedTensorsWNA16MoEMethod
(Worker_TP0 pid=22312) WARNING 12-11 20:20:47 [compressed_tensors.py:742] Acceleration for non-quantized schemes is not supported by Compressed Tensors. Falling back to UnquantizedLinearMethod
(Worker_TP0 pid=22312) INFO 12-11 20:20:47 [layer.py:379] Enabled separate cuda stream for MoE shared_experts
(Worker_TP0 pid=22312) INFO 12-11 20:20:47 [compressed_tensors_moe.py:189] Using CompressedTensorsWNA16MoEMethod
(Worker_TP1 pid=22313) INFO 12-11 20:20:47 [rocm.py:320] Using Triton Attention backend on V1 engine.
(Worker_TP0 pid=22312) INFO 12-11 20:20:47 [rocm.py:320] Using Triton Attention backend on V1 engine.
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 0% Completed | 0/10 [00:00<?, ?it/s]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 10% Completed | 1/10 [00:03<00:29, 3.32s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 20% Completed | 2/10 [00:06<00:26, 3.30s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 30% Completed | 3/10 [00:10<00:23, 3.43s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 40% Completed | 4/10 [00:13<00:21, 3.52s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 50% Completed | 5/10 [00:16<00:16, 3.33s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 60% Completed | 6/10 [00:20<00:13, 3.36s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 70% Completed | 7/10 [00:23<00:10, 3.37s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 80% Completed | 8/10 [00:27<00:06, 3.38s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 100% Completed | 10/10 [00:30<00:00, 2.64s/it]
(Worker_TP0 pid=22312)
Loading safetensors checkpoint shards: 100% Completed | 10/10 [00:30<00:00, 3.07s/it]
(Worker_TP0 pid=22312)
(Worker_TP0 pid=22312) INFO 12-11 20:21:18 [default_loader.py:308] Loading weights took 30.72 seconds
(Worker_TP0 pid=22312) INFO 12-11 20:21:19 [gpu_model_runner.py:3626] Model loading took 23.5020 GiB memory and 31.787280 seconds
(Worker_TP0 pid=22312) INFO 12-11 20:21:23 [backends.py:616] Using cache directory: /home/kyuz0/.cache/vllm/torch_compile_cache/b8cc2bf5b2/rank_0_0/backbone for vLLM's torch.compile
(Worker_TP0 pid=22312) INFO 12-11 20:21:23 [backends.py:676] Dynamo bytecode transform time: 4.21 s
(Worker_TP1 pid=22313) INFO 12-11 20:21:25 [backends.py:243] Cache the graph of compile range (1, 2048) for later use
(Worker_TP0 pid=22312) INFO 12-11 20:21:25 [backends.py:243] Cache the graph of compile range (1, 2048) for later use
(Worker_TP0 pid=22312) WARNING 12-11 20:21:26 [fused_moe.py:888] Using default MoE config. Performance might be sub-optimal! Config file not found at ['/opt/venv/lib/python3.13/site-packages/vllm/model_executor/layers/fused_moe/configs/E=512,N=256,device_name=AMD-gfx1201,dtype=int4_w4a16.json']
(Worker_TP1 pid=22313) WARNING 12-11 20:21:26 [fused_moe.py:888] Using default MoE config. Performance might be sub-optimal! Config file not found at ['/opt/venv/lib/python3.13/site-packages/vllm/model_executor/layers/fused_moe/configs/E=512,N=256,device_name=AMD-gfx1201,dtype=int4_w4a16.json']
(Worker_TP0 pid=22312) INFO 12-11 20:21:31 [backends.py:260] Compiling a graph for compile range (1, 2048) takes 5.37 s
(Worker_TP0 pid=22312) INFO 12-11 20:21:31 [monitor.py:34] torch.compile takes 9.58 s in total
(Worker_TP0 pid=22312) WARNING 12-11 20:21:31 [decorators.py:509] Cannot save aot compilation to path /home/kyuz0/.cache/vllm/torch_aot_compile/b81e0a6e575f55d4e244cb6958a5419624d6ffb67d44abd043f8d86e42bb593b/rank_0_0/model, error:
(Worker_TP1 pid=22313) WARNING 12-11 20:21:31 [decorators.py:509] Cannot save aot compilation to path /home/kyuz0/.cache/vllm/torch_aot_compile/b81e0a6e575f55d4e244cb6958a5419624d6ffb67d44abd043f8d86e42bb593b/rank_1_0/model, error:
(Worker_TP0 pid=22312) INFO 12-11 20:21:33 [gpu_worker.py:364] Available KV cache memory: 7.09 GiB
(EngineCore_DP0 pid=22230) INFO 12-11 20:21:34 [kv_cache_utils.py:1287] GPU KV cache size: 154,496 tokens
(EngineCore_DP0 pid=22230) INFO 12-11 20:21:34 [kv_cache_utils.py:1292] Maximum concurrency for 16,384 tokens per request: 33.47x
(Worker_TP0 pid=22312)
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 0%| | 0/11 [00:00<?, ?it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 9%|▉ | 1/11 [00:00<00:05, 1.78it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 18%|█▊ | 2/11 [00:01<00:04, 1.80it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 27%|██▋ | 3/11 [00:01<00:04, 1.86it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 36%|███▋ | 4/11 [00:02<00:03, 1.89it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 45%|████▌ | 5/11 [00:02<00:03, 1.88it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 55%|█████▍ | 6/11 [00:03<00:02, 1.98it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 64%|██████▎ | 7/11 [00:03<00:01, 2.00it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 73%|███████▎ | 8/11 [00:04<00:01, 2.03it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 82%|████████▏ | 9/11 [00:04<00:00, 2.05it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 91%|█████████ | 10/11 [00:05<00:00, 2.07it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|██████████| 11/11 [00:05<00:00, 2.11it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|██████████| 11/11 [00:05<00:00, 2.00it/s]
(Worker_TP0 pid=22312)
Capturing CUDA graphs (decode, FULL): 0%| | 0/7 [00:00<?, ?it/s]
Capturing CUDA graphs (decode, FULL): 14%|█▍ | 1/7 [00:00<00:03, 1.85it/s]
Capturing CUDA graphs (decode, FULL): 29%|██▊ | 2/7 [00:01<00:02, 1.90it/s]
Capturing CUDA graphs (decode, FULL): 43%|████▎ | 3/7 [00:01<00:01, 2.04it/s]
Capturing CUDA graphs (decode, FULL): 57%|█████▋ | 4/7 [00:01<00:01, 2.13it/s]
Capturing CUDA graphs (decode, FULL): 71%|███████▏ | 5/7 [00:02<00:00, 2.17it/s]
Capturing CUDA graphs (decode, FULL): 86%|████████▌ | 6/7 [00:02<00:00, 2.19it/s]
Capturing CUDA graphs (decode, FULL): 100%|██████████| 7/7 [00:03<00:00, 2.14it/s]
Capturing CUDA graphs (decode, FULL): 100%|██████████| 7/7 [00:03<00:00, 2.10it/s]
(Worker_TP0 pid=22312) INFO 12-11 20:21:43 [gpu_model_runner.py:4548] Graph capturing finished in 10 secs, took 0.87 GiB
(EngineCore_DP0 pid=22230) INFO 12-11 20:21:43 [core.py:256] init engine (profile, create kv cache, warmup model) took 24.45 seconds
(APIServer pid=22068) INFO 12-11 20:21:45 [api_server.py:1099] Supported tasks: ['generate']
(APIServer pid=22068) WARNING 12-11 20:21:45 [model.py:1581] Default sampling parameters have been overridden by the model's Hugging Face generation config recommended from the model creator. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
(APIServer pid=22068) INFO 12-11 20:21:45 [serving_responses.py:197] Using default chat sampling params from model: {'temperature': 0.7, 'top_k': 20, 'top_p': 0.8}
(APIServer pid=22068) INFO 12-11 20:21:45 [serving_chat.py:133] Using default chat sampling params from model: {'temperature': 0.7, 'top_k': 20, 'top_p': 0.8}
(APIServer pid=22068) INFO 12-11 20:21:45 [serving_completion.py:73] Using default completion sampling params from model: {'temperature': 0.7, 'top_k': 20, 'top_p': 0.8}
(APIServer pid=22068) INFO 12-11 20:21:45 [serving_chat.py:133] Using default chat sampling params from model: {'temperature': 0.7, 'top_k': 20, 'top_p': 0.8}
(APIServer pid=22068) INFO 12-11 20:21:45 [api_server.py:1425] Starting vLLM API server 0 on http://127.0.0.1:8000
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:38] Available routes are:
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /openapi.json, Methods: GET, HEAD
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /docs, Methods: GET, HEAD
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: GET, HEAD
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /redoc, Methods: GET, HEAD
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /tokenize, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /detokenize, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /pause, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /resume, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /is_paused, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /metrics, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /health, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /load, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/models, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /version, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/responses, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/messages, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/completions, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/audio/transcriptions, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/audio/translations, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /ping, Methods: GET
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /ping, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /invocations, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /classify, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/embeddings, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /score, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/score, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /rerank, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v1/rerank, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /v2/rerank, Methods: POST
(APIServer pid=22068) INFO 12-11 20:21:45 [launcher.py:46] Route: /pooling, Methods: POST
(APIServer pid=22068) INFO: Started server process [22068]
(APIServer pid=22068) INFO: Waiting for application startup.
(APIServer pid=22068) INFO: Application startup complete.
(APIServer pid=22068) INFO: 127.0.0.1:52410 - "GET /v1/models HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(Worker_TP0 pid=22312) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (12) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP1 pid=22313) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (12) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP0 pid=22312) return fn(*contiguous_args, **contiguous_kwargs)
(Worker_TP1 pid=22313) return fn(*contiguous_args, **contiguous_kwargs)
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53396 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33798 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33802 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:22:16 [loggers.py:248] Engine 000: Avg prompt throughput: 2.4 tokens/s, Avg generation throughput: 18.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.4%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33850 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37948 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:22:26 [loggers.py:248] Engine 000: Avg prompt throughput: 218.4 tokens/s, Avg generation throughput: 17.8 tokens/s, Running: 12 reqs, Waiting: 0 reqs, GPU KV cache usage: 4.2%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33850 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33802 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:22:36 [loggers.py:248] Engine 000: Avg prompt throughput: 207.4 tokens/s, Avg generation throughput: 244.9 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33798 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53396 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37948 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:22:46 [loggers.py:248] Engine 000: Avg prompt throughput: 215.3 tokens/s, Avg generation throughput: 201.8 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53396 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:22:56 [loggers.py:248] Engine 000: Avg prompt throughput: 216.4 tokens/s, Avg generation throughput: 191.4 tokens/s, Running: 6 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.4%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33798 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37948 - "POST /v1/completions HTTP/1.1" 200 OK
(Worker_TP0 pid=22312) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (14) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP1 pid=22313) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (14) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP0 pid=22312) return fn(*contiguous_args, **contiguous_kwargs)
(Worker_TP1 pid=22313) return fn(*contiguous_args, **contiguous_kwargs)
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(Worker_TP0 pid=22312) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (11) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP0 pid=22312) return fn(*contiguous_args, **contiguous_kwargs)
(Worker_TP1 pid=22313) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (11) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP1 pid=22313) return fn(*contiguous_args, **contiguous_kwargs)
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33798 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53396 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(Worker_TP0 pid=22312) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (13) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP0 pid=22312) return fn(*contiguous_args, **contiguous_kwargs)
(Worker_TP1 pid=22313) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (13) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP1 pid=22313) return fn(*contiguous_args, **contiguous_kwargs)
(APIServer pid=22068) INFO: 127.0.0.1:33802 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53396 - "POST /v1/completions HTTP/1.1" 200 OK
(Worker_TP1 pid=22313) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (15) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP0 pid=22312) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (15) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP1 pid=22313) return fn(*contiguous_args, **contiguous_kwargs)
(Worker_TP0 pid=22312) return fn(*contiguous_args, **contiguous_kwargs)
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:23:06 [loggers.py:248] Engine 000: Avg prompt throughput: 379.7 tokens/s, Avg generation throughput: 174.2 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.9%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37948 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37948 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37948 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53396 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37948 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:23:16 [loggers.py:248] Engine 000: Avg prompt throughput: 452.9 tokens/s, Avg generation throughput: 186.8 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:23:26 [loggers.py:248] Engine 000: Avg prompt throughput: 174.4 tokens/s, Avg generation throughput: 197.4 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.1%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52834 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:23:36 [loggers.py:248] Engine 000: Avg prompt throughput: 287.3 tokens/s, Avg generation throughput: 211.8 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:23:46 [loggers.py:248] Engine 000: Avg prompt throughput: 333.3 tokens/s, Avg generation throughput: 253.0 tokens/s, Running: 14 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52834 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:23:56 [loggers.py:248] Engine 000: Avg prompt throughput: 131.0 tokens/s, Avg generation throughput: 273.9 tokens/s, Running: 15 reqs, Waiting: 0 reqs, GPU KV cache usage: 5.6%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:24:06 [loggers.py:248] Engine 000: Avg prompt throughput: 111.2 tokens/s, Avg generation throughput: 274.1 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:24:16 [loggers.py:248] Engine 000: Avg prompt throughput: 243.1 tokens/s, Avg generation throughput: 225.3 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54532 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33842 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54546 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54546 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:33806 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54546 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53392 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:24:26 [loggers.py:248] Engine 000: Avg prompt throughput: 131.6 tokens/s, Avg generation throughput: 242.1 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.3%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:52856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52834 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54546 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:24:36 [loggers.py:248] Engine 000: Avg prompt throughput: 186.6 tokens/s, Avg generation throughput: 219.6 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52834 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49492 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54546 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:24:46 [loggers.py:248] Engine 000: Avg prompt throughput: 46.1 tokens/s, Avg generation throughput: 149.6 tokens/s, Running: 5 reqs, Waiting: 0 reqs, GPU KV cache usage: 1.9%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:47880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54546 - "POST /v1/completions HTTP/1.1" 200 OK
(Worker_TP1 pid=22313) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (10) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP0 pid=22312) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (10) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP1 pid=22313) return fn(*contiguous_args, **contiguous_kwargs)
(Worker_TP0 pid=22312) return fn(*contiguous_args, **contiguous_kwargs)
(APIServer pid=22068) INFO: 127.0.0.1:54532 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37226 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37238 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37250 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47870 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:24:56 [loggers.py:248] Engine 000: Avg prompt throughput: 170.7 tokens/s, Avg generation throughput: 185.5 tokens/s, Running: 7 reqs, Waiting: 0 reqs, GPU KV cache usage: 2.6%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:52840 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37250 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:37250 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39590 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39600 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:25:06 [loggers.py:248] Engine 000: Avg prompt throughput: 240.5 tokens/s, Avg generation throughput: 181.1 tokens/s, Running: 10 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:52834 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39600 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39600 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39600 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46826 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39608 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:25:16 [loggers.py:248] Engine 000: Avg prompt throughput: 88.7 tokens/s, Avg generation throughput: 250.8 tokens/s, Running: 8 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.1%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:25:26 [loggers.py:248] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 143.2 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(Worker_TP0 pid=22312) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (12) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP0 pid=22312) return fn(*contiguous_args, **contiguous_kwargs)
(Worker_TP1 pid=22313) /opt/venv/lib64/python3.13/site-packages/vllm/model_executor/layers/fla/ops/utils.py:113: UserWarning: Input tensor shape suggests potential format mismatch: seq_len (12) < num_heads (16). This may indicate the inputs were passed in head-first format [B, H, T, ...] when head_first=False was specified. Please verify your input tensor format matches the expected shape [B, T, H, ...].
(Worker_TP1 pid=22313) return fn(*contiguous_args, **contiguous_kwargs)
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39164 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:25:36 [loggers.py:248] Engine 000: Avg prompt throughput: 4.8 tokens/s, Avg generation throughput: 23.8 tokens/s, Running: 2 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:39166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39172 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39188 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39190 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39204 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39248 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39274 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39288 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39188 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44732 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44740 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44748 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44750 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44740 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44740 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:25:46 [loggers.py:248] Engine 000: Avg prompt throughput: 853.4 tokens/s, Avg generation throughput: 216.3 tokens/s, Running: 22 reqs, Waiting: 0 reqs, GPU KV cache usage: 8.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:44750 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39172 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44780 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39172 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44792 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44800 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44750 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44828 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44850 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44828 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44732 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44850 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44854 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44862 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44868 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44750 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44862 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44888 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44902 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44916 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44918 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44928 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44932 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44862 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44828 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40686 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40692 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40696 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40710 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40712 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40718 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40730 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40746 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40754 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40772 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:25:56 [loggers.py:248] Engine 000: Avg prompt throughput: 830.1 tokens/s, Avg generation throughput: 399.2 tokens/s, Running: 32 reqs, Waiting: 20 reqs, GPU KV cache usage: 11.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:44918 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44916 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40792 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40796 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44902 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40812 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40820 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40824 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40836 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40844 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40872 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44854 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44932 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44748 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39274 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44862 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40692 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39248 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44740 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40876 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39172 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40892 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40896 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40730 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40746 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44828 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:26:06 [loggers.py:248] Engine 000: Avg prompt throughput: 493.8 tokens/s, Avg generation throughput: 422.4 tokens/s, Running: 31 reqs, Waiting: 30 reqs, GPU KV cache usage: 11.5%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:40772 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44888 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44800 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39204 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36454 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36464 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40792 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36472 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36476 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36488 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36498 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44780 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40836 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44916 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44928 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36508 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40872 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39188 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44854 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40718 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44932 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39274 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44748 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36524 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44750 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36536 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36538 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47072 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47082 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44732 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47098 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47108 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47124 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47126 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:26:16 [loggers.py:248] Engine 000: Avg prompt throughput: 438.2 tokens/s, Avg generation throughput: 441.6 tokens/s, Running: 32 reqs, Waiting: 48 reqs, GPU KV cache usage: 12.1%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:40856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47142 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40696 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47156 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47170 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40876 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47184 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47192 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47202 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47210 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39172 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47232 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47238 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40754 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47276 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40730 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47292 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40710 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47304 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47308 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47312 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47316 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47332 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47334 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47348 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44792 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54000 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54002 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54016 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54020 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54030 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54040 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54054 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54056 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54058 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54064 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54078 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54086 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54088 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54100 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:26:26 [loggers.py:248] Engine 000: Avg prompt throughput: 261.3 tokens/s, Avg generation throughput: 480.0 tokens/s, Running: 32 reqs, Waiting: 84 reqs, GPU KV cache usage: 12.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:54114 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44862 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39190 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39164 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54130 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54138 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54140 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54150 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54154 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54162 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39288 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40692 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44868 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44800 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36454 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36476 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36464 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44918 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54176 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54182 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54190 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54194 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54206 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44740 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54222 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40772 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44928 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44902 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36176 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36186 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36202 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36232 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36244 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36498 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:26:36 [loggers.py:248] Engine 000: Avg prompt throughput: 315.7 tokens/s, Avg generation throughput: 454.4 tokens/s, Running: 31 reqs, Waiting: 106 reqs, GPU KV cache usage: 11.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:40712 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40812 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39188 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39204 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40892 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39274 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36508 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40820 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36538 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44854 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44748 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36472 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40824 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44932 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36272 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40796 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40836 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36276 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44828 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40896 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47142 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:26:46 [loggers.py:248] Engine 000: Avg prompt throughput: 499.9 tokens/s, Avg generation throughput: 425.5 tokens/s, Running: 32 reqs, Waiting: 107 reqs, GPU KV cache usage: 12.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:40686 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47098 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44850 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47184 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44240 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44246 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44260 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40718 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44266 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44276 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44278 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40844 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44282 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44288 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47232 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44916 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47082 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44732 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44312 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44318 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44332 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47124 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44336 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44346 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44356 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44358 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44360 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44374 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44384 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44386 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40730 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39248 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36536 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40872 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57672 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57674 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57688 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57698 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57712 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57728 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:26:56 [loggers.py:248] Engine 000: Avg prompt throughput: 278.9 tokens/s, Avg generation throughput: 464.0 tokens/s, Running: 32 reqs, Waiting: 137 reqs, GPU KV cache usage: 12.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:57744 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47156 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:58032 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57756 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57766 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47334 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40876 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47332 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47312 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47348 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57778 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47304 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57794 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57808 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40746 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57824 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57866 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57882 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57892 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57898 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40710 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57910 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54002 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54016 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57924 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57930 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57944 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57954 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57964 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57970 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54054 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47238 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57984 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57990 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49822 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54000 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49854 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:27:06 [loggers.py:248] Engine 000: Avg prompt throughput: 410.1 tokens/s, Avg generation throughput: 444.8 tokens/s, Running: 32 reqs, Waiting: 163 reqs, GPU KV cache usage: 12.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:49858 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49866 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54086 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49868 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49878 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49888 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49902 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54088 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44780 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49912 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47072 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49916 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44862 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47170 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47202 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39164 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39190 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49926 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40696 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49936 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49942 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49952 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49954 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47192 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49968 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49978 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49990 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50006 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50016 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54162 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42046 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47108 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54078 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42060 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:27:16 [loggers.py:248] Engine 000: Avg prompt throughput: 461.3 tokens/s, Avg generation throughput: 441.6 tokens/s, Running: 32 reqs, Waiting: 183 reqs, GPU KV cache usage: 12.5%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:42076 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42088 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42104 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42118 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42120 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42134 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42144 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54138 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42146 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40692 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42156 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42162 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42178 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42186 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54020 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39288 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47126 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47276 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36524 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54056 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54100 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44816 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36476 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42198 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42206 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42218 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42240 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42242 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54182 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54114 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42244 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54206 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44888 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44792 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54222 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47292 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:53200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54040 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:27:26 [loggers.py:248] Engine 000: Avg prompt throughput: 463.6 tokens/s, Avg generation throughput: 422.4 tokens/s, Running: 31 reqs, Waiting: 200 reqs, GPU KV cache usage: 11.6%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:44918 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36186 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36488 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39172 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54130 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36176 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40792 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47210 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36244 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40754 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39188 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47316 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40892 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36498 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54064 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54140 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36464 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54154 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44748 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54058 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54150 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44928 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47308 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36272 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44800 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40836 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44854 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:27:36 [loggers.py:248] Engine 000: Avg prompt throughput: 682.2 tokens/s, Avg generation throughput: 409.6 tokens/s, Running: 32 reqs, Waiting: 200 reqs, GPU KV cache usage: 11.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:36276 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39274 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44750 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54194 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36454 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42224 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40820 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42226 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54176 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42236 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42246 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42248 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44880 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42252 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42260 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40896 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42274 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40856 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54030 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42276 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42282 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42292 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42308 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42318 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54190 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42324 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42336 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42338 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42342 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42350 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42354 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40712 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46226 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46240 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:27:46 [loggers.py:248] Engine 000: Avg prompt throughput: 123.6 tokens/s, Avg generation throughput: 489.6 tokens/s, Running: 31 reqs, Waiting: 231 reqs, GPU KV cache usage: 11.3%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:44762 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40718 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40772 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46268 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44266 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46280 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46296 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44276 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46308 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46324 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46332 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46346 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44868 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36508 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46360 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46370 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46384 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40812 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36264 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44740 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46400 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44288 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46402 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40824 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46414 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44814 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44902 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46416 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46424 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39204 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46440 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46452 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46468 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46484 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46500 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46504 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36232 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:46516 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44318 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42552 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36538 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:27:56 [loggers.py:248] Engine 000: Avg prompt throughput: 349.4 tokens/s, Avg generation throughput: 457.6 tokens/s, Running: 32 reqs, Waiting: 254 reqs, GPU KV cache usage: 11.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:42566 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44828 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44246 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42582 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44358 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44360 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44240 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44346 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47098 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42584 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44384 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42596 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42600 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40686 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42602 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44278 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42614 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36202 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47142 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42626 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40796 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42638 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42652 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42658 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40872 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47082 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:42670 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47374 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44260 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47388 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47398 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:28:06 [loggers.py:248] Engine 000: Avg prompt throughput: 307.0 tokens/s, Avg generation throughput: 454.4 tokens/s, Running: 32 reqs, Waiting: 270 reqs, GPU KV cache usage: 11.9%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:44932 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47402 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47404 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47408 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47412 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47414 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47430 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47440 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47446 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47458 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47462 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47466 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44282 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44332 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47482 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57766 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47486 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47496 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44850 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44356 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44916 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57698 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39248 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47312 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39254 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39214 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47348 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:28:16 [loggers.py:248] Engine 000: Avg prompt throughput: 450.5 tokens/s, Avg generation throughput: 438.4 tokens/s, Running: 31 reqs, Waiting: 279 reqs, GPU KV cache usage: 11.4%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:47304 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57808 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44336 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47232 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40844 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57852 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57838 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40746 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40730 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57794 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:36472 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44732 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57672 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57756 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44386 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44374 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44312 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50688 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50704 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50718 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57898 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50728 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39166 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50730 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50736 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50752 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50766 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50776 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:40876 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57674 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50780 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50796 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:50800 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57910 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57984 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54002 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54940 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44300 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54944 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54960 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:28:26 [loggers.py:248] Engine 000: Avg prompt throughput: 206.5 tokens/s, Avg generation throughput: 470.4 tokens/s, Running: 32 reqs, Waiting: 299 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO: 127.0.0.1:54972 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54976 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54984 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54986 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55002 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55012 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55014 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55028 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47124 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57744 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57712 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54000 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49858 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47184 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55034 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39234 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55042 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55050 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55060 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55068 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57778 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55076 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49854 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57688 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55090 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55096 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55110 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55122 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55128 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55140 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55148 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55164 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55172 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57970 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55186 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55196 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55202 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55216 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55230 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55246 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47238 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57728 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55262 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:39200 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:55272 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:57892 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:47072 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:60898 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:60914 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:60924 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:60932 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:54088 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:44862 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO: 127.0.0.1:49912 - "POST /v1/completions HTTP/1.1" 200 OK
(APIServer pid=22068) INFO 12-11 20:28:36 [loggers.py:248] Engine 000: Avg prompt throughput: 470.7 tokens/s, Avg generation throughput: 448.0 tokens/s, Running: 32 reqs, Waiting: 331 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:28:46 [loggers.py:248] Engine 000: Avg prompt throughput: 496.4 tokens/s, Avg generation throughput: 441.6 tokens/s, Running: 32 reqs, Waiting: 306 reqs, GPU KV cache usage: 11.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:28:56 [loggers.py:248] Engine 000: Avg prompt throughput: 849.8 tokens/s, Avg generation throughput: 403.2 tokens/s, Running: 32 reqs, Waiting: 279 reqs, GPU KV cache usage: 12.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:29:06 [loggers.py:248] Engine 000: Avg prompt throughput: 371.2 tokens/s, Avg generation throughput: 448.0 tokens/s, Running: 32 reqs, Waiting: 258 reqs, GPU KV cache usage: 12.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:29:16 [loggers.py:248] Engine 000: Avg prompt throughput: 324.2 tokens/s, Avg generation throughput: 454.4 tokens/s, Running: 32 reqs, Waiting: 236 reqs, GPU KV cache usage: 11.5%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:29:26 [loggers.py:248] Engine 000: Avg prompt throughput: 506.1 tokens/s, Avg generation throughput: 438.4 tokens/s, Running: 32 reqs, Waiting: 215 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:29:36 [loggers.py:248] Engine 000: Avg prompt throughput: 117.4 tokens/s, Avg generation throughput: 464.0 tokens/s, Running: 32 reqs, Waiting: 199 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:29:46 [loggers.py:248] Engine 000: Avg prompt throughput: 476.2 tokens/s, Avg generation throughput: 428.8 tokens/s, Running: 32 reqs, Waiting: 175 reqs, GPU KV cache usage: 11.9%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:29:56 [loggers.py:248] Engine 000: Avg prompt throughput: 851.7 tokens/s, Avg generation throughput: 387.2 tokens/s, Running: 32 reqs, Waiting: 144 reqs, GPU KV cache usage: 12.0%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:30:06 [loggers.py:248] Engine 000: Avg prompt throughput: 287.1 tokens/s, Avg generation throughput: 441.6 tokens/s, Running: 32 reqs, Waiting: 127 reqs, GPU KV cache usage: 11.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:30:16 [loggers.py:248] Engine 000: Avg prompt throughput: 560.8 tokens/s, Avg generation throughput: 409.6 tokens/s, Running: 31 reqs, Waiting: 102 reqs, GPU KV cache usage: 11.2%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:30:26 [loggers.py:248] Engine 000: Avg prompt throughput: 264.6 tokens/s, Avg generation throughput: 457.6 tokens/s, Running: 32 reqs, Waiting: 90 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:30:36 [loggers.py:248] Engine 000: Avg prompt throughput: 286.2 tokens/s, Avg generation throughput: 460.8 tokens/s, Running: 32 reqs, Waiting: 70 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:30:46 [loggers.py:248] Engine 000: Avg prompt throughput: 866.7 tokens/s, Avg generation throughput: 400.0 tokens/s, Running: 32 reqs, Waiting: 39 reqs, GPU KV cache usage: 11.8%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:30:56 [loggers.py:248] Engine 000: Avg prompt throughput: 340.5 tokens/s, Avg generation throughput: 448.0 tokens/s, Running: 31 reqs, Waiting: 16 reqs, GPU KV cache usage: 11.3%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:31:06 [loggers.py:248] Engine 000: Avg prompt throughput: 170.5 tokens/s, Avg generation throughput: 475.3 tokens/s, Running: 28 reqs, Waiting: 0 reqs, GPU KV cache usage: 10.1%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:31:16 [loggers.py:248] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 359.4 tokens/s, Running: 9 reqs, Waiting: 0 reqs, GPU KV cache usage: 3.7%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:31:26 [loggers.py:248] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 133.5 tokens/s, Running: 1 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.4%, Prefix cache hit rate: 0.0%
(APIServer pid=22068) INFO 12-11 20:31:27 [launcher.py:110] Shutting down FastAPI HTTP server.
(Worker_TP0 pid=22312) INFO 12-11 20:31:27 [multiproc_executor.py:709] Parent process exited, terminating worker
(Worker_TP1 pid=22313) INFO 12-11 20:31:27 [multiproc_executor.py:709] Parent process exited, terminating worker
(APIServer pid=22068) INFO: Shutting down
(APIServer pid=22068) INFO: Waiting for application shutdown.
(APIServer pid=22068) INFO: Application shutdown complete.