chore: update the vLLM inference impl to use OpenAIMixin for openai-compat functions

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-10-16 06:53:47 +00:00

inference recordings from Qwen3-0.6B and vLLM 0.8.3 -
```
docker run --gpus all -v ~/.cache/huggingface:/root/.cache/huggingface -p 8000:8000 --ipc=host \
    vllm/vllm-openai:latest \
    --model Qwen/Qwen3-0.6B --enable-auto-tool-choice --tool-call-parser hermes
```

test with -

```
./scripts/integration-tests.sh --stack-config server:ci-tests --setup vllm --subdirs inference
```

This commit is contained in:

Matthew Farrellee

2025-09-10 10:10:10 -04:00

parent c86e45496e

commit c2a9c65fff

33 changed files with 51813 additions and 203 deletions

5747

tests/integration/recordings/responses/a5163d8d2d01.json Normal file

View file

File diff suppressed because it is too large Load diff

Rows
Columns

chore: update the vLLM inference impl to use OpenAIMixin for openai-compat functions

5747 tests/integration/recordings/responses/a5163d8d2d01.json Normal file View file

5747

tests/integration/recordings/responses/a5163d8d2d01.json Normal file

View file