llama-stack-mirror/llama_stack/providers/remote
Ben Browning 48fdbf7188 fix: ollama chat completion needs unique ids
The chat completion ids generated by Ollama are not unique enough to
use with stored chat completions as they rely on only 3 numbers of
randomness to give unique values - ie `chatcmpl-373`. This causes
frequent collisions in id values of chat completions in Ollama, which
creates issues in our SQL storage of chat completions by id where it
expects ids to actually be unique.

So, this adjusts Ollama responses to use uuids as unique ids. This
does mean we're replacing the ids generated natively by Ollama. If we
don't wish to do this, we'll either need to relax the unique
constraint on our chat completions id field in the inference storage
or convince Ollama upstream to use something closer to uuid values
here.

Closes #2315

I tested by running the openai completion / chat completion
integration tests in a loop. Without this change, I regularly get
unique id collisions. With this change, I do not.

```
INFERENCE_MODEL="meta-llama/Llama-3.2-3B-Instruct" \
llama stack run llama_stack/templates/ollama/run.yaml

while true; do; \
  INFERENCE_MODEL="meta-llama/Llama-3.2-3B-Instruct" \
  pytest -s -v \
    tests/integration/inference/test_openai_completion.py \
    --stack-config=http://localhost:8321 \
    --text-model="meta-llama/Llama-3.2-3B-Instruct"; \
done
```

Signed-off-by: Ben Browning <bbrownin@redhat.com>
2025-06-02 19:07:42 -04:00
..
agents test: add unit test to ensure all config types are instantiable (#1601) 2025-03-12 22:29:58 -07:00
datasetio chore(refact): move paginate_records fn outside of datasetio (#2137) 2025-05-12 10:56:14 -07:00
eval chore: enable pyupgrade fixes (#1806) 2025-05-01 14:23:50 -07:00
inference fix: ollama chat completion needs unique ids 2025-06-02 19:07:42 -04:00
post_training fix: Pass model parameter as config name to NeMo Customizer (#2218) 2025-05-20 09:51:39 -07:00
safety feat(providers): sambanova safety provider (#2221) 2025-05-21 15:33:02 -07:00
tool_runtime fix: match mcp headers in provider data to Responses API shape (#2263) 2025-05-25 14:33:10 -07:00
vector_io feat(sqlite-vec): enable keyword search for sqlite-vec (#1439) 2025-05-21 15:24:24 -04:00
__init__.py impls -> inline, adapters -> remote (#381) 2024-11-06 14:54:05 -08:00