llama-stack-mirror

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-12-03 09:53:45 +00:00

History

Ben Browning 941f505eb0 feat: File search tool for Responses API (#2426 ) # What does this PR do? This is an initial working prototype of wiring up the `file_search` builtin tool for the Responses API to our existing rag knowledge search tool. This is me seeing what I could pull together on top of the bits we already have merged. This may not be the ideal way to implement this, and things like how I shuffle the vector store ids from the original response API tool request to the actual tool execution feel a bit hacky (grep for `tool_kwargs["vector_db_ids"]` in `_execute_tool_call` to see what I mean). ## Test Plan I stubbed in some new tests to exercise this using text and pdf documents. Note that this is currently under tests/verification only because it sometimes flakes with tool calling of the small Llama-3.2-3B model we run in CI (and that I use as an example below). We'd want to make the test a bit more robust in some way if we moved this over to tests/integration and ran it in CI. ### OpenAI SaaS (to verify test correctness) ``` pytest -sv tests/verifications/openai_api/test_responses.py \ -k 'file_search' \ --base-url=https://api.openai.com/v1 \ --model=gpt-4o ``` ### Fireworks with faiss vector store ``` llama stack run llama_stack/templates/fireworks/run.yaml pytest -sv tests/verifications/openai_api/test_responses.py \ -k 'file_search' \ --base-url=http://localhost:8321/v1/openai/v1 \ --model=meta-llama/Llama-3.3-70B-Instruct ``` ### Ollama with faiss vector store This sometimes flakes on Ollama because the quantized small model doesn't always choose to call the tool to answer the user's question. But, it often works. ``` ollama run llama3.2:3b INFERENCE_MODEL="meta-llama/Llama-3.2-3B-Instruct" \ llama stack run ./llama_stack/templates/ollama/run.yaml \ --image-type venv \ --env OLLAMA_URL="http://0.0.0.0:11434" pytest -sv tests/verifications/openai_api/test_responses.py \ -k'file_search' \ --base-url=http://localhost:8321/v1/openai/v1 \ --model=meta-llama/Llama-3.2-3B-Instruct ``` ### OpenAI provider with sqlite-vec vector store ``` llama stack run ./llama_stack/templates/starter/run.yaml --image-type venv pytest -sv tests/verifications/openai_api/test_responses.py \ -k 'file_search' \ --base-url=http://localhost:8321/v1/openai/v1 \ --model=openai/gpt-4o-mini ``` ### Ensure existing vector store integration tests still pass ``` ollama run llama3.2:3b INFERENCE_MODEL="meta-llama/Llama-3.2-3B-Instruct" \ llama stack run ./llama_stack/templates/ollama/run.yaml \ --image-type venv \ --env OLLAMA_URL="http://0.0.0.0:11434" LLAMA_STACK_CONFIG=http://localhost:8321 \ pytest -sv tests/integration/vector_io \ --text-model "meta-llama/Llama-3.2-3B-Instruct" \ --embedding-model=all-MiniLM-L6-v2 ``` --------- Signed-off-by: Ben Browning <bbrownin@redhat.com>		2025-06-13 14:32:48 -04:00
..
access_control	feat: fine grained access control policy (#2264 )	2025-06-03 14:51:12 -07:00
routers	feat: File search tool for Responses API (#2426 )	2025-06-13 14:32:48 -04:00
routing_tables	feat: fine grained access control policy (#2264 )	2025-06-03 14:51:12 -07:00
server	feat(auth): allow token to be provided for use against jwks endpoint (#2394 )	2025-06-13 10:13:41 +02:00
store	fix(tools): do not index tools, only index toolgroups (#2261 )	2025-05-25 13:27:52 -07:00
ui	chore: more mypy fixes (#2029 )	2025-05-06 09:52:31 -07:00
utils	refactor: remove container from list of run image types (#2178 )	2025-06-02 09:57:55 +02:00
__init__.py	API Updates (#73 )	2024-09-17 19:51:35 -07:00
build.py	feat: add deps dynamically based on metastore config (#2405 )	2025-06-05 14:07:25 -07:00
build_conda_env.sh	chore: remove straggler references to llama-models (#1345 )	2025-03-01 14:26:03 -08:00
build_container.sh	feat: refactor external providers dir (#2049 )	2025-05-15 20:17:03 +02:00
build_venv.sh	chore: remove straggler references to llama-models (#1345 )	2025-03-01 14:26:03 -08:00
client.py	chore: make cprint write to stderr (#2250 )	2025-05-24 23:39:57 -07:00
common.sh	feat(pre-commit): enhance pre-commit hooks with additional checks (#2014 )	2025-04-30 11:35:49 -07:00
configure.py	feat: refactor external providers dir (#2049 )	2025-05-15 20:17:03 +02:00
datatypes.py	feat: fine grained access control policy (#2264 )	2025-06-03 14:51:12 -07:00
distribution.py	ci: fix external provider test (#2438 )	2025-06-12 16:14:32 +02:00
inspect.py	chore: use starlette built-in Route class (#2267 )	2025-05-28 09:53:33 -07:00
library_client.py	refactor: unify stream and non-stream impls for responses (#2388 )	2025-06-05 17:48:09 +02:00
providers.py	fix: catch TimeoutError in place of asyncio.TimeoutError (#2131 )	2025-05-12 11:49:59 +02:00
request_headers.py	feat: fine grained access control policy (#2264 )	2025-06-03 14:51:12 -07:00
resolver.py	feat: OpenAIVectorIOMixin for vector_stores common logic (#2427 )	2025-06-11 15:40:57 -07:00
stack.py	feat: fine grained access control policy (#2264 )	2025-06-03 14:51:12 -07:00
start_stack.sh	refactor: remove container from list of run image types (#2178 )	2025-06-02 09:57:55 +02:00