llama-stack-mirror

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-12-03 18:00:36 +00:00

History

Matthew Farrellee d23607483f chore: update the groq inference impl to use openai-python for openai-compat functions (#3348 ) # What does this PR do? update Groq inference provider to use OpenAIMixin for openai-compat endpoints changes on api.groq.com - - json_schema is now supported for specific models, see https://console.groq.com/docs/structured-outputs#supported-models - response_format with streaming is now supported for models that support response_format - groq no longer returns a 400 error if tools are provided and tool_choice is not "required" ## Test Plan ``` $ GROQ_API_KEY=... uv run llama stack build --image-type venv --providers inference=remote::groq --run ... $ LLAMA_STACK_CONFIG=http://localhost:8321 uv run --group test pytest -v -ra --text-model groq/llama-3.3-70b-versatile tests/integration/inference/test_openai_completion.py -k 'not store' ... SKIPPED [3] tests/integration/inference/test_openai_completion.py:44: Model groq/llama-3.3-70b-versatile hosted by remote::groq doesn't support OpenAI completions. SKIPPED [3] tests/integration/inference/test_openai_completion.py:94: Model groq/llama-3.3-70b-versatile hosted by remote::groq doesn't support vllm extra_body parameters. SKIPPED [4] tests/integration/inference/test_openai_completion.py:73: Model groq/llama-3.3-70b-versatile hosted by remote::groq doesn't support n param. SKIPPED [1] tests/integration/inference/test_openai_completion.py💯 Model groq/llama-3.3-70b-versatile hosted by remote::groq doesn't support chat completion calls with base64 encoded files. ======================= 8 passed, 11 skipped, 8 deselected, 2 warnings in 5.13s ======================== ``` --------- Co-authored-by: raghotham <rsm@meta.com>		2025-09-06 15:36:27 -07:00
..
anthropic	feat(starter)!: simplify starter distro; litellm model registry changes (#2916 )	2025-07-25 15:02:04 -07:00
bedrock	feat(starter)!: simplify starter distro; litellm model registry changes (#2916 )	2025-07-25 15:02:04 -07:00
cerebras	feat(starter)!: simplify starter distro; litellm model registry changes (#2916 )	2025-07-25 15:02:04 -07:00
databricks	feat(starter)!: simplify starter distro; litellm model registry changes (#2916 )	2025-07-25 15:02:04 -07:00
fireworks	refactor(logging): rename llama_stack logger categories (#3065 )	2025-08-21 17:31:04 -07:00
gemini	chore: update the gemini inference impl to use openai-python for openai-compat functions (#3351 )	2025-09-06 12:22:20 -07:00
groq	chore: update the groq inference impl to use openai-python for openai-compat functions (#3348 )	2025-09-06 15:36:27 -07:00
llama_openai_compat	chore: indicate to mypy that InferenceProvider.rerank is concrete (#3238 )	2025-08-22 12:02:13 -07:00
nvidia	docs: add VLM NIM example (#3277 )	2025-08-29 16:23:52 -07:00
ollama	feat(tests): auto-merge all model list responses and unify recordings (#3320 )	2025-09-03 11:33:03 -07:00
openai	refactor(logging): rename llama_stack logger categories (#3065 )	2025-08-21 17:31:04 -07:00
passthrough	chore(rename): move llama_stack.distribution to llama_stack.core (#2975 )	2025-07-30 23:30:53 -07:00
runpod	ci: test safety with starter (#2628 )	2025-07-09 16:53:50 +02:00
sambanova	chore: update the sambanova inference impl to use openai-python for openai-compat functions (#3345 )	2025-09-06 12:25:13 -07:00
tgi	refactor(logging): rename llama_stack logger categories (#3065 )	2025-08-21 17:31:04 -07:00
together	refactor(logging): rename llama_stack logger categories (#3065 )	2025-08-21 17:31:04 -07:00
vertexai	feat: Add Google Vertex AI inference provider support (#2841 )	2025-08-11 08:22:04 -04:00
vllm	chore: indicate to mypy that InferenceProvider.batch_completion/batch_chat_completion is concrete (#3239 )	2025-08-22 14:17:30 -07:00
watsonx	chore(python-deps): replace ibm_watson_machine_learning with ibm_watsonx_ai (#3302 )	2025-09-03 11:33:35 +02:00
__init__.py	`impls` -> `inline`, `adapters` -> `remote` (#381 )	2024-11-06 14:54:05 -08:00