mirror of https://github.com/meta-llama/llama-stack.git synced 2025-12-04 18:13:44 +00:00

History

Ben Browning 8dfce2f596 feat: OpenAI Responses API (#1989 ) # What does this PR do? This provides an initial [OpenAI Responses API](https://platform.openai.com/docs/api-reference/responses) implementation. The API is not yet complete, and this is more a proof-of-concept to show how we can store responses in our key-value stores and use them to support the Responses API concepts like `previous_response_id`. ## Test Plan I've added a new `tests/integration/openai_responses/test_openai_responses.py` as part of a test-driven development for this new API. I'm only testing this locally with the remote-vllm provider for now, but it should work with any of our inference providers since the only API it requires out of the inference provider is the `openai_chat_completion` endpoint. ``` VLLM_URL="http://localhost:8000/v1" \ INFERENCE_MODEL="meta-llama/Llama-3.2-3B-Instruct" \ llama stack build --template remote-vllm --image-type venv --run ``` ``` LLAMA_STACK_CONFIG="http://localhost:8321" \ python -m pytest -v \ tests/integration/openai_responses/test_openai_responses.py \ --text-model "meta-llama/Llama-3.2-3B-Instruct" ``` --------- Signed-off-by: Ben Browning <bbrownin@redhat.com> Co-authored-by: Ashwin Bharambe <ashwin.bharambe@gmail.com>		2025-04-28 14:06:00 -07:00
..
conf	feat: OpenAI Responses API (#1989 )	2025-04-28 14:06:00 -07:00
openai_api	feat: OpenAI Responses API (#1989 )	2025-04-28 14:06:00 -07:00
test_results	test: add multi_image test (#1972 )	2025-04-17 12:51:42 -07:00
__init__.py	feat: adds test suite to verify provider's OAI compat endpoints (#1901 )	2025-04-08 21:21:38 -07:00
conftest.py	feat(verification): various improvements (#1921 )	2025-04-10 10:26:19 -07:00
generate_report.py	feat: OpenAI Responses API (#1989 )	2025-04-28 14:06:00 -07:00
openai-api-verification-run.yaml	feat: OpenAI Responses API (#1989 )	2025-04-28 14:06:00 -07:00
README.md	chore(verification): update README and reorganize generate_report.py (#1978 )	2025-04-17 10:41:22 -07:00
REPORT.md	test: add multi_image test (#1972 )	2025-04-17 12:51:42 -07:00

README.md

Llama Stack Verifications

Llama Stack Verifications provide standardized test suites to ensure API compatibility and behavior consistency across different LLM providers. These tests help verify that different models and providers implement the expected interfaces and behaviors correctly.

Overview

This framework allows you to run the same set of verification tests against different LLM providers' OpenAI-compatible endpoints (Fireworks, Together, Groq, Cerebras, etc., and OpenAI itself) to ensure they meet the expected behavior and interface standards.

Features

The verification suite currently tests the following in both streaming and non-streaming modes:

Basic chat completions
Image input capabilities
Structured JSON output formatting
Tool calling functionality

Report

The lastest report can be found at REPORT.md.

To update the report, ensure you have the API keys set,

export OPENAI_API_KEY=<your_openai_api_key>
export FIREWORKS_API_KEY=<your_fireworks_api_key>
export TOGETHER_API_KEY=<your_together_api_key>

then run

uv run --with-editable ".[dev]" python tests/verifications/generate_report.py --run-tests

Running Tests

To run the verification tests, use pytest with the following parameters:

cd llama-stack
pytest tests/verifications/openai_api --provider=<provider-name>

Example:

# Run all tests
pytest tests/verifications/openai_api --provider=together

# Only run tests with Llama 4 models
pytest tests/verifications/openai_api --provider=together -k 'Llama-4'

Parameters

--provider: The provider name (openai, fireworks, together, groq, cerebras, etc.)
--base-url: The base URL for the provider's API (optional - defaults to the standard URL for the specified provider)
--api-key: Your API key for the provider (optional - defaults to the standard API_KEY name for the specified provider)

Supported Providers

The verification suite supports any provider with an OpenAI compatible endpoint.

See tests/verifications/conf/ for the list of supported providers.

To run on a new provider, simply add a new yaml file to the conf/ directory with the provider config. See tests/verifications/conf/together.yaml for an example.

Adding New Test Cases

To add new test cases, create appropriate JSON files in the openai_api/fixtures/test_cases/ directory following the existing patterns.

Structure

__init__.py - Marks the directory as a Python package
conf/ - Provider-specific configuration files
openai_api/ - Tests specific to OpenAI-compatible APIs
- fixtures/ - Test fixtures and utilities
  - fixtures.py - Provider-specific fixtures
  - load.py - Utilities for loading test cases
  - test_cases/ - JSON test case definitions
- test_chat_completion.py - Tests for chat completion APIs