llama-stack-mirror

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-08-15 14:08:00 +00:00

History

Matthew Farrellee 8e678912ec feat: add batches API with OpenAI compatibility Add complete batches API implementation with protocol, providers, and tests: Core Infrastructure: - Add batches API protocol using OpenAI Batch types directly - Add Api.batches enum value and protocol mapping in resolver - Add OpenAI "batch" file purpose support - Include proper error handling (ConflictError, ResourceNotFoundError) Reference Provider: - Add ReferenceBatchesImpl with full CRUD operations (create, retrieve, cancel, list) - Implement background batch processing with configurable concurrency - Add SQLite KVStore backend for persistence - Support /v1/chat/completions endpoint with request validation Comprehensive Test Suite: - Add unit tests for provider implementation with validation - Add integration tests for end-to-end batch processing workflows - Add error handling tests for validation, malformed inputs, and edge cases Configuration: - Add max_concurrent_batches and max_concurrent_requests_per_batch options - Add provider documentation with sample configurations Test with - ``` $ uv run llama stack build --image-type venv --providers inference=YOU_PICK,files=inline::localfs,batches=inline::reference --run & $ LLAMA_STACK_CONFIG=http://localhost:8321 uv run pytest tests/unit/providers/batches tests/integration/batches --text-model YOU_PICK ```		2025-08-08 08:08:08 -04:00
..
agent	fix: remove @pytest.mark.asyncio from test_get_raw_document_text.py (#2840 )	2025-07-21 09:14:34 -07:00
agents	chore(rename): move llama_stack.distribution to llama_stack.core (#2975 )	2025-07-30 23:30:53 -07:00
batches	feat: add batches API with OpenAI compatibility	2025-08-08 08:08:08 -04:00
inference	feat: Add clear error message when API key is missing (#2992 )	2025-07-31 16:33:16 -04:00
nvidia	chore(rename): move llama_stack.distribution to llama_stack.core (#2975 )	2025-07-30 23:30:53 -07:00
utils	fix(openai-compat): restrict developer/assistant/system/tool messages to text-only content (#2932 )	2025-07-28 10:36:34 -07:00
vector_io	feat: Implement hybrid search in Milvus (#2644 )	2025-08-07 09:42:03 +02:00
test_configs.py	chore(rename): move llama_stack.distribution to llama_stack.core (#2975 )	2025-07-30 23:30:53 -07:00