llama-stack-mirror

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-12-03 09:53:45 +00:00

History

Francisco Arceo 82f13fe83e feat: Add ChunkMetadata to Chunk (#2497 ) # What does this PR do? Adding `ChunkMetadata` so we can properly delete embeddings later. More specifically, this PR refactors and extends the chunk metadata handling in the vector database and introduces a distinction between metadata used for model context and backend-only metadata required for chunk management, storage, and retrieval. It also improves chunk ID generation and propagation throughout the stack, enhances test coverage, and adds new utility modules. ```python class ChunkMetadata(BaseModel): """ `ChunkMetadata` is backend metadata for a `Chunk` that is used to store additional information about the chunk that will NOT be inserted into the context during inference, but is required for backend functionality. Use `metadata` in `Chunk` for metadata that will be used during inference. """ document_id: str \| None = None chunk_id: str \| None = None source: str \| None = None created_timestamp: int \| None = None updated_timestamp: int \| None = None chunk_window: str \| None = None chunk_tokenizer: str \| None = None chunk_embedding_model: str \| None = None chunk_embedding_dimension: int \| None = None content_token_count: int \| None = None metadata_token_count: int \| None = None ``` Eventually we can migrate the document_id out of the `metadata` field. I've introduced the changes so that `ChunkMetadata` is backwards compatible with `metadata`. <!-- If resolving an issue, uncomment and update the line below --> Closes https://github.com/meta-llama/llama-stack/issues/2501 ## Test Plan Added unit tests --------- Signed-off-by: Francisco Javier Arceo <farceo@redhat.com>		2025-06-25 15:55:23 -04:00
..
agents	feat: drop python 3.10 support (#2469 )	2025-06-19 12:07:14 +05:30
batch_inference	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
benchmarks	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
common	feat: Add url field to PaginatedResponse and populate it using route … (#2419 )	2025-06-16 11:19:48 +02:00
datasetio	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
datasets	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
eval	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
files	feat: openai files api (#2321 )	2025-06-02 11:45:53 -07:00
inference	feat: drop python 3.10 support (#2469 )	2025-06-19 12:07:14 +05:30
inspect	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
models	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
post_training	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
providers	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
safety	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
scoring	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
scoring_functions	feat: drop python 3.10 support (#2469 )	2025-06-19 12:07:14 +05:30
shields	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
synthetic_data_generation	chore: enable pyupgrade fixes (#1806 )	2025-05-01 14:23:50 -07:00
telemetry	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
tools	chore: bump python supported version to 3.12 (#2475 )	2025-06-24 09:22:04 +05:30
vector_dbs	chore: more API validators (#2165 )	2025-05-15 11:22:51 -07:00
vector_io	feat: Add ChunkMetadata to Chunk (#2497 )	2025-06-25 15:55:23 -04:00
__init__.py	API Updates (#73 )	2024-09-17 19:51:35 -07:00
datatypes.py	chore: enable pyupgrade fixes (#1806 )	2025-05-01 14:23:50 -07:00
resource.py	feat: drop python 3.10 support (#2469 )	2025-06-19 12:07:14 +05:30
version.py	llama-stack version alpha -> v1	2025-01-15 05:58:09 -08:00