llama-stack-mirror

mirror of https://github.com/meta-llama/llama-stack.git synced 2025-06-27 18:50:41 +00:00

History

Ben Browning 6d20b720b8 feat: Propagate W3C trace context headers from clients (#2153 ) # What does this PR do? This extracts the W3C trace context headers (traceparent and tracestate) from incoming requests, stuffs them as attributes on the spans we create, and uses them within the tracing provider implementation to actually wrap our spans in the proper context. What this means in practice is that when a client (such as an OpenAI client) is instrumented to create these traces, we'll continue that distributed trace within Llama Stack as opposed to creating our own root span that breaks the distributed trace between client and server. It's slightly awkward to do this in Llama Stack because our Tracing API knows nothing about opentelemetry, W3C trace headers, etc - that's only knowledge the specific provider implementation has. So, that's why the trace headers get extracted by in the server code but not actually used until the provider implementation to form the proper context. This also centralizes how we were adding the `__root__` and `__root_span__` attributes, as those two were being added in different parts of the code instead of from a single place. Closes #2097 ## Test Plan This was tested manually using the helpful scripts from #2097. I verified that Llama Stack properly joined the client's span when the client was instrumented for distributed tracing, and that Llama Stack properly started its own root span when the incoming request was not part of an existing trace. Here's an example of the joined spans: ![Screenshot 2025-05-13 at 8 46 09 AM](https://github.com/user-attachments/assets/dbefda28-9faa-4339-a08d-1441efefc149) Signed-off-by: Ben Browning <bbrownin@redhat.com>		2025-05-19 18:56:54 -07:00
..
apis	feat: introduce APIs for retrieving chat completion requests (#2145 )	2025-05-18 21:43:19 -07:00
cli	fix: Pass external_config_dir to BuildConfig (#2190 )	2025-05-19 14:01:28 +02:00
distribution	feat: Propagate W3C trace context headers from clients (#2153 )	2025-05-19 18:56:54 -07:00
models	fix: llama4 tool use prompt fix (#2103 )	2025-05-06 22:18:31 -07:00
providers	feat: Propagate W3C trace context headers from clients (#2153 )	2025-05-19 18:56:54 -07:00
strong_typing	chore: enable pyupgrade fixes (#1806 )	2025-05-01 14:23:50 -07:00
templates	feat: add huggingface post_training impl (#2132 )	2025-05-16 14:41:28 -07:00
ui	feat: Adding dark mode, cleaning the UI a small bit, adding a link to the API documentation, and linting the code. (#2182 )	2025-05-16 10:48:26 -07:00
__init__.py	export LibraryClient	2024-12-13 12:08:00 -08:00
env.py	refactor(test): move tools, evals, datasetio, scoring and post training tests (#1401 )	2025-03-04 14:53:47 -08:00
log.py	chore: enable pyupgrade fixes (#1806 )	2025-05-01 14:23:50 -07:00
schema_utils.py	chore: enable pyupgrade fixes (#1806 )	2025-05-01 14:23:50 -07:00