feat(tests): auto-merge all model list responses and unify recordings (#3320)

One needed to specify record-replay related environment variables for running integration tests. We could not use defaults because integration tests could be run against Ollama instances which could be running different models. For example, text vs vision tests needed separate instances of Ollama because a single instance typically cannot serve both of these models if you assume the standard CI worker configuration on Github. As a result, `client.list()` as returned by the Ollama client would be different between these runs and we'd end up overwriting responses. This PR "solves" it by adding a small amount of complexity -- we store model list responses specially, keyed by the hashes of the models they return. At replay time, we merge all of them and pretend that we have the union of all models available. ## Test Plan Re-recorded all the tests using `scripts/integration-tests.sh --inference-mode record`, including the vision tests.
2025-10-15 06:37:58 +00:00 · 2025-09-03 11:33:03 -07:00 · 2025-09-03 11:33:03 -07:00 · c3d3a0b833
commit c3d3a0b833
parent d948e63340
122 changed files with 17460 additions and 16129 deletions
--- a/tests/integration/recordings/responses/9e7a83d3d596.json
+++ b/tests/integration/recordings/responses/9e7a83d3d596.json
@ -15,23 +15,23 @@
    "body": {
      "__type__": "openai.types.completion.Completion",
      "__data__": {
-        "id": "cmpl-43",
+        "id": "cmpl-775",
        "choices": [
          {
            "finish_reason": "stop",
            "index": 0,
            "logprobs": null,
-            "text": "Blue.\n\nMy response is based on the traditional English rhyme that pairs the colors of roses (red) with violets in a poetic and somewhat whimsical way. This specific version of the rhyme goes like this:\n\n\"Roses are red,\nViolets are blue,\nSugar is sweet,\nAnd so are you.\"\n\nIn modern times, variations of this rhyme can deviate from the original \"blue\" for violets, but in my complete sentence as requested, sticking with a widely recognized completion adds an air of timelessness and familiarity to the phrase."
+            "text": "Blue.\n\nMy response is based on the traditional rhyme \"Roses are Red, Violets are Blue,\" which is a well-known poem or phrase often used as a greeting or way to express affection. The exact wording may vary slightly depending on the source, but the general meaning remains the same: violets are typically depicted as blue-colored flowers in this rhyme."
          }
        ],
-        "created": 1754348148,
+        "created": 1756921025,
        "model": "llama3.2:3b-instruct-fp16",
        "object": "text_completion",
        "system_fingerprint": "fp_ollama",
        "usage": {
-          "completion_tokens": 113,
+          "completion_tokens": 75,
          "prompt_tokens": 50,
-          "total_tokens": 163,
+          "total_tokens": 125,
          "completion_tokens_details": null,
          "prompt_tokens_details": null
        }