Decisions API: shared state counted per question in usage.input_tokens

Summary

The Decisions API reports shared state tokens once per question in a batch. With eight questions, extending the state by 1,000 tokens increases usage.input_tokens by 8,000 rather than 1,000. Identical requests to Jev increase by 1,000 regardless of question count.

Reproduced twice with direct HTTP requests and synthetic text, without an SDK, retries, an application adapter, or a response cache. All requests succeeded with HTTP 200.

Expected behavior

I expected the shared state to contribute once per request, plus the tokens for each question. If repeating the state in billable usage for every question is intentional, please explicitly document that accounting rule and the resulting cost of batching. Otherwise, could you investigate the input-token accounting?

The pricing documentation says billing uses the response usage.input_tokens. I have verified the response counts, but have not independently reconciled them against an invoice.

Actual behavior

State Questions Perplexity pplx-decider-v1.1-27b Jev jev-1.13.0
Short 1 99 288
Short 8 792 393
Long 1 1,099 1,288
Long 8 8,792 1,393

Subtracting short-state usage from long-state usage controls for question overhead and tokenizer differences:

  • Perplexity: 1,000 extra tokens with one question; 8,000 with eight questions.
  • Jev: 1,000 extra tokens with one question; 1,000 with eight questions.

Perplexity also multiplies both full-request counts exactly by eight: 99 → 792 and 1,099 → 8,792. The state occurs only once in each JSON request body.

Minimal reproduction

Save as reproduce.py, set PERPLEXITY_API_KEY and TYPESAFE_API_KEY in the environment, and run python3 reproduce.py. No packages are needed. It makes four paid requests to each provider and prints usage and request IDs. Jev is a comparison control; the first four calls reproduce the Perplexity behavior independently.

import json
import os
from urllib.request import Request, urlopen

providers = [
    ("Perplexity", "https://api.perplexity.ai/v1/decisions",
     "pplx-decider-v1.1-27b", "PERPLEXITY_API_KEY"),
    ("Jev", "https://api.typesafe.ai/v1/systemone",
     "jev-1.13.0", "TYPESAFE_API_KEY"),
]
short = "The warehouse shipment arrived on Tuesday. "
states = {
    "short": {"document": short},
    "long": {"document": short +
             "The inventory contains numbered blue boxes and green labels. " * 100},
}

for provider, endpoint, model, key in providers:
    counts = {}
    for length, state in states.items():
        for n in (1, 8):
            body = {
                "model": model,
                "state": state,
                "questions": {
                    f"q{i}": {
                        "type": "noul",
                        "instructions": "Does the document say the shipment arrived?",
                    }
                    for i in range(n)
                },
            }
            request = Request(
                endpoint, data=json.dumps(body).encode(),
                headers={"Authorization": "Bearer " + os.environ[key],
                         "Content-Type": "application/json"},
            )
            with urlopen(request, timeout=45) as response:
                payload = json.load(response)
                request_id = response.headers.get("x-request-id")
            counts[length, n] = payload["usage"]["input_tokens"]
            print(provider, length, n, payload["usage"], request_id)
    print(provider, "extra state tokens:",
          "one question =", counts["long", 1] - counts["short", 1],
          "eight questions =", counts["long", 8] - counts["short", 8])

Request details

  • Endpoint: POST https://api.perplexity.ai/v1/decisions
  • Model: pplx-decider-v1.1-27b
  • Language: Python 3.13.13; standard-library HTTP client
  • Approximate time: October 8, 2026, 17:56 UTC (19:56 Africa/Johannesburg)
  • Consistency: identical counts in two independent executions
  • Content: synthetic warehouse text only; no images or private data

Perplexity request IDs from the second execution:

State Questions Request ID
Short 1 682bfd05-10aa-4ec2-9736-c6be3d577a73
Short 8 f9017169-e57a-4f2f-a7ad-6eab52e22b31
Long 1 4b30c5e9-f6b3-455c-b8a6-4f3863c85d71
Long 8 bd8553b7-236e-47d7-a993-cf9d92868f11

The long-state, eight-question response returned this usage excerpt (HTTP 200):

{"usage": {"input_tokens": 8792, "output_tokens": 8}}

Impact

Batching many questions against a shared document multiplies the state component of reported billable usage. This materially changes cost estimates compared with accounting that amortizes shared state across questions. Could you confirm whether this is intended billing behavior or a usage-reporting bug?

Hello @Urs_de_Swardt,

Thank you for the report. I have sent it over to Engineering for review. I will confirm if this is the intended behavior.

Andrew

I’ve confirmed with engineering that the model is working as expected.

Each question runs its own inference with the full state, so the state is counted once per question in usage. input_tokens. Billing uses that same value.

I’ve passed your request to engineering: count the shared state once per request as a feature request.

Thank you,

Andrew

The billing count reflects how each sub-request is processed rather than bugged reporting. When batching questions against shared state, the engine treats each question as a separate context evaluation, so the full document tokens are counted for each prompt. If you need to amortize input tokens, sending a single structured prompt that asks for all evaluations at once avoids the per-question state multiplication.