Summary
The Decisions API reports shared state tokens once per question in a batch. With eight questions, extending the state by 1,000 tokens increases usage.input_tokens by 8,000 rather than 1,000. Identical requests to Jev increase by 1,000 regardless of question count.
Reproduced twice with direct HTTP requests and synthetic text, without an SDK, retries, an application adapter, or a response cache. All requests succeeded with HTTP 200.
Expected behavior
I expected the shared state to contribute once per request, plus the tokens for each question. If repeating the state in billable usage for every question is intentional, please explicitly document that accounting rule and the resulting cost of batching. Otherwise, could you investigate the input-token accounting?
The pricing documentation says billing uses the response usage.input_tokens. I have verified the response counts, but have not independently reconciled them against an invoice.
Actual behavior
| State | Questions | Perplexity pplx-decider-v1.1-27b |
Jev jev-1.13.0 |
|---|---|---|---|
| Short | 1 | 99 | 288 |
| Short | 8 | 792 | 393 |
| Long | 1 | 1,099 | 1,288 |
| Long | 8 | 8,792 | 1,393 |
Subtracting short-state usage from long-state usage controls for question overhead and tokenizer differences:
- Perplexity: 1,000 extra tokens with one question; 8,000 with eight questions.
- Jev: 1,000 extra tokens with one question; 1,000 with eight questions.
Perplexity also multiplies both full-request counts exactly by eight: 99 → 792 and 1,099 → 8,792. The state occurs only once in each JSON request body.
Minimal reproduction
Save as reproduce.py, set PERPLEXITY_API_KEY and TYPESAFE_API_KEY in the environment, and run python3 reproduce.py. No packages are needed. It makes four paid requests to each provider and prints usage and request IDs. Jev is a comparison control; the first four calls reproduce the Perplexity behavior independently.
import json
import os
from urllib.request import Request, urlopen
providers = [
("Perplexity", "https://api.perplexity.ai/v1/decisions",
"pplx-decider-v1.1-27b", "PERPLEXITY_API_KEY"),
("Jev", "https://api.typesafe.ai/v1/systemone",
"jev-1.13.0", "TYPESAFE_API_KEY"),
]
short = "The warehouse shipment arrived on Tuesday. "
states = {
"short": {"document": short},
"long": {"document": short +
"The inventory contains numbered blue boxes and green labels. " * 100},
}
for provider, endpoint, model, key in providers:
counts = {}
for length, state in states.items():
for n in (1, 8):
body = {
"model": model,
"state": state,
"questions": {
f"q{i}": {
"type": "noul",
"instructions": "Does the document say the shipment arrived?",
}
for i in range(n)
},
}
request = Request(
endpoint, data=json.dumps(body).encode(),
headers={"Authorization": "Bearer " + os.environ[key],
"Content-Type": "application/json"},
)
with urlopen(request, timeout=45) as response:
payload = json.load(response)
request_id = response.headers.get("x-request-id")
counts[length, n] = payload["usage"]["input_tokens"]
print(provider, length, n, payload["usage"], request_id)
print(provider, "extra state tokens:",
"one question =", counts["long", 1] - counts["short", 1],
"eight questions =", counts["long", 8] - counts["short", 8])
Request details
- Endpoint:
POST https://api.perplexity.ai/v1/decisions - Model:
pplx-decider-v1.1-27b - Language: Python 3.13.13; standard-library HTTP client
- Approximate time: October 8, 2026, 17:56 UTC (19:56 Africa/Johannesburg)
- Consistency: identical counts in two independent executions
- Content: synthetic warehouse text only; no images or private data
Perplexity request IDs from the second execution:
| State | Questions | Request ID |
|---|---|---|
| Short | 1 | 682bfd05-10aa-4ec2-9736-c6be3d577a73 |
| Short | 8 | f9017169-e57a-4f2f-a7ad-6eab52e22b31 |
| Long | 1 | 4b30c5e9-f6b3-455c-b8a6-4f3863c85d71 |
| Long | 8 | bd8553b7-236e-47d7-a993-cf9d92868f11 |
The long-state, eight-question response returned this usage excerpt (HTTP 200):
{"usage": {"input_tokens": 8792, "output_tokens": 8}}
Impact
Batching many questions against a shared document multiplies the state component of reported billable usage. This materially changes cost estimates compared with accounting that amortizes shared state across questions. Could you confirm whether this is intended billing behavior or a usage-reporting bug?