# Decisions API: shared state counted per question in usage.input\_tokens

**URL:** <https://community.perplexity.ai/t/decisions-api-shared-state-counted-per-question-in-usage-input-tokens/6316>\
**Category:** Bug Reports\
**Tags:** api\
**Created:** [October 9, 2026, 10:39pm UTC](https://community.perplexity.ai/t/decisions-api-shared-state-counted-per-question-in-usage-input-tokens/6316 "2026-10-09T22:39:59Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Urs\_de\_Swardt](https://sea1.discourse-cdn.com/flex001/user_avatar/community.perplexity.ai/urs_de_swardt/32/3668_2.png) [@Urs\_de\_Swardt](https://community.perplexity.ai/u/Urs_de_Swardt)\
**Post date:** [October 9, 2026, 10:39pm UTC](https://community.perplexity.ai/t/decisions-api-shared-state-counted-per-question-in-usage-input-tokens/6316/1 "2026-10-09T22:39:59Z")

</div>

## Summary

The Decisions API reports shared `state` tokens once per question in a batch. With eight questions, extending the state by 1,000 tokens increases `usage.input_tokens` by 8,000 rather than 1,000. Identical requests to Jev increase by 1,000 regardless of question count.

Reproduced twice with direct HTTP requests and synthetic text, without an SDK, retries, an application adapter, or a response cache. All requests succeeded with HTTP 200.

## Expected behavior

I expected the shared state to contribute once per request, plus the tokens for each question. If repeating the state in billable usage for every question is intentional, please explicitly document that accounting rule and the resulting cost of batching. Otherwise, could you investigate the input-token accounting?

The [pricing documentation](https://docs.perplexity.ai/docs/decisions/quickstart#pricing) says billing uses the response `usage.input_tokens`. I have verified the response counts, but have not independently reconciled them against an invoice.

## Actual behavior

| State | Questions | Perplexity `pplx-decider-v1.1-27b` | Jev `jev-1.13.0` |
| --- | --- | --- | --- |
| Short | 1 | 99 | 288 |
| Short | 8 | 792 | 393 |
| Long | 1 | 1,099 | 1,288 |
| Long | 8 | 8,792 | 1,393 |

Subtracting short-state usage from long-state usage controls for question overhead and tokenizer differences:

- Perplexity: **1,000** extra tokens with one question; **8,000** with eight questions.
- Jev: **1,000** extra tokens with one question; **1,000** with eight questions.

Perplexity also multiplies both full-request counts exactly by eight: 99 → 792 and 1,099 → 8,792. The state occurs only once in each JSON request body.

## Minimal reproduction

Save as `reproduce.py`, set `PERPLEXITY_API_KEY` and `TYPESAFE_API_KEY` in the environment, and run `python3 reproduce.py`. No packages are needed. It makes four paid requests to each provider and prints usage and request IDs. Jev is a comparison control; the first four calls reproduce the Perplexity behavior independently.

```python
import json
import os
from urllib.request import Request, urlopen

providers = [
    ("Perplexity", "https://api.perplexity.ai/v1/decisions",
     "pplx-decider-v1.1-27b", "PERPLEXITY_API_KEY"),
    ("Jev", "https://api.typesafe.ai/v1/systemone",
     "jev-1.13.0", "TYPESAFE_API_KEY"),
]
short = "The warehouse shipment arrived on Tuesday. "
states = {
    "short": {"document": short},
    "long": {"document": short +
             "The inventory contains numbered blue boxes and green labels. " * 100},
}

for provider, endpoint, model, key in providers:
    counts = {}
    for length, state in states.items():
        for n in (1, 8):
            body = {
                "model": model,
                "state": state,
                "questions": {
                    f"q{i}": {
                        "type": "noul",
                        "instructions": "Does the document say the shipment arrived?",
                    }
                    for i in range(n)
                },
            }
            request = Request(
                endpoint, data=json.dumps(body).encode(),
                headers={"Authorization": "Bearer " + os.environ[key],
                         "Content-Type": "application/json"},
            )
            with urlopen(request, timeout=45) as response:
                payload = json.load(response)
                request_id = response.headers.get("x-request-id")
            counts[length, n] = payload["usage"]["input_tokens"]
            print(provider, length, n, payload["usage"], request_id)
    print(provider, "extra state tokens:",
          "one question =", counts["long", 1] - counts["short", 1],
          "eight questions =", counts["long", 8] - counts["short", 8])

```

## Request details

- Endpoint: `POST https://api.perplexity.ai/v1/decisions`
- Model: `pplx-decider-v1.1-27b`
- Language: Python 3.13.13; standard-library HTTP client
- Approximate time: October 8, 2026, 17:56 UTC (19:56 Africa/Johannesburg)
- Consistency: identical counts in two independent executions
- Content: synthetic warehouse text only; no images or private data

Perplexity request IDs from the second execution:

| State | Questions | Request ID |
| --- | --- | --- |
| Short | 1 | 682bfd05-10aa-4ec2-9736-c6be3d577a73 |
| Short | 8 | f9017169-e57a-4f2f-a7ad-6eab52e22b31 |
| Long | 1 | 4b30c5e9-f6b3-455c-b8a6-4f3863c85d71 |
| Long | 8 | bd8553b7-236e-47d7-a993-cf9d92868f11 |

The long-state, eight-question response returned this usage excerpt (HTTP 200):

```json
{"usage": {"input_tokens": 8792, "output_tokens": 8}}

```

## Impact

Batching many questions against a shared document multiplies the state component of reported billable usage. This materially changes cost estimates compared with accounting that amortizes shared state across questions. Could you confirm whether this is intended billing behavior or a usage-reporting bug?

---

<div class="post-metadata">

**Author:** ![andrewmadson-pplx](https://sea1.discourse-cdn.com/flex001/user_avatar/community.perplexity.ai/andrewmadson-pplx/32/3428_2.png) [@andrewmadson-pplx](https://community.perplexity.ai/u/andrewmadson-pplx)\
**Post date:** [October 9, 2026, 11:10pm UTC](https://community.perplexity.ai/t/decisions-api-shared-state-counted-per-question-in-usage-input-tokens/6316/2 "2026-10-09T23:10:41Z")

</div>

Hello @Urs_de_Swardt,

Thank you for the report. I have sent it over to Engineering for review. I will confirm if this is the intended behavior.

Andrew

---

<div class="post-metadata">

**Author:** ![andrewmadson-pplx](https://sea1.discourse-cdn.com/flex001/user_avatar/community.perplexity.ai/andrewmadson-pplx/32/3428_2.png) [@andrewmadson-pplx](https://community.perplexity.ai/u/andrewmadson-pplx)\
**Post date:** [October 10, 2026, 1:44pm UTC](https://community.perplexity.ai/t/decisions-api-shared-state-counted-per-question-in-usage-input-tokens/6316/3 "2026-10-10T13:44:17Z")

</div>

I’ve confirmed with engineering that the model is working as expected.

Each question runs its own inference with the full state, so the state is counted once per question in usage. input\_tokens. Billing uses that same value.

I’ve passed your request to engineering: count the shared state once per request as a feature request.

Thank you,

Andrew

---

<div class="post-metadata">

**Author:** ![lukasmueller99](https://avatars.discourse-cdn.com/v4/letter/l/ecb155/32.png) [@lukasmueller99](https://community.perplexity.ai/u/lukasmueller99)\
**Post date:** [October 10, 2026, 5:25pm UTC](https://community.perplexity.ai/t/decisions-api-shared-state-counted-per-question-in-usage-input-tokens/6316/4 "2026-10-10T17:25:54Z")

</div>

The billing count reflects how each sub-request is processed rather than bugged reporting. When batching questions against shared state, the engine treats each question as a separate context evaluation, so the full document tokens are counted for each prompt. If you need to amortize input tokens, sending a single structured prompt that asks for all evaluations at once avoids the per-question state multiplication.
