# Add a batch query endpoint to Sonar

**URL:** <https://community.perplexity.ai/t/add-a-batch-query-endpoint-to-sonar/6217>\
**Category:** Feature Requests\
**Tags:** sonar-api, search-api, api, sonar\
**Created:** [September 27, 2026, 1:52am UTC](https://community.perplexity.ai/t/add-a-batch-query-endpoint-to-sonar/6217 "2026-09-27T01:52:53Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![emfedsci](https://sea1.discourse-cdn.com/flex001/user_avatar/community.perplexity.ai/emfedsci/32/3619_2.png) [@emfedsci](https://community.perplexity.ai/u/emfedsci)\
**Post date:** [September 27, 2026, 1:52am UTC](https://community.perplexity.ai/t/add-a-batch-query-endpoint-to-sonar/6217/1 "2026-09-27T01:52:53Z")

</div>

What I’m requesting: A single API call that accepts an array of independent queries and returns an array of results — e.g. POST /sonar/batch with { “queries”: […] } → { “results”: [{ query, answer, citations }, …] }, with per-item status so one failed query doesn’t fail the whole batch.

Use case: Some user actions in our app trigger several independent Sonar queries at the same time rather than a single query. This produces short bursts of concurrent calls instead of steady traffic.

Why it matters: Batching these into one request would make request volume predictable, reduce per-call overhead and round-trips, simplify retry/backoff handling, and make latency and cost easier to forecast as usage grows. Bursty concurrency is the hardest part of capacity planning today.

Current limitation or workaround: We currently issue N separate concurrent calls and manage our own concurrency, retries, and rate-limit backoff around the burst. It works but is brittle and makes throughput and cost hard to predict.

Suggested API behavior:

POST /sonar/batch

{ “model”: “sonar”, “queries”: [“…”, “…”] }

→ 200 { “results”: [{ “query”: “…”, “answer”: “…”, “citations”: […], "

status": “ok” } ] }

Optional shared parameters (e.g. recency filter) applied across the batch.

Example: Request: { “queries”: [“\<query 1\>”, “\<query 2\>”] } Response: { “results”: [{ query:“\<query 1\>”, answer:“…”, citations:[…] }, { query:“\<query 2\>”, answer:“…”, citations:[…] } ] }

Expected vs current: Expected — one call returns all results with per-item citations and status. Current — one call per query; bursty concurrency we orchestrate ourselves.
