Pre-flight checklist
- This is a Perplexity API feature request, not a Perplexity app, Comet, or web UI request.
- I have checked whether this already exists in the docs or another forum topic.
- I have described the developer workflow or API use case this would improve.
Feature type
- New endpoint or API capability
- New model, preset, or model behavior
- New parameter or request option
- New response field or metadata
- SDK improvement
- Dashboard, API key, or auth improvement
- Billing, credits, rate limit, or usage reporting improvement
- Docs or examples improvement
- Other API improvement
Affected API area
- Agent API
- Search API
- Sonar API
- Embeddings API
- SDKs
- Dashboard, API keys, or auth
- Billing or credits
- Not sure
Summary
Add support for Anthropic Claude models in the Agent API, including reasoning controls (effort/budget) and thinking summaries in the response.
Problem
Claude models support extended thinking, but there’s currently no way to control it or see what it produced when going through Perplexity. Developers can’t set how much reasoning a request should use, so they can’t trade off latency and cost against answer quality per request. There’s also no way to get a summary of the model’s reasoning back, which makes it hard to debug, audit, or show users how an answer was reached.
Proposed API behavior
- Claude model compatibility: allow Claude models to be selected in the Agent/Responses API with the same request/response shape as other supported models.
- Reasoning control: a request option to set reasoning effort and/or a thinking token budget (for the models that still support it), e.g.
reasoning.effort(low|medium|high) and an optionalreasoning.budget_tokens. Reasoning can be turned off entirely for fast, cheap calls. - Reasoning summaries: a request option to return a summary of the model’s reasoning, e.g.
reasoning.summary(auto|concise|detailed), with the summary returned as a dedicated field or output item in the response, separate from the final answer text. - Usage metadata: report reasoning tokens separately in the usage object so billing and cost tracking stay transparent.
- Streaming: stream reasoning summary deltas separately from answer deltas.
Example usage
{
"model": "anthropic/claude-sonnet-4-5",
"input": "Compare the tradeoffs of these two database designs.",
"reasoning": {
"effort": "medium",
"summary": "auto"
}
}
Example response shape:
{
"output": [
{
"type": "reasoning",
"summary": [{ "type": "summary_text", "text": "Compared read/write patterns and..." }]
},
{
"type": "message",
"content": [{ "type": "output_text", "text": "..." }]
}
],
"usage": {
"input_tokens": 120,
"output_tokens": 640,
"reasoning_tokens": 310
}
}
Workarounds tried
- Prompting the model to “think step by step” and show its work. This is unreliable, inflates the visible output, and gives no control over reasoning depth.
- Calling Claude directly through Anthropic’s API, which loses Perplexity’s search and tool integration, so I can’t use both in one workflow.
Impact
- Developers building agents/research tools get predictable control over cost and latency per request.
- Debugging and trust: reasoning summaries make outputs easier to inspect, audit, and explain to end users.
Additional context
Anthropic’s extended thinking and OpenAI’s reasoning parameters both follow a similar pattern (effort/budget plus summarized output), so a shared reasoning object could cover multiple providers cleanly.