POST /responses accepts requests in OpenAI’s Responses API format and returns typed output items: reasoning, message, function_call, and web_search_call. It works with every Venice text model and accepts an API key or x402 wallet auth.
When to use it
Use/responses when your client already speaks the Responses format, for example coding agents and SDKs built on OpenAI’s responses.create. For everything else, /chat/completions is the most complete way to use Venice: it supports structured outputs, file inputs, E2EE models, and every venice_parameters option.
Quick start
output array of typed items and a usage object:
Conversations are stateless
Venice does not store responses. Send the whole conversation ininput on every request, appending the previous output items and any tool results.
previous_response_id, store, and conversation are accepted but ignored, so they do not give the model any earlier context.
Streaming
Setstream: true to receive server-sent events: response.created, response.output_item.added, response.content_part.added, response.output_text.delta, response.function_call_arguments.delta, response.content_part.done, response.output_item.done, and finally response.completed, response.incomplete, or response.failed. The stream ends with data: [DONE].
Two events differ from OpenAI’s:
- Reasoning text streams as
response.reasoning.delta, not as OpenAI’s reasoning summary events. - When web search runs, a
response.web_search.doneevent carries the search results before the answer streams.
Supported parameters
Web search is billed as search augmentation, as on
/chat/completions.
Limitations
These are the current gaps between Venice’s/responses and OpenAI’s Responses API. Unless noted, the request still succeeds and the field is ignored.