Text and streaming
OpenAI-shaped chat completions with durable replay.
Read as Markdown ↗POST /api/generate/text/v1/chat/completions
Requires generate scope and Idempotency-Key. Discover a reviewed text model before submitting. This is a subset of the OpenAI chat-completions shape; it is not a Responses, tools, or multimodal chat endpoint.
| Field | Type | Limits / default |
|---|---|---|
model | string, required | Enabled reviewed text model ID, 1–150 characters. |
messages | array, required | 1–60 objects with role and string content. |
messages[].role | string | system, user, or assistant. |
messages[].content | string | Up to 30,000 characters each. |
stream | boolean | false. |
max_tokens | integer | 1–8,192; default 4,096. |
temperature | number | Optional, 0–2. |
top_p | number | Optional, 0–1. |
response_format | object | Optional {type: "text"} or {type: "json_object"}. |
maxGems | integer | Optional conservative admission ceiling, 0–1,000,000. |
Request bytes are bounded to 160,000; serialized conversation bytes to 128,000. Unknown fields, tools, provider-routing overrides, remote references, and array-valued message content are rejected.
Non-streaming example
curl https://wegen.art/api/generate/text/v1/chat/completions \
-H "Authorization: Bearer $LETSGEN_API_KEY" \
-H 'Content-Type: application/json' \
-H 'Idempotency-Key: my-project-text-001' \
--data '{
"model": "TEXT_MODEL_ID_FROM_DISCOVERY",
"messages": [{"role": "user", "content": "Write a short scene about a moonlit harbor."}],
"max_tokens": 256,
"maxGems": 10
}'HTTP 200 returns id, object: "chat.completion", Unix created, model, choices with assistant content and finish_reason, and usage token counts.
Streaming
Set stream: true and consume Content-Type: text/event-stream. Use curl -N to disable buffering. SSE data: frames contain chat.completion.chunk objects. Accumulate choices[].delta.content; the final successful chunk includes usage, followed by data: [DONE].
A stream can also contain {error: {code: "REQUEST_UNCERTAIN", message}}. Treat an error or a stream ending without [DONE] as incomplete; never assume HTTP 200 alone proves success. Disconnecting does not cancel provider spending.
Budgets and replay
The API reserves a conservative whole-Gem maximum from the reviewed rates, input size, and requested output limit before inference. Measured usage settles once through the account LLM Gem allowance and consent-gated paid Gems. The charge cannot exceed the admitted reservation.
Repeat the exact body, including stream, with the same originating key and identity to retrieve a saved response. A completed streaming replay may return the whole saved answer in one chunk instead of reproducing the original chunk timings. Pending or ambiguous responses return 409 REQUEST_UNCERTAIN; they never dispatch again. See Budgets and retries.